mirror of
https://github.com/CopilotKit/CopilotKit.git
synced 2026-09-14 16:26:20 +08:00
ef1ca9808c
Result of 10 parallel QA agents auditing all 30 active demos against
langgraph-python (north-star). Each agent ported drift back to LP-verbatim
across three axes:
1. Agent layer
- tool_rendering_common.py: rebuilt to LP's surface — get_weather,
search_flights(origin, destination), get_stock_price, roll_d20,
roll_dice. Removed the ADK-only query_data.
- tool_rendering_*_agent.py (4 variants): ported LP's travel/concierge
prompt; reasoning-chain variant got LP's chain-two-tools prompt.
- beautiful_chat_agent.py: ported LP's per-tool system prompt; added
manage_sales_todos / get_sales_todos / generate_a2ui; dropped the
redundant schedule_meeting (frontend HITL handles it).
- open_gen_ui_agents.py: ported LP's full SYSTEM_PROMPT for both
variants, including the Websandbox.connection.remote.* contract
for the advanced sandbox demo (was `window.sandbox.*`, which the
LP frontend's Websandbox bridge silently no-ops).
- byoc_agents.py: fused LP's hashbrown + json-render prompts so the
single ADK byoc_agent emits both wire shapes. Aliases exported for
a future per-route split.
- declarative_gen_ui_agent.py: ported LP's a2ui_dynamic SYSTEM_PROMPT.
- a2ui_fixed_agent.py: picked up LP's #4734 regression guard
("exactly ONCE", "do NOT call again").
- agent_config_agent.py: rewrote to read useAgentContext (was
state["config"]); reconciled schema to LP's 3-field camelCase
{tone, expertise, responseLength} with LP's value enums.
- subagents_agent.py: dropped the "running" placeholder; returns
plain str so the LP-verbatim frontend's `result?.trim()` works.
- hitl_in_app_agent.py / hitl_in_chat_book_call_agent.py: prompts +
tool-result shape ({approved, reason}) aligned to LP.
- AGUIToolset() added wherever it was missing on the bespoke agents
(multimodal, mcp_apps, a2ui_fixed) so frontend-registered tools
reach the model.
2. Dedicated runtime routes
- copilotkit-multimodal/route.ts (new) — mirrors LP shape with
ADK's HttpAgent + AGENT_URL pattern.
- copilotkit-agent-config/route.ts (new) — same pattern.
- copilotkit-mcp-apps/route.ts — refreshed.
3. Frontend ports (ADK frontend brought to LP-verbatim where it had
drifted from the parity blitz state)
- tool-rendering family (4 demos): full re-port — WeatherCard,
FlightListCard, StockCard, D20Card, ReasoningBlock, CatchallRenderer,
suggestions, and the page wiring with all useRenderTool /
useDefaultRenderTool / reasoningMessage registrations.
- a2ui-fixed-schema, mcp-apps, multimodal: full frontend re-ports
with their _components/ Tailwind primitives.
- frontend-tools, frontend-tools-async, agent-config: ported LP's
component structure (separate Background, NotesCard with query_notes,
config-context-relay).
- shared-state-read, shared-state-read-write, readonly-state-agent-context:
ported LP's demo-layout + _components + suggestions. recipe-card.tsx
pulled directly from LP (one QA agent had adapted to Unicode glyphs
thinking ADK lacked lucide-react — it doesn't, after the parity blitz).
- shared-state-streaming, subagents, hitl-in-app: ported LP's
DocumentView / supervisor-activity / TicketsPanel structure.
hitl-in-app/page.tsx pulled directly from LP to keep the hyphenated
agent slug aligned with the renamed registry key.
- auth, hitl-in-chat: ported LP's SignInCard-first auth UX and the
time-picker Tailwind port.
- prebuilt-popup: pulled LP's main-content + suggestions split.
4. Test fixtures
- 30 tests/e2e/<slug>.spec.ts ported from LP, several overwriting
stale stubs (shared-state-streaming, subagents, auth, hitl-in-chat,
shared-state-read, agent-config).
- 30 qa/<slug>.md ported from LP with ADK env-var and registry
references substituted (GOOGLE_API_KEY, AGENT_URL, registry.py).
- QA3's byoc-hashbrown / byoc-json-render specs renamed to
declarative-hashbrown / declarative-json-render with internal
URL references substituted (the orchestrator pass had already
renamed the demo dirs + manifest entries).
Frontend changes from QA agents were filtered: kept where they ported
LP-verbatim into ADK, replaced with direct LP pulls where the agent
had made ADK-specific adaptations (one Unicode-glyph case, one
stale-registry-slug case).
Not touched per blitz rules: shared_chat.py, registry.py, manifest.yaml,
src/app/api/copilotkit/route.ts.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
165 lines
5.2 KiB
TypeScript
165 lines
5.2 KiB
TypeScript
import { test, expect } from "@playwright/test";
|
|
|
|
// QA reference: qa/tool-rendering.md
|
|
// Demo source: src/app/demos/tool-rendering/page.tsx
|
|
//
|
|
// Pill-driven 6-test plan. The cell registers a per-tool useRenderTool
|
|
// for every "interesting" tool (get_weather, search_flights,
|
|
// get_stock_price, roll_d20) plus a wildcard catch-all. Each pill
|
|
// drives the corresponding tool path and asserts on the per-tool
|
|
// branded card's stable testid plus deterministic fixture values.
|
|
//
|
|
// Aimock fixtures live in showcase/aimock/d5-all.json and pin every
|
|
// pill prompt to a deterministic tool-call sequence.
|
|
|
|
const SUGGESTION_TIMEOUT = 15000;
|
|
const TOOL_TIMEOUT = 60000;
|
|
|
|
test.describe("Tool Rendering", () => {
|
|
test.beforeEach(async ({ page }) => {
|
|
await page.goto("/demos/tool-rendering");
|
|
await expect(page.getByPlaceholder("Type a message")).toBeVisible({
|
|
timeout: SUGGESTION_TIMEOUT,
|
|
});
|
|
});
|
|
|
|
test("page loads with composer and 5 suggestion pills", async ({ page }) => {
|
|
const suggestions = page.locator('[data-testid="copilot-suggestion"]');
|
|
for (const title of [
|
|
"Weather in SF",
|
|
"Find flights",
|
|
"Stock price",
|
|
"Roll a d20",
|
|
"Chain tools",
|
|
]) {
|
|
await expect(suggestions.filter({ hasText: title }).first()).toBeVisible({
|
|
timeout: SUGGESTION_TIMEOUT,
|
|
});
|
|
}
|
|
});
|
|
|
|
test("Weather in SF pill renders the SF weather card", async ({ page }) => {
|
|
await page
|
|
.locator('[data-testid="copilot-suggestion"]')
|
|
.filter({ hasText: "Weather in SF" })
|
|
.first()
|
|
.click();
|
|
|
|
const card = page.locator('[data-testid="weather-card"]').first();
|
|
await expect(card).toBeVisible({ timeout: TOOL_TIMEOUT });
|
|
await expect(card.locator('[data-testid="weather-city"]')).toContainText(
|
|
"San Francisco",
|
|
{ timeout: TOOL_TIMEOUT },
|
|
);
|
|
await expect(
|
|
card.locator('[data-testid="weather-humidity"]'),
|
|
).toContainText("55%", { timeout: TOOL_TIMEOUT });
|
|
await expect(card.locator('[data-testid="weather-wind"]')).toContainText(
|
|
"10",
|
|
{ timeout: TOOL_TIMEOUT },
|
|
);
|
|
});
|
|
|
|
test("Find flights pill renders the flights card with deterministic flights", async ({
|
|
page,
|
|
}) => {
|
|
await page
|
|
.locator('[data-testid="copilot-suggestion"]')
|
|
.filter({ hasText: "Find flights" })
|
|
.first()
|
|
.click();
|
|
|
|
const card = page.locator('[data-testid="flights-card"]').first();
|
|
await expect(card).toBeVisible({ timeout: TOOL_TIMEOUT });
|
|
await expect(card.locator('[data-testid="flight-origin"]')).toContainText(
|
|
"SFO",
|
|
{ timeout: TOOL_TIMEOUT },
|
|
);
|
|
await expect(
|
|
card.locator('[data-testid="flight-destination"]'),
|
|
).toContainText("JFK", { timeout: TOOL_TIMEOUT });
|
|
|
|
// At least 2 flight rows from the deterministic fixture (NOT the
|
|
// a2ui beautiful-chat boilerplate — a2ui shows a different shell).
|
|
const rows = card.locator('[data-testid="flight-row"]');
|
|
await expect
|
|
.poll(async () => rows.count(), { timeout: TOOL_TIMEOUT })
|
|
.toBeGreaterThanOrEqual(2);
|
|
});
|
|
|
|
test("Stock price pill renders the AAPL stock card", async ({ page }) => {
|
|
await page
|
|
.locator('[data-testid="copilot-suggestion"]')
|
|
.filter({ hasText: "Stock price" })
|
|
.first()
|
|
.click();
|
|
|
|
const card = page.locator('[data-testid="stock-card"]').first();
|
|
await expect(card).toBeVisible({ timeout: TOOL_TIMEOUT });
|
|
await expect(card.locator('[data-testid="stock-ticker"]')).toHaveText(
|
|
"AAPL",
|
|
{ timeout: TOOL_TIMEOUT },
|
|
);
|
|
await expect(card.locator('[data-testid="stock-price"]')).toContainText(
|
|
"$338.37",
|
|
{ timeout: TOOL_TIMEOUT },
|
|
);
|
|
await expect(card.locator('[data-testid="stock-change"]')).toContainText(
|
|
"-2.96%",
|
|
{ timeout: TOOL_TIMEOUT },
|
|
);
|
|
});
|
|
|
|
test("Roll a d20 pill produces exactly 5 d20 cards, last result is 20", async ({
|
|
page,
|
|
}) => {
|
|
await page
|
|
.locator('[data-testid="copilot-suggestion"]')
|
|
.filter({ hasText: "Roll a d20" })
|
|
.first()
|
|
.click();
|
|
|
|
const cards = page.locator('[data-testid="d20-card"]');
|
|
|
|
// Wait for all 5 sequential rolls to land.
|
|
await expect
|
|
.poll(async () => cards.count(), { timeout: TOOL_TIMEOUT })
|
|
.toBe(5);
|
|
|
|
// Final roll is a 20.
|
|
await expect(cards.nth(4).locator('[data-testid="d20-value"]')).toHaveText(
|
|
"20",
|
|
{ timeout: TOOL_TIMEOUT },
|
|
);
|
|
|
|
// First 4 rolls are non-20.
|
|
for (let i = 0; i < 4; i++) {
|
|
const value = await cards
|
|
.nth(i)
|
|
.locator('[data-testid="d20-value"]')
|
|
.innerText();
|
|
expect(value.trim()).not.toBe("20");
|
|
}
|
|
});
|
|
|
|
test("Chain tools pill renders weather + flights + d20 cards in one turn", async ({
|
|
page,
|
|
}) => {
|
|
await page
|
|
.locator('[data-testid="copilot-suggestion"]')
|
|
.filter({ hasText: "Chain tools" })
|
|
.first()
|
|
.click();
|
|
|
|
await expect(
|
|
page.locator('[data-testid="weather-card"]').first(),
|
|
).toBeVisible({ timeout: TOOL_TIMEOUT });
|
|
await expect(
|
|
page.locator('[data-testid="flights-card"]').first(),
|
|
).toBeVisible({ timeout: TOOL_TIMEOUT });
|
|
await expect(page.locator('[data-testid="d20-card"]').first()).toBeVisible({
|
|
timeout: TOOL_TIMEOUT,
|
|
});
|
|
});
|
|
});
|