mirror of
https://github.com/CopilotKit/CopilotKit.git
synced 2026-09-14 16:26:20 +08:00
ef1ca9808c
Result of 10 parallel QA agents auditing all 30 active demos against
langgraph-python (north-star). Each agent ported drift back to LP-verbatim
across three axes:
1. Agent layer
- tool_rendering_common.py: rebuilt to LP's surface — get_weather,
search_flights(origin, destination), get_stock_price, roll_d20,
roll_dice. Removed the ADK-only query_data.
- tool_rendering_*_agent.py (4 variants): ported LP's travel/concierge
prompt; reasoning-chain variant got LP's chain-two-tools prompt.
- beautiful_chat_agent.py: ported LP's per-tool system prompt; added
manage_sales_todos / get_sales_todos / generate_a2ui; dropped the
redundant schedule_meeting (frontend HITL handles it).
- open_gen_ui_agents.py: ported LP's full SYSTEM_PROMPT for both
variants, including the Websandbox.connection.remote.* contract
for the advanced sandbox demo (was `window.sandbox.*`, which the
LP frontend's Websandbox bridge silently no-ops).
- byoc_agents.py: fused LP's hashbrown + json-render prompts so the
single ADK byoc_agent emits both wire shapes. Aliases exported for
a future per-route split.
- declarative_gen_ui_agent.py: ported LP's a2ui_dynamic SYSTEM_PROMPT.
- a2ui_fixed_agent.py: picked up LP's #4734 regression guard
("exactly ONCE", "do NOT call again").
- agent_config_agent.py: rewrote to read useAgentContext (was
state["config"]); reconciled schema to LP's 3-field camelCase
{tone, expertise, responseLength} with LP's value enums.
- subagents_agent.py: dropped the "running" placeholder; returns
plain str so the LP-verbatim frontend's `result?.trim()` works.
- hitl_in_app_agent.py / hitl_in_chat_book_call_agent.py: prompts +
tool-result shape ({approved, reason}) aligned to LP.
- AGUIToolset() added wherever it was missing on the bespoke agents
(multimodal, mcp_apps, a2ui_fixed) so frontend-registered tools
reach the model.
2. Dedicated runtime routes
- copilotkit-multimodal/route.ts (new) — mirrors LP shape with
ADK's HttpAgent + AGENT_URL pattern.
- copilotkit-agent-config/route.ts (new) — same pattern.
- copilotkit-mcp-apps/route.ts — refreshed.
3. Frontend ports (ADK frontend brought to LP-verbatim where it had
drifted from the parity blitz state)
- tool-rendering family (4 demos): full re-port — WeatherCard,
FlightListCard, StockCard, D20Card, ReasoningBlock, CatchallRenderer,
suggestions, and the page wiring with all useRenderTool /
useDefaultRenderTool / reasoningMessage registrations.
- a2ui-fixed-schema, mcp-apps, multimodal: full frontend re-ports
with their _components/ Tailwind primitives.
- frontend-tools, frontend-tools-async, agent-config: ported LP's
component structure (separate Background, NotesCard with query_notes,
config-context-relay).
- shared-state-read, shared-state-read-write, readonly-state-agent-context:
ported LP's demo-layout + _components + suggestions. recipe-card.tsx
pulled directly from LP (one QA agent had adapted to Unicode glyphs
thinking ADK lacked lucide-react — it doesn't, after the parity blitz).
- shared-state-streaming, subagents, hitl-in-app: ported LP's
DocumentView / supervisor-activity / TicketsPanel structure.
hitl-in-app/page.tsx pulled directly from LP to keep the hyphenated
agent slug aligned with the renamed registry key.
- auth, hitl-in-chat: ported LP's SignInCard-first auth UX and the
time-picker Tailwind port.
- prebuilt-popup: pulled LP's main-content + suggestions split.
4. Test fixtures
- 30 tests/e2e/<slug>.spec.ts ported from LP, several overwriting
stale stubs (shared-state-streaming, subagents, auth, hitl-in-chat,
shared-state-read, agent-config).
- 30 qa/<slug>.md ported from LP with ADK env-var and registry
references substituted (GOOGLE_API_KEY, AGENT_URL, registry.py).
- QA3's byoc-hashbrown / byoc-json-render specs renamed to
declarative-hashbrown / declarative-json-render with internal
URL references substituted (the orchestrator pass had already
renamed the demo dirs + manifest entries).
Frontend changes from QA agents were filtered: kept where they ported
LP-verbatim into ADK, replaced with direct LP pulls where the agent
had made ADK-specific adaptations (one Unicode-glyph case, one
stale-registry-slug case).
Not touched per blitz rules: shared_chat.py, registry.py, manifest.yaml,
src/app/api/copilotkit/route.ts.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
3.9 KiB
3.9 KiB
QA: Tool Rendering (Default Catch-all) — Google ADK
Prerequisites
- Demo is deployed and accessible
- Agent backend is healthy (check /api/health)
- Agent slug
tool-rendering-default-catchallis registered at/api/copilotkit
Test Steps
1. Basic Functionality
- Navigate to the
tool-rendering-default-catchalldemo page - Verify the chat interface loads in a centered full-height layout (max-width 4xl,
rounded-2xl) - Verify the chat input placeholder "Type a message" is visible
- Send a basic message (e.g. "Hi")
- Verify the agent responds with a text message
2. Feature-Specific Checks — Built-in Default Tool-Call UI
The frontend calls useDefaultRenderTool() with NO config — it registers CopilotKit's package-provided DefaultToolCallRenderer as the * wildcard. The frontend adds ZERO custom per-tool or custom wildcard renderers. Every tool call must paint via this one built-in card.
Suggestions
- Verify "Weather in SF" suggestion pill is visible
- Verify "Find flights" suggestion pill is visible
- Verify "Roll a d20" suggestion pill is visible
- Click a suggestion and verify it either populates the input or sends the message
get_weather renders via the built-in default card
- Click the "Weather in SF" suggestion (or type "What's the weather in San Francisco?")
- Verify a default tool-call card appears with the tool name
get_weathervisible - Verify a status pill transitions through
Runningand lands onDone - Expand the card's "Arguments" section and verify it shows
{ "location": "San Francisco" }(or similar) - Expand the card's "Result" section and verify it shows the mock payload with
city,temperature: 68,humidity: 55,wind_speed: 10,conditions: "Sunny" - Verify NO custom-branded card appears (no
data-testid="custom-catchall-card", nodata-testid="weather-card")
search_flights renders via the SAME built-in default card
- Click the "Find flights" suggestion (or type "Find flights from SFO to JFK")
- Verify a tool-call card appears with tool name
search_flights - Verify the card has the identical visual style/structure as the
get_weathercard — same header layout, same status pill, same Arguments/Result sections - Verify the Result section contains three mock flights (United UA231, Delta DL412, JetBlue B6722)
roll_dice renders via the SAME built-in default card
- Click the "Roll a d20" suggestion (or type "Roll a 20-sided die")
- Verify a tool-call card for
roll_diceappears with the same default visual style - Verify the Result section shows
{ "sides": 20, "result": <1-20> }
get_stock_price renders via the SAME built-in default card
- Type "How is AAPL doing?"
- Verify a
get_stock_pricetool-call card appears with the default built-in style - Verify the Result shows
ticker: "AAPL", aprice_usd, and achange_pct
Chained tool calls
- Ask "What's the weather in Tokyo?" — the system prompt instructs the agent to chain tools
- Verify at least TWO default tool-call cards render in succession (e.g.
get_weatherthensearch_flights) - Verify every card uses the identical default built-in UI — visually indistinguishable apart from the tool name and payload
3. Error Handling
- Send an empty message (should be handled gracefully)
- Verify no console errors during normal usage
- Verify no unhandled-promise warnings when a tool call streams
Expected Results
- Chat loads within 3 seconds
- Every tool invocation paints via the built-in
DefaultToolCallRenderercard (tool name + live status pill + Arguments + Result) - All four distinct tools (
get_weather,search_flights,get_stock_price,roll_dice) render via the SAME default card — zero visual variance beyond payload - No custom-branded renderer appears anywhere
- No UI errors or broken layouts