Files
Tyler Slaton ef1ca9808c feat(showcase/google-adk): QA-blitz parity ports — agent prompts, tool surface, e2e specs
Result of 10 parallel QA agents auditing all 30 active demos against
langgraph-python (north-star). Each agent ported drift back to LP-verbatim
across three axes:

1. Agent layer
   - tool_rendering_common.py: rebuilt to LP's surface — get_weather,
     search_flights(origin, destination), get_stock_price, roll_d20,
     roll_dice. Removed the ADK-only query_data.
   - tool_rendering_*_agent.py (4 variants): ported LP's travel/concierge
     prompt; reasoning-chain variant got LP's chain-two-tools prompt.
   - beautiful_chat_agent.py: ported LP's per-tool system prompt; added
     manage_sales_todos / get_sales_todos / generate_a2ui; dropped the
     redundant schedule_meeting (frontend HITL handles it).
   - open_gen_ui_agents.py: ported LP's full SYSTEM_PROMPT for both
     variants, including the Websandbox.connection.remote.* contract
     for the advanced sandbox demo (was `window.sandbox.*`, which the
     LP frontend's Websandbox bridge silently no-ops).
   - byoc_agents.py: fused LP's hashbrown + json-render prompts so the
     single ADK byoc_agent emits both wire shapes. Aliases exported for
     a future per-route split.
   - declarative_gen_ui_agent.py: ported LP's a2ui_dynamic SYSTEM_PROMPT.
   - a2ui_fixed_agent.py: picked up LP's #4734 regression guard
     ("exactly ONCE", "do NOT call again").
   - agent_config_agent.py: rewrote to read useAgentContext (was
     state["config"]); reconciled schema to LP's 3-field camelCase
     {tone, expertise, responseLength} with LP's value enums.
   - subagents_agent.py: dropped the "running" placeholder; returns
     plain str so the LP-verbatim frontend's `result?.trim()` works.
   - hitl_in_app_agent.py / hitl_in_chat_book_call_agent.py: prompts +
     tool-result shape ({approved, reason}) aligned to LP.
   - AGUIToolset() added wherever it was missing on the bespoke agents
     (multimodal, mcp_apps, a2ui_fixed) so frontend-registered tools
     reach the model.

2. Dedicated runtime routes
   - copilotkit-multimodal/route.ts (new) — mirrors LP shape with
     ADK's HttpAgent + AGENT_URL pattern.
   - copilotkit-agent-config/route.ts (new) — same pattern.
   - copilotkit-mcp-apps/route.ts — refreshed.

3. Frontend ports (ADK frontend brought to LP-verbatim where it had
   drifted from the parity blitz state)
   - tool-rendering family (4 demos): full re-port — WeatherCard,
     FlightListCard, StockCard, D20Card, ReasoningBlock, CatchallRenderer,
     suggestions, and the page wiring with all useRenderTool /
     useDefaultRenderTool / reasoningMessage registrations.
   - a2ui-fixed-schema, mcp-apps, multimodal: full frontend re-ports
     with their _components/ Tailwind primitives.
   - frontend-tools, frontend-tools-async, agent-config: ported LP's
     component structure (separate Background, NotesCard with query_notes,
     config-context-relay).
   - shared-state-read, shared-state-read-write, readonly-state-agent-context:
     ported LP's demo-layout + _components + suggestions. recipe-card.tsx
     pulled directly from LP (one QA agent had adapted to Unicode glyphs
     thinking ADK lacked lucide-react — it doesn't, after the parity blitz).
   - shared-state-streaming, subagents, hitl-in-app: ported LP's
     DocumentView / supervisor-activity / TicketsPanel structure.
     hitl-in-app/page.tsx pulled directly from LP to keep the hyphenated
     agent slug aligned with the renamed registry key.
   - auth, hitl-in-chat: ported LP's SignInCard-first auth UX and the
     time-picker Tailwind port.
   - prebuilt-popup: pulled LP's main-content + suggestions split.

4. Test fixtures
   - 30 tests/e2e/<slug>.spec.ts ported from LP, several overwriting
     stale stubs (shared-state-streaming, subagents, auth, hitl-in-chat,
     shared-state-read, agent-config).
   - 30 qa/<slug>.md ported from LP with ADK env-var and registry
     references substituted (GOOGLE_API_KEY, AGENT_URL, registry.py).
   - QA3's byoc-hashbrown / byoc-json-render specs renamed to
     declarative-hashbrown / declarative-json-render with internal
     URL references substituted (the orchestrator pass had already
     renamed the demo dirs + manifest entries).

Frontend changes from QA agents were filtered: kept where they ported
LP-verbatim into ADK, replaced with direct LP pulls where the agent
had made ADK-specific adaptations (one Unicode-glyph case, one
stale-registry-slug case).

Not touched per blitz rules: shared_chat.py, registry.py, manifest.yaml,
src/app/api/copilotkit/route.ts.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 17:58:06 -07:00

5.4 KiB

QA: MCP Apps — Google ADK

Prerequisites

  • Demo is deployed and accessible at /demos/mcp-apps on the dashboard host
  • Agent backend is healthy; GOOGLE_API_KEY is set on Railway; AGENT_URL points at the ADK agent server exposing the mcp_apps endpoint (registered as agent name mcp-apps — see src/app/api/copilotkit-mcp-apps/route.ts)
  • MCP server target: the public Excalidraw MCP app at https://mcp.excalidraw.com (override via MCP_SERVER_URL). Pinned serverId: "excalidraw" so URL changes don't silently break persisted activities
  • Note: the demo source contains no data-testid attributes and registers no custom activity renderer — CopilotKit's built-in MCPAppsActivityRenderer handles the sandboxed iframe automatically. Checks below rely on verbatim visible text, network traffic, and the iframe DOM

Test Steps

1. Basic Functionality

  • Navigate to /demos/mcp-apps; verify the page renders within 3s and a single CopilotChat pane is centered (max-width ~896px, rounded-2xl, full-height)
  • Verify the chat is wired to runtimeUrl="/api/copilotkit-mcp-apps" and agent="mcp-apps" (DevTools → Network: sending a message hits that endpoint)
  • Verify both suggestion pills are visible with verbatim titles:
    • "Draw a flowchart"
    • "Sketch a system diagram"
  • Send "Hello" and verify an assistant text response appears within 10s (no MCP activity iframe for plain text)

2. Feature-Specific Checks

MCP Server Connection (runtime mcpApps.servers)

  • Send the first flow-chart prompt; in DevTools → Network, verify the POST to /api/copilotkit-mcp-apps succeeds (status 200) and the server-side runtime resolves tools from https://mcp.excalidraw.com (watch the server logs for the MCP Apps middleware attaching the Excalidraw tool set — notably create_view)
  • Verify no console errors mentioning the MCP server URL, auth, or tool-schema parse failures

MCP Tool Invocation (create_view)

  • Click "Draw a flowchart"; within 60s verify the agent calls the create_view MCP tool exactly ONCE (per SYSTEM_PROMPT in src/agents/mcp_apps_agent.py: "Call create_view ONCE with 3-5 elements total") — confirm via DevTools → Network stream or backend logs
  • Verify the tool payload contains 3-5 Excalidraw elements (shapes + arrows + optional title text), each with a unique string id, and ends with ONE cameraUpdate sized 600x450 or 800x600

Activity Renderer (built-in MCPAppsActivityRenderer)

  • Within 60s of the tool call, verify a sandboxed <iframe> renders inline in the chat transcript (activity-message slot) pointed at the Excalidraw MCP UI resource
  • Verify the iframe has a sandbox attribute (CopilotKit's built-in renderer always sandboxes MCP UI resources)
  • Verify the iframe paints a flow-chart-shaped diagram: at least 3 shape nodes (rectangles, ellipses, or diamonds with text labels) connected by arrows, framed within the viewport (camera-update step from the system prompt)
  • Verify the assistant text below the iframe is a single short sentence describing what was drawn (per system prompt)

Server-Driven UI Update (second prompt, same thread)

  • Without reloading, send the second suggestion "Sketch a system diagram"; within 60s verify a new activity iframe renders in-transcript containing a client → server → database layout (3 labeled shapes + 2 arrows)
  • Verify the previous flow-chart iframe is still present and un-stale in the scrollback (activity messages persist, matching the rationale for the pinned serverId: "excalidraw" in the runtime config)

End-to-End MCP Interaction (concrete, single-case)

  • Send an explicit prompt: "Use Excalidraw to draw exactly 2 rectangles labelled 'A' and 'B' connected by one arrow from A to B."
  • Within 60s verify: (1) create_view is called ONCE with exactly 3 elements (2 rectangles + 1 arrow) plus the trailing cameraUpdate; (2) an iframe renders showing two labelled rectangles with a connecting arrow; (3) the assistant reply is one short sentence; (4) no duplicate create_view invocations or retries appear in network / logs

3. Error Handling

  • Send an empty message; verify it is a no-op (no user bubble, no assistant response)
  • Send "What is 2+2?"; verify the agent replies in plain text without invoking create_view (no iframe, no MCP activity in the stream)
  • DevTools → Console: walk through all flows above; verify no uncaught errors, no CORS failures referencing mcp.excalidraw.com, and no "sandbox" / iframe-permission warnings

Expected Results

  • Chat loads within 3s; plain-text response within 10s; MCP-backed iframe renders within 60s of prompt (bias is "correct-enough diagram fast" per system prompt, one create_view call)
  • MCP server connection to https://mcp.excalidraw.com succeeds and the Excalidraw tool set (including create_view) is advertised to the agent at request time
  • At least one concrete end-to-end MCP interaction completes: user prompt → create_view tool call → activity event → sandboxed iframe painting the requested diagram
  • The built-in MCPAppsActivityRenderer is used (no app-side useRenderActivityMessage / renderActivityMessages registration exists in page.tsx — per the @region[no-frontend-renderer-needed] contract)
  • No UI layout breaks, no uncaught console errors, no duplicate create_view invocations within a single prompt turn