Files
copilotkit__copilotkit/showcase/integrations/ms-agent-python/tests/e2e/tool-rendering-custom-catchall.spec.ts
Alem Tuzlak 5cd77233b7 feat(showcase/ms-agent-python): LGP parity sweep — 33/37 cells green
Brings ms-agent-python to one-to-one parity with langgraph-python (the D5
north star). Playwright e2e suite goes from 49/108 (~26%) → 164/178 (~92%),
33 of 37 cells fully green.

Manifest parity:
- Drop 4 MAF-only cells with no LGP analog: agentic-chat-reasoning,
  hitl-in-chat-booking, shared-state-write, reasoning-default-render.
  Reasoning is handled by reasoning-default + reasoning-custom (LGP);
  booking pill folds into hitl-in-chat; shared-state-write was a TODO stub.
- Rename byoc-hashbrown → declarative-hashbrown and byoc-json-render →
  declarative-json-render. Demo dir, API route dir, and frontend agent id
  follow LGP's naming. Python module files retain the legacy `byoc_*`
  prefix and FastAPI paths stay `/byoc-hashbrown` / `/byoc-json-render`
  (matches LGP's "module name retains legacy graph id" convention).
- Port LGP `_shared/`, `_shared/interrupt-fallback-slots.ts`, and
  `demos/layout.tsx` for one-to-one parity.

Cells ported verbatim from LGP (page + spec):
- agentic-chat, auth, beautiful-chat, chat-customization-css, chat-slots,
  declarative-gen-ui, declarative-hashbrown, declarative-json-render,
  frontend-tools, frontend-tools-async, gen-ui-agent, gen-ui-interrupt,
  gen-ui-tool-based, headless-complete, headless-simple, hitl-in-app,
  hitl-in-chat, shared-state-read, shared-state-read-write,
  shared-state-streaming, subagents, tool-rendering, plus all four
  tool-rendering* variants, a2ui-fixed-schema, agent-config, mcp-apps,
  multimodal, open-gen-ui, open-gen-ui-advanced, prebuilt-popup,
  prebuilt-sidebar, readonly-state-agent-context, reasoning-default,
  reasoning-custom, voice.

Backend infrastructure:
- Swap shared `OpenAIChatClient` (Responses API) → `OpenAIChatCompletionClient`
  (ChatCompletions). Root cause of the cross-cell post-tool ChatClientException
  family: Responses API is stateful and only sends NEW items per leg,
  relying on `previous_response_id` for history. aimock has no view of
  that server-side state, so second-leg requests arrived without the
  user message — fixture matchers keyed on `userMessage` couldn't fire
  and the run fell through to real OpenAI. ChatCompletions sends full
  history every leg, matching the LGP wire shape.
- Bump @ag-ui/client ^0.0.43 → ^0.0.53 (matches google-adk/LGP). Fixes
  the REASONING_* Zod discriminator trap on the catch-all agent.
- Regenerate package-lock.json in isolation outside the pnpm monorepo so
  npm-arborist doesn't resolve transitives against pnpm's hoisted
  symlinks (avoid 40+ `../../../node_modules/.pnpm/...` paths in the
  lockfile that break `npm ci` inside Docker).
- Add `yaml` (^2.8.4) for the new `src/app/demos/layout.tsx` that reads
  manifest.yaml for per-cell page titles (LGP parity).

New / re-added MAF agent backends with LGP-equivalent behavior:
- reasoning_agent.py (uses Responses API explicitly — the only chat
  client that emits AG-UI REASONING_MESSAGE_* events; rest of the
  integration stays on ChatCompletions).
- tool_rendering_agent.py (non-reasoning sibling of the existing
  reasoning_chain variant; shares tool surface via direct imports so
  they can never drift apart; routes the three catchall cells to a
  non-reasoning backend so the default renderer spec stops failing on
  leaked reasoning blocks).
- gen_ui_agent.py — `set_steps` tool + `steps` state schema +
  `predict_state_config` mirrors LGP's StateStreamingMiddleware shape.
- shared_state_streaming.py — `write_document` tool with
  `predict_state_config` that streams the `document` arg into
  `state.document` per-token.
- readonly_state_agent_context.py — minimal agent that consumes
  frontend-provided `useAgentContext` entries; no tools.
- headless_complete_agent.py — three deterministic tools (`get_weather`,
  `get_stock_price`, `get_revenue_chart`) mounted at /headless-complete
  on the mcp-apps runtime (was routing to catch-all sales agent, which
  returned seeded-random weather instead of the deterministic 68°F the
  test asserts on).

Wiring:
- copilotkit/route.ts: register the new agents, drop the stale
  shared-state-write entry, route all three tool-rendering variants to
  the non-reasoning backend (the reasoning-chain cell keeps its own
  dedicated path), register reasoning-default + reasoning-custom on
  /reasoning, register gen-ui-agent on /gen-ui-agent,
  shared-state-streaming on /shared-state-streaming,
  readonly-state-agent-context on its dedicated path.
- copilotkit-mcp-apps/route.ts: register headless-complete agent (was
  missing — the strict useAgent runtime sync in the newer
  @copilotkit/react-core surfaced the gap).
- copilotkit-declarative-hashbrown/route.ts + copilotkit-declarative-json-render/route.ts:
  new dedicated runtimes; agent IDs and runtime URLs follow LGP.
- copilotkit-declarative-gen-ui/route.ts: drop non-LGP `openGenerativeUI:
  false` for parity.

A2UI tool rename — `render_a2ui` → `_design_a2ui_surface`:
- Ported LGP's `tools/generate_a2ui.py` (LGP renamed the secondary-LLM
  tool to `_design_a2ui_surface` to avoid the A2UI middleware's bypass;
  shared d5-all.json fixtures key the response on this name).
- Renamed every `render_a2ui` occurrence in src/agents/{a2ui_dynamic,
  agent,beautiful_chat}.py and `tools/__init__.py`.
- Updated 4 declarative-gen-ui aimock fixtures to pass `context` arg in
  the first-leg `generate_a2ui` tool call (agent_framework doesn't
  auto-inject AgentSession into our @tool function so `session=None` and
  the secondary-LLM `user_content` was defaulting to a catch-all string
  containing "KPI dashboard" — every pill matched the KPI fixture).

Aimock router patch persisted alongside the integration changes:
hasToolResult matcher restricted to scan only messages after the last
user message (was global). The patch lives in F:/projects/cpk/aimock —
upstream PR pending.

Test infrastructure:
- playwright.config.ts: cap local workers at 4 + retries at 1. CI keeps
  workers=1, retries=2. `agent_framework.Agent` is reused across requests
  and the shared OpenAI HTTP client serialises concurrent SSE streams;
  >4 workers makes 30s timeouts inevitable on a few cells. Confirmed
  with hard data: workers=1 = 164 passed (16.8 min), workers=4+retries=1
  = 164 passed (7.2 min), workers=undefined = 159 passed. Same green
  set, ~2x faster. Long-term upstream fix is per-request Agent
  instantiation in agent_framework_ag_ui.

Remaining 14 failures across 4 cells documented per-cell in the Notion
D5 sweep doc (declarative-gen-ui A2UI surface mounting, multimodal
attachment forwarding, tool-rendering-default-catchall multi-pill chain,
tool-rendering-reasoning-chain multi-leg chains). Each has a specific
next-pass action.
2026-05-19 18:36:01 +02:00

254 lines
8.7 KiB
TypeScript
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
import { test, expect } from "@playwright/test";
// QA reference: qa/tool-rendering-custom-catchall.md
// Demo source: src/app/demos/tool-rendering-custom-catchall/page.tsx
// Renderer source: src/app/demos/tool-rendering-custom-catchall/custom-catchall-renderer.tsx
//
// This cell registers a SINGLE branded wildcard renderer via
// `useDefaultRenderTool`. Every tool call must paint via the same
// `[data-testid="custom-wildcard-card"]` shell — no per-tool
// specialization. Test 6 is the load-bearing assertion: every card on
// the page after each pill click shares the same testid signature.
const SUGGESTION_TIMEOUT = 15000;
const TOOL_TIMEOUT = 60000;
const PILLS = ["Weather in SF", "Find flights", "Roll a d20", "Chain tools"];
test.describe("Tool Rendering — Custom Catch-all (branded wildcard)", () => {
test.beforeEach(async ({ page }) => {
await page.goto("/demos/tool-rendering-custom-catchall");
await expect(page.getByPlaceholder("Type a message")).toBeVisible({
timeout: SUGGESTION_TIMEOUT,
});
});
test("page loads with composer and 4 suggestion pills", async ({ page }) => {
const suggestions = page.locator('[data-testid="copilot-suggestion"]');
for (const title of PILLS) {
await expect(suggestions.filter({ hasText: title }).first()).toBeVisible({
timeout: SUGGESTION_TIMEOUT,
});
}
// Sanity: per-tool branded testids from sibling cells stay at zero.
await expect(page.locator('[data-testid="weather-card"]')).toHaveCount(0);
await expect(page.locator('[data-testid="flights-card"]')).toHaveCount(0);
await expect(page.locator('[data-testid="stock-card"]')).toHaveCount(0);
await expect(page.locator('[data-testid="d20-card"]')).toHaveCount(0);
// Sanity: the OOTB default-renderer testid does NOT appear here.
await expect(
page.locator('[data-testid="copilot-tool-render"]'),
).toHaveCount(0);
});
test("Weather in SF pill paints the branded wildcard card for get_weather", async ({
page,
}) => {
await page
.locator('[data-testid="copilot-suggestion"]')
.filter({ hasText: "Weather in SF" })
.first()
.click();
const card = page
.locator(
'[data-testid="custom-wildcard-card"][data-tool-name="get_weather"]',
)
.first();
await expect(card).toBeVisible({ timeout: TOOL_TIMEOUT });
await expect(
card.locator('[data-testid="custom-wildcard-tool-name"]'),
).toHaveText("get_weather");
await expect(
card.locator('[data-testid="custom-wildcard-args"]'),
).toContainText("San Francisco", { timeout: TOOL_TIMEOUT });
});
test("Find flights pill paints the SAME branded wildcard card for search_flights", async ({
page,
}) => {
await page
.locator('[data-testid="copilot-suggestion"]')
.filter({ hasText: "Find flights" })
.first()
.click();
const card = page
.locator(
'[data-testid="custom-wildcard-card"][data-tool-name="search_flights"]',
)
.first();
await expect(card).toBeVisible({ timeout: TOOL_TIMEOUT });
await expect(
card.locator('[data-testid="custom-wildcard-tool-name"]'),
).toHaveText("search_flights");
// Result block surfaces the deterministic flights from our fixture
// (NOT the a2ui beautiful-chat boilerplate).
await expect(
card.locator('[data-testid="custom-wildcard-result"]'),
).toContainText(/United|Delta|JetBlue/, { timeout: TOOL_TIMEOUT });
});
test("Roll a d20 pill paints exactly 5 wildcard cards, last result is 20", async ({
page,
}) => {
await page
.locator('[data-testid="copilot-suggestion"]')
.filter({ hasText: "Roll a d20" })
.first()
.click();
const cards = page.locator(
'[data-testid="custom-wildcard-card"][data-tool-name="roll_d20"]',
);
await expect
.poll(async () => cards.count(), { timeout: TOOL_TIMEOUT })
.toBe(5);
// 5th card's result is 20.
await expect(
cards.nth(4).locator('[data-testid="custom-wildcard-result"]'),
).toContainText(/"value":\s*20|"result":\s*20/, { timeout: TOOL_TIMEOUT });
// First 4 are non-20.
for (let i = 0; i < 4; i++) {
const txt = await cards
.nth(i)
.locator('[data-testid="custom-wildcard-result"]')
.innerText();
expect(txt).not.toMatch(/"value":\s*20|"result":\s*20/);
}
});
test("Chain tools pill paints 3 wildcard cards (one per tool)", async ({
page,
}) => {
await page
.locator('[data-testid="copilot-suggestion"]')
.filter({ hasText: "Chain tools" })
.first()
.click();
await expect(
page
.locator(
'[data-testid="custom-wildcard-card"][data-tool-name="get_weather"]',
)
.first(),
).toBeVisible({ timeout: TOOL_TIMEOUT });
await expect(
page
.locator(
'[data-testid="custom-wildcard-card"][data-tool-name="search_flights"]',
)
.first(),
).toBeVisible({ timeout: TOOL_TIMEOUT });
await expect(
page
.locator(
'[data-testid="custom-wildcard-card"][data-tool-name="roll_d20"]',
)
.first(),
).toBeVisible({ timeout: TOOL_TIMEOUT });
});
test("every rendered card shares the same wildcard testid signature", async ({
page,
}) => {
// Cross-tool sanity: drive Chain tools (3 distinct tools → 3
// cards) and assert every card matches the same wildcard shell.
await page
.locator('[data-testid="copilot-suggestion"]')
.filter({ hasText: "Chain tools" })
.first()
.click();
const cards = page.locator('[data-testid="custom-wildcard-card"]');
await expect
.poll(async () => cards.count(), { timeout: TOOL_TIMEOUT })
.toBeGreaterThanOrEqual(3);
const total = await cards.count();
await expect(
page.locator('[data-testid="custom-wildcard-tool-name"]'),
).toHaveCount(total);
await expect(
page.locator('[data-testid="custom-wildcard-args"]'),
).toHaveCount(total);
// All cards expose distinct tool names but the SAME shell.
const toolNames = await cards.evaluateAll((nodes) =>
nodes.map((n) => n.getAttribute("data-tool-name")),
);
const uniqueNames = new Set(toolNames);
expect(uniqueNames.size).toBeGreaterThanOrEqual(3);
for (const name of toolNames) {
expect(["get_weather", "search_flights", "roll_d20"]).toContain(name);
}
// The OOTB default-renderer testid stays at zero — proves the
// single custom wildcard is what painted, not the framework
// fallback.
await expect(
page.locator('[data-testid="copilot-tool-render"]'),
).toHaveCount(0);
});
// Regression for the aimock multi-pill bug:
// The d20 and Chain-tools fixtures used global thread state
// (`turnIndex`, `hasToolResult`) to drive sequencing. After clicking
// Find flights, the d20 loop entered at turnIndex=2 (rendering only 3
// cards instead of 5) and the Chain-tools tool-emitting fixture was
// skipped entirely (no tool cards, just the final "Done — Tokyo is
// sunny…" content). Fix: chain all follow-ups via `toolCallId`. This
// test drives Find flights → Roll a d20 → Chain tools in one thread
// and asserts the wildcard renderer paints the full card sequence for
// every pill (1 + 5 + 3 = 9 cards).
test("sequential pills in one thread render full card sequences for each", async ({
page,
}) => {
// Three sequential pills × multi-tool chains × LLM-mock latency easily
// exceeds Playwright's 30s default. Bumped to cover the worst case.
test.setTimeout(240_000);
await page
.locator('[data-testid="copilot-suggestion"]')
.filter({ hasText: "Find flights" })
.first()
.click();
const cards = page.locator('[data-testid="custom-wildcard-card"]');
await expect
.poll(async () => cards.count(), { timeout: TOOL_TIMEOUT })
.toBe(1);
await page
.locator('[data-testid="copilot-suggestion"]')
.filter({ hasText: "Roll a d20" })
.first()
.click();
// 1 (flights) + 5 (d20) = 6 cards once d20 chain finishes.
await expect
.poll(async () => cards.count(), { timeout: TOOL_TIMEOUT })
.toBe(6);
await expect(page.getByText("Rolled the d20 five times")).toBeVisible({
timeout: TOOL_TIMEOUT,
});
await page
.locator('[data-testid="copilot-suggestion"]')
.filter({ hasText: "Chain tools" })
.first()
.click();
// 1 + 5 + 3 = 9 once chain tools mounts get_weather + search_flights +
// roll_d20 cards.
await expect
.poll(async () => cards.count(), { timeout: TOOL_TIMEOUT })
.toBe(9);
await expect(page.getByText("Done — Tokyo is sunny")).toBeVisible({
timeout: TOOL_TIMEOUT,
});
});
});