Files
copilotkit__copilotkit/showcase/integrations/ms-agent-python/tests/e2e/frontend-tools-async.spec.ts
Alem Tuzlak 5cd77233b7 feat(showcase/ms-agent-python): LGP parity sweep — 33/37 cells green
Brings ms-agent-python to one-to-one parity with langgraph-python (the D5
north star). Playwright e2e suite goes from 49/108 (~26%) → 164/178 (~92%),
33 of 37 cells fully green.

Manifest parity:
- Drop 4 MAF-only cells with no LGP analog: agentic-chat-reasoning,
  hitl-in-chat-booking, shared-state-write, reasoning-default-render.
  Reasoning is handled by reasoning-default + reasoning-custom (LGP);
  booking pill folds into hitl-in-chat; shared-state-write was a TODO stub.
- Rename byoc-hashbrown → declarative-hashbrown and byoc-json-render →
  declarative-json-render. Demo dir, API route dir, and frontend agent id
  follow LGP's naming. Python module files retain the legacy `byoc_*`
  prefix and FastAPI paths stay `/byoc-hashbrown` / `/byoc-json-render`
  (matches LGP's "module name retains legacy graph id" convention).
- Port LGP `_shared/`, `_shared/interrupt-fallback-slots.ts`, and
  `demos/layout.tsx` for one-to-one parity.

Cells ported verbatim from LGP (page + spec):
- agentic-chat, auth, beautiful-chat, chat-customization-css, chat-slots,
  declarative-gen-ui, declarative-hashbrown, declarative-json-render,
  frontend-tools, frontend-tools-async, gen-ui-agent, gen-ui-interrupt,
  gen-ui-tool-based, headless-complete, headless-simple, hitl-in-app,
  hitl-in-chat, shared-state-read, shared-state-read-write,
  shared-state-streaming, subagents, tool-rendering, plus all four
  tool-rendering* variants, a2ui-fixed-schema, agent-config, mcp-apps,
  multimodal, open-gen-ui, open-gen-ui-advanced, prebuilt-popup,
  prebuilt-sidebar, readonly-state-agent-context, reasoning-default,
  reasoning-custom, voice.

Backend infrastructure:
- Swap shared `OpenAIChatClient` (Responses API) → `OpenAIChatCompletionClient`
  (ChatCompletions). Root cause of the cross-cell post-tool ChatClientException
  family: Responses API is stateful and only sends NEW items per leg,
  relying on `previous_response_id` for history. aimock has no view of
  that server-side state, so second-leg requests arrived without the
  user message — fixture matchers keyed on `userMessage` couldn't fire
  and the run fell through to real OpenAI. ChatCompletions sends full
  history every leg, matching the LGP wire shape.
- Bump @ag-ui/client ^0.0.43 → ^0.0.53 (matches google-adk/LGP). Fixes
  the REASONING_* Zod discriminator trap on the catch-all agent.
- Regenerate package-lock.json in isolation outside the pnpm monorepo so
  npm-arborist doesn't resolve transitives against pnpm's hoisted
  symlinks (avoid 40+ `../../../node_modules/.pnpm/...` paths in the
  lockfile that break `npm ci` inside Docker).
- Add `yaml` (^2.8.4) for the new `src/app/demos/layout.tsx` that reads
  manifest.yaml for per-cell page titles (LGP parity).

New / re-added MAF agent backends with LGP-equivalent behavior:
- reasoning_agent.py (uses Responses API explicitly — the only chat
  client that emits AG-UI REASONING_MESSAGE_* events; rest of the
  integration stays on ChatCompletions).
- tool_rendering_agent.py (non-reasoning sibling of the existing
  reasoning_chain variant; shares tool surface via direct imports so
  they can never drift apart; routes the three catchall cells to a
  non-reasoning backend so the default renderer spec stops failing on
  leaked reasoning blocks).
- gen_ui_agent.py — `set_steps` tool + `steps` state schema +
  `predict_state_config` mirrors LGP's StateStreamingMiddleware shape.
- shared_state_streaming.py — `write_document` tool with
  `predict_state_config` that streams the `document` arg into
  `state.document` per-token.
- readonly_state_agent_context.py — minimal agent that consumes
  frontend-provided `useAgentContext` entries; no tools.
- headless_complete_agent.py — three deterministic tools (`get_weather`,
  `get_stock_price`, `get_revenue_chart`) mounted at /headless-complete
  on the mcp-apps runtime (was routing to catch-all sales agent, which
  returned seeded-random weather instead of the deterministic 68°F the
  test asserts on).

Wiring:
- copilotkit/route.ts: register the new agents, drop the stale
  shared-state-write entry, route all three tool-rendering variants to
  the non-reasoning backend (the reasoning-chain cell keeps its own
  dedicated path), register reasoning-default + reasoning-custom on
  /reasoning, register gen-ui-agent on /gen-ui-agent,
  shared-state-streaming on /shared-state-streaming,
  readonly-state-agent-context on its dedicated path.
- copilotkit-mcp-apps/route.ts: register headless-complete agent (was
  missing — the strict useAgent runtime sync in the newer
  @copilotkit/react-core surfaced the gap).
- copilotkit-declarative-hashbrown/route.ts + copilotkit-declarative-json-render/route.ts:
  new dedicated runtimes; agent IDs and runtime URLs follow LGP.
- copilotkit-declarative-gen-ui/route.ts: drop non-LGP `openGenerativeUI:
  false` for parity.

A2UI tool rename — `render_a2ui` → `_design_a2ui_surface`:
- Ported LGP's `tools/generate_a2ui.py` (LGP renamed the secondary-LLM
  tool to `_design_a2ui_surface` to avoid the A2UI middleware's bypass;
  shared d5-all.json fixtures key the response on this name).
- Renamed every `render_a2ui` occurrence in src/agents/{a2ui_dynamic,
  agent,beautiful_chat}.py and `tools/__init__.py`.
- Updated 4 declarative-gen-ui aimock fixtures to pass `context` arg in
  the first-leg `generate_a2ui` tool call (agent_framework doesn't
  auto-inject AgentSession into our @tool function so `session=None` and
  the secondary-LLM `user_content` was defaulting to a catch-all string
  containing "KPI dashboard" — every pill matched the KPI fixture).

Aimock router patch persisted alongside the integration changes:
hasToolResult matcher restricted to scan only messages after the last
user message (was global). The patch lives in F:/projects/cpk/aimock —
upstream PR pending.

Test infrastructure:
- playwright.config.ts: cap local workers at 4 + retries at 1. CI keeps
  workers=1, retries=2. `agent_framework.Agent` is reused across requests
  and the shared OpenAI HTTP client serialises concurrent SSE streams;
  >4 workers makes 30s timeouts inevitable on a few cells. Confirmed
  with hard data: workers=1 = 164 passed (16.8 min), workers=4+retries=1
  = 164 passed (7.2 min), workers=undefined = 159 passed. Same green
  set, ~2x faster. Long-term upstream fix is per-request Agent
  instantiation in agent_framework_ag_ui.

Remaining 14 failures across 4 cells documented per-cell in the Notion
D5 sweep doc (declarative-gen-ui A2UI surface mounting, multimodal
attachment forwarding, tool-rendering-default-catchall multi-pill chain,
tool-rendering-reasoning-chain multi-leg chains). Each has a specific
next-pass action.
2026-05-19 18:36:01 +02:00

187 lines
7.7 KiB
TypeScript
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
import { test, expect } from "@playwright/test";
// QA reference: qa/frontend-tools-async.md
// Demo source: src/app/demos/frontend-tools-async/{page.tsx, notes-card.tsx}
//
// The demo registers ONE async frontend tool via `useFrontendTool`:
// `query_notes(keyword: string)`. The handler sleeps 500ms (simulated local
// DB latency) then returns up to 5 matches from an in-memory 7-note DB.
// A custom `render` mounts `NotesCard` which exposes:
// - `data-testid="notes-card"` (outer container)
// - `data-testid="notes-keyword"` (heading: `Matching "<keyword>"`)
// - `data-testid="notes-list"` (the <ul> of matches)
// - `data-testid="note-n1"` … `note-n7` per-note rows
//
// Genuine-pass strategy: the deterministic aimock fixtures match each pill's
// verbatim prompt with a dedicated `query_notes(keyword=…)` tool call so the
// async handler runs against the real client-side NOTES_DB. The card's
// `keyword` heading is then the keyword we asserted in the fixture, and the
// `notes-list` rows reflect the actual handler-filtered results — proving
// the async tool round-trip end-to-end.
test.describe("Frontend Tools (async query_notes)", () => {
test.setTimeout(120_000);
test.beforeEach(async ({ page }) => {
await page.goto("/demos/frontend-tools-async");
});
test("page loads with composer and 3 pills", async ({ page }) => {
await expect(page.getByPlaceholder("Type a message")).toBeVisible();
await expect(
page.getByRole("button", { name: /Find project-planning notes/i }),
).toBeVisible({ timeout: 15_000 });
await expect(
page.getByRole("button", { name: /Search for 'auth'/i }),
).toBeVisible({ timeout: 15_000 });
await expect(
page.getByRole("button", { name: /What do I have about reading\?/i }),
).toBeVisible({ timeout: 15_000 });
});
test("project-planning pill → Notes DB card with project-planning notes", async ({
page,
}) => {
await page
.getByRole("button", { name: /Find project-planning notes/i })
.click();
const notesCard = page.locator('[data-testid="notes-card"]').first();
await expect(notesCard).toBeVisible({ timeout: 60_000 });
// The keyword heading proves the async handler resolved against the
// fixture-emitted `query_notes(keyword="project planning")` call.
await expect(notesCard.locator('[data-testid="notes-keyword"]')).toHaveText(
/Matching\s+["“]project planning["”]/i,
{ timeout: 30_000 },
);
// The async handler matches notes n1 ("Q2 project planning kickoff")
// and n5 ("Project planning retrospective notes") from NOTES_DB.
const list = notesCard.locator('[data-testid="notes-list"]');
await expect(list).toBeVisible({ timeout: 30_000 });
await expect(notesCard.locator('[data-testid="note-n1"]')).toBeVisible();
await expect(notesCard.locator('[data-testid="note-n5"]')).toBeVisible();
// Anti-regression: the generic-plan boilerplate from the cross-cell
// catch-all fixture must NOT appear. If it does, the d5-all.json
// fixture lost match priority to feature-parity.json's "plan" entry.
await expect(
page.getByText("Research the topic, Outline key points"),
).toHaveCount(0);
});
test("auth pill → Notes DB card with auth-related notes", async ({
page,
}) => {
await page.getByRole("button", { name: /Search for 'auth'/i }).click();
const notesCard = page.locator('[data-testid="notes-card"]').first();
await expect(notesCard).toBeVisible({ timeout: 60_000 });
await expect(notesCard.locator('[data-testid="notes-keyword"]')).toHaveText(
/Matching\s+["“]auth["”]/i,
{ timeout: 30_000 },
);
// The async handler matches note n2 ("Planning: migrate auth to
// passkeys") on the "auth" tag.
const list = notesCard.locator('[data-testid="notes-list"]');
await expect(list).toBeVisible({ timeout: 30_000 });
await expect(notesCard.locator('[data-testid="note-n2"]')).toBeVisible();
// Anti-regression: the showcase-assistant catch-all from
// feature-parity.json must NOT have intercepted this prompt.
await expect(page.getByText("I'm your showcase assistant")).toHaveCount(0);
});
test("reading pill → Notes DB card with Book recommendations + locked narration", async ({
page,
}) => {
await page
.getByRole("button", { name: /What do I have about reading\?/i })
.click();
const notesCard = page.locator('[data-testid="notes-card"]').first();
await expect(notesCard).toBeVisible({ timeout: 60_000 });
// Keyword heading + match count + per-note testid + content +
// tag chip — the full canonical shape per spec test #4.
await expect(notesCard.locator('[data-testid="notes-keyword"]')).toHaveText(
/Matching\s+["“]reading["”]/i,
{ timeout: 30_000 },
);
await expect(notesCard.getByText("1 match", { exact: false })).toBeVisible({
timeout: 30_000,
});
const note = notesCard.locator('[data-testid="note-n4"]');
await expect(note).toBeVisible({ timeout: 30_000 });
await expect(note.getByText("Book recommendations")).toBeVisible();
await expect(note.getByText(/Thinking Fast and Slow/i)).toBeVisible();
await expect(
note.getByText(/The Design of Everyday Things/i),
).toBeVisible();
await expect(note.getByText("reading", { exact: true })).toBeVisible();
// Locked narration leading phrase — proves the deterministic 2nd-turn
// fixture wired correctly through the async tool result.
await expect(
page
.locator('[data-testid="copilot-assistant-message"]')
.filter({
hasText:
'You have a note titled "Book recommendations" that is tagged with "reading',
})
.first(),
).toBeVisible({ timeout: 60_000 });
});
// Regression for the aimock multi-pill bug:
// The three frontend-tools-async fixtures used `hasToolResult: false/true`
// gates to split first-turn (emit `query_notes`) vs. follow-up (narration).
// After the user clicked a tool-using pill earlier in the same thread, the
// first-turn fixture was skipped (the thread already had a prior tool
// result), the follow-up fixture fired immediately with just narration,
// and the Notes DB card never rendered. Fix: chain via `toolCallId`, drop
// the gates. This test drives all three pills in a single thread and
// asserts every pill renders its own Notes DB card.
test("sequential pills in one thread each render their own Notes DB card", async ({
page,
}) => {
// Three pills × async-handler latency × LLM mock chain; the existing
// describe-level 120s is not enough once we drive all three in one test.
test.setTimeout(240_000);
const cards = page.locator('[data-testid="notes-card"]');
await page
.getByRole("button", { name: /Find project-planning notes/i })
.click();
await expect.poll(() => cards.count(), { timeout: 60_000 }).toBe(1);
await expect(
page.locator('[data-testid="notes-keyword"]', {
hasText: /Matching\s+[""“]project planning[""”]/i,
}),
).toBeVisible({ timeout: 60_000 });
await page.getByRole("button", { name: /Search for 'auth'/i }).click();
await expect.poll(() => cards.count(), { timeout: 60_000 }).toBe(2);
await expect(
page.locator('[data-testid="notes-keyword"]', {
hasText: /Matching\s+[""“]auth[""”]/i,
}),
).toBeVisible({ timeout: 60_000 });
await page
.getByRole("button", { name: /What do I have about reading\?/i })
.click();
await expect.poll(() => cards.count(), { timeout: 60_000 }).toBe(3);
await expect(
page.locator('[data-testid="notes-keyword"]', {
hasText: /Matching\s+[""“]reading[""”]/i,
}),
).toBeVisible({ timeout: 60_000 });
});
});