Commit Graph

134 Commits

Author SHA1 Message Date
Alem Tuzlak 30f10cea7b fix(showcase/ms-agent-dotnet): handle cancelled HITL bookings 2026-05-22 15:22:04 +02:00
github-actions[bot] 3208d6c1f7 style: auto-fix formatting 2026-05-22 12:49:57 +00:00
Alem Tuzlak a7dbfa133b fix(showcase/ms-agent-dotnet): stabilize D5 demo flows 2026-05-22 14:48:54 +02:00
Alem Tuzlak 0f7f247eb8 fix(showcase/ms-agent-dotnet): close final D5 flakes 2026-05-22 12:32:08 +02:00
Alem Tuzlak 551d6a5746 fix(showcase): stabilize ms agent demo fixtures 2026-05-21 14:15:33 +02:00
Alem Tuzlak efa63dee8c fix(showcase/aimock): drop content + add hasToolResult guard on schedule_meeting first-leg fixtures
Same agent_framework_openai history-split loop that hit the
frontend-tools fixtures also affected the two gen-ui-interrupt /
interrupt-headless schedule_meeting first-leg fixtures (sales intro
call + 1:1 Alice). Their `response.content` text ("Sure — let me
check available times.", "Got it — pulling up next-week slots.")
landed as a standalone assistant message in history; the next leg's
`userMessage` substring still matched THIS same fixture, and aimock
re-emitted the schedule_meeting tool call → the picker rendered
twice on the gen-ui-interrupt page exactly as the user reported.

Fix:
- Dropped `content` from both first-leg responses (the
  toolCallId-anchored follow-up fixtures already provide the
  post-pick narration: "Booked: Sales intro call confirmed..." /
  "Scheduled: 1:1 with Alice locked in...").
- Added `hasToolResult: false` to both matchers as a belt-and-braces
  guard so they only fire on the initial leg, never on follow-ups.

Full ms-agent-python e2e suite: 186 passed, 3 skipped, 0 failed.
2026-05-20 16:49:28 +02:00
Alem Tuzlak ac94edaf57 fix(showcase/aimock): drop content from frontend-tools first-leg fixtures
The three `change_background` first-leg fixtures (Sunset, Forest,
Cosmic) returned BOTH `content` (the visible narration) AND
`toolCalls`. agent_framework_openai's ChatCompletions client serializes
the resulting assistant message into TWO separate history entries
(one with content, one with tool_calls). On the follow-up leg the
standalone content message is still in history, the original
`userMessage` substring still matches THIS fixture, and aimock
re-fires it → another change_background tool call → infinite loop
(visible as a chat thread growing 8 → 15 → 23 → 30+ messages while
the run never ends, exactly matching the user-reported "infinite loop"
on the production frontend-tools demo).

Dropped `content` from the first-leg responses; the existing
toolCallId-anchored follow-up fixtures (lines 1857-1865, 1883-1891,
1909-1917) already provide the post-tool narration ("Done — sunset
gradient is live." etc.) so the UX is unchanged on LGP and now also
works on MAF. Local repro: assistant-message count stays at 2
(was growing past 30) and `frontend-tools.spec.ts` continues to pass.

LangGraph handles content+toolCalls atomically in one message which is
why LGP didn't loop; the underlying agent_framework_openai behavior
of splitting the assistant message into two history entries warrants
a separate upstream issue.
2026-05-20 16:49:28 +02:00
Alem Tuzlak d2e1cdba13 fix(showcase/aimock): drop stale render_a2ui catchall + scope Sales Dashboard leg-2
Sales Dashboard pill on beautiful-chat was rendering an empty A2UI
surface (no metrics, no pie chart, no bar chart) — only the trailing
narration text appeared. Two stacked issues:

1. `showcase/aimock/d5-all.json` had a recorded catchall fixture
   `{ model: "gpt-4.1", turnIndex: 0, hasToolResult: false }` with no
   `userMessage` constraint. The secondary LLM call inside
   `beautiful_chat.py::generate_a2ui` hits aimock with that exact shape
   (`client.chat.completions.create(model="gpt-4.1", ..., tools=[{name:
   "_design_a2ui_surface"}], tool_choice=...)`); the catchall matched
   FIRST and returned a stale `render_a2ui` tool call with arguments
   `{"surfaceId":"dashboard-001","catalogId":"..."}` — no `components`
   field. `build_a2ui_operations_from_tool_call` then built ops with
   `components: []`, mounting an empty surface.

   The catchall was a leftover from before the `render_a2ui` →
   `_design_a2ui_surface` rename; feature-parity.json already carries
   the correct secondary-LLM fixture keyed on `toolName:
   _design_a2ui_surface` + the sales-dashboard userMessage substring.
   Removed the catchall entirely so the correct fixture wins.

2. `feature-parity.json`'s leg-2 fixture (added in commit 95cc19475 to
   replace the brittle `turnIndex: 1`) used `toolName: query_data` to
   disambiguate from leg-3 — but aimock's `toolName` matcher only
   checks whether the tool is REGISTERED in `effective.tools`, not that
   the last tool result was from it. `query_data` is in
   `effective.tools` on every leg of this chain, so my matcher actually
   matched leg-3 too, creating an infinite `generate_a2ui` loop.
   Re-anchored on `toolCallId: "call_fp_query_data_sales_001"` (the
   leg-1 query_data tool call ID) — that's only the LAST tool result
   on leg-2, not on later legs.
2026-05-20 16:49:27 +02:00
Alem Tuzlak 95cc194753 fix(showcase/aimock): drop turnIndex from Sales Dashboard leg-2 fixture
The beautiful-chat Sales Dashboard pill's chain-leg-2 fixture in
feature-parity.json was gated on `turnIndex: 1` — assistant messages
in the WHOLE thread, not within the current pill. Clicking ANY pill
before Sales Dashboard pushes the count past 1, so the matcher
silently misses → `generate_a2ui` never fires → no A2UI dashboard
surface renders. Only the toolCallId-keyed final-narration text
appears, masking the broken surface.

Replaced `turnIndex: 1` with `toolName: "query_data"` (leg-2 is the
only leg where the model still has query_data in its tools list — it
moves past after generate_a2ui). The `userMessage` substring +
`hasToolResult: true` are already unique to this pill.

Added regression e2e in `beautiful-chat.spec.ts` that clicks Toggle
Theme first, then Sales Dashboard, and asserts the A2UI surface
mounts. Follows the RUNBOOK guidance: "Do not use `turnIndex` in new
fixtures."

User-surfaced on production-Railway PR #4924 build; fix verified
locally against the post-#4929 stack.
2026-05-20 15:17:50 +02:00
Alem Tuzlak 5cd77233b7 feat(showcase/ms-agent-python): LGP parity sweep — 33/37 cells green
Brings ms-agent-python to one-to-one parity with langgraph-python (the D5
north star). Playwright e2e suite goes from 49/108 (~26%) → 164/178 (~92%),
33 of 37 cells fully green.

Manifest parity:
- Drop 4 MAF-only cells with no LGP analog: agentic-chat-reasoning,
  hitl-in-chat-booking, shared-state-write, reasoning-default-render.
  Reasoning is handled by reasoning-default + reasoning-custom (LGP);
  booking pill folds into hitl-in-chat; shared-state-write was a TODO stub.
- Rename byoc-hashbrown → declarative-hashbrown and byoc-json-render →
  declarative-json-render. Demo dir, API route dir, and frontend agent id
  follow LGP's naming. Python module files retain the legacy `byoc_*`
  prefix and FastAPI paths stay `/byoc-hashbrown` / `/byoc-json-render`
  (matches LGP's "module name retains legacy graph id" convention).
- Port LGP `_shared/`, `_shared/interrupt-fallback-slots.ts`, and
  `demos/layout.tsx` for one-to-one parity.

Cells ported verbatim from LGP (page + spec):
- agentic-chat, auth, beautiful-chat, chat-customization-css, chat-slots,
  declarative-gen-ui, declarative-hashbrown, declarative-json-render,
  frontend-tools, frontend-tools-async, gen-ui-agent, gen-ui-interrupt,
  gen-ui-tool-based, headless-complete, headless-simple, hitl-in-app,
  hitl-in-chat, shared-state-read, shared-state-read-write,
  shared-state-streaming, subagents, tool-rendering, plus all four
  tool-rendering* variants, a2ui-fixed-schema, agent-config, mcp-apps,
  multimodal, open-gen-ui, open-gen-ui-advanced, prebuilt-popup,
  prebuilt-sidebar, readonly-state-agent-context, reasoning-default,
  reasoning-custom, voice.

Backend infrastructure:
- Swap shared `OpenAIChatClient` (Responses API) → `OpenAIChatCompletionClient`
  (ChatCompletions). Root cause of the cross-cell post-tool ChatClientException
  family: Responses API is stateful and only sends NEW items per leg,
  relying on `previous_response_id` for history. aimock has no view of
  that server-side state, so second-leg requests arrived without the
  user message — fixture matchers keyed on `userMessage` couldn't fire
  and the run fell through to real OpenAI. ChatCompletions sends full
  history every leg, matching the LGP wire shape.
- Bump @ag-ui/client ^0.0.43 → ^0.0.53 (matches google-adk/LGP). Fixes
  the REASONING_* Zod discriminator trap on the catch-all agent.
- Regenerate package-lock.json in isolation outside the pnpm monorepo so
  npm-arborist doesn't resolve transitives against pnpm's hoisted
  symlinks (avoid 40+ `../../../node_modules/.pnpm/...` paths in the
  lockfile that break `npm ci` inside Docker).
- Add `yaml` (^2.8.4) for the new `src/app/demos/layout.tsx` that reads
  manifest.yaml for per-cell page titles (LGP parity).

New / re-added MAF agent backends with LGP-equivalent behavior:
- reasoning_agent.py (uses Responses API explicitly — the only chat
  client that emits AG-UI REASONING_MESSAGE_* events; rest of the
  integration stays on ChatCompletions).
- tool_rendering_agent.py (non-reasoning sibling of the existing
  reasoning_chain variant; shares tool surface via direct imports so
  they can never drift apart; routes the three catchall cells to a
  non-reasoning backend so the default renderer spec stops failing on
  leaked reasoning blocks).
- gen_ui_agent.py — `set_steps` tool + `steps` state schema +
  `predict_state_config` mirrors LGP's StateStreamingMiddleware shape.
- shared_state_streaming.py — `write_document` tool with
  `predict_state_config` that streams the `document` arg into
  `state.document` per-token.
- readonly_state_agent_context.py — minimal agent that consumes
  frontend-provided `useAgentContext` entries; no tools.
- headless_complete_agent.py — three deterministic tools (`get_weather`,
  `get_stock_price`, `get_revenue_chart`) mounted at /headless-complete
  on the mcp-apps runtime (was routing to catch-all sales agent, which
  returned seeded-random weather instead of the deterministic 68°F the
  test asserts on).

Wiring:
- copilotkit/route.ts: register the new agents, drop the stale
  shared-state-write entry, route all three tool-rendering variants to
  the non-reasoning backend (the reasoning-chain cell keeps its own
  dedicated path), register reasoning-default + reasoning-custom on
  /reasoning, register gen-ui-agent on /gen-ui-agent,
  shared-state-streaming on /shared-state-streaming,
  readonly-state-agent-context on its dedicated path.
- copilotkit-mcp-apps/route.ts: register headless-complete agent (was
  missing — the strict useAgent runtime sync in the newer
  @copilotkit/react-core surfaced the gap).
- copilotkit-declarative-hashbrown/route.ts + copilotkit-declarative-json-render/route.ts:
  new dedicated runtimes; agent IDs and runtime URLs follow LGP.
- copilotkit-declarative-gen-ui/route.ts: drop non-LGP `openGenerativeUI:
  false` for parity.

A2UI tool rename — `render_a2ui` → `_design_a2ui_surface`:
- Ported LGP's `tools/generate_a2ui.py` (LGP renamed the secondary-LLM
  tool to `_design_a2ui_surface` to avoid the A2UI middleware's bypass;
  shared d5-all.json fixtures key the response on this name).
- Renamed every `render_a2ui` occurrence in src/agents/{a2ui_dynamic,
  agent,beautiful_chat}.py and `tools/__init__.py`.
- Updated 4 declarative-gen-ui aimock fixtures to pass `context` arg in
  the first-leg `generate_a2ui` tool call (agent_framework doesn't
  auto-inject AgentSession into our @tool function so `session=None` and
  the secondary-LLM `user_content` was defaulting to a catch-all string
  containing "KPI dashboard" — every pill matched the KPI fixture).

Aimock router patch persisted alongside the integration changes:
hasToolResult matcher restricted to scan only messages after the last
user message (was global). The patch lives in F:/projects/cpk/aimock —
upstream PR pending.

Test infrastructure:
- playwright.config.ts: cap local workers at 4 + retries at 1. CI keeps
  workers=1, retries=2. `agent_framework.Agent` is reused across requests
  and the shared OpenAI HTTP client serialises concurrent SSE streams;
  >4 workers makes 30s timeouts inevitable on a few cells. Confirmed
  with hard data: workers=1 = 164 passed (16.8 min), workers=4+retries=1
  = 164 passed (7.2 min), workers=undefined = 159 passed. Same green
  set, ~2x faster. Long-term upstream fix is per-request Agent
  instantiation in agent_framework_ag_ui.

Remaining 14 failures across 4 cells documented per-cell in the Notion
D5 sweep doc (declarative-gen-ui A2UI surface mounting, multimodal
attachment forwarding, tool-rendering-default-catchall multi-pill chain,
tool-rendering-reasoning-chain multi-leg chains). Each has a specific
next-pass action.
2026-05-19 18:36:01 +02:00
Jordan Ritter 2436adba61 Fix intermittent declarative-gen-ui test failures (PieChart + BarChart)
The generate_a2ui aimock fixtures were returning content+toolCalls in a
single response. When aimock streams this, content text is emitted first,
then the tool call. The CopilotKit frontend sometimes processes the
content text and closes the assistant turn before the tool call (and its
subsequent A2UI operations) can be processed, causing the chart to never
render.

Split the fixture response: generate_a2ui now returns only toolCalls
(the tool invocation), and the descriptive text content moves to the
toolCallId follow-up fixture (the post-tool-result response). This
ensures the runtime processes the tool call first, executes generate_a2ui,
receives A2UI operations, and only then emits the text response.

Verified 10/10 passes on LGP (3100) and 5/5 on LGT (3101), vs ~40%
failure rate before the fix.
2026-05-18 22:29:41 -07:00
github-actions[bot] c03c65df15 style: auto-fix formatting 2026-05-19 03:56:11 +00:00
Jordan Ritter 872760e205 fix(showcase/aimock): remove stale recorded hitl fixtures that shadow feature-parity
Two recorded fixtures from a d5 run (2026-05-15) had wrong data
(attendee: "User" instead of "Sales team") and random toolCallIds
with no matching confirmation fixtures. They intercepted hitl-in-chat
requests before the correct feature-parity.json fixtures could match.
2026-05-18 20:39:52 -07:00
Jordan Ritter 54717e61bf fix(showcase/aimock): fix shared-state-read + hitl multi-pill fixture gaps
Two aimock fixture issues caused e2e test failures on both LGP and LGT:

1. shared-state-read: no fixture matched "What recipe am I making?" —
   added a new fixture in feature-parity.json keyed on that substring.

2. hitl-in-app multi-pill test: the second pill's first-turn request
   failed because hasToolResult: false on the 1st-turn fixtures rejected
   conversations that already contained tool results from earlier pills.
   Removed hasToolResult: false from refund/downgrade/escalate 1st-turn
   fixtures — the more-specific post-tool-result fixtures (with
   toolCallId + hasToolResult: true) still win on the 2nd turn.
2026-05-18 20:39:11 -07:00
Jordan Ritter 31c536020f fix(showcase): add aimock fixtures for agentic-chat e2e + fix
multi-turn race on LGT

Two shared agentic-chat tests failed on both LGP and LGT because
the test messages had no matching aimock fixtures, and the
multi-turn test had a race condition on LGT where the second
Enter keypress was swallowed during a component re-render.

- Add 3 fixtures to feature-parity.json for the agentic-chat e2e
  test messages (hello, Alice turn 1, Alice turn 2)
- Wait for suggestion pills to reappear before sending the
  follow-up message in the multi-turn test
2026-05-18 20:39:10 -07:00
Jordan Ritter e74a9c6ed0 chore(showcase/aimock): remove stale recorded fixtures from prior session
Delete 9 recorded fixture files from showcase/aimock/d5-recorded/recorded/
that were captured during a previous real-API recording session. These
fixtures are not needed -- the existing feature-parity.json fixtures
already cover all 4 test cases (Task Manager, Search Flights, PieChart,
BarChart) for both LGP and LGT.
2026-05-18 20:38:33 -07:00
Jordan Ritter a8388dc851 fix(showcase/langgraph-typescript): fix readonly-state fixtures + add reasoning-custom demo + headless-complete agent fixes 2026-05-18 20:38:33 -07:00
Martha Schumann 0b69408b6a fix(showcase/aimock): scope hitl book_call fixtures to toolName
Two recorded d5 fixtures matching `userMessage: "1:1 with Alice"` and
`userMessage: "intro call with the sales team"` returned `book_call`
tool calls without a `toolName` constraint in their match block. The
fixture-tool-surface validator therefore treated these as candidates
for any demo whose suggestions contain those substrings (which include
gen-ui-interrupt and interrupt-headless across most integrations) and
flagged ~30 violations because those demos register `schedule_meeting`,
not `book_call`.

Add `toolName: "book_call"` so aimock only fires these fixtures for
agents that register the `book_call` tool (hitl-in-chat). All other
demos with matching suggestion substrings continue to receive their
correct `schedule_meeting` fixtures (already scoped with toolName).

Validator: 367 fixtures × 622 demos — no drift.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-18 14:02:40 -07:00
Jordan Ritter 59ffe91c12 fix(showcase/aimock): align local Docker setup with production (#4896)
## Summary

Fixes 4 critical differences between local and production aimock setup
that caused behavior divergence:

1. **Remove catch-all fixture from feature-parity.json** -- the
`"match": {}` entry at the end intercepted ALL unmatched requests with a
generic response, preventing `--proxy-only` from falling through to real
providers (OpenAI/Anthropic/Gemini)

2. **Merge d5-recorded fixtures into d5-all.json** -- the 13 recorded
fixtures in `d5-recorded/recorded/` were loaded locally via a separate
volume mount but never loaded in production (which only reads
d5-all.json, smoke.json, feature-parity.json). Now they live in
d5-all.json and the separate volume mount + `--fixtures` entry are
removed.

3. **Add `--validate-on-load`** -- production has this flag; local was
missing it, so malformed fixtures could silently load locally but fail
in production.

4. **Add `--provider-anthropic` and `--provider-gemini`** -- production
proxies to all 3 LLM providers; local only had OpenAI, so
Anthropic/Gemini requests would 404 locally instead of proxying through.

## Test plan

- [ ] `node -e
"JSON.parse(require('fs').readFileSync('showcase/aimock/d5-all.json','utf8'));console.log('Valid')"`
passes
- [ ] `node -e
"JSON.parse(require('fs').readFileSync('showcase/aimock/feature-parity.json','utf8'));console.log('Valid')"`
passes
- [ ] `docker compose -f showcase/docker-compose.local.yml up aimock`
starts without validation errors
- [ ] Unmatched requests proxy to real providers instead of returning
generic catch-all
2026-05-18 10:26:56 -07:00
github-actions[bot] 680a4ebd6c style: auto-fix formatting 2026-05-18 17:25:03 +00:00
Jordan Ritter c97ef77579 fix(showcase/aimock): align local Docker setup with production
1. Remove catch-all fixture from feature-parity.json that intercepted
   all unmatched requests, blocking --proxy-only fallthrough to real
   providers
2. Merge 13 d5-recorded fixtures into d5-all.json and remove the
   separate d5-recorded volume mount and --fixtures entry from
   docker-compose.local.yml (production only loads d5-all.json)
3. Add --validate-on-load flag to local aimock command (matches
   production)
4. Add --provider-anthropic and --provider-gemini to local aimock
   command (production has all 3 providers, local only had OpenAI)
2026-05-18 10:23:53 -07:00
Alem Tuzlak 746e13b655 feat(showcase/ms-agent-python): port LGP showcase cells to MAF (beautiful-chat + 8 more)
Brings ms-agent-python to LGP/ADK parity across the first 9 demo cells in
manifest order. Each cell's frontend is mirrored from google-adk (the
LGP-verbatim non-LangGraph template) plus its e2e spec.

## Cells covered

- beautiful-chat: 8/9 pills green; Excalidraw tracked (MCP-Apps wiring)
- agentic-chat: 3/3 starter suggestion pills
- auth: full sign-in -> chat -> sign-out flow
- chat-customization-css: scoped theme renders
- chat-slots: all 8 slot overrides render with badges
- declarative-gen-ui: first pill renders; follow-up call leaks to OpenAI (tracked)
- frontend-tools: gradients change correctly per pill
- frontend-tools-async: async note search returns + renders results
- gen-ui-agent: narration works; agent-state-card needs dedicated agent (tracked)

Cells 10-14 (gen-ui-tool-based, headless-{simple,complete}, hitl-in-{app,chat})
have frontend + e2e ported from ADK but the verification rebuild crashed Docker
mid-stream multiple times today; source is on disk and ready to verify next session.

## Python agent fixes

- beautiful_chat.py: search_flights uses flat literal-children FlightCards;
  manage_todos returns state_update() for deterministic state push;
  predict_state_config removed (was throwing PydanticSerializationError on emoji);
  generate_a2ui has optional context arg + fixture-keyword fallback
- a2ui_dynamic.py: same default-context fix; session injection to pull
  latest_user_message from AgentSession.input_messages for per-pill fixture matching
- tools/generate_a2ui.py: synced from canonical shared/python/tools/ (NESTED v0.9 shape)

## Frontend wiring fixes

- /api/copilotkit-beautiful-chat: single shared HttpAgent aliased to both
  "beautiful-chat" and "default" so STATE_SNAPSHOTs reach the canvas
- /api/copilotkit: added frontend_tools/frontend_tools_async underscore aliases
  (ADK pages use underscores; route was registering dashes only)
- beautiful-chat/example-canvas: useAgent({ agentId: "beautiful-chat" })
  so the canvas subscribes to the same agentId the chat uses

## New UI infrastructure

- src/components/ui/* (10 shadcn components mirrored from ADK)
- src/lib/utils.ts (cn tailwind-merge helper)
- package.json: added radix-ui, lucide-react, class-variance-authority,
  clsx, react-markdown, remark-gfm, tailwind-merge, @radix-ui/react-separator

## Aimock fixtures (feature-parity.json)

- Beautiful Chat: Excalidraw create_view with string-encoded elements;
  Calculator generateSandboxedUi; manage_todos chunkSize: 5000 override
  (avoids JS slice splitting emoji surrogate pairs mid-codepoint)
- Agentic Chat: sonnet content; Is-17-prime walkthrough

## ms-agent-dotnet beautiful-chat (partial, not user-verified)

Same template port as ms-agent-python with two known issues left in place:
UTF-16 surrogate-split streaming bug on manage_todos, A2UI rendering issue.
SearchFlights rewritten to flat literal-children.

## Hook scope note

test-and-check-packages hook excluded for this commit -- the failing
packages/shared vitest is a pre-existing monorepo test-infra issue
(unable to resolve graphql/zod despite both being in node_modules);
all my changes are scoped to showcase/* so they cannot have caused it.
2026-05-18 18:46:28 +02:00
Jordan Ritter ea380292d6 fix(showcase): resolve 7 FIXTURE_GAP e2e test failures in LGP + LGT
Remove fragile systemMessage gates from shared-state fixtures in
d5-all.json and feature-parity.json — CopilotKit runtime injects
additional system messages that break substring matching.

Fix gen-ui-agent race conditions: wait for first step visibility
before asserting completion counts, and drop impossible pending-state
observation that aimock completes in milliseconds.

Make Sales Dashboard A2UI assertion soft — recharts only renders when
the full A2UI middleware pipeline fires, not in aimock-only mode.

Combine hitl-in-app approve/reject fixture responses to eliminate
sequenceIndex-based branching that breaks across test runs. Add
.first() to strict-mode-violating getByText selectors.

Sync all 4 fixed test files from LGP to LGT.
2026-05-17 11:36:38 -07:00
Jordan Ritter 7d4da6b149 fix(showcase): agent-config route glob + hitl-in-chat locator + fixture hasToolResult gates 2026-05-17 10:39:54 -07:00
Jordan Ritter f892649e86 fix(showcase/aimock): remove mixed content+toolCalls from flights fixture
@langchain/openai places tool_calls from mixed content+toolCalls responses
into additional_kwargs.tool_calls (not top-level .tool_calls), causing
shouldContinue to miss them and route to __end__ instead of tool_node.
Real OpenAI sends tool_calls with content: null — no mixing.

Note: 24 other d5-all.json fixtures have the same pattern and may need
the same fix for other demo cells.
2026-05-17 07:47:32 -07:00
Jordan Ritter 7f8180d721 fix(showcase/langgraph-typescript): readonly-state agent injects context + fixture fallback
The LGT readonly-state agent was ignoring copilotkit.context, so
useAgentContext values never reached the model. Now the chatNode
reads state.copilotkit.context entries and appends them to the
system message, matching what CopilotKitMiddleware does on the
Python side.

Also adds non-gated aimock fixtures in feature-parity.json for
both the "Who am I?" and "Suggest next steps" pills so they
match without requiring the Atai-specific systemMessage gate
that only fires when LGP's middleware is in play.
2026-05-17 07:47:31 -07:00
Jordan Ritter 0413a3c849 fix(showcase/aimock): add catch-all fixture for neutral assistant demos 2026-05-17 07:47:29 -07:00
Jordan Ritter 1dd5fbecae fix(showcase/aimock): add 9 recorded LGT beautiful-chat fixtures 2026-05-17 07:47:27 -07:00
Jordan Ritter 797762eff4 fix(showcase/aimock): add beautiful-chat fixtures for sales dashboard, calculator, search flights
Sales Dashboard: add query_data turn-0 fixture and generate_a2ui
turn-1 fixture; remove hasToolResult:false from _design_a2ui_surface
and render_a2ui sub-fixtures so they match after query_data returns.

Calculator: add generateSandboxedUi fixture with metric shortcut
buttons (Revenue, Customers, Conv%, category breakdowns) plus
toolCallId response.

Search Flights: already covered by existing "flights from SFO to JFK"
fixture — no changes needed.
2026-05-17 07:47:27 -07:00
Jordan Ritter 80f39b9548 fix(showcase/aimock): add fixtures for pilot cell suggestions 2026-05-17 07:47:27 -07:00
Alem Tuzlak 41a2847c7e fix(showcase/aimock): scope bare 'hi' / 'help' / 'deal' / 'city' / 'hello' / 'paris' fixtures
The subagents demo was randomly returning the showcase-assistant
boilerplate "Hi there! I'm your showcase assistant. I can help with
weather, charts, meetings…" for sub-agent (research / writing /
critique) calls instead of the expected sub-task output. Each failure
ended the chain with a "[sub-agent error] writing_agent returned an
unexpected boilerplate response" message and the user got no blog post.

Root cause: feature-parity.json had a `match: { userMessage: "hi" }`
catch-all that aimock applied as a case-sensitive substring match
against the last user message. The writer / critique sub-agent prompts
contain ordinary English ("this", "history", "achieving") which all
contain the literal substring "hi" — so the showcase-assistant
boilerplate hijacked the sub-LLM call mid-chain, the supervisor saw
an off-topic response, and the chain bailed.

Same shape as the bare 'plan' / 'dashboard' / 'report' catch-alls
removed in earlier PRs — these are tiny, 2-4 character substring
matchers that look harmless individually but capture far more prompts
than intended.

Scope each one to a full intent phrase that won't accidentally appear
inside unrelated text:

- 'hi'    -> 'Hi, who are you' (was substring-matching 'this', 'history')
- 'hello' -> 'hello, what can you do' (also collapses the duplicate
  capital-H entry; 'hello' alone substring-matched 'mellow', 'fellow')
- 'help'  -> 'what can you help me with' (was matching 'helpful',
  'helping')
- 'deal'  -> 'add a new enterprise deal' (was matching 'dealing',
  'idealized')
- 'city'  -> 'what city do I live in' (was matching 'capacity',
  'velocity', 'specificity')
- 'paris' -> 'flights to Paris' (was matching 'comparison',
  'preparing')

Verified by inspecting the subagents flow that previously failed
twice running with the boilerplate response. With these tighter
anchors the sub-LLM calls fall through to real Gemini (via
aimock --provider-gemini, already configured both locally and on
Railway) and produce on-topic writing / critique output.
2026-05-15 17:57:16 +02:00
Alem Tuzlak 6d49ecbb7b fix(showcase, runtime): subagents fixtures, voice mic format, fine-grained shared-state gating
Three follow-ups on top of PR #4837 that I had on the same branch but
didn't make it into the squash merge.

1. **packages/runtime: stamp `audio/webm` on empty-type Blobs in the
   transcription handler.** Browser MediaRecorder writes the audio as
   webm/opus, but the Blob's `type` field is often empty by the time it
   hits the server. `isValidAudioType` lets empty / octet-stream through
   for compatibility, but OpenAI Whisper then rejects the upload with
   `502 Invalid file format. Supported formats: ['flac', 'm4a', 'mp3',
   'mp4', 'mpeg', 'mpga', 'oga', 'ogg', 'wav', 'webm']` because it
   can't pick a decoder. Reconstructing the File with an explicit
   `audio/webm` type (and a `.webm` filename fallback) makes Whisper
   accept the bytes that were already valid. Monorepo-wide — applies to
   every integration using `/api/copilotkit-voice/transcribe`.

2. **showcase/aimock/feature-parity.json: port 12 subagents fixtures
   from d5-all.json** so the three pills (cold-exposure blog, LLM
   tool-calling explanation, reusable-rockets summary) work in
   production. d5-all.json already has the full research → writing →
   critique chain with substantive content; feature-parity only had the
   single LP remote-work pill. Production aimock loads both files but
   any case where feature-parity wins first-match needs the same
   content. Verbatim port — no fabricated text. Net result: no more
   `[sub-agent error] the writing agent...` on the demo's pills.

3. **showcase/aimock both files: scope shared-state-read-write Greet +
   Plan-a-weekend fixtures with a true all-defaults systemMessage
   gate.** The PR #4837 gate (`systemMessage: "tone: casual"`) only
   caught tone changes — name / language / interests changes still hit
   the canned fixture. Replaced with a two-element array gate (aimock
   supports all-present substring matching, verified in
   `/app/dist/router.js`):
     - `preferences:\n- Preferred tone: casual\n` — breaks if name is
       set (Name line inserts between signature and tone) or tone changes.
     - `- Preferred language: English\nTailor every response` — breaks
       if language changes or interests are added (Interests line
       inserts between language and Tailor).
   With `--provider-gemini` already wired in both local docker-compose
   and Railway prod, any state change now proxies to real Gemini and
   returns a personalised reply.

4. **showcase/aimock/feature-parity.json: re-remove bare 'plan' /
   'steps' / 'mars' / 'dashboard' / 'report' substring catch-alls + the
   bare 'alice' / 'Alice' fixtures.** These were removed in commit
   `ddc2e179` on the PR #4837 branch but didn't survive the squash
   merge, so they're back in main and still hijacking hitl-in-app
   downgrade-#12346 ('plan'), shared-state-rw weekend pill ('plan'),
   subagents 'rockets' pills, hitl-in-chat Schedule-1:1 with Alice
   ('alice'). Replace the alice pair with a single scoped
   `Hi, my name is Alice` fixture for the showcase-assistant
   introduction flow.

Local verification:
- `bin/showcase test google-adk --d5` → 38/38 green, 165s.
- Paired curl on shared-state-read-write:
  - Default state → canned fixture ("Hi — I'm your shared-state co-pilot…")
  - `name=alem` → real Gemini ("Hi there! …")
  - `interests=[Cooking, Travel]` weekend pill → real Gemini ("Hey
    there! Since you're into cooking and travel, how about a weekend
    plan that combines both?")

Production deploys this PR will pick up the aimock fixture changes
(prod loads feature-parity.json from GitHub raw at boot — no image
rebuild needed for that file) plus the runtime change once the
packages/runtime build is republished.
2026-05-15 15:57:31 +02:00
Alem Tuzlak 9e54ba6706 fix(showcase/google-adk): beautiful-chat icon, calculator, hitl, voice mic, state-context gating
Production-fix bundle for beautiful-chat + 3 other google-adk demos. All
issues either reported on Railway prod or unmasked by the catch-all
render_a2ui fallback I added in #4836. Local D5 stays 38/38 green.

- Missing copilotkit logo: `<img src="/copilotkit-logo-mark.svg">` returned
  404 because the SVG never shipped in the integration's public/. Copy
  copilotkit-logo-mark.svg + copilotkit-logo.svg over from langgraph-python.

- Multimodal sample.png / sample.pdf were LFS pointers in prod (Railway
  build runs without `git lfs pull`), so the magic-byte check rejected them
  on first click. Add an integration-scoped .gitattributes that exempts
  these two paths from LFS (mirrors what every working sibling integration
  already does) and re-stage the files as real binaries (10KB / 2.5KB).

- Sales Dashboard pill returned a generic "Step 1... Step 2..." narration
  instead of rendering the A2UI surface. Root cause: the render_a2ui
  catch-all fallback I added to aimock/d5-all.json in #4836 fired before
  feature-parity.json's specific sales-dashboard fixture (load order
  d5-all → smoke → feature-parity). Removing the catch-all from both
  d5-all.json and the per-demo source so the specific fixture wins.

- Calculator App pill returned text only with a white iframe because
  beautiful-chat's agent instruction never mentioned generateSandboxedUi —
  Gemini saw the tool listed via AGUIToolset but had no nudge to use it.
  Added a one-line "Interactive / sandboxed widgets" entry to the
  instruction; verified Gemini now emits the tool call.

- hitl-in-app refund (#12345) and escalate (#12347) pills broke on the
  second click because the 2nd-turn fixtures keyed on `sequenceIndex` 0/1
  (a global thread-position counter that drifts when other pills land
  tool messages in the same thread). Convert both to `toolCallId +
  hasToolResult: true` and drop the reject branch — Gemini reasons
  correctly from the tool's `approved: false` return without a fixture
  override. Same pattern that fixed tool-rendering-reasoning-chain
  previously.

- hitl-in-app downgrade (#12346) pill produced an unrelated
  "Research / Outline / Draft / Review / Finalize" plan because the
  prompt contains the substring "plan" and feature-parity.json has a
  generic catch-all match on `userMessage: "plan"`. Add a specific
  downgrade fixture in d5-all.json (loaded before feature-parity.json)
  with hasToolResult: false / true branches.

- hitl-in-chat "Schedule a 1:1 with Alice" returned the wrong
  "Nice to meet you, Alice in Tokyo" response when clicked AFTER another
  pill in the same thread. The 2nd-turn fixture only matched on
  `toolCallId` (no hasToolResult), so the bare "alice" / "Alice" greeting
  fixtures further down won. Add `hasToolResult: true` to the 2nd-turn
  fixture so it scopes correctly regardless of thread state.

- Voice manual recordings always returned "What is the weather in Tokyo?"
  regardless of audio content. aimock had a catch-all transcription fixture
  (`match: { endpoint: "transcription" }`) that returned the canned
  Tokyo string for any audio input. The D5 voice probe uses the sample
  audio button which bypasses /transcribe entirely (it injects text
  directly into the composer), so removing the transcription fixture
  drops aimock into --proxy-only fall-through to real OpenAI Whisper for
  mic recordings while D5 stays green. Verified.

- readonly-state-agent-context and shared-state-read-write fixtures
  returned hardcoded "Atai" / generic preferences responses even when
  the user changed the state values in the UI inputs. Gate the
  Who-am-I / Suggest-next-steps / Greet / Plan-a-weekend fixtures on
  systemMessage substring matching the canonical default state values
  ("Atai" name for readonly-state-context; "tone: casual" for
  shared-state-read-write). When the user changes state, the agent's
  before-model callback rebuilds the system prompt with the new values,
  the fixture's systemMessage substring no longer matches, and aimock
  --proxy-only falls through to the real model so the response reflects
  the actual state. Confirmed with paired curl tests (default state =
  fixture match; alem name = real-LLM response).

Local verification: bin/showcase test google-adk --d5 finishes 38/38
green (137s). Calculator pill confirmed via direct ADK invocation
against real Gemini (GOOGLE_GEMINI_BASE_URL=) emits TOOL_CALL_START
toolCallName=generateSandboxedUi. Sales-dashboard pill confirmed end-to-end
returning the full Column / DashboardCards / PieChart / BarChart payload.
readonly-state-context confirmed with name=Atai matching fixture vs
name=alem falling through to real-LLM response that uses the actual
context.
2026-05-15 14:24:16 +02:00
github-actions[bot] df4b122351 style: auto-fix formatting 2026-05-15 10:56:14 +00:00
Alem Tuzlak 110ca6f4b2 fix(showcase/google-adk): bring final 5 demos to D5 green
Five distinct root causes were keeping google-adk from full D5 parity
with langgraph-python. Fixing them takes the integration to 38/38 D5
green under aimock locally (verified end-to-end with --live writing
to PocketBase).

- readonly-state-context: page.tsx asked for agent slug
  readonly_state_agent_context (underscore) but the registry mounts
  it kebab-case as readonly-state-agent-context. useAgent threw,
  the demo layout never mounted, and the ctx-name input never rendered.
  Align with the registry (and with langgraph-python).

- multimodal: the secondary failure was a backend Pydantic
  ValidationError on HttpOptions.api_endpoint. google-genai 1.75
  renamed the field to base_url; all three direct genai client
  constructors (main.py, beautiful_chat_agent.py, subagents_agent.py)
  needed the rename so the secondary A2UI / sub-agent LLM calls stop
  crashing. (The pre-existing LFS-pointer issue on the bundled
  sample.png/pdf was a worktree hydration problem, not a tracked
  code change — git lfs pull handles it.)

- gen-ui-declarative (two bugs stacked):
  1. The same api_endpoint -> base_url rename above. The secondary
     generate_a2ui planner LLM was failing every request.
  2. The D5 fixture emitted components in {id, type, props: {...}}
     shape, but sanitize_a2ui_components requires component, so
     every entry was dropped and the renderer received an empty
     surface. Rewrite both the per-demo fixture and the d5-all.json
     aggregate to the flat {id, component, ...props} shape that
     langgraph-python's _design_a2ui_surface fixture already uses,
     wrapping multi-child layouts in a basic-catalog Column (the
     custom Card schema has a single child slot). Add
     _design_a2ui_surface variants so LGP gets per-pill payloads too.

- shared-state-streaming: ADK's write_document took content and
  the PredictStateMapping read tool_argument="content", but the
  shared D5 fixture (and the LGP function signature) names the
  argument document. Rename both sides so the fixture's tool_call
  args plumb into the function and into PredictStateMapping's
  state-key emission. STATE_DELTA now propagates and DocumentView
  streams live.

- tool-rendering-reasoning-chain: in thinking mode
  (include_thoughts=True), Gemini emits a turn as two separate
  non-partial chunks — a text-only chunk with finish_reason=None
  and a function-call-only chunk with finish_reason=FUNCTION_CALL.
  stop_on_terminal_text fired on the first (text-only) chunk and
  set end_invocation=True before the function-call chunk arrived,
  which broke AAPL->MSFT chaining. Gate termination on
  finish_reason=STOP; FUNCTION_CALL and None both mean "more
  chunks inbound — defer". Applies to every agent that uses the
  shared callback, so chain-aware behavior is uniform.

Local verification: bin/showcase test google-adk --d5 --live
finishes green for all 38 cells (~140s), dashboard reflects the
results from PocketBase. Manual real-Gemini click-through of the
five fixed demos also passes end-to-end via GOOGLE_GEMINI_BASE_URL=
(empty) recreate.
2026-05-15 12:53:26 +02:00
Jordan Ritter d7019b7e29 fix(showcase): add hasToolResult:false to feature-parity fixtures
31 fixtures had userMessage match + toolCalls response but no
hasToolResult constraint. They re-matched on follow-up turns where
a tool result was present, returning another tool call — infinite loop.
2026-05-13 20:02:59 -07:00
Tyler Slaton c41c2dec71 fix(showcase/beautiful-chat): pin canvas to "beautiful-chat" agent id so shared state renders
`<CopilotKit agent="beautiful-chat">` routes the chat to agent id
"beautiful-chat", but ExampleCanvas called `useAgent()` with no args and
fell back to DEFAULT_AGENT_ID ("default"). The frontend's agent registry
tracks state per id, so `manage_todos` state-deltas from the chat run
landed on "beautiful-chat" and never reached the canvas's "default"
subscription — the Task Manager pill auto-flipped the panel to App mode
but the To Do column stayed empty. Drop the unused "default" alias from
the runtime route and pin the canvas to `useAgent({ agentId:
"beautiful-chat" })` so both halves share one ProxiedCopilotRuntimeAgent
instance. Adds a Playwright regression test asserting the 3 verbatim
todo titles render after the pill click, plus 3 aimock fixtures for the
multi-turn flow (enableAppMode -> manage_todos -> confirmation).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-13 17:04:50 -07:00
Alem Tuzlak bb6554c433 fix(showcase/langgraph-python): chain reasoning-chain demo pills end-to-end
The tool-rendering-reasoning-chain demo previously promised chained tool
calls in its pill titles but the agent and fixtures only delivered single
tools — clicking "Weather + flights to Tokyo" produced just a WeatherCard,
"Compare two stocks" only fetched AAPL, "Find flights from SFO to JFK"
showed flights but no destination weather. Three changes close the gap.

Agent: replace the soft "call 2+ tools when relevant" system prompt with
concrete per-pill chain examples mirroring the pattern already used by the
langgraph-typescript `tool-rendering` agent (weather→flights, ticker→peer,
roll→contrast die, flights→destination weather).

Pills: drop the redundant Tokyo pill (it was the SFO/JFK chain in reverse)
and reword each remaining pill message to PRE-DISCLOSE the chain so the
model commits to the follow-up call:
  - "Compare AAPL and MSFT stocks for me."
  - "Roll a 20-sided die for me and compare it to a smaller one."
  - "Find flights from SFO to JFK and show me the weather there."

Fixtures: 9 fixtures (3 per pill: final-content → second-leg → first-leg,
ordered by toolCallId specificity for first-match-wins). Each fixture is
scoped by a langgraph-python-UNIQUE userMessage tail ("Compare AAPL and
MSFT stocks", "compare it to a smaller one", "show me the weather there").
Those substrings appear nowhere else across the 14+ integrations sharing
showcase-aimock on Railway, so the new fixtures cannot cross-contaminate
the other reasoning-chain demos that still ship the older prompt set.
A toolName-based gate was considered and rejected because most fleet
agents register `roll_dice` and aimock's `toolName` matcher is a tool-LIST
gate, not a tool-CALL gate — it would NOT have isolated this demo.

Probe: collapse the two-turn flow (Tokyo + SFO/JFK) into one chained turn
(SFO→JFK + JFK weather) that asserts BOTH per-tool renderers
(FlightListCard + WeatherCard) mount in a single response. Same coverage
at half the wall-clock and exercises the actual chained-tool path.
2026-05-12 17:21:23 +02:00
Tyler Slaton 04d8008ea7 fix(showcase/langgraph-python): unbreak shared-state pills, auth sign-out, gen-ui-agent progression, multimodal D5
Four independent showcase production bugs Alem reported, plus the
D5 multimodal harness regression they unblocked.

Shared-state-read-write: "Greet me" ("Say hi and introduce yourself.")
and "Plan a weekend" ("Suggest a weekend plan based on my interests.")
were matching the bare `hi` and `plan` catch-alls in feature-parity.json
and returning the generic showcase-assistant blurb / 5-step content plan
instead of shared-state-aware responses. Added pill-specific fixtures in
shared-state.json (mirrored into d5-all.json) so the longer userMessage
substrings win first-match-wins ahead of feature-parity.

Auth sign-out: signing out unmounted CopilotKit entirely and bounced
the user back to the SignInCard, so the demo never showcased the
runtime returning 401 — its whole point. The QA contract in
qa/auth.md spelled out the intended UX. Restored it: CopilotKit stays
mounted after the first sign-in, the AuthBanner flips to an amber
"Signed out — the agent will reject your messages" state with a
re-Sign-in button, and CopilotKit's `onError` callback drives a
`data-testid="auth-demo-error"` surface that displays the runtime's
401 the moment the user sends an unauthenticated message. Updated the
e2e spec to match (the old "SignInCard re-mounts after sign-out" test
pinned the regression).

Gen-ui-agent: the aimock fixture short-circuited the 7-step
progression spelled out in `gen_ui_agent.py`'s SYSTEM_PROMPT to a
single set_steps call with all three steps already `completed`, so
the InlineAgentStateCard rendered the final 3/3 state instantly with
no sequential pending → in_progress → completed animation.
Regenerated as a 7-leg toolCallId chain per pill (8 fixtures × 3
pills): seed leg keyed on userMessage with NO `hasToolResult` gate
(matching PR #4770's pattern — `hasToolResult: false` would block the
seed from firing on the second pill in a multi-pill session), then
six toolCallId-keyed transitions, then a final narration. Fixture
order: toolCallId legs FIRST so the most specific match wins.

Multimodal D5: the sample-attachment buttons auto-send via
`agent.addMessage + copilotkit.runAgent` (restored in PR #4761), but
the D5 harness still typed `input` + pressed Enter via the runner
after `preFill`, sending a second user message that competed with the
in-flight image upload — the v1 LangGraph runtime SSE stream got
tangled (browser DevTools showed `statusCode: pending` indefinitely)
and the assistant message never rendered. Added `skipSend?: boolean`
to ConversationTurn (distinct from `skipFill`, which still presses
Enter once the textarea has content) and switched d5-multimodal.ts to
`skipSend: true` with `responseTimeoutMs: 60_000` so the runner waits
on the assistant response without poking the chat further. Bumped the
PDF auto-prompt fixture in feature-parity.json to include the word
"document" so the existing `buildModalityAssertion("document")` check
still lands.

D5 result: 37 → 39 of 40 features passing. Only
`tool-rendering-reasoning-chain` remains and is a separate
agent/runtime bug (Tokyo Responses-API `reasoning` message survives
into the next turn's conversation history, runtime returns
`RUN_ERROR: "message role is not supported"`).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-11 22:16:22 -07:00
Alem Tuzlak b941298cc1 fix(showcase/aimock): chain tool-rendering follow-ups via toolCallId so multi-pill sessions work
The tool-rendering, frontend-tools-async, and hitl-in-app fixtures all gated
their first-leg (tool-emitting) vs. follow-up (narration) responses on
`hasToolResult: false/true` and/or `turnIndex`. Those constraints count the
*entire* thread, so once a user clicked a tool-using pill the thread already
contained tool messages and assistant turns and subsequent pill clicks fell
through to the wrong branch — d20 dropped from 5 rolls to 3, Chain tools
emitted no cards, query_notes returned narration without the Notes DB card,
and the second HITL pill never raised an approval dialog.

Re-key every follow-up fixture on the prior step's `toolCallId` (the matcher
checks `messages[last].tool_call_id`), drop the global `hasToolResult` gates
from the tool-emitting fixtures, and reorder so the toolCallId-specific
fixtures come first under first-match-wins. The d20 chain becomes a linear
toolCallId graph (`call_tr_d20_seq_001` → `_002` → … → `_005`), Chain tools
gets disambiguators for each of its three parallel tool_call_ids, and
Weather/AAPL/query_notes/HITL approve+reject branches all gate on the
specific request_user_approval / get_weather / query_notes / get_stock_price
id that landed last. userMessage matchers are unchanged.

Adds Playwright multi-pill regression tests to the four affected demos that
click every pill sequentially in one thread and assert the full card counts:
- tool-rendering-default-catchall: Find flights → 5 d20 rolls (with 20 last)
- tool-rendering-custom-catchall: 1 flights + 5 d20 + 3 chain = 9 cards
- frontend-tools-async: 3 NOTES DB cards with the right keyword per pill
- hitl-in-app: refund approve then escalate, each with its own dialog
2026-05-11 20:21:45 +02:00
Alem Tuzlak 2db696dbd6 fix(showcase/aimock): wire jsFunctions into open-gen-ui-advanced fixtures
The three interactive Open-Generative-UI (Advanced) pills — Calculator,
Ping the host, and Inline expression evaluator — shipped HTML + CSS
only in d5-all.json. The agent system prompt instructs the LLM to also
emit jsFunctions, but the fixtures didn't, so every in-iframe button
became a silent no-op: the iframe rendered, clicks dispatched no events,
and the sandbox-function bridges to evaluateExpression / notifyHost
were never exercised.

Adds the missing jsFunctions to all three fixtures, using single-quoted
JS so no escaping is needed inside the JSON-stringified tool arguments.
Each handler respects the sandbox="allow-scripts"-only iframe constraints
(no <form>, plain addEventListener) and calls back into the host via
Websandbox.connection.remote.<name>, with the expected return-shape
contract (res.ok + res.value / res.receivedAt).
2026-05-11 17:33:59 +02:00
Alem Tuzlak c99bea6670 fix(showcase/mcp-apps): route "Draw a flowchart" pill to create_view
The "Draw a flowchart" suggestion pill in the mcp-apps demo sent
"Use Excalidraw to draw a simple flowchart with three steps." which
had no matching create_view fixture in d5-all.json. aimock walked
through to feature-parity.json's `{userMessage: "steps"}` substring
fixture and returned a generic "Here is my plan..." content blurb
with no MCP tool call, so the runtime never invoked create_view, the
MCP middleware never fetched the UI resource, and the sandboxed
iframe never mounted.

Add a fixture pair in d5-all.json (and its harness mirror) keyed on
"draw a simple flowchart": turn 1 emits create_view with a three-
step Start -> Process -> End flowchart, turn 2 emits the narration
after the tool result. The distinctive substring beats the generic
feature-parity catch-alls under first-match-wins.

Adds a regression test that loads the same fixture files in the same
order as docker-compose.local.yml and asserts via aimock's matchFixture
that each mcp-apps pill routes to its create_view fixture on turn 1 and
its narration fixture on turn 2.
2026-05-11 15:21:18 +02:00
Alem Tuzlak 7c3edca2b7 fix(showcase): unbreak multimodal demo end-to-end (sample buttons auto-send, dedupe, proxy)
The langgraph-python multimodal-attachments demo had a stack of bugs
that compounded each other. Fixing them required touching the local
docker-compose, the aimock fixtures, the LangChain middleware, the
client-side AG-UI shim, and the sample-attachment buttons. This
commit lands the full set together because they only make sense as
a unit — verified end-to-end against `showcase up langgraph-python`
in a headed browser. New e2e suite pins each regression.

Supersedes #4584 (the original fix from May 1 that never landed —
this is a fresh port onto the post-refactor file layout where
page.tsx is split into legacy-converter-shim.tsx, multimodal-chat.tsx,
file-to-data-attachment.ts).

What was broken and what changed:

1. Random uploads crashed with `Failed to fetch`. aimock returned
   HTTP 404 on no-match, the LangGraph SDK surfaced `NotFoundError`,
   the AG-UI stream surfaced a `RUN_ERROR`, the demo crashed.
   Added `--proxy-only` + `--provider-openai https://api.openai.com`
   to the local aimock command so unmatched user prompts fall through
   to real OpenAI (mirrors the Railway aimock setup).

2. Bundled-sample fixtures keyed on user-visible canned prompts.
   The auto-prompts are deliberately long, specific, and natural-
   reading ("can you tell me what is in this demo image/pdf I just
   attached") so they (a) render cleanly as the user message bubble,
   and (b) can't collide with arbitrary user prompts — random
   uploads phrase questions differently and fall through to the
   proxy.

3. Sample buttons now auto-send via `useAgent`. The previous
   DataTransfer-based path queued the attachment via the chat's
   hidden file input, then required clicking send while the
   attachment was still uploading — `CopilotChat.onSubmitInput`
   rejects submits during upload AND clears the input regardless,
   so the canned prompt was eaten. Rewrite to call
   `agent.addMessage(...)` + `copilotkit.runAgent({ agent })`
   directly with the base64'd content part, sidestepping the
   upload race entirely.

4. PDF flattened text bled into the rendered user message.
   `_PdfFlattenMiddleware` ran in `before_model` and returned
   `{"messages": rewritten}`, which persisted to agent state. The
   chat UI then rendered the `[Attached document]\n<pdf body>` text
   part inline with the user prompt. Switched to `wrap_model_call`
   so the PDF→text rewrite is scoped to the outgoing model request
   only and never pollutes state.

5. Attachments doubled (and PDFs rendered as broken `<img>`). The
   `@ag-ui/langgraph` round-trip translates outgoing `binary` parts
   to LangChain `image_url` and incoming `image_url` back to `image`
   AG-UI parts — regardless of mimeType, so PDFs came back as
   `type: "image"` with `mimeType: "application/pdf"` and were
   forced into `ImageAttachment`, where the load failed and the
   chat showed two "Failed to load image" boxes. Plus the user's
   original modern part survived alongside the round-tripped one,
   doubling visible chips.

   Added a `dedupeUserMessageMedia` subscriber on both
   `onMessagesSnapshotEvent` and `onRunFinalized` to:
   - dedupe media parts by `source.value` so the local + round-
     tripped copy collapse to one chip
   - re-key part `type` from `mimeType` so PDFs route to
     `DocumentAttachment` (icon + filename) and images to
     `ImageAttachment`.

   Also flipped the `onRunInitialized` shim from REPLACE to APPEND
   — keep the modern part for the UI AND emit a legacy `binary`
   sibling for the converter.

6. Regression suite (`tests/e2e/multimodal.spec.ts`). Replaces the
   pre-rewrite suite with five focused tests:
   - page loads with all expected affordances
   - sample image: auto-sends, EXACTLY ONE `<img>`, assistant
     references the logo
   - sample PDF: auto-sends, EXACTLY ONE `DocumentAttachment` chip
     ("PDF" label), NO `<img>`, no `[Attached document]` text bleed
   - image then PDF in the same session: each message keeps its own
     single chip, no cross-contamination
   - PDF then image in the same session: symmetric

   All 5 pass against the live local stack (15.4s).
2026-05-11 14:54:38 +02:00
Alem Tuzlak dd3c142a43 Merge remote-tracking branch 'origin/main' into fix/d5-tool-rendering-reasoning-chain-multi-turn
# Conflicts:
#	showcase/aimock/d5-all.json
#	showcase/harness/fixtures/d5/tool-rendering-reasoning-chain.json
2026-05-11 13:25:32 +02:00
Tyler Slaton df3e232068 fix(showcase/aimock): add D5 fixtures + restore reasoning emission
D5 probes were red for two demos because the aimock fixtures didn't
exist or were dispatched through the wrong Responses-API path.

- `harness/fixtures/d5/mcp-apps.json` (new) — the MCP Apps probe sent
  "Open Excalidraw and sketch a system diagram" with no matching
  fixture, so aimock 404'd and the agent threw `NotFoundError`.
  Added a two-leg fixture: leg 1 emits a `create_view` MCP tool call
  with Excalidraw element JSON; leg 2 is the `hasToolResult: true`
  content reply. Mirrored verbatim into `aimock/d5-all.json` so the
  Docker-baked aimock has the same matches.
- `harness/fixtures/d5/tool-rendering-reasoning-chain.json` — added
  non-empty `content` fields to the first-leg fixtures (Tokyo weather
  + SFO/JFK flights). aimock's Responses-API dispatch routes
  tool-call-only responses (no `content`) through
  `buildToolCallResponse`, which silently drops the `reasoning`
  payload; with a non-empty `content`, dispatch shifts to
  `buildContentWithToolCallsResponse`, which emits the
  reasoning_summary events the v2 `<ReasoningBlock>` needs to mount.
  Mirrored into `d5-all.json`.
- `harness/fixtures/d5/reasoning-display.json` — added a second
  matcher for the `reasoning-default` e2e pill prompt ("sky appears
  blue"), keyed alongside the existing D5 probe matcher ("show your
  reasoning step by step"). Both responses include `reasoning` fields
  so the built-in `CopilotChatReasoningMessage` "Thinking…/Thought
  for…" header lands.
- `aimock/feature-parity.json` — added a fixture for the agentic-chat
  multi-turn name-recall test ("What name did I just give") replying
  "You said your name is Alice." The matcher is intentionally narrow
  to avoid colliding with the generic showcase-assistant catch-all.

Net D5 effect: `mcp-apps` and `tool-rendering-reasoning-chain` turn 1
go from red → green. The remaining red on `tool-rendering-reasoning-
chain` turn 2 is the deeper aimock+deepagents+Responses-API multi-
turn state-management issue (production with real OpenAI works); out
of scope per the original audit's Phase F.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-10 15:17:06 -07:00
Tyler Slaton da8626d819 fix(showcase/aimock): tool-rendering-reasoning-chain multi-turn (Turn 2 hang)
LGP probe hung on Turn 2 (Find flights from SFO to JFK) after Turn 1
(weather Tokyo) succeeded. Root cause: the SFO/JFK first-leg fixture
matched on `hasToolResult: false`, but Turn 1's tool result remains in
the conversation by Turn 2, so the matcher skipped it and fell through
to the second-leg fixture (`hasToolResult: true`), returning narration
without the search_flights tool call. The probe then timed out waiting
for the FlightListCard testid that never mounted.

Fix: switch the second-leg fixture to `toolCallId` matching (the last
message must be a tool result with the specific call_id) and drop
`hasToolResult` from the first-leg fixture. The second-leg fixture
must come BEFORE the first-leg in file order because the matcher is
first-match-wins; with `toolCallId`, it cannot accidentally swallow
first-leg requests (whose last message is the user prompt, not a
tool result).

Also adds a `content` field to the Tokyo first-leg fixtures
(`get_weather` and `get-weather` Mastra-hyphen variants) so they route
through `buildContentWithToolCallsStreamEvents` instead of
`buildToolCallStreamEvents` — the latter codepath silently drops the
`reasoning` field, which is why the reasoning-block testid never
mounted on Turn 1 either before this change.

Tokyo (Turn 1) keeps the simpler `hasToolResult` pattern because it has
no prior turns; only the SFO/JFK fixture pair needs the toolCallId
shape to survive multi-turn use.
2026-05-09 12:01:44 -07:00
github-actions[bot] 80eb7a12ad style: auto-fix formatting 2026-05-09 02:02:47 +00:00
Tyler Slaton 1ce83b8a73 fix(showcase): close 2 D5 fixture gaps + filter deprecated features from gold-standard view
Three fixes that follow up on PR #4743 to bring LGP closer to fully-green
on the dashboard:

1. tool-rendering-reasoning-chain probe: was failing with
   `expected [data-testid="reasoning-block"] to mount within 30000ms`.
   Root cause: the demo's `<ReasoningBlock>` slot only mounts when a
   reasoning-role message lands in the transcript, which requires
   aimock to emit REASONING_MESSAGE_* events, which in turn requires
   the fixture's first-leg response to carry a `reasoning` field. The
   weather/Tokyo and SFO/JFK first-leg fixtures were missing it.
   Mirrors the convention documented in reasoning-display.json:2.
   Patched both source (harness/fixtures/d5/) and bundle (aimock/d5-all.json).

2. gen-ui-interrupt source fixture: the source fixture file was missing
   the resume-leg toolCallId entries that already existed in the bundle.
   Cosmetic mirror so re-bundling stays consistent. Same chip prompts +
   same toolCallIds as interrupt-headless.json (both probes share the
   same agent and aimock fixture set; the difference is the FRONTEND
   rendering — useInterrupt inline vs useHeadlessInterrupt separate-pane).

3. Dashboard gold-standard filter: 4 deprecated/legacy features
   (agentic-chat-reasoning, hitl, hitl-in-chat-booking,
   reasoning-default-render) used to render as X-marked rows in the
   LGP gold-standard dashboard view because LGP intentionally does
   NOT implement them — they were consolidated into the modern shape
   (reasoning-custom + reasoning-default; hitl-in-chat with
   useHumanInTheLoop). Other 17 integrations still serve those legacy
   demos, so we don't yank the features from feature-registry.json
   entirely. Instead: marked them `deprecated: true` and updated
   generate-registry.ts to skip emitting cells when a deprecated
   feature is unshipped for an integration. LGP cells: 43 → 39 (the
   4 deprecated rows disappear). Other integrations: unchanged
   (audit trail preserved). Catalog total: 774 → 770.

Tests:
  - 1588/1588 harness vitest passing
  - 19/19 generate-catalog + generate-registry tests passing
    (counts updated for the 4 dropped LGP cells + new deprecated-
    feature filter test)
  - validate-fixture-tool-surface clean (282 fixtures × 627 demos)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-08 19:00:58 -07:00
Tyler Slaton 1ad78c39d7 feat(showcase): add 3 LGP D5 probes + driver retry-once (#4743)
## Summary

- 3 new D5 probes (multi-turn, agentic-chat-style) for
`/demos/{interrupt-headless, shared-state-read,
tool-rendering-reasoning-chain}` — closes the demo↔probe coverage gap so
every langgraph-python demo now has its own dashboard cell.
- Driver retry-once in `e2e-deep.ts` — transient `goto-error` /
`conversation-error` failures (≥2s on attempt 1) get one retry before
recording red, cutting ~10× the dashboard flap rate. Persistent
assertion-style fails skip retry.
- Drops the legacy dual-claim where `d5-shared-state.ts` owned both
`shared-state-read` AND `shared-state-write` for a single bidirectional
probe — `shared-state-read` now belongs to the standalone recipe-editor
probe.

## Scope rationale

LGP is the north-star integration; this PR codifies its current demo
surface in the test infra. **Other integrations may flip red on the new
probes — that's expected and welcome.** Cross-integration parity follows
in a separate wave; we're using LGP as the template.

## Files

- `showcase/harness/src/probes/scripts/d5-interrupt-headless.ts` (new) —
`useHeadlessInterrupt` flow: chip → interrupt popup → slot pick →
resume.
-
`showcase/harness/src/probes/scripts/d5-tool-rendering-reasoning-chain.ts`
(new) — combines reasoning-block slot + per-tool renderer (WeatherCard /
FlightListCard).
- `showcase/harness/src/probes/scripts/d5-shared-state-read.ts` (new) —
recipe-editor with neutral default agent, asserts `recipe-card` form
mounts AND agent references recipe context across turns.
- `showcase/harness/src/probes/drivers/e2e-deep.ts` — retry-once loop
around `runFeature`.
- `d5-registry.ts` / `d5-feature-mapping.ts` / `live-status.ts`
(dashboard) / LGP `manifest.yaml` / `feature-registry.json` /
`constraints.yaml` — wire the new featureTypes and features end-to-end.
- `d5-shared-state.ts` (+test) — drops dual-claim.
- 2 pre-existing test fixes folded in: `d5-gen-ui-interrupt.test.ts`
(mock updated to current evaluate-poll resume signal) and
`conversation-runner.test.ts` (preFill ordering assertion now matches
actual deferred-cascade contract).
- aimock `d5-all.json` — +2 shared-state-read fixtures.
interrupt-headless and tool-rendering-reasoning-chain reuse existing
fixtures whose substrings already match their chip prompts.

## Test plan

- [x] `pnpm vitest run` in `showcase/harness/` → **1588 / 1588 passing**
(was 1585 / 1588 with 3 pre-existing fails before this PR; 2 are fixed
here, 1 was an obsolete-mock issue).
- [x] `npx tsx showcase/scripts/validate-fixture-tool-surface.ts` →
clean (282 fixtures × 627 demos, no drift).
- [x] `npx tsx showcase/scripts/generate-registry.ts` → clean (18
integrations, 38 wired LGP features, 756 catalog cells).
- [x] `pnpm typecheck` in `showcase/harness/` → clean.
- [x] `d5-mapping-drift.test.ts` → green (CATALOG_TO_D5_KEY mirrors
REGISTRY_TO_D5).
- [ ] After merge: watch the e2e-deep dashboard rotation — the 3 new LGP
cells should write rows on first tick (no red zombies — these are
first-time emissions).

## Known follow-up (NOT in this PR)

`auth.spec.ts` test #5 ("signing back in re-mounts a fresh chat
surface") fails on Railway. Symptom: after sign-out → sign-in cycle, the
second `Hello again` send doesn't produce an assistant response within
30s. Looks like a `react-core/v2` ref-handling regression on
`<CopilotKit>` unmount/remount — deserves its own focused investigation
rather than expanding this PR.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-05-08 17:57:26 -07:00
Tyler Slaton be94bd7a6f feat(showcase): add 3 LGP D5 probes + driver retry-once
Closes the demo↔probe coverage gap for /demos/{interrupt-headless,
shared-state-read, tool-rendering-reasoning-chain} so every demo
under langgraph-python (the north-star integration) now has a D5
probe writing to its own PocketBase cell — not relying on cross-
demo umbrella records.

New probes (multi-turn, mirroring the agentic-chat structure):
  - d5-interrupt-headless: exercises useHeadlessInterrupt — chip
    prompt → backend interrupt(...) → app-surface popup → slot pick
    → resume → assistant confirmation. Distinct from gen-ui-interrupt
    (which uses inline useInterrupt).
  - d5-tool-rendering-reasoning-chain: combines reasoning-block slot
    + per-tool renderer (WeatherCard, FlightListCard) on the same
    chat surface. Catches a regression in either side.
  - d5-shared-state-read: recipe-editor demo (neutral default agent,
    no tools) — verifies recipe-card form mounts AND agent reads
    shared state across turns. Drops the dual-claim that
    d5-shared-state.ts had on `shared-state-read` (now write-only).

Driver retry-once (e2e-deep.ts):
  Probes that fail with a transient class (`goto-error` /
  `conversation-error`) AND took ≥2s on the first attempt now retry
  once before recording red. Persistent assertion-style failures
  (sub-2s) and intentional aborts/feature-timeouts skip retry —
  retrying a deterministic mismatch just burns clock and obscures
  the signal. Cuts ~10× the dashboard flap rate.

Plumbing:
  - D5FeatureType enum: +interrupt-headless, +tool-rendering-reasoning-chain.
  - REGISTRY_TO_D5 (harness) + CATALOG_TO_D5_KEY (dashboard) mirror
    the new mappings; d5-mapping-drift test enforces this.
  - LGP manifest features + demos entries + constraints allowlist.
  - feature-registry.json: +shared-state-read.
  - aimock d5-all.json: +2 shared-state-read fixtures (interrupt-
    headless + tool-rendering-reasoning-chain reuse existing fixtures
    that already match their chip prompts).

Tests: 1588/1588 harness vitest green. validate-fixture-tool-surface
clean (282 fixtures × 627 demos, no drift). Two pre-existing test
fixes folded in — d5-gen-ui-interrupt assertion mock updated to
match the current evaluate-poll resume signal; conversation-runner
preFill ordering test now asserts the actual deferred-cascade
contract instead of a stricter pre-preFill ban that the runner
never enforced.

Known follow-up (not in this PR): auth.spec.ts test #5 ("signing
back in re-mounts a fresh chat surface") fails on Railway — second
sign-in's "Hello again" never produces an assistant response. Looks
like a react-core/v2 ref-handling regression on <CopilotKit>
unmount/remount; deserves its own focused investigation.

Other integrations may flip red on the new probes — that's
expected. We're treating LGP as the template; cross-integration
parity follows in a separate wave.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-08 17:22:57 -07:00