The declarative-gen-ui demo across the three langgraph integrations now
relies on the middleware to inject and execute generate_a2ui — the agents
collapse to create_agent + CopilotKitMiddleware with no hand-rolled tool.
Adds render_a2ui fixtures for the new tool path and pins the integrations
to the A2UI alpha SDKs (copilotkit 0.1.94a1, @copilotkit/sdk-js 1.59.3-alpha.1).
Audit follow-up to #5232, which fixed the gen-ui-interrupt d6 fixture
leg mis-order for langgraph-python + langgraph-typescript. aimock's
matchFixture is first-match-wins in array order and the loader dedups
on userMessage, so a toolCallId resume-leg placed BEFORE its
toolName:schedule_meeting first-leg shadows the tool-emitting leg on
turn-2 — the interrupt never fires and time-picker-card never mounts.
Fleet audit of every showcase/aimock/d6/*/gen-ui-interrupt.json found
ag2 as the only remaining mis-ordered fixture (both toolCallId
resume-legs preceded their toolName first-legs). Reorder ag2 so each
pill's toolName first-leg precedes its toolCallId resume-leg, matching
the proven-correct mastra/pydantic-ai/langgraph pattern. Reorder only;
response payloads and match keys are unchanged.
All other gen-ui-interrupt-supporting integrations (built-in-agent,
claude-sdk-*, langgraph-fastapi, langroid, ms-agent-dotnet/python,
pydantic-ai, spring-ai, strands) were already correctly ordered.
The gen-ui-interrupt d6 probe drives two interrupt turns in one thread.
Turn-2 failed because the langgraph-python/typescript fixtures ordered each
pill's toolCallId resume-leg BEFORE its toolName first-leg. aimock dedups on
userMessage (first-match-wins), so the text-only resume-leg shadowed the
tool-emitting first-leg on the duplicate userMessage — the agent emitted a
plain text reply with no schedule_meeting tool call, the interrupt never
fired, and the time-picker-card never mounted.
Reorder each pill so the toolName:schedule_meeting first-leg precedes its
toolCallId resume-leg, matching the passing mastra/pydantic-ai ordering. The
resume-leg still only matches when the last message is its tool result
(toolCallId guard), so confirmation still works.
The gen-ui-headless-complete probe
(showcase/harness/src/probes/scripts/d5-gen-ui-headless-complete.ts)
references fixtureFile "gen-ui-headless-complete.json", but that file
existed in no D6 slug, so the probe's first leg 503'd under strict.
Add gen-ui-headless-complete.json for all 18 D6 slugs, modeled on the
green langgraph-python headless-complete.json pattern: the four gen-UI
pills (weather/stock/highlight/revenue) with narration (toolCallId)
fixtures FIRST and toolcall (userMessage+context) fixtures AFTER, no
turnIndex gate, so every one of the probe's four sequential turns in one
chat thread matches regardless of prior assistant/tool history. Interrupt
fixtures are intentionally omitted (interrupt-headless is a separate
cluster item).
These 8 fixtures per slug share match keys with the pre-existing
headless-complete.json for the same context (the demos share pills and
are disambiguated at runtime by probe path), which raises the
aimock-fixtures collision-detection exact-duplicate count by 46. Bump
KNOWN_DUPLICATE_CEILING 230 -> 276 to match, consistent with how prior
per-integration feature fixtures bumped the baseline; the substring-shadow
ceiling is unchanged (no new shadows).
Validated: all 732 aimock-fixtures schema/collision tests pass, and a
local aimock --strict run returns 200 for each of the four pills across
turns (langgraph-python + pydantic-ai spearheads, plus spring-ai).
D6 drives each feature's pills as sequential turns in one chat thread.
aimock matches turnIndex against the count of assistant messages in
history, so a "turnIndex": 0 gate on a per-pill tool-call leg only
matches the FIRST pill — pills 2+ (turnIndex 1/2/3) match no fixture,
503 under strict, and the harness reports "timeout: assistant did not
respond".
Remove the turnIndex:0 gate from the multi-turn tool-call legs of
frontend-tools (sunset/forest/cosmic) and tool-rendering-reasoning-chain
(the three chained pills: Compare AAPL/MSFT, compare-to-smaller dice,
weather-there flights) so they match on their distinct userMessage
substrings + context, mirroring the already-green langgraph-python
fixtures. Each pill's substring is unambiguous and its toolCallId-keyed
follow-up fixture is ordered before the de-gated first leg (first-match
wins), so follow-up narration still wins and there is no cross-matching.
The reasoning-chain "Find flights from SFO to JFK." pill keeps its
turnIndex:0 gate to match the green langgraph-python reference exactly.
Proven locally with aimock --strict: multi-turn requests for these pills
returned 503 with the gate (negative control) and now return 200 across
all turns with the correct fixture/narration.
The aimock matcher gates turnIndex against the request's assistant-message
count (assistantCount !== turnIndex → skip). The agno, spring-ai, and
langgraph-fastapi agentic-chat fixtures baked turnIndex:0 on all three
goldfish conversation turns. Turn 2+ carries >=1 assistant message, so
turnIndex:0 could never match those turns — aimock returned no-match and the
request fell through to proxy (503), failing the d6 cell.
Mirror the canonical clean pattern used by the other 15 frameworks
(langgraph-python et al.): omit turnIndex on the multi-turn conversation
turns and disambiguate purely by userMessage. Single-turn fixtures in the
same files keep turnIndex:0 (unchanged). Verified locally against the aimock
matcher: turn-2 goes 404→200 for all three frameworks, turn-1 unregressed.
Address the four residual D6 failures on the ms-agent-dotnet integration
(baseline was 177/5).
route.ts toolCallId strip completeness
- Extend stripReplaySafeToolCallIdsFromMessage to also clean the
snake_case `tool_calls[].id` array and any nested OpenAI-style
`function.tool_call_id`. AG-UI canonical uses `toolCalls`, but some
runtime / message-converter paths emit the OpenAI shape and the
replay-safe `__ck_run_<uuid>` suffix was leaking through to aimock on
those paths. Apply the same coverage in applyToolResultDecisionSuffix
so decision-suffix routing (`__approved` / `__rejected` /
`__cancelled`) lands on every tool-call shape.
- Add a universal strip middleware in createAgent itself so EVERY
registered agent (not only the replay-safe ones) clears the suffix
before the request reaches the backend / aimock. Decision suffixing
remains scoped to createReplaySafeAgent because the suffix is
non-idempotent and the inner middleware re-runs the same logic.
- Drop the now-redundant strip calls inside createGenUiAgent,
createReadonlyContextAgent, createSharedStateReadWriteAgent, and
createReasoningAgent — the outer createAgent middleware already
canonicalised inbound messages by the time these run.
chat-slots fixture
- Mirror LGP's turnIndex:1 'Give me a fun fact' entry so the
chat-slots e2e's second-turn assistant slot has a deterministic
reply (the bare 'Give me a fun fact' fixture in headless-simple.json
is turnIndex:0-gated and was never matching chat-slots' turn 1).
HITL reject + interrupt-headless cancel branches
- Add a reject-branch fixture in render-a2ui.json keyed on
toolCallId `call_d5_generate_steps_001__rejected` so the hitl.spec
reject-flow's 'will not execute the Mars trip plan' assertion lands
the right narration. Keep the legacy hasToolResult:true fixture as a
fallback for paths that don't apply a decision suffix. Add an
`__approved` variant for symmetry.
- Add cancel-branch fixtures in interrupt-headless.json for both the
sales-intro and 1:1-with-alice pills, keyed on `__cancelled`
toolCallIds, returning the Denied/not-booked narration the
interrupt-headless cancel spec expects.
hitl-in-app + hitl-in-chat demo pages already mirror LGP (only an extra
README.md per directory in this integration), so no page changes were
needed.
Port the replay-safe toolCallId stripping middleware from the ms-agent-dotnet
sibling route.ts so HITL/interrupt demos match toolCallId-keyed aimock
fixtures across the 2nd-turn request. The strip walks every inbound message
(role=tool, role=assistant.toolCalls[].id, toolCallId, tool_call_id) before
the AG-UI HttpAgent forwards them onto the FastAPI backend, and the outbound
event stream rewrites toolCallId on TOOL_CALL_* events to embed the
deterministic per-run suffix so the next turn's fixture matcher still
keys on the original (suffix-stripped) id.
Wraps the human_in_the_loop, interrupt-adapted, hitl-in-app, and
hitl-in-chat agents with the new replay-safe middleware. Adds gen-ui-agent
set_steps state-snapshot synthesis, readonly-state-agent-context system
message injection, and shared-state-read-write preference-as-system
injection — all ported verbatim from the dotnet sibling so the per-agent
shaping behavior is identical across the MAF runtimes.
chat-slots.json: add the turnIndex:1 "Give me a fun fact" fixture mirrored
from langgraph-python/chat-slots.json so the chat-slots.spec.ts second-
turn assertion ("second assistant turn is also wrapped in the custom
slot") gets a deterministic reply instead of falling through to headless-
simple's fun-fact fixture or the live proxy.
hitl-in-app and hitl-in-chat demo pages were already mirrored from LGP
(only a minor consumerAgentId addition on the in-app suggestions module).
No frontend page changes needed.
The B2 header-forwarding conveyance in src/agents/_header_forwarding.py
is untouched and verified intact (httpx hook + Starlette HTTP middleware
plus ContextVar bridge).
Multimodal diagnosis: the 5-test multimodal.spec.ts timeout is NOT a
routing/wiring gap. The agent_server.py mounts /multimodal FIRST (before
the catch-all "/"), the dedicated Next.js route /api/copilotkit-multimodal
registers the HttpAgent under "multimodal-demo", and the page wires
runtimeUrl + agent correctly. The multimodal fixture's userMessage
"describe the sample image" is the same stale phrasing LGP uses (which
passes at 185/0/2 — the test asserts a /image/i regex on the assistant
transcript, so fallthrough to the proxy still satisfies it). The
remaining suspects are (a) agent_framework_ag_ui's AG-UI -> AF adapter
mishandling inbound `binary` content parts, (b) the dual chat_client
init pattern (multimodal_chat_client built after the global httpx hook
is installed — the hook is idempotent so this should be safe), or
(c) >30s real-OpenAI vision latency under D6 record-replay. Capture
agent_server stderr during a single multimodal run to confirm.
## Status
WIP / not ready to merge. Preserves in-flight D6 work so it isn't lost
mid-rollout. LGP is fixture-complete; other integrations are
mid-rollout.
### Latest banked work
- **langgraph-python — 185 / 0 / 2** (green). Achieved by narrowing a d4
chat matcher that was shadowing the d6 beautiful-chat search_flights
fixture (load order is shared -> d4 -> d6, first-match-wins) plus
refreshing the d6 tool-rendering and tool-rendering-custom-catchall AAPL
fixtures (turnIndex:0 -> hasToolResult:false so the first leg fires in
multi-pill threads). 2 skips are the by-design mcp-apps iframe gap.
- **langgraph-typescript — 185 / 0 / 2** (green). Mirrored the LGP d4
narrowing on the LGT side (3 matchers) and across the
d6/langgraph-typescript suite: replaced fragile turnIndex:0 gates with
hasToolResult:false, fixed em-dash escaping that broke literal matches
in multi-pill threads, and added jsFunctions payloads to the three
sandboxed-ui fixtures (`_from-feature-parity`, `headless-complete`,
`gen-ui-open-advanced`). The final fix removed the chain-tools
`hasToolResult` match gate (which checked the whole thread and made the
chain pill fall through to the broad weather matcher mid-thread),
mirroring LGP.
- **google-adk: 174/7/2 (was 134/52/5)** — conveyance + test-parity +
fixtures + pill-wiring rebuild; 7 residual (6 default-catchall framework
default-renderer testid version question, 1 beautiful-chat fixture).
- **Pill-parity staged across 13 integrations** (ag2, agno, mastra,
pydantic-ai, claude-sdk-python, claude-sdk-typescript, llamaindex,
langroid, strands, spring-ai, built-in-agent, crewai-crews,
langgraph-fastapi). The canonical LGP suggestion pill set is now
mirrored as `src/app/demos/*/suggestions.ts` files in each integration,
with targeted edits to existing `open-gen-ui-advanced` and
`byoc-hashbrown` files. **These new files are currently UNWIRED** — each
integration's `page.tsx` still defines its pill list inline via
`useConfigureSuggestions`. Banked so the canonical source survives; a
follow-up will rewire `page.tsx` to import from `suggestions.ts` and
drop the inline copies.
- **Fleet test-parity sweep**: 576 e2e specs across 15 integrations
aligned to LGP canonical (SHA-verified); 2 orphan specs removed.
- ms-agent-dotnet 177/5/7, ms-agent-python 174/11/2 (post
fixture-mirror); default-catchall green (page-level renderer, not
react-core-gated).
## Scope
### Conveyance (foundation)
Inbound `x-aimock-context` (and friends) must ride along on outbound LLM
HTTP calls so aimock fixture matching sees the inflight test's context.
Without this the call lands on the default project's aimock and silently
picks the wrong fixture. New per-integration
`_header_forwarding.{py,ts}` shim plus matching `agent_server` / route /
factory wiring covers: ag2, agno, built-in-agent, claude-sdk-python,
claude-sdk-typescript, crewai-crews, google-adk, langgraph-fastapi,
langgraph-python, langgraph-typescript, langroid, llamaindex, mastra,
ms-agent-python, pydantic-ai, strands.
For ADK/Gemini the global httpx hook is installed BEFORE any `agents.*`
import (google-genai constructs its client at module-import time).
### langgraph-python — 185 / 0 / 2
Fixture-complete via the conveyance shim + refreshed d6/langgraph-python
fixtures + copilotkit 0.1.93 bump + the latest d4-matcher-narrowing fix
(see banked work above).
### Per-integration fixtures
Mid-rollout snapshot of d6 fixtures across the cohort plus narrowing of
`aimock/shared/common.json`'s generic 'hello' fixture to 'hello world'
so it no longer shadows D6 pills whose prompts contain 'hello' as a
substring.
### Harness `--isolate` patch
`scripts/cli/_common.sh apply_isolation` now rewrites compose-file
relative paths to absolute (build/context/dockerfile/volumes/env_file),
enforces the docker compose `[a-z0-9_-]` project-name rule, and exports
`SHOWCASE_COMPOSE_FILE` / `SHOWCASE_INFRA_PORT_OFFSET` plus offset host
URLs. The TS harness CLI (`aimock-rebuild` / `config` / `doctor` /
`lifecycle`) honors the new env so concurrent isolated stacks stop
reporting each other's services as healthy.
## Lockfile decision flagged
`showcase/integrations/langgraph-python/pnpm-lock.yaml` was deleted in
this branch. Decision: keep the deletion. Rationale:
- 03bed3b76 (fix(showcase): regenerate 18 lockfiles in isolation; switch
to npm ci) migrated all showcase integrations off pnpm onto npm ci.
- The integration's Dockerfile uses `npm ci --legacy-peer-deps`.
- Every sibling integration committed only `package-lock.json` after
03bed3b76.
- The orphan pnpm-lock.yaml only risks tooling drift.
If anyone wants it restored: `git checkout origin/main --
showcase/integrations/langgraph-python/pnpm-lock.yaml`.
## Commits
- feat(showcase): D6 conveyance — forward x-aimock-context headers to
LLM clients
- feat(showcase/langgraph-python): D6 conveyance shim + copilotkit
0.1.93 bump
- feat(showcase/langgraph-typescript): D6 conveyance — propagate request
headers into ChatOpenAI
- feat(showcase/built-in-agent): D6 conveyance — header-forwarding shim
+ factory wiring
- feat(showcase/harness): support concurrent --isolate runs
- test(showcase): D6 langgraph-python fixtures — drive to 180/5/2
- test(showcase): D6 per-integration aimock fixtures + shared narrowing
- docs(showcase): GOTCHAS entry for D6 conveyance + --isolate notes
- feat(showcase): D6 conveyance — wire header-forwarding shims into
remaining entrypoints
- fix(showcase): unblock LGP D6 beautiful-chat + custom-catchall via d4
matcher narrowing
- fix(showcase): narrow LGT D6 d4 shadows + wire sandboxed-ui
jsFunctions
- chore(showcase): copy LGP canonical suggestion pills into 13
integrations
- fix(showcase): forward x-aimock-context per-request in google-adk
routes
- test(showcase): align google-adk e2e specs to langgraph-python
canonical
- fix(showcase): align google-adk D6 fixtures to LGP contract
- fix(showcase): wire google-adk default-catchall to shared 4-pill
suggestions
pydantic-ai had never received the fleet D6 parity sweep — its e2e specs and demo pages
were a pre-sweep, integration-specific set (only 3/26 suggestion files; missing canonical
demos; non-canonical byoc-*/agentic-chat-reasoning/reasoning-default-render variants).
Sitting at 69/109/2.
This change mirrors langgraph-python's canonical frontend (demos + specs + aimock
fixtures) into pydantic-ai, preserving pydantic-ai's Python backend untouched. The
per-demo agent.py files that pydantic-ai carries inside demo directories are preserved.
Changes:
- tests/e2e/: rsync LGP canonical 37-spec set over pydantic-ai (byte-identical). Removes
non-canonical byoc-hashbrown.spec.ts, byoc-json-render.spec.ts, shared-state-write.spec.ts.
Adds canonical declarative-hashbrown.spec.ts, declarative-json-render.spec.ts,
reasoning-custom.spec.ts, reasoning-default.spec.ts.
- src/app/demos/: rsync LGP demos over pydantic-ai. Removes non-canonical demos
(byoc-hashbrown, byoc-json-render, agentic-chat-reasoning, reasoning-default-render,
shared-state-write). Adds canonical demos (declarative-hashbrown, declarative-json-render,
reasoning-default, reasoning-custom) and the _shared/ helpers + demos/layout.tsx
pydantic-ai was missing. Restores pydantic-ai-specific agent.py files into the 9 demo
dirs that survived the mirror.
- src/app/demos/frontend-tools/page.tsx: patched agent slug from "frontend_tools" (LGP)
to "frontend-tools" (matches pydantic-ai's main route.ts registry).
- src/app/api/copilotkit-byoc-{hashbrown,json-render}/ renamed to copilotkit-declarative-*
to match the canonical frontend wiring. Internals still use HttpAgent against the
pydantic backend's /byoc_hashbrown/ + /byoc_json_render/ mounts (Python backend
untouched per scope). copilotkit-declarative-hashbrown/route.ts updates the registered
agent slug from "byoc-hashbrown-demo" to "declarative-hashbrown-demo" to match the
canonical demo. copilotkit-declarative-json-render/route.ts updates only the endpoint
path string (the agent slug "byoc_json_render" is the canonical LGP convention).
- src/app/api/copilotkit/route.ts: renamed reasoning agent registrations from
agentic-chat-reasoning + reasoning-default-render to reasoning-custom + reasoning-default
to match canonical demo slugs. Both still proxy to the same /reasoning/ backend mount.
- manifest.yaml: features[] + demos[] updated to reflect the canonical demo set
(byoc-* + agentic-chat-reasoning + reasoning-default-render removed; declarative-* +
reasoning-default + reasoning-custom added).
- aimock/d6/pydantic-ai/: added gen-ui-custom.json (mirrored from LGP with
context-swap + copiedFrom marker, per established fixture convention). Removed
orphan gen-ui-open-advanced.json (no LGP counterpart in the canonical set).
The Python backend (agent.py / src/agent_server.py / src/agents/) is unchanged.
Some pydantic-ai backend mounts continue to exist that the mirrored frontend no longer
references (e.g. /reasoning/ remains, the deleted demos' agent slugs are still
registered in route.ts but harmlessly orphaned) — these are intentional carry-overs
to avoid touching Python backend code per scope.
Same pattern as ms-agent-dotnet: full d4 chat.json rewrite + d6 mirrors. Dropped stale
turnIndex gates and broad shadow matchers. Default-catchall now green via page-level
shadcn-catchall-renderer. 11 residual: declarative-gen-ui charts, multimodal conveyance,
tool-rendering, reasoning-chain.
Full d4 chat.json rewrite + d6 mirrors. Dropped stale turnIndex gates and broad shadow
matchers that no longer reflect LGP-canonical conveyance. Default-catchall now green via
page-level shadcn-catchall-renderer (not react-core-gated). 5 residual: chat-slots, hitl,
interrupt-headless, readonly-state.
The whole-thread `hasToolResult: false` gate on the Chain-tools first-turn fixture caused
the chain pill to fall through to the broad "weather in Tokyo" matcher mid-thread.
Removing the gate (mirroring LGP) yields LGT 185/0/2.
Mirrors the LGP first-match-wins fix on the langgraph-typescript side. Three broad
matchers in d4/langgraph-typescript/chat.json were shadowing d6 fixtures; narrowed
them so the d6 pills win.
Across the d6/langgraph-typescript suite: dropped fragile turnIndex:0 gates in favour
of hasToolResult:false for first-leg tool emissions, fixed em-dash escaping that broke
literal string matches in multi-pill threads, and added jsFunctions payloads to the
three sandboxed-ui fixtures (_from-feature-parity, headless-complete,
gen-ui-open-advanced) so the sandbox renderer has executable handlers.
LGT D6 now passes 184/1/2 locally. Residual 1 fail is the custom-catchall multi-pill
follow-up case; tracked separately.
Aimock fixture load order is shared -> d4 -> d6, and matching is first-match-wins. A
broad d4 langgraph-python chat fixture for search_flights was shadowing the d6
beautiful-chat fixture and preventing the intended response from firing. Neutered the
d4 matcher to a non-matching sentinel so the d6 fixture wins.
Also refreshed d6/tool-rendering-custom-catchall.json and d6/tool-rendering.json:
replaced fragile turnIndex:0 gates with hasToolResult:false so the AAPL first-leg
fixture fires correctly when D5 probes run a prior 'weather in Tokyo' pill in the
same thread (multi-pill turnIndex>=2).
LGP D6 now passes 185/0/2 locally (2 skips are the by-design mcp-apps iframe gap).
Refresh d6 fixtures across the rollout cohort: ag2, built-in-agent,
claude-sdk-{python,typescript}, crewai-crews, google-adk, langgraph-fastapi,
langgraph-typescript, langroid, llamaindex, mastra, ms-agent-{dotnet,python},
pydantic-ai, strands. Companion d4/{langgraph-typescript,mastra,ms-agent-dotnet}
chat.json refreshes. Add the missing ms-agent-python/gen-ui-custom.json
to bring the integration up to the standard pill set.
Also narrow aimock/shared/common.json's generic 'hello' fixture to
'hello world' so it no longer shadows D6 pills whose prompts contain
'hello' as a substring (e.g. langgraph-python headless-simple sends
'Say hello in one short sentence.'). 'hello world' is unused by any
current demo pill, so the fixture remains a manual-typing fallback
without poisoning fixture matching.
This is a mid-rollout snapshot — fixture coverage is uneven across
integrations and rides alongside the conveyance shims landed earlier
in this branch.
Refresh d6/langgraph-python fixtures (beautiful-chat, chat-slots,
frontend-tools-async, gen-ui-{agent,custom,declarative,interrupt,open},
headless-complete, hitl-in-app, interrupt-headless, shared-state-streaming,
tool-rendering, tool-rendering-reasoning-chain, _from-feature-parity) plus
d4/langgraph-python/chat.json. Combined with the conveyance shim landed
in this branch, LGP is fixture-complete at 180 pass / 5 fail / 2 skip.
The residual 5 fails trace to react-core/v2's custom-wildcard-renderer
bug surfaced through gen-ui-interrupt (shared useInterrupt hook misbehaves
on the second interrupt; cross-integration D6 blocker, not a fixture or
conveyance defect). The 2 skips are the by-design mcp-apps iframe gap.
agno uses useFrontendTool (Strategy B) — async Promise handler in the
frontend — rather than LangGraph's native interrupt() primitive. The D5
probe asserts via useInterrupt hook which routes through the LangGraph
interrupt event path. agno doesn't emit those events; the D5 probe
fundamentally cannot pass against agno's HITL architecture without
per-integration probe-code divergence (which would violate the D6
apples-to-apples invariant).
Excluding these two features at the manifest level (matching google-adk
precedent) lets the D6 probe skip them cleanly. Removes the
corresponding D6 aimock fixtures since they're no longer reachable.
If agno gains LangGraph-style interrupt() support, or if aimock gains
AG-UI-event fixture authoring, this exclusion can be reverted.
Adds per-integration D6 fixtures for langgraph-python, langgraph-typescript,
and langgraph-fastapi. Each fixture is keyed by match.context for cross-
integration isolation and uses hasToolResult:false on toolCall responses
to prevent re-match loops.
The four fixture routing regression tests still referenced the old
monolithic d5-all.json / smoke.json / feature-parity.json files that
were reorganized into per-integration d4/ d6/ shared/ directories.
Changes:
- Load fixtures via glob from d6/langgraph-python/ (reference
integration) instead of deleted monolithic files
- Add _context to test requests (aimock 1.26.1 checks req._context
against match.context for per-integration scoping)
- Adapt subagents test from toolCallId-based chaining to
turnIndex-based chaining (matches new D6 fixture structure)
- Adapt state-context test to D6 context-scoping model (no
systemMessage matching; routing happens via X-AIMock-Context)
- Remove _migrated-from-*.json shared files (all fixtures already
exist in per-integration D6 dirs; the migration files caused
620 shared-vs-scoped collisions)
- Add KNOWN_DUPLICATE_CEILING=11 ratchet for pre-existing D6
intra-feature duplicate match keys
Move monolithic d5-all.json + feature-parity.json + smoke.json into
per-integration directories under d4/<slug>/, d6/<slug>/, and shared/.
Every fixture file is now context-scoped to enable server-side aimock
routing via match.context. Migrated 12 HITL fixtures from main's
d5-all.json additions into shared/_migrated-from-d5-all-hitl.json
for follow-up distribution into per-integration files.
Same agent_framework_openai history-split loop that hit the
frontend-tools fixtures also affected the two gen-ui-interrupt /
interrupt-headless schedule_meeting first-leg fixtures (sales intro
call + 1:1 Alice). Their `response.content` text ("Sure — let me
check available times.", "Got it — pulling up next-week slots.")
landed as a standalone assistant message in history; the next leg's
`userMessage` substring still matched THIS same fixture, and aimock
re-emitted the schedule_meeting tool call → the picker rendered
twice on the gen-ui-interrupt page exactly as the user reported.
Fix:
- Dropped `content` from both first-leg responses (the
toolCallId-anchored follow-up fixtures already provide the
post-pick narration: "Booked: Sales intro call confirmed..." /
"Scheduled: 1:1 with Alice locked in...").
- Added `hasToolResult: false` to both matchers as a belt-and-braces
guard so they only fire on the initial leg, never on follow-ups.
Full ms-agent-python e2e suite: 186 passed, 3 skipped, 0 failed.
The three `change_background` first-leg fixtures (Sunset, Forest,
Cosmic) returned BOTH `content` (the visible narration) AND
`toolCalls`. agent_framework_openai's ChatCompletions client serializes
the resulting assistant message into TWO separate history entries
(one with content, one with tool_calls). On the follow-up leg the
standalone content message is still in history, the original
`userMessage` substring still matches THIS fixture, and aimock
re-fires it → another change_background tool call → infinite loop
(visible as a chat thread growing 8 → 15 → 23 → 30+ messages while
the run never ends, exactly matching the user-reported "infinite loop"
on the production frontend-tools demo).
Dropped `content` from the first-leg responses; the existing
toolCallId-anchored follow-up fixtures (lines 1857-1865, 1883-1891,
1909-1917) already provide the post-tool narration ("Done — sunset
gradient is live." etc.) so the UX is unchanged on LGP and now also
works on MAF. Local repro: assistant-message count stays at 2
(was growing past 30) and `frontend-tools.spec.ts` continues to pass.
LangGraph handles content+toolCalls atomically in one message which is
why LGP didn't loop; the underlying agent_framework_openai behavior
of splitting the assistant message into two history entries warrants
a separate upstream issue.
Sales Dashboard pill on beautiful-chat was rendering an empty A2UI
surface (no metrics, no pie chart, no bar chart) — only the trailing
narration text appeared. Two stacked issues:
1. `showcase/aimock/d5-all.json` had a recorded catchall fixture
`{ model: "gpt-4.1", turnIndex: 0, hasToolResult: false }` with no
`userMessage` constraint. The secondary LLM call inside
`beautiful_chat.py::generate_a2ui` hits aimock with that exact shape
(`client.chat.completions.create(model="gpt-4.1", ..., tools=[{name:
"_design_a2ui_surface"}], tool_choice=...)`); the catchall matched
FIRST and returned a stale `render_a2ui` tool call with arguments
`{"surfaceId":"dashboard-001","catalogId":"..."}` — no `components`
field. `build_a2ui_operations_from_tool_call` then built ops with
`components: []`, mounting an empty surface.
The catchall was a leftover from before the `render_a2ui` →
`_design_a2ui_surface` rename; feature-parity.json already carries
the correct secondary-LLM fixture keyed on `toolName:
_design_a2ui_surface` + the sales-dashboard userMessage substring.
Removed the catchall entirely so the correct fixture wins.
2. `feature-parity.json`'s leg-2 fixture (added in commit 95cc19475 to
replace the brittle `turnIndex: 1`) used `toolName: query_data` to
disambiguate from leg-3 — but aimock's `toolName` matcher only
checks whether the tool is REGISTERED in `effective.tools`, not that
the last tool result was from it. `query_data` is in
`effective.tools` on every leg of this chain, so my matcher actually
matched leg-3 too, creating an infinite `generate_a2ui` loop.
Re-anchored on `toolCallId: "call_fp_query_data_sales_001"` (the
leg-1 query_data tool call ID) — that's only the LAST tool result
on leg-2, not on later legs.
The beautiful-chat Sales Dashboard pill's chain-leg-2 fixture in
feature-parity.json was gated on `turnIndex: 1` — assistant messages
in the WHOLE thread, not within the current pill. Clicking ANY pill
before Sales Dashboard pushes the count past 1, so the matcher
silently misses → `generate_a2ui` never fires → no A2UI dashboard
surface renders. Only the toolCallId-keyed final-narration text
appears, masking the broken surface.
Replaced `turnIndex: 1` with `toolName: "query_data"` (leg-2 is the
only leg where the model still has query_data in its tools list — it
moves past after generate_a2ui). The `userMessage` substring +
`hasToolResult: true` are already unique to this pill.
Added regression e2e in `beautiful-chat.spec.ts` that clicks Toggle
Theme first, then Sales Dashboard, and asserts the A2UI surface
mounts. Follows the RUNBOOK guidance: "Do not use `turnIndex` in new
fixtures."
User-surfaced on production-Railway PR #4924 build; fix verified
locally against the post-#4929 stack.
Brings ms-agent-python to one-to-one parity with langgraph-python (the D5
north star). Playwright e2e suite goes from 49/108 (~26%) → 164/178 (~92%),
33 of 37 cells fully green.
Manifest parity:
- Drop 4 MAF-only cells with no LGP analog: agentic-chat-reasoning,
hitl-in-chat-booking, shared-state-write, reasoning-default-render.
Reasoning is handled by reasoning-default + reasoning-custom (LGP);
booking pill folds into hitl-in-chat; shared-state-write was a TODO stub.
- Rename byoc-hashbrown → declarative-hashbrown and byoc-json-render →
declarative-json-render. Demo dir, API route dir, and frontend agent id
follow LGP's naming. Python module files retain the legacy `byoc_*`
prefix and FastAPI paths stay `/byoc-hashbrown` / `/byoc-json-render`
(matches LGP's "module name retains legacy graph id" convention).
- Port LGP `_shared/`, `_shared/interrupt-fallback-slots.ts`, and
`demos/layout.tsx` for one-to-one parity.
Cells ported verbatim from LGP (page + spec):
- agentic-chat, auth, beautiful-chat, chat-customization-css, chat-slots,
declarative-gen-ui, declarative-hashbrown, declarative-json-render,
frontend-tools, frontend-tools-async, gen-ui-agent, gen-ui-interrupt,
gen-ui-tool-based, headless-complete, headless-simple, hitl-in-app,
hitl-in-chat, shared-state-read, shared-state-read-write,
shared-state-streaming, subagents, tool-rendering, plus all four
tool-rendering* variants, a2ui-fixed-schema, agent-config, mcp-apps,
multimodal, open-gen-ui, open-gen-ui-advanced, prebuilt-popup,
prebuilt-sidebar, readonly-state-agent-context, reasoning-default,
reasoning-custom, voice.
Backend infrastructure:
- Swap shared `OpenAIChatClient` (Responses API) → `OpenAIChatCompletionClient`
(ChatCompletions). Root cause of the cross-cell post-tool ChatClientException
family: Responses API is stateful and only sends NEW items per leg,
relying on `previous_response_id` for history. aimock has no view of
that server-side state, so second-leg requests arrived without the
user message — fixture matchers keyed on `userMessage` couldn't fire
and the run fell through to real OpenAI. ChatCompletions sends full
history every leg, matching the LGP wire shape.
- Bump @ag-ui/client ^0.0.43 → ^0.0.53 (matches google-adk/LGP). Fixes
the REASONING_* Zod discriminator trap on the catch-all agent.
- Regenerate package-lock.json in isolation outside the pnpm monorepo so
npm-arborist doesn't resolve transitives against pnpm's hoisted
symlinks (avoid 40+ `../../../node_modules/.pnpm/...` paths in the
lockfile that break `npm ci` inside Docker).
- Add `yaml` (^2.8.4) for the new `src/app/demos/layout.tsx` that reads
manifest.yaml for per-cell page titles (LGP parity).
New / re-added MAF agent backends with LGP-equivalent behavior:
- reasoning_agent.py (uses Responses API explicitly — the only chat
client that emits AG-UI REASONING_MESSAGE_* events; rest of the
integration stays on ChatCompletions).
- tool_rendering_agent.py (non-reasoning sibling of the existing
reasoning_chain variant; shares tool surface via direct imports so
they can never drift apart; routes the three catchall cells to a
non-reasoning backend so the default renderer spec stops failing on
leaked reasoning blocks).
- gen_ui_agent.py — `set_steps` tool + `steps` state schema +
`predict_state_config` mirrors LGP's StateStreamingMiddleware shape.
- shared_state_streaming.py — `write_document` tool with
`predict_state_config` that streams the `document` arg into
`state.document` per-token.
- readonly_state_agent_context.py — minimal agent that consumes
frontend-provided `useAgentContext` entries; no tools.
- headless_complete_agent.py — three deterministic tools (`get_weather`,
`get_stock_price`, `get_revenue_chart`) mounted at /headless-complete
on the mcp-apps runtime (was routing to catch-all sales agent, which
returned seeded-random weather instead of the deterministic 68°F the
test asserts on).
Wiring:
- copilotkit/route.ts: register the new agents, drop the stale
shared-state-write entry, route all three tool-rendering variants to
the non-reasoning backend (the reasoning-chain cell keeps its own
dedicated path), register reasoning-default + reasoning-custom on
/reasoning, register gen-ui-agent on /gen-ui-agent,
shared-state-streaming on /shared-state-streaming,
readonly-state-agent-context on its dedicated path.
- copilotkit-mcp-apps/route.ts: register headless-complete agent (was
missing — the strict useAgent runtime sync in the newer
@copilotkit/react-core surfaced the gap).
- copilotkit-declarative-hashbrown/route.ts + copilotkit-declarative-json-render/route.ts:
new dedicated runtimes; agent IDs and runtime URLs follow LGP.
- copilotkit-declarative-gen-ui/route.ts: drop non-LGP `openGenerativeUI:
false` for parity.
A2UI tool rename — `render_a2ui` → `_design_a2ui_surface`:
- Ported LGP's `tools/generate_a2ui.py` (LGP renamed the secondary-LLM
tool to `_design_a2ui_surface` to avoid the A2UI middleware's bypass;
shared d5-all.json fixtures key the response on this name).
- Renamed every `render_a2ui` occurrence in src/agents/{a2ui_dynamic,
agent,beautiful_chat}.py and `tools/__init__.py`.
- Updated 4 declarative-gen-ui aimock fixtures to pass `context` arg in
the first-leg `generate_a2ui` tool call (agent_framework doesn't
auto-inject AgentSession into our @tool function so `session=None` and
the secondary-LLM `user_content` was defaulting to a catch-all string
containing "KPI dashboard" — every pill matched the KPI fixture).
Aimock router patch persisted alongside the integration changes:
hasToolResult matcher restricted to scan only messages after the last
user message (was global). The patch lives in F:/projects/cpk/aimock —
upstream PR pending.
Test infrastructure:
- playwright.config.ts: cap local workers at 4 + retries at 1. CI keeps
workers=1, retries=2. `agent_framework.Agent` is reused across requests
and the shared OpenAI HTTP client serialises concurrent SSE streams;
>4 workers makes 30s timeouts inevitable on a few cells. Confirmed
with hard data: workers=1 = 164 passed (16.8 min), workers=4+retries=1
= 164 passed (7.2 min), workers=undefined = 159 passed. Same green
set, ~2x faster. Long-term upstream fix is per-request Agent
instantiation in agent_framework_ag_ui.
Remaining 14 failures across 4 cells documented per-cell in the Notion
D5 sweep doc (declarative-gen-ui A2UI surface mounting, multimodal
attachment forwarding, tool-rendering-default-catchall multi-pill chain,
tool-rendering-reasoning-chain multi-leg chains). Each has a specific
next-pass action.
The generate_a2ui aimock fixtures were returning content+toolCalls in a
single response. When aimock streams this, content text is emitted first,
then the tool call. The CopilotKit frontend sometimes processes the
content text and closes the assistant turn before the tool call (and its
subsequent A2UI operations) can be processed, causing the chart to never
render.
Split the fixture response: generate_a2ui now returns only toolCalls
(the tool invocation), and the descriptive text content moves to the
toolCallId follow-up fixture (the post-tool-result response). This
ensures the runtime processes the tool call first, executes generate_a2ui,
receives A2UI operations, and only then emits the text response.
Verified 10/10 passes on LGP (3100) and 5/5 on LGT (3101), vs ~40%
failure rate before the fix.
Two recorded fixtures from a d5 run (2026-05-15) had wrong data
(attendee: "User" instead of "Sales team") and random toolCallIds
with no matching confirmation fixtures. They intercepted hitl-in-chat
requests before the correct feature-parity.json fixtures could match.
Two aimock fixture issues caused e2e test failures on both LGP and LGT:
1. shared-state-read: no fixture matched "What recipe am I making?" —
added a new fixture in feature-parity.json keyed on that substring.
2. hitl-in-app multi-pill test: the second pill's first-turn request
failed because hasToolResult: false on the 1st-turn fixtures rejected
conversations that already contained tool results from earlier pills.
Removed hasToolResult: false from refund/downgrade/escalate 1st-turn
fixtures — the more-specific post-tool-result fixtures (with
toolCallId + hasToolResult: true) still win on the 2nd turn.
multi-turn race on LGT
Two shared agentic-chat tests failed on both LGP and LGT because
the test messages had no matching aimock fixtures, and the
multi-turn test had a race condition on LGT where the second
Enter keypress was swallowed during a component re-render.
- Add 3 fixtures to feature-parity.json for the agentic-chat e2e
test messages (hello, Alice turn 1, Alice turn 2)
- Wait for suggestion pills to reappear before sending the
follow-up message in the multi-turn test
Delete 9 recorded fixture files from showcase/aimock/d5-recorded/recorded/
that were captured during a previous real-API recording session. These
fixtures are not needed -- the existing feature-parity.json fixtures
already cover all 4 test cases (Task Manager, Search Flights, PieChart,
BarChart) for both LGP and LGT.