The ag2 declarative-gen-ui route pointed its HttpAgent at the root
catch-all mount (agents/agent.py) instead of the dedicated
/declarative-gen-ui mount (a2ui_dynamic.py), and generate_a2ui declared
a required context arg that the model emits as {}. pydantic rejected
every call with "context Field required" and AG2 retried without bound —
a 630-iteration hot loop per pill that flooded logs and starved the
frontend.
Fix: route to the dedicated mount with injectA2UITool:false (the
dedicated agent owns generate_a2ui and emits a2ui_operations itself);
make generate_a2ui a no-arg tool matching the D6 fixtures and the
langgraph-python gold standard, with a constant inner system prompt
(per-pill distinctness comes from the captured user message). Regenerated
the gen-ui-declarative fixture and ported the LP definitions/renderers
catalog (all 7 driver testids) for parity. Eliminates the validation
loop: runsFinished=1, zero validation errors.
Bring the TypeScript AWS Strands integration to parity with the Python
strands sibling now that the @ag-ui/aws-strands TS adapter is confirmed to
support the same feature surface (per its examples/server):
- Restore A2UI: the declarative-gen-ui + a2ui-fixed-schema demos, their
routes, qa, specs, the @copilotkit/a2ui-renderer dep, beautiful-chat's
A2UI catalog, and the manifest entries (generative_ui / features / demos /
a2ui_pattern). manifest now matches strands-python feature-for-feature.
- Header forwarding: attach `x-aimock-context: strands-typescript` as a
static defaultHeader on the OpenAI client (model-factory + sub-agent
client) — the TS analog of the Python integration's _header_forwarding
shim — so aimock matches this integration's fixtures.
- aimock fixtures: add d6/strands-typescript + d4/strands-typescript
(ported from the Python sibling, context retargeted).
- playwright.config: X-AIMock-Context → strands-typescript.
Note: the raw tests/e2e Playwright suite is flaky and not a CI merge gate
(demo e2e / `/eval` D5 are comment-triggered, not required) — it fails the
same specs for strands-python too. The auto-gates (build, validate-
constraints, oxlint/oxfmt, unit) are green.
Mirrors LGP's `generate_a2ui` outer-emit fixture into 8 non-LGP slugs to close
the production 503 gap on the "Show me my sales dashboard for this quarter."
userMessage. Per /tmp/staging-journal-diff.md, aimock-staging returns 503 on
shape-B traffic (model=gpt-4.1, stream=true, tools=[generate_a2ui],
UA=AsyncOpenAI/Python) for all 8 slugs because no fixture matched.
Slugs patched (one entry each in d6/<slug>/gen-ui-declarative.json):
- llamaindex
- built-in-agent
- ag2
- langroid
- claude-sdk-typescript
- claude-sdk-python
- ms-agent-dotnet
- ms-agent-python
Each entry matches `userMessage` + `context: "<slug>"` and emits a
`generate_a2ui` toolcall with no args, mirroring LGP's sales-dashboard outer
entry. Per-slug unique toolCallId.
Red-green proof (local aimock, ghcr.io/copilotkit/aimock:latest):
RED (baseline main, 8 slugs): HTTP=404 no_fixture_match
GREEN (this branch, 8 slugs): HTTP=200 tool=generate_a2ui id=call_d6_decl_dash_outer_<slug>_001
LGP regression (baseline+fix): HTTP=200 (unchanged)
aimock fixture validation: 737/737 tests pass.
PR scope is intentionally narrow per CLAUDE.md "Scope PRs to flagged
findings": this closes ONE userMessage gap. Shape-A 503s (no tools key) and
other userMessage gaps remain as separate follow-ups.
Staging strict-mode replay confirmed ag2 was the only integration missing a
fixture for `Open Excalidraw and sketch a system diagram...` + `create_view`
(the D5 mcp-apps probe's turn-1). All 11 other contexts (langgraph-python,
google-adk, ms-agent-{dotnet,python}, strands, llamaindex, built-in-agent,
claude-sdk-{typescript,python}, langroid, mastra) returned 200 against
aimock-staging with the same payload; ag2 returned 503
`{"code":"no_fixture_match"}`.
Add the missing turn-1 entry to ag2/mcp-apps.json, mirroring the LGP
gold-standard at langgraph-python/tool-rendering-reasoning-chain.json (same
fixture id `call_d5_mcp_apps_create_view_001`, same arguments payload, same
chunkSize). Match keys: userMessage + toolName + context.
Local GREEN proof: iso9 aimock (port 6010) with new fixture mounted returns
200 with the canonical tool_call payload for the exact staging-replay request
body. LGP regression check (same probe with X-AIMock-Context: langgraph-python)
stays green.
Scope intentionally narrow per "scope PRs to the originally-flagged findings":
this PR closes only the one genuine fixture gap identified by the staging
verification audit. The 9 other 503s in that audit (sales-dashboard across 9
slugs + google-adk Excalidraw) are NOT fixture gaps — they are backend
tool-array forwarding issues and are tracked separately.
Re-tier the showcase docs tree to be an agent entry point: README.md
opens with a 'when X, see Y' fanout table that routes to the right
procedural doc; each procedural doc gets a one-line tagline answering
'what does this answer'.
Consolidation:
- DELETE showcase/RUNBOOK.md — operational content merged into DEBUGGING.md
(Integration Patterns, Docker Compose Environment, Production Debugging,
Anti-Patterns, Aimock Fixture Deployment, Dev Iteration Speed). The
--isolate mechanics + CLI rules were already duplicated in DEBUGGING.md.
- DELETE showcase/QA-COVERAGE.md — per-demo coverage matrix + starter hero
matrix + probe depth + infra locations + gaps folded into TESTING.md as
the 'Per-Demo Coverage Matrix' section.
Taglines added (no behavioral change to content): TESTING.md, DEBUGGING.md,
GOTCHAS.md, INTEGRATION-CHECKLIST.md, STYLING-GUIDE.md, FRONTEND-STRATEGY.md,
RAILWAY.md, bin/README.md, aimock/README.md, aimock/RAILWAY.md,
harness/README.md, harness/docs/rotation-drill.md.
Cross-link fixups: FRONTEND-STRATEGY.md (was QA-COVERAGE.md →
TESTING.md#per-demo-coverage-matrix), TESTING.md (removed dangling RUNBOOK
companion reference), README.md (rewritten as fanout entry + retained
from-scratch setup + dashboard SOPs below the fanout).
PARITY_NOTES.md × 12 left alone (per-slug context, not redundant).
(cherry picked from commit 75c9d9755c9118c8abc1fa52deda2012b768cab1)
(cherry picked from commit b64189bae0fe2c9e3a5e3ca440013deb4121f23b)
Mirrors A19b's BIA fix pattern for the Anthropic-family csdkts integration.
Root cause: csdkts uses Anthropic SDK which generates its own toolCallIds (toolu_*) rather than echoing aimock's prescribed call_d6_cc_*. The fixture's toolCallId-gated narration entries never matched on turn-2, causing fall-through to less-specific entries (or 503/no-match).
Fix: replace toolCallId discriminator with turnIndex (count of role:assistant messages). turnIndex is backend-id-invariant — it works regardless of how the backend rewrites tool_call_id values. Same shape as A19b BIA fix.
- Tokyo narration: toolCallId → turnIndex: 1
- AAPL narration: toolCallId → turnIndex: 3
- AAPL emit: added turnIndex: 2
- Tokyo emit: turnIndex: 0
response.content + canonical phrase ("rendered through the custom wildcard catchall") and response.toolCalls UNTOUCHED.
Verified locally on cr5495/fix-a20-csdkts-green at HEAD d178e6730 (post-A21b):
- /tmp/cr/a20v6-green-csdkts.log: 1 passed, INNER_EXIT=0
- iso2 slot, full infra healthy (aimock+pocketbase+dashboard+csdkts)
(cherry picked from commit e66e0eb0ce72c970348183eeb4f4b57c3f5b1d29)
Replace per-leg toolCallId pin with turnIndex (assistant-count) +
userMessage on the two narration fixtures, and add explicit turnIndex
to the AAPL-emit fixture, so first-match-wins partitions the four
request shapes BIA produces against the OpenAI Responses API.
Root cause:
- BIA uses @tanstack/ai-openai openaiText('gpt-4o') which calls
/v1/responses. aimock converts each /v1/responses request to a
chat-completions-shaped completionReq via responsesInputToMessages()
and matches with the same router. The matcher's toolCallId check is
strict equality against the last message's tool_call_id.
- BIA's TanStack runtime auto-generates tool_call_id at request time
(e.g. 'fc-fCgLtvquOtRpCJTM'), so the fixture-side literal
'call_d6_cc_weather_001' / 'call_d6_cc_stock_001' never matched.
Result: 503 STRICT no-fixture-match on the narration turns, BIA
agent looped on AAPL emit indefinitely.
Fix shape:
- Tokyo narration: toolCallId -> turnIndex: 1
- AAPL narration: toolCallId -> turnIndex: 3 (was off-by-one until I
accounted for the Tokyo-narration assistant message itself adding
to the assistant-count tally seen at AAPL emit time)
- AAPL emit: add turnIndex: 2 so first-match-wins partitions emit vs
narration on the second prompt's two turns
Verification (worktree wt-5495-a19-bia-record, slot iso6):
- RED: bin/showcase test built-in-agent:tool-rendering-custom-catchall
--d5 --isolate -> state=red, 0 passed/1 failed, INNER_EXIT=1
(aimock journal: 3x 503 'No fixture matched' on turn-2 narration
request; AAPL emit fixture matched repeatedly = infinite loop)
- GREEN: same command, post-fix and aimock-restart so the container
reloads the fixture -> state=green, 1 passed, INNER_EXIT=0; aimock
journal: 4 requests, all 200, clean progression
Tokyo-emit (asstCount=0) -> Tokyo-narrate (asstCount=1) ->
AAPL-emit (asstCount=2) -> AAPL-narrate (asstCount=3)
- LGP regression: bin/showcase test
langgraph-python:tool-rendering-custom-catchall --d5 --isolate ->
state=green, 1 passed, INNER_EXIT=0 (uses its own fixture under
aimock/d6/langgraph-python/ — untouched by this change)
Constraints honored: response.content and response.toolCalls
preserved verbatim; canonical narration phrases unchanged; only the
match keys (and their explanatory _comment fields) were modified;
other integrations' fixtures and the harness probe were not touched.
(cherry picked from commit 60d027a1376ba72309b5b3fcf94cc26c64757307)
Parent commit 9491b8934 (fix(showcase): disjoint catchall userMessages + content-asserting probes) changed the d5-tool-rendering-custom-catchall probe userMessages to 'Forecast Tokyo through the wildcard renderer' / 'Quote AAPL through the wildcard renderer' but did not add matching entries to LGP-gold's fixture. Result: aimock no-match -> agent_run_error_event -> SSE-missing -> probe RED on LGP with zero bubbles mounted.
This commit adds 4 entries (2 emit + 2 toolCallId-gated narration) following the canonical pattern used by the other 17 integrations. After this commit, LGP bubbles mount and narrations settle with the canonical phrase in bubble.textContent, matching BIA's end-state.
NOTE: A separate fleet-wide probe-layer bug (validateCustomCatchall's customContentPhrasePresent page-wide DOM scan returns false even when phrase is in bubble.textContent) keeps the probe RED in this session. That's a separate concern, to be fixed in a follow-on commit. This commit is a strict improvement: pre-fix LGP got zero bubbles + agent_run_error_event; post-fix LGP renders both bubbles and both narrations settle with the phrase.
(cherry picked from commit 62e04976c3b292a27e1b1cc0bf2bb6fda47db786)
After A2 (commit 6c596d8d6) stripped toolName from the primary emit entries, sibling turnIndex:0 fallback entries became unreachable under first-match-wins. The 7 A2-target siblings (ag2/google-adk/lgf/lgts/mastra/msdotnet/csdkts toolName strip) had equivalent dead fallbacks deleted in that commit; csdkts was inconsistently treated. Removing the 2 dead entries restores cross-fleet consistency. Probe contract unaffected — toolCallId-gated narration entries still emit the required content phrase.
(cherry picked from commit 9720636519f4cd858fcdc08ed84597be05604a2e)
crewai-crews/ms-agent-python/pydantic-ai catchall AAPL entries still carried stale $189.42/up 1.27% narration and toolCall args lacking price_usd/change_pct. Round 1 CR finding #3 was identified but never given a fix agent. Brought all 3 to canonical shape matching built-in-agent (A5) and csdkts narration (A6): price_usd=338.37, change_pct=-2.96, narration 'AAPL is trading at $338.37, down 2.96% on the day — rendered through the custom wildcard catchall.'
(cherry picked from commit df84657310451500278d0b3d0125c9c490042d2b)
ag2/claude-sdk-typescript/google-adk/langgraph-fastapi/langgraph-typescript/mastra/ms-agent-dotnet: AAPL is turn-1 so turnIndex:0 fallback unreachable; toolName gate fails on wildcard-renderer integrations that don't register get_stock_price. Aligned to LGP-gold pattern (userMessage+context discriminator, no toolName, no turnIndex:0 fallback). Preserved legitimate multi-pill matchers (SF/flights/d20/chain) on the 4 multi-pill integrations.
(cherry picked from commit c10821b5d904b31bee2ab2a39db3b1565aecba26)
built-in-agent shipped $189.42 vs the rest of the fleet's $338.37. Drift makes any content-asserting test on AAPL price brittle. Aligned.
(cherry picked from commit c0796da2a63abfeaa1d8ea06200df0754d30a2e6)
spring-ai retained hasToolResult:false on the Tokyo weather emit fixture — outlier vs LGP-gold and the other 16 catchall fixtures. Aligned: userMessage+context discriminator only.
(cherry picked from commit d268a88c4a3f405dfcc15b2aaccc190e61cd7ac7)
built-in-agent, crewai-crews, ms-agent-python, pydantic-ai: Tokyo turn-1 tool result makes hasToolResult permanently true → AAPL fixture never matched → 30s timeout. Aligned with agno/langroid/llamaindex/strands/claude-sdk-python pattern: rely on userMessage+context (and toolCallId where relevant) as the gate.
(cherry picked from commit 8e313cd1b043cf1c61efb82143171a28d7699c49)
Closed PR #5465 introduced cross-fixture leakage by stripping the
toolName discriminator on shared {userMessage, context} keys: with
aimock's alphabetical first-match-wins ordering,
tool-rendering-custom-catchall.json sorts before
tool-rendering-default-catchall.json, so default-catchall page requests
were served the custom file's content. The probe was structurally blind
because it asserted only DOM testids (copilot-tool-render +
data-tool-name=get_weather) — both fixtures emit get_weather, so the
testid signal passed regardless of which fixture won.
Real LGP-gold pattern is disjoint userMessages between default and
custom catchall fixtures (NOT shared keys discriminated by toolName).
This PR ports the LGP-gold pattern to the other 17 integrations and
adds page-text content assertions to both d5 catchall probes so this
class of regression can't recur silently:
- Probe prompts disjoint between default-catchall and custom-catchall
- Fixture userMessages updated to match the new disjoint prompts
- Page-text content assertions in both d5 catchall probes
(default negatively asserts the custom-catchall leak phrase;
custom positively asserts it)
Local gold-standard red-green proof captured:
- Step A: live staging Playwright baseline (RED-OF-RECORD)
- Step B: local control-plane on origin/main reproduces structural
fragility (probes pass at testid level, custom fixture wins on
default's userMessage path)
- Step C: local control-plane on this branch — both catchall probes
GREEN with the new content-asserting assertions
tool-rendering-default-catchall: pass=true (5443ms)
tool-rendering-custom-catchall: pass=true (8956ms)
cross-tool signature passed
LGP (langgraph-python) was already disjoint; it remains untouched.
PR #5459 added userMessage matchers ("forecast for Tokyo" + "current price
of AAPL") on the catchall fixtures, which flipped most integrations from
red to green. Six cells stayed red on staging because the FIXTURE SHAPE
itself was broken on the matched fixtures, not just the userMessage key.
Failing cells (all custom-catchall except strands default):
d5:agno/tool-rendering-custom-catchall
d5:llamaindex/tool-rendering-custom-catchall
d5:langroid/tool-rendering-custom-catchall
d5:claude-sdk-python/tool-rendering-custom-catchall
d5:strands/tool-rendering-custom-catchall
d5:strands/tool-rendering-default-catchall
PocketBase confirms the failure mode: turn 1 (Tokyo) completes; turn 2
(AAPL) times out at 30s or renders the wrong content. E.g. agno:
errorDesc: "timeout: assistant did not respond within 30000ms"
failure_turn: 2
turns_completed: 1 / 2
Root cause: the failing tool-emit fixtures used `hasToolResult: false`
(or `toolName: "<tool>"`) as their gate. aimock's hasToolResult check is
`messages.some(m => m.role === 'tool')` over the WHOLE thread — so once
turn 1's Tokyo tool result lands in the conversation, hasToolResult is
permanently true and `hasToolResult:false` can never match turn 2 → no
fixture → 30s timeout. `toolName:get_stock_price` likewise fails when an
integration backend doesn't forward the tool definition on turn 2.
The same trap is documented in
showcase/aimock/d6/langgraph-python/tool-rendering.json:
"_comment": "Gated on toolName:get_stock_price rather than
hasToolResult:false. The D5 tool-rendering-custom-catchall probe runs
'weather in Tokyo' first, which leaves a get_weather tool result in
the thread; the aimock router implements hasToolResult as
messages.some(m=>m.role==='tool'), so hasToolResult is permanently
true on the AAPL turn and a hasToolResult:false gate could never
match (→ no_fixture_match → 503 → 30s timeout)."
Fix: align all 6 files to the canonical pattern used by mastra/spring-
ai/built-in-agent on this probe:
1. toolCallId-keyed narration fixture FIRST
2. tool-emit fixture SECOND with ONLY `userMessage` + `context`
(no hasToolResult / toolName gate)
The toolCallId narration uses aimock's
`messages[last].role === 'tool' && tool_call_id === ...` check, so it
correctly wins on iteration 2 (post-tool-result) without being affected
by older turns' tool results. The tool-emit fixture matches turn 1 (last
message is user) and re-emits only when the narration above hasn't
matched.
Strands' two cells additionally needed REORDERING — they had the tool-
emit fixture before the toolCallId fixture, defeating first-match-wins.
Strands' staging backend was also returning 502 during testing; once it
recovers, the corrected fixtures should let the probe pass. The fixture
changes are necessary but may not be sufficient for strands if backend
remains down.
No probe-side, harness, or backend changes — pure fixture-content
alignment. Six fixture files modified; line totals: -87 / +68.
The catchall userMessage rename (#5453) only renamed STALE strings to 'forecast for Tokyo'. Integrations whose catchall fixtures used PILL-ALIGNED userMessages ('What's the weather in San Francisco?', 'Find flights from SFO to JFK.', 'Chain a few tools in this single turn', etc.) had no 'forecast for Tokyo' matcher to begin with — so the D5 catchall probe still fell through to live LLM and the cells stayed red on staging.
This PR ADDS the canonical 'forecast for Tokyo' (default+custom catchall) and 'What's the current price of AAPL?' (custom catchall only) emit+narrate fixture pairs to each integration that was missing them. Existing pill-aligned fixtures are preserved (pure prepend at the start of the fixtures array; first-match-wins means new fixtures match the D5 probe inputs without colliding with existing pill prompts).
Integrations touched:
- langgraph-typescript: both default + custom
- langgraph-fastapi: both default + custom
- google-adk: custom only (default already green)
- ms-agent-dotnet: custom only (default already green)
After merge, fleet-cp's e2e-deep probe will re-run within ≤6h staleness window and flip cells GREEN.
## Summary
- D5 e2e-deep probes for `tool-rendering-{default,custom}-catchall` send
`"forecast for Tokyo"` as the test input (see
`showcase/harness/src/probes/scripts/d5-tool-rendering-{default,custom}-catchall.ts`).
- aimock uses substring match on `userMessage`.
- The catchall fixtures on main had a stale `"check Tokyo weather
forecast"` string that couldn't substring-match the probe input →
fixture miss → probe falls through to live LLM → CV ✗ red D4 on the
dashboard.
- This PR renames the userMessage to the canonical `"forecast for
Tokyo"` across 28 catchall fixture files in 16 integrations.
## Scope
28 files × ~2 userMessage occurrences each = 67 line changes. **No code,
no agent, no page.tsx changes.** Pure fixture-data alignment.
Integrations covered (default-catchall and/or custom-catchall): ag2,
agno, built-in-agent, claude-sdk-python, claude-sdk-typescript,
crewai-crews, google-adk, langgraph-python, langroid, llamaindex,
mastra, ms-agent-dotnet, ms-agent-python, pydantic-ai, spring-ai,
strands.
## Commit history note
Commit b1f19bdc8 changes the rename target from the original `"weather
in Tokyo"` (commit 388c69e68) to `"forecast for Tokyo"`. This was a CR
Round 1 catch: `"weather in Tokyo"` was a substring of the chain pill
prompt (`"weather forecast chain in Tokyo"` / similar), which would have
caused the catchall fixture to incorrectly match chain-pill traffic.
`"forecast for Tokyo"` has no such substring collision with any other
probe input.
## Verification
- **Live red-green proof** was performed on the prior `"weather in
Tokyo"` rename (commit 388c69e68) against `ms-agent-python` via
`showcase/bin/showcase test ms-agent-python --d6 --isolate`:
- Before (origin/main): catchall featureTypes red (fixture miss → live
LLM → flaky)
- After: `d6:ms-agent-python/tool-rendering-default-catchall=green`,
`d6:ms-agent-python/tool-rendering-custom-catchall=green`
- **The current HEAD's `"forecast for Tokyo"` rename (b1f19bdc8) has NOT
been re-run live.** It is verified by static analysis only: substring
math (no collision with any known D5 probe input or chain pill prompt)
and a clean CR Round 2 across all reviewing agents.
- Post-merge dashboard re-probe will be the final runtime verification.
## Out of scope (separate follow-ups)
- LG-TS / LG-FastAPI catchall fixtures don't have the stale string — use
pill-aligned userMessages; need different fix
- `tool-rendering` (non-catchall) and `tool-rendering-reasoning-chain`
featureTypes have separate failure modes
- Other red cells in the dashboard (frontend-tools-cosmic timeout,
agent-config, auth, etc) are unrelated
The previous 'weather in Tokyo' rename (388c69e68) was a substring of the
main tool-rendering chain pill prompt 'Chain a few tools in this single
turn: get the weather in Tokyo, search flights from SFO to Tokyo, and roll
a d20.' Because aimock loads fixtures alphabetically per integration dir
and uses substring match with first-match-wins, the catchall fixture
(loaded before tool-rendering.json) was intercepting chain pill matches
across 16 integrations.
Rename catchall fixture userMessage to 'forecast for Tokyo' — a phrase
not contained in any other pill prompt. Update the corresponding D5
catchall probe inputs in d5-tool-rendering-{default,custom}-catchall.ts
and the test assertions that pin those inputs.
Call-Site Enumeration: 'weather in Tokyo' remains intentionally in
page.tsx suggestions.ts pills (user-visible UX) and inside the chain
pill prompt itself — neither is in the substring-match path now.
The D5 e2e-deep probe for tool-rendering-{default,custom}-catchall sends
"weather in Tokyo" as the test input (harness/src/probes/scripts/
d5-tool-rendering-{default,custom}-catchall.ts). The fixture userMessage
matcher uses substring match. Main's catchall fixtures had a stale
"check Tokyo weather forecast" string that could not substring-match
the probe input, causing the fixture to miss and the probe to fall
through to the live LLM — surfacing as CV x red D4 on the dashboard.
Rename to the canonical "weather in Tokyo" string across affected
integrations' tool-rendering-{default,custom}-catchall.json files.
No agent or page.tsx changes; fixture content otherwise unchanged.
The custom-catchall probe sends 'weather in Tokyo' then 'AAPL'. After the
weather tool runs in turn 1, hasToolResult is true across the rest of the
thread — which fires tool-rendering.json's AAPL 'hasToolResult:true' narration
prematurely on turn 2 iteration 1, returning prose without ever emitting the
get_stock_price tool. The custom-catchall assertion (both tools rendered
through the wildcard testid) then fails with missing get_stock_price.
Replace the (hasToolResult:true narration + hasToolResult:false/turnIndex:0
emitter) layered fallbacks with a sequenceIndex:0 emitter ordered before a
bare userMessage+context narration. The per-test fixture-match counter resets
each run, so the emitter fires exactly once on iteration 1 regardless of
prior pills' tool history, then falls through to the narration on
iteration 2. Applied symmetrically to the two AAPL blocks in
tool-rendering.json (the 'What\'s the current price of AAPL?' block at the
top and the legacy 'current price of AAPL' alias block lower down). The
toolCallId-keyed narration above each block is retained for the non-BIA
fast path.
The earlier partial fix to tool-rendering-custom-catchall.json is kept (it
adds toolCallId-scoped narration legs ordered before the existing
hasToolResult:true narrations); those fixtures never match real probe
traffic (the probe sends 'weather in Tokyo' / 'current price of AAPL', not
the unique 'check Tokyo weather forecast' substring in this file) but the
reordering is consistent with the cross-file pattern and harmless.
Verified locally: built-in-agent:tool-rendering and
built-in-agent:tool-rendering-custom-catchall both green via
`./bin/showcase test ... --d6 --direct`; built-in-agent:tool-rendering-default-catchall
also green; aimock-fixtures.test.ts (collision/shadow ceilings) unchanged.
Hero loses its surrounding card (bare KPI strip over the chart cards,
pinned to all six months); team performance pairs the rep table with a
quota-attainment bar chart; top account pairs the fact card with a
product-line pie (new dataset entry); at-risk becomes a risk panel — KPI
strip (ARR at risk / accounts / biggest exposure) over three side-by-side
severity cards with reason + next action. Fixtures re-captured from live
responses; D5 probe drops declarative-card from the hero set; e2e asserts
the accompanying charts and the risk panel; QA docs updated.
Replaces the hand-authored surface payloads with real gpt-5.4 responses
captured via the langgraph dev threads API during OSS-136 prompt
iteration (catalogId injected, since the replay path resolves the catalog
from recorded tool args). Verified: fixture suite 738/738, LGP container
e2e 6/6, ADK container e2e 6/6.
Same three-call choreography per pill (outer generate_a2ui, inner design
toolcall discriminated by toolName, narration matched by toolCallId), now
keyed to the sales-analyst prompts and emitting Vantage Threads payloads:
composed hero dashboard, rep DataTable, at-risk StatusBadge cards, and
top-account InfoRows. Verified by the LGP Playwright run against the
aimock-backed container (6/6).
BIA registers headless-complete's tools (get_weather, get_stock_price,
get_revenue_chart, highlight_note) as server-executed via TanStack's
chat() engine. After the LLM returns a tool call, TanStack runs the
server tool and reprompts the LLM with the result. The userMessage-keyed
toolcall fixtures fired again on every reprompt because the original
user pill text stays in conversation history — and the toolCallId-keyed
narration fallback never matched because BIA's /v1/responses endpoint
rewrites the assistant tool_call_id to a runtime-generated fc-… value.
Net effect: turns 3+4 (highlight, revenue chart) ballooned to 70+
assistant messages within the 60s timeout window.
Restructured each pill as a (sequenceIndex:0 emitter, narration
fallback) pair. The emitter matches the FIRST request for the pill
prompt (counter starts at 0) and emits the tool call; subsequent BIA
reprompt iterations fall through the now-exhausted emitter to the
narration fallback (no tool call), so the loop converges. sequenceIndex
is chosen over hasToolResult:false because hasToolResult is computed
across the entire thread — any earlier pill's tool result would
permanently disable a hasToolResult:false emitter, breaking multi-turn
sessions where turn 2+ would have hasToolResult=true globally.
The legacy bare 'AAPL' userMessage matchers in tool-rendering.json were
substrings of gen-ui-headless-complete's 'price of AAPL right now'
pill prompt, so when the BIA TanStack server-tool reprompt loop hit
iteration 2 (toolCallId chain broken by /v1/responses id rewriting),
the request fell through to tool-rendering.json's loose 'AAPL'
matchers and leaked wrong-card content. Tightened all four bare 'AAPL'
match keys to 'current price of AAPL' — preserves matching for the
tool-rendering 'Stock price' pill ("What's the current price of
AAPL?") and the custom-catchall probe's second prompt ("What's the
current price of AAPL?") while no longer matching the headless
probe's "What's the price of AAPL right now?".
The injected/streamed a2ui fixtures all included catalogId, so aimock
replay never exercised the basic-catalog fallback that broke production
(real models omit catalogId per the tool-usage guide). Strip catalogId
from the langgraph-python sales-dashboard secondary-call fixtures and
hard-assert "Catalog not found" is absent outside the charts-rendered
soft branch, so the spec fails without a route defaultCatalogId.
Also repoint the on-demand e2e workflow at the d4/d5-recorded/d6/shared
fixture dirs — it still referenced feature-parity.json, deleted in the
1e66a5f8d fixture reorg, so every /test-aimock run died at aimock start.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two aimock fixture updates to unblock built-in-agent D6 verification:
- gen-ui-a2ui-fixed.json: new fixture so display_flight pill routes
the Card mount through the a2ui-fixed-schema renderer host
- gen-ui-declarative.json: reshaped to the built-in-agent two-call
sequence with the flat catalog shape the BIA route emits
Adds the missing gen-ui-a2ui-fixed aimock fixture and reshapes the
declarative-gen-ui fixture so the claude-sdk-python D6 probes have
deterministic LLM replay for both gen-ui demos.
- SUPERSEDED annotations on unreachable recorded calculator entries (load-order shadowing is
the only guarantee; model gate is not a safety net)
- GOTCHAS sequenceIndex rewritten to per-X-Test-Id semantics with co-increment/eviction caveats
- hasToolResult paragraph corrected (omission = no gate, thread-global predicate)
- statelessness claim reconciled with sequence counters
- replace Websandbox host-bridge evaluateExpression (never registered in beautiful-chat) with
in-sandbox allowlist-gated Function() eval
- reorder specific calculator pair before generic (first-match-wins)
- allow exponential notation (eE) so chained evaluation of large results works
- drop thread-global hasToolResult gate that broke the pill after other tool-producing pills,
reorder toolCallId follow-ups before leg-1
- mint distinct tool_call_ids per repeat click via sequenceIndex variants + non-sequenced
fallback (fixes second-widget collapse)
- restore google-adk trailing newline
- document ordering invariants, gate tradeoffs, and sequence-counter scoping caveats in
fixture comments
Verified via local Playwright red-green (render, 7+8=15, chained exponent, interleaved pills,
repeat clicks).
The declarative-gen-ui demo across the three langgraph integrations now
relies on the middleware to inject and execute generate_a2ui — the agents
collapse to create_agent + CopilotKitMiddleware with no hand-rolled tool.
Adds render_a2ui fixtures for the new tool path and pins the integrations
to the A2UI alpha SDKs (copilotkit 0.1.94a1, @copilotkit/sdk-js 1.59.3-alpha.1).
Audit follow-up to #5232, which fixed the gen-ui-interrupt d6 fixture
leg mis-order for langgraph-python + langgraph-typescript. aimock's
matchFixture is first-match-wins in array order and the loader dedups
on userMessage, so a toolCallId resume-leg placed BEFORE its
toolName:schedule_meeting first-leg shadows the tool-emitting leg on
turn-2 — the interrupt never fires and time-picker-card never mounts.
Fleet audit of every showcase/aimock/d6/*/gen-ui-interrupt.json found
ag2 as the only remaining mis-ordered fixture (both toolCallId
resume-legs preceded their toolName first-legs). Reorder ag2 so each
pill's toolName first-leg precedes its toolCallId resume-leg, matching
the proven-correct mastra/pydantic-ai/langgraph pattern. Reorder only;
response payloads and match keys are unchanged.
All other gen-ui-interrupt-supporting integrations (built-in-agent,
claude-sdk-*, langgraph-fastapi, langroid, ms-agent-dotnet/python,
pydantic-ai, spring-ai, strands) were already correctly ordered.
The gen-ui-interrupt d6 probe drives two interrupt turns in one thread.
Turn-2 failed because the langgraph-python/typescript fixtures ordered each
pill's toolCallId resume-leg BEFORE its toolName first-leg. aimock dedups on
userMessage (first-match-wins), so the text-only resume-leg shadowed the
tool-emitting first-leg on the duplicate userMessage — the agent emitted a
plain text reply with no schedule_meeting tool call, the interrupt never
fired, and the time-picker-card never mounted.
Reorder each pill so the toolName:schedule_meeting first-leg precedes its
toolCallId resume-leg, matching the passing mastra/pydantic-ai ordering. The
resume-leg still only matches when the last message is its tool result
(toolCallId guard), so confirmation still works.