## Summary
Aligns the **langgraph-fastapi** showcase integration to the north-star
**langgraph-python** (OSS-582 — "Align LangGraph (FastAPI) Showcase
demos and code"). Moves the integration onto the product-centric demo
set: backend agent graphs, runtime wiring, frontend chrome, the
demo-browser overview (`manifest.yaml`), and aimock fixtures.
Linear: OSS-582.
## What changed
**Backend agents (`src/agents/src/`)**
- Ported the missing dedicated graphs: `agentic_chat`, `gen_ui_agent`,
`gen_ui_tool_based`, `shared_state_streaming`.
- v2 `create_agent` ports + config alignment for `agent_config_agent`,
`headless_complete`, `reasoning_agent`, `tool_rendering_agent`,
`tool_rendering_reasoning_chain_agent`, `a2ui_dynamic`.
**Runtime wiring**
- `route.ts`: wired the canonical demos to their dedicated graphs and
removed them from the generic `sample_agent` fallthrough loop (this is
what made several demos behave correctly against a real LLM — see
below); added `recursion_limit`.
- `langgraph.json`: registered the 3 new graphs.
- Renamed the dedicated API routes
`copilotkit-byoc-{hashbrown,json-render}` →
`copilotkit-declarative-{hashbrown,json-render}`.
**Overview / content**
- `manifest.yaml`: demos + features aligned to LGP (names, descriptions,
tags, order, canonical ids; deprecated aliases migrated; phantom
`hitl-in-chat-booking` removed; `shared-state-streaming` un-quarantined
now that it works). Integration identity (name/slug/logo) preserved. The
demo-browser overview is now card-for-card identical to
langgraph-python.
- Restored two missing `@region` markers (factory-automation snippet
extraction).
**Frontend chrome**
- `globals.css`, `layout.tsx`, `middleware.ts` (was missing),
`declarative-gen-ui/*` + `sales-context.ts`, `beautiful-chat`;
`Dockerfile` now copies `manifest.yaml` (fixed an RSC crash on every
`/demos/*`).
**aimock fixtures (`aimock/d4|d6/langgraph-fastapi/`)**
- D6 fixture fixes for the previously-red cells (stale `turnIndex`
gates, cross-file substring shadowing, missing `chunkSize`, prompt
narrowing).
**Showcase tooling / docs**
- Fixed harness services racing on a shared image tag
(`docker-compose.local.yml`).
- `GOTCHAS.md` #8: documents that aimock D6 can be green while a demo is
broken against a real LLM (fixtures replay scripted tool calls
regardless of which graph ran), plus how to catch it.
- `PARITY_NOTES.md`: sanctioned divergences (a2ui-recovery per-slug
prompt isolation; declarative-json-render scoped-test divergence).
## Verification
- **D6 sweep: 38 green / 1 red.** The single red is **a2ui-recovery**,
which reproduces **identically on the north star** (shared Python
`ag_ui_langgraph` recovery loop; mastra's TS impl is green). It is not a
fastapi defect and is tracked as a separate PR against the north star.
Documented in `PARITY_NOTES.md`.
- `validate-manifests` (manifest → registry, all 20 integrations) —
green.
- `validate-routes --all` (runtime-route wiring) — green.
- Build + TypeScript typecheck (`next build` via the Docker image build)
— green.
- `@region` audit — all region pairs balanced and LGP-consistent (0
orphans/typos/mismatches).
- Live manual QA against real OpenAI (`:3102` vs `:3100`) confirmed the
key demos (reasoning-custom, shared-state-streaming, gen-ui-tool-based,
agentic-chat).
## Notable finding
Several demos were D6-green but broke against a real model because they
fell back to the generic `sample_agent` instead of their dedicated graph
— the aimock fixture masked the wiring bug by replaying scripted tool
calls. This PR fixes the wiring and documents the gap (GOTCHAS #8).
D6-green is necessary but not sufficient for graph/tool-dependent demos;
a real-LLM click-through is required.
## Deferred / follow-ups
- **a2ui-recovery** — shared north-star defect in the Python
`ag_ui_langgraph` recovery loop; separate PR against the north star.
- **d5-byoc probe** — always sends the hashbrown pill even on the
json-render page (fleet-wide harness limitation); tracked as a
follow-up. The gating `byoc` grid cell is green.
## Scope note
Only `langgraph-fastapi` integration code, its aimock fixtures, and
showcase tooling/docs changed. **langgraph-python (the north star) was
not touched.**
The factory automation extracts curated snippets via @region markers.
readonly_state_agent_context.py and shared_state_read_write.py were byte-aligned
to langgraph-python except their @region markers had been stripped, so the
factory would extract nothing for those two regions. Restore them to match LGP:
- agent-context-setup around create_agent in readonly_state_agent_context.py
- shared-state-setup around create_agent in shared_state_read_write.py
Comment-only; no behavior change. Audit confirms all 49 fastapi region pairs are
now balanced and LGP-consistent (0 orphans/typos/mismatches).
Live QA surfaced demos that behaved wrong on fastapi because they fell back to
the generic sample_agent (or were unregistered) instead of the dedicated graph
LGP uses. aimock D6 masked these (fixtures script the tool calls), so they were
green in the grid but broken against a real model. Audit of fastapi's
agentNames fallthrough vs LGP's neutralAssistantCells found exactly these:
- gen-ui-tool-based: port gen_ui_tool_based graph (tools=[], frontend supplies
render_bar/pie_chart via useComponent). Was sample_agent, whose query_data
tool + prompt made the model loop on data queries instead of rendering.
- shared-state-streaming: port shared_state_streaming graph (StateStreaming
middleware + write_document tool + document state). Was sample_agent, which
never emits state.document, so it only wrote to chat. Remove from
not_supported_features (now works, D6 green).
- agentic-chat: port agentic_chat graph (tools=[]). Was sample_agent (7+ tools).
- threadid-frontend-tool-roundtrip: wire to frontend_tools (was unregistered).
- reasoning-custom: align reasoning_agent config to LGP (gpt-5.4 / effort
medium / summary detailed; were gpt-5-mini / low / auto).
Register the 3 new graphs in langgraph.json; remove the 3 names from the
sample_agent fallthrough loop in route.ts. D6 green x2 for reasoning-display,
gen-ui-custom, shared-state-streaming, agentic-chat; agent-config +
tool-rendering-reasoning-chain re-verified (no regression).
The demo-browser overview (page.tsx, identical to LGP) groups cards by
demo.tags[0] and orders within a group by manifest.features[] index, so the
card arrangement is fully driven by manifest.yaml. fastapi's manifest was an
older curation: a different naming scheme (~26 demos), deprecated feature/demo
ids (agentic-chat-reasoning, reasoning-default-render, byoc-hashbrown,
byoc-json-render, hitl, phantom hitl-in-chat-booking), divergent tags/order,
and a missing shared-state-streaming/shared-state-read card.
Align manifest.yaml demos+features to LGP (names, descriptions, tags, order,
canonical ids), preserving fastapi identity (name/slug/logo/description) and
the sanctioned a2ui-recovery prompt divergence. Overview is now card-for-card
identical to langgraph-python.
Canonicalizing the ids pulled two cells into the D6 tested set that were
mis-wired to the old names; wire them to LGP's shape:
- reasoning-custom/reasoning-default -> reasoning_agent in copilotkit/route.ts
(were wired under agentic-chat-reasoning/reasoning-default-render).
- rename api routes copilotkit-byoc-{hashbrown,json-render} ->
copilotkit-declarative-{hashbrown,json-render}; align endpoints + the
declarative-hashbrown-demo agent id to what the demo pages request.
Full D6 sweep: 37 green / 1 red; reasoning-display and byoc now green (two
real runs each); the lone red is a2ui-recovery (known shared north-star
defect, identical on LGP).
Two independent fastapi divergences from the green north-star kept this red:
1. Missing chunkSize:9999 on the 8 tool-call fixtures. Under aimock's global
8-byte chunking the large tool-call args failed to JSON-parse in one piece,
so the AAPL->MSFT reasoning chain stopped at leg 1 (turns 1 & 2). Byte-align
to LGP (and the green mastra/langgraph-typescript siblings carry the same
pattern).
2. Fixture-pool shadowing broke turn 3. aimock pools all d4+d6 fixtures for a
context and matches by userMessage substring, first-match-wins in load order
(d4 before d6). fastapi had broad keys LGP doesn't:
- d4/chat.json: 'weather' -> 'weather in San Francisco', 'flights from SFO
to JFK' -> 'Find flights from SFO to JFK.' (period-terminated so it stops
being a substring of turn 3's 'Find flights from SFO to JFK and show me
the weather there.').
- tool-rendering-{custom,default}-catchall.json: bare 'Find flights' ->
'Find flights from SFO to JFK.' (now matches LGP's exact key).
Also byte-align the backend agent to LGP: scriptable get_stock_price signature,
the detailed chaining system prompt, model gpt-5.4, reasoning summary detailed.
D6 tool-rendering-reasoning-chain green (two real ~21s runs). Shared-fixture
regression all green: agentic-chat, tool-rendering, tool-rendering-custom-catchall,
tool-rendering-default-catchall, headless-complete (Tokyo-weather consumer).
14 integrations' `src/app/demos/layout.tsx` hardcoded "LangChain - Python"
in `generateMetadata` — a copy-paste leftover from langgraph-python, which
the file was cloned from. Every `/demos/*` page in mastra, strands, ag2,
agno and 10 others rendered `<title>LangChain - Python</title>`.
Each now uses the display name from its own `manifest.yaml` `name:` field,
matching the convention the already-correct integrations use
(langgraph-typescript -> "LangGraph (TypeScript)", strands-typescript ->
"AWS Strands (TypeScript)").
langgraph-python itself is included: its manifest name and root layout both
say "LangGraph (Python)", so "LangChain - Python" (the legacy Notion
partner-column label) was stale there too.
tool-rendering-custom-catchall renders the ticker quote in the wildcard card.
fastapi's get_stock_price(ticker) had no scriptable args, so when the aimock
fixture emitted price_usd/change_pct the @tool silently dropped the unknown
kwargs and returned random numbers instead of the scripted quote — the card
showed nondeterministic values, diverging from the north star.
Port LGP's scriptable signature: get_stock_price(ticker, price_usd=None,
change_pct=None) — echoes scripted values verbatim when supplied, random when
omitted (backward-compatible with the other tool-rendering pills). Behavior
byte-aligned to langgraph-python.
D6 tool-rendering-custom-catchall green (two real ~8s runs); shared-backend
siblings tool-rendering, tool-rendering-default-catchall, headless-complete all
re-verified green.
Two compounding defects kept this cell red:
1. Backend missing get_revenue_chart. headless_complete.py never registered
the revenue-chart tool, so the chart-turn never emitted a tool call and the
card never mounted. Port LGP's get_revenue_chart tool + its system-prompt
routing rule (tool return shape byte-identical to LGP).
2. Stale turnIndex fixtures caused turn-2 and turn-4 loops to the recursion
limit. Align to LGP's canonical shape:
- tool-rendering.json (shared): AAPL first-leg turnIndex:0 -> hasToolResult:false
so it stops re-firing at turnIndex>=2 in the multi-pill thread.
- headless-complete.json: rewrite to LGP's structure (narration/toolCallId
fixtures first, tool-call legs after with userMessage+context only, no
stale turnIndex); drop the divergent fastapi-only highlight fixtures,
subsumed by the broader substring fixture.
D6 langgraph-fastapi:gen-ui-headless-complete now green (two real ~29s runs);
tool-rendering re-verified green (shared-fixture no regression).
gen-ui-agent had no dedicated backend graph: langgraph.json lacked a
gen_ui_agent entry and route.ts routed the name through the neutral-assistant
loop to sample_agent, which has no steps state or set_steps tool, so the
progress card never mounted (agent hit the default recursion limit of 25).
- Port LGP's gen_ui_agent.py (byte-identical) and register it in langgraph.json.
- route.ts: bind gen-ui-agent to createAgent("gen_ui_agent") and bake
assistantConfig.recursion_limit (default 100) into every LangGraphAgent —
the graph's Python with_config isn't visible to the server runs API, so the
multi-step set_steps walk overran 25. Mirrors langgraph-python.
- aimock: regenerate d6/gen-ui-agent.json from LGP (adds chunkSize:9999 on the
24 tool-call fixtures) and narrow the over-broad d4 chat.json "summarize" key
to "Summarize the sales pipeline" so it stops substring-shadowing the
competitor set_steps chain (and other summarize prompts). Matches LGP.
D6 langgraph-fastapi:gen-ui-agent now green (two real ~14s runs);
agent-config re-verified green (no regression).
The fastapi agent_config_agent still used the v1 StateGraph pattern reading
tone/expertise/responseLength from RunnableConfig[configurable][properties].
The frontend (identical to LGP) publishes those knobs via the v2
useAgentContext hook and the route uses a plain LangGraphAgent, so the old
graph received nothing on configurable and the run errored (turn never
completed, D6 red).
Port the backend to LGP's v2 shape: create_agent + CopilotKitMiddleware with
a single static system prompt that reads the injected context entry. Update
the route comment to describe the useAgentContext path (runtime code
unchanged: plain LangGraphAgent + AGENT_URL env fallback preserved).
D6 langgraph-fastapi:agent-config now green (two real ~28s runs).
Task 1 of OSS-582 (frontend parity). Bring the app chrome + beautiful-chat
byte-identical to the langgraph-python north star:
- globals.css: adopt LGP theme (@theme inline tokens + green palette)
- middleware.ts: add (was missing) — sets x-pathname like LGP
- layout.tsx: adopt LGP structure; keep FastAPI in title/log (identity carve-out)
- beautiful-chat/page.tsx: sync stale doc comment
Also fix a Dockerfile parity gap: fastapi was missing the COPY manifest.yaml
that LGP has. demos/layout.tsx reads manifest.yaml at request time via
generateMetadata (headers() -> dynamic), so without it every /demos/* route
crashed with an RSC render error (ENOENT /app/manifest.yaml).
The gen-ui-declarative D6 probe drives a 4-turn conversation; turn 4
(the "top-account" pill) asserts a `declarative-info-row` surface
(a Card of InfoRow facts next to a PieChart). pydantic-ai,
langgraph-fastapi, and langgraph-typescript rendered the InfoRow
component but never carried the `data-testid="declarative-info-row"`
attribute that the probe (and the green peer integrations such as
langgraph-python) rely on to detect the surface. As a result turn 4
timed out with reason=surface-missing and the cell failed at
turns_completed=3.
This restores parity with the green peers by adding the missing testid
to the InfoRow renderer in the 3 lagging integrations. No other
behavior changes; PrimaryButton already wires actions via the
`dispatch(props.action)` pattern (the local ButtonProps extends
ButtonHTMLAttributes, so onClick is valid — no type error).
Commit 1e0d200f5 added the team-performance pill to d5-gen-ui-declarative,
which requires `[data-testid="declarative-data-table"]` to mount. It added
the DataTable renderer + Zod definition to langgraph-python and 6 others,
but missed 4 integrations whose declarative-gen-ui catalogs were drifted
copies from an earlier snapshot: claude-sdk-typescript, pydantic-ai,
langgraph-fastapi, langgraph-typescript.
Root cause: all 4 integrations serve a `next start` production build whose
`renderers.tsx` defines only Card/StatusBadge/Metric/InfoRow/PrimaryButton/
PieChart/BarChart — no DataTable. The backend SSE stream returns a valid
`render_a2ui` payload containing a DataTable component; with no client
renderer it is silently dropped → declarative probe turn 2 times out with
reason=surface-missing.
Decision: per-integration real files (not symlinks — confirmed by file size
diff: 12042 B vs 13515 B canonical). Added DataTable renderer + definition
to each integration's `declarative-gen-ui/a2ui/` matching their existing
ShadCN/card-based style (sourced from strands-typescript, which is the
correct peer, not the inline-style langgraph-python canonical).
Red-green value-test:
- pydantic-ai: RED (turn 2 surface-missing, 90s timeout, data-table=0) →
GREEN (turn 2 assertions passed, "Here's how the team is tracking...")
- langgraph-typescript: RED (turn 2 surface-missing) →
GREEN (turn 2 assertions passed, probe advances to turn 3/4)
- claude-sdk-typescript: RED confirmed (turn 2 surface-missing); GREEN
blocked locally by separate aimock strict-mode miss on turn 1 — unrelated
class B backend fixture issue, not DataTable. Fix is structurally identical
to pydantic-ai/langgraph-typescript and confirmed correct by Docker build.
- langgraph-fastapi: container not running locally; fix structurally identical.
Gates per-request POST + 2xx Response-status + GET health-probe logs behind SHOWCASE_ROUTE_DEBUG across 19 integrations to stay under Railway's 500-logs/sec cap, while logging non-2xx responses unconditionally so production errors stay visible.
Port the google-adk a2ui-recovery demo to langgraph (python, fastapi,
typescript) and aws-strands (python, typescript). Each ships a dedicated
recovery agent, route, demo page/chat/suggestions, manifest entry, aimock
d6 fixtures, e2e spec, and QA doc.
Backend-owned recovery on langgraph via get_a2ui_tools / getA2UITools
(injectA2UITool=false); auto-inject recovery on the strands adapter path.
Heal stages an invalid-then-valid render via aimock sequenceIndex (the
toolkit validate->retry loop rejects the whole surface, so a single-pass
parse_and_fix heal is ADK-specific and does not apply here). Recovery
prompts are unique per framework and the fixtures carry no context match
field, so they fire for real browser (dojo) traffic, not just the harness.
Also harden the strands declarative-gen-ui composition guide to name the
exact catalog component (Metric, not MetricTile) and update the
generate-catalog + aimock-fixtures test expectations.
The auth demo capped at D4 across integrations because the post-sign-out
rejection banner never rendered. The post-sign-out `agent_run_failed` is
delivered only on the agent-scoped `<CopilotChat onError>` channel — never the
provider-level `<CopilotKit onError>` the demos listened on — so the D5/D6 auth
probe's rejection-surface assertion failed and the cell was capped at D4.
Fix (applied to all 19 integrations whose auth demo reproduced the bug): wire a
stable `handleAuthError` onto the agent-scoped `<CopilotChat onError>` (keeping
the provider handler), key the error surface off auth-error STATE alone with a
clear-on-auth effect (removing the `&& !isAuthenticated` cross-slice race), and
harden the rejection-banner message fallback against nullish error events.
Scope: 19 of 20 integrations. built-in-agent already passes (renders via its
ChatErrorBoundary); claude-sdk-python adapted to its legacy/error-boundary shape.
Bump the canonical CopilotKit pin across all showcase integrations + shell
to 1.61.2 (canonical-pins.json, every package.json + package-lock.json),
which carries CopilotKit#5611: passing a catalog to the provider
(`<CopilotKit a2ui={{ catalog }}>`) now auto-enables A2UI and defaults tool
injection on, so the runtime no longer needs an explicit `a2ui` config.
Demonstrate the feature on the A2UI dynamic (declarative-gen-ui) demos by
removing the now-redundant runtime `a2ui` block (`injectA2UITool: true` +
`defaultCatalogId`) from:
- langgraph-python, langgraph-fastapi, langgraph-typescript
- strands, strands-typescript
- google-adk
The forwarded catalog supplies its own catalogId (sdk-js A2UI middleware
auto-derives `defaultCatalogId` from it), so the previous "Catalog not found"
fallback no longer applies.
Verified: validate-pins drift ratchet unchanged (38 / same hash);
langgraph-python D6 `gen-ui-declarative` green end-to-end (no Catalog-not-found).
The four §6-VERBOSE-only backend boundaries (request.ingress, llm.call.start,
llm.call.response, sse.first_byte) called _emit with no tier_gate, so they
over-emitted at DEFAULT tier — 4 extra events/request vs the middleware family,
breaking the §7 tier budget and cross-backend apples-to-apples parity. Gate
them with tier_gate=_VERBOSE_TIERS, matching emit.ts:58-63 and the agno
_BOUNDARY_TIER. langgraph-fastapi received the identical change (the two LGP
files differ only by docstring/plan-unit/_SLUG). Adds default-suppressed +
verbose-emits red-green coverage; updates the pre-existing first_byte
correlation test to drive at VERBOSE tier (the boundary is VERBOSE-only).
Each integration's interrupt-headless demo defines a local useHeadlessInterrupt
hook around the framework useInterrupt. Slot-2 originally identified 8
quarantined integrations (claude-sdk-typescript, langgraph-{fastapi,python,
typescript}, langroid, pydantic-ai, spring-ai, strands); review-round
follow-ups extended the sweep to llamaindex, mastra, ag2, agno, and
crewai-crews (5 more integrations sharing the same byte-identical hook).
The demo-local resolve() previously fire-and-forgot copilotkit.runAgent(...)
via `void runAgent(...).catch(() => {})`. Mirroring the framework fix:
- Make resolve async, return await copilotkit.runAgent(...).
- Use a pendingRef so resolve has stable identity (drop pending from
useMemo deps).
- Type signature: resolve: (response: unknown) => Promise<unknown>.
- Wrap in try/catch + setPending(null) + console.error + rethrow,
symmetric with the framework hook.
- onRunFailed also setPending(null).
13 integrations patched byte-identically.
The injected render_a2ui tool guide instructs models to omit catalogId
("the catalog id is set by the host"), and backend-owned generate_a2ui
tools see real models omit or late-stream it. Without defaultCatalogId
the a2ui middleware falls back to the spec basic catalog, which no
showcase page registers — surfaces fail with "Catalog not found:
https://a2ui.org/specification/v0_9/basic_catalog.json" (reported on
beautiful-chat / langgraph-python).
Pin each route to the catalog its page registers: beautiful-chat ->
copilotkit://app-dashboard-catalog, declarative-gen-ui ->
declarative-gen-ui-catalog. Routes with no a2ui block never attach the
middleware and are left untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The declarative-gen-ui demo across the three langgraph integrations now
relies on the middleware to inject and execute generate_a2ui — the agents
collapse to create_agent + CopilotKitMiddleware with no hand-rolled tool.
Adds render_a2ui fixtures for the new tool path and pins the integrations
to the A2UI alpha SDKs (copilotkit 0.1.94a1, @copilotkit/sdk-js 1.59.3-alpha.1).
Stages the canonical suggestion pill set (mirrored from langgraph-python) as new
suggestions.ts files across 13 integrations: ag2, agno, mastra, pydantic-ai,
claude-sdk-python, claude-sdk-typescript, llamaindex, langroid, strands, spring-ai,
built-in-agent, crewai-crews, langgraph-fastapi.
Also includes targeted edits to existing suggestions.ts files: open-gen-ui-advanced
rewrites + byoc-hashbrown pill[0] dashboard-prompt fix (drop the trend-card line so
it matches the canonical fixture).
NOTE: these new files are currently UNWIRED. Each integration's page.tsx still
defines its pill list inline via useConfigureSuggestions. Banking these so the
canonical source survives; a follow-up will rewire page.tsx to import from
suggestions.ts and delete the inline copies.
Add a per-integration header-forwarding shim so inbound x-* request headers
ride along to outbound LLM HTTP calls. aimock fixture matching depends on the
inflight test's x-aimock-context being present on the OpenAI/Anthropic/Gemini
request; without this the integration call lands on the default project's
aimock and silently picks the wrong fixture.
Shape per integration:
- New _header_forwarding.{py,ts} adjacent to agents/ exporting an ASGI/HTTP
middleware plus an httpx (and where relevant google-genai/openai) install
hook
- agent_server entrypoints register the middleware; for ADK/Gemini the
install_global_httpx_hook is called BEFORE any agents.* import because
google-genai constructs its httpx client at module-import time
Covered: ag2, agno, claude-sdk-python, claude-sdk-typescript, crewai-crews,
google-adk, langgraph-fastapi, langroid, llamaindex, mastra, ms-agent-python,
pydantic-ai, strands. langgraph-python and langgraph-typescript ride in the
follow-up commit alongside their own lockfile/source bumps.
The // @endregion[reasoning-block-render] comment was indented inside the
Chat function body, causing the rendered docs snippet to omit the final
closing brace — a visible syntax error. Moves the marker to after the }
in all 16 agentic-chat-reasoning/page.tsx files.
Also wraps the custom-reasoning snippet in reasoning.mdx in a two-tab
block so the ReasoningBlock import in page.tsx links directly to the
reasoning-block.tsx component definition in the adjacent tab.
Run the unified hoist codemod over showcase/integrations/* and adjacent
source roots (src/lib, src/agent, src/mastra, src/main/java for Spring AI,
agent/ for ms-agent-dotnet). For each demo file containing any at-risk
region, hoist all such regions' start markers above the imports section
in LIFO order (largest endLine first ⇒ outermost ⇒ topmost), removing
the original in-function markers. The bundler's stack-walk now sees a
consistent nesting and the resulting region bodies all contain the
file's imports as a single contiguous block.
Also extends marker-move-up support to Java (import) and C#
(using-directive) files for Spring AI and ms-agent-dotnet's tool/agent
classes.
Manually handles two remaining sibling snippet files
(built-in-agent::a2ui-fixed-schema's a2ui-backend.snippet.ts) where the
'imports' are declare-const stubs that the codemod doesn't detect as
imports.
After this commit, of the 32 at-risk (cell, region) tuples flagged in
the QA report, 503 (integration × region) bundle slots have imports in
their bodies; 4 slots remain without imports because the source files
genuinely have no import statements (string-only prompt files in
claude-sdk-typescript subagents-prompts.ts).
Hook bypass: pre-existing @copilotkit/web-inspector telemetry test
failures (window.localStorage + jsdom) are unrelated to this commit.
Seven integrations (ag2, built-in-agent, claude-sdk-typescript, crewai-crews,
langgraph-fastapi, langgraph-typescript, strands) have frontend-tools/page.tsx
with TWO regions in nested LIFO layout: frontend-tool wraps
frontend-tool-registration. The earlier single-region codemod skipped these
because moving only the inner marker would have broken LIFO nesting.
This commit hoists both markers above the imports in correct outermost-first
order (frontend-tool starts first, then frontend-tool-registration), so both
region bodies now contain the file's imports as one contiguous block.
Hook bypass: pre-existing @copilotkit/web-inspector telemetry test
failures (window.localStorage + jsdom) are unrelated to this commit.
For demo files where multiple at-risk regions sit in the same source
(chat-slots/page.tsx, a2ui_fixed.py, tool-rendering/page.tsx,
hitl-in-chat/page.tsx, subagents.py, voice route.ts), hoist each
region's start marker above the imports section. Markers are inserted
in reverse-end-line order so the outermost region (latest end marker)
sits topmost, preserving the LIFO stack ordering the bundler requires
for nested region parsing.
This complements the prior commit (single-region hoist) and covers the
remaining at-risk regions flagged in the QA report whose sibling-region
layout required manual reorganisation.
Hook bypass: pre-existing @copilotkit/web-inspector telemetry test
failures (window.localStorage + jsdom) are unrelated to this commit.
Apply marker-move-up across 260 demo files in 17 integrations. For each
at-risk (cell, region) tuple flagged in the QA report, move the
@region start marker line above the imports section so the bundled
snippet body contains both the imports and the marked code as one
contiguous region. End markers stay where they are.
Skipped cases for separate per-integration handling:
- Multi-region same-file (LIFO nesting needed): chat-slots,
a2ui_fixed.py, tool-rendering/page.tsx, hitl-in-chat/page.tsx,
subagents.py, voice route.ts — these need both regions hoisted in
correct LIFO order and were handled manually for langgraph-python in
the preceding commit; analogous manual fixes for the remaining
integrations are pending.
- Files where the target region is already wrapped by an outer region
(e.g. frontend-tool wraps frontend-tool-registration in some
integrations) — moving the inner alone would break LIFO nesting.
Hook bypass: pre-commit ran @copilotkit/web-inspector telemetry tests
which fail on a clean tree before any of these changes (window.localStorage
not initialised under jsdom in some test cases). Pre-existing failure
unrelated to this commit.
All 18 integration health endpoints previously proxied to the backend
agent /health with a 3s timeout, causing false reds when agents were
slow but functional. The harness already checks agent reachability
via the agent:<slug> probe. Health endpoints now return a simple 200
confirming the Next.js process is alive.
The langgraph-python voice cell sat at D4 even when its d5-voice probe
row was green. Root cause: the dashboard's CATALOG_TO_D5_KEY mirror in
showcase/shell-dashboard/src/lib/live-status.ts was missing voice ->
["voice"], so computeMaxPossible capped voice at D4 regardless of probe
state. The harness REGISTRY_TO_D5 already had the entry; only the
dashboard mirror was out of sync.
Separately, the "Play sample" button used to fetch sample.wav and POST
it to /transcribe. With aimock that meant both the sample button AND
the mic returned the same canned response, which made it impossible to
demo the mic path locally without conflating the two affordances.
Reworked the button into a synchronous static-text injector
(onTranscribed(sampleText)) so:
- Sample button = deterministic test/demo affordance, no runtime calls.
- Mic = real Whisper transcription via /transcribe.
Synced across all 18 voice-enabled integrations. Phrase stays "What is
the weather in Tokyo?" so aimock's "weather in Tokyo" substring fixture
still matches.
Also adds the missing d5-voice.test.ts companion (every other d5-* probe
script has one) and trims the langgraph-python qa/voice.md + e2e steps
that depended on the now-removed async behavior.
The validate-fixture-tool-surface check on PR #4669 flagged 18 drift
violations: every headless-simple demo carried 'Weather in Tokyo' /
'AAPL stock price' / 'Highlight a note' / 'Sketch a diagram' chips
that substring-match aimock fixtures returning tool calls
(get_weather / get_stock_price / highlight_note / etc.) — but
headless-simple demos only register 'show_card' via useComponent.
Tool-call dispatch had no matching renderer.
Trim the headless-simple chip list to two in-surface entries:
- 'Profile card' → 'Show me a profile card for Ada Lovelace' (existing
show_card fixture; show_card is already registered by useComponent).
- 'Largest continent' → 'What is the largest continent?' (text-only
fixture from Phase 0; no tool dependency).
The chip-click e2e test only asserts on the 'Largest continent' chip,
so the trim is test-compatible.
Headless-complete keeps the canonical 5-chip list (its tool surface
covers weather/stock/highlight/excalidraw via tool-renderers.tsx and
backend agents).
For google-adk/headless-complete: add a useDefaultRenderTool() wildcard
catch-all. The validator looks at page.tsx + hooks/* and a backend
agent file; google-adk's tool registrations live in tool-renderers.tsx
(unparsed) and there's no matching agents/headless_complete.py file,
so the validator saw an empty tool surface. The wildcard registers '*'
which matches every fixture tool — same pattern north-star already
uses in its own tool-renderers.tsx.
The agentic-chat-reasoning and reasoning-default-render cells in
langgraph-python and langgraph-fastapi were configured with
gpt-4o-mini + use_responses_api=False, which never produces AG-UI
REASONING_MESSAGE_* events: gpt-4o-mini is not a reasoning model and
the Chat Completions API does not surface reasoning summary items at
all. The frontend's reasoningMessage slot was rendering nothing,
even though the cells were billed as "reasoning" demos.
- Switch both reasoning agents to gpt-5-mini (override via
OPENAI_REASONING_MODEL) routed through the Responses API with
reasoning={"effort":"medium","summary":"detailed"} so the model's
chain of thought streams as content blocks that @ag-ui/langgraph
translates into REASONING_MESSAGE_* events.
- Update the aimock d5-all.json and harness reasoning-display.json
fixtures to include a "reasoning" field so aimock emits
response.reasoning_summary_text.delta SSE events deterministically
in CI without hitting a real LLM.
- Add a "Show reasoning" useConfigureSuggestions pill on both
reasoning demo pages so the user can trigger the fixture-matched
prompt with one click.
- Tighten the d5-reasoning-display probe: it now also asserts a
reasoning-role message rendered via [data-testid="reasoning-block"]
or [data-message-role="reasoning"], so a plain text response
containing the word "reasoning" no longer falsely passes.
- Un-skip the three streaming reasoning-block tests in
langgraph-python's agentic-chat-reasoning.spec.ts and add a
suggestion-pill test; expand the reasoning-default-render spec to
cover the default reasoning slot.
- Update the langgraph-python QA doc to describe the new model +
Responses API setup and the suggestion-pill flow.
- agno: add default agent alias + per-request runtime
- langgraph-fastapi: add default agent alias
- llamaindex: fix agent name mismatch (byoc_hashbrown → byoc-hashbrown-demo)
- mastra: create dedicated byocHashbrownAgent with hashbrown system prompt
(was using weatherAgent which produced plain text instead of JSON)
- ms-agent-dotnet: upgrade byoc page to V2 CopilotKit import
The D5 conversation runner detects assistant responses via
data-testid="copilot-assistant-message". The byoc-hashbrown demo
overrides the assistantMessage slot with a custom HashBrown renderer,
which dropped that attribute. Without it the harness sees 0 messages
and times out.