## Summary
- Spring-ai StreamingToolAgent now classifies tool calls as frontend vs
backend, emitting TOOL_CALL events without TOOL_CALL_RESULT for frontend
tools so the CopilotKit runtime's processAgentResult handles them
- On HITL re-invocation (tool result sent back), sends full AG-UI
message history to aimock/LLM via convertMessages() for correct fixture
matching
- Fixes 3 HITL features: hitl-text-input, hitl-approve-deny, hitl-steps
(spring-ai now 6/11 D5 green, up from 3/11)
## Remaining 5 failures (pre-existing, not caused by this PR)
| Feature | Root cause |
|---------|-----------|
| tool-rendering | CopilotKit frontend re-render drops assistant
messages (count goes 1→0) |
| shared-state-read | SharedStateReadWriteController — separate
controller, not StreamingToolAgent |
| shared-state-write | Same as shared-state-read |
| subagents | SubagentsController — aimock returns "No fixture matched"
(404) |
| mcp-apps | Depends on external MCP server (excalidraw.com) unreachable
from Docker |
All 5 fail identically on baseline (main) without this PR's changes.
## Test plan
- [x] `showcase test spring-ai --d5` — 6 green, 5 red (same 5 red on
baseline)
- [x] Direct curl to Java backend confirms correct AG-UI event sequences
for all HITL flows
- [x] No regressions: agentic-chat, gen-ui-custom, gen-ui-headless still
pass
Follow-up to the docs cross-link sweep — visual QA surfaced four more sets
of broken links that the framework-scope middleware was rewriting into 404s
(or that resolved against the wrong path due to relative-link ambiguity).
- agentic-protocols/index.mdx: switch ./ag-ui, ./mcp, and ./a2a to absolute
/agentic-protocols/* paths so they resolve regardless of trailing slash.
- faq.mdx: drop link wrapping on the three "Rich agentic experiences" row
labels (Deep support for LangChain, Human-in-the-loop, Shared state).
The /langgraph/* slugs were docs.copilotkit.ai legacy paths that don't
exist in shell-docs; matches the playbook used for the V1 reference rows.
- integrations/langgraph/agent-app-context.mdx: retarget the "Frontend Data
documentation" Callout link from /langgraph/agent-app-context to
/langgraph-python/agent-app-context (the working slug).
- inspector.mdx: retarget the useAgentContext link in the Context row from
/langgraph/agent-app-context to /langgraph-python/agent-app-context.
The Helm chart's `migrations.enabled` value defaults to `false`, but
the install walkthrough and overview prose both implied the
pre-install migrations Job always runs. Realign the docs with what the
chart actually does:
- intelligence-platform: rewrite the "Install" prose to describe the
Job as conditional on `migrations.enabled: true`.
- self-hosting: add a new step in the install walkthrough between
"Create secrets" and "Install the chart" that explains how to opt
into migrations and when to leave them disabled. Update the
"Verify the install" prose so it only claims a Completed Job if
the reader opted in, and adjust the `--timeout` rationale to match.
The model-name allowlist (`docs/model-allowlist.json`) ships
`gpt-5.4` and `gpt-5.4-mini` but never `gpt-5.2*` — the latter slipped
in during a model-bump cycle and was never caught because the CI
validator only scanned the legacy `docs/` tree.
- Sweep replace `gpt-5.2-mini` -> `gpt-5.4-mini` and `gpt-5.2` ->
`gpt-5.4` across `showcase/shell-docs/src/content/` (~17 files).
- Extend `scripts/validate-doc-model-names.ts` with an
`EXTRA_DOCS_DIRS` list so the validator now scans the shell-docs
content tree alongside the legacy Nextra tree under `docs/`,
preventing the same drift in future.
Sweep of dead links surfaced in the QA triage:
- agentic-protocols/index: retarget AG-UI / MCP / A2A links to the
actual sibling slugs (`./ag-ui`, `./mcp`, `./a2a`); update the
Generative UI table to point at the real `/generative-ui/*`
pages and drop the link to the open-json-ui spec page (now hidden).
- multimodal-attachments: drop the orphaned migration Callout — the
`/migration-guides/migrate-attachments` page does not exist.
- faq: remove `/reference/v1/*` links from the "What's available?"
table; engineering direction is V2-only outside the reference area.
- agent-app-context: retarget the cross-link in the LangGraph guide
and the Inspector reference to `/langgraph/agent-app-context` so
they resolve.
- troubleshooting/common-issues: fix two broken `../X` links to
point at the real `/backend/copilot-runtime` and
`/built-in-agent/model-selection` slugs.
Quickstart pages and a handful of high-traffic guides had their
`import { ... }` headers stripped, leaving bare identifiers above an
orphan `} from "..."` line — copy-pasting the snippets failed to
compile. Restore the missing import statements and add full imports
to layout.tsx / page.tsx blocks that were previously empty so each
block is independently copy-pasteable.
Also fixes the BIA-family quickstart V1 import path: `CopilotKit` is
exported only from the `/v2` entrypoint, so swap
`@copilotkit/react-core` for `@copilotkit/react-core/v2` in the BIA,
agent-spec, and Microsoft Agent Framework quickstarts.
Sidebar / IA / HITL cleanup from the shell-docs QA triage:
- Add `absent` mode to `<WhenFrameworkHas>` so MDX pages can declare a
fallback branch for frameworks where a flag is null/missing, instead
of collapsing to an empty middle.
- Use the new `absent` branch on `useInterrupt.mdx` and `headless.mdx`
to point readers without `interrupt_pattern` at `useHumanInTheLoop`.
- Wrap the `useHeadlessInterrupt`-using "Driving it from plain UI"
section in `headless.mdx` inside the native gate where the symbol is
actually defined.
- Add `multi-agent/meta.json` so breadcrumbs / section labelling for
`/multi-agent/subagents` use the explicit "Multi-Agent" title.
- Add a shared lead-in between `<InlineDemo>` and the gated branches in
`agent-config.mdx`.
- Move `ag-ui-middleware.mdx` into `agentic-protocols/`, register it in
the section's `meta.json`, link to the upstream AG-UI guide, and add
a 302 from the old `/ag-ui-middleware` path.
## Summary
Updates incorrect/outdated documentation about real APIs. Reference
pages aligned with current package source; observability page
mechanically translated V1->V2; contributor onboarding rewritten to
point at shell-docs/Fumadocs/Nx instead of the legacy Nextra tree.
**Items addressed (from [triage
plan](https://app.notion.com/p/3523aa3818528128bcb0ee9e137cfff0)):**
- 9.3 — `LangChainAdapter.mdx`: `gpt-5.4` -> `gpt-4o`, drop the
misleading "auto-generated" header comment
- 19.1 — V2 hook reference pages: add missing `threadId` on `useAgent`,
fix `throttleMs` default cascade, add `lastRunAt` on `useThreads` Thread
shape
- 5.1 — `observability-connectors.mdx`: mechanical V1->V2 translation
(`<CopilotKit>` -> `<CopilotKitProvider>`, V2 import path, updated
`onError` event shape, optional server-side `CopilotObservabilityConfig`
section)
- 20.1 — `docs-contributions.mdx`: rewrite for shell-docs / Fumadocs /
Nx with real dev commands and ports
## Visual inspection
1. Start the dev server:
```bash
nx run shell-docs:dev
```
Open http://localhost:3003.
2. **Reference pages — accuracy.** Visit each page and verify the
documented shape matches the source:
- `/reference/v1/classes/llm-adapters/LangChainAdapter` — `## Example`
should show `model: "gpt-4o"`, not `gpt-5.4`. Page header should no
longer claim it's auto-generated.
- `/reference/v2/hooks/useAgent` — Parameters section should now list
`threadId`. `throttleMs` description should mention the provider-cascade
default.
- `/reference/v2/hooks/useThreads` — `threads` Return Value's Thread
shape should now list `lastRunAt`.
3. **Observability page — V2 translation.** Visit
`/troubleshooting/observability-connectors`:
- Code blocks should use `<CopilotKitProvider>`, not `<CopilotKit>`.
- Imports should be from `@copilotkit/react-core/v2`.
- The `onError` event shape should be `{ error, code, context }`, not
the V1 `CopilotErrorEvent` fields.
- `publicApiKey` and `publicLicenseKey` props should still appear (they
carry over to V2).
- A server-side section was added referencing
`CopilotObservabilityConfig` from the runtime.
4. **Contributor docs — read-through.** Visit the rendered
"Documentation Contributions" page (under the `(other)/contributing`
group). Walk through as if you're a new contributor:
- Clone instructions should point at `CopilotKit` root with Nx commands
run from there.
- Dev port should be 3003.
- Should mention Fumadocs (not Nextra).
- Should mention `<Snippet>`-region authoring at a high level.
- Should mention the pre-commit hook expectation.
5. **Anti-checks (should NOT have changed):**
- The actual JSDoc source in `packages/` is unchanged.
- Other reference pages (e.g. `/reference/v2/hooks/useCapabilities`) are
unchanged.
- The legacy `docs/` tree (the Nextra one) is unchanged.
## Summary
Hide and clean up stale shell-docs pages. Drops orphaned/broken/AI-slop
pages from nav, adds 302 redirects, deletes dead per-framework
overrides.
**Items addressed (from [triage
plan](https://app.notion.com/p/3523aa3818528128bcb0ee9e137cfff0)):**
- 1.1 — Tutorials (broken end-to-end; rewrite post-launch)
- 6.1 — `coding-agent-setup.mdx` (rename straggler)
- 10.1 — `copilot-suggestions.mdx` (orphaned broken stub)
- 11.1 — `generative-ui/open-json-ui.mdx` (AI-slop placeholder; rewrite
post-launch)
- 21.1 — `migrate/1.10.X.mdx` (~1-year-old migration target, no longer
supported)
- 3.4 — Legacy per-framework HITL overrides — see Concerns
- 16.1 — 3 orphan `state-inputs-outputs` / `workflow-execution` files in
adk/langgraph/llamaindex
All redirects are 302 (not 301) — we'll restore at the same URLs when
the affected pages are properly authored post-launch.
## Concerns
**3.4 (HITL overrides) intentionally NOT executed in this PR.** Audit
found that every per-framework HITL override carries substantive
framework-specific content beyond the canonical (e.g. `IframeSwitcher`
blocks pointing at fw-specific feature-viewer URLs, `CTACards` linking
to fw-specific sub-pages like
`/mastra/human-in-the-loop/interrupt-flow`,
`/crewai-flows/human-in-the-loop/flow`, etc.). Per the task's "be
conservative" guidance, none were deleted. If we want to land 3.4 the
right way, it likely needs a follow-up that either (a) merges the
fw-specific iframe demos into the canonical via per-fw branching, or (b)
explicitly designates these as legitimate per-fw overrides and accepts
them. Flagging for a separate decision.
## Visual inspection
1. Start the dev server:
```bash
nx run shell-docs:dev
```
Open http://localhost:3003.
2. **Sidebar disappearance.** Confirm the following are no longer in the
left sidebar:
- "Tutorials" section (entire dividing block + sub-pages)
- `Open-JSON-UI` under Generative UI > Declarative
- `Migrate to 1.10.X` under Migrate
- `Coding Agent Setup` and `Chat Suggestions` (these were already not in
nav — visual check just confirms nothing changed)
3. **Redirect check.** Visit these URLs directly; each should 302 to the
destination:
- `http://localhost:3003/tutorials/ai-todo-app/overview` → `/`
- `http://localhost:3003/tutorials/multi-conversation-chat` → `/`
- `http://localhost:3003/coding-agent-setup` → `/coding-agents`
- `http://localhost:3003/copilot-suggestions` → `/`
- `http://localhost:3003/generative-ui/open-json-ui` → `/generative-ui`
- `http://localhost:3003/migrate/1.10.X` → `/migrate`
4. **Shared-state per-fw check.** For adk, langgraph, llamaindex:
navigate to `/shared-state` (or the corresponding shared-state sidebar
entries). Confirm only the wired version (per `meta.json`) appears in
nav. The orphan URLs should 404 cleanly.
5. **Anti-checks (these should NOT have changed):**
- All other docs pages render normally.
- The canonical `human-in-the-loop` content (root) is unchanged.
- Existing redirect rules in `next.config.ts` still work.
StreamingToolAgent now classifies tool calls as frontend vs backend by
comparing input.tools() (CopilotKit-injected frontend tools) against
the registered toolCallbacks (backend tools). Frontend-only tool calls
emit TOOL_CALL_START/ARGS/END without TOOL_CALL_RESULT so the CopilotKit
runtime's processAgentResult detects the missing result and executes the
frontend handler (useHumanInTheLoop, useFrontendTool).
On re-invocation (when the runtime sends back the tool result), the
agent now sends the full AG-UI message history to aimock/LLM via
convertMessages() so the fixture matcher sees the tool result and
returns a follow-up text response instead of repeating the tool call.
Fixes 3 HITL D5 features: hitl-text-input, hitl-approve-deny,
hitl-steps. Spring-ai now passes 6/11 D5 features locally (up from 3).
The remaining 5 failures (tool-rendering, shared-state-read/write,
subagents, mcp-apps) are pre-existing and unrelated to StreamingToolAgent
— they involve separate controllers, missing aimock fixtures, external
MCP servers, and CopilotKit frontend rendering issues.
## Summary
Expands D5 (e2e-deep multi-turn) coverage for the LangGraph Python (LGP)
showcase integration from the existing 11 `D5FeatureType` literals to
**31 total** (20 new). 24 new D5 scripts cover features that previously
max'd out at D4 on the dashboard.
The plan and per-feature design lives in
[`.claude/specs/lgp-d5-coverage.md`](https://github.com/CopilotKit/CopilotKit/blob/blitz/lgp-d5-coverage-design/integration/.claude/specs/lgp-d5-coverage.md)
— start there for the bigger picture and per-feature turns sketches.
### What landed (8 commits)
- `docs:` — design plan with §1 summary, §2 new literal taxonomy, §3
per-feature plan table, §4 wave proposal, §5 open design questions, §6
anti-scope
- `S0` — scaffold 20 new `D5FeatureType` literals in
`harness/src/probes/helpers/d5-registry.ts`, mirror entries in
`d5-feature-mapping.ts` (`REGISTRY_TO_D5`) and
`shell-dashboard/src/lib/live-status.ts` (`CATALOG_TO_D5_KEY`). Adds
`beautiful-chat` and `shared-state-read` as additional registry-id
mappings to existing literals.
- `B1` chat-surface family — `chat-slots`, `chat-css`,
`prebuilt-sidebar`, `prebuilt-popup`
- `B2` platform family — `auth` (sign-out flow), `multimodal` (image+PDF
via sample buttons), `agent-config`
- `B3` frontend-tools + reasoning — `frontend-tools`,
`frontend-tools-async`, `reasoning-display` (covers
`agentic-chat-reasoning` + `reasoning-default-render`),
`tool-rendering-reasoning-chain`
- `B4` state — `shared-state-streaming`, `readonly-state-context`
- `B5` gen-UI — `gen-ui-declarative`, `gen-ui-a2ui-fixed`, `gen-ui-open`
(covers `open-gen-ui` + `open-gen-ui-advanced`), `gen-ui-agent`
- `B6` interrupt + BYOC — `interrupt-headless`, `gen-ui-interrupt`,
`byoc` (covers `byoc-hashbrown` + `byoc-json-render`)
- `F` — regenerated `aimock/d5-all.json` bundle (29 → 52 fixtures),
DOM-typing fix on chat-css probe
User-locked semantics encoded verbatim:
- **`auth`** — turn 1 chat works → click
`[data-testid=auth-sign-out-button]` → assert
`[data-testid=auth-demo-error]` or
`[data-testid=auth-demo-chat-boundary]` appears
- **`multimodal`** — click "Try with sample image/PDF" buttons → ask
"what is in this?" → assert assistant references attachment content
- **`chat-slots`** — send message → assert
`[data-testid=custom-assistant-message]` slot rendered
- **`prebuilt-sidebar`** / **`prebuilt-popup`** — surface root visible
(`.copilotKitSidebar` / `.copilotKitPopup`) → message → response inside
surface
- **`chat-customization-css`** — assert user-bubble bg contains `255, 0,
110` (hot pink) AND assistant-bubble bg `253, 224, 71` (amber)
`voice` is excluded (mic input not aimockable, deferred to a future
probe family).
## What's verified
- ✅ All 188 harness unit tests pass (30 test files including 24 new
ones)
- ✅ `pnpm typecheck` clean on `showcase/harness/`
- ✅ `aimock/d5-all.json` bundle regenerated cleanly with all 52 fixtures
- ✅ Pre-commit hooks pass (lint, format, package tests, commitlint)
- ✅ Each new D5 script self-registers via `registerD5Script()` and is
picked up by the e2e-deep driver's dynamic loader at boot
## What's NOT verified yet (verification gap — please read before
merging)
- ❌ **End-to-end run against the LGP integration was NOT executed.** A
full `showcase test langgraph-python --d5 --verbose --cycle` requires
the local docker compose to be brought up against this branch's
`aimock/d5-all.json` bundle, which would interrupt other running
services. We expect the first e2e pass to surface real issues — the most
likely failure modes per the design plan §5:
- `multimodal`: pre-fill hook for sample-button click is folded into the
assertion, which races the runner's automatic fill on turn 1. Likely
needs a runner-level `preFill` hook OR the demo to grow a query-string
attach trigger.
- `auth`: the post-sign-out 401 surfaces via `/info` refetch — if the
refetch isn't auto-triggered, the assertion will time out.
- `chat-css`: computed-style `background` shorthand resolution varies by
browser; the gradient parse may not flatten to the expected RGB
substring on all engines.
- `agent-config`: the `tone`/`expertise`/`responselength` keyword check
assumes the canned response surfaces them all — actual agent output may
differ.
- All transcript-keyword assertions are loose by design and will pass on
any response containing the keywords. They detect "no response" / "wrong
agent" but won't catch subtle behavioral regressions.
- ❌ **`cr-loop` was NOT run.** Sandbox blocks subagent file writes in
this environment, which would stall the cr-loop fix cycle. Recommend
running `cr-loop` in a fresh local session OR accepting reviewer
feedback on this PR directly.
- ❌ **8 open design questions** in `.claude/specs/lgp-d5-coverage.md` §5
— most are family-collapse vs split decisions (Q1
`beautiful-chat`/`agentic-chat`, Q3 `gen-ui-open` vs split, Q4 `byoc` vs
split, Q5 `reasoning-display` vs split). Defaults applied; flag in
review if any need reversing.
## Test plan (for reviewer / merger)
- [ ] Pull this branch locally
- [ ] `cd showcase && ./bin/showcase down && ./bin/showcase up`
(rebuilds aimock with the new `d5-all.json`)
- [ ] `./bin/showcase test langgraph-python --d5 --verbose --cycle` —
expect failures on first run; iterate on script/fixture/agent fixes
- [ ] Verify dashboard: features that were strikethrough `~~D5~~` should
now show D5 chips after the next 15-min probe tick on Railway
## Anti-scope
- D6 parity coverage is separate (post-D5).
- Other framework integrations (CrewAI, Mastra, LangGraph-TS,
Pydantic-AI) — not in this PR. The new `D5FeatureType` literals are
registry-wide; per-framework implementation is each framework's own
work.
- Voice probe family — needs a separate non-aimock probe shape,
deferred.
Follow-up to the matrix memoization PR. The initial dashboard load still
freezes for a few seconds because (a) the initial PB fetch is 10
sequential getList round-trips before any data arrives, and (b) when it
does arrive, the empty-map → populated-map transition forces every cell
in the matrix to re-render in one synchronous commit.
- Parallelize initial fetch: pull page 1 sequentially to learn
totalItems, then fire pages 2..N concurrently via Promise.all. Wall
time drops from `sum(rtt_per_page)` to roughly `max(rtt_per_page)`
modulo network parallelism. PocketBase reads are independent so this
is safe.
- Wrap initial setRows in startTransition. The first commit with real
data is unavoidably a full-matrix re-render (per-key memo checks
invalidate on every cell when the map flips from empty to populated);
marking it as a transition lets React 19 yield to user input mid-walk
instead of blocking the main thread for the duration. setStatus stays
urgent so the "connecting → live" indicator still flips immediately.
- Wrap the SSE flush setRows in startTransition for the same reason —
bursts that the 16ms coalescer can't fully absorb still need to yield.
## Summary
The shell-dashboard matrix tab can spike CPU and freeze the browser on
load and during update bursts. Two compounding causes:
1. **Every PB SSE delta forces a full grid re-render.** The matrix
renders ~720 cells (18 integrations × ~40 features), and nothing in the
cell tree was memoized — a single status row update propagated a new
`ctx.liveStatus` Map identity through every cell, and identity-based
React.memo couldn't help because the Map ref changed.
2. **PocketBase fires the subscribe callback once per record.** A probe
finishing dozens of services, or the initial-state replay on reconnect,
produced N consecutive `setRows` calls and N React commits across
separate microtasks (auto-batching only catches synchronous bursts).
## Changes
- **`ComposedCell`** wrapped in `React.memo` with a custom
`arePropsEqual`. Compares `overlays`/`catalogCell` refs, ctx scalars,
and — when the `liveStatus` Map identity changes — only the 5 row keys
this cell actually reads (`health:slug`, `e2e:slug/feature`,
`smoke:slug`, `d5:slug/feature`, `d6:slug/feature`). Because
`upsertByKey` preserves row identity for unchanged keys, deltas that
don't touch a cell's `slug/featureId` short-circuit at the memo
boundary.
- **`useLiveStatus`** buffers SSE callbacks into a per-key `Map<string,
PendingOp>` and flushes via a single 16ms `setTimeout`. A burst of N
deltas now produces 1 React commit instead of N. Last-write-wins per
key. Buffer is cleared on effect teardown and on reconnect kickoff so
post-reconnect initial fetches never land on top of stale buffered rows.
## Why these two together
Throttling alone would still re-render every cell on each flush.
Memoization alone would still get invalidated on every delta because
each delta produces a new Map. Together: a burst of unrelated row
updates → 1 React commit, and within that commit only cells whose
specific rows changed actually re-render.
## What I have not changed
- The three independent `useNowTick` 1s timers on the Ops tab. They only
matter if the freeze reproduces on Ops, not Matrix; happy to consolidate
as a follow-up if needed.
- `OverlayColumnHeader` (rendered 18× in the header, calls `LevelStrip`
which does 4 Map lookups each). Same memo treatment would apply if these
still show up after this lands.
- No virtualization. Should not be necessary if the memo + throttle
combo works as expected.
## Test plan
- [x] Existing unit tests in the touched files pass —
`useLiveStatus.test.tsx` 15/15, `composed-cell.test.tsx` 9/10 (the 1
failure is a pre-existing test-vs-impl drift in the "renders 3 layers
when all active — docs deduped by health" case, verified by stashing my
changes and re-running).
- [x] No new TypeScript errors in the touched files.
- [ ] **Manual perf trace before/after on the live dashboard** — Chrome
DevTools → Performance tab → record while loading the dashboard against
a real PB feed. Expected: scripting time during initial load and during
probe-finish bursts drops sharply.
- [ ] **Functional sanity** — confirm cells still update on real status
changes (e.g. trigger a probe and watch the matching cell flip tone
within ~1 frame).
The matrix tab can freeze the browser on load and during update bursts
because every PB SSE delta triggers a full re-render of all ~720 cells
(18 integrations × ~40 features), and PocketBase fires the subscribe
callback once per record — a probe finishing dozens of services or an
initial-state replay produces N consecutive React commits.
- ComposedCell: wrap in React.memo with a custom equality check that
compares overlays/catalogCell refs, ctx scalars, and (when the
liveStatus Map identity changes) only the 5 row keys this cell
actually reads. Because upsertByKey preserves row identity for
unchanged keys, deltas that don't touch a cell's slug/featureId
short-circuit at the memo boundary.
- useLiveStatus: buffer SSE callbacks into a per-key Map and flush via
a single 16ms setTimeout. A burst of N deltas now produces 1 React
commit instead of N. Last-write-wins per key. Buffer is cleared on
teardown and on reconnect kickoff so post-reconnect initial fetches
never land on top of stale buffered rows.
## Summary
Bumps `@copilotkit/aimock` to 1.16.4 to pick up the router fix from
<https://github.com/CopilotKit/aimock/pull/148>. After this lands and
Railway rebuilds `ghcr.io/copilotkit/aimock:latest`, the showcase demos
pick up the fix automatically — no fixture changes required.
**The bug it fixes:** `match.toolCallId` was scanning the entire
conversation for the most recent `tool` message, so once any prior tool
result was in history, every subsequent request still had a "last tool
message" buried in the array. A stale `toolCallId` fixture could win and
shadow `userMessage` matchers for new user turns.
**User-visible symptom:** in `beautiful-chat`, clicking the first
suggestion (e.g. pie chart) worked, but clicking any subsequent
suggestion replayed the previous chart's "Pie chart rendered above —
Electronics is the largest slice…" content fixture instead of producing
a new tool call. Looked broken to anyone clicking through demos.
**Upstream fix:** the matcher now requires the tool message to be the
**last** message in the request — the only state in which the LLM is
being asked to respond to a tool result. Two regression tests in aimock
cover the "new user turn after tool" and "assistant content reply after
tool" cases.
## Changes
- `pnpm-lock.yaml` — refresh resolutions for `@copilotkit/runtime`'s
`@copilotkit/aimock` devDep (`1.16.2` → `1.16.4`). Mostly mechanical;
lockfile is slightly more compact than before.
- `showcase/scripts/package-lock.json` — refresh to `1.16.4` (was
`1.14.3`).
- `.github/workflows/test_e2e-showcase-on-demand.yml` — bump the
known-good floor from `@copilotkit/aimock@^1.14.3` to
`@copilotkit/aimock@^1.16.4` so the `/test-aimock <slug>` PR-comment
workflow always installs a build that contains the fix.
`packages/runtime/package.json` and `showcase/scripts/package.json` keep
their `"latest"` specifier per existing convention; the lockfile pins
are the reproducibility layer.
## Test plan
- [x] `pnpm nx run @copilotkit/runtime:test` — 1412/1412 pass (covers
`LLMock` / `MCPMock` imports from `@copilotkit/aimock` in the v2 MCP
integration tests)
- [x] `npm test -- aimock-fixtures` in `showcase/scripts` — 18/18 pass
(loadFixtureFile + validateFixtures schema validation)
- [x] Lefthook pre-commit (`check-binaries`, `sync-lockfile`,
`lint-fix`, `test-and-check-packages`) green
- [x] commitlint conventional-commit format green
- [ ] After merge: confirm Railway picks up the new
`ghcr.io/copilotkit/aimock:latest` image on next service restart and
beautiful-chat suggestions work end-to-end
The 71 unrelated test failures in `showcase/scripts` are pre-existing
Windows path-separator issues on `main` (audit CLI subprocess tests,
integration registry tests) — verified identical counts on origin/main
without this branch's changes. Out of scope for this bump.
## Related
- aimock PR: <https://github.com/CopilotKit/aimock/pull/148>
- aimock release:
<https://www.npmjs.com/package/@copilotkit/aimock/v/1.16.4>
Picks up the router fix from CopilotKit/aimock#148 — `toolCallId` matchers
now only fire when the tool message is the *last* message in the request,
preventing stale tool_call_ids from history shadowing `userMessage`
matchers on new user turns.
Surfaced as: in beautiful-chat, clicking a second suggestion replayed the
prior chart's "Pie chart rendered above…" content fixture instead of
producing a new tool call. Once Railway rebuilds `ghcr.io/copilotkit/aimock:latest`
and restarts the service, demos will pick up the fix automatically.
- Refresh `pnpm-lock.yaml` resolutions (workspace `@copilotkit/runtime` devDep)
- Refresh `showcase/scripts/package-lock.json` to 1.16.4
- Bump the floor in `test_e2e-showcase-on-demand.yml` from `^1.14.3` → `^1.16.4`
so the `/test-aimock` PR-comment workflow always installs a build that
contains the fix
Three fixes for the built-in-agent TanStack integration:
1. **Custom stream converter** — the runtime's `convertTanStackStream`
(PR #4476) blocks all events after the first `RUN_FINISHED`, which
breaks TanStack's multi-turn agent loop. Server tools like
`get_weather` and `set_notes` need the loop to execute the tool,
emit TOOL_CALL_RESULT, and re-prompt for the text response. Switch
to `type: "custom"` with a local converter that skips RUN_FINISHED
without blocking subsequent events, and deduplicates tool-call
events (TanStack's buildToolResultChunks re-emits START/ARGS/END
for server tool results).
2. **Frontend tool forwarding** — register AG-UI frontend tools
(useHumanInTheLoop, useRenderTool, useFrontendTool) as TanStack
definition-only declarations so the LLM can call them. Without
this, tools like `book_call` were unknown to TanStack and silently
dropped.
3. **HITL status case mismatch** — CopilotKit v2's ToolCallStatus uses
lowercase strings ("executing") but the TimePickerCard checked for
PascalCase ("Executing"). Normalize with toLowerCase() so buttons
enable correctly.
Also migrates hitl-in-chat demo from CopilotKitProvider (v1) to
CopilotKit (v2) with proper agentId wiring.
## Summary
- Register `request_user_approval` FunctionTool stub in the llamaindex
`hitl_in_app_agent.py` so that `AGUIChatWorkflow` emits the necessary
`TOOL_CALL_CHUNK` AG-UI event for CopilotKit to intercept the tool call
and open the approval dialog
- Same pattern as the existing `book_call` stub in
`hitl_in_chat_agent.py`
## Why
`AGUIChatWorkflow` only emits AG-UI tool-call events for tools in its
`frontend_tools` registry. Without the stub, the aimock's
`request_user_approval` tool call was silently dropped by the workflow,
so CopilotKit never opened the approval dialog and the
`hitl-approve-deny` D5 test timed out.
## Test plan
- [x] `showcase/bin/showcase build llamaindex` succeeds
- [x] `showcase/bin/showcase test llamaindex --d5 --verbose` — all 11
features green (was 10/11)
- [x] Verified the fix follows the same pattern as `book_call` in
`hitl_in_chat_agent.py`
AGUIChatWorkflow only emits TOOL_CALL_CHUNK events for tools registered
via the frontend_tools constructor argument. Without a backend stub for
request_user_approval the workflow silently dropped the tool call from
the aimock response, so CopilotKit never intercepted it and the approval
dialog never opened.
Add a FunctionTool stub (same pattern as book_call in hitl_in_chat_agent)
and pass it in frontend_tools so the AG-UI event chain fires correctly.
All 11 llamaindex D5 features now pass locally.
Adds an optional `preFill` callback to `ConversationTurn` that runs
before the runner fills the chat input and presses Enter for each turn.
Failure semantics mirror the post-settle `assertions` callback: a thrown
error records the turn as failed and stops the conversation.
Rewires `d5-multimodal.ts` to use `preFill` to click the sample image /
PDF buttons before each turn — fixing the false-green where the
multimodal probe's transcript-keyword assertion passed without the
attachment ever being sent.
Adds unit tests covering preFill ordering, failure mode, and a
no-preFill regression guard, plus tests verifying the multimodal
script wires `preFill` to the right sample-button selectors.
Three changes that together restore D3 (e2e-readiness per-cell) emission for
half the dashboard cells after the pool sat on dead chromium instances and
the deploy webhook started 400-ing:
1. BrowserPool now detects dead browsers proactively. Once a chromium
process died (OOM, crash, network blip), the dead Browser instance
stayed in `available[]` and every probe that drew the slot failed with
"browser.newContext: Target page, context or browser has been closed"
— for hours/days, until the harness restarted or 100 release-cycles
tripped contextCount-based recycle. Adds a `disconnected` listener per
slot (registered in init() and recycleSlot()), an isConnected() check
at the top of acquire() that skips zombies and recycles them, a
release-time isConnected() check that catches the disconnect-event-
pending race, and a per-slot recyclingSlots guard preventing double-
relaunch when both paths fire concurrently. New unit tests use
fake browsers via the existing test-injection point on the constructor.
2. createPooledE2eDemosLauncher now honours the driver's abort signal,
mirroring the createPooledE2eDeepLauncher fix from ed0933e5c. Without
this, an outer-timeout kept the pooled browser held until the orphaned
driver promise drained all remaining demos — pool starvation across
ticks. Tracks open contexts so abort closes them before releasing,
uses a forceReleased flag so the driver's normal `browser.close()` in
the finally block doesn't double-release, and wires the logger
through the orchestrator registration.
3. /webhooks/deploy now accepts buildRunId / buildRunUrl. PR #4471 split
build and deploy into separate workflows and started co-sending these
fields, but `deployPayloadSchema.strict()` rejected them as unknown
keys, returning 400. Every Showcase: Verify Deploy run has 400'd
since. Adds both as optional, mirrors runUrl's http(s)-only
refinement on buildRunUrl, and propagates them onto DeployResultEvent
so downstream consumers can link a red deploy back to the build that
produced its images.
Pre-push gates: oxfmt --check clean; tsc --noEmit clean; harness vitest
suite passes the 1367 platform-portable tests (the 19 Windows-specific
pre-existing failures — path separators in test fixtures + boot-wiring
test timeouts — are present on origin/main with the same shape, untouched
by this change); tsc -p tsconfig.build.json clean.
## Summary
Install pkg-pr-new build of @copilotkit/runtime (from PR #4482) in
built-in-agent, and align all demo page wiring with the LangGraph-Python
D5 fixtures.
**Runtime fix** — The `@copilotkit/runtime` dependency points to a
`pkg.pr.new` URL instead of the `next` npm tag. This MUST be reverted to
`"next"` once a proper @copilotkit/runtime release includes the TanStack
RUN_FINISHED fix from PR #4482.
**D5 wiring fixes** — Seven D5 features failed because built-in-agent
demo pages used different tool names and hook types than the LGP
reference agent:
- **tool-rendering**: switched from `useComponent("weather")` to
`useRenderTool("get_weather")` with correct `{parameters, result,
status}` render shape
- **gen-ui-tool-based**: renamed `useComponent("haiku")` to
`useComponent("generate_haiku")`, rewrote HaikuCard to accept args as
direct props
- **hitl-in-chat**: created `/demos/hitl-in-chat` route (was missing);
reuses TimePickerCard from hitl-in-chat-booking with `book_call` tool
- **shared-state**: added `set_notes` server tool so the fixture's
`set_notes` call has a backend handler
- **subagents**: renamed `delegate_to_planner`/`delegate_to_researcher`
to `research_agent`/`writing_agent`/`critique_agent` matching LGP
fixture tool names
- **subagents page**: updated `useComponent` registrations and
DelegationCard for the renamed tools
## Results
- Before: 3/10 D5 features pass
- After: targeting 10/10 (pending D5 probe verification)
## Test plan
- [x] Docker image builds
- [x] TypeScript check — no new errors (only pre-existing Zod version
mismatch)
- [x] All 6 files modified/created verified
- [ ] D5 probe run against deployed built-in-agent
- [ ] Team review of runtime change scope
- [ ] Revert runtime dep to `"next"` after proper release
## Summary
Fixes three D5 (e2e-deep) probe failures for the llamaindex integration:
- **tool-rendering**: `get_weather` was registered as a backend tool, so
the LlamaIndex AG-UI workflow emitted `TOOL_CALL_CHUNK` but never
`TOOL_CALL_RESULT`. CopilotKit's `useRenderTool` stayed stuck in loading
state and the WeatherCard never rendered with
`data-testid="weather-card"`. Fix: move `get_weather` to
`frontend_tools`, add `ToolCallResultWorkflowEvent` (subclasses
`ToolCallEndWorkflowEvent` to pass the AG_UI_EVENTS isinstance filter),
override `aggregate_tool_calls` to emit it for render-only tools, and
make WeatherCard always render its testid wrapper even during loading.
- **gen-ui-headless**: the shared agent was missing a `show_card`
frontend tool stub, so the workflow never emitted `TOOL_CALL_CHUNK` for
it. Fix: add `show_card` stub and register in `frontend_tools`.
- **hitl-steps**: `human_in_the_loop` was routed to the hitl-in-chat
specialized agent (which only has `book_call`), not the shared agent
(which has `generate_task_steps`). Fix: move `human_in_the_loop` to
`sharedAgentNames`. Introduce `render_only_tool_names` set to
distinguish render-only tools from interactive tools -- only render-only
tools get premature `TOOL_CALL_RESULT`; interactive tools let CopilotKit
manage the result lifecycle.
Also fixes two harness probe bugs:
- esbuild's `keepNames` transform injected `__name()` wrappers in
`page.evaluate()` browser context, causing `ReferenceError: __name is
not defined`. Replaced with string-based `new Function()` construction.
- `querySelector` only returned the first match; on headless chat pages
the first assistant message is an empty wrapper. Switched to
`querySelectorAll` with iteration.
## Test plan
- [x] `showcase/bin/showcase test llamaindex --d5 --verbose` passes
10/11 (hitl-approve-deny is pre-existing)
- [x] Two consecutive stable runs confirm no flapping
- [x] LangGraph-Python regression check confirms no breakage from
harness changes
## Summary
- **StreamingToolAgent Phase 1 (streaming)**: Set
`internalToolExecutionEnabled=false` via `OpenAiChatOptions` so Spring
AI detects `tool_calls` in the stream without attempting execution
through the global `ToolCallingManager`. Previously,
`OpenAiChatModel.internalStream()` would auto-execute tool calls, and
when the model called frontend tools (like `generate_task_steps`,
`show_card`) that aren't registered on the Java backend,
`StaticToolCallbackResolver` returned null causing
`DefaultToolCallingManager` to throw `IllegalStateException`.
- **BoundedToolCallingManagerConfig**: Wrap `StaticToolCallbackResolver`
with `LenientToolCallbackResolver` that returns a
`FrontendToolPlaceholder` for unknown tools instead of letting the
resolver return null. This prevents crashes during Phase 2 (`.call()`)
when the model calls frontend tools injected by the CopilotKit runtime.
- **AG-UI event ordering**: Reorder event emission so tool call events
are emitted BEFORE `textMessageEnd`, which is required for the
frontend's `useRenderTool` to see them while the message is still open.
## Why
PR #4475 introduced `StreamingToolAgent` which uses Spring AI's
`.stream()` for real-time text delivery. However, Spring AI's
`OpenAiChatModel.internalStream()` auto-executes tool calls through the
global `ToolCallingManager`, which only knows about backend tools.
Frontend tools injected by the CopilotKit runtime caused
`IllegalStateException` crashes, breaking all D5 tests (0/11).
This fix recovers agentic-chat, gen-ui-custom, and gen-ui-headless D5
tests (3/11). The remaining 8 failures are pre-existing issues in other
controllers (subagents, shared-state, hitl) unrelated to the streaming
regression.
## Test plan
- [ ] `showcase test spring-ai --d5 --verbose` passes agentic-chat,
gen-ui-custom, gen-ui-headless
- [ ] Docker build succeeds locally
- [ ] No `IllegalStateException` in Spring AI container logs during D5
runs
- [ ] Frontend tool calls (generate_task_steps, show_card) handled
gracefully via placeholder
## Summary
- Remove trailing slash from the `hitl-in-app` agent URL in
ms-agent-python's CopilotKit route handler
## Why
The `hitl-in-app` agent was the **only** agent registered with a
trailing slash in the URL (`/hitl-in-app/`). The FastAPI backend mounts
the endpoint at `/hitl-in-app` (no slash). FastAPI's default
`redirect_slashes=True` returns a **307 redirect** for POST requests to
the trailing-slash variant, and the AG-UI `HttpAgent` does not follow
POST redirects during streaming. This caused the agent to appear
completely unresponsive — the D5 `hitl-approve-deny` probe timed out at
60s with `baseline=0, current=0` (zero assistant messages).
Verified via container: `POST /hitl-in-app/` returns 307 → `POST
/hitl-in-app` returns 422 (correct routing, body validation).
The fix uses the shared `createAgent("/hitl-in-app")` helper (which does
not append a trailing slash) for consistency with every other agent
registration in the file.
## Test plan
- [ ] D5 `hitl-approve-deny` passes for ms-agent-python (`showcase test
ms-agent-python --d5`)
- [ ] No regression in other ms-agent-python D5 features (10/11 → 11/11)
Drop orphaned/broken/AI-slop pages from nav, add 302 redirects, and
delete dead per-framework override stragglers. Items addressed:
- 1.1: Tutorials section hidden (broken end-to-end; rewrite post-launch)
- 6.1: coding-agent-setup.mdx (rename straggler -> /coding-agents)
- 10.1: copilot-suggestions.mdx (orphaned broken stub)
- 11.1: generative-ui/open-json-ui.mdx (AI-slop placeholder)
- 21.1: migrate/1.10.X.mdx (~1yr-old migration target)
- 16.1: 3 orphan shared-state files in adk/langgraph/llamaindex
(each meta.json wires only one of state-inputs-outputs vs
workflow-execution; the other was a dead duplicate)
All redirects use permanent: false (302) so URLs can be restored at
the same paths once the affected pages are properly authored.
agent-config: drop the AgentConfigLangGraphAgent subclass and use plain
LangGraphAgent. The subclass repacked CopilotKit provider properties
into forwardedProps.config.configurable.properties so the Python graph
could read them via RunnableConfig.configurable.properties — but
@ag-ui/langgraph@0.0.31 builds the LangGraph SDK request as
{ ..., config, context: { ...input.context, ...config.configurable } }
which merges configurable INTO context. LangGraph 0.6.0+ then rejects
with HTTP 400 'Cannot specify both configurable and context' on every
chat round-trip. Net effect: chat sent the user message, runtime 400'd,
no assistant response ever rendered. Removing the subclass unbreaks
the round-trip; the Python agent falls back to its DEFAULT_* constants
so the demo's frontend toggles no longer steer the system prompt
(known regression, tracked separately pending @ag-ui/langgraph fix
that decouples context from configurable).
byoc:
- D5 probe now sends the 'Sales dashboard' pill prompt (matches the
fixtures added in main:f0a89b843 in feature-parity.json) instead of
the previous generic 'render a byoc hashbrown' prompt that had no
matching JSON-shaped fixture. Removed the now-obsolete byoc.json D5
fixture file and regenerated the d5-all.json bundle (52 -> 50
fixtures).
- Added data-testid='copilot-assistant-message' + data-message-role=
'assistant' to the byoc-hashbrown and byoc-json-render renderer
wrapper divs. The CopilotChat default assistantMessage slot includes
these markers; overriding the slot with a custom JSON-rendering
component dropped them, so the e2e-deep conversation runner's
settle-detection cascade (which counts these selectors) never saw
the response and timed out at 30s. Re-attaching the markers is a
purely additive change that doesn't affect the renderers'
behavior.
- D5 byoc assertion now waits for [data-testid='metric-card'] AND a
chart (bar-chart or pie-chart) to render — a structural check on
the BYOC contract output, not a transcript-keyword check that the
custom renderer would never produce.
E2E status: 31/31 passing locally against
./bin/showcase up langgraph-python aimock with this branch's bundle.
Brings local D5 pass rate for langgraph-python from 24/31 to 29/31.
- preNavigateRoute added to chat-css, gen-ui-declarative,
gen-ui-a2ui-fixed, readonly-state-context — featureType literals
differ from registry IDs so default /demos/<featureType> 404'd.
- auth: <cpk-web-inspector> intercepts the sign-out click — added
force: true. The demo doesn't auto-refetch /info on header change,
so the post-sign-out 401 surface only appears after a probe send
— assertion now triggers one before polling.
- chat-css: inline-only page.evaluate body to avoid esbuild's __name
helper emit (undefined in the browser).
Remaining 2/31 known issues (documented in PR body for follow-up):
- byoc: agent runs but demo's hashbrown renderer expects streaming
JSON; aimock plain-text canned response leaves nothing to show.
- agent-config: runtime route never forwards to LangGraph (no
graph_id=agent_config_agent runs in the backend logs).
## Summary
- All six BYOC suggestion pills (3 per demo on `byoc-json-render` and
`byoc-hashbrown`) were getting eaten by older generic substring fixtures
in `showcase/aimock/feature-parity.json`. The `dashboard` rule swallowed
the Sales pills with placeholder text, and the `revenue by category as a
pie chart` / `monthly expenses as a bar chart` rules redirected the
json-render Revenue/Expense pills to a `render_pie_chart` /
`render_bar_chart` tool call from a different demo — leaving the BYOC
renderers staring at either canned text or empty content.
- Adds one fixture per pill, keyed on the **full pill prompt sentence**,
placed before the conflicting generic rules so first-fixture-wins picks
the right one. Each fixture returns the exact content shape the demo's
renderer expects: `{root, elements}` flat spec for json-render and
`{ui:[{<tag>:{props:{...}}}]}` for hashbrown.
- Removes two pre-existing duplicate fixtures from the bottom of the
file that had the right content but were placed after the conflicting
rules and never matched.
## Pill → fixture map
| Demo / pill | Match (full prompt) | Response shape |
|---|---|---|
| json-render — Sales dashboard | `Show me the sales dashboard with
metrics and a revenue chart` | flat spec, `MetricCard` with nested
`BarChart` |
| json-render — Revenue by category | `Break down revenue by category as
a pie chart` | flat spec, `PieChart` (4 segments) |
| json-render — Expense trend | `Show me monthly expenses as a bar
chart` | flat spec, `BarChart` (3 months) |
| hashbrown — Sales dashboard | `Show me a Q4 sales dashboard. Include a
total-revenue metric card, …` | ui kit: Markdown + metric + pieChart +
barChart |
| hashbrown — Revenue by category | `Break down Q4 revenue by product
category as a pie chart. …` | ui kit: pieChart, 4 segments |
| hashbrown — Expense trend | `Show me monthly operating expenses for
the last six months as a bar chart …` | ui kit: barChart, 6 months |
## Test plan
- [ ]
`https://showcase.copilotkit.ai/integrations/langgraph-python/byoc-json-render/preview`
— click each of `Sales dashboard`, `Revenue by category`, `Expense
trend`. Assert a chart/metric renders inside the assistant bubble (not
empty, not "Working on your request…").
- [ ]
`https://showcase.copilotkit.ai/integrations/langgraph-python/byoc-hashbrown/preview`
— same three pills. Sales should now render the multi-component
dashboard (Markdown + metric + pie + bar). Revenue/Expense should render
their single chart.
- [ ] Console: no `No renderer for component type:` warnings from
`@json-render/react`.
- [ ] `node -e "require('./showcase/aimock/feature-parity.json')"`
parses; no duplicate `userMessage` warnings from the aimock fixture
loader.
## Summary
Two related Beautiful Chat bugs in the integration showcases:
### Bug 1 — Infinite tool-call loop on every suggestion click
In `showcase/aimock/feature-parity.json`, 10 tool-calling fixtures
(`pieChart`, `barChart`, `render_pie_chart`×3, `render_bar_chart`×2,
`scheduleTime`, `search_flights`, `toggleTheme`) lacked an explicit `id`
on the `toolCall` and lacked a paired `toolCallId` followup fixture.
aimock's matcher (`router.js:36-46`) matches `userMessage` by
**substring** against the **last** user message. After the agent ran the
returned toolCall and re-prompted the LLM with the tool result attached,
the last user message was unchanged → the same fixture matched → the
same toolCall was returned forever. Repro: click any of the Pie Chart /
Bar Chart / Schedule Meeting / Search Flights / Toggle Theme suggestions
on `langgraph-python`'s `beautiful-chat` demo and watch the chart
re-render until manually stopped.
Fix follows the existing convention in the same file (compare the `Ada
Lovelace` / `show_card` pair):
- Added explicit `id` to each broken `toolCall`
(`call_fp_pie_chart_001`, `call_fp_bar_chart_001`, …)
- Added 11 new `toolCallId` followup fixtures at the top of the file
that match each id and respond with a brief `content` summary instead of
another toolCall
This shared file is loaded by every integration's `beautiful-chat` demo,
so the fix benefits all of them.
### Bug 2 — Black/white background split + broken CopilotKit logo
8 integrations use the full `ExampleLayout` pattern with `ThemeProvider`
and an `<img src="/copilotkit-logo.svg" />` (crewai-crews,
langgraph-fastapi, langgraph-python, langgraph-typescript, mastra,
ms-agent-dotnet, ms-agent-python, pydantic-ai). Two issues:
1. `globals.css` hardcoded `body { background: #fafaf9 }` and never
defined the brand tokens (`--background`, `--foreground`, `--card`,
`--primary`, `--secondary`, `--muted`, `--accent`, `--destructive`,
`--border`, `--input`, `--ring`, `--radius`) that the layout,
mode-toggle, todo card/column, and chart components reference via
Tailwind 4 arbitrary values. `ThemeProvider` was also adding `dark` to
`<html>` from system preference, so CopilotKit's chat went dark via
`[data-copilotkit] { background: var(--background) }` while body stayed
cream — that's the visible split the user reported.
2. `example-layout/index.tsx` renders `<img src="/copilotkit-logo.svg"
/>` but the file did not exist in any integration's `public/`.
Fix per integration:
- Added the full token set (light + dark) under `:root` and `:root.dark,
.dark`, registered the Tailwind 4 dark variant with `@custom-variant
dark (&:where(.dark, .dark *));`, switched `body` to
`var(--background)`/`var(--foreground)` so it tracks the same `dark`
class CopilotKit's chat is already responding to.
- Copied `copilotkit-logo.svg` and `copilotkit-logo-mark.svg` into each
integration's `public/` from
`examples/integrations/langgraph-python/public/`.
The 10 simplified-layout integrations (ag2, agno, built-in-agent,
claude-sdk-*, google-adk, langroid, llamaindex, spring-ai, strands)
don't render the logo or canvas split, so they don't need either fix.
## Verification
- `feature-parity.json`, `d5-all.json`, `smoke.json` parse cleanly
- `pnpm tsx showcase/scripts/validate-fixture-tool-surface.ts` → `104
fixtures × 603 demos — no drift`
- Programmatic check: every `toolCalls[].id` in `feature-parity.json`
now has a paired `toolCallId` followup fixture (no orphans)
- All 8 updated `globals.css` files have balanced braces and the
`:root.dark` block
## Test plan
- [ ] Click each of the 9 suggestion pills in `langgraph-python`
`beautiful-chat`; confirm tool-using suggestions no longer loop and end
with a text summary
- [ ] Toggle system theme → confirm chat side and canvas/body side stay
in sync (both light or both dark) on `langgraph-python` `beautiful-chat`
- [ ] CopilotKit logo renders in the chat header (no broken-image
placeholder)
- [ ] Spot-check one Pattern A integration (`langgraph-fastapi` or
`pydantic-ai`) and one Pattern B (`mastra` or `crewai-crews`) for the
same behavior
## Summary
Aligns `langroid` and `claude-sdk-typescript` with the rest of the
showcase fleet for how they expose unsupported HITL/interrupt demos on
the dashboard matrix.
Today these two integrations list `gen-ui-interrupt` and
`interrupt-headless` in BOTH `not_supported_features` AND `demos[]`
(with stub pages that explain why the demo doesn't work). Every other
integration that marks these features unsupported (ag2, agno,
built-in-agent, claude-sdk-python, crewai-crews, google-adk, llamaindex,
mastra, pydantic-ai, spring-ai, strands) lists them only in
`not_supported_features` — the dashboard then renders a single no-entry
icon with a hover tooltip and no demo link.
This PR moves langroid and claude-sdk-typescript onto that same
convention.
## Changes
- `showcase/integrations/langroid/manifest.yaml` — drop
`gen-ui-interrupt` and `interrupt-headless` entries from `demos:` (kept
in `not_supported_features`)
- `showcase/integrations/claude-sdk-typescript/manifest.yaml` — same
- Delete the orphan stub `page.tsx` + `README.md` files under
`src/app/demos/gen-ui-interrupt/` and
`src/app/demos/interrupt-headless/` for both integrations (8 files
total)
## Test plan
- [x] `pnpm run validate-manifests` — all 18 integrations pass
- [x] `pnpm exec oxfmt --check <both manifests>` — clean
- [x] `pnpm exec nx run @copilotkit/react-core:test` — passes (one flaky
retry from the pre-commit hook resolved on second run, per Nx flaky-task
warning)
- [ ] Verify on the dashboard matrix that langroid and
claude-sdk-typescript render an icon-only cell with hover tooltip for
`In-Chat HITL (useInterrupt — low-level primitive)` and `Interrupt
(Headless)`, matching the other 12 integrations
Pre-existing drift on main: count held at 134, hash shifted (one
FAIL healed, another regressed). Captured the new sorted-FAIL hash
locally with the same algorithm CI uses (sort -u | shasum -a 256)
and updated showcase/scripts/fail-baseline.json to match.
Two bugs introduced in 534cd1efa (D5 integration fixes) when the
langroid adapter switched backend tool results from TEXT_MESSAGE_*
triples to TOOL_CALL_RESULT events:
1. Two leftover `logger.warning("DEBUG ...")` calls in agui_adapter.py
that fired on every chat turn — exactly the noise pattern the
`test_plain_text_turn_does_not_warn` test was written to prevent.
2. Three stale tests in test_agui_adapter.py (only run on 3.11+, so
the failure was 3.12-only):
- test_backend_tool_execution_happy_path: still expected the old
TEXT_MESSAGE_START/CONTENT/END inner triple instead of a single
TOOL_CALL_RESULT event after TOOL_CALL_END.
- test_backend_tool_exception_returns_sanitized_error: still read
the sanitized error JSON from TEXT_MESSAGE_CONTENT.delta — it
now rides on TOOL_CALL_RESULT.content.
- test_plain_text_turn_does_not_warn: caught the DEBUG leak above.
Tests are skipped on Python 3.10 (langroid imports `typing.Self`),
which is why this only shows up on the 3.12 matrix entry.
Both BYOC demo pills were getting eaten by older generic substring
fixtures ("dashboard", "revenue by category as a pie chart", "monthly
expenses as a bar chart") — the json-render pills came back as foreign
tool calls (render_pie_chart) and the hashbrown sales pill came back as
the "Working on your request..." placeholder, leaving both renderers
empty.
Add a dedicated fixture per pill keyed on the full pill prompt sentence,
placed before the conflicting generic rules so first-fixture-wins lands
on the right one. Each fixture returns the exact content shape the
demo's renderer expects:
byoc-json-render -> {root, elements} flat spec for <Renderer />
byoc-hashbrown -> {ui:[{<tag>:{props:{...}}}]} for useJsonParser
Removes two pre-existing duplicates that had been added at the bottom of
the file but never matched because the conflicting generic rules came
first.