Replaces the hand-authored surface payloads with real gpt-5.4 responses
captured via the langgraph dev threads API during OSS-136 prompt
iteration (catalogId injected, since the replay path resolves the catalog
from recorded tool args). Verified: fixture suite 738/738, LGP container
e2e 6/6, ADK container e2e 6/6.
Same three-call choreography per pill (outer generate_a2ui, inner design
toolcall discriminated by toolName, narration matched by toolCallId), now
keyed to the sales-analyst prompts and emitting Vantage Threads payloads:
composed hero dashboard, rep DataTable, at-risk StatusBadge cards, and
top-account InfoRows. Verified by the LGP Playwright run against the
aimock-backed container (6/6).
Hero pill asserts the composed dashboard conjunctively (card + metric +
pie + bar); pills 2-4 each require a testid the hero is steered not to
mount (data-table / status-badge / info-row), preserving the
newly-mounted anti-masking gate.
The demo now plays an embedded sales analyst for a fictional company:
suggestion pills are natural business questions (chart-type steering moved
from user prompts into the system prompt), the hero pill composes a full
dashboard (KPI metrics + pie + bar in one surface) modelled on
beautiful-chat's sales dashboard, and the catalog gains DataTable,
gap-aware Row/Column, Metric trendValue, and the beautiful-chat palette.
Dataset + composition rules ship as frontend agent context
(sales-context.ts) so they reach both the primary agent and the secondary
A2UI planner in LGP and ADK alike. E2E specs and QA docs updated to the
new pill set.
* duplicate ceiling 290→291: tool-rendering.json's tightened 'current
price of AAPL' matchers now share two match keys with the existing
tool-rendering-custom-catchall.json entries in the same BIA context,
runtime-disambiguated by feature route.
* shadow ceiling 134→132 (ratchet down): the bare 'AAPL' vs 'current
price of AAPL' shadow pair on the tool-rendering.json side is gone.
* PARITY_NOTES: replaces the 'headless-complete turns 3+4 server-tool
reprompt loop' known-issue section with a resolved-via-sequenceIndex
description; the architectural reprompt loop now converges via the
sequenceIndex-gated emitter + narration-fallback pattern in
gen-ui-headless-complete.json.
BIA registers headless-complete's tools (get_weather, get_stock_price,
get_revenue_chart, highlight_note) as server-executed via TanStack's
chat() engine. After the LLM returns a tool call, TanStack runs the
server tool and reprompts the LLM with the result. The userMessage-keyed
toolcall fixtures fired again on every reprompt because the original
user pill text stays in conversation history — and the toolCallId-keyed
narration fallback never matched because BIA's /v1/responses endpoint
rewrites the assistant tool_call_id to a runtime-generated fc-… value.
Net effect: turns 3+4 (highlight, revenue chart) ballooned to 70+
assistant messages within the 60s timeout window.
Restructured each pill as a (sequenceIndex:0 emitter, narration
fallback) pair. The emitter matches the FIRST request for the pill
prompt (counter starts at 0) and emits the tool call; subsequent BIA
reprompt iterations fall through the now-exhausted emitter to the
narration fallback (no tool call), so the loop converges. sequenceIndex
is chosen over hasToolResult:false because hasToolResult is computed
across the entire thread — any earlier pill's tool result would
permanently disable a hasToolResult:false emitter, breaking multi-turn
sessions where turn 2+ would have hasToolResult=true globally.
The legacy bare 'AAPL' userMessage matchers in tool-rendering.json were
substrings of gen-ui-headless-complete's 'price of AAPL right now'
pill prompt, so when the BIA TanStack server-tool reprompt loop hit
iteration 2 (toolCallId chain broken by /v1/responses id rewriting),
the request fell through to tool-rendering.json's loose 'AAPL'
matchers and leaked wrong-card content. Tightened all four bare 'AAPL'
match keys to 'current price of AAPL' — preserves matching for the
tool-rendering 'Stock price' pill ("What's the current price of
AAPL?") and the custom-catchall probe's second prompt ("What's the
current price of AAPL?") while no longer matching the headless
probe's "What's the price of AAPL right now?".
## Problem
Asking for a sales dashboard in the beautiful-chat demo (reported on
`dojo.showcase.copilotkit.ai/?integration=langgraph-python&demo=beautiful-chat`)
fails with:
> A2UI render error: Catalog not found:
https://a2ui.org/specification/v0_9/basic_catalog.json
## Root cause
**Introduced by #5245** ("chore(showcase): upgrade langgraph A2UI to
stable releases + rewrite dynamic-schema docs", merged 2026-06-05): it
flipped the beautiful-chat routes to `injectA2UITool: true` — switching
them onto the runtime-injected `render_a2ui` tool — without adding the
`defaultCatalogId` the middleware contract expects the host to provide.
/cc @ranst91
The runtime-injected `render_a2ui` tool guide explicitly tells the model
**not** to send a `catalogId` ("the catalog id is set by the host, not
by you" — `@ag-ui/a2ui-middleware` `tools.ts`), and integrations whose
backend owns `generate_a2ui` see real models omit or late-stream the
argument. When the streamed `createSurface` is emitted without a
configured `defaultCatalogId`, the middleware falls back to the spec
basic catalog URL — which no showcase page registers, so the renderer
throws "Catalog not found".
The middleware contract expects the host to pin the catalog via
`a2ui.defaultCatalogId`; none of the showcase routes did. The pinned
`@ag-ui/a2ui-middleware@0.0.6` (resolved via
`@copilotkit/runtime@1.59.4` in the integration lockfiles) already
prefers the configured value over streamed args, so this is config-only.
## Fix
Add `defaultCatalogId` matching the catalog each page registers:
- 17 × `copilotkit-beautiful-chat` routes →
`copilotkit://app-dashboard-catalog` (every beautiful-chat page
registers this id)
- 14 × `copilotkit-declarative-gen-ui` routes →
`declarative-gen-ui-catalog` (all declarative catalogs use this id)
Left untouched: the 3 declarative routes with no `a2ui` block (ag2,
claude-sdk-typescript, strands) — `isA2UIEnabled` is false there, the
middleware never attaches, so the fallback path can't fire.
`copilotkit-a2ui-fixed-schema` routes are also untouched (direct tools
carry their catalog in the result envelope).
## Verification
- RUNBOOK-canonical D6 isolate run (`bin/showcase test langgraph-python
--d6 --isolate ...`) on this branch: **all A2UI pills green**
(`gen-ui-declarative`, `gen-ui-a2ui-fixed`, all five beautiful-chat
pills) — these stream through the middleware with the new
`defaultCatalogId` active (config preference). Aggregate green; the only
non-green features are pre-existing and unrelated (2×
`crypto.randomUUID` insecure-context in headless demos, 1× agent-config
fixture-length assertion).
- All 31 patched files parse-checked clean; full `tsc --noEmit` on
langgraph-python shows zero errors in the patched routes.
- Confirmed published middleware 0.0.6 (what the integration lockfiles
resolve) contains the host-preference logic for `defaultCatalogId`.
- Caveat: the Sales Dashboard pill itself could not be manually verified
end-to-end — its aimock fixture chain is broken by a pre-existing
fixture-shadowing collision from the `1e66a5f8d` fixture reorg (a
`userMessage: "First"` substring catch-all shadows the pill prompt;
three overlapping fixture chains mismatch on `toolCallId`
mid-conversation → `404 no_fixture_match` → `RUN_ERROR`). Broken
identically on `main`; tracked as a follow-up issue.
## Follow-ups (not in this PR)
- The canonical `examples/integrations/langgraph-python` route has the
same gap (`a2ui: { injectA2UITool: false }`, no `defaultCatalogId`,
pages register `copilotkit://app-dashboard-catalog`) — needs the same
fix through the parity sync flow.
- Test gap that let this ship: the D5 declarative fixture is a text-only
stub and beautiful-chat's D5 mapping reuses the agentic-chat
conversation, so no fixture exercises a `generate_a2ui` call that omits
`catalogId` the way real models do.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Adds a **React Native** SDK to the reference docs at
`/reference/react-native`, covering the full `@copilotkit/react-native`
public API. Modeled on the existing React/Core/Bots reference sections.
**Nav:** register `react-native` in `reference-items.ts`, add the SDK
selector label + an overview card.
**Content (22 pages, written from package source):** index, 8
components, 13 hooks. `CopilotChat` and `CopilotModal` exist in both a
headless (root) and prebuilt-UI (`/components`) form; each page
documents both, split by import path.
**Verified:** `next build` passes, all pages render, `oxlint` + `tsc`
clean, and content appears in `llms.txt` / `llms-full.txt`.
Out of scope per the issue: guide content and new demos.
---
## Preview
**Overview + SDK nav** — `/reference/react-native` (left sidebar shows
the new “React Native” SDK selector and all 8 components + 13 hooks)

**Component page — `CopilotChat`** —
`/reference/react-native/components/CopilotChat` (headless + prebuilt-UI
split by import path)

**Hook page — `useAgent`** — `/reference/react-native/hooks/useAgent`
(re-export callout, react-native code examples)

The injected/streamed a2ui fixtures all included catalogId, so aimock
replay never exercised the basic-catalog fallback that broke production
(real models omit catalogId per the tool-usage guide). Strip catalogId
from the langgraph-python sales-dashboard secondary-call fixtures and
hard-assert "Catalog not found" is absent outside the charts-rendered
soft branch, so the spec fails without a route defaultCatalogId.
Also repoint the on-demand e2e workflow at the d4/d5-recorded/d6/shared
fixture dirs — it still referenced feature-parity.json, deleted in the
1e66a5f8d fixture reorg, so every /test-aimock run died at aimock start.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adds a "React Native" SDK to the reference section, documenting the full
public surface of @copilotkit/react-native at /reference/react-native.
Navigation wiring:
- Register `react-native` in REFERENCE_VERSIONS / VERSION_SUBDIRS
- Add the "React Native" label to the SDK version selector
- Add a React Native card to the reference overview page
Content (22 pages, sourced from package source for accuracy):
- index: install, polyfills, provider setup, and the headless vs prebuilt
two-tier model
- components: CopilotKitProvider, CopilotChat, CopilotModal, CopilotSidebar,
CopilotPopup, CopilotMarkdown, AssistantMessage, UserMessage — disambiguating
the headless (root) and prebuilt-UI (/components) variants of CopilotChat and
CopilotModal by import path
- hooks: RN-specific useAttachments and useRenderTool, plus the platform-
agnostic hooks re-exported from react-core/v2 adapted to RN imports and
primitives (useAgent, useCopilotKit, useFrontendTool, useAgentContext,
useThreads, useCapabilities, useComponent, useHumanInTheLoop, useInterrupt,
useSuggestions, useConfigureSuggestions)
Content is picked up automatically by llms.txt / llms-full.txt via the
reference content walker.
OSS-250
The injected render_a2ui tool guide instructs models to omit catalogId
("the catalog id is set by the host"), and backend-owned generate_a2ui
tools see real models omit or late-stream it. Without defaultCatalogId
the a2ui middleware falls back to the spec basic catalog, which no
showcase page registers — surfaces fail with "Catalog not found:
https://a2ui.org/specification/v0_9/basic_catalog.json" (reported on
beautiful-chat / langgraph-python).
Pin each route to the catalog its page registers: beautiful-chat ->
copilotkit://app-dashboard-catalog, declarative-gen-ui ->
declarative-gen-ui-catalog. Routes with no a2ui block never attach the
middleware and are left untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
## What does this PR do?
Fixes#5282.
Provider header updates already reach the core instance via
`setHeaders`, but React consumers of `useCopilotKit()` were only
re-rendered for runtime connection status changes. That left
`useThreads()` with stale context after the provider `headers` prop
changed, so subsequent `/threads` requests could miss headers such as
`X-CSRF`.
This subscribes the React context hook to `onHeadersChanged` and adds a
provider-level regression test that verifies `/threads` is refetched
with the updated header.
## Related PRs and Issues
- Fixes https://github.com/CopilotKit/CopilotKit/issues/5282
## Testing
- `pnpm install --frozen-lockfile`
- `pnpm --dir packages/react-core exec vitest run
src/v2/hooks/__tests__/use-threads-provider-headers.e2e.test.tsx`
(failed before the production change, passed after)
- `pnpm --dir packages/react-core exec vitest run
src/v2/hooks/__tests__/use-threads.test.tsx
src/v2/hooks/__tests__/use-threads-provider-headers.e2e.test.tsx`
- `pnpm exec oxlint packages/react-core/src/v2/context.ts
packages/react-core/src/v2/hooks/__tests__/use-threads-provider-headers.e2e.test.tsx`
- `pnpm exec oxfmt --check packages/react-core/src/v2/context.ts
packages/react-core/src/v2/hooks/__tests__/use-threads-provider-headers.e2e.test.tsx`
- `pnpm nx run @copilotkit/react-core:build`
- `pnpm nx run @copilotkit/react-core:test`
- `git commit -m "fix(react-core): refresh thread headers on provider
updates"` (pre-commit ran `pnpm run test && pnpm run check:packages`;
commit-msg ran `pnpm commitlint --edit`)
Also checked: `pnpm nx run @copilotkit/react-core:check-types` currently
fails before checking this package because dependency package type/build
errors are reported in `@copilotkit/shared` and
`@copilotkit/a2ui-renderer`; I did not change those packages.
Note: I used Codex while preparing this change, reviewed the final diff,
and ran the listed checks locally.
## Checklist
- [x] I have read the [Contribution
Guide](https://github.com/copilotkit/copilotkit/blob/master/CONTRIBUTING.md)
- [ ] If the PR changes or adds functionality, I have updated the
relevant documentation (not applicable: internal header propagation bug
fix)
- [x] "Allow edits by maintainers" is checked (lets us help iterate on
your PR directly - faster turnaround for everyone)
Addresses review feedback on #4215 / OSS-192.
- Thread an explicit `data-testid` prop through `AutoResizingTextarea` so it
reaches the rendered <textarea>. The component destructures a fixed prop set
with no `{...rest}` spread, so the id passed from Input.tsx was silently
dropped and never landed in the DOM.
- Align selector names with the V2 components in @copilotkit/react-core:
`copilot-chat-textarea` on the textarea and `copilot-send-button` on the send
control. The legacy `data-test-id` values are preserved for back-compat.
- Add a source-level test asserting the Input wiring and that Textarea forwards
the prop (the dropped-prop guard); the package's vitest runs in a node env
with no DOM harness, matching the existing testids.test.ts convention.
Stable selectors let tests locate the controls; the headless input-driving
issue in #4215 remains a separate follow-up and should stay open.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Ports the usePinToSend half of #5386 to the Vue package, which is a
documented 1:1 parity port and retained the shrink-only ResizeObserver.
When content below the anchored user message loses height (suggestions
swap, input resize), the spacer now grows back so total scrollable space
below the bubble stays constant and the message stays pinned (#5355).
Vue has no scroll-to-bottom button, so the listener half of #5386 does
not apply.
## Symptom
Dashboard banner false alarm: *"Worker family E2E demos has not
completed successfully since 2h 40m ago"* (and the same on D5/D6) firing
while workers were healthy and emitting results every cycle.
## Triage (already done)
The D5/D6 family-silence banner fires because:
- `lastSuccessAt` is computed via `maxFinishedAtIso(completed)` where
`completed` requires the §5.2.1 `deriveOutcome` to return `"completed"`
(all jobs `done`, zero cell failures) — see
`showcase/harness/src/fleet/control-plane/run-view.ts` (pre-fix
`findCompletedBatch` + the `lastSuccessAt` line in `projectFamily`).
- With chronic content-reds present, every batch lands with
`jobs.done=1, jobs.failed=17` (cells failing inside the probes). The
outcome derives `"failed"`, no batch satisfies `outcome ===
"completed"`, and `findCompletedBatch` returns null.
- The family-silence monitor falls back to oldest-batch-`enqueuedAt` and
fires after `now − oldest > 3 × period` (see
`family-silence-monitor.ts`).
- The worker is fine: per-job `commError: null`, jobs completing with
real cell results, just `<100%` pass-rate.
This is a DEFINITIONAL bug. The worker is healthy; the definition of
"success" was conflating worker-completion with cells-all-green.
## The fix
Redefine "success" as **terminal-state accounting**, not pass-rate.
A batch counts as a terminal completion when:
1. Every job is in a terminal status (`done` or `failed`), AND
2. No job carries a `result.commError` (i.e. no worker-crashed-mid-job,
lease-expired, or other comm-level outage signal).
Cells red without a `commError` IS a terminal completion — the worker
reached the pool, ran the probe, and returned a result. Chronic content
reds become "known bad" instead of "stalled."
The strict §5.2.1 `outcome=="completed"` all-green semantic is unchanged
— it still drives `lastRun.outcome` and §6.2 dashboard rendering. No
external caller of run-view consumed the old strict `lastSuccessAt` for
logic beyond the banner and monitor (grep verified — every reference
outside this file is either a test fixture or the silence-banner display
surface), so renaming was unnecessary; only the definition was
redirected.
The fallback path in `family-silence-monitor.ts`
(oldest-batch-enqueuedAt when `lastSuccessAt` is null) is intentionally
**untouched** — with the new semantics, `lastSuccessAt` advances every
sweep on a healthy family, so the fallback only fires on true
never-completed envs (the rule it was designed for).
## Affected files
- `showcase/harness/src/fleet/control-plane/run-view.ts` — new
`isTerminalCompletionBatch` + `findTerminalCompletionBatch` predicates;
`projectFamily` now sources `lastSuccessAt` from terminal completion;
doc comment on the type updated.
- `showcase/harness/src/fleet/control-plane/run-view.test.ts` — 3 new
tests; 1 existing walk-back fixture updated to add a `commError` so its
original intent (skip non-completion batches) survives the new
predicate.
-
`showcase/harness/src/fleet/control-plane/family-silence-monitor.test.ts`
— 2 new tests: chronic-reds stays silent; real outage still fires.
- `showcase/harness/src/http/fleet-runs.test.ts` — T5 multi-page
walk-back fixture updated to flag its 20 "outage" batches with a
`commError` (so the walk-back still has to extend to page 2);
`batchFixture` gained an optional `resultFor` parameter.
## Tests (RED → GREEN)
Confirmed the 2 net-new predicate tests fail under the old
implementation and pass under the new:
- `lastSuccessAt advances when all jobs reach a terminal state with no
commError, even if cells failed` — old: `null`, new: newest finishedAt
- `lastSuccessAt does NOT count a batch where any job carries a
commError (worker-outage signal)` — old: `null` (walk-back to prior
worked but for the wrong reason — strict all-green semantics), new:
walks back past the comm-erroring batch to the terminal-completion batch
Negative cases preserved:
- `lastSuccessAt is null when every batch in the capped window has a
commError` — real outage stays loud
- `lastSuccessAt does NOT count a batch with non-terminal jobs` —
stall/abandon still excluded
- Family-silence monitor still alerts on real outages
Final test counts: harness `2700/2700` green, fleet/control-plane +
fleet-runs slice `355/355` green; format / lint / typecheck / build all
clean.
## Follow-up
Per the minimize-PRs directive, the D5+D6 banner remediation and
`redsCleared` semantic refresh are tracked separately in the coordinate
channel handoff — this PR is scoped to the `lastSuccessAt` definitional
fix.
## What does this PR do?
Fixes `pin-to-send` scrolling in the v2 chat view.
- Re-attaches the non-autoscroll scroll listener after the real scroll
element mounts by depending on `nonAutoScrollEl`, not the stable
`scrollRef` object.
- Lets the `usePinToSend` spacer adjust in both directions as content
below the pinned user message changes, so the user message stays
anchored after streaming finishes and layout height changes.
- Adds regression coverage for the scroll-to-bottom button and spacer
adjustment behavior.
## Related PRs and Issues
Fixes#5355
## Tests
- `corepack pnpm -C packages/react-core exec vitest run
src/v2/hooks/__tests__/use-pin-to-send.test.tsx
src/v2/components/chat/__tests__/CopilotChatView.pinToSend.test.tsx`
- `corepack pnpm exec oxfmt --check
packages/react-core/src/v2/components/chat/CopilotChatView.tsx
packages/react-core/src/v2/hooks/use-pin-to-send.ts
packages/react-core/src/v2/hooks/__tests__/use-pin-to-send.test.tsx
packages/react-core/src/v2/components/chat/__tests__/CopilotChatView.pinToSend.test.tsx`
- `git diff --check`
Attempted:
- `corepack pnpm -C packages/react-core run check-types`
- This failed in the local workspace on existing/type-resolution issues
outside this diff, including `react-markdown` JSX namespace errors,
missing `@copilotkit/runtime-client-gql` declarations, and existing e2e
mock `AbstractAgent` private member mismatches.
## Checklist
- [x] I have read the [Contribution
Guide](https://github.com/copilotkit/copilotkit/blob/master/CONTRIBUTING.md)
- [x] If the PR changes or adds functionality, I have updated the
relevant documentation (N/A: bug fix only, no API/docs change)
- [x] "Allow edits by maintainers" is checked (lets us help iterate on
your PR directly — faster turnaround for everyone)
## Summary
Small followups to PR #5413 (BIA D6 readiness, merged 2026-06-12).
- **`PARITY_NOTES.md` NSF naming**: corrected `hitl` (which is supported
in BIA via Strategy-B `useFrontendTool`) to call out the actual
NSF-quarantined demos `gen-ui-interrupt` and `shared-state-streaming`
per the manifest. PR #5413's NSF banner commit (`3585c33b8`) mounted the
banners on the correct demos.
- **Sub-Agents TODO removed**: post-merge staging dashboard showed
sub-agents RED; diagnose confirmed code on `5e828aed9` is correct (local
`--direct` D6 PASSES). The staging RED was stale dashboard image /
fixture cache, not code. PARITY_NOTES no longer claims sub-agents is a
known-issue.
- **gen-ui-agent STATE_DELTA op fix**: BIA's `set_steps` tool emission
used RFC-6902 `op: "replace"` on path `/steps`, but the agent's initial
state is `{}` and `fast-json-patch` rejects unresolvable paths in strict
mode (`@ag-ui/client@0.0.57` swallows the throw with `console.warn`).
Changing to `op: "add"` creates the path on first emission and
idempotently overwrites on subsequent calls. Single-line fix in
`tanstack-factory.ts`.
## Out-of-scope (tracked separately)
- **a2ui-fixed-schema + declarative-gen-ui** (PR #5413 documented
known-issues): root cause is in `@copilotkit/runtime/v2` v2-stack
pipeline OR `@ag-ui/a2ui-middleware@0.0.8` JSON.parse gap — one upstream
fix likely unblocks both. Diagnose reports at
`/tmp/cr/bia-rxr-diagnose/{a2ui-fixed-schema,declarative-gen-ui}.md`.
Needs separate packages PR.
- **tool-rendering port + headless-complete missing components**: BIA
needs StockCard, ChartCard, `get_revenue_chart` server tool, UI
primitives. Separate component-port PR.
- **shadcn catchall**: product/design call pending (escalate to PM).
## Test plan
- [x] Local `--direct` D6 verify on `gen-ui-agent` after the JSON-Patch
op fix
- [x] CI green on PR
Chronic content-reds in the D5/D6 worker families left `lastSuccessAt`
pinned to null on /api/runs, which tripped the dashboard "worker family X
has not completed successfully since Yh ago" silence banner even though
the workers were healthy. Each batch landed with jobs.done=1,
jobs.failed=17 (cells failing inside the probes), the §5.2.1 outcome
precedence derived "failed", findCompletedBatch returned null, and the
§9 family-silence monitor fell back to oldest-batch-enqueuedAt and fired
after 3 x period.
The fix is definitional: a batch counts as a "terminal completion" for
lastSuccessAt purposes when every job reaches a terminal state (done |
failed) with no `result.commError` on any row. Cell-level reds (rollup
counts) no longer block the timestamp; only worker-level outages
(crashed / reclaimed / lease-expired, surfaced via commError) do.
The strict §5.2.1 outcome=="completed" all-green semantic is unchanged
(it still drives lastRun.outcome and §6.2 dashboard rendering). No
caller of run-view consumed the old strict lastSuccessAt for logic
beyond the banner and monitor (grep verified — every external reference
is either a test fixture or the silence-banner display surface).
Tests added (RED -> GREEN):
- run-view.test.ts: lastSuccessAt advances when all jobs reach a terminal
state with no commError, even if cells failed (the regression case)
- run-view.test.ts: lastSuccessAt does NOT count a batch where any job
carries a commError (worker-outage signal stays loud)
- run-view.test.ts: lastSuccessAt does NOT count a batch with non-terminal
jobs (stalled / abandoned still excluded)
- family-silence-monitor.test.ts: does not alert when the family runs every
cycle with cell-level reds but no commError
- family-silence-monitor.test.ts: still alerts when the family stops
emitting results entirely (real outage)
Two existing tests updated to add commError to their "no completion"
fixtures so their original intent (walk-back across non-completion
batches) survives the new predicate:
- run-view.test.ts: lastSuccessAt walk-back through a commError batch
- fleet-runs.test.ts: T5 multi-page walk-back coverage
Fixes#5416
## Problem
`AgentStore` mirrored the agent's messages into a signal via
`this.#messages.set(abstractAgent.messages)`. `AbstractAgent.addMessage`
pushes **in place** and notifies with the **same array reference**.
Angular signals compare with `Object.is`, so `set(sameRef)` is a no-op:
the signal never notifies, the OnPush `<copilot-chat>` view is never
marked dirty, and a freshly added message — including the user's own
bubble on submit — does not render until the run pipeline reassigns
`messages` to a new array at run completion.
## Fix
```ts
this.#messages.set([...abstractAgent.messages]);
```
A shallow copy gives the array a new identity, so the signal notifies
and OnPush views re-render immediately. The `Message` objects stay
referentially stable, so `trackBy` still avoids re-creating existing
bubbles.
`onStateChanged` is intentionally left unchanged: `AbstractAgent`
reassigns `this.state` to a freshly cloned reference on every change
(`setState(e) { this.state = clone(e) }`) and never mutates state in
place, so that signal already notifies correctly. Scoping the fix to
messages keeps it minimal.
## Tests
Added two regression tests in `agent.spec.ts`, both exercising the
in-place mutation path that the existing tests never hit (they assign a
fresh array each emit):
1. **Reference guard** — the store must not hand back the agent's live
`messages` array.
2. **OnPush re-render** — a count rendered in an OnPush *descendant* (a
root `detectChanges()` would force-check the root and hide the bug; only
a child whose dirty flag depends on the signal exposes it). Two in-place
pushes; the second is the same-reference no-op the fix repairs.
Both tests fail without the fix (`'1' to be '2' // Object.is equality`)
and pass with it.
## Verification
- `agent.spec.ts`: 6/6 pass
- full `@copilotkitnext/angular` suite: 50/50 pass
- `tsc --noEmit`: clean
## Summary
- export the telemetry lambda client as a named binding instead of
default-reexporting it through the shared barrel
- update shared telemetry internals/tests to consume the named binding
- add a built-package smoke check to ensure both CommonJS and ESM expose
`lambdaClient.send`
## Root Cause
`@copilotkit/shared` re-exported `lambdaClient` from a default export.
The unbundled CommonJS build emitted `exports.lambdaClient =
require_lambda_client`, so CommonJS consumers received the module
namespace object instead of the `{ send }` client. `@copilotkit/runtime`
then called `lambdaClient.send(...)` and crashed because `send` was
nested under `lambdaClient.default`.
## Validation
- `pnpm nx run @copilotkit/shared:build --skip-nx-cache`
- `node packages/shared/scripts/verify-cjs-exports.cjs`
- `pnpm nx run @copilotkit/shared:test --skip-nx-cache`
- `pnpm nx run @copilotkit/shared:check-types --skip-nx-cache`
- `pnpm nx run @copilotkit/shared:publint --skip-nx-cache`
- `pnpm nx run @copilotkit/runtime:build --skip-nx-cache`
- runtime CJS telemetry capture smoke test
- `pnpm nx run @copilotkit/runtime:test --skip-nx-cache --
src/v2/runtime/__tests__/telemetry.test.ts`
- pre-commit hook: `pnpm run test && pnpm run check:packages`
## Summary
Fixes#5072
Fix unstyled Streamdown markdown in non-Tailwind apps that import
`@copilotkit/react-core/v2/styles.css`.
Streamdown renders plain utility class names, but CopilotKit’s CSS
bundle is generated with the `cpk:` Tailwind prefix. That left default
markdown elements like headings, lists, strong text, tables,
blockquotes, inline code, and code blocks without matching styles unless
the host app also shipped unprefixed Tailwind utilities.
## Changes
-> Add scoped fallback styles for Streamdown’s stable `data-streamdown`
elements under `[data-copilotkit]`.
-> Use CopilotKit theme variables and prefixed Tailwind utilities so
styles stay inside the CopilotKit surface.
-> Remove the ineffective Streamdown `@source` scan, since it only
generated prefixed utilities that Streamdown does not emit.
-> Add a regression test to ensure the scoped Streamdown selectors
remain present.
## Testing
-> `node node_modules\nx\bin\nx.js run @copilotkit/react-core:build:css
--excludeTaskDependencies`
-> `node node_modules\nx\bin\nx.js run @copilotkit/react-core:test
--excludeTaskDependencies --skip-nx-cache --
src/v2/styles/__tests__/streamdown-styles.test.ts`
AgentStore mirrored the agent's messages into a signal via
`this.#messages.set(abstractAgent.messages)`. AbstractAgent.addMessage
pushes in place and notifies with the same array reference, so the
signal's Object.is equality check treats set(sameRef) as a no-op: it
never notifies, the OnPush <copilot-chat> view is never marked dirty,
and a freshly added message (including the user's own on submit) does
not render until the run pipeline reassigns messages to a new array at
run completion.
Copy into a fresh array so the reference changes and the signal
notifies. Message objects stay referentially stable, so trackBy still
avoids re-creating existing bubbles. State is unaffected: AbstractAgent
reassigns this.state to a cloned reference on every change, so its
signal already notifies.
Adds regression tests covering the in-place mutation path: a reference
guard and an OnPush descendant re-render check. Both fail without the
fix.
Fixes#5416
## Summary
Follow-up to #5383 (per-agent A2UI scoping). Two small,
backwards-compatible runtime changes:
1. **One shared `isA2UIEnabled()` predicate.** Before, "is a2ui on?" was
computed independently in two places — `if (runtime.a2ui)` in the run
path (`agent-utils.ts`) and `!!runtime.a2ui` in the `/info` response
(`get-runtime-info.ts`). Two separate checks can drift apart, and that
drift between the run path and the info path is the root of #5369. Both
now call the same predicate, so they can't disagree.
2. **An explicit `enabled` off-switch** on the runtime `a2ui` config.
Previously, once `a2ui` was configured there was no way to turn it off
short of deleting the whole block. Now `a2ui: { enabled: false }`
disables it while keeping the rest of the config (e.g.
`schema`/`catalog`) in place.
```ts
export function isA2UIEnabled(a2ui) {
return !!a2ui && a2ui.enabled !== false;
}
```
## Backwards compatibility
No breaking changes. `enabled` is a new optional field; every existing
config behaves identically:
- omitted `a2ui` → off (unchanged)
- `a2ui: {}` → on (unchanged)
- `a2ui: { injectA2UITool: true }` / `{ agents: [...] }` → on
(unchanged)
- only an explicit `a2ui: { enabled: false }` is new behavior — and it's
an opt-out nobody could have relied on.
Enablement stays entirely server-side; no client/provider API change.
## Scope
This does **not** fix the #5369 leak — #5383 does that via per-agent
scoping. This is complementary hardening: it adds the missing off-switch
and removes the run-path/info-path divergence that allowed the leak to
exist in the first place.
## Tests
- `get-runtime-info`: reports `a2uiEnabled: false` when `a2ui: {
enabled: false }`.
- run path (`handle-run`): does not apply `A2UIMiddleware` when
`a2ui.enabled` is `false`.
https://claude.ai/code/session_01TYohiEJyhsU3mJS4jabdv6
---
_Generated by [Claude
Code](https://claude.ai/code/session_01TYohiEJyhsU3mJS4jabdv6)_
## Summary
Content-red dashboard remediation across CopilotKit showcase
integrations (spring-ai, ag2, claude-sdk-typescript, llamaindex). Aligns
aimock fixtures (d4 and d6) with real backend output shapes, ports A2UI
v0.9 declarative legs, quarantines unsupported features as NSF with the
canonical d5-feature-mapping literals, adds route aliases the harness
expects, and fixes backend bugs in the per-family agent stacks.
## Verification (local `bin/showcase test <family> --d6`
baseline-vs-branch comparison)
**13 dashboard cells go from RED on `main` to GREEN on this branch:**
| Family | Cells fixed |
|--------|-------------|
| spring-ai | auth, gen-ui-open, gen-ui-open-advanced, headless-simple,
subagents, tool-rendering-custom-catchall |
| ag2 | gen-ui-a2ui-fixed, gen-ui-agent, tool-rendering |
| claude-sdk-typescript | beautiful-chat-search-flights,
gen-ui-a2ui-fixed, gen-ui-agent, tool-rendering |
| built-in-agent | headless-simple ✓ (verified — was the targeted A20
fix) |
**Zero regressions on spring-ai, claude-sdk-typescript.** Pre-existing
reds unchanged.
## ⚠️ Known regression — DOCUMENTED, not yet fixed
- **ag2 / `mcp-apps`**: was GREEN on `main`, RED on this branch.
Per-pill snapshot confirms (`/tmp/cr/baseline-ag2-pills.json` vs
`/tmp/cr/ag2-final-pills.json`). Harness signature: `assistant did not
respond within 30000ms (turns_completed=0)` → "Run ended without
emitting a terminal event". Root cause not localized; suspect
interaction between R7-A1 (`http.disconnect` synthesis dropped from
request-context middleware) and the mcp-apps SSE shutdown path, but the
middleware fix was deliberate ASGI semantics correction and reverting it
re-opens the SSE-abort it was meant to address.
Possible remediations (for follow-up):
1. Investigate the mcp-apps `/api/copilotkit-mcp-apps` SSE close
sequence under R7-A1's "always await real disconnect" semantics. If the
inner FastAPI app holds the channel waiting for a real disconnect, the
harness may time out before uvicorn signals it.
2. Re-evaluate R7-A1: emit a synthetic disconnect ONCE per request, but
ONLY in response to a specific signal (e.g. `more_body=False` for the
inner app's final receive), not unconditionally.
This regression was caught BY the verification gate, NOT bypassed.
Documented here so reviewers can decide ship-as-is + follow-up vs
hold-for-fix.
## Pipeline summary
- 8 cr-loop review rounds, 1000+ findings deduped per round
- 5 fix cycles (R1: 24 items; R2: 10 R2-A regressions + 5 in-diff
bucket-(a); R3: 7 items; R5: 3 items; R7: 4 items; R9: 1 byoc demo
cleanup)
- Procedure 3 bucket-(c) promotion audit: 0 PROMOTE_TO_A
- Rebased onto `origin/main` (203-commit gap caught at pre-push)
- `pre-push-quality` clean: formatter (oxfmt + ruff) green, lint green,
typecheck green, tests green, build green, docker build green
(`showcase-spring-ai:prepush-test` 3.62GB), actionlint green on new
workflow
- Local `bin/showcase test --d6` verification per the user's hard gate —
5 families verified + baseline comparison
## Coordinated work
Coordinated via `coordinate` channel with peer session
`dashboard-cell-fixes` (PR #5407, merged: D6 manifest cleanup +
claude-sdk-python testid backfill, 13 manifests excluding ag2 +
spring-ai). Manifests ag2 + spring-ai are this PR's territory;
non-overlap confirmed.
## Test plan — what improved
Per-family local `bin/showcase test <family> --d6` baseline-on-main vs
branch (per-pill PocketBase snapshots in
`/tmp/cr/{baseline,our}-*-pills.json`). The change attributable to this
PR:
| Family | Δ on D6 dashboard | Cells flipped red → green |
|--------|-------------------|----------------------------|
| **spring-ai** | **+6** | `auth`, `gen-ui-open`,
`gen-ui-open-advanced`, `headless-simple`, `subagents`,
`tool-rendering-custom-catchall` |
| **ag2** | **+2 net** (+3 fixed, −1 regression) | fixed:
`gen-ui-a2ui-fixed`, `gen-ui-agent`, `tool-rendering` · regression:
`mcp-apps` (see ⚠️ above) · also: dangling `byoc` declarative-* demos
removed from manifest |
| **claude-sdk-typescript** | **+4** | `beautiful-chat-search-flights`,
`gen-ui-a2ui-fixed`, `gen-ui-agent`, `tool-rendering` |
| **llamaindex** | **+1** | `gen-ui-agent` (R1-A4 target) —
`shared-state-write` confirmed green |
| **built-in-agent** | **+1** | `headless-simple` (A20 targeted fix) |
**Net dashboard impact: 13 cells red → green, 1 regression
(ag2/`mcp-apps`) = +12 net.**
What landed per family (commit per concern, in stack order):
- **`975d5d80b` spring-ai** — manifest NSF reasoning canon (closes
spring-ai's reasoning-default-render + agentic-chat-reasoning +
tool-rendering-reasoning-chain pills against AG-UI Java SDK's missing
`REASONING_MESSAGE_*` union) + Jackson `@Primary` ObjectMapper bean-race
fix (deterministic JSON wire shape) + `RunErrorEvent` mixin renaming
`error` → `message` for AG-UI core wire compatibility + error-banner
pattern across mcp-apps routes + Java tool JSON hardening + `.ag-ui-sha`
centralized pin source + new `test_unit-spring-ai.yml` workflow +
Dockerfile explicit `git fetch origin $SHA` guard. → fixes the 6
spring-ai cells above by closing the wire-shape gap between spring-ai
and the harness.
- **`2b0ce1065` ag2** — manifest NSF reasoning canon + new ASGI
middleware `_request_context.py` with `_latest_user_message`
`ContextVar` (unblocks per-request user-message threading the harness
asserts on) + async OpenAI client + tool serialization fix + R9-A2
manifest cleanup of stale declarative-* demos. → fixes the 3 ag2 cells.
- **`4e79e3bd1` claude-sdk-typescript** — per-content-block array
refactor for text/thinking/tool_use replay order on Anthropic
`/v1/messages` shape (the core fix) + per-block `randomUUID` + reasoning
route + alias coherence. → fixes the 4 cst cells by aligning Anthropic
content-block sequencing with the harness's strict replay-order check.
- **`5fd4028b4` llamaindex** — `gen-ui` + beautiful-chat agent updates +
mcp-apps + copilotkit route alias coherence + demo-files assets. → fixes
`gen-ui-agent`.
- **`aec5b82ed` per-integration adjacencies** — `@ts-expect-error` swaps
+ agno comment + built-in-agent `headless-simple` UUID fallback + error
banner across agno/crewai-crews/langroid/mastra/pydantic-ai/strands
mcp-apps routes. → fixes built-in-agent `headless-simple`; prevents
banner-render regressions across 6 other families.
- **`a02cfb7a7` aimock** — d4 broad-matcher narrowing + d6 per-family
content-shape alignment + chunkSize emoji surrogate-pair guard +
threadid context gates. → the fixture realignment that makes all the
above backend fixes verifiable end-to-end through the harness.
- **`1f5b90095` Validate Showcase fix** — `KNOWN_SHADOW_CEILING` 123 →
128 in `aimock-fixtures.test.ts` (the d6 content-shape alignment adds
tracked, runtime-disambiguated substring overlaps; this is the test
author's documented escape hatch — see the comment around the constant).
Quality gates run:
- [x] `pre-push-quality` — oxfmt+ruff, lint, typecheck, tests, monorepo
build, `docker build` on `showcase-spring-ai:prepush-test` (3.62 GB
image), actionlint on the new CI workflow — all green
- [x] CI: 43/43 green on `1f5b90095` (Validate Showcase needed attempt
2; `generate-registry.test.ts:211` is a known sentinel-cleanup flake
unrelated to this PR's diff)
**Decision needed**: ship-as-is and roll `ag2/mcp-apps` regression into
the consolidated follow-up PR (alongside the bucket-(c)/(d) deferred
items in `## Follow-up` below), OR hold this PR and fix `mcp-apps` in-PR
first. Two candidate remediations are described in the ⚠️ block above.
## Upstream gaps (flagged, not in scope)
- AG-UI Java SDK lacks `REASONING_MESSAGE_*` typed event union —
spring-ai uses raw-JSON workaround in `ReasoningController` (manifest
declares `reasoning-default-render` + `agentic-chat-reasoning` +
`tool-rendering-reasoning-chain` NSF accordingly)
- aimock has no `customEvents` key — blocks headless-interrupt D6 replay
(interrupt-headless red across all families)
- llama-index `ag-ui` lacks `REASONING_MESSAGE_*` ingestion gap —
manifest declares accordingly
- ag2 image-content rejection — multimodal pill red across ag2 (Python
ValueError in agent.py); separate from this PR
## Follow-up
A consolidated follow-up handoff file documents the deferred
bucket-(c)/(d) items (R2-D1 stack-leak hardening sweep across 11 routes,
R2-D2 customEvents, R2-D3 @copilotkit/runtime agents-generic, R2-D4
mcp.excalidraw.com default URL, R3 misc adjacencies, etc.) for a single
consolidated follow-up PR per the "minimize PRs by aggregating related
work" directive.
The 5 new substring shadows are in d6 fixtures landed by this PR:
- d6/ag2/gen-ui-declarative.json: 'Show me a quick KPI dashboard' inner-call mirror overlaps with pre-existing 'KPI dashboard' entry in render-a2ui.json (same toolName=render_a2ui, context=ag2). Runtime-disambiguated by load order (inner-call mirrors ordered BEFORE outer generate_a2ui fixtures per the _meta._note in that file).
- d6/claude-sdk-typescript/gen-ui-declarative.json: same pattern as ag2.
- d6/claude-sdk-typescript/tool-rendering.json: dropped userMessage/turnIndex gates on toolCallId-keyed follow-up fixtures (per _shape_note in that file — Anthropic /v1/messages shape needs toolCallId-only gating for multi-pill loop safety). The remaining ungated userMessage fixtures ('weather in Tokyo', 'AAPL') now substring-overlap with gen-ui-headless-complete and tool-rendering-reasoning-chain prompts in the same context. Runtime-disambiguated by toolCallId chains and first-match-wins ordering.
These overlaps are exactly the runtime-disambiguated pattern the test comment endorses for ceiling bumps.