## Summary
- Added `data-testid="copilot-assistant-message"` and
`data-message-role="assistant"` to the BYOC hashbrown and json-render
custom assistant message renderers in claude-sdk-typescript
- These attributes were already present in the langgraph-python
gold-standard; this aligns the TS port
## Why
The byoc-hashbrown and byoc-json-render demos override CopilotChat's
`messageView.assistantMessage` slot with custom renderers. The e2e-deep
conversation runner counts assistant messages via
`[data-testid="copilot-assistant-message"]` to detect "response
settled." Without that attribute on the custom renderers, the count
stayed at 0 forever and the byoc feature timed out -- the only D5
failure blocking claude-sdk-typescript from green.
## Test plan
- `showcase test claude-sdk-typescript --d5 --verbose` goes from 28/29
(byoc timeout) to 29/29 green
- No demo functionality changed -- only test-harness attributes added to
wrapper divs
The byoc-hashbrown and byoc-json-render demos override the CopilotChat
assistantMessage slot with custom renderers, but the overrides were
missing the data-testid="copilot-assistant-message" attribute that the
e2e-deep conversation runner uses to detect settled assistant responses.
Without it, the runner's readMessageCount always returned 0 and the
byoc feature timed out at D5.
The langgraph-python gold-standard already had these attributes; this
aligns claude-sdk-typescript to match.
## Summary
- docs-only features (e.g. cli-start) exist purely for documentation
coverage tracking and have no route, depth probes, or health signals
- The catalog metadata was counting their 18 stub cells in the headline
wired/stub/unshipped/unsupported breakdown, inflating the total and
making the stats bar misleading
- Exclude docs-only cells from the headline counts; add a new
`docs_only` field to track them separately
**Before:** total_cells=720 wired=673 stub=18 unsupported=29
**After:** total_cells=702 wired=673 stub=0 unsupported=29 docs_only=18
## Test plan
- [x] `showcase/scripts` test suite: all 574 tests pass (including 13
catalog tests)
- [x] `showcase/shell-dashboard` test suite: same 3 pre-existing
failures, no regressions
- [x] TypeScript check clean (no new errors)
- [x] Invariant holds: wired + stub + unshipped + unsupported +
docs_only == cells.length
docs-only features (e.g. cli-start) exist purely for documentation
coverage tracking -- they have no route, no depth probes, and no
health signals. The catalog metadata was counting their 18 stub cells
in the headline wired/stub/unshipped/unsupported breakdown, inflating
the total and making the stats bar misleading.
Exclude docs-only cells from the headline counts. A new docs_only
field tracks the excluded count separately so the invariant
(wired + stub + unshipped + unsupported + docs_only == cells.length)
holds.
Before: total_cells=720 wired=673 stub=18 unsupported=29
After: total_cells=702 wired=673 stub=0 unsupported=29 docs_only=18
## Summary
- Add `CustomAssistantMessage` and `CustomDisclaimer` components to
llamaindex chat-slots demo, matching langgraph-python gold standard
- Wire `input.disclaimer` and `messageView.assistantMessage` slot props
into `<CopilotChat>` alongside the existing `welcomeScreen` slot
- Update e2e spec to D5 coverage (welcome screen, suggestions, assistant
message slot, disclaimer slot, multi-turn slot persistence)
- Update manifest.yaml highlight list with new component files
## Why
The llamaindex chat-slots demo only had the `welcomeScreen` slot
override, while the D5 probe (`d5-chat-slots.ts`) asserts
`data-testid="custom-assistant-message"` is present after the assistant
responds. This left the feature stuck at D4.
## Test plan
- [x] `showcase build llamaindex` -- builds clean
- [x] `showcase test llamaindex --d5 --verbose` -- chat-slots passes
(was the only feature stuck at D4)
- [x] Verified Docker container renders all three slot testids
Replaces the 501 stubs in handleGetThreadEvents and handleGetThreadState
with real delegation to CopilotKitIntelligence. Adds two new client
methods (`getThreadEvents`, `getThreadState`) on intelligence-platform/
client.ts that hit the new `/api/_inspect/threads/:id/{events,state}`
routes shipped in Intelligence PR #144 (CPK-7453). Auth flows through
the existing `resolveIntelligenceUser` path; threadId scoping happens
server-side via the API key's org/project resolution.
Wire shapes match the in-memory branch so the inspector consumes both
runtimes identically:
- events: `{ events }` (platform-internal `decodeErrorRowIds` and
`truncated` flags are stripped at the runtime boundary)
- state: `{ state }` where the platform's discriminated `ThreadStateResult`
flattens to the snapshot value for `kind: "snapshot"` and to `null` for
both `no-snapshot` and `snapshot-decode-error`
Replaces the 501 regression-protection tests with real delegation tests
that mock the platform methods, assert call args, and exercise the
no-snapshot / decode-error / throw paths.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The llamaindex chat-slots demo only had the welcomeScreen slot,
missing the input.disclaimer and messageView.assistantMessage
overrides that langgraph-python (gold standard) provides. The D5
probe checks for data-testid="custom-assistant-message" after the
assistant responds, so the feature was stuck at D4.
Add CustomAssistantMessage and CustomDisclaimer components (matching
langgraph-python), wire them into CopilotChat props, update the e2e
spec to cover all three slots, and update manifest highlights.
The first reactivity fix triggered a `_loadingMessages` flicker between
streaming chunks because every live re-fetch toggled the loading state
and replaced the conversation array wholesale. The silent flag suppresses
both: live re-fetches keep the loading indicator off and preserve the
last-good conversation on transient fetch errors. Initial threadId-change
fetches still show a real loading state.
Adds the same staleness guard the other tab fetches use (skip applying
results if `threadId` changed mid-flight) so a quick thread switch can't
leave the wrong thread's data on screen.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
## Summary
- Widen dialog from `w-72` (288px) to `w-[480px]` so error details are
not cramped
- Extract key signal fields (`errorDesc`, `error`, `failureSummary`,
`backendUrl`, `apiRequestCount`, `step`) as readable key-value pairs
instead of raw JSON
- Deduplicate `errorDesc` and `error` when both present (first match
wins)
- Make raw signal payload collapsible ("Raw Signal" toggle) for
debugging
- Remove duplicate tooltip text that repeated status/timing info already
conveyed by badge color and failure metadata
- Compact fail_count and first_failure_at into a single inline row
## Test plan
- [x] All 11 cell-drilldown tests pass (4 new tests added for extracted
fields, collapsible toggle, deduplication, and wider width)
- [x] No new TS errors introduced
- [x] Pre-existing test/TS failures in other files are unrelated
Widen dialog from w-72 to w-[480px], extract key signal fields
(errorDesc, error, failureSummary, backendUrl) as readable key-value
pairs instead of buried JSON, make raw signal collapsible, remove
duplicate tooltip text that repeated badge color/status info.
## Summary
- Suppress depth chip rendering for docs-only features in the RefDepth
column (parity overlay) -- show `--` instead, since docs-only features
have no probe coverage and should not display a depth level
- Add depth distribution counts (D6, D5, D4, D3, D2, D1) to the
AdaptiveStatsBar when the depth overlay is active, giving operators
at-a-glance visibility into how many wired cells sit at each depth level
## Test plan
- [ ] Verify docs-only features (e.g. `cli-start`) show `--` in the Ref
Depth column instead of a depth chip
- [ ] Verify non-docs-only features still show their correct depth chip
in Ref Depth
- [ ] Toggle the depth overlay on and confirm the depth distribution
section appears in the stats bar
- [ ] Toggle depth overlay off and confirm the distribution section
disappears
Suppress RefDepthCell for docs-only features in the parity overlay,
showing '--' instead of a depth chip since docs-only features have no
probe coverage. Add depth distribution counts (D6..D1) to the
AdaptiveStatsBar when the depth overlay is active, giving operators
at-a-glance visibility into how many wired cells are at each depth.
The conversation view in cpk-thread-details only re-fetched on threadId
change, so live agent output during streaming wasn't visible until the
user switched threads (or unmounted/remounted the element by clicking
out and back in). Restoring the original conversationOverride channel
isn't viable because the parent's agentMessages map is keyed by agentId
(not threadId) and ConversationItem mapping lives in the child — passing
mapped messages would either leak across threads or duplicate the
runtime's AG-UI → ConversationItem conversion in the parent.
Instead, the parent now tracks each agent's currently-running threadId
(from `onAgentRunStarted`) and ticks a per-thread liveMessageVersion
counter every time `syncAgentMessages` fires for that agent. The counter
is passed to cpk-thread-details, which watches it in `updated()` and
re-fetches `/threads/:id/messages` when it changes for the same threadId.
The runtime endpoint is the single source of truth for conversation
shape, so streaming output flows in without any client-side mapping.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
## Summary
- Adds `docs-only` as a new `FeatureKind` alongside `primary` and
`testing`
- Features with this kind show only the docs row in the dashboard -- no
Dx badges (API/RT/CV), no links, no depth chips
- Row label is muted/italic with a "docs-only" tag, matching the visual
treatment of "testing" rows
- `cli-start` is the first feature changed to `docs-only` since it has
no runnable demo
## Changes
| File | What |
|------|------|
| `showcase/shared/feature-registry.json` | `cli-start` kind: `primary`
-> `docs-only` |
| `showcase/shell-dashboard/src/lib/registry.ts` | `FeatureKind` union
extended |
| `showcase/shell-dashboard/src/components/cell-pieces.tsx` |
`CellStatus` returns null for docs-only |
| `showcase/shell-dashboard/src/components/composed-cell.tsx` | Only
renders DocsLayer for docs-only |
| `showcase/shell-dashboard/src/components/feature-grid.tsx` | Muted row
styling + tag |
| `cell-pieces.test.tsx` | New test for docs-only hiding all badges |
| `composed-cell.test.tsx` | 3 new tests for docs-only overlay behavior
|
## Test plan
- [x] New unit tests pass for all three behavioral changes
- [x] Pre-existing test failures confirmed unrelated (useOverlays URL
hash, composed-cell docs-dedup, missing generated data files)
Features with kind "docs-only" show only the docs row in the
dashboard -- no Dx badges, no status badges, no links or depth
layers. The row label is muted and tagged like "testing" rows.
- feature-registry.json: cli-start changed from primary to docs-only
- registry.ts: FeatureKind union extended with "docs-only"
- cell-pieces.tsx: CellStatus returns null for docs-only (hides all
badges)
- composed-cell.tsx: docs-only features render only DocsLayer, skip
links/depth/health
- feature-grid.tsx: docs-only rows use muted italic styling with a
"docs-only" tag
- Tests added for all three behavioral changes
## Summary
Adds a clickable "Run Evaluation" link in a bot comment on every PR. One
click triggers the showcase eval — no Checks tab navigation needed.
### How it works
1. `showcase_eval_check.yml` creates the Check Run AND posts a bot
comment with a signed link
2. The link points to
`eval-webhook-production.up.railway.app/trigger/eval?pr=N&check_run_id=M&sig=HMAC`
3. User clicks → webhook service verifies signature, dispatches eval
workflow, updates Check Run to in_progress, redirects user back to the
PR
4. Eval results appear in the Check Run (already implemented in #4517)
### Security
- Link is HMAC-signed using the webhook secret — can't be forged
- Signature covers `pr:check_run_id` — changing either invalidates the
link
- `EVAL_WEBHOOK_SECRET` added as repo secret for the signing step
### Changes
- `showcase/eval-webhook/src/server.ts` — new `GET /trigger/eval` route
with signature verification + redirect UX
- `.github/workflows/showcase_eval_check.yml` — posts bot comment with
signed link (updates on push, idempotent via HTML marker)
## Test plan
- [ ] Open a test PR → verify bot comment appears with "Run Evaluation"
link
- [ ] Click the link → verify redirect page shows, eval dispatches,
Check Run updates
- [ ] Tamper with the sig parameter → verify 403 response
deriveDepth() was computing D2 for not_supported_features because D1/D2
are integration-scoped. Add unsupported boolean to DepthResult, guard in
deriveDepth, update consumer components. 24/24 depth-utils tests pass.
Fix chat-slots agent name (dashes to underscores), prebuilt-popup
defaultOpen prop, auth demo defaults to unauthenticated state, and
add missing data-testid attributes to byoc renderers.
## Summary
Fixes 15 ag2 showcase files to bring all 29 ag2 features to D5-green
(verified locally with the D5 probe runner).
### Backend fixes
- **Starlette mount order**: `/open-gen-ui-advanced` must mount before
`/open-gen-ui` — Starlette resolves mounts via prefix matching in
registration order, so the shorter prefix was shadowing the longer one
- **`@tool` -> `@tool()` decorator**: bare `@tool` in
`agent_config_agent.py` silently fails to register the function as a
CopilotKit tool; needs the call form `@tool()`
- **headless-complete route wiring**: wired the headless-complete agent
into the `copilotkit-mcp-apps` runtime route so MCP Apps activity
messages render in the headless cell
- **Auth route cleanup**: removed stale `@ts-ignore`
- **Manifest update**: corrected headless-complete's route reference
from `copilotkit` to `copilotkit-mcp-apps`
### Demo page fixes
- **agent-config**: replaced broken `useAgent`/`setState` pattern with
CopilotKit `properties` prop (the correct way to pass config to AG2
ContextVariables)
- **multimodal**: rewrote the legacy-binary-shape converter to a content
flattener — AG2's AGUIStream validates message content as a plain string
and rejects content-part arrays with a 400
- **headless-complete**: changed `runtimeUrl` from `/api/copilotkit` to
`/api/copilotkit-mcp-apps`
- **byoc-hashbrown, byoc-json-render**: re-attached
`data-testid="copilot-assistant-message"` on custom slot overrides
(harness conversation runner needs this to count settled responses)
- **agentic-chat-reasoning, beautiful-chat**: added demo-level
`data-testid` wrappers
- **tool-rendering-reasoning-chain, voice**: fixed CopilotKit import
paths
### Railway rebuild note
Several demos (chat-slots, shared-state-*, subagents) were already
code-correct on `main` but showed D5 failures due to a stale Railway
image. A Railway rebuild (no code changes) will bring those green as
well.
## Test plan
- [ ] 29/29 ag2 features pass D5 probes locally (verified before PR)
- [ ] CI passes
- [ ] After merge + Railway rebuild, ag2 column should be fully D5-green
Rewrite the multimodal demo's message converter from a legacy-binary-shape
rewriter to a content flattener. AG2's AGUIStream validates message content
as a plain string and rejects arrays of content parts with a 400. The new
ContentFlattenerShim extracts text from multipart user messages before the
AG-UI run dispatches them to the AG2 backend.
Also adds defensive validation to sample-attachment-buttons: magic-byte
checks, LFS pointer detection, and actionable error messages.
- Fix Starlette mount order: /open-gen-ui-advanced before /open-gen-ui to
prevent prefix shadowing (open-gen-ui-advanced was unreachable)
- Fix @tool -> @tool() decorator in agent_config_agent.py (bare @tool
silently fails to register the function as a CopilotKit tool)
- Wire headless-complete agent into copilotkit-mcp-apps route so the
headless-complete demo can exercise MCP Apps rendering
- Update manifest to reference copilotkit-mcp-apps route for headless-complete
- Remove stale @ts-ignore in auth route
Add GET /trigger/eval route to eval-webhook with HMAC-signed URLs.
The showcase_eval_check.yml workflow now posts a bot comment with a
clickable "Run Evaluation" link. Clicking triggers the eval and
redirects back to the PR. The link is signed so it can't be forged.
## Summary
- Fix TOOL_CALL_RESULT events reusing parent assistant message ID across
StreamingToolAgent, SharedStateReadWriteController, and
SubagentsController —
React's deduplicateMessages() Map overwrote assistant messages with tool
messages when they shared the same ID
- Add `data-testid="copilot-assistant-message"` to BYOC hashbrown
renderer so
the D5 test harness can detect assistant messages
- Register AG-UI Jackson mixins via JacksonConfig postConfigurer,
normalize
array-format content via ContentNormalizingModule, remove ChatMemory
beans
that interfered with aimock fixture matching, inject Spring-managed
ObjectMapper into AgentConfigController
## Test plan
- [x] All 28/28 spring-ai D5 features pass locally (verified stable
across
multiple consecutive runs)
- [x] Docker build succeeds
- [x] No regressions — previously passing 22 features remain green
Root cause: TOOL_CALL_RESULT events reused the parent assistant
message ID, causing React deduplicateMessages() to overwrite
assistant messages with tool messages (Map keyed by ID). Fix uses
unique UUIDs for each tool result message across all three
controllers. Also adds data-testid to the BYOC hashbrown renderer
so the D5 harness can detect assistant messages, registers AG-UI
Jackson mixins via JacksonConfig postConfigurer, normalizes
array-format content via ContentNormalizingModule, removes ChatMemory
beans that interfered with aimock fixture matching, and injects
Spring-managed ObjectMapper into AgentConfigController.
Three new test areas covering surfaces this PR introduced or relies on:
- fetch-router: matchRoute tests for `/threads/:id/events`,
`/threads/:id/state`, and `/threads/clear` (with and without URL
encoding). Critically pins that "/threads/clear" resolves to
`threads/clear` and does NOT fall through to the more permissive
`threads/update` arm with threadId="clear" — the explicit guard in
the router exists for this reason.
- fetch-handler validation: 405 enforcement for POST/PATCH/DELETE on
the read-only `/threads/:id/events` and `/threads/:id/state`
endpoints, with `Allow: GET` header. Complementary positive case
asserts GET is NOT a 405.
- handle-run: end-to-end test that handle-run.ts:40
(`agent.agentId = agentId`) propagates the registry key onto historic
runs. Without this stamp, InMemoryAgentRunner falls back to "default"
and the agentId filter on `GET /threads?agentId=...` breaks for the
local-dev fallback.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Pins three integration cases for the thread-store auto-unregister branch
in CopilotKitCore's internal onAgentsChanged subscriber:
1. agent removed from agents → store IS unregistered; subscriber receives
the previous store via prevStore.
2. FIRST onAgentsChanged({ agents: {} }) on a published-style core where
the store was registered before any agents arrived → store SURVIVES.
Reproduces the race that the "previously had" guard exists to
prevent.
3. add → register → remove cycle → store IS unregistered. Complements (2)
by exercising the same code path's positive branch.
Test (1) surfaced a real bug in the seed: agentRegistry.initialize does
NOT emit onAgentsChanged, so the constructor's internal subscriber never
saw the agents__unsafe_dev_only set, and `previousAgentIds` stayed empty.
A later removeAgent__unsafe_dev_only call would then be guarded into a
no-op because the agentId looked "new". Fixed by seeding
`previousAgentIds = new Set(Object.keys(agents__unsafe_dev_only))` right
before wiring the subscriber.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Type quality:
- intelligence/threads.ts handleGetThreadMessages: switch on the Message
discriminant (role) and read narrowed fields directly. Removes
`as Record<string, unknown>` laundering and chained `as` casts on
toolCalls/function/arguments. AssistantMessage's toolCalls always have
`function: { name, arguments }`, so the prior fallbacks (`tc.name`,
`tc.args`) were dead branches.
- in-memory.ts getThreadState: import StateSnapshotEvent and use it
instead of `(event as { snapshot?: unknown }).snapshot`.
Comment / API alignment:
- intelligence/threads.ts handleClearThreads JSDoc no longer claims the
inspector calls this; the actual caller is the demo button.
- in-memory.ts clearThreads JSDoc updated to match.
- in-memory.ts getThreadEvents JSDoc no longer references a SQLite
runner that does not exist; just describes the compaction logic.
- web-inspector lint fix: rename unused `changed` parameter to
`_changed` per oxc no-unused-vars.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Covers: failure triage (text-only vs tool features, probe text extraction,
message disappearing, chatMemory pollution), framework-specific gotchas
(LlamaIndex, Spring AI, Built-in Agent, MS Agent, AG2), and production
vs local parity checklist.
CopilotKitCore subscribes to onAgentsChanged and unregisters thread
stores for any agentId not in the new agents map. For published cores,
core.agents is asynchronously populated, so the FIRST
onAgentsChanged({ agents: {} }) notification fires BEFORE published
agents are merged in. Without a guard, that empty notification rips out
a thread store that a consumer (e.g. useThreads) just registered.
Track previousAgentIds and only unregister an agentId that was present
in the previous snapshot AND missing from the new one. The first
empty-agents notification (where the agentId was never previously
present) becomes a no-op.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Five behavioural fixes on cpk-thread-details / WebInspectorElement:
- fetchEvents/fetchState: mirror the AbortController pattern fetchMessages
already uses. Without this, switching threads quickly (A→B) can leave
the user looking at thread B with thread A's events/state when A's
request resolves last.
- mapMessages: when JSON.parse fails on tool-call args or tool result
content, log via console.error and attach __parseError + __raw on the
parsed object instead of silently substituting `{}`. The inspector is a
debugging surface; hiding malformed payloads defeats its purpose.
- fetchAnnouncement: keep the captured error and console.warn it. The
prior `catch {}` swallowed Malformed-payload throws, JSON parse
failures, and convertMarkdownToHtml exceptions silently.
- subscribeToThreadStore: also subscribe to ɵselectThreadsError, store
per-agent in `_threadsErrorByAgent`, and surface in renderThreadsView /
cpk-thread-list as an error state branch alongside empty/no-results.
Previously a thread-store load failure (REST list rejection, Phoenix
subscribe failure, retry exhaustion) left the user with stale data and
no indication of failure.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The button hit DELETE /threads, which the runtime router resolves to
"threads/list" (a GET-only route) and returns 405. The intended endpoint
is POST /threads/clear (defined at fetch-router.ts:175). Add try/catch
and surface non-ok HTTP status via console.error so the click no longer
silently appears to succeed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
## Summary
Replaces the `/eval` PR comment trigger with a professional GitHub Check
Run button UX, and adds native (no-Docker) CI execution infrastructure.
### Check Run Button (Phase 1)
- **showcase_eval_check.yml** — creates a Check Run with "Run Showcase
Eval" action button on every PR open/push
- **showcase/eval-webhook/** — tiny Hono relay service on Railway that
receives `check_run.requested_action` webhooks, authenticates as devops
bot, and dispatches `showcase_eval.yml` via `workflow_dispatch`
- **showcase_eval.yml** — adds `workflow_dispatch` trigger with
`dispatch-gate` job, Check Run lifecycle updates (neutral → in_progress
→ completed with results), devops bot token for Checks API
### Native Execution (Phase 2)
- **`--ci` flag** on eval orchestrator — skips Docker lifecycle
(`up/down/isRunning/docker-inspect`), assumes services already running
- **ci-native-eval.sh** — standalone helper that installs deps, starts
`next dev --turbopack` + Python agents natively, health-waits, runs
`showcase eval --ci`
- **test_e2e-showcase-on-demand.yml** — fixed and extended with
`langgraph-python` support, agent-type detection (uvicorn vs
langgraph_cli), relaxed aimock_toggle.py requirement
### Setup required (post-merge)
1. Add `checks:write` permission to the copilotkit-devops-bot GitHub App
2. Deploy eval-webhook to Railway with app secrets
3. Configure GitHub App webhook URL to point to the Railway service
## Test plan
- [x] showcase-harness: 1480 tests pass
- [ ] CI green
- [ ] Deploy eval-webhook, verify `/health` endpoint
- [ ] Open test PR, verify Check Run appears with button
- [ ] Click button, verify eval triggers and Check Run updates