The aggregated multi-turn probe from #4672 hit a CopilotKit v2 quirk on
/demos/beautiful-chat: only the FIRST useComponent tool call in a
conversation paints its component. Subsequent tool calls emit (the
agent's followup content arrives) but the component never mounts.
Reproduced cleanly without any frontend tool involvement —
pie-chart turn 1 paints 5 svg circles in seconds, bar-chart turn 2
emits "Bar chart rendered above..." but paints zero recharts elements.
The runner can't sidestep this from inside one conversation without a
page.reload() between turns, which the structural Page type doesn't
expose. Splitting into per-pill scripts means each probe gets its own
browser launch — fresh page state, fresh conversation, no useComponent
ordering pollution. CATALOG_TO_D5_KEY maps `beautiful-chat` to all
listed literals; isD5Green requires every key green for the cell to
advance to D5, and per-pill failure isolation surfaces in PB row names.
Coverage in this PR (5 pills):
- beautiful-chat-toggle-theme (frontend tool, html.dark flip)
- beautiful-chat-pie-chart (controlled gen-UI useComponent)
- beautiful-chat-bar-chart (controlled gen-UI useComponent)
- beautiful-chat-search-flights (A2UI fixed-schema FlightCards)
- beautiful-chat-schedule-meeting (HITL with slot-click resolution)
All 5 verified locally end-to-end (5/5 pass against the local stack).
Out of scope, intentionally (track in follow-up):
- Excalidraw — depends on mcp.excalidraw.com reachability
- Calculator — sandboxed iframe; dup of d5-gen-ui-open
- Sales Dashboard — generate_a2ui → render_a2ui chain renders
Metric labels but Row-bound charts don't paint
recharts containers under aimock fixtures (live
pill against same fixture chain shows the
inverse symptom). Suggests aimock's
non-progressive arg streaming differs from a
live LLM in a way the A2UI binder is sensitive
to. Needs separate aimock/binder investigation.
- Task Manager — manage_todos dispatches and agent emits closing
content, but StateStreamingMiddleware's
state.todos propagation doesn't populate the
App pane TodoList through aimock — same suspected
root cause as Sales Dashboard.
Architecture details:
- _beautiful-chat-shared.ts factors DOM helpers + per-pill
assertions, mirroring _hitl-shared.ts's pattern for an extended
Page type with click() + a runtime guard
- Each fixture uses unique D5-prefixed userMessage substrings; the
multi-stage Schedule Meeting flow uses hasToolResult false→true
for round disambiguation (no toolCallId leakage since each probe
runs in its own fresh page session)
## Summary
- **Dashboard mapping fix.** `CATALOG_TO_D5_KEY` in
`showcase/shell-dashboard/src/lib/live-status.ts` was missing `voice →
["voice"]`, so `computeMaxPossible` capped the langgraph-python voice
cell at D4 even when the d5-voice probe row was green. The harness
`REGISTRY_TO_D5` already had the entry; only the dashboard mirror was
out of sync.
- **Sample-button decoupled from `/transcribe`.** The "Play sample"
button used to fetch `sample.wav` and POST it to the runtime's
transcription endpoint, which made the sample button and the mic
indistinguishable under aimock (both returned the canned transcription).
Reworked it into a synchronous static-text injector — sample button is
now a deterministic test/demo affordance, and the mic is the only path
that exercises real Whisper transcription. Synced across all 18
voice-enabled integrations. Phrase stays `"What is the weather in
Tokyo?"` so aimock's `weather in Tokyo` substring fixture still matches.
- **Probe-test parity.** Added the missing `d5-voice.test.ts` companion
(every other `d5-*.ts` script has one) — 9 tests covering registration,
`buildTurns`, `preFill` (sample-button click + textarea-poll path), and
the weather/Tokyo assertion.
- **QA + e2e cleanup** for langgraph-python: dropped the
no-longer-applicable "Transcribing…" mid-flight assertion and the `block
/demo-audio/sample.wav` error-state subsection. Other 16 integrations'
qa/e2e files follow in a parity sync PR.
## Test plan
- [x] `nx test @copilotkit/showcase-harness -- --run d5-voice` → 9/9
pass
- [x] `npm test` in `showcase/shell-dashboard` → 509/510 pass (1
skipped, 0 failed)
- [x] `nx build @copilotkit/showcase-harness` → clean
- [x] Local boot: `langgraph-cli dev` (port 8123) + `next dev` (port
3000) + dashboard (port 3002) — voice page at `/demos/voice` renders,
"Play sample" injects the canned phrase instantly, send → agent returns
weather, mic → real Whisper transcription with `OPENAI_API_KEY` set
- [ ] Reviewer: confirm the langgraph-python voice cell on the live
dashboard advances to D5 once the next d5-voice probe tick lands a green
row
The langgraph-python voice cell sat at D4 even when its d5-voice probe
row was green. Root cause: the dashboard's CATALOG_TO_D5_KEY mirror in
showcase/shell-dashboard/src/lib/live-status.ts was missing voice ->
["voice"], so computeMaxPossible capped voice at D4 regardless of probe
state. The harness REGISTRY_TO_D5 already had the entry; only the
dashboard mirror was out of sync.
Separately, the "Play sample" button used to fetch sample.wav and POST
it to /transcribe. With aimock that meant both the sample button AND
the mic returned the same canned response, which made it impossible to
demo the mic path locally without conflating the two affordances.
Reworked the button into a synchronous static-text injector
(onTranscribed(sampleText)) so:
- Sample button = deterministic test/demo affordance, no runtime calls.
- Mic = real Whisper transcription via /transcribe.
Synced across all 18 voice-enabled integrations. Phrase stays "What is
the weather in Tokyo?" so aimock's "weather in Tokyo" substring fixture
still matches.
Also adds the missing d5-voice.test.ts companion (every other d5-* probe
script has one) and trims the langgraph-python qa/voice.md + e2e steps
that depended on the now-removed async behavior.
Promotes /demos/headless-complete to its own D5 feature type so the
dashboard cell can reach D5 instead of riding on the headless-simple
probe (which was navigating to /demos/headless-simple regardless of
which catalog feature triggered it).
- New gen-ui-headless-complete D5 feature type + script that clicks
each suggestion chip via preFill and asserts the right surface
renders: WeatherCard (get_weather), StockCard (get_stock_price),
HighlightNote (frontend useComponent), Excalidraw best-effort, and
the canonical "Asia is the largest continent" text reply.
- Existing gen-ui-headless script now drives both turns by chip
click (Profile card + Largest continent) instead of typing.
- Fixtures pin narration legs with both userMessage AND toolCallId
and order them before the bare userMessage toolCall fixture —
aimock's toolCallId matcher reads the LAST tool message in the
request, but in a multi-turn probe that "last tool" stays on a
previous turn's id until a new tool runs, which would otherwise
hijack a later turn's prompt with a stale narration.
- headless-complete UserBubble + AssistantBubble now carry
data-message-role so the harness conversation runner can detect
message arrivals (mirrors the headless-simple convention).
- Mappings updated in lockstep:
- REGISTRY_TO_D5: headless-complete -> ["gen-ui-headless-complete"]
- CATALOG_TO_D5_KEY (dashboard): same.
Beautiful Chat was capped at D4 in the dashboard because it had no
dedicated D5 probe and was deliberately excluded from CATALOG_TO_D5_KEY
(commit 974494ecb stripped the freeloading "agentic-chat" alias). PR
#4668 fixed the A2UI surface rendering and added e2e tests, but those
land at the D3 tier — D5 is a separate probe with its own driver.
Changes:
- New d5-beautiful-chat probe asserts the A2UI fixed-schema FlightCard
surface renders with literal United/Delta/$349/$289 fingerprints from
the search_flights tool. 60s budget on first card, 5s on siblings.
- New harness/fixtures/d5/beautiful-chat.json with two-stage fixture
(hasToolResult false→true) mirroring the gen-ui-headless pattern.
Fixture spliced into the bundled aimock/d5-all.json.
- New "beautiful-chat" D5FeatureType literal in the registry's union +
runtime mirror.
- d5-feature-mapping.ts: replace "beautiful-chat": ["agentic-chat"]
alias with ["beautiful-chat"] so the probe targets its own dedicated
PB key instead of freeloading agentic-chat's green status.
- live-status.ts CATALOG_TO_D5_KEY: re-add "beautiful-chat":
["beautiful-chat"] so computeMaxPossible lifts the D4 cap to D5.
Two follow-ups to the parity/cell-matrix fix:
1. The live cell view (`composed-cell.tsx`, `feature-grid.tsx` →
`ref-depth-column.tsx`) had the same regression-vs-ceiling collision —
they passed `regression={isRegression}` to DepthChip and `regression`
short-circuits the chip to red regardless of `maxDepth`. Switched
them to pass `maxDepth={maxPossible}` only, mirroring the matrix fix.
`RefDepthCellProps` also flips `regression?` → `maxDepth?` to stop
propagating the dead prop.
2. `computeMaxPossible` returned 6 whenever a D5 mapping existed, on the
assumption D6 was structurally reachable. There is no per-feature D6
registry, so D6 is a stretch goal in practice and every D5 cell was
rendered amber as "below ceiling". Cap at 5; D6-green cells still
render green because `depthColorClass` treats `depth >= maxDepth`
as at-ceiling. Tests updated accordingly.
`page.tsx` healthStats lost its `isRegression` short-circuit so it now
mirrors the same graduated logic the chips use.
The matrix chips were hard-coded red at D5 because parity-matrix and
cell-matrix passed `regression={depth.isRegression}` to DepthChip but
never `maxDepth`. DepthChip short-circuits to red when `regression` is
true, regardless of depth. Since `isRegression` was redefined as
`achieved < maxPossible` and `computeMaxPossible` returns 6 whenever a
D5 mapping exists, every D5 cell without a green D6 probe (the common
case — D6 probes are rare) ended up red.
Pass `maxDepth={depth.maxPossible}` and drop the `regression` prop. This
re-engages the chip's intended graduated coloring: green at ceiling,
amber 1-2 below, red 3+ below. The `isRegression` field is still used
by the cell-matrix `filter="regressions"` row filter — that's unchanged.
Two object literals in parity-matrix.tsx ParityCategorySection were missing
the required maxPossible field on DepthResult, causing 'next build' to fail
with TS2322 and blocking the shell-dashboard production rebuild. Setting
maxPossible: 0 matches achieved: 0 for unwired cells (no probe → no ceiling),
which is consistent with how DepthChip renders an unshipped cell.
Cherry-picked from 1db2fbc7b on fix/dashboard-polish (Jordan Ritter); landing
standalone so the showcase deploy can rebuild from main without waiting on
the larger dashboard-polish branch.
DocsLayer rendered unconditionally for docs-only features whenever
any overlay was active. Now respects the Docs toggle — docs-only
cells show Links (command) when Links is on, Docs when Docs is on,
empty otherwise.
D3/D4 was blue (accent), now amber/yellow to signal "not yet D5".
D1/D2 was amber, now red to signal "needs work".
Removed D6 references — D6 no longer exists.
The docs-only early return only rendered DocsLayer, skipping LinksLayer
which is responsible for rendering CommandCell when a demo has a
command field. The npx starter command disappeared from the matrix.
The CLI Start Command feature (kind: "docs-only") was rendering empty
cells in the dashboard grid because the docs row only displayed when the
"docs" overlay was explicitly toggled on. Since docs are the only content
for docs-only features, show the docs row whenever any content-producing
overlay is active (links, depth, health, or docs), not just docs alone.
## Summary
Follow-up to #4502. With matrix cells memoized and SSE deltas coalesced,
the dashboard is much smoother — but the **initial dashboard load**
still freezes for a few seconds because:
1. **Initial PB fetch is 10 sequential round-trips.** `fetchInitial`
walks `getList(page=1, 200) → getList(page=2, 200) → …` up to 10 pages,
each awaiting the previous. Network wall time stacks.
2. **First commit with real data is unavoidably a full-matrix
re-render.** The empty-map → populated-map transition invalidates
per-key memo checks on every cell — that's 720 cell renders in one
synchronous React commit, blocking the main thread.
## Changes
- **Parallelize the initial fetch.** Pull page 1 sequentially to learn
`totalItems`, then fire pages 2..N concurrently via `Promise.all`. Wall
time drops from `sum(rtt_per_page)` to roughly `max(rtt_per_page)`
modulo network parallelism. PocketBase reads are independent so this is
safe.
- **`startTransition` around the initial `setRows(initial)`.** The first
big commit is unavoidable, but marking it as a transition lets React 19
yield to user input mid-walk instead of blocking for the entire render.
`setStatus` stays urgent so the "connecting → live" indicator still
flips immediately.
- **`startTransition` around the SSE flush.** Bursts that the 16ms
coalescer can't fully absorb (large reconnect replays, simultaneous
probe completions) still need to yield rather than block.
## Test plan
- [x] `useLiveStatus.test.tsx` 15/15 pass with parallelized fetch (the
mock handles arbitrary page numbers).
- [x] `composed-cell.test.tsx` 9/10 pass (1 pre-existing failure, same
as #4502).
- [x] No new TypeScript errors in `useLiveStatus.ts`.
- [ ] **Manual perf trace before/after on the live dashboard.**
Expected: initial-load wall time drops sharply (sequential → parallel
pages), and any remaining heavy commit no longer blocks the main thread
(transition lets React yield).
- [ ] **Functional sanity** — confirm rows still arrive correctly when
pages return out of order, and the "connecting → live" status transition
still fires before the heavy render.
## Summary
- docs-only features (e.g. cli-start) exist purely for documentation
coverage tracking and have no route, depth probes, or health signals
- The catalog metadata was counting their 18 stub cells in the headline
wired/stub/unshipped/unsupported breakdown, inflating the total and
making the stats bar misleading
- Exclude docs-only cells from the headline counts; add a new
`docs_only` field to track them separately
**Before:** total_cells=720 wired=673 stub=18 unsupported=29
**After:** total_cells=702 wired=673 stub=0 unsupported=29 docs_only=18
## Test plan
- [x] `showcase/scripts` test suite: all 574 tests pass (including 13
catalog tests)
- [x] `showcase/shell-dashboard` test suite: same 3 pre-existing
failures, no regressions
- [x] TypeScript check clean (no new errors)
- [x] Invariant holds: wired + stub + unshipped + unsupported +
docs_only == cells.length
docs-only features (e.g. cli-start) exist purely for documentation
coverage tracking -- they have no route, no depth probes, and no
health signals. The catalog metadata was counting their 18 stub cells
in the headline wired/stub/unshipped/unsupported breakdown, inflating
the total and making the stats bar misleading.
Exclude docs-only cells from the headline counts. A new docs_only
field tracks the excluded count separately so the invariant
(wired + stub + unshipped + unsupported + docs_only == cells.length)
holds.
Before: total_cells=720 wired=673 stub=18 unsupported=29
After: total_cells=702 wired=673 stub=0 unsupported=29 docs_only=18
Widen dialog from w-72 to w-[480px], extract key signal fields
(errorDesc, error, failureSummary, backendUrl) as readable key-value
pairs instead of buried JSON, make raw signal collapsible, remove
duplicate tooltip text that repeated badge color/status info.
Suppress RefDepthCell for docs-only features in the parity overlay,
showing '--' instead of a depth chip since docs-only features have no
probe coverage. Add depth distribution counts (D6..D1) to the
AdaptiveStatsBar when the depth overlay is active, giving operators
at-a-glance visibility into how many wired cells are at each depth.
Features with kind "docs-only" show only the docs row in the
dashboard -- no Dx badges, no status badges, no links or depth
layers. The row label is muted and tagged like "testing" rows.
- feature-registry.json: cli-start changed from primary to docs-only
- registry.ts: FeatureKind union extended with "docs-only"
- cell-pieces.tsx: CellStatus returns null for docs-only (hides all
badges)
- composed-cell.tsx: docs-only features render only DocsLayer, skip
links/depth/health
- feature-grid.tsx: docs-only rows use muted italic styling with a
"docs-only" tag
- Tests added for all three behavioral changes
deriveDepth() was computing D2 for not_supported_features because D1/D2
are integration-scoped. Add unsupported boolean to DepthResult, guard in
deriveDepth, update consumer components. 24/24 depth-utils tests pass.
The cell status row rendered RT (D4/e2e) and CV (D5) badges but was
missing the D2/API badge despite the legend already documenting it.
- Add `d2` field to `CellState` sourced from `agent:<slug>` rows
- Render API LiveBadge before RT in CellStatus (D2 < D4 ordering)
- Add API (Agent) to CellDrilldown DIMENSIONS array
- Fix stale FP/D6 references in tests left over from #4513
Reorder legend items so depth-related explanations (D0-D4, L1-L4 Strip,
D4 RT, D5 CV, D6 FP) are grouped together on the left, followed by the
regression indicator, color chips, and status symbols.
The DepthChip (colored badge) correctly shows D4/D5/D6 depth numbers.
The descriptive text row below each cell's chip was also showing D4/D5/D6
instead of the human-readable RT (Round Trip) / CV (Conversation) /
FP (Feature Parity) labels. Fix the LiveBadge name props in CellStatus,
the DIMENSIONS labels in CellDrilldown, and all related tests.
Partially reverts #4506: chips display depth-layer numbers (D4, D5, D6)
instead of abbreviations (RT, CV, FP). The descriptive labels Round Trip,
Conversation, and Feature Parity now appear in parentheses in the cell
drilldown dimensions and as legend explanations, keeping the abbreviations
in prose only.
- cell-pieces: chip names D4/D5/D6 (were RT/CV/FP)
- cell-drilldown: dimension labels "D4 (Round Trip)" etc.
- adaptive-legend: chip refs use D4/D5/D6, descriptive text uses RT/CV/FP
- feature-grid tooltip: D4 per feature
- tests updated to match
Rename per-cell badge chip labels for clarity:
- E2E -> RT (Round Trip): single message, full-stack response verification
- D5 -> CV (Conversation): multi-turn scripted dialogue with tool calls
- D6 -> FP (Feature Parity): cross-framework behavioral consistency check
Reorganize the adaptive legend to group chip color indicators first,
followed by descriptive labels explaining what RT, CV, and FP mean.
Update cell-drilldown dimension labels to match the new naming.
Update all affected tests to use the new label strings.
## Summary
Expands D5 (e2e-deep multi-turn) coverage for the LangGraph Python (LGP)
showcase integration from the existing 11 `D5FeatureType` literals to
**31 total** (20 new). 24 new D5 scripts cover features that previously
max'd out at D4 on the dashboard.
The plan and per-feature design lives in
[`.claude/specs/lgp-d5-coverage.md`](https://github.com/CopilotKit/CopilotKit/blob/blitz/lgp-d5-coverage-design/integration/.claude/specs/lgp-d5-coverage.md)
— start there for the bigger picture and per-feature turns sketches.
### What landed (8 commits)
- `docs:` — design plan with §1 summary, §2 new literal taxonomy, §3
per-feature plan table, §4 wave proposal, §5 open design questions, §6
anti-scope
- `S0` — scaffold 20 new `D5FeatureType` literals in
`harness/src/probes/helpers/d5-registry.ts`, mirror entries in
`d5-feature-mapping.ts` (`REGISTRY_TO_D5`) and
`shell-dashboard/src/lib/live-status.ts` (`CATALOG_TO_D5_KEY`). Adds
`beautiful-chat` and `shared-state-read` as additional registry-id
mappings to existing literals.
- `B1` chat-surface family — `chat-slots`, `chat-css`,
`prebuilt-sidebar`, `prebuilt-popup`
- `B2` platform family — `auth` (sign-out flow), `multimodal` (image+PDF
via sample buttons), `agent-config`
- `B3` frontend-tools + reasoning — `frontend-tools`,
`frontend-tools-async`, `reasoning-display` (covers
`agentic-chat-reasoning` + `reasoning-default-render`),
`tool-rendering-reasoning-chain`
- `B4` state — `shared-state-streaming`, `readonly-state-context`
- `B5` gen-UI — `gen-ui-declarative`, `gen-ui-a2ui-fixed`, `gen-ui-open`
(covers `open-gen-ui` + `open-gen-ui-advanced`), `gen-ui-agent`
- `B6` interrupt + BYOC — `interrupt-headless`, `gen-ui-interrupt`,
`byoc` (covers `byoc-hashbrown` + `byoc-json-render`)
- `F` — regenerated `aimock/d5-all.json` bundle (29 → 52 fixtures),
DOM-typing fix on chat-css probe
User-locked semantics encoded verbatim:
- **`auth`** — turn 1 chat works → click
`[data-testid=auth-sign-out-button]` → assert
`[data-testid=auth-demo-error]` or
`[data-testid=auth-demo-chat-boundary]` appears
- **`multimodal`** — click "Try with sample image/PDF" buttons → ask
"what is in this?" → assert assistant references attachment content
- **`chat-slots`** — send message → assert
`[data-testid=custom-assistant-message]` slot rendered
- **`prebuilt-sidebar`** / **`prebuilt-popup`** — surface root visible
(`.copilotKitSidebar` / `.copilotKitPopup`) → message → response inside
surface
- **`chat-customization-css`** — assert user-bubble bg contains `255, 0,
110` (hot pink) AND assistant-bubble bg `253, 224, 71` (amber)
`voice` is excluded (mic input not aimockable, deferred to a future
probe family).
## What's verified
- ✅ All 188 harness unit tests pass (30 test files including 24 new
ones)
- ✅ `pnpm typecheck` clean on `showcase/harness/`
- ✅ `aimock/d5-all.json` bundle regenerated cleanly with all 52 fixtures
- ✅ Pre-commit hooks pass (lint, format, package tests, commitlint)
- ✅ Each new D5 script self-registers via `registerD5Script()` and is
picked up by the e2e-deep driver's dynamic loader at boot
## What's NOT verified yet (verification gap — please read before
merging)
- ❌ **End-to-end run against the LGP integration was NOT executed.** A
full `showcase test langgraph-python --d5 --verbose --cycle` requires
the local docker compose to be brought up against this branch's
`aimock/d5-all.json` bundle, which would interrupt other running
services. We expect the first e2e pass to surface real issues — the most
likely failure modes per the design plan §5:
- `multimodal`: pre-fill hook for sample-button click is folded into the
assertion, which races the runner's automatic fill on turn 1. Likely
needs a runner-level `preFill` hook OR the demo to grow a query-string
attach trigger.
- `auth`: the post-sign-out 401 surfaces via `/info` refetch — if the
refetch isn't auto-triggered, the assertion will time out.
- `chat-css`: computed-style `background` shorthand resolution varies by
browser; the gradient parse may not flatten to the expected RGB
substring on all engines.
- `agent-config`: the `tone`/`expertise`/`responselength` keyword check
assumes the canned response surfaces them all — actual agent output may
differ.
- All transcript-keyword assertions are loose by design and will pass on
any response containing the keywords. They detect "no response" / "wrong
agent" but won't catch subtle behavioral regressions.
- ❌ **`cr-loop` was NOT run.** Sandbox blocks subagent file writes in
this environment, which would stall the cr-loop fix cycle. Recommend
running `cr-loop` in a fresh local session OR accepting reviewer
feedback on this PR directly.
- ❌ **8 open design questions** in `.claude/specs/lgp-d5-coverage.md` §5
— most are family-collapse vs split decisions (Q1
`beautiful-chat`/`agentic-chat`, Q3 `gen-ui-open` vs split, Q4 `byoc` vs
split, Q5 `reasoning-display` vs split). Defaults applied; flag in
review if any need reversing.
## Test plan (for reviewer / merger)
- [ ] Pull this branch locally
- [ ] `cd showcase && ./bin/showcase down && ./bin/showcase up`
(rebuilds aimock with the new `d5-all.json`)
- [ ] `./bin/showcase test langgraph-python --d5 --verbose --cycle` —
expect failures on first run; iterate on script/fixture/agent fixes
- [ ] Verify dashboard: features that were strikethrough `~~D5~~` should
now show D5 chips after the next 15-min probe tick on Railway
## Anti-scope
- D6 parity coverage is separate (post-D5).
- Other framework integrations (CrewAI, Mastra, LangGraph-TS,
Pydantic-AI) — not in this PR. The new `D5FeatureType` literals are
registry-wide; per-framework implementation is each framework's own
work.
- Voice probe family — needs a separate non-aimock probe shape,
deferred.
Follow-up to the matrix memoization PR. The initial dashboard load still
freezes for a few seconds because (a) the initial PB fetch is 10
sequential getList round-trips before any data arrives, and (b) when it
does arrive, the empty-map → populated-map transition forces every cell
in the matrix to re-render in one synchronous commit.
- Parallelize initial fetch: pull page 1 sequentially to learn
totalItems, then fire pages 2..N concurrently via Promise.all. Wall
time drops from `sum(rtt_per_page)` to roughly `max(rtt_per_page)`
modulo network parallelism. PocketBase reads are independent so this
is safe.
- Wrap initial setRows in startTransition. The first commit with real
data is unavoidably a full-matrix re-render (per-key memo checks
invalidate on every cell when the map flips from empty to populated);
marking it as a transition lets React 19 yield to user input mid-walk
instead of blocking the main thread for the duration. setStatus stays
urgent so the "connecting → live" indicator still flips immediately.
- Wrap the SSE flush setRows in startTransition for the same reason —
bursts that the 16ms coalescer can't fully absorb still need to yield.
The matrix tab can freeze the browser on load and during update bursts
because every PB SSE delta triggers a full re-render of all ~720 cells
(18 integrations × ~40 features), and PocketBase fires the subscribe
callback once per record — a probe finishing dozens of services or an
initial-state replay produces N consecutive React commits.
- ComposedCell: wrap in React.memo with a custom equality check that
compares overlays/catalogCell refs, ctx scalars, and (when the
liveStatus Map identity changes) only the 5 row keys this cell
actually reads. Because upsertByKey preserves row identity for
unchanged keys, deltas that don't touch a cell's slug/featureId
short-circuit at the memo boundary.
- useLiveStatus: buffer SSE callbacks into a per-key Map and flush via
a single 16ms setTimeout. A burst of N deltas now produces 1 React
commit instead of N. Last-write-wins per key. Buffer is cleared on
teardown and on reconnect kickoff so post-reconnect initial fetches
never land on top of stale buffered rows.
When a dashboard cell has no probe data for a depth (D5, D6), render
the dimension name with CSS line-through instead of appending a "?"
label. Strikethrough communicates "this depth doesn't exist / hasn't
run" more clearly than "D5 ?".
Changes:
- Badge component: when label is "?", render name with line-through
styling and suppress the question mark
- CellDrilldown BadgeRow: render "n/a" with line-through instead of "?"
- Update cell-drilldown test to verify strikethrough rendering
## Summary
Fixes the root causes for **30+ live-dashboard cells stuck at D2**
across 31 integration rows on https://dashboard.showcase.copilotkit.ai/.
Five surgical fixes plus one structural depth-walk fix; deploy-stale
cells (~40) are flagged separately for ops since they are not
code-fixable.
## Why these fixes ship together
A scrape of the live dashboard found 100 D2 chips. Bucketing them by
failure mode revealed the pattern is **structural + per-framework**, not
100 independent failures. Five frameworks ship genuine probe/integration
bugs we can fix in code; three frameworks are deploy-stale (Railway
image is older than the recent commit landings); one cell is a
per-feature slot-API contract bug. The dashboard itself also has a
structural bug that under-reports cells whose D5 is green when D3 has no
row yet.
## What ships
### Slot-agent fixes (Phase A/B/C of the blitz)
| Commit | What |
|---|---|
| `1de1556be` D5-green bypass for missing/red D3 in depth walk |
`deriveDepth` in
`showcase/shell-dashboard/src/components/depth-utils.ts` was
contiguous-walk-only, so 13 cells whose `d5:<slug>/<feature>` row is
green but whose `e2e:<slug>/<feature>` row is missing or red were stuck
at D2. Per the harness model, the e2e-deep driver only emits D5 after D3
+ D4 have passed at some point; D5-green is sufficient evidence. After
the D2 (agent) gate passes, `deriveDepth` now checks D5 first; if green,
achieved jumps to 5 and proceeds to D6. D1+D2 remain hard gates - D5
cannot bypass health/agent. 7 new tests pin every asymmetry. |
| `e0d89e0fe` fail loud on slug missing from registry in e2e-demos
driver | `E2eDemosResolver` previously returned `E2eDemoEntry[]`,
conflating "registry has the slug with zero demos" (legitimate brand-new
package) with "slug absent from registry" (operational fault - manifest
gap, stale mount, slug rename). Both collapsed silently into "aggregate
green / no per-feature rows" - exactly the symptom shape on the live
dashboard for several frameworks (`mastra`, `langroid`, `llamaindex`,
`langgraph-typescript`, `langgraph-fastapi`,
`ms-agent-{python,dotnet}`). Resolver now returns `{ present: boolean,
entries: E2eDemoEntry[] }`; when `present === false`, the driver emits a
synthetic red `e2e:<slug>/__missing-registry` side row + flips the
aggregate red with `errorClass: "registry-missing"`. Read-failure path
still goes through the existing green/silent branch to avoid swamping
the dashboard with N noisy red dots when the registry isn't mounted at
all. |
| `00e801209` claude-sdk-python missing openai dep crashed agent on
import | `src/agents/a2ui_dynamic.py` had a top-level `import openai`
but `openai` was never declared in `requirements.txt`. `agent_server.py`
imports `a2ui_dynamic` at module load, so `ModuleNotFoundError` cascaded
through `entrypoint.sh`'s "Agent failed to start - exiting" gate,
preventing Next.js from booting at all. Result: every `/demos/<id>`
route across the integration was unreachable - explaining all 9 red E2E
cells in the `claude-sdk-python` framework column uniformly. Two-line
fix: add `openai>=1.50.0` to `requirements.txt` and move `import openai`
to a lazy import inside `_generate_a2ui` (mirrors the existing pattern
in `agents/agent.py`). |
| `cb2ef9411` remove __future__ annotations from ag2 agent_config_agent
| Same class of bug as commit `e38fab4d6` (Apr 28) which fixed
`shared_state_read_write.py` and `subagents.py`: `from __future__ import
annotations` turns `ContextVariables` in the `@tool`-decorated
`get_current_config(context_variables: ContextVariables) -> str`
signature into a ForwardRef. Autogen's `TypeAdapter` can't resolve it,
registration raises `PydanticUserError` at module import,
`agent_server.py` fails to load, the whole AG2 integration goes
unreachable. Drop the `__future__` import. |
| `b377b78a6` render input slot in google-adk chat-slots welcome screen
| The custom welcome screen at
`src/app/demos/chat-slots/custom-welcome-screen.tsx` did not accept or
render the `input` ReactElement that `CopilotChatView` passes into the
`welcomeScreen` slot. `CopilotChatView` only mounts the chat composer
inside the welcome screen on the empty-thread path
(`CopilotChatView.tsx:301-348`), so when the empty-state welcome was
active no `<textarea>` ever reached the DOM. The e2e-readiness probe's
structural selectors all timed out, flipping the D2 chat-slots cell red.
Mirrors the canonical pattern used by 16+ other integrations (strands,
langgraph-python, claude-sdk-python). |
### cr-loop fixes (Round 1+2 review findings, applied during
convergence)
| Commit | What |
|---|---|
| `b1768f5db` return red aggregate when e2e-readiness resolver throws |
Round 1 finding (3 of 7 reviewers converged independently). When
`demosResolver(slug)` throws, the catch path was emitting a `__resolver`
red side row but `slugPresentInRegistry` defaulted to `true` and was
never flipped, so the fail-loud branch was skipped and the aggregate
fell through to green with `note: "no demos declared"`. Catch now
returns early with a red aggregate (`errorDesc: "resolver-error"`),
mirroring the missing-registry early-return shape. New C12b test pins
the aggregate state, since the existing C12 test only asserted the side
row. |
| `44037b9ce` align claude-sdk-python requirements.txt comment with
lazy-import implementation | Round 1 trivial - the comment block above
the `openai` dep claimed it was "imported at module top". The
`00e801209` fix moved the import to a lazy import; the comment was now
falsified by the diff. Updated to describe the actual two-layer
protection (lazy import as runtime safety net, requirements.txt as
authoritative dep declaration). |
| `070ecd051` empty in-band input.demos no longer suppresses
missing-registry red | Round 2 finding. The in-band fallback `if
(demos.length === 0 && Array.isArray(input.demos))` triggered even when
`input.demos === []`, forcing `slugPresentInRegistry = true` and
silently bypassing the fail-loud branch. A misconfigured probe YAML with
`demos: []` for a slug not in the registry would have rendered green.
Tightened to also require `input.demos.length > 0`. |
## Live-dashboard impact
After deploy:
- **13 cells** whose `d5:<slug>/<feature>` is already green flip to D5
immediately (purely from the depth-walk fix; no probe/agent change
required).
- **30+ cells** across `mastra` / `langroid` / `llamaindex` /
`langgraph-typescript` / `langgraph-fastapi` / `ms-agent-python` /
`ms-agent-dotnet` will surface as red `__missing-registry` rows on the
next probe tick (the previously-silent missing-registry fault becomes
loud and actionable).
- **9 cells** in `claude-sdk-python` flip green once the integration
container rebuilds with the openai dep.
- **15 cells** in `ag2` framework column should flip green once the
agent server boots cleanly (the `agent_config_agent` import crash was
preventing `agent_server.py` from loading at all).
- **1 cell** `google-adk` × `chat-slots` flips green on next deploy
(welcome screen now renders the chat composer).
## Deploy-stale cohort (NOT code-fixable, flagged for ops)
| Framework | D2 cells | Status |
|---|---|---|
| `built-in-agent` | 28 | **Needs Railway redeploy** - local `next
build` succeeds for all 53 routes; live
`https://showcase-built-in-agent-production.up.railway.app` returns HTTP
404 for every demo added in commits `e18ee1d25 e4afffd4f7abb1d07674e57483869cb5f51c7c44b71f6 53455bfeb`. Source healthy; deploy is
older than the demo-landing commits. |
| `claude-sdk-typescript` | 10 | **Needs Railway redeploy** - same
pattern: local build green, live URL returns 404 for newly-added demos
(`b9dbacd99 / 70522b1a6 / 417e3b397`). |
| `ms-agent-python` | 2 | Probable snapshot-stale - live HTML now
renders the canonical `data-testid="copilot-chat-textarea"` selector;
the readiness scan was probably running before the redeploy completed.
Should self-heal on next probe tick. |
Slack thread to Jordan with the `built-in-agent` deploy-stale evidence
already sent earlier in the run.
## Test plan
- [x] `pnpm --filter @copilotkit/showcase-harness exec vitest run
e2e-readiness` - 34/34 pass (32 pre-existing + 1 new C12b resolver-throw
aggregate-red + 1 new empty-in-band missing-registry)
- [x] `npx vitest run depth-utils` (in `showcase/shell-dashboard/`) -
30/30 pass (23 pre-existing + 7 new D5-bypass cohort)
- [x] `pnpm typecheck` on `showcase/harness` - clean
- [x] `npx tsc --noEmit` on `showcase/shell-dashboard` - touched files
(`depth-utils.ts`, `depth-utils.test.ts`) clean. Pre-existing TS errors
in `compute-tally-detail.test.tsx` are on `main`, not introduced by this
PR.
- [x] cr-loop converged on Round 3 (7 unbiased reviewers, 0 bucket (a)
findings, byte-identical verbatim prompts)
- [ ] CI green on PR HEAD
- [ ] After merge: ops triggers Railway redeploy of
`showcase-built-in-agent-production` and
`showcase-claude-sdk-typescript-production` - those 38 cells flip green
without any code change.
## Out of scope (deferred to follow-up PRs)
A 7-agent unbiased CR surfaced ~50 pre-existing concerns in files this
PR happens to touch but whose subject is distinct (bucket (d) per
cr-loop's classification rules). They have coherent, named subjects and
belong in their own PRs:
- **a2ui_dynamic.py agent observability audit** (claude-sdk-python):
broad `except Exception` leaks tracebacks to chat content; sync `openai`
call in async generator blocks the event loop; Anthropic-format messages
forwarded to OpenAI's `chat.completions.create` will 400 on the second
tool round; no max-iteration cap on the outer `while True` loop;
empty-string `ANTHROPIC_API_KEY` fallback masks misconfig;
`tool_call_id` falls back to empty string.
- **e2e-readiness probe robustness audit**: `Math.max(0, deadline -
Date.now())` returns `0` after deadline expiry, which Playwright treats
as "no timeout" instead of "fail immediately"; whitespace-only `route`
strings bypass the `config-invalid` guard; URL concatenation does not
normalize trailing/leading slashes; no abort check between `newPage()`
and `page.goto()`; `setTimeout` overflow on huge `E2E_DEMOS_TIMEOUT_MS`
env values; `__resolver` / `__missing-registry` sentinel namespacing
speculation.
- **e2e-readiness test hygiene**: the C7 "aborts mid-selector-loop
without walking all 6 selectors" test asserts behavior the
implementation no longer has (single compound selector replaced the
sequential loop); the assertion
`expect(selectorsTried.length).toBeLessThan(5)` is now trivially true.
- **depth-walk semantics audit**: D6 unreachable when `isD5Green` is
false (D6 was always gated on D5-green even before this PR; design
discussion on whether D6-green should imply D5-green - the same way
D5-green now implies D3/D4-pass).
- **agent_config_agent observability**: invalid frontend values
(`tone="snarky"`) silently coerce to defaults with no log; `logger`
instance imported but never used.
## Worktrees / cleanup
- Integration worktree at
`<repo>/.claude/worktrees/blitz-d2-to-d4-integration` retained alive
across cr-loop rounds per blitz protocol; will be removed after this PR
merges.
- 8 per-slot blitz worktrees and 2 ephemeral cr-fix worktrees were
created; 7 cleaned up successfully, 3 had Windows file-lock issues
(built-in-agent, claude-sdk-typescript, ms-agent-python) - git metadata
is gone, just stale dirs that need a manual `rm -rf`.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
historical probe services
Running panel and historical run chips now show result-specific
icons: green=checkmark, yellow=warning, red=X. Previously all
completed services showed a checkmark regardless of result, making
it look like failing tests passed. Also fix inflight tally to count
yellow (degraded) results as failures alongside red.
overlays, and probe drilldown
Replace replaceState with pushState for user-initiated navigation
(tab switches, overlay toggles, preset changes, probe selection).
Add popstate listener to sync React state on back/forward. Encode
probe detail drilldown in URL hash (#ops:probe=<id>) so the back
button closes the detail panel and returns to the ops overview.