DocsLayer rendered unconditionally for docs-only features whenever
any overlay was active. Now respects the Docs toggle — docs-only
cells show Links (command) when Links is on, Docs when Docs is on,
empty otherwise.
D3/D4 was blue (accent), now amber/yellow to signal "not yet D5".
D1/D2 was amber, now red to signal "needs work".
Removed D6 references — D6 no longer exists.
The docs-only early return only rendered DocsLayer, skipping LinksLayer
which is responsible for rendering CommandCell when a demo has a
command field. The npx starter command disappeared from the matrix.
The CLI Start Command feature (kind: "docs-only") was rendering empty
cells in the dashboard grid because the docs row only displayed when the
"docs" overlay was explicitly toggled on. Since docs are the only content
for docs-only features, show the docs row whenever any content-producing
overlay is active (links, depth, health, or docs), not just docs alone.
## Summary
Follow-up to #4502. With matrix cells memoized and SSE deltas coalesced,
the dashboard is much smoother — but the **initial dashboard load**
still freezes for a few seconds because:
1. **Initial PB fetch is 10 sequential round-trips.** `fetchInitial`
walks `getList(page=1, 200) → getList(page=2, 200) → …` up to 10 pages,
each awaiting the previous. Network wall time stacks.
2. **First commit with real data is unavoidably a full-matrix
re-render.** The empty-map → populated-map transition invalidates
per-key memo checks on every cell — that's 720 cell renders in one
synchronous React commit, blocking the main thread.
## Changes
- **Parallelize the initial fetch.** Pull page 1 sequentially to learn
`totalItems`, then fire pages 2..N concurrently via `Promise.all`. Wall
time drops from `sum(rtt_per_page)` to roughly `max(rtt_per_page)`
modulo network parallelism. PocketBase reads are independent so this is
safe.
- **`startTransition` around the initial `setRows(initial)`.** The first
big commit is unavoidable, but marking it as a transition lets React 19
yield to user input mid-walk instead of blocking for the entire render.
`setStatus` stays urgent so the "connecting → live" indicator still
flips immediately.
- **`startTransition` around the SSE flush.** Bursts that the 16ms
coalescer can't fully absorb (large reconnect replays, simultaneous
probe completions) still need to yield rather than block.
## Test plan
- [x] `useLiveStatus.test.tsx` 15/15 pass with parallelized fetch (the
mock handles arbitrary page numbers).
- [x] `composed-cell.test.tsx` 9/10 pass (1 pre-existing failure, same
as #4502).
- [x] No new TypeScript errors in `useLiveStatus.ts`.
- [ ] **Manual perf trace before/after on the live dashboard.**
Expected: initial-load wall time drops sharply (sequential → parallel
pages), and any remaining heavy commit no longer blocks the main thread
(transition lets React yield).
- [ ] **Functional sanity** — confirm rows still arrive correctly when
pages return out of order, and the "connecting → live" status transition
still fires before the heavy render.
## Summary
- docs-only features (e.g. cli-start) exist purely for documentation
coverage tracking and have no route, depth probes, or health signals
- The catalog metadata was counting their 18 stub cells in the headline
wired/stub/unshipped/unsupported breakdown, inflating the total and
making the stats bar misleading
- Exclude docs-only cells from the headline counts; add a new
`docs_only` field to track them separately
**Before:** total_cells=720 wired=673 stub=18 unsupported=29
**After:** total_cells=702 wired=673 stub=0 unsupported=29 docs_only=18
## Test plan
- [x] `showcase/scripts` test suite: all 574 tests pass (including 13
catalog tests)
- [x] `showcase/shell-dashboard` test suite: same 3 pre-existing
failures, no regressions
- [x] TypeScript check clean (no new errors)
- [x] Invariant holds: wired + stub + unshipped + unsupported +
docs_only == cells.length
docs-only features (e.g. cli-start) exist purely for documentation
coverage tracking -- they have no route, no depth probes, and no
health signals. The catalog metadata was counting their 18 stub cells
in the headline wired/stub/unshipped/unsupported breakdown, inflating
the total and making the stats bar misleading.
Exclude docs-only cells from the headline counts. A new docs_only
field tracks the excluded count separately so the invariant
(wired + stub + unshipped + unsupported + docs_only == cells.length)
holds.
Before: total_cells=720 wired=673 stub=18 unsupported=29
After: total_cells=702 wired=673 stub=0 unsupported=29 docs_only=18
Widen dialog from w-72 to w-[480px], extract key signal fields
(errorDesc, error, failureSummary, backendUrl) as readable key-value
pairs instead of buried JSON, make raw signal collapsible, remove
duplicate tooltip text that repeated badge color/status info.
Suppress RefDepthCell for docs-only features in the parity overlay,
showing '--' instead of a depth chip since docs-only features have no
probe coverage. Add depth distribution counts (D6..D1) to the
AdaptiveStatsBar when the depth overlay is active, giving operators
at-a-glance visibility into how many wired cells are at each depth.
Features with kind "docs-only" show only the docs row in the
dashboard -- no Dx badges, no status badges, no links or depth
layers. The row label is muted and tagged like "testing" rows.
- feature-registry.json: cli-start changed from primary to docs-only
- registry.ts: FeatureKind union extended with "docs-only"
- cell-pieces.tsx: CellStatus returns null for docs-only (hides all
badges)
- composed-cell.tsx: docs-only features render only DocsLayer, skip
links/depth/health
- feature-grid.tsx: docs-only rows use muted italic styling with a
"docs-only" tag
- Tests added for all three behavioral changes
deriveDepth() was computing D2 for not_supported_features because D1/D2
are integration-scoped. Add unsupported boolean to DepthResult, guard in
deriveDepth, update consumer components. 24/24 depth-utils tests pass.
The cell status row rendered RT (D4/e2e) and CV (D5) badges but was
missing the D2/API badge despite the legend already documenting it.
- Add `d2` field to `CellState` sourced from `agent:<slug>` rows
- Render API LiveBadge before RT in CellStatus (D2 < D4 ordering)
- Add API (Agent) to CellDrilldown DIMENSIONS array
- Fix stale FP/D6 references in tests left over from #4513
Reorder legend items so depth-related explanations (D0-D4, L1-L4 Strip,
D4 RT, D5 CV, D6 FP) are grouped together on the left, followed by the
regression indicator, color chips, and status symbols.
The DepthChip (colored badge) correctly shows D4/D5/D6 depth numbers.
The descriptive text row below each cell's chip was also showing D4/D5/D6
instead of the human-readable RT (Round Trip) / CV (Conversation) /
FP (Feature Parity) labels. Fix the LiveBadge name props in CellStatus,
the DIMENSIONS labels in CellDrilldown, and all related tests.
Partially reverts #4506: chips display depth-layer numbers (D4, D5, D6)
instead of abbreviations (RT, CV, FP). The descriptive labels Round Trip,
Conversation, and Feature Parity now appear in parentheses in the cell
drilldown dimensions and as legend explanations, keeping the abbreviations
in prose only.
- cell-pieces: chip names D4/D5/D6 (were RT/CV/FP)
- cell-drilldown: dimension labels "D4 (Round Trip)" etc.
- adaptive-legend: chip refs use D4/D5/D6, descriptive text uses RT/CV/FP
- feature-grid tooltip: D4 per feature
- tests updated to match
Rename per-cell badge chip labels for clarity:
- E2E -> RT (Round Trip): single message, full-stack response verification
- D5 -> CV (Conversation): multi-turn scripted dialogue with tool calls
- D6 -> FP (Feature Parity): cross-framework behavioral consistency check
Reorganize the adaptive legend to group chip color indicators first,
followed by descriptive labels explaining what RT, CV, and FP mean.
Update cell-drilldown dimension labels to match the new naming.
Update all affected tests to use the new label strings.
## Summary
Expands D5 (e2e-deep multi-turn) coverage for the LangGraph Python (LGP)
showcase integration from the existing 11 `D5FeatureType` literals to
**31 total** (20 new). 24 new D5 scripts cover features that previously
max'd out at D4 on the dashboard.
The plan and per-feature design lives in
[`.claude/specs/lgp-d5-coverage.md`](https://github.com/CopilotKit/CopilotKit/blob/blitz/lgp-d5-coverage-design/integration/.claude/specs/lgp-d5-coverage.md)
— start there for the bigger picture and per-feature turns sketches.
### What landed (8 commits)
- `docs:` — design plan with §1 summary, §2 new literal taxonomy, §3
per-feature plan table, §4 wave proposal, §5 open design questions, §6
anti-scope
- `S0` — scaffold 20 new `D5FeatureType` literals in
`harness/src/probes/helpers/d5-registry.ts`, mirror entries in
`d5-feature-mapping.ts` (`REGISTRY_TO_D5`) and
`shell-dashboard/src/lib/live-status.ts` (`CATALOG_TO_D5_KEY`). Adds
`beautiful-chat` and `shared-state-read` as additional registry-id
mappings to existing literals.
- `B1` chat-surface family — `chat-slots`, `chat-css`,
`prebuilt-sidebar`, `prebuilt-popup`
- `B2` platform family — `auth` (sign-out flow), `multimodal` (image+PDF
via sample buttons), `agent-config`
- `B3` frontend-tools + reasoning — `frontend-tools`,
`frontend-tools-async`, `reasoning-display` (covers
`agentic-chat-reasoning` + `reasoning-default-render`),
`tool-rendering-reasoning-chain`
- `B4` state — `shared-state-streaming`, `readonly-state-context`
- `B5` gen-UI — `gen-ui-declarative`, `gen-ui-a2ui-fixed`, `gen-ui-open`
(covers `open-gen-ui` + `open-gen-ui-advanced`), `gen-ui-agent`
- `B6` interrupt + BYOC — `interrupt-headless`, `gen-ui-interrupt`,
`byoc` (covers `byoc-hashbrown` + `byoc-json-render`)
- `F` — regenerated `aimock/d5-all.json` bundle (29 → 52 fixtures),
DOM-typing fix on chat-css probe
User-locked semantics encoded verbatim:
- **`auth`** — turn 1 chat works → click
`[data-testid=auth-sign-out-button]` → assert
`[data-testid=auth-demo-error]` or
`[data-testid=auth-demo-chat-boundary]` appears
- **`multimodal`** — click "Try with sample image/PDF" buttons → ask
"what is in this?" → assert assistant references attachment content
- **`chat-slots`** — send message → assert
`[data-testid=custom-assistant-message]` slot rendered
- **`prebuilt-sidebar`** / **`prebuilt-popup`** — surface root visible
(`.copilotKitSidebar` / `.copilotKitPopup`) → message → response inside
surface
- **`chat-customization-css`** — assert user-bubble bg contains `255, 0,
110` (hot pink) AND assistant-bubble bg `253, 224, 71` (amber)
`voice` is excluded (mic input not aimockable, deferred to a future
probe family).
## What's verified
- ✅ All 188 harness unit tests pass (30 test files including 24 new
ones)
- ✅ `pnpm typecheck` clean on `showcase/harness/`
- ✅ `aimock/d5-all.json` bundle regenerated cleanly with all 52 fixtures
- ✅ Pre-commit hooks pass (lint, format, package tests, commitlint)
- ✅ Each new D5 script self-registers via `registerD5Script()` and is
picked up by the e2e-deep driver's dynamic loader at boot
## What's NOT verified yet (verification gap — please read before
merging)
- ❌ **End-to-end run against the LGP integration was NOT executed.** A
full `showcase test langgraph-python --d5 --verbose --cycle` requires
the local docker compose to be brought up against this branch's
`aimock/d5-all.json` bundle, which would interrupt other running
services. We expect the first e2e pass to surface real issues — the most
likely failure modes per the design plan §5:
- `multimodal`: pre-fill hook for sample-button click is folded into the
assertion, which races the runner's automatic fill on turn 1. Likely
needs a runner-level `preFill` hook OR the demo to grow a query-string
attach trigger.
- `auth`: the post-sign-out 401 surfaces via `/info` refetch — if the
refetch isn't auto-triggered, the assertion will time out.
- `chat-css`: computed-style `background` shorthand resolution varies by
browser; the gradient parse may not flatten to the expected RGB
substring on all engines.
- `agent-config`: the `tone`/`expertise`/`responselength` keyword check
assumes the canned response surfaces them all — actual agent output may
differ.
- All transcript-keyword assertions are loose by design and will pass on
any response containing the keywords. They detect "no response" / "wrong
agent" but won't catch subtle behavioral regressions.
- ❌ **`cr-loop` was NOT run.** Sandbox blocks subagent file writes in
this environment, which would stall the cr-loop fix cycle. Recommend
running `cr-loop` in a fresh local session OR accepting reviewer
feedback on this PR directly.
- ❌ **8 open design questions** in `.claude/specs/lgp-d5-coverage.md` §5
— most are family-collapse vs split decisions (Q1
`beautiful-chat`/`agentic-chat`, Q3 `gen-ui-open` vs split, Q4 `byoc` vs
split, Q5 `reasoning-display` vs split). Defaults applied; flag in
review if any need reversing.
## Test plan (for reviewer / merger)
- [ ] Pull this branch locally
- [ ] `cd showcase && ./bin/showcase down && ./bin/showcase up`
(rebuilds aimock with the new `d5-all.json`)
- [ ] `./bin/showcase test langgraph-python --d5 --verbose --cycle` —
expect failures on first run; iterate on script/fixture/agent fixes
- [ ] Verify dashboard: features that were strikethrough `~~D5~~` should
now show D5 chips after the next 15-min probe tick on Railway
## Anti-scope
- D6 parity coverage is separate (post-D5).
- Other framework integrations (CrewAI, Mastra, LangGraph-TS,
Pydantic-AI) — not in this PR. The new `D5FeatureType` literals are
registry-wide; per-framework implementation is each framework's own
work.
- Voice probe family — needs a separate non-aimock probe shape,
deferred.
Follow-up to the matrix memoization PR. The initial dashboard load still
freezes for a few seconds because (a) the initial PB fetch is 10
sequential getList round-trips before any data arrives, and (b) when it
does arrive, the empty-map → populated-map transition forces every cell
in the matrix to re-render in one synchronous commit.
- Parallelize initial fetch: pull page 1 sequentially to learn
totalItems, then fire pages 2..N concurrently via Promise.all. Wall
time drops from `sum(rtt_per_page)` to roughly `max(rtt_per_page)`
modulo network parallelism. PocketBase reads are independent so this
is safe.
- Wrap initial setRows in startTransition. The first commit with real
data is unavoidably a full-matrix re-render (per-key memo checks
invalidate on every cell when the map flips from empty to populated);
marking it as a transition lets React 19 yield to user input mid-walk
instead of blocking the main thread for the duration. setStatus stays
urgent so the "connecting → live" indicator still flips immediately.
- Wrap the SSE flush setRows in startTransition for the same reason —
bursts that the 16ms coalescer can't fully absorb still need to yield.
The matrix tab can freeze the browser on load and during update bursts
because every PB SSE delta triggers a full re-render of all ~720 cells
(18 integrations × ~40 features), and PocketBase fires the subscribe
callback once per record — a probe finishing dozens of services or an
initial-state replay produces N consecutive React commits.
- ComposedCell: wrap in React.memo with a custom equality check that
compares overlays/catalogCell refs, ctx scalars, and (when the
liveStatus Map identity changes) only the 5 row keys this cell
actually reads. Because upsertByKey preserves row identity for
unchanged keys, deltas that don't touch a cell's slug/featureId
short-circuit at the memo boundary.
- useLiveStatus: buffer SSE callbacks into a per-key Map and flush via
a single 16ms setTimeout. A burst of N deltas now produces 1 React
commit instead of N. Last-write-wins per key. Buffer is cleared on
teardown and on reconnect kickoff so post-reconnect initial fetches
never land on top of stale buffered rows.
When a dashboard cell has no probe data for a depth (D5, D6), render
the dimension name with CSS line-through instead of appending a "?"
label. Strikethrough communicates "this depth doesn't exist / hasn't
run" more clearly than "D5 ?".
Changes:
- Badge component: when label is "?", render name with line-through
styling and suppress the question mark
- CellDrilldown BadgeRow: render "n/a" with line-through instead of "?"
- Update cell-drilldown test to verify strikethrough rendering
## Summary
Fixes the root causes for **30+ live-dashboard cells stuck at D2**
across 31 integration rows on https://dashboard.showcase.copilotkit.ai/.
Five surgical fixes plus one structural depth-walk fix; deploy-stale
cells (~40) are flagged separately for ops since they are not
code-fixable.
## Why these fixes ship together
A scrape of the live dashboard found 100 D2 chips. Bucketing them by
failure mode revealed the pattern is **structural + per-framework**, not
100 independent failures. Five frameworks ship genuine probe/integration
bugs we can fix in code; three frameworks are deploy-stale (Railway
image is older than the recent commit landings); one cell is a
per-feature slot-API contract bug. The dashboard itself also has a
structural bug that under-reports cells whose D5 is green when D3 has no
row yet.
## What ships
### Slot-agent fixes (Phase A/B/C of the blitz)
| Commit | What |
|---|---|
| `1de1556be` D5-green bypass for missing/red D3 in depth walk |
`deriveDepth` in
`showcase/shell-dashboard/src/components/depth-utils.ts` was
contiguous-walk-only, so 13 cells whose `d5:<slug>/<feature>` row is
green but whose `e2e:<slug>/<feature>` row is missing or red were stuck
at D2. Per the harness model, the e2e-deep driver only emits D5 after D3
+ D4 have passed at some point; D5-green is sufficient evidence. After
the D2 (agent) gate passes, `deriveDepth` now checks D5 first; if green,
achieved jumps to 5 and proceeds to D6. D1+D2 remain hard gates - D5
cannot bypass health/agent. 7 new tests pin every asymmetry. |
| `e0d89e0fe` fail loud on slug missing from registry in e2e-demos
driver | `E2eDemosResolver` previously returned `E2eDemoEntry[]`,
conflating "registry has the slug with zero demos" (legitimate brand-new
package) with "slug absent from registry" (operational fault - manifest
gap, stale mount, slug rename). Both collapsed silently into "aggregate
green / no per-feature rows" - exactly the symptom shape on the live
dashboard for several frameworks (`mastra`, `langroid`, `llamaindex`,
`langgraph-typescript`, `langgraph-fastapi`,
`ms-agent-{python,dotnet}`). Resolver now returns `{ present: boolean,
entries: E2eDemoEntry[] }`; when `present === false`, the driver emits a
synthetic red `e2e:<slug>/__missing-registry` side row + flips the
aggregate red with `errorClass: "registry-missing"`. Read-failure path
still goes through the existing green/silent branch to avoid swamping
the dashboard with N noisy red dots when the registry isn't mounted at
all. |
| `00e801209` claude-sdk-python missing openai dep crashed agent on
import | `src/agents/a2ui_dynamic.py` had a top-level `import openai`
but `openai` was never declared in `requirements.txt`. `agent_server.py`
imports `a2ui_dynamic` at module load, so `ModuleNotFoundError` cascaded
through `entrypoint.sh`'s "Agent failed to start - exiting" gate,
preventing Next.js from booting at all. Result: every `/demos/<id>`
route across the integration was unreachable - explaining all 9 red E2E
cells in the `claude-sdk-python` framework column uniformly. Two-line
fix: add `openai>=1.50.0` to `requirements.txt` and move `import openai`
to a lazy import inside `_generate_a2ui` (mirrors the existing pattern
in `agents/agent.py`). |
| `cb2ef9411` remove __future__ annotations from ag2 agent_config_agent
| Same class of bug as commit `e38fab4d6` (Apr 28) which fixed
`shared_state_read_write.py` and `subagents.py`: `from __future__ import
annotations` turns `ContextVariables` in the `@tool`-decorated
`get_current_config(context_variables: ContextVariables) -> str`
signature into a ForwardRef. Autogen's `TypeAdapter` can't resolve it,
registration raises `PydanticUserError` at module import,
`agent_server.py` fails to load, the whole AG2 integration goes
unreachable. Drop the `__future__` import. |
| `b377b78a6` render input slot in google-adk chat-slots welcome screen
| The custom welcome screen at
`src/app/demos/chat-slots/custom-welcome-screen.tsx` did not accept or
render the `input` ReactElement that `CopilotChatView` passes into the
`welcomeScreen` slot. `CopilotChatView` only mounts the chat composer
inside the welcome screen on the empty-thread path
(`CopilotChatView.tsx:301-348`), so when the empty-state welcome was
active no `<textarea>` ever reached the DOM. The e2e-readiness probe's
structural selectors all timed out, flipping the D2 chat-slots cell red.
Mirrors the canonical pattern used by 16+ other integrations (strands,
langgraph-python, claude-sdk-python). |
### cr-loop fixes (Round 1+2 review findings, applied during
convergence)
| Commit | What |
|---|---|
| `b1768f5db` return red aggregate when e2e-readiness resolver throws |
Round 1 finding (3 of 7 reviewers converged independently). When
`demosResolver(slug)` throws, the catch path was emitting a `__resolver`
red side row but `slugPresentInRegistry` defaulted to `true` and was
never flipped, so the fail-loud branch was skipped and the aggregate
fell through to green with `note: "no demos declared"`. Catch now
returns early with a red aggregate (`errorDesc: "resolver-error"`),
mirroring the missing-registry early-return shape. New C12b test pins
the aggregate state, since the existing C12 test only asserted the side
row. |
| `44037b9ce` align claude-sdk-python requirements.txt comment with
lazy-import implementation | Round 1 trivial - the comment block above
the `openai` dep claimed it was "imported at module top". The
`00e801209` fix moved the import to a lazy import; the comment was now
falsified by the diff. Updated to describe the actual two-layer
protection (lazy import as runtime safety net, requirements.txt as
authoritative dep declaration). |
| `070ecd051` empty in-band input.demos no longer suppresses
missing-registry red | Round 2 finding. The in-band fallback `if
(demos.length === 0 && Array.isArray(input.demos))` triggered even when
`input.demos === []`, forcing `slugPresentInRegistry = true` and
silently bypassing the fail-loud branch. A misconfigured probe YAML with
`demos: []` for a slug not in the registry would have rendered green.
Tightened to also require `input.demos.length > 0`. |
## Live-dashboard impact
After deploy:
- **13 cells** whose `d5:<slug>/<feature>` is already green flip to D5
immediately (purely from the depth-walk fix; no probe/agent change
required).
- **30+ cells** across `mastra` / `langroid` / `llamaindex` /
`langgraph-typescript` / `langgraph-fastapi` / `ms-agent-python` /
`ms-agent-dotnet` will surface as red `__missing-registry` rows on the
next probe tick (the previously-silent missing-registry fault becomes
loud and actionable).
- **9 cells** in `claude-sdk-python` flip green once the integration
container rebuilds with the openai dep.
- **15 cells** in `ag2` framework column should flip green once the
agent server boots cleanly (the `agent_config_agent` import crash was
preventing `agent_server.py` from loading at all).
- **1 cell** `google-adk` × `chat-slots` flips green on next deploy
(welcome screen now renders the chat composer).
## Deploy-stale cohort (NOT code-fixable, flagged for ops)
| Framework | D2 cells | Status |
|---|---|---|
| `built-in-agent` | 28 | **Needs Railway redeploy** - local `next
build` succeeds for all 53 routes; live
`https://showcase-built-in-agent-production.up.railway.app` returns HTTP
404 for every demo added in commits `e18ee1d25 e4afffd4f7abb1d07674e57483869cb5f51c7c44b71f6 53455bfeb`. Source healthy; deploy is
older than the demo-landing commits. |
| `claude-sdk-typescript` | 10 | **Needs Railway redeploy** - same
pattern: local build green, live URL returns 404 for newly-added demos
(`b9dbacd99 / 70522b1a6 / 417e3b397`). |
| `ms-agent-python` | 2 | Probable snapshot-stale - live HTML now
renders the canonical `data-testid="copilot-chat-textarea"` selector;
the readiness scan was probably running before the redeploy completed.
Should self-heal on next probe tick. |
Slack thread to Jordan with the `built-in-agent` deploy-stale evidence
already sent earlier in the run.
## Test plan
- [x] `pnpm --filter @copilotkit/showcase-harness exec vitest run
e2e-readiness` - 34/34 pass (32 pre-existing + 1 new C12b resolver-throw
aggregate-red + 1 new empty-in-band missing-registry)
- [x] `npx vitest run depth-utils` (in `showcase/shell-dashboard/`) -
30/30 pass (23 pre-existing + 7 new D5-bypass cohort)
- [x] `pnpm typecheck` on `showcase/harness` - clean
- [x] `npx tsc --noEmit` on `showcase/shell-dashboard` - touched files
(`depth-utils.ts`, `depth-utils.test.ts`) clean. Pre-existing TS errors
in `compute-tally-detail.test.tsx` are on `main`, not introduced by this
PR.
- [x] cr-loop converged on Round 3 (7 unbiased reviewers, 0 bucket (a)
findings, byte-identical verbatim prompts)
- [ ] CI green on PR HEAD
- [ ] After merge: ops triggers Railway redeploy of
`showcase-built-in-agent-production` and
`showcase-claude-sdk-typescript-production` - those 38 cells flip green
without any code change.
## Out of scope (deferred to follow-up PRs)
A 7-agent unbiased CR surfaced ~50 pre-existing concerns in files this
PR happens to touch but whose subject is distinct (bucket (d) per
cr-loop's classification rules). They have coherent, named subjects and
belong in their own PRs:
- **a2ui_dynamic.py agent observability audit** (claude-sdk-python):
broad `except Exception` leaks tracebacks to chat content; sync `openai`
call in async generator blocks the event loop; Anthropic-format messages
forwarded to OpenAI's `chat.completions.create` will 400 on the second
tool round; no max-iteration cap on the outer `while True` loop;
empty-string `ANTHROPIC_API_KEY` fallback masks misconfig;
`tool_call_id` falls back to empty string.
- **e2e-readiness probe robustness audit**: `Math.max(0, deadline -
Date.now())` returns `0` after deadline expiry, which Playwright treats
as "no timeout" instead of "fail immediately"; whitespace-only `route`
strings bypass the `config-invalid` guard; URL concatenation does not
normalize trailing/leading slashes; no abort check between `newPage()`
and `page.goto()`; `setTimeout` overflow on huge `E2E_DEMOS_TIMEOUT_MS`
env values; `__resolver` / `__missing-registry` sentinel namespacing
speculation.
- **e2e-readiness test hygiene**: the C7 "aborts mid-selector-loop
without walking all 6 selectors" test asserts behavior the
implementation no longer has (single compound selector replaced the
sequential loop); the assertion
`expect(selectorsTried.length).toBeLessThan(5)` is now trivially true.
- **depth-walk semantics audit**: D6 unreachable when `isD5Green` is
false (D6 was always gated on D5-green even before this PR; design
discussion on whether D6-green should imply D5-green - the same way
D5-green now implies D3/D4-pass).
- **agent_config_agent observability**: invalid frontend values
(`tone="snarky"`) silently coerce to defaults with no log; `logger`
instance imported but never used.
## Worktrees / cleanup
- Integration worktree at
`<repo>/.claude/worktrees/blitz-d2-to-d4-integration` retained alive
across cr-loop rounds per blitz protocol; will be removed after this PR
merges.
- 8 per-slot blitz worktrees and 2 ephemeral cr-fix worktrees were
created; 7 cleaned up successfully, 3 had Windows file-lock issues
(built-in-agent, claude-sdk-typescript, ms-agent-python) - git metadata
is gone, just stale dirs that need a manual `rm -rf`.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
historical probe services
Running panel and historical run chips now show result-specific
icons: green=checkmark, yellow=warning, red=X. Previously all
completed services showed a checkmark regardless of result, making
it look like failing tests passed. Also fix inflight tally to count
yellow (degraded) results as failures alongside red.
overlays, and probe drilldown
Replace replaceState with pushState for user-initiated navigation
(tab switches, overlay toggles, preset changes, probe selection).
Add popstate listener to sync React state on back/forward. Encode
probe detail drilldown in URL hash (#ops:probe=<id>) so the back
button closes the detail panel and returns to the ops overview.
## Summary
- Persist per-service probe results from the tracker snapshot into
PocketBase `summary.services` when a run finishes
- Dashboard historical run rows with service data show a chevron —
clicking expands a detail row with service chips matching the inflight
panel pattern
- Existing runs (pre-deploy) render without the expand affordance; only
new runs populate service data
## Test plan
- [x] Verify new runs after harness deploy populate `summary.services`
in PocketBase
- [x] Verify historical run rows with service data show expand chevron
- [x] Verify clicking chevron expands/collapses the service chip grid
- [x] Verify old runs (no service data) render without chevron as before
Persist per-service probe results from the tracker snapshot into
PocketBase summary.services when a run finishes. Dashboard runs-list
rows with service data now show a chevron — clicking expands a detail
row with a service chip grid matching the inflight panel's pattern.
Existing runs (before this deploy) won't have service data and will
render without the expand affordance.
Two bugs in DocsRow's link construction surfaced after PR #4395 + #4419:
1. **No feature-default fallback for shell_docs_path.** The component
only read `integration.docs_links.features[fid].shell_docs_path`. If
no per-integration override existed, the link rendered as 'missing'
even when the feature-registry had a valid default. The og_docs_url
path already had this fallback; shell_docs_path didn't.
Fix: same `hasOverride ? override : feature_default` chain that
og_docs_url uses. Explicit `null` overrides still mean opt-out
(rendered with the 'opt-out' tooltip variant). Missing override
inherits the feature default.
2. **Obsolete /unselected/ prefix in URL construction.** The shell-docs
IA restructure (commit c11976819) retired the `unselected/`
directory. The dashboard was still building URLs as
`<framework>/unselected<path>`. Shell-docs's route handler tolerates
the legacy shape (it strips the prefix), but the URLs are
misleading. Dropped the SHELL_UNSELECTED_PATH constant and emit
canonical `<framework><path>` URLs.
After: 0 broken (404) shell-docs links across 720 cells (was 8). 577
cells now resolve cleanly to a docs page; the remaining 143 are
features that genuinely don't have a shell-docs page yet (some have
upstream docs.copilotkit.ai pages, some don't — see follow-up ticket).
The contiguous depth walk in deriveDepth previously short-circuited at
D2 whenever the D3 (e2e) row was missing or red, even when the D5
(e2e-deep) row was green. Per the harness model the e2e-deep driver
only emits a green d5:<slug>/<featureType> row after D3 + D4 probes
have passed, so a green D5 is sufficient evidence that D3/D4 implicitly
passed at some point.
After the D2 (agent) gate now passes, deriveDepth checks D5 first. If
D5 is green it sets achieved = 5 and proceeds to the D6 check; otherwise
it falls through to the existing D3 -> D4 -> D5 walk. D1 and D2 are
still hard gates -- D5 cannot bypass health/agent.
Adds 7 new tests covering: D3-missing+D5-green, D3-red+D5-green,
D3-missing+D5-missing (unchanged D2), D1-red+D5-green (must stay D0),
D2-red+D5-green (must stay D1), D6 via D5-bypass, and isRegression
when bypass lifts cell from D2 to D5 at max_depth=5.
Resolves the 13 live-dashboard cells (8 with e2e=? + 5 with e2e=red)
that should have shown D5 but were stuck at D2.
## Summary
- When a probe is running, the summary table now computes a live tally
from `inflight.services` instead of showing stale results from the last
completed run.
- Last Run and Duration columns also update to reflect the in-progress
run's start time and elapsed duration.
- Result shows "X/Y pass — running" while inflight, with amber tone for
in-progress and red if any failures detected.
## Test plan
- Trigger a probe run, watch the summary row update in real-time as
services complete.
- Once the run finishes, verify the row switches back to showing the
completed run's final results.
## Summary
- Previous 94%/6% color-mix was invisible (white vs near-white). Changed
to 50/50 mix between `--bg-surface` and `--bg-muted` for actually
visible alternating rows.
## Test plan
- Alternating rows should have a subtle but visible tint difference.
## Summary
- Adds subtle alternating row backgrounds to the feature matrix using
`color-mix(in srgb, var(--bg-surface) 94%, var(--bg-muted))` on odd
rows. Sticky feature-name column matches the stripe. Theme-aware via CSS
variables.
## Test plan
- Visual: alternating rows should have a barely-noticeable tint
difference.
- Sticky column background should match the row stripe when scrolling
horizontally.