Commit Graph

191 Commits

Author SHA1 Message Date
Tyler Slaton 04f77586f3 style: fix formatting failures on main
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
2026-05-04 13:46:32 -07:00
Jordan Ritter 22cef3b073 Remove outer black border from showcase-dashboard favicon
Keep the black fill inside the shield shape, only make the area
outside the shield transparent. Revert og-image.png to original.
2026-05-01 23:30:00 -07:00
Jordan Ritter b1b7bb5cb8 Add Baseline tab, Coverage perf + depth fixes, dark mode toggle
Baseline tab (new):
- 33-feature x 25-partner correctness matrix from Notion Partner Hub
- E3 cell treatment: status emoji + colored letter badges (C/A/I/▶/D/T/✱)
- View/Edit toggle with commit/cancel accumulator (batch saves to PB)
- PocketBase baseline collection with public read/write, SSE live updates
- Stats bar, collapsible categories, fixed legend, 200ms CSS tooltips
- Full-row hover highlight, zebra stripes matching Coverage tab

Coverage tab fixes:
- Depth chip: green = at max achievable (not hardcoded per level)
- maxPossible computed from probe existence (CATALOG_TO_D5_KEY)
- D5 false positives: removed shared key aliases
- Stats bar derives green/amber/red from same logic as depth chips
- Hide badges for non-existent tests (was strikethrough)
- Render perf: memoize featuresByCategory, React.memo CategorySection
- Load perf: getFullList single call (was 7 sequential round-trips)

Both tabs:
- Dark/light/system theme toggle (upper right)
- Column order: Coverage first (1-18), Baseline extras (19-25)
- LangGraph naming (orchestration layer, not model router)
- Row hover highlight including sticky column
2026-05-01 22:51:23 -07:00
Jordan Ritter 92f5a0ec83 fix(showcase): only show docs in CLI Start row when Docs toggle is on
DocsLayer rendered unconditionally for docs-only features whenever
any overlay was active. Now respects the Docs toggle — docs-only
cells show Links (command) when Links is on, Docs when Docs is on,
empty otherwise.
2026-05-01 18:29:33 -07:00
Jordan Ritter dc517e505b fix(showcase): order drilldown dimensions highest to lowest
Was: API (D3), CV (D5), RT (D4), Health, Smoke
Now: CV (D5), RT (D4), API (D3), Health, Smoke
2026-05-01 00:08:06 -07:00
Jordan Ritter d0bb0eca2d fix(showcase): restore depth 6 in DepthChipProps type union
AchievedDepth includes 6 — removing it from the chip prop broke the
build. Keep 6 in the type (renders as emerald, same as D5).
2026-04-30 23:59:25 -07:00
Jordan Ritter 660aae0d8d fix(showcase): depth chip colors — D3/D4 amber, D1/D2 red, remove D6
D3/D4 was blue (accent), now amber/yellow to signal "not yet D5".
D1/D2 was amber, now red to signal "needs work".
Removed D6 references — D6 no longer exists.
2026-04-30 23:46:02 -07:00
Jordan Ritter 59f8a54c80 fix(showcase): render CLI Start command in docs-only cells
The docs-only early return only rendered DocsLayer, skipping LinksLayer
which is responsible for rendering CommandCell when a demo has a
command field. The npx starter command disappeared from the matrix.
2026-04-30 17:18:39 -07:00
Jordan Ritter 73833a368b fix(showcase): show docs links for docs-only features under any active overlay
The CLI Start Command feature (kind: "docs-only") was rendering empty
cells in the dashboard grid because the docs row only displayed when the
"docs" overlay was explicitly toggled on. Since docs are the only content
for docs-only features, show the docs row whenever any content-producing
overlay is active (links, depth, health, or docs), not just docs alone.
2026-04-30 17:04:57 -07:00
Jordan Ritter d439b61c68 Revert "perf(showcase-dashboard): yield during fetch + parallelize initial pages (#4504)"
This reverts commit d17ea911e0, reversing
changes made to a0770ea0cf.
2026-04-30 16:31:37 -07:00
Jordan Ritter d17ea911e0 perf(showcase-dashboard): yield during fetch + parallelize initial pages (#4504)
## Summary

Follow-up to #4502. With matrix cells memoized and SSE deltas coalesced,
the dashboard is much smoother — but the **initial dashboard load**
still freezes for a few seconds because:

1. **Initial PB fetch is 10 sequential round-trips.** `fetchInitial`
walks `getList(page=1, 200) → getList(page=2, 200) → …` up to 10 pages,
each awaiting the previous. Network wall time stacks.
2. **First commit with real data is unavoidably a full-matrix
re-render.** The empty-map → populated-map transition invalidates
per-key memo checks on every cell — that's 720 cell renders in one
synchronous React commit, blocking the main thread.

## Changes

- **Parallelize the initial fetch.** Pull page 1 sequentially to learn
`totalItems`, then fire pages 2..N concurrently via `Promise.all`. Wall
time drops from `sum(rtt_per_page)` to roughly `max(rtt_per_page)`
modulo network parallelism. PocketBase reads are independent so this is
safe.
- **`startTransition` around the initial `setRows(initial)`.** The first
big commit is unavoidable, but marking it as a transition lets React 19
yield to user input mid-walk instead of blocking for the entire render.
`setStatus` stays urgent so the "connecting → live" indicator still
flips immediately.
- **`startTransition` around the SSE flush.** Bursts that the 16ms
coalescer can't fully absorb (large reconnect replays, simultaneous
probe completions) still need to yield rather than block.

## Test plan

- [x] `useLiveStatus.test.tsx` 15/15 pass with parallelized fetch (the
mock handles arbitrary page numbers).
- [x] `composed-cell.test.tsx` 9/10 pass (1 pre-existing failure, same
as #4502).
- [x] No new TypeScript errors in `useLiveStatus.ts`.
- [ ] **Manual perf trace before/after on the live dashboard.**
Expected: initial-load wall time drops sharply (sequential → parallel
pages), and any remaining heavy commit no longer blocks the main thread
(transition lets React yield).
- [ ] **Functional sanity** — confirm rows still arrive correctly when
pages return out of order, and the "connecting → live" status transition
still fires before the heavy render.
2026-04-30 16:02:43 -07:00
Jordan Ritter 8e36bb8b2d fix(showcase): reorder legend — L1-L4 first, add D3, remove D0-D4 2026-04-30 14:50:39 -07:00
Jordan Ritter 0c4ad13460 fix(showcase): exclude docs-only features from catalog metadata counts (#4538)
## Summary

- docs-only features (e.g. cli-start) exist purely for documentation
coverage tracking and have no route, depth probes, or health signals
- The catalog metadata was counting their 18 stub cells in the headline
wired/stub/unshipped/unsupported breakdown, inflating the total and
making the stats bar misleading
- Exclude docs-only cells from the headline counts; add a new
`docs_only` field to track them separately

**Before:** total_cells=720 wired=673 stub=18 unsupported=29
**After:** total_cells=702 wired=673 stub=0 unsupported=29 docs_only=18

## Test plan

- [x] `showcase/scripts` test suite: all 574 tests pass (including 13
catalog tests)
- [x] `showcase/shell-dashboard` test suite: same 3 pre-existing
failures, no regressions
- [x] TypeScript check clean (no new errors)
- [x] Invariant holds: wired + stub + unshipped + unsupported +
docs_only == cells.length
2026-04-30 14:39:50 -07:00
Jordan Ritter 79c10dbce2 fix(showcase): exclude docs-only features from catalog metadata counts
docs-only features (e.g. cli-start) exist purely for documentation
coverage tracking -- they have no route, no depth probes, and no
health signals. The catalog metadata was counting their 18 stub cells
in the headline wired/stub/unshipped/unsupported breakdown, inflating
the total and making the stats bar misleading.

Exclude docs-only cells from the headline counts. A new docs_only
field tracks the excluded count separately so the invariant
(wired + stub + unshipped + unsupported + docs_only == cells.length)
holds.

Before: total_cells=720 wired=673 stub=18 unsupported=29
After:  total_cells=702 wired=673 stub=0  unsupported=29 docs_only=18
2026-04-30 14:38:18 -07:00
Jordan Ritter 9ec8b90516 fix(showcase): remove D6 from depth stats, tighten grid spacing 2026-04-30 14:23:58 -07:00
Jordan Ritter 7556b0d244 fix(showcase): improve cell drilldown dialog width, signal readability, and deduplication
Widen dialog from w-72 to w-[480px], extract key signal fields
(errorDesc, error, failureSummary, backendUrl) as readable key-value
pairs instead of buried JSON, make raw signal collapsible, remove
duplicate tooltip text that repeated badge color/status info.
2026-04-30 14:15:46 -07:00
Jordan Ritter accf9b5050 fix(showcase): hide depth chip for docs-only features, add depth distribution to stats bar
Suppress RefDepthCell for docs-only features in the parity overlay,
showing '--' instead of a depth chip since docs-only features have no
probe coverage. Add depth distribution counts (D6..D1) to the
AdaptiveStatsBar when the depth overlay is active, giving operators
at-a-glance visibility into how many wired cells are at each depth.
2026-04-30 14:02:57 -07:00
Jordan Ritter 1d817577f2 fix(showcase): add favicon and og-image to shell-dashboard 2026-04-30 13:53:57 -07:00
Jordan Ritter 03aa2f3756 feat(showcase/dashboard): add docs-only feature kind
Features with kind "docs-only" show only the docs row in the
dashboard -- no Dx badges, no status badges, no links or depth
layers. The row label is muted and tagged like "testing" rows.

- feature-registry.json: cli-start changed from primary to docs-only
- registry.ts: FeatureKind union extended with "docs-only"
- cell-pieces.tsx: CellStatus returns null for docs-only (hides all
  badges)
- composed-cell.tsx: docs-only features render only DocsLayer, skip
  links/depth/health
- feature-grid.tsx: docs-only rows use muted italic styling with a
  "docs-only" tag
- Tests added for all three behavioral changes
2026-04-30 13:49:05 -07:00
Jordan Ritter 1541cbc96f fix(showcase/dashboard): unsupported cells show indicator instead of numeric depth
deriveDepth() was computing D2 for not_supported_features because D1/D2
are integration-scoped. Add unsupported boolean to DepthResult, guard in
deriveDepth, update consumer components. 24/24 depth-utils tests pass.
2026-04-30 13:34:57 -07:00
Tyler Slaton 26e245c009 chore: run pnpm format
Signed-off-by: Tyler Slaton <tyler@copilotkit.ai>
2026-04-30 12:32:31 -07:00
Jordan Ritter d7a9a2efca feat(showcase): add D2/API badge to dashboard per-cell status row
The cell status row rendered RT (D4/e2e) and CV (D5) badges but was
missing the D2/API badge despite the legend already documenting it.

- Add `d2` field to `CellState` sourced from `agent:<slug>` rows
- Render API LiveBadge before RT in CellStatus (D2 < D4 ordering)
- Add API (Agent) to CellDrilldown DIMENSIONS array
- Fix stale FP/D6 references in tests left over from #4513
2026-04-30 10:49:06 -07:00
Jordan Ritter 60e8da1a75 fix(showcase): remove FP/D6 badge from cells and drilldown 2026-04-30 10:17:59 -07:00
Jordan Ritter 5d81f6c381 fix(showcase): reorder legend D2 before D4/D5, remove stale D6 comment 2026-04-30 10:07:49 -07:00
Jordan Ritter 37c854fb8c fix(showcase): group depth/level explanations together in dashboard legend
Reorder legend items so depth-related explanations (D0-D4, L1-L4 Strip,
D4 RT, D5 CV, D6 FP) are grouped together on the left, followed by the
regression indicator, color chips, and status symbols.
2026-04-30 10:01:19 -07:00
Jordan Ritter 472e7d1031 fix(showcase): use RT/CV/FP labels in dashboard health badge text
The DepthChip (colored badge) correctly shows D4/D5/D6 depth numbers.
The descriptive text row below each cell's chip was also showing D4/D5/D6
instead of the human-readable RT (Round Trip) / CV (Conversation) /
FP (Feature Parity) labels. Fix the LiveBadge name props in CellStatus,
the DIMENSIONS labels in CellDrilldown, and all related tests.
2026-04-30 09:47:03 -07:00
Jordan Ritter 2744b8c801 fix(showcase): use D4/D5/D6 depth numbers on chips, RT/CV/FP in descriptive text only
Partially reverts #4506: chips display depth-layer numbers (D4, D5, D6)
instead of abbreviations (RT, CV, FP). The descriptive labels Round Trip,
Conversation, and Feature Parity now appear in parentheses in the cell
drilldown dimensions and as legend explanations, keeping the abbreviations
in prose only.

- cell-pieces: chip names D4/D5/D6 (were RT/CV/FP)
- cell-drilldown: dimension labels "D4 (Round Trip)" etc.
- adaptive-legend: chip refs use D4/D5/D6, descriptive text uses RT/CV/FP
- feature-grid tooltip: D4 per feature
- tests updated to match
2026-04-30 09:25:40 -07:00
Jordan Ritter 539bb27c94 fix(showcase): rename depth chip labels E2E/D5/D6 to RT/CV/FP
Rename per-cell badge chip labels for clarity:
- E2E -> RT (Round Trip): single message, full-stack response verification
- D5 -> CV (Conversation): multi-turn scripted dialogue with tool calls
- D6 -> FP (Feature Parity): cross-framework behavioral consistency check

Reorganize the adaptive legend to group chip color indicators first,
followed by descriptive labels explaining what RT, CV, and FP mean.

Update cell-drilldown dimension labels to match the new naming.
Update all affected tests to use the new label strings.
2026-04-30 09:15:01 -07:00
Alem Tuzlak c34a28acf8 feat(showcase): D5 coverage for LangGraph Python (24 features, 20 new D5FeatureType literals) (#4491)
## Summary

Expands D5 (e2e-deep multi-turn) coverage for the LangGraph Python (LGP)
showcase integration from the existing 11 `D5FeatureType` literals to
**31 total** (20 new). 24 new D5 scripts cover features that previously
max'd out at D4 on the dashboard.

The plan and per-feature design lives in
[`.claude/specs/lgp-d5-coverage.md`](https://github.com/CopilotKit/CopilotKit/blob/blitz/lgp-d5-coverage-design/integration/.claude/specs/lgp-d5-coverage.md)
— start there for the bigger picture and per-feature turns sketches.

### What landed (8 commits)

- `docs:` — design plan with §1 summary, §2 new literal taxonomy, §3
per-feature plan table, §4 wave proposal, §5 open design questions, §6
anti-scope
- `S0` — scaffold 20 new `D5FeatureType` literals in
`harness/src/probes/helpers/d5-registry.ts`, mirror entries in
`d5-feature-mapping.ts` (`REGISTRY_TO_D5`) and
`shell-dashboard/src/lib/live-status.ts` (`CATALOG_TO_D5_KEY`). Adds
`beautiful-chat` and `shared-state-read` as additional registry-id
mappings to existing literals.
- `B1` chat-surface family — `chat-slots`, `chat-css`,
`prebuilt-sidebar`, `prebuilt-popup`
- `B2` platform family — `auth` (sign-out flow), `multimodal` (image+PDF
via sample buttons), `agent-config`
- `B3` frontend-tools + reasoning — `frontend-tools`,
`frontend-tools-async`, `reasoning-display` (covers
`agentic-chat-reasoning` + `reasoning-default-render`),
`tool-rendering-reasoning-chain`
- `B4` state — `shared-state-streaming`, `readonly-state-context`
- `B5` gen-UI — `gen-ui-declarative`, `gen-ui-a2ui-fixed`, `gen-ui-open`
(covers `open-gen-ui` + `open-gen-ui-advanced`), `gen-ui-agent`
- `B6` interrupt + BYOC — `interrupt-headless`, `gen-ui-interrupt`,
`byoc` (covers `byoc-hashbrown` + `byoc-json-render`)
- `F` — regenerated `aimock/d5-all.json` bundle (29 → 52 fixtures),
DOM-typing fix on chat-css probe

User-locked semantics encoded verbatim:
- **`auth`** — turn 1 chat works → click
`[data-testid=auth-sign-out-button]` → assert
`[data-testid=auth-demo-error]` or
`[data-testid=auth-demo-chat-boundary]` appears
- **`multimodal`** — click "Try with sample image/PDF" buttons → ask
"what is in this?" → assert assistant references attachment content
- **`chat-slots`** — send message → assert
`[data-testid=custom-assistant-message]` slot rendered
- **`prebuilt-sidebar`** / **`prebuilt-popup`** — surface root visible
(`.copilotKitSidebar` / `.copilotKitPopup`) → message → response inside
surface
- **`chat-customization-css`** — assert user-bubble bg contains `255, 0,
110` (hot pink) AND assistant-bubble bg `253, 224, 71` (amber)

`voice` is excluded (mic input not aimockable, deferred to a future
probe family).

## What's verified

- ✅ All 188 harness unit tests pass (30 test files including 24 new
ones)
- ✅ `pnpm typecheck` clean on `showcase/harness/`
- ✅ `aimock/d5-all.json` bundle regenerated cleanly with all 52 fixtures
- ✅ Pre-commit hooks pass (lint, format, package tests, commitlint)
- ✅ Each new D5 script self-registers via `registerD5Script()` and is
picked up by the e2e-deep driver's dynamic loader at boot

## What's NOT verified yet (verification gap — please read before
merging)

- ❌ **End-to-end run against the LGP integration was NOT executed.** A
full `showcase test langgraph-python --d5 --verbose --cycle` requires
the local docker compose to be brought up against this branch's
`aimock/d5-all.json` bundle, which would interrupt other running
services. We expect the first e2e pass to surface real issues — the most
likely failure modes per the design plan §5:
- `multimodal`: pre-fill hook for sample-button click is folded into the
assertion, which races the runner's automatic fill on turn 1. Likely
needs a runner-level `preFill` hook OR the demo to grow a query-string
attach trigger.
- `auth`: the post-sign-out 401 surfaces via `/info` refetch — if the
refetch isn't auto-triggered, the assertion will time out.
- `chat-css`: computed-style `background` shorthand resolution varies by
browser; the gradient parse may not flatten to the expected RGB
substring on all engines.
- `agent-config`: the `tone`/`expertise`/`responselength` keyword check
assumes the canned response surfaces them all — actual agent output may
differ.
- All transcript-keyword assertions are loose by design and will pass on
any response containing the keywords. They detect "no response" / "wrong
agent" but won't catch subtle behavioral regressions.
- ❌ **`cr-loop` was NOT run.** Sandbox blocks subagent file writes in
this environment, which would stall the cr-loop fix cycle. Recommend
running `cr-loop` in a fresh local session OR accepting reviewer
feedback on this PR directly.
- ❌ **8 open design questions** in `.claude/specs/lgp-d5-coverage.md` §5
— most are family-collapse vs split decisions (Q1
`beautiful-chat`/`agentic-chat`, Q3 `gen-ui-open` vs split, Q4 `byoc` vs
split, Q5 `reasoning-display` vs split). Defaults applied; flag in
review if any need reversing.

## Test plan (for reviewer / merger)

- [ ] Pull this branch locally
- [ ] `cd showcase && ./bin/showcase down && ./bin/showcase up`
(rebuilds aimock with the new `d5-all.json`)
- [ ] `./bin/showcase test langgraph-python --d5 --verbose --cycle` —
expect failures on first run; iterate on script/fixture/agent fixes
- [ ] Verify dashboard: features that were strikethrough `~~D5~~` should
now show D5 chips after the next 15-min probe tick on Railway

## Anti-scope

- D6 parity coverage is separate (post-D5).
- Other framework integrations (CrewAI, Mastra, LangGraph-TS,
Pydantic-AI) — not in this PR. The new `D5FeatureType` literals are
registry-wide; per-framework implementation is each framework's own
work.
- Voice probe family — needs a separate non-aimock probe shape,
deferred.
2026-04-30 17:23:10 +02:00
Alem Tuzlak fa4cc06b18 perf(showcase-dashboard): yield during fetch + parallelize initial pages
Follow-up to the matrix memoization PR. The initial dashboard load still
freezes for a few seconds because (a) the initial PB fetch is 10
sequential getList round-trips before any data arrives, and (b) when it
does arrive, the empty-map → populated-map transition forces every cell
in the matrix to re-render in one synchronous commit.

- Parallelize initial fetch: pull page 1 sequentially to learn
  totalItems, then fire pages 2..N concurrently via Promise.all. Wall
  time drops from `sum(rtt_per_page)` to roughly `max(rtt_per_page)`
  modulo network parallelism. PocketBase reads are independent so this
  is safe.
- Wrap initial setRows in startTransition. The first commit with real
  data is unavoidably a full-matrix re-render (per-key memo checks
  invalidate on every cell when the map flips from empty to populated);
  marking it as a transition lets React 19 yield to user input mid-walk
  instead of blocking the main thread for the duration. setStatus stays
  urgent so the "connecting → live" indicator still flips immediately.
- Wrap the SSE flush setRows in startTransition for the same reason —
  bursts that the 16ms coalescer can't fully absorb still need to yield.
2026-04-30 17:18:11 +02:00
github-actions[bot] 8b5c6d2d6d style: auto-fix formatting 2026-04-30 14:43:44 +00:00
Alem Tuzlak 61b8c663e3 perf(showcase-dashboard): memoize matrix cells + coalesce SSE deltas
The matrix tab can freeze the browser on load and during update bursts
because every PB SSE delta triggers a full re-render of all ~720 cells
(18 integrations × ~40 features), and PocketBase fires the subscribe
callback once per record — a probe finishing dozens of services or an
initial-state replay produces N consecutive React commits.

- ComposedCell: wrap in React.memo with a custom equality check that
  compares overlays/catalogCell refs, ctx scalars, and (when the
  liveStatus Map identity changes) only the 5 row keys this cell
  actually reads. Because upsertByKey preserves row identity for
  unchanged keys, deltas that don't touch a cell's slug/featureId
  short-circuit at the memo boundary.

- useLiveStatus: buffer SSE callbacks into a per-key Map and flush via
  a single 16ms setTimeout. A burst of N deltas now produces 1 React
  commit instead of N. Last-write-wins per key. Buffer is cleared on
  teardown and on reconnect kickoff so post-reconnect initial fetches
  never land on top of stale buffered rows.
2026-04-30 16:40:19 +02:00
Alem Tuzlak e2065dc98f feat(showcase): scaffold 20 new D5FeatureType literals for LGP coverage 2026-04-30 13:06:33 +02:00
Jordan Ritter 7e5da38471 fix(showcase): replace ? with strikethrough for unavailable probe depths
When a dashboard cell has no probe data for a depth (D5, D6), render
the dimension name with CSS line-through instead of appending a "?"
label. Strikethrough communicates "this depth doesn't exist / hasn't
run" more clearly than "D5 ?".

Changes:
- Badge component: when label is "?", render name with line-through
  styling and suppress the question mark
- CellDrilldown BadgeRow: render "n/a" with line-through instead of "?"
- Update cell-drilldown test to verify strikethrough rendering
2026-04-30 00:08:34 -07:00
Jordan Ritter 497b205d1e fix: auto-format 16 files with pre-existing oxfmt violations
These files accumulated formatting drift across recent PRs. Fixes the
format CI check on main.
2026-04-29 19:12:22 -07:00
Jordan Ritter 5a28b778cd Revert "fix(showcase): unblock D2 cells (depth walk + framework probe fixes) (#4433)"
This reverts commit 44241a89d1, reversing
changes made to 0475a74b17.
2026-04-29 13:46:31 -07:00
Jordan Ritter 44241a89d1 fix(showcase): unblock D2 cells (depth walk + framework probe fixes) (#4433)
## Summary

Fixes the root causes for **30+ live-dashboard cells stuck at D2**
across 31 integration rows on https://dashboard.showcase.copilotkit.ai/.
Five surgical fixes plus one structural depth-walk fix; deploy-stale
cells (~40) are flagged separately for ops since they are not
code-fixable.

## Why these fixes ship together

A scrape of the live dashboard found 100 D2 chips. Bucketing them by
failure mode revealed the pattern is **structural + per-framework**, not
100 independent failures. Five frameworks ship genuine probe/integration
bugs we can fix in code; three frameworks are deploy-stale (Railway
image is older than the recent commit landings); one cell is a
per-feature slot-API contract bug. The dashboard itself also has a
structural bug that under-reports cells whose D5 is green when D3 has no
row yet.

## What ships

### Slot-agent fixes (Phase A/B/C of the blitz)

| Commit | What |
|---|---|
| `1de1556be` D5-green bypass for missing/red D3 in depth walk |
`deriveDepth` in
`showcase/shell-dashboard/src/components/depth-utils.ts` was
contiguous-walk-only, so 13 cells whose `d5:<slug>/<feature>` row is
green but whose `e2e:<slug>/<feature>` row is missing or red were stuck
at D2. Per the harness model, the e2e-deep driver only emits D5 after D3
+ D4 have passed at some point; D5-green is sufficient evidence. After
the D2 (agent) gate passes, `deriveDepth` now checks D5 first; if green,
achieved jumps to 5 and proceeds to D6. D1+D2 remain hard gates - D5
cannot bypass health/agent. 7 new tests pin every asymmetry. |
| `e0d89e0fe` fail loud on slug missing from registry in e2e-demos
driver | `E2eDemosResolver` previously returned `E2eDemoEntry[]`,
conflating "registry has the slug with zero demos" (legitimate brand-new
package) with "slug absent from registry" (operational fault - manifest
gap, stale mount, slug rename). Both collapsed silently into "aggregate
green / no per-feature rows" - exactly the symptom shape on the live
dashboard for several frameworks (`mastra`, `langroid`, `llamaindex`,
`langgraph-typescript`, `langgraph-fastapi`,
`ms-agent-{python,dotnet}`). Resolver now returns `{ present: boolean,
entries: E2eDemoEntry[] }`; when `present === false`, the driver emits a
synthetic red `e2e:<slug>/__missing-registry` side row + flips the
aggregate red with `errorClass: "registry-missing"`. Read-failure path
still goes through the existing green/silent branch to avoid swamping
the dashboard with N noisy red dots when the registry isn't mounted at
all. |
| `00e801209` claude-sdk-python missing openai dep crashed agent on
import | `src/agents/a2ui_dynamic.py` had a top-level `import openai`
but `openai` was never declared in `requirements.txt`. `agent_server.py`
imports `a2ui_dynamic` at module load, so `ModuleNotFoundError` cascaded
through `entrypoint.sh`'s "Agent failed to start - exiting" gate,
preventing Next.js from booting at all. Result: every `/demos/<id>`
route across the integration was unreachable - explaining all 9 red E2E
cells in the `claude-sdk-python` framework column uniformly. Two-line
fix: add `openai>=1.50.0` to `requirements.txt` and move `import openai`
to a lazy import inside `_generate_a2ui` (mirrors the existing pattern
in `agents/agent.py`). |
| `cb2ef9411` remove __future__ annotations from ag2 agent_config_agent
| Same class of bug as commit `e38fab4d6` (Apr 28) which fixed
`shared_state_read_write.py` and `subagents.py`: `from __future__ import
annotations` turns `ContextVariables` in the `@tool`-decorated
`get_current_config(context_variables: ContextVariables) -> str`
signature into a ForwardRef. Autogen's `TypeAdapter` can't resolve it,
registration raises `PydanticUserError` at module import,
`agent_server.py` fails to load, the whole AG2 integration goes
unreachable. Drop the `__future__` import. |
| `b377b78a6` render input slot in google-adk chat-slots welcome screen
| The custom welcome screen at
`src/app/demos/chat-slots/custom-welcome-screen.tsx` did not accept or
render the `input` ReactElement that `CopilotChatView` passes into the
`welcomeScreen` slot. `CopilotChatView` only mounts the chat composer
inside the welcome screen on the empty-thread path
(`CopilotChatView.tsx:301-348`), so when the empty-state welcome was
active no `<textarea>` ever reached the DOM. The e2e-readiness probe's
structural selectors all timed out, flipping the D2 chat-slots cell red.
Mirrors the canonical pattern used by 16+ other integrations (strands,
langgraph-python, claude-sdk-python). |

### cr-loop fixes (Round 1+2 review findings, applied during
convergence)

| Commit | What |
|---|---|
| `b1768f5db` return red aggregate when e2e-readiness resolver throws |
Round 1 finding (3 of 7 reviewers converged independently). When
`demosResolver(slug)` throws, the catch path was emitting a `__resolver`
red side row but `slugPresentInRegistry` defaulted to `true` and was
never flipped, so the fail-loud branch was skipped and the aggregate
fell through to green with `note: "no demos declared"`. Catch now
returns early with a red aggregate (`errorDesc: "resolver-error"`),
mirroring the missing-registry early-return shape. New C12b test pins
the aggregate state, since the existing C12 test only asserted the side
row. |
| `44037b9ce` align claude-sdk-python requirements.txt comment with
lazy-import implementation | Round 1 trivial - the comment block above
the `openai` dep claimed it was "imported at module top". The
`00e801209` fix moved the import to a lazy import; the comment was now
falsified by the diff. Updated to describe the actual two-layer
protection (lazy import as runtime safety net, requirements.txt as
authoritative dep declaration). |
| `070ecd051` empty in-band input.demos no longer suppresses
missing-registry red | Round 2 finding. The in-band fallback `if
(demos.length === 0 && Array.isArray(input.demos))` triggered even when
`input.demos === []`, forcing `slugPresentInRegistry = true` and
silently bypassing the fail-loud branch. A misconfigured probe YAML with
`demos: []` for a slug not in the registry would have rendered green.
Tightened to also require `input.demos.length > 0`. |

## Live-dashboard impact

After deploy:

- **13 cells** whose `d5:<slug>/<feature>` is already green flip to D5
immediately (purely from the depth-walk fix; no probe/agent change
required).
- **30+ cells** across `mastra` / `langroid` / `llamaindex` /
`langgraph-typescript` / `langgraph-fastapi` / `ms-agent-python` /
`ms-agent-dotnet` will surface as red `__missing-registry` rows on the
next probe tick (the previously-silent missing-registry fault becomes
loud and actionable).
- **9 cells** in `claude-sdk-python` flip green once the integration
container rebuilds with the openai dep.
- **15 cells** in `ag2` framework column should flip green once the
agent server boots cleanly (the `agent_config_agent` import crash was
preventing `agent_server.py` from loading at all).
- **1 cell** `google-adk` × `chat-slots` flips green on next deploy
(welcome screen now renders the chat composer).

## Deploy-stale cohort (NOT code-fixable, flagged for ops)

| Framework | D2 cells | Status |
|---|---|---|
| `built-in-agent` | 28 | **Needs Railway redeploy** - local `next
build` succeeds for all 53 routes; live
`https://showcase-built-in-agent-production.up.railway.app` returns HTTP
404 for every demo added in commits `e18ee1d25 e4afffd4f 7abb1d076
74e574838 69cb5f51c 7c44b71f6 53455bfeb`. Source healthy; deploy is
older than the demo-landing commits. |
| `claude-sdk-typescript` | 10 | **Needs Railway redeploy** - same
pattern: local build green, live URL returns 404 for newly-added demos
(`b9dbacd99 / 70522b1a6 / 417e3b397`). |
| `ms-agent-python` | 2 | Probable snapshot-stale - live HTML now
renders the canonical `data-testid="copilot-chat-textarea"` selector;
the readiness scan was probably running before the redeploy completed.
Should self-heal on next probe tick. |

Slack thread to Jordan with the `built-in-agent` deploy-stale evidence
already sent earlier in the run.

## Test plan

- [x] `pnpm --filter @copilotkit/showcase-harness exec vitest run
e2e-readiness` - 34/34 pass (32 pre-existing + 1 new C12b resolver-throw
aggregate-red + 1 new empty-in-band missing-registry)
- [x] `npx vitest run depth-utils` (in `showcase/shell-dashboard/`) -
30/30 pass (23 pre-existing + 7 new D5-bypass cohort)
- [x] `pnpm typecheck` on `showcase/harness` - clean
- [x] `npx tsc --noEmit` on `showcase/shell-dashboard` - touched files
(`depth-utils.ts`, `depth-utils.test.ts`) clean. Pre-existing TS errors
in `compute-tally-detail.test.tsx` are on `main`, not introduced by this
PR.
- [x] cr-loop converged on Round 3 (7 unbiased reviewers, 0 bucket (a)
findings, byte-identical verbatim prompts)
- [ ] CI green on PR HEAD
- [ ] After merge: ops triggers Railway redeploy of
`showcase-built-in-agent-production` and
`showcase-claude-sdk-typescript-production` - those 38 cells flip green
without any code change.

## Out of scope (deferred to follow-up PRs)

A 7-agent unbiased CR surfaced ~50 pre-existing concerns in files this
PR happens to touch but whose subject is distinct (bucket (d) per
cr-loop's classification rules). They have coherent, named subjects and
belong in their own PRs:

- **a2ui_dynamic.py agent observability audit** (claude-sdk-python):
broad `except Exception` leaks tracebacks to chat content; sync `openai`
call in async generator blocks the event loop; Anthropic-format messages
forwarded to OpenAI's `chat.completions.create` will 400 on the second
tool round; no max-iteration cap on the outer `while True` loop;
empty-string `ANTHROPIC_API_KEY` fallback masks misconfig;
`tool_call_id` falls back to empty string.
- **e2e-readiness probe robustness audit**: `Math.max(0, deadline -
Date.now())` returns `0` after deadline expiry, which Playwright treats
as "no timeout" instead of "fail immediately"; whitespace-only `route`
strings bypass the `config-invalid` guard; URL concatenation does not
normalize trailing/leading slashes; no abort check between `newPage()`
and `page.goto()`; `setTimeout` overflow on huge `E2E_DEMOS_TIMEOUT_MS`
env values; `__resolver` / `__missing-registry` sentinel namespacing
speculation.
- **e2e-readiness test hygiene**: the C7 "aborts mid-selector-loop
without walking all 6 selectors" test asserts behavior the
implementation no longer has (single compound selector replaced the
sequential loop); the assertion
`expect(selectorsTried.length).toBeLessThan(5)` is now trivially true.
- **depth-walk semantics audit**: D6 unreachable when `isD5Green` is
false (D6 was always gated on D5-green even before this PR; design
discussion on whether D6-green should imply D5-green - the same way
D5-green now implies D3/D4-pass).
- **agent_config_agent observability**: invalid frontend values
(`tone="snarky"`) silently coerce to defaults with no log; `logger`
instance imported but never used.

## Worktrees / cleanup

- Integration worktree at
`<repo>/.claude/worktrees/blitz-d2-to-d4-integration` retained alive
across cr-loop rounds per blitz protocol; will be removed after this PR
merges.
- 8 per-slot blitz worktrees and 2 ephemeral cr-fix worktrees were
created; 7 cleaned up successfully, 3 had Windows file-lock issues
(built-in-agent, claude-sdk-typescript, ms-agent-python) - git metadata
is gone, just stale dirs that need a manual `rm -rf`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-04-29 13:05:40 -07:00
Jordan Ritter f69697b4c8 fix(showcase): result-aware probe icons + browser history navigation (#4438)
## Summary

- **Inflight/historical icons now differentiate pass vs fail**: Running
panel and historical run chips show result-specific icons (green=✅,
yellow=⚠️, red=❌). Previously all completed services showed ✅ regardless
of result, making failing tests look like passes.
- **Inflight tally counts yellow (degraded) as failure**: The summary
table's inflight branch now counts `result === "yellow"` services as
failures alongside red, fixing undercounted fail tallies.
- **Browser back/forward navigation works**: All tab switches, overlay
toggles, preset changes, and probe drilldowns now create browser history
entries via pushState. Added popstate listener to sync state. Probe
detail encoded in URL hash (`#ops:probe=<id>`).

## Test plan

- [x] Trigger an e2e-deep probe run with mixed results; verify inflight
panel shows ❌ for red services (not ✅)
- [x] Verify summary tally matches inflight panel counts
- [x] Switch tabs Matrix→Ops, press Back, verify returns to Matrix
- [x] Open probe detail, press Back, verify returns to Ops overview
- [x] Toggle overlays, press Back, verify previous overlay state
restores
- [x] Direct URL navigation to `#ops:probe=<id>` opens detail panel
2026-04-29 11:47:58 -07:00
Jordan Ritter 728ebae0a1 fix(showcase): result-aware icons and tally for inflight and
historical probe services

Running panel and historical run chips now show result-specific
icons: green=checkmark, yellow=warning, red=X. Previously all
completed services showed a checkmark regardless of result, making
it look like failing tests passed. Also fix inflight tally to count
yellow (degraded) results as failures alongside red.
2026-04-29 11:47:31 -07:00
Jordan Ritter b9a511560c fix(showcase): browser history navigation for dashboard tabs,
overlays, and probe drilldown

Replace replaceState with pushState for user-initiated navigation
(tab switches, overlay toggles, preset changes, probe selection).
Add popstate listener to sync React state on back/forward. Encode
probe detail drilldown in URL hash (#ops:probe=<id>) so the back
button closes the detail panel and returns to the ops overview.
2026-04-29 11:47:25 -07:00
Jordan Ritter 430fd995e0 fix(showcase): expandable historical run rows with per-service results (#4437)
## Summary

- Persist per-service probe results from the tracker snapshot into
PocketBase `summary.services` when a run finishes
- Dashboard historical run rows with service data show a chevron —
clicking expands a detail row with service chips matching the inflight
panel pattern
- Existing runs (pre-deploy) render without the expand affordance; only
new runs populate service data

## Test plan

- [x] Verify new runs after harness deploy populate `summary.services`
in PocketBase
- [x] Verify historical run rows with service data show expand chevron
- [x] Verify clicking chevron expands/collapses the service chip grid
- [x] Verify old runs (no service data) render without chevron as before
2026-04-29 11:37:21 -07:00
Jordan Ritter dfa35f4286 fix(showcase): expandable historical run rows with per-service results
Persist per-service probe results from the tracker snapshot into
PocketBase summary.services when a run finishes. Dashboard runs-list
rows with service data now show a chevron — clicking expands a detail
row with a service chip grid matching the inflight panel's pattern.

Existing runs (before this deploy) won't have service data and will
render without the expand affordance.
2026-04-29 11:36:45 -07:00
Sam Julien e20df1f86c fix(showcase/shell-dashboard): cell docs links fall back to feature defaults; drop /unselected/ prefix
Two bugs in DocsRow's link construction surfaced after PR #4395 + #4419:

1. **No feature-default fallback for shell_docs_path.** The component
   only read `integration.docs_links.features[fid].shell_docs_path`. If
   no per-integration override existed, the link rendered as 'missing'
   even when the feature-registry had a valid default. The og_docs_url
   path already had this fallback; shell_docs_path didn't.

   Fix: same `hasOverride ? override : feature_default` chain that
   og_docs_url uses. Explicit `null` overrides still mean opt-out
   (rendered with the 'opt-out' tooltip variant). Missing override
   inherits the feature default.

2. **Obsolete /unselected/ prefix in URL construction.** The shell-docs
   IA restructure (commit c11976819) retired the `unselected/`
   directory. The dashboard was still building URLs as
   `<framework>/unselected<path>`. Shell-docs's route handler tolerates
   the legacy shape (it strips the prefix), but the URLs are
   misleading. Dropped the SHELL_UNSELECTED_PATH constant and emit
   canonical `<framework><path>` URLs.

After: 0 broken (404) shell-docs links across 720 cells (was 8). 577
cells now resolve cleanly to a docs page; the remaining 143 are
features that genuinely don't have a shell-docs page yet (some have
upstream docs.copilotkit.ai pages, some don't — see follow-up ticket).
2026-04-29 09:22:02 -07:00
Alem Tuzlak 1de1556be3 fix(showcase): D5-green bypass for missing/red D3 in depth walk
The contiguous depth walk in deriveDepth previously short-circuited at
D2 whenever the D3 (e2e) row was missing or red, even when the D5
(e2e-deep) row was green. Per the harness model the e2e-deep driver
only emits a green d5:<slug>/<featureType> row after D3 + D4 probes
have passed, so a green D5 is sufficient evidence that D3/D4 implicitly
passed at some point.

After the D2 (agent) gate now passes, deriveDepth checks D5 first. If
D5 is green it sets achieved = 5 and proceeds to the D6 check; otherwise
it falls through to the existing D3 -> D4 -> D5 walk. D1 and D2 are
still hard gates -- D5 cannot bypass health/agent.

Adds 7 new tests covering: D3-missing+D5-green, D3-red+D5-green,
D3-missing+D5-missing (unchanged D2), D1-red+D5-green (must stay D0),
D2-red+D5-green (must stay D1), D6 via D5-bypass, and isRegression
when bypass lifts cell from D2 to D5 at max_depth=5.

Resolves the 13 live-dashboard cells (8 with e2e=? + 5 with e2e=red)
that should have shown D5 but were stuck at D2.
2026-04-29 17:27:00 +02:00
Jordan Ritter e5ec2066ff fix(showcase): live inflight tally in Ops probe summary table (#4429)
## Summary
- When a probe is running, the summary table now computes a live tally
from `inflight.services` instead of showing stale results from the last
completed run.
- Last Run and Duration columns also update to reflect the in-progress
run's start time and elapsed duration.
- Result shows "X/Y pass — running" while inflight, with amber tone for
in-progress and red if any failures detected.

## Test plan
- Trigger a probe run, watch the summary row update in real-time as
services complete.
- Once the run finishes, verify the row switches back to showing the
completed run's final results.
2026-04-29 08:21:20 -07:00
Jordan Ritter d7e6429c6a fix(showcase): show live inflight tally in probe summary table instead of stale last-run 2026-04-29 08:20:54 -07:00
Jordan Ritter 82ae68d5f7 fix(showcase): increase zebra stripe visibility (#4428)
## Summary
- Previous 94%/6% color-mix was invisible (white vs near-white). Changed
to 50/50 mix between `--bg-surface` and `--bg-muted` for actually
visible alternating rows.

## Test plan
- Alternating rows should have a subtle but visible tint difference.
2026-04-29 08:13:13 -07:00
Jordan Ritter e391f19c4f fix(showcase): increase zebra stripe contrast from 6% to 50% mix 2026-04-29 08:12:50 -07:00
Jordan Ritter 360cfb9a73 fix(showcase): subtle zebra striping on feature matrix (#4427)
## Summary
- Adds subtle alternating row backgrounds to the feature matrix using
`color-mix(in srgb, var(--bg-surface) 94%, var(--bg-muted))` on odd
rows. Sticky feature-name column matches the stripe. Theme-aware via CSS
variables.

## Test plan
- Visual: alternating rows should have a barely-noticeable tint
difference.
- Sticky column background should match the row stripe when scrolling
horizontally.
2026-04-29 07:57:29 -07:00
Jordan Ritter 22e72af6ab fix(showcase): subtle zebra striping on feature matrix rows 2026-04-29 07:57:03 -07:00