Per-category subtitle counts overlapped (a feature can be both deprecated
and unique), overstating hidden rows. Show the distinct total instead;
per-category counts remain on the toggle badges. CR round 1 finding F2.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Tooltip said 'only one framework' but uniqueCount is frameworkCount<2,
which includes zero-framework demos. CR round 1 finding F3.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Hides single-framework (non-common) demo rows by default, mirroring the
Show deprecated toggle. Common = shipped by >=2 frameworks (demos[]).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
generate-registry.ts imports the catalog cross-join/flatten fold from
../harness/src/shared/catalog/catalog-flatten.ts, which does
`import yaml from "js-yaml"`. The generator's build/test environments did
not stage that file (or its module-resolution scope), so the fold could
not resolve.
- Dockerfiles (shell, shell-dashboard, shell-docs, shell-dojo): COPY the
shared catalog source + harness/package.json (its `"type":"module"` is
required so catalog-flatten resolves as ESM and its named exports bind)
and provide a node_modules for js-yaml resolution.
- generate-registry-pattern.test.ts (makeHarness): stage catalog-flatten.ts
and harness/package.json at the exact relative path the generator
resolves, and symlink the scripts node_modules onto the harness tree so
the ESM `import yaml from "js-yaml"` resolves.
- js-yaml + @types/js-yaml added to showcase/scripts (package.json and the
npm package-lock.json), and the root pnpm-lock.yaml regenerated to add
the matching importer entries for showcase/scripts (js-yaml >=4.1.1 via
the root override, @types/js-yaml ^4.0.9) so `pnpm install
--frozen-lockfile` stays in sync.
deriveDepth becomes a thin adapter over the shared buildCellModel engine so the
dashboard and API render from one ladder; page-stats routes through the shared
catalog input; dashboard-page per-cell render try/catch isolates a bad cell.
Turbopack has no resolve.extensionAlias parity (Next #82945), so
'next dev --turbopack' can't resolve the shared cell-model fold's
.js->.ts specifiers and fails with Can't resolve './live-status.js'.
The extensionAlias in next.config.ts (added in #5955) is honoured by
webpack, which next build already uses. Drop --turbopack from the dev
script so dev runs on webpack too and resolves the fold.
Deploy path unaffected: the Dockerfile builds with 'next build' (webpack)
and serves with 'next start' -- the dev script is never in the build or
runtime path.
PR #5952 (9a8cf615) added explicit `.js` extensions to the relative imports
inside the harness's shared cell-model fold
(showcase/harness/src/shared/cell-model/{cell-model,live-status,staleness}.ts)
— REQUIRED for the harness's pure-Node-ESM runtime and correct as-is.
But the dashboard re-exports that fold via shims
(showcase/shell-dashboard/src/lib/{cell-model,live-status,staleness,format-ts}.ts
`export * from "../../../harness/src/shared/cell-model/*"`), pulling the fold
into the dashboard's `next build`. `export *` does not rewrite the fold's
INTERNAL `.js` edges, and the dashboard's empty next.config.ts had no
extensionAlias, so webpack resolved `./live-status.js` literally, found only
the `.ts` source, and failed:
Module not found: Can't resolve './live-status.js'
Module not found: Can't resolve './staleness.js'
Module not found: Can't resolve './format-ts.js'
> Build failed because of webpack errors
Two-part fix (one coherent subject):
1. Resolution: add `webpack.resolve.extensionAlias` to
showcase/shell-dashboard/next.config.ts so `.js`/`.mjs` specifiers resolve
to `.ts`/`.tsx`/`.mts` sources — the bundler complement to TS NodeNext's
`.js`-import convention. Covers the `next build` (webpack) path CI uses.
The harness fold's `.js` imports are left untouched (they are correct).
2. CI gap: the dashboard build did not run on #5952 because the build matrix
is path-filtered and #5952 only touched `showcase/harness/**`, which
selects `showcase_harness` but not `shell_dashboard`. Add
`showcase/harness/src/shared/**` to the `shell_dashboard` paths-filter so
any change to the shared fold the dashboard compiles in also selects the
dashboard build — a fold change can never again ship an unbuilt dashboard.
Local red-green proof:
- RED (main, before fix): `next build` in showcase/shell-dashboard emitted the
4 fold-resolve errors above.
- GREEN (after extensionAlias): same build → 0 fold-resolve errors; the fold
resolves. Remaining `@/data/*.json` errors are the prebuild-generated files
(generate-registry/probe-docs) skipped in the local repro, produced in CI's
Docker build — unrelated to this fix.
The shell-dashboard Docker build context (repo root) copies scripts,
shared, integrations, shell-docs content, and shell-dashboard — but not
showcase/harness. The dashboard's src/lib/{cell-model,live-status,
staleness,format-ts}.ts re-export barrels forward to
../../../harness/src/shared/cell-model/*, so the isolated next build
failed with Module-not-found (CI build-check (shell-dashboard)). Local
next build passed only because the full monorepo is present.
Copy just the 4-file cell-model fold to the exact relative path the
shims expect, preserving the single-shared-fold invariant (harness
monitor still imports the same canonical copy) without dragging the
whole harness package into the dashboard image.
Bucket-(a) fixes, each with a local red-green test:
- A1: replace the substring `:${slug}` onset match with an anchored
exact slug-segment match (`keyBelongsToSlug`) so a prefix-colliding
sibling (`strands` vs `strands-typescript`) no longer mis-attributes
the earlier sibling's red onset. RED: strands' sinceAt was pulled to
the strands-typescript onset; GREEN: each slug gets its own onset.
- A2: recovery/CLOSE is now SYMMETRIC with OPEN — a recovery requires a
second agreeing fresh-healthy read (confirm scan). RED: a single
transient healthy read fired a false "recovered"; GREEN: held until
two reads agree.
- A3: guard an empty/degenerate schedule set — longestPeriodMs 0/NaN
would make idleWindowMs 0 → isProducerLive permanently false → the
monitor SUSPENDS forever and never pages. Falls back to a 45m default
window (DEFAULT_IDLE_WINDOW_MS) and logs at error.
- A4: bound readStatusRows — guard NaN/undefined totalPages (a `page >=
NaN` break never trips) and add a hard MAX_STATUS_PAGES cap so a full
page + bad totalPages cannot infinite-loop/OOM. RED: OOM; GREEN:
terminates at the cap.
- A5: a registry-load failure logs at error with a stable errorId (not a
silent warn-once permanent no-op), and the monitor accepts a loader
thunk so it re-reads registry.json each tick while the wired-cell set
is empty — a transiently-missing file self-heals without a redeploy.
- A6 (verified, no code change): createSlackWebhookTarget already throws
on every non-2xx (4xx/5xx/429/3xx/network-exhausted); added a test
asserting the monitor does NOT delete recovery state when the post
throws.
Bucket-(b): stamp the recovery message with the confirm-scan instant
(evidenceMs) not tick-start; wrap the scheduler tick handler in a
catch+errorId; log the prod env-gate skip at warn with a reason; add a
clarifying comment that the aggregate lastAlertAt reset is intentional
one-message-one-clock cadence; simplify the three dashboard barrel-shim
comments (drop the rot-prone enumerated symbol lists).
Add a harness-native monitor that pages #oss-alerts when a whole
integration column collapses to red-D0 ("completely gone" / backend
unreachable) in production — the incident class the per-cell alert rules
miss (LGT went fully gone on 2026-07-13 and nothing paged).
Detection runs the dashboard's OWN buildCellModel fold (the shared
cell-model module both the dashboard and the monitor import) over the
same PocketBase status rows and applies a column-gone predicate over the
resulting CellModel fields, so the monitor's verdict equals the DepthChip
the dashboard renders by construction — no parallel re-derivation.
- d0-gone-predicate.ts: pure cellGone/columnGone/columnFreshHealthy over
buildCellModel outputs + registry-derived wired-cell enumeration
(mirrors the dashboard page-stats iteration / determineCellStatus rule).
- d0-gone-monitor.ts: createD0GoneMonitor factory — producer-liveness
SUSPENDED gate (reuses the family-silence inflight-aware /api/runs
reasoning, 3x-longest-period idle window), 60s confirm re-read (never a
re-probe), 15m-detect vs 1h-repost state machine, positive-fresh-healthy
CLOSE gate, ONE aggregated outage / consolidated recovery Slack message,
durable per-slug JSON map in alert_state (getSet/putSet).
- orchestrator.ts: register internal:prod-d0-gone-monitor @ */15, gated on
SHOWCASE_ENV ?? RAILWAY_ENVIRONMENT_NAME === production + kill-switch,
control-plane-only (inside runControlPlane), reusing the oss_alerts
webhook target + shared memoized family summary.
- unified-cell.test.tsx: add the required isStaleCell/observedAtAgeMs
fields to the CellModel test literal (Phase-1 dashboard tsc gate).
Red-green: a frozen test-only naiveGone (achievedDepth===0 alone)
mislabels gray-D0-no-data and stale columns as gone on committed
fixtures (RED); the real predicate fires only on red-D0-fresh and matches
buildCellModel's own outputs (GREEN). Producer-idle SUSPENDED proven
load-bearing (disabling the gate flips both F1 tests red). Plus
confirm-scan blip-rejection, hourly dedup, recovery-clear, failure modes,
and the prod-only/kill-switch registration gate.
Move the pure cell-classification fold cluster (cell-model, live-status,
staleness, format-ts) out of showcase/shell-dashboard/src/lib/ into
showcase/harness/src/shared/cell-model/ so BOTH the dashboard and a new
harness monitor import ONE copy with zero duplication and no behavior change.
The harness builds via tsc -p tsconfig.build.json with rootDir:"src" and
cannot import outside its own src/, so the harness is the correct library
home. The dashboard consumes the cluster via relative path across the package
boundary (established precedent, e.g. d5-cadence-banner.redgreen.test.ts).
- git mv the four files into harness shared/cell-model/; their intra-cluster
relative imports stay valid (they move together, no external coupling).
- Replace the four original shell-dashboard paths with thin export-* barrels
so all ~51 existing dashboard import sites resolve unchanged.
- Repoint commError-contract-drift.test.ts's source-text drift parse at the
new canonical harness location (the barrels carry no derivation body).
- Add cell-model.equivalence.test.ts + committed fixtures + a pre-move
baseline JSON (generated from the original git-HEAD code) proving the move
is byte-identical across a red-D0, gray no-data, stale, mixed, all-green,
and unsupported column.
Wire the starter-smoke probe's keyed errorClass into the dashboard cell-state
flip logic so transient SOFT failures (transport-error / aborted) get two-miss
tolerance: a single soft miss renders amber ~ ("transient, not yet actionable")
instead of flapping the cell red, and only flips red on a second consecutive
miss. HARD failures (smoke-failed) and untagged reds flip immediately.
The flip gate reuses the producer-maintained fail_count (the persisted
consecutive-red counter: 1 on green->red, incremented on sustained red, 0 on
red->green) so the dashboard stays a pure function of the current row — no
dashboard-side counter to thread or reset. Tolerance is applied as a
state->degraded downgrade in buildStarterBadge (same pattern as the existing
stale-green fold), keeping the change additive and self-contained.
Adds STARTER_FAILURE_CLASSES as a dashboard-side mirror of the harness
StarterFailureClass union (the dashboard imports only @/*), guarded by a new
starter-error-class-drift.test.ts set-equality lint against the harness source.
A normal harness deploy rebuilds the shared showcase-harness image and
bounces the pool workers (PR #5715). Immediately after the bounce the
workers re-register, the producers re-arm, and every family is mid-sweep:
lastSuccessAt still points at the pre-bounce success, so it reads stale
against the silence thresholds (banner 2x period, Slack alert 3x period +
3 consecutive ticks). The result was a FALSE "worker family X has not
completed successfully" banner AND Slack family-silence alert during the
expected post-bounce drain window.
Fix: a bounce-keyed grace window. The freshest worker registered_at across
the /api/runs workers strip is the fleet's most-recent bounce instant
(independent of CP boot — a worker can bounce while the CP stays up). While
now - bounce < 2 x period, a family with no success yet is DRAINING, not
silent, so neither the §7.4 banner, the §7.3 cell glyph, nor the §9 Slack
alert flags it. Beyond the window with still-no-success, genuine silence
fires exactly as before.
The determination lives in two surfaces (server monitor for the Slack
alert; client isFamilySilent for the banner + glyph), so both now consume
the same new SSOT field (WorkerView.registeredAt) and the same 2x-period
grace constant, keeping them consistent.
- run-view.ts: project registered_at -> WorkerView.registeredAt (server)
- family-silence-monitor.ts: BOUNCE_GRACE_PERIOD_MULTIPLIER + freshest-bounce
grace gate, keyed off body.workers
- worker-runs-context.tsx: freshestBounceMs + bounceAtMs grace arg on
isFamilySilent; banner + cell glyph pass it
- ops-api.ts: WorkerView.registeredAt on the client DTO
A GREEN coverage cell flapped green -> grey -> green every time its probe
job's worker lease lapsed and the control-plane sweeper re-queued the job
(worker-reclaimed-pending). The data layer already preserves the
last-known-good colour in chipColor while surfaceState flips to "pending",
but depth-chip.tsx's `if (pending)` branch did a destructive grey
early-return that never read chipColor.
Split the pending branch:
- prior-good (chipColor is a real colour, or depth > 0): render the normal
coloured chip via the default branch's exact ternary (threading regression
through) PLUS a non-destructive refreshing affordance -- a corner ⟳ glyph
and a subtle pulsing ring in the chip's own colour, with
data-refreshing="true" / data-has-prior="true". Colour preserved.
- no-prior (never-run / first load: gray + depth 0): keep today's honest grey
⟳ chip, now tagged data-has-prior="false".
The unreachable red ⚡ overlay and the "failure never masked" gate are
untouched; red/regression cells pass through with no spinner. The refreshing
cue is conveyed by shape (⟳) + motion (ring), never by colour alone, and the
pulse/spin respect prefers-reduced-motion (motion-reduce:animate-none), so the
static ⟳ carries the meaning for colour-blind and reduced-motion operators.
Renderer-only change; unified-cell.tsx is the sole caller passing `pending`.
Add a d5-a2ui-recovery probe so the A2UI error-recovery demo runs on every
PR via the d5/d6 fleet harness, not only the manual on-demand workflow.
- New probe d5-a2ui-recovery.ts drives both pills in one session: HEAL
asserts >=2 newly-mounted declarative-metric tiles and no hard-failure
card; EXHAUST asserts the "Couldn't generate the UI" card appears and
no surface paints. Deltas (vs a pre-send baseline) keep the two
mutually-exclusive negatives correct across the shared session. The
transient "Retrying..." label is not asserted (timing-flaky).
- Prompts are sent as typed input, keyed per integration slug, mirroring
each slug's suggestions.ts message verbatim. The recovery prompts are
unique per slug because the inner render_a2ui calls carry no
x-aimock-context; a typed message is byte-identical to the pill
dispatch, so it matches the same fixture. Sending via input (not a
preFill pill click) lets the runner snapshot its run-lifecycle baseline
first, avoiding a false done-signal-missing failure.
- Register a2ui-recovery in d5-registry, map it in d5-feature-mapping,
add its representative fixture, and mirror the mapping in the dashboard
CATALOG_TO_D5_KEY (kept in lock-step via the drift test).
Verified green locally on both recovery paths: langgraph-python
(backend-owned get_a2ui_tools) and strands (auto-inject middleware).
Add strands-typescript to BASELINE_PARTNERS so it gets its own coverage
column alongside its Python sibling (mirroring how langgraph-typescript sits
beside langgraph-python). Bump the partner-count assertion 26 -> 27.
Add a new node/TypeScript-backed AWS Strands showcase integration at
showcase/integrations/strands-typescript.
Backend: a node/TS agent server (src/agent/) built on @strands-agents/sdk
`Agent`/`tool` wrapped in @ag-ui/aws-strands `StrandsAgent` and served via
@ag-ui/aws-strands/server (`createStrandsApp`/`addStrandsExpressEndpoint`),
modeled on the upstream ag-ui aws-strands TS example server and the
langgraph-typescript infra. A single shared agent at "/" serves most demos
(tools, shared state via toolBehaviors/stateContextBuilder, HITL,
sub-agents), with tool-free specialized agents mounted at /voice,
/byoc-hashbrown, /byoc-json-render. model-factory targets OpenAI chat
completions and honors OPENAI_API_KEY / OPENAI_BASE_URL so it works behind
the showcase aimock proxy. Node-based Dockerfile + entrypoint run the agent
server (:8000) alongside the Next.js frontend.
Frontend mirrors the strands (Python) sibling's demo set and the
langgraph-typescript conventions, with HttpAgent routes proxying to the TS
agent server.
Scope: base integration + standard demos only. A2UI / declarative-gen-ui /
a2ui-fixed-schema is intentionally excluded (no A2UI agents, routes, demos,
or deps) and layered on later.
Platform wiring (mirrors langgraph-typescript): docker-compose local/dev
services on host port 3119, local-ports.json, packages.json, slug-map.ts
(born-in-showcase), showcase_build.yml matrix + path filter + metadata,
shell-docs/dashboard registries, and a logo asset. The python strands
integration is untouched.
The ms-agent-harness-dotnet slug was excluded from per-cell D6/BE/smoke
probe enumeration by a placeholder fence added 2026-06-07, before the
real column existed. The column shipped in PR #5569 and its d6/d4 aimock
fixtures landed on main today (e10df0b4), so the fence is now stale.
Remove the slug from all 8 exclude SSOT sites so the column populates.
The internal counter vars pctBeAgent and totalBeAgent were introduced in
PR #5498 when the display label happened to be "BE (Agent)". The label
has since changed (#5505 made it "API (HTTP)"). Coupling internal variable
identifiers to display labels is fragile — labels move; the underlying
Status enum string "wired" is a persisted contract that does not.
Revert the var identifiers to mirror the enum:
pctBeAgent -> pctWired (coverage-bar.tsx)
totalBeAgent -> totalWired (cells-view.tsx, parity-view.tsx)
Display labels ("API (HTTP)", etc.) are unchanged — this is internal
identifiers only.
PR #5498 renamed the L1 row "Wired" → "BE (Agent)" for the D2 agent-liveness
dimension (transport/Railway up). PR #5503 renamed the per-cell D4 badge
"RT" → "BE", whose long form was already "BE (Round Trip)". The result was a
same-label-different-concept collision: "BE (Agent)" referred to D2 in some
places while the D4 per-cell badge used "BE (Round Trip)", and the per-cell
D2 badge separately used "API (Agent)" — three names for two concepts.
Normalize the long-form labels so each user-facing layer has exactly one name:
D2 (transport / Railway up, HTTP-reachable) = "API (HTTP)"
D4 (agent chat round-trip, end-to-end) = "BE (Agent)"
That cleanly distinguishes by layer: HTTP transport vs Agent message handling.
Sites touched (visible labels + matching legend prose / test assertions only):
- stats-bar.tsx, adaptive-stats-bar.tsx — wired count label
- filter-chips.tsx — chip label (id "wired" preserved)
- packages-section.tsx + .test.tsx — UWCT legend mnemonic B → A
- level-strip.tsx + .test.tsx — agent-dimension badge label (and
derived first letter B → A)
- cell-drilldown.tsx + .test.tsx +
cell-drilldown.lazy-signal.test.tsx — D4 label and D2 label, plus
testid-derivation drift
- adaptive-legend.tsx — D2 / D4 prose
Stable contracts preserved (NOT changed):
- keyFor("agent" | "d2" | "d4", …) and the "agent" LiveDimension value
- Filter-chip id "wired"
- Variable names (`wired`, etc.) — internal; pure label rename is in scope
- Status enum values "wired" | "stub" | "unshipped" | "unsupported"
- Driver names, probe registry keys, harness API contracts
Verified: typecheck clean, 1089 pass / 1 skip / 0 fail, build clean.
The per-cell health badge previously labelled "CV" (for "Conversation") is
renamed to "1P" — Single Pill. The new label tells operators what the badge
covers in scope terms (one pill out of N), which is the actual contrast the
D5/D6 ladder draws: D5 driver is `d5-single-pill.ts` (one canonical scripted
conversation), D6 driver is `d6-all-pills.ts` (the full suite). "CV" was
opaque — operators had to remember what "Conversation" meant and how it
differed from D6's full run. "1P vs D6 all-pills" is self-describing.
Mirrors PR #5503's RT → BE pass: only the badge LABEL changes; every stable
contract is preserved.
Stable contracts preserved:
- Dimension level identifier `model.d5` / `cell.d5` / `level={model.d5}`
- Drilldown dimension key `keyFor("d5", ...)` / `key: "d5"`
- Probe registry key `d5:<slug>/<featureId>`
- PocketBase row keys unchanged
- LiveDimension union / Status enum strings unchanged
- Driver file names `d5-single-pill.ts`, `d6-all-pills.ts`
- `e2e-deep` producer name unchanged (separate from CV → 1P label)
Updates:
- Source: badge label in `unified-cell.tsx`, `cell-pieces.tsx`, drilldown
label in `cell-drilldown.tsx` ("CV (Conversation)" → "1P (Single Pill)"),
legend text in `adaptive-legend.tsx`, comment refs in `composed-cell.tsx`,
`cell-model.ts`, `page-stats.ts`, `depth-utils.ts`, `live-status.ts`.
- Tests: testid strings (`mock-badge-CV` → `mock-badge-1P`), label-text
assertions, type unions, and comment refs in the affected component +
drilldown + integration + lib tests.
Verification: typecheck clean, `npm test` = 1089 pass / 1 skip / 0 fail
(same shape as #5503), `npm run build` clean, no remaining `\bCV\b` in
`showcase/shell-dashboard/src/`.
The testid generator in stats-bar.tsx used `label.toLowerCase()` directly,
producing fragile selectors with spaces and parens for compound labels.
After this PR's rename of "Wired" to "BE (Agent)", the testid became
`stat-be (agent)` — a CSS-hostile selector. A pre-existing
`stat-max depth` (from "Max Depth") had the same problem but predated
this PR.
Fix: replace with a proper slugifier:
label.toLowerCase().replace(/[^a-z0-9]+/g, "-").replace(/(^-|-$)/g, "")
Resulting testids:
"BE (Agent)" -> stat-be-agent (was stat-be (agent), broken)
"Max Depth" -> stat-max-depth (was stat-max depth, pre-existing fix)
"Stub" -> stat-stub (unchanged)
"Unshipped" -> stat-unshipped (unchanged)
"Unsupported" -> stat-unsupported (unchanged)
"Regressions" -> stat-regressions (unchanged)
"Failures" -> stat-failures (unchanged)
Call-site enumeration receipt (mandatory):
rg 'stat-wired|stat-be|stat-max|stat-stub|stat-unshipped|stat-unsupported|stat-regressions|stat-failures' showcase/shell-dashboard/
-> no matches (zero hardcoded references anywhere)
rg 'data-testid=.stat-|"stat-|`stat-' showcase/shell-dashboard/
-> only one hit: stats-bar.tsx:29 (the generator itself)
rg 'stat-' showcase/shell-dashboard/tests/
-> no matches (no Playwright/visual tests depend on these testids)
No test updates required.
Verification:
npm run typecheck -> clean
npm test -> 1088 pass / 1 pre-existing unrelated failure
(useLiveStatus.test.tsx R5 F5.2 — does not
reference stats-bar, exists on PR HEAD)
npm run build -> clean (Next.js lint inline; no lint script)
NUL byte sweep -> empty
Renames the internal display-counter variables that fed the
stats-bar / coverage-bar "Wired" labels to match the new "BE (Agent)"
taxonomy: totalWired → totalBeAgent in cells-view.tsx and
parity-view.tsx, pctWired → pctBeAgent in coverage-bar.tsx. These
are local render-time aggregates, not part of any persisted
contract.
The status enum literal "wired" (the .filter((c) => c.status ===
"wired") guard, the coverage-segment-wired test ID, and the
wiredByCategory Map's local name) is intentionally preserved — those
all key directly off the persisted catalog Status enum and changing
them would expand scope into the catalog contract.
Unifies the catalog integration status display label with the same
"BE (Agent)" taxonomy as the live-probe agent dimension. The
underlying Status enum literal ("wired"), the catalog.metadata.wired
data field, and the filter chip id ("wired" → cell-matrix.tsx:311
status === "wired" filter key) are all PRESERVED — they are
persisted catalog data + filter state contracts that must not move.
Only the user-facing display label on stats-bar, adaptive-stats-bar,
and the filter chip flips.
This collapses the two distinct dashboard "Wired" surfaces (the L1
live-probe dimension and the per-cell build-state count) under one
unified label, matching the user-confirmed Path B.
Unifies the L1 "agent" live-probe display label with the taxonomy
convention established by #5473 (UI (Frontend), E2E, CV, D6 — layer
descriptor in parentheses). The dimension name stays "agent" in code
(PocketBase row keys agent:<slug>, LiveDimension union, keyFor
lookups are all unchanged stable contracts) — only the visible label
changes.
Updates the level-strip L1 badge label and the packages-section
L1-L4 header legend (W → B, "Wired" → "BE (Agent)") so the
level-strip's ToneChip first-letter abbreviation matches the legend
key. Test assertions covering the rendered letter, the legend text,
and the degraded-tone test's local variable follow suit.
The smoke probe was the same HTTP contract as /health on the same
service (200-OK JSON body), so every tick paid two HTTP calls for the
same liveness signal. Drop the /smoke GET + the smoke:<slug> primary
ProbeResult; the driver now emits health:<slug> as the primary and
agent:<slug> via writer side-emit (half the per-tick cost).
The driver's registry kind stays "smoke" so existing YAML configs and
orchestrator family wiring keep routing to this driver — the
emission key (health:<slug>) is the taxonomy contract that matters.
Dashboard:
- Drilldown: D3/e2e row labelled "UI (Frontend)" (was "E2E (Demo)");
Smoke row dropped (CellState.smoke field retained for back-compat).
- Cell badges: short label "UI" (was "E2E") in cell-pieces + unified-cell.
- Legend: D3 = "UI (Frontend): demo page renders in browser (Playwright)"
plus an explicit Health row at the top.
Tests updated to assert new labels:
- cell-drilldown.test.tsx: 6 dimensions (no Smoke); UI (Frontend) label.
- cell-drilldown.lazy-signal.test.tsx: drilldown-badge-ui--frontend- testid.
- cell-pieces.test.tsx + .signal-degrade.test.tsx: badge name "UI".
- unified-cell.test.tsx: mock-badge-UI.
- overlay-selector-integration.test.tsx: "UI" in place of "E2E".
- dashboard-color-matrix.test.tsx: badge: "UI" case names.
- liveness.test.ts: two-call contract (health + agent), regression guard
asserting no smoke:<slug> ever emitted.
Underlying probe key (e2e:<slug>/<feature>) preserved on PocketBase so
historical rows render correctly during the rename window.
Adds the dashboard Ops worker-runs section — family table, worker strip,
run-history drill-down, D0-from-staleness vs D0-from-failure family
annotation with clock glyph, and the per-family silence banner on the
coverage tab. Wires the data layer: DTOs, /api/runs fetchers, polling
hook, and a worker-runs context provider. cell-drilldown / cell-pieces
gain family-aware rendering.
The cold-load comm-error supplemental fetch runs CONCURRENTLY with the bulk
pages, so the bulk copy of an aggregate row can be NEWER than the supplemental
snapshot (the row's state changed between the two reads). The previous merge
replaced the bulk row unconditionally — regressing state/observed_at and
potentially fail_count back to the older supplemental values until the row's
next SSE delta (long for slow-cadence aggregates).
Add a freshness guard: when the supplemental row is strictly older (by
observed_at), keep the newer bulk row INTACT — signal-less rather than
chimera (newer core + stale signal). A chimera row would be silently swallowed
by the reducer's signal-PRESENCE no-op check; a signal-less bulk row lets
the next SSE delta restore the real current signal via the
undefined→defined presence flip.
Equal timestamps and unparseable timestamps both prefer the supplemental
(signal-bearing) row — only POSITIVELY-stale supplemental is suppressed,
preserving the cold-load comm-error overlay intent of CF7-F3 #1.
The harness contract gained statusSignalHasCommErrorKey (REQ-B
version-skew observability) with its dashboard sibling explicitly left
to the dashboard owner. Add the byte-identity-safe companion to
live-status.ts — placed OUTSIDE the commErrorFromStatusSignal region
pinned by commError-contract-drift.test.ts (only the decode function
source is mirrored; proven by the drift suite staying green) — so
dashboard consumers can distinguish a present-but-undecodable overlay
from a genuinely absent one. Unit-pinned: unknown future kind,
well-formed, absent, and array-expando wire shapes.
- mergeRowsToMap's disjoint-key divergence warn no longer fires on a
signal undefined⇄defined flip between row groups (live-status.ts:
coreRowFieldsEqual split out of rowsAreNoop): the initial fetch
projects signal away while SSE deltas deliver full rows, so the flip
is expected provenance, not a keyspace violation. upsertByKey's
reducer keeps treating the flip as observable (existing tests pin it).
- __tests__/cell-model.test.ts destructure comment no longer claims
noUncheckedIndexedAccess is enabled (it is not in this package's
tsconfig) nor that the sibling matched (it did not).
- src/lib/cell-model.test.ts row() helper aligned to the
destructure-with-fallback shape the comment describes.
- cell-model.ts:29-31 re-export rationale now cites the actual
importer (__tests__/cell-model.test.ts).
staleness.ts isStale treats an unparseable observed_at as NOT stale
(false-stale would downgrade a live green row), while the FF7 gate in
decodeCellCommError treats unparseable as stale/skip (false-not-stale
would pin an uncleared overlay forever). isStale predates this branch
(blame: d0da6e357, pre-existing on main; untouched here), so per the
blame gate it is left as-is and the divergence is documented at the
FF7 site instead. Reported for the ledger.
decodeCellCommError's staleness gate (cell-model.ts:697) compared only
`now - parsed > staleAfterMs`, which is never true for a future-dated
observedAt — clock skew or a corrupt producer timestamp would pin the
unreachable/pending overlay indefinitely, the same permanent-phantom
failure mode the FF7 unparseable-timestamp skip prevents. A timestamp
more than COMM_ERROR_FUTURE_SKEW_TOLERANCE_MS (5min) ahead of now is
now skipped like an unparseable one; skew within tolerance still
surfaces (pinned by a companion test so the guard can't over-correct).
The D1-D4 gate fires only on d3.exists/d4.exists, so a cell with ONLY
green D5/D6 rows (no e2e/chat/tools rows at all) slipped past it and
rendered a green chip + green d6Effective at achievedDepth=0/
ceilingDepth=0 — a false top-of-ladder claim contradicting the
strictness doctrine (PRESENT-but-null D4 grays; D5-no-data grays the
ladder). A wholly absent D3/D4 family now collapses to the gray
"unverified" chip (same shape as the d4NoData collapse) with red-D5/D6
dominance preserved, and d6Effective stays blocked (null).
cell-model.ts:847-852 (gate) / :905-921 (d6Effective).
One existing fixture (amber pass-through under reclaimed-pending) built
its amber from an absent-D3/D4 map; it now carries green e2e/chat rows
so the chip is genuinely amber through an intact ladder — the test's
never-mask assertion is unchanged.
decodeCellCommError (cell-model.ts) derives the REQ-B unreachable/
pending overlay from row.signal, but the bulk initial fetch projects
STATUS_LIST_FIELDS, which omits signal (useLiveStatus.ts /
live-status.ts:STATUS_LIST_FIELDS) — so on every page refresh rows
materialized with signal undefined and ACTIVE overlays vanished until
an SSE delta happened to re-deliver that row.
Fix option (a) — matching the projection's data-volume rationale (the
signal blob is ~61% of the bulk payload): useLiveStatus now issues a
SUPPLEMENTAL initial fetch, concurrent with the bulk pages, of ONLY
the comm-error candidate aggregate rows (key has no /<featureId>
segment) for the four mirror dimensions, WITH signal, and merges them
over their projected bulk twins by key. That is ~4 rows per
integration, so the bulk projection's first-paint win is preserved.
Per-cell rows under the same dimensions carry heavy parity-diff
signals and are deliberately not re-fetched (a stale per-cell comm
error is only a same/lower-severity tie-break candidate against the
aggregate mirror).
- live-status.ts: new FLEET_COMM_AGGREGATE_DIMENSIONS single source of
truth (d6/d4/e2e-demos/d5-single-pill-e2e); STATUS_LIST_FIELDS doc
updated to the truth (buildCellModel reads signal per cell at render;
the old "only ever read in the drilldown" claim was false).
- cell-model.ts: decodeCellCommError derives its aggregate candidates
from the shared constant (scan order preserved; pinned by the
equal-timestamp tie-break tests).
- useLiveStatus.ts: fetchCommAggregateRows + by-key merge; skipped
entirely for dimension scopes outside the aggregate set; supplemental
failure retries through the same connect() chain (fail loud).
- useLiveStatus.test.tsx: the PB mock now honours the fields projection
(returning full rows for the projected bulk fetch is exactly how this
bug stayed invisible to the suite) and serves the supplemental fetch
from a dedicated fixture; new cold-fetch red-green tests assert the
overlay renders end-to-end with NO SSE delta, the narrow filter
shape, the out-of-set skip, and the in-set narrowing.
- live-status.test.ts: formatter pass on the CF7-F3 #5 test added in
the previous commit (oxfmt).