## Summary
- Bumps `@ag-ui/client`, `@ag-ui/core`, and `@ag-ui/encoder` to `0.0.57`
across all packages under `packages/`
- Updates `pnpm-lock.yaml` to reflect the new versions
## Test plan
- [ ] Verify packages build successfully
- [ ] Confirm no runtime regressions in ag-ui communication
Prevents a file-parallelism race where this suite's beforeAll generator
call atomically renamed catalog.json while generate-registry.test.ts was
mid-sentinel-test, clobbering the appended sentinel.
Documentation companion to the BIA D6 readiness work:
- PARITY_NOTES.md: tracks built-in-agent feature parity vs LGP +
a 'Known Issues' section documenting 3 D6 demos that remain red
due to downstream renderer/state gaps in @copilotkit/react-core
(a2ui renderer host doesn't mount Card; STATE_DELTA → useAgent
state-subscription gap; declarative-gen-ui renderer host shares
the same class). All three require fixes outside this PR's scope.
- feature-registry.json: adds threadid-frontend-tool-roundtrip
feature entry covering the new gen-ui-agent STATE_DELTA wiring.
Adds the set_steps tool definition to state-tools and wires it through
the tanstack factory so the gen-ui-agent demo emits STATE_DELTA events
on its agent-plan stream the same way the LGP integration does.
Two aimock fixture updates to unblock built-in-agent D6 verification:
- gen-ui-a2ui-fixed.json: new fixture so display_flight pill routes
the Card mount through the a2ui-fixed-schema renderer host
- gen-ui-declarative.json: reshaped to the built-in-agent two-call
sequence with the flat catalog shape the BIA route emits
Extracts the inline subagents page DelegationCard into three reusable
components that mirror the LGP integration layout, enabling D6 testid
coverage:
- subagent-activity-card.tsx: per-subagent activity panel
- delegation-log.tsx: cumulative delegation event log
- supervisor-activity-banner.tsx: supervisor-level status banner
The subagents page now composes these components instead of inlining
the markup, with data-testid hooks for D6 selectors.
## Summary
D6 reds remediation across the staging dashboard. Consolidates
manifest-level NSF/hygiene fixes for 13 integrations with two targeted
claude-sdk-python fixes that surfaced during the per-integration audit.
### What's in scope
**Manifest cleanup** (12-14 integrations touched):
- Added `not_supported_features` entries for `reasoning-default-render`,
`agentic-chat-reasoning`, `tool-rendering-reasoning-chain`, and
`interrupt-headless` per backend capability (e.g. ms-agent-dotnet gets
only `interrupt-headless` since its `ReasoningAgent.cs` natively emits
REASONING_MESSAGE_*; langgraph-fastapi natively supports both → no
additions).
- Deduped `shared-state-streaming` and related entries that previously
appeared in BOTH `features:` AND `not_supported_features:`.
- Added `shared-state-read` and `threadid-frontend-tool-roundtrip` to
`features:` where the demo exists but wasn't declared. **Note**:
`threadid-frontend-tool-roundtrip` was REVERTED — feature isn't
registered in `showcase/shared/feature-registry.json`; the registry
needs the canonical entry before manifests can declare it. Tracked as
follow-up.
**claude-sdk-python**:
- `hitl/page.tsx`: replaced `useLangGraphInterrupt` (from
`@copilotkit/react-core`) with `useInterrupt` (from
`@copilotkit/react-core/v2`) — wrong-framework import on a non-LG
backend.
- `gen-ui-interrupt`: added `data-testid="time-picker-cancel"` to the
cancel button.
### Out of scope (deferred / handled elsewhere)
- **ag2 + spring-ai manifests**: owned by parallel session
`fix/content-reds` — those changes land via that PR.
- **11 of 12 audit-flagged claude-sdk-python testid gaps**: investigated
and refused — they're missing FEATURES (custom-catchall renderers,
sign-in cards, headless composer, etc.), not missing testid attributes
on existing JSX. Adding them would be a behavior change requiring
backend-capability audit per demo. Separate follow-up if the team wants
the structural gaps closed.
- **built-in-agent 17-item testid backfill**: per D analysis, the
single-`"default"`-agent pattern is INTENTIONAL architecture (documented
in README + earliest port commit + uniform across 12+ routes); per-demo
dimension comes from `src/lib/factory/*` keyed off route URL. Track 1
(PARITY_NOTES.md doc) + Track 2 (testid-only backfill, separate PR)
tracked as follow-ups.
### Validation
- `npm run validate-manifests` passes; all 15 affected manifests parse
valid YAML.
- claude-sdk-python typecheck: no new errors (pre-existing
module-resolution noise unchanged).
- No vitest tests exist for the touched claude-sdk-python demos.
- `npx oxfmt --write` on changed files.
### Provenance
Driven by a per-integration parity audit (LGP canon, 17 red
integrations, 11-slot CR fan-out, 39-demo canon enumerated; reports
archived to /tmp/cr/audit/ during the investigation). Coordinated with
parallel `fix/content-reds` session via the `coordinate` skill to avoid
file-overlap with their ag2 + spring-ai manifest work and other deeper
changes.
Closes#5406 (claude-sdk-python testid/hook fixes — folded into this
PR).
## Additional fixes (folded in after re-audit)
Re-audit of the original claude-sdk-python audit revealed 16 ADD-EMITTER
fixes the prior agent missed. All components existed; only `data-testid`
attributes were absent.
| demo | testids | commit |
|---|---|---|
| auth | 1 (`auth-demo-error-message`) | 1238046fc |
| headless-simple | 3 (composer + 2 message bubbles) | 5793a3652 |
| a2ui-fixed-schema | 1 (`a2ui-fixed-card` via Card renderer override +
matching definitions entry, mirroring LGP) | 4436f3afa |
| headless-complete | 6 (composer + 2 bubbles + stock/weather/highlight
cards) | 083eb09a7 |
| declarative-gen-ui | 5 (card + status-badge + metric + pie-chart +
bar-chart) | 6187040b5 |
Not added (genuinely missing components, deferred):
- `auth-sign-in-card`, `auth-sign-out-button`, `auth-profile-card`
- `headless-revenue-chart`
- tool-rendering testids (15 components), chat-slots (2), etc.
Adds inline comments next to the data-testid="headless-message-{user,assistant}"
markers in assistant-bubble.tsx, user-bubble.tsx, and headless-simple/page.tsx
explaining that the testids intentionally repeat once per message — mirroring
the canonical LGP implementation — and that role discrimination for the D6
conversation-runner is via data-message-role, not unique testids.
Backfills the data-testid markers the D6 probes assert against across the
auth, headless-simple, headless-complete, a2ui-fixed-schema,
declarative-gen-ui, and gen-ui-interrupt demos; aligns the python
a2ui_fixed agent + a2ui definitions/renderers with the canonical LGP
shapes; and switches the gen-ui-interrupt CopilotKit provider import to
@copilotkit/react-core/v2 so the demo mounts under the V2 runtime that
the D6 probe drives.
Adds the chart-card renderer and wires it into the headless-complete
tool-renderers map so the D6 probe for the headless-revenue-chart feature
can mount and assert against a real chart component.
Adds the missing gen-ui-a2ui-fixed aimock fixture and reshapes the
declarative-gen-ui fixture so the claude-sdk-python D6 probes have
deterministic LLM replay for both gen-ui demos.
Replaces useLangGraphInterrupt with useInterrupt (LangGraph-specific hook
not exported from the V2 React core) and switches the CopilotKit provider
import to @copilotkit/react-core/v2 so the demo mounts under the V2
runtime that the D6 probe drives.
Annotates 14 integration manifests with d6_not_supported_features entries
and aligns d6_supported_features with what each backend actually implements,
so the D6 fleet probe only enumerates demos that the backend can serve.
## Summary
Adds end-to-end **run-visibility** to the showcase fleet — a new data
path from queue enqueue → per-family projection → dashboard + Slack
alerting:
- **Contracts + PB schema**: `EnqueueJobInput.family`,
`FleetQueueClient.pruneAged`, run-id/family/worker-id columns on
`probe_jobs` / `resource_snapshots`, fleet-claim hook stamps at claim
time.
- **Queue + aggregator**: queue-client stamps run-id/family on enqueue;
d6 producer owns `pruneAged` retention; aggregator computes
`redsIntroduced`/`redsCleared` into `probe_runs`.
- **Run-view projection (`run-view.ts`)**: §5.1 `FLEET_FAMILIES`
registry + §5.2.1 family-summary projection.
`createMemoizedFamilySummary` bounds PB load at ~one fan-out per TTL
regardless of viewer count.
- **Producer wiring**: `PRODUCER_FAMILY_WIRING` drift-locks the four
producers' family ids set-equal to the registry (unit-tested);
CLI-triggered runs route through `familyForLevel` so a registry rename
breaks loudly.
- **Family-silence monitor (§9)**: Slack alerting for the one incident
class transition-keyed rules are blind to (a silent family produces no
row transitions). Rides the existing fleet-health interval (no extra
timer), per-family 6h rate-limit + recovered one-shot + boot-grace,
fails open on PB-down for the meta-alert. Closed-vocabulary redaction.
- **HTTP `/api/runs`** + **`/health.fleetRuns.lastEvaluatedAt`**:
read-only fleet-runs routes mounted unconditionally on the CP role,
backed by the SAME memoized summary the monitor uses. `/health` stamp is
the §9 compensating control so an external poll can detect a wedged
monitor.
- **Orchestrator wiring**: boot-resolved `workerStaleAfterMs` threaded
through both fleet-health and the projection so they judge staleness
against the same window; test queues across the harness gain a no-op
`pruneAged`.
- **Dashboard**: worker-runs Ops section (family table, worker strip,
run-history drill-down), D0-from-staleness vs D0-from-failure family
annotation, per-family silence banner on the coverage tab; data layer
(DTOs, fetchers, polling hook, context provider).
- **Integration test**: 689-line end-to-end test exercising the full
queue lifecycle → /api/runs projection across all four families.
This is the merged scope of the **runviz blitz** lanes T1-T15.
## Test plan
- [x] `pnpm typecheck` on `showcase/harness` and
`showcase/shell-dashboard` — clean
- [x] `vitest run` on full harness suite — 2276 pass / 3 pre-existing
probe-pool timing flakes (NOT in diff; subject is "probe-pool timing
fragility", a different PR's concern)
- [x] `oxfmt --check showcase/` — all green
- [x] Integration test `run-visibility.integration.test.ts` passes —
exercises queue → projection across all four families
- [ ] CI on PR HEAD — pending push, monitored after open
## Known follow-up (bucket-d)
- Probe-pool timing-sensitive flakes in
`src/probes/helpers/browser-pool.test.ts` (FIX#4a self-heal) and
`src/probes/loader/probe-invoker.test.ts` (timeout assertions) —
different test fails on each run; subject is probe-pool timing
fragility, not runviz.
## Deviation note
The runviz ship-finisher session ran in a Claude Agent SDK harness
without an Agent dispatch tool, so the standard 7-agent CR-loop could
not be dispatched. Inline review was performed against the cr-loop
subject-manifest + four-bucket partition: T1-T15 each landed via their
own per-task review on the blitz integration branch before reaching this
PR, and T8/T15 (the two final-merged streams) were validated via
typecheck-clean + targeted-tests-pass + cross-module wiring inspection.
A heavyweight 7-agent CR can be run against this PR if desired.
- queue-client.ts: add missing closing brace at EOF (TS1005 after rebase)
- job-producer.test.ts: add required `family: "d6"` to producer fixtures
and remove duplicate `logger` key in startedProducer
- result-aggregator.test.ts: align with per-row try/catch + dedup-lookup
behavior introduced in 0b2f613e0 — add `persisted: true` and
`writeOverlay` to test writers, update B6 contract expectations
- queue-client.test.ts: drop a stray blank line
Adds a 689-line integration test that exercises the full queue lifecycle
to the /api/runs projection (enqueue → claim → terminal → projection)
across all four families. Updates the railway-envs golden + verify-deploy
drivers regression test to account for the new fleet-runs route surface.
Adds the dashboard Ops worker-runs section — family table, worker strip,
run-history drill-down, D0-from-staleness vs D0-from-failure family
annotation with clock glyph, and the per-family silence banner on the
coverage tab. Wires the data layer: DTOs, /api/runs fetchers, polling
hook, and a worker-runs context provider. cell-drilldown / cell-pieces
gain family-aware rendering.
PRODUCER_FAMILY_WIRING (drift-locked set-equal to FLEET_FAMILIES via a unit
test) drives every buildJobProducer call site; the boot-resolved
worker-stale-after window is threaded through BOTH fleet-health and the
shared family-summary projection so they judge staleness against the same
window. Triggered CLI control-plane runs route through familyForLevel so a
registry rename breaks loudly instead of silently enqueueing jobs invisible
to the projection. Test queues across the harness gain a no-op pruneAged
for the new contract.
Read-only /api/runs routes mounted on the control-plane role, backed by
the SHARED memoized family-summary instance (one PB fan-out per TTL).
Bounds: per-route memo, request rate-limit. /health gains the
fleetRuns.lastEvaluatedAt stamp from the family-silence monitor as the
§9 compensating control for a wedged monitor (an external poll detects
the wedge — the monitor cannot report its own host's death).
job-producer takes a family option and stamps it on every enqueue; the
prune-ownership key is the d6 producer's family. family-silence-monitor
rides the existing fleet-health interval (no extra timer to tear down),
keys 6 h rate-limit + recovered one-shot per family, fails open on
PB-down (the meta-alert path must still fire), and renders alert text
from closed-vocabulary parts only (§5.2.1 redaction). control-plane
fire-and-forgets familySilence.tick(now) each fleet-health cycle.
The §5.1 FLEET_FAMILIES registry and the §5.2.1 family-summary projection
that derives per-family outcome / inflight / lastRun / lastSuccessAt from
the PB-backed batches. The memoized variant fan-outs PB reads at most once
per TTL regardless of viewer count — the SAME instance is shared by the
/api/runs routes and the family-silence monitor so a dashboard poll and a
monitor evaluation inside the same TTL cost one PB fan-out total.
queue-client stamps run_id/family onto every enqueue so downstream
projections can attribute jobs to a family-scoped batch; pruneAged
retention legs land here (the d6 producer owns the call, per §4.2).
result-aggregator computes redsIntroduced/redsCleared from claimed/
sequenced job state into probe_runs summary.
Adds run-id/family/worker-id columns to probe_jobs and resource_snapshots,
plus the EnqueueJobInput.family + FleetQueueClient.pruneAged contracts and
the hoisted deriveHealth primitive that downstream projections share. The
fleet-claim PB hook stamps run-id/family at claim time so every later
projection has a stable join key. probes/run-history is updated to read
the new columns.
## Summary
Bundles two related dashboard/harness lanes that landed together once
both reached green:
1. **CF #18 fairness / claim-fair lane** — the 23-commit base
(`fix/fleet-claim-fairness` rebased onto its CF7 integration tip
`80c5c9402`) carrying claim-spike, cf3/cf4/cf5/cf6/cf7 hardening waves,
plus the **CF8 micro-fix** for the round-8 supplemental-merge finding:
- **CF8 F3 supplemental-merge freshness guard** (`useLiveStatus.ts`):
the cold-load comm-error supplemental fetch runs CONCURRENTLY with the
bulk pages, so the bulk copy of an aggregate row can be NEWER than the
supplemental snapshot. The previous merge replaced the bulk row
unconditionally, regressing `state`/`observed_at` to stale values until
the row's next SSE delta (long for slow-cadence aggregates). Added
`supplementalRowIsOlder` — when the supplemental row is strictly older
the newer bulk row stays INTACT (signal-less) rather than being grafted
into a chimera (newer core + stale signal) that the reducer's
signal-PRESENCE no-op check would silently swallow.
- **Procedure 3 promotions** (bucket-c/d audit over the CF round-8
ledger): one functional fix (escape `workerId` in fleet-health's reclaim
list filter — matches the `JSON.stringify` pattern used at
orchestrator.ts:3240 for the same field; a `"`-bearing worker_id would
otherwise break out of the literal) plus five doc reconciliations on the
producer/contract surfaces flagged across slot1/slot2/slot4/slot5 (stale
queue-client cross-reference, WarmHealthConfig doc vs `gate.specs`
reality, TickResult.reclaimedIndeterminate "Of reclaimed" → "In addition
to reclaimed", TickResult.skippedForBacklog poisoned-count contributor,
SweepResult.commErrors equation scoping).
2. **Dashboard drilldown D4 parity** — three commits making the
dashboard drilldown surface the D4 rung the same way the grid does:
- `feat(showcase): add resolveD4Row + CellState.d4 so the drilldown can
see the D4 rung` — exposes the D4 row on `CellState`, modeled after
`resolveD3Row`.
- `fix(showcase): drilldown shows the D4 row, de-crosses the e2e label,
and scopes the rollup line honestly` — renders D4 in the drilldown
panel, fixes the crossed e2e label, scopes the rollup line to its real
source set.
- `fix(showcase): unify dimension naming on the legend taxonomy across
grid, legacy cells, and legend` — taxonomy cleanup so legend, grid, and
legacy cells use one set of labels.
## Verification
- Dashboard `npx tsc --noEmit`: clean (only the pre-existing
missing-`@/data/*.json` errors that exist on `main`).
- Dashboard `vitest run`: 59 files / 1012 tests passing, 1 skipped.
- Harness `npx tsc --noEmit`: clean.
- Harness `vitest run`: 121/122 files passing; 1 failed = `probe-invoker
times out invoker-level even when driver ignores abortSignal` — the
NAMED known wall-clock flake (per CF7 integration verification), re-run
in isolation: 67/67 PASS.
- Rebase of drilldown-parity onto the updated `fix/cf8-m1` tip: ZERO
conflicts (the two lanes touch disjoint files).
## Test plan
- [ ] CI green on the PR
- [ ] Dashboard drilldown shows D4 rung in addition to D3/D5/D6
- [ ] Legend taxonomy reads the same in legend, grid, drilldown
- [ ] (post-merge, in showcase) cold-load comm-error overlay still
paints; a stale supplemental no longer regresses a freshly-failed
aggregate row
The auth-middleware-presence regex in queue-client.test.ts hook-parity
suite was anchored to the single-line `}, $apis.requireAdminAuth());`
closer. After oxfmt rewrote fleet-claim.pb.js to its multi-line form
(`},\n $apis.requireAdminAuth(),\n);`) the regex stopped matching
and the test asserted 0 routes were guarded — a false alarm.
Broaden the regex to accept both the single-line closer and the
formatter's split form; the structural intent (each routerAdd
handler-end is followed by the requireAdminAuth middleware) is
unchanged.
PROMOTE_TO_A (defense-in-depth, exploit-class):
- fleet-health.ts reclaim list interpolated workerId raw into the PB filter
literal. workerId is DB-sourced (read back from the workers roster row),
not a compile-time constant, and the same field is escaped via JSON.stringify
at orchestrator.ts:3240 — but the reclaim path was missing the same
hardening. A double-quote in worker_id (corrupt row, buggy self-registration)
would either throw the list (silently skipping this worker's reclaim every
cycle) or widen the filter to claim other workers' jobs. Match the sibling
escape pattern.
PROMOTE_TO_B (doc reconciliations exposed by the CF round-8 audit):
- queue-client.ts COUNT-NAME CAVEAT was stale — the cross-referenced
contracts.ts doc was already updated (commit 80c5c940) but the queue-client
side still asked a future maintainer to make the edit that had already
landed.
- WarmHealthConfig + JobProducerOptions.warmHealth docs said the producer
warms 'every enumerated backend' / 'each enumerated spec'; the implementation
warms gate.specs (post-validation, post-backlog-gate). A fully-backlogged
tick warms nothing. Updated both interface-level docs to match.
- TickResult.reclaimedIndeterminate said 'Of reclaimed, the reclaims...' —
but the field is DISJOINT from reclaimed (sibling SweepResult contract +
the queue-client both state a thrown release lands here exclusively).
Rewrote to 'In addition to reclaimed, ...'.
- TickResult.skippedForBacklog doc named only the dedupe-gate contributor;
the in-function comment correctly documents the fail-CLOSED poisoned-count
fold-in. Expanded the field doc to name both contributors, and to clarify
that the fail-OPEN leg lands in backlogGateFailedOpen separately.
- SweepResult.commErrors pairing equation
(commErrors.length === reclaimed + reclaimedIndeterminate) was asserted
unconditionally on the shared contract, but reclaimedIndeterminate is
optional and fakes may not report the split. Scoped to 'implementations
that report the split'.
The cold-load comm-error supplemental fetch runs CONCURRENTLY with the bulk
pages, so the bulk copy of an aggregate row can be NEWER than the supplemental
snapshot (the row's state changed between the two reads). The previous merge
replaced the bulk row unconditionally — regressing state/observed_at and
potentially fail_count back to the older supplemental values until the row's
next SSE delta (long for slow-cadence aggregates).
Add a freshness guard: when the supplemental row is strictly older (by
observed_at), keep the newer bulk row INTACT — signal-less rather than
chimera (newer core + stale signal). A chimera row would be silently swallowed
by the reducer's signal-PRESENCE no-op check; a signal-less bulk row lets
the next SSE delta restore the real current signal via the
undefined→defined presence flip.
Equal timestamps and unparseable timestamps both prefer the supplemental
(signal-bearing) row — only POSITIVELY-stale supplemental is suppressed,
preserving the cold-load comm-error overlay intent of CF7-F3 #1.
Sibling of the queue-client prose fix (CF7 #10), which flagged this
contract doc as describing only the drain phase: despite the name, the
lease phase's long-expired carve-out also claim-deletes claimed/running
rows (stale created-age, long-expired or unparseable lease) into this
count — no re-queue, no comm error, no reclaimed increment.
The harness contract gained statusSignalHasCommErrorKey (REQ-B
version-skew observability) with its dashboard sibling explicitly left
to the dashboard owner. Add the byte-identity-safe companion to
live-status.ts — placed OUTSIDE the commErrorFromStatusSignal region
pinned by commError-contract-drift.test.ts (only the decode function
source is mirrored; proven by the drift suite staying green) — so
dashboard consumers can distinguish a present-but-undecodable overlay
from a genuinely absent one. Unit-pinned: unknown future kind,
well-formed, absent, and array-expando wire shapes.