The previous probe relied on a fragile cross-pill-difference heuristic:
it fingerprinted the rendered step text and used inter-pill differences
as a settling signal, accepting any render with >=2 rows at the deadline.
That heuristic false-greened a real pill-3 stale-card regression, because
a stale render (pill 3 still showing pill 1's content) still satisfied
">=2 rows present" and the dedup was only a wait mechanism, never a hard
gate.
This rewrites the assertion to check EXPECTED CONTENT per pill. Each pill
carries a small set of low-brittleness content markers derived from its
step titles in the d5 fixture (product-launch -> launch/marketing,
team-offsite -> venue/agenda, competitor-research -> competitor/weakness).
The assertion polls the swap window until the card shows >=2 NON-EMPTY
step rows whose joined text contains ALL of that pill's markers; otherwise
it hard-fails. Marker matching is partial and case-insensitive so it stays
robust to live-LLM (--direct) nondeterminism while still proving the RIGHT
pill's content rendered.
Verified with a dynamic-fake red-green plus three retained false-green
guards: identical-across-pills canned steps, stale non-adjacent content
(pill 3 showing pill 1), and empty/whitespace-only rows all turn RED;
distinct-per-pill content passes.
NOTE: this EXPOSES a genuine pill-3 stale-state regression on
agno/langroid/crewai-crews. Those cells are legitimately RED under the
corrected probe until that backend/frontend bug is fixed (separate
follow-up). The langgraph-python reference passes.
The aimock_wiring:global probe went red on the residual-6 live services
(harness-workers + 5 starters). Root cause is EXCLUDE naming drift after
the egress/private-networking migration: EXCLUDE_SERVICES keyed starters
as showcase-starter-<framework> (matching only starter-<framework>), but
live Railway names are bare starter-<framework>[-lang] (e.g.
starter-strands-python, starter-langgraph-js) — no match, so they fell
through to being checked, landed in unwired, and kept the probe red.
harness-workers had no exclude entry at all.
Fix: exclude the whole starter-* family by prefix in isExcluded (starters
are contributor scaffolds, categorically not wired through aimock; safe
because no showcase-* backend name starts with starter-, so it never
over-excludes a real backend), and add bare harness-workers to the infra
exclude set. Superseded showcase-starter-* literals removed; inert
showcase-shell-* legacy literals retained to keep the diff minimal.
The 20 showcase-* LLM backends were already re-wired via Railway
AIMOCK_URL; this is a naming/exclusion fix only (no starter is repointed).
## Summary
Routes the ~20 showcase demo backends to **aimock** (the record/replay
LLM proxy) over Railway **private networking** (`*.railway.internal`)
instead of aimock's **public** `*.up.railway.app` host.
Railway bills traffic to a public domain as **egress even
intra-project**, while `*.railway.internal` private networking is
**free** and **env-scoped**. The 240-concurrent-browser harness fleet
drives every demo continuously, so every LLM SSE stream from aimock back
to a demo backend is currently billed egress.
- aimock ≈ **89% of showcase egress**, ≈ **92% of the 13TB→78TB/mo
increase**.
- Estimated impact: avoids the ≈ **$602/mo → $3,856/mo** growth on the
aimock path.
## Change (config-only, reversible; SSOT-driven)
1. **SSOT** (`showcase/scripts/railway-envs.ts`): add an env-scoped
`internalDomain: "showcase-aimock.railway.internal"` to the aimock entry
in **both** envs. The public `domain` is **kept** (health probes /
external reachability).
2. **Emitter** (`showcase/scripts/emit-railway-envs-json.ts`): emit
`internalDomains` (additive, after `domains`) into the generated JSON.
Every non-aimock service keeps its frozen shape.
3. **Generated JSON** regenerated (oxfmt-canonical; 4-line additive
diff, only the aimock entry).
4. **Promote preflight** (`showcase/bin/railway`): `ssot_target_host`
now **prefers** the private `internalDomains[env]` over the public
`domains[env]`, so the Stage-2 (U5) serviceRef assertion requires demo
backends' `OPENAI_BASE_URL`/etc. to point at the private host.
Non-aimock targets (no `internalDomains`) fall back to their public host
unchanged.
5. **Harness** wiring probe needs **no code change** (it matches on
hostname); added a discriminating test pair + updated the drift-alert
Fix text to the private host.
**Target:** aimock binds `0.0.0.0:4010` (per
`showcase/aimock/RAILWAY.md`); demo backends resolve to
`http://showcase-aimock.railway.internal:4010`.
Deployed env vars, both envs (before → after):
| key | before (public, billed egress) | after (private, free) |
|---|---|---|
| `OPENAI_BASE_URL` | `https://<aimock>.up.railway.app/v1` |
`http://showcase-aimock.railway.internal:4010/v1` |
| `ANTHROPIC_BASE_URL` | `https://<aimock>.up.railway.app` |
`http://showcase-aimock.railway.internal:4010` |
| `GOOGLE_GEMINI_BASE_URL` | `https://<aimock>.up.railway.app` |
`http://showcase-aimock.railway.internal:4010` |
| `AIMOCK_URL` | `https://<aimock>.up.railway.app` |
`http://showcase-aimock.railway.internal:4010` |
`<aimock>` = `aimock-staging` (staging) / `showcase-aimock-production`
(prod). `railway.internal` is env-scoped, so staging demos reach the
staging aimock and prod demos reach prod aimock automatically — the same
private DNS name in both envs.
---
## Red-green proof (verbatim)
### RED — live staging today (billed public egress)
Deployed `showcase-langgraph-fastapi` (staging, service `06cccb5c-…`)
via Railway `variables(...)` GraphQL:
```
OPENAI_BASE_URL = https://aimock-staging.up.railway.app/v1
ANTHROPIC_BASE_URL = https://aimock-staging.up.railway.app
GOOGLE_GEMINI_BASE_URL = https://aimock-staging.up.railway.app
AIMOCK_URL = https://aimock-staging.up.railway.app
```
Pre-fix generated JSON aimock entry — **no** `internalDomains`:
```json
{ "domains": { "staging": "aimock-staging.up.railway.app",
"prod": "showcase-aimock-production.up.railway.app" },
"internalDomains": "ABSENT" }
```
### RED — Ruby U5 serviceref resolver, with the resolver reverted to
public-only
The three new U5 tests FAIL when `ssot_target_host` returns the public
host:
```
7 runs, 15 assertions, 3 failures
1) test_serviceref_prod_pointing_at_public_aimock_host_refuses:
expected REFUSE for prod serviceRef on the public egress host, got []
2) test_serviceref_prod_pointing_at_private_aimock_passes:
prod private aimock ref must not REFUSE, got ["REFUSE: §5.2 (showcase-ag2): prod
OPENAI_BASE_URL="http://showcase-aimock.railway.internal:4010/v1" does NOT point at
aimock's env-LOCAL prod host "showcase-aimock-production.up.railway.app" ..."]
3) test_ssot_target_host_prefers_internal_over_public:
expected "showcase-aimock.railway.internal",
actual "showcase-aimock-production.up.railway.app"
```
### GREEN — after the fix
Post-fix generated JSON aimock entry:
```json
{ "domains": { "staging": "aimock-staging.up.railway.app",
"prod": "showcase-aimock-production.up.railway.app" },
"internalDomains": { "staging": "showcase-aimock.railway.internal",
"prod": "showcase-aimock.railway.internal" } }
```
Ruby U5 serviceref tests (fixed resolver — prefers `internalDomains`):
```
7 runs, 20 assertions, 0 failures, 0 errors, 0 skips
```
Full Ruby spec suite:
```
184 runs, 715 assertions, 0 failures, 0 errors, 0 skips
```
Harness aimock-wiring probe (hostname-match; internal host with `:4010`
+ `/v1` → green, demo still on public host while harness on private →
red):
```
src/probes/aimock-wiring.test.ts 29 passed (was 27; +2 new: internal-host green, public-host drift red)
src/probes/drivers/aimock-wiring.test.ts 16 passed
src/rules/rule-loader.test.ts 61 passed (aimock-wiring-drift.yml parses after Fix-text update)
renderer + render-red-tick + orchestrator 157 passed (no alert-text snapshot broke)
```
Scripts test suite (emitter golden + everything): `2147 passed, 7
skipped` (one pre-existing `/tmp` lockfile flake in
`integration-smoke-registry.test.ts`, green on rerun after clearing the
stale lock). `emit --check` idempotent + oxfmt-canonical. Harness `tsc
--noEmit`: clean.
### GREEN — live infra confirmation
- aimock **staging** deployment status = `SUCCESS` (running), binds
`0.0.0.0:4010` — so `showcase-aimock.railway.internal:4010` resolves to
a live listener for any peer in the staging env.
- aimock serving LLM-shaped responses on `:4010`: `GET /health` → `200`;
`GET /v1/models` → `200` `{gpt-4o, gpt-4o-mini}`.
## What was vs wasn't live-validated
**Validated live:** the RED (deployed staging vars still on the public
egress host); aimock staging is deployed/running and serving on `:4010`;
the full unit/wiring/promote-preflight test surface passes with the new
internal-host values.
**NOT live-validated in-session:** the in-Railway-network DNS resolution
of `showcase-aimock.railway.internal:4010` from a peer service, and a
full staging deploy that flips the four keys + redeploys a demo backend.
Reason: the in-network vantage needs `railway ssh` (requires registering
a persistent account SSH key — a stateful, human-gated change I declined
to make unsupervised) or a staging deploy (the local Railway access
token was expired; the CLI refreshed it for read/GraphQL but a deploy is
a separate gated action). Railway private networking
(`*.railway.internal`) is a standard platform feature; the local
`docker-compose.local.yml` already runs the identical
`http://aimock:4010` internal-host pattern, and the wiring probe's
hostname match is exercised by the new tests. The staging deploy +
in-network curl is the first step of the rollout plan below and must be
run before prod.
## Irreducible egress remains
This does **not** zero showcase egress. Still billed: real browse users
hitting the public demo/shell domains; and aimock in **record mode**
proxying to real providers (the outbound prompt to
OpenAI/Anthropic/Google still bills).
## Rollout plan (reversible config change, staging-first, user-gated)
1. Land this branch (SSOT + generated JSON + assertions).
2. **Staging first:** set the four keys on staging demo backends +
`AIMOCK_URL` on the harness to
`http://showcase-aimock.railway.internal:4010` (`/v1` on
`OPENAI_BASE_URL`); redeploy one demo backend + aimock; from inside a
staging service curl
`http://showcase-aimock.railway.internal:4010/health` (expect 200) and
run a real demo LLM turn / aimock-wiring probe (expect green); confirm
the aimock egress path stops accruing
(`usage(measurements:[NETWORK_TX_GB])`).
3. **User-gated** promote to prod (staging→prod), same key flip.
4. **Rollback** = flip the keys back to the public host (no code revert
needed).
## Follow-ups (out of scope — do NOT bundle)
- Fleet right-sizing (240-concurrent-browser harness).
- `OPENAI_API_KEY` consolidation.
---
Draft — do not merge. Do not deploy to prod.
The aimock-wiring probe matches on hostname, so it needs no code change for
the private-networking migration. Add a discriminating test pair proving the
internal host (http://showcase-aimock.railway.internal:4010, with :4010 port
and /v1 suffix) resolves green while a demo still on the public egress host
goes red. Update the aimock-wiring-drift.yml Fix text to point operators at
the private host instead of the public production URL.
The harness runs as pure Node ESM (package.json "type":"module", built
with tsc moduleResolution:"bundler" which preserves extensionless import
specifiers at emit, launched via node dist/orchestrator.js). Under pure
Node ESM, relative import specifiers must carry the .js extension — a
convention the harness already honors everywhere (79/79 relative imports
in orchestrator.ts end in .js).
The relocated shared/cell-model fold broke that convention: cell-model.ts,
live-status.ts, staleness.ts, and the equivalence fixtures/test imported
sibling modules extensionless ("./live-status", "./staleness", etc). tsc,
vitest, and tsx all resolve those fine, so it built and tested green — but
at container boot node threw ERR_MODULE_NOT_FOUND on
dist/shared/cell-model/live-status and crash-looped the orchestrator,
breaking the staging auto-deploy.
Add the .js extension to every offending relative import to match the
harness convention. Minimal fix — no tsconfig change.
- recoveryMessage multi-slug branch now wraps sinceAt in renderSince() so a
corrupt-but-shaped persisted sinceAt renders "unknown", not raw garbage
(matches outage path and single-slug recovery guard).
- summary.get() read failure now logs ERROR with errorId d0-monitor-summary-read
instead of a low-signal WARN (silently blinds the detector, same family as the
other silent-disable guards).
- classifyProducer inflight short-circuit now also requires anyWorkerOnline, so a
stale/orphaned inflight from a dead worker cannot force a blind live scan.
- Extended C6 recovery test to cover the multi-slug arm; split the inflight
predicate test into online/offline-worker cases; fixed C1(ii) rotation comment.
C1 (core): select the shown/named outage slugs by re-post-due-ness + rotation
instead of an alphabetical prefix slice, gate the aggregate post on a DUE slug
actually being named, and advance lastAlertAt only for named-and-due slugs. This
stops a wide (>maxSlugs) outage from re-posting every 15m (overflow slugs whose
clock never advanced stayed perpetually "due") and stops a newly-opened overflow
slug from forcing a per-tick re-post; every open slug is now named within a
bounded number of re-posts.
C2: derive outage onset from the gate-failing (non-green) winner rung, not only a
literal red row — a degraded winner no longer strands earliest at NaN and
re-stamps sinceAt to now.
C3: log a loud errorId when readStatusRows breaks on a finite totalPages while the
last page was full (short/inconsistent read) instead of silently truncating.
C4: floor repostMinutes at min 1 in resolveConfig (0 → repostMs 0 → every-tick
re-post).
C5: classifyProducer distinguishes a fresh deploy (workers online, no run history)
as "no-data / not-yet" from a paused "idle" producer; both HOLD (never page
without data) but the fresh case is no longer a misleading permanent SUSPEND.
C6: validate a persisted sinceAt is a parseable ISO before interpolating into the
outage message (renderSince) so a corrupt-but-shaped state blob renders "unknown",
not garbage.
C7: fix the mis-annotated `unsupported` fixture flag and add a symmetric assertion
pinning every naiveMislabels flag == (naiveGone != expectedGone).
Bucket-b: stamp the outage-duration line with evidenceMs (consistent with the
recovery post + lastAlertAt); skip the alert_state write on a pure no-op tick.
Central structural lever: derive every monitor state from a single per-cell
classifier (classifyCell → gone | healthy | unknown). Treat UNKNOWN (gray /
no-data / stale / amber / comm-error) as UNKNOWN everywhere — never "gone",
never positive-healthy.
- B-F1: recovery/CLOSE requires POSITIVE green cells (cellHealthy: chipColor
green, achievedDepth>=3, fresh), not the mere absence of red. A gone column
decaying to no-data no longer auto-recovers. Fixed the healthyRows test
fixture to emit a genuine green D5/D6 ladder.
- B-onset: derive sinceAt from the folded verdict's contributing ladder rows,
not a raw row.state==="red" re-scan (removed keyBelongsToSlug).
- B-A5gap: guard self-heal/loud-log on "no slug has any wired cell", not
map.size (every integration slug is keyed even with zero wired cells).
- B-env: normalize the prod gate (trim+lowercase, empty-as-unset) so an empty
SHOWCASE_ENV no longer shadows a prod Railway env and a mis-cased/padded
value no longer silently disables the monitor. Added resolveMonitorEnv +
shouldRegister; orchestrator + gate test both use the real predicate.
- B-flap: an already-open outage keeps its hourly re-post even when a later
confirm scan is inconclusive (confirm gates only OPEN and CLOSE).
- B-cadence: advance lastAlertAt only for slugs actually named in the message
(respect maxSlugsInMessage); overflow slugs keep their clock.
- Cheap: no-wired-cells logs once per tick; MAX_SLUGS floors at 1; lastAlertAt
stamps evidenceMs; resolveConfig negative/NaN/empty coverage.
Bucket-(a) fixes, each with a local red-green test:
- A1: replace the substring `:${slug}` onset match with an anchored
exact slug-segment match (`keyBelongsToSlug`) so a prefix-colliding
sibling (`strands` vs `strands-typescript`) no longer mis-attributes
the earlier sibling's red onset. RED: strands' sinceAt was pulled to
the strands-typescript onset; GREEN: each slug gets its own onset.
- A2: recovery/CLOSE is now SYMMETRIC with OPEN — a recovery requires a
second agreeing fresh-healthy read (confirm scan). RED: a single
transient healthy read fired a false "recovered"; GREEN: held until
two reads agree.
- A3: guard an empty/degenerate schedule set — longestPeriodMs 0/NaN
would make idleWindowMs 0 → isProducerLive permanently false → the
monitor SUSPENDS forever and never pages. Falls back to a 45m default
window (DEFAULT_IDLE_WINDOW_MS) and logs at error.
- A4: bound readStatusRows — guard NaN/undefined totalPages (a `page >=
NaN` break never trips) and add a hard MAX_STATUS_PAGES cap so a full
page + bad totalPages cannot infinite-loop/OOM. RED: OOM; GREEN:
terminates at the cap.
- A5: a registry-load failure logs at error with a stable errorId (not a
silent warn-once permanent no-op), and the monitor accepts a loader
thunk so it re-reads registry.json each tick while the wired-cell set
is empty — a transiently-missing file self-heals without a redeploy.
- A6 (verified, no code change): createSlackWebhookTarget already throws
on every non-2xx (4xx/5xx/429/3xx/network-exhausted); added a test
asserting the monitor does NOT delete recovery state when the post
throws.
Bucket-(b): stamp the recovery message with the confirm-scan instant
(evidenceMs) not tick-start; wrap the scheduler tick handler in a
catch+errorId; log the prod env-gate skip at warn with a reason; add a
clarifying comment that the aggregate lastAlertAt reset is intentional
one-message-one-clock cadence; simplify the three dashboard barrel-shim
comments (drop the rot-prone enumerated symbol lists).
Add a harness-native monitor that pages #oss-alerts when a whole
integration column collapses to red-D0 ("completely gone" / backend
unreachable) in production — the incident class the per-cell alert rules
miss (LGT went fully gone on 2026-07-13 and nothing paged).
Detection runs the dashboard's OWN buildCellModel fold (the shared
cell-model module both the dashboard and the monitor import) over the
same PocketBase status rows and applies a column-gone predicate over the
resulting CellModel fields, so the monitor's verdict equals the DepthChip
the dashboard renders by construction — no parallel re-derivation.
- d0-gone-predicate.ts: pure cellGone/columnGone/columnFreshHealthy over
buildCellModel outputs + registry-derived wired-cell enumeration
(mirrors the dashboard page-stats iteration / determineCellStatus rule).
- d0-gone-monitor.ts: createD0GoneMonitor factory — producer-liveness
SUSPENDED gate (reuses the family-silence inflight-aware /api/runs
reasoning, 3x-longest-period idle window), 60s confirm re-read (never a
re-probe), 15m-detect vs 1h-repost state machine, positive-fresh-healthy
CLOSE gate, ONE aggregated outage / consolidated recovery Slack message,
durable per-slug JSON map in alert_state (getSet/putSet).
- orchestrator.ts: register internal:prod-d0-gone-monitor @ */15, gated on
SHOWCASE_ENV ?? RAILWAY_ENVIRONMENT_NAME === production + kill-switch,
control-plane-only (inside runControlPlane), reusing the oss_alerts
webhook target + shared memoized family summary.
- unified-cell.test.tsx: add the required isStaleCell/observedAtAgeMs
fields to the CellModel test literal (Phase-1 dashboard tsc gate).
Red-green: a frozen test-only naiveGone (achievedDepth===0 alone)
mislabels gray-D0-no-data and stale columns as gone on committed
fixtures (RED); the real predicate fires only on red-D0-fresh and matches
buildCellModel's own outputs (GREEN). Producer-idle SUSPENDED proven
load-bearing (disabling the gate flips both F1 tests red). Plus
confirm-scan blip-rejection, hourly dedup, recovery-clear, failure modes,
and the prod-only/kill-switch registration gate.
Move the pure cell-classification fold cluster (cell-model, live-status,
staleness, format-ts) out of showcase/shell-dashboard/src/lib/ into
showcase/harness/src/shared/cell-model/ so BOTH the dashboard and a new
harness monitor import ONE copy with zero duplication and no behavior change.
The harness builds via tsc -p tsconfig.build.json with rootDir:"src" and
cannot import outside its own src/, so the harness is the correct library
home. The dashboard consumes the cluster via relative path across the package
boundary (established precedent, e.g. d5-cadence-banner.redgreen.test.ts).
- git mv the four files into harness shared/cell-model/; their intra-cluster
relative imports stay valid (they move together, no external coupling).
- Replace the four original shell-dashboard paths with thin export-* barrels
so all ~51 existing dashboard import sites resolve unchanged.
- Repoint commError-contract-drift.test.ts's source-text drift parse at the
new canonical harness location (the barrels carry no derivation body).
- Add cell-model.equivalence.test.ts + committed fixtures + a pre-move
baseline JSON (generated from the original git-HEAD code) proving the move
is byte-identical across a red-D0, gray no-data, stale, mixed, all-green,
and unsupported column.
The aggregate e2e-smoke:<slug> red signal omitted errorDesc on the normal
return path, so an abort/timeout/send-budget-exhausted red that runLevel
RETURNS (not throws) showed on the PRIMARY dashboard tick as an unclassified
content-shaped red — only the side chat:/tools: rows kept the classifier.
Thread the failing level's errorDesc (L3 precedence, L4 fallback) onto the
aggregate so the primary tick matches the side row and the launcher-phase
abort path. Does not change red/green — only carries the classifier.
Also reorder the aborted-and-empty short-circuit ABOVE the alternate-content
/ raw-byte evaluate reads: an aborted run's page is tearing down, so those
reads were swallowed against a dead page and emitted an ambiguous empty
histogram. Non-aborted runs still perform the alternate-content salvage.
A mid-poll abort — the external ctx.abortSignal firing, or the driver's
own hard-timeout landing during the first-token poll — makes runAttempt
return empty WITHOUT throwing. The retry loop breaks and control falls to
the clean-exit path, where the level was misclassified as a generic
content red ("empty assistant response", probe.exit outcome "err", no
errorDesc). That masqueraded a teardown/abort/timeout as a CONTENT
failure on the dashboard + CVDIAG.
Add an aborted-AND-empty guard before the content-red gate that
short-circuits to the same abort classification the other paths use
(errorDesc "abort", probe.exit outcome "timeout"). Discriminator is
abortSignal.aborted, not emptiness alone: a genuinely-completed-empty
turn (not aborted) stays the content-red "empty assistant response".
Harmonize the three readTurnState() consumers in the d4 chat-roundtrip probe
through one guarded safeReadTurnState() wrapper so a mid-poll readTurnState()
throw is handled consistently everywhere: it means "no reliable signal" ->
degraded widen + observable telemetry, never a silent false-red (the prior
readDegraded swallow) nor a spurious level-error (the prior unguarded
readBaseline/readTurnComplete escape). A genuinely-empty degraded turn still
reds at the ceiling.
Folds completing the PR's own items:
- item-1 first-send cap: guard the in-send press against SEND_PRESS_MIN_BUDGET_MS
so a near-hang type can't floor press to ~1ms and produce a generic
level-error; classify distinctly as send-budget-exhausted. Only press is
guarded (type opens the envelope), so a legitimately-small pageTimeoutMs still
issues a healthy first send.
- item-3 null-header: suppress the finally-block fallback probe.message.send once
a real-header boundary already fired, so a retry whose winning resend lands no
POST no longer emits a second null-header boundary (mis-attributed
edge_interference_signal).
Also add "abort" to the errorDesc JSDoc enumeration (zero-risk).
Red-green covered for all three behavioral items against the real
runLevel/readTurnComplete path with a faithful fake.
Tests for the #5882 bucket-(b) follow-up hardening:
- Budget-exhaustion retry guard (red-green): a near-exhausted-budget retry no
longer attempts a doomed ~1ms-floored resend (type invoked exactly once).
- Degraded-path floor (red-green): a degraded page (sseAttachFailed, no
completion signal) with a late-but-present token resolves GREEN instead of a
base-floor false-red; a genuinely-empty degraded run still reds.
- Retry telemetry re-attribution (focused test): a retry-rescued GREEN turn
records the winning attempt's edge headers on probe.message.send.
- Coverage: FIFO-cap CVDIAG_MAX_OUTSTANDING_STARTS_PER_URL eviction backstop;
DEBUG-auto-disarm fail-closed (disarmed => no raw-byte capture);
alternate-content / raw-byte block SKIPPED on the container-success path.
- Fake fix: makeLateTokenBrowser now mints per-page state so L3 and L4 each run
an independent stall+retry cycle (was a shared-page singleton that leaked
sendCount from L3 into L4, so the L4 retry path was never genuinely exercised).
Follow-up to #5882 (bucket-(b) CR items). Three behavioral/telemetry fixes
plus one documentation clarification, all in d4-chat-roundtrip.ts:
- Budget-exhaustion retry guard: skip a non-completion retry resend when the
remaining wall-clock budget is below RETRY_MIN_BUDGET_MS (750ms). A late
resend previously floored its type/press action timeout to ~1ms, throwing a
page-fault-shaped error that mis-classified the stall as a generic red — a
spurious-red flap source. The stall now reds on its own terms.
- Degraded-path floor: when the SSE interceptor silently no-ops
(sseAttachFailed), no completion signal ever arrives, so the poll could only
fall into the never-observed branch and pin the deadline to the base floor —
reintroducing the slow-first-token false-red #5882 targets. Consult
sseAttachFailed to WIDEN the never-observed wait to the per-attempt ceiling so
a late-but-present token on a degraded page is still captured.
- Retry edge-header re-attribution: on a retry-rescued turn, re-arm the
message-POST edge-header capture (messageSendEdge / lastMessagePostResp /
emitMessageSend latch) so probe.message.send / edge_interference_signal / the
DEBUG raw-byte sample reflect the WINNING attempt, not the stalled first one.
- Document why lastStoppedAtMs is retained on the d4 TurnState (write-only in
d4; part of the shared attachSseInterceptor global shape the d6 run-signal
snapshot also mirrors) so it does not read as dead code.
The sseOnlyStaleStop guard in makeLateTokenBrowser was ineffective: the
driver reads the per-attempt baseline via readTurnState() BEFORE the
first send, and at that point lastSendAtMs===0 made elapsed (~epoch ms)
exceed completeAt, spuriously flipping complete=true at the baseline
read. That inflated the baseline runsFinished to prior+1, so the current
turn's real finished edge never rose PAST the baseline and the driver's
sseDone = runsFinished > baseline.runsFinished could never fire. The
SSE-only-stale-grace path was therefore never entered — the test passed
only because the token rendered directly, so it did NOT guard the
Node-stamp fix.
Gate the fake's complete on started (a turn has been sent) so the
pre-send baseline read is a TRUE baseline (runsFinished = prior +
priorSendsDone). The finished edge is now a genuine THIS-turn transition
the driver observes via sseDone, and the SSE-only completion path with a
stale lastStoppedAtMs is genuinely exercised.
Proof the guard now bites (temporary production revert, not committed):
- pre-fix stamp (completeAtMs = stoppedAtMs > 0 ? stoppedAtMs : now):
test FAILS, expected 'red' to be 'green' (grace collapses).
- restored Node-stamp (completeAtMs = Date.now()): test PASSES.
Also corrects two production comments (3 reviewers flagged): the
completed-empty deadline comment claimed the base floor is always
respected, but Math.min(..., fastFailEnd, attemptCeiling) intentionally
clamps below the floor (fast-fail); and the FIRST_TOKEN_FAST_FAIL_MS doc
now spells out the full Math.min term. Comment-only, no logic change.
On an SSE-only completion (turn detected complete via runsFinished>baseline
with no fresh DOM stop-edge for THIS turn), readTurnState().lastStoppedAtMs
still held a stale prior-run value, and it is stamped on the browser-page
clock while graceEnd/deadline math runs on the Node clock. Feeding it into
Node-clock arithmetic pushed graceEnd into the past and collapsed the
FIRST_TOKEN_GRACE_MS window to the base floor, false-REDing a late-but-present
first token.
Stamp completeAtMs from Date.now() (Node) at the first poll that observes
THIS turn complete; readTurnComplete no longer threads stoppedAtMs. Also DRY
the attempt-0 baseline onto the existing readBaseline helper and fix stale
comments (fallback emit is in finally; eviction rides the onResponse wiring;
lastStoppedAtMs doc). completed-empty still fast-reds; base floor and
hardCeiling caps preserved.
Adds a red-green unit test modelling an SSE-only completion with a stale
lastStoppedAtMs (grace collapses pre-fix, honored post-fix).
The first-token poll keyed turn-complete off the page-GLOBAL monotonic
`runsFinished >= 1` / latched `sawRunningTrue`. A PRIOR run on the page
(auto-greeting / initial-mount run) leaves those already satisfied when the
user's turn starts, so the poll treated THIS turn as already complete, saw the
still-empty container, and spuriously fast-failed RED — the a1 false-red.
Fix: capture a per-attempt BASELINE (`runsFinished` + `runStartCount`) at send
time and treat the turn complete only on a NEW edge past that baseline
(`runsFinished > baseline` / a new `runStartCount` DOM run-start). A fresh
baseline is taken before each retry resend, so a stale prior edge can no longer
defeat the retry. The grace window is now stamped from the REAL finished edge
(`lastStoppedAtMs`) rather than the poll's local clock (fixes the ~500ms-short
grace). `TurnState` / `readTurnState` are extended to surface `runStartCount`
and `lastStoppedAtMs` (the sse-interceptor already latches them).
Also:
- Bound the retry resend's type/press action timeouts by the remaining budget
to `hardCeiling` (was flat `pageTimeoutMs`, letting a stalled resend push
poll-phase wall-clock to ~2x past the ceiling).
- Surface an interceptor-attach fault (`wirePlaywrightPage.goto`) via an
injectable `onAttachFault` marker + `TurnState.sseAttachFailed` so a silent
regression to the inert base-floor path is detectable, not invisible.
- Move the `probe.message.send` fallback emit into the `finally` (idempotent)
so it fires on nav/send throw paths too.
- Label external `ctx.abortSignal` aborts as `"abort"` (not `"driver-error"`)
in the aggregate, matching the per-level classification.
- Correct the fast-fail-floor / hardCeiling / poll-deadline doc comments.
Tests: a1 regression (prior finished run + in-flight turn → not false-red) at
L3 and L4; completed-empty does exactly ONE send (retry does not fire); L3
coverage for the grace/fast-fail/retry path; attach-fault telemetry surfaces.
The prior D4 first-token fix (ceae0c2c9) was INERT in production: it keyed the
turn-complete decision off the `onSseEvent` Node-side seam, which the real
launchers never wire (Playwright has no per-SSE-event signal). `sseObserved`
therefore stayed false on the real path and the whole extension collapsed to the
pre-fix base budget floor. Its red-green used a fake page that invoked the seam
synthetically, so the deadness was never caught.
Root cause of the flap: a STALLED turn (RUN_FINISHED served by aimock but the
page never rendered it — the real 20:16:52Z failure), NOT a mere client render
race. So we need a real completion signal AND a retry for never-completed turns.
Three-part fix (mirrors d6-all-pills' production-wired signal):
1. Wire the real signal. `wirePlaywrightPage.goto` now calls
`attachSseInterceptor(page)` before navigation (injectable for tests), seeding
the page-side `__hk_runsFinished` / `__hk_copilotRunning` turn-lifecycle
globals at document_start. A new `readTurnState()` E2ePage seam reads them via
`page.evaluate`; the first-token poll keys off THAT — the same
transport-level + DOM run-stop edge d6 trusts — not the dead onSseEvent seam.
2. Fast-fail genuinely-empty turns. With a real turn-complete edge, a turn that
completes with empty assistant text reds in ~completion+grace (bounded by
FIRST_TOKEN_FAST_FAIL_MS ~15s) instead of burning the flat 60s. A
completed-empty turn still reds (no masking) and is never retried.
3. Retry-on-non-completion. A turn OBSERVED in-flight that never signals
completion within budget (stalled/dropped stream) retries once before red.
Never-observed (dead/no-turn) runs stop at the base floor, no retry. Total
wall-clock is bounded by pageTimeoutMs (per-attempt budget split).
CR findings resolved: abortSignal.aborted checked inside the poll loop; the
body-scrape fallback keeps fromAssistantContainer=false AND no longer clears
cvdiagResponseEmpty, so a fallback-salvaged red can't emit terminal_outcome=ok;
red/green never gated on cvdiag (telemetry-only); the no-turn budget is capped by
the per-attempt ceiling; the fallback `tail.length>20` floor and
`split("\n")[0]` truncation removed (false-red on short/multiline answers);
stale "not started" / dead-seam comments corrected.
Real-surface red-green (real chromium + real attachSseInterceptor + real driver
against a local fixture serving SSE /api/copilotkit with injected stream delay):
- RED (pre-fix, interceptor unwired): late-token turn -> red "empty assistant
response" in ~1.4s; sseObserved-on-real-path = FALSE (dead seam proven).
- GREEN (post-fix, interceptor wired): same late-token -> green;
readTurnState on real path = {attrPresent:true,sawRunningTrue:true,...};
sseObserved-on-real-path = TRUE.
- completed-empty (wired): still RED, fast-fail ~2.5s/level, runsFinished:1.
- recoverable-stall (wired): attempt 1 never completes -> retry -> green.
Unit suite (56 tests) rewritten to exercise the real readTurnState path plus the
retry/fast-fail behaviors; tsc + build + vitest all pass.
D4's L4 "tools" probe read the assistant-message container by polling
textContent for a fixed textPollTimeoutMs. On a run where the first token
rendered into the DOM slightly later than that budget — on a turn that
genuinely produced content — the poll exhausted and read the container as
empty, yielding a spurious "L4: empty assistant response" red (a client-side
first-token render race, not a real-LLM/fixture issue).
Harden the wait to key off the AG-UI SSE turn lifecycle rather than a fixed
timeout: track RUN_FINISHED/RUN_ERROR on the already-wired onSseEvent seam,
keep polling while a turn is in-flight (up to the pageTimeoutMs hard ceiling),
and after completion allow a small bounded first-token grace window for the
DOM to paint. A turn that completes with no content ever still fails, and when
no SSE stream is observed at all the poll falls back to the base budget floor
(unchanged pre-fix behavior, no hangs).
Adds red-green tests exercising the real runLevel wait path: a late-but-present
first token now passes; a genuinely-empty completed turn still fails.
Brings the 499-commit-stale foundations branch up to date with main so #5761
has a clean diff and no stale reverts (e.g. forwardHeaders). Conflicts:
- CopilotThreadsDrawer.tsx: took main's (main renamed CopilotDrawer -> ThreadsDrawer
+ added the collapse feature; the branch's edit was a no-op import-type split).
- pnpm-lock.yaml: regenerated with the pinned pnpm 10.33.4 (adds @copilotkit/bot-intelligence).
The fleet pnpm-workspace.yaml carries a multi-segment glob
(`examples/v2/*/apps/*`) that the strict matcher rejects with a
SchemaError. Because that throw happens during enumeration — before the
probe's `pathPrefix` filter applies — it aborted the entire version_drift
discovery, surfacing as probe.discovery-enumerate-failed / discoveryFailed
with 0 PB rows.
Skip patterns whose static (wildcard-free) prefix cannot intersect the
requested `pathPrefix` BEFORE validating their glob shape, so a deep-glob
for an unrelated subtree no longer aborts a probe that only wants
`packages/`. An unsupported pattern that DOES overlap the requested prefix
still surfaces the strict-shape SchemaError, and behavior with no
pathPrefix is unchanged.
Commit 7c3edca changed sample-attachment-buttons.tsx across all integrations
to auto-send via agent.addMessage with autoPrompt strings:
- "can you tell me what is in this demo image I just attached"
- "can you tell me what is in this demo pdf I just attached"
But the d5 harness fixture and all 19 d6 per-integration multimodal.json
fixtures still matched on the old strings:
- "describe the sample image"
- "summarize the sample document"
Aimock received requests with the new prompts, found no match, returned
a STRICT 404, and the agent emitted a streaming error back to the UI
(exact symptom: "An internal error has occurred while streaming events").
Also update agentic-chat.json across all 20 integrations (those files had
duplicate fallback entries for the old prompts) and fix split-fixtures.ts
to route the new strings to the "multimodal" feature bucket.
Local RED: ms-agent-python and crewai-crews both fail with fixture-miss
status=miss before this change.
Local GREEN: langgraph-typescript passes after this change (both turns
settle with "image" / "document" keywords confirmed in transcript).
Remaining failures after this fix are pre-existing Python backend issues
(ChatClientException on binary content parts in ms-agent-python; CrewAI
flow failure on binary content in crewai-crews) — unrelated to fixture
keys and tracked separately in the pydantic-ai multimodal work.
Three stale, pre-layer-(b) prose sites still claimed drain() ABANDONS the
in-flight run and the loop skips queue.report: the drainFleetWorker JSDoc, the
runWorker stop() inline comment, AND the case-worker SIGTERM-handler block.
Layer (b) changed this — on SIGTERM the worker stops claiming, lets its
in-flight cell FINISH within the 90s grace, and REPORTS its real terminal
result via the runAbort/abortedWithoutResult discriminator; only a run that
OVERRUNS the grace is abandoned -> reclaimed by layer (a). Preserved the still-
accurate drainReason=shutdown red-side-emit soft-wind-down + deregister-vs-crash
distinction. Also refreshed the WHY-THIS-ORDER preamble (90s finish-and-report
grace, 180s platform stop / layer-c drainingSeconds). Comments only — no
behavior change.
CR P2 round on layer-(b) graceful worker drain.
P2-A (contested TOCTOU, reconciled empirically): a run that resolves with a
valid result in the same flush the grace setTimeout fires is REPORTED, not
spuriously abandoned. runAbort.abort() lives only in stop()'s Promise.race
TIMEOUT leg, which loses the race once `done` is resolvable — so a finished
run never trips the abortedWithoutResult discriminator. Reviewer crb6 was
correct; the TOCTOU is a non-bug, so no logic changed. Added a deterministic
fake-timer regression pin forcing both the run-completion timer and the grace
timer due in one advance (verified to BITE under an over-aggressive
abandon mutation), plus a clarifying comment at the discriminator.
P2-B: corrected the requestDrain() JSDoc — post-B2 the report-skip,
heartbeat-stop, and driver-cancel key on the grace-expiry signal
runAbort.signal, not stopAbort.signal (which now only stops claiming).
B5 — set the production drain grace T to 90s (was 6s). The grace stopped being
a bare TEARDOWN budget and is now the FINISH-AND-REPORT budget stop() waits for
an in-flight run to finish-and-report (layer b) before firing runAbort -> abandon
-> layer-(a) reclaim. T must BOUND a typical cell-job (so a normal in-flight job
finishes within grace) yet stay SHORTER than the platform SIGTERM->SIGKILL window
with headroom. Sized from in-repo cell-job signal: a single-service cell-job runs
~15s (light e2e-deep) up to ~200s (heavy d6-all-pills under contention); the
per-job lease ceiling is 300s. 90s covers the bulk and stays well under the lease
so a finishing job's lease never lapses. The tail (a job that cannot finish in
grace) falls back to layer (a) — grace is deliberately FINITE. 90s is a defensible
default; B-VAL confirms/retunes from staging p95. Env-overridable via
WORKER_DRAIN_GRACE_MS.
Introduce PLATFORM_STOP_GRACE_MS (180s) documenting the C3 requirement: layer-(c)
must set Railway terminationGracePeriodSeconds = 180 so the composed serial budget
DRAIN_DEREGISTER_TIMEOUT_MS (3s) + DEFAULT_WORKER_DRAIN_GRACE_MS (90s) fits with
>=30s headroom for the health-server-close + pool-shutdown remainder. The
composed-budget test now pins the relation 3s + grace < PLATFORM_STOP_GRACE_MS
(was hardcoded < 10s) and the concrete numbers.
Add a behavioral fake-clock test of the composed-budget invariant: a run that
finishes WITHIN T is reported (finish-and-report); a run still running AT
grace-expiry fires runAbort -> abandons WITHOUT a usable result -> not reported
(layer-a reclaim backstop). The raise also broke the existing wedged-driver
default-grace test (advancing a 90s fake span re-armed the 0ms-yielding heartbeat
into a runaway cascade); fixed by giving it a clock-honoring sleep so the
heartbeat stays quiet across the window and only the grace timer is crossed.
Post-B2/B3 a run can finish-and-report after a graceful drain (drain no
longer hard-cancels the run; only grace-expiry runAbort does, which is
ctx.abortSignal). The d6 red-suppression already AND-s drainReason with
ctx.abortSignal.aborted and errorClass==="abort", which makes it
aborted-only by construction — a finished run whose legitimate red is a
genuine failure (errorClass!="abort") is NOT suppressed. This adds a
pinning test for that aborted-only contract (a finished goto-error red
under drainReason=shutdown with an un-fired abort signal must still be
reported). Verified RED via a temporary mutation to the over-broad
"drainReason alone" suppression form, GREEN on the real code.
B3 (layer-b graceful drain): the heartbeat-abort previously keyed on the
DRAIN signal, stopping lease renewal at drain-start. With B2 decoupling drain
from run-abort, a still-finishing job now runs past drain-start until grace
expiry, so killing renewal at drain-start lets the lease lapse mid-finish and
the layer-(a) reaper could reclaim the row out from under the worker
(double-run / report-after-reclaim). Gate the heartbeat-abort on the
grace-expiry signal (runAbortSignal) instead: a finishing job keeps renewing
until it reports terminal; a genuinely-abandoned job (runAbort fired at
grace-expiry) still stops renewing so its lease lapses and the reaper reclaims
it. runAbortSignal defaults to drainSignal, preserving direct-call unit-test
semantics.
Decouple 'stop claiming new jobs' from 'abort the in-flight run': drain()
no longer fires the run's abortSignal. The run's hard cancel is now a
separate grace-expiry signal (runAbort), so a run seconds from done
finishes within grace and is reported instead of being abandoned to the
sweeper. The abandon break is now conditional on abortedWithoutResult
(runAbort fired at grace-expiry → no usable result); a finished run falls
through to the report path. A wedged run that overruns grace is still cut
(runAbort.abort() in stop()'s timeout leg) and abandons → layer (a)
reclaim is the backstop.
Reconcile the drain tests that encoded the old abort-on-drain coupling to
the new contract (drain no longer fires ctx.abortSignal; the grace-expiry
abort fires at grace, exercised via a short WORKER_DRAIN_GRACE_MS).
## What
Replaces the queue reaper's delete-wins behavior for stale long-expired
orphaned in-flight rows with **reclaim-wins-until-cap**: an orphaned row
is re-queued (not deleted) up to a bounded number of CONSECUTIVE
re-orphans before final claim-delete. This is layer (a) of the
worker-reclamation + graceful-rollover redesign — it makes a worker
bounce mid-column non-lossy (the column's in-flight work is reclaimed by
surviving workers instead of dropped).
## How
- **Reclaim-wins-until-cap** in `runSweepExpired`
(`showcase/harness/src/fleet/queue-client.ts`): stale long-expired
orphaned in-flight rows are re-queued; the prior delete-wins carve-out
is inverted.
- **Consecutive-orphan budget** (`consecutive_orphan_count`, migration
`1779990400`): the reclaim cap (`MAX_RECLAIM_ATTEMPTS=3`) is scoped to
*consecutive* sweeper re-orphans — incremented only on the sweeper
re-queue path, **reset to 0 on terminal `done`/`failed`**. Peer-worker
expired-lease *steals* bump the lifetime `reclaim_count` (retained as a
dashboard diagnostic) but do NOT consume the reclaim budget.
- **Stale-age re-anchoring** (`requeued_at ?? created`, migration
`1779990300`): a reclaimed row's staleness clock restarts so it isn't
immediately re-expired.
## Review
7-agent CR + 7-agent confirmation round → **0 P0 / 0 P1**. Highlights:
- **Red-green GENUINE (empirically re-verified):** P1-A — a long-lived
job with lifetime `reclaim_count=3` but `consecutive_orphan_count=0` is
RE-QUEUED on a fresh orphan (pre-fix: wrongly deleted); P1-B — low
boundary at `MAX-1=2` re-queues, mutation-killed. Both non-tautological
(the test fake bumps `consecutive_orphan_count` only on reclaim, never
on steal; a dedicated pin test confirms steals don't consume budget).
- **No regression / terminates / fails-safe:** 248/248 tests, tsc +
oxlint clean; reclaim loop terminates at 3 consecutive orphans →
claim-delete; queue cannot grow unbounded; SweepResult metrics +
`reclaim_count` dashboard consumers (run-view `jobs.reclaimed`,
family-silence) unbroken.
- **Migration `1779990400`** additive, nullable, default-0;
pre-migration rows evaluate to 0 → safe re-queue path.
## Follow-on (not in this PR)
- Layer (b) graceful worker drain and layer (c) rolling restart are the
remaining phases of the redesign — see the Notion proposal:
https://app.notion.com/p/38b3aa381852817bacf5c9cda1f11cc0🤖 Generated with [Claude Code](https://claude.com/claude-code)
## What
Tightens the **D4/BE chat-roundtrip probe gate** from `text.length > 0`
to a **provenance check**. A turn now only counts green if the assistant
text came from the `[data-testid="copilot-assistant-message"]` container
— not from the `<body>` fallback scrape (which picks up static page
chrome).
## Why
This is the gate weakness that **masked the BIA prod outage**: a dead
agent (RUN_STARTED→RUN_FINISHED with zero `TEXT_MESSAGE`) left the page
chrome on screen, the `<body>` scrape returned non-empty text, and the
old `text.length > 0` gate reported **green** while the agent was
actually broken. The new gate threads a `fromAssistantContainer` flag
(set true only on the testid-container read path) so a body-scrape-only
result fails (`!fromAssistantContainer` → red). The same guard is
applied at L4 (tool-rendering).
## Review
**7-agent CR → 0 P0 / 0 P1.** Highlights:
- **Red-green genuine** (empirical): old gate false-passes BIA (1/52
RED); post-fix 52/52. 3-way guard (green /
empty-via-unchanged-length-guard / BIA-via-provenance).
- **Blast radius: no false-fail risk** — all 20 integrations'
`agentic-chat` + `tool-rendering` demos use the shared v2 `CopilotChat`
which mounts the testid unconditionally; aimock `d4/<slug>/chat.json`
fixtures are text-bearing for all 20. Verdict: **ship as a blocking
gate, no soak.**
- **No lost coverage** — the `<body>`-scrape fallback *code* is retained
(diagnostics); only the gate *semantics* narrowed. No D4-probed cell
relied on the fallback to pass.
- **No cross-family regression** — change is confined to
`d4-chat-roundtrip.ts`; other probe drivers untouched; 169/0 tests pass;
tsc clean.
## Notes / follow-ups (non-blocking, P2)
- A *future* probed demo using a `CopilotChatAssistantMessage`
children-render override (which doesn't emit the testid) would
false-fail — worth a one-line note near the fallback comment if that
pattern is ever added to a D4 route.
Part of the 2026-06-26 incident remediation (companion to #5733); this
gate would have surfaced the BIA failure that the dashboard masked.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
P1-A (cap semantics): the MAX_RECLAIM_ATTEMPTS cap was keyed on
`reclaim_count`, a LIFETIME tally bumped by BOTH the sweeper re-queue
path AND the peer-worker expired-lease steal (claim CAS). A long-lived
job that accrues benign peer steals could exhaust its 3-budget and then
get claim-DELETED on its first real orphan rather than re-queued.
Fix: introduce a dedicated `consecutive_orphan_count` column (migration
1779990400) that is bumped ONLY by the sweeper re-queue path in the
fleet-claim release CAS, and reset to 0 on every terminal done|failed
release. The peer-worker steal (claim CAS wasExpiredSteal branch) does
NOT touch this counter. The reaper's cap check now uses
`consecutive_orphan_count` instead of `reclaim_count`. `reclaim_count`
is left intact as the lifetime dashboard diagnostic (jobs.reclaimed).
P1-B (low boundary): adds a test at consecutive_orphan_count = MAX-1
(= 2) asserting the row is RE-QUEUED, not deleted. The off-by-one
mutation `>= MAX` -> `>= MAX-1` causes this test to go RED.
Test-fake honesty: `makeReclaimClaim`'s claimJob now explicitly models
the steal-bump on `reclaim_count` (matching the real hook) while
intentionally NOT bumping `consecutive_orphan_count`, and adds an
explicit pin test confirming steals do not consume the reclaim budget.
JSDoc on MAX_RECLAIM_ATTEMPTS updated to describe the correct semantics:
consecutive re-orphans scoped by sweeper re-queue, reset on terminal.
Red-green proof:
- P1-A RED: revert cap to reclaim_count → "P1-A CAP SCOPE" fails with
`expect(undefined).toBeDefined()` (job deleted instead of re-queued)
- P1-A GREEN: consecutive_orphan_count cap → test passes (re-queued)
- P1-B RED: mutate `>= MAX` to `>= MAX-1` → low-boundary test fails
- P1-B GREEN: revert mutation → low-boundary test passes
Suite: 145 queue-client + 103 producer = 248 total, all green.
Layer (a) of the worker reclamation+rollover redesign: make a worker bounce
non-lossy. The reaper's long-expired carve-out (G1d) used to claim-DELETE an
orphaned in-flight (claimed/running) row whose lease expired beyond its
family's stale window AND whose created-age was past that window —
`reclaimed=0, expiredPending++` — silently dropping work an abrupt bounce
(SIGKILL past grace / OOM / crash) left mid-flight.
Invert it to RECLAIM-WINS-UNTIL-CAP: re-queue the orphan to pending (it
re-runs; idempotent probes make at-least-once safe) until its durable
`reclaim_count` reaches MAX_RECLAIM_ATTEMPTS (3), only then claim-deleting a
row that keeps re-orphaning so a poison job cannot loop forever.
The carve-out existed to dodge an honesty bind: re-queueing a `created`-stale
row emitted a "back in flight" gray the next sweep falsified by claim-deleting
it off the renewal-immune `created` age. Dissolve the bind with a new
`requeued_at` column (migration 1779990300) the release CAS stamps on every
pending re-queue; both stale phases now age off `staleAgeAnchorMs`
(`requeued_at ?? created`), so a reclaimed row is genuinely young again and the
next sweep does not delete it. `reclaim_count` (migration 1779990200) is reused
as the attempt counter — no second tally to drift.
Red-green proven on the real reaper (queue-client.test.ts): a stale-aged
long-expired orphan below the cap goes RED (deleted, reclaimed=0,
expiredPending=1) on delete-wins and GREEN (re-queued, reclaimed=1,
expiredPending=0, requeued_at stamped) on reclaim-wins; plus an attempt-cap
deletion test and a next-sweep no-falsification test.
The L3 chat-roundtrip gate only asserted `text.length > 0`, which the
`<body>` fallback scrape can satisfy with static page text (nav links,
footer copy, demo blurb) trailing the sent message. When the agent run
finished with ZERO assistant content (RUN_STARTED -> RUN_FINISHED, no
TEXT_MESSAGE) the assistant-message bubble never rendered, so the probe
fell back to scraping <body> and false-PASSED on incidental page chrome —
the mechanism that masked the BIA outage (dead agent reported green).
Track the provenance of the captured response and require it to come from
the [data-testid="copilot-assistant-message"] container (a genuine
assistant turn — the DOM-layer equivalent of RUN_FINISHED + a non-empty
TEXT_MESSAGE), not the body fallback. Apply the same guard to L4 so weather
content must also come from a real assistant turn. The instrumented probe
sends the aimock fixture headers and renders real content into the
container, so this does not affect a working agent; it only rejects the
empty-turn false-pass.
Adds a red-green regression test reproducing the BIA false-pass.
A normal harness deploy rebuilds the shared showcase-harness image and
bounces the pool workers (PR #5715). Immediately after the bounce the
workers re-register, the producers re-arm, and every family is mid-sweep:
lastSuccessAt still points at the pre-bounce success, so it reads stale
against the silence thresholds (banner 2x period, Slack alert 3x period +
3 consecutive ticks). The result was a FALSE "worker family X has not
completed successfully" banner AND Slack family-silence alert during the
expected post-bounce drain window.
Fix: a bounce-keyed grace window. The freshest worker registered_at across
the /api/runs workers strip is the fleet's most-recent bounce instant
(independent of CP boot — a worker can bounce while the CP stays up). While
now - bounce < 2 x period, a family with no success yet is DRAINING, not
silent, so neither the §7.4 banner, the §7.3 cell glyph, nor the §9 Slack
alert flags it. Beyond the window with still-no-success, genuine silence
fires exactly as before.
The determination lives in two surfaces (server monitor for the Slack
alert; client isFamilySilent for the banner + glyph), so both now consume
the same new SSOT field (WorkerView.registeredAt) and the same 2x-period
grace constant, keeping them consistent.
- run-view.ts: project registered_at -> WorkerView.registeredAt (server)
- family-silence-monitor.ts: BOUNCE_GRACE_PERIOD_MULTIPLIER + freshest-bounce
grace gate, keyed off body.workers
- worker-runs-context.tsx: freshestBounceMs + bounceAtMs grace arg on
isFamilySilent; banner + cell glyph pass it
- ops-api.ts: WorkerView.registeredAt on the client DTO
Add a d5-a2ui-recovery probe so the A2UI error-recovery demo runs on every
PR via the d5/d6 fleet harness, not only the manual on-demand workflow.
- New probe d5-a2ui-recovery.ts drives both pills in one session: HEAL
asserts >=2 newly-mounted declarative-metric tiles and no hard-failure
card; EXHAUST asserts the "Couldn't generate the UI" card appears and
no surface paints. Deltas (vs a pre-send baseline) keep the two
mutually-exclusive negatives correct across the shared session. The
transient "Retrying..." label is not asserted (timing-flaky).
- Prompts are sent as typed input, keyed per integration slug, mirroring
each slug's suggestions.ts message verbatim. The recovery prompts are
unique per slug because the inner render_a2ui calls carry no
x-aimock-context; a typed message is byte-identical to the pill
dispatch, so it matches the same fixture. Sending via input (not a
preFill pill click) lets the runner snapshot its run-lifecycle baseline
first, avoiding a false done-signal-missing failure.
- Register a2ui-recovery in d5-registry, map it in d5-feature-mapping,
add its representative fixture, and mirror the mapping in the dashboard
CATALOG_TO_D5_KEY (kept in lock-step via the drift test).
Verified green locally on both recovery paths: langgraph-python
(backend-owned get_a2ui_tools) and strands (auto-inject middleware).
## Summary
Makes the showcase harness probe's turn-done signal **reliable**,
killing the dominant class of dashboard false-red flaps without ever
hiding a real failure.
`waitForTurnComplete` previously relied on a fragile SSE fetch-counter
conjunct that false-reds healthy demos whenever the page-side fetch
wrapper missed the runtime URL/transport. This change makes the
**`data-copilot-running` DOM attribute** (driven directly by the agent
run lifecycle, `RUN_STARTED`→true / `RUN_FINISHED`→false,
transport-independent) the **PRIMARY** done-signal, with the SSE counter
demoted to a **headless-only fallback** (headless demos never render
`CopilotChatView`, so the attribute is absent).
Design (all three preserved — no false-green, no false-red, hangs still
red):
- **Primary signal** = the `data-copilot-running` true→false
**transition** with a **stayed-stopped quiescence window** (a stop must
persist on the same run-start count for `settleMs`; a new sub-run resets
it) — so it cannot complete on an intermediate stop in a multi-step
turn.
- **SSE counter** = headless fallback only; never an OR-trigger when the
DOM signal is present.
- **`done-signal-missing` backstop** (gated on `attrPresent===true` +
`runningNow!==true`) reds a genuine painted-but-never-finished DOM turn
before the hard timeout; headless turns use their full timeout for their
only signal.
## How it was reviewed
A full 4-round `cr-loop` (7 unbiased agents/round + confirmation rounds
+ a Procedure-3 promotion audit) caught and fixed **5 distinct
correctness defects** in the implementation before merge:
- **F1** — SSE OR-trigger could complete a multi-step turn early on an
intermediate stop (false-GREEN), in both the loop and the post-loop
classifier.
- **F2** — the run-start baseline was captured *after* the message send,
killing the primary signal on fast turns (false-RED).
- **F3** — non-atomic double `surfaceReady` read per poll (latent hazard
+ wasted round-trip).
- **F4** — the surface-mount (`completeOnMount`) path had no quiescence
window (false-GREEN on intermediate stop + false-RED on a still-running
gen-UI turn).
- **F5** — the early backstop false-redded slow-but-healthy **headless**
turns (now gated on the DOM signal).
Bidirectional red-green tests for F1–F5 plus a systematic `{DOM,
headless} × {completes, lagging-recovers, genuine-hang} × {text,
surface}` completion/backstop matrix. Full harness unit suite: **3173
passed / 18 skipped / 0 failed**; `tsc --noEmit` clean; lint 0 errors;
build clean.
## Known follow-ups (NOT in this PR — pre-existing / non-blocking)
- **Theoretical edge (not reachable on real or realistically-streamed
turns):** if a run completed within a single synchronous microtask
(zero-duration), the page-side MutationObserver could miss the true edge
while `attrPresent===true` → false-red. Real LLM turns and aimock
realistic-streaming hold the attribute true across many event-loop
ticks, so the observer reliably latches it. A naive "re-add SSE fallback
for DOM-present" fix would reintroduce F1's multi-step false-green, so
it's intentionally not done here.
- **Recommended quick follow-up (latency only, no wrong verdict):**
capture `baselineBannerText` pre-`sendTurnMessage` (mirroring the
run-start/count baselines) so a fast-erroring cold-start turn fast-fails
(#5142) instead of burning the full timeout.
- **Pre-existing sse-interceptor capture/counter internals** (none
load-bearing for the new done-signal; verified STAY_IN_C by the
Procedure-3 audit): page-side counter soft-nav/multi-capture reset,
`__hk_fetchWrapped` pattern reuse + hardcoded fallback, g/y-flag
stateful RegExp, TextDecoder end-of-stream flush, bare-catch
reader-error swallow, framenav payload discard/TOCTOU,
CDP-wallTime-vs-Date.now TTFT, addInitScript/close-listener
re-registration accumulation.
## Test plan
- [x] `pnpm test` (harness) — 3173 passed / 18 skipped / 0 failed
- [x] `tsc --noEmit` exit 0, lint 0 errors, build exit 0
- [ ] Verify on staging that auth / prebuilt-sidebar / claude-sdk-tools
(and other previously-flapping cells) stop false-redding while
genuinely-broken cells stay red
Please review the replay/primary-signal approach. Not auto-merging.