Commit 7c3edca changed sample-attachment-buttons.tsx across all integrations
to auto-send via agent.addMessage with autoPrompt strings:
- "can you tell me what is in this demo image I just attached"
- "can you tell me what is in this demo pdf I just attached"
But the d5 harness fixture and all 19 d6 per-integration multimodal.json
fixtures still matched on the old strings:
- "describe the sample image"
- "summarize the sample document"
Aimock received requests with the new prompts, found no match, returned
a STRICT 404, and the agent emitted a streaming error back to the UI
(exact symptom: "An internal error has occurred while streaming events").
Also update agentic-chat.json across all 20 integrations (those files had
duplicate fallback entries for the old prompts) and fix split-fixtures.ts
to route the new strings to the "multimodal" feature bucket.
Local RED: ms-agent-python and crewai-crews both fail with fixture-miss
status=miss before this change.
Local GREEN: langgraph-typescript passes after this change (both turns
settle with "image" / "document" keywords confirmed in transcript).
Remaining failures after this fix are pre-existing Python backend issues
(ChatClientException on binary content parts in ms-agent-python; CrewAI
flow failure on binary content in crewai-crews) — unrelated to fixture
keys and tracked separately in the pydantic-ai multimodal work.
Three stale, pre-layer-(b) prose sites still claimed drain() ABANDONS the
in-flight run and the loop skips queue.report: the drainFleetWorker JSDoc, the
runWorker stop() inline comment, AND the case-worker SIGTERM-handler block.
Layer (b) changed this — on SIGTERM the worker stops claiming, lets its
in-flight cell FINISH within the 90s grace, and REPORTS its real terminal
result via the runAbort/abortedWithoutResult discriminator; only a run that
OVERRUNS the grace is abandoned -> reclaimed by layer (a). Preserved the still-
accurate drainReason=shutdown red-side-emit soft-wind-down + deregister-vs-crash
distinction. Also refreshed the WHY-THIS-ORDER preamble (90s finish-and-report
grace, 180s platform stop / layer-c drainingSeconds). Comments only — no
behavior change.
CR P2 round on layer-(b) graceful worker drain.
P2-A (contested TOCTOU, reconciled empirically): a run that resolves with a
valid result in the same flush the grace setTimeout fires is REPORTED, not
spuriously abandoned. runAbort.abort() lives only in stop()'s Promise.race
TIMEOUT leg, which loses the race once `done` is resolvable — so a finished
run never trips the abortedWithoutResult discriminator. Reviewer crb6 was
correct; the TOCTOU is a non-bug, so no logic changed. Added a deterministic
fake-timer regression pin forcing both the run-completion timer and the grace
timer due in one advance (verified to BITE under an over-aggressive
abandon mutation), plus a clarifying comment at the discriminator.
P2-B: corrected the requestDrain() JSDoc — post-B2 the report-skip,
heartbeat-stop, and driver-cancel key on the grace-expiry signal
runAbort.signal, not stopAbort.signal (which now only stops claiming).
B5 — set the production drain grace T to 90s (was 6s). The grace stopped being
a bare TEARDOWN budget and is now the FINISH-AND-REPORT budget stop() waits for
an in-flight run to finish-and-report (layer b) before firing runAbort -> abandon
-> layer-(a) reclaim. T must BOUND a typical cell-job (so a normal in-flight job
finishes within grace) yet stay SHORTER than the platform SIGTERM->SIGKILL window
with headroom. Sized from in-repo cell-job signal: a single-service cell-job runs
~15s (light e2e-deep) up to ~200s (heavy d6-all-pills under contention); the
per-job lease ceiling is 300s. 90s covers the bulk and stays well under the lease
so a finishing job's lease never lapses. The tail (a job that cannot finish in
grace) falls back to layer (a) — grace is deliberately FINITE. 90s is a defensible
default; B-VAL confirms/retunes from staging p95. Env-overridable via
WORKER_DRAIN_GRACE_MS.
Introduce PLATFORM_STOP_GRACE_MS (180s) documenting the C3 requirement: layer-(c)
must set Railway terminationGracePeriodSeconds = 180 so the composed serial budget
DRAIN_DEREGISTER_TIMEOUT_MS (3s) + DEFAULT_WORKER_DRAIN_GRACE_MS (90s) fits with
>=30s headroom for the health-server-close + pool-shutdown remainder. The
composed-budget test now pins the relation 3s + grace < PLATFORM_STOP_GRACE_MS
(was hardcoded < 10s) and the concrete numbers.
Add a behavioral fake-clock test of the composed-budget invariant: a run that
finishes WITHIN T is reported (finish-and-report); a run still running AT
grace-expiry fires runAbort -> abandons WITHOUT a usable result -> not reported
(layer-a reclaim backstop). The raise also broke the existing wedged-driver
default-grace test (advancing a 90s fake span re-armed the 0ms-yielding heartbeat
into a runaway cascade); fixed by giving it a clock-honoring sleep so the
heartbeat stays quiet across the window and only the grace timer is crossed.
Post-B2/B3 a run can finish-and-report after a graceful drain (drain no
longer hard-cancels the run; only grace-expiry runAbort does, which is
ctx.abortSignal). The d6 red-suppression already AND-s drainReason with
ctx.abortSignal.aborted and errorClass==="abort", which makes it
aborted-only by construction — a finished run whose legitimate red is a
genuine failure (errorClass!="abort") is NOT suppressed. This adds a
pinning test for that aborted-only contract (a finished goto-error red
under drainReason=shutdown with an un-fired abort signal must still be
reported). Verified RED via a temporary mutation to the over-broad
"drainReason alone" suppression form, GREEN on the real code.
B3 (layer-b graceful drain): the heartbeat-abort previously keyed on the
DRAIN signal, stopping lease renewal at drain-start. With B2 decoupling drain
from run-abort, a still-finishing job now runs past drain-start until grace
expiry, so killing renewal at drain-start lets the lease lapse mid-finish and
the layer-(a) reaper could reclaim the row out from under the worker
(double-run / report-after-reclaim). Gate the heartbeat-abort on the
grace-expiry signal (runAbortSignal) instead: a finishing job keeps renewing
until it reports terminal; a genuinely-abandoned job (runAbort fired at
grace-expiry) still stops renewing so its lease lapses and the reaper reclaims
it. runAbortSignal defaults to drainSignal, preserving direct-call unit-test
semantics.
Decouple 'stop claiming new jobs' from 'abort the in-flight run': drain()
no longer fires the run's abortSignal. The run's hard cancel is now a
separate grace-expiry signal (runAbort), so a run seconds from done
finishes within grace and is reported instead of being abandoned to the
sweeper. The abandon break is now conditional on abortedWithoutResult
(runAbort fired at grace-expiry → no usable result); a finished run falls
through to the report path. A wedged run that overruns grace is still cut
(runAbort.abort() in stop()'s timeout leg) and abandons → layer (a)
reclaim is the backstop.
Reconcile the drain tests that encoded the old abort-on-drain coupling to
the new contract (drain no longer fires ctx.abortSignal; the grace-expiry
abort fires at grace, exercised via a short WORKER_DRAIN_GRACE_MS).
## What
Replaces the queue reaper's delete-wins behavior for stale long-expired
orphaned in-flight rows with **reclaim-wins-until-cap**: an orphaned row
is re-queued (not deleted) up to a bounded number of CONSECUTIVE
re-orphans before final claim-delete. This is layer (a) of the
worker-reclamation + graceful-rollover redesign — it makes a worker
bounce mid-column non-lossy (the column's in-flight work is reclaimed by
surviving workers instead of dropped).
## How
- **Reclaim-wins-until-cap** in `runSweepExpired`
(`showcase/harness/src/fleet/queue-client.ts`): stale long-expired
orphaned in-flight rows are re-queued; the prior delete-wins carve-out
is inverted.
- **Consecutive-orphan budget** (`consecutive_orphan_count`, migration
`1779990400`): the reclaim cap (`MAX_RECLAIM_ATTEMPTS=3`) is scoped to
*consecutive* sweeper re-orphans — incremented only on the sweeper
re-queue path, **reset to 0 on terminal `done`/`failed`**. Peer-worker
expired-lease *steals* bump the lifetime `reclaim_count` (retained as a
dashboard diagnostic) but do NOT consume the reclaim budget.
- **Stale-age re-anchoring** (`requeued_at ?? created`, migration
`1779990300`): a reclaimed row's staleness clock restarts so it isn't
immediately re-expired.
## Review
7-agent CR + 7-agent confirmation round → **0 P0 / 0 P1**. Highlights:
- **Red-green GENUINE (empirically re-verified):** P1-A — a long-lived
job with lifetime `reclaim_count=3` but `consecutive_orphan_count=0` is
RE-QUEUED on a fresh orphan (pre-fix: wrongly deleted); P1-B — low
boundary at `MAX-1=2` re-queues, mutation-killed. Both non-tautological
(the test fake bumps `consecutive_orphan_count` only on reclaim, never
on steal; a dedicated pin test confirms steals don't consume budget).
- **No regression / terminates / fails-safe:** 248/248 tests, tsc +
oxlint clean; reclaim loop terminates at 3 consecutive orphans →
claim-delete; queue cannot grow unbounded; SweepResult metrics +
`reclaim_count` dashboard consumers (run-view `jobs.reclaimed`,
family-silence) unbroken.
- **Migration `1779990400`** additive, nullable, default-0;
pre-migration rows evaluate to 0 → safe re-queue path.
## Follow-on (not in this PR)
- Layer (b) graceful worker drain and layer (c) rolling restart are the
remaining phases of the redesign — see the Notion proposal:
https://app.notion.com/p/38b3aa381852817bacf5c9cda1f11cc0🤖 Generated with [Claude Code](https://claude.com/claude-code)
## What
Tightens the **D4/BE chat-roundtrip probe gate** from `text.length > 0`
to a **provenance check**. A turn now only counts green if the assistant
text came from the `[data-testid="copilot-assistant-message"]` container
— not from the `<body>` fallback scrape (which picks up static page
chrome).
## Why
This is the gate weakness that **masked the BIA prod outage**: a dead
agent (RUN_STARTED→RUN_FINISHED with zero `TEXT_MESSAGE`) left the page
chrome on screen, the `<body>` scrape returned non-empty text, and the
old `text.length > 0` gate reported **green** while the agent was
actually broken. The new gate threads a `fromAssistantContainer` flag
(set true only on the testid-container read path) so a body-scrape-only
result fails (`!fromAssistantContainer` → red). The same guard is
applied at L4 (tool-rendering).
## Review
**7-agent CR → 0 P0 / 0 P1.** Highlights:
- **Red-green genuine** (empirical): old gate false-passes BIA (1/52
RED); post-fix 52/52. 3-way guard (green /
empty-via-unchanged-length-guard / BIA-via-provenance).
- **Blast radius: no false-fail risk** — all 20 integrations'
`agentic-chat` + `tool-rendering` demos use the shared v2 `CopilotChat`
which mounts the testid unconditionally; aimock `d4/<slug>/chat.json`
fixtures are text-bearing for all 20. Verdict: **ship as a blocking
gate, no soak.**
- **No lost coverage** — the `<body>`-scrape fallback *code* is retained
(diagnostics); only the gate *semantics* narrowed. No D4-probed cell
relied on the fallback to pass.
- **No cross-family regression** — change is confined to
`d4-chat-roundtrip.ts`; other probe drivers untouched; 169/0 tests pass;
tsc clean.
## Notes / follow-ups (non-blocking, P2)
- A *future* probed demo using a `CopilotChatAssistantMessage`
children-render override (which doesn't emit the testid) would
false-fail — worth a one-line note near the fallback comment if that
pattern is ever added to a D4 route.
Part of the 2026-06-26 incident remediation (companion to #5733); this
gate would have surfaced the BIA failure that the dashboard masked.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
P1-A (cap semantics): the MAX_RECLAIM_ATTEMPTS cap was keyed on
`reclaim_count`, a LIFETIME tally bumped by BOTH the sweeper re-queue
path AND the peer-worker expired-lease steal (claim CAS). A long-lived
job that accrues benign peer steals could exhaust its 3-budget and then
get claim-DELETED on its first real orphan rather than re-queued.
Fix: introduce a dedicated `consecutive_orphan_count` column (migration
1779990400) that is bumped ONLY by the sweeper re-queue path in the
fleet-claim release CAS, and reset to 0 on every terminal done|failed
release. The peer-worker steal (claim CAS wasExpiredSteal branch) does
NOT touch this counter. The reaper's cap check now uses
`consecutive_orphan_count` instead of `reclaim_count`. `reclaim_count`
is left intact as the lifetime dashboard diagnostic (jobs.reclaimed).
P1-B (low boundary): adds a test at consecutive_orphan_count = MAX-1
(= 2) asserting the row is RE-QUEUED, not deleted. The off-by-one
mutation `>= MAX` -> `>= MAX-1` causes this test to go RED.
Test-fake honesty: `makeReclaimClaim`'s claimJob now explicitly models
the steal-bump on `reclaim_count` (matching the real hook) while
intentionally NOT bumping `consecutive_orphan_count`, and adds an
explicit pin test confirming steals do not consume the reclaim budget.
JSDoc on MAX_RECLAIM_ATTEMPTS updated to describe the correct semantics:
consecutive re-orphans scoped by sweeper re-queue, reset on terminal.
Red-green proof:
- P1-A RED: revert cap to reclaim_count → "P1-A CAP SCOPE" fails with
`expect(undefined).toBeDefined()` (job deleted instead of re-queued)
- P1-A GREEN: consecutive_orphan_count cap → test passes (re-queued)
- P1-B RED: mutate `>= MAX` to `>= MAX-1` → low-boundary test fails
- P1-B GREEN: revert mutation → low-boundary test passes
Suite: 145 queue-client + 103 producer = 248 total, all green.
Layer (a) of the worker reclamation+rollover redesign: make a worker bounce
non-lossy. The reaper's long-expired carve-out (G1d) used to claim-DELETE an
orphaned in-flight (claimed/running) row whose lease expired beyond its
family's stale window AND whose created-age was past that window —
`reclaimed=0, expiredPending++` — silently dropping work an abrupt bounce
(SIGKILL past grace / OOM / crash) left mid-flight.
Invert it to RECLAIM-WINS-UNTIL-CAP: re-queue the orphan to pending (it
re-runs; idempotent probes make at-least-once safe) until its durable
`reclaim_count` reaches MAX_RECLAIM_ATTEMPTS (3), only then claim-deleting a
row that keeps re-orphaning so a poison job cannot loop forever.
The carve-out existed to dodge an honesty bind: re-queueing a `created`-stale
row emitted a "back in flight" gray the next sweep falsified by claim-deleting
it off the renewal-immune `created` age. Dissolve the bind with a new
`requeued_at` column (migration 1779990300) the release CAS stamps on every
pending re-queue; both stale phases now age off `staleAgeAnchorMs`
(`requeued_at ?? created`), so a reclaimed row is genuinely young again and the
next sweep does not delete it. `reclaim_count` (migration 1779990200) is reused
as the attempt counter — no second tally to drift.
Red-green proven on the real reaper (queue-client.test.ts): a stale-aged
long-expired orphan below the cap goes RED (deleted, reclaimed=0,
expiredPending=1) on delete-wins and GREEN (re-queued, reclaimed=1,
expiredPending=0, requeued_at stamped) on reclaim-wins; plus an attempt-cap
deletion test and a next-sweep no-falsification test.
The L3 chat-roundtrip gate only asserted `text.length > 0`, which the
`<body>` fallback scrape can satisfy with static page text (nav links,
footer copy, demo blurb) trailing the sent message. When the agent run
finished with ZERO assistant content (RUN_STARTED -> RUN_FINISHED, no
TEXT_MESSAGE) the assistant-message bubble never rendered, so the probe
fell back to scraping <body> and false-PASSED on incidental page chrome —
the mechanism that masked the BIA outage (dead agent reported green).
Track the provenance of the captured response and require it to come from
the [data-testid="copilot-assistant-message"] container (a genuine
assistant turn — the DOM-layer equivalent of RUN_FINISHED + a non-empty
TEXT_MESSAGE), not the body fallback. Apply the same guard to L4 so weather
content must also come from a real assistant turn. The instrumented probe
sends the aimock fixture headers and renders real content into the
container, so this does not affect a working agent; it only rejects the
empty-turn false-pass.
Adds a red-green regression test reproducing the BIA false-pass.
A normal harness deploy rebuilds the shared showcase-harness image and
bounces the pool workers (PR #5715). Immediately after the bounce the
workers re-register, the producers re-arm, and every family is mid-sweep:
lastSuccessAt still points at the pre-bounce success, so it reads stale
against the silence thresholds (banner 2x period, Slack alert 3x period +
3 consecutive ticks). The result was a FALSE "worker family X has not
completed successfully" banner AND Slack family-silence alert during the
expected post-bounce drain window.
Fix: a bounce-keyed grace window. The freshest worker registered_at across
the /api/runs workers strip is the fleet's most-recent bounce instant
(independent of CP boot — a worker can bounce while the CP stays up). While
now - bounce < 2 x period, a family with no success yet is DRAINING, not
silent, so neither the §7.4 banner, the §7.3 cell glyph, nor the §9 Slack
alert flags it. Beyond the window with still-no-success, genuine silence
fires exactly as before.
The determination lives in two surfaces (server monitor for the Slack
alert; client isFamilySilent for the banner + glyph), so both now consume
the same new SSOT field (WorkerView.registeredAt) and the same 2x-period
grace constant, keeping them consistent.
- run-view.ts: project registered_at -> WorkerView.registeredAt (server)
- family-silence-monitor.ts: BOUNCE_GRACE_PERIOD_MULTIPLIER + freshest-bounce
grace gate, keyed off body.workers
- worker-runs-context.tsx: freshestBounceMs + bounceAtMs grace arg on
isFamilySilent; banner + cell glyph pass it
- ops-api.ts: WorkerView.registeredAt on the client DTO
Add a d5-a2ui-recovery probe so the A2UI error-recovery demo runs on every
PR via the d5/d6 fleet harness, not only the manual on-demand workflow.
- New probe d5-a2ui-recovery.ts drives both pills in one session: HEAL
asserts >=2 newly-mounted declarative-metric tiles and no hard-failure
card; EXHAUST asserts the "Couldn't generate the UI" card appears and
no surface paints. Deltas (vs a pre-send baseline) keep the two
mutually-exclusive negatives correct across the shared session. The
transient "Retrying..." label is not asserted (timing-flaky).
- Prompts are sent as typed input, keyed per integration slug, mirroring
each slug's suggestions.ts message verbatim. The recovery prompts are
unique per slug because the inner render_a2ui calls carry no
x-aimock-context; a typed message is byte-identical to the pill
dispatch, so it matches the same fixture. Sending via input (not a
preFill pill click) lets the runner snapshot its run-lifecycle baseline
first, avoiding a false done-signal-missing failure.
- Register a2ui-recovery in d5-registry, map it in d5-feature-mapping,
add its representative fixture, and mirror the mapping in the dashboard
CATALOG_TO_D5_KEY (kept in lock-step via the drift test).
Verified green locally on both recovery paths: langgraph-python
(backend-owned get_a2ui_tools) and strands (auto-inject middleware).
## Summary
Makes the showcase harness probe's turn-done signal **reliable**,
killing the dominant class of dashboard false-red flaps without ever
hiding a real failure.
`waitForTurnComplete` previously relied on a fragile SSE fetch-counter
conjunct that false-reds healthy demos whenever the page-side fetch
wrapper missed the runtime URL/transport. This change makes the
**`data-copilot-running` DOM attribute** (driven directly by the agent
run lifecycle, `RUN_STARTED`→true / `RUN_FINISHED`→false,
transport-independent) the **PRIMARY** done-signal, with the SSE counter
demoted to a **headless-only fallback** (headless demos never render
`CopilotChatView`, so the attribute is absent).
Design (all three preserved — no false-green, no false-red, hangs still
red):
- **Primary signal** = the `data-copilot-running` true→false
**transition** with a **stayed-stopped quiescence window** (a stop must
persist on the same run-start count for `settleMs`; a new sub-run resets
it) — so it cannot complete on an intermediate stop in a multi-step
turn.
- **SSE counter** = headless fallback only; never an OR-trigger when the
DOM signal is present.
- **`done-signal-missing` backstop** (gated on `attrPresent===true` +
`runningNow!==true`) reds a genuine painted-but-never-finished DOM turn
before the hard timeout; headless turns use their full timeout for their
only signal.
## How it was reviewed
A full 4-round `cr-loop` (7 unbiased agents/round + confirmation rounds
+ a Procedure-3 promotion audit) caught and fixed **5 distinct
correctness defects** in the implementation before merge:
- **F1** — SSE OR-trigger could complete a multi-step turn early on an
intermediate stop (false-GREEN), in both the loop and the post-loop
classifier.
- **F2** — the run-start baseline was captured *after* the message send,
killing the primary signal on fast turns (false-RED).
- **F3** — non-atomic double `surfaceReady` read per poll (latent hazard
+ wasted round-trip).
- **F4** — the surface-mount (`completeOnMount`) path had no quiescence
window (false-GREEN on intermediate stop + false-RED on a still-running
gen-UI turn).
- **F5** — the early backstop false-redded slow-but-healthy **headless**
turns (now gated on the DOM signal).
Bidirectional red-green tests for F1–F5 plus a systematic `{DOM,
headless} × {completes, lagging-recovers, genuine-hang} × {text,
surface}` completion/backstop matrix. Full harness unit suite: **3173
passed / 18 skipped / 0 failed**; `tsc --noEmit` clean; lint 0 errors;
build clean.
## Known follow-ups (NOT in this PR — pre-existing / non-blocking)
- **Theoretical edge (not reachable on real or realistically-streamed
turns):** if a run completed within a single synchronous microtask
(zero-duration), the page-side MutationObserver could miss the true edge
while `attrPresent===true` → false-red. Real LLM turns and aimock
realistic-streaming hold the attribute true across many event-loop
ticks, so the observer reliably latches it. A naive "re-add SSE fallback
for DOM-present" fix would reintroduce F1's multi-step false-green, so
it's intentionally not done here.
- **Recommended quick follow-up (latency only, no wrong verdict):**
capture `baselineBannerText` pre-`sendTurnMessage` (mirroring the
run-start/count baselines) so a fast-erroring cold-start turn fast-fails
(#5142) instead of burning the full timeout.
- **Pre-existing sse-interceptor capture/counter internals** (none
load-bearing for the new done-signal; verified STAY_IN_C by the
Procedure-3 audit): page-side counter soft-nav/multi-capture reset,
`__hk_fetchWrapped` pattern reuse + hardcoded fallback, g/y-flag
stateful RegExp, TextDecoder end-of-stream flush, bare-catch
reader-error swallow, framenav payload discard/TOCTOU,
CDP-wallTime-vs-Date.now TTFT, addInitScript/close-listener
re-registration accumulation.
## Test plan
- [x] `pnpm test` (harness) — 3173 passed / 18 skipped / 0 failed
- [x] `tsc --noEmit` exit 0, lint 0 errors, build exit 0
- [ ] Verify on staging that auth / prebuilt-sidebar / claude-sdk-tools
(and other previously-flapping cells) stop false-redding while
genuinely-broken cells stay red
Please review the replay/primary-signal approach. Not auto-merging.
The D6 fleet worker drives each integration over its insecure Docker
origin (http://<slug>:10000), where crypto.randomUUID is undefined so
hand-rolled headless chats threw and never mounted (sse-missing). tsx/
esbuild also wraps named inner functions in __name(...) calls that leak
into page.evaluate and throw __name is not defined.
Add installBrowserContextShims (init-scripts.ts): a __name no-op helper
and a crypto.randomUUID secure-context polyfill, registered via
addInitScript at document_start of every D6 page; wired into the
d6-all-pills newPage goto path.
Cover the data-copilot-running turn-done signal in waitForTurnComplete:
true->false transition completion, stayed-stopped quiescence, the
attr-gated early backstop, pre-send run-start baseline, and the
integration wait-for-turn-complete behavior.
Make waitForTurnComplete use the page-side data-copilot-running attribute
as the primary turn-done signal: detect the running true->false transition,
require stayed-stopped quiescence, and gate the early backstop on
attrPresent + runningNow to avoid headless false-RED. Capture a pre-send
run-start baseline so fast turns keep the primary signal alive, read
surfaceReady once per poll, and add computeMaxTurnDurationMs.
Add buildCopilotRunningObserverScript to sse-interceptor.ts and wire it
via addInitScript so the page exposes a data-copilot-running attribute
that the harness can observe for turn-completion signaling.
Adds a durable persistence layer for @copilotkit/bot, replacing the
in-memory-only ActionStore with a pluggable StateStore.
- StateStore interface (kv/list/lock/dedup/queue) with a shared
conformance suite; MemoryStore default plus @copilotkit/bot-store-redis
and @copilotkit/bot-store-postgres backends.
- createBot({ store }): typed per-thread state via Standard Schema,
action snapshots persisted through the store, per-conversation turn
lock (onLockConflict drop|force), and inbound-event dedup keyed on a
stable eventId. ActionStore is kept as a deprecated alias.
- Cross-platform transcripts (bot.transcripts + identity resolver) with
age-bounded retention (prune on append + filter on read), and
runAgent({ transcript: true }) to auto-inject history and capture the
reply.
- createBot({ components }) re-registers components so durable actions
re-fire after a restart; restart-durability demo in examples/slack.
- Dedup is marked seen only after the turn lock is acquired, so a turn
dropped on lock-conflict does not burn its eventId (no lost retries).
- Release lockstep: bot-store-redis/postgres version with bot + bot-ui.
- open a per-feature CvdiagProbeSession for each d5/d6 pill probe
- emit exactly-once probe.exit and failure_classifier per session
- join probe-session output to its run via the X-Test-Id header
- thread cvdiagPbWriter through the orchestrator and CLI runner
Call-Site Enumeration: FAILURE_CLASSIFIER_SET is exported from
cvdiag/probe-session and consumed by d6-all-pills (classifier validation
against the canonical set). The export has no other call sites; any future
classifier addition must update the canonical set in probe-session and the
validation in d6-all-pills together.
Behavior-preserving extraction of the CvdiagProbeSession lifecycle from the
d4 chat-roundtrip driver into a shared cvdiag/probe-session module, so the
d5/d6 probe path can reuse the same session boundaries. d4-chat-roundtrip
now imports the extracted session instead of defining it inline.
schema.json regenerated from the merged canonical schema.ts (failure_classifier
probe.exit additions UNION backend request.ingress/sse.first_byte/llm.call.*
boundaries + test_id adoption). Per-integration staged schema.ts copies
re-derived via 'showcase cvdiag-stage-ts' so codegen --check and stage --check
are both in sync. No hand-merge of generated artifacts.
## What
Adds **cvdiag** — a permanent, always-available observability subsystem
for the showcase, built to diagnose the red↔green cell flap on the
staging dashboard and to make that diagnosis a dashboard query rather
than a multi-day forensic hunt in the future.
Captures the full request path with `X-Test-Id` correlation across
**probe → backend → aimock → edge**, across every integration
(TypeScript, Python, Java/spring-ai, .NET):
- Per-language backend emitters (canonical + staged/compile-linked
mirrors), all sharing one schema (`schema.json`, closed-world
`additionalProperties:false`).
- CREATE-only writes to two new PocketBase collections: `cvdiag_events`
and `cvdiag_raw_byte_samples` (additive migrations — no existing data
touched).
- An 8-class flap classifier mapping to the observed failure signatures
(`sse-missing` / `text-unstable` / `dom-missing`).
- DEBUG-tier raw-byte capture (secret-scrubbed) and HMAC-guarded A/B
edge-interference detection.
## Why
The runId flap-fix (`cdc1e90e`, 2026-06-09) did **not** fully resolve
the flap — it was still observed 2026-06-19. cvdiag exists so the
*remaining* cause is observed live with full correlation instead of
inferred.
## Safety / enablement
- **Inert by default.** With `CVDIAG_BACKEND_EMITTER` unset the
subsystem performs zero host mutation (no logging-config changes, no
threads/tasks, no stdout) — verified by
`test_cvdiag_inert_when_disabled`. **To accumulate data, set
`CVDIAG_BACKEND_EMITTER=1` on the showcase services.**
- All per-language scrubbers match the canonical `scrubSecrets`
(sk-/base64url, Bearer, colon-less URL userinfo, size-guard) — verified
with real toolchains (vitest / mvn / dotnet).
- Merged latest `main` (only conflict: a clean `.csproj` include union).
## Verification
- harness `tsc --noEmit` ✓ · `src/cvdiag` vitest 251/251 ✓ ·
`cvdiag-stage-ts --check` in-sync ✓
- Java MessageScrubber 17/17 (mvn) ✓ · .NET CvdiagBackend 5/5 (dotnet
sdk:9.0) ✓ · Python emitters 93/93 (3.12) ✓
🤖 Generated with [Claude Code](https://claude.com/claude-code)
> **DRAFT / WIP — not reviewed, not ready to merge.** Checkpoint per
request. The mandatory 7-agent cr-loop + CI-green gate runs before this
leaves draft. LGP and ADK ship together in this PR.
## Problem (a false-D6 in both directions)
Declarative A2UI demos wire `a2ui.injectA2UITool: true`, so the response
is a rendered `render_a2ui` surface with **no assistant text bubble**.
The D6 conversation-runner's turn-completion gate required the assistant
**text** to stabilize — so on a working declarative demo the run
finished and the dashboard painted, but text never settled →
`waitForTurnComplete` timed out (`reason=text-unstable`) **before the
render assertion ran**.
Result: `langgraph-python:declarative-gen-ui` (the gold standard)
reported **false-RED while rendering correctly** (all 4 pills verified
live on staging), while `google-adk` reported **false-GREEN**.
## Fix
Opt-in `ConversationTurn.completeOnMount` (set only by
`d5-gen-ui-declarative.ts`). For those turns the text-stability
completion conjunct is **replaced** by a surface-mount predicate:
run-finished (sseOk) + a new assistant bubble + the expected declarative
testids **newly mounting**. A non-rendering surface now yields a new
`surface-missing` failure reason (truthful RED). Text-based demos are
byte-for-byte unchanged (opt-in, per-turn).
## Proof (both directions, live D6)
- RED (before): `text-unstable` timeout; dashboard text painted.
- GREEN (after): passes in ~5s; `buildDeclarativeAssertion` actually
runs and verifies testids mount for all 4 pills.
- INTEGRITY: forced a broken render (renamed testids) → test goes
**red** (`surface-missing`). Not "always green now."
- Unit: 89/89 + 3 new (green-on-mount, red-on-surface-missing).
## Scope: LGP + ADK (ship together) — both truthful GREEN
- **LGP** gold-standard cell: false-RED → truthful GREEN (all 4 pills
assert).
- **ADK** realignment: **test-only, complete.** The shared-script fix
auto-applies; verified all 4 ADK pills truthfully GREEN (each surface
mounts from baseline 0 via surface-mount completion). Prior false-green
closed; **no ADK backend gap**.
## Out of scope (someone else's problem)
This shared-script change re-evaluates **every** declarative-gen-ui cell
truthfully. Integrations beyond LGP/ADK that don't actually render will
flip to **truthful RED** — e.g. `langgraph-typescript` pill 2
(team-performance `declarative-data-table` doesn't mount). Those are
real per-demo render gaps for their owners; **not fixed here.**
## Follow-up (not in this PR)
The `render_a2ui` call returns a ~5.9 MB SSE for a ~2 KB surface (LGT
worse) — a separate runtime amplification concern in the
`injectA2UITool:true` middleware path.
## Before ready/merge
- [x] ADK empirical verdict — all 4 pills truthful GREEN, realignment
test-only, no backend gap
- [ ] mandatory cr-loop → zero findings
- [ ] CI green
Opts each declarative-gen-ui pill into the new `completeOnMount` turn
completion so these surface-rendering demos are gated on their expected
declarative testids mounting rather than assistant-text stability.
Declarative A2UI demos render a surface (mounted testids) with no assistant
text bubble, so the assistant-text-stability completion gate never settled and
timed out on working demos — a false-RED.
This adds an opt-in `completeOnMount` turn-completion path that replaces the
assistant-text-stability conjunct with a surface-mount predicate: run-finished
+ a new assistant bubble + the expected declarative testids newly mounting.
A new `surface-missing` failure reason reports when the run finishes but the
expected surface never mounts. Turns that do not opt in keep the existing
text-stability behavior unchanged.