Commit Graph

653 Commits

Author SHA1 Message Date
Jordan Ritter b2811f4feb fix(showcase): update multimodal fixture match keys to match actual autoPrompts
Commit 7c3edca changed sample-attachment-buttons.tsx across all integrations
to auto-send via agent.addMessage with autoPrompt strings:
  - "can you tell me what is in this demo image I just attached"
  - "can you tell me what is in this demo pdf I just attached"

But the d5 harness fixture and all 19 d6 per-integration multimodal.json
fixtures still matched on the old strings:
  - "describe the sample image"
  - "summarize the sample document"

Aimock received requests with the new prompts, found no match, returned
a STRICT 404, and the agent emitted a streaming error back to the UI
(exact symptom: "An internal error has occurred while streaming events").

Also update agentic-chat.json across all 20 integrations (those files had
duplicate fallback entries for the old prompts) and fix split-fixtures.ts
to route the new strings to the "multimodal" feature bucket.

Local RED: ms-agent-python and crewai-crews both fail with fixture-miss
  status=miss before this change.
Local GREEN: langgraph-typescript passes after this change (both turns
  settle with "image" / "document" keywords confirmed in transcript).

Remaining failures after this fix are pre-existing Python backend issues
(ChatClientException on binary content parts in ms-agent-python; CrewAI
flow failure on binary content in crewai-crews) — unrelated to fixture
keys and tracked separately in the pydantic-ai multimodal work.
2026-07-06 15:31:51 -07:00
Ran Shemtov 4cc25b56bf Merge branch 'main' into claude/jolly-brown-77c87b 2026-06-29 18:53:01 +02:00
Jordan Ritter af306b52c5 fix(showcase): slow D5 deep sweep cadence to 30min to stop staleness banner flap 2026-06-28 00:15:27 -07:00
Jordan Ritter 400d90e87e fix(showcase): make probe-invoker timeout test deterministic (kill timing flake) 2026-06-27 22:54:01 -07:00
Jordan Ritter 19962ba5c7 fix(showcase): flip stale orchestrator drain test to finish-and-report contract (layer-b cleanup #5743 missed) 2026-06-27 22:40:48 -07:00
Jordan Ritter aaed2dc545 fix(showcase): correct stale drain comments in orchestrator (post-layer-b)
Three stale, pre-layer-(b) prose sites still claimed drain() ABANDONS the
in-flight run and the loop skips queue.report: the drainFleetWorker JSDoc, the
runWorker stop() inline comment, AND the case-worker SIGTERM-handler block.
Layer (b) changed this — on SIGTERM the worker stops claiming, lets its
in-flight cell FINISH within the 90s grace, and REPORTS its real terminal
result via the runAbort/abortedWithoutResult discriminator; only a run that
OVERRUNS the grace is abandoned -> reclaimed by layer (a). Preserved the still-
accurate drainReason=shutdown red-side-emit soft-wind-down + deregister-vs-crash
distinction. Also refreshed the WHY-THIS-ORDER preamble (90s finish-and-report
grace, 180s platform stop / layer-c drainingSeconds). Comments only — no
behavior change.
2026-06-27 22:22:11 -07:00
Jordan Ritter 04177348d1 style(showcase): oxfmt graceful-drain touched files 2026-06-26 23:40:14 -07:00
Jordan Ritter dfb59c3576 fix(showcase): pin same-turn drain race as safe + correct stale requestDrain JSDoc
CR P2 round on layer-(b) graceful worker drain.

P2-A (contested TOCTOU, reconciled empirically): a run that resolves with a
valid result in the same flush the grace setTimeout fires is REPORTED, not
spuriously abandoned. runAbort.abort() lives only in stop()'s Promise.race
TIMEOUT leg, which loses the race once `done` is resolvable — so a finished
run never trips the abortedWithoutResult discriminator. Reviewer crb6 was
correct; the TOCTOU is a non-bug, so no logic changed. Added a deterministic
fake-timer regression pin forcing both the run-completion timer and the grace
timer due in one advance (verified to BITE under an over-aggressive
abandon mutation), plus a clarifying comment at the discriminator.

P2-B: corrected the requestDrain() JSDoc — post-B2 the report-skip,
heartbeat-stop, and driver-cancel key on the grace-expiry signal
runAbort.signal, not stopAbort.signal (which now only stops claiming).
2026-06-26 23:29:34 -07:00
Jordan Ritter 09b78f419a fix(showcase): raise drain grace to bound one cell-job; re-derive composed SIGKILL-window budget
B5 — set the production drain grace T to 90s (was 6s). The grace stopped being
a bare TEARDOWN budget and is now the FINISH-AND-REPORT budget stop() waits for
an in-flight run to finish-and-report (layer b) before firing runAbort -> abandon
-> layer-(a) reclaim. T must BOUND a typical cell-job (so a normal in-flight job
finishes within grace) yet stay SHORTER than the platform SIGTERM->SIGKILL window
with headroom. Sized from in-repo cell-job signal: a single-service cell-job runs
~15s (light e2e-deep) up to ~200s (heavy d6-all-pills under contention); the
per-job lease ceiling is 300s. 90s covers the bulk and stays well under the lease
so a finishing job's lease never lapses. The tail (a job that cannot finish in
grace) falls back to layer (a) — grace is deliberately FINITE. 90s is a defensible
default; B-VAL confirms/retunes from staging p95. Env-overridable via
WORKER_DRAIN_GRACE_MS.

Introduce PLATFORM_STOP_GRACE_MS (180s) documenting the C3 requirement: layer-(c)
must set Railway terminationGracePeriodSeconds = 180 so the composed serial budget
DRAIN_DEREGISTER_TIMEOUT_MS (3s) + DEFAULT_WORKER_DRAIN_GRACE_MS (90s) fits with
>=30s headroom for the health-server-close + pool-shutdown remainder. The
composed-budget test now pins the relation 3s + grace < PLATFORM_STOP_GRACE_MS
(was hardcoded < 10s) and the concrete numbers.

Add a behavioral fake-clock test of the composed-budget invariant: a run that
finishes WITHIN T is reported (finish-and-report); a run still running AT
grace-expiry fires runAbort -> abandons WITHOUT a usable result -> not reported
(layer-a reclaim backstop). The raise also broke the existing wedged-driver
default-grace test (advancing a 90s fake span re-armed the 0ms-yielding heartbeat
into a runaway cascade); fixed by giving it a clock-honoring sleep so the
heartbeat stays quiet across the window and only the grace timer is crossed.
2026-06-26 23:15:30 -07:00
Jordan Ritter 2b741f39df fix(showcase): pin aborted-only drain-red suppression so a finished-on-drain run reports its real terminal result
Post-B2/B3 a run can finish-and-report after a graceful drain (drain no
longer hard-cancels the run; only grace-expiry runAbort does, which is
ctx.abortSignal). The d6 red-suppression already AND-s drainReason with
ctx.abortSignal.aborted and errorClass==="abort", which makes it
aborted-only by construction — a finished run whose legitimate red is a
genuine failure (errorClass!="abort") is NOT suppressed. This adds a
pinning test for that aborted-only contract (a finished goto-error red
under drainReason=shutdown with an un-fired abort signal must still be
reported). Verified RED via a temporary mutation to the over-broad
"drainReason alone" suppression form, GREEN on the real code.
2026-06-26 23:00:38 -07:00
Jordan Ritter f1cc1c8f08 fix(showcase): keep the lease renewing for a finishing-on-drain job so the reaper cannot double-claim it
B3 (layer-b graceful drain): the heartbeat-abort previously keyed on the
DRAIN signal, stopping lease renewal at drain-start. With B2 decoupling drain
from run-abort, a still-finishing job now runs past drain-start until grace
expiry, so killing renewal at drain-start lets the lease lapse mid-finish and
the layer-(a) reaper could reclaim the row out from under the worker
(double-run / report-after-reclaim). Gate the heartbeat-abort on the
grace-expiry signal (runAbortSignal) instead: a finishing job keeps renewing
until it reports terminal; a genuinely-abandoned job (runAbort fired at
grace-expiry) still stops renewing so its lease lapses and the reaper reclaims
it. runAbortSignal defaults to drainSignal, preserving direct-call unit-test
semantics.
2026-06-26 22:54:45 -07:00
Jordan Ritter 0aacd72d66 fix(showcase): finish-and-report a completed run on graceful drain instead of abandoning it
Decouple 'stop claiming new jobs' from 'abort the in-flight run': drain()
no longer fires the run's abortSignal. The run's hard cancel is now a
separate grace-expiry signal (runAbort), so a run seconds from done
finishes within grace and is reported instead of being abandoned to the
sweeper. The abandon break is now conditional on abortedWithoutResult
(runAbort fired at grace-expiry → no usable result); a finished run falls
through to the report path. A wedged run that overruns grace is still cut
(runAbort.abort() in stop()'s timeout leg) and abandons → layer (a)
reclaim is the backstop.

Reconcile the drain tests that encoded the old abort-on-drain coupling to
the new contract (drain no longer fires ctx.abortSignal; the grace-expiry
abort fires at grace, exercised via a short WORKER_DRAIN_GRACE_MS).
2026-06-26 22:49:38 -07:00
Jordan Ritter 5afb474c7f test(showcase): pin desired finish-and-report drain semantics (red) 2026-06-26 22:44:13 -07:00
Jordan Ritter 31abb3eff3 fix(showcase): reclaimable leases — reclaim-wins-until-cap reaper with consecutive-orphan budget (#5738)
## What
Replaces the queue reaper's delete-wins behavior for stale long-expired
orphaned in-flight rows with **reclaim-wins-until-cap**: an orphaned row
is re-queued (not deleted) up to a bounded number of CONSECUTIVE
re-orphans before final claim-delete. This is layer (a) of the
worker-reclamation + graceful-rollover redesign — it makes a worker
bounce mid-column non-lossy (the column's in-flight work is reclaimed by
surviving workers instead of dropped).

## How
- **Reclaim-wins-until-cap** in `runSweepExpired`
(`showcase/harness/src/fleet/queue-client.ts`): stale long-expired
orphaned in-flight rows are re-queued; the prior delete-wins carve-out
is inverted.
- **Consecutive-orphan budget** (`consecutive_orphan_count`, migration
`1779990400`): the reclaim cap (`MAX_RECLAIM_ATTEMPTS=3`) is scoped to
*consecutive* sweeper re-orphans — incremented only on the sweeper
re-queue path, **reset to 0 on terminal `done`/`failed`**. Peer-worker
expired-lease *steals* bump the lifetime `reclaim_count` (retained as a
dashboard diagnostic) but do NOT consume the reclaim budget.
- **Stale-age re-anchoring** (`requeued_at ?? created`, migration
`1779990300`): a reclaimed row's staleness clock restarts so it isn't
immediately re-expired.

## Review
7-agent CR + 7-agent confirmation round → **0 P0 / 0 P1**. Highlights:
- **Red-green GENUINE (empirically re-verified):** P1-A — a long-lived
job with lifetime `reclaim_count=3` but `consecutive_orphan_count=0` is
RE-QUEUED on a fresh orphan (pre-fix: wrongly deleted); P1-B — low
boundary at `MAX-1=2` re-queues, mutation-killed. Both non-tautological
(the test fake bumps `consecutive_orphan_count` only on reclaim, never
on steal; a dedicated pin test confirms steals don't consume budget).
- **No regression / terminates / fails-safe:** 248/248 tests, tsc +
oxlint clean; reclaim loop terminates at 3 consecutive orphans →
claim-delete; queue cannot grow unbounded; SweepResult metrics +
`reclaim_count` dashboard consumers (run-view `jobs.reclaimed`,
family-silence) unbroken.
- **Migration `1779990400`** additive, nullable, default-0;
pre-migration rows evaluate to 0 → safe re-queue path.

## Follow-on (not in this PR)
- Layer (b) graceful worker drain and layer (c) rolling restart are the
remaining phases of the redesign — see the Notion proposal:
https://app.notion.com/p/38b3aa381852817bacf5c9cda1f11cc0

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-26 15:46:05 -07:00
Jordan Ritter d4d363d940 fix(showcase): tighten D4/BE probe gate to a provenance check (catch the BIA-style dead-agent false-pass) (#5737)
## What

Tightens the **D4/BE chat-roundtrip probe gate** from `text.length > 0`
to a **provenance check**. A turn now only counts green if the assistant
text came from the `[data-testid="copilot-assistant-message"]` container
— not from the `<body>` fallback scrape (which picks up static page
chrome).

## Why

This is the gate weakness that **masked the BIA prod outage**: a dead
agent (RUN_STARTED→RUN_FINISHED with zero `TEXT_MESSAGE`) left the page
chrome on screen, the `<body>` scrape returned non-empty text, and the
old `text.length > 0` gate reported **green** while the agent was
actually broken. The new gate threads a `fromAssistantContainer` flag
(set true only on the testid-container read path) so a body-scrape-only
result fails (`!fromAssistantContainer` → red). The same guard is
applied at L4 (tool-rendering).

## Review

**7-agent CR → 0 P0 / 0 P1.** Highlights:
- **Red-green genuine** (empirical): old gate false-passes BIA (1/52
RED); post-fix 52/52. 3-way guard (green /
empty-via-unchanged-length-guard / BIA-via-provenance).
- **Blast radius: no false-fail risk** — all 20 integrations'
`agentic-chat` + `tool-rendering` demos use the shared v2 `CopilotChat`
which mounts the testid unconditionally; aimock `d4/<slug>/chat.json`
fixtures are text-bearing for all 20. Verdict: **ship as a blocking
gate, no soak.**
- **No lost coverage** — the `<body>`-scrape fallback *code* is retained
(diagnostics); only the gate *semantics* narrowed. No D4-probed cell
relied on the fallback to pass.
- **No cross-family regression** — change is confined to
`d4-chat-roundtrip.ts`; other probe drivers untouched; 169/0 tests pass;
tsc clean.

## Notes / follow-ups (non-blocking, P2)
- A *future* probed demo using a `CopilotChatAssistantMessage`
children-render override (which doesn't emit the testid) would
false-fail — worth a one-line note near the fallback comment if that
pattern is ever added to a D4 route.

Part of the 2026-06-26 incident remediation (companion to #5733); this
gate would have surfaced the BIA failure that the dashboard masked.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-26 15:46:02 -07:00
Jordan Ritter c3a13ac988 fix(showcase): scope reclaim cap to consecutive orphans, not lifetime steals (P1-A + P1-B)
P1-A (cap semantics): the MAX_RECLAIM_ATTEMPTS cap was keyed on
`reclaim_count`, a LIFETIME tally bumped by BOTH the sweeper re-queue
path AND the peer-worker expired-lease steal (claim CAS). A long-lived
job that accrues benign peer steals could exhaust its 3-budget and then
get claim-DELETED on its first real orphan rather than re-queued.

Fix: introduce a dedicated `consecutive_orphan_count` column (migration
1779990400) that is bumped ONLY by the sweeper re-queue path in the
fleet-claim release CAS, and reset to 0 on every terminal done|failed
release. The peer-worker steal (claim CAS wasExpiredSteal branch) does
NOT touch this counter. The reaper's cap check now uses
`consecutive_orphan_count` instead of `reclaim_count`. `reclaim_count`
is left intact as the lifetime dashboard diagnostic (jobs.reclaimed).

P1-B (low boundary): adds a test at consecutive_orphan_count = MAX-1
(= 2) asserting the row is RE-QUEUED, not deleted. The off-by-one
mutation `>= MAX` -> `>= MAX-1` causes this test to go RED.

Test-fake honesty: `makeReclaimClaim`'s claimJob now explicitly models
the steal-bump on `reclaim_count` (matching the real hook) while
intentionally NOT bumping `consecutive_orphan_count`, and adds an
explicit pin test confirming steals do not consume the reclaim budget.

JSDoc on MAX_RECLAIM_ATTEMPTS updated to describe the correct semantics:
consecutive re-orphans scoped by sweeper re-queue, reset on terminal.

Red-green proof:
- P1-A RED: revert cap to reclaim_count → "P1-A CAP SCOPE" fails with
  `expect(undefined).toBeDefined()` (job deleted instead of re-queued)
- P1-A GREEN: consecutive_orphan_count cap → test passes (re-queued)
- P1-B RED: mutate `>= MAX` to `>= MAX-1` → low-boundary test fails
- P1-B GREEN: revert mutation → low-boundary test passes

Suite: 145 queue-client + 103 producer = 248 total, all green.
2026-06-26 15:21:41 -07:00
Jordan Ritter 159de7b1ae feat(showcase): reclaimable leases — invert long-expired carve-out to reclaim-wins until cap
Layer (a) of the worker reclamation+rollover redesign: make a worker bounce
non-lossy. The reaper's long-expired carve-out (G1d) used to claim-DELETE an
orphaned in-flight (claimed/running) row whose lease expired beyond its
family's stale window AND whose created-age was past that window —
`reclaimed=0, expiredPending++` — silently dropping work an abrupt bounce
(SIGKILL past grace / OOM / crash) left mid-flight.

Invert it to RECLAIM-WINS-UNTIL-CAP: re-queue the orphan to pending (it
re-runs; idempotent probes make at-least-once safe) until its durable
`reclaim_count` reaches MAX_RECLAIM_ATTEMPTS (3), only then claim-deleting a
row that keeps re-orphaning so a poison job cannot loop forever.

The carve-out existed to dodge an honesty bind: re-queueing a `created`-stale
row emitted a "back in flight" gray the next sweep falsified by claim-deleting
it off the renewal-immune `created` age. Dissolve the bind with a new
`requeued_at` column (migration 1779990300) the release CAS stamps on every
pending re-queue; both stale phases now age off `staleAgeAnchorMs`
(`requeued_at ?? created`), so a reclaimed row is genuinely young again and the
next sweep does not delete it. `reclaim_count` (migration 1779990200) is reused
as the attempt counter — no second tally to drift.

Red-green proven on the real reaper (queue-client.test.ts): a stale-aged
long-expired orphan below the cap goes RED (deleted, reclaimed=0,
expiredPending=1) on delete-wins and GREEN (re-queued, reclaimed=1,
expiredPending=0, requeued_at stamped) on reclaim-wins; plus an attempt-cap
deletion test and a next-sweep no-falsification test.
2026-06-26 15:02:32 -07:00
Jordan Ritter 7364bdfe43 fix(showcase): tighten D4/BE chat gate to require a real assistant turn
The L3 chat-roundtrip gate only asserted `text.length > 0`, which the
`<body>` fallback scrape can satisfy with static page text (nav links,
footer copy, demo blurb) trailing the sent message. When the agent run
finished with ZERO assistant content (RUN_STARTED -> RUN_FINISHED, no
TEXT_MESSAGE) the assistant-message bubble never rendered, so the probe
fell back to scraping <body> and false-PASSED on incidental page chrome —
the mechanism that masked the BIA outage (dead agent reported green).

Track the provenance of the captured response and require it to come from
the [data-testid="copilot-assistant-message"] container (a genuine
assistant turn — the DOM-layer equivalent of RUN_FINISHED + a non-empty
TEXT_MESSAGE), not the body fallback. Apply the same guard to L4 so weather
content must also come from a real assistant turn. The instrumented probe
sends the aimock fixture headers and renders real content into the
container, so this does not affect a working agent; it only rejects the
empty-turn false-pass.

Adds a red-green regression test reproducing the BIA false-pass.
2026-06-26 14:59:55 -07:00
Jordan Ritter ca49bc4bcd test(showcase): re-anchor family-silence grace-edge test to genuinely pin the 2×period boundary 2026-06-26 14:17:22 -07:00
Jordan Ritter edb8f8cbe8 fix(showcase): CR polish — soften family-silence comments, prune tautological tests, a11y label, grace-edge test 2026-06-26 13:55:38 -07:00
Jordan Ritter 6c5fc73cba fix(showcase): grace-window suppresses false family-silence on post-deploy worker bounce
A normal harness deploy rebuilds the shared showcase-harness image and
bounces the pool workers (PR #5715). Immediately after the bounce the
workers re-register, the producers re-arm, and every family is mid-sweep:
lastSuccessAt still points at the pre-bounce success, so it reads stale
against the silence thresholds (banner 2x period, Slack alert 3x period +
3 consecutive ticks). The result was a FALSE "worker family X has not
completed successfully" banner AND Slack family-silence alert during the
expected post-bounce drain window.

Fix: a bounce-keyed grace window. The freshest worker registered_at across
the /api/runs workers strip is the fleet's most-recent bounce instant
(independent of CP boot — a worker can bounce while the CP stays up). While
now - bounce < 2 x period, a family with no success yet is DRAINING, not
silent, so neither the §7.4 banner, the §7.3 cell glyph, nor the §9 Slack
alert flags it. Beyond the window with still-no-success, genuine silence
fires exactly as before.

The determination lives in two surfaces (server monitor for the Slack
alert; client isFamilySilent for the banner + glyph), so both now consume
the same new SSOT field (WorkerView.registeredAt) and the same 2x-period
grace constant, keeping them consistent.

- run-view.ts: project registered_at -> WorkerView.registeredAt (server)
- family-silence-monitor.ts: BOUNCE_GRACE_PERIOD_MULTIPLIER + freshest-bounce
  grace gate, keyed off body.workers
- worker-runs-context.tsx: freshestBounceMs + bounceAtMs grace arg on
  isFamilySilent; banner + cell glyph pass it
- ops-api.ts: WorkerView.registeredAt on the client DTO
2026-06-26 13:41:36 -07:00
Ran Shem Tov b706b9e84c test(showcase): wire a2ui-recovery into the d5/d6 harness fleet
Add a d5-a2ui-recovery probe so the A2UI error-recovery demo runs on every
PR via the d5/d6 fleet harness, not only the manual on-demand workflow.

- New probe d5-a2ui-recovery.ts drives both pills in one session: HEAL
  asserts >=2 newly-mounted declarative-metric tiles and no hard-failure
  card; EXHAUST asserts the "Couldn't generate the UI" card appears and
  no surface paints. Deltas (vs a pre-send baseline) keep the two
  mutually-exclusive negatives correct across the shared session. The
  transient "Retrying..." label is not asserted (timing-flaky).
- Prompts are sent as typed input, keyed per integration slug, mirroring
  each slug's suggestions.ts message verbatim. The recovery prompts are
  unique per slug because the inner render_a2ui calls carry no
  x-aimock-context; a typed message is byte-identical to the pill
  dispatch, so it matches the same fixture. Sending via input (not a
  preFill pill click) lets the runner snapshot its run-lifecycle baseline
  first, avoiding a false done-signal-missing failure.
- Register a2ui-recovery in d5-registry, map it in d5-feature-mapping,
  add its representative fixture, and mirror the mapping in the dashboard
  CATALOG_TO_D5_KEY (kept in lock-step via the drift test).

Verified green locally on both recovery paths: langgraph-python
(backend-owned get_a2ui_tools) and strands (auto-inject middleware).
2026-06-26 18:07:33 +02:00
Jordan Ritter 6b04bb08a9 fix(showcase/harness): reliable data-copilot-running turn-done signal (kill probe false-red flaps) (#5649)
## Summary

Makes the showcase harness probe's turn-done signal **reliable**,
killing the dominant class of dashboard false-red flaps without ever
hiding a real failure.

`waitForTurnComplete` previously relied on a fragile SSE fetch-counter
conjunct that false-reds healthy demos whenever the page-side fetch
wrapper missed the runtime URL/transport. This change makes the
**`data-copilot-running` DOM attribute** (driven directly by the agent
run lifecycle, `RUN_STARTED`→true / `RUN_FINISHED`→false,
transport-independent) the **PRIMARY** done-signal, with the SSE counter
demoted to a **headless-only fallback** (headless demos never render
`CopilotChatView`, so the attribute is absent).

Design (all three preserved — no false-green, no false-red, hangs still
red):
- **Primary signal** = the `data-copilot-running` true→false
**transition** with a **stayed-stopped quiescence window** (a stop must
persist on the same run-start count for `settleMs`; a new sub-run resets
it) — so it cannot complete on an intermediate stop in a multi-step
turn.
- **SSE counter** = headless fallback only; never an OR-trigger when the
DOM signal is present.
- **`done-signal-missing` backstop** (gated on `attrPresent===true` +
`runningNow!==true`) reds a genuine painted-but-never-finished DOM turn
before the hard timeout; headless turns use their full timeout for their
only signal.

## How it was reviewed

A full 4-round `cr-loop` (7 unbiased agents/round + confirmation rounds
+ a Procedure-3 promotion audit) caught and fixed **5 distinct
correctness defects** in the implementation before merge:
- **F1** — SSE OR-trigger could complete a multi-step turn early on an
intermediate stop (false-GREEN), in both the loop and the post-loop
classifier.
- **F2** — the run-start baseline was captured *after* the message send,
killing the primary signal on fast turns (false-RED).
- **F3** — non-atomic double `surfaceReady` read per poll (latent hazard
+ wasted round-trip).
- **F4** — the surface-mount (`completeOnMount`) path had no quiescence
window (false-GREEN on intermediate stop + false-RED on a still-running
gen-UI turn).
- **F5** — the early backstop false-redded slow-but-healthy **headless**
turns (now gated on the DOM signal).

Bidirectional red-green tests for F1–F5 plus a systematic `{DOM,
headless} × {completes, lagging-recovers, genuine-hang} × {text,
surface}` completion/backstop matrix. Full harness unit suite: **3173
passed / 18 skipped / 0 failed**; `tsc --noEmit` clean; lint 0 errors;
build clean.

## Known follow-ups (NOT in this PR — pre-existing / non-blocking)

- **Theoretical edge (not reachable on real or realistically-streamed
turns):** if a run completed within a single synchronous microtask
(zero-duration), the page-side MutationObserver could miss the true edge
while `attrPresent===true` → false-red. Real LLM turns and aimock
realistic-streaming hold the attribute true across many event-loop
ticks, so the observer reliably latches it. A naive "re-add SSE fallback
for DOM-present" fix would reintroduce F1's multi-step false-green, so
it's intentionally not done here.
- **Recommended quick follow-up (latency only, no wrong verdict):**
capture `baselineBannerText` pre-`sendTurnMessage` (mirroring the
run-start/count baselines) so a fast-erroring cold-start turn fast-fails
(#5142) instead of burning the full timeout.
- **Pre-existing sse-interceptor capture/counter internals** (none
load-bearing for the new done-signal; verified STAY_IN_C by the
Procedure-3 audit): page-side counter soft-nav/multi-capture reset,
`__hk_fetchWrapped` pattern reuse + hardcoded fallback, g/y-flag
stateful RegExp, TextDecoder end-of-stream flush, bare-catch
reader-error swallow, framenav payload discard/TOCTOU,
CDP-wallTime-vs-Date.now TTFT, addInitScript/close-listener
re-registration accumulation.

## Test plan

- [x] `pnpm test` (harness) — 3173 passed / 18 skipped / 0 failed
- [x] `tsc --noEmit` exit 0, lint 0 errors, build exit 0
- [ ] Verify on staging that auth / prebuilt-sidebar / claude-sdk-tools
(and other previously-flapping cells) stop false-redding while
genuinely-broken cells stay red

Please review the replay/primary-signal approach. Not auto-merging.
2026-06-24 09:57:13 -07:00
Tyler Slaton a13c3ee663 chore: merge main into PR 5480 2026-06-23 20:50:16 -07:00
Ran Shem Tov 405d78d8e6 fix(showcase): shim __name + crypto.randomUUID in D6 harness page contexts
The D6 fleet worker drives each integration over its insecure Docker
origin (http://<slug>:10000), where crypto.randomUUID is undefined so
hand-rolled headless chats threw and never mounted (sse-missing). tsx/
esbuild also wraps named inner functions in __name(...) calls that leak
into page.evaluate and throw __name is not defined.

Add installBrowserContextShims (init-scripts.ts): a __name no-op helper
and a crypto.randomUUID secure-context polyfill, registered via
addInitScript at document_start of every D6 page; wired into the
d6-all-pills newPage goto path.
2026-06-23 19:46:43 -07:00
Tyler Slaton 75611b272c chore: merge main into PR 5480 2026-06-23 15:32:09 -07:00
Jordan Ritter d0497649c7 test(showcase/harness): turn-done signal completion/backstop matrix
Cover the data-copilot-running turn-done signal in waitForTurnComplete:
true->false transition completion, stayed-stopped quiescence, the
attr-gated early backstop, pre-send run-start baseline, and the
integration wait-for-turn-complete behavior.
2026-06-23 14:32:05 -07:00
Jordan Ritter 494d836005 feat(showcase/harness): reliable data-copilot-running turn-done signal in waitForTurnComplete
Make waitForTurnComplete use the page-side data-copilot-running attribute
as the primary turn-done signal: detect the running true->false transition,
require stayed-stopped quiescence, and gate the early backstop on
attrPresent + runningNow to avoid headless false-RED. Capture a pre-send
run-start baseline so fast turns keep the primary signal alive, read
surfaceReady once per poll, and add computeMaxTurnDurationMs.
2026-06-23 14:31:59 -07:00
Jordan Ritter 1696445da3 feat(showcase/harness): page-side data-copilot-running MutationObserver
Add buildCopilotRunningObserverScript to sse-interceptor.ts and wire it
via addInitScript so the page exposes a data-copilot-running attribute
that the harness can observe for turn-completion signaling.
2026-06-23 14:31:50 -07:00
Alem Tuzlak 5ecdee36b8 feat(bot): pluggable StateStore persistence + cross-platform transcripts
Adds a durable persistence layer for @copilotkit/bot, replacing the
in-memory-only ActionStore with a pluggable StateStore.

- StateStore interface (kv/list/lock/dedup/queue) with a shared
  conformance suite; MemoryStore default plus @copilotkit/bot-store-redis
  and @copilotkit/bot-store-postgres backends.
- createBot({ store }): typed per-thread state via Standard Schema,
  action snapshots persisted through the store, per-conversation turn
  lock (onLockConflict drop|force), and inbound-event dedup keyed on a
  stable eventId. ActionStore is kept as a deprecated alias.
- Cross-platform transcripts (bot.transcripts + identity resolver) with
  age-bounded retention (prune on append + filter on read), and
  runAgent({ transcript: true }) to auto-inject history and capture the
  reply.
- createBot({ components }) re-registers components so durable actions
  re-fire after a restart; restart-durability demo in examples/slack.
- Dedup is marked seen only after the turn lock is acquired, so a turn
  dropped on lock-conflict does not burn its eventId (no lost retries).
- Release lockstep: bot-store-redis/postgres version with bot + bot-ui.
2026-06-23 18:33:38 +02:00
Jordan Ritter a618bdce35 feat(harness): emit cvdiag probe-session boundaries from the d5/d6 probe path
- open a per-feature CvdiagProbeSession for each d5/d6 pill probe
- emit exactly-once probe.exit and failure_classifier per session
- join probe-session output to its run via the X-Test-Id header
- thread cvdiagPbWriter through the orchestrator and CLI runner

Call-Site Enumeration: FAILURE_CLASSIFIER_SET is exported from
cvdiag/probe-session and consumed by d6-all-pills (classifier validation
against the canonical set). The export has no other call sites; any future
classifier addition must update the canonical set in probe-session and the
validation in d6-all-pills together.
2026-06-22 16:42:43 -07:00
Jordan Ritter 6929ad1c33 refactor(harness): extract CvdiagProbeSession into shared cvdiag/probe-session
Behavior-preserving extraction of the CvdiagProbeSession lifecycle from the
d4 chat-roundtrip driver into a shared cvdiag/probe-session module, so the
d5/d6 probe path can reuse the same session boundaries. d4-chat-roundtrip
now imports the extracted session instead of defining it inline.
2026-06-22 16:42:31 -07:00
Jordan Ritter 9a6780b853 fix(showcase/harness): break fleet control-plane import cycle — FLEET_PRODUCER_SCHEDULE_ID TDZ crashed the harness on boot (regression from #5616) 2026-06-22 11:01:48 -07:00
github-actions[bot] 931e046f6d style: auto-fix formatting 2026-06-22 17:35:50 +00:00
Jordan Ritter 1a80d28f4e chore(cvdiag): regenerate schema.json + re-stage TS emitter after attribution-fix union merge
schema.json regenerated from the merged canonical schema.ts (failure_classifier
probe.exit additions UNION backend request.ingress/sse.first_byte/llm.call.*
boundaries + test_id adoption). Per-integration staged schema.ts copies
re-derived via 'showcase cvdiag-stage-ts' so codegen --check and stage --check
are both in sync. No hand-merge of generated artifacts.
2026-06-22 10:31:04 -07:00
Jordan Ritter b7576114dc feat(showcase/harness): on-demand trigger route for ALL probes incl fleet/D6 (OPS_TRIGGER_TOKEN-gated enqueue) 2026-06-22 10:30:23 -07:00
Jordan Ritter 1258c077f4 fix(cvdiag): probe.exit stamps the waitForTurnComplete failure classifier (sse-missing/text-unstable/dom-missing) so reds are labeled in cvdiag 2026-06-22 10:30:14 -07:00
Jordan Ritter e5450f7722 fix(cvdiag): probe records the forwarded X-Test-Id as its cvdiag test_id (sanitizeJoinTestId) so probe↔backend rows join 2026-06-22 10:30:09 -07:00
Jordan Ritter 4925d0fc60 fix(cvdiag): backend adopts inbound x-test-id as cross-layer test_id + emits request.ingress/sse.first_byte/llm.call.* boundaries (TS integrations) 2026-06-22 10:30:09 -07:00
Jordan Ritter 154ffb969c style(cvdiag): oxfmt the new writer-auth TS files + re-stage (preempt auto-format bot) 2026-06-21 12:52:44 -07:00
Jordan Ritter 2efc99fa5b fix(cvdiag): bundle a concrete writer-role PB writer into TS integration backends (auth-with-password→Bearer fetch) so backend events persist; was a type-only no-op seam 2026-06-21 12:49:05 -07:00
Jordan Ritter ca12d09c36 cvdiag: permanent showcase observability subsystem (probe→backend→aimock→edge) (#5591)
## What

Adds **cvdiag** — a permanent, always-available observability subsystem
for the showcase, built to diagnose the red↔green cell flap on the
staging dashboard and to make that diagnosis a dashboard query rather
than a multi-day forensic hunt in the future.

Captures the full request path with `X-Test-Id` correlation across
**probe → backend → aimock → edge**, across every integration
(TypeScript, Python, Java/spring-ai, .NET):
- Per-language backend emitters (canonical + staged/compile-linked
mirrors), all sharing one schema (`schema.json`, closed-world
`additionalProperties:false`).
- CREATE-only writes to two new PocketBase collections: `cvdiag_events`
and `cvdiag_raw_byte_samples` (additive migrations — no existing data
touched).
- An 8-class flap classifier mapping to the observed failure signatures
(`sse-missing` / `text-unstable` / `dom-missing`).
- DEBUG-tier raw-byte capture (secret-scrubbed) and HMAC-guarded A/B
edge-interference detection.

## Why

The runId flap-fix (`cdc1e90e`, 2026-06-09) did **not** fully resolve
the flap — it was still observed 2026-06-19. cvdiag exists so the
*remaining* cause is observed live with full correlation instead of
inferred.

## Safety / enablement

- **Inert by default.** With `CVDIAG_BACKEND_EMITTER` unset the
subsystem performs zero host mutation (no logging-config changes, no
threads/tasks, no stdout) — verified by
`test_cvdiag_inert_when_disabled`. **To accumulate data, set
`CVDIAG_BACKEND_EMITTER=1` on the showcase services.**
- All per-language scrubbers match the canonical `scrubSecrets`
(sk-/base64url, Bearer, colon-less URL userinfo, size-guard) — verified
with real toolchains (vitest / mvn / dotnet).
- Merged latest `main` (only conflict: a clean `.csproj` include union).

## Verification
- harness `tsc --noEmit` ✓ · `src/cvdiag` vitest 251/251 ✓ ·
`cvdiag-stage-ts --check` in-sync ✓
- Java MessageScrubber 17/17 (mvn) ✓ · .NET CvdiagBackend 5/5 (dotnet
sdk:9.0) ✓ · Python emitters 93/93 (3.12) ✓

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-19 20:45:52 -07:00
Tyler Slaton db2fd6539b fix: address merge conflicts and run formatter 2026-06-19 15:51:03 -07:00
Jordan Ritter 478afc0f78 fix(showcase): declarative-gen-ui D6 completes on surface-mount, not text-stability (#5590)
> **DRAFT / WIP — not reviewed, not ready to merge.** Checkpoint per
request. The mandatory 7-agent cr-loop + CI-green gate runs before this
leaves draft. LGP and ADK ship together in this PR.

## Problem (a false-D6 in both directions)
Declarative A2UI demos wire `a2ui.injectA2UITool: true`, so the response
is a rendered `render_a2ui` surface with **no assistant text bubble**.
The D6 conversation-runner's turn-completion gate required the assistant
**text** to stabilize — so on a working declarative demo the run
finished and the dashboard painted, but text never settled →
`waitForTurnComplete` timed out (`reason=text-unstable`) **before the
render assertion ran**.

Result: `langgraph-python:declarative-gen-ui` (the gold standard)
reported **false-RED while rendering correctly** (all 4 pills verified
live on staging), while `google-adk` reported **false-GREEN**.

## Fix
Opt-in `ConversationTurn.completeOnMount` (set only by
`d5-gen-ui-declarative.ts`). For those turns the text-stability
completion conjunct is **replaced** by a surface-mount predicate:
run-finished (sseOk) + a new assistant bubble + the expected declarative
testids **newly mounting**. A non-rendering surface now yields a new
`surface-missing` failure reason (truthful RED). Text-based demos are
byte-for-byte unchanged (opt-in, per-turn).

## Proof (both directions, live D6)
- RED (before): `text-unstable` timeout; dashboard text painted.
- GREEN (after): passes in ~5s; `buildDeclarativeAssertion` actually
runs and verifies testids mount for all 4 pills.
- INTEGRITY: forced a broken render (renamed testids) → test goes
**red** (`surface-missing`). Not "always green now."
- Unit: 89/89 + 3 new (green-on-mount, red-on-surface-missing).

## Scope: LGP + ADK (ship together) — both truthful GREEN
- **LGP** gold-standard cell: false-RED → truthful GREEN (all 4 pills
assert).
- **ADK** realignment: **test-only, complete.** The shared-script fix
auto-applies; verified all 4 ADK pills truthfully GREEN (each surface
mounts from baseline 0 via surface-mount completion). Prior false-green
closed; **no ADK backend gap**.

## Out of scope (someone else's problem)
This shared-script change re-evaluates **every** declarative-gen-ui cell
truthfully. Integrations beyond LGP/ADK that don't actually render will
flip to **truthful RED** — e.g. `langgraph-typescript` pill 2
(team-performance `declarative-data-table` doesn't mount). Those are
real per-demo render gaps for their owners; **not fixed here.**

## Follow-up (not in this PR)
The `render_a2ui` call returns a ~5.9 MB SSE for a ~2 KB surface (LGT
worse) — a separate runtime amplification concern in the
`injectA2UITool:true` middleware path.

## Before ready/merge
- [x] ADK empirical verdict — all 4 pills truthful GREEN, realignment
test-only, no backend gap
- [ ] mandatory cr-loop → zero findings
- [ ] CI green
2026-06-19 14:19:10 -07:00
Jordan Ritter 3f9d8a4826 fix(showcase/harness): opt declarative-gen-ui pills into surface-mount completion
Opts each declarative-gen-ui pill into the new `completeOnMount` turn
completion so these surface-rendering demos are gated on their expected
declarative testids mounting rather than assistant-text stability.
2026-06-19 14:08:47 -07:00
Jordan Ritter 93a65dbb5b fix(showcase/harness): complete tool-rendered turns on surface-mount, not text-stability
Declarative A2UI demos render a surface (mounted testids) with no assistant
text bubble, so the assistant-text-stability completion gate never settled and
timed out on working demos — a false-RED.

This adds an opt-in `completeOnMount` turn-completion path that replaces the
assistant-text-stability conjunct with a surface-mount predicate: run-finished
+ a new assistant bubble + the expected declarative testids newly mounting.
A new `surface-missing` failure reason reports when the run finishes but the
expected surface never mounts. Turns that do not opt in keep the existing
text-stability behavior unchanged.
2026-06-19 14:08:40 -07:00
github-actions[bot] 691c036789 style: auto-fix formatting 2026-06-19 20:54:20 +00:00
Jordan Ritter fada109b72 Merge remote-tracking branch 'origin/main' into blitz/cvdiag-observability/integration
# Conflicts:
#	showcase/integrations/ms-agent-harness-dotnet/agent/BeautifulChatAgent.csproj
2026-06-19 13:51:02 -07:00
Jordan Ritter 0a826bf17f fix(cvdiag): re-stage TS cvdiag emitter from canonical — staged copies were stale (leaked sk-ant-/colon-less URL userinfo + half-missing emit.ts) (M6) 2026-06-19 12:33:00 -07:00
Jordan Ritter c3c7b7908b feat(showcase): cross-env pin-drift probe + Ops routing + bring starters under the image-ref gate 2026-06-19 12:23:23 -07:00