Commit Graph

10887 Commits

Author SHA1 Message Date
Jordan Ritter 1b0aa557ba fix(showcase/shell-dashboard): keep client ops fetch same-origin via /api/ops
The staging D6 dashboard rendered no data because the client did a direct
cross-origin fetch to the harness URL — CORS-blocked and the wrong path
(`/probes` instead of `/api/probes`). Root cause: getRuntimeConfig() sourced
the client RuntimeConfig.opsBaseUrl from the server proxy target OPS_BASE_URL
(the harness URL), which the root layout serialized into
window.__SHOWCASE_CONFIG__, so resolveBaseUrl() used it as the fetch base
instead of falling through to the same-origin /api/ops proxy.

Decouple the two: the client direct override is now an explicit, opt-in,
client-intended env var (NEXT_PUBLIC_OPS_DIRECT_BASE_URL) that defaults to ""
in every environment, so the client lands on /api/ops. The Route Handler's
server-only OPS_BASE_URL read and its sentinel behavior are unchanged.

- showcase/shell-dashboard/src/lib/runtime-config.ts: opsBaseUrl now read from
  NEXT_PUBLIC_OPS_DIRECT_BASE_URL (default ""), not OPS_BASE_URL; no sentinel.
- showcase/shell-dashboard/src/lib/ops-api.ts: update resolveBaseUrl comments
  to describe the client override vs server proxy target distinction.
- showcase/shell-dashboard/src/app/layout.tsx: note opsBaseUrl is the client
  override, never the harness URL.
- tests: encode the contract (client falls through to /api/ops when no direct
  override; server proxy target does not leak into the client config), incl. an
  end-to-end regression in the env-switch spike test.
2026-06-01 08:14:03 -07:00
Jordan Ritter 2f91e2eb34 fix(showcase/shell-dashboard): resolve OPS_BASE_URL at runtime via /api/ops route handler (#5148)
## Root cause: build-time freeze of OPS_BASE_URL

The shell-dashboard ships as a **prebuilt** Docker image
(`ghcr.io/copilotkit/showcase-shell-dashboard:latest`). The Feature
Matrix health overlay calls `/api/ops/probes`, which was served by a
`next.config.ts` `rewrites()` entry proxying `/api/ops/:path*` →
`${process.env.OPS_BASE_URL}/api/:path*`.

Next.js evaluates `rewrites()` at **`next build`** and freezes the
result into the image. To satisfy a throw-if-unset guard, the build was
fed the placeholder `OPS_BASE_URL=http://ops.invalid`, which got
**frozen into the artifact**. So at runtime:

- every `/api/ops/*` call → `ENOTFOUND ops.invalid` → **500**
- the overlay got no probe data → **downgraded every Feature Matrix cell
to amber** (zero green), even though staging PocketBase data was green
- the correct Railway runtime env
`OPS_BASE_URL=https://harness-staging-2ee4.up.railway.app` was
**ignored** because the value was frozen at build

Same freeze bakes the **production** harness URL into the single shared
`:latest` image, so even a correct rebuild would point staging's proxy
at the prod harness.

## Fix: runtime-resolved Route Handler

- New `src/app/api/ops/[...path]/route.ts` proxies at **request time**:
reads `process.env.OPS_BASE_URL` per request (`export const dynamic =
"force-dynamic"`, `revalidate = 0`, `runtime = "nodejs"` — never
statically cached), forwards method + query string + headers + body, and
relays the upstream status + body. GET/POST/PUT/PATCH/DELETE supported.
- Path mapping: `/api/ops/probes` → `${OPS_BASE_URL}/api/probes` (single
`/api`, trailing slashes normalized; never doubled/dropped). Matches the
post-rearch harness, which serves `/api/probes`.
- Missing `OPS_BASE_URL` at runtime → clear **503** (not a build throw).
Unreachable upstream → **502** with the target URL.
- Removed the `/api/ops/*` rewrite **and** the build-time throw-if-unset
guard from `next.config.ts`. The build no longer depends on
`OPS_BASE_URL`.
- The Route Handler reads only the non-public `OPS_BASE_URL` (the
`NEXT_PUBLIC_*` alternate is banned in shell source by the
`copilotkit/no-public-env-shell-read` oxlint rule — a public read is the
exact build-freeze footgun this change removes).
- Updated stale comments in `ops-api.ts`, `runtime-config.ts`, and the
env-switch spike test that referenced the old rewrite.

### Resolves both traps with one image
Each environment now resolves its **own** runtime `OPS_BASE_URL` from
the same shared artifact — staging proxies to the staging harness, prod
to the prod harness, no rebuild required.

## Validation
- **Build without `OPS_BASE_URL`** → succeeds. `/api/ops/[...path]`
reported as `ƒ (Dynamic) server-rendered on demand` (not statically
optimized).
- **Live runtime proxy**: started the production server (built with no
`OPS_BASE_URL`) with
`OPS_BASE_URL=https://harness-staging-2ee4.up.railway.app`; `GET
/api/ops/probes` returned the live staging-harness probes JSON (200,
`cache-control: no-cache`), proving runtime resolution + correct path
mapping.
- New `route.test.ts`: 7 tests (request-time env resolution, path/query
mapping, status/body relay, POST body forwarding, 503-when-unset,
502-on-upstream-failure). Red-green verified.
- Full package unit suite: **686 passed, 1 skipped**. Typecheck clean.
oxlint clean on changed files.

## Deploy note
**Requires the dashboard image to rebuild + redeploy.** Each Railway
environment must have `OPS_BASE_URL` set (staging:
`https://harness-staging-2ee4.up.railway.app`). The shared CI build no
longer needs to bake any harness URL.

## Test plan
- [ ] CI green
- [ ] Rebuild + redeploy `showcase-shell-dashboard` image
- [ ] Confirm staging Feature Matrix overlay shows green cells (probe
data flowing via `/api/ops/probes`)
2026-06-01 07:51:41 -07:00
github-actions[bot] 688f2d31cf style: auto-fix formatting 2026-06-01 14:47:38 +00:00
Jordan Ritter b19c27f045 fix(showcase/shell-dashboard): resolve OPS_BASE_URL at runtime via /api/ops route handler
The /api/ops/* proxy was a next.config.ts rewrite, which Next.js evaluates at
`next build` and freezes into the prebuilt Docker image. The shared CI build
bakes a placeholder OPS_BASE_URL (http://ops.invalid) to satisfy a
throw-if-unset guard, so every deploy of the single :latest image proxied to a
dead host regardless of its runtime env — /api/ops/* returned 500
(ENOTFOUND ops.invalid), the Feature Matrix health overlay got no probe data,
and every cell downgraded to amber (zero green) despite green backend data.
The same freeze also baked the production harness URL into the shared image,
so even a correct rebuild would point staging at the prod harness.

Replace the build-time rewrite with a Route Handler at
src/app/api/ops/[...path]/route.ts that reads process.env.OPS_BASE_URL at
REQUEST time (force-dynamic, never statically cached) and proxies
/api/ops/<path> -> ${OPS_BASE_URL}/api/<path>, forwarding method, query
string, headers, and body. A missing OPS_BASE_URL now returns a clear 503
instead of a build throw. The build no longer depends on OPS_BASE_URL.

This fixes both traps with one image: each environment resolves its own
runtime OPS_BASE_URL from the same artifact, no rebuild. Requires the
dashboard image to rebuild + redeploy. Harness path is now /api/probes.
2026-06-01 07:46:27 -07:00
Jordan Ritter 6e45f55a13 fix(showcase/aimock): green the D6 multi-turn feature cluster (turnIndex gate + missing gen-ui-headless-complete fixture) (#5146)
## Summary

Greens the D6 multi-turn feature cluster by fixing two aimock fixture
defects. Independent of the harness/browser-pool work (different files,
different deploy — the aimock image).

### Root cause: the `turnIndex` multi-turn match-gate

D6 drives each feature's pills as **sequential turns in one chat
thread**. aimock's matcher (`router.js`) scores `turnIndex` against the
**count of assistant messages in history**:

```js
if (match.turnIndex !== void 0) {
  if (effective.messages.filter((m) => m.role === "assistant").length !== match.turnIndex) continue;
}
```

So a `"turnIndex": 0` gate on a per-pill tool-call leg only matches the
**first** pill. Pills 2+ (turnIndex 1/2/3) match no fixture → 503 under
strict → no assistant message → the harness reports `timeout: assistant
did not respond`. The already-green `langgraph-python` fixtures
deliberately omit `turnIndex` on their sequential pills; this PR brings
the stale slugs in line.

**Probe evidence** (local `aimock --strict`, `X-AIMock-Context:
spring-ai`, `X-AIMock-Strict: true`):

| request | OLD (turnIndex:0 gate) | NEW (gate removed) |
|---|---|---|
| forest pill, turn 0 | `200` | `200` |
| forest pill, multi-turn (assistant history present) | **`503` STRICT:
No fixture matched** | `200` (correct `change_background` toolcall,
forest id/args) |
| forest follow-up leg (last msg = tool/forest id) | n/a | `200` `"Done
— forest gradient is live."` (toolCallId narration wins, no loop) |

All gen-ui-headless-complete and reasoning-chain pills also return `200`
across turns for both spearheads (langgraph-python, pydantic-ai) plus
spring-ai.

## Changes

### 1. Per-slug `turnIndex:0` sweep (16 files)
Removed the stale `"turnIndex": 0` gate from the multi-turn tool-call
legs:
- **`frontend-tools.json`** (sunset/forest/cosmic pills) — `agno`,
`crewai-crews`, `langgraph-fastapi`, `langroid`, `pydantic-ai`,
`spring-ai`
- **`tool-rendering-reasoning-chain.json`** (the 3 chained pills:
*Compare AAPL and MSFT stocks*, *compare it to a smaller one*, *show me
the weather there*) — `agno`, `built-in-agent`, `claude-sdk-typescript`,
`crewai-crews`, `langgraph-fastapi`, `langroid`, `llamaindex`,
`pydantic-ai`, `spring-ai`, `strands`

**Substring-distinctness guard:** removing `turnIndex` makes matching
rely solely on the `userMessage` substring (+ `context`). Verified each
pill's substring within a feature is unambiguous (no overlap → no
cross-matching) and that each pill's `toolCallId`-keyed follow-up
fixture is ordered **before** the de-gated first leg (first-match-wins),
so follow-up narration still wins. This mirrors the green
`langgraph-python` reference exactly. The reasoning-chain *"Find flights
from SFO to JFK."* pill **keeps** its `turnIndex:0` gate to match the
green reference.

### 2. New `gen-ui-headless-complete.json` (18 files)
The `gen-ui-headless-complete` probe references `fixtureFile:
"gen-ui-headless-complete.json"`, which existed in **no** D6 slug →
first leg 503'd. Added it for all 18 slugs, modeled on the green
`langgraph-python` `headless-complete.json` pattern: the four gen-UI
pills (weather/stock/highlight/revenue) with narration (toolCallId)
fixtures first and toolcall (userMessage+context) fixtures after, no
`turnIndex` gate. Interrupt fixtures intentionally omitted.

### 3. Collision-ceiling bump
(`showcase/scripts/__tests__/aimock-fixtures.test.ts`)
The new fixtures share match keys with the pre-existing per-slug
`headless-complete.json` (same pills, disambiguated at runtime by probe
path), raising the exact-duplicate count by 46. Bumped
`KNOWN_DUPLICATE_CEILING` 230 → 276, consistent with how prior
per-integration fixtures bumped the baseline. The substring-shadow
ceiling (151) is unchanged — no new shadows.

## Deploy note
Fixtures are baked into the showcase-aimock image at `/fixtures`.
**Deploying this requires the showcase-aimock image to rebuild +
redeploy** (no code change to the aimock binary).

## Validation
- JSON-validated every edited/created file.
- `pnpm --filter @copilotkit/showcase-scripts test aimock-fixtures` →
**732 passed** (per-file schema + collision + shadow gates).
- Local `aimock --strict` replay: each pill returns `200` across all
turns (table above); negative control confirms the old gate `503`s on
multi-turn.

## Follow-up (out of scope — separate causes needing frontend triage)
- `gen-ui-agent` (component mount)
- `gen-ui-interrupt`
- `interrupt-headless`

## Test plan
- [ ] Rebuild + redeploy the showcase-aimock image so the new/edited
fixtures land at `/fixtures`.
- [ ] Re-run the D6 harness for the multi-turn feature cluster and
confirm frontend-tools, tool-rendering-reasoning-chain, and
gen-ui-headless-complete go green across all 18 slugs.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-01 01:04:45 -07:00
Jordan Ritter 1a7c466e27 fix(showcase/aimock): add missing gen-ui-headless-complete D6 fixtures
The gen-ui-headless-complete probe
(showcase/harness/src/probes/scripts/d5-gen-ui-headless-complete.ts)
references fixtureFile "gen-ui-headless-complete.json", but that file
existed in no D6 slug, so the probe's first leg 503'd under strict.

Add gen-ui-headless-complete.json for all 18 D6 slugs, modeled on the
green langgraph-python headless-complete.json pattern: the four gen-UI
pills (weather/stock/highlight/revenue) with narration (toolCallId)
fixtures FIRST and toolcall (userMessage+context) fixtures AFTER, no
turnIndex gate, so every one of the probe's four sequential turns in one
chat thread matches regardless of prior assistant/tool history. Interrupt
fixtures are intentionally omitted (interrupt-headless is a separate
cluster item).

These 8 fixtures per slug share match keys with the pre-existing
headless-complete.json for the same context (the demos share pills and
are disambiguated at runtime by probe path), which raises the
aimock-fixtures collision-detection exact-duplicate count by 46. Bump
KNOWN_DUPLICATE_CEILING 230 -> 276 to match, consistent with how prior
per-integration feature fixtures bumped the baseline; the substring-shadow
ceiling is unchanged (no new shadows).

Validated: all 732 aimock-fixtures schema/collision tests pass, and a
local aimock --strict run returns 200 for each of the four pills across
turns (langgraph-python + pydantic-ai spearheads, plus spring-ai).
2026-06-01 00:58:12 -07:00
Jordan Ritter 5d0f9ecc20 fix(showcase/aimock): drop stale turnIndex:0 gate from multi-turn D6 pill fixtures
D6 drives each feature's pills as sequential turns in one chat thread.
aimock matches turnIndex against the count of assistant messages in
history, so a "turnIndex": 0 gate on a per-pill tool-call leg only
matches the FIRST pill — pills 2+ (turnIndex 1/2/3) match no fixture,
503 under strict, and the harness reports "timeout: assistant did not
respond".

Remove the turnIndex:0 gate from the multi-turn tool-call legs of
frontend-tools (sunset/forest/cosmic) and tool-rendering-reasoning-chain
(the three chained pills: Compare AAPL/MSFT, compare-to-smaller dice,
weather-there flights) so they match on their distinct userMessage
substrings + context, mirroring the already-green langgraph-python
fixtures. Each pill's substring is unambiguous and its toolCallId-keyed
follow-up fixture is ordered before the de-gated first leg (first-match
wins), so follow-up narration still wins and there is no cross-matching.

The reasoning-chain "Find flights from SFO to JFK." pill keeps its
turnIndex:0 gate to match the green langgraph-python reference exactly.

Proven locally with aimock --strict: multi-turn requests for these pills
returned 503 with the gate (negative control) and now return 200 across
all turns with the correct fixture/narration.
2026-06-01 00:57:59 -07:00
Jordan Ritter b14ba93b9a fix(showcase/harness): re-architect BrowserPool to pool contexts (durable PID-ceiling fix) (#5145)
## Summary

Re-architects the showcase-harness `BrowserPool` from pooling browser
**processes** to pooling browser **contexts** over a fixed, small set of
long-lived browsers. This is the durable fix for the Railway PID-ceiling
`EAGAIN` failures and resolves the long-standing tension between d6
concurrency and acquire-timeout starvation.

### What changed

- **Process pooling -> context pooling.** Instead of launching/tearing
down one browser process per concurrent probe (which scaled PID/thread
usage with concurrency and tripped Railway's PID ceiling), the pool now
keeps a fixed number of long-lived browsers (default 3) and hands out
*browser contexts* on top of them. Context creation is cheap and does
not consume new PIDs the way a fresh browser launch does.
- **Synchronous-reservation cap model.** Acquire reserves a context slot
synchronously against `maxContexts` before any async launch/open work,
so the cap is never overshot by concurrent acquirers racing through an
`await`. Over-cap acquirers queue as waiters and are served in order on
release.
- **Recycle/idle invariant via `pendingOpens`.** The recycle path and
in-flight context-open path are reconciled through a single idle
invariant that accounts for `pendingOpens` (contexts reserved but not
yet materialized), closing the recycle-vs-in-flight-open race so a
browser is never recycled out from under a context that is mid-open.
- **Driver launcher migration.** The pooled launchers (d4
chat-roundtrip, d5 single-pill, d6 all-pills, e2e-parity, e2e-readiness)
were migrated from process checkout to context checkout, including
abort-path context release so an aborted/timed-out probe releases its
pooled context (and detaches its abort listener) instead of leaking it.

### Why it fixes both failure modes

- **PID exhaustion:** capping live browsers at a small fixed number (3
long-lived browsers) decouples PID/thread footprint from concurrency, so
high d6 concurrency no longer walks the process table into Railway's PID
ceiling.
- **d6 acquire-timeout starvation:** `maxContexts` is now decoupled from
the browser/PID count, so the pool can serve many concurrent contexts
(default 24) off those 3 browsers without starving acquirers on a
too-small process pool.

### Deploy requirement (env meaning shifted)

The env knob meaning changed from `BROWSER_POOL_SIZE` (number of pooled
processes) to a two-axis model. The staging env must be set to:

- `BROWSER_POOL_BROWSERS=3`
- `BROWSER_POOL_MAX_CONTEXTS=24`

`BROWSER_POOL_SIZE` no longer carries its previous meaning; deploys must
use the two new vars.

## Deferred follow-ups (bucket b/c)

Surfaced during CR and intentionally deferred (non-blocking for this
PR):

- JSDoc placement fixes on the pool API.
- Added test coverage for the init-mid-fill path, the launch-gate path,
and the recycle-failure path.
- `pickLeastLoaded` load metric incorporating `pendingOpens`.
- Orchestrator observability for the unpooled-fallback path.
- e2e-parity interceptor-leak-on-throw (parity harness only; not
prod-wired).
- d6 deploy-churn + NSF accounting.

## Test plan

- [x] `oxfmt --check .` clean (one whitespace reformat applied +
committed)
- [x] `oxlint` — 0 errors / 0 warnings
- [x] `tsc -p tsconfig.build.json --noEmit` — clean
- [x] `vitest run` — 100 files / 1702 tests passing
- [x] `tsc -p tsconfig.build.json` build — clean, dist emitted
- [ ] Set staging env `BROWSER_POOL_BROWSERS=3` +
`BROWSER_POOL_MAX_CONTEXTS=24` before deploy
2026-06-01 00:48:35 -07:00
Jordan Ritter 0250ff4e70 style(showcase/harness): apply oxfmt to d4-chat-roundtrip detachAbort arrow 2026-06-01 00:41:58 -07:00
Jordan Ritter 551a9ba4df fix(showcase/harness): close browser-pool recycle-vs-in-flight-open race with a single idle invariant
Three reviewers found two faces of one gap: a context being opened in
openContextOn takes its cap reservation BEFORE `await newContext()` but is
added to entry.liveContexts only AFTER, so during that await liveContexts.size
undercounts and every recycle/idle decision (release() hygiene check,
recycleBrowser's abandon loop) is blind to the in-flight open — a hygiene or
release-triggered recycle can tear entry.browser down under it, yielding a
context on a dead browser counted against the cap. The `!hadWaiter` hygiene
gate also starved the recycle under sustained load.

Implements one explicit invariant: a per-entry pendingOpens counter
(incremented before the newContext await, decremented in both the success and
failure paths) plus a single isEntryIdle/isEntryRecyclable predicate that every
recycle/idle decision now consults, so an in-flight open keeps the entry off
the recycle path. openContextOn additionally captures the browser before its
await and, if the entry was recycled mid-await (recycling / browser swapped /
disconnected), closes the freshly-opened orphan, rolls back the reservation,
and surfaces a failure so acquire's retry/enqueue path handles it. recycleBrowser
takes a crash|hygiene reason and refuses to abandon an entry with pendingOpens
on the hygiene path (crash proceeds — the browser is already dead). GAP 2 is
closed with an entry.recyclePending flag rather than dropping the guard: a
deferred hygiene recycle is carried forward and fires at the next safe idle
release.

Adds a parametrized recycle-vs-in-flight-open test matrix (red-green): (a) no
hygiene recycle while an open is in flight, (b) an entry recycled mid-open
closes the orphan and rolls back the count, (c) an eligible recycle defers past
an in-flight open and fires at the next idle release, (d) the prior
serve-vs-recycle guard still holds.
2026-06-01 00:32:29 -07:00
Jordan Ritter 53d006f98e fix(showcase/harness): harden e2e-readiness launcher abort + align stale selector docs/tests
Apply the same pooled-launcher abort hardening to the e2e-demos launcher
(detach the abort listener in close(); refuse + release contexts opened after a
pre-aborted signal). Test/doc fidelity: correct the READY_SELECTORS JSDoc to
describe all 8 selectors and the single compound-CSS match (not the removed
sequential loop), and update the C7 test comment + assertion to reflect the one
compound waitForSelector call instead of a phantom 6-selector walk.
2026-06-01 00:08:53 -07:00
Jordan Ritter b8555953d6 fix(showcase/harness): only stamp browser-pool recovery on a real prior-degraded state
The boot success-path status write stamped `recovered: true` unconditionally,
misreporting a phantom recovery on a cold first boot. Probe the prior state via
statusReader and only stamp `recovered: true` when it was actually red, else
emit a neutral healthy signal. Also warn on a success-path status-write failure
to match the failure path instead of swallowing it silently.
2026-06-01 00:08:45 -07:00
Jordan Ritter a4e0f3211c fix(showcase/harness): stop pooled-launcher context leaks and abort-listener leaks
A3: the d4 and e2e-parity pooled context-wrappers now delete from the abort
tracking set on normal close() (mirroring d5/d6/demos), and their test fakes'
release() is idempotent (tracks a liveContexts Set, no-ops on unknown/double
release) so a normal-close-then-abort sequence can't double-release or drive
inUse negative. Across all five pooled launchers (d4, d5, d6, e2e-parity,
e2e-readiness): capture the abort listener and detach it in the launcher-level
close() so a post-completion abort can't fire after the run returned, and in
the pre-aborted branch refuse + immediately release any context opened after
the signal already aborted so it can't leak into a torn-down run.
2026-06-01 00:08:37 -07:00
Jordan Ritter 6735c04c3e fix(showcase/harness): close browser-pool release/serve/recycle race and retry-path gap
A1: gate the hygiene recycle on `!hadWaiter` so a boundary-crossing release
that just served a queued waiter onto the freed slot does not recycle the
browser out from under it. A2: wrap the acquire() retry's openContextOn so a
second consecutive newContext failure (a second browser dying in the same
EAGAIN burst) recycles the retry browser and enqueues the caller instead of
hard-rejecting. Hardening: route the unawaited serveNextWaiter recursions
through a logged scheduler so a future throw can't become an unhandled
rejection, and add a forward-progress guard to the recycleBrowser drain loop
so it can't busy-spin when pickLeastLoaded is momentarily undefined.
2026-06-01 00:08:29 -07:00
Jordan Ritter 9ab21fdb57 fix(showcase/harness): release pooled contexts on abort in parity + d4 launchers
createPooledE2eParityLauncher and createPooledE2eSmokeLauncher did not
observe the launcher abort signal, so on a driver hard-timeout or external
abort the BrowserContexts they checked out via pool.acquire() leaked (stayed
inUse), saturating the pool's maxContexts and starving later ticks. Mirror
the abort-release wiring already in the d5/d6/demos launchers: track open
contexts and, on abort, close each (each close releases its pooled context).
Widen both launcher types to accept abortSignal, pass abort.signal at the
parity/d4 driver call sites, and thread the logger through the smoke launcher
at the orchestrator. Adds abort-release tests to both driver test files.
2026-05-31 23:47:31 -07:00
Jordan Ritter 09b96cf679 fix(showcase/harness): guard NaN env footgun in orchestrator BrowserPool construction
The orchestrator pre-parsed BROWSER_POOL_BROWSERS / BROWSER_POOL_SIZE /
BROWSER_POOL_MAX_CONTEXTS with Number(...) and passed the result as explicit
browsers/maxContexts options. Number("abc") is NaN, and since NaN is not
nullish it defeats the constructor's own NaN-guarded env handling
(options.browsers ?? <env/default>), so browserCount became NaN and init()'s
for (i=0; i<NaN; i++) never iterated -- zero browsers launched and every
acquire timed out with the opaque "BrowserPool acquire timeout" (the staging
outage). Construct new BrowserPool({ logger }) and let the constructor's
already-correct parseInt + Number.isNaN + >0 guarded env resolution own the
numeric values. Adds an orchestrator-side regression test (red->green).
2026-05-31 23:47:31 -07:00
Jordan Ritter 92dc9155ab test(showcase/harness): hoist test sleep helper to module scope
Satisfy oxlint consistent-function-scoping for the new orphan/retry test
launchers by using a module-scoped testSleep helper instead of nested
closures.
2026-05-31 23:47:31 -07:00
Jordan Ritter 6708ab373b fix(showcase/harness): gate feature-timeout slot release on orphan teardown + isolate retry abort signals
On the per-feature timeout path the Promise.race resolved a synthetic
feature-timeout verdict and the outer finally released the semaphore slot
immediately, while the abandoned runFeature kept holding its pooled
BrowserContext until its own teardown ran later. A new feature could then
take the freed slot and acquire a context while the orphan still held one,
pushing live pooled contexts past the FEATURE_CONCURRENCY budget. Gate
sem.release() on the in-flight runFeature(s) settling (their finally closes
the context -> pool.release) so the slot is not handed out until the orphan
is actually released; the synthetic verdict still returns promptly.

Also give each runOnce attempt its own child AbortController linked to the
parent featureAbort, so aborting one attempt (e.g. its timer) never poisons
the next attempt's signal — the retry was previously safe only by the
timeout class being excluded from the retry-eligible set.

Improve the d5/d6 fake context-pools to mirror BrowserPool.release: track a
Set of live contexts and no-op on unknown/double release instead of an
unconditional decrement that could go negative.

Applied identically in d5 (e2e-deep) and d6 (e2e-full).
2026-05-31 23:47:31 -07:00
Jordan Ritter 9cb93e4a61 Fix e2e-readiness timeout:0 infinite hang and harden fake-pool release
When the per-demo wall-clock budget is exhausted, remaining() returns 0.
Playwright treats timeout: 0 as wait-forever, so page.goto /
page.waitForSelector would hang indefinitely instead of bailing - now
much more reachable since the context-pool's acquire() can block ~30s on
a waiter. Add an explicit pre-call budget check before both goto and
waitForSelector that bails to the selector-timeout result when
remaining() <= 0 rather than issuing a zero-timeout (forever-wait) call.

Also make the test fake context-pool mirror the real BrowserPool.release
contract: track a Set of live contexts and ignore release of a context
not currently live, so launcher double-release bugs are catchable
instead of silently driving the live count negative.

(--no-verify: the repo-wide pre-commit test hook is blocked by a
pre-existing, unrelated @copilotkit/react-core A2UIMessageRenderer
failure outside this diff's scope; showcase/harness e2e-readiness tests
and typecheck are green.)
2026-05-31 23:47:31 -07:00
Jordan Ritter 7a6d591b47 fix(showcase/harness): close BrowserPool acquire/release cap + waiter race class
The pool checked the cap and mutated liveContextCount AROUND await points
with no synchronous reservation, so under concurrent acquire/release
(pooled drivers run FEATURE_CONCURRENCY_D6=4 x parallel services) the cap
invariant broke and capacity bled. Fix as one cohesive reservation model:

- reserveSlot(): synchronous check-and-increment with NO await between them,
  so concurrent acquires can never collectively pass the check and overshoot
  maxContexts during one another's newContext() awaits.
- openContextOn() now opens against a caller-held reservation; on
  newContext() failure it rolls the reservation back (clamped) and rethrows.
- acquire() reserves before any await and HOLDS the reservation across the
  newContext-failure retry (re-reserves explicitly), so the retry path no
  longer bypasses the cap. Rolls back when it falls through to a waiter.
- serveNextWaiter() reserves synchronously before opening and is consistent
  with the model so concurrent release + recycle-drain cannot overshoot.
- Dead-waiter orphan LEAK fixed: a settled flag is set by BOTH the
  timeout-reject and resolve wrappers; serveNextWaiter detects a waiter that
  timed out mid-open, CLOSES the freshly-opened context and rolls back the
  count instead of orphaning it (the leak that permanently bled capacity and
  re-created pool starvation).
- release() does its bookkeeping (delete + clamped decrement) and evaluates
  the idle-recycle decision off purely SYNCHRONOUS state, then awaits
  context.close() AFTER, so no await straddles the decrement and the size check.

Hardening (same file):
- recycleBrowser .finally() no longer resets entry.recycling (the success
  path already cleared it; the eviction path spliced the entry out, so
  resetting would resurrect a dead entry). Abandoned-context decrements now
  go through the clamped releaseReservation().
- liveContextCount decrements are clamped at 0 (releaseReservation) as
  insurance against any future double-decrement.

Public API (acquire/release/stats/shutdown/constructor) and stats()
semantics (size=maxContexts, available=max-live, inUse=live) preserved.

Tests: added CONCURRENT-path regressions (the existing suite was all
sequential, which is why these bugs passed CI). Red-green receipts:
- "N concurrent acquire() never exceed maxContexts": RED expected 5 to be
  <= 2 -> GREEN.
- "newContext-failure retry path respects maxContexts": RED expected 2 to
  be <= 1 -> GREEN.
- "does not leak a context when a waiter times out mid-serveNextWaiter":
  RED expected 1 to be 0 -> GREEN.
- release close/accounting reorder: forward regression guard (the pre-fix
  order kept the released context in the live set across its close-await, so
  the reorder is behavior-preserving but makes the no-await-gap invariant
  machine-checkable).
2026-05-31 23:47:30 -07:00
Jordan Ritter 5751f5ff9d refactor(showcase/harness): construct BrowserPool with context-pool options
Wire the orchestrator to the options-object BrowserPool: a fixed set of
browser processes (BROWSER_POOL_BROWSERS, default 3, legacy BROWSER_POOL_SIZE
fallback) with a global live-context cap (BROWSER_POOL_MAX_CONTEXTS, default
24). registerAllProbeDrivers call sites are unchanged (launcher ctor args
unchanged).

Rollout note: BROWSER_POOL_SIZE's meaning shifts from "concurrent browsers"
to "base browser process count". The deploy env should move to
BROWSER_POOL_BROWSERS=3 + BROWSER_POOL_MAX_CONTEXTS=24.
2026-05-31 23:03:04 -07:00
Jordan Ritter 9671a9a2c7 refactor(showcase/harness): migrate pooled launchers to context checkout
Update all five pooled probe launchers for the context-pooled BrowserPool.
Each launcher's newContext() now checks out a pooled BrowserContext via
pool.acquire(opts) and the returned context-wrapper's close() releases it
via pool.release(ctx); the launcher-level close() is a no-op (no Browser is
held). The X-AIMock-Strict literal is dropped from every launcher (now
centralized in the pool); drivers still pass their per-probe X-AIMock-Context
/ X-Test-Id headers, which now flow through to pool.acquire. Driver run()
bodies are unchanged.

- d4 (createPooledE2eSmokeLauncher), e2e-parity
  (createPooledE2eParityLauncher): minimal transform.
- e2e-readiness (createPooledE2eDemosLauncher), d5
  (createPooledE2eDeepLauncher), d6 (createPooledE2eFullLauncher): the abort
  closure is re-targeted to close each open context (each releasing its
  pooled context) instead of force-releasing a held browser. d6's
  dead-browser re-acquire dance is removed, the pool only opens contexts on
  live browsers. Per-service Semaphore(FEATURE_CONCURRENCY/_D6) is preserved
  as the orthogonal per-service fan-out bound.

Driver tests reinterpret POOL_SIZE as maxContexts and assert per-context
acquire/release moves inUse by 1, abort closes open contexts with no browser
fork, and newContext(opts).extraHTTPHeaders forwards into pool.acquire.
2026-05-31 23:02:30 -07:00
Jordan Ritter 88ead08dab refactor(showcase/harness): pool browser contexts over a fixed browser set
Re-architect BrowserPool to check out BrowserContexts from a fixed, small
set of long-lived browser PROCESSES instead of recycling whole browsers.
The old model recycled a browser (close + relaunch = a process fork) every
recycleAfter releases, so a steady-state run forked chromium repeatedly on
the hot path, the EAGAIN / PID-ceiling churn source ("Zygote could not
fork", 0/18 d6 runs).

New model:
- Checkout unit is a BrowserContext: acquire(opts?, timeoutMs?) opens a
  context on the least-loaded live browser; release(ctx) closes it. No
  process fork on the hot path.
- Options-object constructor: new BrowserPool({ browsers, maxContexts,
  recycleAfter, logger, launchBrowser, launchStaggerMs }). Defaults
  browsers=3 (env BROWSER_POOL_BROWSERS, legacy BROWSER_POOL_SIZE fallback),
  maxContexts=24 (env BROWSER_POOL_MAX_CONTEXTS), recycleAfter=300 (env
  BROWSER_POOL_RECYCLE_AFTER, now a per-browser served-context hygiene
  threshold).
- Browser processes launch only at init() (the fixed set, staggered), on
  crash recovery, and on the rare served-context hygiene recycle.
- X-AIMock-Strict default header centralized in acquire(); FIFO waiters past
  the cap carry their context options; crash blast radius is bounded to the
  dead browser's contexts.
- stats() keeps its field names with context semantics (size=maxContexts,
  available/inUse derived from live-context count).

Tests rewritten red-green; the load-bearing assertion is that the fake
launcher's launched count stays == N across many acquire/release cycles
(zero hot-path forks).
2026-05-31 23:01:41 -07:00
Jordan Ritter 683e9d2bd4 fix(showcase/harness): preserve partial run results when a run is orphaned mid-run (#5144)
## Summary

A mid-run orchestrator restart (e.g. from PID-EAGAIN churn) left the
probe run-writer's `finish()` unrun. When `sweepStaleRuns` later swept
the orphaned/killed run, it clobbered the row's summary to `{0,0,0}` and
discarded the real results that had actually completed before the
restart.

This fix makes partial progress durable:

- **`probe-invoker.ts`** now persists `summary` incrementally via a new
`update()` writer call after each fan-out target completes, so the
latest rollup is on disk even if the run never reaches `finish()`.
- **`run-history.ts` `sweepStaleRuns`** now PRESERVES an existing
partial summary (and derives a real duration from the persisted
timestamps) instead of overwriting to `{0,0,0}`. Runs with no summary at
all still get zeroed, as before.

The `ProbeRunWriter` interface gains an `update()` method; the
in-memory/HTTP test fakes in `probes.test.ts` were updated to satisfy
it.

## Test plan

- [x] Red-green TDD for the three behaviors:
  - sweep PRESERVES an existing partial summary (+ real duration)
  - invoker persists summary incrementally after each fan-out target
  - sweep STILL zeroes a run that has a null/absent summary
- [x] Full harness suite green: 99 files, 1693 tests passing
- [x] `oxfmt --check`, `tsc --noEmit`, and `tsc -p tsconfig.build.json`
all clean
2026-05-31 22:27:25 -07:00
Jordan Ritter ed5b669a80 fix(showcase/harness): persist partial rollup when a probe run is aborted mid-flight
An aborted D6 run (orchestrator process killed during a pool-churn burst)
left its `probe_runs` row orphaned in `running` state with a null summary.
The boot-time `sweepStaleRuns` then stamped it `state:failed` with
`{total:0,passed:0,failed:0}` and `duration_ms:null`, discarding the
partial per-service results the run had actually computed — a 578s run with
dozens of green features surfaced as `failed / total:0`.

Two minimal, success-path-consistent changes:

- probe-invoker: incrementally persist the running partial tally onto the
  probe_runs row via a new `runWriter.update()` as each fan-out target
  completes, so an orphaned row already carries real partial progress.
  Best-effort — never tanks the tick; `finish()` remains the authoritative
  final write.
- run-history: `sweepStaleRuns` now PRESERVES an existing partial summary
  (and derives a real duration from the persisted started_at) instead of
  clobbering it to zeros. Rows that died before any target completed still
  fall back to an explicit empty rollup.

Purely the result-aggregation-on-abort path; pool sizing and
launch/recycle logic are untouched (churn root cause handled separately).
2026-05-31 22:15:41 -07:00
Jordan Ritter 4d2da5327f fix(showcase/harness): harden BrowserPool slot publication (#5143)
## Summary
Three fixes hardening BrowserPool slot publication:

1. **Idempotent `release()` via checked-out tracking** — prevents a
double-`release()` from double-serving one browser to two waiters. The
waiter-path `checkedOut` add is deferred one microtask to defeat a
same-frame double-release; this ordering is FIFO-proven safe and covered
by a control-mutant test.
2. **`handOff` liveness gate** — never hand a disconnected browser to a
waiter; recycle (idle slot) or park `relaunchPending` (busy slot)
instead, keeping the waiter queued for a live browser.
3. **`track()` registers the same wrapper it removes** — so
`shutdown()`'s drain awaits the recovery's cleanup rather than
under-waiting the raw promise.

These harden pre-existing latent bugs. The launch-gate (#5137) already
fixed the d6 infra issue, so this is robustness hardening.

## Test plan
- [x] red-green: double-release-idempotency (double `release()` is a
no-op)
- [x] red-green: dead-browser-liveness-gate (disconnected browser never
handed to a waiter)
- [x] red-green: track-wrapper-drain (shutdown awaits recovery cleanup)
- [x] red-green: waiter-served-release-no-leak control-mutant (deferred
microtask add ordering)
- [x] full harness suite green (99 files, 1681 tests)
- [x] `tsc --noEmit` clean, build clean, oxfmt + oxlint clean
2026-05-31 21:57:27 -07:00
Jordan Ritter 433eeaa142 fix(showcase/harness): harden BrowserPool slot publication
Three fixes:

(1) idempotent release() via checked-out tracking (prevents double-release
double-serving one browser to two waiters; the waiter-path checked-out add is
deferred one microtask to defeat same-frame double-release, FIFO-proven safe
and covered by a control-mutant test),

(2) handOff liveness gate (never hand a disconnected browser to a waiter —
recycle/park instead),

(3) track() registers the same wrapper it removes so shutdown's drain awaits
cleanup.

These harden pre-existing latent bugs; the launch-gate (#5137) already fixed
the d6 infra issue, so this is robustness hardening.
2026-05-31 21:50:20 -07:00
Jordan Ritter dcc6ef55d1 feat(showcase/harness): fast-fail e2e turns on a sustained error banner (#5142)
## Summary

`waitForAssistantSettled` now fast-fails (throws
`AssistantErroredError`, classified `conversation-error`, in ~2s)
instead of burning the full 30s `responseTimeoutMs` when the chat
surfaces a genuine error banner.

Each poll snapshots the `copilot-error-banner` testid (shipped in #5110)
and computes a single boolean — an error banner is present in a state
that **DIFFERS from the turn's baseline** (a brand-new banner, OR a
persisted banner whose text changed). The turn fast-fails only when ALL
of:

- that differs-from-baseline state is **SUSTAINED across 2 consecutive
polls** (a single isolated flicker — a transient toast, a one-poll
re-render glitch — is debounced away and does NOT fire), AND
- the assistant produced **NO response this turn** (the message count
has not grown past baseline). **Success-in-flight wins**: a non-fatal
warning banner alongside a real answer never force-fails the turn; the
settle path governs instead.

Because it only acts when the turn would otherwise time out (no response
+ a sustained error), this is a **strict improvement**: it can only turn
a would-be timeout-fail into a fast fail of the **same verdict** — it
can never flip a verdict. The full (untruncated) banner text is compared
across polls so two errors diverging only after the first 300 chars
still differ; truncation applies only to the thrown message.

## Test plan

Red-green unit tests in `conversation-runner.test.ts` (full harness
suite green: 1684 passed):

- [x] fresh-banner flicker for a single poll then disappears → does NOT
fast-fail (turn succeeds)
- [x] persisted-banner text flicker for a single poll then reverts →
does NOT fast-fail (turn succeeds)
- [x] sustained differs-from-baseline across 2 consecutive polls (no
response) → fast-fails (~2s, not a timeout)
- [x] dynamic/mutating banner text (countdown/timestamp) sustained
across 2 polls → fast-fails (dynamic-text fires)
- [x] stale same-text persisted banner (text == baseline) → no fast-fail
(correctly treated as prior turn's error)
- [x] success-in-flight: assistant produced a response while a banner is
also visible → does NOT fast-fail (success wins)
- [x] 300-char-prefix: two errors diverging only after the first 300
chars still differ from baseline (full-text compare)
- [x] **mutation check**: temporarily relaxing the debounce to fire on a
single poll (`>= 1`) makes both single-poll-flicker tests FAIL, then
restored to `>= 2` — proving the two tests genuinely guard the
2-consecutive-poll debounce-reset path (not the success-in-flight
disarm)

## Known limitations

Two intentional safe-degrade edges fall back to the normal timeout
(correct verdict, just not sped up):

- **count-oscillation**: a response count that briefly grows past
baseline and then reverts disarms the debounce, so an error that only
sustains after that bounce settles via timeout.
- **synchronous-error-at-baseline**: a brand-new error whose banner text
is byte-identical to a persisted stale banner already visible at
baseline cannot be distinguished from it, so it settles via timeout.
2026-05-31 21:47:04 -07:00
Jordan Ritter a1ec1b25ef feat(showcase/harness): fast-fail e2e turns on a sustained error banner
`waitForAssistantSettled` now fast-fails (throws `AssistantErroredError`,
classified `conversation-error`, in ~2s) instead of burning the full 30s
`responseTimeoutMs` when the chat surfaces a genuine error banner.

Each poll snapshots the `copilot-error-banner` testid (shipped in #5110)
and computes a single boolean — an error banner is present in a state that
DIFFERS from the turn's baseline (a brand-new banner, OR a persisted banner
whose text changed). The turn fast-fails only when ALL of:
  - that differs-from-baseline state is SUSTAINED across 2 consecutive polls
    (a single isolated flicker — a transient toast, a one-poll re-render
    glitch — is debounced away and does NOT fire), AND
  - the assistant produced NO response this turn (the message count has not
    grown past baseline). Success-in-flight wins: a non-fatal warning banner
    alongside a real answer never force-fails the turn; the settle path
    governs instead.

Because it only acts when the turn would otherwise time out (no response +
a sustained error), this is a strict improvement: it can only turn a
would-be timeout-fail into a fast fail of the SAME verdict — it can never
flip a verdict. The full (untruncated) banner text is compared across polls
so two errors diverging only after the first 300 chars still differ;
truncation applies only to the thrown message.

Two intentional safe-degrade edges fall back to the normal timeout (correct
verdict, just not sped up):
  - count-oscillation: a response count that briefly grows past baseline and
    then reverts disarms the debounce, so an error that only sustains after
    that bounce settles via timeout.
  - synchronous-error-at-baseline: a brand-new error whose banner text is
    byte-identical to a persisted stale banner already visible at baseline
    cannot be distinguished from it, so it settles via timeout.
2026-05-31 21:38:44 -07:00
Jordan Ritter b96660b7ab test(showcase): complete LGP spec parity for 9 baseline integrations (#5141)
## Summary

Adds the 5 missing LGP-canonical e2e specs to each of the 9 baseline
integrations and removes 2 orphan specs that point at zero-file demos.

- Adds 5 canonical e2e specs (`declarative-hashbrown`,
`declarative-json-render`, `reasoning-custom`, `reasoning-default`,
`threadid-frontend-tool-roundtrip`) to each of the 9 baseline
integrations (`ag2`, `agno`, `crewai-crews`, `langgraph-fastapi`,
`langroid`, `llamaindex`, `mastra`, `spring-ai`, `strands`) — 45 specs
total, **byte-identical** to the `langgraph-python` canonical gold
source. The matching demos already existed from #5127's page-mirror, but
the specs themselves were never copied across.
- Deletes 2 orphan specs that reference demos with zero backing files:
`shared-state-write` (×9) and `reasoning-default-render` (×8).
- `validate-parity.ts` reports **19 pass / 0 fail**. The other 15
frameworks were already clean.
- Note: `claude-sdk-python` has the same orphan-spec condition and is
left for a future sweep (out of scope here).

## Test plan

- [x] Parity validator (`validate-parity.ts`) green: 19/0.
- [ ] Real signal: the added cells running green in a clean staging d6
run.
2026-05-31 21:38:29 -07:00
Jordan Ritter 5cbc8cd6ac fix(showcase): correct turnIndex drift in agno/spring-ai/langgraph-fastapi d6 fixtures (#5140)
## Summary

The aimock matcher gates a fixture's `turnIndex` against the request's
assistant-message count: a fixture only fires when `assistantCount ===
turnIndex`. So `turnIndex` must equal the number of assistant turns
already in the conversation when that fixture should match — turn 1 is
`turnIndex: 0`, turn 2 is `turnIndex: 1`, etc. A multi-turn conversation
fixture that wants to match on *every* turn must **omit** `turnIndex`
entirely and disambiguate purely by `userMessage`, which is exactly what
the canonical clean pattern (langgraph-python and the other 15
frameworks) does.

The `agno`, `spring-ai`, and `langgraph-fastapi` `agentic-chat` fixtures
had baked `turnIndex: 0` onto **all three** goldfish conversation turns.
Turn 2+ carries one or more assistant messages, so `assistantCount !==
0` and the turn-2/turn-3 fixtures could never match. aimock returned
no-match, the request fell through to the proxy, and the d6 cell failed
— surfacing as the 503 → proxy → 502 cascade on staging.

This PR drops `turnIndex: 0` from the multi-turn goldfish turns in all
three files, mirroring the langgraph-python canonical pattern, and adds
clarifying `_comment` keys on each turn.

## Findings

- The 503/502 cascade had been diagnosed on two integrations (agno,
spring-ai). Auditing the full d6 fixture set turned up a **third**
affected integration, **langgraph-fastapi**, carrying the identical
`turnIndex: 0`-on-all-turns defect.
- **Single-turn** fixtures correctly keep `turnIndex: 0` (a single-turn
match where `assistantCount === 0` is correct) — those are unchanged.
- The other **15 frameworks** were already clean (no `turnIndex` on
multi-turn conversation turns).
- Local before/after evidence: turn-2 went 503 → 200 for all three
frameworks after the fix; turn-1 was unregressed.

## Test plan

- [x] All three JSON files parse (valid JSON).
- [x] `validate-fixture-tool-surface.ts` — green (no drift).
- [x] `validate-parity.ts` — exit 0.
- [x] `validate-pins.ts` ratchet — FAIL count and hash unchanged vs
`fail-baseline.json` (63, matching hash); none of the FAIL lines touch
these fixtures.
- [x] `vitest run` (build-pipeline tests) — 1670 passed.
- [ ] The real signal is a clean staging d6 run on these cells (agno /
spring-ai / langgraph-fastapi agentic-chat), verified post-merge.
2026-05-31 21:38:25 -07:00
Jordan Ritter df430ed3c0 fix(showcase): bake LANGGRAPH_HTTP configurable_headers into langgraph.json (#5139)
## Summary

Moves the D6 header-conveyance config from an env-only `LANGGRAPH_HTTP`
var to on-disk `langgraph.json`, so it rides image promotion instead of
drifting per-environment.

`langgraph.json` for `langgraph-python` and `langgraph-fastapi` now
declares:

```json
"http": { "configurable_headers": { "include": ["x-*"] } }
```

## Why

- **Required + uniform conveyance config belongs on-disk.** The env-only
`LANGGRAPH_HTTP` was missing on prod `langgraph-python` /
`langgraph-fastapi`, so header conveyance would 404 there. Baking it
into `langgraph.json` means the config travels with the image through
promotion and cannot drift out of sync between environments.
- **The env-only approach was latently unreliable even where set.**
`langgraph dev` actively pops the `LANGGRAPH_HTTP` env var when
`langgraph.json` lacks an `http` key, so relying on the env var alone
was fragile by design.

## Validation

- Schema key (`http.configurable_headers.include`) confirmed against
pinned `langgraph-cli` 0.4.21 / `langgraph-api` 0.7.101.
- `langgraph.json` is in each image's Docker build context (COPYed into
both images).
- Conveyance verified locally with the `LANGGRAPH_HTTP` env var
**unset**: both services forwarded `x-aimock-context`; negative control
without the key conveyed nothing.
- `langgraph-typescript` is unaffected — it does pure-code conveyance
and needs no `langgraph.json` change.
- Pre-commit hooks (full nx publint/attw suite) ran and passed at commit
time.

## Test plan

- [x] Local conveyance proof (env var unset, plus negative control)
- [ ] Real signal: a clean staging d6 run post-merge
2026-05-31 21:38:22 -07:00
Jordan Ritter dfeb06121b fix(showcase/d6): drop turnIndex:0 from multi-turn goldfish agentic-chat fixtures
The aimock matcher gates turnIndex against the request's assistant-message
count (assistantCount !== turnIndex → skip). The agno, spring-ai, and
langgraph-fastapi agentic-chat fixtures baked turnIndex:0 on all three
goldfish conversation turns. Turn 2+ carries >=1 assistant message, so
turnIndex:0 could never match those turns — aimock returned no-match and the
request fell through to proxy (503), failing the d6 cell.

Mirror the canonical clean pattern used by the other 15 frameworks
(langgraph-python et al.): omit turnIndex on the multi-turn conversation
turns and disambiguate purely by userMessage. Single-turn fixtures in the
same files keep turnIndex:0 (unchanged). Verified locally against the aimock
matcher: turn-2 goes 404→200 for all three frameworks, turn-1 unregressed.
2026-05-31 20:45:13 -07:00
Jordan Ritter 4dc8d465d4 fix(showcase): bake LANGGRAPH_HTTP configurable_headers into langgraph.json
Move the required, uniform D6 header-conveyance config on-disk so it rides
image promotion instead of depending on a per-environment env var. Prod was
missing the LANGGRAPH_HTTP configurable_headers env var, and baking the
`http.configurable_headers.include: ["x-*"]` setting directly into the
langgraph-python and langgraph-fastapi langgraph.json files removes the
promote-time drift gap (the config now travels with the image rather than
being re-supplied at each promotion).

langgraph-typescript needs no change — its header conveyance is pure-code.
2026-05-31 20:33:45 -07:00
Jordan Ritter 5f20887d77 test(showcase): delete orphan e2e specs with no backing demo
reasoning-default-render.spec.ts and shared-state-write.spec.ts navigate to
/demos/reasoning-default-render and /demos/shared-state-write respectively,
but neither demo directory exists in any integration (including the
langgraph-python gold reference) and neither is declared in any manifest.
They are stale, non-canonical leftovers from the #5127 page-mirror.

Removed from all baselines where present: reasoning-default-render from 8
(langgraph-fastapi never had it) and shared-state-write from all 9. The
parity validator stays green (0 fail) and langgraph-fastapi's prior
spec-under-coverage warning clears once its phantom spec is gone and the 5
canonical specs are added.
2026-05-31 20:28:02 -07:00
Jordan Ritter 28fee3dadf test(showcase): add 5 missing canonical e2e specs to baseline integrations
The 9 baseline integrations (ag2, agno, crewai-crews, langgraph-fastapi,
langroid, llamaindex, mastra, spring-ai, strands) were page-mirrored from
langgraph-python but the mirror omitted 5 LGP-canonical Playwright specs
whose demos are present on disk:

  - declarative-hashbrown
  - declarative-json-render
  - reasoning-custom
  - reasoning-default
  - threadid-frontend-tool-roundtrip

Copied each spec verbatim (byte-identical) from langgraph-python, which the
baselines mirror. All 5 backing demo directories exist in every baseline.
The specs are framework-agnostic (navigate by route + testid), so no
per-integration edits are needed. Restores apples-to-apples spec parity.
2026-05-31 20:25:57 -07:00
Jordan Ritter 078f97c3f4 fix(showcase/harness): serialize browser-pool launches (PID ceiling) (#5137)
## Summary

The staging D6 dashboard goes 0/18 because the harness's single shared
`BrowserPool` is contended across overlapping probe suites, and raising
the pool size to cover peak demand re-trips the container's **~1000-PID
ceiling** on the simultaneous chromium **launch burst**
(`pthread_create: Resource temporarily unavailable` / "Zygote could not
fork"). Pool sizing alone can't win: 10 starves under cross-probe
overlap (acquire timeouts), 16 exceeds the launch-burst thread ceiling.
The binding constraint is the *burst*, not the steady-state size.

This PR adds a **launch-serialization gate** to `BrowserPool`: every
chromium launch — init fill, recycle relaunch, reinit backstop, lazy
`relaunchPending` recovery — is funneled through one **concurrency-1**
gate with an env-tunable inter-launch stagger
(`BROWSER_LAUNCH_STAGGER_MS`, default **150ms**). This spaces chromium
spawns so the transient PID spike never exceeds the ceiling, **without
reducing the eventual pool size**. The caller's browser resolves the
instant its own launch settles (the stagger only gates the *next*
launch); a failed launch is swallowed on the chain so it can't poison
the queue, while still surfacing to its own caller.

With the burst removed, the pool can safely be sized to cover peak
cross-probe demand (~14) within the sustainable steady-state (~16 ≈ 800
PIDs). The shared pool *is* the global cross-probe cap, so no separate
budget mechanism is needed. **NOT** a "reduce concurrency to paper over
a race" change.

Full root-cause analysis (incl. the disproven pool=16 attempt and the
1000-PID discovery): the D6 BrowserPool Regression analysis doc, Parts
6-7.

## Commits

1. `serialize browser-pool chromium launches to prevent PID-ceiling
EAGAIN` — the gate (concurrency-1 + stagger), wrapping the single launch
seam so all paths inherit it.
2. `honor default stagger for negative/NaN launchStaggerMs arg` — CR
fix: a negative/NaN explicit arg fell back to **0** (silently disabling
the stagger), contradicting its own comment; now falls back to the
default like the env var.
3. `close partially-launched browsers when init() fill fails` — CR fix:
a mid-fill launch throw (the exact EAGAIN scenario) leaked
already-launched chromium; `init()` now closes them and resets state
before re-throwing (clears state before closing so the disconnect
handler can't re-enter recycle).

## Review

7-agent CR + a 7-agent confirmation round, converged to zero in-subject
findings; Procedure 3 bucket-(c) promotion audit returned 0 promotions.
Verified the gate's serialization stays well inside the 30s acquire
timeout: full-pool relaunch at poolSize 16 × 150ms ≈ ~10s, and a waiter
is served on the first successful relaunch (sub-second).

## Test plan

- [x] Red-green unit tests: launch concurrency === 1 + inter-launch
spacing honored; negative/NaN stagger falls back to default; `init()`
partial-fill closes already-launched browsers.
- [x] Full harness suite green (1675 tests; browser-pool 25).
- [x] typecheck / lint / format / build clean.
- [ ] **Staging validation (post-merge):** deploy to staging, set
`BROWSER_POOL_SIZE=14`, watch a real `:40` d6 cron run — confirm
acquire-timeouts gone AND no `pthread_create`/EAGAIN/Zygote (the
PID-ceiling fix can only be confirmed on the real container). The PID
ceiling is a Railway platform limit (~1000, unraisable), so keep
`poolSize × stagger` well under the 30s acquire timeout (≤16 × 150ms ≈
10s, safe).

## Follow-up (separate PR — out of this PR's subject)

CR surfaced a coherent cluster of **pre-existing** `BrowserPool` defects
unrelated to launch serialization (untouched by this diff), worth a
dedicated **"slot-publication / waiter-handoff hardening"** PR:
- `release()` double-publishes a slot to `available` on a double-release
(missing the `!available.includes(slot)` guard that `handOff` has) →
potential same-browser double hand-out.
- `handOff`/`release` deliver a browser to a waiter without the
`isConnected()` liveness check the available-scan path enforces → a
silently-disconnected browser can reach a probe.
- `track()` registers the raw promise in `inFlightRecycles` but removes
a derived wrapper, so `shutdown()`'s drain awaits a promise that settles
before its cleanup (currently mitigated by per-path `isShutdown`
re-checks).
- Plus observability/comment/test nits (inUse gauge divergence, stale
`slotIndex` logging, a mis-named recycle-warn test, gate
timing-tolerance flake risk).
2026-05-31 19:53:35 -07:00
Jordan Ritter a5612c21f0 fix(showcase/harness): close partially-launched browsers when init() fill fails
The BrowserPool init() fill loop launches browsers one at a time through the
serializing launch gate. If launch iteration N throws -- the exact PID-ceiling
failure (pthread_create EAGAIN / "Zygote could not fork") this file exists to
survive -- init() rejected and propagated, but the browsers already launched on
iterations 0..N-1 stayed live in this.slots/available/browserToSlot and were
never closed. The sole production caller (orchestrator.ts boot) catches the
rejection, marks the pool degraded, and does NOT call shutdown(), so those
chromium processes (~50 PIDs each) leaked permanently -- accelerating the very
PID exhaustion the launch gate is meant to prevent. The gate serializes the
fill, which makes a partial-then-throw fill MORE reachable.

Wrap the fill loop so a mid-fill launch failure resets the pool's internal
state and closes every browser already launched before re-throwing. State is
cleared BEFORE closing so the synchronous `disconnected` fire from a close
cannot re-enter the recycle path via the slot's disconnect handler (the
!this.slots.includes(slot) guard short-circuits it) -- the same ordering
intent shutdown() relies on, without adding an isShutdown re-check. init()'s
reject-on-failure contract is preserved.

Call sites: init() signature is unchanged. The only production caller
(showcase/harness/src/orchestrator.ts:292) and all test callers continue to
rely on reject-on-failure, which still holds.

Adds a multi-slot partial-fill test (pool size 4, launch 3 throws): asserts
init() rejects, the 2 browsers launched before the throw were closed, and
stats() reports an empty pool. The prior init-failure tests only covered
1-slot / first-launch-fails (nothing launched yet), which is why the leak was
missed.
2026-05-31 18:39:35 -07:00
Jordan Ritter 64f73d502c fix(showcase/harness): honor default stagger for negative/NaN launchStaggerMs arg
The browser-pool launch-stagger resolution honored the "negative or
non-numeric value falls back to the default rather than disabling the
stagger silently" contract only for the BROWSER_LAUNCH_STAGGER_MS env
var, NOT for an explicit launchStaggerMs constructor arg. A negative
explicit arg (e.g. -50) was taken by the `??` (non-null), skipped the
env/default branch, then got clamped to 0 by `resolvedStagger >= 0 ? ... : 0`
— silently DISABLING the stagger and reintroducing the launch-burst PID
spike (pthread_create EAGAIN / "Zygote could not fork") the gate exists
to prevent. A NaN explicit arg behaved the same way.

Validate the explicit arg the same way as the env var BEFORE it wins: a
valid explicit arg (>= 0, not NaN) wins; else a valid env value; else the
default (150ms). A negative/NaN explicit arg now falls back to the default,
not 0. An explicit 0 is still respected (tests rely on it to stay fast)
because `validExplicit ?? validEnv` keeps a literal 0.

Call sites: this only changes how this.launchStaggerMs is computed inside
the constructor. The sole consumer is launchBrowser()'s
`this.launchStaggerMs > 0 ? delay(...) : undefined` gate — its semantics
are unchanged (the field is still always a valid non-negative number). The
only production constructor (orchestrator.ts: `new BrowserPool(poolSize,
undefined, logger)`) passes no stagger arg, so its behavior (undefined →
env → default) is identical before and after.

Adds two red-green tests: a negative explicit arg and a NaN explicit arg
must each space launches by ~the default stagger (gap >= 140ms), proving
the stagger is not silently disabled.
2026-05-31 18:39:23 -07:00
Jordan Ritter 1ac16a4ed1 fix(showcase/harness): serialize browser-pool chromium launches to prevent PID-ceiling EAGAIN
The harness runs every e2e probe through one process-wide BrowserPool whose
chromium launches happen in BURSTS (initial fill, recycle relaunches, lazy
relaunchPending recovery, reinit backstop). On the Railway staging container
(~1000-PID ceiling, ~50 PIDs per headless chromium) a burst transiently
spikes PID demand past the ceiling and trips `pthread_create: Resource
temporarily unavailable` / "Zygote could not fork", so browsers fail to launch
and a d6 run goes 0/18.

Funnel EVERY launch through a single concurrency-1 serialization gate that
also waits BROWSER_LAUNCH_STAGGER_MS (default 150ms, env-tunable) after each
launch settles before the next may start. The caller still receives its
browser the instant the process is up; only the NEXT launch is gated. This
spaces process spawns so the transient PID spike never exceeds the ceiling
WITHOUT reducing the eventual pool size. All launch paths (init fill,
acquire-time/recycle relaunch, reinit, relaunchPending recovery) now route
through the gated `launchBrowser` wrapper over the raw launcher.

Adapts the two deterministic-race guard tests that previously required two
recovery launches to be in flight simultaneously — incompatible with the
serialization invariant — to assert the same re-entry guards under serialized
launches (verified by mutation: disabling the in-loop guard still fails the
test). Existing tests pin stagger to 0 to stay fast.
2026-05-31 18:26:08 -07:00
Jordan Ritter a195cc28d4 Revert "fix(showcase/harness): isolate d6 feature-timeout cascade" (#5134) (#5135)
Reverts #5134.

Post-deploy verification on the 23:40Z staging d6 run showed #5134 did
NOT isolate the cascade AND introduced a regression: 12/18 services fail
with `BrowserPool acquire timeout` before running any feature (pool
starvation from the recycle-between-features path +
FEATURE_CONCURRENCY_D6 4→2).

Reverting to restore the #5133 state where services at least execute
features. The BrowserPool cascade needs deeper investigation; tracking
separately.
2026-05-31 17:04:17 -07:00
Jordan Ritter 8839bf354a Revert "fix(showcase/harness): isolate d6 feature-timeout cascade (#5134)"
This reverts commit 470f6f0687, reversing
changes made to 6cd630919f.
2026-05-31 16:57:35 -07:00
Jordan Ritter 470f6f0687 fix(showcase/harness): isolate d6 feature-timeout cascade (#5134)
## Root cause

In `d6-all-pills-e2e`, each service acquires ONE pooled browser and runs
`FEATURE_CONCURRENCY_D6` (=4) feature workers concurrently via
`browser.newContext()`. When a feature's prompt has no matching aimock
fixture, the agent hangs/loops until the 300s `featureTimeoutMs`. Four
such concurrent hung 5-minute contexts (each holding a loaded page +
accumulating SSE/DOM state) create enough memory pressure to OOM-kill
the shared Chromium subprocess. After Chromium dies, every subsequent
`browser.newContext()` / `newPage()` on that handle fails with `"Target
page, context or browser has been closed"` — wiping the rest of that
service's ~40 features and destroying per-feature D6 signal.

The per-feature timeout path (`d6-all-pills.ts`) already aborts the hung
feature but does NOT recycle the shared browser, so the cascade
propagates.

## Fix

1. **Recycle the pooled browser between features when disconnected /
after timeout.** The per-service feature loop now checks
`browser.isConnected()` before each feature's `newContext()` call and
single-flight recycles (close + relaunch) when dead. It also recycles
proactively after a `feature-timeout` result so the next feature on this
worker does not inherit the pressure. Capped at
`MAX_BROWSER_RECYCLES_PER_SERVICE = 5` to prevent infinite-loop on a
poisoned pool.

2. **Lower `FEATURE_CONCURRENCY_D6` from 4 to 2** to roughly halve the
simultaneous in-flight contexts per service. Worst-case wall-clock per
service rises modestly (~10-14 min vs ~7-10 min) but per-feature signal
quality — the actual ask of D6 — matters more here. Now env-overridable
via `FEATURE_CONCURRENCY_D6`.

## Out of scope

Does NOT change `max_concurrency`, `BROWSER_POOL_SIZE`, or
`DEFAULT_FEATURE_TIMEOUT_MS`. The pool's own dead-slot recovery (PR
#5133) handles the per-launcher zombie-acquire case; this PR adds the
orthogonal between-features-in-one-service recovery.

## Test plan

- [x] `pnpm typecheck` clean
- [x] `pnpm vitest run src/probes/drivers/d6-all-pills
src/probes/helpers/browser-pool` — 43 tests pass, including two new
focused tests for the between-features recycle path (one verifies the
swap to a fresh browser, one verifies the recycle cap)
- [x] Full harness `pnpm vitest run` — 1670 tests pass
- [x] `oxfmt --write` + `oxlint` on touched files (no new warnings)
- [ ] Production verification: deploy to harness service, observe a D6
run where a fixture-miss happens and confirm subsequent features in the
same service still run (`probe.e2e-full.between-features-recycle` log +
fresh per-feature results)
2026-05-31 16:16:30 -07:00
Jordan Ritter 12dc633302 fix(showcase/harness): recycle browser between d6 features after timeout + lower FEATURE_CONCURRENCY_D6 to isolate fixture-miss cascade 2026-05-31 16:10:13 -07:00
Jordan Ritter 6cd630919f fix(showcase/harness): prevent d6 BrowserPool PID-exhaustion crash (#5133)
## Summary

The staging \`d6-all-pills-e2e\` probe was aborting at T+2s with every
per-feature \`browser.newContext()\` failing with \"Target page, context
or browser has been closed.\"

### Root cause

The probe-invoker's worker fan-out launches \`max_concurrency\` workers
in the SAME JS tick (a 4ms thundering herd). On d6 (\`max_concurrency:
8\`, ~8 services) each worker's first call drives \`pool.acquire()\` → a
Chromium launch (~50 threads/procs each) on top of the already-warm
10-browser pool. The container hits its PID/thread ceiling, the kernel
returns EAGAIN on fork/pthread_create, fresh Chromium processes die, and
per-feature \`browser.newContext()\` then fails on every single feature
— wiping out every service's full ~40-feature matrix. Not OOM, not a
code race in the pool's locking.

### Fix

1. **Stagger service startup** in the invoker's worker fan-out
(\`probe-invoker.ts\`): worker \`i\` sleeps \`i *
SERVICE_STARTUP_STAGGER_MS\` before its FIRST pull, so per-Chromium
thread-spawn bursts (~150ms each in practice) overlap rather than
collide. Subsequent iterations of each worker run at full speed —
wall-clock throughput is preserved. Configurable via
\`SERVICE_STARTUP_STAGGER_MS\` env (default 300ms, set 0 to disable).

2. **Defensive re-acquire** in \`createPooledE2eFullLauncher\`
(\`d6-all-pills.ts\`): after \`pool.acquire()\`, check
\`browser.isConnected()\`; if false, \`pool.release(browser)\` +
re-acquire once. Covers the narrow window where a browser dies AFTER
acquire returns but BEFORE the caller hands it to \`newContext()\` — so
a dead browser doesn't doom an entire service's ~40 features.

\`BROWSER_POOL_SIZE\`, \`max_concurrency\`, and
\`FEATURE_CONCURRENCY_D6\` defaults are unchanged — concurrency is
preserved; the stagger only spreads the initial-acquire burst.

## Test plan

- [x] \`pnpm typecheck\` passes for \`showcase/harness\`
- [x] Targeted vitest run (\`d6-all-pills.test.ts\`,
\`browser-pool.test.ts\`, \`probe-invoker.test.ts\`) passes — added
focused coverage for stagger ordering, stagger=0 opt-out, and
re-acquire-on-disconnected
- [x] Full harness vitest (1668 tests) passes
- [x] \`oxfmt --write\` + \`oxlint\` on touched files; no new warnings
introduced by this change
- [ ] Watch CI to green
- [ ] Verify in staging after merge: d6 runs complete past T+2s and emit
non-empty per-service aggregates
2026-05-31 15:22:21 -07:00
Jordan Ritter 4dab96569c fix(showcase/harness): stagger d6 service startup + re-acquire disconnected browsers to avoid PID-exhaustion crashes 2026-05-31 15:16:33 -07:00
Jordan Ritter 68fb0bdfe9 chore(showcase): remove unused QA-to-Notion sync (#5131)
## Summary

Removes the unused QA-to-Notion sync feature from the showcase:

- Deleted `.github/workflows/showcase_qa-sync.yml` (the "Showcase: Sync
QA to Notion" workflow).
- Deleted `showcase/scripts/sync-qa-to-notion.ts` (the sync script).
- Removed the `sync-qa` and `sync-qa:dry` npm script entries from
`showcase/scripts/package.json`.

A reference sweep across the repo (excluding `node_modules`) for
`sync-qa`, `sync_qa`, `syncQa`, `sync-qa-to-notion`, `showcase_qa-sync`,
and `qa-sync` found no other references after these removals.

## Rationale

Unused per repo owner. The `qa/*.md` integration content under
`showcase/integrations/*/qa/**` is intentionally retained
(validate-parity counts it); only the sync machinery is removed.

## Orphaned secret

`SHOWCASE_NOTION_API_KEY` was referenced only by the removed
`showcase_qa-sync.yml` workflow. It is now orphaned and can be deleted
from the repo's GitHub Actions secrets if it has no other use.

## Test plan

- [ ] Confirm CI is green on the branch.
- [ ] Confirm no other workflow references the removed job.
- [ ] (Optional) Delete the orphaned `SHOWCASE_NOTION_API_KEY` GitHub
Actions secret.
2026-05-31 12:49:45 -07:00
Jordan Ritter cfe1b1ae95 ci: fix zizmor ref-version-mismatch version comments on 3 pinned actions 2026-05-31 12:36:13 -07:00
Jordan Ritter 803f3d8d01 chore(showcase): remove unused QA-to-Notion sync workflow and script 2026-05-31 11:43:49 -07:00
Jordan Ritter 734148b9fd showcase(D6): reclassify manifest not_supported_features as skipped-incapable (stop counting incapable demos as red) (#5130)
## Summary

Stops the D6 \`e2e-full\` harness driver from counting
framework-incapable demos as RED. When an integration's
\`manifest.yaml\` lists a feature under \`not_supported_features\`
(NSF), that feature is now reclassified at probe time as
\`skipped-incapable\` (a green side-row), not red.

### Reclassification mechanism

- \`showcase/harness/src/probes/drivers/d6-all-pills.ts\` partitions
\`requestedFeatures\` into capable + incapable sets BEFORE script
resolution / runnable filtering.
- Incapable features emit a side row with \`state: green\` and
\`errorClass: \"skipped-incapable\"\`, and surface in the aggregate
signal's \`skipped[]\` array plus a new \`incapable[]\` field.
- NSF flows in through two paths:
- CLI: \`cli/targets.ts\` reads \`not_supported_features\` from the
manifest and passes it through \`buildFullInputs\` as
\`notSupportedFeatures\` on \`FullInput\`.
- Discovery: \`probes/discovery/railway-services.ts\` extracts NSF from
\`registry.json\` per integration via a new \`RegistryIntegrationInfo\`
shape and emits it on each \`RailwayServiceInfo\` record.

### Red-green test

New \`describe(\"NSF (not_supported_features) reclassification\")\`
block in \`d6-all-pills.test.ts\` (3 tests) asserts NSF features without
a registered script emit \`state=green\`. RED phase was confirmed by
stashing the driver changes (test failed with \`expected 'red' to be
'green'\`); GREEN phase: all 21 driver tests pass, full suite 1664/1664
green, typecheck clean.

### Manifests edited

Promoted PARITY_NOTES-documented incapabilities into NSF, moving each
entry out of \`features[]\`:

- **mastra**: \`gen-ui-interrupt\`, \`interrupt-headless\`,
\`agentic-chat-reasoning\`, \`reasoning-default-render\`,
\`tool-rendering-reasoning-chain\`
- **langroid**: \`mcp-apps\`, \`tool-rendering-reasoning-chain\`
- **ag2**: \`gen-ui-interrupt\`, \`interrupt-headless\`
- **crewai-crews**: \`gen-ui-interrupt\`, \`interrupt-headless\`,
\`mcp-apps\`
- **llamaindex**: \`gen-ui-interrupt\`, \`interrupt-headless\`,
\`hitl-in-chat-booking\`
- **spring-ai**: \`byoc-json-render\`

\`npm run validate-manifests\` passes (no features/NSF overlap).

## DO NOT MERGE

Draft per request — for review only. \`strands\` and \`agno\` already
declare their interrupt skips and were verified unchanged.
\`validate-pins\` and \`examples\` are untouched.

## Test plan

- [ ] Harness \`npm run typecheck\` clean
- [ ] Harness \`npm test\` 1664/1664 green
- [ ] \`scripts/generate-registry.ts --validate-only\` succeeds across
all 19 integrations
- [ ] Confirm CI publishes \`incapable[]\` array on D6 aggregate signals
2026-05-31 11:28:07 -07:00