## Problem
The `format` job sets `ref: github.head_ref` but leaves `repository:`
unset, so it checks the head branch out of the **base** repo. For fork
PRs the head branch only exists on the fork, so checkout fails before
formatting even runs.
Seen on #5099 (fork PR, branch `fix/5072`):
```
##[error]A branch or tag with the name 'fix/5072' could not be found
```
## Fix
Resolve `repository` to the head repo on PRs, matching `ref`:
```yaml
ref: ${{ github.event_name == 'pull_request' && github.head_ref || github.ref }}
repository: ${{ github.event_name == 'pull_request' && github.event.pull_request.head.repo.full_name || github.repository }}
```
Same-repo PRs are unchanged (head repo == base repo), so the auto-format
push-back still works.
The format job checked out github.head_ref without setting repository,
so it defaulted to the base repo. For fork PRs the head branch only
exists on the fork, making checkout fail with 'a branch or tag with the
name <branch> could not be found' (e.g. PR #5099). Resolve repository to
the head repo on pull_request events; same-repo PRs are unchanged, so
the same-repo-guarded auto-format push-back still works.
The Feature Matrix rendered every D6 cell in an integration's column from
the aggregate PocketBase row `d6:<slug>` (red whenever ANY cell fails),
so genuinely-green cells (e.g. langgraph-python/voice) rendered red. The
harness already emits per-cell `d6:<slug>/<featureType>` rows over the same
featureType keyspace as D5.
Resolve D6 per-cell through the same CATALOG_TO_D5_KEY bridge D5 uses,
mirroring resolveD5/resolveD5Row exactly: per-sub-row stale-green→degraded
fold before the worst-state fold, and strict handling of a missing mapped
sub-row (no-data unless a present sub-row is red). The aggregate `d6:<slug>`
row no longer drives per-cell rendering. Updated the two memo comparators
(composed-cell, unified-cell) to diff per-cell D6 sub-keys, and the stale
doc-comments that described D6 as an integration-level aggregate.
Introduces a third SDK in the reference docs alongside React v2 and v1:
- reference-items.ts: add the 'core' version, generalize root-vs-nested
routing, add 'types'/'enums' subdirs + categories, recognize a core/
slug prefix (literal strip), and emit its static params
- reference-version-selector.tsx: relabel the picker as an SDK switch
(React v2 / React v1 / Core (TypeScript)), import ReferenceVersion from
reference-items, give listbox options role=option/aria-selected
- app/reference/page.tsx: rename 'API Reference' to 'Overview' and add a
'Choose your SDK' card chooser
## Summary
A new showcase at `examples/showcases/a2ui-pdf-analyst` - an agent that
answers questions about a PDF by composing UI from your own design
system.
PDF text is extracted client-side via `pdfjs-dist` and inlined into the
message using CopilotKit's multimodal attachment support.
`multimodal_middleware.py` patches `ag-ui-langgraph` so it survives
serialization and arrives intact at the agent.
Two rendering strategies on the same 21-component catalog:
- `/fixed` : you author the layout once as a JSON tree. The agent
extracts KPIs, trend, and table rows and fills the slots via
`update_data_model` ops. One LLM call per turn, predictable,
brand-locked.
- `/dynamic` : no pre-written layout. The agent calls `query_pdf` to
extract a structured answer, then `generate_a2ui` — which spawns a
secondary LLM bound to a `render_a2ui` shim via `tool_choice`, forcing
structured component output that becomes A2UI `create_surface` +
`update_components` ops. A stat question lands as a StatCard; a
breakdown becomes a DonutChart; a research paper becomes Heading +
Callout + BulletList.
- `/catalog` : all 21 components live, filterable by group.
Frontend: Next.js 16 · React 19 · Tailwind v4 ·
`@copilotkit/a2ui-renderer` · `pdfjs-dist`.
Backend: Python 3.12 · FastAPI · LangChain · LangGraph.
Demo video (with all use cases), sample PDFs with prompts and full
architecture in `README.md`.
## Test plan
- [ ] `pnpm install` completes (`uv sync` runs via postinstall)
- [ ] `pnpm dev` boots web on :3000 and agent on :8123
- [ ] `/fixed` — attach Apple Q4 PDF, ask `Render the dashboard.` →
KPIs, trend chart, donut, and table all render. Click a scope chip →
re-renders without re-uploading.
- [ ] `/dynamic` with Apple Q4 PDF — `What was net income last quarter?`
→ StatCard, `Break iPhone vs Mac vs iPad as a donut.` → DonutChart
- [ ] `/dynamic` with Tesla Q3 PDF — `Plot quarterly production against
deliveries as a scatter chart` → ScatterChart
- [ ] `/catalog` — all 21 components visible, group filters work
## What
Disables the "CopilotKit Unlicensed" watermark in the Angular SDK. When
no
CopilotCloud license key is provided, the SDK no longer:
- injects the fixed-position `copilotkit-license-watermark` element into
the DOM, and
- logs the "License Required" console warning.
## How
The watermark is **not deleted** — it's gated behind a single
`LICENSE_WATERMARK_ENABLED` flag (now `false`) in
`packages/angular/src/lib/license-watermark.ts`.
`ensureLicenseWatermark()`
early-returns when disabled, and `provideCopilotKit()` skips the console
warning on the same flag. Re-enabling is a one-line flip back to `true`.
The `licenseKey` option and its `X-CopilotCloud-Public-Api-Key` header
injection are unchanged — supplying a valid key still sets the Cloud
header.
## Tests
Updated the two specs that asserted the active-watermark behavior to
assert
the now-dormant behavior (no watermark element, no console warning).
Full
Angular suite passes (48 tests).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Gate the Angular SDK's "CopilotKit Unlicensed" watermark and "License
Required" console warning behind a LICENSE_WATERMARK_ENABLED flag set to
false. The watermark implementation is retained for easy re-enablement;
the licenseKey option and its X-CopilotCloud-Public-Api-Key header
injection are unchanged.
## Summary
Follow-on to #5148. The staging D6 dashboard rendered **no data** (empty
Ops tab, no green cells in the matrix overlay) even though the backend
was healthy. Browser-console ground truth showed the **client** doing a
direct cross-origin fetch to
`https://harness-staging-2ee4.up.railway.app/probes` — CORS-blocked
(harness has no `Access-Control-Allow-Origin`) and the wrong path
(`/probes` vs `/api/probes`).
### Root cause
`getRuntimeConfig()` sourced the client `RuntimeConfig.opsBaseUrl` from
the **server proxy target** `OPS_BASE_URL` (the harness URL on staging)
and `layout.tsx` serialized it into `window.__SHOWCASE_CONFIG__`. So
`ops-api.ts:resolveBaseUrl()` used the harness URL as the fetch base and
bypassed the same-origin `/api/ops` Route Handler that #5148 added. The
Route Handler reads `process.env.OPS_BASE_URL` itself, so the server
proxy worked — but the client never called it.
### Fix
Decouple the two env surfaces:
- **Server proxy target** `OPS_BASE_URL` — server-only, read by the
`/api/ops/[...path]` Route Handler; **never** injected into the client.
- **Client direct override** `NEXT_PUBLIC_OPS_DIRECT_BASE_URL` — opt-in
escape hatch for direct cross-origin dev calls; **defaults to `""`** so
the client falls through to the same-origin `/api/ops` proxy.
With the override empty (the normal production/staging case), the
browser fetches `/api/ops/probes` (same-origin) → Route Handler →
`${OPS_BASE_URL}/api/probes`. No CORS, correct path.
## Test plan
- [x] Red→green regression tests: `opsBaseUrl` defaults to `""` even
when `OPS_BASE_URL` is set; sourced only from
`NEXT_PUBLIC_OPS_DIRECT_BASE_URL`; empty override → client uses
`/api/ops/probes`, never the harness directly; server proxy target does
NOT leak into the client config.
- [x] `tsc --noEmit` clean; `vitest run` 48 files / 693 passing; `next
build` green (incl. `/api/ops/[...path]` dynamic route).
- [ ] Post-merge: dashboard image rebuilds → staging redeploy → verify
cells render green and Ops tab populates.
The staging D6 dashboard rendered no data because the client did a direct
cross-origin fetch to the harness URL — CORS-blocked and the wrong path
(`/probes` instead of `/api/probes`). Root cause: getRuntimeConfig() sourced
the client RuntimeConfig.opsBaseUrl from the server proxy target OPS_BASE_URL
(the harness URL), which the root layout serialized into
window.__SHOWCASE_CONFIG__, so resolveBaseUrl() used it as the fetch base
instead of falling through to the same-origin /api/ops proxy.
Decouple the two: the client direct override is now an explicit, opt-in,
client-intended env var (NEXT_PUBLIC_OPS_DIRECT_BASE_URL) that defaults to ""
in every environment, so the client lands on /api/ops. The Route Handler's
server-only OPS_BASE_URL read and its sentinel behavior are unchanged.
- showcase/shell-dashboard/src/lib/runtime-config.ts: opsBaseUrl now read from
NEXT_PUBLIC_OPS_DIRECT_BASE_URL (default ""), not OPS_BASE_URL; no sentinel.
- showcase/shell-dashboard/src/lib/ops-api.ts: update resolveBaseUrl comments
to describe the client override vs server proxy target distinction.
- showcase/shell-dashboard/src/app/layout.tsx: note opsBaseUrl is the client
override, never the harness URL.
- tests: encode the contract (client falls through to /api/ops when no direct
override; server proxy target does not leak into the client config), incl. an
end-to-end regression in the env-switch spike test.
## Root cause: build-time freeze of OPS_BASE_URL
The shell-dashboard ships as a **prebuilt** Docker image
(`ghcr.io/copilotkit/showcase-shell-dashboard:latest`). The Feature
Matrix health overlay calls `/api/ops/probes`, which was served by a
`next.config.ts` `rewrites()` entry proxying `/api/ops/:path*` →
`${process.env.OPS_BASE_URL}/api/:path*`.
Next.js evaluates `rewrites()` at **`next build`** and freezes the
result into the image. To satisfy a throw-if-unset guard, the build was
fed the placeholder `OPS_BASE_URL=http://ops.invalid`, which got
**frozen into the artifact**. So at runtime:
- every `/api/ops/*` call → `ENOTFOUND ops.invalid` → **500**
- the overlay got no probe data → **downgraded every Feature Matrix cell
to amber** (zero green), even though staging PocketBase data was green
- the correct Railway runtime env
`OPS_BASE_URL=https://harness-staging-2ee4.up.railway.app` was
**ignored** because the value was frozen at build
Same freeze bakes the **production** harness URL into the single shared
`:latest` image, so even a correct rebuild would point staging's proxy
at the prod harness.
## Fix: runtime-resolved Route Handler
- New `src/app/api/ops/[...path]/route.ts` proxies at **request time**:
reads `process.env.OPS_BASE_URL` per request (`export const dynamic =
"force-dynamic"`, `revalidate = 0`, `runtime = "nodejs"` — never
statically cached), forwards method + query string + headers + body, and
relays the upstream status + body. GET/POST/PUT/PATCH/DELETE supported.
- Path mapping: `/api/ops/probes` → `${OPS_BASE_URL}/api/probes` (single
`/api`, trailing slashes normalized; never doubled/dropped). Matches the
post-rearch harness, which serves `/api/probes`.
- Missing `OPS_BASE_URL` at runtime → clear **503** (not a build throw).
Unreachable upstream → **502** with the target URL.
- Removed the `/api/ops/*` rewrite **and** the build-time throw-if-unset
guard from `next.config.ts`. The build no longer depends on
`OPS_BASE_URL`.
- The Route Handler reads only the non-public `OPS_BASE_URL` (the
`NEXT_PUBLIC_*` alternate is banned in shell source by the
`copilotkit/no-public-env-shell-read` oxlint rule — a public read is the
exact build-freeze footgun this change removes).
- Updated stale comments in `ops-api.ts`, `runtime-config.ts`, and the
env-switch spike test that referenced the old rewrite.
### Resolves both traps with one image
Each environment now resolves its **own** runtime `OPS_BASE_URL` from
the same shared artifact — staging proxies to the staging harness, prod
to the prod harness, no rebuild required.
## Validation
- **Build without `OPS_BASE_URL`** → succeeds. `/api/ops/[...path]`
reported as `ƒ (Dynamic) server-rendered on demand` (not statically
optimized).
- **Live runtime proxy**: started the production server (built with no
`OPS_BASE_URL`) with
`OPS_BASE_URL=https://harness-staging-2ee4.up.railway.app`; `GET
/api/ops/probes` returned the live staging-harness probes JSON (200,
`cache-control: no-cache`), proving runtime resolution + correct path
mapping.
- New `route.test.ts`: 7 tests (request-time env resolution, path/query
mapping, status/body relay, POST body forwarding, 503-when-unset,
502-on-upstream-failure). Red-green verified.
- Full package unit suite: **686 passed, 1 skipped**. Typecheck clean.
oxlint clean on changed files.
## Deploy note
**Requires the dashboard image to rebuild + redeploy.** Each Railway
environment must have `OPS_BASE_URL` set (staging:
`https://harness-staging-2ee4.up.railway.app`). The shared CI build no
longer needs to bake any harness URL.
## Test plan
- [ ] CI green
- [ ] Rebuild + redeploy `showcase-shell-dashboard` image
- [ ] Confirm staging Feature Matrix overlay shows green cells (probe
data flowing via `/api/ops/probes`)
The /api/ops/* proxy was a next.config.ts rewrite, which Next.js evaluates at
`next build` and freezes into the prebuilt Docker image. The shared CI build
bakes a placeholder OPS_BASE_URL (http://ops.invalid) to satisfy a
throw-if-unset guard, so every deploy of the single :latest image proxied to a
dead host regardless of its runtime env — /api/ops/* returned 500
(ENOTFOUND ops.invalid), the Feature Matrix health overlay got no probe data,
and every cell downgraded to amber (zero green) despite green backend data.
The same freeze also baked the production harness URL into the shared image,
so even a correct rebuild would point staging at the prod harness.
Replace the build-time rewrite with a Route Handler at
src/app/api/ops/[...path]/route.ts that reads process.env.OPS_BASE_URL at
REQUEST time (force-dynamic, never statically cached) and proxies
/api/ops/<path> -> ${OPS_BASE_URL}/api/<path>, forwarding method, query
string, headers, and body. A missing OPS_BASE_URL now returns a clear 503
instead of a build throw. The build no longer depends on OPS_BASE_URL.
This fixes both traps with one image: each environment resolves its own
runtime OPS_BASE_URL from the same artifact, no rebuild. Requires the
dashboard image to rebuild + redeploy. Harness path is now /api/probes.
## Summary
Greens the D6 multi-turn feature cluster by fixing two aimock fixture
defects. Independent of the harness/browser-pool work (different files,
different deploy — the aimock image).
### Root cause: the `turnIndex` multi-turn match-gate
D6 drives each feature's pills as **sequential turns in one chat
thread**. aimock's matcher (`router.js`) scores `turnIndex` against the
**count of assistant messages in history**:
```js
if (match.turnIndex !== void 0) {
if (effective.messages.filter((m) => m.role === "assistant").length !== match.turnIndex) continue;
}
```
So a `"turnIndex": 0` gate on a per-pill tool-call leg only matches the
**first** pill. Pills 2+ (turnIndex 1/2/3) match no fixture → 503 under
strict → no assistant message → the harness reports `timeout: assistant
did not respond`. The already-green `langgraph-python` fixtures
deliberately omit `turnIndex` on their sequential pills; this PR brings
the stale slugs in line.
**Probe evidence** (local `aimock --strict`, `X-AIMock-Context:
spring-ai`, `X-AIMock-Strict: true`):
| request | OLD (turnIndex:0 gate) | NEW (gate removed) |
|---|---|---|
| forest pill, turn 0 | `200` | `200` |
| forest pill, multi-turn (assistant history present) | **`503` STRICT:
No fixture matched** | `200` (correct `change_background` toolcall,
forest id/args) |
| forest follow-up leg (last msg = tool/forest id) | n/a | `200` `"Done
— forest gradient is live."` (toolCallId narration wins, no loop) |
All gen-ui-headless-complete and reasoning-chain pills also return `200`
across turns for both spearheads (langgraph-python, pydantic-ai) plus
spring-ai.
## Changes
### 1. Per-slug `turnIndex:0` sweep (16 files)
Removed the stale `"turnIndex": 0` gate from the multi-turn tool-call
legs:
- **`frontend-tools.json`** (sunset/forest/cosmic pills) — `agno`,
`crewai-crews`, `langgraph-fastapi`, `langroid`, `pydantic-ai`,
`spring-ai`
- **`tool-rendering-reasoning-chain.json`** (the 3 chained pills:
*Compare AAPL and MSFT stocks*, *compare it to a smaller one*, *show me
the weather there*) — `agno`, `built-in-agent`, `claude-sdk-typescript`,
`crewai-crews`, `langgraph-fastapi`, `langroid`, `llamaindex`,
`pydantic-ai`, `spring-ai`, `strands`
**Substring-distinctness guard:** removing `turnIndex` makes matching
rely solely on the `userMessage` substring (+ `context`). Verified each
pill's substring within a feature is unambiguous (no overlap → no
cross-matching) and that each pill's `toolCallId`-keyed follow-up
fixture is ordered **before** the de-gated first leg (first-match-wins),
so follow-up narration still wins. This mirrors the green
`langgraph-python` reference exactly. The reasoning-chain *"Find flights
from SFO to JFK."* pill **keeps** its `turnIndex:0` gate to match the
green reference.
### 2. New `gen-ui-headless-complete.json` (18 files)
The `gen-ui-headless-complete` probe references `fixtureFile:
"gen-ui-headless-complete.json"`, which existed in **no** D6 slug →
first leg 503'd. Added it for all 18 slugs, modeled on the green
`langgraph-python` `headless-complete.json` pattern: the four gen-UI
pills (weather/stock/highlight/revenue) with narration (toolCallId)
fixtures first and toolcall (userMessage+context) fixtures after, no
`turnIndex` gate. Interrupt fixtures intentionally omitted.
### 3. Collision-ceiling bump
(`showcase/scripts/__tests__/aimock-fixtures.test.ts`)
The new fixtures share match keys with the pre-existing per-slug
`headless-complete.json` (same pills, disambiguated at runtime by probe
path), raising the exact-duplicate count by 46. Bumped
`KNOWN_DUPLICATE_CEILING` 230 → 276, consistent with how prior
per-integration fixtures bumped the baseline. The substring-shadow
ceiling (151) is unchanged — no new shadows.
## Deploy note
Fixtures are baked into the showcase-aimock image at `/fixtures`.
**Deploying this requires the showcase-aimock image to rebuild +
redeploy** (no code change to the aimock binary).
## Validation
- JSON-validated every edited/created file.
- `pnpm --filter @copilotkit/showcase-scripts test aimock-fixtures` →
**732 passed** (per-file schema + collision + shadow gates).
- Local `aimock --strict` replay: each pill returns `200` across all
turns (table above); negative control confirms the old gate `503`s on
multi-turn.
## Follow-up (out of scope — separate causes needing frontend triage)
- `gen-ui-agent` (component mount)
- `gen-ui-interrupt`
- `interrupt-headless`
## Test plan
- [ ] Rebuild + redeploy the showcase-aimock image so the new/edited
fixtures land at `/fixtures`.
- [ ] Re-run the D6 harness for the multi-turn feature cluster and
confirm frontend-tools, tool-rendering-reasoning-chain, and
gen-ui-headless-complete go green across all 18 slugs.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
The gen-ui-headless-complete probe
(showcase/harness/src/probes/scripts/d5-gen-ui-headless-complete.ts)
references fixtureFile "gen-ui-headless-complete.json", but that file
existed in no D6 slug, so the probe's first leg 503'd under strict.
Add gen-ui-headless-complete.json for all 18 D6 slugs, modeled on the
green langgraph-python headless-complete.json pattern: the four gen-UI
pills (weather/stock/highlight/revenue) with narration (toolCallId)
fixtures FIRST and toolcall (userMessage+context) fixtures AFTER, no
turnIndex gate, so every one of the probe's four sequential turns in one
chat thread matches regardless of prior assistant/tool history. Interrupt
fixtures are intentionally omitted (interrupt-headless is a separate
cluster item).
These 8 fixtures per slug share match keys with the pre-existing
headless-complete.json for the same context (the demos share pills and
are disambiguated at runtime by probe path), which raises the
aimock-fixtures collision-detection exact-duplicate count by 46. Bump
KNOWN_DUPLICATE_CEILING 230 -> 276 to match, consistent with how prior
per-integration feature fixtures bumped the baseline; the substring-shadow
ceiling is unchanged (no new shadows).
Validated: all 732 aimock-fixtures schema/collision tests pass, and a
local aimock --strict run returns 200 for each of the four pills across
turns (langgraph-python + pydantic-ai spearheads, plus spring-ai).
D6 drives each feature's pills as sequential turns in one chat thread.
aimock matches turnIndex against the count of assistant messages in
history, so a "turnIndex": 0 gate on a per-pill tool-call leg only
matches the FIRST pill — pills 2+ (turnIndex 1/2/3) match no fixture,
503 under strict, and the harness reports "timeout: assistant did not
respond".
Remove the turnIndex:0 gate from the multi-turn tool-call legs of
frontend-tools (sunset/forest/cosmic) and tool-rendering-reasoning-chain
(the three chained pills: Compare AAPL/MSFT, compare-to-smaller dice,
weather-there flights) so they match on their distinct userMessage
substrings + context, mirroring the already-green langgraph-python
fixtures. Each pill's substring is unambiguous and its toolCallId-keyed
follow-up fixture is ordered before the de-gated first leg (first-match
wins), so follow-up narration still wins and there is no cross-matching.
The reasoning-chain "Find flights from SFO to JFK." pill keeps its
turnIndex:0 gate to match the green langgraph-python reference exactly.
Proven locally with aimock --strict: multi-turn requests for these pills
returned 503 with the gate (negative control) and now return 200 across
all turns with the correct fixture/narration.
## Summary
Re-architects the showcase-harness `BrowserPool` from pooling browser
**processes** to pooling browser **contexts** over a fixed, small set of
long-lived browsers. This is the durable fix for the Railway PID-ceiling
`EAGAIN` failures and resolves the long-standing tension between d6
concurrency and acquire-timeout starvation.
### What changed
- **Process pooling -> context pooling.** Instead of launching/tearing
down one browser process per concurrent probe (which scaled PID/thread
usage with concurrency and tripped Railway's PID ceiling), the pool now
keeps a fixed number of long-lived browsers (default 3) and hands out
*browser contexts* on top of them. Context creation is cheap and does
not consume new PIDs the way a fresh browser launch does.
- **Synchronous-reservation cap model.** Acquire reserves a context slot
synchronously against `maxContexts` before any async launch/open work,
so the cap is never overshot by concurrent acquirers racing through an
`await`. Over-cap acquirers queue as waiters and are served in order on
release.
- **Recycle/idle invariant via `pendingOpens`.** The recycle path and
in-flight context-open path are reconciled through a single idle
invariant that accounts for `pendingOpens` (contexts reserved but not
yet materialized), closing the recycle-vs-in-flight-open race so a
browser is never recycled out from under a context that is mid-open.
- **Driver launcher migration.** The pooled launchers (d4
chat-roundtrip, d5 single-pill, d6 all-pills, e2e-parity, e2e-readiness)
were migrated from process checkout to context checkout, including
abort-path context release so an aborted/timed-out probe releases its
pooled context (and detaches its abort listener) instead of leaking it.
### Why it fixes both failure modes
- **PID exhaustion:** capping live browsers at a small fixed number (3
long-lived browsers) decouples PID/thread footprint from concurrency, so
high d6 concurrency no longer walks the process table into Railway's PID
ceiling.
- **d6 acquire-timeout starvation:** `maxContexts` is now decoupled from
the browser/PID count, so the pool can serve many concurrent contexts
(default 24) off those 3 browsers without starving acquirers on a
too-small process pool.
### Deploy requirement (env meaning shifted)
The env knob meaning changed from `BROWSER_POOL_SIZE` (number of pooled
processes) to a two-axis model. The staging env must be set to:
- `BROWSER_POOL_BROWSERS=3`
- `BROWSER_POOL_MAX_CONTEXTS=24`
`BROWSER_POOL_SIZE` no longer carries its previous meaning; deploys must
use the two new vars.
## Deferred follow-ups (bucket b/c)
Surfaced during CR and intentionally deferred (non-blocking for this
PR):
- JSDoc placement fixes on the pool API.
- Added test coverage for the init-mid-fill path, the launch-gate path,
and the recycle-failure path.
- `pickLeastLoaded` load metric incorporating `pendingOpens`.
- Orchestrator observability for the unpooled-fallback path.
- e2e-parity interceptor-leak-on-throw (parity harness only; not
prod-wired).
- d6 deploy-churn + NSF accounting.
## Test plan
- [x] `oxfmt --check .` clean (one whitespace reformat applied +
committed)
- [x] `oxlint` — 0 errors / 0 warnings
- [x] `tsc -p tsconfig.build.json --noEmit` — clean
- [x] `vitest run` — 100 files / 1702 tests passing
- [x] `tsc -p tsconfig.build.json` build — clean, dist emitted
- [ ] Set staging env `BROWSER_POOL_BROWSERS=3` +
`BROWSER_POOL_MAX_CONTEXTS=24` before deploy
Three reviewers found two faces of one gap: a context being opened in
openContextOn takes its cap reservation BEFORE `await newContext()` but is
added to entry.liveContexts only AFTER, so during that await liveContexts.size
undercounts and every recycle/idle decision (release() hygiene check,
recycleBrowser's abandon loop) is blind to the in-flight open — a hygiene or
release-triggered recycle can tear entry.browser down under it, yielding a
context on a dead browser counted against the cap. The `!hadWaiter` hygiene
gate also starved the recycle under sustained load.
Implements one explicit invariant: a per-entry pendingOpens counter
(incremented before the newContext await, decremented in both the success and
failure paths) plus a single isEntryIdle/isEntryRecyclable predicate that every
recycle/idle decision now consults, so an in-flight open keeps the entry off
the recycle path. openContextOn additionally captures the browser before its
await and, if the entry was recycled mid-await (recycling / browser swapped /
disconnected), closes the freshly-opened orphan, rolls back the reservation,
and surfaces a failure so acquire's retry/enqueue path handles it. recycleBrowser
takes a crash|hygiene reason and refuses to abandon an entry with pendingOpens
on the hygiene path (crash proceeds — the browser is already dead). GAP 2 is
closed with an entry.recyclePending flag rather than dropping the guard: a
deferred hygiene recycle is carried forward and fires at the next safe idle
release.
Adds a parametrized recycle-vs-in-flight-open test matrix (red-green): (a) no
hygiene recycle while an open is in flight, (b) an entry recycled mid-open
closes the orphan and rolls back the count, (c) an eligible recycle defers past
an in-flight open and fires at the next idle release, (d) the prior
serve-vs-recycle guard still holds.
Apply the same pooled-launcher abort hardening to the e2e-demos launcher
(detach the abort listener in close(); refuse + release contexts opened after a
pre-aborted signal). Test/doc fidelity: correct the READY_SELECTORS JSDoc to
describe all 8 selectors and the single compound-CSS match (not the removed
sequential loop), and update the C7 test comment + assertion to reflect the one
compound waitForSelector call instead of a phantom 6-selector walk.
The boot success-path status write stamped `recovered: true` unconditionally,
misreporting a phantom recovery on a cold first boot. Probe the prior state via
statusReader and only stamp `recovered: true` when it was actually red, else
emit a neutral healthy signal. Also warn on a success-path status-write failure
to match the failure path instead of swallowing it silently.
A3: the d4 and e2e-parity pooled context-wrappers now delete from the abort
tracking set on normal close() (mirroring d5/d6/demos), and their test fakes'
release() is idempotent (tracks a liveContexts Set, no-ops on unknown/double
release) so a normal-close-then-abort sequence can't double-release or drive
inUse negative. Across all five pooled launchers (d4, d5, d6, e2e-parity,
e2e-readiness): capture the abort listener and detach it in the launcher-level
close() so a post-completion abort can't fire after the run returned, and in
the pre-aborted branch refuse + immediately release any context opened after
the signal already aborted so it can't leak into a torn-down run.
A1: gate the hygiene recycle on `!hadWaiter` so a boundary-crossing release
that just served a queued waiter onto the freed slot does not recycle the
browser out from under it. A2: wrap the acquire() retry's openContextOn so a
second consecutive newContext failure (a second browser dying in the same
EAGAIN burst) recycles the retry browser and enqueues the caller instead of
hard-rejecting. Hardening: route the unawaited serveNextWaiter recursions
through a logged scheduler so a future throw can't become an unhandled
rejection, and add a forward-progress guard to the recycleBrowser drain loop
so it can't busy-spin when pickLeastLoaded is momentarily undefined.
createPooledE2eParityLauncher and createPooledE2eSmokeLauncher did not
observe the launcher abort signal, so on a driver hard-timeout or external
abort the BrowserContexts they checked out via pool.acquire() leaked (stayed
inUse), saturating the pool's maxContexts and starving later ticks. Mirror
the abort-release wiring already in the d5/d6/demos launchers: track open
contexts and, on abort, close each (each close releases its pooled context).
Widen both launcher types to accept abortSignal, pass abort.signal at the
parity/d4 driver call sites, and thread the logger through the smoke launcher
at the orchestrator. Adds abort-release tests to both driver test files.
The orchestrator pre-parsed BROWSER_POOL_BROWSERS / BROWSER_POOL_SIZE /
BROWSER_POOL_MAX_CONTEXTS with Number(...) and passed the result as explicit
browsers/maxContexts options. Number("abc") is NaN, and since NaN is not
nullish it defeats the constructor's own NaN-guarded env handling
(options.browsers ?? <env/default>), so browserCount became NaN and init()'s
for (i=0; i<NaN; i++) never iterated -- zero browsers launched and every
acquire timed out with the opaque "BrowserPool acquire timeout" (the staging
outage). Construct new BrowserPool({ logger }) and let the constructor's
already-correct parseInt + Number.isNaN + >0 guarded env resolution own the
numeric values. Adds an orchestrator-side regression test (red->green).
Satisfy oxlint consistent-function-scoping for the new orphan/retry test
launchers by using a module-scoped testSleep helper instead of nested
closures.
On the per-feature timeout path the Promise.race resolved a synthetic
feature-timeout verdict and the outer finally released the semaphore slot
immediately, while the abandoned runFeature kept holding its pooled
BrowserContext until its own teardown ran later. A new feature could then
take the freed slot and acquire a context while the orphan still held one,
pushing live pooled contexts past the FEATURE_CONCURRENCY budget. Gate
sem.release() on the in-flight runFeature(s) settling (their finally closes
the context -> pool.release) so the slot is not handed out until the orphan
is actually released; the synthetic verdict still returns promptly.
Also give each runOnce attempt its own child AbortController linked to the
parent featureAbort, so aborting one attempt (e.g. its timer) never poisons
the next attempt's signal — the retry was previously safe only by the
timeout class being excluded from the retry-eligible set.
Improve the d5/d6 fake context-pools to mirror BrowserPool.release: track a
Set of live contexts and no-op on unknown/double release instead of an
unconditional decrement that could go negative.
Applied identically in d5 (e2e-deep) and d6 (e2e-full).
When the per-demo wall-clock budget is exhausted, remaining() returns 0.
Playwright treats timeout: 0 as wait-forever, so page.goto /
page.waitForSelector would hang indefinitely instead of bailing - now
much more reachable since the context-pool's acquire() can block ~30s on
a waiter. Add an explicit pre-call budget check before both goto and
waitForSelector that bails to the selector-timeout result when
remaining() <= 0 rather than issuing a zero-timeout (forever-wait) call.
Also make the test fake context-pool mirror the real BrowserPool.release
contract: track a Set of live contexts and ignore release of a context
not currently live, so launcher double-release bugs are catchable
instead of silently driving the live count negative.
(--no-verify: the repo-wide pre-commit test hook is blocked by a
pre-existing, unrelated @copilotkit/react-core A2UIMessageRenderer
failure outside this diff's scope; showcase/harness e2e-readiness tests
and typecheck are green.)
The pool checked the cap and mutated liveContextCount AROUND await points
with no synchronous reservation, so under concurrent acquire/release
(pooled drivers run FEATURE_CONCURRENCY_D6=4 x parallel services) the cap
invariant broke and capacity bled. Fix as one cohesive reservation model:
- reserveSlot(): synchronous check-and-increment with NO await between them,
so concurrent acquires can never collectively pass the check and overshoot
maxContexts during one another's newContext() awaits.
- openContextOn() now opens against a caller-held reservation; on
newContext() failure it rolls the reservation back (clamped) and rethrows.
- acquire() reserves before any await and HOLDS the reservation across the
newContext-failure retry (re-reserves explicitly), so the retry path no
longer bypasses the cap. Rolls back when it falls through to a waiter.
- serveNextWaiter() reserves synchronously before opening and is consistent
with the model so concurrent release + recycle-drain cannot overshoot.
- Dead-waiter orphan LEAK fixed: a settled flag is set by BOTH the
timeout-reject and resolve wrappers; serveNextWaiter detects a waiter that
timed out mid-open, CLOSES the freshly-opened context and rolls back the
count instead of orphaning it (the leak that permanently bled capacity and
re-created pool starvation).
- release() does its bookkeeping (delete + clamped decrement) and evaluates
the idle-recycle decision off purely SYNCHRONOUS state, then awaits
context.close() AFTER, so no await straddles the decrement and the size check.
Hardening (same file):
- recycleBrowser .finally() no longer resets entry.recycling (the success
path already cleared it; the eviction path spliced the entry out, so
resetting would resurrect a dead entry). Abandoned-context decrements now
go through the clamped releaseReservation().
- liveContextCount decrements are clamped at 0 (releaseReservation) as
insurance against any future double-decrement.
Public API (acquire/release/stats/shutdown/constructor) and stats()
semantics (size=maxContexts, available=max-live, inUse=live) preserved.
Tests: added CONCURRENT-path regressions (the existing suite was all
sequential, which is why these bugs passed CI). Red-green receipts:
- "N concurrent acquire() never exceed maxContexts": RED expected 5 to be
<= 2 -> GREEN.
- "newContext-failure retry path respects maxContexts": RED expected 2 to
be <= 1 -> GREEN.
- "does not leak a context when a waiter times out mid-serveNextWaiter":
RED expected 1 to be 0 -> GREEN.
- release close/accounting reorder: forward regression guard (the pre-fix
order kept the released context in the live set across its close-await, so
the reorder is behavior-preserving but makes the no-await-gap invariant
machine-checkable).
Wire the orchestrator to the options-object BrowserPool: a fixed set of
browser processes (BROWSER_POOL_BROWSERS, default 3, legacy BROWSER_POOL_SIZE
fallback) with a global live-context cap (BROWSER_POOL_MAX_CONTEXTS, default
24). registerAllProbeDrivers call sites are unchanged (launcher ctor args
unchanged).
Rollout note: BROWSER_POOL_SIZE's meaning shifts from "concurrent browsers"
to "base browser process count". The deploy env should move to
BROWSER_POOL_BROWSERS=3 + BROWSER_POOL_MAX_CONTEXTS=24.
Update all five pooled probe launchers for the context-pooled BrowserPool.
Each launcher's newContext() now checks out a pooled BrowserContext via
pool.acquire(opts) and the returned context-wrapper's close() releases it
via pool.release(ctx); the launcher-level close() is a no-op (no Browser is
held). The X-AIMock-Strict literal is dropped from every launcher (now
centralized in the pool); drivers still pass their per-probe X-AIMock-Context
/ X-Test-Id headers, which now flow through to pool.acquire. Driver run()
bodies are unchanged.
- d4 (createPooledE2eSmokeLauncher), e2e-parity
(createPooledE2eParityLauncher): minimal transform.
- e2e-readiness (createPooledE2eDemosLauncher), d5
(createPooledE2eDeepLauncher), d6 (createPooledE2eFullLauncher): the abort
closure is re-targeted to close each open context (each releasing its
pooled context) instead of force-releasing a held browser. d6's
dead-browser re-acquire dance is removed, the pool only opens contexts on
live browsers. Per-service Semaphore(FEATURE_CONCURRENCY/_D6) is preserved
as the orthogonal per-service fan-out bound.
Driver tests reinterpret POOL_SIZE as maxContexts and assert per-context
acquire/release moves inUse by 1, abort closes open contexts with no browser
fork, and newContext(opts).extraHTTPHeaders forwards into pool.acquire.
Re-architect BrowserPool to check out BrowserContexts from a fixed, small
set of long-lived browser PROCESSES instead of recycling whole browsers.
The old model recycled a browser (close + relaunch = a process fork) every
recycleAfter releases, so a steady-state run forked chromium repeatedly on
the hot path, the EAGAIN / PID-ceiling churn source ("Zygote could not
fork", 0/18 d6 runs).
New model:
- Checkout unit is a BrowserContext: acquire(opts?, timeoutMs?) opens a
context on the least-loaded live browser; release(ctx) closes it. No
process fork on the hot path.
- Options-object constructor: new BrowserPool({ browsers, maxContexts,
recycleAfter, logger, launchBrowser, launchStaggerMs }). Defaults
browsers=3 (env BROWSER_POOL_BROWSERS, legacy BROWSER_POOL_SIZE fallback),
maxContexts=24 (env BROWSER_POOL_MAX_CONTEXTS), recycleAfter=300 (env
BROWSER_POOL_RECYCLE_AFTER, now a per-browser served-context hygiene
threshold).
- Browser processes launch only at init() (the fixed set, staggered), on
crash recovery, and on the rare served-context hygiene recycle.
- X-AIMock-Strict default header centralized in acquire(); FIFO waiters past
the cap carry their context options; crash blast radius is bounded to the
dead browser's contexts.
- stats() keeps its field names with context semantics (size=maxContexts,
available/inUse derived from live-context count).
Tests rewritten red-green; the load-bearing assertion is that the fake
launcher's launched count stays == N across many acquire/release cycles
(zero hot-path forks).
## Summary
A mid-run orchestrator restart (e.g. from PID-EAGAIN churn) left the
probe run-writer's `finish()` unrun. When `sweepStaleRuns` later swept
the orphaned/killed run, it clobbered the row's summary to `{0,0,0}` and
discarded the real results that had actually completed before the
restart.
This fix makes partial progress durable:
- **`probe-invoker.ts`** now persists `summary` incrementally via a new
`update()` writer call after each fan-out target completes, so the
latest rollup is on disk even if the run never reaches `finish()`.
- **`run-history.ts` `sweepStaleRuns`** now PRESERVES an existing
partial summary (and derives a real duration from the persisted
timestamps) instead of overwriting to `{0,0,0}`. Runs with no summary at
all still get zeroed, as before.
The `ProbeRunWriter` interface gains an `update()` method; the
in-memory/HTTP test fakes in `probes.test.ts` were updated to satisfy
it.
## Test plan
- [x] Red-green TDD for the three behaviors:
- sweep PRESERVES an existing partial summary (+ real duration)
- invoker persists summary incrementally after each fan-out target
- sweep STILL zeroes a run that has a null/absent summary
- [x] Full harness suite green: 99 files, 1693 tests passing
- [x] `oxfmt --check`, `tsc --noEmit`, and `tsc -p tsconfig.build.json`
all clean
An aborted D6 run (orchestrator process killed during a pool-churn burst)
left its `probe_runs` row orphaned in `running` state with a null summary.
The boot-time `sweepStaleRuns` then stamped it `state:failed` with
`{total:0,passed:0,failed:0}` and `duration_ms:null`, discarding the
partial per-service results the run had actually computed — a 578s run with
dozens of green features surfaced as `failed / total:0`.
Two minimal, success-path-consistent changes:
- probe-invoker: incrementally persist the running partial tally onto the
probe_runs row via a new `runWriter.update()` as each fan-out target
completes, so an orphaned row already carries real partial progress.
Best-effort — never tanks the tick; `finish()` remains the authoritative
final write.
- run-history: `sweepStaleRuns` now PRESERVES an existing partial summary
(and derives a real duration from the persisted started_at) instead of
clobbering it to zeros. Rows that died before any target completed still
fall back to an explicit empty rollup.
Purely the result-aggregation-on-abort path; pool sizing and
launch/recycle logic are untouched (churn root cause handled separately).
## Summary
Three fixes hardening BrowserPool slot publication:
1. **Idempotent `release()` via checked-out tracking** — prevents a
double-`release()` from double-serving one browser to two waiters. The
waiter-path `checkedOut` add is deferred one microtask to defeat a
same-frame double-release; this ordering is FIFO-proven safe and covered
by a control-mutant test.
2. **`handOff` liveness gate** — never hand a disconnected browser to a
waiter; recycle (idle slot) or park `relaunchPending` (busy slot)
instead, keeping the waiter queued for a live browser.
3. **`track()` registers the same wrapper it removes** — so
`shutdown()`'s drain awaits the recovery's cleanup rather than
under-waiting the raw promise.
These harden pre-existing latent bugs. The launch-gate (#5137) already
fixed the d6 infra issue, so this is robustness hardening.
## Test plan
- [x] red-green: double-release-idempotency (double `release()` is a
no-op)
- [x] red-green: dead-browser-liveness-gate (disconnected browser never
handed to a waiter)
- [x] red-green: track-wrapper-drain (shutdown awaits recovery cleanup)
- [x] red-green: waiter-served-release-no-leak control-mutant (deferred
microtask add ordering)
- [x] full harness suite green (99 files, 1681 tests)
- [x] `tsc --noEmit` clean, build clean, oxfmt + oxlint clean
Three fixes:
(1) idempotent release() via checked-out tracking (prevents double-release
double-serving one browser to two waiters; the waiter-path checked-out add is
deferred one microtask to defeat same-frame double-release, FIFO-proven safe
and covered by a control-mutant test),
(2) handOff liveness gate (never hand a disconnected browser to a waiter —
recycle/park instead),
(3) track() registers the same wrapper it removes so shutdown's drain awaits
cleanup.
These harden pre-existing latent bugs; the launch-gate (#5137) already fixed
the d6 infra issue, so this is robustness hardening.
## Summary
`waitForAssistantSettled` now fast-fails (throws
`AssistantErroredError`, classified `conversation-error`, in ~2s)
instead of burning the full 30s `responseTimeoutMs` when the chat
surfaces a genuine error banner.
Each poll snapshots the `copilot-error-banner` testid (shipped in #5110)
and computes a single boolean — an error banner is present in a state
that **DIFFERS from the turn's baseline** (a brand-new banner, OR a
persisted banner whose text changed). The turn fast-fails only when ALL
of:
- that differs-from-baseline state is **SUSTAINED across 2 consecutive
polls** (a single isolated flicker — a transient toast, a one-poll
re-render glitch — is debounced away and does NOT fire), AND
- the assistant produced **NO response this turn** (the message count
has not grown past baseline). **Success-in-flight wins**: a non-fatal
warning banner alongside a real answer never force-fails the turn; the
settle path governs instead.
Because it only acts when the turn would otherwise time out (no response
+ a sustained error), this is a **strict improvement**: it can only turn
a would-be timeout-fail into a fast fail of the **same verdict** — it
can never flip a verdict. The full (untruncated) banner text is compared
across polls so two errors diverging only after the first 300 chars
still differ; truncation applies only to the thrown message.
## Test plan
Red-green unit tests in `conversation-runner.test.ts` (full harness
suite green: 1684 passed):
- [x] fresh-banner flicker for a single poll then disappears → does NOT
fast-fail (turn succeeds)
- [x] persisted-banner text flicker for a single poll then reverts →
does NOT fast-fail (turn succeeds)
- [x] sustained differs-from-baseline across 2 consecutive polls (no
response) → fast-fails (~2s, not a timeout)
- [x] dynamic/mutating banner text (countdown/timestamp) sustained
across 2 polls → fast-fails (dynamic-text fires)
- [x] stale same-text persisted banner (text == baseline) → no fast-fail
(correctly treated as prior turn's error)
- [x] success-in-flight: assistant produced a response while a banner is
also visible → does NOT fast-fail (success wins)
- [x] 300-char-prefix: two errors diverging only after the first 300
chars still differ from baseline (full-text compare)
- [x] **mutation check**: temporarily relaxing the debounce to fire on a
single poll (`>= 1`) makes both single-poll-flicker tests FAIL, then
restored to `>= 2` — proving the two tests genuinely guard the
2-consecutive-poll debounce-reset path (not the success-in-flight
disarm)
## Known limitations
Two intentional safe-degrade edges fall back to the normal timeout
(correct verdict, just not sped up):
- **count-oscillation**: a response count that briefly grows past
baseline and then reverts disarms the debounce, so an error that only
sustains after that bounce settles via timeout.
- **synchronous-error-at-baseline**: a brand-new error whose banner text
is byte-identical to a persisted stale banner already visible at
baseline cannot be distinguished from it, so it settles via timeout.
`waitForAssistantSettled` now fast-fails (throws `AssistantErroredError`,
classified `conversation-error`, in ~2s) instead of burning the full 30s
`responseTimeoutMs` when the chat surfaces a genuine error banner.
Each poll snapshots the `copilot-error-banner` testid (shipped in #5110)
and computes a single boolean — an error banner is present in a state that
DIFFERS from the turn's baseline (a brand-new banner, OR a persisted banner
whose text changed). The turn fast-fails only when ALL of:
- that differs-from-baseline state is SUSTAINED across 2 consecutive polls
(a single isolated flicker — a transient toast, a one-poll re-render
glitch — is debounced away and does NOT fire), AND
- the assistant produced NO response this turn (the message count has not
grown past baseline). Success-in-flight wins: a non-fatal warning banner
alongside a real answer never force-fails the turn; the settle path
governs instead.
Because it only acts when the turn would otherwise time out (no response +
a sustained error), this is a strict improvement: it can only turn a
would-be timeout-fail into a fast fail of the SAME verdict — it can never
flip a verdict. The full (untruncated) banner text is compared across polls
so two errors diverging only after the first 300 chars still differ;
truncation applies only to the thrown message.
Two intentional safe-degrade edges fall back to the normal timeout (correct
verdict, just not sped up):
- count-oscillation: a response count that briefly grows past baseline and
then reverts disarms the debounce, so an error that only sustains after
that bounce settles via timeout.
- synchronous-error-at-baseline: a brand-new error whose banner text is
byte-identical to a persisted stale banner already visible at baseline
cannot be distinguished from it, so it settles via timeout.
## Summary
Adds the 5 missing LGP-canonical e2e specs to each of the 9 baseline
integrations and removes 2 orphan specs that point at zero-file demos.
- Adds 5 canonical e2e specs (`declarative-hashbrown`,
`declarative-json-render`, `reasoning-custom`, `reasoning-default`,
`threadid-frontend-tool-roundtrip`) to each of the 9 baseline
integrations (`ag2`, `agno`, `crewai-crews`, `langgraph-fastapi`,
`langroid`, `llamaindex`, `mastra`, `spring-ai`, `strands`) — 45 specs
total, **byte-identical** to the `langgraph-python` canonical gold
source. The matching demos already existed from #5127's page-mirror, but
the specs themselves were never copied across.
- Deletes 2 orphan specs that reference demos with zero backing files:
`shared-state-write` (×9) and `reasoning-default-render` (×8).
- `validate-parity.ts` reports **19 pass / 0 fail**. The other 15
frameworks were already clean.
- Note: `claude-sdk-python` has the same orphan-spec condition and is
left for a future sweep (out of scope here).
## Test plan
- [x] Parity validator (`validate-parity.ts`) green: 19/0.
- [ ] Real signal: the added cells running green in a clean staging d6
run.