The pool checked the cap and mutated liveContextCount AROUND await points
with no synchronous reservation, so under concurrent acquire/release
(pooled drivers run FEATURE_CONCURRENCY_D6=4 x parallel services) the cap
invariant broke and capacity bled. Fix as one cohesive reservation model:
- reserveSlot(): synchronous check-and-increment with NO await between them,
so concurrent acquires can never collectively pass the check and overshoot
maxContexts during one another's newContext() awaits.
- openContextOn() now opens against a caller-held reservation; on
newContext() failure it rolls the reservation back (clamped) and rethrows.
- acquire() reserves before any await and HOLDS the reservation across the
newContext-failure retry (re-reserves explicitly), so the retry path no
longer bypasses the cap. Rolls back when it falls through to a waiter.
- serveNextWaiter() reserves synchronously before opening and is consistent
with the model so concurrent release + recycle-drain cannot overshoot.
- Dead-waiter orphan LEAK fixed: a settled flag is set by BOTH the
timeout-reject and resolve wrappers; serveNextWaiter detects a waiter that
timed out mid-open, CLOSES the freshly-opened context and rolls back the
count instead of orphaning it (the leak that permanently bled capacity and
re-created pool starvation).
- release() does its bookkeeping (delete + clamped decrement) and evaluates
the idle-recycle decision off purely SYNCHRONOUS state, then awaits
context.close() AFTER, so no await straddles the decrement and the size check.
Hardening (same file):
- recycleBrowser .finally() no longer resets entry.recycling (the success
path already cleared it; the eviction path spliced the entry out, so
resetting would resurrect a dead entry). Abandoned-context decrements now
go through the clamped releaseReservation().
- liveContextCount decrements are clamped at 0 (releaseReservation) as
insurance against any future double-decrement.
Public API (acquire/release/stats/shutdown/constructor) and stats()
semantics (size=maxContexts, available=max-live, inUse=live) preserved.
Tests: added CONCURRENT-path regressions (the existing suite was all
sequential, which is why these bugs passed CI). Red-green receipts:
- "N concurrent acquire() never exceed maxContexts": RED expected 5 to be
<= 2 -> GREEN.
- "newContext-failure retry path respects maxContexts": RED expected 2 to
be <= 1 -> GREEN.
- "does not leak a context when a waiter times out mid-serveNextWaiter":
RED expected 1 to be 0 -> GREEN.
- release close/accounting reorder: forward regression guard (the pre-fix
order kept the released context in the live set across its close-await, so
the reorder is behavior-preserving but makes the no-await-gap invariant
machine-checkable).
Wire the orchestrator to the options-object BrowserPool: a fixed set of
browser processes (BROWSER_POOL_BROWSERS, default 3, legacy BROWSER_POOL_SIZE
fallback) with a global live-context cap (BROWSER_POOL_MAX_CONTEXTS, default
24). registerAllProbeDrivers call sites are unchanged (launcher ctor args
unchanged).
Rollout note: BROWSER_POOL_SIZE's meaning shifts from "concurrent browsers"
to "base browser process count". The deploy env should move to
BROWSER_POOL_BROWSERS=3 + BROWSER_POOL_MAX_CONTEXTS=24.
Update all five pooled probe launchers for the context-pooled BrowserPool.
Each launcher's newContext() now checks out a pooled BrowserContext via
pool.acquire(opts) and the returned context-wrapper's close() releases it
via pool.release(ctx); the launcher-level close() is a no-op (no Browser is
held). The X-AIMock-Strict literal is dropped from every launcher (now
centralized in the pool); drivers still pass their per-probe X-AIMock-Context
/ X-Test-Id headers, which now flow through to pool.acquire. Driver run()
bodies are unchanged.
- d4 (createPooledE2eSmokeLauncher), e2e-parity
(createPooledE2eParityLauncher): minimal transform.
- e2e-readiness (createPooledE2eDemosLauncher), d5
(createPooledE2eDeepLauncher), d6 (createPooledE2eFullLauncher): the abort
closure is re-targeted to close each open context (each releasing its
pooled context) instead of force-releasing a held browser. d6's
dead-browser re-acquire dance is removed, the pool only opens contexts on
live browsers. Per-service Semaphore(FEATURE_CONCURRENCY/_D6) is preserved
as the orthogonal per-service fan-out bound.
Driver tests reinterpret POOL_SIZE as maxContexts and assert per-context
acquire/release moves inUse by 1, abort closes open contexts with no browser
fork, and newContext(opts).extraHTTPHeaders forwards into pool.acquire.
Re-architect BrowserPool to check out BrowserContexts from a fixed, small
set of long-lived browser PROCESSES instead of recycling whole browsers.
The old model recycled a browser (close + relaunch = a process fork) every
recycleAfter releases, so a steady-state run forked chromium repeatedly on
the hot path, the EAGAIN / PID-ceiling churn source ("Zygote could not
fork", 0/18 d6 runs).
New model:
- Checkout unit is a BrowserContext: acquire(opts?, timeoutMs?) opens a
context on the least-loaded live browser; release(ctx) closes it. No
process fork on the hot path.
- Options-object constructor: new BrowserPool({ browsers, maxContexts,
recycleAfter, logger, launchBrowser, launchStaggerMs }). Defaults
browsers=3 (env BROWSER_POOL_BROWSERS, legacy BROWSER_POOL_SIZE fallback),
maxContexts=24 (env BROWSER_POOL_MAX_CONTEXTS), recycleAfter=300 (env
BROWSER_POOL_RECYCLE_AFTER, now a per-browser served-context hygiene
threshold).
- Browser processes launch only at init() (the fixed set, staggered), on
crash recovery, and on the rare served-context hygiene recycle.
- X-AIMock-Strict default header centralized in acquire(); FIFO waiters past
the cap carry their context options; crash blast radius is bounded to the
dead browser's contexts.
- stats() keeps its field names with context semantics (size=maxContexts,
available/inUse derived from live-context count).
Tests rewritten red-green; the load-bearing assertion is that the fake
launcher's launched count stays == N across many acquire/release cycles
(zero hot-path forks).
An aborted D6 run (orchestrator process killed during a pool-churn burst)
left its `probe_runs` row orphaned in `running` state with a null summary.
The boot-time `sweepStaleRuns` then stamped it `state:failed` with
`{total:0,passed:0,failed:0}` and `duration_ms:null`, discarding the
partial per-service results the run had actually computed — a 578s run with
dozens of green features surfaced as `failed / total:0`.
Two minimal, success-path-consistent changes:
- probe-invoker: incrementally persist the running partial tally onto the
probe_runs row via a new `runWriter.update()` as each fan-out target
completes, so an orphaned row already carries real partial progress.
Best-effort — never tanks the tick; `finish()` remains the authoritative
final write.
- run-history: `sweepStaleRuns` now PRESERVES an existing partial summary
(and derives a real duration from the persisted started_at) instead of
clobbering it to zeros. Rows that died before any target completed still
fall back to an explicit empty rollup.
Purely the result-aggregation-on-abort path; pool sizing and
launch/recycle logic are untouched (churn root cause handled separately).
## Summary
Three fixes hardening BrowserPool slot publication:
1. **Idempotent `release()` via checked-out tracking** — prevents a
double-`release()` from double-serving one browser to two waiters. The
waiter-path `checkedOut` add is deferred one microtask to defeat a
same-frame double-release; this ordering is FIFO-proven safe and covered
by a control-mutant test.
2. **`handOff` liveness gate** — never hand a disconnected browser to a
waiter; recycle (idle slot) or park `relaunchPending` (busy slot)
instead, keeping the waiter queued for a live browser.
3. **`track()` registers the same wrapper it removes** — so
`shutdown()`'s drain awaits the recovery's cleanup rather than
under-waiting the raw promise.
These harden pre-existing latent bugs. The launch-gate (#5137) already
fixed the d6 infra issue, so this is robustness hardening.
## Test plan
- [x] red-green: double-release-idempotency (double `release()` is a
no-op)
- [x] red-green: dead-browser-liveness-gate (disconnected browser never
handed to a waiter)
- [x] red-green: track-wrapper-drain (shutdown awaits recovery cleanup)
- [x] red-green: waiter-served-release-no-leak control-mutant (deferred
microtask add ordering)
- [x] full harness suite green (99 files, 1681 tests)
- [x] `tsc --noEmit` clean, build clean, oxfmt + oxlint clean
Three fixes:
(1) idempotent release() via checked-out tracking (prevents double-release
double-serving one browser to two waiters; the waiter-path checked-out add is
deferred one microtask to defeat same-frame double-release, FIFO-proven safe
and covered by a control-mutant test),
(2) handOff liveness gate (never hand a disconnected browser to a waiter —
recycle/park instead),
(3) track() registers the same wrapper it removes so shutdown's drain awaits
cleanup.
These harden pre-existing latent bugs; the launch-gate (#5137) already fixed
the d6 infra issue, so this is robustness hardening.
## Summary
`waitForAssistantSettled` now fast-fails (throws
`AssistantErroredError`, classified `conversation-error`, in ~2s)
instead of burning the full 30s `responseTimeoutMs` when the chat
surfaces a genuine error banner.
Each poll snapshots the `copilot-error-banner` testid (shipped in #5110)
and computes a single boolean — an error banner is present in a state
that **DIFFERS from the turn's baseline** (a brand-new banner, OR a
persisted banner whose text changed). The turn fast-fails only when ALL
of:
- that differs-from-baseline state is **SUSTAINED across 2 consecutive
polls** (a single isolated flicker — a transient toast, a one-poll
re-render glitch — is debounced away and does NOT fire), AND
- the assistant produced **NO response this turn** (the message count
has not grown past baseline). **Success-in-flight wins**: a non-fatal
warning banner alongside a real answer never force-fails the turn; the
settle path governs instead.
Because it only acts when the turn would otherwise time out (no response
+ a sustained error), this is a **strict improvement**: it can only turn
a would-be timeout-fail into a fast fail of the **same verdict** — it
can never flip a verdict. The full (untruncated) banner text is compared
across polls so two errors diverging only after the first 300 chars
still differ; truncation applies only to the thrown message.
## Test plan
Red-green unit tests in `conversation-runner.test.ts` (full harness
suite green: 1684 passed):
- [x] fresh-banner flicker for a single poll then disappears → does NOT
fast-fail (turn succeeds)
- [x] persisted-banner text flicker for a single poll then reverts →
does NOT fast-fail (turn succeeds)
- [x] sustained differs-from-baseline across 2 consecutive polls (no
response) → fast-fails (~2s, not a timeout)
- [x] dynamic/mutating banner text (countdown/timestamp) sustained
across 2 polls → fast-fails (dynamic-text fires)
- [x] stale same-text persisted banner (text == baseline) → no fast-fail
(correctly treated as prior turn's error)
- [x] success-in-flight: assistant produced a response while a banner is
also visible → does NOT fast-fail (success wins)
- [x] 300-char-prefix: two errors diverging only after the first 300
chars still differ from baseline (full-text compare)
- [x] **mutation check**: temporarily relaxing the debounce to fire on a
single poll (`>= 1`) makes both single-poll-flicker tests FAIL, then
restored to `>= 2` — proving the two tests genuinely guard the
2-consecutive-poll debounce-reset path (not the success-in-flight
disarm)
## Known limitations
Two intentional safe-degrade edges fall back to the normal timeout
(correct verdict, just not sped up):
- **count-oscillation**: a response count that briefly grows past
baseline and then reverts disarms the debounce, so an error that only
sustains after that bounce settles via timeout.
- **synchronous-error-at-baseline**: a brand-new error whose banner text
is byte-identical to a persisted stale banner already visible at
baseline cannot be distinguished from it, so it settles via timeout.
`waitForAssistantSettled` now fast-fails (throws `AssistantErroredError`,
classified `conversation-error`, in ~2s) instead of burning the full 30s
`responseTimeoutMs` when the chat surfaces a genuine error banner.
Each poll snapshots the `copilot-error-banner` testid (shipped in #5110)
and computes a single boolean — an error banner is present in a state that
DIFFERS from the turn's baseline (a brand-new banner, OR a persisted banner
whose text changed). The turn fast-fails only when ALL of:
- that differs-from-baseline state is SUSTAINED across 2 consecutive polls
(a single isolated flicker — a transient toast, a one-poll re-render
glitch — is debounced away and does NOT fire), AND
- the assistant produced NO response this turn (the message count has not
grown past baseline). Success-in-flight wins: a non-fatal warning banner
alongside a real answer never force-fails the turn; the settle path
governs instead.
Because it only acts when the turn would otherwise time out (no response +
a sustained error), this is a strict improvement: it can only turn a
would-be timeout-fail into a fast fail of the SAME verdict — it can never
flip a verdict. The full (untruncated) banner text is compared across polls
so two errors diverging only after the first 300 chars still differ;
truncation applies only to the thrown message.
Two intentional safe-degrade edges fall back to the normal timeout (correct
verdict, just not sped up):
- count-oscillation: a response count that briefly grows past baseline and
then reverts disarms the debounce, so an error that only sustains after
that bounce settles via timeout.
- synchronous-error-at-baseline: a brand-new error whose banner text is
byte-identical to a persisted stale banner already visible at baseline
cannot be distinguished from it, so it settles via timeout.
## Summary
Adds the 5 missing LGP-canonical e2e specs to each of the 9 baseline
integrations and removes 2 orphan specs that point at zero-file demos.
- Adds 5 canonical e2e specs (`declarative-hashbrown`,
`declarative-json-render`, `reasoning-custom`, `reasoning-default`,
`threadid-frontend-tool-roundtrip`) to each of the 9 baseline
integrations (`ag2`, `agno`, `crewai-crews`, `langgraph-fastapi`,
`langroid`, `llamaindex`, `mastra`, `spring-ai`, `strands`) — 45 specs
total, **byte-identical** to the `langgraph-python` canonical gold
source. The matching demos already existed from #5127's page-mirror, but
the specs themselves were never copied across.
- Deletes 2 orphan specs that reference demos with zero backing files:
`shared-state-write` (×9) and `reasoning-default-render` (×8).
- `validate-parity.ts` reports **19 pass / 0 fail**. The other 15
frameworks were already clean.
- Note: `claude-sdk-python` has the same orphan-spec condition and is
left for a future sweep (out of scope here).
## Test plan
- [x] Parity validator (`validate-parity.ts`) green: 19/0.
- [ ] Real signal: the added cells running green in a clean staging d6
run.
## Summary
The aimock matcher gates a fixture's `turnIndex` against the request's
assistant-message count: a fixture only fires when `assistantCount ===
turnIndex`. So `turnIndex` must equal the number of assistant turns
already in the conversation when that fixture should match — turn 1 is
`turnIndex: 0`, turn 2 is `turnIndex: 1`, etc. A multi-turn conversation
fixture that wants to match on *every* turn must **omit** `turnIndex`
entirely and disambiguate purely by `userMessage`, which is exactly what
the canonical clean pattern (langgraph-python and the other 15
frameworks) does.
The `agno`, `spring-ai`, and `langgraph-fastapi` `agentic-chat` fixtures
had baked `turnIndex: 0` onto **all three** goldfish conversation turns.
Turn 2+ carries one or more assistant messages, so `assistantCount !==
0` and the turn-2/turn-3 fixtures could never match. aimock returned
no-match, the request fell through to the proxy, and the d6 cell failed
— surfacing as the 503 → proxy → 502 cascade on staging.
This PR drops `turnIndex: 0` from the multi-turn goldfish turns in all
three files, mirroring the langgraph-python canonical pattern, and adds
clarifying `_comment` keys on each turn.
## Findings
- The 503/502 cascade had been diagnosed on two integrations (agno,
spring-ai). Auditing the full d6 fixture set turned up a **third**
affected integration, **langgraph-fastapi**, carrying the identical
`turnIndex: 0`-on-all-turns defect.
- **Single-turn** fixtures correctly keep `turnIndex: 0` (a single-turn
match where `assistantCount === 0` is correct) — those are unchanged.
- The other **15 frameworks** were already clean (no `turnIndex` on
multi-turn conversation turns).
- Local before/after evidence: turn-2 went 503 → 200 for all three
frameworks after the fix; turn-1 was unregressed.
## Test plan
- [x] All three JSON files parse (valid JSON).
- [x] `validate-fixture-tool-surface.ts` — green (no drift).
- [x] `validate-parity.ts` — exit 0.
- [x] `validate-pins.ts` ratchet — FAIL count and hash unchanged vs
`fail-baseline.json` (63, matching hash); none of the FAIL lines touch
these fixtures.
- [x] `vitest run` (build-pipeline tests) — 1670 passed.
- [ ] The real signal is a clean staging d6 run on these cells (agno /
spring-ai / langgraph-fastapi agentic-chat), verified post-merge.
The aimock matcher gates turnIndex against the request's assistant-message
count (assistantCount !== turnIndex → skip). The agno, spring-ai, and
langgraph-fastapi agentic-chat fixtures baked turnIndex:0 on all three
goldfish conversation turns. Turn 2+ carries >=1 assistant message, so
turnIndex:0 could never match those turns — aimock returned no-match and the
request fell through to proxy (503), failing the d6 cell.
Mirror the canonical clean pattern used by the other 15 frameworks
(langgraph-python et al.): omit turnIndex on the multi-turn conversation
turns and disambiguate purely by userMessage. Single-turn fixtures in the
same files keep turnIndex:0 (unchanged). Verified locally against the aimock
matcher: turn-2 goes 404→200 for all three frameworks, turn-1 unregressed.
Move the required, uniform D6 header-conveyance config on-disk so it rides
image promotion instead of depending on a per-environment env var. Prod was
missing the LANGGRAPH_HTTP configurable_headers env var, and baking the
`http.configurable_headers.include: ["x-*"]` setting directly into the
langgraph-python and langgraph-fastapi langgraph.json files removes the
promote-time drift gap (the config now travels with the image rather than
being re-supplied at each promotion).
langgraph-typescript needs no change — its header conveyance is pure-code.
reasoning-default-render.spec.ts and shared-state-write.spec.ts navigate to
/demos/reasoning-default-render and /demos/shared-state-write respectively,
but neither demo directory exists in any integration (including the
langgraph-python gold reference) and neither is declared in any manifest.
They are stale, non-canonical leftovers from the #5127 page-mirror.
Removed from all baselines where present: reasoning-default-render from 8
(langgraph-fastapi never had it) and shared-state-write from all 9. The
parity validator stays green (0 fail) and langgraph-fastapi's prior
spec-under-coverage warning clears once its phantom spec is gone and the 5
canonical specs are added.
The 9 baseline integrations (ag2, agno, crewai-crews, langgraph-fastapi,
langroid, llamaindex, mastra, spring-ai, strands) were page-mirrored from
langgraph-python but the mirror omitted 5 LGP-canonical Playwright specs
whose demos are present on disk:
- declarative-hashbrown
- declarative-json-render
- reasoning-custom
- reasoning-default
- threadid-frontend-tool-roundtrip
Copied each spec verbatim (byte-identical) from langgraph-python, which the
baselines mirror. All 5 backing demo directories exist in every baseline.
The specs are framework-agnostic (navigate by route + testid), so no
per-integration edits are needed. Restores apples-to-apples spec parity.
The BrowserPool init() fill loop launches browsers one at a time through the
serializing launch gate. If launch iteration N throws -- the exact PID-ceiling
failure (pthread_create EAGAIN / "Zygote could not fork") this file exists to
survive -- init() rejected and propagated, but the browsers already launched on
iterations 0..N-1 stayed live in this.slots/available/browserToSlot and were
never closed. The sole production caller (orchestrator.ts boot) catches the
rejection, marks the pool degraded, and does NOT call shutdown(), so those
chromium processes (~50 PIDs each) leaked permanently -- accelerating the very
PID exhaustion the launch gate is meant to prevent. The gate serializes the
fill, which makes a partial-then-throw fill MORE reachable.
Wrap the fill loop so a mid-fill launch failure resets the pool's internal
state and closes every browser already launched before re-throwing. State is
cleared BEFORE closing so the synchronous `disconnected` fire from a close
cannot re-enter the recycle path via the slot's disconnect handler (the
!this.slots.includes(slot) guard short-circuits it) -- the same ordering
intent shutdown() relies on, without adding an isShutdown re-check. init()'s
reject-on-failure contract is preserved.
Call sites: init() signature is unchanged. The only production caller
(showcase/harness/src/orchestrator.ts:292) and all test callers continue to
rely on reject-on-failure, which still holds.
Adds a multi-slot partial-fill test (pool size 4, launch 3 throws): asserts
init() rejects, the 2 browsers launched before the throw were closed, and
stats() reports an empty pool. The prior init-failure tests only covered
1-slot / first-launch-fails (nothing launched yet), which is why the leak was
missed.
The browser-pool launch-stagger resolution honored the "negative or
non-numeric value falls back to the default rather than disabling the
stagger silently" contract only for the BROWSER_LAUNCH_STAGGER_MS env
var, NOT for an explicit launchStaggerMs constructor arg. A negative
explicit arg (e.g. -50) was taken by the `??` (non-null), skipped the
env/default branch, then got clamped to 0 by `resolvedStagger >= 0 ? ... : 0`
— silently DISABLING the stagger and reintroducing the launch-burst PID
spike (pthread_create EAGAIN / "Zygote could not fork") the gate exists
to prevent. A NaN explicit arg behaved the same way.
Validate the explicit arg the same way as the env var BEFORE it wins: a
valid explicit arg (>= 0, not NaN) wins; else a valid env value; else the
default (150ms). A negative/NaN explicit arg now falls back to the default,
not 0. An explicit 0 is still respected (tests rely on it to stay fast)
because `validExplicit ?? validEnv` keeps a literal 0.
Call sites: this only changes how this.launchStaggerMs is computed inside
the constructor. The sole consumer is launchBrowser()'s
`this.launchStaggerMs > 0 ? delay(...) : undefined` gate — its semantics
are unchanged (the field is still always a valid non-negative number). The
only production constructor (orchestrator.ts: `new BrowserPool(poolSize,
undefined, logger)`) passes no stagger arg, so its behavior (undefined →
env → default) is identical before and after.
Adds two red-green tests: a negative explicit arg and a NaN explicit arg
must each space launches by ~the default stagger (gap >= 140ms), proving
the stagger is not silently disabled.
The harness runs every e2e probe through one process-wide BrowserPool whose
chromium launches happen in BURSTS (initial fill, recycle relaunches, lazy
relaunchPending recovery, reinit backstop). On the Railway staging container
(~1000-PID ceiling, ~50 PIDs per headless chromium) a burst transiently
spikes PID demand past the ceiling and trips `pthread_create: Resource
temporarily unavailable` / "Zygote could not fork", so browsers fail to launch
and a d6 run goes 0/18.
Funnel EVERY launch through a single concurrency-1 serialization gate that
also waits BROWSER_LAUNCH_STAGGER_MS (default 150ms, env-tunable) after each
launch settles before the next may start. The caller still receives its
browser the instant the process is up; only the NEXT launch is gated. This
spaces process spawns so the transient PID spike never exceeds the ceiling
WITHOUT reducing the eventual pool size. All launch paths (init fill,
acquire-time/recycle relaunch, reinit, relaunchPending recovery) now route
through the gated `launchBrowser` wrapper over the raw launcher.
Adapts the two deterministic-race guard tests that previously required two
recovery launches to be in flight simultaneously — incompatible with the
serialization invariant — to assert the same re-entry guards under serialized
launches (verified by mutation: disabling the in-loop guard still fails the
test). Existing tests pin stagger to 0 to stay fast.
Move framework-incapable features from features[] into not_supported_features[]
so the D6 harness reclassifies them as skipped-incapable instead of red:
- mastra: gen-ui-interrupt, interrupt-headless, agentic-chat-reasoning,
reasoning-default-render, tool-rendering-reasoning-chain
- langroid: mcp-apps, tool-rendering-reasoning-chain
- ag2: gen-ui-interrupt, interrupt-headless
- crewai-crews: gen-ui-interrupt, interrupt-headless, mcp-apps
- llamaindex: gen-ui-interrupt, interrupt-headless, hitl-in-chat-booking
- spring-ai: byoc-json-render
All entries already documented as incapable in each integration's PARITY_NOTES.md.
Validated via 'npm run validate-manifests' (no features/NSF overlap).
- cli/targets.ts: read not_supported_features from manifest YAML, pass through
buildFullInputs as notSupportedFeatures on FullInput
- probes/discovery/railway-services.ts: extract not_supported_features from
registry.json per integration via new RegistryIntegrationInfo type and emit
on RailwayServiceInfo records so the D6 driver receives it at probe time
When an integration's manifest lists a feature in not_supported_features (NSF),
the D6 e2e-full driver now partitions requestedFeatures into capable + incapable
sets BEFORE script resolution. Incapable features get a green side-row with
errorClass='skipped-incapable' and surface in the aggregate skipped[] list plus
a new incapable[] field. They no longer count as red.
Includes red-green test that asserts NSF features without a registered script
emit state=green (was red prior to this change).
The mirrored LGP homepage prerenders / and reads manifest.yaml at build
time; these 3 baselines use an explicit COPY of config files that omitted
manifest.yaml, breaking SSG export (same fix as pydantic-ai in #5125). The
other 6 baselines copy the dir wholesale and already include it.
## Summary
D6 showcase-parity wave for the priority bloc. Brings four integrations
closer to the langgraph-python (LGP) gold reference and unblocks
pydantic's production Docker build.
| Integration | D6 result | Change |
|---|---|---|
| **claude-sdk-typescript** | **130/55** (was 79/106, **+51**) | Full
LGP demo-page mirror + shadcn primitives + `copilotkit-beautiful-chat`
route (was 404) |
| **ms-agent-dotnet** | **180/2** (was 177/5) | toolCallId-strip via
universal `createAgent` `FunctionMiddleware` + reasoning-chain injection
+ fixtures |
| **ms-agent-python** | toolCallId-strip failures resolved | Ported the
universal strip middleware + snake_case `tool_calls` / nested
`function.tool_call_id` to match dotnet. Residual non-toolCallId
failures remain (interrupt timeouts, multimodal endpoint, declarative
pie-chart) — out of scope here |
| **pydantic-ai** | **Docker build fixed** (SSG export) + page-mirror |
Aligned deps to LGP; mirrored demo pages; **Dockerfile now copies
`manifest.yaml`** into the frontend stage (the homepage prerenders `/`
and reads it at build time, previously crashing the SSG export with
ENOENT). Residual gen-ui/timeout failures remain |
## Notes
- **Scope:** showcase integrations only (no shipped SDK code). CST and
dotnet are clear wins; pydantic build-fix is essential; python lands
correct middleware parity with documented residuals on the
de-prioritized Python bloc.
- **Page-mirrors** are verbatim copies of the LGP gold reference (demo
pages, `components/ui/*`, `lib/utils.ts`, the LGP dep set) with only the
backend/runtime swap retained at the API layer.
- **Code review:** 7-agent CR loop run against `origin/main`; **zero
functional (bucket-a) findings**. Flagged items were verbatim-LGP
dep-pin parity, pre-existing foundation machinery (not in this diff), or
internal-tool-bar non-functional (security/style) — deferred.
- **CI is the authoritative package-test gate** (the isolated worktree's
pnpm/lefthook SDK suite can't run locally). Integration builds verified
green via local docker builds.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Count is correct at 104 (down from 106 — net improvement from the
@copilotkit 1.59.2 exact-pin + manifest highlight fixes). The prior hash
was computed in a local env whose FAIL-line set differed from CI's
canonical pipeline; sync to CI's printed actual hash.
Previously pydantic-ai and claude-sdk-typescript used 'latest' for all
@copilotkit/* dependencies (and the @copilotkit/web-inspector pnpm
overrides on @copilotkit/core), which the showcase validate-pins ratchet
counts as non-exact pin drift. Intended design is to pin to exact 1.59.2.
Changes:
- integrations/pydantic-ai/package.json: @copilotkit/{a2ui-renderer,react-core,runtime,shared,voice} 'latest' -> '1.59.2'; npm + pnpm overrides on @copilotkit/web-inspector>@copilotkit/core 'latest' -> '1.59.2'
- integrations/claude-sdk-typescript/package.json: same set as above
- Regenerated both package-lock.json files via npm install --legacy-peer-deps --package-lock-only
- scripts/fail-baseline.json: ratcheted DOWN validatePinsFailCount 106 -> 104; updated validatePinsFailHash to dde7950e8d691de5a7b2c0c16ca64b3e550221cb6072d2c29e24dcb497515cf6 (matches local sort-uniq + shasum-256 of the new [FAIL] set)
Verified locally: npx tsx validate-pins.ts reports Summary FAIL=104
(2 fewer than baseline because pydantic-ai's @copilotkit/react-core and
@copilotkit/runtime moved from 'latest' (non-exact) to '1.59.2' (exact);
CST already had unrelated FAILs that remain). Non-@copilotkit deps
(lucide-react, cmdk, openai, @ag-ui/*) intentionally left unchanged.
## Summary
Unifies how the showcase status dashboard decides whether a demo's data
is "too old to trust," and removes two ways a chip/badge could show
green when the underlying data was stale or absent.
- **Shared staleness helper** (`lib/staleness.ts`): one `isStale` +
per-driver windows (e2e 6h, D4 1h, liveness 45m), so the chip, badges,
depth, and filters all apply the same freshness rule instead of three
private copies.
- **Stale-green → amber**: `resolveCell` and `resolveD5Row` downgrade
frozen-green rows so a driver that stopped reporting no longer reads as
healthy.
- **Honest D5 coverage**: `resolveD5Row` now returns `null` for
unmapped/empty-map features, matching the chip (`resolveD5`) and depth
(`isD5Green`). Previously an unmapped feature with a stray
`d5:<slug>/<featureId>` row rendered a green badge while the chip showed
gray — a visible contradiction.
- **No phantom regressions**: `isRegression` requires emitted data on
the rung above the achieved depth before flagging a regression.
- **D4 worst-state-wins**: `deriveDepth` folds chat/tools to the worst
state rather than OR-ing green.
- **One clock per render**: a single `now` (memoized on `liveStatus`) is
threaded through the cell-matrix render path so the chip, the
regressions/gaps filter, and the badges all judge staleness against the
same instant; previously the render path used a fresh `Date.now()` per
cell and could disagree across a window boundary.
## Test plan
- [x] `npm run typecheck` — clean
- [x] `npm run test` — 681 passed, 1 skipped
- [x] `npm run build` — clean
- [x] New red-green tests: stale-green downgrade (order-independent),
STRICT missing-sub-row, unmapped-feature → gray (not green), and
shared-`now` agreement across a staleness-window boundary
The homepage prerenders / at build time and reads manifest.yaml via the
filesystem, but the frontend stage only copied next.config/tsconfig/postcss.
Without the manifest the SSG export of /page failed with ENOENT, breaking the
Docker build. Copy manifest.yaml alongside the other config files.
The claude-sdk-typescript beautiful-chat demo page sets
runtimeUrl="/api/copilotkit-beautiful-chat" but no such Next.js route
existed, so the cell 404'd on every request.
Add a dedicated runtime mirroring pydantic-ai's beautiful-chat route:
openGenerativeUI + a2ui (injectA2UITool: false) + mcpApps middleware
configured together (the canonical LGP combined-runtime shape). Wire
the HttpAgent to CST's pass-through / mount since agent_server.ts has
no dedicated beautiful_chat backend graph — CST is frontend +
middleware driven, matching how mcp-apps is wired.
Also defensively register beautiful-chat in the shared /api/copilotkit
agentNames list so probe requests against the default runtime resolve
cleanly, matching the pattern used by the other dedicated-runtime
demos.
Hoist a single now (memoized on liveStatus) and thread it into the
render-path buildCellModel so the chip, the regressions/gaps filter, and
the badges all judge staleness against one instant. Previously the render
path defaulted to a fresh Date.now() per cell, so a cell the filter
included could render a different staleness state across a window boundary.
Drop the direct-key d5 fallback in resolveD5Row so the coverage badge
agrees with the chip (cell-model resolveD5) and depth (depth-utils
isD5Green), which both treat an unmapped feature as no-D5. Previously an
unmapped feature with a stray d5:<slug>/<featureId> row rendered a green
badge while the chip showed gray. Also guards the empty-map case.