## Summary
- **Fix per-cell D6 resolution**: the dashboard's `resolveD6` now reads
per-cell ENUM keys (`d6:<slug>/<featureType>` via `CATALOG_TO_D5_KEY`
fan-out) instead of the integration aggregate — D6 cells render real
per-cell green/red instead of all-gray.
- **Depth + D6 surfaced by default**: `DEFAULT_OVERLAYS = [links,
health, depth]` (the per-cell D6 badge rides on the Health layer; the
depth chip folds D6).
- **Correct the aggregate stats bar** (`page-stats.ts` extraction): D6
cells were dropped from the depth distribution (`dist["d6"]++` → `NaN`,
masked by an `as` cast); gray/no-data cells were counted as green;
`d6Stats` swallowed amber into gray. Now renders the reachable buckets
**D0/D3/D4/D5/D6** (dropped permanently-zero D1/D2), tracks `noData` and
`degraded` distinctly, and validates `parity_tier` fail-loud.
- **Test coverage**: new table-driven cell color/rollup matrix
(`dashboard-color-matrix.test.tsx`) + `page-stats.test.ts` +
order-independence/comment hardening.
- **CR fixes**: effective (stale-downgraded) `.row` from cell-model
resolvers, `WORST_STATE_RANK` rename, composed-cell memo keys,
overlay-types re-export dedup, `setTab` persistence + `window.location`
test stub, `page.tsx` shared-`now` threading + 60s staleness re-render.
## Verification
- **776 tests pass / 1 skip**, `tsc --noEmit` clean, Next.js production
build OK.
- **Local visual proof**: seeded enum D6 rows (30✓/8✗/0 gray) into a
local PocketBase, rebuilt this branch, screenshotted bare `#matrix` —
per-cell D6 green/red renders, and the depth distribution shows `D6:30
D5:8 D4:0 D3:0 D0:588`.
- **7-agent code review converged** over 3 confirmation cycles.
## Notes
- The **#5152 `d6-all-pills` driver must stay on enum keys** (matches
this dashboard's `resolveD6` + the `d5-mapping-drift` test). Do not
revert it to raw catalog keys.
- **Follow-up (separate PR)**: stats-bar single-source-of-truth
hardening (`computeD6Stats`→`buildCellModel`; reconcile
`resolveD6Row`/`resolveD5Row` to the effective row; thread `connection`
into stats for SSE-offline; docs-only stats exclusion); delete
deprecated `composed-cell.tsx` + `deriveDepth`; `useOverlays` mount-hash
deep-link clobber; Notion visualization-doc sync to per-cell
`resolveD6`.
## Test plan
- [ ] CI green
- [ ] After staging deploy, dashboard renders per-cell D6 green/red (not
all-gray) for LangGraph-Python
The depth-distribution row showed permanent-zero D2/D1 rows while the computed
D0 bucket (wired-but-unverified cells) was never rendered, so wired cells
vanished from the row and it never summed to the "Wired" count.
buildCellModel().achievedDepth is typed 0|3|4|5|6 and can never be 1 or 2.
- Remove unreachable d1/d2 from the DepthDistribution type, from
computeDepthDistribution's init, and from the rendered levels array.
- Add a D0 row to the rendered levels so wired-unverified cells are visible
and the distribution sums to the wired-cell count (D6,D5,D4,D3,D0).
- Reconcile section wrapper keys: use each section's stable key instead of the
array index so overlay-toggle reconciliation is correct.
- Document that health/depth/d6 rollups need no dedup: catalogData.cells is
one row per (integration, feature) grid cell (verified: 0 duplicate pairs),
so these per-cell signals are counted exactly once.
Extract the AdaptiveStatsBar aggregate computations out of page.tsx into a
unit-testable pure module (src/lib/page-stats.ts), mirroring the
computeColumnTally pattern, and fix a cluster of correctness bugs:
1. D6 cells were dropped from the depth distribution: DepthDistribution
lacked a `d6` key and the `\`d${depth}\` as keyof` cast produced
`dist["d6"]++ === NaN`. Add `d6` to the type, render a D6 row in
DepthDistributionSection, and replace the cast with an exhaustive
Record<0|3|4|5|6, keyof DepthDistribution> map the compiler checks.
2. d6Stats folded amber (stale/degraded D6) into gray. Count degraded
distinctly and surface it in D6Section.
3. healthStats counted gray (no-data) cells as green, contradicting the
"stats bar matches the matrix" invariant. Track no-data separately and
render it in HealthSection.
4. isSupported is correctly hardcoded true for wired-cell stats: a wired
catalog cell can never be in not_supported_features (generate-registry
resolves those to status "unsupported" before "wired"), so stats and
renderCell cannot diverge. Comment updated to state the invariant.
5. parity_tier was indexed via an unchecked cast (unknown tier →
`undefined++ === NaN`). Validate against the known tier set and skip +
log loud on unknown.
The matrix render path (renderCell/buildCellModel) is untouched.
Add showcase/scripts/sync-promote-service-options.ts: generates the
promote workflow's service `choice` options from the SSOT
(railway-envs.ts), spliced between BEGIN/END markers in
showcase_promote.yml. Fail-loud throughout — every emitted token must
resolve to exactly one service under the resolve-step predicate
(name|dispatchName match AND probe.prod), tokens are YAML-safe, args are
strict (a typo'd flag cannot trigger a destructive write), and markers
are validated before any rewrite.
Wire it into a lefthook pre-commit hook (regenerate + restage; set -e so
a failed regen blocks the commit) and an advisory (never-failing) drift
check in showcase_validate.yml. Vitest coverage for ordering, exclusion,
collision/ambiguity guards, marker errors, exit codes, idempotency, and
the import-side-effect guard.
Add a D5 any-fail fan-out case where the red sub-key is NOT first (greens
then a trailing red) and assert the rollup still resolves red, proving the
worst-state fold is order-independent. Also tighten three imprecise/misleading
comments: clarify the absent-tools D4 worst-state skip, annotate the d6:lgp
aggregate as a distractor not consulted by per-cell resolveD6, and drop the
inapplicable amber-path clause from the aggregate-only gray case.
Stub window.location explicitly via Object.defineProperty (instead of the
global) so the hook's window.location.hash read-path is genuinely exercised
and robust even if window !== globalThis. Add coverage for setTab persistence,
the #baseline parse resolving to the baseline tab, and selectProbe leaving the
tab and hash consistent (ops probe drilldown).
setTab wrote the URL hash but, unlike toggle/updateOverlays, never called
saveToStorage, so a tab switch dropped overlay persistence asymmetrically.
Persist the current overlay set in setTab so the user's overlay selection
survives tab switches and reloads. Covered by a red-green test in the
useOverlays suite.
OverlayToggleBar redefined ALL_OVERLAYS, PRESETS, and OverlayPreset as local
literals duplicating src/lib/overlay-types.ts (drift hazard if an overlay or
preset is added to one but not the other). Replace the local definitions with
re-exports from overlay-types so there is a single source of truth. Behavior
identical; existing component-level imports keep working via the re-export.
ComposedCell.arePropsEqual watched the wrong liveStatus keys: it omitted
agent/chat/tools (which its only consumer, deriveDepth, reads for D2/D4) and
watched an unused smoke key. Watch exactly deriveDepth's reads and fix the
stale "keep in sync with resolveCell" comment.
UnifiedCell rendered a blank cell for a {d6}-only overlay set because d6 is
consumed only by AdaptiveStatsBar and produced no per-cell content. Treat d6
as content-bearing: surface the depth chip + health row (which renders the
per-cell D6 badge). Default overlay set (links/health/depth) is unchanged.
- Rename module-private D5_STATE_RANK to WORST_STATE_RANK (used by both
resolveD5Row and resolveD6Row) and move the misplaced resolveD5Row doc
block onto resolveD5Row.
- Add a test asserting every CATALOG_TO_D5_KEY mapping value is free of the
':' / '/' key delimiters (keyFor's guard did not cover mapping values).
- mergeRowsToMap: fire the collision warning only on genuine state
divergence (rowsAreNoop content compare) instead of reference inequality,
eliminating noisy false warnings.
- isD6Green JSDoc: note the caller only invokes it after D5 is green
(contiguous ladder), so D6 can never be credited over a broken D5.
- computeMaxPossible: treat stub like unshipped (maxPossible=0) so a
not-yet-wired stub cell no longer false-positives isRegression.
resolveD4/D5/D6 returned the RAW status row in `.row` while `.status` was
derived from the stale-downgraded effective state, so a stale-green fold
reported `.row.state === "green"` but `.status === "amber"`. Store the
effective (downgraded) row instead so `.row.state` agrees with `.status`,
mirroring the invariant in live-status.ts `buildBadge`. Also replace
resolveD4's `worstState!`/`winner!` non-null assertions with a guard like
resolveD5/D6, and pin the resolveD3 producer-invariant (D1/D2 gate is
enforced upstream; buildCellModel never reads health:/agent: rows) with a
characterization test and comment.
Make LOCAL_SERVICES_JSON a permanent, env-gated feature of the railway-services
discovery source so the harness (especially the d6-all-pills-e2e probe driver)
can run against LOCAL backends instead of querying Railway.
When LOCAL_SERVICES_JSON is set, enumerate() builds the IDENTICAL
RailwayServiceInfo[] shape from the injected static list and returns it without
consulting Railway creds — only the service URLs differ (local container
hostnames vs Railway public domains). The same namePrefix/nameExcludes filter is
applied, shape is recomputed from the name via classifyShape (single source of
truth), and malformed JSON throws DiscoverySourceSchemaError (same taxonomy as a
Railway shape failure). When the var is unset OR empty, the Railway discovery
path is byte-identical to before.
Critically, the injected demos array is plumbed end-to-end
(demos: svc.demos ?? []) so the d6-all-pills driver's
demosToFeatureTypes(input.demos) produces a real feature matrix; an empty demos
would short-circuit the driver to a zero-cell false-green.
The reference route rendered fenced code blocks as bare, unstyled
<pre> (no highlighting, no copy button) because its MDXRemote call
omitted the rehypeCode plugin and the pre: MdxCodeBlock override that
the main docs pipeline uses. Wire both in (verbatim from the framework
route) so reference code blocks match the rest of the docs. Fixes
rendering for all reference SDKs (React v2/v1 + Core).
Rewrite the two stale live-status keyFor tests that still asserted the old
integration-scoped D6 model (and the wrong e2e-full driver) to assert the
per-cell d6:<slug>/<featureId> shape; expand the two beautiful-chat D5 fan-out
tests to emit all 5 sub-keys so the worst-state fold (not the missing-row path)
is what's exercised; drop the dead unmapped-feature else branch in unified-cell
arePropsEqual (resolvers never read a direct key for unmapped features).
## What & why
[OSS-137](https://linear.app/copilotkit/issue/OSS-137/controlled-gen-ui-demo-optimize-2nd-suggestion-prompt-rename-sidebar)
— the **Controlled Generative UI** demo (`gen-ui-tool-based`) had two
issues, scoped here to **LangGraph-Python** and **Google ADK** (per the
ticket; other 16 integrations roll out later).
### 1. 2nd suggestion didn't reliably render UI
The "Traffic pie chart" chip (`"Show me a pie chart of website traffic
by source."`) names a subject but supplies no numbers, so the agent
**asked the user for data** instead of rendering a chart.
**Fix:** a system-prompt directive (both LGP + ADK agents) instructing
the agent to invent plausible illustrative sample values, call
`render_*` immediately, and **never** reply with a clarifying question.
The suggestion copy stays clean — behavior is carried by the system
prompt, not by leaking "(use sample data)" hints into the UI.
### 2. Sidebar tag → product language
Retagged the demo from `generative-ui` → `controlled-generative-ui` (LGP
+ ADK), so the dojo sidebar pill reads **"Controlled Generative UI"** —
the established taxonomy already used in `shared/feature-registry.json`
and the dashboard catalog.
## Tests
Added D5 aimock fixture entries mirroring all three suggestion chips
(bar / traffic-pie / market-share) so the suggestion-click path has
deterministic coverage. The existing `"revenue by category"` probe
message is **preserved**, so the
[dashboard](https://dashboard.showcase.copilotkit.ai/#matrix:links,health)
D5 row for the edited row stays green.
## Acceptance check
- [x] 2nd suggestion renders UI without asking for data (system-prompt
directive; verified locally against the live agent)
- [x] Sidebar entry tagged "Controlled Generative UI"
- [x] Sample/hallucinated data supplied via system prompt
- [x] Scoped to LGP + ADK
- [x] Tests augmented (D5 fixtures for every chip)
- [x] D5 still shown for the edited row (probe message unchanged)
## Out of scope (left out deliberately)
`package-lock.json` churn from a local reinstall (un-pins `latest`) was
**not** committed.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
The Feature Matrix rendered every D6 cell in an integration's column from
the aggregate PocketBase row `d6:<slug>` (red whenever ANY cell fails),
so genuinely-green cells (e.g. langgraph-python/voice) rendered red. The
harness already emits per-cell `d6:<slug>/<featureType>` rows over the same
featureType keyspace as D5.
Resolve D6 per-cell through the same CATALOG_TO_D5_KEY bridge D5 uses,
mirroring resolveD5/resolveD5Row exactly: per-sub-row stale-green→degraded
fold before the worst-state fold, and strict handling of a missing mapped
sub-row (no-data unless a present sub-row is red). The aggregate `d6:<slug>`
row no longer drives per-cell rendering. Updated the two memo comparators
(composed-cell, unified-cell) to diff per-cell D6 sub-keys, and the stale
doc-comments that described D6 as an integration-level aggregate.
Introduces a third SDK in the reference docs alongside React v2 and v1:
- reference-items.ts: add the 'core' version, generalize root-vs-nested
routing, add 'types'/'enums' subdirs + categories, recognize a core/
slug prefix (literal strip), and emit its static params
- reference-version-selector.tsx: relabel the picker as an SDK switch
(React v2 / React v1 / Core (TypeScript)), import ReferenceVersion from
reference-items, give listbox options role=option/aria-selected
- app/reference/page.tsx: rename 'API Reference' to 'Overview' and add a
'Choose your SDK' card chooser
The staging D6 dashboard rendered no data because the client did a direct
cross-origin fetch to the harness URL — CORS-blocked and the wrong path
(`/probes` instead of `/api/probes`). Root cause: getRuntimeConfig() sourced
the client RuntimeConfig.opsBaseUrl from the server proxy target OPS_BASE_URL
(the harness URL), which the root layout serialized into
window.__SHOWCASE_CONFIG__, so resolveBaseUrl() used it as the fetch base
instead of falling through to the same-origin /api/ops proxy.
Decouple the two: the client direct override is now an explicit, opt-in,
client-intended env var (NEXT_PUBLIC_OPS_DIRECT_BASE_URL) that defaults to ""
in every environment, so the client lands on /api/ops. The Route Handler's
server-only OPS_BASE_URL read and its sentinel behavior are unchanged.
- showcase/shell-dashboard/src/lib/runtime-config.ts: opsBaseUrl now read from
NEXT_PUBLIC_OPS_DIRECT_BASE_URL (default ""), not OPS_BASE_URL; no sentinel.
- showcase/shell-dashboard/src/lib/ops-api.ts: update resolveBaseUrl comments
to describe the client override vs server proxy target distinction.
- showcase/shell-dashboard/src/app/layout.tsx: note opsBaseUrl is the client
override, never the harness URL.
- tests: encode the contract (client falls through to /api/ops when no direct
override; server proxy target does not leak into the client config), incl. an
end-to-end regression in the env-switch spike test.
The /api/ops/* proxy was a next.config.ts rewrite, which Next.js evaluates at
`next build` and freezes into the prebuilt Docker image. The shared CI build
bakes a placeholder OPS_BASE_URL (http://ops.invalid) to satisfy a
throw-if-unset guard, so every deploy of the single :latest image proxied to a
dead host regardless of its runtime env — /api/ops/* returned 500
(ENOTFOUND ops.invalid), the Feature Matrix health overlay got no probe data,
and every cell downgraded to amber (zero green) despite green backend data.
The same freeze also baked the production harness URL into the shared image,
so even a correct rebuild would point staging at the prod harness.
Replace the build-time rewrite with a Route Handler at
src/app/api/ops/[...path]/route.ts that reads process.env.OPS_BASE_URL at
REQUEST time (force-dynamic, never statically cached) and proxies
/api/ops/<path> -> ${OPS_BASE_URL}/api/<path>, forwarding method, query
string, headers, and body. A missing OPS_BASE_URL now returns a clear 503
instead of a build throw. The build no longer depends on OPS_BASE_URL.
This fixes both traps with one image: each environment resolves its own
runtime OPS_BASE_URL from the same artifact, no rebuild. Requires the
dashboard image to rebuild + redeploy. Harness path is now /api/probes.
The gen-ui-headless-complete probe
(showcase/harness/src/probes/scripts/d5-gen-ui-headless-complete.ts)
references fixtureFile "gen-ui-headless-complete.json", but that file
existed in no D6 slug, so the probe's first leg 503'd under strict.
Add gen-ui-headless-complete.json for all 18 D6 slugs, modeled on the
green langgraph-python headless-complete.json pattern: the four gen-UI
pills (weather/stock/highlight/revenue) with narration (toolCallId)
fixtures FIRST and toolcall (userMessage+context) fixtures AFTER, no
turnIndex gate, so every one of the probe's four sequential turns in one
chat thread matches regardless of prior assistant/tool history. Interrupt
fixtures are intentionally omitted (interrupt-headless is a separate
cluster item).
These 8 fixtures per slug share match keys with the pre-existing
headless-complete.json for the same context (the demos share pills and
are disambiguated at runtime by probe path), which raises the
aimock-fixtures collision-detection exact-duplicate count by 46. Bump
KNOWN_DUPLICATE_CEILING 230 -> 276 to match, consistent with how prior
per-integration feature fixtures bumped the baseline; the substring-shadow
ceiling is unchanged (no new shadows).
Validated: all 732 aimock-fixtures schema/collision tests pass, and a
local aimock --strict run returns 200 for each of the four pills across
turns (langgraph-python + pydantic-ai spearheads, plus spring-ai).
D6 drives each feature's pills as sequential turns in one chat thread.
aimock matches turnIndex against the count of assistant messages in
history, so a "turnIndex": 0 gate on a per-pill tool-call leg only
matches the FIRST pill — pills 2+ (turnIndex 1/2/3) match no fixture,
503 under strict, and the harness reports "timeout: assistant did not
respond".
Remove the turnIndex:0 gate from the multi-turn tool-call legs of
frontend-tools (sunset/forest/cosmic) and tool-rendering-reasoning-chain
(the three chained pills: Compare AAPL/MSFT, compare-to-smaller dice,
weather-there flights) so they match on their distinct userMessage
substrings + context, mirroring the already-green langgraph-python
fixtures. Each pill's substring is unambiguous and its toolCallId-keyed
follow-up fixture is ordered before the de-gated first leg (first-match
wins), so follow-up narration still wins and there is no cross-matching.
The reasoning-chain "Find flights from SFO to JFK." pill keeps its
turnIndex:0 gate to match the green langgraph-python reference exactly.
Proven locally with aimock --strict: multi-turn requests for these pills
returned 503 with the gate (negative control) and now return 200 across
all turns with the correct fixture/narration.
Three reviewers found two faces of one gap: a context being opened in
openContextOn takes its cap reservation BEFORE `await newContext()` but is
added to entry.liveContexts only AFTER, so during that await liveContexts.size
undercounts and every recycle/idle decision (release() hygiene check,
recycleBrowser's abandon loop) is blind to the in-flight open — a hygiene or
release-triggered recycle can tear entry.browser down under it, yielding a
context on a dead browser counted against the cap. The `!hadWaiter` hygiene
gate also starved the recycle under sustained load.
Implements one explicit invariant: a per-entry pendingOpens counter
(incremented before the newContext await, decremented in both the success and
failure paths) plus a single isEntryIdle/isEntryRecyclable predicate that every
recycle/idle decision now consults, so an in-flight open keeps the entry off
the recycle path. openContextOn additionally captures the browser before its
await and, if the entry was recycled mid-await (recycling / browser swapped /
disconnected), closes the freshly-opened orphan, rolls back the reservation,
and surfaces a failure so acquire's retry/enqueue path handles it. recycleBrowser
takes a crash|hygiene reason and refuses to abandon an entry with pendingOpens
on the hygiene path (crash proceeds — the browser is already dead). GAP 2 is
closed with an entry.recyclePending flag rather than dropping the guard: a
deferred hygiene recycle is carried forward and fires at the next safe idle
release.
Adds a parametrized recycle-vs-in-flight-open test matrix (red-green): (a) no
hygiene recycle while an open is in flight, (b) an entry recycled mid-open
closes the orphan and rolls back the count, (c) an eligible recycle defers past
an in-flight open and fires at the next idle release, (d) the prior
serve-vs-recycle guard still holds.
Apply the same pooled-launcher abort hardening to the e2e-demos launcher
(detach the abort listener in close(); refuse + release contexts opened after a
pre-aborted signal). Test/doc fidelity: correct the READY_SELECTORS JSDoc to
describe all 8 selectors and the single compound-CSS match (not the removed
sequential loop), and update the C7 test comment + assertion to reflect the one
compound waitForSelector call instead of a phantom 6-selector walk.
The boot success-path status write stamped `recovered: true` unconditionally,
misreporting a phantom recovery on a cold first boot. Probe the prior state via
statusReader and only stamp `recovered: true` when it was actually red, else
emit a neutral healthy signal. Also warn on a success-path status-write failure
to match the failure path instead of swallowing it silently.
A3: the d4 and e2e-parity pooled context-wrappers now delete from the abort
tracking set on normal close() (mirroring d5/d6/demos), and their test fakes'
release() is idempotent (tracks a liveContexts Set, no-ops on unknown/double
release) so a normal-close-then-abort sequence can't double-release or drive
inUse negative. Across all five pooled launchers (d4, d5, d6, e2e-parity,
e2e-readiness): capture the abort listener and detach it in the launcher-level
close() so a post-completion abort can't fire after the run returned, and in
the pre-aborted branch refuse + immediately release any context opened after
the signal already aborted so it can't leak into a torn-down run.
A1: gate the hygiene recycle on `!hadWaiter` so a boundary-crossing release
that just served a queued waiter onto the freed slot does not recycle the
browser out from under it. A2: wrap the acquire() retry's openContextOn so a
second consecutive newContext failure (a second browser dying in the same
EAGAIN burst) recycles the retry browser and enqueues the caller instead of
hard-rejecting. Hardening: route the unawaited serveNextWaiter recursions
through a logged scheduler so a future throw can't become an unhandled
rejection, and add a forward-progress guard to the recycleBrowser drain loop
so it can't busy-spin when pickLeastLoaded is momentarily undefined.
createPooledE2eParityLauncher and createPooledE2eSmokeLauncher did not
observe the launcher abort signal, so on a driver hard-timeout or external
abort the BrowserContexts they checked out via pool.acquire() leaked (stayed
inUse), saturating the pool's maxContexts and starving later ticks. Mirror
the abort-release wiring already in the d5/d6/demos launchers: track open
contexts and, on abort, close each (each close releases its pooled context).
Widen both launcher types to accept abortSignal, pass abort.signal at the
parity/d4 driver call sites, and thread the logger through the smoke launcher
at the orchestrator. Adds abort-release tests to both driver test files.
The orchestrator pre-parsed BROWSER_POOL_BROWSERS / BROWSER_POOL_SIZE /
BROWSER_POOL_MAX_CONTEXTS with Number(...) and passed the result as explicit
browsers/maxContexts options. Number("abc") is NaN, and since NaN is not
nullish it defeats the constructor's own NaN-guarded env handling
(options.browsers ?? <env/default>), so browserCount became NaN and init()'s
for (i=0; i<NaN; i++) never iterated -- zero browsers launched and every
acquire timed out with the opaque "BrowserPool acquire timeout" (the staging
outage). Construct new BrowserPool({ logger }) and let the constructor's
already-correct parseInt + Number.isNaN + >0 guarded env resolution own the
numeric values. Adds an orchestrator-side regression test (red->green).
Satisfy oxlint consistent-function-scoping for the new orphan/retry test
launchers by using a module-scoped testSleep helper instead of nested
closures.
On the per-feature timeout path the Promise.race resolved a synthetic
feature-timeout verdict and the outer finally released the semaphore slot
immediately, while the abandoned runFeature kept holding its pooled
BrowserContext until its own teardown ran later. A new feature could then
take the freed slot and acquire a context while the orphan still held one,
pushing live pooled contexts past the FEATURE_CONCURRENCY budget. Gate
sem.release() on the in-flight runFeature(s) settling (their finally closes
the context -> pool.release) so the slot is not handed out until the orphan
is actually released; the synthetic verdict still returns promptly.
Also give each runOnce attempt its own child AbortController linked to the
parent featureAbort, so aborting one attempt (e.g. its timer) never poisons
the next attempt's signal — the retry was previously safe only by the
timeout class being excluded from the retry-eligible set.
Improve the d5/d6 fake context-pools to mirror BrowserPool.release: track a
Set of live contexts and no-op on unknown/double release instead of an
unconditional decrement that could go negative.
Applied identically in d5 (e2e-deep) and d6 (e2e-full).
When the per-demo wall-clock budget is exhausted, remaining() returns 0.
Playwright treats timeout: 0 as wait-forever, so page.goto /
page.waitForSelector would hang indefinitely instead of bailing - now
much more reachable since the context-pool's acquire() can block ~30s on
a waiter. Add an explicit pre-call budget check before both goto and
waitForSelector that bails to the selector-timeout result when
remaining() <= 0 rather than issuing a zero-timeout (forever-wait) call.
Also make the test fake context-pool mirror the real BrowserPool.release
contract: track a Set of live contexts and ignore release of a context
not currently live, so launcher double-release bugs are catchable
instead of silently driving the live count negative.
(--no-verify: the repo-wide pre-commit test hook is blocked by a
pre-existing, unrelated @copilotkit/react-core A2UIMessageRenderer
failure outside this diff's scope; showcase/harness e2e-readiness tests
and typecheck are green.)