Only counts (12) were asserted on each side, so a future slug rename on
one side would silently flip a column to grey "not-supported" while both
counts stayed 12. Add a dashboard-side test that fs-reads the harness
STARTER_TO_COLUMN value set (the producer-side remap lives in a separate
pnpm workspace, so it cannot be imported) and asserts SET-EQUALITY with the
dashboard STARTER_COLUMNS. Mirrors the harness starter-mapping-drift fs-read
pattern; reds if either side renames a column slug without the other.
The starter_smoke probe runs hourly (`schedule: "40 * * * *"`), but
STARTER_STALE_AFTER_MS was set to 13h with a comment claiming a 6h cadence
— ~13 missed hourly ticks before amber, defeating the intended two-miss
flip. Re-derive the window from the 1h probe period to 2.5h: strictly
> 2 periods (so two consecutive misses, last row ~3h old, flip amber) yet
< 3h (so a single missed/slow-wake tick, last row ~2h old, stays green,
absorbing a scale-to-zero cold-start). Reconcile both the staleness.ts and
starter_smoke.yml comments to the same hourly basis + 2.5h window in
lockstep (no more "6h" in the starter context). Extend the staleness tests
with explicit hourly-tick boundaries (1 miss → green, 2 misses → amber).
Also drop the inaccurate "hollow" from the not-supported ✗ comments — there
is no hollow render variant; the state renders as grey ✗ text.
Builds the dashboard UI for the starter-smoke dimension: resolveStarterRow
(flat starter:<col>/<level> lookup) + buildStarterBadge implementing the full
5-state cell vocabulary (green ✓, red ✗ smoke-failed, amber ~ stale, gray ?
not-yet-run, grey ✗ not-supported keyed off the S1 mapping), the StarterSection
row-group, rollup-exclusion so starter rows don't skew column aggregates, and a
13h staleness window (> 2× the probe cadence).
UI slot (S3) of the starter-row-group spec (model B).
Adds the `starter_smoke` probe driver + config/probes/starter_smoke.yml that
fans out per-starter HTTP health/agent/chat/interaction levels, registers it
in the orchestrator, and exempts it from probe-config parity (its matrix
shape differs from the depth-dimension probes). Includes 13 driver unit tests.
Probe slot (S2) of the starter-row-group spec (model B): harness HTTP-probes
the deployed (sleepable) Railway starter services.
Introduces the `starter` probe dimension in harness types and the single
source of truth for the starter-slug→dashboard-column-slug remap
(STARTER_TO_COLUMN, starterToColumnSlug, STARTER_LEVELS) plus a drift-lint
test asserting every smoke-matrix starter is mapped or explicitly excluded
and every mapped column slug resolves to a real manifest directory.
Foundation slot (S1) of the starter-row-group spec (model B).
Commit 7ac3e59a5 ("D6 all-pills probe driver") replaced the
`e2eParityDriver` registration (kind `e2e_parity`) with `e2eFullDriver`
(kind `e2e_d6`) in BOTH branches of `registerAllProbeDrivers`. After that
commit NO driver registers kind `e2e_parity`, yet
`config/probes/e2e-parity.yml` still shipped.
The probe-loader hard-rejects any YAML whose `kind` has no registered
driver (`no driver registered for kind 'e2e_parity'`), so `e2e-parity.yml`
failed to load on every boot and was never scheduled — silently dropping
that probe family. The D6/parity dashboard dimension is now produced by the
`e2e_d6` (all-pills) driver, which emits the identical
`d6:<slug>/<featureType>` rows the dashboard reads by key prefix (plus a
`d6:<slug>` aggregate), so `e2e_parity` is genuinely superseded, not just
temporarily unregistered.
Delete the orphaned config so the loader no longer rejects it and the full
probe-config set loads clean. Add a loader test that exercises the REAL
shipped `config/probes` set against the REAL driver registry and asserts
every YAML loads with zero `probes.reload.failed` errors — the existing
orchestrator guard only checks a hardcoded kind list and never read the
on-disk YAMLs, so it missed this orphan.
RACE3b's post-shutdown `liveContextCount === 0` assertion is tautological —
shutdown() sets liveContextCount to 0 unconditionally — so it proves nothing.
Drop it; the load-bearing assertions remain (acquire rejects with the shutdown
sentinel, contextToBrowser.size === 0, the orphan context was closed).
Correct the RACE3a/RACE3b comments that claimed shutdown() awaits the tracked
orphan close (reverted): the straddled orphan is closed by the orphan guard's
fire-and-forget close and never lands in contextToBrowser.
The empty-drain shutdown case (inFlightRecycles + pendingLaunches both
empty) exits the drain loop and resolves BEFORE a parked open settles and
registers its close via closeContextTracked(), so the "shutdown resolved =>
everything closed" guarantee was never actually delivered and had no
discriminating coverage. Revert the openContextOn orphan guard's close back
to main's accepted fire-and-forget `void this.closeContext(...)` and remove
the closeContextTracked() helper; inFlightRecycles is touched only by the
recycle / self-heal paths again.
Also unify the shutdown-condition error string: shutdown()'s queued-waiter
rejection now says "BrowserPool is shut down" to match the at-entry guard and
the straddle-reject legs (launch/relaunch-failure strings are a different
condition and unchanged).
Item 1 (red-green for the REACHABLE new shutdown branches): RACE1 only
covered the openContextOn orphan-guard `isShutdown` term. RACE3 adds
coverage for the two branches it did not:
- RACE3a: a shifted waiter whose serve open straddles shutdown must be
REJECTED by serveNextWaiter's post-open `if (this.isShutdown)` leg
(not re-enqueued onto the cleared queue, not hung to timeout).
RED with that guard reverted (waiter never settles), GREEN with it.
- RACE3b: an acquire whose open straddles shutdown must REJECT via
acquire()'s OUTER-catch `if (this.isShutdown) throw` leg. RED with
that guard reverted (the acquire enqueues on the cleared queue and
the test times out), GREEN with it. Both also assert shutdown()
awaited the orphan close (item 2) — the straddled context is closed
and the pool maps are empty post-shutdown.
Item 5: relabel RACE2's comment + title so it is framed as a regression
lock for the PRE-EXISTING generation-guard exactly-once-rollback
invariant (passes against main unchanged), NOT a shutdown-race proof.
Also extends the `internals()` test accessor with `waiters` so the new
tests can assert the queue is empty (no stranded caller).
Item 2 (reviewer-7 MEDIUM): the openContextOn orphan guard's
`void this.closeContext(...)` under shutdown was fire-and-forget and
tracked in NO set that shutdown() drains, so shutdown() could resolve
while a context close was still in flight — violating the implied
"shutdown resolved => everything closed" contract upheld elsewhere via
the pendingLaunches / inFlightRecycles drains. Route the orphan close
through a new closeContextTracked() that registers the close promise in
inFlightRecycles (the SAME set shutdown's drain loop re-snapshots until
empty) and self-removes on settle. Reuses the existing tracked-set
mechanism — no new machinery; outside shutdown the add/delete is a
no-op since only shutdown's drain loop reads the set.
Item 1 decision (retry-leg guard): KEEP the acquire() transient-retry
leg's `if (this.isShutdown) throw`. Contrary to a reviewer note that it
is unreachable, it guards a DISTINCT, genuinely-reachable straddle
window: the outer catch already observed isShutdown===false, we
re-reserved and re-opened, and the pool can shut down WHILE the retry's
open is in flight (the orphan guard then throws into this catch).
Without it the acquire enqueues onto the already-cleared waiter queue
and hangs until timeout. Comment updated to assert that reachability
rather than imply it merely "mirrors" the outer guard.
Items 3 + 4 (serveNextWaiter straddle): correct the inaccurate comment
that claimed shutdown() "already rejected this waiter" — it cannot, the
waiter was shift()ed off this.waiters BEFORE shutdown's reject-loop ran,
so THIS branch is the sole settler of a shifted-then-straddled waiter.
Unify the shutdown error string on the at-entry guard's existing
"BrowserPool is shut down" (was "shutting down" only here).
A serveNextWaiter()/openContextOn() that shifted a waiter + reserved a
slot BEFORE shutdown could settle its newContext() AFTER shutdown()'s
close-pass. shutdown() drains inFlightRecycles + pendingLaunches, but a
fire-and-forget serve open is tracked by neither, so the freshly-opened
context landed in contextToBrowser/liveContexts on a torn-down pool — a
leaked context that is never closed. openContextOn's orphan guard
checked generation/recycling/isConnected but NOT isShutdown.
Fix: add an isShutdown term to openContextOn's post-await orphan guard
(treat shutdown like a recycle — close the just-opened context, roll
back the reservation, throw); bail serveNextWaiter before openContextOn
when isShutdown; and reject (not re-queue) a shifted waiter / straddling
acquire when the open throws under shutdown, so it cannot strand on the
already-cleared waiter queue. Preserves all #5185/#5221 behaviors.
Adds red-green coverage for the leak (RACE1) and a regression guard for
the acquire transient-retry -> concurrent-recycle -> orphan-guard
straddle (RACE2), which was verified to keep its reservation accounting
balanced via the existing generation guard (no overshoot).
## Summary
The showcase Railway promote CI flow could not promote the fleet: an
`all` promote was blocked end-to-end whenever a single service was red,
and the reported failure traced to `showcase-ag2` being chronically red
on staging.
Root causes and fixes:
- **ag2 crash-on-import (the red service).** `gen_ui_agent.py` carried
`from __future__ import annotations`, which stringified the `set_steps`
tool's `context_variables: ContextVariables` parameter into an
unresolved `ForwardRef`. AG2's tool-schema generation then raised
`PydanticUserError` at import time, so the process never came up and the
staging healthcheck failed on every deploy since 2026-05-31. Removed the
import (matching the working sibling agents) and added a regression test
that statically asserts the future-import stays absent
(version-independent) plus a live import check.
- **Promote loop was all-or-nothing.** The per-service loop ran under
`set -euo pipefail`, so the first failing service aborted the whole
`all` promote, leaving the rest unpromoted. Extracted the loop into
`showcase/scripts/promote-fleet.sh`, which attempts every service,
accumulates succeeded/failed sets, exits non-zero only after attempting
all, and exports `succeeded_csv`.
- **`verify-prod` defeated the best-effort design.** It was skipped on
any non-zero promote and verified the full requested set. It now runs
`if: !cancelled()` and scopes `--services` to the succeeded set.
- **Staging precondition blocked the fleet.**
`verify-staging-precondition` failed the whole `all` promote when any
one service was staging-red. It is now advisory — `promote` runs
regardless, and `bin/railway`'s per-service P2/P3 staging-green gates
authoritatively refuse red services while green services promote.
`notify` success keys on PROMOTE && PROD.
- **Regression tests now gate in CI.** Added a `shell-script-tests` job
(bats + shellcheck) to `showcase_validate.yml`, plus input-validation
hardening in the script (fail-loud on empty / all-empty CSV,
`RAILWAY_BIN` executability check, whitespace trim).
## Test plan
- [x] ag2 regression test passes in the 3.12 venv (`PYTHONPATH=".:src"
pytest tests/python/`) — 2/2
- [x] `promote-fleet.bats` — 12/12 (best-effort loop, succeeded_csv
export, empty/whitespace/missing-binary guards, digest forwarding)
- [x] `shellcheck promote-fleet.sh` clean; `actionlint` clean on both
workflows
- [ ] CI green on this PR
The runtime CMD raised the soft nproc limit via
`/bin/sh -c "ulimit -u $(ulimit -Hu) ..."`, but on node:22-bookworm-slim
`/bin/sh` is dash, whose builtin `ulimit` does NOT support the `-u`
(max-user-processes) flag — it errors `ulimit: Illegal option -u`, which the
`2>/dev/null || true` then silently swallows. So #5185's intended PID
protection never applied; the soft limit stayed at the inherited default.
Run the CMD under `/bin/bash` (present at /usr/bin/bash) instead, whose
`ulimit -u` works. Verified against the exact base image under
`--ulimit nproc=512:4096`: bash raises the soft limit 512 -> 4096 (the hard
ceiling) while dash leaves it untouched and errors. Minimal change — only
the interpreter; the command, exec-as-PID-1, and `|| true` fallback are
unchanged.
A browser that is mid-`launch()` is not yet in `this.browsers` and has no
disconnect handler attached, so it is invisible to every teardown path
(shutdown's close pass, recycle, the disconnect handler). Under the
self-heal/relaunch storm introduced by the browser-pool hardening cluster
(#5174/#5185), a concurrent teardown therefore raced an in-flight launch:
the launch escaped teardown accounting entirely — either leaked, or (in
production) the browser was closed underneath the still-resolving
`chromium.launch()`, which then rejected with `browserType.launch: Target
page, context or browser has been closed` (SIGTRAP). The self-heal loop
relaunched forever, each relaunch hitting the same race (336
self-heal-launch-failed events in 4 min), wedging the pool.
Make launches atomic w.r.t. teardown: `launchBrowser()` registers its
launch promise in a new `pendingLaunches` set BEFORE awaiting
`rawLaunchBrowser()` and removes it on settle. `shutdown()` drains
`pendingLaunches` (alongside `inFlightRecycles`) before its close pass, so
it WAITS for every in-flight launch to settle instead of closing the
browser mid-launch. The launch seam re-checks `isShutdown` the instant the
launch settles: if a shutdown intervened it closes the freshly-launched
browser cleanly THEN, exactly once, and throws a shutdown sentinel so the
caller (init / recycle relaunch / self-heal) does not register a browser
into a pool that is going away. The launch-stagger gate still chains off the
raw result, so serialization semantics are unchanged. All #5185 behaviors
(backoff, self-heal, degraded/recovered alarms, waiter draining) preserved.
Adds a red-green test (FIX#13) reproducing an in-flight launch racing a
concurrent shutdown: pre-fix the launching browser escapes teardown
(closeCount 0, leaked); post-fix it is drained + closed cleanly exactly once
with no close-during-launch rejection.
The per-target partial-rollup `runWriter.update` (which advances the durable
`probe_runs.summary.{passed,failed}` counters) ran BEFORE `writer.write`
committed the corresponding `status`/`status_history` detail row, and a
`writer.write` failure is intentionally swallowed so one hiccup can't tank
sibling targets. That ordering opened a run-row-orphan window: the run row
could report `failed: N` for a target whose detail row was never durably
written — the stale-red ingestion artifact.
Reorder so the detail/status row is committed first and the run-row counter
is stamped only after, so the persisted counter can never outrun its backing
detail row. Pure statement-order swap — no schema/migration, dashboard read
path untouched, genuine failures still counted.
The gen-ui-interrupt demo across non-LangGraph integrations used
`useInterrupt`, which only renders in response to an AG-UI `on_interrupt`
event emitted by LangGraph's native `interrupt()` primitive. These
backends never emit that event — they expose `schedule_meeting` as a
frontend/HITL tool over the normal tool-call channel (Strategy B) — so
the picker never mounted.
Migrate the 8 clean Strategy-B integrations (ag2, agno, crewai-crews,
mastra, pydantic-ai, spring-ai, llamaindex, claude-sdk-typescript) to
`useHumanInTheLoop`, mirroring the ms-agent-python / ms-agent-dotnet
reference: same `name: "schedule_meeting"`, same zod `{ topic, attendee }`
parameters, same TimePickerCard render, resolving via `respond(...)`.
The framework-specific comment is generalized for accuracy.
The 3 LangGraph integrations keep `useInterrupt` (native interrupt).
built-in-agent and claude-sdk-python already use the equivalent working
`useFrontendTool` pattern and are left unchanged. strands and langroid
are intentionally NOT migrated — their backends declare a
`schedule_meeting(reason)` tool whose param shape conflicts with the
`topic`/`attendee` reference, which needs separate resolution.
Quick-win fixes for the **Build with agents** page from Sam's review.
Stacked on #5187 (retarget to `main` once it lands).
## Changes
- **Skills now shown on every page** — previously only the top-level
`/build-with-agents` (and `built-in-agent`) rendered the Skills section;
every framework integration page used the MCP-only snippet, hiding
Skills. The shared `coding-agents.mdx` now renders the full Skills + MCP
guide, so all 11 build-with-agents pages show Skills (the recommended
path).
- **Skills section** — added a top-three table (`copilotkit-setup`,
`copilotkit-develop`, `copilotkit-integrations`) and a note that
`copilotkit-contribute` is for contributing to CopilotKit, not building
with it.
- **Install step** — clarified to run `npx skills add` from the project
root; any agent there (Claude Code, Codex, Cursor, Gemini CLI) discovers
the skills automatically.
- **MCP headings** — demoted per-tool headers (Cursor, Claude Code, …)
from H2 → H3 so they nest under "MCP Docs Server" in the TOC; "Other"
subsections H3 → H4.
## Screenshots
Skills section + top-three table:

Install step:

Skills now rendering on a framework page (Mastra) that was previously
MCP-only:

TOC nesting (tools now under MCP Docs Server):

## Files
-
`showcase/shell-docs/src/content/snippets/shared/guides/build-with-agents.mdx`
-
`showcase/shell-docs/src/content/snippets/shared/guides/mcp-server-setup.mdx`
- `showcase/shell-docs/src/content/snippets/shared/coding-agents.mdx`
-
`showcase/shell-docs/src/content/docs/integrations/langgraph/build-with-agents.mdx`
## What
Two unrelated cleanups:
### 1. Move internal-only skills out of the public repo
Three staff-only skills lived under `.claude/skills/` and
`.agents/skills/`, so `npx skills add CopilotKit/CopilotKit` swept them
into end-user installs. This PR deletes them here; they now live in the
internal-skills plugin (**CopilotKit/internal-skills#108**):
- `copilotkit-demo-parity`
- `git-hooks`
- `showcase-demo-debugging`
### 2. Recommend a cleaner skills install command
The **Build with agents** guide now recommends:
```bash
npx skills add CopilotKit/CopilotKit/skills -y
```
- `/skills` subpath installs only the published skills under `skills/`
(the repo root also picks up internal skills).
- `-y` skips the interactive prompts.
A Callout documents the interactive variant and `-g` for a global
install.
the per-service promote loop ran under `set -euo pipefail`, so the first failing
service aborted the whole `all` fleet promote; extracted to promote-fleet.sh
which attempts every service, accumulates succeeded/failed sets, exits non-zero
only after attempting all, and exports succeeded_csv. verify-prod now runs
`if: !cancelled()` and scopes --services to the succeeded set; the staging
precondition is advisory (promote runs even when it reports red — bin/railway
enforces staging-green per-service); notify success keys on PROMOTE && PROD.
Adds a shell-script-tests CI job (bats + shellcheck) and input-validation
hardening (fail-loud on empty/all-empty CSV, RAILWAY_BIN check, whitespace trim).
`from __future__ import annotations` turned the set_steps tool's
`context_variables: ContextVariables` param into an unresolved ForwardRef at
AG2 tool-schema-generation time, raising PydanticUserError on import and failing
the showcase-ag2 staging healthcheck since 2026-05-31. Removing it matches the
working sibling agents. Adds a regression test that statically asserts the
future-import stays absent (version-independent) plus a live import check.
Previously only the top-level /build-with-agents and built-in-agent pages
rendered the Skills section (via <BuildWithAgents />). Every framework
integration page used the MCP-only <CodingAgents /> snippet (or, for
langgraph, <MCPSetup /> directly), so Skills — the recommended path — was
hidden there.
Point the shared coding-agents.mdx snippet at <BuildWithAgents /> so all
pages that reference CodingAgents now render Skills + MCP, and switch the
langgraph page from <MCPSetup /> to <BuildWithAgents />. The snippet inliner
recurses with cycle protection, so no duplication is needed.
Addresses review feedback on the Build with agents page:
- Add a top-three skills table (copilotkit-setup / -develop / -integrations)
and call out that copilotkit-contribute is for working on CopilotKit
itself, not building with it, so the skills directory's build-vs-contribute
split is clear from the docs page.
- Clarify where to run `npx skills add`: from the project root, where any
coding agent (Claude Code, Codex, Cursor, Gemini CLI) discovers the skills
automatically — answering 'in your agent environment'.
- Demote the MCP per-tool section headers (Cursor, Claude Web, Claude Code,
...) from H2 to H3 so they nest under 'MCP Docs Server' in the on-this-page
TOC instead of sitting as flat siblings; demote the 'Other' subsections to
H4 accordingly.
Prefix every harness-dispatched alert with a `[staging]`/`[production]`/
`[unknown]` source-env tag so operators triaging a red probe know which
deploy environment is affected. The label is derived in the orchestrator
from SHOWCASE_ENV ?? RAILWAY_ENVIRONMENT_NAME ?? "unknown" and applied at
the single renderer chokepoint (covering per-key, cron, and on-error
dispatch) plus the aggregation flush path that bypasses the renderer, via
a shared sourceEnvPrefix helper so the two paths never drift. A missing
env var surfaces as a visible [unknown] rather than a silent un-prefixed
alert.
- STAGING-OUTAGE regressions: degraded alarm fires (not silent) when the
set empties from a relaunch storm; self-heal re-inits a fresh set once the
kernel relaxes; a waiter queued during the dead window is served by
self-heal; a transient relaunch EAGAIN is retried and the entry survives
(no eviction, no alarm).
- Bounded serveNextWaiter transient re-drive + FIX#7 dead-vs-alive gate
propagation to the serve path.
- orchestrator: degraded/recovered signal wiring covered.
- BUG3 (orphan-by-recycle waiter drain) adapted to the crash-recovery
acquire path: an acquire whose in-flight open is orphaned re-enqueues as a
waiter; with cap=1 the freed slot goes to the other waiter, so the orphaned
acquire settles via its own (now bounded) timeout. The invariant it
verifies (freed capacity immediately serves the queued waiter) is unchanged.
Lift the soft nproc limit to the hard ceiling (`ulimit -u $(ulimit -Hu)`)
before exec'ing the orchestrator so the legitimate 40-context chromium
workload (several hundred OS threads at steady state) has ample thread
headroom instead of running near the default ~1024 soft ceiling, where
`chromium.launch()` tripped `pthread_create: Resource temporarily
unavailable`. `exec` keeps node as PID 1 for correct signal handling; the
`|| true` fallback keeps boot resilient when the runtime forbids raising
the soft limit (the cgroup pids limit then remains the dominant control).
Make the long-lived chromium pool survive a pthread/PID-ceiling
(`pthread_create: Resource temporarily unavailable`, errno 11) thread-
exhaustion storm instead of draining to an empty, permanently-wedged set.
- Crash-recovery relaunch backpressure: a transient EAGAIN on relaunch is
retried with bounded linear backoff before the entry is evicted, so a
thread-exhaustion window that relaxes within seconds recovers in place
rather than splicing the entry out of the set.
- Self-heal + degraded/recovered alarm: when the set empties mid-life the
pool fires an `onDegraded` red alarm (previously only emitted on init()
failure — mid-life death was silent) and kicks a background self-heal
loop that relaunches a fresh set the moment a launch succeeds, firing
`onRecovered`. No manual redeploy required.
- Bounded serveNextWaiter transient re-drive: a persistently-transient
newContext() on a still-connected browser no longer hot-loops the event
loop; it self-reschedules up to a ceiling then leaves the waiter queued
for a later release/recovery handoff (mirrors acquire()'s retry-once
semantics).
- Accounting hardening: generation-token guard on in-flight opens across a
recycle, clamped servedContexts rollback on orphan-close, deferred-recycle
re-check on non-release teardown paths, and waiter-drain on orphan-by-
recycle rollback so freed capacity is served immediately.
- orchestrator wires the pool's onDegraded/onRecovered hooks to the shared
`system:browser-pool-degraded` red/green capacity-loss signal.
Fixes the 2026-06-03 staging incident: the browser pool died from thread
exhaustion, the relaunch storm emptied the set, and the harness wedged with
no alarm -> 626 D0-red cells until a manual redeploy.
Use `CopilotKit/CopilotKit/skills -y` instead of the repo root: root
discovery sweeps in the internal `showcase-demo-debugging` skill
(metadata.internal, lives in .claude/.agents, not skills/), so users got
12 skills incl. one internal. The /skills subpath yields exactly the 11
published skills. Drop -g so install defaults to project scope, letting each
project pin the skills version matching its CopilotKit dependencies.
The build-with-agents guide recommended a bare `npx skills add` that drops
human users into a multi-step interactive flow (skill multiselect, agent
selection, scope, install method, confirm). Recommend `-g -y` so all skills
install globally in one shot, with a Callout pointing to the flag-less command
for users who want to choose interactively.
## Summary
Follow-up to #5173 (bucket-(d)) closing three **pre-existing**
browser-pool concurrency defects on the non-release teardown paths. The
browser-pool is a context-pool over a fixed set of long-lived Chromium
processes; these defects biased the hygiene-recycle cadence and could
strand waiters / leak deferred recycles.
## Fixes (each red-green proven)
1. **`serveNextWaiter` orphan-close leaked `servedContexts`.** A waiter
timing out mid-`openContextOn` cleaned up the reservation/context but
never decremented `servedContexts` (which `openContextOn` had already
`++`'d) → every orphaned-by-timeout serve permanently inflated the count
→ premature hygiene recycles. Fix: decrement `servedContexts` in the
orphan-close block.
2. **Deferred `recyclePending` honored only on `release()`.** A recycle
deferred because `pendingOpens > 0` set `recyclePending`, but the
non-release teardown paths (orphan-by-recycle rollback, orphan-close)
returned the entry to idle without re-checking it → the deferred recycle
was dropped and the browser exceeded `recycleAfter` indefinitely. Fix:
shared `maybeFireDeferredRecycle(entry)` helper called on both
non-release teardown paths.
3. **`openContextOn` rollback didn't drain waiters.** The
orphan-by-recycle rollback freed a reservation but never
`scheduleServeNextWaiter()` → queued waiters could stall with free
capacity until an unrelated release. Fix: `scheduleServeNextWaiter()`
after the rollback.
## Verification
- Red-green for all three (reproduced each bug, then green).
- Full harness vitest: **1702–1705 passed**; browser-pool suite
**27/27** (mutation-tested — reverting fix#3 fails its guard). `tsc
--noEmit` exit 0.
- Public API + `BROWSER_POOL_MAX_CONTEXTS` default untouched.
## Reviewed
7-agent unbiased CR (cr-loop): the three fixes confirmed sound; one
ordering concern on the new code investigated and **refuted**
(sole-browser relaunch-failure rejects the waiter rather than stranding
it).
## Known further hardening (separate effort — NOT in this PR)
The CR surfaced additional **pre-existing** browser-pool reliability
bugs that warrant a dedicated hardening pass, independently flagged by
multiple reviewers:
- `shutdown()` vs in-flight `openContextOn` → leaked context (no
`isShutdown` re-check post-`newContext`); recycles added to
`inFlightRecycles` after shutdown's snapshot not awaited.
- crash-reason `recycleBrowser` abandons live contexts without
`.close()` — leaks on the non-dead `acquire`-retry path.
- `acquire` retry treats a transient `newContext` failure as a full
crash → recycles the whole browser, tearing down unrelated live contexts
(correlated flakiness).
- `launchChain` launch gate has no timeout → a single hung
`chromium.launch()` permanently deadlocks all relaunches.
- `parseInt` env parsing silently accepts trailing garbage
(`MAX_CONTEXTS=24x` → 24); no warning.
- context-close failures swallowed via bare `.catch(() => {})`
(inconsistent with `closeBrowser`'s logged path).
- relaunch-failure eviction only rejects waiters when the pool is fully
empty (doesn't redistribute onto surviving browsers); `pickLeastLoaded`
tie-break concentrates load on browser 0.
Bug 1: serveNextWaiter orphan-close (timed-out waiter mid-open) now mirrors
openContextOn's servedContexts++ with a decrement, so an orphaned serve no
longer permanently inflates servedContexts and biases the hygiene recycle to
fire early.
Bug 2: a hygiene recycle deferred via the release-path shouldRecycle&&hadWaiter
guard is now re-checked on the NON-release teardown paths (openContextOn
orphan-by-recycle rollback and serveNextWaiter orphan-close) via a shared
maybeFireDeferredRecycle helper, so the deferred recycle still fires when the
entry's last activity ends without a release() — previously it was dropped and
the browser exceeded recycleAfter indefinitely.
Bug 3: openContextOn's orphan-by-recycle rollback now calls
scheduleServeNextWaiter() so freed capacity immediately drains queued waiters
instead of stalling them until the next unrelated release/recycle handoff.
Adds three red-green regression tests (BUG1/BUG2/BUG3) to the browser-pool
suite. Public API and MAX_CONTEXTS default unchanged.
## Fixes
- **Gate LGP tool-rendering AAPL + Find-flights fixtures on `toolName`**
(not `hasToolResult`): the AAPL tool-rendering and Find-flights
first-leg fixtures now key off `toolName` so the right tool renders.
Local proof: tool-rendering AAPL + Find-flights run **local D6 green**.
- **Restore sandboxed-UI `jsFunctions` in `gen-ui-open-advanced`
fixtures** (agno, crewai-crews, langgraph-fastapi, langgraph-python,
mastra): the sandboxed Calculator/Ping `jsFunctions` were missing, so
the calculator never computed. Local proof: calc **`=` → 4
browser-verified**.
- **Resolve dashboard links to the real shell host via server-threaded
`shellUrl`**: the dashboard tree was entirely `"use client"`, so
`getRuntimeConfig()` returned the `ssr-placeholder.invalid` SSR sentinel
and baked dead hrefs into every Demo/Code link. `page.tsx` is now a
server component that reads the real host server-side and threads it
into the client `DashboardPage`. Local proof: **SSR links click → real
demo, verified**.
- **Fail `verify-deploy` on env-unset config sentinel + robust config
extractor**: when `SHELL_URL` is unset the server config returns the
`about:blank#shell-url-missing` sentinel; the deploy guard now fails
loud on it rather than shipping dead links, with a hardened config
extractor. Local proof: **deploy-guard red-green**.
- **Gate Coverage D6 badge + stats by the depth ladder; gated indicator
only on genuine lower-rung failure**: D6 is the top of the verification
ladder, so a green D6 claim is only valid when the ladder through D5 is
intact. New `d6Effective` collapses to gated (`—`) when a lower rung
genuinely fails (never on no-data), keeping the badge, stat, regression
flag, and chip in agreement. Local proof: **D6-gating full-suite 790
green incl dashboard-color-matrix 54/54**.
- **Raise browser-pool default `MAX_CONTEXTS` to 40 + correct pool
docs**: contexts (not chromium processes) are the scaling knob since the
PID ceiling of 1000 is the binding constraint; D6 peak 32 + D5 peak 8 =
40. Probe cadence/docs corrected to match. Local proof: **pool
MAX_CONTEXTS=40 locally proven, 50 PIDs ≪ 1000**.
## CR
Converged via 4 unbiased 7-agent cr-loop rounds + 2 fix rounds.
## Known follow-ups (not in this PR)
- **(d) browser-pool concurrency hardening** — `servedContexts`
inflation on `serveNextWaiter` orphan-close, `recyclePending` deferral
on non-release teardown, and waiter-drain on `openContextOn` rollback.
These are pre-existing pool internals; separate PR.
- **(b/c) minor cosmetic / naming items** — `DocsRow` unused `shellUrl`
prop; `computeColumnTallyDetail` labels a D6-absent amber as `"e2e"`;
the agno `gen-ui-open-advanced` `_meta` note is misleading but is the
SOLE source of agno Calculator/Ping fixtures (do NOT delete); `API=d3`
vs `d2` naming; `resolveD3` has no effective stale row (pre-existing);
`e2e-deep.yml` stale primary-key comment.
- **react-core consecutive-interrupt run-state fix** — a SEPARATE
pending branch; the `gen-ui-interrupt` cell needs it.
## Summary
Carries [ag-ui PR
#1784](https://github.com/ag-ui-protocol/ag-ui/pull/1784)
("fix(langgraph): skip regeneration check when `command.resume` is set")
into the three langgraph showcase stacks, greening the
`gen-ui-interrupt` D6 cell.
- ag-ui `@ag-ui/langgraph` 0.0.35 (and earlier) incorrectly ran the
regenerate path on a **resumed** run, tripping the regeneration trap and
breaking gen-ui-interrupt. `0.0.36` adds the `command.resume` guard so a
resume skips the regeneration check.
- Pins `@ag-ui/langgraph` `0.0.36` via npm `overrides` (it is a
**transitive** dep of `@copilotkit/runtime`) in `langgraph-python`,
`langgraph-typescript`, and `langgraph-fastapi`.
- Regenerates each integration's `package-lock.json` (python/fastapi
regenerated with `--legacy-peer-deps`, matching their Dockerfile `npm ci
--legacy-peer-deps`).
## Files changed
-
`showcase/integrations/langgraph-typescript/{package.json,package-lock.json}`
-
`showcase/integrations/langgraph-python/{package.json,package-lock.json}`
-
`showcase/integrations/langgraph-fastapi/{package.json,package-lock.json}`
Lockfiles flip `@ag-ui/langgraph` 0.0.34 → 0.0.36 and pull in 0.0.36's
new transitive dep `@ag-ui/a2ui-toolkit@0.0.1-alpha.3`.
## Verification
- `@ag-ui/langgraph@0.0.36` is published to npm `latest`; its
`dist/index.js` contains the `!command?.resume` regeneration guard from
#1784.
- All three regenerated lockfiles resolve
`node_modules/@ag-ui/langgraph` to `0.0.36`.
## Test plan
- [ ] CI green
- [ ] gen-ui-interrupt D6 cell green for langgraph-python,
langgraph-typescript, langgraph-fastapi after showcase rebuild + re-run
Carries ag-ui PR #1784 (skip regeneration check when command.resume is
set) into the three langgraph showcase stacks. ag-ui 0.0.35 incorrectly
ran the regenerate path on a resumed run, breaking the gen-ui-interrupt
D6 cell. 0.0.36 adds the command.resume guard.
Pins @ag-ui/langgraph 0.0.36 via npm overrides (transitive dep of
@copilotkit/runtime) in langgraph-python, langgraph-typescript, and
langgraph-fastapi, and regenerates each per-integration package-lock.json.
## Problem
Shell-docs had conflicting v2 guidance around the provider import path.
Some migration/reference/quickstart pages either recommended
`CopilotKitProvider` or kept `CopilotKit` examples on the root
`@copilotkit/react-core` package even though v2 docs should import the
`CopilotKit` component from `@copilotkit/react-core/v2`.
## Why
The correct recommendation is the `CopilotKit` component name, imported
from the v2 entrypoint. Leaving root-package imports in v2-facing docs
makes the migration and reference guidance contradict the v2 package
layout.
## Fix
- Recommend `CopilotKit` from `@copilotkit/react-core/v2`, not
`CopilotKitProvider`.
- Update v2 migration, reference, and quickstart examples to use the v2
provider/style entrypoints.
- Leave root `@copilotkit/react-core` imports only in v1 docs and
explicit migration “Before” examples.
- Add regression coverage for stale provider/style package paths.
- Fix the shell-docs SignupLink SSR test typing exposed by typecheck.
Closes#5153
## Summary
The showcase dashboard's Ops tab fetches `/api/ops/*` as a same-origin
path, which the Route Handler at
`shell-dashboard/src/app/api/ops/[...path]/route.ts` forwards at request
time to `${OPS_BASE_URL}/api/*` on the showcase-harness HTTP origin (the
service that serves `/api/probes`).
In the local compose stack, `OPS_BASE_URL` was set to
`http://localhost:3200` — the dashboard's own host. The proxy therefore
looped back into the dashboard instead of reaching the harness, so the
probe-trigger endpoint failed (self-referential 500/503) and the Ops
live-probe grid could not resolve.
This points `OPS_BASE_URL` at the harness origin over the compose
network: `http://showcase-harness:8080`. The harness `Dockerfile`
EXPOSEs `8080` and `orchestrator.ts` binds `PORT ?? 8080`, so the
dashboard reaches `/api/probes` by container name on the internal port.
This mirrors staging, where the dashboard's `OPS_BASE_URL` likewise
points at the harness origin rather than at itself.
Scope: a single build-arg value in `showcase/docker-compose.local.yml`
(plus an updated explanatory comment). `OPS_BASE_URL` is read at request
time by the Route Handler, so this only seeds the runtime default — no
build-time resolution required.
## Test plan
- [ ] Dashboard Ops tab loads without a 500/503 from the ops proxy
- [ ] Probe-trigger endpoint (`/api/ops/probes` POST) returns 2xx,
forwarded to the harness `/api/probes`
- [ ] Ops live-probe grid renders harness data (harness running on the
compose network as `showcase-harness`)
## Summary
Adds a permanent, env-gated injection seam to the showcase harness
service discovery
(`showcase/harness/src/probes/discovery/railway-services.ts`). When
`LOCAL_SERVICES_JSON` is set, the harness (especially the
d6-all-pills-e2e driver) runs against **LOCAL** backend services instead
of performing Railway discovery — enabling apples-to-apples LOCAL D6
verification without Railway credentials.
- **Zero behavior change when unset/empty.** An unset or empty
`LOCAL_SERVICES_JSON` takes the byte-identical Railway discovery path —
the seam is fully transparent in the default configuration.
- **`demos` plumbed end-to-end (load-bearing).** The injected service
records carry `demos` all the way through. This matters: empty `demos`
would short-circuit the D6 driver into a false 15ms zero-cell "green,"
masking real failures. Plumbing `demos` end-to-end is what makes the
LOCAL path a faithful stand-in for Railway discovery.
- **Enables apples-to-apples LOCAL D6 verification** — run the full pill
suite against local services with the same code path shape as staging.
## Notes
- 8 new tests covering the `LOCAL_SERVICES_JSON injection` path (66
tests total in `railway-services.test.ts`, all passing).
- Env-gated: no Railway credentials required when services are injected.
- The injection seam executes the real discovery code (not mocked).
## Test plan
- [ ] `tsc --noEmit` (typecheck) exits 0
- [ ] `railway-services.test.ts` passes (66 tests, incl. all 8
`LOCAL_SERVICES_JSON injection` tests)
- [ ] Injection seam executes against the real code path (verified via
`discovery.railway-services.local-injection` log emission, not a mock)
- [ ] Unset/empty `LOCAL_SERVICES_JSON` produces a byte-identical
Railway discovery path (zero behavior change)
- [ ] `demos` plumbed end-to-end through injected service records
(guards against false zero-cell D6 green)