Quick-win fixes for the **Build with agents** page from Sam's review.
Stacked on #5187 (retarget to `main` once it lands).
## Changes
- **Skills now shown on every page** — previously only the top-level
`/build-with-agents` (and `built-in-agent`) rendered the Skills section;
every framework integration page used the MCP-only snippet, hiding
Skills. The shared `coding-agents.mdx` now renders the full Skills + MCP
guide, so all 11 build-with-agents pages show Skills (the recommended
path).
- **Skills section** — added a top-three table (`copilotkit-setup`,
`copilotkit-develop`, `copilotkit-integrations`) and a note that
`copilotkit-contribute` is for contributing to CopilotKit, not building
with it.
- **Install step** — clarified to run `npx skills add` from the project
root; any agent there (Claude Code, Codex, Cursor, Gemini CLI) discovers
the skills automatically.
- **MCP headings** — demoted per-tool headers (Cursor, Claude Code, …)
from H2 → H3 so they nest under "MCP Docs Server" in the TOC; "Other"
subsections H3 → H4.
## Screenshots
Skills section + top-three table:

Install step:

Skills now rendering on a framework page (Mastra) that was previously
MCP-only:

TOC nesting (tools now under MCP Docs Server):

## Files
-
`showcase/shell-docs/src/content/snippets/shared/guides/build-with-agents.mdx`
-
`showcase/shell-docs/src/content/snippets/shared/guides/mcp-server-setup.mdx`
- `showcase/shell-docs/src/content/snippets/shared/coding-agents.mdx`
-
`showcase/shell-docs/src/content/docs/integrations/langgraph/build-with-agents.mdx`
## What
Two unrelated cleanups:
### 1. Move internal-only skills out of the public repo
Three staff-only skills lived under `.claude/skills/` and
`.agents/skills/`, so `npx skills add CopilotKit/CopilotKit` swept them
into end-user installs. This PR deletes them here; they now live in the
internal-skills plugin (**CopilotKit/internal-skills#108**):
- `copilotkit-demo-parity`
- `git-hooks`
- `showcase-demo-debugging`
### 2. Recommend a cleaner skills install command
The **Build with agents** guide now recommends:
```bash
npx skills add CopilotKit/CopilotKit/skills -y
```
- `/skills` subpath installs only the published skills under `skills/`
(the repo root also picks up internal skills).
- `-y` skips the interactive prompts.
A Callout documents the interactive variant and `-g` for a global
install.
On consecutive interrupts, pressing Enter for the second turn (turn-2) while
the resumed run from turn-1 was still in flight routed the keystroke to the
STOP action, aborting the in-flight resume instead of sending the new message.
The fix gates the Enter handler on `canSend` so a pending/running state no
longer maps Enter to STOP, and `onSubmitInput` now awaits the in-flight run's
completion before dispatching the queued message — the message is sent after
the current run finishes rather than aborting it.
Also hardens the queuing and attachment tests to cover the consecutive-
interrupt path and the send-after-run-completes behavior.
Previously only the top-level /build-with-agents and built-in-agent pages
rendered the Skills section (via <BuildWithAgents />). Every framework
integration page used the MCP-only <CodingAgents /> snippet (or, for
langgraph, <MCPSetup /> directly), so Skills — the recommended path — was
hidden there.
Point the shared coding-agents.mdx snippet at <BuildWithAgents /> so all
pages that reference CodingAgents now render Skills + MCP, and switch the
langgraph page from <MCPSetup /> to <BuildWithAgents />. The snippet inliner
recurses with cycle protection, so no duplication is needed.
## Summary
Two body links on the **Self-Hosting Intelligence** page
(`/<framework>/premium/self-hosting`, rendered from the shared snippet
`docs/snippets/shared/premium/self-hosting.mdx`) were broken in
production:
- **"How the Intelligence Platform Works"** (Prerequisites + Next steps)
linked `/learn/intelligence-platform`. The `NavigationLink` rewriter
(`docs/components/react/subdocs-menu.tsx`) makes absolute links
section-relative, so under each framework it became
`/<framework>/learn/intelligence-platform` → **404**. The page is served
at `/premium/intelligence-platform` in every section (verified live
across built-in-agent, langgraph, mastra, crewai, agno, and root).
Changed both occurrences to `/premium/intelligence-platform`, matching
the sibling premium links (`/premium/overview`).
- **"chart releases" (GHCR)** used the repo-scoped package URL for the
private `CopilotKit/Intelligence` repo → **404** for public readers.
Switched to the public org-scoped package URL
`github.com/orgs/CopilotKit/packages/container/package/charts%2Fintelligence`
(200). The `oci://ghcr.io/...` pull itself was already fine — only the
web link was broken.
The other two body links (`/premium/overview`, `/threads`) and all
in-page anchors were already valid.
## Why this shipped broken
`scripts/check-broken-links.js` only scans `content/docs` and
`components` — it never looks at `snippets/`, so links in shared
snippets are completely unvalidated. Extending the checker to cover
`snippets/` (and to model the section-relative rewriting) would catch
this class of bug. Not included here to keep this PR focused — happy to
follow up.
## Test plan
- [x] `/built-in-agent/premium/intelligence-platform` returns 200
(verified across built-in-agent, langgraph, mastra, crewai, agno, and
root prefixes)
- [x] Org-scoped GHCR package URL returns 200
- [x] Pre-commit hooks pass (`test-and-check-packages`, `commitlint`)
- [ ] Confirm on docs preview deploy that both links resolve from the
rendered page
The shared self-hosting snippet renders under every framework section
(built-in-agent, langgraph, ...), and the NavigationLink rewriter makes
absolute links section-relative.
- "Intelligence Platform Works" linked /learn/intelligence-platform,
which became /<framework>/learn/intelligence-platform (404). The page
is served at /premium/intelligence-platform in every section.
- The GHCR chart-releases link used the repo-scoped URL for the private
Intelligence repo (404 for public readers); switched to the public
org-scoped package URL.
Addresses review feedback on the Build with agents page:
- Add a top-three skills table (copilotkit-setup / -develop / -integrations)
and call out that copilotkit-contribute is for working on CopilotKit
itself, not building with it, so the skills directory's build-vs-contribute
split is clear from the docs page.
- Clarify where to run `npx skills add`: from the project root, where any
coding agent (Claude Code, Codex, Cursor, Gemini CLI) discovers the skills
automatically — answering 'in your agent environment'.
- Demote the MCP per-tool section headers (Cursor, Claude Web, Claude Code,
...) from H2 to H3 so they nest under 'MCP Docs Server' in the on-this-page
TOC instead of sitting as flat siblings; demote the 'Other' subsections to
H4 accordingly.
The static/quality format job runs in check mode on push and was failing
on main: `ruff format --check .` flagged 5 unformatted Python files under
examples/showcases/a2ui-pdf-analyst/agent (main.py, src/dynamic_agent.py,
src/fixed_agent.py, src/multimodal_middleware.py, src/pdf_tools.py). The
oxfmt JS/TS check already passes, so this is ruff-only drift. Applied
`ruff format` (pinned 0.15.13, matching CI); diff is formatting-only.
The langgraph-js agent `dev` script ran `npx @langchain/langgraph-cli@1.2.1`,
which resolves its own isolated dependency tree. That tree pulled a 1.2.x
`@langchain/langgraph` to satisfy the `@langchain/langgraph-api` peer
dependency, but the API hard-imports `STREAM_EVENTS_V3_MODES` from
`@langchain/langgraph/web` — a symbol only present in langgraph 1.3.0+ —
crashing the agent at startup with a SyntaxError.
Install `@langchain/langgraph-cli@1.2.4` as a devDependency and invoke the
local `langgraphjs` binary so the `@langchain/langgraph` peer resolves from
the agent's own pinned 1.3.0, which exports the symbol. Verified the agent
boots cleanly (graph registered, API on :8123, no SyntaxError).
The v1→v2 migration left tool `parameters` in the old v1 array shape
(`[{ name, type, description, required }]`), which does not satisfy the
v2 `FrontendTool.parameters?: StandardSchemaV1` contract. This produced a
`Property '"~standard"' is missing` TypeScript error, failing the
Next.js build for starters whose Docker smoke build typechecks (mastra,
ms-agent-framework-python, adk).
Convert each tool's parameters to a `z.object({...})` schema so typed
`args`/handler arguments resolve correctly. Add `zod` as a dependency to
the two starters (adk, a2a-middleware) that lacked it.
Affected starters: adk, agno, llamaindex, mastra, pydantic-ai,
ms-agent-framework-dotnet, ms-agent-framework-python, a2a-middleware.
The starter smoke-test Slack alert had no indication of where it came
from, which is ambiguous when the same workflow runs across multiple
repos (e.g. the public CopilotKit/CopilotKit repo vs the internal
testybara fork). Prepend a `[ci:<owner/repo>]` tag derived from
github.repository so triage is unambiguous about the source. These CI
alerts test example source and carry no staging/production dimension,
so [ci] is the meaningful source axis. Existing message format is
otherwise preserved.
Prefix every harness-dispatched alert with a `[staging]`/`[production]`/
`[unknown]` source-env tag so operators triaging a red probe know which
deploy environment is affected. The label is derived in the orchestrator
from SHOWCASE_ENV ?? RAILWAY_ENVIRONMENT_NAME ?? "unknown" and applied at
the single renderer chokepoint (covering per-key, cron, and on-error
dispatch) plus the aggregation flush path that bypasses the renderer, via
a shared sourceEnvPrefix helper so the two paths never drift. A missing
env var surfaces as a visible [unknown] rather than a silent un-prefixed
alert.
- STAGING-OUTAGE regressions: degraded alarm fires (not silent) when the
set empties from a relaunch storm; self-heal re-inits a fresh set once the
kernel relaxes; a waiter queued during the dead window is served by
self-heal; a transient relaunch EAGAIN is retried and the entry survives
(no eviction, no alarm).
- Bounded serveNextWaiter transient re-drive + FIX#7 dead-vs-alive gate
propagation to the serve path.
- orchestrator: degraded/recovered signal wiring covered.
- BUG3 (orphan-by-recycle waiter drain) adapted to the crash-recovery
acquire path: an acquire whose in-flight open is orphaned re-enqueues as a
waiter; with cap=1 the freed slot goes to the other waiter, so the orphaned
acquire settles via its own (now bounded) timeout. The invariant it
verifies (freed capacity immediately serves the queued waiter) is unchanged.
Lift the soft nproc limit to the hard ceiling (`ulimit -u $(ulimit -Hu)`)
before exec'ing the orchestrator so the legitimate 40-context chromium
workload (several hundred OS threads at steady state) has ample thread
headroom instead of running near the default ~1024 soft ceiling, where
`chromium.launch()` tripped `pthread_create: Resource temporarily
unavailable`. `exec` keeps node as PID 1 for correct signal handling; the
`|| true` fallback keeps boot resilient when the runtime forbids raising
the soft limit (the cgroup pids limit then remains the dominant control).
Make the long-lived chromium pool survive a pthread/PID-ceiling
(`pthread_create: Resource temporarily unavailable`, errno 11) thread-
exhaustion storm instead of draining to an empty, permanently-wedged set.
- Crash-recovery relaunch backpressure: a transient EAGAIN on relaunch is
retried with bounded linear backoff before the entry is evicted, so a
thread-exhaustion window that relaxes within seconds recovers in place
rather than splicing the entry out of the set.
- Self-heal + degraded/recovered alarm: when the set empties mid-life the
pool fires an `onDegraded` red alarm (previously only emitted on init()
failure — mid-life death was silent) and kicks a background self-heal
loop that relaunches a fresh set the moment a launch succeeds, firing
`onRecovered`. No manual redeploy required.
- Bounded serveNextWaiter transient re-drive: a persistently-transient
newContext() on a still-connected browser no longer hot-loops the event
loop; it self-reschedules up to a ceiling then leaves the waiter queued
for a later release/recovery handoff (mirrors acquire()'s retry-once
semantics).
- Accounting hardening: generation-token guard on in-flight opens across a
recycle, clamped servedContexts rollback on orphan-close, deferred-recycle
re-check on non-release teardown paths, and waiter-drain on orphan-by-
recycle rollback so freed capacity is served immediately.
- orchestrator wires the pool's onDegraded/onRecovered hooks to the shared
`system:browser-pool-degraded` red/green capacity-loss signal.
Fixes the 2026-06-03 staging incident: the browser pool died from thread
exhaustion, the relaunch storm emptied the set, and the harness wedged with
no alarm -> 626 D0-red cells until a manual redeploy.
Removes the three internal-only skills (copilotkit-demo-parity, git-hooks,
showcase-demo-debugging) from .claude/skills/ and .agents/skills/. These are
staff-only and now live in the internal-skills plugin. Removing them at the
source means root skill discovery no longer sweeps internal skills into a
user's install.
## Problem
The top-level `docs/` app is retired, but nothing in the repo said so,
and contributors (and agents) kept editing it. Two parallel docs trees
plus a one-directional legacy sync script made it ambiguous where
documentation should be authored:
- `docs/content/docs/` — the old Fumadocs app, no longer publishing
- `showcase/shell-docs/src/content/docs/` — the live source for
docs.copilotkit.ai
Two instruction surfaces actively pointed the wrong way: `CLAUDE.md`
said nothing about docs at all, and `.claude/docs/hooks.md` told
contributors to "add a docs page under `/docs`" (the retired location).
## Change
Establish one canonical rule and reduce the other surfaces to pointers:
- **`.claude/docs/documentation.md`** (new) — source of truth.
CopilotKit docs are authored in `showcase/shell-docs/src/content/`
(`docs/`, `reference/`, `snippets/`, `framework-overviews/`); the
top-level `docs/` folder is retired; AG-UI protocol docs are authored
upstream in `ag-ui-protocol/ag-ui` (publishing to docs.ag-ui.com) and
mirrored into `content/ag-ui/`.
- **`CLAUDE.md`** — adds an Essentials hard-rule and a Reference link.
- **`docs/README.md`** — replaces boilerplate with a retired/STOP
banner; legacy README retained under a `<details>`.
- **`.claude/docs/hooks.md`** — fixes the stale `/docs` pointer and
clarifies that a hook's API reference page lives in
`reference/hooks/<hookName>.mdx`, where v2 reference navigation is
generated automatically from frontmatter (no `meta.json`); conceptual
guide pages under `docs/` still use `meta.json`.
- **`CONTRIBUTING.md`** — adds a two-domain documentation section for
human contributors.
## Notes
- Two docs domains: **CopilotKit docs** → shell-docs; **AG-UI protocol
docs** → upstream `ag-ui-protocol/ag-ui`, then synced into the in-repo
mirror.
- Instructions-only change; no enforcement hook or sync-process change.
- Markdown only; no package code touched.
Use `CopilotKit/CopilotKit/skills -y` instead of the repo root: root
discovery sweeps in the internal `showcase-demo-debugging` skill
(metadata.internal, lives in .claude/.agents, not skills/), so users got
12 skills incl. one internal. The /skills subpath yields exactly the 11
published skills. Drop -g so install defaults to project scope, letting each
project pin the skills version matching its CopilotKit dependencies.
The build-with-agents guide recommended a bare `npx skills add` that drops
human users into a multi-step interactive flow (skill multiselect, agent
selection, scope, install method, confirm). Recommend `-g -y` so all skills
install globally in one shot, with a Callout pointing to the flag-less command
for users who want to choose interactively.
The top-level docs/ app is retired but nothing said so, and two
instruction surfaces still pointed contributors there. Establish a
single canonical rule and reduce the other surfaces to pointers.
- Add .claude/docs/documentation.md as the source of truth: CopilotKit
docs are authored in showcase/shell-docs/src/content/; the top-level
docs/ folder is retired; AG-UI protocol docs are authored upstream in
ag-ui-protocol/ag-ui and mirrored here.
- CLAUDE.md: add an Essentials rule and a Reference link.
- docs/README.md: replace boilerplate with a retired/STOP banner.
- .claude/docs/hooks.md: fix the stale /docs pointer; document that a
hook's API reference page lives in reference/hooks/ and that v2
reference nav is generated from frontmatter (no meta.json).
- CONTRIBUTING.md: add a two-domain documentation section.
## Summary
- Select generated thread titles only from assistant text messages
returned by the naming run.
- Reject JSON-shaped title output unless it parses to an object with a
string title.
- Add Runtime regression coverage for tool-result suffixes and invalid
JSON title payloads.
## Validation
- pnpm nx run @copilotkit/runtime:test --
src/v2/runtime/__tests__/thread-names.test.ts
src/v2/runtime/__tests__/handle-run.test.ts (passed, 56 tests)
- pnpm nx run @copilotkit/runtime:build (passed)
- Pre-commit package checks passed
- pnpm nx run @copilotkit/runtime:check-types (blocked in dependency
task @copilotkit/shared:check-types: missing
LicenseContextValue/LicenseMode exports from
@copilotkit/license-verifier and telemetry index error)
- NODE_OPTIONS=--max-old-space-size=8192 pnpm nx run
@copilotkit/runtime:check-types --excludeTaskDependencies (failed: tsc
heap out of memory near 8 GB)
## Summary
Follow-up to #5173 (bucket-(d)) closing three **pre-existing**
browser-pool concurrency defects on the non-release teardown paths. The
browser-pool is a context-pool over a fixed set of long-lived Chromium
processes; these defects biased the hygiene-recycle cadence and could
strand waiters / leak deferred recycles.
## Fixes (each red-green proven)
1. **`serveNextWaiter` orphan-close leaked `servedContexts`.** A waiter
timing out mid-`openContextOn` cleaned up the reservation/context but
never decremented `servedContexts` (which `openContextOn` had already
`++`'d) → every orphaned-by-timeout serve permanently inflated the count
→ premature hygiene recycles. Fix: decrement `servedContexts` in the
orphan-close block.
2. **Deferred `recyclePending` honored only on `release()`.** A recycle
deferred because `pendingOpens > 0` set `recyclePending`, but the
non-release teardown paths (orphan-by-recycle rollback, orphan-close)
returned the entry to idle without re-checking it → the deferred recycle
was dropped and the browser exceeded `recycleAfter` indefinitely. Fix:
shared `maybeFireDeferredRecycle(entry)` helper called on both
non-release teardown paths.
3. **`openContextOn` rollback didn't drain waiters.** The
orphan-by-recycle rollback freed a reservation but never
`scheduleServeNextWaiter()` → queued waiters could stall with free
capacity until an unrelated release. Fix: `scheduleServeNextWaiter()`
after the rollback.
## Verification
- Red-green for all three (reproduced each bug, then green).
- Full harness vitest: **1702–1705 passed**; browser-pool suite
**27/27** (mutation-tested — reverting fix#3 fails its guard). `tsc
--noEmit` exit 0.
- Public API + `BROWSER_POOL_MAX_CONTEXTS` default untouched.
## Reviewed
7-agent unbiased CR (cr-loop): the three fixes confirmed sound; one
ordering concern on the new code investigated and **refuted**
(sole-browser relaunch-failure rejects the waiter rather than stranding
it).
## Known further hardening (separate effort — NOT in this PR)
The CR surfaced additional **pre-existing** browser-pool reliability
bugs that warrant a dedicated hardening pass, independently flagged by
multiple reviewers:
- `shutdown()` vs in-flight `openContextOn` → leaked context (no
`isShutdown` re-check post-`newContext`); recycles added to
`inFlightRecycles` after shutdown's snapshot not awaited.
- crash-reason `recycleBrowser` abandons live contexts without
`.close()` — leaks on the non-dead `acquire`-retry path.
- `acquire` retry treats a transient `newContext` failure as a full
crash → recycles the whole browser, tearing down unrelated live contexts
(correlated flakiness).
- `launchChain` launch gate has no timeout → a single hung
`chromium.launch()` permanently deadlocks all relaunches.
- `parseInt` env parsing silently accepts trailing garbage
(`MAX_CONTEXTS=24x` → 24); no warning.
- context-close failures swallowed via bare `.catch(() => {})`
(inconsistent with `closeBrowser`'s logged path).
- relaunch-failure eviction only rejects waiters when the pool is fully
empty (doesn't redistribute onto surviving browsers); `pickLeastLoaded`
tie-break concentrates load on browser 0.
Bug 1: serveNextWaiter orphan-close (timed-out waiter mid-open) now mirrors
openContextOn's servedContexts++ with a decrement, so an orphaned serve no
longer permanently inflates servedContexts and biases the hygiene recycle to
fire early.
Bug 2: a hygiene recycle deferred via the release-path shouldRecycle&&hadWaiter
guard is now re-checked on the NON-release teardown paths (openContextOn
orphan-by-recycle rollback and serveNextWaiter orphan-close) via a shared
maybeFireDeferredRecycle helper, so the deferred recycle still fires when the
entry's last activity ends without a release() — previously it was dropped and
the browser exceeded recycleAfter indefinitely.
Bug 3: openContextOn's orphan-by-recycle rollback now calls
scheduleServeNextWaiter() so freed capacity immediately drains queued waiters
instead of stalling them until the next unrelated release/recycle handoff.
Adds three red-green regression tests (BUG1/BUG2/BUG3) to the browser-pool
suite. Public API and MAX_CONTEXTS default unchanged.
## Fixes
- **Gate LGP tool-rendering AAPL + Find-flights fixtures on `toolName`**
(not `hasToolResult`): the AAPL tool-rendering and Find-flights
first-leg fixtures now key off `toolName` so the right tool renders.
Local proof: tool-rendering AAPL + Find-flights run **local D6 green**.
- **Restore sandboxed-UI `jsFunctions` in `gen-ui-open-advanced`
fixtures** (agno, crewai-crews, langgraph-fastapi, langgraph-python,
mastra): the sandboxed Calculator/Ping `jsFunctions` were missing, so
the calculator never computed. Local proof: calc **`=` → 4
browser-verified**.
- **Resolve dashboard links to the real shell host via server-threaded
`shellUrl`**: the dashboard tree was entirely `"use client"`, so
`getRuntimeConfig()` returned the `ssr-placeholder.invalid` SSR sentinel
and baked dead hrefs into every Demo/Code link. `page.tsx` is now a
server component that reads the real host server-side and threads it
into the client `DashboardPage`. Local proof: **SSR links click → real
demo, verified**.
- **Fail `verify-deploy` on env-unset config sentinel + robust config
extractor**: when `SHELL_URL` is unset the server config returns the
`about:blank#shell-url-missing` sentinel; the deploy guard now fails
loud on it rather than shipping dead links, with a hardened config
extractor. Local proof: **deploy-guard red-green**.
- **Gate Coverage D6 badge + stats by the depth ladder; gated indicator
only on genuine lower-rung failure**: D6 is the top of the verification
ladder, so a green D6 claim is only valid when the ladder through D5 is
intact. New `d6Effective` collapses to gated (`—`) when a lower rung
genuinely fails (never on no-data), keeping the badge, stat, regression
flag, and chip in agreement. Local proof: **D6-gating full-suite 790
green incl dashboard-color-matrix 54/54**.
- **Raise browser-pool default `MAX_CONTEXTS` to 40 + correct pool
docs**: contexts (not chromium processes) are the scaling knob since the
PID ceiling of 1000 is the binding constraint; D6 peak 32 + D5 peak 8 =
40. Probe cadence/docs corrected to match. Local proof: **pool
MAX_CONTEXTS=40 locally proven, 50 PIDs ≪ 1000**.
## CR
Converged via 4 unbiased 7-agent cr-loop rounds + 2 fix rounds.
## Known follow-ups (not in this PR)
- **(d) browser-pool concurrency hardening** — `servedContexts`
inflation on `serveNextWaiter` orphan-close, `recyclePending` deferral
on non-release teardown, and waiter-drain on `openContextOn` rollback.
These are pre-existing pool internals; separate PR.
- **(b/c) minor cosmetic / naming items** — `DocsRow` unused `shellUrl`
prop; `computeColumnTallyDetail` labels a D6-absent amber as `"e2e"`;
the agno `gen-ui-open-advanced` `_meta` note is misleading but is the
SOLE source of agno Calculator/Ping fixtures (do NOT delete); `API=d3`
vs `d2` naming; `resolveD3` has no effective stale row (pre-existing);
`e2e-deep.yml` stale primary-key comment.
- **react-core consecutive-interrupt run-state fix** — a SEPARATE
pending branch; the `gen-ui-interrupt` cell needs it.
## Summary
Carries [ag-ui PR
#1784](https://github.com/ag-ui-protocol/ag-ui/pull/1784)
("fix(langgraph): skip regeneration check when `command.resume` is set")
into the three langgraph showcase stacks, greening the
`gen-ui-interrupt` D6 cell.
- ag-ui `@ag-ui/langgraph` 0.0.35 (and earlier) incorrectly ran the
regenerate path on a **resumed** run, tripping the regeneration trap and
breaking gen-ui-interrupt. `0.0.36` adds the `command.resume` guard so a
resume skips the regeneration check.
- Pins `@ag-ui/langgraph` `0.0.36` via npm `overrides` (it is a
**transitive** dep of `@copilotkit/runtime`) in `langgraph-python`,
`langgraph-typescript`, and `langgraph-fastapi`.
- Regenerates each integration's `package-lock.json` (python/fastapi
regenerated with `--legacy-peer-deps`, matching their Dockerfile `npm ci
--legacy-peer-deps`).
## Files changed
-
`showcase/integrations/langgraph-typescript/{package.json,package-lock.json}`
-
`showcase/integrations/langgraph-python/{package.json,package-lock.json}`
-
`showcase/integrations/langgraph-fastapi/{package.json,package-lock.json}`
Lockfiles flip `@ag-ui/langgraph` 0.0.34 → 0.0.36 and pull in 0.0.36's
new transitive dep `@ag-ui/a2ui-toolkit@0.0.1-alpha.3`.
## Verification
- `@ag-ui/langgraph@0.0.36` is published to npm `latest`; its
`dist/index.js` contains the `!command?.resume` regeneration guard from
#1784.
- All three regenerated lockfiles resolve
`node_modules/@ag-ui/langgraph` to `0.0.36`.
## Test plan
- [ ] CI green
- [ ] gen-ui-interrupt D6 cell green for langgraph-python,
langgraph-typescript, langgraph-fastapi after showcase rebuild + re-run
Carries ag-ui PR #1784 (skip regeneration check when command.resume is
set) into the three langgraph showcase stacks. ag-ui 0.0.35 incorrectly
ran the regenerate path on a resumed run, breaking the gen-ui-interrupt
D6 cell. 0.0.36 adds the command.resume guard.
Pins @ag-ui/langgraph 0.0.36 via npm overrides (transitive dep of
@copilotkit/runtime) in langgraph-python, langgraph-typescript, and
langgraph-fastapi, and regenerates each per-integration package-lock.json.
## Problem
Shell-docs had conflicting v2 guidance around the provider import path.
Some migration/reference/quickstart pages either recommended
`CopilotKitProvider` or kept `CopilotKit` examples on the root
`@copilotkit/react-core` package even though v2 docs should import the
`CopilotKit` component from `@copilotkit/react-core/v2`.
## Why
The correct recommendation is the `CopilotKit` component name, imported
from the v2 entrypoint. Leaving root-package imports in v2-facing docs
makes the migration and reference guidance contradict the v2 package
layout.
## Fix
- Recommend `CopilotKit` from `@copilotkit/react-core/v2`, not
`CopilotKitProvider`.
- Update v2 migration, reference, and quickstart examples to use the v2
provider/style entrypoints.
- Leave root `@copilotkit/react-core` imports only in v1 docs and
explicit migration “Before” examples.
- Add regression coverage for stale provider/style package paths.
- Fix the shell-docs SignupLink SSR test typing exposed by typecheck.
Closes#5153
## Summary
The showcase dashboard's Ops tab fetches `/api/ops/*` as a same-origin
path, which the Route Handler at
`shell-dashboard/src/app/api/ops/[...path]/route.ts` forwards at request
time to `${OPS_BASE_URL}/api/*` on the showcase-harness HTTP origin (the
service that serves `/api/probes`).
In the local compose stack, `OPS_BASE_URL` was set to
`http://localhost:3200` — the dashboard's own host. The proxy therefore
looped back into the dashboard instead of reaching the harness, so the
probe-trigger endpoint failed (self-referential 500/503) and the Ops
live-probe grid could not resolve.
This points `OPS_BASE_URL` at the harness origin over the compose
network: `http://showcase-harness:8080`. The harness `Dockerfile`
EXPOSEs `8080` and `orchestrator.ts` binds `PORT ?? 8080`, so the
dashboard reaches `/api/probes` by container name on the internal port.
This mirrors staging, where the dashboard's `OPS_BASE_URL` likewise
points at the harness origin rather than at itself.
Scope: a single build-arg value in `showcase/docker-compose.local.yml`
(plus an updated explanatory comment). `OPS_BASE_URL` is read at request
time by the Route Handler, so this only seeds the runtime default — no
build-time resolution required.
## Test plan
- [ ] Dashboard Ops tab loads without a 500/503 from the ops proxy
- [ ] Probe-trigger endpoint (`/api/ops/probes` POST) returns 2xx,
forwarded to the harness `/api/probes`
- [ ] Ops live-probe grid renders harness data (harness running on the
compose network as `showcase-harness`)
## Summary
Adds a permanent, env-gated injection seam to the showcase harness
service discovery
(`showcase/harness/src/probes/discovery/railway-services.ts`). When
`LOCAL_SERVICES_JSON` is set, the harness (especially the
d6-all-pills-e2e driver) runs against **LOCAL** backend services instead
of performing Railway discovery — enabling apples-to-apples LOCAL D6
verification without Railway credentials.
- **Zero behavior change when unset/empty.** An unset or empty
`LOCAL_SERVICES_JSON` takes the byte-identical Railway discovery path —
the seam is fully transparent in the default configuration.
- **`demos` plumbed end-to-end (load-bearing).** The injected service
records carry `demos` all the way through. This matters: empty `demos`
would short-circuit the D6 driver into a false 15ms zero-cell "green,"
masking real failures. Plumbing `demos` end-to-end is what makes the
LOCAL path a faithful stand-in for Railway discovery.
- **Enables apples-to-apples LOCAL D6 verification** — run the full pill
suite against local services with the same code path shape as staging.
## Notes
- 8 new tests covering the `LOCAL_SERVICES_JSON injection` path (66
tests total in `railway-services.test.ts`, all passing).
- Env-gated: no Railway credentials required when services are injected.
- The injection seam executes the real discovery code (not mocked).
## Test plan
- [ ] `tsc --noEmit` (typecheck) exits 0
- [ ] `railway-services.test.ts` passes (66 tests, incl. all 8
`LOCAL_SERVICES_JSON injection` tests)
- [ ] Injection seam executes against the real code path (verified via
`discovery.railway-services.local-injection` log emission, not a mock)
- [ ] Unset/empty `LOCAL_SERVICES_JSON` produces a byte-identical
Railway discovery path (zero behavior change)
- [ ] `demos` plumbed end-to-end through injected service records
(guards against false zero-cell D6 green)
## Summary
- **Fix per-cell D6 resolution**: the dashboard's `resolveD6` now reads
per-cell ENUM keys (`d6:<slug>/<featureType>` via `CATALOG_TO_D5_KEY`
fan-out) instead of the integration aggregate — D6 cells render real
per-cell green/red instead of all-gray.
- **Depth + D6 surfaced by default**: `DEFAULT_OVERLAYS = [links,
health, depth]` (the per-cell D6 badge rides on the Health layer; the
depth chip folds D6).
- **Correct the aggregate stats bar** (`page-stats.ts` extraction): D6
cells were dropped from the depth distribution (`dist["d6"]++` → `NaN`,
masked by an `as` cast); gray/no-data cells were counted as green;
`d6Stats` swallowed amber into gray. Now renders the reachable buckets
**D0/D3/D4/D5/D6** (dropped permanently-zero D1/D2), tracks `noData` and
`degraded` distinctly, and validates `parity_tier` fail-loud.
- **Test coverage**: new table-driven cell color/rollup matrix
(`dashboard-color-matrix.test.tsx`) + `page-stats.test.ts` +
order-independence/comment hardening.
- **CR fixes**: effective (stale-downgraded) `.row` from cell-model
resolvers, `WORST_STATE_RANK` rename, composed-cell memo keys,
overlay-types re-export dedup, `setTab` persistence + `window.location`
test stub, `page.tsx` shared-`now` threading + 60s staleness re-render.
## Verification
- **776 tests pass / 1 skip**, `tsc --noEmit` clean, Next.js production
build OK.
- **Local visual proof**: seeded enum D6 rows (30✓/8✗/0 gray) into a
local PocketBase, rebuilt this branch, screenshotted bare `#matrix` —
per-cell D6 green/red renders, and the depth distribution shows `D6:30
D5:8 D4:0 D3:0 D0:588`.
- **7-agent code review converged** over 3 confirmation cycles.
## Notes
- The **#5152 `d6-all-pills` driver must stay on enum keys** (matches
this dashboard's `resolveD6` + the `d5-mapping-drift` test). Do not
revert it to raw catalog keys.
- **Follow-up (separate PR)**: stats-bar single-source-of-truth
hardening (`computeD6Stats`→`buildCellModel`; reconcile
`resolveD6Row`/`resolveD5Row` to the effective row; thread `connection`
into stats for SSE-offline; docs-only stats exclusion); delete
deprecated `composed-cell.tsx` + `deriveDepth`; `useOverlays` mount-hash
deep-link clobber; Notion visualization-doc sync to per-cell
`resolveD6`.
## Test plan
- [ ] CI green
- [ ] After staging deploy, dashboard renders per-cell D6 green/red (not
all-gray) for LangGraph-Python
The depth-distribution row showed permanent-zero D2/D1 rows while the computed
D0 bucket (wired-but-unverified cells) was never rendered, so wired cells
vanished from the row and it never summed to the "Wired" count.
buildCellModel().achievedDepth is typed 0|3|4|5|6 and can never be 1 or 2.
- Remove unreachable d1/d2 from the DepthDistribution type, from
computeDepthDistribution's init, and from the rendered levels array.
- Add a D0 row to the rendered levels so wired-unverified cells are visible
and the distribution sums to the wired-cell count (D6,D5,D4,D3,D0).
- Reconcile section wrapper keys: use each section's stable key instead of the
array index so overlay-toggle reconciliation is correct.
- Document that health/depth/d6 rollups need no dedup: catalogData.cells is
one row per (integration, feature) grid cell (verified: 0 duplicate pairs),
so these per-cell signals are counted exactly once.
Extract the AdaptiveStatsBar aggregate computations out of page.tsx into a
unit-testable pure module (src/lib/page-stats.ts), mirroring the
computeColumnTally pattern, and fix a cluster of correctness bugs:
1. D6 cells were dropped from the depth distribution: DepthDistribution
lacked a `d6` key and the `\`d${depth}\` as keyof` cast produced
`dist["d6"]++ === NaN`. Add `d6` to the type, render a D6 row in
DepthDistributionSection, and replace the cast with an exhaustive
Record<0|3|4|5|6, keyof DepthDistribution> map the compiler checks.
2. d6Stats folded amber (stale/degraded D6) into gray. Count degraded
distinctly and surface it in D6Section.
3. healthStats counted gray (no-data) cells as green, contradicting the
"stats bar matches the matrix" invariant. Track no-data separately and
render it in HealthSection.
4. isSupported is correctly hardcoded true for wired-cell stats: a wired
catalog cell can never be in not_supported_features (generate-registry
resolves those to status "unsupported" before "wired"), so stats and
renderCell cannot diverge. Comment updated to state the invariant.
5. parity_tier was indexed via an unchecked cast (unknown tier →
`undefined++ === NaN`). Validate against the known tier set and skip +
log loud on unknown.
The matrix render path (renderCell/buildCellModel) is untouched.
Convert the promote workflow's `service` input from a freeform string
(default "all", the accidental-fleet-promote footgun) to a generated
`type: choice` dropdown whose first/default option is a rejected sentinel
so a blind "Run workflow" aborts instead of promoting.
resolve-targets: reject the sentinel, fail loud on ambiguous matches
(no silent head -n1), independently re-filter probe.prod, and reject
--digest combined with `all`. promote: run unattended in CI
(--yes --non-interactive) — the manual dispatch + service selection is
the human authorization. notify: empty-webhook guard + fallback warning,
neutral state for sentinel-abort and any cancellation, and surface
resolve-targets.result in the failure alert for triage.
Add showcase/scripts/sync-promote-service-options.ts: generates the
promote workflow's service `choice` options from the SSOT
(railway-envs.ts), spliced between BEGIN/END markers in
showcase_promote.yml. Fail-loud throughout — every emitted token must
resolve to exactly one service under the resolve-step predicate
(name|dispatchName match AND probe.prod), tokens are YAML-safe, args are
strict (a typo'd flag cannot trigger a destructive write), and markers
are validated before any rewrite.
Wire it into a lefthook pre-commit hook (regenerate + restage; set -e so
a failed regen blocks the commit) and an advisory (never-failing) drift
check in showcase_validate.yml. Vitest coverage for ordering, exclusion,
collision/ambiguity guards, marker errors, exit codes, idempotency, and
the import-side-effect guard.