There is no app/integrations/ route, so the reservation was causing
/integrations/... URLs to 404. Without it, these paths fall through to
UnscopedDocsPage via the non-integration fallthrough in [framework].
Strip snippet_framework: langgraph-python from frontmatter and
integration="langgraph-python" from all InlineDemo tags in the
main docs tree (27 files). On framework-scoped pages frameworkOverride
drives Snippet and InlineDemo; on unscoped pages FrameworkGuardedContent
already hides the body until a framework is selected, so no default
is needed.
buildNavTree now checks meta.root before including a subdirectory,
so unselected/ (and any other root:true directory) is excluded from
the main nav tree. This eliminates the flicker where SidebarLink
generated /langgraph-python/unselected/coding-agents and the
[framework] route had to server-redirect to the correct scoped URL.
On framework-scoped pages, MDX links like /quickstart rendered as plain
<a href="/quickstart"> causing RouterPivot redirect flicker. Override the
MDX `a` component when frameworkOverride is set to prepend the framework
prefix to root-relative internal links.
Next.js routes /quickstart to [framework]/[[...slug]] (dynamic segment
beats optional catch-all), where "quickstart" is not a registered
integration, causing notFound(). Fix by:
- Extracting the unscoped doc rendering logic into UnscopedDocsPage
- Falling through to it in [framework] when the slug is not a framework
- Simplifying [[...slug]]/page.tsx to handle only the root overview
SidebarLink now uses storedFramework as fallback so links on the
overview page go directly to /<framework>/<slug> without a RouterPivot
redirect. OverviewNavItem and section cards now use SidebarLink for the
same reason. FrameworkSelector "Clear selection" was navigating to
/docs/<slug> (broken); fixed to /<slug>. hrefFor now preserves the
current slug when switching frameworks from an unscoped page.
SidebarLink fallback, framework-scoped backLink, "framework-agnostic
version" banner link, FrameworkLandingPage sidebar, and brand-nav
all still pointed at /docs/* which routes to 404 via the [framework]
catch-all. Switch all to / or /<slug>.
Override InlineDemo in DocsPageView's MDX component map to substitute
defaultFramework for the hardcoded integration prop when a framework
is selected. MDX files don't need to change — the override happens at
the render layer, matching how Snippet already handles this.
DOCS_SECTIONS, OverviewNavItem, CopilotKit Docs link, backLink, and
slugHrefPrefix all generated /docs/<slug> paths that 404 — the
[framework] catch-all intercepted "docs" as a framework slug and
returned notFound(). Strip the /docs prefix so all root-route links
resolve correctly.
Page-by-page audit of showcase/shell-docs MDX content. All changes are
in src/content/docs/.
## Broken snippet region fixed
- generative-ui/a2ui/dynamic-schema.mdx: Step 5 referenced
`runtime-inject-tool` which does not exist in the declarative-gen-ui
cell. Replaced with a hardcoded code block showing `injectA2UITool: true`
(same content already shown on the parent a2ui.mdx page).
## MDX syntax fix
- frontend-actions.mdx: stray closing ``` at end of file (would cause
a parse/render failure).
## Internal link fixes — /docs/ prefix removed (×14 files)
The docs app routing does not use a /docs/ prefix — pages live at
/<slug>. Links using /docs/<slug> hit the reserved-slug guard in the
[framework] route and 404. Fixed across:
generative-ui/index.mdx (7 links)
generative-ui/tool-based.mdx (3 links)
learn/index.mdx (6 links)
multi-agent/subagents.mdx (2 links)
shared-state.mdx (2 links)
shared-state/streaming.mdx (2 links)
shared-state/agent-readonly.mdx (4 links)
troubleshooting/debug-mode.mdx
troubleshooting/error-debugging.mdx
troubleshooting/migrate-to-1.10.X.mdx
troubleshooting/migrate-to-1.8.2.mdx (2 links)
troubleshooting/observability-connectors.mdx
## Internal link fixes — /unselected/ removed (×2 files)
backend/copilot-runtime.mdx
backend/custom-agent.mdx
Add per-slug Dockerfile emitters that pair with the new
AGENT_BUILD_STEPS / AGENT_BUILD_COPY tokens in Dockerfile.typescript:
- getAgentBuildSteps(fw): runs in the builder stage, after `npm run build`.
Emits `npx tsc` for claude-sdk-typescript (compiles agent/index.ts →
/app/dist/agent/index.js with flags that match the sibling package
Dockerfile), and `npx mastra build --dir src/mastra` for mastra
(bundles the server into .mastra/output/index.mjs). Returns "" for
every other slug so their Dockerfile cache stays unchanged.
- getAgentBuildCopy(fw): runs in the runner stage, after the agent-code
COPY. Moves /app/dist (claude-sdk-ts) or /app/.mastra (mastra) from
the frontend stage into the runner.
- getEntrypointBlock() prod-mode updates: mastra now boots via
`node /app/.mastra/output/index.mjs` (not `npx mastra dev`) and
claude-sdk-typescript via `node /app/dist/agent/index.js` (not
`npx tsx agent/index.ts`). Cold start is a straight `node`
invocation on Railway — mirrors PR #4132's fix for langgraph-ts.
- Wire AGENT_BUILD_STEPS / AGENT_BUILD_COPY into the `vars` map in
generateStarterImpl so the template substitution picks up the new
tokens, and export the two helpers so the test suite can guard them.
Also refresh the generator header comment to describe the multi-stage
shape (builder toolchain vs. minimal runtime) and the prod-mode emitter
wiring.
Tests:
- Replace the legacy `mastra dev` / `npx tsx` entrypoint expectations
with prod-mode assertions (`node /app/.mastra/output/index.mjs`,
`node /app/dist/agent/index.js`), plus not-to-contain guards so a
future refactor can't accidentally re-enable the tsx/dev path.
- Add a dedicated describe block for getAgentBuildSteps /
getAgentBuildCopy covering the two opted-in slots, the ""
fallthrough for langgraph-typescript (which has its own server.mjs
migration path), and the "" fallthrough for every Python slug
(Python prod-mode is shared-template, not per-slug).
Full `vitest run` in showcase/scripts is green (1085 tests).
Starter Dockerfile regeneration is deferred to a follow-up commit
block once Task 1's template wiring lands.
Restructure the shared starter Dockerfile templates so the runtime stage
ships only a minimal base plus built artifacts — no dev toolchains (pip,
pnpm, tsx, *-cli dev) run in the runner stage.
Dockerfile.typescript:
- Add AGENT_BUILD_STEPS / AGENT_BUILD_COPY template tokens so per-slug
prod-mode compiles (claude-sdk-typescript tsc, mastra `mastra build`)
can be emitted by the generator. Mirrors PR #4132's langgraph-typescript
fix for the rest of the TS starter slots — cold start becomes a straight
`node` call instead of `npx tsx` / `npx mastra dev` on Railway.
Dockerfile.python:
- Split pip install into a dedicated agent-builder stage that carries
full `python:3.12` (with compilers + dev headers needed for wheels
that don't ship pre-built for -slim), and produces a self-contained
/opt/venv. The runner then COPYs /opt/venv as a single opaque tree —
no pip, no build tools in the final image.
Dockerfile.java and Dockerfile.dotnet already implement the true
multi-stage shape (temurin-21 JDK / dotnet-9 SDK builder + JRE / aspnet
runtime); no structural changes needed.
Regenerating the 17 starter Dockerfiles happens in the next commit
block — this commit is template-authoring only.
PR #4132 migrated away from `langgraph-cli dev` to the @langchain/langgraph-api
`startServer` Hono server. That server does bind `/ok` via its meta routes, but
only AFTER `registerFromEnv` imports the graph and spawns the N workers — on
Railway that window has been exceeding the watchdog's 180s startup grace plus
the 90s strike budget, producing an infinite restart loop (deploy ddc88ef9 is
currently flapping ~every 4m30s).
Move liveness off the heavy Hono path entirely:
- server.mjs now spins up a bare node:http server on port 8124 (configurable
via HEALTH_PORT) *synchronously before any await*, answering GET /ok with
{"status":"ok"}. Always live within ms of `node` boot.
- entrypoint.sh watchdog now probes 127.0.0.1:8124/ok instead of :8123/ok.
The liveness port is independent of whether :8123 has finished graph import,
so the probe reflects what the watchdog actually cares about: is the Node
process alive and its event loop responsive. A true hang (event loop pinned)
still fails the probe and triggers container restart, so the existing
hang-detection behavior is preserved.
Local Docker smoke (amd64): curl /ok returned 200 within 20s of container
start; LangGraph API logged "listening on 0.0.0.0:8123" ~15s later as normal.
## Summary
- The shell app's Docker build ran `generate-search-index.ts` against
`/app/shell-docs/src/content/{reference,ag-ui,docs}`, but the shell
Dockerfile never copied that directory into the build context. The
script's fail-loud guard (refuses to emit a near-empty index when all
scan roots are missing) then aborted the build. Fix copies
`showcase/shell-docs/src/content/` into the builder stage so the scan
finds its input.
- Removed `COPY --from=builder /app/shell/src/content ./src/content`
from the runner stage — that path was left over from the pre-split
layout and never existed under the current tree, so the runner stage
failed before this could ever succeed.
Main was red on `Showcase: Build & Deploy` for both the `shell` and
`shell-dashboard` matrix entries after #4127 merged. `shell-dashboard`
was unblocked separately (GHCR package linkage). `shell` needed this
source fix.
## Test plan
- [x] `depot build --platform linux/amd64 -f showcase/shell/Dockerfile
.` — succeeds end-to-end (prior failure was inside the builder RUN;
runner stage COPY also failed before the fix).
- [ ] After merge: `Showcase: Build & Deploy` `build (shell, ...)` job
green.
- [ ] After merge: Railway redeploys `showcase-shell` from the new
`:latest`; `https://showcase.copilotkit.ai/` returns 200.
The shell build ran generate-search-index.ts against shell-docs/src/content
but the Dockerfile never copied that directory into the build context,
so the script failed loudly and the whole shell build aborted. Also drop
the runner-stage COPY of shell/src/content — that path was left over from
the pre-split layout and never existed under the current tree.
The all-dirs-missing guard added in c91c7c567 fataled the shell Docker
build, which intentionally does not COPY shell-docs/src/content/. The
shell only needs the static-pages stub so its header search modal has
something to render (links resolve across to docs.showcase.copilotkit.ai).
Downgrade the fatal to a loud warn + emit the 5-entry static-pages stub.
A misconfigured full build is still visible in logs.
Regression from PR #4127 (8e6991cea) on top of c91c7c567 / debfa6600.
## Summary
- **shell-dashboard:** dashboard.showcase.copilotkit.ai rendered every
"demo" and "code" link as `http://localhost:3000/...`. Root cause:
`NEXT_PUBLIC_SHELL_URL` was never provided at build time and the source
fell back to `localhost:3000`. Next.js inlines `NEXT_PUBLIC_*` at `next
build`, so a runtime Railway env var could not rescue a bad build. Fix
plumbs the value through as a Docker build arg from
`showcase_deploy.yml` and fails loudly at build if it's unset so this
can't regress silently.
- **shell-dojo:** dojo.showcase.copilotkit.ai was missing items in the
langgraph column (langgraph-python showed 9 demos vs 20+ in the
manifests). Root cause: `shell-dojo/src/data/registry.json` was stale —
the generator only wrote to `shell/`, the dojo Dockerfile never ran the
generator at build, and the CI path filter didn't rebuild the dojo when
manifests changed. Fix dual-emits from `generate-registry.ts` to
`shell/`, `shell-dojo/`, and `shell-docs/`, runs the generator in the
dojo Dockerfile, expands the workflow's path filter to include
`packages/**` and `shared/**`, and refreshes the committed JSON so it
matches what the generator produces today. Langgraph-python demo count 9
→ 32.
## Test plan
- [x] `shell-dashboard` Docker build succeeds with
`NEXT_PUBLIC_SHELL_URL` build arg (Depot `36wlvzkgp1`).
- [x] `shell-dashboard` Docker build fails loudly when the build arg is
omitted (Depot `mbh41c3qtk`).
- [x] `shell-dojo` Docker build succeeds; generator+bundler run at
build; 159 demos bundled (Depot `h7bbq8f8jt`).
- [x] `showcase_deploy.yml` passes YAML validation.
- [ ] After merge + deploy: verify `dashboard.showcase.copilotkit.ai`
links point to `https://showcase.copilotkit.ai/...`.
- [ ] After merge + deploy: verify `dojo.showcase.copilotkit.ai` shows
the full langgraph column (langgraph-python ≥ 20 items).
## Summary
Root cause of the 4+ day `langgraph-typescript` flap: `npx
@langchain/langgraph-cli dev` is a dev-mode CLI that compiles TypeScript
on-demand per HTTP request, blocking the single-thread event loop ~5s
per request. On Railway, Studio IPC + `langgraph-api` JIT spawn + graph
compile routinely exceeded the 90s watchdog budget; PR #4123's 180s
grace papered over cold-start but didn't fix steady-state `/schemas`
event-loop blocking.
This PR replaces the CLI with a direct import of `startServer` from
`@langchain/langgraph-api/server` (the same Hono server `dev` wraps,
minus the tsx watcher, Studio handshake, and LangSmith tenant lookup).
Schema extraction runs in a `worker_thread` and is pre-warmed in
parallel with server startup.
Container-tested on `node:22-slim`:
- `/ok` live in <3s (vs. 5-30s dev)
- `/ok` steady-state <10ms
- `/assistants/search` returns the registered `starterAgent` graph
- Pre-warm log: `[server] Pre-warmed schemas in 3402ms`
## Caveats
- Uses `--legacy-peer-deps` for `@langchain/langgraph-sdk` version skew.
- `@langchain/langgraph-api` exports `./server` but it's not officially
public API — exact version pin (`1.1.17`) required.
- `/runs` streaming not end-to-end tested (dummy key);
`/assistants/search` and schema extraction verified.
## Test plan
- [ ] `docker build` of package succeeds (verified locally on
`node:22-slim`, arm64)
- [ ] Railway deploy passes healthcheck within watchdog budget
- [ ] Live probe: `curl
https://showcase-langgraph-typescript-production.up.railway.app/api/health`
returns 200 + `agent: ok`
- [ ] Smoke test run through Studio / demo client end-to-end (real key)
## Related
- Prior attempts: #4116 (watchdog), #4123 (180s grace) — neither
addressed the structural issue
- Proposal context:
https://www.notion.so/3493aa381852814d968df05fc461206d §2c
Root cause of the 4+ day langgraph-typescript Railway flap: `npx @langchain/langgraph-cli dev` is a dev-mode CLI that compiles TypeScript on-demand per HTTP request, blocking the single-thread event loop ~5s per /schemas call. On Railway, Studio IPC handshake + langgraph-api JIT spawn + graph compile routinely exceeded the 90s watchdog budget; PR #4123's 180s grace papered over cold-start but didn't fix steady-state event-loop blocking.
Replace the CLI with a direct import of startServer from @langchain/langgraph-api/server (the same Hono server dev wraps, minus the tsx watcher, Studio handshake, and LangSmith tenant lookup). Schema extraction runs in a worker_thread and is pre-warmed in parallel with server startup.
Container-tested on node:22-slim:
- /ok live in <3s (vs. 5-30s dev)
- /ok steady-state <10ms
- /assistants/search returns the registered graph
- Pre-warm log: [server] Pre-warmed schemas in ~1.3-3.4s
Caveats:
- Uses --legacy-peer-deps for @langchain/langgraph-sdk version skew.
- @langchain/langgraph-api exports ./server but it's not officially public API — exact version pin (1.1.17) required.
- /runs streaming not end-to-end tested (dummy key); /assistants/search and schema extraction verified.
Drops the inline TypeScript typecast from the Mastra agent-app-context
example and uses optional chaining + direct access instead, so the doc
snippet is easier to read and copy. Keeps optional chaining on `.find`
so the example stays safe when the AG-UI context is absent. Also fixes
the `[!code highlight:N]` count after the comment line was removed.
Ports @Abubakar-01's changes from #4125 so they can ship together with
the `requestContext` rename, targeting the new `showcase/shell-docs/`
path after the shell restructure on main.
Co-authored-by: Muhammad Abubakar <abubakaran102025@gmail.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The dojo app was missing items under the langgraph column because
shell-dojo shipped a stale committed registry.json. The generator
only wrote to shell/, the dojo Dockerfile didn't run the generator
at build, and the CI path filter didn't rebuild the dojo when
manifest files changed.
Fix: emit from generate-registry.ts to shell, shell-dojo, and
shell-docs; add the generator step to shell-dojo's Dockerfile;
expand the deploy workflow's path filter to include packages/**
and shared/**; and refresh the committed registry/demo-content
JSON so files on disk match what the generator produces today.
The shell-dashboard app baked http://localhost:3000 into every demo and code link because NEXT_PUBLIC_SHELL_URL was never provided at build time and the source defaulted to localhost. Next.js inlines NEXT_PUBLIC_* at next build, so setting the value on Railway at runtime does nothing.
Fix: remove the silent localhost fallback, pass NEXT_PUBLIC_SHELL_URL as a Docker build arg from showcase_deploy.yml, and fail loudly if it's unset at build so this can't regress silently.
The upstream docs sync added 24 new integration pages (`threads.mdx`,
`premium/self-hosting.mdx`) plus learn/reference content, but didn't
land the runtime wiring they depend on — so the pages would ship blank.
It also accepted upstream changes in tool-rendering.mdx over local
modifications that needed to be restored.
- Register `Threads` / `SelfHosting` in SNIPPET_MAP so the shared
`<Threads />` and `<SelfHosting />` tags inline on integration pages
- Register `ThreadsEarlyAccess` wrapper + `MessageSquareMore` /
`Network` / `Newspaper` icons in mdx-registry for the new threads
content and the rebuilt `learn/index` card grid
- Revert tool-rendering.mdx upstream-wins regressions: framework-neutral
`/unselected/server-tools` link and `components={props.components}`
forwarding for the `<SharedContent>` override pathway
- Add `threads` + `...premium` to every integration's root meta.json
plus 11 new `premium/meta.json` files so the new pages appear in the
sidebar; update `docs/meta.json`, `docs/learn/meta.json`, and
`docs/premium/meta.json` the same way
- Fix `convertMarkdownTableToHtml` to emit JSX-valid tags. It was
generating `style="..."` string attributes, which next-mdx-remote
rejects; this path was dormant until the new threads page put a
markdown table inside a `<Step>`
Regenerated via `cd showcase/scripts && npx tsx generate-starters.ts`
after the getWatchdogGraceSeconds() mapping update in the previous
commit. Only these two starters change — every other framework still
returns 0s from the grace mapping, so their entrypoints are byte-for-
byte identical to main.
Two more starters are restart-looping with the same pattern #4123 fixed
for langgraph-*: the 90s (3-strike) watchdog budget is shorter than the
cold-start path on a fresh Railway container, so the agent gets killed
before it ever reports healthy, and the loop never breaks.
Railway evidence:
- claude-sdk-typescript (package, :8000): restart-looping since 04-20
16:54 UTC. Runs compiled `node /app/agent_server.js` which spins up
the full @anthropic-ai/claude-agent-sdk. Package entrypoint is hand-
written (not generator-emitted), so add the same grace block inline
matching #4123's shape.
- mastra (starter, :8123/api): restart-looping since 04-20 18:18 UTC.
Runs `mastra dev` on :8123 alongside Next.js on ${PORT:-10000}.
`mastra dev` performs a tsx build + Mastra server boot on first
request — legitimate supervised process, grace is the right fix.
Package-level mastra has no watchdog (the entire package IS the Next.js
app; no separate agent to watchdog) — PR #4116's classification of
"N/A" was correct for the package. The starter has a separate Mastra
dev server, so the generator-emitted watchdog is legitimate there.
Per-framework grace (not universal) preserves the #4123 design: uvicorn
and express agents are responsive within the 2-3s sleep before
AGENT_HEALTH_CHECK, so adding grace elsewhere would only delay
legitimate restart on true hangs.
This commit covers the generator mapping + the hand-written package
entrypoint. The regenerated starter entrypoints land in the next commit.
## Incident
`langgraph-typescript` is in a restart loop on Railway as of **04-20
17:05 UTC** (deployment `58bbebe8-7a94-4f99-b6e4-ffcbb4eb78b9`),
returning 502 to production traffic.
## Root cause
PR #4116 generalized the silent-hang watchdog from `crewai-crews` to
every showcase starter: poll agent health every 30s, kill after 3
consecutive failures (~90s). That's fine for uvicorn/express agents
(responsive in <5s), but `langgraph-cli dev` does a heavy cold-start:
- Studio browser IPC handshake (`createIpcServer`)
- `@langchain/langgraph-api` JIT spawn (`spawnServer`)
- Graph compile
On cold Railway containers this routinely exceeds 90s — the watchdog was
killing the process before `/ok` ever became reachable, producing a
kill-loop every ~90s.
## Phase 1 findings
Verified the health path is correct — not a path bug:
- `@langchain/langgraph-cli@1.1.17/dist/cli/up.mjs:27` uses `/ok` as its
own probe.
- `@langchain/langgraph-api@1.1.17/dist/api/meta.mjs:64` registers
`api.get("/ok", ...)`.
- `langgraph_cli` (Python) serves `/ok` from the same api surface.
The regression is purely a **timing** issue: 90s strike budget is
insufficient for langgraph cold-start on Railway.
## Fix — Option B (startup grace)
Added per-framework `getWatchdogGraceSeconds()` in the generator:
- `langgraph-*` starters: **180s grace**
- All other starters: **0s grace** (matches pre-#4116 behavior)
The grace loop waits up to 180s for the first healthy `/ok` probe before
arming the strike counter:
- First success → fall through immediately and arm the counter.
- 180s elapsed without success → arm the counter anyway (steady-state
watchdog then handles true hangs on the normal 90s schedule).
- Agent dies during grace → exit grace loop, `wait -n` in main shell
handles it.
**Not a revert of #4116.** The silent-hang vulnerability class remains
covered (still kills after 90s of steady-state failures); the grace only
defers the first strike.
## Per-starter changes
| Starter | Grace | Behavioral change |
|---|---|---|
| langgraph-typescript | 180s | **Fixes restart loop** |
| langgraph-python | 180s | Preemptive |
| langgraph-fastapi | 180s | Preemptive |
| crewai-crews, ag2, agno, claude-sdk-*, google-adk, llamaindex,
langroid, mastra, ms-agent-*, pydantic-ai, spring-ai, strands | 0s |
None (comment-only diff) |
## Files touched
- `showcase/scripts/generate-starters.ts` — new
`getWatchdogGraceSeconds()` + grace block in `getWatchdogBlock()`
- `showcase/starters/*/entrypoint.sh` — regenerated (17 files; grace
block for langgraph-*, comment-only for the other 14)
- `showcase/packages/langgraph-typescript/entrypoint.sh` — hand-edited
grace block (this package uses its own hand-edited entrypoint, not the
template)
- `showcase/packages/langgraph-fastapi/entrypoint.sh` — same
- `showcase/packages/langgraph-python/entrypoint.sh` — same
## Validation
- `bash -n` passes on all 19 edited `entrypoint.sh` files
- 1079/1079 `showcase-scripts` vitest pass (`pnpm --filter
@copilotkit/showcase-scripts test`)
## References
- Regression source: #4116
- Railway deployment: `58bbebe8-7a94-4f99-b6e4-ffcbb4eb78b9`
- Incident window: 04-20 17:05 UTC — ongoing at PR creation
## Test plan
- [ ] CI green
- [ ] Deploy to Railway — watch first boot log for `[watchdog] Startup
grace: waiting up to 180s...` followed by `[watchdog] Agent healthy
after Ns — arming strike counter`
- [ ] Confirm no restart loop on `langgraph-typescript`
PR #4116 introduced a generalized silent-hang watchdog that polls the
agent health endpoint every 30s and kills the agent after 3 consecutive
failures (~90s). This worked for fast-starting agents (uvicorn,
express) but regressed langgraph-typescript on Railway: `langgraph-cli
dev` does a heavy cold-start (Studio IPC setup + @langchain/langgraph-api
JIT spawn + graph compile) that routinely exceeds 90s on a fresh
container, so the watchdog was killing the process before /ok ever
became reachable — producing the 04-20 17:05 UTC restart loop on
deployment 58bbebe8-7a94-4f99-b6e4-ffcbb4eb78b9.
Phase 1 verification:
- `@langchain/langgraph-cli@1.1.17/dist/cli/up.mjs:27` uses /ok as
its own health probe, and `@langchain/langgraph-api@1.1.17/dist/
api/meta.mjs:64` registers `api.get("/ok", ...)` — so the watchdog
path is correct. The regression is purely timing.
- Railway logs show no successful /ok probe before kill-loop starts.
Fix (Option B — startup grace): add a per-framework startup-grace
window in getWatchdogGraceSeconds(). For langgraph-* starters, the
watchdog now waits up to 180s for the first healthy /ok probe before
arming the strike counter. If /ok comes up sooner, fall through
immediately. If 180s elapses without success, arm the counter anyway —
the steady-state watchdog will then handle a true hang on the normal
90s schedule.
Non-langgraph starters (crewai-crews, ag2, agno, etc.) have grace=0 and
behave exactly as before PR #4116 — the 2–3s `sleep` before the health
check is sufficient for uvicorn-based agents.
Files changed:
- showcase/scripts/generate-starters.ts — new getWatchdogGraceSeconds()
+ grace block in getWatchdogBlock()
- showcase/starters/*/entrypoint.sh — regenerated (grace block for
langgraph-*, comment-only for all others)
- showcase/packages/langgraph-typescript/entrypoint.sh — hand-edited
grace block (this package uses its own hand-edited entrypoint, not
the template)
- showcase/packages/langgraph-fastapi/entrypoint.sh — same
- showcase/packages/langgraph-python/entrypoint.sh — same
Not a revert of #4116: the silent-hang vulnerability class remains
covered (the watchdog still kills after 90s of steady-state failures;
the grace only defers the first strike).
Validated:
- bash -n on all 19 edited entrypoint.sh files
- 1079/1079 showcase-scripts vitest pass
Regression from the 2026-04-21 incident: 18 production Railway services
were found with malformed image refs of the form
`ghcr.io/copilotkit/showcase-<slug>atest` (missing the `:` before
`latest`, so Docker treats `...atest` as the tag). Root cause was an
out-of-band MCP/manual mutation — no committed code touched those refs,
so the data has been fixed but no source-controlled guardrail exists.
Add a standalone script that queries Railway's GraphQL API for every
service in the CopilotKit Showcase project and asserts each image ref
matches the canonical shape `ghcr.io/copilotkit/<service-name>:latest`.
Wire it into showcase_deploy.yml as a pre-build job so any drift aborts
the workflow before the build matrix fans out.
On violation the script prints the service name, the current image, the
expected shape, and the reason, so the fix is obvious in the run log.
Slack classification in the notify job distinguishes a drift failure
from other pre-build failures.
Verified locally: 41 services pass against current Railway state; the
exported `validateImage` function rejects the exact `...atest`
corruption, mismatched service/image names, missing tags, wrong
registries, wrong tag values, and null sources (9/9 simulated cases).
Hand-applied watchdog + unbuffered-stdout block for every Bucket-B package
entrypoint where the shape couldn't come from the starter generator:
* ag2, agno, claude-sdk-python, claude-sdk-typescript, google-adk,
langgraph-fastapi, langgraph-python, langgraph-typescript, langroid,
llamaindex, ms-agent-dotnet, ms-agent-python, pydantic-ai, strands
-> watchdog polls the framework-native health endpoint; Python variants
also get `python -u` + awk fflush log prefixing.
* spring-ai (Bucket C) -> watchdog extends the existing startup probe;
polls :8000/health every 30s and kills PID on 3 consecutive failures.
Added WATCHDOG_PID cleanup to the post-wait block.
* mastra is single-`exec next start` on the package side — no agent
process to probe, so no watchdog needed (N/A).
Reference shape: showcase/packages/crewai-crews/entrypoint.sh (unchanged
from origin/main; proven in production via PRs #4114 + #4115).
Regenerates all 16 starter entrypoint.sh files from the new template. Every
starter now ships:
- $WATCHDOG_PID tracked in cleanup + kill block
- PYTHONUNBUFFERED=1 export (harmless for Java/Node/.NET)
- Explicit `wait -n $AGENT_PID $NEXTJS_PID` (excludes the watchdog so its
exit doesn't short-circuit the agent's true status)
- Per-framework health probe (see prior commit for path matrix)
Also reflects the crewai-crews root-cause fix (already on origin/main via
commit 9379b8855): ag-ui-crewai >=0.2.0 bump and drop of the crewai.cli
monkey-patch shim. These files are regenerated from showcase/packages/
crewai-crews source.
Generalize the watchdog shape proven in showcase/packages/crewai-crews/entrypoint.sh
(PRs #4114 + #4115) so that every Bucket-B starter gets:
- PYTHONUNBUFFERED=1 export (harmless for non-Python frameworks)
- A backgrounded watchdog subshell that polls the agent health endpoint
every 30s and kills the agent after 3 consecutive failures, letting
wait -n + container runtime handle the restart through the normal path.
- python -u on uvicorn / langgraph_cli invocations (Python frameworks).
- awk ... fflush() log prefixing (replaces the previous sed pipe; keeps $!
pointing at the real agent process).
Per-framework agent health paths:
* FastAPI / uvicorn agents -> /health
* langgraph-python, langgraph-fastapi, langgraph-typescript -> /ok
* claude-sdk-typescript, ms-agent-dotnet, spring-ai -> /health
* mastra -> /api
Spring Boot starter also gets a 60s /health startup probe (replacing the
blind sleep 5) because JVM warmup + context refresh can exceed 30s.
Root-cause fix for the 04-21 silent-hang incident on the crewai-crews
Railway deploy. Three tightly-coupled changes:
1. Bump ag-ui-crewai pin from `>=0.1.4,<0.1.6` to `>=0.2.0,<0.3.0`.
0.1.5 had three defects that wedged the agent: unguarded
`source.state.messages` access, an orphan `asyncio.create_task` with
no cancel ref, and a sync `completion()` call that pinned the event
loop. All three are fixed in ag-ui PR #1550, released as 0.2.0 on
2026-04-18. Our upper ceiling was blocking the fix.
2. Remove the pre-bind LLM crash hardening shim in `agent_server.py`.
The shim monkey-patched `crewai.cli.crew_chat.generate_*_description_with_ai`
to static strings so that `ChatWithCrewFlow.__init__` — which
ag-ui-crewai <= 0.1.5 invoked at endpoint-registration time, BEFORE
uvicorn bound its port — could not crash the process before the HTTP
server was listening. 0.2.0 defers `ChatWithCrewFlow` construction to
first request via a module-scoped `_cached_flow` + `asyncio.Lock`
inside `add_crewai_crew_fastapi_endpoint`. Any LLM hiccup now
surfaces as a 5xx on the first request instead of a startup crash,
which is what the shim was reaching for. The shim is dead code on
0.2.0 and has been removed (with `logging` import dropped as it was
only used by the shim).
3. Add `python -u` to the uvicorn invocation in `entrypoint.sh` as a
belt-and-suspenders complement to the existing `PYTHONUNBUFFERED=1`
export. The env var can in principle be un-exported by a child;
`-u` forces unbuffered stdout/stderr at the interpreter level and
is not overridable by user code. Combined with `awk '{...; fflush()}'`
in the pipe (already in place), this guarantees uvicorn request
lines reach Railway's log stream line-at-a-time. During the 04-21
incident Railway saw only ~15 log lines over 9h of uptime because
of buffering through a previous `sed` formulation.
Also updates `showcase/scripts/fail-baseline.json`'s `validatePinsFailHash`
to match the new `ag-ui-crewai` spec string. Pin-drift FAIL count is
unchanged (110); the hash changed only because the `ag-ui-crewai` line
in the FAIL set went from `>=0.1.4,<0.1.6` to `>=0.2.0,<0.3.0`.
Verified locally:
- `pip install -r requirements.txt` resolves `ag-ui-crewai-0.2.0` cleanly.
- `python -u -m uvicorn agent_server:app` starts; `/health` returns
200 `{"status":"ok"}`; request lines appear in real-time logs.
- `pytest tests/python/` — 94/94 pass.
- `pnpm -C showcase/scripts test` (vitest) — 1079/1079 pass.
- `validate-pins.ts` — count=110 matches baseline; hash updated.
Upstream refs:
- crewAI issue: https://github.com/crewAIInc/crewAI/issues/5510
- ag-ui PR #1550: https://github.com/ag-ui-protocol/ag-ui/pull/1550
Intentionally NOT in this PR:
- `showcase/starters/crewai-crews/` parity backport (the starter still
carries the 0.1.5 pin and the shim).
- The 14-starter watchdog generalisation.
Both belong to the silent-hang vulnerability-class work tracked
separately.
Production deploy of showcase-crewai-crews was stuck in agent:"down" for
~4-5h on 2026-04-21. Railway deployment logs show the FastAPI agent on
:8000 was reachable at startup and handled hundreds of requests, then
stopped responding at ~06:56-07:01 UTC after an
unhandledRejection: UND_ERR_BODY_TIMEOUT from the Next.js runtime. The
Python process did not exit — it hung — so bash's wait -n never fired,
the container never restarted, and every subsequent request stacked up
on undici's default 5-minute headers timeout (UND_ERR_HEADERS_TIMEOUT).
Two surgical fixes:
1. entrypoint.sh: add a watchdog goroutine that polls
http://127.0.0.1:8000/health every 30s with a 5s curl timeout. After
3 consecutive failures (~90s unreachable) it SIGKILLs the agent PID,
causing wait -n to return and the container to restart. This converts
a silent agent hang into a fast, loggable restart. Also adds structured
[entrypoint] / [agent] / [nextjs] / [watchdog] log prefixes, a startup
PID-is-alive check, and PYTHONUNBUFFERED=1 so crashes surface before
the log pipe closes — mirroring the starter entrypoint pattern.
2. src/app/api/smoke/route.ts: split the upstream fetch timeout
(UPSTREAM_TIMEOUT_MS = 25_000) from the route's overall budget
(maxDuration = 60). Previously the inner fetch had AbortSignal.timeout(45000)
which, on a hung agent, burned nearly all of the 60s before Next.js
killed the request — producing HTTP 000 at the caller instead of a
structured stage: "timeout" JSON response. With the shorter inner
timeout the route can report a clean timeout within ~25s even when
the agent is wedged.
Both fixes are defensive against upstream CrewAI ChatWithCrewFlow
deadlocks (which the logs show happening repeatedly via the pre-existing
"Cannot send 'RUN_FINISHED' while steps are still active" errors). The
watchdog is scoped to this package only; the starter already has a
richer entrypoint with similar intent.
Directory-plus-index.mdx layout produced two reference entries for
the same logical page: one at `foo/` (the index) and one at `foo`
(a flat sibling if it existed). Collapse index.mdx into the parent
slug so the generated list has a single canonical entry, and bail
out with a clear error when a flat-file collision would shadow the
index. Prevents silently dropping one of the two pages at build
time.
Correctness and portability fixes in the demo-content bundler:
- Track contributor snippets per-file so edits to multiple files in
one commit all get attributed, not just the last one walked.
- Extend endLine when the same file is seen again rather than
dropping the earlier slice — previously a later, smaller region
overwrote a larger one.
- Warn when the watch flag is set on Linux without the recursive
fs.watch support matrix, so the user sees why nothing is firing
instead of assuming silent success.
- Normalize path separators for Windows so bundle manifests use
POSIX paths regardless of the host OS.
The log effect depended on the error object identity, which changes
on every render even when the underlying error is the same — so the
effect fired repeatedly and spammed the log. Depend on
error.message and error.digest (primitive, stable across renders
of the same error) instead. React's exhaustive-deps check is still
satisfied because those are the fields the effect actually reads.
A cluster of UX correctness fixes across the page handlers and
the shared brand nav:
- Filter out undeployed frameworks before rendering so the route
doesn't produce blank pages that users can reach via stale links.
- Strip the leading body H1 with a CRLF-safe regex that only matches
when the body H1 equals the frontmatter title — mirrors ag-ui
route behavior so two stacked titles never render.
- Reference routes now titleCase the slug and resolve via the
index.mdx fallback, matching the directory-plus-index layout the
docs source uses.
- ag-ui title resolver gains a fallback so deep slugs without a
matching registry entry still produce a reasonable title instead
of crashing.
- brand-nav builds the mobile link href dynamically so it points at
the current framework rather than a hard-coded placeholder.
The selector dropdown previously rendered every framework entry as
a clickable option, including ones that had been marked as not
deployed. Selecting an undeployed entry routed to a blank page.
Disable the button for those entries so the dropdown matches the
surrounding tab behavior, which already hides them.
localStorage can hold a value from a previous deploy that no longer
exists in the current frameworks list (package renames, undeploy,
etc.), which pins the provider to a bogus framework identity until
the user manually picks another. Validate the stored value against
the current list both at mount and on cross-tab `storage` events,
and fall back to the default when it doesn't match.
Framework-tabs assumed the items array always matched the frameworks
prop one-to-one. A count mismatch wedged the active index at a stale
value that could fall outside the new frameworks array, rendering a
blank tab. Add a length-mismatch guard that re-syncs state when the
frameworks prop changes between renders, and remove props that were
threaded through but never read so the surface matches what the
component actually uses.