Commit Graph

4085 Commits

Author SHA1 Message Date
Jordan Ritter dc49520513 feat(showcase/starters): regenerate 17 starter Dockerfiles for multi-stage builds 2026-04-21 18:14:00 -07:00
github-actions[bot] b241d75652 style: auto-fix formatting 2026-04-22 00:37:17 +00:00
Sam Julien 8c4d0c36f9 fix(shell-docs): remove stale 'integrations' from RESERVED_ROUTE_SLUGS
There is no app/integrations/ route, so the reservation was causing
/integrations/... URLs to 404. Without it, these paths fall through to
UnscopedDocsPage via the non-integration fallthrough in [framework].
2026-04-21 17:35:01 -07:00
Sam Julien b37a3043e6 fix(shell-docs): remove hardcoded langgraph-python from main docs MDX
Strip snippet_framework: langgraph-python from frontmatter and
integration="langgraph-python" from all InlineDemo tags in the
main docs tree (27 files). On framework-scoped pages frameworkOverride
drives Snippet and InlineDemo; on unscoped pages FrameworkGuardedContent
already hides the body until a framework is selected, so no default
is needed.
2026-04-21 17:35:01 -07:00
github-actions[bot] a043aa5ba0 style: auto-fix formatting 2026-04-21 17:35:01 -07:00
Sam Julien a265cc7d48 fix(shell-docs): skip root-flagged dirs from parent nav tree
buildNavTree now checks meta.root before including a subdirectory,
so unselected/ (and any other root:true directory) is excluded from
the main nav tree. This eliminates the flicker where SidebarLink
generated /langgraph-python/unselected/coding-agents and the
[framework] route had to server-redirect to the correct scoped URL.
2026-04-21 17:35:01 -07:00
Sam Julien 64aa05d1b8 fix(showcase/shell-docs): rewrite MDX body links to framework-scoped URLs
On framework-scoped pages, MDX links like /quickstart rendered as plain
<a href="/quickstart"> causing RouterPivot redirect flicker. Override the
MDX `a` component when frameworkOverride is set to prepend the framework
prefix to root-relative internal links.
2026-04-21 17:35:01 -07:00
Sam Julien 87fb9e8260 fix(showcase/shell-docs): fix /<slug> 404 via UnscopedDocsPage fallthrough
Next.js routes /quickstart to [framework]/[[...slug]] (dynamic segment
beats optional catch-all), where "quickstart" is not a registered
integration, causing notFound(). Fix by:

- Extracting the unscoped doc rendering logic into UnscopedDocsPage
- Falling through to it in [framework] when the slug is not a framework
- Simplifying [[...slug]]/page.tsx to handle only the root overview
2026-04-21 17:35:01 -07:00
Sam Julien bb0b14702c fix(showcase/shell-docs): eliminate URL flicker on unscoped navigation
SidebarLink now uses storedFramework as fallback so links on the
overview page go directly to /<framework>/<slug> without a RouterPivot
redirect. OverviewNavItem and section cards now use SidebarLink for the
same reason. FrameworkSelector "Clear selection" was navigating to
/docs/<slug> (broken); fixed to /<slug>. hrefFor now preserves the
current slug when switching frameworks from an unscoped page.
2026-04-21 17:35:01 -07:00
github-actions[bot] 5e7e1be672 style: auto-fix formatting 2026-04-21 17:35:01 -07:00
Sam Julien 3218c54cc1 fix(showcase/shell-docs): remove remaining /docs/ hardcoded hrefs
SidebarLink fallback, framework-scoped backLink, "framework-agnostic
version" banner link, FrameworkLandingPage sidebar, and brand-nav
all still pointed at /docs/* which routes to 404 via the [framework]
catch-all. Switch all to / or /<slug>.
2026-04-21 17:35:01 -07:00
Sam Julien 671ecd1091 fix(showcase/shell-docs): make InlineDemo framework-aware via frameworkOverride
Override InlineDemo in DocsPageView's MDX component map to substitute
defaultFramework for the hardcoded integration prop when a framework
is selected. MDX files don't need to change — the override happens at
the render layer, matching how Snippet already handles this.
2026-04-21 17:35:01 -07:00
Sam Julien 0057d6beb4 fix(showcase/shell-docs): fix broken /docs/* route links in root page
DOCS_SECTIONS, OverviewNavItem, CopilotKit Docs link, backLink, and
slugHrefPrefix all generated /docs/<slug> paths that 404 — the
[framework] catch-all intercepted "docs" as a framework slug and
returned notFound(). Strip the /docs prefix so all root-route links
resolve correctly.
2026-04-21 17:35:01 -07:00
Sam Julien c730330d0d docs(shell-docs): content audit — fix broken snippets, links, and syntax
Page-by-page audit of showcase/shell-docs MDX content. All changes are
in src/content/docs/.

## Broken snippet region fixed
- generative-ui/a2ui/dynamic-schema.mdx: Step 5 referenced
  `runtime-inject-tool` which does not exist in the declarative-gen-ui
  cell. Replaced with a hardcoded code block showing `injectA2UITool: true`
  (same content already shown on the parent a2ui.mdx page).

## MDX syntax fix
- frontend-actions.mdx: stray closing ``` at end of file (would cause
  a parse/render failure).

## Internal link fixes — /docs/ prefix removed (×14 files)
The docs app routing does not use a /docs/ prefix — pages live at
/<slug>. Links using /docs/<slug> hit the reserved-slug guard in the
[framework] route and 404. Fixed across:
  generative-ui/index.mdx (7 links)
  generative-ui/tool-based.mdx (3 links)
  learn/index.mdx (6 links)
  multi-agent/subagents.mdx (2 links)
  shared-state.mdx (2 links)
  shared-state/streaming.mdx (2 links)
  shared-state/agent-readonly.mdx (4 links)
  troubleshooting/debug-mode.mdx
  troubleshooting/error-debugging.mdx
  troubleshooting/migrate-to-1.10.X.mdx
  troubleshooting/migrate-to-1.8.2.mdx (2 links)
  troubleshooting/observability-connectors.mdx

## Internal link fixes — /unselected/ removed (×2 files)
  backend/copilot-runtime.mdx
  backend/custom-agent.mdx
2026-04-21 17:35:01 -07:00
Jordan Ritter 59a09e3e65 feat(showcase/scripts/generate-starters): encode prod-mode conditionals per framework
Add per-slug Dockerfile emitters that pair with the new
AGENT_BUILD_STEPS / AGENT_BUILD_COPY tokens in Dockerfile.typescript:

- getAgentBuildSteps(fw): runs in the builder stage, after `npm run build`.
  Emits `npx tsc` for claude-sdk-typescript (compiles agent/index.ts →
  /app/dist/agent/index.js with flags that match the sibling package
  Dockerfile), and `npx mastra build --dir src/mastra` for mastra
  (bundles the server into .mastra/output/index.mjs). Returns "" for
  every other slug so their Dockerfile cache stays unchanged.

- getAgentBuildCopy(fw): runs in the runner stage, after the agent-code
  COPY. Moves /app/dist (claude-sdk-ts) or /app/.mastra (mastra) from
  the frontend stage into the runner.

- getEntrypointBlock() prod-mode updates: mastra now boots via
  `node /app/.mastra/output/index.mjs` (not `npx mastra dev`) and
  claude-sdk-typescript via `node /app/dist/agent/index.js` (not
  `npx tsx agent/index.ts`). Cold start is a straight `node`
  invocation on Railway — mirrors PR #4132's fix for langgraph-ts.

- Wire AGENT_BUILD_STEPS / AGENT_BUILD_COPY into the `vars` map in
  generateStarterImpl so the template substitution picks up the new
  tokens, and export the two helpers so the test suite can guard them.

Also refresh the generator header comment to describe the multi-stage
shape (builder toolchain vs. minimal runtime) and the prod-mode emitter
wiring.

Tests:
- Replace the legacy `mastra dev` / `npx tsx` entrypoint expectations
  with prod-mode assertions (`node /app/.mastra/output/index.mjs`,
  `node /app/dist/agent/index.js`), plus not-to-contain guards so a
  future refactor can't accidentally re-enable the tsx/dev path.
- Add a dedicated describe block for getAgentBuildSteps /
  getAgentBuildCopy covering the two opted-in slots, the ""
  fallthrough for langgraph-typescript (which has its own server.mjs
  migration path), and the "" fallthrough for every Python slug
  (Python prod-mode is shared-template, not per-slug).

Full `vitest run` in showcase/scripts is green (1085 tests).
Starter Dockerfile regeneration is deferred to a follow-up commit
block once Task 1's template wiring lands.
2026-04-21 17:20:50 -07:00
Jordan Ritter a32d5752d4 feat(showcase/template): true multi-stage Dockerfile templates with prod-mode tokens
Restructure the shared starter Dockerfile templates so the runtime stage
ships only a minimal base plus built artifacts — no dev toolchains (pip,
pnpm, tsx, *-cli dev) run in the runner stage.

Dockerfile.typescript:
- Add AGENT_BUILD_STEPS / AGENT_BUILD_COPY template tokens so per-slug
  prod-mode compiles (claude-sdk-typescript tsc, mastra `mastra build`)
  can be emitted by the generator. Mirrors PR #4132's langgraph-typescript
  fix for the rest of the TS starter slots — cold start becomes a straight
  `node` call instead of `npx tsx` / `npx mastra dev` on Railway.

Dockerfile.python:
- Split pip install into a dedicated agent-builder stage that carries
  full `python:3.12` (with compilers + dev headers needed for wheels
  that don't ship pre-built for -slim), and produces a self-contained
  /opt/venv. The runner then COPYs /opt/venv as a single opaque tree —
  no pip, no build tools in the final image.

Dockerfile.java and Dockerfile.dotnet already implement the true
multi-stage shape (temurin-21 JDK / dotnet-9 SDK builder + JRE / aspnet
runtime); no structural changes needed.

Regenerating the 17 starter Dockerfiles happens in the next commit
block — this commit is template-authoring only.
2026-04-21 17:20:33 -07:00
Jordan Ritter 229b3e1da9 fix(showcase/langgraph-typescript): liveness probe on sibling port to break watchdog restart loop
PR #4132 migrated away from `langgraph-cli dev` to the @langchain/langgraph-api
`startServer` Hono server. That server does bind `/ok` via its meta routes, but
only AFTER `registerFromEnv` imports the graph and spawns the N workers — on
Railway that window has been exceeding the watchdog's 180s startup grace plus
the 90s strike budget, producing an infinite restart loop (deploy ddc88ef9 is
currently flapping ~every 4m30s).

Move liveness off the heavy Hono path entirely:

- server.mjs now spins up a bare node:http server on port 8124 (configurable
  via HEALTH_PORT) *synchronously before any await*, answering GET /ok with
  {"status":"ok"}. Always live within ms of `node` boot.
- entrypoint.sh watchdog now probes 127.0.0.1:8124/ok instead of :8123/ok.

The liveness port is independent of whether :8123 has finished graph import,
so the probe reflects what the watchdog actually cares about: is the Node
process alive and its event loop responsive. A true hang (event loop pinned)
still fails the probe and triggers container restart, so the existing
hang-detection behavior is preserved.

Local Docker smoke (amd64): curl /ok returned 200 within 20s of container
start; LangGraph API logged "listening on 0.0.0.0:8123" ~15s later as normal.
2026-04-21 17:12:23 -07:00
Jordan Ritter 9bcc20b3a1 fix(showcase): restore shell build by copying shell-docs content into scan dir (#4136)
## Summary

- The shell app's Docker build ran `generate-search-index.ts` against
`/app/shell-docs/src/content/{reference,ag-ui,docs}`, but the shell
Dockerfile never copied that directory into the build context. The
script's fail-loud guard (refuses to emit a near-empty index when all
scan roots are missing) then aborted the build. Fix copies
`showcase/shell-docs/src/content/` into the builder stage so the scan
finds its input.
- Removed `COPY --from=builder /app/shell/src/content ./src/content`
from the runner stage — that path was left over from the pre-split
layout and never existed under the current tree, so the runner stage
failed before this could ever succeed.

Main was red on `Showcase: Build & Deploy` for both the `shell` and
`shell-dashboard` matrix entries after #4127 merged. `shell-dashboard`
was unblocked separately (GHCR package linkage). `shell` needed this
source fix.

## Test plan

- [x] `depot build --platform linux/amd64 -f showcase/shell/Dockerfile
.` — succeeds end-to-end (prior failure was inside the builder RUN;
runner stage COPY also failed before the fix).
- [ ] After merge: `Showcase: Build & Deploy` `build (shell, ...)` job
green.
- [ ] After merge: Railway redeploys `showcase-shell` from the new
`:latest`; `https://showcase.copilotkit.ai/` returns 200.
2026-04-21 16:02:07 -07:00
Jordan Ritter 4deeb4d0c4 fix(showcase): restore shell build by copying shell-docs content into scan dir
The shell build ran generate-search-index.ts against shell-docs/src/content
but the Dockerfile never copied that directory into the build context,
so the script failed loudly and the whole shell build aborted. Also drop
the runner-stage COPY of shell/src/content — that path was left over from
the pre-split layout and never existed under the current tree.
2026-04-21 15:54:57 -07:00
Jordan Ritter 53273b96cd fix(showcase/scripts): allow generate-search-index to emit stub when scan dirs missing
The all-dirs-missing guard added in c91c7c567 fataled the shell Docker
build, which intentionally does not COPY shell-docs/src/content/. The
shell only needs the static-pages stub so its header search modal has
something to render (links resolve across to docs.showcase.copilotkit.ai).

Downgrade the fatal to a loud warn + emit the 5-entry static-pages stub.
A misconfigured full build is still visible in logs.

Regression from PR #4127 (8e6991cea) on top of c91c7c567 / debfa6600.
2026-04-21 15:39:41 -07:00
Jordan Ritter 8e6991cea1 Fix showcase dashboard links and dojo langgraph column (#4127)
## Summary

- **shell-dashboard:** dashboard.showcase.copilotkit.ai rendered every
"demo" and "code" link as `http://localhost:3000/...`. Root cause:
`NEXT_PUBLIC_SHELL_URL` was never provided at build time and the source
fell back to `localhost:3000`. Next.js inlines `NEXT_PUBLIC_*` at `next
build`, so a runtime Railway env var could not rescue a bad build. Fix
plumbs the value through as a Docker build arg from
`showcase_deploy.yml` and fails loudly at build if it's unset so this
can't regress silently.
- **shell-dojo:** dojo.showcase.copilotkit.ai was missing items in the
langgraph column (langgraph-python showed 9 demos vs 20+ in the
manifests). Root cause: `shell-dojo/src/data/registry.json` was stale —
the generator only wrote to `shell/`, the dojo Dockerfile never ran the
generator at build, and the CI path filter didn't rebuild the dojo when
manifests changed. Fix dual-emits from `generate-registry.ts` to
`shell/`, `shell-dojo/`, and `shell-docs/`, runs the generator in the
dojo Dockerfile, expands the workflow's path filter to include
`packages/**` and `shared/**`, and refreshes the committed JSON so it
matches what the generator produces today. Langgraph-python demo count 9
→ 32.

## Test plan

- [x] `shell-dashboard` Docker build succeeds with
`NEXT_PUBLIC_SHELL_URL` build arg (Depot `36wlvzkgp1`).
- [x] `shell-dashboard` Docker build fails loudly when the build arg is
omitted (Depot `mbh41c3qtk`).
- [x] `shell-dojo` Docker build succeeds; generator+bundler run at
build; 159 demos bundled (Depot `h7bbq8f8jt`).
- [x] `showcase_deploy.yml` passes YAML validation.
- [ ] After merge + deploy: verify `dashboard.showcase.copilotkit.ai`
links point to `https://showcase.copilotkit.ai/...`.
- [ ] After merge + deploy: verify `dojo.showcase.copilotkit.ai` shows
the full langgraph column (langgraph-python ≥ 20 items).
2026-04-21 15:27:13 -07:00
Jordan Ritter 6bdbb062ff fix(showcase/langgraph-typescript): replace langgraph-cli dev with direct @langchain/langgraph-api server (#4132)
## Summary

Root cause of the 4+ day `langgraph-typescript` flap: `npx
@langchain/langgraph-cli dev` is a dev-mode CLI that compiles TypeScript
on-demand per HTTP request, blocking the single-thread event loop ~5s
per request. On Railway, Studio IPC + `langgraph-api` JIT spawn + graph
compile routinely exceeded the 90s watchdog budget; PR #4123's 180s
grace papered over cold-start but didn't fix steady-state `/schemas`
event-loop blocking.

This PR replaces the CLI with a direct import of `startServer` from
`@langchain/langgraph-api/server` (the same Hono server `dev` wraps,
minus the tsx watcher, Studio handshake, and LangSmith tenant lookup).
Schema extraction runs in a `worker_thread` and is pre-warmed in
parallel with server startup.

Container-tested on `node:22-slim`:
- `/ok` live in <3s (vs. 5-30s dev)
- `/ok` steady-state <10ms
- `/assistants/search` returns the registered `starterAgent` graph
- Pre-warm log: `[server] Pre-warmed schemas in 3402ms`

## Caveats

- Uses `--legacy-peer-deps` for `@langchain/langgraph-sdk` version skew.
- `@langchain/langgraph-api` exports `./server` but it's not officially
public API — exact version pin (`1.1.17`) required.
- `/runs` streaming not end-to-end tested (dummy key);
`/assistants/search` and schema extraction verified.

## Test plan

- [ ] `docker build` of package succeeds (verified locally on
`node:22-slim`, arm64)
- [ ] Railway deploy passes healthcheck within watchdog budget
- [ ] Live probe: `curl
https://showcase-langgraph-typescript-production.up.railway.app/api/health`
returns 200 + `agent: ok`
- [ ] Smoke test run through Studio / demo client end-to-end (real key)

## Related

- Prior attempts: #4116 (watchdog), #4123 (180s grace) — neither
addressed the structural issue
- Proposal context:
https://www.notion.so/3493aa381852814d968df05fc461206d §2c
2026-04-21 15:13:00 -07:00
Jordan Ritter dce445e6bc fix(showcase/langgraph-typescript): replace langgraph-cli dev with direct @langchain/langgraph-api server
Root cause of the 4+ day langgraph-typescript Railway flap: `npx @langchain/langgraph-cli dev` is a dev-mode CLI that compiles TypeScript on-demand per HTTP request, blocking the single-thread event loop ~5s per /schemas call. On Railway, Studio IPC handshake + langgraph-api JIT spawn + graph compile routinely exceeded the 90s watchdog budget; PR #4123's 180s grace papered over cold-start but didn't fix steady-state event-loop blocking.

Replace the CLI with a direct import of startServer from @langchain/langgraph-api/server (the same Hono server dev wraps, minus the tsx watcher, Studio handshake, and LangSmith tenant lookup). Schema extraction runs in a worker_thread and is pre-warmed in parallel with server startup.

Container-tested on node:22-slim:
- /ok live in <3s (vs. 5-30s dev)
- /ok steady-state <10ms
- /assistants/search returns the registered graph
- Pre-warm log: [server] Pre-warmed schemas in ~1.3-3.4s

Caveats:
- Uses --legacy-peer-deps for @langchain/langgraph-sdk version skew.
- @langchain/langgraph-api exports ./server but it's not officially public API — exact version pin (1.1.17) required.
- /runs streaming not end-to-end tested (dummy key); /assistants/search and schema extraction verified.
2026-04-21 14:34:22 -07:00
Martha Schumann bb0a9191bb docs(mastra): simplify AG-UI context access in Agent instructions example
Drops the inline TypeScript typecast from the Mastra agent-app-context
example and uses optional chaining + direct access instead, so the doc
snippet is easier to read and copy. Keeps optional chaining on `.find`
so the example stays safe when the AG-UI context is absent. Also fixes
the `[!code highlight:N]` count after the comment line was removed.

Ports @Abubakar-01's changes from #4125 so they can ship together with
the `requestContext` rename, targeting the new `showcase/shell-docs/`
path after the shell restructure on main.

Co-authored-by: Muhammad Abubakar <abubakaran102025@gmail.com>
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-21 13:44:58 -07:00
Martha Schumann bcd5a92d32 Merge remote-tracking branch 'origin/main' into fix/restore-mastra-readables-content 2026-04-21 13:16:01 -07:00
github-actions[bot] d0e99e66be style: auto-fix formatting 2026-04-21 20:13:15 +00:00
Jordan Ritter e2a6cc2bfe Regenerate shell-dojo registry at build and expand CI trigger paths
The dojo app was missing items under the langgraph column because
shell-dojo shipped a stale committed registry.json. The generator
only wrote to shell/, the dojo Dockerfile didn't run the generator
at build, and the CI path filter didn't rebuild the dojo when
manifest files changed.

Fix: emit from generate-registry.ts to shell, shell-dojo, and
shell-docs; add the generator step to shell-dojo's Dockerfile;
expand the deploy workflow's path filter to include packages/**
and shared/**; and refresh the committed registry/demo-content
JSON so files on disk match what the generator produces today.
2026-04-21 13:10:52 -07:00
Jordan Ritter aa4af554a0 Fix shell-dashboard shell links by plumbing NEXT_PUBLIC_SHELL_URL through build
The shell-dashboard app baked http://localhost:3000 into every demo and code link because NEXT_PUBLIC_SHELL_URL was never provided at build time and the source defaulted to localhost. Next.js inlines NEXT_PUBLIC_* at next build, so setting the value on Railway at runtime does nothing.

Fix: remove the silent localhost fallback, pass NEXT_PUBLIC_SHELL_URL as a Docker build arg from showcase_deploy.yml, and fail loudly if it's unset at build so this can't regress silently.
2026-04-21 13:09:47 -07:00
github-actions[bot] 82f3a7d9bd style: auto-fix formatting 2026-04-21 19:30:23 +00:00
Sam Julien 217ad734c8 fix(showcase/shell-docs): wire up threads + self-hosting from docs sync
The upstream docs sync added 24 new integration pages (`threads.mdx`,
`premium/self-hosting.mdx`) plus learn/reference content, but didn't
land the runtime wiring they depend on — so the pages would ship blank.
It also accepted upstream changes in tool-rendering.mdx over local
modifications that needed to be restored.

- Register `Threads` / `SelfHosting` in SNIPPET_MAP so the shared
  `<Threads />` and `<SelfHosting />` tags inline on integration pages
- Register `ThreadsEarlyAccess` wrapper + `MessageSquareMore` /
  `Network` / `Newspaper` icons in mdx-registry for the new threads
  content and the rebuilt `learn/index` card grid
- Revert tool-rendering.mdx upstream-wins regressions: framework-neutral
  `/unselected/server-tools` link and `components={props.components}`
  forwarding for the `<SharedContent>` override pathway
- Add `threads` + `...premium` to every integration's root meta.json
  plus 11 new `premium/meta.json` files so the new pages appear in the
  sidebar; update `docs/meta.json`, `docs/learn/meta.json`, and
  `docs/premium/meta.json` the same way
- Fix `convertMarkdownTableToHtml` to emit JSX-valid tags. It was
  generating `style="..."` string attributes, which next-mdx-remote
  rejects; this path was dormant until the new threads page put a
  markdown table inside a `<Step>`
2026-04-21 12:21:36 -07:00
Jordan Ritter ed45135100 chore(showcase): regenerate claude-sdk-typescript + mastra starters
Regenerated via `cd showcase/scripts && npx tsx generate-starters.ts`
after the getWatchdogGraceSeconds() mapping update in the previous
commit. Only these two starters change — every other framework still
returns 0s from the grace mapping, so their entrypoints are byte-for-
byte identical to main.
2026-04-21 11:48:02 -07:00
Jordan Ritter c3ddd92be9 fix(showcase): extend watchdog 180s grace to claude-sdk-typescript + mastra
Two more starters are restart-looping with the same pattern #4123 fixed
for langgraph-*: the 90s (3-strike) watchdog budget is shorter than the
cold-start path on a fresh Railway container, so the agent gets killed
before it ever reports healthy, and the loop never breaks.

Railway evidence:
- claude-sdk-typescript (package, :8000): restart-looping since 04-20
  16:54 UTC. Runs compiled `node /app/agent_server.js` which spins up
  the full @anthropic-ai/claude-agent-sdk. Package entrypoint is hand-
  written (not generator-emitted), so add the same grace block inline
  matching #4123's shape.
- mastra (starter, :8123/api): restart-looping since 04-20 18:18 UTC.
  Runs `mastra dev` on :8123 alongside Next.js on ${PORT:-10000}.
  `mastra dev` performs a tsx build + Mastra server boot on first
  request — legitimate supervised process, grace is the right fix.

Package-level mastra has no watchdog (the entire package IS the Next.js
app; no separate agent to watchdog) — PR #4116's classification of
"N/A" was correct for the package. The starter has a separate Mastra
dev server, so the generator-emitted watchdog is legitimate there.

Per-framework grace (not universal) preserves the #4123 design: uvicorn
and express agents are responsive within the 2-3s sleep before
AGENT_HEALTH_CHECK, so adding grace elsewhere would only delay
legitimate restart on true hangs.

This commit covers the generator mapping + the hand-written package
entrypoint. The regenerated starter entrypoints land in the next commit.
2026-04-21 11:47:57 -07:00
Jordan Ritter 8708095d34 fix(showcase): langgraph watchdog regression — startup grace period (#4123)
## Incident

`langgraph-typescript` is in a restart loop on Railway as of **04-20
17:05 UTC** (deployment `58bbebe8-7a94-4f99-b6e4-ffcbb4eb78b9`),
returning 502 to production traffic.

## Root cause

PR #4116 generalized the silent-hang watchdog from `crewai-crews` to
every showcase starter: poll agent health every 30s, kill after 3
consecutive failures (~90s). That's fine for uvicorn/express agents
(responsive in <5s), but `langgraph-cli dev` does a heavy cold-start:

- Studio browser IPC handshake (`createIpcServer`)
- `@langchain/langgraph-api` JIT spawn (`spawnServer`)
- Graph compile

On cold Railway containers this routinely exceeds 90s — the watchdog was
killing the process before `/ok` ever became reachable, producing a
kill-loop every ~90s.

## Phase 1 findings

Verified the health path is correct — not a path bug:

- `@langchain/langgraph-cli@1.1.17/dist/cli/up.mjs:27` uses `/ok` as its
own probe.
- `@langchain/langgraph-api@1.1.17/dist/api/meta.mjs:64` registers
`api.get("/ok", ...)`.
- `langgraph_cli` (Python) serves `/ok` from the same api surface.

The regression is purely a **timing** issue: 90s strike budget is
insufficient for langgraph cold-start on Railway.

## Fix — Option B (startup grace)

Added per-framework `getWatchdogGraceSeconds()` in the generator:

- `langgraph-*` starters: **180s grace**
- All other starters: **0s grace** (matches pre-#4116 behavior)

The grace loop waits up to 180s for the first healthy `/ok` probe before
arming the strike counter:

- First success → fall through immediately and arm the counter.
- 180s elapsed without success → arm the counter anyway (steady-state
watchdog then handles true hangs on the normal 90s schedule).
- Agent dies during grace → exit grace loop, `wait -n` in main shell
handles it.

**Not a revert of #4116.** The silent-hang vulnerability class remains
covered (still kills after 90s of steady-state failures); the grace only
defers the first strike.

## Per-starter changes

| Starter | Grace | Behavioral change |
|---|---|---|
| langgraph-typescript | 180s | **Fixes restart loop** |
| langgraph-python | 180s | Preemptive |
| langgraph-fastapi | 180s | Preemptive |
| crewai-crews, ag2, agno, claude-sdk-*, google-adk, llamaindex,
langroid, mastra, ms-agent-*, pydantic-ai, spring-ai, strands | 0s |
None (comment-only diff) |

## Files touched

- `showcase/scripts/generate-starters.ts` — new
`getWatchdogGraceSeconds()` + grace block in `getWatchdogBlock()`
- `showcase/starters/*/entrypoint.sh` — regenerated (17 files; grace
block for langgraph-*, comment-only for the other 14)
- `showcase/packages/langgraph-typescript/entrypoint.sh` — hand-edited
grace block (this package uses its own hand-edited entrypoint, not the
template)
- `showcase/packages/langgraph-fastapi/entrypoint.sh` — same
- `showcase/packages/langgraph-python/entrypoint.sh` — same

## Validation

- `bash -n` passes on all 19 edited `entrypoint.sh` files
- 1079/1079 `showcase-scripts` vitest pass (`pnpm --filter
@copilotkit/showcase-scripts test`)

## References

- Regression source: #4116
- Railway deployment: `58bbebe8-7a94-4f99-b6e4-ffcbb4eb78b9`
- Incident window: 04-20 17:05 UTC — ongoing at PR creation

## Test plan

- [ ] CI green
- [ ] Deploy to Railway — watch first boot log for `[watchdog] Startup
grace: waiting up to 180s...` followed by `[watchdog] Agent healthy
after Ns — arming strike counter`
- [ ] Confirm no restart loop on `langgraph-typescript`
2026-04-21 11:01:19 -07:00
Jordan Ritter 008d4e58e8 chore(showcase): ratchet validate-pins baseline after Dojo pin bump 2026-04-21 10:51:55 -07:00
Jordan Ritter f5b7a65b3d fix(showcase): langgraph watchdog regression — add startup grace period
PR #4116 introduced a generalized silent-hang watchdog that polls the
agent health endpoint every 30s and kills the agent after 3 consecutive
failures (~90s). This worked for fast-starting agents (uvicorn,
express) but regressed langgraph-typescript on Railway: `langgraph-cli
dev` does a heavy cold-start (Studio IPC setup + @langchain/langgraph-api
JIT spawn + graph compile) that routinely exceeds 90s on a fresh
container, so the watchdog was killing the process before /ok ever
became reachable — producing the 04-20 17:05 UTC restart loop on
deployment 58bbebe8-7a94-4f99-b6e4-ffcbb4eb78b9.

Phase 1 verification:
  - `@langchain/langgraph-cli@1.1.17/dist/cli/up.mjs:27` uses /ok as
    its own health probe, and `@langchain/langgraph-api@1.1.17/dist/
    api/meta.mjs:64` registers `api.get("/ok", ...)` — so the watchdog
    path is correct. The regression is purely timing.
  - Railway logs show no successful /ok probe before kill-loop starts.

Fix (Option B — startup grace): add a per-framework startup-grace
window in getWatchdogGraceSeconds(). For langgraph-* starters, the
watchdog now waits up to 180s for the first healthy /ok probe before
arming the strike counter. If /ok comes up sooner, fall through
immediately. If 180s elapses without success, arm the counter anyway —
the steady-state watchdog will then handle a true hang on the normal
90s schedule.

Non-langgraph starters (crewai-crews, ag2, agno, etc.) have grace=0 and
behave exactly as before PR #4116 — the 2–3s `sleep` before the health
check is sufficient for uvicorn-based agents.

Files changed:
  - showcase/scripts/generate-starters.ts — new getWatchdogGraceSeconds()
    + grace block in getWatchdogBlock()
  - showcase/starters/*/entrypoint.sh — regenerated (grace block for
    langgraph-*, comment-only for all others)
  - showcase/packages/langgraph-typescript/entrypoint.sh — hand-edited
    grace block (this package uses its own hand-edited entrypoint, not
    the template)
  - showcase/packages/langgraph-fastapi/entrypoint.sh — same
  - showcase/packages/langgraph-python/entrypoint.sh — same

Not a revert of #4116: the silent-hang vulnerability class remains
covered (the watchdog still kills after 90s of steady-state failures;
the grace only defers the first strike).

Validated:
  - bash -n on all 19 edited entrypoint.sh files
  - 1079/1079 showcase-scripts vitest pass
2026-04-21 10:48:07 -07:00
copilotkit-devops-bot[bot] c3867ff6cd chore: docs sync from main — needs review (2026-04-21) 2026-04-21 17:17:17 +00:00
Jordan Ritter 7dd930cf32 fix(showcase): add Railway image-ref drift assertion
Regression from the 2026-04-21 incident: 18 production Railway services
were found with malformed image refs of the form
`ghcr.io/copilotkit/showcase-<slug>atest` (missing the `:` before
`latest`, so Docker treats `...atest` as the tag). Root cause was an
out-of-band MCP/manual mutation — no committed code touched those refs,
so the data has been fixed but no source-controlled guardrail exists.

Add a standalone script that queries Railway's GraphQL API for every
service in the CopilotKit Showcase project and asserts each image ref
matches the canonical shape `ghcr.io/copilotkit/<service-name>:latest`.
Wire it into showcase_deploy.yml as a pre-build job so any drift aborts
the workflow before the build matrix fans out.

On violation the script prints the service name, the current image, the
expected shape, and the reason, so the fix is obvious in the run log.
Slack classification in the notify job distinguishes a drift failure
from other pre-build failures.

Verified locally: 41 services pass against current Railway state; the
exported `validateImage` function rejects the exact `...atest`
corruption, mismatched service/image names, missing tags, wrong
registries, wrong tag values, and null sources (9/9 simulated cases).
2026-04-21 10:02:35 -07:00
Jordan Ritter 6f7dd1b8b1 fix(showcase/packages): extend silent-hang watchdog across Bucket-B + spring-ai
Hand-applied watchdog + unbuffered-stdout block for every Bucket-B package
entrypoint where the shape couldn't come from the starter generator:

  * ag2, agno, claude-sdk-python, claude-sdk-typescript, google-adk,
    langgraph-fastapi, langgraph-python, langgraph-typescript, langroid,
    llamaindex, ms-agent-dotnet, ms-agent-python, pydantic-ai, strands
    -> watchdog polls the framework-native health endpoint; Python variants
       also get `python -u` + awk fflush log prefixing.

  * spring-ai (Bucket C) -> watchdog extends the existing startup probe;
    polls :8000/health every 30s and kills PID on 3 consecutive failures.
    Added WATCHDOG_PID cleanup to the post-wait block.

  * mastra is single-`exec next start` on the package side — no agent
    process to probe, so no watchdog needed (N/A).

Reference shape: showcase/packages/crewai-crews/entrypoint.sh (unchanged
from origin/main; proven in production via PRs #4114 + #4115).
2026-04-21 09:20:58 -07:00
Jordan Ritter c17914d480 fix(showcase/starters): regenerate entrypoints + crewai-crews content from upgraded generator
Regenerates all 16 starter entrypoint.sh files from the new template. Every
starter now ships:
  - $WATCHDOG_PID tracked in cleanup + kill block
  - PYTHONUNBUFFERED=1 export (harmless for Java/Node/.NET)
  - Explicit `wait -n $AGENT_PID $NEXTJS_PID` (excludes the watchdog so its
    exit doesn't short-circuit the agent's true status)
  - Per-framework health probe (see prior commit for path matrix)

Also reflects the crewai-crews root-cause fix (already on origin/main via
commit 9379b8855): ag-ui-crewai >=0.2.0 bump and drop of the crewai.cli
monkey-patch shim. These files are regenerated from showcase/packages/
crewai-crews source.
2026-04-21 09:20:50 -07:00
Jordan Ritter de3db5e72e fix(showcase/scripts): add silent-hang watchdog + unbuffered-stdout scaffolding to generator + template
Generalize the watchdog shape proven in showcase/packages/crewai-crews/entrypoint.sh
(PRs #4114 + #4115) so that every Bucket-B starter gets:
  - PYTHONUNBUFFERED=1 export (harmless for non-Python frameworks)
  - A backgrounded watchdog subshell that polls the agent health endpoint
    every 30s and kills the agent after 3 consecutive failures, letting
    wait -n + container runtime handle the restart through the normal path.
  - python -u on uvicorn / langgraph_cli invocations (Python frameworks).
  - awk ... fflush() log prefixing (replaces the previous sed pipe; keeps $!
    pointing at the real agent process).

Per-framework agent health paths:
  * FastAPI / uvicorn agents                                     -> /health
  * langgraph-python, langgraph-fastapi, langgraph-typescript    -> /ok
  * claude-sdk-typescript, ms-agent-dotnet, spring-ai            -> /health
  * mastra                                                        -> /api

Spring Boot starter also gets a 60s /health startup probe (replacing the
blind sleep 5) because JVM warmup + context refresh can exceed 30s.
2026-04-21 09:20:36 -07:00
Jordan Ritter 9379b8855a fix(showcase/crewai-crews): upgrade ag-ui-crewai to 0.2.0 + drop shim + -u flag
Root-cause fix for the 04-21 silent-hang incident on the crewai-crews
Railway deploy. Three tightly-coupled changes:

1. Bump ag-ui-crewai pin from `>=0.1.4,<0.1.6` to `>=0.2.0,<0.3.0`.
   0.1.5 had three defects that wedged the agent: unguarded
   `source.state.messages` access, an orphan `asyncio.create_task` with
   no cancel ref, and a sync `completion()` call that pinned the event
   loop. All three are fixed in ag-ui PR #1550, released as 0.2.0 on
   2026-04-18. Our upper ceiling was blocking the fix.

2. Remove the pre-bind LLM crash hardening shim in `agent_server.py`.
   The shim monkey-patched `crewai.cli.crew_chat.generate_*_description_with_ai`
   to static strings so that `ChatWithCrewFlow.__init__` — which
   ag-ui-crewai <= 0.1.5 invoked at endpoint-registration time, BEFORE
   uvicorn bound its port — could not crash the process before the HTTP
   server was listening. 0.2.0 defers `ChatWithCrewFlow` construction to
   first request via a module-scoped `_cached_flow` + `asyncio.Lock`
   inside `add_crewai_crew_fastapi_endpoint`. Any LLM hiccup now
   surfaces as a 5xx on the first request instead of a startup crash,
   which is what the shim was reaching for. The shim is dead code on
   0.2.0 and has been removed (with `logging` import dropped as it was
   only used by the shim).

3. Add `python -u` to the uvicorn invocation in `entrypoint.sh` as a
   belt-and-suspenders complement to the existing `PYTHONUNBUFFERED=1`
   export. The env var can in principle be un-exported by a child;
   `-u` forces unbuffered stdout/stderr at the interpreter level and
   is not overridable by user code. Combined with `awk '{...; fflush()}'`
   in the pipe (already in place), this guarantees uvicorn request
   lines reach Railway's log stream line-at-a-time. During the 04-21
   incident Railway saw only ~15 log lines over 9h of uptime because
   of buffering through a previous `sed` formulation.

Also updates `showcase/scripts/fail-baseline.json`'s `validatePinsFailHash`
to match the new `ag-ui-crewai` spec string. Pin-drift FAIL count is
unchanged (110); the hash changed only because the `ag-ui-crewai` line
in the FAIL set went from `>=0.1.4,<0.1.6` to `>=0.2.0,<0.3.0`.

Verified locally:
- `pip install -r requirements.txt` resolves `ag-ui-crewai-0.2.0` cleanly.
- `python -u -m uvicorn agent_server:app` starts; `/health` returns
  200 `{"status":"ok"}`; request lines appear in real-time logs.
- `pytest tests/python/` — 94/94 pass.
- `pnpm -C showcase/scripts test` (vitest) — 1079/1079 pass.
- `validate-pins.ts` — count=110 matches baseline; hash updated.

Upstream refs:
- crewAI issue: https://github.com/crewAIInc/crewAI/issues/5510
- ag-ui PR #1550: https://github.com/ag-ui-protocol/ag-ui/pull/1550

Intentionally NOT in this PR:
- `showcase/starters/crewai-crews/` parity backport (the starter still
  carries the 0.1.5 pin and the shim).
- The 14-starter watchdog generalisation.
Both belong to the silent-hang vulnerability-class work tracked
separately.
2026-04-21 08:49:21 -07:00
Jordan Ritter 9ce6513306 fix(showcase/crewai-crews): watchdog agent process + decouple smoke upstream timeout
Production deploy of showcase-crewai-crews was stuck in agent:"down" for
~4-5h on 2026-04-21. Railway deployment logs show the FastAPI agent on
:8000 was reachable at startup and handled hundreds of requests, then
stopped responding at ~06:56-07:01 UTC after an
unhandledRejection: UND_ERR_BODY_TIMEOUT from the Next.js runtime. The
Python process did not exit — it hung — so bash's wait -n never fired,
the container never restarted, and every subsequent request stacked up
on undici's default 5-minute headers timeout (UND_ERR_HEADERS_TIMEOUT).

Two surgical fixes:

1. entrypoint.sh: add a watchdog goroutine that polls
   http://127.0.0.1:8000/health every 30s with a 5s curl timeout. After
   3 consecutive failures (~90s unreachable) it SIGKILLs the agent PID,
   causing wait -n to return and the container to restart. This converts
   a silent agent hang into a fast, loggable restart. Also adds structured
   [entrypoint] / [agent] / [nextjs] / [watchdog] log prefixes, a startup
   PID-is-alive check, and PYTHONUNBUFFERED=1 so crashes surface before
   the log pipe closes — mirroring the starter entrypoint pattern.

2. src/app/api/smoke/route.ts: split the upstream fetch timeout
   (UPSTREAM_TIMEOUT_MS = 25_000) from the route's overall budget
   (maxDuration = 60). Previously the inner fetch had AbortSignal.timeout(45000)
   which, on a hung agent, burned nearly all of the 60s before Next.js
   killed the request — producing HTTP 000 at the caller instead of a
   structured stage: "timeout" JSON response. With the shorter inner
   timeout the route can report a clean timeout within ~25s even when
   the agent is wedged.

Both fixes are defensive against upstream CrewAI ChatWithCrewFlow
deadlocks (which the logs show happening repeatedly via the pre-existing
"Cannot send 'RUN_FINISHED' while steps are still active" errors). The
watchdog is scoped to this package only; the starter already has a
richer entrypoint with similar intent.
2026-04-21 05:39:18 -07:00
Jordan Ritter 4e3a23199d chore(showcase): regenerate demo-content bundles
Regenerate after the bundler fix so the committed bundle reflects
the multi-file contributor tracking and endLine extension behavior.
2026-04-20 22:21:14 -07:00
Jordan Ritter bfe0ff1e7f fix(shell-docs/reference-items): collapse index.mdx to parent slug with flat-file collision guard
Directory-plus-index.mdx layout produced two reference entries for
the same logical page: one at `foo/` (the index) and one at `foo`
(a flat sibling if it existed). Collapse index.mdx into the parent
slug so the generated list has a single canonical entry, and bail
out with a clear error when a flat-file collision would shadow the
index. Prevents silently dropping one of the two pages at build
time.
2026-04-20 22:21:03 -07:00
Jordan Ritter ec8430a939 fix(scripts/bundle-demo-content): multi-file tracking, endLine extension, Linux watch warning, Windows paths
Correctness and portability fixes in the demo-content bundler:

- Track contributor snippets per-file so edits to multiple files in
  one commit all get attributed, not just the last one walked.
- Extend endLine when the same file is seen again rather than
  dropping the earlier slice — previously a later, smaller region
  overwrote a larger one.
- Warn when the watch flag is set on Linux without the recursive
  fs.watch support matrix, so the user sees why nothing is firing
  instead of assuming silent success.
- Normalize path separators for Windows so bundle manifests use
  POSIX paths regardless of the host OS.
2026-04-20 22:20:56 -07:00
Jordan Ritter 272deca9d5 fix(shell-docs/error-boundary): stable effect deps via error.message and error.digest
The log effect depended on the error object identity, which changes
on every render even when the underlying error is the same — so the
effect fired repeatedly and spammed the log. Depend on
error.message and error.digest (primitive, stable across renders
of the same error) instead. React's exhaustive-deps check is still
satisfied because those are the fields the effect actually reads.
2026-04-20 22:20:49 -07:00
Jordan Ritter e84f1f34cd fix(shell-docs): route-level content handling and navigation correctness
A cluster of UX correctness fixes across the page handlers and
the shared brand nav:

- Filter out undeployed frameworks before rendering so the route
  doesn't produce blank pages that users can reach via stale links.
- Strip the leading body H1 with a CRLF-safe regex that only matches
  when the body H1 equals the frontmatter title — mirrors ag-ui
  route behavior so two stacked titles never render.
- Reference routes now titleCase the slug and resolve via the
  index.mdx fallback, matching the directory-plus-index layout the
  docs source uses.
- ag-ui title resolver gains a fallback so deep slugs without a
  matching registry entry still produce a reasonable title instead
  of crashing.
- brand-nav builds the mobile link href dynamically so it points at
  the current framework rather than a hard-coded placeholder.
2026-04-20 22:20:44 -07:00
Jordan Ritter a609132295 fix(shell-docs/framework-selector): disable undeployed option buttons in dropdown
The selector dropdown previously rendered every framework entry as
a clickable option, including ones that had been marked as not
deployed. Selecting an undeployed entry routed to a blank page.
Disable the button for those entries so the dropdown matches the
surrounding tab behavior, which already hides them.
2026-04-20 22:20:30 -07:00
Jordan Ritter 2dc2da4eb2 fix(shell-docs/framework-provider): validate stored framework on mount and cross-tab storage events
localStorage can hold a value from a previous deploy that no longer
exists in the current frameworks list (package renames, undeploy,
etc.), which pins the provider to a bogus framework identity until
the user manually picks another. Validate the stored value against
the current list both at mount and on cross-tab `storage` events,
and fall back to the default when it doesn't match.
2026-04-20 22:20:25 -07:00
Jordan Ritter 206ec9e6bc fix(shell-docs/framework-tabs): guard count mismatch, resync on frameworks prop change, drop unused props
Framework-tabs assumed the items array always matched the frameworks
prop one-to-one. A count mismatch wedged the active index at a stale
value that could fall outside the new frameworks array, rendering a
blank tab. Add a length-mismatch guard that re-syncs state when the
frameworks prop changes between renders, and remove props that were
threaded through but never read so the surface matches what the
component actually uses.
2026-04-20 22:20:19 -07:00