The LangGraph agent server was started without --host, defaulting to
'localhost' which resolves IPv6-first (::1) on modern Node. The Next.js
frontend's /api/health probe hits 'localhost:8123' which can resolve to
IPv4 127.0.0.1 — unbound — producing a 503 with agent:"error".
Add --host 0.0.0.0 so the agent listens on both IPv4 and IPv6 loopback.
Verified locally: /api/health returns 200 with agent:"ok".
parseManifest (lib/manifest.ts):
- Shape errors for demos[i].route: number / null / object / empty string
- Shape error for route not starting with /demos/
- Happy path persists frozen demo.route on the parsed entry
- Backward-compat: demos[i] without route is accepted with undefined route
validate-parity:
- Negative-case regression: missing-demo-dir's expectedDir is derived from
demo.route (not demo.id) when route is present; mismatched id + route
was silently hiding drift.
- Fallback-case: demo with no route resolves expectedDir from demo.id
- routeToDirName unit tests: undefined / bare /demos/ / normal tail segment
TDD verified: mutating the /demos/ prefix guard failed the relevant test
(RED), restoring the guard passed it (GREEN). Mutating expectedDir to
demo.id-only failed the negative-case test (RED), restoring passed
(GREEN).
demo.id is the CATALOG identifier (matched to qa/spec filenames and
shell registry entries). demo.route is the URL + filesystem path
(/demos/<dir> → src/app/demos/<dir>/). They are deliberately separate
— a manifest with id: hitl-in-chat and route: /demos/hitl lives at
src/app/demos/hitl/.
validate-parity.ts previously resolved the demo directory from
demo.id, producing a spurious missing-demo-dir MUST for every such
split. Fix:
- lib/manifest.ts: add optional route field to ManifestDemo; if present,
parser requires it to be a non-empty string beginning with "/demos/".
- validate-parity.ts: introduce routeToDirName helper (matches
bundle-demo-content.ts idiom); loop over demos resolving expected
dir from route and falling back to id. missing-demo-dir PackageIssue
now carries both demoId and expectedDir so deriveMessage can flag
route-resolved paths distinctly.
- __tests__/validate-parity.test.ts: red-green regression test — a
package with id: hitl-in-chat, route: /demos/hitl, and dir
src/app/demos/hitl/ must PASS (no missing-demo-dir error).
Review pass across all 18 langgraph-python demo columns. For each
demo, verified strict minimality (no cross-contamination from other
features) and extracted any inline view JSX >= 15 lines into dedicated
sibling files. All 18 demos validated end-to-end via Playwright
screenshot + visual review.
View components extracted to dedicated files:
- `a2ui-fixed-schema/flight-card.tsx` + `catalog.ts`
- `agentic-chat-reasoning/reasoning-block.tsx`
- `chat-slots/custom-welcome-screen.tsx`, `custom-assistant-message.tsx`,
`custom-disclaimer.tsx`
- `gen-ui-agent/InlineAgentStateCard.tsx`
- `gen-ui-interrupt/InterruptCard.tsx`
- `headless-complete/message-list.tsx`, `user-bubble.tsx`,
`assistant-bubble.tsx`, `input-bar.tsx`, `typing-indicator.tsx`
- `hitl-in-chat/approval-card.tsx`
- `open-gen-ui/sandbox-functions.ts`, `suggestions.ts`
- `tool-rendering/weather-card.tsx`
Backend fixes:
- `mcp_apps_agent.py`: tighten system prompt to prevent the LLM from
hallucinating unrelated framework names in its reply.
- `open_gen_ui_agent.py`: short-circuit subsequent runs when a
`ToolMessage` is already in state. `generateSandboxedUi` is
registered with `followUp: true` by the provider, which caused the
agent to re-emit the tool call in a loop after each sandbox handler
response.
`agentic-chat/README.md`: removed stale references to tools that no
longer exist in the minimized demo.
Regenerated `shell/src/data/demo-content.json`.
Each demo's page.tsx is now pure composition; presentational code lives
in dedicated sibling files. Zero cross-contamination between demos
(audited via grep for forbidden-feature imports).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Parallel review subagents each need a fresh
showcase-langgraph-python build to screenshot. Running
`dev-local.sh up langgraph-python` N times in parallel serializes on
Docker and wastes time rebuilding identical source.
`scripts/agent-notes/rebuild-coord.sh` aligns rebuild attempts to a
30-second clock tick with 0-5s jitter. First agent on a tick acquires
a `mkdir`-based lock, rebuilds, and writes a per-tick done marker at
`/tmp/showcase-rebuild.done-<tick>`. Peers on the same tick skip their
own build and wait up to 30s for that marker. If no marker appears,
they retry the next tick (up to 3 ticks).
Agents call `bash scripts/agent-notes/rebuild-coord.sh` instead of
`./scripts/dev-local.sh up langgraph-python` when they need the live
container refreshed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Move `agentic-chat-reasoning` from `chat-ui` category to
`generative-ui`, positioned after `tool-rendering`.
- Rename display "Agentic Chat (Reasoning)" -> "Reasoning".
- Update langgraph-python manifest demo entry's name, description,
and tag to match.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
These 7 feature ids had zero manifest references, zero demo dirs, and
zero matching packages backing them in the showcase. They added noise
to the feature matrix without carrying substance:
- mobile-react-native
- mobile-swiftui
- mobile-android
- web-svelte
- web-vue
- web-tanstack — conceptually misplaced anyway; TanStack Start is a
React meta-framework, not a React alternative
- web-angular — `packages/angular/` exists as a published SDK, but the
showcase has zero integration for it. Angular support is really its
own matrix/story; removing this placeholder row until there's a real
showcase integration.
Also drops the `mobile` and `web-frameworks` categories now that they
have no members. Regenerates registry.json + constraints.json.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Generated via `generate-registry.ts` and `bundle-demo-content.ts` from
the new manifest + feature-registry + demo sources.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Minimal demo pages covering the Chat & UI category:
- `agentic-chat/page.tsx` — stripped to pure chat (CopilotKit +
CopilotChat, single "Write a sonnet" suggestion). No tool calls.
- `chat-slots/page.tsx` — three slot overrides: custom welcome screen,
custom disclaimer (input sub-slot), and a custom assistant message
bubble so the slot is visible during active chat.
- `headless-simple/page.tsx` — bare-bones custom chat built on
`useAgent` (no CopilotChat) + a single `useComponent` for inline
card rendering.
- `headless-complete/page.tsx` — fuller custom chat: scrollable
message list, user/assistant bubbles, auto-scroll, disabled state
while running, inline tool-call rendering, cancel button.
- `agentic-chat-reasoning/page.tsx` — custom `messageView.reasoningMessage`
slot renders an amber "Reasoning" block when the agent emits
`role: "reasoning"` messages.
- `prebuilt-chat/page.tsx` — minimal `<CopilotChat />`.
- `prebuilt-sidebar/page.tsx` — dummy main content + `<CopilotSidebar />`
(defaultOpen).
- `prebuilt-popup/page.tsx` — dummy main content + `<CopilotPopup />`
(defaultOpen).
Each outer function wraps provider + layout; inner `Chat` holds only
the registrations for that demo's feature. No cross-contamination
(no a2ui/HITL/interrupt/tool-rendering in demos that don't demonstrate
them).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- manifest.yaml: add 12 new demo entries (chat-slots, headless-simple,
headless-complete, agentic-chat-reasoning, hitl-in-chat,
gen-ui-interrupt, declarative-gen-ui, a2ui-fixed-schema, mcp-apps,
open-gen-ui, prebuilt-chat/sidebar/popup).
- langgraph.json: register dedicated graphs for demos that need a
non-default backend (tool_rendering, interrupt_agent, reasoning_agent,
a2ui_dynamic, a2ui_fixed, mcp_apps, open_gen_ui).
- route.ts: register each new agent name; demos with dedicated graphs
are mapped explicitly outside the shared-graph loop. A2UI runtime
middleware scoped to `declarative-gen-ui` + `a2ui-fixed-schema` so
non-A2UI demos keep their `useComponent` renderers intact. MCP Apps
middleware wraps only the `mcp-apps` agent and synthesizes an
`ACTIVITY_SNAPSHOT` event when the backend `show_mcp_app` tool fires.
- New `/api/copilotkit-ogui/route.ts`: dedicated isolated runtime for
the `open-gen-ui` demo. `openGenerativeUI: {...}` advertises the
`openGenerativeUIEnabled` flag globally on the probe response, which
causes the CopilotKit provider's setTools effect to wipe per-demo
`useFrontendTool`/`useComponent` registrations; isolating to its own
endpoint keeps the default runtime clean.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Adds rows for the demos being built in this branch:
- Split `chat-prebuilt` into `prebuilt-chat`, `prebuilt-sidebar`,
`prebuilt-popup` (CopilotChat / CopilotSidebar / CopilotPopup).
- Reorder `generative-ui` rows and add: `hitl-in-chat` (In-Chat HITL),
`gen-ui-interrupt`, `declarative-gen-ui` (Dynamic Schema),
`a2ui-fixed-schema` (new), `mcp-apps`, `open-gen-ui`, `gen-ui-agent`,
`tool-rendering`.
- `constraints.yaml`: expand `constrained-explicit` allowlist so the
langgraph-python manifest validates against the new feature ids.
- Bump expected langgraph-python feature/demo count in
`generate-registry.test.ts` from 10 -> 22 to match the manifest.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Run 24602985507 (singleFork: true) was strictly worse than the prior
fork-per-file run (24602657301):
singleFork (24602985507): 1/14 files completed, 64.64s, then timeout
fork-per-file (24602657301): 14/14 files passed, 169.56s, then teardown timeout
The RPC timeout under Node 20 is a hardcoded 60 s DEFAULT_TIMEOUT in
birpc (DEFAULT_TIMEOUT = 6e4 in vitest's index.B521nVV-.js, not tunable
via config). validate-pins.test.ts alone takes ~60 s because of its
134 `npx tsx` subprocess invocations, so singleFork blows its RPC
budget before a single file finishes. Fork-per-file gives each file
its own fresh 60 s RPC channel, which is why the prior run got all
14/14 files to complete before the final teardown hiccup.
Revert `poolOptions.forks.singleFork: true`. Keep the teardownTimeout
/ hookTimeout bumps to 30 s — they don't affect the RPC timeout but
do protect against slow per-hook teardowns under the same load. Note
in comments that the RPC timeout is upstream-hardcoded and tracked at
https://github.com/vitest-dev/vitest/issues/6129.
`sorts integrations by sort_order` read registry.json without invoking
the generator, making it dependent on whichever run last left disk
state in place. afterEach(dataRestorer.restore()) restores the data
files to HEAD between tests, so test 2 was reading the committed
baseline registry — not the live generator output — and would have
silently agreed with whatever state happened to be on main.
Add a `runGenerator()` call at the top of test 2 mirroring the
`runBundlerAndRead()` pattern that tests 2-5 of bundle-demo-content
already use. Test now exercises the actual generator's sort
behavior under live conditions.
The PR description advertised a drifted-baseline guard on CI for
restoreFromGitHead, but the implementation never actually ran a
post-heal `git diff --quiet HEAD -- <tracked>` check — the only
diff was the off-CI pre-checkout guard against clobbering developer
edits. A racing external mutator (or a parallel test suite we haven't
accounted for) could rewrite a tracked file between our `git checkout
HEAD --` and our subsequent snapshot, and we'd silently bake the drift
into the baseline and the afterEach restore loop would maintain it
forever.
This commit adds the missing post-heal check immediately after the
`git checkout HEAD --`. On CI it throws with `drifted-baseline guard:
post-heal diff failed` citing every offending path; off-CI it warns so
developers iterating on a dirty tree aren't blocked.
Also:
- Rephrases the `isBenignPathspec` catch comment to describe the
realistic case (belt-and-braces against a rm race, not the normal
flow) now that partitionTrackedPaths pre-filters untracked paths.
The branch is intentionally kept — cheap to tolerate, and it
guards against a race that's hard to rule out in a shared worktree.
- Red-green unit coverage: new tests use a counter-based git shim
that fails the N-th `diff --quiet` invocation, so we can target
the post-heal diff independently of the off-CI pre-checkout diff.
- CI throws on drift with the advertised message
- off-CI warns (does not throw) with the same message
- no false-positive on a clean tracked path
Verified red without the guard, green with it.
Run 24602657301 showed unit (20.x) still emitting
`Timeout calling "onTaskUpdate"` during @copilotkit/showcase-scripts:test
teardown — 1008/1008 tests passing, then ELIFECYCLE. The previous fix
switched the pool from threads to forks but left the default fork-per-file
behavior, so every file teardown tore down and reestablished the
parent↔child RPC channel. Under Node 20 the channel occasionally fails
to reestablish within the 10s default teardownTimeout — a known vitest
issue (https://github.com/vitest-dev/vitest/issues/6129).
Two defenses:
- poolOptions.forks.singleFork: true — one long-lived fork across all
test files (combined with the existing fileParallelism: false, files
still run sequentially). The RPC channel stays warm for the whole
run instead of being torn down and reestablished between every file,
which eliminates the per-file teardown as a surface for the race.
- teardownTimeout / hookTimeout bumped from the 10s default to 30s,
matching testTimeout. A slow teardown can't be the bottleneck that
kills the run.
This is the most conservative setting short of dropping to a single-
thread pool entirely. Node 22 / 24 unaffected.
Three small CR8 fixes bundled by area (crewai-crews Python):
* aimock_toggle.py: prod-guard WARNING now uses a grep-friendly prefix
`aimock_toggle: DISABLED (prod-guard: VAR=value)`. Rationale in-file:
`ENV=production` is a common shell-profile default that can silently
disable the toggle in dev/CI without the operator realizing; a
consistent greppable prefix lets them find unexpected refusals without
reading every WARNING line. Strengthened the refusal test to assert
exactly one WARNING is emitted with that prefix (regression guard
against future iteration of _PROD_ENV_VARS firing past the first match).
* agent_server.py: added a `CORSMiddleware` comment documenting why
`allow_origins=["*"]` is intentional for this LOCAL DEMO STARTER — the
agent binds to localhost:8000 (or :8123 inside a generated starter
container) and is only reached by the Next.js frontend during
development. Pointed at `.env.example` / a future `CORS_ORIGIN` env var
as the path for real deployments. No behavior change.
* test_aimock_toggle.py: `test_default_arg_path_mutates_os_environ` now
`monkeypatch.delenv`s every entry in `_PROD_ENV_VARS` before asserting
`enabled is True`. A CI runner or dev shell that exports e.g.
`ENV=production` or `RAILWAY_ENVIRONMENT=production` would otherwise
silently flip this test from "toggle enabled" to "prod-guard refused"
and the enabled-True assertion would fail for the wrong reason.
monkeypatch restores on teardown so other tests are unaffected.
Regenerated crewai-crews starter to mirror the package changes.
The R6 "scope NODE_ENV to Next.js invocation only" fix landed for the
Python Dockerfile template but missed the .NET / Java / TypeScript
templates — they still set `ENV NODE_ENV=production` at the image level,
which leaks into every child process. That's surprising for the TS
templates specifically (mastra dev, langgraph-cli dev, tsx) because those
frameworks gate HMR / source maps / telemetry on `NODE_ENV` and should
NOT see "production" when they're running as the agent backend behind
Next.js.
Fix: remove image-level `ENV NODE_ENV=production` from all three non-Python
Dockerfile templates. The entrypoint template already binds
`env NODE_ENV=production npx next start` so Next.js still sees the value —
only the child agent processes get their environment back. Added a lang-
specific comment to each template explaining the intent and calling out
the affected child processes.
Regenerated the 5 affected starter Dockerfiles
(claude-sdk-typescript, langgraph-typescript, mastra, ms-agent-dotnet,
spring-ai) via `tsx generate-starters.ts`; diffs are byte-identical to the
template change.
Propagates the crewai-crews aimock_toggle.py (adds RAILWAY_ENVIRONMENT_NAME
to `_PROD_ENV_VARS` + updated header comment) and .env.example (full
prod-guard var enumeration) from packages/ to starters/ via
generate-starters.ts. Only crewai-crews ships aimock_toggle.py today so
only that starter changes; other starters re-generate byte-identical.
Verified: `tsx generate-starters.ts --check` OK across all 17 starters.
- Add RAILWAY_ENVIRONMENT_NAME to `_PROD_ENV_VARS`. Railway exposes BOTH
the legacy `RAILWAY_ENVIRONMENT` and the newer canonical
`RAILWAY_ENVIRONMENT_NAME`; both can hold "production" on a prod service,
so both must be guarded. Missing this one was a production-aimock-leak
footgun for anyone hosting on newer Railway service templates.
- Update the parametrized prod-guard tests (both "refuses" + "silent
no-op" variants) to cover RAILWAY_ENVIRONMENT_NAME so new vars added to
`_PROD_ENV_VARS` can't drift from the test matrix.
- Sync .env.example comment block to enumerate ALL 8 prod-guard vars
(NODE_ENV, ENV, ENVIRONMENT, APP_ENV, DEPLOY_ENV, PYTHON_ENV,
RAILWAY_ENVIRONMENT, RAILWAY_ENVIRONMENT_NAME). The old comment said
"NODE_ENV or ENV" which was misleading — operators reading the .env
scaffold had no way to discover the full guard surface.
- Update outdated "Local dev on macOS ships Python 3.9" header comment
in aimock_toggle.py. CI matrix is 3.10/3.12 and requirements.txt
enforces 3.10+ via marker-gated typing_extensions.
Tests: 94 pass (was 88, +6 from RAILWAY_ENVIRONMENT_NAME parametrize entries
across both refuse + silent-no-op variants).
- Rewrite restoreFromGitHead JSDoc to match actual behavior for tracked
vs untracked paths (including the on-CI throw / off-CI warn for an
entirely-untracked input list).
- Rephrase internal review-round markers ("CR4/CR5 HIGH/MEDIUM") in
test-cleanup.ts and test-cleanup.test.ts to describe the behavior or
invariant being guarded instead of the review history.
- Point the WINDOWS comment at the sibling test files that actually
invoke npx (this module itself doesn't).
- Drop the "belt-and-braces" afterAll disclaimer from
bundle-demo-content.test.ts and generate-registry.test.ts — the
afterAll(restore) pattern is self-explanatory.
unit (20.x) CI was failing with 'Timeout calling onTaskUpdate' during
the showcase-scripts test run — an unhandled vitest error that causes
ELIFECYCLE after otherwise-green tests. Reproducible on Node 20 only;
22.x and 24.x are green on the same code.
Root cause: vitest's default thread-based worker pool times out on the
parent-worker RPC channel when a test file spawns many subprocesses
(validate-pins.test.ts runs 134 tests each invoking a subprocess;
create-integration / generate-registry / bundle-demo-content each
spawn npx tsx via execFileSync). Under Node 20 the stdio / signal
traffic from these children contends with the worker-thread RPC
channel and surfaces as an unhandled timeout mid-suite.
Switch to pool: 'forks' — the fork pool uses node IPC for the RPC
rather than worker-thread messageports and is robust under the same
load. fileParallelism: false keeps files sequential so shared-env /
tmp-dir mutations don't race, but each file now gets its own fresh
fork so one file's subprocess churn can't stall the RPC for
subsequent files. Node 22/24 unaffected either way.
Also pipes stdio explicitly on every git subprocess in
test-cleanup.test.ts — inherited stdio on a fork vitest worker
interleaves with the worker's own stdout/stderr and is another input
to the RPC contention under Node 20.
generate-registry.ts writes to shell/src/data/registry.json and
shell/src/data/constraints.json. bundle-demo-content.ts writes to
shell/src/data/demo-content.json. All three files are tracked, and
without restoration every run leaks regenerated JSON into the working
tree.
This commit:
- replaces execSync-with-path-interpolation invocations with
execFileSync argv form via runGenerator() / runBundler() helpers;
eliminates any shell-parser involvement (clean hygiene even when
the interpolated constant happens to be safe today).
- snapshots the three data files in beforeAll, restores them in
afterEach + afterAll, adds a regression-guard test that proves the
hooks actually heal drift (sentinel append + bit-exact in-memory
comparison), and a terminal safety-net bit-for-bit check.
- drops a redundant bundler pre-run in beforeAll (test 1 exercises
the bundler itself), and replaces byte-length sentinel checks with
content-level Buffer.concat assertions so the regression guard
survives any hypothetical fs shim that updates stat but not bytes.
create-integration generates an integration package in
showcase/test-integration-tmp and mutates existing workflow YAMLs in
.github/workflows/ (showcase_deploy.yml, showcase_drift-detection.yml,
starter-smoke.yml) to register the new slug. Without restoration, both
the generated tmp package and the workflow-YAML drift leak into the
working tree on every run, and the workflow-YAML drift in particular
breaks pnpm run check on subsequent test invocations.
This commit wires FileSnapshotRestorer + restoreFromGitHead into the
test so:
- tracked workflow files are restored to HEAD before each test run
- the tmp output dir is deleted eagerly in afterEach (independent of
the generator — no reliance on its internal cleanup)
- a regression-guard test proves the snapshot/restore hooks actually
heal drift via a sentinel append + bit-exact assertion against
the in-memory snapshot (not a re-read of disk, which would
silently agree with a buggy restore()).
- a terminal safety net re-asserts every snapshotted file is
byte-identical to its baseline at the end of the suite.
Introduces FileSnapshotRestorer + restoreFromGitHead helpers used by the
showcase test suites to snapshot tracked files in beforeAll and restore
them in afterEach / afterAll. Several of our test scripts invoke real
generators (create-integration, generate-registry, bundle-demo-content)
that write to tracked files outside any tmp dir: .github/workflows/ and
showcase/shell/src/data/*.json. Without explicit restoration these writes
leak into the working tree on every nx run-many -t test and, under Node
20 + vitest worker pools, the accumulated drift races the worker-RPC
channel surfacing as 'Timeout calling onTaskUpdate' -> ELIFECYCLE on CI.
Highlights:
- FileSnapshotRestorer captures bytes at snapshot time, rewrites only
drifted files via atomic temp+rename, and sweeps leftover
.<basename>.<hex>.tmp stragglers scoped to the snapshotted basenames
(no more whole-directory unlink).
- restoreFromGitHead uses execFileSync with a frozen env (GIT_*
scrubbed, PATH/HOME preserved) to heal a working tree left dirty by
a crashed prior run before we snapshot.
- On CI, a baseline that drifts after the pre-snapshot heal is a hard
error (git binary missing, sandbox, etc.); off-CI it warns instead
of blocking local iteration.
- Narrow catch in the git path partitioner: only genuine 'not in
index' pathspec errors are treated as untracked; ENOENT / EACCES /
non-exit-1 failures re-raise so sandbox and missing-binary cases
don't get silently swallowed and lock in a drifted baseline.
- test-cleanup.test.ts itself strips GIT_* from child env when it
creates tmp repos — pre-commit hooks (lefthook) run with GIT_DIR /
GIT_INDEX_FILE set, which would cause tmp-repo 'git commit' calls
to ignore cwd and write to the HOST working-tree HEAD. Without the
scrub, running 'git commit' itself silently accumulates 'initial' /
'init' commits on the real repo.
Also pulls scripts-dir / repo-root / data-dir constants into a shared
paths.ts so future layout changes flip in one place.
Downstream of the template refactor: Dockerfile + entrypoint.sh across
all 17 starters regenerated to pick up the generic NODE_ENV comment,
the `env NODE_ENV=production` invocation form, and the langgraph-only
gating of `/app/.langgraph_api`. Non-langgraph starters lose the
spurious mkdir layer they never used.
- Reword the NODE_ENV=production comment in Dockerfile + entrypoint
templates to generic language. The previous comment mentioned
`aimock_toggle.py refuses to apply` — that text got rendered
verbatim into .NET / TypeScript / Java starters via the shared
template, misleading readers who don't have a Python toggle. New
comment describes the general behavior (image-level NODE_ENV leaks
into every child process, most of which don't interpret it like
Next.js does).
- Use `env NODE_ENV=production npx next start` instead of bare
`NODE_ENV=production npx next start` in entrypoint.sh. `env` prefix
is the syntactically robust form across shells.
- Gate the `RUN mkdir -p /app/.langgraph_api` layer to langgraph-*
starters only via a new `{{LANGGRAPH_MKDIR}}` template variable.
Creating that directory in crewai / agno / pydantic-ai / claude-sdk
/ etc. was copy-paste residue — langgraph_cli is the only thing
that ever writes there.
- Apply the same NODE_ENV rewording to the crewai-crews SOURCE
Dockerfile + entrypoint (the templates' downstream consumers
regenerate via generate-starters, but the source package is hand-
maintained).
- Extend `_PROD_ENV_VARS` from {NODE_ENV, ENV} to also include
ENVIRONMENT, APP_ENV, DEPLOY_ENV, PYTHON_ENV, RAILWAY_ENVIRONMENT.
Railway is this project's primary hosting target (per CLAUDE.md:
"Always Railway for hosting") so its env var was a prominent gap.
The other additions cover common PaaS / framework conventions —
keep the guard broad (cheap) vs narrow (production-aimock-leak
footgun if someone deploys via a convention we don't recognize).
- Parametrize the prod-guard tests symmetrically across all prod vars.
Both the "refuses with warning" and "silent when URL unset" contracts
now run on NODE_ENV / ENV / ENVIRONMENT / APP_ENV / DEPLOY_ENV /
PYTHON_ENV / RAILWAY_ENVIRONMENT so a future var add that forgets
the silent-path fix would fail here.
- Collapse `_classify`'s "empty" + "whitespace" buckets into a single
"blank" label. The two never produced different behavior at any
use site; keeping them distinct was noise in the state machine.
Diagnostic strings still use "empty or whitespace-only" for the
operator-facing message (the input distinction doesn't matter; the
fix-your-env-value message is identical for both).
Test run: 88 passed (was 78 before; +10 from new prod-var parametrize).
Regenerate all 17 starters via `tsx generate-starters.ts` to propagate:
- aimock_toggle.py: prod-guard only fires when AIMOCK_URL is set;
dummy key renamed to `sk-aimock-dev-ci-only` (from previous file).
- agent_server.py: removed dead `main()`/RELOAD entrypoint (non-langgraph
Python starters).
- Dockerfile + entrypoint.sh: `NODE_ENV=production` moved from global
image ENV onto the `npx next start` invocation so the Python agent
doesn't inherit the prod-guard that would silently disable aimock.
- langgraph-fastapi/.env.example: source-package port fix already propagated
(port was already correct at :8123 for this starter; no content change
beyond reflowing the comment from the source package).
The starter regen is a single atomic commit (separate from the source
edits) so reviewers can inspect which input edit produced which starter
output without wading through unrelated edits in the same diff.
LOW CR5-2 L4: Drop the `test:python` npm script. CI runs pytest directly
(.github/workflows/showcase_validate.yml invokes `python -m pytest
tests/python/ -v` in each package dir) and the script isn't wired into
any Nx target — it's an orphan knob that would only cause drift between
"what CI runs" and "what a contributor gets from `npm test:python`".
LOW CR5-6 L1 (docs): Document the AIMOCK_URL + NODE_ENV=production
interaction in .env.example so operators know the prod-guard refuses to
apply the toggle even when AIMOCK_URL is set, and understand why.
M2: The imported `AGENT_URL_LOCALHOST_8000_RE` carries the /g flag, so
`.test()` / `.exec()` advance lastIndex — sharing the exact instance
across the describe-loop iterations coupled any two iterations that
happened to touch it. Clone the regex per-call via a `re8000()` factory
that returns `new RegExp(source, flags)` each time. Same treatment for
the companion `re8123()`.
LOW: Replace the conditional `it()` registration (only created when the
package .env.example contained `:8000`) with an UNCONDITIONAL `it()`
that internally early-returns when there's nothing to rewrite. Vitest's
reporter registers the test either way so a future regression where every
package suddenly stopped matching the pattern would surface as "test
skipped" rather than silently vanishing from CI output.
M9: `ENV NODE_ENV=production` at the Dockerfile level propagates to every
process spawned by the entrypoint — including the Python agent. The
aimock toggle (aimock_toggle.py) refuses to apply when `NODE_ENV=production`
is set, so a Docker-based smoke test with `AIMOCK_URL` exported would
silently get a no-op toggle and hit real OpenAI instead of aimock.
Move `NODE_ENV=production` out of the global Dockerfile ENV blocks and
inline it on the `npx next start` invocation in entrypoint.sh. Next.js
still treats the server as production (that's all Next looks at), but the
Python sibling process is no longer subject to a production-guard that
only ever made sense for the JS runtime.
Affected: crewai-crews package Dockerfile + entrypoint.sh, plus the
starter template (Dockerfile.python + entrypoint.template.sh) — the
starter regen in a follow-up commit propagates this to all 17 starters.
M7: Every execution path (package `pnpm dev`, starter `pnpm dev`, Docker
entrypoint, CI workflow) invokes `python -m uvicorn agent_server:app ...`
directly via CLI. The `main()` wrapper reading PORT/RELOAD from env was
dead code that:
- was never exercised by any CI path;
- defaulted PORT=8000 while starter dev scripts bound :8123 (drift);
- carried a RELOAD gate whose only purpose was to guard `reload=True`
in a function that never ran.
Remove `main()` and the `if __name__ == "__main__"` dispatch entirely, plus
the now-unused `os` import. A comment at the bottom of the module documents
why no CLI entrypoint lives here so a future contributor doesn't add one
"for consistency" with other showcase packages.
M6: The prod-guard (`NODE_ENV=production` / `ENV=production` refuses to apply
the toggle) was entering its refusal branch even when AIMOCK_URL was unset
— returning a misleading `"aimock toggle refused"` reason when in fact
nothing was being applied. Move the `aimock_url` non-empty check BEFORE the
prod-guard block; the unset-URL path now falls through to the normal "not
configured" diagnostic, reserving the refusal message for actual refusals.
Regression pinned: `test_production_env_without_aimock_url_is_silent_no_op`
now asserts (a) no WARNING is logged, (b) reason text does not contain
"refused", and (c) neither OPENAI_BASE_URL nor LITELLM_API_BASE is mutated.
LOW: Rename `_AIMOCK_DUMMY_KEY` from `"sk-aimock-DO-NOT-SET-AIMOCK-URL-IN-PROD"`
to `"sk-aimock-dev-ci-only"`. The old name embedded a directive aimed at
operators but the value is never user-provided — the toggle itself injects
it when AIMOCK_URL is set and the prod-guard refuses production anyway.
Short, honest, still satisfies SDK versions that validate `sk-` prefix.
Test `test_dummy_key_does_not_mislead_with_replace_in_prod_suffix` extended
to also reject the old "DO-NOT-SET" directive.
LOW: Pin `LITELLM_API_BASE not in env` in
`test_production_env_refuses_toggle_even_with_aimock_url` so a future
regression that mutates LITELLM_API_BASE in the refusal branch (symmetric
with the existing OPENAI_BASE_URL check) fails the test.
The package's own `pnpm dev` script binds langgraph_cli on port 8123
(via package.json "dev"), but the .env.example documented port 8000.
A contributor running `cp .env.example .env && pnpm dev` would hit
ECONNREFUSED out of the box because nothing binds :8000 in this package.
Generated starters (showcase/starters/langgraph-fastapi/.env.example)
already point at :8123; this aligns the package-side .env.example with
both its own dev script and the scaffolded starter output.
Regenerated from showcase/scripts/generate-starters.ts with the updated
propagation rules. Every starter whose package ships .env.example now carries
one (port-rewritten to :8123 where AGENT_URL points at localhost), and the
crewai-crews starter additionally carries aimock_toggle.py alongside the
agent_server.
OPENAI_API_KEY placeholder text in every package + starter .env.example now
reads 'replace-with-your-key' (not a fake sk-... value), so users who skip
editing get a loud 401 from real OpenAI instead of a silent aimock success
masking misconfig.
Extends generate-starters.ts so every starter whose package ships .env.example
gets the file copied through (not just non-langgraph Python), and ships the
aimock_toggle.py alongside agent_server.py for any Python package that has one.
The .env.example copy rewrites AGENT_URL=http://(localhost|127.0.0.1):8000 to
:8123 during propagation because starter dev scripts bind the agent on 8123
while package dev scripts bind on 8000.
The port-rewrite regex is now exported as AGENT_URL_LOCALHOST_8000_RE and
imported by the starter-consistency test, so the two sides cannot drift.
Previously the test used a broader [^:\/]+ host pattern while the generator
correctly narrowed to localhost/127.0.0.1 — a future package documenting
AGENT_URL=https://api.corp.example:8000 would have tripped the test while the
generator correctly preserved the non-localhost hostname.
Adds an env-controlled aimock toggle for the CrewAI Python agent so dev/CI
traffic can be redirected to a local aimock server while production (AIMOCK_URL
unset) is unchanged. When AIMOCK_URL is set, the toggle populates
OPENAI_BASE_URL + LITELLM_API_BASE and injects a dummy OPENAI_API_KEY if one
isn't already set.
- aimock_toggle.py: configure_aimock() with production-env guard — NODE_ENV /
ENV set to prod/production refuses to apply the toggle (fail-safe WARNING)
so a misconfigured deploy can't silently redirect real traffic through a
mock. Dummy key renamed 'sk-aimock-DO-NOT-SET-AIMOCK-URL-IN-PROD' (was
'sk-aimock-dev-only-REPLACE-IN-PROD' — the old name implied operators
should substitute, but the value is never user-facing). Override warning
now symmetric across OPENAI_BASE_URL and LITELLM_API_BASE (both are
overwritten; both must report overrides).
- agent_server.py: uvicorn reload flag gated on RELOAD env (defaults false)
so 'python agent_server.py' in prod doesn't spawn the reloader.
- requirements.txt: typing_extensions>=4.6; python_version < "3.11" so the
NotRequired fallback only installs on older Pythons.
- Dockerfile: updated COPY comment to match reality (two COPY lines, not
three — agents/ is kept separate because it churns independently).
- pytest.ini + conftest.py cleanup: src/ now on the import path via
pythonpath= declarative config (pytest 7+), removing the sys.path.insert
mutation that leaked past test collection.
- 55 pytest cases (was 38), including red-green proofs for: prod-env guard,
LITELLM override warning, dummy-key rename, whitespace-wrapped falsy
USE_AIMOCK values, and the typing_extensions fallback import path.
## Showcase validation tooling (Bundle 3)
Ships three CLI validators, a shared parsing lib, and two CI workflows
that enforce consistency across the 17 showcase packages and detect
drift before it lands on main. Consolidates four earlier tooling PRs
(#3985, #3987, #3995, #3996).
## What's in the box
### `showcase/scripts/` — three validators
| Tool | Purpose | Exit codes |
|------|---------|-----------|
| `audit.ts` | Cross-checks manifest-declared demos against
`tests/e2e/*.spec.ts` and `qa/*.md`, plus `examples/integrations/`
provenance via `SLUG_TO_EXAMPLES` / `FALLBACK_MAP` | 0 ok, 1 anomalies,
2 invalid-input, 3 unreadable, 4 internal, 5 strict-warnings |
| `validate-pins.ts` | Framework-dep pin-drift between
`showcase/packages/*/` and their dojo `examples/integrations/*/`
counterparts. Parses package.json, requirements.txt, pyproject.toml
(Poetry + PEP 621) | 0 ok, 1 drift, 2 internal, 3 unreadable |
| `validate-parity.ts` | Enforces demo ↔ spec ↔ qa coverage per package
with a monotonic demo-count floor | 0 ok, 1 warnings, 2 invalid-input, 3
unreadable, 4 internal, 5 must-failure |
### `showcase/scripts/lib/` — shared primitives
- **`slug-map.ts`** — single-source-of-truth `ENTRIES` for the showcase
slug taxonomy;
`BORN_IN_SHOWCASE`/`SLUG_MAP`/`SLUG_TO_EXAMPLES`/`FALLBACK_MAP` derived
and frozen at module load. `SlugEntry` is a discriminated union that
makes illegal states (born-in-showcase with non-empty examples)
unrepresentable. `freezeSet`/`freezeMap` helpers install throwing
replacements via `Object.defineProperty({writable:false,
configurable:false})` so `Set.add` / `Map.set` truly fail at runtime.
- **`manifest.ts`** — `parseManifest` returns a tagged `ParsedManifest`
union (`ok` | `missing` | `malformed{subkind: "syntax"|"shape"}` |
`unreadable`) with a never-throws content contract. Uses `statSync` +
errno inspection (not `existsSync`, which conflates ENOENT with EACCES).
`DemoId` is a branded string minted only via `createDemoId`.
### `.github/workflows/` — CI enforcement
- **`showcase_validate.yml`** — runs on PR and push-to-main. Enforces
the e2e-spec floor, runs the validators, and drives the pin-drift
ratchet.
- **`showcase_drift-report.yml`** — weekly Monday 10:00 UTC +
workflow_dispatch. Computes `set_status` (OK / SET DRIFTED / COUNT
DRIFTED) and posts to Slack.
Both workflows run on `depot-ubuntu-24.04-4` (Startup plan, unlimited)
for persistent pnpm/npm cache across runs.
## The pin-drift ratchet
`validate-pins.ts` currently finds **111 existing pin-drift failures**
across 12 showcase packages. Rather than block the PR on those, we
baseline them in `showcase/scripts/fail-baseline.json` and ratchet:
- `validatePinsFailCount` must not increase; CI tells you to ratchet
down when it decreases.
- `validatePinsFailHash` is SHA-256 of the sorted-uniqued `[FAIL]` set.
When the count is equal but the hash differs, a fail healed AND a new
one regressed — CI prints the diff and fails.
- `baselineDemoCount` (9) is the single source of truth for the e2e-spec
floor; consumed by both the workflow and `validate-parity.ts` with sync
enforced by a dedicated regression test.
Tracked in #4047. The 111 failures are mostly showcase packages pinning
`@copilotkit/*` to the `next` dist-tag while dojo pins concrete
versions; direction of fix (align showcase → dojo vs. bump dojo →
showcase) is a separate versioning decision outside this PR.
## Correctness posture
- **977 tests**, 13 files, covering every `Anomaly` / `PackageIssue` /
`ParsedManifest` variant in-process and via subprocess CLI for every
exit code. EACCES/ENOTDIR/TOCTOU paths are exercised via chmod probes
(with `it.skipIf` fallback when CI runs as root) and path-filtered
`vi.spyOn` fall-throughs.
- **`fs.statSync` + errno everywhere** — `fs.existsSync` silently
collapses ENOENT with EACCES and is a known anti-pattern in validation
tooling; the codebase uses structured errno discrimination throughout.
- **Tagged discriminated unions with exhaustive `switch` + `never`
guards** — `bucketFor` in `audit.ts`, `deriveMessage` in
`validate-parity.ts`. Adding a new variant without wiring every site is
a compile error.
- **Partial-report preservation** — when an infra error hits
mid-slug-loop, `UnreadableInputError.partialReport` carries
already-collected drift findings so the top-level catch prints them
before exiting 3. One bad package never orphans signal for the rest.
- **Per-slug isolation** — in `validate-parity.ts runParityImpl`, each
slug's audit is wrapped; a crash surfaces as a `crashed` `PackageIssue`
and forces `EXIT_INTERNAL` without aborting siblings.
- **Pipefail + scoped `|| true`** — every workflow step uses `set -euo
pipefail` with grep's no-match tolerance wrapped in `{ grep || true; }`
so producer failures (sort, shasum, cut) still surface.
## Diff
+16,455 / −18 across 34 files (26 source + 5 fixture trees + 2 workflows
+ 1 baseline).
Commits grouped by purpose:
1. `chore(showcase/scripts)`: vitest config + test deps
2. `feat(showcase/scripts)`: shared slug-map and manifest parsing lib
3. `feat(showcase/scripts)`: audit.ts coverage auditor
4. `feat(showcase/scripts)`: validate-pins.ts pin-drift validator
5. `feat(showcase/scripts)`: validate-parity.ts demo/spec/qa parity
validator
6. `ci(showcase)`: validation + weekly drift-report workflows (Depot
runners)
## Test plan
- [x] `pnpm vitest run` in `showcase/scripts/` — 977/977 green
- [x] Exit-code taxonomy verified end-to-end via subprocess tests for
every documented code
- [x] EACCES/ENOENT/ENOTDIR routing verified in all three validators
- [x] Partial-report preservation verified in both in-process and
subprocess paths
- [x] Per-slug crash isolation verified (one broken slug does not orphan
siblings)
- [x] Baseline sync contract (`BASELINE_DEMO_COUNT` ↔
`fail-baseline.json.baselineDemoCount`) pinned by test
- [ ] First CI run on Depot to confirm cold-cache timing (expected 5–8m
vs. 18–20m on ubuntu-latest)
Refs: [Full Action
Inventory](https://www.notion.so/3443aa38185281b5a1dfc6e0890264e1),
#4047
Audits each showcase package's manifest-declared demos for matching
spec (tests/e2e/*.spec.ts) and qa/*.md coverage, and enforces a
monotonic demo-count baseline via fail-baseline.json.
Key design:
- PackageIssue tagged union (13 variants) cleanly separates MUST
errors from warnings; deriveMessage is the single renderer so new
variants cannot emit mismatched prose.
- ProbeResult tagged union (missing | ok | unreadable) driven by
statSync + errno inspection, distinguishing ENOENT from EACCES and
surfacing ENOTDIR as a misconfiguration rather than a silent miss.
- runParityImpl isolates each slug's audit in try/catch; a crash in
one slug surfaces as a crashed PackageIssue variant and forces
EXIT_INTERNAL without aborting siblings.
- runParity never throws for content errors. InvalidBaselineError
covers coerceBaseline failures; unknown errors route to
EXIT_INTERNAL via formatErrorChain (walks .cause with cycle
guard + depth cap).
- parseMainArgs rejects unrecognised flags and duplicate --baseline
with EXIT_INVALID_INPUT (2), mirroring audit.ts parseArgs
discipline.
- BASELINE_DEMO_COUNT default must match
fail-baseline.json.baselineDemoCount; enforced by
__tests__/baseline-sync.test.ts.
- Exit codes: 0 ok, 1 should-warnings-only, 2 invalid-input, 3
unreadable, 4 internal, 5 must-failure.
Tests cover every PackageIssue variant, every exit code
(in-process + subprocess), per-slug crash isolation, EACCES routing,
ENOTDIR classification, cascade suppression when tests/e2e or qa
dirs are unreadable, and baseline coercion edge cases (leading
zeros, negative, float, hex, non-numeric).
Compares framework dependency pins across showcase/packages/*/ and
the corresponding dojo examples/integrations/* trees, flagging drift
between the two and rejecting non-exact specs on the showcase side.
Key design:
- Parses package.json, requirements.txt, and pyproject.toml
(including Poetry's [tool.poetry.dependencies] and PEP 621
[project.dependencies] / optional-dependencies). Separate jsDeps
and pythonDeps maps prevent cross-ecosystem name collisions.
- isExactSpec enforces exact-version pins per ecosystem (npm: no
operators, workspace refs, or ranges; Python: ==X / ===X / ~=X
with PEP 440 body). Symmetric rejection of bare MAJOR-only forms.
- parseRequirementsTxt and parsePyprojectToml thin wrappers throw
when the detailed form produced skipped[] or dropped[] entries,
preventing silent data loss in simpler callers.
- canonicalizeDepMap canonicalises names per PEP 503 and surfaces
same-file collisions with differing specs as warnings.
- First-writer-wins at both file and package levels.
- UnreadableInputError carries an optional partialReport so an
infra failure mid-slug-loop preserves already-collected drift
findings for other slugs.
- Exit codes: 0 ok, 1 drift, 2 internal, 3 unreadable.
fail-baseline.json is the single source of truth for the CI ratchet
(validatePinsFailCount + validatePinsFailHash) and the demo-count
floor (baselineDemoCount, cross-checked against validate-parity.ts
in a dedicated sync test).
Test coverage spans every parser variant, EACCES routing via chmod
probe + fs spies, exit-code taxonomy subprocess tests, partial-report
preservation on mid-loop infra throws, and Poetry/PEP 503 edge cases
via committed fixture files under __tests__/fixtures/pins/.
Cross-checks each showcase package's manifest-declared demos against
the spec (tests/e2e/) and qa/ directory contents, plus
examples/integrations provenance via SLUG_TO_EXAMPLES / FALLBACK_MAP.
Key design:
- Discriminated Anomaly union with nine variants
(count-mismatch, not-deployed, missing-examples, missing-manifest,
malformed-manifest, unreadable-dir, unreadable-manifest,
unreadable-examples, mapped-candidate-not-directory). bucketFor uses
an exhaustive switch with a never guard so a new variant cannot
silently escape routing.
- CountState tagged union separates known-count, legitimate-missing,
and unreadable cases so an EACCES on tests/e2e/ cannot be
misclassified as a real zero count.
- ExamplesSourceResult carries structured unreadableForSlug /
nonDirectoryForSlug flags; classification never substring-matches
the human-readable warning text.
- SHOWCASE_AUDIT_ROOT env var is validated with statSync + distinct
error messages for ENOENT vs ENOTDIR vs EACCES.
- Text and --json output modes. Exit-code taxonomy: 0 ok, 1 anomalies,
2 invalid-input, 3 unreadable, 4 internal, 5 strict-warnings.
- Deep-freezes AuditReport.packages and anomalies before return.
Tests cover every Anomaly variant, every exit code (in-process + CLI
subprocess), buildReport bucket exhaustiveness, EACCES routing via
path-filtered fs spies, TOCTOU ENOENT races, and the --columns filter
surface.
Two foundational modules consumed by all three validators:
- lib/slug-map.ts: single source of truth for the showcase slug
taxonomy. ENTRIES array is the sole declaration; BORN_IN_SHOWCASE,
SLUG_MAP, SLUG_TO_EXAMPLES, and FALLBACK_MAP are derived at module
load and frozen via freezeSet/freezeMap/freezeMap2D helpers
(defineProperty-based to block Set.add / Map.set at runtime).
SlugEntry is a tagged union: born-in-showcase variants have empty
examples and no fallback; non-born variants carry a non-empty
tuple. Each slug passes isShowcaseSlug at module load.
- lib/manifest.ts: parseManifest returns a tagged ParsedManifest
union (ok | missing | malformed | unreadable) with never-throws
content contract. Uses statSync + errno inspection rather than
existsSync to distinguish ENOENT from EACCES/ENOTDIR. DemoId is a
branded string minted only through createDemoId. Empty-string
dirSlug is rejected as a caller bug; undefined opts out of the
slug-match check. Deep-freezes the returned Manifest.
Configure file-level isolation (fileParallelism: false) to prevent
cross-test env-var contamination when the three validators mutate
process.env.VALIDATE_PARITY_REPO_ROOT / VALIDATE_PINS_REPO_ROOT /
SHOWCASE_AUDIT_ROOT for fixture tmpdirs. Adds scripts test deps to
showcase/scripts/package.json.
## Summary
Two showcase MDX files reference `<TailoredContent>` +
`<TailoredContentOption>` from
`@/components/react/tailored-content.tsx`, but the component file was
missing in showcase. This PR:
- Adds `showcase/shell/src/components/react/tailored-content.tsx`
(copied from upstream `docs/components/`)
- Inlines a tiny `cn()` helper to avoid pulling `classnames` into
showcase/shell deps
- Hardens BOTH copies (docs + showcase, kept byte-identical) against
real bugs found in CR review
## Bugs fixed in TailoredContent (both copies)
- **Build-breaker:** `useSearchParams()` in Next.js 14+ App Router
requires `<Suspense>` wrapper or `next build` fails. Added internal
Suspense wrapper so consumers don't need to add one.
- **State/URL desync:** `selectedIndex` was stored in useState
initialized once — back/forward nav didn't update. Now derived from
`searchParams` each render.
- **Keyboard a11y broken:** `role="tab"` + `tabIndex={0}` with no
keyboard handler. Added standard ARIA tab pattern: `role="tablist"`,
`role="tabpanel"`, Enter/Space to select, ArrowLeft/Right + Home/End to
navigate, roving tabindex.
- **`cloneElement` clobbered caller's icon className:** now merges via
`cn()`.
- **`TailoredContentOption` rendered `<div>` despite JSDoc saying "won't
render":** now returns `null`.
- **Unvalidated `defaultOptionIndex`:** clamped to `[0, options.length -
1]`.
- **Empty options silently produced broken UI:** now `return null`
(hooks-rules-safe; no throw mid-render).
- **Duplicate option IDs silently collided:** `console.warn` in dev mode
(via `useEffect`, not render body).
- **Hooks rules violation:** conditional hook ordering under state
changes (fixed by running all hooks unconditionally).
- **Side effects during render:** `console.warn` mutation moved to
`useEffect`.
- **`options`/`optionIds` unstable identities:** memoized via `useMemo`.
## Known remaining (non-blocking)
- The 2 copies of `tailored-content.tsx` must be kept in sync manually.
Proper fix is a shared package — out of scope for this PR.
- Stale `searchParams` race on rapid concurrent clicks across multiple
TailoredContent widgets on the same page (pre-existing upstream).
- `useMemo([children])` is ineffective since React.Children identity
changes per parent render (minor perf).
## Test plan
- [ ] CI green
- [ ] showcase-shell builds without Suspense errors
- [ ] Visit docs pages with `<TailoredContent>` — tabs render, keyboard
nav works, URL reflects selection
## Summary
Adds a new internal-facing showcase app — `showcase/shell-internal` —
that renders a **feature × integration grid**. Each cell links to one of
two new **canonical standalone routes** on the main `shell` app, or
shows a red ✗ when the feature isn't supported.
## What's new
### 1. Two canonical standalone routes in `showcase/shell`
These give every (integration × feature) pair a single, embeddable URL
for each artifact — useful for docs, marketing, and tooling.
- **`/integrations/[slug]/[demo]/preview`** — iframe-only hosted demo,
no chrome
- **`/integrations/[slug]/[demo]/code`** — code viewer only. Supports
URL params for future refinements:
- `?file=<filename>` — which file tab to show
- `?lines=10-20` or `?lines=10-20,35` — highlight specific line ranges
Example:
`/integrations/langgraph-python/agentic-chat/code?file=page.tsx&lines=15-22`
### 2. `showcase/shell-internal` — a new Next.js app on port 3002
- Single grid page: **rows = features**, **columns =
integrations/frameworks** (transpose of shell's existing `/matrix` page,
which has integrations as rows)
- Each cell has **two mini-links** — green `▶ demo` and blue `</> code`
— pointing at the canonical `shell` routes, or a red `✗` if not
supported
- Reads `showcase/shell/src/data/registry.json` directly via relative
import — single source of truth, no duplicate data
- `NEXT_PUBLIC_SHELL_URL` env var (default `http://localhost:3000`) to
point the cells at a deployed `shell` in non-local environments
### 3. Small fix: drop `--turbopack` from shell's dev script
`showcase/shell`'s Next.js 15.4.10 turbopack panics (`"Next.js package
not found"`) on this repo's multi-lockfile layout. Switching to webpack
dev resolves it; production builds (which don't use turbopack) are
unaffected.
## Why two apps instead of one
`shell-internal` could have hosted the demo and code pages itself, but
keeping them in `shell`:
- Makes the canonical URLs reusable outside internal ops (docs,
marketing, linking into product)
- Avoids duplicating the demo-rendering and code-viewer plumbing across
two apps
Internal shell stays a pure overview.
## Test plan
- [ ] `cd showcase/shell && npm run dev` — confirm shell starts on :3000
(webpack, no turbopack panic)
- [ ] `cd showcase/shell-internal && npm install && npm run dev` —
confirm shell-internal starts on :3002
- [ ] Open http://localhost:3002 — verify the feature × integration grid
renders
- [ ] Click a `▶ demo` cell — verify it opens
`http://localhost:3000/integrations/<slug>/<feature>/preview` with only
the iframe demo
- [ ] Click a `</> code` cell — verify it opens
`http://localhost:3000/integrations/<slug>/<feature>/code` with the code
viewer
- [ ] In the code route, try `?file=<name>` and `?lines=10-20` URL
params — verify file switches and lines highlight
- [ ] Verify red ✗ shows for unsupported (integration × feature)
combinations
🤖 Generated with [Claude Code](https://claude.com/claude-code)
⚠️ **Docs sync — MANUAL REVIEW REQUIRED**
This PR was auto-opened because the docs-sync script detected
showcase-local modifications overlapping with upstream changes.
The script attempted a best-effort 3-way merge:
- Where `git merge-file` produced a clean merge, the merged content was
written.
- Where `git merge-file` produced conflict markers, **upstream content
was written as-is** and showcase-local modifications were overridden.
**Manual review required.**
### Review items
```
Files where 3-way merge FAILED — upstream content written as-is, local modifications overridden. Manual review REQUIRED before merging this PR:
- docs/content/docs/integrations/agent-spec/quickstart.mdx
- docs/content/docs/integrations/aws-strands/generative-ui/state-rendering.mdx
Files auto-merged via 3-way merge (clean, no conflict markers — still worth a glance):
- docs/content/docs/integrations/aws-strands/shared-state/in-app-agent-read.mdx
```
### Source
- Upstream ref:
[`3b8e457a1`](https://github.com/CopilotKit/CopilotKit/commit/3b8e457a1)
- Workflow run:
https://github.com/CopilotKit/CopilotKit/actions/runs/24593085026
**Review before merging.** Auto-merge is intentionally disabled
for `needs-review` PRs — confirm the upstream-wins sections
preserve any intentional showcase-local divergence you want to
keep, then merge manually.