Port the 4084 scripts-layer enhancements so 4085's showcase toolchain
matches the new feature shape:
- lib/manifest.ts: ManifestDemo gains optional `command` field; parser
accepts + validates it (non-empty string, frozen).
- bundle-demo-content.ts: inline `@region[name]` / `@endregion[name]`
comment-marker extraction; informational-only demos (no route, e.g.
cli-start) are skipped; markers stripped from bundled content;
regions: { file, startLine, endLine, code, language } emitted per
demo. External-highlight-file merging (4085-specific) preserved, so
backend agents under src/agents/*.py still flow into the bundle.
- validate-parity.ts: accepts demos at BOTH demos/<cell>/ (4084 layout)
and src/app/demos/<cell>/ (4085 layout); informational demos
(command field) are excluded from the parity audit.
- tests: bundle-demo-content.test.ts expectedDemos updated for the
shared-state rename; generate-registry.test.ts feature count 25→32;
validate-parity.test.ts missing-demo-dir message updated to match
the new dual-location wording.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Port the 4084 internal-showcase-followups state of the shared showcase
schema to 4085 (no Docker restructure):
- feature-registry.json: new Dev Ex category, category reorder
(Dev Ex → Chat & UI → Platform → Controlled → Declarative → Open →
Operational → Interactivity → Agent State → Multi-Agent → BYOC);
+8 new features (beautiful-chat, cli-start, interrupt-headless,
reasoning-default-render, tool-rendering-reasoning-chain,
open-gen-ui-advanced, declarative-gen-ui-hardcoded,
readonly-state-agent-context); renames (shared-state-write →
shared-state-read-write, shared-state-agent-readonly →
readonly-state-agent-context); removals (shared-state-read,
state-rendering, a2ui-*, shared-state-io, shared-state, readables,
a2a-chat, deep-agents, vnext-chat, byoc-a2ui, byoc-tambo);
gen-ui-agent → kind: testing.
- constraints.yaml: allowlists updated to reference the new IDs.
- manifest.schema.json: `route` is optional; `command` field added for
informational demos (e.g. cli-start).
- All 17 column manifests: shared-state-read removed,
shared-state-write renamed to shared-state-read-write.
- langgraph-python manifest: full 4084 features list + demos entries
with 4085-appropriate highlight paths (src/agents/*.py +
src/app/demos/<cell>/* rather than backend/ + frontend/).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Port the full docs routing + rendering infrastructure from the 4084
branch (atai/2026-04-17/internal-showcase-followups):
- New [framework]/[[...slug]] catch-all that validates the first
segment against the registry and renders framework-scoped docs
with a "not available for this framework" banner when the page's
defaultCell is not tagged in the framework's cells.
- FrameworkProvider + selector: framework is strictly URL-derived,
storedFramework is an advisory localStorage signal. No
auto-redirect from localStorage.
- /docs/[[...slug]] rewritten as a thin wrapper around DocsPageView
with a new DocsOverview landing: framework picker + topic cards.
- New lib/docs-render.tsx centralising snippet inlining,
frontmatter parsing, nav tree + breadcrumb builders.
- New lib/mdx-registry.tsx consolidating the ~700-line component
shim used by both routes.
- Server <Snippet> component with demo-content regions + syntax
highlighting via highlight.js.
- Per-feature <InlineDemo> iframe embedder pointing at the
integration's backend (single-container shape
http://localhost:3100/demos/<cell>).
- FrameworkTabs, docs-callout, docs-steps, docs-tabs, router-pivot,
stored-framework-highlight components for the MDX registry.
- Global typography + color pivot in globals.css: Plus Jakarta
Sans body, Spline Sans Mono code, cool-gray background, purple
accent.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Drop the 5 experimental variant-presentation routes + SingleCell helper;
the variant layouts no longer match the enriched cell chrome.
- Rename variant-pieces.tsx -> cell-pieces.tsx; expose a shared
<CellStatus ctx={ctx} /> that renders the docs-og/docs-shell row +
E2E/Smoke/QA/health badges so both the runnable-demo cell and the
informational CommandCell share the same bottom section.
- Add CommandCell (client component): takes the full CellContext, renders
<code> + Copy button for ctx.demo.command, then the same CellStatus row.
- page.tsx branches on ctx.demo.command so the CLI Start row renders a
copy-pasteable command in place of Demo/Code links while keeping the
matrix visually consistent.
- registry.ts: Feature gains kind + og_docs_url/shell_docs_url; Demo.route
is optional and Demo.command is the new informational field. Feature-grid
passes the full demo onto CellContext.
- status.ts: drop variant helpers / types (no longer used).
- Add docs-status.ts reader + a stub docs-status.json so the enriched
DocsRow renders without blowing up the build before the schema agent
lands the generated bundle.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Root cause: `WORKDIR /app` at the top of the runner stage creates /app
owned by root BEFORE the unprivileged `app` user is created. The
`COPY --chown=app:app` lines added in #4092 only reassign the files
being copied, NOT the /app directory itself.
Symptom: starter-mastra CRASHED on Railway with
EACCES: permission denied, mkdir '/app/.mastra'
when `npx mastra dev` tried to create its cache dir under CWD as user
`app`. Likely co-victim class: any starter whose runtime invokes a Node
CLI that writes to CWD (e.g. langgraph-typescript via @langchain/langgraph-cli dev).
Fix: add `RUN chown app:app /app` immediately before `USER app` in all
four Dockerfile templates (typescript, python, dotnet, java), applied
uniformly to all 17 generated starters via the generator.
Verified locally:
- docker build + run starter-mastra → container stays up, /app owned
by app:app, /app/.mastra created successfully, /api/health returns 200.
- docker build + run starter-langgraph-typescript → container stays up
(8+ min), write to /app succeeds, /api/health returns 503 degraded
(expected on dummy OPENAI_API_KEY) — NOT a crash.
All 17 starter Dockerfiles regenerated via
pnpm run generate-starters
and the existing showcase/scripts vitest suite stays green (1073 tests).
- Regenerate agno starter to pick up agno>=2.5.17 (from #4095)
- Ratchet validate-pins fail-baseline hash to match new FAIL set
(count unchanged at 110; hash rotates because agno Dojo/showcase
pair now reflects the SDK upgrade)
agent_server.py now calls app.add_middleware(HealthMiddleware) to serve
/health above the routing layer (restores starter health accuracy after
create_strands_app's catch-all at "/" shadowed @app.get("/health")).
Both _FakeFastAPI stubs — in tests/python/conftest.py and
tests/python/test_instrumentor_patch.py::test_agent_server_module_installs_patch —
only covered the decorator surface (get/post/put/delete/patch). When the
test execs agent_server.py verbatim under the stub, add_middleware raises
AttributeError and fails Python unit tests (3.12). The 3.10 matrix skips
strands (ag_ui_strands requires >=3.12) so the regression only showed
under 3.12.
Mirror the stub's API with a no-op add_middleware so the module loads
cleanly; the patched ThreadingInstrumentor assertion is still the
invariant under test.
generate-starters.ts rewrites relative imports to absolute for langgraph
starters because langgraph_cli loads modules standalone rather than as
packages. The rewrite was flat:
from .X import ... -> from <agentDir>.X import ...
For a file at <agentDir>/tools/get_weather.py, `from .types import ...`
resolves to `tools.types` — the CURRENT package — not `<agentDir>.types`.
The flat rewrite dropped the `tools` segment and produced:
ModuleNotFoundError: No module named 'src.agents.types'
at startup. The agent crashed during module import; the entrypoint
pipe swallowed the traceback (see previous commit for the pipe bug);
the 2-3s sleep guard happened to fire while `sed` was still alive; and
Railway's /api/health probe reported `agent: "error"`.
Make the rewrite subdir-aware: compute the file's containing Python
package from its relative path under agentDest, and prepend that to the
relative import target. So `from .types import ...` inside
`src/agents/tools/get_weather.py` becomes
`from src.agents.tools.types import ...`.
langgraph-python has the same broken imports in its tools/__init__.py
but doesn't crash at runtime because main.py doesn't import from tools
(dead path). Regenerating fixes the dead code too.
The entrypoint.sh template used `cmd 2>&1 | sed 's/^/[agent] /' &`
followed by `AGENT_PID=\$!`. After a pipeline, `\$!` points to the LAST
command in the pipe (the `sed` process), not the agent. Every subsequent
`kill -0 \$AGENT_PID` and `wait -n \$AGENT_PID` was therefore monitoring
`sed`, which stays alive until its stdin closes — long after the agent has
crashed. Railway restarts the container mid-loop; the health probe sees
`{status: "degraded", agent: "down"}` for a few seconds during each cycle.
A second, compounding bug: `sed` buffers by default, and Python agents
buffer their own stdout, so a stack trace emitted during module import
could sit in userspace memory until the pipe closed — by which point the
log was discarded and the real cause of the crash was lost.
Fix both by switching to bash process substitution:
cmd &> >(awk '{print "[agent] " \$0; fflush()}') &
AGENT_PID=\$!
Process substitution does not create a pipeline, so `\$!` remains the
agent's PID. `awk` with `fflush()` flushes each prefixed line to the
container log immediately. Also export `PYTHONUNBUFFERED=1` at the
entrypoint level so Python-based agents don't buffer before awk.
Applies to all 17 starters (python, langgraph-python, langgraph-fastapi,
langgraph-typescript, mastra, typescript, java/spring-ai, csharp/
ms-agent-dotnet). Done once in generate-starters.ts + the template +
regenerated entrypoint.sh files.
Python FastAPI showcase starters were registering an @app.get("/health")
handler BEFORE app.mount("/", ...) (or the adapter's equivalent catch-all).
Starlette's Mount at "/" is a prefix match that swallows every subsequent
request including /health, so the decorated handler never fired and the
Next.js /api/health probe saw a non-JSON agent response and flagged the
starter as "degraded".
Install a HealthMiddleware (BaseHTTPMiddleware) that short-circuits
GET /health above the routing layer, which runs before any Mount can
claim the path. The agent mount is kept intact so AG-UI traffic at / is
unaffected.
Affects 10 Python agent starters: ag2, agno, claude-sdk-python,
crewai-crews, google-adk, langroid, llamaindex, ms-agent-python,
pydantic-ai, strands.
## Summary
Bumps the `agno` Python SDK pin from `>=1.7.8` to `>=2.5.17` (latest
stable on PyPI) in `showcase/packages/agno/requirements.txt`.
## Why
The agno 1.x AGUI interface had a known bug where the terminal SSE event
(`RUN_FINISHED` / `TEXT_MESSAGE_END`) was never emitted, causing
consumers that block on stream close (such as `/api/smoke` calling
`await res.text()`) to hang until AbortSignal fired (25-45s).
Agno 2.5.17 emits the terminal `RUN_FINISHED` frame and closes the
stream cleanly, even when the underlying LLM call fails (dummy API key
-> 401).
## Version
| | Before | After |
|---|---|---|
| Pin | `agno>=1.7.8` | `agno>=2.5.17` |
| Installed | 1.7.x | 2.5.17 |
## Breaking changes encountered
None in this package. All existing imports continue to resolve without
change:
- `from agno.os import AgentOS`
- `from agno.os.interfaces.agui import AGUI`
- `from agno.agent.agent import Agent`
- `from agno.models.openai import OpenAIChat`
- `from agno.tools import tool`
No API adaptation was required.
## Relationship to PR #4093
PR #4093 works around the 1.x SSE bug in the consumer (`/api/smoke`) by
reading the stream incrementally and bailing on `TEXT_MESSAGE_CONTENT
"OK"`. With this upstream fix, that workaround becomes defense-in-depth
rather than load-bearing.
**Recommendation:** leave #4093 in place after this merges. It's cheap,
it's already written, and it protects against future agno regressions in
stream termination.
## Local verification evidence
All tests ran against a local `docker build -t agno-test .` of this
package (`agno==2.5.17`, Python 3.12.13).
### Imports
```
$ docker exec agno-test python -c "import agno; print(agno.__version__)"
2.5.17
$ docker exec agno-test python -c "from agno.os.interfaces.agui import AGUI; ..."
all imports OK
```
### /api/smoke timing (three consecutive runs)
```
run 1: time=0.256s code=200
run 2: time=0.157s code=200
run 3: time=0.151s code=200
{"status":"ok","integration":"agno","latency_ms":12,"timestamp":"2026-04-19T15:03:36.057Z"}
```
Previously: 25-45s hang ending in `AbortError`.
### Raw SSE stream inspection (dummy OpenAI key -> 401)
```
data: {"type":"RUN_STARTED","threadId":"test-thread","runId":"test-run",...}
data: {"type":"RUN_FINISHED","threadId":"test-thread","runId":"test-run"}
```
Stream closes immediately after `RUN_FINISHED`. Terminal event is
present.
### Demos (homepage + 9 demo routes)
```
/ 200 0.008s
/demos/agentic-chat 200 0.010s
/demos/hitl 200 0.005s
/demos/tool-rendering 200 0.005s
/demos/gen-ui-tool-based 200 0.006s
/demos/subagents 200 0.004s
/demos/shared-state-read 200 0.005s
/demos/shared-state-write 200 0.006s
/demos/shared-state-streaming 200 0.005s
/demos/gen-ui-agent 200 0.005s
```
### Agent object
```
agent loaded: Agent
tool count: 8
model: gpt-4o
```
## Test plan
- [x] `docker build` succeeds with new pin
- [x] `import agno` reports 2.5.17
- [x] All existing `from agno...` imports resolve unchanged
- [x] `/api/health` returns 200
- [x] `/api/smoke` returns 200 in <1s (was 25-45s hang)
- [x] Raw SSE stream contains `RUN_FINISHED` terminal event
- [x] Homepage + all 9 demo routes return 200
- [x] Agent object loads with 8 tools
- [ ] CI green on PR
Bumps agno from >=1.7.8 to >=2.5.17 (latest stable on PyPI).
The agno 1.x AGUI interface had a known bug where the terminal SSE
event (RUN_FINISHED) was never emitted, causing consumers that block
on stream close (such as /api/smoke reading via res.text()) to hang
until AbortSignal fired at ~25-45s.
Verified locally against a docker build of this package that 2.5.17
emits the terminal RUN_FINISHED frame and closes the stream cleanly,
even when the underlying LLM call fails (dummy API key -> 401).
Measured /api/smoke round-trip: ~150-400ms (was 25-45s timeout).
All existing imports (agno.os.AgentOS, agno.os.interfaces.agui.AGUI,
agno.agent.agent.Agent, agno.models.openai.OpenAIChat, agno.tools.tool)
continue to resolve without changes — no API adaptation required.
The google-adk package's primary agent uses Gemini via google-adk's
LlmAgent, but the secondary `generate_a2ui` tool was calling OpenAI's
gpt-4.1 through openai.OpenAI(). Cross-provider dependency with no
architectural justification — the A2UI planner's function-calling +
forced tool-choice pattern works equally well via google.genai.
Rewrite to use google.genai.Client.models.generate_content with a
ToolConfig forcing the render_a2ui function call (mode="ANY" +
allowed_function_names). Drop the openai + httpx direct dependency
entirely; google-genai is already pulled in transitively via google-adk
and is now pinned explicitly at >=0.8.0.
Replaces short-term hotfix PR #4090 (which pinned openai in
requirements.txt after post-merge Railway deploy of #4083 crashed
with ModuleNotFoundError).
Changes:
- src/agents/main.py: `_get_openai_client` → `_get_genai_client`
(lru_cached google.genai.Client). Rewrote generate_a2ui to build a
`types.Tool(function_declarations=[FunctionDeclaration])`, force the
call via `types.ToolConfig(function_calling_config=...)` and read
the response from `candidates[0].content.parts[*].function_call.args`.
Narrowed except clause to `genai_errors.APIError`/`ClientError`/
`ServerError` + `ValueError` (for construction-time failures).
A2UI_MODEL env var overrides the default gemini-2.5-flash model.
- entrypoint.sh: `OPENAI_API_KEY` guard → `GOOGLE_API_KEY` guard;
`REQUIRE_OPENAI_API_KEY` → `REQUIRE_GOOGLE_API_KEY`.
- requirements.txt: explicit `google-genai>=0.8.0` pin; no openai, no
httpx.
- src/app/api/health|debug/route.ts: report GOOGLE_API_KEY status, not
OpenAI/Anthropic/LangSmith (this is a Gemini-only package).
- playwright.config.ts: inject GOOGLE_API_KEY to the dev server env, not
OPENAI_API_KEY / OPENAI_BASE_URL.
- tests/python/test_generate_a2ui.py: rewrote 19 tests to mock
google.genai.Client in place of openai.OpenAI. Preserves every error
branch (a2ui_llm_error, a2ui_empty_response, a2ui_no_tool_call,
a2ui_invalid_arguments) plus happy path. Added A2UI_MODEL override
coverage.
- tests/python/test_entrypoint_env_guards.py: migrated to exercise the
GOOGLE_API_KEY / REQUIRE_GOOGLE_API_KEY branch.
Verification:
- Unit tests: 53 passed locally against Python 3.12 (google-adk + google-
genai + all package deps installed).
- Docker build: clean (showcase-google-adk:local) via
scripts/dev-local.sh.
- Container startup: `docker compose up -d google-adk` launches cleanly
(no ModuleNotFoundError); uvicorn reports healthy; entrypoint emits
the new WARN when GOOGLE_API_KEY is unset.
- L1+L2 smoke against local compose: both pass for google-adk.
- `import agents.main` inside the container no longer pulls in any
`openai.*` module.
## Summary
Replaces `RUN chown -R app:app /app` with `COPY --chown=app:app` across
all 17 starter Dockerfiles + 4 shared templates.
Every starter ended with a recursive chown over `/app`, which walks ~50k
files (Next.js `node_modules` dominates) to fix ownership after the
fact. Under the 23-way Depot runner fan-out used by
`.github/workflows/deploy-showcase-services.yml`, that step consistently
ran 5+ minutes under I/O contention — busting the **15-minute GH Actions
job budget** before images could finish pushing.
Failed run this fixes:
https://github.com/CopilotKit/CopilotKit/actions/runs/24621469277 (22/23
showcase services cancelled mid-push).
## What changed
- Create the `app` user right after `WORKDIR /app` in every runner stage
so `--chown=app:app` resolves by name.
- Add `--chown=app:app` to every runner-stage COPY (including
multi-stage `COPY --from=frontend` and `COPY --from=<agent-builder>`).
- Drop the trailing `RUN chown -R app:app /app` (or the tail of the
compound RUN in the TS starters).
- For langgraph starters, fold `chown app:app /app/.langgraph_api` into
the same RUN as the `mkdir`, so `langgraph_cli`'s scratch dir remains
writable by the runtime user.
- Update `showcase/scripts/generate-starters.ts` so the shared templates
(`Dockerfile.python/typescript/dotnet/java`) emit the new pattern.
Starters are regenerated from templates.
## Local timing (Docker Desktop, single starter, no contention)
| Starter | Pre-fix | Post-fix | `chown -R` step |
| --- | --- | --- | --- |
| ag2 (Python) | 3:03 | 1:39 | 50.1s → removed |
| langgraph-fastapi | n/a | 1:41 | replaced with 0.2s targeted chown on
`.langgraph_api` |
| mastra (TS) | n/a | 2:16 | removed |
The real-world win on the 23-way Depot fan-out is substantially larger
than the local 50s baseline — recursive chown degrades super-linearly
with concurrent I/O pressure, which is exactly what the 5+ min Depot
step demonstrated.
## Smoke tests
- ag2 image: container starts clean as `app` user, all files under
`/app` owned by `app:app`, `/api/health` returns 200.
- langgraph-fastapi image: `/app/.langgraph_api` exists and is owned by
`app:app`.
- mastra image: builds cleanly, ownership correct.
## Test plan
- [ ] CI green (existing showcase starter smoke suite covers startup)
- [ ] Watch the next showcase deploy workflow run — expect jobs to
finish well under 15m
## Not touched
No workflow files, no `examples/` Dockerfiles (none matched the problem
pattern), no `chmod -R` offenders (none found).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
## Summary
Two related defects in the showcase deploy pipeline let stale images sit
live on Railway while Slack stayed green. This PR fixes both.
### Defect 1 — Drift detector skipped all starter services
`.github/workflows/showcase_smoke-monitor.yml` listed only 19
**package** slugs in its `SERVICES=(...)` array (ag2, mastra,
llamaindex, ...). Zero **starter** slugs. As a result:
- GHCR `showcase-starter-<svc>` tags were never checked for drift.
- `gh workflow run showcase_deploy.yml -f service=starter-*` was never
auto-dispatched.
- Starter services could run with weeks-old images and no alert would
fire.
`showcase_deploy.yml` already supports `starter-*` dispatch names and
already calls `serviceInstanceRedeploy` for any service with a
`railway_id`, so no change is required there. The fix is extending
`SERVICES=(...)` to include all 17 starter slugs via a
sparse-checkout-driven filesystem enumeration (no more literal
duplication between workflow and `showcase/starters/`).
### Defect 2 — Silent deploy failures reported green
`.github/workflows/showcase_deploy.yml` emitted `::warning::` and exited
0 when a service never returned 200 on its health path within 360s. The
legacy justification (`# Don't fail — sleep-on-idle services take time
to wake`) no longer applies: Railway is on the Pro tier with no
sleep-on-idle, so a 6-minute failure to become healthy is a real
failure. Changed to `::error::` + `exit 1`.
## Round 2 fixes
Round 2 CR raised six findings against the original smoke-monitor +
validator changes. All fixed in this PR:
- **BLOCKING 1/2 — smoke-monitor guard.** Replaced the magic `-eq 19`
sentinel with `grep -c '^starter-'` so adds/removes to the literal
non-starter list can't silently disable the guard. Added `shopt -s
nullglob` around the `showcase/starters/*/` loop so an empty starters
tree no longer corrupts `SERVICES` with a `starter-*` literal.
- **BLOCKING 3 — GHCR stderr isolation.** Dropped `2>&1` on the `gh api
-i` call; captured stderr to a temp file and surfaced it only when `gh`
returns a non-zero RC with no HTTP status. Auth / rate-limit / network
noise can no longer splice into the HTTP header block and poison
`HTTP_STATUS` / `API_BODY` parsing.
- **BLOCKING 4 — validator tests.** Added
`showcase/scripts/__tests__/validate-workflow-starters.test.ts` (12
specs): happy path, missing-from-options-only, missing-from-matrix-only,
missing-from-both, empty starters dir (exit 3), template/ excluded,
substring-spoof (starter-ag2 vs starter-ag2-extended), missing workflow
file (exit 3). Also extended `VALIDATE_WORKFLOW_STARTERS_REPO_ROOT` to
re-home the starters dir for testability.
- **MEDIUM 1 — YAML parsing.** Replaced the fragile regex-over-YAML
options scanner with a real `yaml.parse()` + typed navigation down
`on.workflow_dispatch.inputs.service.options`. ALL_SERVICES stays
regex-scanned (embedded JSON in a bash heredoc, with `${{ ... }}`
interpolations that aren't valid JSON pre-execution), but the
surrounding step is now located via YAML.
- **MEDIUM 2 — Slack list truncation.** Replaced `cut -c1-200` with a
`truncate_csv` helper that drops whole comma-separated entries until
under budget and appends `…` when truncated. No more
`starter-claude-sdk-pyth` mid-slug corruption.
- **MEDIUM 3 — template/ exclusion cross-references.** Both the TS
validator's `EXCLUDED_DIRS` and `showcase_smoke-monitor.yml`'s `[
"$slug" = "template" ] && continue` now carry `# keep in sync with ...`
comments pointing at each other.
- **NIT 2 — entry-check simplification.** Dropped the
belt-and-suspenders `import.meta.url === \`file://${argv[1]}\`` branch;
kept only the canonical `fileURLToPath(import.meta.url)` form.
- **NIT 3 — jq pipeline collapse.** Single-pass `.jobs[]? | select |
"\(...)"` replaces the three-pass `map | map | .[]` chain in the notify
step.
## Files changed
- `.github/workflows/showcase_deploy.yml` — warning → error + exit 1 on
unhealthy deploy; `truncate_csv` replaces `cut -c1-200` (3 sites);
single-pass jq pipeline in notify step.
- `.github/workflows/showcase_smoke-monitor.yml` — filesystem-driven
`SERVICES=(...)`, starter-count guard, `nullglob` loop, stderr-isolated
`gh api` call, cross-reference comment.
- `.github/workflows/showcase_validate.yml` — wires
`validate-workflow-starters` into CI.
- `showcase/scripts/validate-workflow-starters.ts` — YAML-aware presence
checks; env-var override homes both starters dir and workflow path.
- `showcase/scripts/tsconfig.json` — scripts-local tsconfig for LSP type
resolution.
- `showcase/scripts/__tests__/validate-workflow-starters.test.ts` — 12
specs covering the full matrix of drift scenarios.
## Test plan
- [ ] Next scheduled `showcase_smoke-monitor` run includes starter
services in its drift scan.
- [ ] A deliberately-unhealthy deploy (simulate by pointing health_path
at a 404) fails the job and fires the Slack alert.
- [ ] `showcase_validate` CI job runs `validate-workflow-starters` and
`npx vitest run scripts/__tests__/validate-workflow-starters.test.ts`
green.
Every starter Dockerfile ended with `RUN chown -R app:app /app`, which
recursively walks the full `/app` tree (Next.js `node_modules` dominates
— ~50k files) to fix ownership after the fact. On the 23-way Depot
runner fan-out used by .github/workflows/deploy-showcase-services.yml
this step consistently took 5+ minutes under I/O contention, busting
the 15-minute GH Actions job budget before the image could even finish
pushing (run 24621469277 — 22/23 services cancelled).
The standard Docker idiom is `COPY --chown=<user>:<group>` which applies
ownership during the copy step itself — no extra layer, no full-tree
traversal, zero runtime cost.
Changes:
- Create the `app` user right after `WORKDIR /app` in every runner
stage so `--chown=app:app` resolves by name.
- Add `--chown=app:app` to every COPY instruction that lands files
under `/app` in the runner stage (including `COPY --from=frontend`
and `COPY --from=<agent-builder>` multi-stage copies).
- Drop the trailing `RUN chown -R app:app /app` (or the tail of the
compound RUN in the TS starters).
- For langgraph starters, fold `chown app:app /app/.langgraph_api`
into the mkdir RUN so langgraph_cli's scratch dir is still writable
by the runtime user.
- Update showcase/scripts/generate-starters.ts so the shared templates
(Dockerfile.python/typescript/dotnet/java) emit the new pattern and
the framework-specific COPY lines (`COPY ${dest} ./` for extra files
and `COPY agent_server.py ./` for Python starters) include `--chown`.
Local verification (Docker Desktop, single starter, no contention):
ag2 pre-fix: 3:03 total, `RUN chown -R app:app /app` = 50.1s
ag2 post-fix: 1:39 total, no chown step
langgraph-fastapi post-fix: 1:41 (langgraph_api chown is 0.2s)
mastra post-fix: 2:16
Smoke test on ag2 image: container starts clean, all files under /app
owned by app:app, /api/health returns 200.
Under the 23-way Depot fan-out the real-world win is substantially
larger than the local 50s baseline because recursive chown degrades
super-linearly with concurrent I/O pressure.
Failed run that motivated this: https://github.com/CopilotKit/CopilotKit/actions/runs/24621469277
## Summary
The Deployed Starters smoke test is falsely reporting `status=404
path=/health` for every starter whose agent process is degraded. The
reported path is misleading — it's just the last path probed before the
test gave up. The real upstream is a 503 on `/api/health`.
## Root cause
`checkHealth()` in `showcase/tests/e2e/helpers.ts` defaulted to probing
`[\"/api/health\", \"/health\"]` in sequence. All deployed starters and
showcase backends serve their health endpoint at `/api/health` (standard
Next.js convention). `/health` returns a Next.js 404 HTML page.
When `/api/health` returns a legitimate 5xx (e.g. 503 `\"agent
degraded\"`), the helper's `res.ok()` check is false, so it falls
through to probe `/health`, gets a 404, and reports that as
`lastResult`. The Slack alert then reads `path=/health` and implies a
route-missing bug, hiding the real 503.
Verified with live probes:
- Healthy starter `/api/health` → `200 {\"status\":\"ok\",...}`
- Degraded starter `/api/health` → `503
{\"status\":\"degraded\",\"agent\":\"error\",...}`
- Either starter `/health` → `404` (Next.js not-found page)
## Fix (Shape A — remove the fallback)
Change the default `paths` argument from `[\"/api/health\",
\"/health\"]` to `[\"/api/health\"]`. No fallback. The reported status
and body are the actual `/api/health` response.
- `starter-smoke.spec.ts` passes an explicit `paths` argument and is
unaffected.
- `integration-smoke.spec.ts` L1 health check targets `/api/health` on
all deployed backends (verified live on langgraph-python and mastra).
## Out of scope
15/17 starters currently return 503 because their agent processes are
failing. That's a separate investigation — this PR is purely about
making the smoke alert accurate so that investigation can proceed
per-service.
## Evidence runs (all same failure mode tonight)
- https://github.com/CopilotKit/CopilotKit/actions/runs/24619147643
- https://github.com/CopilotKit/CopilotKit/actions/runs/24618944101
- https://github.com/CopilotKit/CopilotKit/actions/runs/24617315522
## Test plan
- [ ] Next scheduled smoke run surfaces `status=503 path=/api/health`
for the 15 degraded starters (agent-down state, clearly attributable)
- [ ] Healthy starters (langgraph-python, one other) continue to pass L1
- [ ] `starter-smoke.spec.ts` (local Docker starters) unaffected — still
uses explicit `starter.healthPaths`
## Summary
Cold-start 502s on showcase package smoke probes (notably `agno`)
consistently land just over the 25s upstream budget. The Next.js
`/api/smoke` route uses `AbortSignal.timeout(25000)` to call its own
`/api/copilotkit` which proxies to the Python agent which calls the LLM.
On cold start, the end-to-end round-trip takes >25s and the probe aborts
with `latency_ms: 25001, stage: "timeout"`, returning 502.
This PR raises the upstream budget from **25s -> 45s** across all
showcase packages that exercise a full agent round-trip on smoke, and
bumps Next.js `maxDuration` from 30s -> 60s so the route can actually
run that long. The `create-integration` template is bumped too so new
packages inherit the new budget.
The 3 `langgraph-*` routes have a different (lightweight health probe)
pattern and `ms-agent-dotnet` is already at 50s — their `maxDuration` is
bumped for consistency but their abort budgets are left alone.
**Evidence — agno 502 alerts tonight (PDT):**
- 18:40 PDT attempt 1 — HTTP 502 on `/api/smoke`, self-recovered
- 19:42 PDT attempt 2 — HTTP 502 on `/api/smoke`, `/api/health` already
200
Both probes had `latency_ms: 25001, stage: "timeout"` — a textbook
cold-start boundary trip.
## Why 45s (not 60s or 90s)
The alert tells us the cold path lands **just over** 25s. 45s gives ~20s
of headroom — enough to absorb cold starts without making slow failures
look like slow successes. Tail latency beyond 45s would still alert,
which is the intent. If we see tail probes >45s, we'll revisit with a
post-deploy warmup ping.
## Files changed (17)
- 16 `showcase/packages/*/src/app/api/smoke/route.ts` (13 with both
abort + maxDuration bumped, 3 langgraph-* with only maxDuration bumped,
ms-agent-dotnet untouched)
- `showcase/scripts/create-integration/index.ts` template (new packages
inherit the new budget)
## Test plan
- [ ] Deploy to Railway, observe agno `/api/smoke` no longer 502s on
cold start
- [ ] Confirm other showcase services still smoke-pass
- [ ] If cold path drifts beyond 45s for any service, add a post-deploy
warmup step (deferred)
Do not merge yet — pending user review.
Cold-start on showcase packages with heavy Python agents (agno in particular)
consistently lands just over the 25s budget, producing 502s on first probe
with latency_ms: 25001 and stage: "timeout". Raise the upstream
AbortSignal.timeout on /api/smoke from 25s to 45s across all 16 showcase
packages that exercise the full agent round-trip, and bump Next.js
maxDuration from 30s to 60s so the route can actually run that long.
Also bumps the create-integration template so new packages inherit the
new budget.
Tail-latency beyond 45s will still alert — which is the intent.
Evidence: agno 502 alerts tonight at 18:40 PDT and 19:42 PDT, both with
latency_ms: 25001, stage: "timeout" on /api/smoke. /api/health already
200 on both probes — pure cold-start boundary issue.
The Deployed Starters smoke test in integration-smoke.spec.ts uses
checkHealth() with its default paths list ["/api/health", "/health"].
All deployed starters serve their health endpoint at /api/health
(standard Next.js convention); /health returns a Next.js 404 HTML
page. When /api/health returns a legitimate 5xx (e.g. 503
"agent degraded"), the fallback silently probes /health and the
resulting 404 becomes the reported failure — so the Slack alert
reads "status=404 path=/health" and implies a route-missing bug,
hiding the real upstream 503.
Shape A fix: change the default to ["/api/health"] only. No
fallback. The reported status and body are now the actual
/api/health response.
starter-smoke.spec.ts passes an explicit paths argument and is
unaffected. The integration backend L1 health check also targets
/api/health on all deployed backends (verified).
Evidence runs (all same failure mode):
- https://github.com/CopilotKit/CopilotKit/actions/runs/24619147643
- https://github.com/CopilotKit/CopilotKit/actions/runs/24618944101
- https://github.com/CopilotKit/CopilotKit/actions/runs/24617315522
ms-agent-dotnet starter was hand-patched with bin/, obj/, *.user, *.suo
but the template wasn't updated — drift-check kept flagging it on regen.
Add the entries to the template so all 17 starters pick them up uniformly
(harmless for non-.NET starters).
R5 added BoundedToolCallingManagerConfig, OpenAiApiKeyValidator, and
WebClientConfig to showcase/packages/spring-ai/agent/ plus updated
application.properties and pom.xml. Starters were out of sync so
drift-check failed. Ran generate-starters.ts to sync.
Also add bin/ and obj/ .NET build artifacts to ms-agent-dotnet starter
.gitignore so generator output stays clean.
- Narrow the mid-stream except around agent.llm_response_async from a
broad Exception to (openai.APIError, httpx.HTTPError,
asyncio.TimeoutError, pydantic.ValidationError). Programmer bugs
(AttributeError, NameError, TypeError) now propagate instead of
being sanitized as 'Agent run failed: <ClassName>' — keeps real
bugs visible in server logs and stack traces, where they belong.
Expected runtime errors still get the sanitized TEXT_MESSAGE +
RUN_FINISHED treatment so the UI never hangs mid-stream.
- Update _parse_tool_args call-site comment in handle_run: it used
to say 'returns None' when args could not be parsed; the function
now returns ParsedArgs with status in {ok, empty, malformed}, and
the caller already pattern-matches on .usable.
- Fix stale _try_parse_tool docstring that referenced the old
agent.agent_response() API; the adapter has used
llm_response_async() for a while now.
Tests: add test_llm_response_async_programmer_bug_propagates — mocks
llm_response_async to raise AttributeError('typo') and asserts the
exception propagates out of the SSE generator (not sanitized).
Update the existing mid-stream-failure test to raise
asyncio.TimeoutError (a covered exception) instead of RuntimeError
so it continues to exercise the sanitization path.
Add openai.OpenAIError (SDK base class) to the generate_a2ui except
tuple so config-time failures — e.g. OpenAI() construction with unset
or malformed OPENAI_API_KEY — become a structured a2ui_llm_error dict
instead of bubbling through the ADK tool machinery as an uncaught
exception. APIError and friends are subclasses of OpenAIError; the
constructor raises OpenAIError directly, which the narrowed tuple
previously missed. Parity fix with the strands sibling (R4).
Also:
- _A2uiError docstring now acknowledges 3-way sync with strands AND
langroid (all three carry the same TypedDict).
- search_flights docstring relaxed from 'exactly 2 flights' to '2-3
flights' so google-adk is consistent with ms-agent-dotnet (which
returns 3).
- Added test exercising openai.OpenAIError from the OpenAI()
constructor path; asserts a2ui_llm_error with full error shape.
- entrypoint.sh: bounded 10s SIGTERM grace window + SIGKILL fallback
when terminating the surviving sibling; prevents indefinite hang past
the platform SIGKILL grace period and preserves the structured
who-died log line.
- BoundedToolCallingManagerConfig: make the weakKeys() identity-vs-equals
footgun explicit in the comment; add JUnit concurrent-contention tests
(ExecutorService + CountDownLatch) asserting the cap invariant holds
under N-thread race on the same ChatOptions — no two threads exceed
cap-1 and both increment past N.
- WebClientConfig: switch the @Bean-time keepalive check from log-and-
proceed to fail-fast IllegalStateException unless COPILOTKIT_ALLOW_KEEPALIVE
opt-in is present (matches OpenAiApiKeyValidator philosophy). Narrow
isTruthy to 1/true (drop yes) since Spring/Java canonicalize on
true/false. Update tests to reflect the narrowed vocabulary and
fail-fast throw.
- application.properties: collapse the 8-line rationale block duplicating
OpenAiApiKeyValidator javadoc to a one-line pointer (drift risk).
- main.py: before_model_modifier no longer chops user suffix when the
prefix end_marker is missing (mangled/drifted prefix). Leaving
original_text untouched means worst-case is one duplicated signature
(non-stacking) instead of silent data loss.
- main.py: simple_after_model_modifier now logs llm_response.error_message
at WARNING with agent name before returning. Previously Gemini
quota/safety/context-overflow errors were silently swallowed.
- main.py: _A2uiError TypedDict now carries a sync-pointer comment to
the identical TypedDict in strands/src/agents/agent.py.
- entrypoint.sh: added explicit survivor cleanup with bounded 5s grace
window, mirroring the spring-ai pattern. Also escalated the
'both still running' wait -n race branch to ERROR + exit 1 so it
cannot silently mask the real child death.
- tests: added before_model_modifier suffix-preservation test and
simple_after_model_modifier error_message warning log tests.
Updated the test_generate_a2ui docstring to reflect the getattr()
guard pattern rather than the old AttributeError handler wording.
- _try_parse_tool function_call path now accepts (str, bytes, bytearray)
for arguments, matching _parse_tool_args. Real OpenAI/httpx stacks can
deliver bytes; the old narrow isinstance(args, str) guard fell through
to tool_cls(**args) with raw bytes and silently TypeError'd.
- Wrap agent.llm_response_async in try/except inside event_stream. On
failure after RUN_STARTED, emit a sanitized TEXT_MESSAGE triple and
RUN_FINISHED so the frontend never hangs waiting for stream closure.
- Wrap request.json() and RunAgentInput(**body) with explicit handlers
that return structured JSONResponse errors carrying an errorId for
client-side correlation (400 for bad JSON, 422 for schema mismatch).
- Defensive guard on run_input.messages: skip entries whose role or
content are not strings, with a warning. Prevents f-string coercion
of None/complex values into garbage prompts.
- Wrap the content-JSON-path asyncio.to_thread in try/except that emits
a sanitized error + RUN_FINISHED before re-raising, so scheduler-level
failures do not kill the SSE generator mid-stream.
- Duplicate-tool RuntimeError now includes fully qualified
module.Class identities for each colliding ToolMessage, instead of
just bare request names.
Tests: 30/30 pass (added 5 new tests covering bytes args, mid-stream
failure, malformed JSON body, invalid RunAgentInput, non-string
role/content, collision error identity).
entrypoint.sh:
- Disable errexit before wait -n so the diagnostic kill -0 probes and
which-died log line run regardless of child exit code.
- Kill the surviving sibling after one child dies to avoid orphan-reparent.
BoundedToolCallingManagerConfig.java:
- Use Caffeine.newBuilder().weakKeys() to force identity comparison on
cache keys. Without it, two concurrent turns with equal-but-distinct
ChatOptions references would collapse onto one counter and the cap
would fire early on the second turn.
- Narrow counter-cleanup catch from Throwable to Exception so we do not
run arbitrary cleanup code against a JVM in an undefined state
(OutOfMemoryError, StackOverflowError). Errors now unwind unhandled.
WebClientConfig.java:
- Accept 1 / true / yes (case-insensitive, trimmed) as opt-in values
for COPILOTKIT_ALLOW_KEEPALIVE instead of only the literal '1'.
- Add defensive runtime check in the @Bean method that logs ERROR if
jdk.httpclient.keepalive.timeout is not '0' at bean-construction
time — catches the case where the JVM arg path is dropped AND this
class loads after an HttpClient has already been constructed.
OpenAiApiKeyValidator.java / application.properties:
- Tighten comments to explicitly note that Spring treats
'?OPENAI_API_KEY must be set' as the DEFAULT VALUE (non-blank, would
pass a naive blank check) when the env var is unset.
Tests:
- Add equal-but-not-same ChatOptions counter-independence test.
- Replace Throwable-cleanup test with Error-propagation test.
- Add tests for true/yes/case-insensitive opt-in.
- Add tests for isTruthy helper.
- Add @Bean defensive-ERROR-log test.
- Result: 33 tests, all green (up from 26).
Verified:
- mvn test: 33/33 green
- docker build + smoke (integration-smoke -g spring-ai, LOCAL_PORTS=1): passing
Changes to src/agents/agui_adapter.py:
- Extract _execute_backend_tool / _run_backend_tool helpers so the
oai_tool_calls path and the content-JSON path share one sanitization
contract. The _try_parse_tool backend-tool path no longer runs
.handle() barefoot — it goes through _execute_backend_tool which
catches the narrowed (pydantic.ValidationError, ValueError), logs
the traceback server-side, and returns a sanitized JSON error payload.
- Narrow tool-execution except from
(RuntimeError, ValueError, TypeError, KeyError, pydantic.ValidationError)
to (pydantic.ValidationError, ValueError). RuntimeError/KeyError from
unrelated libraries or real config drift now propagate as real bugs
rather than being misclassified as user-facing tool failures.
- Replace dict | None three-way return of _parse_tool_args with a
ParsedArgs dataclass (args + status: 'ok' / 'empty' / 'malformed' +
.usable helper). Call sites pattern-match on status explicitly.
- Treat empty-string raw_args (legacy function_call path) as 'malformed'
instead of 'ok with {}' — firing a tool with empty args produces a
meaningless UI card, same rationale as the oai path.
- Drop TypeError from the json.loads except in _try_parse_tool (it
signals a programmer bug, not plain text). Add an
isinstance(content, (str, bytes, bytearray)) guard that logs a
WARNING rather than silencing the case.
- Drop pydantic.errors.PydanticUserError from tool-instantiation except
lists — it signals malformed model *definitions*, not runtime data
errors, and must surface loudly at startup.
- Freeze _TOOL_BY_NAME behind types.MappingProxyType so post-import
mutation raises TypeError (e.g. accidental monkeypatch / shadowing).
Changes to tests/python/conftest.py:
- Tighten the docstring cross-reference to src/agent_server.py (the
real entry point confirmed to import agents.agui_adapter identically).
Changes to tests/python/test_agui_adapter.py:
- Update helper tests to assert on ParsedArgs.status / .usable / .args.
- Flip empty-string expectation to 'malformed' per the new contract.
- Add: bytes-input parse, unknown-type 'empty' status, MappingProxyType
frozen-ness of _TOOL_BY_NAME, unified-helper happy path, helper-level
sanitization + traceback capture, helper propagates non-narrowed
RuntimeError, _try_parse_tool non-str/bytes guard warns and returns
None, and str(exc) never appears anywhere in the SSE stream bytes.
- Switch sanitized-error test's raise to ValueError (narrowed except
covers it) — previously RuntimeError leaked through by accident.
24/24 tests pass (16 existing + 8 new).