Commit Graph

4085 Commits

Author SHA1 Message Date
Atai Barkai 18ed9dbdbd chore(showcase/scripts): port bundle regions + parity dual-location
Port the 4084 scripts-layer enhancements so 4085's showcase toolchain
matches the new feature shape:

- lib/manifest.ts: ManifestDemo gains optional `command` field; parser
  accepts + validates it (non-empty string, frozen).
- bundle-demo-content.ts: inline `@region[name]` / `@endregion[name]`
  comment-marker extraction; informational-only demos (no route, e.g.
  cli-start) are skipped; markers stripped from bundled content;
  regions: { file, startLine, endLine, code, language } emitted per
  demo. External-highlight-file merging (4085-specific) preserved, so
  backend agents under src/agents/*.py still flow into the bundle.
- validate-parity.ts: accepts demos at BOTH demos/<cell>/ (4084 layout)
  and src/app/demos/<cell>/ (4085 layout); informational demos
  (command field) are excluded from the parity audit.
- tests: bundle-demo-content.test.ts expectedDemos updated for the
  shared-state rename; generate-registry.test.ts feature count 25→32;
  validate-parity.test.ts missing-demo-dir message updated to match
  the new dual-location wording.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-19 11:24:55 -07:00
Atai Barkai fb062d3750 chore(showcase): port shared schema + feature-registry + constraints
Port the 4084 internal-showcase-followups state of the shared showcase
schema to 4085 (no Docker restructure):

- feature-registry.json: new Dev Ex category, category reorder
  (Dev Ex → Chat & UI → Platform → Controlled → Declarative → Open →
  Operational → Interactivity → Agent State → Multi-Agent → BYOC);
  +8 new features (beautiful-chat, cli-start, interrupt-headless,
  reasoning-default-render, tool-rendering-reasoning-chain,
  open-gen-ui-advanced, declarative-gen-ui-hardcoded,
  readonly-state-agent-context); renames (shared-state-write →
  shared-state-read-write, shared-state-agent-readonly →
  readonly-state-agent-context); removals (shared-state-read,
  state-rendering, a2ui-*, shared-state-io, shared-state, readables,
  a2a-chat, deep-agents, vnext-chat, byoc-a2ui, byoc-tambo);
  gen-ui-agent → kind: testing.
- constraints.yaml: allowlists updated to reference the new IDs.
- manifest.schema.json: `route` is optional; `command` field added for
  informational demos (e.g. cli-start).
- All 17 column manifests: shared-state-read removed,
  shared-state-write renamed to shared-state-read-write.
- langgraph-python manifest: full 4084 features list + demos entries
  with 4085-appropriate highlight paths (src/agents/*.py +
  src/app/demos/<cell>/* rather than backend/ + frontend/).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-19 11:24:33 -07:00
Atai Barkai beb42019d4 docs(showcase/shell): port ~70 MDX pages from 4084 into shell
Port the MDX content delta between the 4085 base branch
(atai/2026-04-18/feature-port-no-docker-restructure) and the 4084
follow-ups branch (atai/2026-04-17/internal-showcase-followups):

- built-in-agent tree (ag-ui, coding-agents, copilot-runtime,
  frontend-tools, inspector, prebuilt-components, programmatic-control,
  custom-look-and-feel/{headless-ui,slots}, generative-ui/{a2ui,
  tool-rendering,your-components/{display-only,interactive}},
  premium/{headless-ui,observability,overview},
  troubleshooting/{common-issues,error-debugging,migrate-to-*})
- top-level feature docs (agentic-chat-ui, frontend-tools, headless,
  human-in-the-loop, inspector, shared-state, programmatic-control,
  ag-ui-middleware, coding-agents, coding-agent-setup)
- custom-look-and-feel (css, reasoning-messages, slots)
- generative-ui (a2ui, a2ui/dynamic-schema, a2ui/fixed-schema,
  mcp-apps, open-generative-ui, reasoning, tool-based, tool-rendering)
- human-in-the-loop (headless, useInterrupt)
- prebuilt-components (index, popup, sidebar) — replaces the single
  prebuilt-components.mdx
- shared-state (agent-readonly, streaming)
- troubleshooting (common-issues, debug-mode, error-debugging,
  migrate-to-*, observability-connectors)
- learn (connect-mcp-servers)
- multi-agent/subagents
- integrations/built-in-agent/custom-agent,
  integrations/langgraph/generative-ui/tool-rendering
- meta.json updates across backend, custom-look-and-feel,
  generative-ui, generative-ui/a2ui, prebuilt-components,
  troubleshooting, root docs

These pages use the new Snippet / InlineDemo / FrameworkTabs
components shipped in the preceding infrastructure commit, plus the
FrameworkGuardedContent pattern that only applies when a page has a
defaultCell frontmatter. Framework-agnostic pages (/learn/*,
/troubleshooting/*, /premium/*) keep rendering their body
unconditionally.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-19 11:22:57 -07:00
Atai Barkai c1b595596e feat(showcase/shell): port /docs routing + framework-scoped routes from 4084
Port the full docs routing + rendering infrastructure from the 4084
branch (atai/2026-04-17/internal-showcase-followups):

- New [framework]/[[...slug]] catch-all that validates the first
  segment against the registry and renders framework-scoped docs
  with a "not available for this framework" banner when the page's
  defaultCell is not tagged in the framework's cells.
- FrameworkProvider + selector: framework is strictly URL-derived,
  storedFramework is an advisory localStorage signal. No
  auto-redirect from localStorage.
- /docs/[[...slug]] rewritten as a thin wrapper around DocsPageView
  with a new DocsOverview landing: framework picker + topic cards.
- New lib/docs-render.tsx centralising snippet inlining,
  frontmatter parsing, nav tree + breadcrumb builders.
- New lib/mdx-registry.tsx consolidating the ~700-line component
  shim used by both routes.
- Server <Snippet> component with demo-content regions + syntax
  highlighting via highlight.js.
- Per-feature <InlineDemo> iframe embedder pointing at the
  integration's backend (single-container shape
  http://localhost:3100/demos/<cell>).
- FrameworkTabs, docs-callout, docs-steps, docs-tabs, router-pivot,
  stored-framework-highlight components for the MDX registry.
- Global typography + color pivot in globals.css: Plus Jakarta
  Sans body, Spline Sans Mono code, cool-gray background, purple
  accent.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-19 11:22:28 -07:00
Atai Barkai 218a356256 feat(showcase/shell-internal): port enriched cell chrome + command cell from 4084
- Drop the 5 experimental variant-presentation routes + SingleCell helper;
  the variant layouts no longer match the enriched cell chrome.
- Rename variant-pieces.tsx -> cell-pieces.tsx; expose a shared
  <CellStatus ctx={ctx} /> that renders the docs-og/docs-shell row +
  E2E/Smoke/QA/health badges so both the runnable-demo cell and the
  informational CommandCell share the same bottom section.
- Add CommandCell (client component): takes the full CellContext, renders
  <code> + Copy button for ctx.demo.command, then the same CellStatus row.
- page.tsx branches on ctx.demo.command so the CLI Start row renders a
  copy-pasteable command in place of Demo/Code links while keeping the
  matrix visually consistent.
- registry.ts: Feature gains kind + og_docs_url/shell_docs_url; Demo.route
  is optional and Demo.command is the new informational field. Feature-grid
  passes the full demo onto CellContext.
- status.ts: drop variant helpers / types (no longer used).
- Add docs-status.ts reader + a stub docs-status.json so the enriched
  DocsRow renders without blowing up the build before the schema agent
  lands the generated bundle.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-19 11:22:02 -07:00
Jordan Ritter 98841aaba5 fix(showcase/starters): chown WORKDIR to app user in Dockerfiles
Root cause: `WORKDIR /app` at the top of the runner stage creates /app
owned by root BEFORE the unprivileged `app` user is created. The
`COPY --chown=app:app` lines added in #4092 only reassign the files
being copied, NOT the /app directory itself.

Symptom: starter-mastra CRASHED on Railway with
  EACCES: permission denied, mkdir '/app/.mastra'
when `npx mastra dev` tried to create its cache dir under CWD as user
`app`. Likely co-victim class: any starter whose runtime invokes a Node
CLI that writes to CWD (e.g. langgraph-typescript via @langchain/langgraph-cli dev).

Fix: add `RUN chown app:app /app` immediately before `USER app` in all
four Dockerfile templates (typescript, python, dotnet, java), applied
uniformly to all 17 generated starters via the generator.

Verified locally:
 - docker build + run starter-mastra → container stays up, /app owned
   by app:app, /app/.mastra created successfully, /api/health returns 200.
 - docker build + run starter-langgraph-typescript → container stays up
   (8+ min), write to /app succeeds, /api/health returns 503 degraded
   (expected on dummy OPENAI_API_KEY) — NOT a crash.

All 17 starter Dockerfiles regenerated via
  pnpm run generate-starters
and the existing showcase/scripts vitest suite stays green (1073 tests).
2026-04-19 10:40:12 -07:00
Jordan Ritter de586ececd chore(showcase): regenerate starters post-#4095 rebase
- Regenerate agno starter to pick up agno>=2.5.17 (from #4095)
- Ratchet validate-pins fail-baseline hash to match new FAIL set
  (count unchanged at 110; hash rotates because agno Dojo/showcase
  pair now reflects the SDK upgrade)
2026-04-19 09:00:11 -07:00
Jordan Ritter 689a4b8da7 fix(showcase/strands): teach _FakeFastAPI stub to accept add_middleware
agent_server.py now calls app.add_middleware(HealthMiddleware) to serve
/health above the routing layer (restores starter health accuracy after
create_strands_app's catch-all at "/" shadowed @app.get("/health")).

Both _FakeFastAPI stubs — in tests/python/conftest.py and
tests/python/test_instrumentor_patch.py::test_agent_server_module_installs_patch —
only covered the decorator surface (get/post/put/delete/patch). When the
test execs agent_server.py verbatim under the stub, add_middleware raises
AttributeError and fails Python unit tests (3.12). The 3.10 matrix skips
strands (ag_ui_strands requires >=3.12) so the regression only showed
under 3.12.

Mirror the stub's API with a no-op add_middleware so the module loads
cleanly; the patched ThreadingInstrumentor assertion is still the
invariant under test.
2026-04-19 08:57:44 -07:00
github-actions[bot] e45d86b8f3 style: auto-fix formatting 2026-04-19 08:57:44 -07:00
Jordan Ritter dfd8346bc1 fix(showcase/starters/langgraph-fastapi): correct src.agents.tools.* imports
generate-starters.ts rewrites relative imports to absolute for langgraph
starters because langgraph_cli loads modules standalone rather than as
packages. The rewrite was flat:

  from .X import ...  ->  from <agentDir>.X import ...

For a file at <agentDir>/tools/get_weather.py, `from .types import ...`
resolves to `tools.types` — the CURRENT package — not `<agentDir>.types`.
The flat rewrite dropped the `tools` segment and produced:

  ModuleNotFoundError: No module named 'src.agents.types'

at startup. The agent crashed during module import; the entrypoint
pipe swallowed the traceback (see previous commit for the pipe bug);
the 2-3s sleep guard happened to fire while `sed` was still alive; and
Railway's /api/health probe reported `agent: "error"`.

Make the rewrite subdir-aware: compute the file's containing Python
package from its relative path under agentDest, and prepend that to the
relative import target. So `from .types import ...` inside
`src/agents/tools/get_weather.py` becomes
`from src.agents.tools.types import ...`.

langgraph-python has the same broken imports in its tools/__init__.py
but doesn't crash at runtime because main.py doesn't import from tools
(dead path). Regenerating fixes the dead code too.
2026-04-19 08:57:44 -07:00
Jordan Ritter de815ae2d5 fix(showcase/starters): capture real agent PID in entrypoint.sh + unbuffer stderr
The entrypoint.sh template used `cmd 2>&1 | sed 's/^/[agent] /' &`
followed by `AGENT_PID=\$!`. After a pipeline, `\$!` points to the LAST
command in the pipe (the `sed` process), not the agent. Every subsequent
`kill -0 \$AGENT_PID` and `wait -n \$AGENT_PID` was therefore monitoring
`sed`, which stays alive until its stdin closes — long after the agent has
crashed. Railway restarts the container mid-loop; the health probe sees
`{status: "degraded", agent: "down"}` for a few seconds during each cycle.

A second, compounding bug: `sed` buffers by default, and Python agents
buffer their own stdout, so a stack trace emitted during module import
could sit in userspace memory until the pipe closed — by which point the
log was discarded and the real cause of the crash was lost.

Fix both by switching to bash process substitution:
  cmd &> >(awk '{print "[agent] " \$0; fflush()}') &
  AGENT_PID=\$!
Process substitution does not create a pipeline, so `\$!` remains the
agent's PID. `awk` with `fflush()` flushes each prefixed line to the
container log immediately. Also export `PYTHONUNBUFFERED=1` at the
entrypoint level so Python-based agents don't buffer before awk.

Applies to all 17 starters (python, langgraph-python, langgraph-fastapi,
langgraph-typescript, mastra, typescript, java/spring-ai, csharp/
ms-agent-dotnet). Done once in generate-starters.ts + the template +
regenerated entrypoint.sh files.
2026-04-19 08:57:44 -07:00
Jordan Ritter 20efc40960 fix(showcase/starters): serve /health via middleware to bypass FastAPI mount shadow
Python FastAPI showcase starters were registering an @app.get("/health")
handler BEFORE app.mount("/", ...) (or the adapter's equivalent catch-all).
Starlette's Mount at "/" is a prefix match that swallows every subsequent
request including /health, so the decorated handler never fired and the
Next.js /api/health probe saw a non-JSON agent response and flagged the
starter as "degraded".

Install a HealthMiddleware (BaseHTTPMiddleware) that short-circuits
GET /health above the routing layer, which runs before any Mount can
claim the path. The agent mount is kept intact so AG-UI traffic at / is
unaffected.

Affects 10 Python agent starters: ag2, agno, claude-sdk-python,
crewai-crews, google-adk, langroid, llamaindex, ms-agent-python,
pydantic-ai, strands.
2026-04-19 08:57:44 -07:00
Jordan Ritter 6092c8f888 fix(showcase/agno): upgrade agno SDK to 2.5.17 (#4095)
## Summary

Bumps the `agno` Python SDK pin from `>=1.7.8` to `>=2.5.17` (latest
stable on PyPI) in `showcase/packages/agno/requirements.txt`.

## Why

The agno 1.x AGUI interface had a known bug where the terminal SSE event
(`RUN_FINISHED` / `TEXT_MESSAGE_END`) was never emitted, causing
consumers that block on stream close (such as `/api/smoke` calling
`await res.text()`) to hang until AbortSignal fired (25-45s).

Agno 2.5.17 emits the terminal `RUN_FINISHED` frame and closes the
stream cleanly, even when the underlying LLM call fails (dummy API key
-> 401).

## Version

| | Before | After |
|---|---|---|
| Pin | `agno>=1.7.8` | `agno>=2.5.17` |
| Installed | 1.7.x | 2.5.17 |

## Breaking changes encountered

None in this package. All existing imports continue to resolve without
change:

- `from agno.os import AgentOS`
- `from agno.os.interfaces.agui import AGUI`
- `from agno.agent.agent import Agent`
- `from agno.models.openai import OpenAIChat`
- `from agno.tools import tool`

No API adaptation was required.

## Relationship to PR #4093

PR #4093 works around the 1.x SSE bug in the consumer (`/api/smoke`) by
reading the stream incrementally and bailing on `TEXT_MESSAGE_CONTENT
"OK"`. With this upstream fix, that workaround becomes defense-in-depth
rather than load-bearing.

**Recommendation:** leave #4093 in place after this merges. It's cheap,
it's already written, and it protects against future agno regressions in
stream termination.

## Local verification evidence

All tests ran against a local `docker build -t agno-test .` of this
package (`agno==2.5.17`, Python 3.12.13).

### Imports
```
$ docker exec agno-test python -c "import agno; print(agno.__version__)"
2.5.17

$ docker exec agno-test python -c "from agno.os.interfaces.agui import AGUI; ..."
all imports OK
```

### /api/smoke timing (three consecutive runs)
```
run 1: time=0.256s code=200
run 2: time=0.157s code=200
run 3: time=0.151s code=200

{"status":"ok","integration":"agno","latency_ms":12,"timestamp":"2026-04-19T15:03:36.057Z"}
```
Previously: 25-45s hang ending in `AbortError`.

### Raw SSE stream inspection (dummy OpenAI key -> 401)
```
data: {"type":"RUN_STARTED","threadId":"test-thread","runId":"test-run",...}

data: {"type":"RUN_FINISHED","threadId":"test-thread","runId":"test-run"}
```
Stream closes immediately after `RUN_FINISHED`. Terminal event is
present.

### Demos (homepage + 9 demo routes)
```
/                          200   0.008s
/demos/agentic-chat        200   0.010s
/demos/hitl                200   0.005s
/demos/tool-rendering      200   0.005s
/demos/gen-ui-tool-based   200   0.006s
/demos/subagents           200   0.004s
/demos/shared-state-read   200   0.005s
/demos/shared-state-write  200   0.006s
/demos/shared-state-streaming 200 0.005s
/demos/gen-ui-agent        200   0.005s
```

### Agent object
```
agent loaded: Agent
tool count: 8
model: gpt-4o
```

## Test plan

- [x] `docker build` succeeds with new pin
- [x] `import agno` reports 2.5.17
- [x] All existing `from agno...` imports resolve unchanged
- [x] `/api/health` returns 200
- [x] `/api/smoke` returns 200 in <1s (was 25-45s hang)
- [x] Raw SSE stream contains `RUN_FINISHED` terminal event
- [x] Homepage + all 9 demo routes return 200
- [x] Agent object loads with 8 tools
- [ ] CI green on PR
2026-04-19 08:41:20 -07:00
github-actions[bot] e3b9defb45 style: auto-fix formatting 2026-04-19 15:08:32 +00:00
Jordan Ritter 0aa95d7e15 chore(showcase): ratchet validate-pins baseline 109->110 for google-genai pin 2026-04-19 08:07:02 -07:00
github-actions[bot] f5b69285c3 style: auto-fix formatting 2026-04-19 15:05:54 +00:00
Jordan Ritter 420254cfdc fix(showcase/agno): upgrade agno SDK to 2.5.17 to fix SSE stream termination
Bumps agno from >=1.7.8 to >=2.5.17 (latest stable on PyPI).

The agno 1.x AGUI interface had a known bug where the terminal SSE
event (RUN_FINISHED) was never emitted, causing consumers that block
on stream close (such as /api/smoke reading via res.text()) to hang
until AbortSignal fired at ~25-45s.

Verified locally against a docker build of this package that 2.5.17
emits the terminal RUN_FINISHED frame and closes the stream cleanly,
even when the underlying LLM call fails (dummy API key -> 401).

Measured /api/smoke round-trip: ~150-400ms (was 25-45s timeout).

All existing imports (agno.os.AgentOS, agno.os.interfaces.agui.AGUI,
agno.agent.agent.Agent, agno.models.openai.OpenAIChat, agno.tools.tool)
continue to resolve without changes — no API adaptation required.
2026-04-19 08:03:56 -07:00
Jordan Ritter e851ea70e6 fix(showcase/starters): regenerate google-adk starter after Gemini rewrite 2026-04-19 07:57:13 -07:00
Jordan Ritter 5bc22a396d fix(showcase/google-adk): use Gemini for A2UI planner instead of OpenAI
The google-adk package's primary agent uses Gemini via google-adk's
LlmAgent, but the secondary `generate_a2ui` tool was calling OpenAI's
gpt-4.1 through openai.OpenAI(). Cross-provider dependency with no
architectural justification — the A2UI planner's function-calling +
forced tool-choice pattern works equally well via google.genai.

Rewrite to use google.genai.Client.models.generate_content with a
ToolConfig forcing the render_a2ui function call (mode="ANY" +
allowed_function_names). Drop the openai + httpx direct dependency
entirely; google-genai is already pulled in transitively via google-adk
and is now pinned explicitly at >=0.8.0.

Replaces short-term hotfix PR #4090 (which pinned openai in
requirements.txt after post-merge Railway deploy of #4083 crashed
with ModuleNotFoundError).

Changes:
- src/agents/main.py: `_get_openai_client` → `_get_genai_client`
  (lru_cached google.genai.Client). Rewrote generate_a2ui to build a
  `types.Tool(function_declarations=[FunctionDeclaration])`, force the
  call via `types.ToolConfig(function_calling_config=...)` and read
  the response from `candidates[0].content.parts[*].function_call.args`.
  Narrowed except clause to `genai_errors.APIError`/`ClientError`/
  `ServerError` + `ValueError` (for construction-time failures).
  A2UI_MODEL env var overrides the default gemini-2.5-flash model.
- entrypoint.sh: `OPENAI_API_KEY` guard → `GOOGLE_API_KEY` guard;
  `REQUIRE_OPENAI_API_KEY` → `REQUIRE_GOOGLE_API_KEY`.
- requirements.txt: explicit `google-genai>=0.8.0` pin; no openai, no
  httpx.
- src/app/api/health|debug/route.ts: report GOOGLE_API_KEY status, not
  OpenAI/Anthropic/LangSmith (this is a Gemini-only package).
- playwright.config.ts: inject GOOGLE_API_KEY to the dev server env, not
  OPENAI_API_KEY / OPENAI_BASE_URL.
- tests/python/test_generate_a2ui.py: rewrote 19 tests to mock
  google.genai.Client in place of openai.OpenAI. Preserves every error
  branch (a2ui_llm_error, a2ui_empty_response, a2ui_no_tool_call,
  a2ui_invalid_arguments) plus happy path. Added A2UI_MODEL override
  coverage.
- tests/python/test_entrypoint_env_guards.py: migrated to exercise the
  GOOGLE_API_KEY / REQUIRE_GOOGLE_API_KEY branch.

Verification:
- Unit tests: 53 passed locally against Python 3.12 (google-adk + google-
  genai + all package deps installed).
- Docker build: clean (showcase-google-adk:local) via
  scripts/dev-local.sh.
- Container startup: `docker compose up -d google-adk` launches cleanly
  (no ModuleNotFoundError); uvicorn reports healthy; entrypoint emits
  the new WARN when GOOGLE_API_KEY is unset.
- L1+L2 smoke against local compose: both pass for google-adk.
- `import agents.main` inside the container no longer pulls in any
  `openai.*` module.
2026-04-19 07:55:27 -07:00
Jordan Ritter a4be366e95 perf(showcase): replace recursive chown in starter Dockerfiles with COPY --chown (#4092)
## Summary

Replaces `RUN chown -R app:app /app` with `COPY --chown=app:app` across
all 17 starter Dockerfiles + 4 shared templates.

Every starter ended with a recursive chown over `/app`, which walks ~50k
files (Next.js `node_modules` dominates) to fix ownership after the
fact. Under the 23-way Depot runner fan-out used by
`.github/workflows/deploy-showcase-services.yml`, that step consistently
ran 5+ minutes under I/O contention — busting the **15-minute GH Actions
job budget** before images could finish pushing.

Failed run this fixes:
https://github.com/CopilotKit/CopilotKit/actions/runs/24621469277 (22/23
showcase services cancelled mid-push).

## What changed

- Create the `app` user right after `WORKDIR /app` in every runner stage
so `--chown=app:app` resolves by name.
- Add `--chown=app:app` to every runner-stage COPY (including
multi-stage `COPY --from=frontend` and `COPY --from=<agent-builder>`).
- Drop the trailing `RUN chown -R app:app /app` (or the tail of the
compound RUN in the TS starters).
- For langgraph starters, fold `chown app:app /app/.langgraph_api` into
the same RUN as the `mkdir`, so `langgraph_cli`'s scratch dir remains
writable by the runtime user.
- Update `showcase/scripts/generate-starters.ts` so the shared templates
(`Dockerfile.python/typescript/dotnet/java`) emit the new pattern.
Starters are regenerated from templates.

## Local timing (Docker Desktop, single starter, no contention)

| Starter | Pre-fix | Post-fix | `chown -R` step |
| --- | --- | --- | --- |
| ag2 (Python) | 3:03 | 1:39 | 50.1s → removed |
| langgraph-fastapi | n/a | 1:41 | replaced with 0.2s targeted chown on
`.langgraph_api` |
| mastra (TS) | n/a | 2:16 | removed |

The real-world win on the 23-way Depot fan-out is substantially larger
than the local 50s baseline — recursive chown degrades super-linearly
with concurrent I/O pressure, which is exactly what the 5+ min Depot
step demonstrated.

## Smoke tests

- ag2 image: container starts clean as `app` user, all files under
`/app` owned by `app:app`, `/api/health` returns 200.
- langgraph-fastapi image: `/app/.langgraph_api` exists and is owned by
`app:app`.
- mastra image: builds cleanly, ownership correct.

## Test plan

- [ ] CI green (existing showcase starter smoke suite covers startup)
- [ ] Watch the next showcase deploy workflow run — expect jobs to
finish well under 15m

## Not touched

No workflow files, no `examples/` Dockerfiles (none matched the problem
pattern), no `chmod -R` offenders (none found).

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-04-19 07:47:17 -07:00
Jordan Ritter 2c73a3480b ci: close showcase deploy automation gap for starters (#4082)
## Summary

Two related defects in the showcase deploy pipeline let stale images sit
live on Railway while Slack stayed green. This PR fixes both.

### Defect 1 — Drift detector skipped all starter services

`.github/workflows/showcase_smoke-monitor.yml` listed only 19
**package** slugs in its `SERVICES=(...)` array (ag2, mastra,
llamaindex, ...). Zero **starter** slugs. As a result:

- GHCR `showcase-starter-<svc>` tags were never checked for drift.
- `gh workflow run showcase_deploy.yml -f service=starter-*` was never
auto-dispatched.
- Starter services could run with weeks-old images and no alert would
fire.

`showcase_deploy.yml` already supports `starter-*` dispatch names and
already calls `serviceInstanceRedeploy` for any service with a
`railway_id`, so no change is required there. The fix is extending
`SERVICES=(...)` to include all 17 starter slugs via a
sparse-checkout-driven filesystem enumeration (no more literal
duplication between workflow and `showcase/starters/`).

### Defect 2 — Silent deploy failures reported green

`.github/workflows/showcase_deploy.yml` emitted `::warning::` and exited
0 when a service never returned 200 on its health path within 360s. The
legacy justification (`# Don't fail — sleep-on-idle services take time
to wake`) no longer applies: Railway is on the Pro tier with no
sleep-on-idle, so a 6-minute failure to become healthy is a real
failure. Changed to `::error::` + `exit 1`.

## Round 2 fixes

Round 2 CR raised six findings against the original smoke-monitor +
validator changes. All fixed in this PR:

- **BLOCKING 1/2 — smoke-monitor guard.** Replaced the magic `-eq 19`
sentinel with `grep -c '^starter-'` so adds/removes to the literal
non-starter list can't silently disable the guard. Added `shopt -s
nullglob` around the `showcase/starters/*/` loop so an empty starters
tree no longer corrupts `SERVICES` with a `starter-*` literal.
- **BLOCKING 3 — GHCR stderr isolation.** Dropped `2>&1` on the `gh api
-i` call; captured stderr to a temp file and surfaced it only when `gh`
returns a non-zero RC with no HTTP status. Auth / rate-limit / network
noise can no longer splice into the HTTP header block and poison
`HTTP_STATUS` / `API_BODY` parsing.
- **BLOCKING 4 — validator tests.** Added
`showcase/scripts/__tests__/validate-workflow-starters.test.ts` (12
specs): happy path, missing-from-options-only, missing-from-matrix-only,
missing-from-both, empty starters dir (exit 3), template/ excluded,
substring-spoof (starter-ag2 vs starter-ag2-extended), missing workflow
file (exit 3). Also extended `VALIDATE_WORKFLOW_STARTERS_REPO_ROOT` to
re-home the starters dir for testability.
- **MEDIUM 1 — YAML parsing.** Replaced the fragile regex-over-YAML
options scanner with a real `yaml.parse()` + typed navigation down
`on.workflow_dispatch.inputs.service.options`. ALL_SERVICES stays
regex-scanned (embedded JSON in a bash heredoc, with `${{ ... }}`
interpolations that aren't valid JSON pre-execution), but the
surrounding step is now located via YAML.
- **MEDIUM 2 — Slack list truncation.** Replaced `cut -c1-200` with a
`truncate_csv` helper that drops whole comma-separated entries until
under budget and appends `…` when truncated. No more
`starter-claude-sdk-pyth` mid-slug corruption.
- **MEDIUM 3 — template/ exclusion cross-references.** Both the TS
validator's `EXCLUDED_DIRS` and `showcase_smoke-monitor.yml`'s `[
"$slug" = "template" ] && continue` now carry `# keep in sync with ...`
comments pointing at each other.
- **NIT 2 — entry-check simplification.** Dropped the
belt-and-suspenders `import.meta.url === \`file://${argv[1]}\`` branch;
kept only the canonical `fileURLToPath(import.meta.url)` form.
- **NIT 3 — jq pipeline collapse.** Single-pass `.jobs[]? | select |
"\(...)"` replaces the three-pass `map | map | .[]` chain in the notify
step.

## Files changed

- `.github/workflows/showcase_deploy.yml` — warning → error + exit 1 on
unhealthy deploy; `truncate_csv` replaces `cut -c1-200` (3 sites);
single-pass jq pipeline in notify step.
- `.github/workflows/showcase_smoke-monitor.yml` — filesystem-driven
`SERVICES=(...)`, starter-count guard, `nullglob` loop, stderr-isolated
`gh api` call, cross-reference comment.
- `.github/workflows/showcase_validate.yml` — wires
`validate-workflow-starters` into CI.
- `showcase/scripts/validate-workflow-starters.ts` — YAML-aware presence
checks; env-var override homes both starters dir and workflow path.
- `showcase/scripts/tsconfig.json` — scripts-local tsconfig for LSP type
resolution.
- `showcase/scripts/__tests__/validate-workflow-starters.test.ts` — 12
specs covering the full matrix of drift scenarios.

## Test plan

- [ ] Next scheduled `showcase_smoke-monitor` run includes starter
services in its drift scan.
- [ ] A deliberately-unhealthy deploy (simulate by pointing health_path
at a 404) fails the job and fires the Slack alert.
- [ ] `showcase_validate` CI job runs `validate-workflow-starters` and
`npx vitest run scripts/__tests__/validate-workflow-starters.test.ts`
green.
2026-04-19 07:43:42 -07:00
Jordan Ritter c591ad71b1 perf(showcase): replace recursive chown with COPY --chown in starter Dockerfiles
Every starter Dockerfile ended with `RUN chown -R app:app /app`, which
recursively walks the full `/app` tree (Next.js `node_modules` dominates
— ~50k files) to fix ownership after the fact. On the 23-way Depot
runner fan-out used by .github/workflows/deploy-showcase-services.yml
this step consistently took 5+ minutes under I/O contention, busting
the 15-minute GH Actions job budget before the image could even finish
pushing (run 24621469277 — 22/23 services cancelled).

The standard Docker idiom is `COPY --chown=<user>:<group>` which applies
ownership during the copy step itself — no extra layer, no full-tree
traversal, zero runtime cost.

Changes:
  - Create the `app` user right after `WORKDIR /app` in every runner
    stage so `--chown=app:app` resolves by name.
  - Add `--chown=app:app` to every COPY instruction that lands files
    under `/app` in the runner stage (including `COPY --from=frontend`
    and `COPY --from=<agent-builder>` multi-stage copies).
  - Drop the trailing `RUN chown -R app:app /app` (or the tail of the
    compound RUN in the TS starters).
  - For langgraph starters, fold `chown app:app /app/.langgraph_api`
    into the mkdir RUN so langgraph_cli's scratch dir is still writable
    by the runtime user.
  - Update showcase/scripts/generate-starters.ts so the shared templates
    (Dockerfile.python/typescript/dotnet/java) emit the new pattern and
    the framework-specific COPY lines (`COPY ${dest} ./` for extra files
    and `COPY agent_server.py ./` for Python starters) include `--chown`.

Local verification (Docker Desktop, single starter, no contention):
  ag2 pre-fix:  3:03 total, `RUN chown -R app:app /app` = 50.1s
  ag2 post-fix: 1:39 total, no chown step
  langgraph-fastapi post-fix: 1:41 (langgraph_api chown is 0.2s)
  mastra post-fix: 2:16

Smoke test on ag2 image: container starts clean, all files under /app
owned by app:app, /api/health returns 200.

Under the 23-way Depot fan-out the real-world win is substantially
larger than the local 50s baseline because recursive chown degrades
super-linearly with concurrent I/O pressure.

Failed run that motivated this: https://github.com/CopilotKit/CopilotKit/actions/runs/24621469277
2026-04-19 07:41:55 -07:00
Jordan Ritter da487b7e8c ci: probe only /api/health in starter smoke test (#4088)
## Summary

The Deployed Starters smoke test is falsely reporting `status=404
path=/health` for every starter whose agent process is degraded. The
reported path is misleading — it's just the last path probed before the
test gave up. The real upstream is a 503 on `/api/health`.

## Root cause

`checkHealth()` in `showcase/tests/e2e/helpers.ts` defaulted to probing
`[\"/api/health\", \"/health\"]` in sequence. All deployed starters and
showcase backends serve their health endpoint at `/api/health` (standard
Next.js convention). `/health` returns a Next.js 404 HTML page.

When `/api/health` returns a legitimate 5xx (e.g. 503 `\"agent
degraded\"`), the helper's `res.ok()` check is false, so it falls
through to probe `/health`, gets a 404, and reports that as
`lastResult`. The Slack alert then reads `path=/health` and implies a
route-missing bug, hiding the real 503.

Verified with live probes:
- Healthy starter `/api/health` → `200 {\"status\":\"ok\",...}`
- Degraded starter `/api/health` → `503
{\"status\":\"degraded\",\"agent\":\"error\",...}`
- Either starter `/health` → `404` (Next.js not-found page)

## Fix (Shape A — remove the fallback)

Change the default `paths` argument from `[\"/api/health\",
\"/health\"]` to `[\"/api/health\"]`. No fallback. The reported status
and body are the actual `/api/health` response.

- `starter-smoke.spec.ts` passes an explicit `paths` argument and is
unaffected.
- `integration-smoke.spec.ts` L1 health check targets `/api/health` on
all deployed backends (verified live on langgraph-python and mastra).

## Out of scope

15/17 starters currently return 503 because their agent processes are
failing. That's a separate investigation — this PR is purely about
making the smoke alert accurate so that investigation can proceed
per-service.

## Evidence runs (all same failure mode tonight)

- https://github.com/CopilotKit/CopilotKit/actions/runs/24619147643
- https://github.com/CopilotKit/CopilotKit/actions/runs/24618944101
- https://github.com/CopilotKit/CopilotKit/actions/runs/24617315522

## Test plan

- [ ] Next scheduled smoke run surfaces `status=503 path=/api/health`
for the 15 degraded starters (agent-down state, clearly attributable)
- [ ] Healthy starters (langgraph-python, one other) continue to pass L1
- [ ] `starter-smoke.spec.ts` (local Docker starters) unaffected — still
uses explicit `starter.healthPaths`
2026-04-19 07:23:38 -07:00
Jordan Ritter 35a7e0038a fix(showcase): raise /api/smoke upstream timeout from 25s to 45s (#4089)
## Summary

Cold-start 502s on showcase package smoke probes (notably `agno`)
consistently land just over the 25s upstream budget. The Next.js
`/api/smoke` route uses `AbortSignal.timeout(25000)` to call its own
`/api/copilotkit` which proxies to the Python agent which calls the LLM.
On cold start, the end-to-end round-trip takes >25s and the probe aborts
with `latency_ms: 25001, stage: "timeout"`, returning 502.

This PR raises the upstream budget from **25s -> 45s** across all
showcase packages that exercise a full agent round-trip on smoke, and
bumps Next.js `maxDuration` from 30s -> 60s so the route can actually
run that long. The `create-integration` template is bumped too so new
packages inherit the new budget.

The 3 `langgraph-*` routes have a different (lightweight health probe)
pattern and `ms-agent-dotnet` is already at 50s — their `maxDuration` is
bumped for consistency but their abort budgets are left alone.

**Evidence — agno 502 alerts tonight (PDT):**
- 18:40 PDT attempt 1 — HTTP 502 on `/api/smoke`, self-recovered
- 19:42 PDT attempt 2 — HTTP 502 on `/api/smoke`, `/api/health` already
200

Both probes had `latency_ms: 25001, stage: "timeout"` — a textbook
cold-start boundary trip.

## Why 45s (not 60s or 90s)

The alert tells us the cold path lands **just over** 25s. 45s gives ~20s
of headroom — enough to absorb cold starts without making slow failures
look like slow successes. Tail latency beyond 45s would still alert,
which is the intent. If we see tail probes >45s, we'll revisit with a
post-deploy warmup ping.

## Files changed (17)

- 16 `showcase/packages/*/src/app/api/smoke/route.ts` (13 with both
abort + maxDuration bumped, 3 langgraph-* with only maxDuration bumped,
ms-agent-dotnet untouched)
- `showcase/scripts/create-integration/index.ts` template (new packages
inherit the new budget)

## Test plan

- [ ] Deploy to Railway, observe agno `/api/smoke` no longer 502s on
cold start
- [ ] Confirm other showcase services still smoke-pass
- [ ] If cold path drifts beyond 45s for any service, add a post-deploy
warmup step (deferred)

Do not merge yet — pending user review.
2026-04-19 07:23:26 -07:00
Jordan Ritter df83668bd0 fix(showcase): raise /api/smoke upstream timeout from 25s to 45s
Cold-start on showcase packages with heavy Python agents (agno in particular)
consistently lands just over the 25s budget, producing 502s on first probe
with latency_ms: 25001 and stage: "timeout". Raise the upstream
AbortSignal.timeout on /api/smoke from 25s to 45s across all 16 showcase
packages that exercise the full agent round-trip, and bump Next.js
maxDuration from 30s to 60s so the route can actually run that long.

Also bumps the create-integration template so new packages inherit the
new budget.

Tail-latency beyond 45s will still alert — which is the intent.

Evidence: agno 502 alerts tonight at 18:40 PDT and 19:42 PDT, both with
latency_ms: 25001, stage: "timeout" on /api/smoke. /api/health already
200 on both probes — pure cold-start boundary issue.
2026-04-18 20:56:30 -07:00
Jordan Ritter 31a7d23d1e ci: probe only /api/health in starter smoke test
The Deployed Starters smoke test in integration-smoke.spec.ts uses
checkHealth() with its default paths list ["/api/health", "/health"].
All deployed starters serve their health endpoint at /api/health
(standard Next.js convention); /health returns a Next.js 404 HTML
page. When /api/health returns a legitimate 5xx (e.g. 503
"agent degraded"), the fallback silently probes /health and the
resulting 404 becomes the reported failure — so the Slack alert
reads "status=404 path=/health" and implies a route-missing bug,
hiding the real upstream 503.

Shape A fix: change the default to ["/api/health"] only. No
fallback. The reported status and body are now the actual
/api/health response.

starter-smoke.spec.ts passes an explicit paths argument and is
unaffected. The integration backend L1 health check also targets
/api/health on all deployed backends (verified).

Evidence runs (all same failure mode):
- https://github.com/CopilotKit/CopilotKit/actions/runs/24619147643
- https://github.com/CopilotKit/CopilotKit/actions/runs/24618944101
- https://github.com/CopilotKit/CopilotKit/actions/runs/24617315522
2026-04-18 20:56:10 -07:00
Jordan Ritter de27e9453f fix(showcase/starters): add .NET build outputs to template .gitignore
ms-agent-dotnet starter was hand-patched with bin/, obj/, *.user, *.suo
but the template wasn't updated — drift-check kept flagging it on regen.
Add the entries to the template so all 17 starters pick them up uniformly
(harmless for non-.NET starters).
2026-04-18 20:47:51 -07:00
Jordan Ritter 51e6f4869d fix(showcase/starters): regenerate after R5 spring-ai template changes
R5 added BoundedToolCallingManagerConfig, OpenAiApiKeyValidator, and
WebClientConfig to showcase/packages/spring-ai/agent/ plus updated
application.properties and pom.xml. Starters were out of sync so
drift-check failed. Ran generate-starters.ts to sync.

Also add bin/ and obj/ .NET build artifacts to ms-agent-dotnet starter
.gitignore so generator output stays clean.
2026-04-18 20:44:39 -07:00
Jordan Ritter b426bcfad7 docs(showcase/spring-ai): clarify OPENAI_BASE_URL /v1 convention and consolidate JVM-arg rationale 2026-04-18 20:44:39 -07:00
Jordan Ritter abbf74d9eb docs(showcase/ms-agent-dotnet): replace line-number refs with symbolic names + rename SingleUseMessages 2026-04-18 20:44:39 -07:00
Jordan Ritter 61c3d8f354 docs(showcase/strands): fix comment accuracy for cross-package sync and dict internals 2026-04-18 20:44:39 -07:00
Jordan Ritter 34b66b7ddf fix(showcase/langroid): narrow handle_run except + comment accuracy
- Narrow the mid-stream except around agent.llm_response_async from a
  broad Exception to (openai.APIError, httpx.HTTPError,
  asyncio.TimeoutError, pydantic.ValidationError). Programmer bugs
  (AttributeError, NameError, TypeError) now propagate instead of
  being sanitized as 'Agent run failed: <ClassName>' — keeps real
  bugs visible in server logs and stack traces, where they belong.
  Expected runtime errors still get the sanitized TEXT_MESSAGE +
  RUN_FINISHED treatment so the UI never hangs mid-stream.

- Update _parse_tool_args call-site comment in handle_run: it used
  to say 'returns None' when args could not be parsed; the function
  now returns ParsedArgs with status in {ok, empty, malformed}, and
  the caller already pattern-matches on .usable.

- Fix stale _try_parse_tool docstring that referenced the old
  agent.agent_response() API; the adapter has used
  llm_response_async() for a while now.

Tests: add test_llm_response_async_programmer_bug_propagates — mocks
llm_response_async to raise AttributeError('typo') and asserts the
exception propagates out of the SSE generator (not sanitized).
Update the existing mid-stream-failure test to raise
asyncio.TimeoutError (a covered exception) instead of RuntimeError
so it continues to exercise the sanitization path.
2026-04-18 20:44:39 -07:00
Jordan Ritter fa2534fb28 fix(showcase/google-adk): catch OpenAIError base + comment accuracy
Add openai.OpenAIError (SDK base class) to the generate_a2ui except
tuple so config-time failures — e.g. OpenAI() construction with unset
or malformed OPENAI_API_KEY — become a structured a2ui_llm_error dict
instead of bubbling through the ADK tool machinery as an uncaught
exception. APIError and friends are subclasses of OpenAIError; the
constructor raises OpenAIError directly, which the narrowed tuple
previously missed. Parity fix with the strands sibling (R4).

Also:
- _A2uiError docstring now acknowledges 3-way sync with strands AND
  langroid (all three carry the same TypedDict).
- search_flights docstring relaxed from 'exactly 2 flights' to '2-3
  flights' so google-adk is consistent with ms-agent-dotnet (which
  returns 3).
- Added test exercising openai.OpenAIError from the OpenAI()
  constructor path; asserts a2ui_llm_error with full error shape.
2026-04-18 20:44:39 -07:00
Jordan Ritter b0095dc792 fix(showcase/mastra): cancel upstream body when wrapStreamingResponse throws 2026-04-18 20:44:38 -07:00
Jordan Ritter 2a08f491c3 fix(showcase/spring-ai): address CR round 4 findings (survivor timeout, isTruthy, fail-fast)
- entrypoint.sh: bounded 10s SIGTERM grace window + SIGKILL fallback
  when terminating the surviving sibling; prevents indefinite hang past
  the platform SIGKILL grace period and preserves the structured
  who-died log line.

- BoundedToolCallingManagerConfig: make the weakKeys() identity-vs-equals
  footgun explicit in the comment; add JUnit concurrent-contention tests
  (ExecutorService + CountDownLatch) asserting the cap invariant holds
  under N-thread race on the same ChatOptions — no two threads exceed
  cap-1 and both increment past N.

- WebClientConfig: switch the @Bean-time keepalive check from log-and-
  proceed to fail-fast IllegalStateException unless COPILOTKIT_ALLOW_KEEPALIVE
  opt-in is present (matches OpenAiApiKeyValidator philosophy). Narrow
  isTruthy to 1/true (drop yes) since Spring/Java canonicalize on
  true/false. Update tests to reflect the narrowed vocabulary and
  fail-fast throw.

- application.properties: collapse the 8-line rationale block duplicating
  OpenAiApiKeyValidator javadoc to a one-line pointer (drift risk).
2026-04-18 20:44:38 -07:00
Jordan Ritter 9134e44a28 fix(showcase/strands): address CR round 4 findings (base OpenAIError, concurrency test, comment drift) 2026-04-18 20:44:38 -07:00
Jordan Ritter 4b6d155d5a fix(showcase/ms-agent-dotnet): address CR round 4 findings (null-text guard, dockerignore, comment accuracy) 2026-04-18 20:44:38 -07:00
Jordan Ritter ed8d5f2deb fix(showcase/google-adk): address CR round 4 findings (suffix preservation, error log, survivor)
- main.py: before_model_modifier no longer chops user suffix when the
  prefix end_marker is missing (mangled/drifted prefix). Leaving
  original_text untouched means worst-case is one duplicated signature
  (non-stacking) instead of silent data loss.
- main.py: simple_after_model_modifier now logs llm_response.error_message
  at WARNING with agent name before returning. Previously Gemini
  quota/safety/context-overflow errors were silently swallowed.
- main.py: _A2uiError TypedDict now carries a sync-pointer comment to
  the identical TypedDict in strands/src/agents/agent.py.
- entrypoint.sh: added explicit survivor cleanup with bounded 5s grace
  window, mirroring the spring-ai pattern. Also escalated the
  'both still running' wait -n race branch to ERROR + exit 1 so it
  cannot silently mask the real child death.
- tests: added before_model_modifier suffix-preservation test and
  simple_after_model_modifier error_message warning log tests.
  Updated the test_generate_a2ui docstring to reflect the getattr()
  guard pattern rather than the old AttributeError handler wording.
2026-04-18 20:44:38 -07:00
Jordan Ritter af8207ea22 fix(showcase/langroid): address CR round 4 findings (bytes args, mid-stream SSE wrap)
- _try_parse_tool function_call path now accepts (str, bytes, bytearray)
  for arguments, matching _parse_tool_args. Real OpenAI/httpx stacks can
  deliver bytes; the old narrow isinstance(args, str) guard fell through
  to tool_cls(**args) with raw bytes and silently TypeError'd.
- Wrap agent.llm_response_async in try/except inside event_stream. On
  failure after RUN_STARTED, emit a sanitized TEXT_MESSAGE triple and
  RUN_FINISHED so the frontend never hangs waiting for stream closure.
- Wrap request.json() and RunAgentInput(**body) with explicit handlers
  that return structured JSONResponse errors carrying an errorId for
  client-side correlation (400 for bad JSON, 422 for schema mismatch).
- Defensive guard on run_input.messages: skip entries whose role or
  content are not strings, with a warning. Prevents f-string coercion
  of None/complex values into garbage prompts.
- Wrap the content-JSON-path asyncio.to_thread in try/except that emits
  a sanitized error + RUN_FINISHED before re-raising, so scheduler-level
  failures do not kill the SSE generator mid-stream.
- Duplicate-tool RuntimeError now includes fully qualified
  module.Class identities for each colliding ToolMessage, instead of
  just bare request names.

Tests: 30/30 pass (added 5 new tests covering bytes args, mid-stream
failure, malformed JSON body, invalid RunAgentInput, non-string
role/content, collision error identity).
2026-04-18 20:44:38 -07:00
Jordan Ritter 410ddf8f04 fix(showcase/mastra): address CR round 4 findings (mid-stream errors, resourceId contract) 2026-04-18 20:44:38 -07:00
Jordan Ritter 5122bbec00 fix(showcase/spring-ai): address CR round 3 findings (entrypoint, cache identity, env parsing)
entrypoint.sh:
- Disable errexit before wait -n so the diagnostic kill -0 probes and
  which-died log line run regardless of child exit code.
- Kill the surviving sibling after one child dies to avoid orphan-reparent.

BoundedToolCallingManagerConfig.java:
- Use Caffeine.newBuilder().weakKeys() to force identity comparison on
  cache keys. Without it, two concurrent turns with equal-but-distinct
  ChatOptions references would collapse onto one counter and the cap
  would fire early on the second turn.
- Narrow counter-cleanup catch from Throwable to Exception so we do not
  run arbitrary cleanup code against a JVM in an undefined state
  (OutOfMemoryError, StackOverflowError). Errors now unwind unhandled.

WebClientConfig.java:
- Accept 1 / true / yes (case-insensitive, trimmed) as opt-in values
  for COPILOTKIT_ALLOW_KEEPALIVE instead of only the literal '1'.
- Add defensive runtime check in the @Bean method that logs ERROR if
  jdk.httpclient.keepalive.timeout is not '0' at bean-construction
  time — catches the case where the JVM arg path is dropped AND this
  class loads after an HttpClient has already been constructed.

OpenAiApiKeyValidator.java / application.properties:
- Tighten comments to explicitly note that Spring treats
  '?OPENAI_API_KEY must be set' as the DEFAULT VALUE (non-blank, would
  pass a naive blank check) when the env var is unset.

Tests:
- Add equal-but-not-same ChatOptions counter-independence test.
- Replace Throwable-cleanup test with Error-propagation test.
- Add tests for true/yes/case-insensitive opt-in.
- Add tests for isTruthy helper.
- Add @Bean defensive-ERROR-log test.
- Result: 33 tests, all green (up from 26).

Verified:
- mvn test: 33/33 green
- docker build + smoke (integration-smoke -g spring-ai, LOCAL_PORTS=1): passing
2026-04-18 20:44:37 -07:00
Jordan Ritter 0b2fc399db fix(showcase/ms-agent-dotnet): address CR round 3 findings (invariants, cancellation, tests) 2026-04-18 20:44:37 -07:00
Jordan Ritter a76d904c07 fix(showcase/google-adk): address CR round 3 findings (immutability, narrow excepts, env guards) 2026-04-18 20:44:37 -07:00
Jordan Ritter 99136fc64d fix(showcase/langroid): address CR round 3 findings (unified sanitization, narrow excepts, ParsedArgs)
Changes to src/agents/agui_adapter.py:
- Extract _execute_backend_tool / _run_backend_tool helpers so the
  oai_tool_calls path and the content-JSON path share one sanitization
  contract. The _try_parse_tool backend-tool path no longer runs
  .handle() barefoot — it goes through _execute_backend_tool which
  catches the narrowed (pydantic.ValidationError, ValueError), logs
  the traceback server-side, and returns a sanitized JSON error payload.
- Narrow tool-execution except from
  (RuntimeError, ValueError, TypeError, KeyError, pydantic.ValidationError)
  to (pydantic.ValidationError, ValueError). RuntimeError/KeyError from
  unrelated libraries or real config drift now propagate as real bugs
  rather than being misclassified as user-facing tool failures.
- Replace dict | None three-way return of _parse_tool_args with a
  ParsedArgs dataclass (args + status: 'ok' / 'empty' / 'malformed' +
  .usable helper). Call sites pattern-match on status explicitly.
- Treat empty-string raw_args (legacy function_call path) as 'malformed'
  instead of 'ok with {}' — firing a tool with empty args produces a
  meaningless UI card, same rationale as the oai path.
- Drop TypeError from the json.loads except in _try_parse_tool (it
  signals a programmer bug, not plain text). Add an
  isinstance(content, (str, bytes, bytearray)) guard that logs a
  WARNING rather than silencing the case.
- Drop pydantic.errors.PydanticUserError from tool-instantiation except
  lists — it signals malformed model *definitions*, not runtime data
  errors, and must surface loudly at startup.
- Freeze _TOOL_BY_NAME behind types.MappingProxyType so post-import
  mutation raises TypeError (e.g. accidental monkeypatch / shadowing).

Changes to tests/python/conftest.py:
- Tighten the docstring cross-reference to src/agent_server.py (the
  real entry point confirmed to import agents.agui_adapter identically).

Changes to tests/python/test_agui_adapter.py:
- Update helper tests to assert on ParsedArgs.status / .usable / .args.
- Flip empty-string expectation to 'malformed' per the new contract.
- Add: bytes-input parse, unknown-type 'empty' status, MappingProxyType
  frozen-ness of _TOOL_BY_NAME, unified-helper happy path, helper-level
  sanitization + traceback capture, helper propagates non-narrowed
  RuntimeError, _try_parse_tool non-str/bytes guard warns and returns
  None, and str(exc) never appears anywhere in the SSE stream bytes.
- Switch sanitized-error test's raise to ValueError (narrowed except
  covers it) — previously RuntimeError leaked through by accident.

24/24 tests pass (16 existing + 8 new).
2026-04-18 20:44:37 -07:00
Jordan Ritter 5cc6dd8731 fix(showcase/strands): address CR round 3 findings (python -O guards, TypedDict, __ror__) 2026-04-18 20:44:37 -07:00
github-actions[bot] e03757b3b6 style: auto-fix formatting 2026-04-18 20:44:37 -07:00
Jordan Ritter f794b7b95e fix(showcase/mastra): restore R2 error-handler and type-narrowing lost during rebase 2026-04-18 20:44:37 -07:00
Jordan Ritter 858119259c fix(showcase/langroid): address CR round 2 findings (logging noise, sanitization, tests) 2026-04-18 20:44:37 -07:00
Jordan Ritter 51df7b26ea fix(showcase/spring-ai): address CR round 2 findings (fail-fast, leak, tests) 2026-04-18 20:44:36 -07:00
Jordan Ritter 14a636595f fix(showcase/ms-agent-dotnet): address CR round 2 findings (enumeration, error handling, types) 2026-04-18 20:44:36 -07:00