- BoundedToolCallingManagerConfig: use ConcurrentHashMap with atomic
compute() for per-turn counter; short-circuit null ChatOptions to
avoid shared null-key contamination; clear counter on delegate
exception + log before rethrow; make MAX iterations configurable via
copilotkit.tool.max-iterations (@Value); rename to
MAX_TOOL_ITERATIONS_BEFORE_RETURN_DIRECT for unambiguous semantics;
pin Spring-AI version in Javadoc and document invariant.
- WebClientConfig: keep static initializer as defensive belt-and-
suspenders but document entrypoint.sh JVM arg as authoritative path;
warn when a non-zero keepalive timeout is already set.
- application.properties: fail fast when OPENAI_API_KEY is unset via
${OPENAI_API_KEY:?…} placeholder.
- entrypoint.sh: drop bash 5.1+ PID args from 'wait -n'; log which
child (java/node) exited with which code; bump health-probe timeout
30s -> 60s to cover cold-start JVM warmup.
- Add JUnit 5 tests (spring-boot-starter-test, test scope): first/Nth
call, fresh-options reset, null-options safety, counter eviction on
cap and exception, HTTP/1.1 connector pin.
Address CR findings against the strands showcase package:
agent_server.py
- Replace throwaway lambda with named _disabled_instrument that returns
self instead of None (fluent callers no longer AttributeError).
- Add module docstring + inline warnings explaining the strict import
ordering invariant (instrumentor patch must precede ag_ui_strands /
strands imports).
- Add runtime assertion that the patch is actually installed.
- Annotate each # noqa: E402 with the concrete reason.
- Note that no upstream strands-agents issue has been filed yet.
agents/agent.py
- Wrap all module-level side effects (model init, agent construction,
per-thread dict patching) in build_showcase_agent() so import
failures are localized.
- Fail fast with RuntimeError when OPENAI_API_KEY is unset rather than
accepting an empty string silently.
- _ToolCallCapHook: protect _count with threading.Lock, document the
single-thread-at-a-time ag_ui_strands concurrency model, and emit a
logger.warning when the cap trips so operators see loops.
- _HookInjectingAgentDict: override update(), setdefault(), __ior__,
and __or__ in addition to __setitem__ so CPython's bulk-update C
paths (which bypass __setitem__) still route through hook injection.
- Preserve any pre-existing entries in _agents_by_thread when swapping
in the injecting dict (copy into the new dict first).
- Guard against double-injection on re-insert: before adding a cap
hook, check whether the Agent already has one attached.
- Narrow sales_state_from_args exception handler from bare Exception to
(json.JSONDecodeError, AttributeError, TypeError) and log a truncated
excerpt of the offending tool input at WARNING.
- Document that schedule_meeting duration is intentionally defaulted
in the showcase.
- Remove unused BaseModel / Field pydantic imports.
- Replace ASCII banner comments with section headers.
tests/python/
- New pytest suite exercising the cap hook counter semantics (fires at
max+1, resets on BeforeInvocationEvent, stop_event_loop sentinel),
the hook-injection dict (all mutation paths inject; existing entries
preserved; no double-injection on re-insert), and the threading
instrumentor patch (returns self, accepts arbitrary args, actually
replaces the method, agent_server module installs the patch).
- conftest.py stubs out strands / ag_ui_strands / uvicorn / dotenv so
the suite runs in environments where the heavy runtime deps aren't
installed.
- agui_adapter: compute thread_id once so RUN_STARTED/RUN_FINISHED never
disagree when the caller omits it.
- Narrow silent `except Exception: pass` sites, log parse/tool failures
with sanitized context, and drop the dead `agent.agent_response` path
(returns ChatDocument, not ToolMessage).
- Extract `_emit_text_block` helper DRYing the empty-delta guard, and a
`_parse_tool_args` helper with explicit warning on malformed input.
- Assert tool-class request-name uniqueness at import so collisions fail
loudly instead of silently shadowing.
- Replace `str(response)` fallback with documented empty default.
- Add unit tests covering the OAI tool-calls happy path, malformed args,
legacy function_call shape, empty-content skip, thread_id stability,
and tool-class uniqueness (red-green verified).
Spring-AI ignores the showcase-wide OPENAI_BASE_URL because its split
base-url + completions-path model would produce a double-/v1 against
aimock. Introduce SPRING_AI_OPENAI_BASE_URL carrying the host without
/v1 (e.g. http://aimock:4010), default to https://api.openai.com.
Additionally, pin WebClient's JDK HttpClient to HTTP/1.1 and disable
connection pooling:
- Without reactor-netty, Spring-AI's auto-configured WebClient sent
`Upgrade: h2c` on every cleartext request, which some fixtures (aimock)
reject with 404.
- Default JDK HttpClient pooling reused half-closed sockets between
fixture responses, tripping `Connection reset` on tool-result follow-ups.
L3 chat now passes locally against the stack. L4 tool-rendering still
needs additional work (tool-execution loop cap WIP in BoundedToolCallingManagerConfig).
The LLM could enter an infinite get_weather tool-call loop (300+ calls per
invocation) because strands has no upstream max-iterations knob, and
aimock's fuzzy fixture matching returns the same tool_call response
whenever the last user message contains the matched text. Direct OpenAI
usage can exhibit similar fixation under certain prompts.
Mitigations in agents/agent.py:
1. _ToolCallCapHook (HookProvider): counts tool calls per Agent
invocation (reset on BeforeInvocationEvent). On BeforeToolCallEvent
past the cap, sets event.cancel_tool so the tool returns a benign
error result. On AfterToolCallEvent past the cap, sets
invocation_state['request_state']['stop_event_loop'] = True which
strands' event_loop checks to terminate recursion.
2. ToolBehavior(stop_streaming_after_result=True) for get_weather: the
frontend weather card IS the response, so halting after the first
tool result also short-circuits any would-be recursion for the
tool-rendering demo.
3. _HookInjectingAgentDict: ag_ui_strands spawns a fresh Agent per
thread_id from the template WITHOUT copying hooks. Subclass the
per-thread dict and attach a fresh _ToolCallCapHook (stateful) on
every agent insertion.
Verified: L3 and L4 strands smoke tests pass 3x locally.
Langroid's OpenAI-backed LLM returns tool calls on
`response.oai_tool_calls` (OpenAI tools API) or `response.function_call`
(legacy function-calling API) with an empty `response.content`. The prior
adapter only inspected `response.content`, so pure tool-call turns
produced no TOOL_CALL events, the frontend never rendered the weather
card, and L4 (tool-rendering) smoke timed out waiting for an assistant
message.
Synthesize TOOL_CALL_START / TOOL_CALL_ARGS / TOOL_CALL_END events from
`oai_tool_calls` (and `function_call` as a fallback), preserving the
OpenAI-assigned call id and JSON-encoded arguments. For backend tools,
instantiate the matching ToolMessage subclass, execute its `.handle()`
off-thread, and stream the result back as a text message (skipping empty
deltas to stay schema-compliant). Frontend tool calls are forwarded as
events only, letting CopilotKit's client render them.
strands-agents 1.35.0 unconditionally calls ThreadingInstrumentor().instrument()
when constructing its Tracer (strands/telemetry/tracer.py). Combined with the
async model client dispatching work onto ThreadPoolExecutor, this wraps
ThreadPoolExecutor.submit such that it re-enters itself recursively, producing
RecursionError during tool-rendering requests and surfacing as an OpenAI
APIConnectionError.
Setting OTEL_PYTHON_DISABLED_INSTRUMENTATIONS=threading does not help because
strands imports and instruments the class directly, bypassing the entry_point
autoloader. Monkey-patch the instrument() method to a no-op before strands is
imported.
Gemini 2.5-flash, given the SalesPipelineAgent's sales-heavy system prompt,
intermittently returned an empty response (FinishReason.STOP with no content
parts) for weather queries — about 70% of requests, causing the L4
tool-rendering smoke test to time out waiting for the assistant message.
Root causes and fixes:
1. System prompt priming: WEATHER guidance was buried after SALES TODOS.
Reordered to lead with a direct weather instruction ("When the user
asks about the weather, call the get_weather tool") and added an
explicit "ALWAYS provide a textual response after any tool call"
guard. The model now reliably invokes get_weather and emits a summary.
2. after_model_modifier over-termination: the old check ended the
invocation whenever parts[0].text was truthy, which fired on
streaming partial events and on mixed text+function_call responses,
stranding pending tool calls. Now skip partial events, and only
terminate when the response is role=model AND has text AND has no
pending function_call.
3. Added ADK_DISABLE_PROGRESSIVE_SSE_STREAMING=1 in entrypoint.sh to
suppress the ADK's progressive-SSE path, which aggregates cleaner
final events (removes the 'The last event is partial' warnings we
were seeing on every run).
Verified: 10/10 direct backend calls now return TOOL_CALL + TEXT_MESSAGE,
5/5 repeated L4 playwright runs pass, and L3 (agentic-chat /Hello/) still
passes.
SharedStateAgent forced ResponseFormat=json-schema for every request whenever
ag_ui_state was present in the options (which it always is for CopilotKit-
wrapped demos). For plain chat prompts like 'hello' the model dutifully
produced a SalesStateSnapshot shell that failed TryDeserialize, the code
hit yield break, and nothing was streamed to the client -- Playwright's L3
smoke waited 60s for [data-testid=copilot-assistant-message] and timed out.
Bypass the two-pass flow (run the agent normally) unless ag_ui_state.todos
is a non-empty array, i.e. there's actual sales pipeline data to sync. The
two-pass logic still runs for sales-pipeline demos that populate todos.
The AG-UI adapter was emitting TextMessageContentEvent with delta=""
whenever the Langroid response had empty content (e.g. a pure tool-call
turn) or a backend tool handler returned nothing. ag_ui.core validates
delta as min-length 1, so the event construction crashed with a pydantic
ValidationError and the SSE stream blew up mid-run, taking out the
tool-rendering L4 smoke test.
Guard both emission sites by skipping the entire
TEXT_MESSAGE_START/CONTENT/END block when the delta would be empty,
rather than padding with whitespace (which would leak a spurious blank
assistant bubble into the UI).
## Summary
Adds a one-command local smoke harness so the full 17-integration suite
can be exercised against Docker on the dev machine instead of Railway.
Useful when Railway is degraded (aimock OOM, rate limits, cold-start
drift) or when validating changes that haven't been deployed yet.
## Usage
```bash
# one-time
cp showcase/.env.example showcase/.env # fill in keys
pnpm --filter @showcase/e2e-smoke install
# full L1-L4 smoke
pnpm --filter @showcase/e2e-smoke smoke:local
# single level / keep containers up between runs
pnpm --filter @showcase/e2e-smoke smoke:local:L1
pnpm --filter @showcase/e2e-smoke smoke:local:keep
pnpm --filter @showcase/e2e-smoke smoke:local:nobuild
```
## What's in here
- **`docker-compose.local.yml`**: `aimock` added as 18th service →
integration containers reach `http://aimock:4010` on the compose
network, mirroring Railway's `showcase-aimock`.
- **`integration-smoke.spec.ts`**: `LOCAL_PORTS=1` env gates URL
rewriting from `https://showcase-<slug>-production.up.railway.app` →
`http://localhost:<port>` via `shared/local-ports.json`. Starters are
skipped under the flag because they're not in `local-ports.json`.
- **`scripts/smoke-local.sh`**: thin orchestrator — `build → up → wait
20s → playwright → down`. Flags: `--level=L1|L2|L3|L4`, `--keep`,
`--no-build`.
- **`tests/package.json`**: `pnpm smoke:local[:L1|:keep|:nobuild]`
wrappers.
- **`.env.example`**: documents optional
`OPENAI_BASE_URL`/`ANTHROPIC_BASE_URL` + `GitHubToken` (ms-agent-dotnet)
and `GOOGLE_API_KEY` (google-adk).
## Verification
Run locally against a fresh checkout of this branch:
- `LOCAL_PORTS=1 SMOKE_ALL=true npx playwright test integration-smoke
--grep @health` → **17/17 pass in 478ms**
- Full L1-L4 against the local stack → **42/51 pass** (9 failures in
L3/L4 for mastra, google-adk, ms-agent-dotnet, strands, langroid,
spring-ai — these are test-data / fixture gaps unrelated to this
infrastructure and will be filed separately)
- `docker compose -f showcase/docker-compose.local.yml config` validates
with 18 services
## Scope
Pure dev-ergonomics addition. No runtime behaviour changes in the
shipped containers. `LOCAL_PORTS` is opt-in; unset = existing
Railway-URL behaviour preserved.
## Test plan
- [ ] `Validate Showcase` CI still green (no package-source changes)
- [ ] No unrelated CI regressions
- [ ] Follow-up PR will investigate and fix the 9 L3/L4 failures
surfaced by local smoke
- test-cleanup.ts: `new Error(msg, { cause })` is ES2022; workspace lib is
ES2020 so the two-arg overload is missing. Replaced with an
`errorWithCause()` helper that assigns `.cause` after construction.
Runtime is identical (Node >=16.9); only the TS signature differs.
- test-cleanup.ts: retyped `SAFE_STDIO` as `StdioOptions` (still frozen at
runtime to keep `test-cleanup.test.ts` freeze assertion green) so
spreading `SAFE_EXEC_OPTS` into `execFileSync(..., opts)` no longer trips
the readonly-vs-mutable-array mismatch on `stdio` (fixes the error at
create-integration.test.ts:132).
- create-integration/index.ts: dropped unused `devCmd` local and unused
`args` parameter on `generateDemoPage` (+ call site); both were dead code
introduced during the parallel-isolation refactor.
- validate-pins.parsers.test.ts: annotated all `withTmp((tmp) => ...)`
callbacks as `(tmp: string)` for robustness under LSP module-resolution
glitches. Matches the contract in `validate-pins.shared.ts`.
Tests: 1061/1061 pass (`pnpm nx run @copilotkit/showcase-scripts:test`).
Prior reasons for `fileParallelism: false` are resolved:
- Env-var mutation races (VALIDATE_PINS_REPO_ROOT, SHOWCASE_AUDIT_ROOT,
VALIDATE_PARITY_REPO_ROOT) are moot under `pool: 'forks'` — every file
already gets its own node process with its own `process.env`.
- `.git/index.lock` races between suites that call `restoreFromGitHead`
are fixed by the cross-process lock in test-cleanup.ts.
- The create-integration vs generate-registry collision on
`showcase/packages/` is fixed by the create-integration tmpdir
isolation.
Under fork-per-file + parallel, each test file also gets a fresh 60s
birpc `onTaskUpdate` budget (vitest #6129), eliminating the cumulative
RPC back-pressure that tripped unit(20.x/22.x/24.x) on #4018/#4068/#4079.
Empirical: local full suite 158s → 12s, 1061/1061 passing across three
consecutive `--skip-nx-cache` runs with zero timeouts, zero ENOENT, zero
index.lock contention.
audit.test.ts is 3034 lines / 119 tests and on Node 22 CI its single-file
runtime grew from 36.7s (PR #4071) to 71.4s (PR #4081) — over the hardcoded
60s birpc onTaskUpdate RPC window (vitest #6129). Same cliff that motivated
the validate-pins split earlier in this PR.
Extract makeTmpTree, makeConfig, writePackage, makeExampleDir, anomalyStrings,
and the AUDIT_SCRIPT path constant into audit.shared.ts so the forthcoming
split files can share them without duplication. No behavior change — the
original audit.test.ts still re-declares its own local copies until the
split commit removes them.
Break the 3567-line validate-pins.test.ts into five smaller test files so
each one fits comfortably under vitest's hardcoded 60s birpc onTaskUpdate
RPC window (upstream vitest #6129: DEFAULT_TIMEOUT = 6e4 in the bundled
birpc). pool: 'forks' + fileParallelism: false already gives each file a
fresh 60s RPC budget; splitting ensures no single file is anywhere near
that cliff even when CI is slow.
Split buckets, chosen for logical cohesion and balanced subprocess load:
- validate-pins.parsers.test.ts - pure parsers (no validateAll, no subprocess)
- validate-pins.validate-all.test.ts - in-process validateAll + drift detection
- validate-pins.cli.test.ts - CLI subprocess exit codes
- validate-pins.eacces.test.ts - chmod/EACCES-routed subprocess tests
- validate-pins.r-scenarios.test.ts - R29/R33 regression scenarios
Test count is identical pre/post-split (134 tests). Behaviour-neutral
refactor: no test logic changes, only file boundaries + shared helper
imports.
Hoist tmpdir/write/withTmp helpers and FIXTURES_DIR/VALIDATE_PINS_SCRIPT
path constants out of validate-pins.test.ts so the forthcoming split files
can share them without duplication. Behaviour-neutral.
create-integration.test.ts invoked the real generator against
`showcase/packages/` and `.github/workflows/`, then healed the mutations
in afterEach. Under `fileParallelism: true` that collided with
generate-registry.test.ts (concurrent readdirSync of `showcase/packages/`
observed partial state → ENOENT) and with every suite that restored
workflow YAMLs from git (`.git/index.lock` contention).
Teach `create-integration/index.ts` to honor two env overrides —
`CREATE_INTEGRATION_PACKAGES_DIR` and `CREATE_INTEGRATION_WORKFLOWS_DIR`
— that redirect its writes to any directory. Production behavior
unchanged (defaults resolve to the real paths as before).
Rewrite the test to create a per-suite `os.tmpdir()`-backed root, seed
it with copies of the three real workflow YAMLs so the generator's
regex-based edits still match, and point both env vars there. The test
now never mutates a tracked file — no restorer, no git invocation, no
cross-suite shared state. generate-registry can scan real
`showcase/packages/` concurrently without collisions.
The regression-guard test still exercises the same cleanup semantics,
just against the tmpdir-backed baseline map instead of
`workflowRestorer.snapshotMap`.
`restoreFromGitHead` runs `git checkout HEAD -- <paths>` inside three
sibling suites (bundle-demo-content, generate-registry, create-integration)
plus concurrent `git` from the pre-commit hook. Every one of those grabs
`.git/index.lock` — parallel callers race for it and flake with
"fatal: Unable to create .git/index.lock: File exists".
Acquire a cross-process advisory lock (atomic `fs.mkdirSync` of
`/tmp/copilotkit-showcase-git-restore.lock`) around every git invocation
in this module: partition, pre-heal diff, checkout, post-heal diff. Held
for the entire sequence so intermediate state is consistent from the
caller's perspective. Stale locks (> 60s) are reaped before the wait loop
so a hard-killed previous run can't wedge subsequent runs.
Unblocks enabling vitest `fileParallelism: true` — the three consumer
suites can now run in parallel forks without stepping on each other's
git operations or on the pre-commit hook's.
Fixes upstream birpc onTaskUpdate timeout (vitest-dev/vitest#8164, fixed
by #8297, v4-only). Root package.json already on ^4.1.3; this aligns
showcase/scripts, which held the only remaining ^3.0.0 pin and was
therefore the only package affected by the bug.
Local verification: 3 consecutive `pnpm nx run
@copilotkit/showcase-scripts:test --skip-nx-cache` runs on Node 20.20.2,
all green, 1061/1061 passing, zero onTaskUpdate errors, zero unhandled
exceptions (wall times: 34s / 30s / 29s).
Pick up the split generative-ui categories, new cells (chat-customization-css,
tool-rendering-default/custom-catchall, tool-rendering-frontend-tools), and
the flat `files` shape (backend files merged in from manifest highlight:).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- Shell /code viewer now builds a recursive file tree with core-only
(★ highlighted) and show-all-files toggle via ?view=all; collapses the
legacy flat files + backend_files arrays into one tree
- bundle-demo-content: strict mode — errors on missing highlight paths;
drop backend_files field; pull in external backend files referenced
by highlight: (column-relative paths) alongside demo-folder contents;
stable page-first ordering
- Update tests to reflect new column-relative filename shape
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- tool-rendering: per-tool WeatherCard + FlightListCard + wildcard
CustomCatchallRenderer (primary variant of 3)
- NEW tool-rendering-default-catchall: no frontend renderers, relies on
CopilotKit's built-in default catch-all
- NEW tool-rendering-custom-catchall: single wildcard renderer via
useDefaultRenderTool
- NEW tool-rendering-frontend-tools: useFrontendTool defines get_weather
client-side; render callback uses `args` (not `parameters`)
- NEW chat-customization-css cell: page.tsx + scoped theme.css
- hitl-in-chat: time-picker-card.tsx + page.tsx using useHumanInTheLoop
with book_call + topic/attendee args
- gen-ui-interrupt: time-picker-card.tsx + page.tsx using useInterrupt
against the schedule_meeting interrupt payload
- a2ui-fixed-schema: update page.tsx comment to point at new backend
- Delete old hitl/ folder
- Update api/copilotkit/route.ts to register new agent names
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- a2ui_fixed: load schemas from JSON via a2ui.load_schema (schemas in
src/agents/a2ui_schemas/); adds search_flight tool with airline + price
- a2ui_dynamic: secondary-LLM bind_tools([render_a2ui], tool_choice=...)
pattern; primary LLM only decides when to render
- tool_rendering_agent: expand to 4 mock tools (get_weather, search_flights,
get_stock_price, roll_dice) — shared by all 3 tool-rendering variants
- interrupt_agent: port schedule_meeting(topic, attendee) time-picker
interrupt payload
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Split generative-ui category into 4 (controlled/declarative/open/
operational), add chat-customization-css + tool-rendering-frontend-tools,
replace old tool-rendering-status/-result IDs with default-catchall /
custom-catchall. Port langgraph-python manifest features + highlight:
paths (column-relative to preserve pre-existing Docker structure).
Update tests for new category/feature counts.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Replace the regex-over-YAML options-block scanner with a real YAML
parse + typed navigation down on.workflow_dispatch.inputs.service.
options. The previous regex depended on a brittle terminator
(`^\S | \n\s*\w+:\s*\n\s{6,}\w+:|\nconcurrency:`) that
would silently break under trivial reformats — adding a top-level key
after `on:` or reordering `concurrency:` would have produced a
false-positive pass.
ALL_SERVICES stays regex-scanned because its matrix is embedded as a
bash heredoc inside a run: step, not as a YAML sub-structure. We do
locate the step via YAML (jobs.detect-changes.steps[*] with run body
referencing ALL_SERVICES) and only then scan `"dispatch_name":"X"`
occurrences inside that step's run body. Cannot JSON.parse the matrix
directly — it interpolates ${{ github.sha }} expressions that aren't
valid JSON pre-execution.
NIT 2: drop the dual `import.meta.url === \`file://${argv[1]}\``
branch and keep only the canonical `fileURLToPath(import.meta.url)`
comparison. The file:// form was belt-and-suspenders for a
pre-fileURLToPath era of Node; Node 20+ handles the modern form
uniformly.
Verified:
- All 12 validator tests pass (substring-spoof, missing-from-both,
template-excluded, empty-starters, missing-workflow).
- Real showcase_deploy.yml returns OK: all 17 starter(s) registered.
Every peer validator under showcase/scripts/ already has a matching
__tests__ suite — this one shipped without. Model the new suite on
validate-parity.test.ts: per-test tmpdir fixtures + CLI subprocess
runs gated by VALIDATE_WORKFLOW_STARTERS_REPO_ROOT.
Extend the existing env-var override to re-home STARTERS_DIR as well
(previously it only redirected .github/). One env var, one root,
matches the pattern used by the other validators.
Coverage:
- happy path (all slugs registered, exit 0)
- slug missing from workflow_dispatch options only (exit 1)
- slug missing from ALL_SERVICES matrix only (exit 1)
- slug missing from both (exit 1, both sources named)
- empty showcase/starters/ (exit 3, refuses trivial pass)
- template/ excluded (not flagged as missing)
- substring-spoof: starter-ag2 missing vs starter-ag2-extended present
— regression guard for the word-boundary anchor
- showcase_deploy.yml absent (exit 3)
All 12 specs pass against the existing regex-based checks. The
substring-spoof case in particular pins the word-boundary contract so
a later refactor can't silently lose it.
The SERVICES=(...) array in showcase_smoke-monitor.yml's 'Check image
drift' step was a hardcoded copy of the 17 starter slugs already
declared by showcase/starters/*/ directory names. Every new starter
required a manual edit in three places; the parity validator caught
drift after the fact but couldn't prevent it.
This commit:
- Adds a sparse actions/checkout step for showcase/starters/ only.
- Replaces the literal starter-* entries with a filesystem enumeration
(for dir in showcase/starters/*/; do ... done), skipping template/.
- Fails loudly if the enumeration produces zero starters, so a broken
checkout can't silently under-check drift.
- Updates validate-workflow-starters.ts to drop the smoke-monitor check
(drift is now structurally impossible) while keeping the two remaining
literal-list checks against showcase_deploy.yml (workflow_dispatch
options must be literal pre-checkout; ALL_SERVICES matrix carries
per-starter deploy metadata like railway_id).
Non-starter services stay literal — they don't live under
showcase/starters/ and are provisioned differently.
The scripts under showcase/scripts/ run via tsx at runtime and had no
tsconfig.json, which left LSP falling back to no-config defaults and
flagging bogus errors ("Cannot find module 'fs' / 'path' / 'url'",
"Cannot find name 'process'"). Adds a minimal tsconfig that matches
tsx's runtime semantics: bundler resolution (so existing extensionless
relative imports keep working), ES2022 + DOM libs, Node types, strict
discriminated-union narrowing. Excludes __tests__/ to avoid surfacing
pre-existing type issues outside this PR's scope.
The starter slug list is duplicated across at least three places:
- .github/workflows/showcase_deploy.yml workflow_dispatch options
- .github/workflows/showcase_deploy.yml ALL_SERVICES matrix entries
- .github/workflows/showcase_smoke-monitor.yml SERVICES bash array
Adding a new starter under `showcase/starters/` but forgetting any of
these leaves the service deployable in theory but invisible to the
dispatch UI and/or drift detection — exactly the failure mode this PR
is trying to close.
Add `showcase/scripts/validate-workflow-starters.ts`. It enumerates
every directory under `showcase/starters/` (excluding `template/`)
and confirms `starter-<slug>` is present in each of the three
workflow locations, emitting a precise "missing from: <source>"
diagnostic per gap.
Wire it into showcase_validate.yml right after `validate-parity` so
it gates every PR + main push touching `showcase/**` or the relevant
workflow files. Also extend that workflow's `on.paths` filter to
include `showcase_deploy.yml` and `showcase_smoke-monitor.yml` so
edits to those files trigger the parity check.
Wires up a single-command path to run the full integration smoke suite
against a local Docker stack instead of Railway. Useful when Railway is
degraded (OOM, rate limits) or when testing changes that haven't been
deployed yet.
Additions:
- showcase/docker-compose.local.yml: add `aimock` as 18th service so
integration containers can reach http://aimock:4010 on the compose
network, mirroring the Railway setup where they call showcase-aimock.
- showcase/tests/e2e/integration-smoke.spec.ts: `LOCAL_PORTS=1` env
rewrites each integration's Railway URL to http://localhost:<port>
via showcase/shared/local-ports.json. Starters are skipped under this
flag because they aren't in local-ports.json.
- showcase/scripts/smoke-local.sh: orchestrates build → up → wait → run
Playwright → tear down. Supports --level=L1/L2/L3/L4, --keep, --no-build.
- showcase/tests/package.json: `pnpm smoke:local[:L1|:keep|:nobuild]`
scripts delegate to the helper.
- showcase/.env.example: document optional OPENAI_BASE_URL +
ANTHROPIC_BASE_URL (route through local aimock) and package-specific
GitHubToken + GOOGLE_API_KEY (ms-agent-dotnet, google-adk).
Verified locally: `pnpm smoke:local:L1` → 17/17 L1 green against the
local stack.
The STARTERS filter in integration-smoke.spec.ts gated on
`i.starter?.deployed === true`, but the manifest→registry bundler
(showcase/scripts/bundle-demo-content.ts + generate-registry.ts)
does not carry that field forward from the source YAML manifests.
At registry regen time (commit 14537d8f3), all `starter.deployed`
values were dropped.
Result: the filter matched zero entries, `Deployed Starters`
yielded no tests, and the starter-deployed-smoke CI workflow has
failed on every run (>15 consecutive reds starting 2026-04-18) with
Playwright erroring `No tests found.` against
`--grep "@starter-health|@starter-agent|@starter-chat"`.
Switch to integration-level `i.deployed`, which is the single
deployment flag the registry actually carries and which the
INTEGRATIONS array in the same file already uses for its own
gating. This keeps both filters anchored to the same source of
truth.
Verified locally: `--grep "@starter-health|@starter-agent|@starter-chat" --list`
now discovers 17 tests (one per deployed integration).
Real pin-fix PRs have landed on `main` since the baseline was last
refreshed, reducing the FAIL set from 111 to 109 and invalidating the
stored hash. The ratchet gate now rejects every PR (including ones
that don't touch pins) with a "ratchet down" instruction.
Refresh the baseline to reflect the actual state of `origin/main`:
validatePinsFailCount: 111 -> 109
validatePinsFailHash: 77b586b7 -> d03716b5
Values computed by running the canonical pipeline from
`.github/workflows/showcase_validate.yml` ("Run validate-pins (ratchet)"
step) against a clean `origin/main` worktree:
pnpm exec tsx showcase/scripts/validate-pins.ts 2> stderr
grep ^Summary stdout -> FAIL=109
grep '^\[FAIL\]' stderr | LC_ALL=C sort -u | shasum -a 256
-> d03716b5...f597e81d
This is a pure ratchet-down to match reality, not a policy change.
No validator behavior, workflow, or pin change is included. Actual
pin drift cleanup (109 -> 0) continues as a separate effort.
Unblocks #4068 and any other PR stalled on the same ratchet.
To support exhaustive E2E testing via multiple variants per feature ×
framework, extend the status model with an optional `variants[]` array
per demo (each variant has the same demo/code/E2E/Smoke/QA/health
breakdown) and mount five different visual treatments so we can compare
side-by-side before committing to one:
- `/variants-stack` — each variant rendered as its own mini-row
in the cell; tall cells, all info visible.
- `/variants-tabs` — tabs at the cell top, click to switch
variant; cell stays compact.
- `/variants-aggregate` — pass/total rollups per signal +
"N variants ▾" expand button to drill down.
- `/variants-grid` — mini-matrix: rows = variants, cols =
demo/code/E2E/Smoke/QA/health.
- `/variants-strip` — one colored chip per variant per signal;
hover chip for variant name, click for URL.
Refactor: the grid chrome moves to `components/feature-grid.tsx`
(accepts a `renderCell` callback). Main `/` keeps the existing
single-variant layout via `components/cell-single.tsx`. Shared badge /
links helpers live in `components/badges.tsx` and
`components/variant-pieces.tsx`.
Mock variant data is seeded on 4 demos (langgraph-python's
agentic-chat, gen-ui-tool-based, hitl-in-chat; langgraph-typescript's
agentic-chat) so each option shows variants in context alongside
ordinary cells.
Variant-specific deep links append `?variant=<name>` to the shell
preview / code / hosted URLs — the shell routes can pick that up
later to highlight variant-specific files or payloads.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
- New `scripts/generate-status.ts` — probes every
`integration.backend_url + demo.route` in parallel (~166 URLs) and
writes health status per demo to `shell/src/data/status.json`. E2E,
Smoke, and QA stay mock with explicit `TODO(wire-*)` comments; real
readers for those ingest from `showcase_aimock-e2e.yml`,
`showcase_smoke-monitor.yml`, and `showcase_qa-sync.yml` later.
`GENERATE_STATUS_MOCK_HEALTH=1` offline override for dev.
- Health badge in the feature-matrix cell is now clickable — opens the
hosted URL (`integration.backend_url + demo.route`) in a new tab,
with a tooltip noting the last probe time + status.
- Ran the probe once against Railway: 10/22 langgraph-python demos up
(pre-merge features), 12/22 down (new demos on this branch, not yet
deployed). Other 16 integrations fully live.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Each grid cell now renders a two-row readiness rollup:
demo · code
E2E ✓ Smoke ✓ QA 3d ● up
Signals:
- E2E / Smoke — pass/fail + freshness (green <6h, amber older, red fail
or no suite, gray when bundle itself is stale)
- QA — days since human sign-off (green <7d, amber <30d, red otherwise
or never)
- Health — live probe dot (up/down/unknown)
Cells with no demo show a single centered ✗ (unsupported) or `—`
(supported but no demo yet). Client-side staleness check: if
`status.json.generated_at` is older than 24h, every signal degrades
to a gray `?` and the header shows a stale-bundle warning — so a
failed cron visibly announces itself instead of silently serving
green badges.
Data layer (`src/lib/status.ts`) is a single source of truth for the
badge color/label logic. The cell component is pure presentation.
Status data is currently mock (deterministic per slug+demo) so the
visuals can be reviewed; the CI/Notion/health-probe pipeline that
writes the real `status.json` is out of scope for this commit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
## Summary
Three showcase test suites leak working-tree drift (workflow YAMLs +
shell data JSONs) on every run. This fixes the leaks and adds a shared
`FileSnapshotRestorer` + `restoreFromGitHead` harness so the suites are
idempotent.
Also moves `showcase/scripts/vitest.config.ts` from the thread pool to
the fork pool, which is required under Node 20 for the test subprocess
churn in `validate-pins` + the three generator-invoking suites.
## What this PR does NOT fix — vitest 3.2.4 RPC timeout on Node 20
(upstream)
`unit (20.x)` still reports `Timeout calling "onTaskUpdate"` ->
ELIFECYCLE **after** all 14/14 test files and 1011/1011 tests pass.
Known upstream bug: https://github.com/vitest-dev/vitest/issues/6129.
The birpc timeout is hardcoded at 60 s (`DEFAULT_TIMEOUT = 6e4` in
vitest's bundled `index.B521nVV-.js`) and is NOT exposed to
`vitest.config.ts`. No `poolOptions.forks.*` / `teardownTimeout` /
`hookTimeout` knob influences it:
- `singleFork: true` made the run STRICTLY WORSE — only 1/14 files
completed (validate-pins consumes the whole 60 s budget on its own, run
24602985507).
- `fork-per-file` (default) gives each file a fresh RPC channel, and
every file passes — but the final pool-teardown RPC still races on Node
20 and surfaces as a process-level exit 1.
Under vitest 3.2.4 the fix requires either (a) upgrading to vitest 4.x
(out of scope — monorepo-wide upgrade), or (b) a pnpm patch against the
bundled `DEFAULT_TIMEOUT` constant (cross-cutting change; declined
here). The observable reality: this PR makes the showcase-scripts suites
idempotent and green; the remaining `unit (20.x)` redness is a known
vitest flake orthogonal to what this PR is trying to fix.
## CI-killer error (verbatim)
```
Error: Package directory already exists: /home/runner/work/CopilotKit/CopilotKit/showcase/packages/test-integration-tmp
⎯⎯⎯⎯⎯⎯ Unhandled Errors ⎯⎯⎯⎯⎯⎯
Error: [vitest-worker]: Timeout calling "onTaskUpdate"
ELIFECYCLE Test failed.
Failed tasks:
- @copilotkit/showcase-scripts:test
```
## Root cause (fixed here)
Three test suites invoke real generator scripts that write to tracked
files OUTSIDE any tmp dir, leaking drift on every `nx run-many -t test`:
- `create-integration.test.ts` scaffolds
`showcase/packages/test-integration-tmp/` AND mutates three CI workflow
YAMLs (`showcase_deploy.yml`, `showcase_drift-detection.yml`,
`starter-smoke.yml`).
- `generate-registry.test.ts` rewrites
`showcase/shell/src/data/registry.json` + `constraints.json`.
- `bundle-demo-content.test.ts` rewrites
`showcase/shell/src/data/demo-content.json`.
## Fix (9 commits, by area of concern)
1. **`test(showcase/scripts): add shared test-cleanup snapshot/restore
helper`** — new `__tests__/test-cleanup.ts` + `__tests__/paths.ts` +
direct unit coverage in `__tests__/test-cleanup.test.ts`:
- `FileSnapshotRestorer` — snapshots file content as `Buffer`
(byte-exact, preserves non-utf8), restores only files that drifted,
writes atomically via temp-file + rename, recreates parent dirs on write
ENOENT. Temp filenames use `crypto.randomBytes(8).toString("hex")` so
concurrent writes can't collide and the `snapshot()` sweep can
unambiguously identify stragglers. On `snapshot()`, sweeps
`.<basename>.<16-hex>.tmp` stragglers **scoped to the snapshotted
basenames only**.
- `restoreFromGitHead(repoRoot, paths)` — partitions the input via `git
ls-files --error-unmatch` BEFORE any destructive op. **Narrow catch**:
only genuine exit-1 pathspec errors are treated as untracked; ENOENT /
EACCES / non-exit-1 failures re-raise with captured stderr. Uses
`execFileSync` (no shell), forces `LC_ALL=C` / `LANG=C`, scrubs all
`GIT_*` environment overrides, frozen exec options.
- `test-cleanup.test.ts` itself strips `GIT_*` from child env when it
creates tmp repos — pre-commit hooks (lefthook) run with `GIT_DIR` /
`GIT_INDEX_FILE` set, which would otherwise cause tmp-repo `git commit`
calls to write to the HOST working-tree HEAD.
2. **`fix(showcase/test-integration): clean up test-integration-tmp
between runs`** — `create-integration.test.ts` wires
`FileSnapshotRestorer` + `restoreFromGitHead` into the suite, wraps
`rmSync` in `try/finally` so workflow restoration always runs, and
migrates `execSync(string)` -> `execFileSync("npx", [...args])` via a
shared `runGenerator()` helper.
3. **`fix(showcase/test-integration): stop generate-registry +
bundle-content leaks`** — same pattern applied to
`generate-registry.test.ts` and `bundle-demo-content.test.ts`; drops a
redundant bundler pre-run in the latter.
4. **`fix(showcase/scripts): switch vitest to forks pool for Node 20
stability`** — `vitest.config.ts`: thread -> fork pool.
5. **`docs(showcase/scripts): tidy test-cleanup comments and JSDoc`** —
documentation cleanup.
6. **`fix(showcase/scripts): pin vitest to a single fork + bump teardown
timeouts`** — SUPERSEDED by commit 9 below (left in history for
auditability).
7. **`fix(showcase/scripts): add post-heal drifted-baseline guard`** —
the PR had claimed a drifted-baseline guard on CI for
`restoreFromGitHead`, but no post-heal `git diff --quiet` was actually
running. Adds the missing check: on CI, any tracked path still drifted
post-heal throws `drifted-baseline guard: post-heal diff failed`; off-CI
warns. Red-green unit coverage via a counter-based git shim that
selectively fails the N-th `diff --quiet` (so the post-heal diff is
targeted independently of the off-CI pre-checkout diff).
8. **`fix(showcase/scripts): decouple generate-registry test 2 from test
1 output`** — `sorts integrations by sort_order` was reading
`registry.json` without invoking the generator, so `afterEach(restore)`
between tests meant it was exercising the committed baseline rather than
live output. Adds a `runGenerator()` call at the top.
9. **`fix(showcase/scripts): revert singleFork — fork-per-file is
strictly better`** — empirical data from run 24602985507 proved
`singleFork: true` was worse than fork-per-file (1/14 vs 14/14 files
completing before RPC timeout). Reverts the
`poolOptions.forks.singleFork` change; keeps the 30 s `teardownTimeout`
/ `hookTimeout` bumps. Comments now accurately reflect that the RPC
timeout is upstream-hardcoded in birpc and NOT tunable via vitest
config.
## Proof of idempotence
```
pnpm nx run @copilotkit/showcase-scripts:test --skip-nx-cache
Run 1: Test Files 14 passed (14), Tests 1011 passed (1011)
Run 2: Test Files 14 passed (14), Tests 1011 passed (1011)
git status after each: only the intentional test file edits.
```
Red-green verified for the HIGH CR-findings:
- narrow `partitionTrackedPaths` catch: unit test with empty PATH
reproduces ENOENT; pre-fix hid it as "untracked", post-fix throws.
- basename-scoped tmp sweep: unit test places both target-basename and
unrelated `.something-else.<hex>.tmp`; pre-fix swept both, post-fix
sweeps only target.
- post-heal drifted-baseline guard: counter-based git shim fails the 2nd
`diff --quiet`; pre-fix tests pass (guard absent), post-fix tests throw
on CI / warn off-CI with the advertised message.
## Test plan
- [x] `pnpm nx run @copilotkit/showcase-scripts:test` passes 1011/1011
twice in a row (locally, Node 25)
- [x] Red-green: disabling `restore()` fails the regression + safety-net
tests
- [x] Red-green: disabling the post-heal drift guard fails the new guard
tests
- [x] Working tree clean after full run
- [x] `prettier --check` + `oxlint` clean on touched files
- [x] `GIT_*` scrub in test harness prevents pre-commit-hook-induced
pollution of real HEAD
## CI status
- **`unit (22.x)`**: pass
- **`unit (24.x)`**: pass
- **`unit (20.x)`**: all 14/14 files + 1011/1011 tests pass; post-suite
`onTaskUpdate` RPC timeout emits exit 1. Upstream bug
https://github.com/vitest-dev/vitest/issues/6129; not fixable at
`vitest.config.ts` level on vitest 3.2.4.
## Caveats
- Local verification ran on Node v25.8.0. No Node 20 binary on this dev
host; Node 20 CI was the final gate.
- The residual `unit (20.x)` failure is orthogonal to this PR. To
resolve it we would need to upgrade vitest to 4.x (monorepo-wide change)
or apply a pnpm patch to bump `DEFAULT_TIMEOUT` in vitest's bundled
`birpc`. Both are tracked separately.
Problem: shell's /code viewer was showing every `.py` file under
`src/agents/` for every demo of a package. Visually contaminating:
opening gen-ui-tool-based (Controlled Gen-UI Display) showed
a2ui_dynamic, a2ui_fixed, mcp_apps_agent, open_gen_ui_agent,
reasoning_agent, interrupt_agent, and tool_rendering_agent in the
file picker even though none of them are relevant to that demo.
Root cause: `bundle-demo-content.ts` ran `discoverBackendFiles()` once
per package and attached the same union of all agent files to every
demo. This was fine when all demos shared one graph, but since we
split demos into dedicated graphs the bundle stopped matching reality.
Fix:
- `manifest.schema.json`: add optional `backend_files` field per demo
(string array, paths relative to the package root).
- `bundle-demo-content.ts`: when `demo.backend_files` is present, bundle
exactly those. Otherwise fall back to the legacy full-package scan
so packages that haven't adopted the field still work as before.
- `langgraph-python/manifest.yaml`: populate `backend_files` for every
demo. Each demo bundles `src/agent_server.py` plus only the agent
file its graph routes to (main.py for shared-graph demos;
reasoning_agent.py / interrupt_agent.py / a2ui_dynamic.py /
a2ui_fixed.py / mcp_apps_agent.py / open_gen_ui_agent.py /
tool_rendering_agent.py for demos with dedicated graphs).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>