Strands mints UUID tool_call_ids for tool results (confirmed via the aimock
journal), so the langgraph-python toolCallId-keyed follow-up fixtures never
matched and the agent re-emitted the tool, looping (text-unstable) across
frontend-tools, gen-ui-agent, gen-ui-open, gen-ui-open-advanced,
gen-ui-headless-complete, reasoning-chain and the weather pills. Re-key the
affected multi-leg fixtures to the id- and thread-history-invariant
sequenceIndex pattern (as built-in-agent does), for both integrations.
Also:
- Add write_document poem/email/quantum fixtures for shared-state-streaming
(was a stale single fixture that 404d).
- Strip content from reasoning-chain tool legs (content+toolCalls in one
fixture is undefined behavior); reasoning rides toolCalls alone.
- Narrow the over-broad d4 summarize catch-all to "Summarize the sales
pipeline" so it stops shadowing gen-ui-agent competitor pill; matches
langgraph-python.
- Commit the real multimodal sample.png/sample.pdf (were git-LFS pointers
the harness could not resolve).
Add a write_document tool + document state hook to both Strands agents
so the shared-state-streaming demo can stream a document into
state.document (mirrors langgraph-python StateStreamingMiddleware). Strands
updates state from the full tool args, which the D6 probe tolerates.
Remove the get_weather stopStreamingAfterResult / stop_streaming_after_result
guard from both agents so weather concludes with a natural follow-up turn,
matching the langgraph-python gold standard; the accompanying multi-turn
id-invariant weather fixtures prevent the aimock replay loop without it.
The D6 fleet worker drives each integration over its insecure Docker
origin (http://<slug>:10000), where crypto.randomUUID is undefined so
hand-rolled headless chats threw and never mounted (sse-missing). tsx/
esbuild also wraps named inner functions in __name(...) calls that leak
into page.evaluate and throw __name is not defined.
Add installBrowserContextShims (init-scripts.ts): a __name no-op helper
and a crypto.randomUUID secure-context polyfill, registered via
addInitScript at document_start of every D6 page; wired into the
d6-all-pills newPage goto path.
0.2.2 ships the adapter fixes for the two upstream defects this hunt surfaced:
empty tool-result content (render-tool demos -> OpenAI 400) and RUN_FINISHED
emitted before parallel tool calls drain (INCOMPLETE_STREAM). With the bump,
beautiful-chat, the chart demos, and the catch-all 'Chain tools' pill complete
cleanly. Verified on real gpt-4o.
Both integrations: a dedicated `a2ui_fixed_schema` backend agent exposes a
`display_flight` tool that returns the A2UI `a2ui_operations` envelope
(createSurface -> updateComponents -> updateDataModel) built from a fixed,
pre-authored flight layout targeting the page's `copilotkit://flight-fixed-catalog`
catalog. Mirrors the upstream ag-ui dojo a2ui_fixed_schema demo (ag-ui#2021)
adapted to the showcase's existing catalog/renderers.
- TS returns the envelope as an object (lands in a json content block); Python
returns it as a JSON string (text block) — per each SDK's tool-return shape.
- Enable the runtime A2UIMiddleware on the a2ui-fixed-schema route
(injectA2UITool: false, defaultCatalogId pinned) so it detects the envelope
and paints; the agent emits the envelope itself, so no generate_a2ui injection.
- Mount the agent on the /a2ui-fixed-schema sub-path; point the route there.
Verified on real gpt-4o: the flight card paints on both integrations.
Add strands-typescript to BASELINE_PARTNERS so it gets its own coverage
column alongside its Python sibling (mirroring how langgraph-typescript sits
beside langgraph-python). Bump the partner-count assertion 26 -> 27.
Apply the same fixes to both the strands (Python) and strands-typescript
integrations for 1:1 parity:
- Lift RunAgentInput.context into the prompt (buildStatePrompt /
build_state_prompt) so useAgentContext values (readonly-state-agent-context)
and the openGenerativeUI design-skill / sandbox-function context actually
reach the model. The adapter does not surface context on its own; this
mirrors langgraph's lift-context-into-prompt pattern.
- open-gen-ui: prepend an imperative to the visualization design skill (and add
a design skill to the advanced cell) so the model calls generateSandboxedUi
instead of answering in plain text, and clarify that sandbox functions are
iframe->host bridges, not LLM tools.
- hitl-in-chat: sharpen the book_call description so it wins scheduling intents
over the shared backend schedule_meeting tool (which renders no picker).
- Add a roll_dice tool (shared python tools + TS tools) so the tool-rendering
catch-all 'Roll a d20' and 'Chain tools' pills work.
## Why
The pool-fleet migration is **complete**. The control-plane harness plus
the prod workers (deployed 2026-06-19, `HARNESS_ROLE=worker`, pool count
2) are live in both envs and cover every probe dimension the interim
`harness-legacy` fleet-migration bridge was holding live.
`harness-legacy` is now dead config.
This PR is the **code-side cleanup only**. The live Railway
`harness-legacy` service is torn down separately (see follow-up below).
## What changed
- **SSOT** (`showcase/scripts/railway-envs.ts`): removed the entire
`harness-legacy` entry and the now-dead `key === "harness-legacy"`
special-case branch in `computePromoteClosure`. Updated stale doc
comments that referenced legacy.
- **Stale `harness-workers` comment fixed**: it claimed the worker is
"STAGING-ONLY". The prod worker is in fact live on Railway (deployed
2026-06-19, `HARNESS_ROLE=worker`, pool count 2). Comment now states
workers run in BOTH envs, while flagging that this SSOT entry still
models only the staging instance — see ambiguity note below.
- **Generated artifact**: regenerated `railway-envs.generated.json` (41
-> 40 services).
- **Golden snapshot** (`__tests__/fixtures/railway-envs.golden.json` +
its test header): dropped the `harness-legacy` block.
- **Fixtures / expectation sets**: removed `harness-legacy` from the
`GATE_IGNORED` sets in `__tests__/verify-railway-image-refs.test.ts`,
the `envsFor` assertion + the `computePromoteClosure` exclusion test in
`railway-envs.test.ts`, the promote-notify fixtures (`success.json`,
`partial.json`, `total-failure.json`), and the `redeploy-env.ts` doc
comments. Updated service-count assertions (41 -> 40) in
`railway-envs.test.ts`, `__tests__/verify-railway-image-refs.test.ts`,
and `__tests__/emit-railway-envs-json.test.ts`.
- **Ruby spec**
(`bin/spec/test_promote_single_service_fleet_invariants.rb`): renamed
the synthetic prod-only fixture name from `harness-legacy` to
`deprecated-prod-only-svc` (a fabricated example name with no SSOT
coupling; renamed to avoid implying the SSOT still carries legacy).
## Red-green proof
This is a behavior change to the SSOT. A new assertion (`railway-envs
SSOT > does NOT contain harness-legacy`) was added and observed FAILING
with the entry still present, then PASSING after removal.
**RED** (assertion added, SSOT entry still present):
```
FAIL railway-envs.test.ts > railway-envs SSOT > does NOT contain harness-legacy (fleet-migration bridge retired)
AssertionError: expected [ Array(41) ] to not include 'harness-legacy'
❯ railway-envs.test.ts:137:23
136| const names = listServiceNames();
137| expect(names).not.toContain("harness-legacy");
| ^
Test Files 1 failed (1)
Tests 1 failed | 103 skipped (104)
```
**GREEN** (SSOT entry removed + fixtures updated):
```
Test Files 4 passed (4)
Tests 171 passed (171)
```
(`railway-envs.test.ts`, `__tests__/verify-railway-image-refs.test.ts`,
`__tests__/railway-envs.golden.test.ts`,
`__tests__/emit-railway-envs-json.test.ts`)
Ruby spec:
```
5 runs, 42 assertions, 0 failures, 0 errors, 0 skips
```
## Pre-push gates
- Formatter (`oxfmt --check`): green
- Lint (`oxlint`): 0 errors (2 pre-existing warnings in
`redeploy-env.ts`, unrelated)
- Typecheck (`tsc -p scripts/tsconfig.json`): green for changed files (1
pre-existing unrelated error in `generate-search-index.ts`, present on
the clean base)
- Tests: full `scripts` suite 2086 passed / 12 skipped. (Two test files
hit a flaky `/tmp/...-generated-data.lock` EEXIST parallel-mkdir race;
they pass when run serially and do not touch `harness-legacy`.)
## Follow-up infra step (NOT in this PR)
- [ ] Delete the live Railway `harness-legacy` service (id
`11279eba-97eb-417e-82a5-7cb4254eb147`) from **both** staging and prod
environments. Manual infra step handled separately by the orchestrator.
## Ambiguity note (prod-worker SSOT entry)
The `harness-workers` SSOT entry currently declares **only** a `staging`
env — there is no `prod` env entry, even though a prod worker is live on
Railway. Per task scope I did **not** invent the prod `serviceInstance`
config (it isn't verifiable from the repo), so I only corrected the
comment to stop asserting "staging-only" and left the entry shape
unchanged. Backfilling a real `prod` env entry is a separate follow-up.
---
## Update: README count sync + code review
Added commit `5859ae4` (`docs(showcase): sync promote-notify README
counts...`): the harness-legacy removal dropped each promote-notify
fixture by one service, so `test-fixtures/promote-notify/README.md` was
updated 29/26/29 → **28/25/28** (counts derived from the fixtures via
`jq`, not hand-set).
**Code review:** 11-agent CR round + 11-agent confirmation round, both
converged to **0 mandatory (bucket a) findings**; Procedure 3 bucket-(c)
promotion audit returned **PROMOTE_TO_A = 0**. The harness-legacy
removal is clean — no dangling SSOT references, generated JSON 41→40
exact, no behavior change.
### Follow-up backlog (NOT this PR — surfaced by CR)
_Distinct-subject (candidate spin-off PRs):_
- **prototype-key hardening** — `verify-railway-image-refs.ts:268`
`findUntrackedServices` uses bare `SERVICES[name]` instead of
`Object.hasOwn`; a Railway service named `constructor`/`toString` is
silently treated as tracked → Railway↔SSOT drift false-negative (real
bug, 2 reviewers).
- **harness-workers prod-backfill** — SSOT declares only a staging env
though a prod worker is live; promote-notify fixtures encode
harness-workers as a promote outcome vs SSOT-skipped;
`serviceId===prodInstanceId` smell.
- **promote-notify fixture fleet completeness** — fixtures model a
28-service fleet, omit the 12 `starter-*` services (live fleet = 40).
_Pre-existing subject-neutral nits:_ `emit-railway-envs-json.ts`
`--check` uncurated-crash edge; `serviceEnvPairs()` unconsumed export;
orphaned `promote-notify/validate.sh`; stale "starter-* staging-only"
rationale in the Ruby fleet spec; truncation-suffix/`failed_count`
README contract with no fixture coverage; unguarded `sibling` nil in the
Ruby spec.
Add the missing "Show me my sales dashboard for this quarter." pill to
the pydantic-ai gen-ui-declarative D6 fixture set: an outer turn
(generate_a2ui, no args) plus the matching _design_a2ui_surface inner
turn carrying the dashboard component payload (KPI metrics row + revenue
pie + monthly-revenue bar), mirrored from the langgraph-python canonical
and the ms-agent-dotnet equivalent.
Closes the staging pydantic-ai D-chat 503 (no_fixture_match): the backend
hits aimock with tools=[generate_a2ui] for this userMessage under
x-aimock-context: pydantic-ai, but only ms-agent-dotnet had the fixture.
Deterministic canonical mirror (no real-LLM recording); contains only the
request-match shape and the A2UI response — no credentials.
## Summary
Two related fixes to the showcase promote pipeline:
**Promote-notify Slack message — say what actually happened.**
Previously a partial promote rendered a message that read as if
*everything* failed, with no indication of which services promoted vs.
failed. Now the message:
- announces `🚂 Promoting showcase → prod (N): <names>` with a legible
count
- on a partial/failed outcome leads with `Promoted:` / `Failed:` (one
header each, then bullets) so you can see exactly which services
succeeded and which didn't
- reports real wall-clock elapsed (integer-coerced; degrades to omitting
the phrase if metadata is unparseable)
- drops the constant `verify-prod:` legend line that preceded the
meaningful data
**Durable healthcheckPath tracking — stop the silent prod drift.** The
promote pin path and the provisioning path could leave a service's
Railway `healthcheckPath` diverged from intent (this is what left aimock
answering on the wrong path in prod). Now:
- `healthcheckPath` is tracked per-service/per-env in the SSOT
(`railway-envs`)
- the promote pin re-asserts it (omit-when-absent — never sends `null`)
- `deploy-to-railway` provisioning routes through `isTrackedService` /
`resolveProvisionHealthcheck`, so a tracked-null service correctly
*omits* the healthcheck while an untracked one keeps the `/api/health`
default
## Test plan
- [x] `test_promote_healthcheck_reassert.rb` — pin re-asserts SSOT
healthcheck, omits when absent, never null (3 runs / 11 assertions)
- [x] `deploy-to-railway.healthcheck.test.ts` + `railway-envs.test.ts` +
`emit-railway-envs-json.test.ts` (121 tests)
- [x] snapshot-ivar lint green (2 runs / 32 assertions)
- [x] notify renderer dry-run mirrors success / partial / total-failure
shapes
The hand-rolled google.genai generate_a2ui planner (and the orphaned
SalesPipelineAgent that consumed it) in main.py are superseded by the
ag_ui_adk 0.7.0 middleware (get_a2ui_tool), now wired backend-owned in
declarative_gen_ui_agent.py / beautiful_chat_agent.py. main.py is reduced to
the shared tool wrappers + before_model/before_agent callbacks still covered
by tests; dead A2UI imports pruned.
- delete tests/python/test_generate_a2ui.py (tested the removed planner)
- manifest declarative-gen-ui: drop stale src/agents/main.py highlight + fix description
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The pool-fleet migration is complete: the control-plane harness plus the
prod workers (deployed 2026-06-19, HARNESS_ROLE=worker, pool count 2) now
cover every probe dimension the interim `harness-legacy` fleet-migration
bridge was holding live, so `harness-legacy` is dead config.
This is the code-side cleanup only:
- Remove the `harness-legacy` entry from the railway-envs SSOT and the
now-dead `key === "harness-legacy"` special-case in computePromoteClosure.
- Regenerate railway-envs.generated.json (41 -> 40 services).
- Drop harness-legacy from the golden snapshot, the gateIgnore expectation
sets, the promote-notify fixtures, and the redeploy-env doc comments;
update the service-count assertions (41 -> 40).
- Fix the stale "STAGING-ONLY" harness-workers comment: prod workers are
live on Railway, though this SSOT entry still models the staging
instance only (no prod env backfilled here yet).
The live Railway `harness-legacy` service is torn down separately as a
follow-up infra step.
Promote-notify Slack message: name the promoted AND failed services (one
Failed: header + bullets), legible "(N): <names>" count, real wall-clock
elapsed (integer-coerced), and drop the constant verify-prod legend line.
Durable healthcheckPath: track it per-service/env in the SSOT (railway-envs),
re-assert it on the promote pin path (omit-when-absent, never null), and route
deploy-to-railway provisioning through isTrackedService/resolveProvisionHealthcheck
so a tracked-null service omits the healthcheck while an untracked one keeps the
/api/health default — fixing the silent prod-healthcheck drift that refused aimock.
Tests: ruby pin-reassert spec + deploy-to-railway healthcheck spec + emit/golden/accessor.
Replace the hand-rolled google.genai A2UI planners (in main.py and
beautiful_chat_agent.py) with the published ag-ui-adk >= 0.7.0 middleware
sub-agent via get_a2ui_tool(), surfacing OSS-158 (forced render_a2ui
sub-agent + toolkit validate->retry recovery loop + recovery-exhausted
hard-fail envelope + render_as_llm_instructions / parse_and_fix healing).
Wiring is BACKEND-OWNED (injectA2UITool: false), matching the AWS Strands /
ag2 external-framework convention rather than langgraph-python's
runtime-driven injectA2UITool: true. Backend-owned is required: the planner
now lives in the ADK middleware, so letting the runtime also inject would
double-bind the tool slot. The explicit is load-bearing post
CopilotKit#5611 (a provider catalog otherwise defaults injectA2UITool to true).
- declarative_gen_ui_agent / beautiful_chat_agent: tools include
get_a2ui_tool({model, default_catalog_id}); beautiful keeps its other tools.
- shared_chat: add get_a2ui_model() to resolve a concrete Gemini BaseLlm for
the sub-agent (mirrors get_model's aimock-proxy wiring).
- routes: injectA2UITool stays false (declarative + beautiful-chat).
- registry/agent_server: no spec-level a2ui config needed (tool is agent-owned).
Verified vs published 0.7.0 + real Gemini: backend-wired generate_a2ui emits
a2ui_operations; OSS-158 gate (subagent + recovery invalid->valid + hard-fail
a2ui_recovery_exhausted) all retained through this exact wiring.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Pins the published middleware that ships the A2UI auto-inject + toolkit
recovery + render_as_llm_instructions/parse_and_fix path. Foundational
step; the A2UI agent re-wire onto that path follows under verification
once 0.7.0 is on PyPI.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
## What this does
Follow-up to #5356. CI never ran `check-types`, so type errors piled up
silently across the monorepo. #5356 fixed shared, a2ui-renderer, and
angular. This PR fixes every remaining package and adds a CI gate so it
cannot happen again.
Almost all of the diff is mechanical type repair: type annotations,
`import type` splits, casts, ambient declarations, and tsconfig
module-resolution bumps. The sections below call out the parts that are
not purely mechanical so review can focus there.
## Where to focus your review
These are the only changes with runtime or public-API impact. Everything
else is type-level and behavior-preserving.
1. **License gating wired to /info** (shared, react-core, vue,
react-native). Behavioral change, details below.
2. **react-native bug fix**: a `catch` binding referenced by
`TypeError.cause` had been lint-stripped and is restored, plus a `{
cause }` is now attached to a parse error.
3. **runtime `Schema` passthrough**: a type-only cap on `tool()` schema
inference. Runtime behavior is unchanged; it only stops tsc from blowing
past 8 GB.
4. **New public type export**: `Anchor` from web-inspector, consumed by
react-core's `inspectorDefaultAnchor`.
5. **New CI job**, see the CI section.
## License context wired to /info
Follow-up to @MikeRyanDev's review on #5356 (wire the inert license
machinery to /info). `createLicenseContextValue` now takes the
server-reported license status instead of a hardwired null.
`checkFeature()` returns false only when the runtime reports `expired`
or `invalid`, and fails open otherwise (null, none, expiring, and valid
all keep features on). The React, Vue, and React Native providers feed
the status they already track from /info. New tests cover the gating in
shared and react-core. Per-feature data is not in /info yet, so gating
is uniform across features.
## Type fixes by package
- **core** (392 errors): bundler module resolution (matches the tsdown
build and the vue package), strict-mode and AG-UI drift in sources and
tests, added `@types/phoenix`.
- **react-core** (255): component, hook, and test drift (now-private
`activeRunCompletionPromise`, RunFinished outcome union, slot prop
types, StandardSchema variance), ambient declarations for `katex` CSS
and the react-markdown JSX namespace, plus the `Anchor` export.
- **runtime** (393, previously hidden behind an OOM crash):
BuiltInAgentConfiguration narrowing, AI SDK v6 and AG-UI drift, the
UserMessage attachment migration, OpenAI v4/v5 surface changes, and the
`Schema` passthrough noted above.
- **react-native** (23): provider/core bridge types, RN 0.85 drift, and
the `catch`/`cause` bug fix noted above.
- **smaller fixes**: runtime-client-gql (UserMessage image narrowing,
pinned `types`), web-inspector (nodenext import extensions, RunFinished
narrowing), react-textarea (es2023 lib for `toReversed`), react-ui,
voice, agentcore-runner, sqlite-runner (one-liners), sdk-js (bundler
resolution plus the missing `@standard-schema/spec` dev dep),
examples/v2/node (drop stale node10 resolution overrides).
## CI
New `check-types` job in `static_quality.yml`: run graphql codegen, then
`nx run-many -t check-types --parallel=1` with a 12 GB heap. The runtime
typecheck alone peaks near 10 GB and takes about 5 minutes on a 16 GB
runner, so it runs serially.
## Testing
- `nx run-many -t check-types` passes for all 24 projects locally and in
CI.
- All package test suites pass: core 431, react-core 1271, runtime 1523,
vue 1001, react-native 246, runtime-client-gql 129, plus the smaller
packages.
- Lockfile diff is only the two added type-only dev dependencies
(`@types/phoenix`, `@standard-schema/spec`).
## Follow-ups (not in this PR)
- react-core's built `.d.mts` files emit extensionless relative imports,
which silently degrade consumer types to `any` under `skipLibCheck`
(found while fixing react-native).
- The AI SDK x zod type-instantiation cost in runtime deserves a real
fix (40M instantiations); trace data available.
- `@ai-sdk/anthropic`'s new `authToken` setting is not forwarded by the
Anthropic adapter.
- `MCPClientProvider.tools()` has diverged from AI SDK v6's `ToolSet`
typing.
## Summary
- Add a dedicated `shell-docs-primary-cta` style for the hero quickstart
link so it keeps the primary CTA color inside reference content.
- Cover the new class usage in the hero and framework overview tests.
## Testing
- Updated unit tests to verify the quickstart CTA class and the matching
global CSS override.
- Updated unit tests to verify the framework overview markup includes
the primary CTA class.
Repairs TypeScript check-types across the monorepo and adds a CI gate so
regressions are caught going forward:
- core: bundler module resolution and strict-mode fixes
- sdk-js: bundler module resolution; keep codegen, formatter, packaging working
- react-core: fixes across components, hooks, and tests
- react-native: restore catch binding referenced by TypeError cause
- runtime: repair check-types and bound AI SDK schema inference
- web-inspector: nodenext import extensions, export Anchor
- remaining packages and node example: assorted check-types repairs
- deps: add missing type-only devDependencies
- license context driven from /info licenseStatus
- ci: run check-types in the static quality workflow
Squashed from 12 commits for a single, easily-revertable change.
CR round 2 follow-ups (no behavior change):
- Add a core-headers regression test proving the agentOwnHeaders baseline
stays pristine across remove + re-add (the WeakMap is intentionally not
cleared on removal; clearing would re-capture polluted headers).
- Correct the stale `headers` config doc ("appended" -> merged on top of each
HttpAgent's own headers, core wins).
- Tighten the e2e no-provider-headers assertion to toEqual.
## Summary
Both the **AG2** and **LlamaIndex** showcase integrations were 100% red
on the staging matrix (every cell `BE ×`, 0 green all day). Root cause
was **not** per-cell logic — each framework's Next.js frontend was
crash-looping under matrix load, driven by an agent-side hot loop. This
PR fixes both root-cause loops.
### LlamaIndex — `sse-missing` (commit forwarding request-time tools)
- **Root cause:** the llama-index AG-UI adapter never forwarded
`RunAgentInput.tools`, so page-injected tools (`toggleTheme`,
`pieChart`, `barChart`, `scheduleTime`, MCP, `generateSandboxedUi`) were
invisible to the LLM → `RUN_ERROR` (404 `no_fixture_match` →
`openai.NotFoundError`) → no `RUN_FINISHED` → `sse-missing`. The retry
storm OOM-crashed the frontend (Node fatal, `next-server` v15.5.18).
- **Fix:** new `RequestAwareAGUIChatWorkflow` (`_request_tools.py`) that
forwards request tools, re-roles the tool-result on an LLM-bound copy
only, and skips duplicate frontend-tool chunk emission (bare snapshot
already carries them). Applied to beautiful_chat / mcp_apps /
open_gen_ui / open_gen_ui_advanced routers; added the `search_flights`
backend tool to beautiful_chat for langgraph-python parity.
- **Red→green:** beautiful-chat **0/5 → 5/5** (toggle-theme,
schedule-meeting, search-flights, pie-chart, bar-chart), freshly
re-confirmed via `--d6 --rebuild`; `mcp-apps` red→green.
### AG2 — `generate_a2ui` empty-arg validation loop
- **Root cause:** the declarative route pointed `HttpAgent` at the root
catch-all agent instead of the dedicated `/declarative-gen-ui` mount,
and `generate_a2ui` required a `context` arg the model emits as `{}` →
pydantic `context Field required` → infinite retry. **630 loop
iterations** observed; the flood starved the ag2 frontend (502s).
- **Fix:** dedicated mount + `injectA2UITool:false` + no-arg
`generate_a2ui` (matches langgraph-python / google-adk) + regenerated
aimock fixture to 4-pill parity + ported the stale ag2
renderers/definitions.
- **Red→green:** **630 → 0** validation errors, `runsFinished=1`, clean
`generate_a2ui → TOOL_CALL_RESULT → RUN_FINISHED`.
## Known remaining (pre-existing, separate, NOT regressions)
- **ag2 `gen-ui-declarative`:** run completes but the a2ui surface
components don't paint (autogen↔AG-UI bridge `TOOL_CALL_RESULT` →
A2UIMiddleware surface-conversion gap). Tracked separately.
- **llamaindex `open-gen-ui`:** `sse-missing` resolved; residual
sandboxed-UI iframe render gap is **pre-existing** (red since
2026-06-01, fc=509 — predates this crash).
## Test plan
- [x] LlamaIndex beautiful-chat `--d6 --rebuild`: 5/5 green (fresh,
isolated)
- [x] LlamaIndex mcp-apps: red→green
- [x] AG2 declarative hot-loop: 630→0 validation errors, runsFinished=1
- [x] Pyright clean on both integrations (verified in-Docker with real
deps)
- [x] oxfmt clean
- [ ] Post-deploy: re-run D6 matrix on staging to confirm ag2 +
llamaindex column recovery
## Deploy note
Staging is **image-sourced, main-only** (the "Showcase: Build & Push"
workflow builds GHCR images and redeploys on push to `main`). This
branch will NOT auto-deploy. To validate on staging before merge: `gh
workflow run showcase_build.yml --ref fix/llamaindex-sse-request-tools
-f service=ag2` (and `service=llamaindex`).
The ag2 declarative-gen-ui route pointed its HttpAgent at the root
catch-all mount (agents/agent.py) instead of the dedicated
/declarative-gen-ui mount (a2ui_dynamic.py), and generate_a2ui declared
a required context arg that the model emits as {}. pydantic rejected
every call with "context Field required" and AG2 retried without bound —
a 630-iteration hot loop per pill that flooded logs and starved the
frontend.
Fix: route to the dedicated mount with injectA2UITool:false (the
dedicated agent owns generate_a2ui and emits a2ui_operations itself);
make generate_a2ui a no-arg tool matching the D6 fixtures and the
langgraph-python gold standard, with a constant inner system prompt
(per-pill distinctness comes from the captured user message). Regenerated
the gen-ui-declarative fixture and ported the LP definitions/renderers
catalog (all 7 driver testids) for parity. Eliminates the validation
loop: runsFinished=1, zero validation errors.
The LlamaIndex AG-UI adapter never forwarded RunAgentInput.tools, so
page-injected frontend tools (useFrontendTool / useComponent /
useHumanInTheLoop) were invisible to the LLM — runs the model couldn't
satisfy ended in RUN_ERROR (no RUN_FINISHED), which the harness reported
as sse-missing.
New RequestAwareAGUIChatWorkflow re-implements the chat step as the
upstream body plus three additions: forward RunAgentInput.tools as no-op
FunctionTool stubs carrying the verbatim injected JSON schema; re-role
tool-result messages to role="tool" on the LLM-bound message copy only
(so the hasToolResult second-leg fixture matches) without mutating
stored/snapshot history; and for frontend tool calls dispatch the
ToolCallEvent but skip the duplicate TOOL_CALL_CHUNK (the bare snapshot
already delivers the call). beautiful-chat 0/5 -> 5/5; mcp-apps fixed.
## Summary
A `service=all` promote runs `verify-staging-precondition` →
`verify-deploy.ts --env staging --services <all>`, which
**hard-crashes** (exit 2) the moment it's handed a service with
`probe.staging=false` (`resolveProbeTargets` throws `not
probe-eligible`). That failed the whole-fleet precondition. This adds an
opt-in `--skip-ineligible` flag (wired into the
`verify-staging-precondition` step **only**) that skips a
non-probe-eligible service with a clear status line instead of throwing;
eligible services are still probed normally, and an explicit
single-service probe / unknown name still hard-fails (strict default
preserved; `verify-prod` untouched).
After #5641 (starters always-on), the only remaining
`probe.staging=false` services are **`harness-workers`** and
**`harness-legacy`** — empirically confirmed to still crash a
`service=all` precondition, so this fix is still needed (the starters
that originally triggered it are now eligible and probed normally).
## Verification
- Red-green (vitest, real surface): a service list containing
`harness-workers` crashes WITHOUT the flag, and is SKIPPED
(`harness-workers (skipped — not probe-eligible for staging,
probe.staging=false)`) WITH it.
- verify-deploy suite: 124/124 pass. `tsc --noEmit` clean on changed
files. `actionlint` on the workflow: exit 0.
## Test plan
- [x] vitest verify-deploy suite green (124)
- [ ] CI green
🤖 Generated with [Claude Code](https://claude.com/claude-code)
## Summary
The 12 `starter-*` services were deliberately `sleepApplication=true`
(sleepable), `probe.staging=false` (held out of the verify-deploy
staging matrix). That class-difference is the root of three recurring
symptoms: the `service=all` promote false-fails them (SLEEPING / no
running-digest), the deployed-starter smoke test 404s a cold container,
and they're excluded from staging validation. Decision: **bring them
into the normal managed-fleet flow — always-on + staging-probed.**
- **`verify-deploy.drivers.starter.ts` (new)** — real `probeStarter`
baseline driver (deployment SUCCESS + HTTP 200 on `/`, mirroring
`probeShell`); replaces the fail-loud `case "starter"` stub.
- **`railway-envs.ts`** — `staging.probe: false→true` for all 12
starters (prod already true); stale "staging probe OFF" comments
rewritten. `railway-envs.generated.json` regenerated (in-sync via
`--check`).
- **`provision-starter-fleet.ts`** — `sleepApplication: true→false` +
comments/logs; test updated.
## ⚠️ Deploy ordering (must hold)
This PR flips `probe.staging=true`, which routes starters into the
staging matrix. It must **not** merge until the live services are
flipped always-on, or a `service=all` promote would probe still-sleeping
starters and fail. Sequence: **(1) live-flip 24 instances
`sleepApplication=false` + redeploy (12 staging + 12 prod), (2) verify
awake + serving `/`, (3) merge this PR.**
## Verification
- Red-green: starter driver (stub→probe `/`), SSOT golden (probe.staging
true), provisioner (sleep false) — all RED→GREEN.
- `showcase/scripts`: 2092 tests pass; `tsc --noEmit` clean on changed
files.
## Test plan
- [x] vitest suites green (2092)
- [ ] CR
- [ ] live 24-instance always-on flip + redeploy (staging + prod)
- [ ] CI green
- [ ] post-flip: `service=all` dry-run shows starters pass staging probe
🤖 Generated with [Claude Code](https://claude.com/claude-code)
## Summary
The built-in-agent `shared-state-read-write` demo wrote agent state via
**direct property assignment** (`agent.state = {...}`). That sets the
value but does **not** fire `onStateChanged`, so components subscribed
via `useAgent` never re-render off a UI write. Switched the seed + title
write to **`agent.setState(...)`**, which updates state *and* notifies
subscribers.
This was the lone divergence: every other framework demo — including the
`langgraph-python` gold standard — already uses `agent.setState(...)`,
and `docs/shared-state.mdx` documents writes going through `setState`.
```diff
- agent.state = { ...defaultRecipe } as unknown as typeof agent.state;
+ agent.setState({ ...defaultRecipe } as unknown as typeof agent.state);
...
- agent.state = { ...(agent.state as object), title: next } as unknown as typeof agent.state;
+ agent.setState({ ...(agent.state as object), title: next } as unknown as typeof agent.state);
```
## Why
`@ag-ui/client`'s `AbstractAgent.setState()` both assigns `this.state`
and notifies all subscribers (`onStateChanged`), which is what drives
the `useAgent` re-render. Direct `agent.state = {...}` is a plain
property write — it skips the notification, so any component relying on
the hook to reflect the change silently goes stale.
## Test plan
- `showcase/bin/showcase test built-in-agent --d5 --verbose --cycle
--isolate` → green after rebuild from the changed source.
- lefthook pre-commit (lint-fix, package checks) + commitlint pass.
## Follow-up (not in this PR)
The built-in-agent demo is recipe-shaped (title/ingredients/steps) while
the other read-write demos are preferences/notes + `DemoLayout`. Per the
showcase 1:1 convention it should mirror `langgraph-python` — larger
change (+ fixture realignment), flagging separately.
## Summary
Adds a rule to `.claude/docs/git.md`: fetch and branch off `origin/main`
before starting work, rather than whatever the local `main` happens to
be.
```sh
git fetch origin && git switch -c <branch-name> origin/main
# worktree equivalent:
git fetch origin && git worktree add -b <branch-name> ../<dir> origin/main
```
Branching off a stale local `main` ("behind origin/main by N commits")
bases the PR on an old commit and invites avoidable merge conflicts. The
rule also documents the recovery path for an already-stale branch: `git
fetch origin && git rebase origin/main`.
Docs-only change to agent guidance; no code impact.
Cover the data-copilot-running turn-done signal in waitForTurnComplete:
true->false transition completion, stayed-stopped quiescence, the
attr-gated early backstop, pre-send run-start baseline, and the
integration wait-for-turn-complete behavior.
Make waitForTurnComplete use the page-side data-copilot-running attribute
as the primary turn-done signal: detect the running true->false transition,
require stayed-stopped quiescence, and gate the early backstop on
attrPresent + runningNow to avoid headless false-RED. Capture a pre-send
run-start baseline so fast turns keep the primary signal alive, read
surfaceReady once per poll, and add computeMaxTurnDurationMs.
Add buildCopilotRunningObserverScript to sse-interceptor.ts and wire it
via addInitScript so the page exposes a data-copilot-running attribute
that the harness can observe for turn-completion signaling.
Add a rule to .claude/docs/git.md to fetch and branch off origin/main
(not stale local main) before starting work, with the worktree
equivalent and the rebase recovery command for an already-stale branch.
Prevents PRs from being based on an old commit.
The built-in-agent shared-state-read-write demo wrote agent state via
direct property assignment (`agent.state = {...}`), which sets the value
but does not fire `onStateChanged` — so subscribed components never
re-render off a UI write. Every other framework demo (including the
langgraph-python gold standard) already uses `agent.setState(...)`, which
both updates state and notifies subscribers. This was the lone divergence.
Switch the seed and the title write to `agent.setState(...)` so the demo
matches the documented pattern (docs/shared-state.mdx says writes go
through `setState`) and the rest of the showcase.
Verified: built-in-agent D5 e2e-deep suite green after rebuild.
## Release monorepo v1.61.1
**Scope:** `monorepo` | **Bump:** `patch`
---
### How this release process works
1. **This PR was created automatically** by the "release / create-pr"
workflow.
It bumped the `monorepo` packages to `1.61.1`
and generated AI-enhanced release notes.
2. **CI runs on this PR** — the full test suite (unit tests, lint, type
checks, build)
must pass before merging. This is the review gate.
3. **Review the release notes** in `release-notes.md` in this PR.
If a Notion draft was created, you can edit the release notes there
before merging.
4. **When this PR is merged**, the `release / publish` workflow
automatically:
- Builds all packages
- Publishes the `monorepo` packages to npm at version `1.61.1`
- Creates git tag `monorepo/v1.61.1`
- Creates a GitHub Release with the final release notes
### Before merging
- [ ] CI is green (tests, lint, types, build)
- [ ] Version bumps look correct
- [ ] Release notes are accurate (edit in Notion if a draft was created)
---
> **Do not merge until CI is fully green.** The full test suite runs
automatically on this PR.
## Problem
Self-hosted users on the v2 SSE runtime never get a `telemetry_id` on
their runtime telemetry events — even with a license token configured.
The root cause is in `packages/runtime/src/v2/runtime/core/runtime.ts`:
- `CopilotIntelligenceRuntime` **did** call
`telemetry.setLicenseToken(...)` in its constructor.
- `BaseCopilotRuntime` and `CopilotSseRuntime` **did not**.
`telemetry_id` is derived only inside `setLicenseToken`
(`parseAndWarnTelemetryId`). If it's never called, every event the
runtime emits is sent anonymously. SSE-mode handlers (`handle-connect`,
`handle-run`, `sse-response`) all emit `oss.runtime.*` events through
the shared telemetry singleton, so all of them went out unattributed for
SSE users.
## Fix
Hoist the license-token resolution (`options.licenseToken ??
COPILOTKIT_LICENSE_TOKEN`) and the `telemetry.setLicenseToken` call
**into `BaseCopilotRuntime`'s constructor**, so SSE and Intelligence
runtimes attribute telemetry identically. The now-redundant duplicate is
removed from `CopilotIntelligenceRuntime` (its `licenseChecker` stays).
The v1 `CopilotRuntime` already set the token and is unchanged.
## Test coverage — every construction path into the endpoints
The token is set at construction time and all endpoints share one
telemetry singleton, so covering every runtime construction path covers
every endpoint.
- **`runtime-license-telemetry.test.ts`** — `CopilotSseRuntime` and
`CopilotIntelligenceRuntime` (direct) + the `CopilotRuntime` shim (both
SSE and Intelligence delegates), each across `{explicit option,
COPILOTKIT_LICENSE_TOKEN fallback, none}`. Asserts the token is set
**exactly once** (guards against a double-set after the hoist).
- **`sse-license-telemetry.integration.test.ts`** — end-to-end: an SSE
runtime built with a license token forwards it all the way to
`lambdaClient.send` on a real Express endpoint request.
- **`copilot-runtime-license-telemetry.test.ts`** — regression guard for
the v1 `CopilotRuntime` path (already worked, previously untested —
exactly the kind of untested path that let this gap appear).
The new tests are **red before the fix** (the SSE-path assertions fail)
and green after.
## Verification
- New + existing v2 runtime suite: **770/770 pass**, no regressions.
- Lint: 0 new warnings. Build: green (types clean). Format: clean
(oxfmt).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Addresses PR review feedback:
- Resolve the license token once (option ?? COPILOTKIT_LICENSE_TOKEN) into a
protected readonly field on BaseCopilotRuntime, and have
CopilotIntelligenceRuntime's licenseChecker reuse it. Collapses the duplicated
resolution and structurally enforces that telemetry attribution and feature
gating can never disagree, instead of relying on a "keep in sync" comment.
- Add an integration test for the env-var-only path (no licenseToken option) —
the exact self-hosted scenario this PR targets — proving the env-resolved
token reaches lambdaClient.send through a real request. Kept in its own file
so the process-wide telemetry singleton (last-write-wins) can't false-pass it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Adds genuine end-to-end coverage beyond the SSE-via-Express case:
- SSE via the Hono adapter
- SSE via the framework-agnostic fetch handler (what node + custom adapters wrap)
- Intelligence mode end-to-end (real CopilotIntelligenceRuntime, WS runner stubbed)
Each constructs a real runtime (so the base-class setLicenseToken runs), drives a
real request through the adapter, and asserts the token reaches lambdaClient.send
on oss.runtime.copilot_request_created.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Only CopilotIntelligenceRuntime called telemetry.setLicenseToken in its
constructor; BaseCopilotRuntime and CopilotSseRuntime did not. As a result,
self-hosted SSE users got anonymous runtime telemetry (no telemetry_id) even
with a license token configured — and those events were additionally throttled
to the 5% anonymous sample rate, leaving runtime telemetry_id stuck at ~1%.
Hoist the licenseToken resolution (option ?? COPILOTKIT_LICENSE_TOKEN env
fallback) and telemetry.setLicenseToken call into BaseCopilotRuntime so SSE and
Intelligence runtimes attribute telemetry identically. Remove the now-redundant
duplicate from CopilotIntelligenceRuntime (its licenseChecker stays).
Tests cover every construction path into the endpoints:
- runtime-license-telemetry.test.ts: SSE/Intelligence direct + CopilotRuntime
shim (both delegates) x {explicit option, env fallback, none}; asserts the
token is set exactly once (guards against a double-set after the hoist).
- sse-license-telemetry.integration.test.ts: end-to-end proof the token rides
to lambdaClient.send through a real Express endpoint request.
- copilot-runtime-license-telemetry.test.ts: regression guard for the v1
CopilotRuntime path (already worked, previously untested).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
## Summary
- Mint a GitHub App token for the stable release workflow and reuse it
for PR creation and follow-up API calls
- Disable lefthook during automation commits so release PR generation
does not depend on local developer hooks
- Relax the CopilotChat perf regression test to assert correctness
without a hard 5s wall-clock check
## Testing
- Unit/UI test updated to allow longer async rendering while still
verifying 100 messages render successfully
- Not run (not requested)
## What does this PR do?
Routes the Gemini canvas stack analyzer through its registered `end`
node instead of directly to LangGraph's `END` sentinel.
This keeps the workflow wiring consistent with
`workflow.set_finish_point("end")` and ensures the cleanup/final state
emission in `end_node` is reachable.
## Related PRs and Issues
- Fixes#5605
## Tests
- `python3.12 -m py_compile examples/canvas/gemini/agent/stack_agent.py`
- `python3.12 - <<'PY' ... PY` (AST check that
`workflow.add_edge("analyze", "end")` exists and
`workflow.add_edge("analyze", END)` does not)
- `. /tmp/oss-pr-pipeline/langgraph-venv/bin/activate && python - <<'PY'
... PY` (LangGraph topology reproduction asserts `analyze -> end ->
__end__` and that `end_node` runs)
- `. /tmp/oss-pr-pipeline/langgraph-venv/bin/activate && python - <<'PY'
... PY` (imports `stack_agent.py` with CopilotKit/Gemini stubs and
asserts compiled `stack_analysis_graph` edges include `analyze -> end`
and not `analyze -> __end__`)
- `git diff --check`
## Checklist
- [x] I have read the [Contribution
Guide](https://github.com/copilotkit/copilotkit/blob/master/CONTRIBUTING.md)
- [x] If the PR changes or adds functionality, I have updated the
relevant documentation (not applicable: example graph wiring bug fix)
- [x] "Allow edits by maintainers" is checked (lets us help iterate on
your PR directly — faster turnaround for everyone)