_extract_user_input in the ag2, crewai-crews, and langroid reasoning agents
documented a str return but passed AG-UI message content straight through.
Multimodal content can be a list of parts, which would flow unmodified into
the single caller in each file (_run_reasoning_agent ->
messages=[{"role": "user", "content": user_input}]) sent to the OpenAI
chat-completions API. Coerce: str passes through, a list joins its text
parts (dict or attr form), anything else falls back to str().
Callers (one per file):
- ag2/src/agents/reasoning_agent.py:104 -> chat message at :119
- crewai-crews/src/agents/reasoning_agent.py:108 -> chat message at :123
- langroid/src/agents/reasoning_agent.py:103 -> chat message at :118
Add a dedicated Spring/Java ReasoningController (/reasoning/) that
reimplements the ag2 reasoning_agent.py BEHAVIOR: it makes a direct
streaming chat-completions call, reads the native delta.reasoning_content
channel (with a <reasoning>...</reasoning> regex fallback), and emits
RUN_STARTED -> REASONING_MESSAGE_START/CONTENT/END -> TEXT_MESSAGE_*
-> RUN_FINISHED so the CopilotKit reasoning slot mounts
[data-testid="reasoning-block"].
Spring AI's ChatClient drops delta.reasoning_content and the AG-UI Java
SDK has no REASONING_MESSAGE_* event types (only THINKING_*, which
@ag-ui/client drops), so the controller manages its own SseEmitter and
writes the reasoning frames as raw JSON matching the @ag-ui/client 0.0.55
wire schema. Header forwarding (x-aimock-context) rides the existing
WebClientConfig exchange filter. Wire route reasoning-custom/-default (plus
legacy aliases) to /reasoning/, mirroring ag2's reasoningAgentNames, and
bump @ag-ui/client ^0.0.43 -> 0.0.55 for REASONING_MESSAGE_* decode support.
(cherry picked from commit 4d183371c489013f0a7bcce1a447078164974aef)
## What
Adopts the new `ag-ui-langgraph` single-arg tool-factory API in
CopilotKit's LangGraph middlewares, and upgrades the `@ag-ui/*`
dependencies to the latest published releases.
Counterpart to ag-ui-protocol/ag-ui#1894 (OSS-248: re-enable A2UI
generation & design guidelines via the shared `A2UIToolParams` bag).
## Commits
1. **`feat(a2ui): use new ag-ui single-arg A2UIToolParams API`**
- sdk-python + sdk-js middleware: `get_a2ui_tools`/`getA2UITools` now
take a single `A2UIToolParams` object (model inside);
`composition_guide` folds into the `guidelines` bag.
- Guarded `A2UIToolParams` import (py); typed import (ts). a2ui test
skip-reason wording.
- Pins: `@ag-ui/langgraph 0.0.40` (js, the release carrying the new
API), `ag-ui-langgraph >=0.0.41` (py).
2. **`chore(deps): use latest @ag-ui packages and langgraph
integration`**
- `@ag-ui/core`, `@ag-ui/client`: `0.0.53 → 0.0.56` across all packages
+ root pnpm override.
- `@ag-ui/langgraph` (runtime): `0.0.39 → 0.0.40`. `ag-ui-protocol
>=0.1.19`.
- `.npmrc`: exclude first-party `@ag-ui/{core,client,encoder,proto}`
from the minimum-release-age gate so the freshly-published `0.0.56` set
installs.
- Regenerated `pnpm-lock.yaml` + `sdk-python/poetry.lock`.
3. **`chore(showcase): align a2ui to single-arg get_a2ui_tools API`**
- `showcase/.../graph.ts`: `getA2UITools(model, opts)` → single-arg.
## Notes
- All `@ag-ui` deps are live on npm/PyPI, so the branch installs
cleanly.
- Showcase **docs** alignment (TS backend tabs for dynamic-schema) is
still being sorted out — the per-integration docs architecture needs
untangling first; deferred to a follow-up.
- ISOLATE_KEEP promoted to a global so --keep survives cmd_test return
into the trap scope
- early-die and default-stack protection: failed --isolate setup no
longer tears down the default stack; half-initialized state is
cleaned up on the way out
- liveness/PID/age reaping signals with a sweep lock: heartbeat
updates, own-pid lock release, tombstones, and a claim-then-verify
duplicate-name guard close slot-registry races (TOCTOU, lock
takeover, reap order)
- teardown robustness: --volumes on every compose down, failed-down
runs preserve state for diagnosis, reap remnants get a compose-down,
path-traversal guard, uniform rm guards under set -e
- name validation: --isolate names must start with a lowercase letter
or digit; reserved name 'showcase' rejected (it aliases the default
stack)
- fail-loud warning before pre-down of an existing stack; help text
updated
Result of an 8-round, 7-agent code-review loop with red-green
verified fixes.
harness-workers runs the SAME showcase-harness GHCR image as the harness
scheduler but has ciBuilt:false (it has no build slot of its own), and the
CI staging redeploy scope was derived purely from ciBuilt — so a main-merge
rebuild of showcase-harness:latest only bounced the scheduler while the
workers silently kept running the stale image (PR #5352's worker-side fixes
never reached staging).
Model image consumption explicitly in the SSOT instead:
- railway-envs.ts: new optional `imageOf` field on ServiceEntry — the SSOT
key of the ciBuilt service whose image this entry runs. Set
`imageOf: "harness"` on harness-workers. New module-load invariant
`assertImageConsumersValid` (fail-loud, same style as
assertDispatchNamesUnique): imageOf must name an existing SSOT key, the
target must be ciBuilt, and the consumer itself must not be ciBuilt.
- redeploy-env.ts: new `expandImageConsumers(names, env)` applied inside
runRedeploy — the redeploy set becomes the resolved scope PLUS any
service whose imageOf points at a service already in scope. Env-aware:
a consumer only joins envs it declares, so the staging-only worker
never enters a prod redeploy (prod behavior unchanged).
No workflow change needed: showcase_build.yml keeps passing the
built-and-successful dispatch_names; the script expands them. Gate
behavior (gateIgnore / image-ref gate), the generated JSON
(emit --check passes byte-identical), and the promote dropdown are all
untouched. harness-legacy deliberately gets no imageOf (pinned pre-fleet
digest; must not follow rebuilds).
## Summary
- fix shell-docs heading anchors so docs headings remain block-level
rows
- prevent adjacent headings, like `## What the runtime provides`
followed by `### Authentication & security`, from rendering on the same
line
- add a `globals.css` regression guard against bringing back
`inline-flex` on `.docs-heading`
## Verification
- `npm run test -- src/app/__tests__/globals-css.test.ts` in
`showcase/shell-docs`
- `npm run lint` in `showcase/shell-docs`
- `npm run test` in `showcase/shell-docs`
- `npm run typecheck` in `showcase/shell-docs`
- `npm run build` in `showcase/shell-docs`
- local visual check at `http://localhost:3003/backend/copilot-runtime`
with Playwright: the H2 and H3 render as separate full-width rows
- lefthook pre-commit passed `check-binaries`, `lint-fix`,
`test-and-check-packages`, and commitlint
## Notes
- `npm run lint` still reports existing warnings unrelated to this
change.
- `npm run build` still emits the existing Turbopack/NFT warning from
`next.config.ts` / `llms-mdx`, but completes successfully.
## Summary
Root-caused and fixed the staging dashboard's red↔green flapping.
Empirical investigation separated it into three mechanisms:
1. **PRIMARY (recurring flap): aimock fixture-sequence desync.** Harness
drivers sent a stable `d6-<slug>` `X-Test-Id` shared across runs.
aimock's stateful strict-mode matching (sequenced fixtures match only
when per-test-id matchCount === sequenceIndex) drifts via replay
exhaustion and 500-test-id FIFO eviction, so whole integrations 503'd
together (~2,200/hr) and recovered — the flap.
2. **SECONDARY (one-off reds): worker restart mid-job.** In-flight
`driver.run` was un-cancellable; redeploys SIGKILL'd workers mid-run,
leaving partial results that fleet-health painted red
(`worker-crashed-mid-job`).
3. **COSMETIC (red floor): dead roster rows.** 22 dead-container
registry rows were never GC'd; the restart hook is a no-op stub, so
`restartsAttempted` over-counted on rows nothing would ever restart.
## Fixes
- **Per-run-unique X-Test-Id** — fold the existing per-run `mintRunId()`
into the test id via exported `buildE2eTestId(slug, runId)` (injectable
`idFactory` for tests) in the d6/d5 and d4 drivers, so every run gets a
fresh aimock sequence cursor.
- **Graceful worker drain** — on SIGTERM the worker aborts in-flight
runs, suppresses red side-emits (`drainReason === "shutdown"` live
getter + abort conjunction), **abandons** partial results (skips
`queue.report`; the 300s lease sweeper synthesizes neutral
`worker-reclaimed-pending`), and deregisters via a latched,
`lastWrite`-chained roster delete so late heartbeats cannot resurrect
the row. Stop bounded by `WORKER_DRAIN_GRACE_MS` (25s default); order
`worker.stop → registration.stop → deregister → pool.shutdown`.
- **Fleet-health GC** — roster rows dead longer than
`DEFAULT_WORKER_GC_AFTER_MS` (24h) are deleted (`gcDeleted` counter),
malformed rows are GC'd by age before the worker_id guard,
`restartsAttempted` is demoted when the restart hook is a no-op, and
`createFleetHealthMonitor` fails loud when `gcAfterMs <= staleAfterMs`.
Companion PR: CopilotKit/aimock#259 (strict-mode no-match diagnostics,
1.30.0) makes the PRIMARY mechanism observable instead of a bare 503.
## Review
- 3 unbiased CR rounds (8 reviewer agents each, verbatim non-leading
prompt) converged to **zero actionable findings**; round-1 criticals
(drain-suppression of genuine timeouts, deregister/heartbeat
resurrection race, vacuous tests, GC-skipped malformed rows) were fixed
with red-green proofs and confirmed in byte-identical follow-up rounds.
- Bucket-(c) promotion audit: 0 promotions to actionable; 3 trivial
consistency items deferred to the follow-up backlog with ~25
pre-existing harness findings.
## Test plan
- [x] Full harness suite: 122 files / 2146 tests green
- [x] `tsc --noEmit` on both tsconfigs; build green
- [x] oxfmt clean on all 14 changed files
- [x] Red-green discipline on every behavioral fix (failures observed
before fixes)
- [x] CI green on this PR
- [ ] Staging promotion (explicitly out of scope here; separate step)
🤖 Generated with [Claude Code](https://claude.com/claude-code)
The global CopyTracker already emits the generic cli_command_copied on
any clipboard copy, but it cannot distinguish which hero card was used.
Each card now also captures hero_command_copied with command_id
(create | onboard), the full command string, and clipboard_blocked, so
create-vs-onboard funnels are queryable per landing page.
Validated locally against a live PostHog client: each click POSTs one
$autocapture, one cli_command_copied, and one hero_command_copied to
/ingest/e (HTTP 200) with the expected properties.
Review follow-up: bring back the quickstart button alongside the two
recommended commands, and render the identical <HeroStartActions> block
on the home hero and every framework landing hero.
- Quickstart returns in its original accent treatment: the framework
dropdown on home, a direct guide link on framework pages; bespoke-init
frameworks (a2a, ms-agent-dotnet) get back the old button + chip row.
- One layout everywhere: cards are always two-up from sm (no more
stacked framework variant); commands wrap balanced at token boundaries
instead of truncating, so long framework-pinned commands stay fully
readable at every width.
Follow-up to #5345 per review feedback:
- Propagate the provider change to all remaining skills: every example,
props table, eval check, and prose mention now recommends CopilotKit
imported from @copilotkit/react-core/v2 (the compatibility bridge and
strict superset) instead of CopilotKitProvider. Migration docs in
copilotkit-upgrade now point at the /v2 import path as the target and
explicitly warn against migrating to CopilotKitProvider.
- Scrub CopilotCloud / Copilot Cloud / CopilotKit Cloud branding from
skills, replacing it with CopilotKit Intelligence where the hosted
platform is meant. Literal endpoint URLs and real identifiers like
MissingPublicApiKeyError are unchanged.
- Fix react-core provider-setup.md which claimed publicApiKey was the
canonical prop; publicLicenseKey is canonical and publicApiKey is a
deprecated alias, matching #5345.
- Edits made in the packages/*/skills source dirs for the three mirrored
skills, with skills/ regenerated via pnpm sync:plugin-skills (this also
re-pins the plugin version fields to 1.59.5).
Per review feedback: CopilotKit from @copilotkit/react-core/v2 is the
compatibility bridge across v1 and v2 and a strict superset of both the
legacy root CopilotKit and CopilotKitProvider. Switch all provider
examples, the props table, references, the page asset, and eval checks
to it, and finish the licenseKey -> publicLicenseKey rename in
telemetry-setup.md.
Previously the EXIT-trap restore_isolation always tore down the isolated
stack, ignoring --keep. Now restore_isolation reads a keep flag (set in
cmd-test.sh when --keep is parsed): when kept it skips compose down, the run-dir
removal, and the slot release, and instead prints a survival notice with the
project, slot, the three offset host ports, and the exact manual teardown
command. The kept stack's live containers keep its slot from being reaped.
Persist the compose project name into each claimed slot dir, and at claim time
reap any slot whose recorded project has no live containers (queried via
docker ps --filter label=com.docker.compose.project). This correctly leaves a
--keep'd stack's slot alone since its containers are still up. The existing
PID/age heuristics remain as a fallback for slots predating the project file.
Move the --isolate slot registry and per-run scratch dir off /tmp (wiped on
reboot, world-writable) to $XDG_STATE_HOME/copilotkit/showcase (slots/ and
runs/<name>). The run dir is now keyed by the finalized project name instead
of the PID so a kept run is locatable for manual teardown. Adds a bats suite
covering the new state-base helper and run-dir location.
Replaces the "Install as a Claude Code plugin" section and its
skill-inventory table with our recommended one-line install:
```bash
npx copilotkit@latest skills install
```
**Why:**
- The hand-maintained skill table goes stale fast. It already listed 9
skills and 6 lifecycle slugs that no longer exist (there are 11 skills
now).
- `npx copilotkit@latest skills install` is our new recommended way to
install the skills. It works across Claude Code, Codex, Cursor, Gemini,
and others, and always pulls the latest.
**Two commits, reviewed separately:**
1. `style: format README with oxfmt`. README has failed `oxfmt` on main
since 1e70dae97 widened the lint-fix glob to cover markdown but never
backfilled existing files. This runs the formatter once so the file
conforms, which also keeps the content change below from being buried
under whole-file reformat noise. Render-neutral: only emphasis markers
and table column padding, which GitHub renders identically.
2. `docs: simplify README agent-skills section`. The actual content
change. Surgical, one section.