Commit Graph

12110 Commits

Author SHA1 Message Date
Jordan Ritter dd35aec38a fix(showcase): mint GHCR bearer via /token exchange (unblocks docs promote) (#5239)
## Summary

`showcase bin/railway promote docs` has been failing in CI on every run
with `REFUSE: P1 (docs): GHCR manifest HEAD 403 for
ghcr.io/copilotkit/showcase-shell-docs:latest`. Root cause:
`GHCR#bearer_for` sent the GitHub/Actions token **raw** as
`Authorization: Bearer <token>` to the GHCR OCI manifest endpoint, which
rejects it with `403`. GHCR requires the token to first be exchanged for
a short-lived registry bearer via its `/token` endpoint.

Verified live against the real public `showcase-shell-docs` package: raw
bearer → `403`, `/token`-minted bearer → `200`, anonymous → `401`.

## Changes

- `bearer_for` now **always** performs the GHCR `/token` exchange. When
a token is present it authenticates the exchange with Basic auth
(`base64("x-access-token:<token>")`) and uses the minted token as the
manifest bearer; public packages still mint anonymously.
- Fail-loud hardening: a failed exchange (non-2xx with a token present,
or an unparseable `/token` body) now raises `GHCR::Error` instead of
silently swallowing to `nil` and degrading to an anonymous request —
restoring `manifest_exists`'s documented "raises on transport/5xx"
contract.
- Tests: new `test_ghcr_bearer.rb` pins the on-the-wire behavior
(Basic-auth exchange, minted-bearer-not-raw-token, anonymous path,
raise-on-failed-exchange, never-anonymous-after-supplied-token-failure).
Existing GHCR specs updated to model the now-mandatory exchange.

## Test plan

- [x] `showcase/bin/spec` suite green: 117 runs, 410 assertions, 0
failures
- [x] `ruby -c showcase/bin/railway` syntax OK
- [x] Red-green verified on the new fail-loud paths
- [ ] CI green

## Out of scope (separate follow-up)

Review surfaced three pre-existing, unrelated `bin/railway` bugs not
touched here: `rollback`'s `find_previous_deployment` selects from an
unsorted deployment list, and `env-diff`'s `--ignore-env-scoped` flag /
advertised custom-domain comparison are no-ops. To be addressed in a
separate PR.
2026-06-04 09:43:36 -07:00
Ran Shem Tov 02ae7895e9 chore: restore copilotkit py version 2026-06-04 18:37:52 +02:00
Ran Shem Tov 3ca0f194b8 feat(sdk): gate auto-A2UI injection on injectA2UITool (opt-in)
The A2UI middleware (@ag-ui/a2ui-middleware) forwards injectA2UITool on
forwardedProps; ag-ui-langgraph surfaces it into agent state at
state["ag-ui"]["inject_a2ui_tool"]. The CopilotKit LangGraph middleware
(py + js) now reads that flag and only injects generate_a2ui when it is
truthy (opt-in), drops the runtime's render_a2ui so the model sees one
A2UI tool, and skips if the agent already defines generate_a2ui. The
catalog only binds surfaces; it is no longer the gate.

Reverts the earlier runtime-forward + context-channel approach.

Deps: ag-ui-langgraph>=0.0.38 (py), @ag-ui/langgraph 0.0.37 (sdk-js),
@ag-ui/a2ui-middleware 0.0.6 + @ag-ui/langgraph 0.0.37 (runtime);
.npmrc min-release-age exclude for @ag-ui/a2ui-middleware. Showcase
langgraph pins bumped to copilotkit==0.1.94a3 / sdk-js 1.59.3-alpha.3 /
@ag-ui/langgraph 0.0.37.
2026-06-04 18:37:03 +02:00
Ran Shem Tov 1654d4efb4 test(showcase): bump fixture duplicate ceiling 276->288 for auto-A2UI
declarative-gen-ui moved to the CopilotKitMiddleware auto-A2UI path across
the 3 langgraph integrations. The middleware's inner forced tool is
render_a2ui, so each integration's gen-ui-declarative.json gained 4
render_a2ui fixtures that share match keys with the pre-existing render_a2ui
entries in that integration's render-a2ui.json (the a2ui_fixed demo). 4
pills x 3 integrations = 12. Same-context cross-demo overlap, disambiguated
at runtime by the probe fixtureFile.
2026-06-04 18:36:12 +02:00
Ran Shem Tov acdd1dac7d chore(showcase): ratchet validate-pins baseline for ag-ui-langgraph >=0.0.37
langgraph-fastapi bumped ag-ui-langgraph[fastapi] 0.0.35->0.0.37 (auto-A2UI
needs get_a2ui_tools). Both non-exact, so the FAIL count stays 63 but the
FAIL-set hash shifted. Update validatePinsFailHash to the new set.
2026-06-04 18:36:12 +02:00
Ran Shem Tov 02b11c985d feat(showcase): drive dynamic A2UI via CopilotKitMiddleware (langgraph)
The declarative-gen-ui demo across the three langgraph integrations now
relies on the middleware to inject and execute generate_a2ui — the agents
collapse to create_agent + CopilotKitMiddleware with no hand-rolled tool.
Adds render_a2ui fixtures for the new tool path and pins the integrations
to the A2UI alpha SDKs (copilotkit 0.1.94a1, @copilotkit/sdk-js 1.59.3-alpha.1).
2026-06-04 18:36:12 +02:00
Ran Shem Tov 1bd944bd43 feat(sdk): source A2UI catalog wherever the frontend registered it
The auto-injected generate_a2ui tool now resolves the A2UI catalog from
both delivery paths: the CopilotKit runtime proxy (a copilotkit.context
entry) and the AG-UI native endpoint (state["ag-ui"].a2ui_schema). The
registered catalog id is extracted and bound to generated surfaces so
BYOC custom catalogs render their own components instead of the basic
catalog. Covers @copilotkit/sdk-js and the copilotkit Python SDK.
2026-06-04 18:36:12 +02:00
Ran Shem Tov ce4e4142b2 feat(sdk): auto-inject A2UI tool in CopilotKitMiddleware
Prebuilt agents get dynamic A2UI with no extra wiring — adding the
middleware is enough. When the frontend registers an A2UI catalog
(surfaced by the runtime into state["ag-ui"].a2ui_schema), the
middleware infers the agent's own model, advertises the generate_a2ui
tool in the model-call hook, and executes it in the tool-call hook.
No catalog → the tool is never advertised.

Covers both @copilotkit/sdk-js and the copilotkit Python SDK. Bumps
the A2UI tool-factory dependency to where get_a2ui_tools ships
(@ag-ui/langgraph 0.0.35, ag-ui-langgraph >=0.0.37).
2026-06-04 18:36:12 +02:00
Jordan Ritter b4664424fa ci(showcase): add #team-showcase webhook-unset fallback log on success
Mirror the failure-path 'no Slack' observability fallback for the new
#team-showcase success post. When SLACK_WEBHOOK_TEAM_SHOWCASE is unset
(current default), a successful prod promote now emits a ::notice:: log
line instead of a silent green.
2026-06-04 09:29:12 -07:00
Jordan Ritter 914dbc5855 fix(showcase/bin): surface GHCR /token exchange failures instead of masking them
bearer_for ended with a blanket `rescue StandardError => nil` plus
`return nil if status >= 400`. Now that bearer_for ALWAYS performs the
/token exchange (even when a token is present), that swallow masked real
failures: a non-2xx /token response, a malformed JSON body, or a transport
error all collapsed to nil, after which manifest_exists issued the manifest
HEAD anonymously (no Authorization). That silently violated manifest_exists's
documented raise-on-transport/5xx contract and conflated "no token supplied"
with "supplied token failed to exchange".

bearer_for now:
- raises GHCR::Error on a >=400 /token response WHEN a token was supplied
  (token absent still returns nil — the legitimate anonymous-public fallback);
- raises GHCR::Error on JSON::ParserError for an unparseable 200 body;
- no longer swallows StandardError, so transport exceptions
  (Errno::ECONNREFUSED, Net::ReadTimeout, ...) propagate.

Call-site enumeration (bearer_for is called only by manifest_exists and
resolve_digest):
- manifest_exists: documents "Raises GHCR::Error on 5xx or transport failure",
  so a raised exchange error is consistent with — and strengthens — its own
  contract. Assumption holds.
- resolve_digest: already raises GHCR::Error on its own manifest HEAD >=400 and
  does not rescue bearer_for, so a raised exchange error propagates exactly as
  its other failures do. Assumption holds.
Neither caller is broken by bearer_for raising; both already propagate
GHCR::Error to their callers.

Tests: add 3 red-green cases in test_ghcr_bearer.rb (token-present 401 raises;
token-present malformed body raises; token-present exchange failure issues NO
anonymous manifest HEAD). The existing manifest_exists/digest fakes were only
green because the old swallow masked the fake's "no fake response" error — they
never modeled the mandatory /token exchange; added a successful /token fake to
each so they exercise the real path. Renumbered the snapshot-ivar-lint
allowlist (+13 lines) to track the bin/railway line drift.
2026-06-04 09:18:08 -07:00
Jordan Ritter 8454191090 fix(showcase/bin): always mint GHCR bearer via /token exchange
bearer_for returned a raw GitHub/Actions token, which GHCR's OCI manifest
endpoint rejects with HTTP 403. Always perform the /token exchange instead:
authenticate with Basic base64("x-access-token:<token>") when a token is
present, fall back to the anonymous exchange for public packages, and use
the minted token as the manifest-read bearer. This fixes the docs promote
that has deterministically 403'd in CI.

Route the exchange through the injectable @http seam (matching resolve_digest
/ manifest_exists) so it is testable. Renumber the snapshot-ivar lint
allowlist for the +21 line shift.
2026-06-04 09:07:45 -07:00
Jordan Ritter 2e3e4b2868 ci(showcase): notify #team-showcase on successful prod promotion
Add a success-only Slack post to #team-showcase in the showcase_promote
notify job, mirroring the existing failure→#oss-alerts step. Guarded by
env.SLACK_WEBHOOK_TS so it safely no-ops until the SLACK_WEBHOOK_TEAM_SHOWCASE
org secret exists. Failures continue to route only to #oss-alerts.
2026-06-04 09:06:52 -07:00
Markus Ecker 4d324338ce feat(react-core): per-turn persistent intelligence indicator + finished tag
Render a per-turn persistent intelligence indicator that stays stable
across multi-step turns and settles into a "finished" tag. Splits the
component into IntelligenceIndicator (logic) + IntelligenceIndicatorView
(presentation), wires it into CopilotChatView / CopilotChatMessageView,
adds the slot styles to globals.css, and a Storybook story plus
timer-free logic tests.
2026-06-04 17:56:30 +02:00
Markus Ecker e5ae1b23b7 feat(react-core): useLearningContainers sets thread learning containers
Add `useLearningContainers` / `...InCurrentThread`, which emit
`set_learning_containers` annotations (defaulting to the `project`
container) for the active thread, and export the new hooks from the
react-core hooks barrel.
2026-06-04 17:56:29 +02:00
Markus Ecker 139ffcdd6a feat(react-core): useLearnFromUserAction records user actions via recordAnnotation
Add the `recordAnnotation` client (posts to the runtime `/annotate`
endpoint; no client-side userId) and the `useLearnFromUserAction` /
`...InCurrentThread` hooks that record `user_action` annotations for the
self-learning loop.
2026-06-04 17:56:29 +02:00
Markus Ecker c3f7961242 feat(runtime): attach enterprise-learning MCP middleware on real agent runs
Move enterprise-learning MCP attachment out of the BuiltInAgent-specific
path and the intelligence run handler into a single request-scoped hook:

- `attachIntelligenceEnterpriseLearning` (agent-utils) attaches
  `@ag-ui/mcp-middleware` via `configureAgentForRequest`, gated on
  `ɵisEnterpriseLearningEnabled()`, resolving the user via `identifyUser`
  and the project apiKey.
- Called from `handleRunAgent`; the old `forwardedProps.auth` MCP plumbing
  in `intelligence/run.ts` and the BuiltInAgent attach in `agent/index.ts`
  are removed.
- Add released `@ag-ui/mcp-middleware@0.0.1` dependency (lockfile +
  `@ag-ui/client` override). Drops the obsolete intelligence-mcp-helper test.
2026-06-04 17:56:29 +02:00
Markus Ecker c68be7ecbe feat(runtime): add /annotate endpoint and generalize intelligence client
Replace the user-actions-specific path with a general thread-event
annotation endpoint:

- Route `PUT /connector/annotate/:clientEventId` via fetch-router /
  fetch-handler / core hooks RouteInfo.
- `handleAnnotate` forwards `{ type, data, learningContainer }` to the
  Intelligence platform.
- `CopilotKitIntelligence.annotate()` (idempotent PUT) supersedes the
  prior `recordUserAction`; `AnnotateParams`/`AnnotateResponse` carry the
  event `type` (`user_action`, `set_learning_containers`, ...).

Also renames the `mcpServer` config flag to `enableEnterpriseLearning`
(`ɵisEnterpriseLearningEnabled()`), consumed by the enterprise-learning
attach in the following commit.
2026-06-04 17:56:29 +02:00
Austin Merrick 14bfd3bf83 docs(shell-docs): reference documentation for @copilotkit/core (#5154)
Adds reference documentation for **`@copilotkit/core`** and turns the
reference landing into a multi-SDK **Overview**.

## What changed
- **3-way SDK picker** — React v2 / React v1 / **Core (TypeScript)**
(was a v1/v2 toggle).
- **"API Reference" → "Overview"**, with a "Choose your SDK" chooser.
- **12 hand-written core reference pages** — index, 2 classes, 6 types,
3 enums — matching the existing v2 React convention, every
signature/default/behavior verified against `@copilotkit/core` source.
- **Fixed reference code-block rendering** — the reference route was
missing syntax highlighting + the styled code block (bare `<pre>`); now
wired to the same pipeline as the main docs. Improves all reference SDKs
(v1/v2/core).

## Screenshots
**Overview + "Choose your SDK"**

<img width="1512" height="805" alt="overview"
src="https://github.com/user-attachments/assets/478d6839-4a21-4ba7-917e-5f8b2b74ea12"
/>


**3-way SDK picker**

<img width="1512" height="805" alt="picker"
src="https://github.com/user-attachments/assets/43a094d4-1f62-476b-82f2-d293714471f0"
/>

**Core reference page (syntax-highlighted code + copy button)**

<img width="1512" height="805" alt="core-page"
src="https://github.com/user-attachments/assets/92b661a9-5859-4719-8a39-aa9ade739950"
/>


## Why hand-written (not auto-generated)
The generator can't extract enums/type-aliases (which dominate core's
surface) and overwrites hand-edits; the existing v2 React docs are
themselves hand-written. The ticket explicitly allows deferring
auto-gen.
2026-06-04 08:50:09 -07:00
Ben Taylor c3faf9d17a feat(integrations): add Intelligence threads to crewai-crews (#5198)
## Summary
- Roll Intelligence/Threads support into
`examples/integrations/crewai-crews` using the #5151/#5189 migration
pattern.
- Add the env-gated `CopilotKitIntelligence` runtime path, REST thread
transport, Threads drawer UI, and local Intelligence env docs.
- Add a focused migration contract test covering the crewai-crews wiring
and stale v2 sidebar prop regression.

Linear:
https://linear.app/copilotkit/issue/ENT-734/roll-intelligencethreads-into-the-remaining-11-framework-examples-ent

## Verification
- `pnpm exec vitest run
scripts/__tests__/integration-intelligence-migration.test.ts`
- `pnpm exec oxfmt --check
scripts/__tests__/integration-intelligence-migration.test.ts
'examples/integrations/crewai-crews/src/app/api/copilotkit/[[...slug]]/route.ts'
examples/integrations/crewai-crews/src/app/page.tsx`
- `git diff --check`
- `./node_modules/.bin/tsc --noEmit` in
`examples/integrations/crewai-crews`
- `npm run build` in `examples/integrations/crewai-crews`
- `pnpm install --frozen-lockfile --ignore-scripts`
- pre-commit hook: `pnpm run lint --fix && pnpm run format`, `pnpm run
test`, `pnpm run check:packages`
- `pnpm run build`

## Manual smoke
- Copied local test env from `/Users/mothra/Projects/test-signups4/.env`
into the ignored example `.env`.
- Replaced missing `OPENAI_API_KEY` with a validated local
company/project key.
- Started the local Intelligence stack from `test-signups4` with `docker
compose up -d --wait`.
- Seeded `1_demo-user` / `demo-user` for organization `casa-de-erlang`.
- Ran the CrewAI agent on `:8000` and the UI on `:3010`, with `.env`
exported into both processes.
- Verified `/`, `/api/copilotkit/info`, and
`/api/copilotkit/threads?agentId=starterAgent&limit=20` returned
successfully; runtime info reported `mode: intelligence` and
`licenseStatus: valid`.
- Verified the licensed Threads UI renders in the browser and the drawer
opens with persisted threads.
- Created/selected a thread, sent `What are the proverbs?`, verified the
agent response rendered, refreshed, selected the saved thread, and
verified the user message + assistant response restored.

Note: this `CopilotSidebar` type surface does not expose visible
suggestion chips in `@copilotkit/react-core@1.59.1`, so the send step
was verified by typing the suggested prompt rather than clicking a
suggestion chip.
2026-06-04 10:26:33 -05:00
Benjamin Taylor 26ea1187e7 Merge ben1/intelligence-threads-examples-rollout into codex/ent-734-crewai-crews
Graft resolution (her branch predates the crewai-crews rebuild on main):
- page.tsx: jpr5's rebuilt full demo (setThemeColor/updateProverb/
  get_weather tools, seed effect, delete buttons, WeatherCard) is the
  base; her threads shell (drawer + gate + provider) grafted around it;
  agent key AGENT_ID=default (was starterAgent on her branch)
- layout/route/package.json: rollout side (default key, AGENT_URL
  normalization, 1.59.3) + her drawer deps + intelligence route block
- contract test: crewai-crews added to the parameterized array (66/66)
- lockfile regenerated; .gitignore !.env.example negation added

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 10:26:23 -05:00
Ben Taylor f19ed53e7b feat(integrations): add strands-python intelligence threads (#5202)
## Summary
- rolls the north-star Intelligence/Threads shell into
`examples/integrations/strands-python`
- wires the Strands route to use CopilotKit Intelligence when
`COPILOTKIT_LICENSE_TOKEN` is present, with in-memory fallback when
unlicensed
- documents the Intelligence env vars and removes the temporary
`strands-python` parity allowances
- adds a contract test for the migration pattern and stabilizes the
existing A2UI renderer test that blocked pre-commit

## Verification
- `pnpm exec vitest run
scripts/__tests__/integration-intelligence-migration.test.ts`
- `pnpm parity:verify --target=strands-python`
- `pnpm parity:check`
- `npm run build` in `examples/integrations/strands-python`
- `pnpm --dir packages/react-core exec vitest run
src/v2/__tests__/A2UIMessageRenderer.test.tsx`
- `NX_TUI=false pnpm nx run @copilotkit/react-core:test`
- `git diff --check`
- pre-commit hooks on both commits: `pnpm run test` and `pnpm run
check:packages`

## Manual smoke
- ran the local Intelligence stack from `test-signups4`
- started the Strands agent and Next UI on `localhost:3001`
- confirmed `/api/copilotkit/info` reported `mode: "intelligence"` and
`licenseStatus: "valid"`
- created a thread, asked the agent to enable app mode and add three
todos, then reloaded and selected the saved thread
- verified the restored thread showed the user prompt, tool calls,
assistant response, and the three persisted todos
2026-06-04 10:23:04 -05:00
Benjamin Taylor 5071060eac Merge ben1/intelligence-threads-examples-rollout into codex/ent-734-strands-python
Conflict resolutions:
- contract test: keep the rollout's parameterized version and add
  strands-python to migratedIntegrations/appRoots (60/60 passing) in
  place of her bespoke MIGRATED_INSTANCES file
- parity manifest: rollout's version with strands' three threads-shield
  allowances removed (mirrors the langgraph-fastapi migration); parity
  verify green — strands now 88 tracked files, zero drift
- package-lock: regenerated at 1.59.3 (a2ui-renderer stays 1.56.5,
  the family-wide pin shared with the north-star)

Also rides: her react-core A2UIMessageRenderer test flake fix
(act -> waitFor), kept intentionally.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 10:22:53 -05:00
Ben Taylor 104674f172 feat(examples): add Intelligence threads to agentcore (#5207)
## Summary
- add env-gated CopilotKit Intelligence wiring to the shared AgentCore
Hono/Lambda runtime bridge
- wire the Vite frontend to REST Threads transport with the Threads
drawer, locked gate, and shared `threadId` chat/canvas context
- add local Docker env propagation for Intelligence vars and the Vite
client-safe Threads gate
- pin AgentCore frontend/runtime packages to the threads-capable
CopilotKit/AG-UI versions and add missing frontend peer deps needed for
a clean build
- extend the integration migration regression suite with
AgentCore-specific coverage

## Verification
- `pnpm exec vitest run
scripts/__tests__/integration-intelligence-migration.test.ts`
- `pnpm exec oxfmt --check
scripts/__tests__/integration-intelligence-migration.test.ts
examples/integrations/agentcore/frontend/src/components/chat/CopilotKit/index.tsx
examples/integrations/agentcore/frontend/src/components/threads-drawer/locked-state.tsx
examples/integrations/agentcore/frontend/src/components/threads-drawer/threads-drawer.tsx
examples/integrations/agentcore/frontend/vite.config.ts
examples/integrations/agentcore/infra-cdk/lambdas/copilotkit-runtime/src/runtime.ts`
- `npm run build` from `examples/integrations/agentcore/frontend`
- `npm run build` from
`examples/integrations/agentcore/infra-cdk/lambdas/copilotkit-runtime`
- pre-commit hook: `check-binaries`, `sync-lockfile`, `lint-fix`,
`test-and-check-packages`, `commitlint`
- local Intelligence stack: copied
`/Users/mothra/Projects/test-signups4/.env` into AgentCore's ignored
`docker/.env`, started `docker compose up -d --wait`, seeded
`demo-user`/`1_demo-user`
- local bridge smoke: `GET
http://localhost:3101/copilotkit/threads?agentId=default` returned `200`
- Playwright smoke with local generated `aws-exports.json` + seeded OIDC
storage: authenticated Vite app rendered, licensed Threads drawer
visible on desktop, mobile floating pill visible, no horizontal
overflow, reload preserved authenticated render

## Notes
- Full live AgentCore response verification was not completed because
this path requires a deployed AWS AgentCore/Cognito stack and real
AgentCore agent runtime. The browser smoke used fake OIDC storage and a
dummy agent URL, so it verified frontend/runtime Threads wiring but not
a real agent response.
- In the fake-auth local smoke, the Intelligence realtime websocket
returned `403`, while the REST Threads list rendered successfully. I did
not treat that as proof of a production websocket issue because the
smoke bypassed real Cognito/auth, but it is worth keeping an eye on in a
deployed AgentCore environment.
2026-06-04 09:57:22 -05:00
Benjamin Taylor 7b4e5d1169 Merge ben1/intelligence-threads-examples-rollout into codex/ent-734-agentcore
Conflict resolution: take the rollout's contract test and append the
agentcore describe block (CDK lambda runtime gate, Vite frontend with
import.meta.env gate, docker env wiring). Bump both agentcore
package.jsons 1.59.1 -> 1.59.3 + regen lockfiles. 54/54.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 09:57:10 -05:00
Ben Taylor 2806cfd7ea feat(examples): add Intelligence threads to a2a-a2ui (#5201)
## Summary
- add Intelligence Threads wiring for `examples/integrations/a2a-a2ui`
- add the reusable threads drawer UI and env-gated Intelligence runtime
config
- add a local A2UI v0.8 renderer so the demo renders the current A2A
operation payloads
- extend the migration verifier coverage for a2a-a2ui

## Verification
- `pnpm exec vitest run
scripts/__tests__/integration-intelligence-migration.test.ts`
- `npm run build` from `examples/integrations/a2a-a2ui`
- pre-commit hook: `pnpm run test && pnpm run check:packages` via
lefthook

Refs ENT-734.
2026-06-04 09:53:28 -05:00
Benjamin Taylor da45029d70 Merge ben1/intelligence-threads-examples-rollout into codex/ent-734-a2a-a2ui
Conflict resolution: take the rollout's parameterized contract test and
append the a2a-a2ui bespoke tests (namespaced helper, 1.59.3 pins).
Bump a2a-a2ui @copilotkit/* 1.59.1 -> 1.59.3 + regen lockfile. 49/49.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 09:53:16 -05:00
Jordan Ritter fe95444314 fix(showcase/aimock): reorder gen-ui-interrupt d6 fixture legs fleet-wide (toolName before toolCallId)
Audit follow-up to #5232, which fixed the gen-ui-interrupt d6 fixture
leg mis-order for langgraph-python + langgraph-typescript. aimock's
matchFixture is first-match-wins in array order and the loader dedups
on userMessage, so a toolCallId resume-leg placed BEFORE its
toolName:schedule_meeting first-leg shadows the tool-emitting leg on
turn-2 — the interrupt never fires and time-picker-card never mounts.

Fleet audit of every showcase/aimock/d6/*/gen-ui-interrupt.json found
ag2 as the only remaining mis-ordered fixture (both toolCallId
resume-legs preceded their toolName first-legs). Reorder ag2 so each
pill's toolName first-leg precedes its toolCallId resume-leg, matching
the proven-correct mastra/pydantic-ai/langgraph pattern. Reorder only;
response payloads and match keys are unchanged.

All other gen-ui-interrupt-supporting integrations (built-in-agent,
claude-sdk-*, langgraph-fastapi, langroid, ms-agent-dotnet/python,
pydantic-ai, spring-ai, strands) were already correctly ordered.
2026-06-04 07:36:43 -07:00
Benjamin Taylor 3de8146d5d chore(integrations): unify all threads examples on CopilotKit 1.59.3
- bump a2a-middleware, mcp-apps, agent-spec from 1.59.1 (their verified
  pre-revert state) to 1.59.3 to match the starters
- regenerate package-lock.json for the 11 examples whose package.json
  changed (drawer deps re-added on starters, version bumps on the three)
- update the migration contract test's version assertions to 1.59.3
  (43/43 passing); oxfmt pass

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 09:33:20 -05:00
Benjamin Taylor 041a561381 Revert "Revert "feat(integrations): add Intelligence threads to agent-spec" (#5216)"
This reverts commit ec856ab314, reversing
changes made to 9e7e7e653d.
2026-06-04 09:25:37 -05:00
Benjamin Taylor 4243bd0fa7 Revert "Revert "feat(integrations): add Intelligence threads to langgraph-fastapi" (#5215)"
This reverts commit 9e7e7e653d, reversing
changes made to 11e18e378a.

# Conflicts:
#	examples/integrations/langgraph-fastapi/package-lock.json
2026-06-04 09:25:37 -05:00
Benjamin Taylor 7762c43863 Reapply the ENT-679 Intelligence threads rollout (revert of #5217)
Restores #5151 (north-star + batch 1 + crewai-flows + llamaindex),
#5196 (pydantic-ai), #5205 (a2a-middleware), #5211 (mcp-apps), reconciled
onto the current main baseline rather than the pre-revert tree:

- keep main's 1.59.3 pins, AGENT_URL normalization, default agent keys,
  useConfigureSuggestions, available:false, useRenderTool
  status/parameters API, call-time agent.state reads, and crewai-crews'
  rebuilt page (not yet threads-migrated)
- graft the threads layer (drawer/gate/provider, env-gated route
  intelligence block, next.config gate, env docs, drawer deps) on top
- drop threads-era sidebar suggestions props where main now registers
  suggestions via useConfigureSuggestions (or omits them)
- fix the stale pydantic-ai doc link main reintroduced in ms-af-dotnet

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-04 09:25:16 -05:00
Jordan Ritter f2059cc65d fix(showcase/harness): --rebuild forces recreate + includes infra profile
`bin/showcase test --rebuild` was either a silent no-op or an error
depending on container state:

- runner.run() only rebuilt services in `autoStarted` (those NOT already
  running), so --rebuild against an already-up service did nothing and a
  stale container was silently reused (the 36h-stale-image false-positive).
- lifecycle.rebuild([slug]) ran `compose --profile <slug> build` WITHOUT
  `--profile infra`, leaving `aimock` undefined → "service <slug> depends
  on undefined service aimock: invalid compose project" when the service
  was down.

Fix both:
- runner.run() now rebuilds every targeted slug when --rebuild is set,
  regardless of running state, and health-checks rebuilt-but-already-running
  services before probing.
- rebuild() includes `--profile infra` alongside the slug profiles (mirrors
  up()) and force-recreates the targeted containers so the freshly-built
  image is actually adopted.

Adds red-green unit coverage asserting the infra profile and force-recreate
in the rebuild compose invocation.
2026-06-04 07:11:58 -07:00
lukasmoschitz 5de391fa27 fix(core): make setRuntimeTransport idempotent on the requested transport mode (#5179)
## Summary

`setRuntimeTransport` was not idempotent on the **requested** transport
mode. Auto-detect resolves the requested `"auto"` to a concrete
transport (`"rest"`/`"single"`) and writes that back to
`_runtimeTransport`; the guard then compared against that **resolved**
value. So re-applying the same requested mode — which the provider
effect does on **every render** — compared unequal and **re-ran the
entire `/info` handshake**, rebuilding the runtime agents mid-session.

When that re-sync lands during a turn, `useAgent` hands the UI a
freshly-rebuilt (empty) agent for a render, **blanking the whole
transcript** (and any per-message UI bound to it — e.g. the intelligence
indicator) until it replays.

## Fix

Track the **requested** mode separately (`_requestedTransport`) and
guard on it. Re-applying an unchanged requested transport — including
`"auto"` after auto-detect has resolved it — is now a no-op, so no
redundant `/info` re-sync fires.

## Tests

`packages/core/src/__tests__/agent-registry-resync.test.ts`: re-applying
`"auto"` after auto-detect resolves it does **not** refetch `/info`.

## Scope / risk

Touches only `packages/core` (`core/agent-registry.ts` + the test). **No
public API change.** Genuine transport *changes* still re-sync exactly
as before; only redundant re-applications of the same requested mode are
skipped.

> **Verified live:** this idempotency guard *alone* eliminates the
mid-turn re-sync in the e-commerce intelligence demo — instrumentation
shows no `/info` handshake fires during a turn, and the
transcript/indicator no longer flickers. (An earlier draft also made the
re-sync itself non-destructive as defense-in-depth; that was dropped
because, with this guard, no mid-turn re-sync occurs for it to protect
against.)
2026-06-04 11:56:47 +02:00
Jordan Ritter 798ed56b03 fix(showcase/aimock): reorder gen-ui-interrupt d6 fixture legs so turn-2 emits the tool call (langgraph-python+typescript)
The gen-ui-interrupt d6 probe drives two interrupt turns in one thread.
Turn-2 failed because the langgraph-python/typescript fixtures ordered each
pill's toolCallId resume-leg BEFORE its toolName first-leg. aimock dedups on
userMessage (first-match-wins), so the text-only resume-leg shadowed the
tool-emitting first-leg on the duplicate userMessage — the agent emitted a
plain text reply with no schedule_meeting tool call, the interrupt never
fired, and the time-picker-card never mounted.

Reorder each pill so the toolName:schedule_meeting first-leg precedes its
toolCallId resume-leg, matching the passing mastra/pydantic-ai ordering. The
resume-leg still only matches when the last message is its tool result
(toolCallId guard), so confirmation still works.
2026-06-04 01:53:16 -07:00
Jordan Ritter be8a5c04ea fix(showcase/dashboard): render Starter row-group in the live FeatureGrid (was in dead CellMatrix)
PR #5226 added the "Starter" pseudo-category row-group (StarterSection,
using resolveStarterRow/buildStarterBadge) only to CellMatrix in
cell-matrix.tsx. But the live DashboardPage renders the matrix tab via
FeatureGrid (feature-grid.tsx); CellMatrix is reached only through
cells-view.tsx, which no live route imports — it is dead/legacy. So the
Starter row shipped in code but could never render (the served prod bundle
had zero StarterSection/starter-row-/"no starter" strings).

Port the StarterSection into FeatureGrid: render the four fixed sub-rows
(health/agent/chat/interaction) after the feature categories, resolving
each cell via the existing S3 helpers (resolveStarterRow + buildStarterBadge,
5-state vocab) and structurally excluded from the feature rollup/column tally
(never calls renderCell/buildCellModel). Aligned to FeatureGrid's table
structure (Feature col + optional parity ref-depth spacer + categoryColSpan).
Remove the duplicate from the dead CellMatrix and move its dedicated tests
to feature-grid.test.tsx, which exercises the live registry-backed grid.

Render proof: production next build now bakes starter-row-/starter-cell-/
"no starter for this integration" into the served page.js (client + server)
bundles. Full dashboard suite green (815 passed, 1 skipped).
2026-06-04 01:17:59 -07:00
Jordan Ritter de6a53304d fix(examples/integrations/crewai-crews): stop @chat crash from non-serializable useFrontendTool dep
The updateProverb useFrontendTool passed the live `agent` object in its
deps array. react-core's use-frontend-tool does JSON.stringify(extraDeps),
which throws "Converting circular structure to JSON" on agent.activeRunDetach$
(an RxJS Subject), crashing the React tree so no assistant bubble renders.

`agent`/`agent.setState` are stable references; only `state` needs to be a
dep. Also bump the docker ag-ui-crewai override to >=0.2.0,<0.3.0 to match
agent/requirements.txt and the showcase backend.
2026-06-04 01:16:24 -07:00
Jordan Ritter 07f0cb2dc8 fix(examples/integrations/mastra): open the CopilotSidebar by default
Add defaultOpen={true} to mastra's CopilotSidebar so it opens on load,
matching every other starter.
2026-06-04 01:16:24 -07:00
Jordan Ritter f992e06737 fix(examples/integrations): replace default create-next-app metadata
Swap the boilerplate "Create Next App" / "Generated by create next app"
title and description in each starter layout for a per-framework
"<Framework> + CopilotKit Starter" title and matching description.
2026-06-04 01:16:24 -07:00
Jordan Ritter 1ce71c8744 fix(examples/integrations): normalize AGENT_URL handling in runtime routes
crewai-flows hardcoded the agent URL with no AGENT_URL env override; add
the fallback to match siblings. agno and llamaindex concatenate a path
suffix onto the base URL, so strip any trailing slash first to avoid a
double slash when AGENT_URL ends in "/".
2026-06-04 01:16:24 -07:00
Jordan Ritter c6ae28ebac fix(examples/integrations/crewai-crews): wire the tools the welcome copy advertises
The welcome message promised theme/proverb/weather tools, but the page
registered none of them (only a debug console.log of agent.state).
Mirror the crewai-flows sibling to wire setThemeColor, the proverbs
shared-state tool, and a weather render tool, remove the leftover
console.log effect, and drop the stray w-screen overflow.
2026-06-04 01:16:24 -07:00
Jordan Ritter 20795fd3a5 fix(examples/integrations): read agent.state at call time in add_proverb
The agno and llamaindex add_proverb frontend tools closed over a stale
`state` snapshot with no dependency array, so rapid successive adds
dropped earlier proverbs. Read agent.state at call time inside the
handler instead. Also drop the stray w-screen on the llamaindex page
that caused horizontal scrollbar overflow.
2026-06-04 01:16:24 -07:00
Jordan Ritter 7d80ce4a9e fix(examples/integrations/langgraph-fastapi): bump ag-ui-langgraph to 0.0.37 in pyproject/uv.lock so uv sync resolves
The smoke build path (docker/Dockerfile.agent) runs `uv sync` against
agent/pyproject.toml + agent/uv.lock, but those still pinned
ag-ui-langgraph==0.0.34 (and the lock still pinned copilotkit==0.1.87).
copilotkit 0.1.93 requires ag-ui-langgraph[fastapi]>=0.0.35, so uv sync
failed with "No solution found" and the langgraph-fastapi smoke job went
red. The prior fix only bumped the production Dockerfile's
`uv pip install`, which the smoke path does not use.

Bump pyproject to ag-ui-langgraph[fastapi]==0.0.37 and regenerate uv.lock
(copilotkit 0.1.87->0.1.93, ag-ui-langgraph 0.0.34->0.0.37). Verified by
building docker/Dockerfile.agent locally: uv sync now resolves and the
agent boots ("Uvicorn running on 0.0.0.0:8123" / "Application startup
complete") serving HTTP 200 on /health.

Also reconcile parity: strands-python's entrypoint.sh legitimately diverges
from the langgraph-python north-star (it boots the strands agent via
`uv run python main.py` and serves Next via `next start`, not serve.py +
standalone server.js). Add entrypoint.sh to strands-python's
allowedDivergence, mirroring how langgraph-js already declares its own
custom entrypoint. parity:check now reports 0 errors.
2026-06-04 01:10:12 -07:00
Jordan Ritter ff9210ba6c fix(examples/integrations/strands-python): boot the strands agent, not langgraph serve.py
The entrypoint was copy-pasted from langgraph-python: it logged
'langgraph-python starter', ran 'python serve.py' (no serve.py exists in this
starter), and started Next via standalone 'node server.js' (this image is a
non-standalone .next + node_modules build, so server.js does not exist either).
The container crash-looped on boot with
"python: can't open file '/app/serve.py'".

Run the Strands agent's self-serving agent/main.py via its uv venv
(AGENT_PORT=8123) and serve the frontend with 'next start'. Verified by booting
the image: agent reaches 'Application startup complete' on :8123 and the
container serves HTTP 200 on /.
2026-06-04 01:10:12 -07:00
Jordan Ritter 9653cdb4f0 fix(examples/integrations): bump ag-ui-langgraph to 0.0.37 for copilotkit 0.1.93
copilotkit 0.1.93 re-exports StateStreamingMiddleware/StateItem FROM
ag_ui_langgraph.middlewares.state_streaming, which only exists in
ag-ui-langgraph>=0.0.35. The Dockerfiles installed copilotkit with --no-deps
and pinned ag-ui-langgraph[fastapi]==0.0.22, so the middlewares submodule was
missing and the agent crash-looped at boot with
'ModuleNotFoundError: No module named ag_ui_langgraph.middlewares'.
Bump the pin to 0.0.37 (the version copilotkit 0.1.93 resolves with deps).

Verified by booting the langgraph-python image: the agent now reaches
'Application startup complete' / 'Uvicorn running on 0.0.0.0:8123' with no
ImportError, and the container serves HTTP 200 on /.
2026-06-04 01:10:12 -07:00
Jordan Ritter 25811cf94f fix(examples/integrations/strands-python): copy showcase.json into build
The strands-python frontend imports `showcaseConfig from "../../showcase.json"`
(via src/hooks/use-example-suggestions.tsx) but the Dockerfile frontend stage
never copied showcase.json into the build context, so `next build` fails with
"Module not found: Can't resolve '../../showcase.json'". Add the COPY step,
mirroring the langgraph-js/langgraph-python Dockerfiles that already copy it.

The @copilotkit/react-core v2 pin (1.59.3) was already correct and matches all
working v2 starters; no package.json change was needed.
2026-06-04 01:10:12 -07:00
Jordan Ritter 0a740a9819 fix(examples/integrations): pin copilotkit python 0.1.93
The build-starters Docker images hardcoded copilotkit==0.1.78 in the
langgraph-python and langgraph-fastapi Dockerfiles, but their agents import
StateStreamingMiddleware and StateItem, which 0.1.78 does not export. This
crash-loops the Python agent on boot with an ImportError. Bump the hardcoded
pin to 0.1.93 (latest stable, re-exports both symbols, matches local
sdk-python). Align all three langgraph-family agent pyproject pins
(0.1.87 -> 0.1.93) for consistency and regenerate the strands uv.lock.
2026-06-04 01:10:12 -07:00
Jordan Ritter c4b66c6eba fix(showcase/harness): size d6-all-pills probe budget for the 18-service fleet
The d6-all-pills-e2e probe header was sized for a legacy "8 services x 4
features" single-round fleet (timeout_ms: 600000 / 10 min). Discovery now
enumerates ~18 showcase demo services, so at max_concurrency: 8 the fan-out
runs ceil(18/8) = 3 serialized rounds. Rounds 2-3 start late and execute under
CPU contention, stretching each service toward ~200s and blowing the 10-min
per-service budget — all 18 services went red with `driver timeout after
600000ms`.

Raise the outer cap to 1200000 (20 min), matching the sibling 18-service
e2e-demos probe (same fleet, same budget). 3 rounds x ~200s fits comfortably.

Timeout-only change: max_concurrency stays at 8, so D6 peak stays at 32
contexts (under BROWSER_POOL_MAX_CONTEXTS=40, leaving 8 for the offset D5
tick). Header comment rewritten to document the real fleet size and the
fan-out math.
2026-06-04 01:00:23 -07:00
github-actions[bot] 6a7199e3da style: auto-fix formatting 2026-06-04 00:24:11 -07:00
Jordan Ritter 2d5e198502 ci(showcase): tolerate starter-* dispatch + keep 6h smoke cron
showcase_build.yml's existing detect-changes job shared the workflow_dispatch
`service` input with the new detect-starter-changes job, so dispatching
`service=starter-<slug>` tripped detect-changes's fail-loud `exit 1` ("did
not match any entry in ALL_SERVICES") and reddened the run even though
build-starters published fine. Scope the showcase fail-loud to non-`starter-*`
inputs via a case statement (mirroring how the starter job scopes its own
fail-loud) so a `starter-*` dispatch resolves to an empty showcase matrix and
SKIPS; typo'd showcase service names still fail loud.

Restore the 6h `schedule` cron in test_smoke-starter.yml. The PR had removed
it, leaving NO post-merge floating-dependency breakage detector for starters
— and its harness-probe replacement depends on S5 Railway services that don't
exist yet. Keep the cron until S5 starter-service probing is confirmed live.
2026-06-04 00:24:11 -07:00
Jordan Ritter 7c560d00e4 fix(showcase/harness): unmapped aggregate count + abort/writer-throw tests
The unmapped-starter aggregate emitted `total: 0` while `failed` listed all
four levels — a broken count invariant. Set `total: STARTER_LEVELS.length`
to match. Add the two missing red-green unit tests: the `aborted` errorClass
branch (already-aborted ctx.abortSignal → levels classified `aborted`, not
`transport-error`) and the makeSideEmit writer-throw resilience (a writer
whose write() rejects is swallowed at error-level and must not fail the
aggregate tick).
2026-06-04 00:24:11 -07:00