Commit Graph

15424 Commits

Author SHA1 Message Date
Ran Shemtov 9b768a0b98 feat(showcase): finalize MAF Python - D6 green on agent-framework 1.0 latest (#5985)
## Finalize MAF Python: D6 green on official agent-framework 1.0 latest

Brings the `ms-agent-python` showcase integration to a clean,
reproducible D6 state on the officially published latest
`agent-framework` packages, with feature parity to `langgraph-python` on
everything buildable today.

### Dependencies (exact pins, official latest)

- `agent-framework-ag-ui==1.0.1`
- `agent-framework-openai==1.12.0`
- `agent-framework-core==1.13.0`

No beta/rc floors, no ranges. Removed two unused `langchain-*` deps. All
framework deps are exact pins; `validate-pins` ratchet baseline moves
down 31 to 27. The only remaining ms-agent-python pin FAIL is the
shared-frontend `openai ^5.9.0`, identical across every integration
(pre-existing baseline).

### D6 result: all green on the published mock

Verified with `showcase test ms-agent-python --d6 --direct --rebuild`
against the actual published `ghcr.io/copilotkit/aimock:latest`
(**v1.38.0**), freshly pulled: 37 distinct cells executed, 37
conversations completed, zero failures, aggregate `d6:ms-agent-python
green (104.2s)`.

`tool-rendering-reasoning-chain` (previously the only red on the
published mock) is now green: it needed `reasoning.encrypted_content`
echoed back on the second Responses request (upstream
microsoft/agent-framework#7233), which the published mock did not
synthesize until
[aimock#342](https://github.com/CopilotKit/aimock/pull/342), shipped in
aimock **v1.38.0**. Fixed upstream, not worked around.

`multimodal` is un-quarantined and now matches langgraph. It had been
wrongly marked unsupported based on a local-only failure: the
`sample.png`/`sample.pdf` demo assets are Git LFS pointers, and without
git-lfs on PATH the attachment send fails before the run starts
(`runStartCount=0`). langgraph-python multimodal fails locally for the
identical reason yet declares the feature supported. Verified the MAF
agent works (D6 cell green with the real assets, 2 turns, assertions
passed); both production deploys serve the real 10KB PNG.
`not_supported_features` now equals langgraph exactly:
`[gen-ui-interrupt, interrupt-headless]` (both a shared
`@copilotkit/react-core/v2` resume-path bug, quarantined in langgraph
too).

### Cells fixed on this branch

- `tool-rendering-custom-catchall` (18-entry fixture +
MESSAGES_SNAPSHOT-drop subclass so narration renders last)
- `shared-state-streaming` (seed `/document` after RUN_STARTED +
`chunkSize` fixtures so replay emits per-token deltas)
- `tool-rendering-reasoning-chain` (un-quarantined; green on aimock
v1.38.0)
- `frontend-tools-async` (removed a stray broad fixture that
shadowed/looped)
- `open-gen-ui` + `open-gen-ui-advanced` (removed six stray fixtures
colliding in the shared gen-ui fixture file)
- `multimodal` (un-quarantined; parity with langgraph)

### Deferred to upstream (not worked around)

- **a2ui-recovery**: langgraph ships a bespoke A2UI validate-and-retry
recovery demo. MAF Python's A2UI is going native via
[microsoft/agent-framework#7423](https://github.com/microsoft/agent-framework/pull/7423),
which delivers progressive streaming, error recovery, and the sub-agent
design built into `agent-framework-ag-ui`, and even includes the same
two bridge fixes hand-rolled here (unanswered-tool-call stripping + A2UI
MESSAGES_SNAPSHOT suppression). Building a bespoke recovery demo now
would be throwaway. When #7423 merges and releases, the showcase A2UI
migrates to the native path and the recovery demo lands with it.

### Validators

- `generate-registry`: OK
- `validate-pins`: 27 fails, hash matches ratcheted baseline
- `validate-parity`: PASS
- `validate-fixture-tool-surface`: clean

### Notes

- `useCoAgent` is deprecated; all demos use `useAgent` from
`@copilotkit/react-core/v2`.
- Kept in draft pending review. No blocking external gates: aimock#342
shipped in v1.38.0.
2026-08-06 12:13:09 +02:00
Ran Shemtov 464550ef68 fix(runtime): skip value-less activity patch for null open-gen-ui params (#6396)
## Problem

In `open-generative-ui-middleware.ts`, when the LLM emits `jsFunctions`
(or `css`) as `null`/empty, `setParam` sets the param to `undefined`,
then `emitParamDelta` produces a JSON Patch op `{op:"add",
path:"/jsFunctions"}` with **no `value` property**.

`fast-json-patch` rejects this client-side with
`OPERATION_VALUE_REQUIRED` and drops the whole activity patch:

```
Failed to apply activity patch ... Operation `value` property is not present ... path /jsFunctions
```

Benign today (the dropped delta carried nothing) but noisy in the
console and fragile.

**Repro:** `open-gen-ui-advanced` on a live LLM (any framework); the
model frequently emits an empty `jsFunctions`.

## Fix

Guard `emitParamDelta` to skip emitting when `value === undefined`. This
covers all three callers (`jsFunctions`, `css`, and the delayed
`initialHeight` delta). Empty arrays (`[]`) and completion markers
(`jsFunctionsComplete: true`) still emit as before.

## Test

Adds a regression test feeding a `null` jsFunctions value and asserting:
- no emitted patch op is missing its `value` property
- the value-less `/jsFunctions` delta is skipped entirely
- the `/jsFunctionsComplete` marker still fires

`nx test runtime` → 22 passed. `nx check-types runtime` clean.

## Release note

This is in the published `@copilotkit/runtime` source; a runtime release
is needed before showcase consumes it.
2026-08-06 12:12:05 +02:00
Murat Sari 89a5c4503c Revert change in unrelated area 2026-08-06 12:05:33 +02:00
Murat Sari 4d5e3da712 feat(chat): implement input height measurement and adjust scroll view styles 2026-08-06 12:04:54 +02:00
Alem Tuzlak 19c0c109f8 fix(showcase): let the Mastra MCP Apps agent self-correct a rejected diagram (#6398)
## Problem

PNI-118: the Mastra **MCP Apps** cell intermittently renders an **empty
iframe** (an empty box or a thin band) on a live endpoint, while passing
under aimock. Reproduced on `showcase-mastra-staging`.

## Root cause

Not the renderer, not a delivery race, not a version pin.

Excalidraw's `create_view` takes `elements` as a **stringified** JSON
array:

```json
{"elements": {"type": "string", "description": "JSON array string of Excalidraw elements. Must be valid JSON ..."}}
```

So the model has to hand-escape nested JSON, and it appends a stray `}`
just past the closing `]`:

```
tail:  ...,"width":800,"height":600}]}
                                    ^ stray brace
```

The MCP server rejects the call, returns `isError: true`, and there is
no diagram to draw, so the iframe paints empty.

The agent's raw tool-call args arrive over SSE as **valid** JSON, so
streaming and arg assembly are healthy. The stray brace sits inside the
`elements` string value, written by the model.

## Fix: let the agent self-correct

Swapping models only moved the failure rate around (gpt-4o-mini ~63% of
diagrams failed, gpt-5.4 ~30%), so this stops depending on one-shot
accuracy.

The server's error already names the exact fault and already comes back
as a tool result, and the agent had no step cap, so a retry was
mechanically possible all along. **What blocked it was our own prompt:**
`"Call create_view ONCE"` and `"do NOT iterate, do NOT make multiple
calls. Ship on the first shot."`

Now the prompt tells the model to read the error and try again, capped
at **2 corrections (3 calls total)**, with `stopWhen: stepCountIs(6)`
bounding the loop if it never converges. This mirrors the
validate-then-retry recovery pattern already used for A2UI on the other
integrations.

The agent also moves to `gpt-5.4` (owner preference for the 5.x line).

## Validation

Against the **real Excalidraw MCP server**, using the agent's prompt
extracted verbatim from this file and the real tool schema, over the
Responses API (the path the AI SDK actually uses):

| scenario | result |
| --- | --- |
| normal runs | **12/12 succeeded**, all on the first call |
| attempt 1 force-corrupted with the real-world stray `}` | **10/10
recovered on the second call** (`err > ok` every trial) |

Model comparison that motivated moving away from one-shot (create_view
against the real server):

| model | OK | isError |
| --- | --- | --- |
| gpt-4o-mini (before) | 3 | 5 |
| gpt-5.4 (no retry) | 7 | 3 |
| gpt-4.1 | 8 | 0 |
| gpt-5.5 | 10 | 0 |

Prompt hardening alone was measured and does **not** fix it (gpt-4o-mini
8/12 to 7/12 invalid; gpt-5.4 still 2/16).

Also verified in the running app (local dev server, real key): valid
JSON, `isError: false`, diagram rendered.

**Not yet verified in-app:** the recovery path itself. No natural
failure occurred during the in-app runs, so the retry is proven at the
API level rather than through the Mastra agent loop.

## Why aimock never caught this

- aimock replaces the LLM and the fixture hard-codes the `create_view`
tool call, so no model-generated JSON is involved. This bug cannot occur
under replay.
- The shared d5/d6 probe (`harness/src/probes/scripts/d5-mcp-apps.ts`)
asserts only that an iframe element exists
(`[data-testid="mcp-app-iframe"] || iframe[sandbox]`) - "the page
renders an iframe shell". It never asserts a diagram rendered, and a
**rejected payload still mounts an iframe**. So the probe is blind to
this class of failure by construction.

The fixture is unaffected by this change:
`showcase/aimock/d6/mastra/mcp-apps.json` contains no model or `gpt`
reference.

## Follow-ups (not fixed here)

1. **The mcp-apps fixture payload is itself schema-invalid.** Its
`elements` is a JSON **array**, but `create_view` requires a **string**.
Sent verbatim to the real server it returns `-32602 Invalid arguments
for tool create_view: expected "string"`; stringified, the same payload
succeeds. No test overrides `MCP_SERVER_URL`, so aimock runs still call
the real server.
2. **`MCPAppsActivityRenderer` ignores `isError`** and still renders an
empty sandbox iframe, which is why a rejected tool call looks like a
silent blank box. Lives in the published `@copilotkit/react-core`, so it
would need a release.
3. **The shared probe could assert the diagram actually drew**, not just
that an iframe exists. That is one shared probe across every integration
(iron rule 1), so it is broader than this ticket.
2026-08-06 10:54:34 +02:00
Ran Shemtov adb2848db7 Merge branch 'main' into claude/jolly-boyd-38b55c 2026-08-06 08:49:55 +02:00
xiaoqinvar ed559fc4c2 fix(core): preserve fresh event run IDs after completed runs 2026-08-06 14:33:32 +08:00
Ran Shemtov a7006cd1ed Merge branch 'main' into claude/elated-snyder-01a8ce 2026-08-06 08:09:08 +02:00
xiaoqinvar 3fce5ecec4 Merge upstream main into fix/state-manager-run-id 2026-08-06 10:54:23 +08:00
Ran Shem Tov 3c4e91e55f test(showcase): ratchet CrewAI fixture aliases 2026-08-06 03:53:28 +03:00
David McKay 2d05114160 fix(core): stop HITL continuations reusing the originating run id on the wire (#6411)
Linear:
[CPK-7786](https://linear.app/copilotkit/issue/CPK-7786/hitl-continuation-reuses-the-originating-run-id-on-the-wire-breaking)

Regression fix. **Nothing from #6296 is reverted** — its goal is kept
and moved one layer up.

## What broke

#6296 (@rodboev) preserved the logical run id across a HITL resolve by
pinning the originating id on the follow-up's agent invocation,
correctly fixing #3456 (external tracing saw one logical run split into
two halves).

Pinning it **on the wire**, though, made the transport treat the
follow-up as a resumption of a run it had already finished:

- it re-delivered that run's already-applied half, duplicating every
tool call on the message — each duplicate carrying **empty arguments**,
because a start event has none and the `TOOL_CALL_ARGS` deltas that
follow are addressed to the first copy; and
- the follow-up's own tool call never reached client state, so its card
never rendered.

In `reskinnable-demo`'s banking skin that killed teach mode: the agent
called `awaitDashboardDemonstration`, the server emitted
`TOOL_CALL_START` for it, and the live "Recording your workflow" card
never appeared — no REC indicator, no step feed, no "I'm done", so a
demonstration could not be finished or saved.

## Evidence

- Intelligence event log shows **two `RUN_STARTED`/`RUN_FINISHED` cycles
under one run id**, with `TOOL_CALL_START [awaitDashboardDemonstration]`
in the second — the server does emit it.
- Instrumented `useRenderToolCall`: **never invoked** for that tool,
despite its renderer being registered.
- Client message state after the click: `assistant
tc:["recall_memory","recall_memory"]`, `assistant
tc:["offerWorkflowRecording","offerWorkflowRecording"]` — prior cycle
re-applied, new call absent.
- Bisect: reverting #6296's four source files makes the card appear;
restoring them breaks it again. No later commit touches those files.
- Ruled out: React StrictMode (disabled in that app),
model/prompt/tool-availability differences (byte-identical between the
working and broken apps), and the view-layer dedupe from #6407.

## The fix

`markNextRunAsContinuation` already accepted an `expectedRunId` that was
never used. It now records it, and the state manager re-stamps the
continuation's events onto that id. So:

- **logical identity is preserved** — state/message association and
external tracing still see one run, which is #6296's whole point;
- **the wire identifies the invocation honestly** — the follow-up no
longer claims to be a run that already finished, so nothing is
re-delivered and the continuation's own tool call lands.

## Test changes — please review this part closely

`core-follow-up`'s run-id test asserted the *mechanism* (both
invocations carry the same wire id), which this deliberately changes. It
now asserts the *goal*: the originating id is pinned on the first
invocation, and the follow-up leaves the id to the transport. Its
sibling assertion — the thread still knows exactly one run — was already
there and passes untouched.

A new `StateManager` test covers the re-stamp directly, verified **red
before green** by dropping the `expectedRunId` lookup (it fails, along
with one of #6296's own tests).

@rodboev — flagging you directly since this touches your change from
today. If the wire-level pinning was load-bearing for something I have
not seen, say so and I will rework it.

## Verification

- `@copilotkit/core` 58 files and `@copilotkit/react-core` 123 files
pass.
- In-browser against a live Intelligence stack: before, the recording
card never rendered; after, it renders with its REC indicator and I'm
done / Cancel controls.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-08-05 16:36:41 -07:00
Maxim 1972599063 docs(reskinnable-demo): document the LOCK_SKIN single-tenant gate
Covers the two things a reader would otherwise assume wrongly: it does NOT pin
dark/light (separate axis), and it does NOT hide the inspector — a locked deploy
still shows it, which is the intended FDE configuration.

Says what the gate actually governs: the UI and routing expose only that skin.
It is a presentation/deploy gate, NOT a security boundary — all four agents stay
registered server-side, so another skin's agent endpoint remains reachable under
a lock.

Also corrects the stale "floating selector at the bottom-left" description; the
switcher is a dropdown at the top of the assistant column.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-06 01:24:34 +02:00
David McKay 6188404f3a test(react-core): assert the legacy HITL follow-up goal, not the wire id
The sibling of the core-follow-up assertion, missed in the previous commit: it
required the follow-up invocation to repeat the originating run id on the wire,
which is exactly what this change stops doing.

It now asserts the follow-up happened and left the id to the transport. Logical
identity is covered where it now lives — StateManager's re-stamp test.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 16:24:28 -07:00
Maxim 3b548cb1c7 docs(reskinnable-demo): note why inspector identity ignores LOCK_SKIN
The inspector's agentId-less /memories and /info requests stay keyed to
defaultSkinId even under a lock, so on a deploy locked to a non-default skin they
resolve a different scope than the running agent. Deliberate: the default
resolver is the one whose scope is seeded, so switching to the locked skin's
resolver would read empty on any skin without seed data. Only banking ships real
durable memory and it is also the default, so the two align in the configuration
that matters.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-06 01:24:28 +02:00
Maxim 779bd501d0 test(reskinnable-demo): pin LOCK_SKIN off for the e2e dev server
The suite visits /airline and asserts all four switcher options, so a developer
with LOCK_SKIN set locally would watch it fail for reasons that look nothing like
the cause.

The pin only covers a server Playwright STARTS. reuseExistingServer means a warm
local run adopts an already-running pnpm dev and skips the whole env block, so
the comment documents both shapes: a non-banking lock 404s the hardcoded
/banking readiness probe and aborts at webServer startup, while a banking lock
gets through and fails the /airline assertions instead. In CI reuseExistingServer
is false, so the pin always applies.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-06 01:24:17 +02:00
Maxim 8b5de10b11 feat(reskinnable-demo): collapse the switcher to a brand badge when locked
The dropdown is not rendered at all — no trigger, no chevron, no options in the
DOM. A disabled dropdown was rejected: it implies a choice that does not exist
and reads as a bug rather than as a single-tenant product. The badge is a div
with cursor:auto, no handler and tabIndex -1, so there is no dead control to
click or tab onto.

The identity block is defined once and rendered into either a button or a plain
div, so the two modes cannot drift apart. Swap-sides, hide, the skin-selector
testid the layout e2e keys off, and useSkinThemeReconcile's root all stay put.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-06 01:23:54 +02:00
Maxim 635900c73f feat(reskinnable-demo): brand the locked deploy's SSR title and description
The tab read "CopilotKit Reskinnable Demo" beside the locked skin's own favicon,
leaking both "CopilotKit" and "demo" on the most visible surface a prospect sees.
generateMetadata brands the title AND the description from the locked skin, so
crawlers and link unfurlers see a coherent product page. A client effect cannot
do this: Next applies route metadata after hydration, so SSR always shipped the
demo strings.

force-dynamic here too. The root layout reads LOCK_SKIN and
PRESENTER_RESET_ENABLED per request and threads both into client gates;
correctness otherwise rested on the implicit invariant that every descendant
route happens to be dynamic. Cost is one dynamically-rendered /_not-found.

Unlocked metadata is byte-identical in both fields.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-06 01:23:51 +02:00
Maxim f113cf69d6 feat(reskinnable-demo): 404 every skin but the locked one
Enforcement is one extra notFound() condition in SkinLayout. No middleware is
needed: notFound() throws during render, so a disowned skin never mounts a
provider, a thread or an agent registration. A redirect WOULD have needed
request-time middleware, and a route matcher there risks intercepting
/api/copilotkit's SSE stream.

Under a lock, /airline is as absent as /nope — uniform 404 semantics. / now
redirects to the locked skin, without which a locked deploy's front door would
land on defaultSkinId and 404.

force-dynamic on / is not optional: reading process.env is not a dynamic API, so
next build otherwise prerenders / and bakes the build-time skin into the
redirect. A deploy built unset then run locked sent / to a 404 front door.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-06 01:23:44 +02:00
Maxim 8a7250bccc feat(reskinnable-demo): thread the LOCK_SKIN gate to client chrome
Same shape as the presenter-reset gate: the server env is read in the root
layout and passed down through a small context. The context default is null
(unlocked) so any subtree without the provider — including SelectorCard's bare
unit tests — behaves exactly as before.

isSkinLockedOut is a named predicate rather than an inline comparison because
inverting it would 404 every skin on an UNLOCKED deploy. Extracted, it gets
exhaustive mutation-sensitive tests without rendering SkinLayout and mounting
CopilotKitProvider.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-06 01:23:21 +02:00
Maxim 35b267dc9e feat(reskinnable-demo): add lockedSkinId(), the LOCK_SKIN reader
A per-deploy SERVER env, deliberately non-NEXT_PUBLIC_ like
PRESENTER_RESET_ENABLED, so one build serves both a locked single-tenant host
and the unlocked four-skin demo.

Throws on an unrecognised id rather than falling back to unlocked: silently
accepting a typo would 404 every skin AND send / to a 404 too, leaving the whole
app dark with nothing pointing at the cause.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-06 01:23:16 +02:00
Maxim ca5c84bb76 feat(reskinnable-demo): expose skin ids and identities as import-free config
LOCK_SKIN must be validated, and a locked deploy's brand and tagline resolved,
from server components. Those cannot import registry.ts — it pulls in four
client skin modules. skins-config.ts is the import-free home for that, so the
data is duplicated there and fenced by a drift guard asserting it matches the
registry, which is what stops the copy rotting.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-08-06 01:23:07 +02:00
Tyler Slaton ac8b3bb6d0 fix(channels): make managed Slack DM replies reliable (#6368)
## Summary

- Start managed Slack DM status in the Slack thread. The server uses
`messageTs` as the native status and streaming anchor.
- Use legacy message create and replace only when native stream start
fails. After native output opens, append and stop errors fail the run.
- Keep a separate provider reference for each legacy long-message chunk.

## Test plan

- Fresh `@copilotkit/channels-slack` package tests: 383 passed.
- Fresh `@copilotkit/channels-intelligence` package tests: 192 passed.
- Type checks passed for both packages.
- Builds passed for both packages.
- Formatter check passed.
- `git diff --check` passed.

No live Slack testing was run.
2026-08-05 16:13:49 -07:00
David McKay 696c44244b fix(core): stop HITL continuations reusing the originating run id on the wire
#6296 preserved the logical run id across a HITL resolve by pinning the
originating id on the follow-up's agent invocation. That fixed #3456 (external
tracing saw one logical run split into two halves), but pinning it on the WIRE
made the transport treat the follow-up as a resumption of a run it had already
finished. It re-delivered that run's already-applied half — duplicating every
tool call on the message, each duplicate carrying empty arguments, since a start
event has none and the TOOL_CALL_ARGS deltas that follow are addressed to the
first copy — and the follow-up's own tool call never reached client state, so
its card never rendered.

In the reskinnable-demo banking skin that broke teach mode outright: the agent
called awaitDashboardDemonstration, the server emitted TOOL_CALL_START for it,
and the live "Recording your workflow" card never appeared, leaving no way to
finish or save the demonstration.

#6296's goal is kept, moved one layer up. The continuation is registered against
the originating id (markNextRunAsContinuation already took an expectedRunId
parameter, previously unused) and the state manager re-stamps the continuation's
events onto it. State/message association and external tracing still see ONE
logical run; the wire is simply allowed to identify the invocation honestly.
Nothing from #6296 is reverted.

core-follow-up's run-id test asserted the mechanism (both invocations carry the
same wire id), which this deliberately changes, so it now asserts the goal: the
originating id is pinned on the first invocation and the follow-up leaves it to
the transport. Its sibling assertion — the thread still knows exactly one run —
was already there and still passes untouched. A new StateManager test covers the
re-stamp directly; verified red before green by dropping the expectedRunId
lookup.

Verified in the browser against a live Intelligence stack: before, the recording
card never rendered; after, it renders with its REC indicator and I'm done /
Cancel controls. `@copilotkit/core` 58 files and `@copilotkit/react-core` 123
files pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 16:12:12 -07:00
Tyler Slaton 0cf899095b fix(channels): take the legacy path when a dropped stream start settles as applied
The gateway settles a direct-message slack.stream.start provider failure
as applied with capabilityError and no provider reference. Parse the
start result inside the fallback try so that shape reaches the legacy
create instead of hard-failing the stream body.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-05 15:44:06 -07:00
Alem Tuzlak 85d30dd327 fix(channels): keep managed Slack DM replies visible 2026-08-05 15:44:06 -07:00
Benjamin Taylor a3c5f079ed fix(react-ui): let sidebar children fill the viewport height
Children of `CopilotSidebar` cannot use `height: 100%`. Both wrappers the
sidebar puts around your app -- `.copilotKitSidebarContentWrapper` and
`.copilotKitModalChildrenWrapper` (which had no CSS rule at all) -- are
auto-height blocks, so a percentage height on a child has no definite
containing block and collapses to content height.

Add an opt-in `fullHeightChildren` prop that gives the content wrapper a
one-viewport height and lets the children wrapper fill it. It is opt-in
because the content wrapper wraps the entire consumer app, and giving
every react-ui sidebar user a flex column with a fixed height would
reflow apps that never asked for it.

The height is a viewport unit, not `100%`: `100%` only resolves when
every ancestor (html/body/#root) also declares a height, which react-ui
neither sets nor can guarantee, so it would silently no-op in a stock
Next.js app. `min-height: 0` on the children wrapper clears the flex-item
`min-height: auto` floor so tall content scrolls inside the child instead
of stretching the wrapper past the viewport.

Fixes #261

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Ashish Shaw <77574570+ashish4143@users.noreply.github.com>
2026-08-05 16:54:47 -05:00
Adrien Pouligny add2a80fd8 Merge branch 'main' into fix/runtime-deprecated-uuid-dependency 2026-08-05 14:39:55 -07:00
Ben Taylor ca9a481efd fix(web-inspector): show step lifecycle events (#6323)
## What does this PR do?

The Web Inspector did not record `STEP_STARTED` or `STEP_FINISHED` from
live agent events.

This PR:

- records both step lifecycle events
- adds both events to the Inspector filter
- adds a regression test

## Related PRs and Issues

- Fixes #6324

## Testing

- `pnpm nx test @copilotkit/web-inspector --skip-nx-cache`
- `pnpm nx run @copilotkit/web-inspector:check-types --skip-nx-cache`
- `pnpm nx run @copilotkit/web-inspector:build --skip-nx-cache`
- repository pre-commit checks

## Checklist

- [x] I have read the [Contribution
Guide](https://github.com/copilotkit/copilotkit/blob/master/CONTRIBUTING.md)
- [x] No documentation update is needed because this fixes existing
behavior without changing the public API
- [x] "Allow edits by maintainers" is checked
2026-08-05 16:34:06 -05:00
David McKay d4d10409f2 fix(react-core): collapse duplicate tool-call ids in the message view (#6407)
Follow-up to #6404. That PR stopped the *duplicate write* from a
resurrected approval card; this one removes the duplicate card itself.

## Symptom

Approving a HITL card in the banking skin left a **second, blank copy**
of the same card in the transcript — "Open policy exception" with no
transaction id and no code — plus a React warning:

```
Encountered two children with the same key, call_QHa801k5PL32WtPKsPItO1pI
```

## Root cause

`@ag-ui/client`'s `TOOL_CALL_START` handler appends to the parent
assistant message's `toolCalls` with no check for an existing entry with
that id:

```js
message.toolCalls ??= [];
message.toolCalls.push({ id, type: "function", function: { name, arguments: "" } });
```

So when a start event is applied twice — which the HITL flow triggers
when the run syncs after `respond()` — the message carries the same call
twice. The second copy has **empty `arguments`**, because a start event
carries none; the args arrive afterwards as `TOOL_CALL_ARGS` deltas
addressed to the first copy. That empty copy is what rendered as the
blank card, and `CopilotChatToolCallsView` uses the call id as its
render key, hence the warning.

Evidence gathered while diagnosing:

- Instrumented the message view: the offending assistant message ends up
with `toolCalls: ["call_X","call_X"]`, and the React warning fires on
the same id in the same tick.
- Queried a local Intelligence stack: the server emits **exactly one**
`TOOL_CALL_START` for that id (one `START`, one `END`, one `RESULT`). So
this is client-side state, not a stream defect.
- React StrictMode is **not** involved — the demo sets `reactStrictMode:
false`.

## Fix

Extends the existing `deduplicateMessages()` — which already collapses
duplicate *message* ids arriving from streaming re-delivery — to also
collapse duplicate *call* ids within a message, preferring whichever
copy actually carries arguments.

Applied outside the merge branch too, because the duplicate also lands
on a message that was never itself duplicated (observed `rawMsgs=5
dedupMsgs=5` with the duplicate still present). Returns the original
array untouched when there is nothing to collapse, so memoized consumers
don't re-render needlessly.

## Scope / known limitation

This fixes what renders. The underlying agent state still holds the
duplicate entry, so a fully correct fix also wants an idempotency guard
in the AG-UI start handler (upstream). Filed separately — flagging here
so the remaining gap is explicit rather than implied-fixed.

The precise reason a single start event gets applied twice is also still
open; the fix is deliberately robust to re-application whatever the
trigger.

## Verification

- 5 regression tests added, **verified red before green** (neutered the
fix → 3 failed with "Expected 1, Received 2"; restored → pass).
- Full `@copilotkit/react-core` suite: **1480 passed / 123 files**.
- `test`, `publint`, `attw` green for react-core and its dependents.
- In-browser against a live Intelligence stack: before, approving left a
blank second card + 3 key warnings; after, **one card and zero
warnings**.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-08-05 14:22:23 -07:00
Ben Taylor 9a86e21d48 feat(runtime): split Channel status into transport and provider legs (refs OSS-739) (#6360)
**Half 2 of 2 for OSS-739.** Gateway half ships first:
CopilotKit/Intelligence#746.

## Why

`status().overall === "online"` proved only that the runtime reached the
Gateway with a valid project API key. It said nothing about whether a
Slack/Teams app was bound to the Channel, so **a Channel with no
provider at all reported `online`** — and every version of our Channels
onboarding guidance used that value to certify end-to-end success.

`setup_required` had **no producer**. The manager set it only when the
activation engine threw `SETUP_REQUIRED`, and the engine stopped doing
that at the 2026-07-29 realtime-boundary cutover (`8f166577ce`). In
published `@copilotkit/channels-intelligence@0.7.0` the string survives
in exactly one file — a shipped *test*. The 15 doc comments describing
the state outlived the mechanism, which is why nobody noticed for a
week.

## Change

- `connectRealtimeGateway` captures the control join reply (it was
**discarded**) and exposes `providerStates()`. Phoenix's `Push.resend`
preserves `recHooks`, so the hook re-fires on every auto-rejoin — a
Channel provisioned while the runtime was disconnected is picked up with
no extra plumbing.
- The launcher and the manager's handle view delegate it as a
**getter**, not a captured snapshot, for that same reason.
- `status()` gains `detail`, reporting `transport` and `provider`
separately so a caller can assert the leg it cares about:

```ts
status() → {
  overall: "setup_required",
  channels: { support: "setup_required" },
  detail: { support: { status: "setup_required", transport: "online", provider: "not_attached" } },
}
```

`channels` keeps its shape — turning its values into objects would break
the CLI's `channels-report` and the starter channel-host — but its
values are now the fold of the two legs, which is what makes `overall`
honest.

- The stale `setup_required` doc comments are corrected, with a note not
to describe the state again without a path that can emit it.

## Back-compat: `unknown` is load-bearing

An older Gateway, a Gateway whose lookup failed, a handle without the
seam, a Channel the Gateway did not mention, an unrecognised state, and
a throwing getter **all** yield `unknown`, which keeps the
transport-derived status — exactly today's behaviour. Only a *positively
reported* absence downgrades a Channel, so no existing deployment turns
amber on upgrade.

The **41 pre-existing channel-manager tests pass unchanged**, which is
that guarantee.

## Testing

- `channel-manager-provider-leg.test.ts` — 14, incl. the regression test
that never existed ("reports setup_required for a joined Channel with no
provider attached") and one case per degradation path
- `realtime-gateway-provider-states.test.ts` — 8 parser cases
- `realtime-gateway.test.ts` — +2 proving the wiring end-to-end through
real Phoenix framing, not just the parser
- Full suites: runtime **1874/1874**, channels-intelligence **192/192**;
`check-types` and `build` clean for both packages

**I mutation-tested the fold** — neutralising it fails 5 tests including
the linchpin — so these assert behaviour rather than passing vacuously.

Worth noting: two type errors (`ChannelsHandle` in `runtime.ts`, a
session mock) were invisible to vitest, which transpiles without
typechecking. The pre-commit build gate caught them.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-08-05 16:06:40 -05:00
David McKay eee580f6da fix(react-core): collapse duplicate tool-call ids in the message view
AG-UI's TOOL_CALL_START handler appends to the parent assistant message's
`toolCalls` without checking whether an entry with that id is already present,
so whenever a start event is applied twice — which the human-in-the-loop flow
triggers when the run syncs after `respond()` — the message carries the same
call twice. The second copy has EMPTY arguments, because a start event carries
none; the args arrive afterwards as TOOL_CALL_ARGS deltas addressed to the first.

Rendering both produced a phantom duplicate card in the transcript (in the
banking skin: a second "Open policy exception" with no transaction id or code)
plus a React "Encountered two children with the same key" warning, since the
call id is the render key in CopilotChatToolCallsView.

Verified against a local Intelligence stack that the server emits exactly ONE
TOOL_CALL_START for the affected id, so this is client-side state, not a stream
defect. React StrictMode is not involved (the demo disables it).

Extends the existing deduplicateMessages() — which already collapses duplicate
message ids from streaming re-delivery — to also collapse duplicate call ids
within a message, preferring whichever copy actually carries arguments. Applied
outside the merge branch too, because the duplicate also lands on a message that
was never itself duplicated. Returns the original array when there is nothing to
collapse, so memoized consumers do not re-render needlessly.

Does not change the underlying agent state, which still holds the duplicate;
that needs an idempotency guard in the AG-UI start handler.

5 regression tests, verified red before green. Full react-core suite: 1480 passed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 14:00:54 -07:00
Ben Taylor 743351cb3f fix(runtime): initialize arrays before /- append in AGUISendStateDelta (#6293)
## Summary

`AGUISendStateDelta` can emit an array append against a state where the
target array has not been initialized. The emitted patch then fails
during event compaction with `OPERATION_PATH_CANNOT_ADD`.

This change keeps the authoritative state represented in emitted events,
initializes a missing array immediately before its first `/-` append,
and applies the same contract across the generic, AI SDK, and TanStack
state-delta paths. Existing arrays and valid deltas retain their current
contents and operation order.

Closes https://github.com/CopilotKit/CopilotKit/issues/5998

## Changes

- Emit the input state before the first delta when no earlier state
event represents it
- Initialize a missing array before normalized classic, AI SDK, and
TanStack state deltas append through `/-`; custom raw-event mode remains
outside this normalizer
- Preserve populated arrays and already-valid patch operations
- Preserve caller-owned state when structured cloning falls back
- Cover missing-array reconstruction, sibling converters, failure
guards, and existing-array preservation
- Follow the current `release.config.json` and Nx release convention; no
Changeset file is added.

## Test plan

- [x] `pnpm -C packages/runtime exec vitest run
src/agent/__tests__/state-tools.test.ts
src/agent/__tests__/converter-aisdk.test.ts
src/agent/__tests__/converter-tanstack.test.ts`
- [x] `pnpm -C packages/runtime exec vitest run`
- [x] `pnpm exec nx run @copilotkit/runtime:check-types`
- [x] `pnpm exec oxfmt --check` on changed runtime files
- [x] `pnpm exec oxlint` on changed runtime files, zero errors with two
pre-existing no-shadow warnings
- [x] Verify no `.changeset` file is added because current main uses Nx
release
- [x] Verify the final diff contains only the private helper, runtime
implementation, sibling converters, and focused tests
2026-08-05 15:40:21 -05:00
Ben Taylor a3d0d2bfab fix(runtime): reject unenforceable mcpApps tool policy instead of silently ignoring it (#6292)
## Summary

`mcpApps.servers` entries that carry `includeTools` or `excludeTools`
are currently accepted even though the pinned
`@ag-ui/mcp-apps-middleware` package has no option for them. The runtime
then ignores the keys, so tools an operator intended to restrict remain
available. This change rejects that configuration instead of allowing a
silent no-op.

## What CopilotKit owns

- `mcpApps.servers` configuration and `agentId` scoping.
- Projection of selected servers into `MCPAppsMiddleware`.
- Reporting unsupported configuration before middleware construction.

Discovery, model-emitted tool execution, frontend-proxied execution,
server identity, and tool provenance belong to
`@ag-ui/mcp-apps-middleware`.

## Changes

- Extract the server projection into `resolveMcpAppsServers`, which
scans all configured entries for defined policy keys, filters by
`agentId`, strips only `agentId`, and forwards other fields unchanged.
- Return a configuration error naming the unsupported key, server,
pinned middleware version, owning package, and issue when a policy key
is supplied.
- Add tests for agent scoping, field forwarding, malformed and empty
values, undefined spread values, constructor avoidance, and the existing
HTTP error path.
- Document the ownership boundary and add a runtime changeset.

## Why the filter stays external

The pinned package is version `0.0.3`. It owns the private server maps,
UI-tool discovery, model-emitted execution, and frontend proxy
execution. A CopilotKit middleware could observe only one of those paths
and would have to duplicate private server identity and tool provenance.
The complete `includeTools` and `excludeTools` implementation belongs in
the external package, where one predicate can cover discovery and both
execution paths.

## Current behavior

Plain JavaScript or JSON configuration can supply `excludeTools:
["delete_account"]` without a TypeScript excess-property check. The
runtime currently accepts the configuration, constructs
`MCPAppsMiddleware`, and leaves the tool available. The new behavior
returns an HTTP 500 through the existing runtime error path, names the
unsupported key and dependency, and does not construct the middleware.

## Follow-up

The counterpart change in `@ag-ui/mcp-apps-middleware` should add the
fields to the per-server configuration, preserve absent versus empty
include lists, resolve server identity through its existing maps, and
apply one predicate after UI-resource discovery and before model-emitted
and proxied tool execution. Once that version is released, CopilotKit
can remove the rejection and pass the fields through unchanged.

## Related issue

Refs #5930.

The cross-repository ownership split follows the proposal in
https://github.com/CopilotKit/CopilotKit/issues/5930#issuecomment-5128722524.
This PR does not close the issue.

## Test plan

- [x] `pnpm -C packages/runtime exec vitest run
src/v2/runtime/__tests__/mcp-apps-servers.test.ts
src/v2/runtime/__tests__/mcp-apps-middleware-integration.test.ts`
passed, 2 files and 20 tests
- [x] `pnpm -C packages/runtime exec vitest run` passed, 129 files and
1,836 tests
- [x] `pnpm exec nx run @copilotkit/runtime:check-types` passed
- [x] `pnpm exec oxlint` and `pnpm exec oxfmt --check` passed on changed
TypeScript files
- [x] `pnpm check:plugin-skills` passed
- [ ] `CI green for static / quality and test / unit on Node 20, 22, and
24`
2026-08-05 15:29:37 -05:00
Benjamin Taylor 9280e71346 refactor(channels): trim redundant provider-leg tests and fix a wrong comment
Self-review of the previous commit.

Corrects a factual error I introduced: the `SETUP_REQUIRED` note claimed such a
Channel "has no transport at all (the launcher never returned a handle)". False
for the `hasDirectAdapter` branch, which starts the developer-owned transport and
assigns a synthetic handle — so there IS a running transport there. Dropped the
wrong reasoning and shortened the note to the part that holds: the misattribution
is cosmetic on a path with no producer, and a future producer should report
through the `providerStates` seam.

Removes three tests that did not earn their place:

- "calls the seam ON the session" — redundant. `ProviderStateGateway.providerStates`
  reads `this`, so the two remaining tests already fail if the launcher ever used
  a detached reference. It died on the same mutation as the first test, for the
  same reason.
- "omits providerStates for a session without the seam" — survived the mutation
  that removes the forward, so it guarded nothing.
- the channel-level-error rejoin case — same `Push.resend` hook as the transport
  drop, so it re-proved one mechanism at ~1s extra wall-clock. Kept the drop
  case: it asserts a genuinely fresh socket, which is the "provisioned while the
  runtime was disconnected" story the design claim is about.
- "keeps the last reported states while a rejoin has not yet succeeded" — pinned
  behaviour with no observable consequence, since the transport leg dominates the
  fold while offline.

Also trims the drift-guard comment: why a guard was NOT added belongs in the PR
discussion, not permanently in source.

Re-mutation-tested after trimming: removing the forward kills both remaining seam
tests; making `providerStates` a snapshot kills the rejoin test while the other
41 gateway tests pass.

Verified: channels-intelligence 195/195, runtime 1874/1874, build + oxfmt clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 15:27:46 -05:00
Ben Taylor a0daff1eab docs(showcase): align A2UI docs with v0.9 helpers (#6288)
## Summary

The A2UI integration docs still use helper names and wire keys from
before the Python SDK's v0.9 API. The examples now match the current
SDK, while the A2A page clearly labels its cloned starter's v0.8
compatibility contract.

## Changes

- Update the generic, DeepAgents, and LangGraph A2UI pages to current
helper names, wire keys, and operation order.
- Reconcile the A2A page to the cloned starter's v0.8 payload and defer
its v0.9 migration.
- Remove unsupported `action_handlers=` and `dataContextPath` claims
from the advanced examples.
- Use the exported `createA2UIMessageRenderer` `onAction` interceptor in
the React guides.

## Out of scope

Migrating `examples/integrations/a2a-a2ui/` from its v0.8 renderer and
operation list requires source and example changes outside this
documentation-only target.

## Related PRs and Issues

Addresses the v0.9 documentation portion of #4821.

The corrected names follow `sdk-python/copilotkit/a2ui.py`; the A2A page
follows the cloned starter's current v0.8 contract. Closed partial work
is tracked in https://github.com/CopilotKit/CopilotKit/pull/5854.

## Test plan

- [x] Documentation search passed. The v0.9 pages contain no stale
Python helper names, current wire names are present, and the A2A
compatibility page is labeled.
- [x] Diff validation passed. Only the eight named MDX files changed.
- [ ] Shell-docs typecheck, lint, and formatting were unavailable
because this worktree has no installed `node_modules`; CI will run them
on the PR.
- [ ] CI green (`static / quality`, `test / unit` on Node 20/22/24).
2026-08-05 15:20:28 -05:00
Benjamin Taylor c71d9ac58b fix(channels): forward the provider seam through the exported launcher helper
Addresses review on #6360.

`startChannelsWithGatewayControl` is public so callers can compose over a
session they manage themselves, but it forwarded only `onClose` and
`onStateChange` — not `providerStates`. A handle without the provider seam makes
`ChannelManager.providerLeg` fall back to `unknown`, which keeps the
transport-derived status and reports `online` for a Channel with no provider
bound: the exact false green OSS-739 removes, still reachable through a public
export. The guard is widened too, so a session exposing only `providerStates`
is no longer dropped on the fall-through path.

Tests the reconnect claim that makes `providerStates` a getter rather than a
snapshot. Nothing exercised it: the gateway tests covered only the initial join
reply, and the manager-side rejoin test proves the manager re-reads on each
`status()` call, not that the session's value ever changes. The fake socket's
join reply can now vary per join, so a drop -> rejoin carrying a different
`channels` map asserts the refresh over real Phoenix framing — via both rejoin
paths (channel-level error on a live socket, and a full transport drop onto a
fresh socket), plus the case where a rejoin has not yet succeeded and the last
known states must persist.

Mutation-tested both: removing the forward kills 3 of 4 seam tests, and making
`providerStates` a captured snapshot kills both rejoin tests while all 3
pre-existing provider-state tests still pass — which is the gap itself.

Docs corrected against their real mechanisms:

- `attached`/`unhealthy`/`not_attached` now state the gateway's actual rule
  (adapter `status == "active"` is part of the predicate; a configured adapter in
  `error` is `unhealthy` with no failed health check), plus the best-of adapter
  fold that keeps a Slack-only Channel `attached`.
- `ready()` no longer promises it rejects on `error`. It awaits activation, so
  it can resolve while `status().overall === "error"` from an `unhealthy`
  provider. Says that instead.
- Notes the legacy `SETUP_REQUIRED` path reports a provider condition on the
  transport leg (dead, cosmetic, left rather than guessed at), and why the
  provider-state set is duplicated across the duck-typed package seam.

Verified: channels-intelligence 199/199, runtime 1874/1874, both builds clean,
oxfmt/oxlint clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 15:16:31 -05:00
Ran Shem Tov ab94c1315e fix(showcase): let the Mastra MCP Apps agent self-correct a rejected diagram
Switching models only moved the failure rate around, it never removed it, so
stop relying on the model getting hand-escaped JSON right on the first try.

`create_view` takes `elements` as a stringified JSON array. When the model
appends a stray `}` past the closing `]`, the MCP server rejects the call and
names the exact fault ("Invalid JSON in elements: Unexpected non-whitespace
character after JSON at position N"). That error already comes back as a tool
result, and the agent had no step cap, so a retry was mechanically possible
all along. What blocked it was our own prompt: "Call create_view ONCE" and
"do NOT iterate, do NOT make multiple calls. Ship on the first shot."

The prompt now tells the model to read the error and try again, capped at 2
corrections (3 calls total), with stopWhen: stepCountIs(6) bounding the loop
if it never converges. This mirrors the validate-then-retry recovery pattern
already used for A2UI on the other integrations.

Validated against the real Excalidraw MCP server, using the agent's prompt
extracted verbatim from this file and the real tool schema:

  normal runs                       12/12 succeeded, all on the first call
  attempt 1 force-corrupted with
  the real-world stray `}`          10/10 recovered on the second call

Also verified in the running app (local dev server, real key): valid JSON,
isError false, diagram rendered.

Not yet verified in-app: the recovery path itself. No natural failure occurred
during the in-app runs, so the retry is proven at the API level rather than
through the Mastra agent loop.
2026-08-05 22:55:49 +03:00
Ran Shemtov 701b03ab96 Merge branch 'main' into claude/jolly-boyd-38b55c 2026-08-05 21:47:52 +02:00
Ran Shem Tov f6c79bed97 chore(showcase): remove CrewAI working notes 2026-08-05 22:39:05 +03:00
Ran Shem Tov 88d9719faf docs(showcase): connect CrewAI full parity docs 2026-08-05 22:38:46 +03:00
Ran Shem Tov ccf979eca8 fix(showcase): stabilize remaining CrewAI D6 cells 2026-08-05 22:38:23 +03:00
Rod Boev 3a9eda8441 fix(runtime): redact sensitive headers from runtime error context 2026-08-05 15:17:25 -04:00
Benjamin Taylor b77ebb435f chore: stop changeset files from reappearing in PRs
The repo migrated off @changesets/* to conventional-commit-driven releases
(scripts/release/ reads commit subjects from git log <lastTag>..HEAD), but
.changeset/ has been removed twice already (5afa55f067, 1e5ba689e0) and five
open PRs currently carry changeset files again. Two mechanisms keep feeding it:
contributor forks whose default branch still has the pre-cleanup .changeset/
debris, and plain convention inference — the repo reads as a Changesets repo
(pnpm monorepo, Changesets-formatted CHANGELOG.md files, "chore: release" PRs)
and nothing anywhere said otherwise.

- CONTRIBUTING.md: explain that we used Changesets, what replaced it, and what
  to do instead (a good conventional commit subject).
- AGENTS.md / CLAUDE.md: same rule for coding agents, which author most of
  these PRs and don't read CONTRIBUTING.md.
- static / check binaries: fail on added .changeset/* files, so this stops
  depending on review catching it. Filters on added/modified only, so a PR
  that deletes stale changesets still passes.
- .oxfmtrc.json: drop the ignore entry for the long-gone vendored
  .github/actions/changesets-action, a stale "we use changesets" signal.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 14:15:44 -05:00
Ran Shemtov 8f24b0373b Merge branch 'main' into claude/framework-d6-integration-validate-7be45e 2026-08-05 21:08:41 +02:00
David McKay 8b46745755 fix(reskinnable-demo): collapse HITL approval buttons on the tool result (#6404)
Ports the banking demo's #6401 fix, which was never carried over to
`reskinnable-demo`. Found while bringing the demo up against a local
Intelligence stack.

## The bug

`ApprovalButtons` collapsed only on local `responded` state, which dies
with the component. These cards **do** get remounted when the run syncs,
which resurrects live Approve/Deny buttons on an action the user already
took — a second click fires a duplicate write against an already-settled
call.

Reproduced in the banking skin: clicking the "Approve the $15,000 AWS
charge" pill and approving the policy exception left a **second card
carrying live Approve/Deny buttons** (with empty args).

## The fix

Adds a durable `resolved` prop, OR-ed with the local state so a click
still collapses without waiting for the round trip. It is passed from
the tool call itself:

```tsx
resolved={status === "complete" || !!result}
```

Applied at the three HITL renders that don't already early-return on
`status === "complete"` (`openPolicyException`,
`finalizePolicyException`, `approveTransaction`).

The other three (`offerWorkflowRecording`,
`awaitDashboardDemonstration`, `saveLearnedWorkflow`) render their own
terminal card when complete, so they never reach the buttons — passing
`resolved` there is both redundant and a type error, since `status` is
narrowed to `ToolCallStatus.Executing`. That is why banking also has
exactly three call sites.

## Verification

- Before: second card had live Approve/Deny. After: it reads *"Response
submitted."*
- `pnpm lint` → exit 0
- `pnpm build` (the type-check gate) → exit 0

Note: the duplicate card still renders — that comes from a separate
pre-existing `Encountered two children with the same key` warning also
present in banking, and is out of scope here. This PR removes the
*duplicate-write* hazard, which is what #6401 addressed.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-08-05 12:05:38 -07:00
Maxim 3fde920ea4 Merge branch 'main' into fix/reskinnable-demo-hitl-resolved 2026-08-05 21:05:31 +02:00
Ran Shemtov 4ad21db755 fix(mastra-showcase): beautiful-chat A2UI — ground dynamic render + fixed-schema flights (PNI-122) (#6387)
Fixes **PNI-122** — beautiful-chat A2UI dynamic renders a varying error
/ no UI, flights render no UI, and (multi-turn) the dashboard paints
loose charts. All fixed and **verified live on a real LLM** (gpt-5.4
outer + gpt-4.1 render, matching gold `beautiful_chat.py`).

## 1. Dynamic A2UI ungrounded (Sales Dashboard) — no UI / varying error
`generate_a2ui` grounded its inner `render_a2ui` subagent from the
tool's `contextEntries` ARG, which the outer model always sends
**empty** (captured live: `contextEntries: []`) → empty system prompt →
invalid/misnamed components (or none) → no UI, nondeterministic. aimock
hid it (fixture returns a valid envelope regardless).

**Fix:** read the catalog schema + generation guidelines the
`@ag-ui/mastra` bridge already forwards onto the Mastra request context
(`requestContext.get("ag-ui").context`) and ground the render there.
Mirrors `readAgUiContext` in `@ag-ui/mastra`'s `getA2UITools`; preserves
per-demo catalogId. New leaf `tools/a2ui-context.ts` + regression test.

## 2. Flights narrated as text — langgraph-python parity
mastra reused the shared `searchFlightsTool` (plain `{flights}`,
rendered by the tool-rendering cells' own frontend `FlightListCard`), so
beautiful-chat produced no A2UI surface. Gold `beautiful_chat.py` wires
a **dedicated fixed-schema `search_flights`** returning an
`a2ui_operations` FlightCard envelope. Mirrored via
`searchFlightsA2uiTool` + a dedicated `beautifulChatAgent` (query_data,
todos, generate_a2ui, the fixed search_flights, flight/dashboard
steering, `parallel_tool_calls=False`, gpt-5.4). Route repointed to it,
keeping the fixed flights + steering out of the shared `weatherAgent` /
tool-rendering cells.

## 3. Dashboard over-called standalone charts (multi-turn)
gpt-4o fired the standalone `pieChart`/`barChart` frontend tools **and**
`generate_a2ui`, painting loose charts next to the dashboard. **Fix:**
gold's `parallel_tool_calls=False` + gpt-5.4 + sharpened steering (a
dashboard / "using A2UI" request calls generate_a2ui ONLY; a
single-chart request still uses the standalone tool). The aimock side
had the same over-call **baked into `recorded.json`** (a real 2026-05-15
gpt-4o session) — dropped that turnIndex-2 leg so the aimock chain is
`query_data → generate_a2ui → narration`, matching live. Flights aimock
fixture realigned to `search_flights`.

## Verification
Ran the built branch locally against a **real OpenAI key**:
- Sales Dashboard (A2UI Dynamic) → full grounded dashboard,
`[query_data, generate_a2ui]` only, single **and** after-flights ✅
- Search Flights pill → two FlightCard surfaces (no A2UI steering
needed) ✅
- Standalone Pie / Bar pills → still call `pieChart`/`barChart` and
render ✅

Test suite: added `a2ui-context.test.ts` (5/5). No new failures
(pre-existing `route.test.ts` + `demoAgentNames.parity` fail on `main`).

✅ **Aimock verified locally too** — ran the aimock image (context-routed
d6 fixtures) at :14010 and pointed the app at it via `OPENAI_BASE_URL`.
Flights pill → `search_flights` → 2 FlightCards (United $349 / Delta
$289). Sales Dashboard (after flights) → `generate_a2ui` only, no
standalone chart over-call; renders Total Revenue + 1 pie + 1 bar
(exactly 2 recharts containers). Matches live.

## Follow-ups (out of PNI-122 scope)
- Tool-call/text render ORDER in multi-turn is the tracked Mastra
one-message-id bug (separate chip).
- The dedicated `a2ui-fixed-schema` **demo** (catalog
`flight-fixed-catalog`) still uses the plain shared `search_flights`;
gold uses a separate `display_flight` there.
2026-08-05 21:03:47 +02:00
Ran Shemtov 21dcd95ce2 Merge branch 'main' into claude/competent-chatelet-7305ee 2026-08-05 21:03:41 +02:00
Ben Taylor 26d352f898 fix(channels): default welcome agent prompt (#6377)
## What does this PR do?

- Gives `thread.runAgent()` inside `onWelcome` the default prompt
`"Introduce yourself to the channel!"`.
- Keeps an explicit `runAgent({ prompt })` ahead of the default.
- Adds regression coverage for the default and override paths.

Welcome deliveries have no inbound message. Before this change,
`thread.runAgent()` passed an empty message list to the AI SDK, which
rejected the run before it reached the model.

## Related PRs and Issues

- No linked issue.

## Testing

- `pnpm nx test @copilotkit/channels-core -- --run src/welcome.test.ts`
- `pnpm nx run-many -t test,check-types,build
--projects=@copilotkit/channels-core`
- `pnpm run lint`
- `pnpm run check-format`
- `pnpm nx run-many -t publint,attw
--projects=@copilotkit/channels-core`

## Checklist

- [x] I have read the [Contribution
Guide](https://github.com/copilotkit/copilotkit/blob/master/CONTRIBUTING.md)
- [x] If the PR changes or adds functionality, I have updated the
relevant documentation
- [x] Maintainers can edit this same-repository branch; GitHub only
shows the "Allow edits by maintainers" control for forks
2026-08-05 14:02:50 -05:00
Austin Merrick ab7a50fa10 feat(inspector): add agent-ready Rich Threads setup 2026-08-05 11:55:53 -07:00