- frontend-tools-async: drop the stray broad 'project planning' fixture entry that
substring-shadowed 'Find my notes about project planning' -> run-loop. Green.
- gen-ui-open + gen-ui-open-advanced: remove 6 gen-ui-open-owned strays from
gen-ui-tool-based.json ('3D axis visualization', 'Inline expression evaluator',
'render an open gen-ui element', 'continue the advanced gen-ui flow') that
collided (same context, different responses) with gen-ui-open.json. Green;
gen-ui-tool-based unaffected.
- multimodal: quarantine. On main + latest deps the browser gets 200 from the
runtime but never starts a run (runStartCount=0) — a shared multimodal
frontend/runtime issue (identical frontend to langgraph; agent works via direct
POST; my #5985 fix doesn't change it). Genuine gap, needs shared-layer work.
Full D6 now 36/36 green on official latest (ag-ui 1.0.1, core 1.13.0, openai 1.12.0).
- shared-state-streaming: replace main's stale 'counter' fixture (drifted from
langgraph canonical) with the write_document document demo (6-entry poem/email/
quantum + chunkSize) + SharedStateStreamingFrameworkAgent seed subclass; frontend
already matches langgraph. Un-quarantine + add to features. Green (3/3 turns).
- tool-rendering-reasoning-chain: un-quarantine + add to features. Green on core
1.13.0 / openai 1.12.0 (latest) via aimock encrypted_content (CopilotKit/aimock#342,
which fixes the Responses reasoning multi-tool regression) + the store:False agent.
CI-green depends on aimock#342 releasing.
- Drop reasoning-default-render / agentic-chat-reasoning from not_supported (no D6
probe featureType — not real cells).
Port the #5985 fix onto main: ported langgraph's full 18-entry custom-catchall
fixture (main's 4-entry set left SF/flights/d20/chain pills leaking to the
default-catchall fixture) + _ToolRenderingFrameworkAgent dropping the divergent
end-of-run MESSAGES_SNAPSHOT so the narration (not the tool card) is the terminal
bubble. Verified green via per-cell --direct. First cell of the #5985->main
re-integration.
Both open-gen-ui cells reused the shared weatherAgent
("You are a helpful assistant.", gpt-4o). On a live LLM that produced
static HTML with no JS wiring - buttons rendered but were dead. aimock
hid it: fixtures match on userMessage/hasToolResult and replay a
fully-wired UI regardless of the agent.
Gold langgraph-python uses dedicated agents whose system prompt mandates
a single interactive generateSandboxedUi call and reads the design-skill
+ sandbox-function descriptors from copilotkit context. On Mastra,
RunAgentInput.context is a read-channel (requestContext.get("ag-ui")),
not injected into the prompt, so that guidance never reached the model.
Add dedicated openGenUiAgent / openGenUiAdvancedAgent porting gold's
prompts, with dynamic instructions that fold the ag-ui context into the
system prompt, and point the ogui route at them.
Verified live (real OpenAI, aimock bypassed) - all three advanced pills
render interactive, host-wired UI with confirmed round-trips:
evaluateExpression 7*8=56, notifyHost, evaluateExpression 5+3*2=11.
aimock matcher keys are untouched, so replay is unaffected.
Updates the runtime agent-runner skill references to match the bounded
in-memory runner: correct the InMemoryAgentRunner store as a process-global
singleton, document its bounds and onConcurrentRun concurrency handling, note
that dedup weakens past the run cap, fix rotted in-memory.ts citations onto
stable symbols, and correct the multi-instance SqliteAgentRunner scaling
guidance. The generated skills/ mirror is regenerated in lockstep so source and
mirror stay in sync.
Co-Authored-By: Claude <noreply@anthropic.com>
Documents the now-bounded in-memory runner for users: the maxThreads /
maxRunsPerThread / maxBytes limits and their defaults, the precise eviction
model (LRU threads, per-thread run-cap, cross-thread byte ceiling enforced at
run completion), and the onConcurrentRun throw/supersede option. Clarifies that
the store is a process-global singleton shared by every runner, that dedup
weakens past the run cap, and points at the first-party SqliteAgentRunner for
durable or multi-instance deployments. Adds a troubleshooting entry for the
in-memory eviction warning.
Co-Authored-By: Claude <noreply@anthropic.com>
Adds a dedicated bounded-thread-store suite and extends the in-memory runner
suite to lock in the behaviors introduced with the bounded, teardown-isolated
runner:
- LRU thread eviction, per-thread run-cap FIFO trimming, and byte-ceiling
eviction (including that a live/running or stop-requested thread is never
evicted, and that a just-appended thread pushes OTHER threads out rather than
self-evicting).
- InMemoryLimits validation/normalization: invalid bounds clamp to defaults
instead of crashing enforceRunCap, and the 0/Infinity disable sentinels are
preserved.
- Thread-level createdAt and message-snapshot decoupling survive run-cap
eviction and interleaved empty-snapshot runs.
- stop() guards and the supersede path: an aborted run finalizes as a clean
RUN_FINISHED against its own captured intent, a superseded run cannot clobber
its replacement's state or history, and an immediate abort-throw that emitted
nothing creates no phantom historic run.
Also restores the shared store's default limits after tests that reconfigure
the process-global store so suites stay isolated, and updates a handle-run
comment for the GLOBAL_STORE -> shared store rename.
Co-Authored-By: Claude <noreply@anthropic.com>
A run's async teardown used to read the shared, mutable store.stopRequested to
decide whether to finalize as a clean RUN_FINISHED or a synthetic RUN_ERROR.
Under supersede (or a stop() immediately followed by a new run()), a later run
resets that shared flag, so the earlier run's fire-and-forget finalization could
be mislabelled — an intentionally stopped run finalized as an error, or vice
versa — and could push history under, or clobber the state of, the newer run
that now owns the thread.
Fix by capturing a per-run RunFinalizeControl when the run starts. stop() and a
superseding run() flip THAT run's captured control (not just the store flag),
and the run's teardown reads its own captured intent, so a later run resetting
store state can never change how an earlier run finalizes.
The teardown itself is unified into a single finalizeRun helper shared by the
success and error paths (they were near-identical and must stay symmetric) and
made ownership-aware:
- It only pushes history / resets shared store state when this run still owns
the thread (store.currentRunId still equals this run's id), so a superseded
run cannot corrupt the successor's history or state.
- The error path additionally requires at least one real (pre-finalize) event,
reviving a guard that had gone dead: an immediate throw that emitted nothing
must not create a phantom historic run holding only the synthetic terminal.
- On completion it releases the run's infinite ReplaySubject buffer via an
identity guard (store.subject === nextSubject), reclaiming the duplicate
buffer on the owning path while leaving a live successor's subject untouched.
The concurrency branch now also triggers on store.stopRequested, not just
isRunning: stop() flips isRunning off the instant it aborts but the run keeps
finalizing, and a run() slipping through that window went entirely unhandled.
The previous-subject bridge is removed: forwarding a dying superseded run's
subject would replay its RUN_STARTED and push its terminal into the live run's
stream, an invalid AG-UI sequence — a superseded run must stay isolated to its
own subscribers.
Co-Authored-By: Claude <noreply@anthropic.com>
Route every storage access in the runner through the shared ɵBoundedThreadStore
instead of the old unbounded GLOBAL_STORE Map: run() acquires threads via
getOrCreate (which applies LRU eviction), and connect/isRunning/stop/
listThreads/getThreadMessages/getThreadEvents/clearThreads read through the
store's touch-aware accessors so reads keep LRU order honest.
The constructor now accepts InMemoryLimits inline alongside onConcurrentRun.
Note the scope difference, called out in the JSDoc: onConcurrentRun is
per-runner, but the limits reconfigure the PROCESS-GLOBAL store shared by every
runner. A partial limits update coalesces each unspecified field against the
store's current effective bounds (not the hardcoded defaults), so tuning one
bound never silently resets its siblings; a genuine clobber of an
already-customized store warns once.
getThreadMessages now returns the thread-level snapshot (a shallow array-level
copy) rather than the last run's snapshot, so run-cap eviction and interleaved
empty-snapshot runs can never lose it. getThreadState is hardened to reject
arrays (which pass `typeof === "object"`) and to return a defensive shallow
copy so callers cannot mutate stored snapshot state.
Co-Authored-By: Claude <noreply@anthropic.com>
The in-memory runner previously kept every thread and run forever in an
unbounded process-global Map, so a long-lived process leaked memory without
limit. Introduce ɵBoundedThreadStore as the single backing store, enforcing
three independent bounds resolved from InMemoryLimits (defaults in
ɵINMEMORY_DEFAULTS):
- maxThreads: LRU eviction of whole threads.
- maxRunsPerThread: FIFO run-cap per thread.
- maxBytes: approximate cross-thread byte ceiling (via ɵestimateBytes),
enforced at run completion by evicting other LRU non-running threads.
Limit values are validated and normalized once (ɵnormalizeLimits /
ɵisValidLimit): only a non-negative integer or +Infinity is well-formed.
Invalid values (negatives, -Infinity, NaN, fractional caps) would otherwise
turn the `count > limit` enforcement guards into infinite loops or a shift()
of undefined; they are instead clamped to the documented default with a single
warning. Clamp-and-warn rather than throw matches this file's best-effort
posture (ɵestimateBytes swallows serialization failures), because constructing
a non-durable convenience runner must never abort — or later surface an
unhandled rejection — on a typo'd bound.
Thread creation time and the latest non-empty message snapshot are held at the
THREAD level (InMemoryEventStore.createdAt / messagesSnapshot), decoupled from
historicRuns so run-cap FIFO eviction can neither drift the reported creation
time forward nor drop the message history. Eviction — whole-thread LRU and
per-thread run-cap trimming alike — is logged once per store (warn-once latch)
so bounded history loss is visible rather than silent.
Also defines the per-run RunFinalizeControl shape and the store's
activeFinalize holder that the run-teardown isolation builds on.
Co-Authored-By: Claude <noreply@anthropic.com>
Fixes the six defects a 12-agent review found in the banking skin, by
fixing the thing that caused most of them: the skin had **three
disagreeing answers to "what did we spend"**.
| Source | Total | Drove |
|---|---|---|
| `charges-data.ts` fixture (45 rows, Apr–Jun) | $632,806 | the Charges
page |
| `seed.json` ledger (4 rows, **Apr–May only**) | $30,089 | the report's
**charts** |
| static `policies[].spent` | $137,000 | the report's **KPI** |
Every one of those numbers appeared in the product, and they were never
reconciled.
## The headline defect
`SpendingTrendChart` substituted a hard-coded `[3200, 4100, 3600, 5200,
4800, 6400]` / Jan–Jun series whenever fewer than three distinct months
were present. Intended as an empty state — but the ledger spanned
exactly **two** months, so the fallback was the **default path**. The
report's "Spend over time" showed six invented figures, roughly 20×
smaller than the total printed directly above them, under a card whose
own docstring reads:
> Every number is computed here from the live ledger … so a report can
never quote a figure the app disagrees with.
Reproduced in the running app before fixing (`POST
/api/banking/v1/reports`, no agent needed), and there's a nasty
interaction worth knowing: attaching an invoice dates a synthetic
transaction *today*, supplying a third month, so **attaching an invoice
masked the bug**. A test written casually would sit in the masked state
and pass.
## The fix: one ledger
The 45 charges now live in `seed.json` as real transactions across
**Apr/May/Jun**, and the Charges page reads them over REST like every
other surface.
**Team and policy are now different axes.** A charge belongs to one of
seven org teams; a policy is one of three budget envelopes (Technology /
Go-to-Market / G&A) and several teams share one. They were a single
`ExpenseRole` enum — which is exactly why covering every team meant
choosing between a seven-slice donut and discarding real charges.
`ExpenseRole` still types a member's own team; `PolicyType` types the
envelopes, joined by `policyForTeam`.
**`policies[].spent` is derived from approved charges on every read.**
It could no longer disagree with the charts — and it also now *moves*:
approving a charge previously left `spent` untouched, so the budget
never reflected the approval and the over-limit gate kept comparing
against a stale figure. Verified live: approving a \$960 charge moved
`spent` by exactly \$960.
**Over-limit is derived-only.** A charge no longer *stores*
`over-limit`; the Charges badge resolves through `withOverLimit`, the
same rule the report uses.
## The six review findings
| | Defect | Fix |
|---|---|---|
| a1 | report charted fabricated spend | charts real months; explicit
empty state at zero |
| a2 | donut its own docs said would collapse | real \$533k base — the
\$900k invoice that hit **89%** now reaches **73%**; 89% would need
~\$2.8M |
| a3 | `report` matched inside "quarterly report" | anchored both sides,
whole words |
| a4 | `as Transaction` laundered a nullable `policyId` | cast removed;
compiler checks it |
| a5 | comment claimed additions "have no policyId" | corrected — three
lines from the code that sets it |
| a6 | `?sort=banana` lit the "active" tint | params validated; unknown
reads as unset |
All six existed **identically in `examples/showcases/banking/`** — they
came from the upstream PRs this skin replayed, not from the port. Scoped
to the skin per review; banking still carries them.
## The scripted demo is unchanged by construction
- the four demo-load-bearing transactions survive byte-identical
- over-limit is still **exactly 3 charges / \$30,000**
- AWS \$15,000 still derives over-limit for the teach-mode pill
- Delta Airlines is still the only Delta charge (the fixture's
near-duplicate "Delta Air Lines" became United Airlines)
- all four status badges still appear (`Amazon Business` is kept
pending, under its policy's headroom, so a plain **Pending** chip
survives)
The donut goes 44/40/16 → **48/27/24** and is relabelled "Spend by
policy", which is what it reads.
## Verification
```
nx build react-core,a2ui-renderer,core,runtime,shared exit 0
tsc --noEmit 0 errors
vitest 239 passed (was 215)
eslint 0 problems
```
Plus live checks against a running server: 49 transactions over 3
months, derived `spent` tracking an approval, over-limit holding at
3/\$30,000.
The 24 new tests are confirmed **red** against the old code —
`parseSort("banana")` returned `"banana"`, `parseTop("-5")` returned
`-5`, the seed spanned two months, no row carried a team, and `spent`
was stored.
## Not in this PR
A "tool replay-guard sweep" (4 items, `navigateToPageAndPerform` and
three approval tools missing the resolved-state guard `showCharges` has)
and ~13 subject-neutral items are captured as follow-ups.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Adversarial review of the two prior commits found the multimodal gate was
only half fixed, plus an overclaim.
Gate: _NATIVE_CONTENT still short-circuited ImageUrl and DocumentUrl. v2's
load_messages builds those from unvetted client input, so url-source
attachments bypassed the gate entirely where v1 had routed them through it:
- an image/heic or image/svg+xml url reached the provider as input_image,
which the Responses API rejects, failing the turn
- an audio/mpeg document url reached it as input_file
- a blank-mime inline PDF was handed to the image sniffer and degraded,
because load_messages collapses ImageInputContent and DocumentInputContent
to the same bare BinaryContent and erases the modality v1 defaulted on
Native content is now gated BEFORE the fixpoint check rather than instead of
it: _classify_native_content returns a tuple only when the gate must act, and
None means provider-safe, which falls through to the fixpoint and preserves
the object's identity. ImageUrl is gated rather than rerouted, so a
provider-safe one keeps its identity and any explicit _media_type.
_url_media_type reads FileUrl.media_type, which INFERS from the url and
raises when it cannot; the explicit _media_type is None unless the client set
one, so reading only the latter would have missed a.heic entirely.
Parity claim: the mount_agent docstring said routing and behaviour were both
unchanged. Routing is; model input is not. v2 defaults
manage_system_prompt='server', so each agent's system_prompt now reaches the
model. On v1 it never did — _agent_graph emitted system parts only if not
messages, and the AG-UI bridge always supplied history — so 18 of 19 agents
had silently dead prompts. Verified by A/B on both versions with the same
agent and request: v1 sends 0 system-prompt parts, v2 sends 1. The new
behaviour is correct and is kept; the docstring now says so instead of
claiming parity.
52/52 python tests pass. Gate verified end-to-end through
load_messages -> flatten -> _map_user_prompt for heic/svg/png urls, inline
heic/png, audio-mime document urls, blank-mime PDFs, and un-inferable urls;
f(f(x)) == f(x) confirmed for ImageUrl, DocumentUrl and BinaryContent.
PydanticAI v2's AGUIAdapter.load_messages converts AG-UI attachments to
native content types before the model boundary; v1 delivered the raw AG-UI
part dicts. _NATIVE_CONTENT listed BinaryContent as a flatten fixpoint, so
under v2 inline attachments were waved straight through and the whole
provider gate was skipped:
- inline PDFs were no longer text-extracted, so raw bytes went to OpenAI
- unsupported image subtypes (HEIC/SVG/TIFF) were no longer degraded and
reached the provider as images, which fails the turn
- missing-mime magic-byte sniffing never ran
- AudioUrl/VideoUrl were neither fixpoints nor classifiable, so they hit
the fail-loud raise
BinaryContent is no longer treated as a fixpoint. _classify_native_content
maps native content onto the same (kind, scheme, mime, value) tuple the
AG-UI classifier produces, so every existing gate — PDF extraction, image
sniffing, the supported-subtype allow list, the binary choke point —
applies unchanged. audio/* and video/* mimes are named explicitly because
_kind_for routes them to "other", and a missing mime defaults to "image"
so the sniffer runs.
Five test assertions checked that state-backing content was still AG-UI
InputContent, which encoded v1's bridging. They now assert the flatten's
output (ImageUrl) never appears in state, which is the guarantee they were
written to protect. Behaviour is otherwise unchanged: a supported inline
image still flattens to an ImageUrl data URI exactly as on v1.
52/52 python tests pass, up from 42/52. Verified against pydantic-ai
2.22.0.
Two new suites, each written so it would have failed against the previous code
rather than merely describing the new behaviour:
- derived-spend: policies[].spent equals the sum of approved charges, excludes
pending/flagged (so the over-limit check cannot double-count the charge it is
gating), and agrees between policies() and findPolicy(). Plus the properties
the demo depends on — at least three distinct months (the condition whose
absence made the trend chart fabricate), AWS $15,000 still deriving
over-limit for the teach-mode pill, exactly one Delta charge, and a team and
category on every ledger row.
- charges-data: parseSort and parseTop reject the values that previously slipped
through as valid, and toChargeRow's over-limit projection only applies to
pending charges.
Confirmed red against the old code: parseSort("banana") returned "banana",
parseTop("-5") returned -5, the seed spanned two months, no row carried a team,
and spent was a stored field.
The four existing fixtures move from ExpenseRole to PolicyType for policy
`type`, following the team/policy split.
The regex that decides whether a transaction note gets a red-alert prefix was
anchored on the left only, so `report` matched inside "reporter" and "quarterly
report" and `disput` inside "disputation". A note reading "attached to the
quarterly report" was served a fraud marker.
Anchored on both sides and switched from stems to whole words, verified against
the four phrasings that previously false-positived.
Four defects in the report's charts, all reported by review and all
reproduced in the running app before fixing.
- SpendingTrendChart substituted a hard-coded [3200, 4100, 3600, 5200, 4800,
6400] Jan-Jun series whenever fewer than three months were present. Intended
as an empty state, it was the DEFAULT path: the seeded ledger spanned two
months, so the report's "Spend over time" always showed six invented numbers
— roughly 20x smaller than the total printed directly above them — under a
card whose own contract says every number comes from the live ledger. It now
charts whatever months exist, with a real empty state at zero.
- SpendBreakdownChart's docstring said the report must use SpendByTeamBars
instead, "because an attached invoice can push one team to ~96% and a donut
cannot survive that", while the report rendered the donut anyway. The warning
was real but its cause was the thin ledger, not the chart: against the old
$137,000 base a $900,000 invoice took one slice to 89%. Against the real
~$533,000 base the same invoice reaches 73%, and 89% would need ~$2.8M.
Robust because the data is real, not because a floor was added to the arc.
- augmentForReport built its synthetic transactions behind an `as Transaction`
cast that was hiding a real hole: `policyId` came from an `?.id` lookup, so it
could be undefined where the field is a required string. The cast is gone and
the compiler checks it. Additions also now resolve their model-authored team
to a policy envelope through `policyForTeam` rather than comparing a team name
to a policy name, with one "Unattributed" segment for unmappable names.
- A comment inside TopChargesChart claimed document-sourced charges "have no
policyId" — three lines above the code that gives them one.
The donut column is relabelled "Spend by policy", which is what it reads.
The skin held three disagreeing answers to "what did we spend": a 45-row
Charges fixture ($632,806), a 4-row seeded ledger ($30,089 across two
months), and static policy totals ($137,000). Each surface read a different
one, so they drifted silently — and because the ledger spanned only two
months, the report's trend chart fell back to a hard-coded series and showed
invented figures under a card that promises live numbers.
Now there is one ledger. The 45 charges live in seed.json as real
transactions across Apr/May/Jun, and the Charges page reads them over REST
like every other surface.
- Splits team from policy. A charge belongs to one of seven org teams; a
policy is one of three budget envelopes (Technology / Go-to-Market / G&A)
and several teams share one. These were a single `ExpenseRole` enum, which
is why the two axes read as one thing and why covering every team meant
either a seven-slice donut or discarding real charges. `ExpenseRole` still
types a member's own team; `PolicyType` types the envelopes, joined by
`policyForTeam`.
- Derives `policies[].spent` from approved charges on every read, so it can
no longer disagree with the charts. It also now MOVES: approving a charge
previously left `spent` untouched, so the budget never reflected the
approval and the over-limit gate kept comparing against a stale figure.
- Makes over-limit derived-only. A charge no longer stores "over-limit"; the
Charges table resolves the badge through `withOverLimit`, the same rule the
report uses, so the two cannot disagree.
- Validates the `?sort=` and `?top=` params. `?sort=banana` used to be cast
straight to a SortKey and lit the control's "active" tint while the table
silently sorted by the default; `?top=-5` reached `slice(0, -5)` and dropped
the LAST five rows, inverting top-N.
The scripted demo is unchanged by construction: the four demo-load-bearing
transactions survive byte-identical, over-limit is still exactly three charges
totalling $30,000, AWS $15,000 still derives over-limit for the teach-mode
pill, and Delta Airlines is still the only Delta charge (the fixture's near
-duplicate "Delta Air Lines" became United Airlines).
Server, agents and pins moved to v2. The /multimodal cell is NOT yet
ported — see below.
- requirements.txt: pydantic-ai-slim 2.22.0, ag-ui-protocol 0.1.19. Drops
the opentelemetry-api<1.44 ceiling; v2 works with otel 1.44.0.
- StateDeps imports move from pydantic_ai.ag_ui to pydantic_ai.ui across
the 8 agent modules and one test.
- agent_server.py: to_ag_ui() was removed in v2, so a mount_agent() helper
builds the equivalent Starlette sub-app (single POST / route, matching
what AGUIApp produced) and mounts it. All 19 paths are byte-identical,
trailing slashes included, so no TS route changes. deps are constructed
per request because v2's adapter assigns client state onto deps.state
instead of replacing the deps object as v1 did.
- test_multimodal_content_mapping.py: importorskip pointed at the removed
pydantic_ai.ag_ui, which would have skipped all 43 tests green under v2.
Now targets pydantic_ai.ui.ag_ui and uses the public
AGUIAdapter.load_messages in place of the v1 private helper.
Verified against pydantic-ai 2.22.0 by driving the real app with
TestClient: 16/19 mounts return 200 SSE with RUN_STARTED..RUN_FINISHED and
no RUN_ERROR. The other 3 (/a2ui_dynamic, /beautiful_chat, /) reach tool
execution and then fail on a raw OpenAI() client constructed inside a tool,
which the harness cannot intercept and aimock handles in CI. Per-request
deps isolation confirmed on /shared_state_read_write: state sent by one
request does not appear in the next.
Known incomplete: multimodal_agent.py does not handle v2's content types.
v2's load_messages emits BinaryContent where v1 emitted ImageUrl, so
_flatten_messages_for_model no longer flattens it, VideoUrl reaches the
unrecognised-part raise, and one binary path raises RuntimeError. 10 of
the 43 multimodal tests fail as a result.
## Release monorepo v1.66.2
**Scope:** `monorepo` | **Bump:** `patch`
---
### How this release process works
1. **This PR was created automatically** by the "release / create-pr"
workflow.
It bumped the `monorepo` packages to `1.66.2`
and generated AI-enhanced release notes.
2. **CI runs on this PR** — the full test suite (unit tests, lint, type
checks, build)
must pass before merging. This is the review gate.
3. **Review the release notes** in `release-notes.md` in this PR.
If a Notion draft was created, you can edit the release notes there
before merging.
4. **When this PR is merged**, the `release / publish` workflow
automatically:
- Builds all packages
- Publishes the `monorepo` packages to npm at version `1.66.2`
- Creates git tag `monorepo/v1.66.2`
- Creates a GitHub Release with the final release notes
### Before merging
- [ ] CI is green (tests, lint, types, build)
- [ ] Version bumps look correct
- [ ] Release notes are accurate (edit in Notion if a draft was created)
---
> **Do not merge until CI is fully green.** The full test suite runs
automatically on this PR.
## Release channels v0.7.3
**Scope:** `channels` | **Bump:** `patch`
---
### How this release process works
1. **This PR was created automatically** by the "release / create-pr"
workflow.
It bumped the `channels` packages to `0.7.3`
and generated AI-enhanced release notes.
2. **CI runs on this PR** — the full test suite (unit tests, lint, type
checks, build)
must pass before merging. This is the review gate.
3. **Review the release notes** in `release-notes.md` in this PR.
If a Notion draft was created, you can edit the release notes there
before merging.
4. **When this PR is merged**, the `release / publish` workflow
automatically:
- Builds all packages
- Publishes the `channels` packages to npm at version `0.7.3`
- Creates git tag `channels/v0.7.3`
- Creates a GitHub Release with the final release notes
### Before merging
- [ ] CI is green (tests, lint, types, build)
- [ ] Version bumps look correct
- [ ] Release notes are accurate (edit in Notion if a draft was created)
---
> **Do not merge until CI is fully green.** The full test suite runs
automatically on this PR.
`showcase/integrations/pydantic-ai` cannot be installed from scratch
today. A clean `pip install -r requirements.txt` produces an agent that
dies on import:
```
ModuleNotFoundError: No module named 'opentelemetry._events'
```
pydantic-ai 1.0.18 imports `opentelemetry._events`. That module was
removed in **opentelemetry-api 1.44.0** (last present in 1.43.0).
pydantic-ai declares an unbounded `opentelemetry-api>=1.28.0`, so a
fresh resolve picks 1.44.0 and the import fails.
## Why nothing was red
`Dockerfile:22-24` copies **only** `requirements.txt` before `pip
install`, so that layer's cache key depends on nothing else. The agent
image has been reusing a pip layer baked before 1.44.0 shipped
(2026-07-16) — including a successful build earlier today. The failure
was masked by Docker layer caching, not absent. Any change to
`requirements.txt`, a cache eviction, or a `--no-cache` build surfaces
it.
Note that this PR busts that cache by definition, so CI's image build is
a genuine fresh-resolve test of the fix rather than a cached pass.
## Verification
Run with `pip` in a clean 3.12 venv, matching how the Dockerfile
installs:
```
RED (main) opentelemetry-api 1.44.0 -> import pydantic_ai raises ModuleNotFoundError
GREEN (this branch) opentelemetry-api 1.43.0 -> import pydantic_ai + pydantic_ai.ag_ui OK
```
Only the ceiling is added; `opentelemetry-sdk` follows to 1.43.0 on its
own, so a second pin isn't needed. No resolution conflict with
`logfire>=4.10.0`.
Docker and the `--d6` probe stack aren't available in my environment, so
the mandatory value-test per `showcase/AGENTS.md` rule 4 has not been
run locally — CI's image build and the dojo cells are the real gate
here.
## Scope
One line plus a comment recording why it exists and when to remove it.
The comment matters: an undocumented pin is what caused the sibling
`starlette==0.45.3` rot in #6363, where a correct-when-written pin
outlived its reason and silently capped pydantic-ai a full major.
The ceiling should come off when this package moves to pydantic-ai v2,
which requires `opentelemetry-api>=1.28.0` without needing `_events`.
## Summary
- replace Phoenix's retained closed transport when a live managed
session receives an unexpected WebSocket close with code 1000
- reconnect and rejoin through the existing Phoenix channel lifecycle
- preserve intentional session disconnects as terminal
- cover the clean-close recovery path with a regression test
## Root cause
Phoenix 1.8.4 does not schedule its reconnect timer for close code 1000.
CopilotKit still transitioned the managed session to `reconnecting`, so
the runtime reported that Phoenix was retrying indefinitely even though
Phoenix never created another transport.
## Impact
A brief gateway interruption that cleanly closes an established socket
can now recover once the gateway is healthy, instead of leaving the
runtime stuck and repeatedly logging the managed-session-down warning.
## Verification
- `pnpm nx run @copilotkit/channels-intelligence:test` — 188 tests
passed
- `pnpm nx run @copilotkit/channels-intelligence:check-types` — passed
- `pnpm nx format:check
--files=packages/channels-intelligence/src/realtime-gateway.ts,packages/channels-intelligence/src/realtime-gateway.test.ts`
— passed
- repository pre-commit test, package validation, and lint gates —
passed
pydantic-ai 1.0.18 imports opentelemetry._events, removed in
opentelemetry-api 1.44.0. Its own floor is an unbounded
opentelemetry-api>=1.28.0, so a fresh resolve of this package's
requirements picks 1.44.0 and import pydantic_ai fails with
ModuleNotFoundError.
The agent image kept building because the Dockerfile copies only
requirements.txt before pip install, so that layer stayed cached from
before 1.44.0 shipped (2026-07-16). Any requirements change or cache
eviction would have surfaced it.
Verified with pip in a clean 3.12 venv, as the Dockerfile installs:
before: opentelemetry-api 1.44.0, import pydantic_ai raises
after: opentelemetry-api 1.43.0, import pydantic_ai OK
## Summary
- serialize Slack Carousel cards through elements and validate Card and
Carousel payloads before provider calls
- keep bounded provider diagnostics from ChannelDeliveryError through
RUN_ERROR and ChannelCanonicalRunError
- log low-cardinality error fields without validation messages or
provider bodies
- move best-effort Slack status cleanup failures to debug logs
## Why
Invalid nested Slack Card payloads failed at the provider with no useful
path at the application boundary. This change catches known Card and
Carousel shape errors locally and preserves safe JSON pointers when
Slack rejects a payload.
## Compatibility
The new error details are optional. Existing consumers keep their
current behavior, so no protocol version bump is needed.
## Validation
- NX_DAEMON=false pnpm nx run-many -t test,check-types,build
--projects=@copilotkit/channels-core,@copilotkit/channels-intelligence,@copilotkit/channels-slack,@copilotkit/runtime
--skip-nx-cache
- pnpm lint
- pnpm check:channel-native-catalogs
- targeted oxfmt and oxlint checks
- git diff --check
- pre-commit test, publint, and attw checks for affected packages
Companion Intelligence PR:
https://github.com/CopilotKit/Intelligence/pull/759
This pull request was posted by Claude Code using claude-opus-5 on
behalf of David. David has not reviewed this diff line by line.
Closes https://github.com/CopilotKit/CopilotKit/issues/6363
`Agent.to_ag_ui()`, `AGUIApp` and the whole `pydantic_ai.ag_ui` module
were removed in Pydantic AI v2. The docs installed pydantic-ai unpinned,
so following the quickstart today gets 2.22.0 and fails twice: first at
resolution (`starlette==0.45.3` conflicts with the `>=0.46.2` the
`ag-ui` extra requires), then at `AttributeError`.
## What changed
**8 doc pages** under
`showcase/shell-docs/src/content/docs/integrations/pydantic-ai/`
(`quickstart.mdx`, `quickstart/pydantic-ai.mdx`,
`human-in-the-loop.mdx`, `human-in-the-loop/agent.mdx`,
`generative-ui/tool-rendering.mdx`, and the three `shared-state/`
pages):
- the agent is served from a Starlette route via
`AGUIAdapter.dispatch_request(request, agent=agent)`
- `StateDeps` imports move from `pydantic_ai.ag_ui` to `pydantic_ai.ui`
- install commands exact-pin `pydantic-ai-slim[ag-ui,openai]==2.22.0`
and `ag-ui-protocol==0.1.19`, matching the starter fleet, plus
`starlette>=0.46.2` since the snippets import Starlette directly
**Per-request deps.** Every stateful snippet builds `StateDeps` inside
the request handler:
```python
async def run_agent(request: Request) -> Response:
return await AGUIAdapter.dispatch_request(
request, agent=agent, deps=StateDeps(AgentState())
)
```
`dispatch_request` validates the client's state into `deps.state`
(`pydantic_ai/ui/_adapter.py`, `run_stream_native`), so a module-level
instance shared across requests lets concurrent runs clobber each other.
The old `to_ag_ui(deps=...)` snippets all did this.
**`examples/canvas/pydantic-ai`** — `requirements.txt` pinned,
`agent/agent.py` ported, README corrected.
**`examples/showcases/pydantic-ai-todos`** — `pyproject.toml` pinned and
`uv.lock` regenerated (it was still resolving 1.0.10), `agent/main.py`
ported, `src/agent.py` and `src/tools.py` imports moved, README and
`src/app/api/copilotkit/route.ts` comments corrected.
**`skills/copilotkit-integrations`** — beyond the issue's file list:
`SKILL.md`, `sources.md` and `references/integrations/pydantic-ai.md`
also taught `to_ag_ui()`. Same rot, same fix.
## Verified by execution
The reason these docs rotted is that nothing runs them, so everything
below was actually run, not read.
- Both install commands were run verbatim in throwaway environments. `uv
add 'pydantic-ai-slim[ag-ui,openai]==2.22.0' 'ag-ui-protocol==0.1.19'
'starlette>=0.46.2' uvicorn` and the `pip install` equivalent both
resolve, landing pydantic-ai-slim 2.22.0, ag-ui-protocol 0.1.19,
starlette 1.3.1.
- Every ```python fence on the 8 doc pages was extracted, `exec`'d, and
driven with a real `RunAgentInput` POST through
`starlette.testclient.TestClient` with the model overridden to
`TestModel`. All 8 return 200 `text/event-stream` with a `RUN_STARTED`
... `RUN_FINISHED` sequence and no `RUN_ERROR`.
- The canvas agent was installed from its `requirements.txt` and driven
the same way: 200, SSE, `RUN_STARTED` ... `TOOL_CALL_*` ...
`STATE_SNAPSHOT` ... `RUN_FINISHED`.
- The todos agent was installed with `uv sync --frozen` from the
regenerated lock and driven the same way. Two sequential requests, one
seeding a todo and one sending empty state, each saw only their own
state, confirming the per-request deps actually isolate.
Not executed: the Next.js frontends and the docs site build (no
`node_modules` in this checkout). The TypeScript edits are comment-only.
## Deliberately out of scope
`showcase/integrations/pydantic-ai` is left on its v1 fleet pin. It is
418 files, 19 mounts and 190 e2e specs, and CopilotKit said they will
take it as https://github.com/CopilotKit/CopilotKit/issues/6364. The
dojo and the docs therefore diverge until that lands.
The CI guard from the issue's last acceptance criterion is not built
here. A proposal for it is posted on
https://github.com/CopilotKit/CopilotKit/issues/6363 for the team to
own.
Two pre-existing malformed code fences were fixed in passing, because
leaving them meant the ported snippets still would not run:
`quickstart/pydantic-ai.mdx` and
`shared-state/predictive-state-updates.mdx` each had TypeScript embedded
inside an unterminated ```python fence. The TypeScript now sits in its
own fence.
Overlaps with https://github.com/CopilotKit/CopilotKit/pull/6355, which
ports `examples/integrations/pydantic-ai`. No file overlap.
## Release channels v0.7.2
**Scope:** `channels` | **Bump:** `patch`
---
### How this release process works
1. **This PR was created automatically** by the "release / create-pr"
workflow.
It bumped the `channels` packages to `0.7.2`
and generated AI-enhanced release notes.
2. **CI runs on this PR** — the full test suite (unit tests, lint, type
checks, build)
must pass before merging. This is the review gate.
3. **Review the release notes** in `release-notes.md` in this PR.
If a Notion draft was created, you can edit the release notes there
before merging.
4. **When this PR is merged**, the `release / publish` workflow
automatically:
- Builds all packages
- Publishes the `channels` packages to npm at version `0.7.2`
- Creates git tag `channels/v0.7.2`
- Creates a GitHub Release with the final release notes
### Before merging
- [ ] CI is green (tests, lint, types, build)
- [ ] Version bumps look correct
- [ ] Release notes are accurate (edit in Notion if a draft was created)
---
> **Do not merge until CI is fully green.** The full test suite runs
automatically on this PR.
## Summary
Retesting the Slack setup end to end found the skill **too cautious to
be useful**. It stopped at almost every step, where the earlier version
stopped only for passwords and got through — and the run was rescued by
the developer telling the agent outright to just control their browser,
which cut human involvement down to typing passwords.
A run that pauses at every control is slower than the manual path it
replaced. Three changes, all pulling the same direction.
## 1. Driving is now the default, not implicit
The skill said *"most of this workflow happens in a browser"* —
descriptive, and it never told the agent to **drive** that browser. It
now does, explicitly, and checks its own capability before Phase 0
rather than assuming either way.
## 2. When there is no browser, ask for one — per harness
Generic advice is useless here, because enabling browser control differs
by harness. The agent now works out which one it is in and names the
single applicable route: Claude Code and Codex each ship their own
support and enable it differently, most other harnesses take a general
browser-use MCP server such as Playwright MCP. **If it isn't sure, it
looks it up rather than guessing.**
It also names the payoff — driving turns this into typing three secrets,
the fallback is roughly fifteen manual browser steps — because that is
what turns a shrug into a yes. The step-by-step walkthrough survives
only for an explicit decline.
## 3. Consent is batched into one authorization
Phase 0 now takes **one** yes naming the whole sequence: the Slack app
from the wizard manifest, its install into a named workspace, the
Channel, the adapter attach, and the project key.
Phase 3 and `references/intelligence-channel.md` previously required
*"state what you are about to change, get an explicit yes"* for
**every** dashboard goal — and the reference said it twice, back to
back. That is the concrete source of the stop-at-every-step behaviour.
Reading the page before acting stays required; it is no longer a reason
to check in.
## What did not change
The secret boundary. The developer still types the bot token, signing
secret, and API key themselves, and those remain the only mandatory
stops alongside anything the authorization did not cover. Batching
consent must not batch away a password — that is called out in the text.
`version` bumped 1.0.0 → 1.1.0.
## Validation
- `tsx scripts/sync-plugin-skills.ts --check` → **plugin skill mirror in
sync**
- `scripts/__tests__/sync-plugin-skills.test.ts` → 10 tests passed
- `prettier --check` clean on both changed files
## Companion
The same shift landed on the hosted big prompt in
[CopilotKit/website#445](https://github.com/CopilotKit/website/pull/445)
— default to driving, one batched authorization, harness-aware
capability request. Both surfaces now say the same thing, which was the
point of keeping them in step.
Worth noting for reviewers: this takes Atai's side on *"where the skill
asks permission to proceed, just do it"*, which was only half-applied
before. It also reverses my own per-step confirmations from the first
pass on #445 — the dogfooding run is the reason.
Batching the Phase 0 authorization removed the per-goal confirmations that were
accidentally serving as decision points. Nothing then asked the developer for the
inputs the agent cannot legitimately choose, and Phase 1 still said "Enter a
Display name" in the imperative — so an autonomous run named the bot itself.
That name is the expensive one. The wizard derives the Channel Code from it, the
Code is what createChannel({ name }) declares and what the developer types as
/invite, and Slack bot names are workspace-wide — Phase 1 already warns that a
collision blocks the install. An agent that settles it has named someone's bot for
them and can fail the install doing it.
Phase 0 now gathers four decisions in one exchange before any browser opens: the
display name, the workspace, the test channel, and whether this is throwaway.
Phase 1 consumes the chosen name instead of inventing one, Phase 1's workspace step
uses the named workspace, and Phase 2's invite names the agreed channel and says
the developer runs it.
States the rule the whole design turns on: the decisions are inputs you cannot
invent, the authorization is permission you need once, and collapsing the second
does not license skipping the first.
## Summary
- stop handling in-flight lock renewal failures after the run has
settled
- keep aborting when a renewal fails during an active run
- cover the completion-before-renewal race with a regression test
## Why
Intelligence releases the thread lock after it accepts a terminal run
event. A renewal that was already in flight can then return a 409.
Clearing the interval stops future renewals, but it does not cancel that
pending promise, so Runtime logged an error and called `abortRun()`
after the run had completed.
The lifecycle guard makes that late rejection a no-op. Active-run
renewal failures still follow the existing abort path.
## Testing
- `pnpm nx test @copilotkit/runtime` — 1,866 tests passed
- `pnpm nx run @copilotkit/runtime:check-types`
- `pnpm nx build @copilotkit/runtime`
- `pnpm exec oxlint
packages/runtime/src/v2/runtime/handlers/intelligence/run.ts
packages/runtime/src/v2/runtime/__tests__/intelligence-lock-heartbeat.test.ts`
- `pnpm exec oxfmt --check
packages/runtime/src/v2/runtime/handlers/intelligence/run.ts
packages/runtime/src/v2/runtime/__tests__/intelligence-lock-heartbeat.test.ts`
The pre-commit hook also passed affected tests, `publint`, and `attw`.
The repo-wide `pnpm check-format` still reports 25 unrelated files
already present on `main`; both changed files pass the focused format
check.
Retesting the Slack setup end to end found the skill too cautious to be useful.
It stopped at almost every step, where the earlier version stopped only for
passwords and got through — and the run was rescued by the developer telling the
agent outright to just control their browser. A run that pauses at every control
is slower than the manual path it replaced.
Three changes, all pulling the same direction.
Driving the browser is now stated as the default rather than left implicit in
"most of this workflow happens in a browser". When the agent has no browser tool
it asks the developer to install one before starting, and names the route for the
harness it is actually running in — Claude Code and Codex enable this differently,
most other harnesses want a browser-use MCP server — with an instruction to look
it up rather than guess. The manual walkthrough survives only for an explicit
decline.
Consent is batched into one Phase 0 authorization naming the whole sequence: the
Slack app from the wizard manifest, its install, the Channel, the adapter attach,
and the API key. Phase 3 and the Intelligence reference previously required "state
what you are about to change, get an explicit yes" for every dashboard goal, and
the reference said it twice. Reading the page before acting stays required; it is
no longer a reason to stop.
The secret boundary is untouched: the developer still types the bot token, signing
secret, and API key themselves, and those remain the only mandatory stops
alongside anything the authorization did not cover.
Follows the maintainer's Correction #2 on issue 6363. An exact version in a
docs install command is the same rot as the starlette==0.45.3 pin it replaced:
it goes stale silently and nobody re-resolves prose. The 2.22.0 the docs shipped
was already a version behind current the day it was written.
- docs install lines use pydantic-ai-slim[ag-ui,openai]>=2,<3, which constrains
the dep the pages actually care about and fails loudly at the v3 boundary
- ag-ui-protocol drops out of the docs lines entirely; no doc snippet imports
ag_ui, so naming it there was the transitive-dep noise the correction is about
- starlette>=0.46.2 stays, because the v2 snippets import Starlette directly.
A floor with no ceiling cannot force a downgrade, so it does not recreate the
silent backtrack
- examples/showcases/pydantic-ai-todos moves to a range in pyproject.toml and
relocks; the uv.lock is what reproduces
- examples/canvas/pydantic-ai keeps exact pins: it has no lockfile, so
requirements.txt is its only reproducibility artifact
Smoke-tested the open question from the issue: starlette 1.x works on
pydantic-ai v2. All 8 doc pages pass on 2.23.0 + starlette 1.3.1 and on
2.23.0 + starlette 0.52.1, so Jordan's <1.0 guard can be dropped rather
than raised.