## Summary
In multi-turn conversations using the Anthropic adapter, a model turn
that contains both assistant text and a tool call can be replayed as two
consecutive `{role: "assistant"}` entries in the Anthropic payload.
Anthropic expects one message object per turn with alternating roles, so
that split payload can blur turn boundaries on the next request. This PR
coalesces same-role Anthropic messages before dispatch so one assistant
turn stays one assistant message.
## Root cause
[`convertMessageToAnthropicMessage`](https://github.com/CopilotKit/CopilotKit/blob/005aebbededbfdd7978b3c1ee221580b74dd088d/packages/runtime/src/service-adapters/anthropic/utils.ts#L142-L219)
maps each CopilotKit message independently. A mixed assistant turn, a
`TextMessage(role=assistant)` followed by an `ActionExecutionMessage`,
therefore becomes two separate assistant entries. The Anthropic Messages
API requires all content blocks for one assistant turn to be sent in a
single message object: https://docs.anthropic.com/en/api/messages
[`AnthropicAdapter.process()`](https://github.com/CopilotKit/CopilotKit/blob/005aebbededbfdd7978b3c1ee221580b74dd088d/packages/runtime/src/service-adapters/anthropic/anthropic-adapter.ts#L282-L366)
already deduplicates `tool_result` blocks on the user side, but it
previously forwarded the mapped assistant-side payload without a
coalescing pass. Rebasing onto current `main` also exposed test-fixture
drift in the two new regression tests, so the final branch updates those
fixtures to the current object-form `TextMessage` constructor and
removes one stale `ActionInput` field while keeping the production fix
unchanged.
## Changes
- `packages/runtime/src/service-adapters/anthropic/utils.ts`: add
`coalesceConsecutiveSameRoleMessages(messages)` to merge adjacent
equal-role
Anthropic messages by concatenating their `content` arrays.
-
`packages/runtime/src/service-adapters/anthropic/anthropic-adapter.ts`:
call
`coalesceConsecutiveSameRoleMessages` before
`limitMessagesToTokenCount`.
-
`packages/runtime/tests/service-adapters/anthropic/anthropic-adapter.test.ts`:
add same-role regression coverage, align the new fixtures with the
current
object-form `TextMessage` constructor, and remove the stale `parameters`
field
from the action fixture.
- `.changeset/coalesce-anthropic-same-role-messages.md`: patch changeset
for
`@copilotkit/runtime`.
## Scope
The `tool_result` allowlist deduplication stays unchanged. Token
trimming and the orphan-removal post-processor still run after
coalescing, so `tool_use` and `tool_result` pairing remains intact. The
OpenAI, Google, LangChain, and Groq adapters are unaffected.
The reported regression and the explicit regression coverage are
assistant-side. The coalescing helper itself is role-agnostic for
adjacent equal-role array content, but this PR does not add a separate
user-side regression case.
## Related PRs and Issues
Prior attempted fix: #2864, which addressed unrelated message callbacks
rather than the adapter payload.
## Test plan
- [x] `pnpm -C packages/runtime exec vitest run
tests/service-adapters/anthropic/anthropic-adapter.test.ts`
10/10 tests pass, including the two same-role coalescing cases.
- [x] `pnpm -C packages/runtime exec vitest run
tests/service-adapters/anthropic/utils-token-trimming.test.ts`
9/9 tests pass.
- [ ] CI green (`static / quality`, `test / unit` on Node 20/22/24)
## Upstream
Closes#2910.
Reported by @alonronin.
## Problem
A frontend/client tool (`useFrontendTool`) whose Zod `parameters` use
`z.discriminatedUnion(...)` — or any schema that serializes to a
JSON-schema `anyOf`/`oneOf` node — silently loses the union-typed field
when calling OpenAI. The tool call arrives with that field missing or
empty. Switching the same tool to a flat object with an enum
discriminant works, which points at schema conversion rather than the
model.
## Root cause
The runtime has **two** JSON-Schema → Zod converters:
- `@copilotkit/shared`'s `convertJsonSchemaToZodSchema` already handles
`anyOf`/`oneOf` as `z.union` (and `$ref`, null-unions, graceful
fallback).
- The **local copy** in `packages/runtime/src/agent/index.ts` — used by
the classic `BuiltInAgent` / AI SDK path via
`convertToolsToVercelAITools` — never did.
A union node carries no top-level `type`, so it hit the empty-schema
guard (`if (!jsonSchema.type)`) and collapsed to `z.object({})`. The
reconstructed tool schema therefore dropped the union entirely — most
visibly for a union nested inside array `items` — so the model was never
offered those fields and could not emit them. (The legacy GraphQL OpenAI
adapter is unaffected: it forwards the JSON schema directly.)
## Fix
Handle `anyOf`/`oneOf` as `z.union` **before** the empty-schema guard,
mirroring the already-proven shared converter. A single-variant union
unwraps to that variant.
## Backward compatibility
Only previously-broken union nodes change behavior (empty object → real
union). Empty `{}` schemas, typed nodes, and the `isJsonSchema` gate are
untouched. Two regression tests added: a direct `anyOf` conversion and
the exact nested-`items` `oneOf` trap.
## Repro
A `useFrontendTool` with
```ts
parameters: z.object({
blocks: z.array(z.discriminatedUnion("type", [
z.object({ type: z.literal("heading"), level: z.number() }),
z.object({ type: z.literal("paragraph"), content: z.string() }),
])),
})
```
against OpenAI: before, `blocks` items arrived empty; after, both
variants survive into the model call.
## Note on OpenAI strict mode
This path does not enable OpenAI strict function-calling, so once the
union survives conversion it serializes back to `anyOf` and OpenAI
accepts it. The data loss was upstream, in CopilotKit's own converter,
not an OpenAI strict-mode limitation.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
## Problem
`resolveModel()` builds each provider with only an `apiKey`, so a
**string model spec** (e.g. `"openai/gpt-4o-mini"`) cannot target an
**OpenAI-compatible endpoint** — Azure OpenAI, OpenRouter, an LLM
gateway, or a local server (vLLM / LM Studio / Ollama).
`OPENAI_BASE_URL` is silently ignored, and the only workaround is to
drop the ergonomic string form and construct a `LanguageModel` instance
yourself. The same gap exists for Anthropic and Google.
## Fix
`resolveModel` now passes `baseURL` from the standard env var for each
provider:
| Provider | Env var |
|---|---|
| OpenAI | `OPENAI_BASE_URL` |
| Anthropic | `ANTHROPIC_BASE_URL` |
| Google | `GOOGLE_GENERATIVE_AI_BASE_URL` |
It is `undefined` when the var is unset, so each provider falls back to
its default endpoint — **fully backward compatible** (no behavior change
unless you opt in).
## Tests
Adds `resolve-model-baseurl.test.ts` (4 cases): each provider forwards
its env var to the SDK factory, and an unset var leaves `baseURL`
undefined. All existing `resolveModel` tests pass unchanged (118 runtime
tests green locally).
## Context
Found while building an OpenAI-compatible agent on the V2 runtime +
`@copilotkit/react-native`, where the string model form couldn't reach
the configured endpoint.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
## Summary
**Plan C: channels are an Intelligence capability.** A valid
Intelligence configuration (API key) is required to run any channel.
There's a **free tier**, so this is "connect your account," not "pay."
The Intelligence runtime runs every channel: managed (Slack/Teams) over
the gateway, and **direct-adapter** channels (Discord/Telegram/WhatsApp,
or self-hosted Slack/Teams) via their own credentialed transport,
started by the runtime. The SSE runtime type-rejects `channels`, and the
handler builds channel activation for an Intelligence runtime.
> Note: the `channel.ɵruntime.{start,stop,addAdapter}` seam is an
internal (`ɵ`) contract the runtime and the managed launcher drive — not
a public API.
Design doc: [Channels = Intelligence-only (Plan
C)](https://app.notion.com/p/3a63aa381852814e83dafedc0742e7cc)
## What changed
- **`refactor(channels-core)`** — relocate the `Channel` lifecycle onto
an internal `channel.ɵruntime.{start,stop,addAdapter}` seam.
- **`feat(channels-core)!`** — **remove** public
`Channel.start()/stop()/addAdapter()`. Channels are runtime-driven only.
- **`feat(runtime)!`** — the `ChannelManager` (built only for an
Intelligence runtime) now **starts direct-adapter channels** via
`ɵruntime.start()` instead of recording them `"unmanaged"` and skipping.
Dead `"unmanaged"` status removed. Direct channels reuse the same
bounded/idempotent/resilient teardown via a synthetic handle.
- **`fix(channels-core)`** — a channel whose adapters **all** fail to
start now **errors** instead of falsely reporting `online` (a partial
start, ≥1 adapter live, still counts as started).
- **`fix(examples)`** — examples bound `channels.ready({ timeoutMs })`
so a wedged adapter start can't hang readiness.
- **`docs`** — 7 channels READMEs + the teams/slack examples reworked to
`new CopilotRuntime({ intelligence, channels })` +
`handler.channels.ready()/stop()`; no `channel.start()`, no DIY.
Multi-platform slack now runs under Intelligence (one `ɵruntime.start()`
starts all its direct adapters).
**Breaking:** `Channel.start()/stop()/addAdapter()` removed
(`channels-core` 0.2.x) — channels are driven by the runtime (`new
CopilotRuntime({ intelligence, channels })`).
## Validation
**Unit** (per-package raw `tsc` + `vitest`): channels-core **156**,
channels-intelligence **182**, runtime channel-manager **38** (incl. a
new `failStart` regression test), channels-integration 8.
**Independent whole-branch review:** direct-channel lifecycle traced
correct for every start/stop interleaving (no wedge, single-stop,
bounded/idempotent teardown shared with the managed path); public-API
removal complete (zero callers); no DIY/`channelRunner`/guard-relaxation
debris; SSE guard intact.
**Adversarial review** (tried to break it) surfaced two real items, both
handled here:
- **False `online`** — `ɵruntime.start()` swallowed adapter start
failures and resolved, so a dead channel read `online`. **Fixed**
(`fix(channels-core)` + `failStart` test).
- **Overstated invariant** — "no standalone path" is enforced by the
runtime gate + `ɵ` convention, not the type system. **Clarified** (see
the note above). Also bounded example `ready()` (`fix(examples)`).
**Live run** against a real Intelligence key (**7/7**): the gate rejects
channels-without-Intelligence; the runtime starts a direct channel via
`ɵruntime.start()` (a real turn invoked the agent, clean `stop()`); and
the real key **authenticated against the real gateway** (a managed join
was rejected by a project feature-flag, not an auth error — so the key
was accepted).
> Examples typecheck only in CI — the worktree can't build every
`@copilotkit/*` sibling; those `TS2307`/`TS2305` are environmental.
## Related
- **Improvements (fast-follow):**
[OSS-599](https://linear.app/copilotkit/issue/OSS-599) — §2 response
policy, four-mode agent binding, run-correctness (Intelligence-side).
- **Managed-provider parity:**
[OSS-600](https://linear.app/copilotkit/issue/OSS-600) +
[601](https://linear.app/copilotkit/issue/OSS-601)/[602](https://linear.app/copilotkit/issue/OSS-602)/[603](https://linear.app/copilotkit/issue/OSS-603)
(Discord/Telegram/WhatsApp) — hosted transport for the remaining
platforms; also closes the direct-channel key-validation gap (managed =
gateway-validated).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Adversarial-review finding: ɵruntime.start() swallowed every adapter start failure
and resolved, so the runtime reported status "online" on a dead channel (revoked
token, port-in-use). Now: if a channel has adapters and NONE started, start() rejects
so ChannelManager surfaces "error" (a partial start, >=1 adapter live, still counts as
started). Adds a channel-manager failStart regression test + updates the telemetry test.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The ChannelManager (only constructed for an Intelligence runtime) now STARTS
direct-adapter channels via channel.ɵruntime.start() instead of recording them
"unmanaged" and skipping — so every channel runs only because Intelligence is
configured; there is no standalone path. Removed the now-dead "unmanaged" status;
direct channels reuse the bounded/idempotent teardown via a synthetic handle.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The vi.fn mock factories were 0-arg but invoked with the provider options in
the vi.mock adapters, tripping tsc (TS2554). vitest transpiles without type-
checking, so it passed vitest locally but failed the check-types job.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
_def.typeName isn't on ZodTypeDef's public type (TS2339 in check-types);
instanceof is the type-safe equivalent.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A frontend/client tool (`useFrontendTool`) whose Zod `parameters` use
`z.discriminatedUnion(...)` — or any schema that serializes to a JSON-schema
`anyOf`/`oneOf` node — silently lost the union-typed field when calling OpenAI.
The tool call arrived with that field missing or empty.
Root cause: the runtime has two JSON-Schema -> Zod converters. The shared one
(`@copilotkit/shared`) already handles `anyOf`/`oneOf` as `z.union`, but the
local copy in `packages/runtime/src/agent/index.ts` — used by the classic
`BuiltInAgent` / AI SDK path via `convertToolsToVercelAITools` — never did.
A union node carries no top-level `type`, so it hit the empty-schema guard and
collapsed to `z.object({})`. The reconstructed tool schema therefore dropped
the union entirely (most visibly for a union nested inside array `items`), so
the model was never offered those fields and could not emit them.
Fix: handle `anyOf`/`oneOf` as `z.union` before the empty-schema guard,
mirroring the shared converter. A single-variant union unwraps to that variant.
Backward compatible: only previously-broken union nodes change behavior
(empty object -> real union). Empty `{}` schemas, typed nodes, and the
`isJsonSchema` gate are untouched.
Repro: a `useFrontendTool` with `parameters: z.object({ blocks: z.array(
z.discriminatedUnion("type", [...])) })` against OpenAI — before, `blocks`
items arrived empty; after, both variants survive into the model call.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
resolveModel() built each provider with only an apiKey, so a string model
spec (e.g. "openai/gpt-4o-mini") could not target an OpenAI-compatible
endpoint — OPENAI_BASE_URL was silently ignored, and the only workaround was
to construct the LanguageModel yourself and pass the instance. Same gap for
Anthropic and Google.
Now resolveModel passes baseURL from the standard env var
(OPENAI_BASE_URL / ANTHROPIC_BASE_URL / GOOGLE_GENERATIVE_AI_BASE_URL). It is
undefined when unset, so the provider falls back to its default endpoint —
fully backward compatible.
Tests: adds resolve-model-baseurl.test.ts (4 cases). Existing resolveModel
tests unchanged (118 runtime tests pass locally).
Discovered while building an OpenAI-compatible agent on the V2 runtime +
@copilotkit/react-native.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The optional peer dependency on @anthropic-ai/sdk was pinned to
^0.57.0, which conflicts with current SDK versions (e.g. 0.109) and
forces consumers to install @copilotkit/runtime with --legacy-peer-deps.
The anthropic adapter only relies on stable @anthropic-ai/sdk APIs
(client construction, messages.create, ephemeral cache_control), so
loosen the range to >=0.57.0 to remove the false peer conflict.
pnpm auto-installs this optional peer, so its specifier is mirrored in
the lockfile; update that specifier line to match. The resolved version
(0.57.0) still satisfies the range, so resolution is otherwise unchanged
and `pnpm install --frozen-lockfile` passes.
Reported by an outside contributor.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
## Summary
- Extract the platform-neutral Channels foundation into
`@copilotkit/channels-core`.
- Make `@copilotkit/channels` the batteries-included consumer entry
point, with adapter and UI subpaths.
- Release the umbrella, core, UI, and all six adapters as one shared
`channels` version scope.
- Verify the packed consumer contract and migrate the Slack and Teams
examples to the umbrella.
## Why
Consumers should be able to install one Channels package without making
the runtime or selective integrations depend on every platform SDK.
Shipping the complete Channels family together prevents adapter/core
version drift and makes the umbrella's exact dependency set release as a
compatible unit.
## How
- Move shared bot/runtime primitives into `channels-core` and invert
adapter/runtime dependencies.
- Add exact workspace dependencies and export subpaths from the umbrella
package.
- Consolidate the existing release configuration and all three release
workflow selectors into one shared `channels` scope containing all nine
packages.
- Use `@copilotkit/channels` as the version source: the next minor
release resolves to `0.2.0` and bumps every Channels package together.
- Publish scoped packages in dependency order for both stable and canary
releases: UI, core, adapters, then the umbrella.
- For a stable Channels release, publish the UI/core/adapters first, run
the registry-backed packed-consumer verifier against those newly
published exact versions, then publish the umbrella.
- Generate Channels release notes from `channels/v*` tags rather than
the monorepo `v*` tags.
- Verify builds, type checks, tests, package artifacts, examples, and
release workflow scope synchronization locally.
### First stable release sequence
1. Bootstrap the currently unpublished `@copilotkit/channels-core`
package on npm and configure npm trusted publishing for every Channels
package against this repository's `release / publish` workflow and `npm`
environment.
2. Create and merge the `channels` minor release PR. The stable workflow
publishes the Channels family in the staged order above and validates
the packed umbrella from the registry before the umbrella is released.
3. Create and merge a subsequent `monorepo` release PR so the published
`@copilotkit/runtime` switches from the historical umbrella dependency
to `@copilotkit/channels-core`.
- [P1] History/files were absent on the NORMAL managed path: defaultActivateChannel
never forwarded the app-api HTTP URL, so the transport (which installs
fetchFile/getHistory/uploadFile only when appApiBaseUrl is set) ran without
them for Channels started by the CopilotRuntime handler — only manual
low-level launcher callers got file/history. Thread intelligence.ɵgetApiUrl()
through ChannelActivationConfig.apiUrl → defaultActivateChannel →
startChannelsOverRealtimeGateway({ appApiBaseUrl }). The launcher + transport
already accepted it.
- [P2] Managed turns exposed the provider profile under a non-public `displayName`
field, leaving PlatformUser.name undefined. Map env.user.displayName -> name
(parity with the direct Slack adapter, which populates `name`).
Tests: deriver returns apiUrl; defaultActivateChannel forwards appApiBaseUrl to
the launcher opts; onMessage sees message.user.name. channel-activation-config +
channels-intelligence (170) green; runtime build type-checks. (channel-manager.test
executes in CI — local vitest hits the known optional-peer-dep resolution flake.)
Two fixes from the pre-merge adversarial CR of this PR (both pre-existing,
flagged as in-subject):
- ChannelManager.status() reported overall "online" for a manager stopped
BEFORE activate() (e.g. SIGTERM during startup): `entries` is empty, so the
empty-set fold returned "online" — a torn-down manager reading healthy. Now
short-circuits to "stopped" when `this.stopped`, matching the documented
status() contract. New red-green test covers the stop()-before-activate() case.
- examples/slack/.env.example: COPILOTKIT_INTELLIGENCE_WS_URL example was
ws://localhost:4401, but derivation is a scheme-only swap of the :4201 API URL
(→ ws://localhost:4201) and 4401 is used nowhere — a user uncommenting it hit a
dead port. Corrected to :4201 and clarified the derivation note.
Three ChannelManager observability/robustness hardenings from the OSS-473 CR.
stop() per-handle timeout: stopEntry awaited handle.stop() unbounded, so a
wedged stop() (e.g. a socket.disconnect that never returns) hung teardown — and
thus SIGTERM shutdown — forever. Each handle.stop() is now bounded by
stopHandleTimeoutMs (default 5000); on timeout it is logged and abandoned so
every other entry still reaches `stopped`.
ready() surfaces the real reason on hang: ready({ timeoutMs }) wrapped the whole
`allSettled` in one timeout, so when one channel settled to `error` while a
sibling hung, it rejected with only a generic timeout and DISCARDED the erroring
channel's reason. The deadline is now applied PER CHANNEL, so the AggregateError
carries both each failed channel's real reason AND a named timeout for each
still-hanging channel.
Log forwarding: the manager's `log` reached activation-level events only; the
default engine (defaultActivateChannel → startChannelsOverRealtimeGateway) never
passed it down, so transport-level drop diagnostics (e.g. a version-skew
missing-leaseToken outage) were silent in the managed path. `log` is now
forwarded down to the launcher/transport.
Note: the reconnect "gave-up → error" escalation from the same CR is already
implemented end-to-end in OSS-473 (realtime-gateway.ts reconnectGiveUpMs +
ChannelManager.registerConnectionObserver) and is not re-done here.
Tests: stop() resolves + logs a timeout when handle.stop() never settles;
ready() aggregate contains both a real activation error and the hung sibling's
named timeout; defaultActivateChannel forwards its log sink to the launcher opts.
The PR's documented snippet — `await handler.channels.ready(...)` with no
`!`/`?.` — did not type-check under strict TS because
`createCopilotRuntimeHandler` always returned `channels?: ChannelsControl`.
Encode channel-presence at the type level:
- runtime.ts: `CopilotRuntime` is now a `const` typed as `CopilotRuntimeConstructor`
(backed by an internal `CopilotRuntimeShim` class; behavior unchanged). A
class constructor cannot vary its return type across overloads, so the two
construct-signature overloads live on the constructor interface: `intelligence`
+ a non-empty `channels` tuple returns a `RuntimeWithDeclaredChannels`-branded
runtime; every other config (SSE, intelligence-without-channels, empty
`channels: []`, or a non-literal `Channel[]` variable) stays unbranded. The
brand is a phantom (compile-time-only) property. `export interface CopilotRuntime`
preserves the name as a type for existing `runtime: CopilotRuntime` / `as
CopilotRuntime` sites.
- fetch-handler.ts: overload `createCopilotRuntimeHandler` — a branded runtime
(unless `activateChannels: false`, constrained to `true | undefined`) returns
the new `CopilotRuntimeFetchHandlerWithChannels` (non-optional `channels`);
everything else keeps the optional shape. Opting out of activation honestly
falls through to the optional overload.
- Added a compile-time type test (checked by `tsc --noEmit`, the `check-types`
gate). It probes the optionality modifier structurally (`{} extends Pick<T,K>`)
rather than for `undefined`, since this package compiles `strict: false`.
Confirmed it fails pre-change on the required-channels assertion and passes
after. Dropped the now-unnecessary `!` in handler-channels.test.ts.
Call sites: the second overload is byte-identical to the former single signature,
so every `createCopilotRuntimeHandler` caller (node/express/hono endpoints,
integration servers, examples) and every `new CopilotRuntime` site resolves
unchanged; only inline non-empty-`channels` construction gains the (strict
supertype-assignable) branded type. Verified via a clean full-package check-types.
P1#2 — reachable setup_required on the PRODUCTION engine path.
connectRealtimeGateway no longer flattens every join rejection into a
generic Error. A join `.receive("error", reason)` whose reason is a known
setup-required code (`channel_declaration_unavailable`, and defensively
`adapter_setup_required` / `not_configured`) now rejects with a
distinguishable `RealtimeGatewaySetupRequiredError` (`code === "SETUP_REQUIRED"`,
raw reason preserved). ChannelManager already detects that code, so an
unconfigured managed provider now degrades to `setup_required` (ready()
resolves) instead of `error`. All other reasons keep the generic error and
the socket-leak teardown is unchanged.
P1#3 — status() reflects real connection health instead of `online` forever.
ConnectedRealtimeGatewaySession exposes `onStateChange(cb)` over
`RealtimeGatewayConnectionState` (`online` | `reconnecting` | `gave_up`),
driven by the real Phoenix seams: an unexpected socket drop → `reconnecting`;
a successful (re)join (the join-push recHooks survive Phoenix `resend`, so
`"ok"` re-fires on every auto-rejoin) → `online`; and a BOUNDED give-up —
Phoenix retries forever, so a `reconnectGiveUpMs` window (default 60000, runs
from the first drop of an episode, cleared on rejoin) elapsing while still
reconnecting → `gave_up` (terminal). Our own disconnect() stays silent.
ChannelManager wires this in place of the log-only onClose breadcrumb:
`reconnecting`→status reconnecting, `online`→online, `gave_up`→error; a
stopped manager/entry ignores late events. computeOverall now ranks
`error > reconnecting > setup_required > connecting > online`. ready() keeps
its one-shot semantics (settles on the initial outcome); later health
transitions move only status(). Docs updated to state `online` means
currently-sendable.
Wording: the direct-adapter skip comment/log now states delivery is
exclusive-per-platform (managed OR direct per platform, not both — attaching
both would double-deliver) with true coexistence tracked in OSS-484. Skip
behavior unchanged.
Call-sites for the changed signatures:
- connectRealtimeGateway error shape: only caller is
startChannelsOverRealtimeGateway (realtime-gateway-launcher.ts:216); it
awaits and lets the rejection propagate, so the setup-required error flows
through unchanged (no branch to update).
- new ConnectedRealtimeGatewaySession.onStateChange: passed through in
startChannelsWithGatewaySession and startChannelsOverRealtimeGateway
(realtime-gateway-launcher.ts); added to ChannelsHandle (runtime.ts) and the
manager's local ChannelsHandle view (channel-manager.ts); exported from
index.ts. RealtimeGatewaySession (base, no observer) consumers
(realtime-gateway-transport.ts) unaffected.
- manager onClose→state transitions: registerOnClose renamed to
registerConnectionObserver; sole caller is the online settle handler in
activate().
Honors the SoT "never infer managed intent from a direct adapter" rule.
Channel.adapters is a new additive read-only member; verified no consumer
constructs Channel literals (only createChannel does).
RC15: getOrCreateChannelManager bridged the manager log as
`logger.warn({ meta }, msg)`, but pino only serializes an Error's
(non-enumerable) message/stack under the `err` key — under `meta` a
failed activation rendered as `{}`, losing the cause. Route an Error to
`err` and keep `meta` for everything else.
LEVER: removed the channel-name FORMAT/length validation block plus the
replicated CHANNEL_NAME_PATTERN / MIN / MAX constants from
channel-activation-config.ts. This was a third copy of managed-specific
rules whose source of truth is channels-intelligence's
assertValidChannelRealtimeScope + assertValidChannelNames, and it kept
drifting (omitted the reserved-name rule). Now that activation failures
are logged, recorded as `error` status, and surfaced via ready(), the
up-front check is not worth cross-package rule parity. Missing/empty
name still throws (the config's own precondition). Deleted the obsolete
"Slack"/"support_bot"/"cs"/65-char rejection tests.
projectId>0: parseProjectIdFromApiKey now throws ChannelConfigError when
the parsed id is <= 0 (`cpk-0_...` matched but failed deep in the
launcher). Parser validating its own output, not a channel-name replica,
so it stays here; reuses the existing key redaction.
adapter default hardening: `adapter ?? "slack"` -> truthiness/trim check
so ""/whitespace falls back to "slack".
activate() stopped-guard: short-circuits on `this.activated || this.stopped`
so a post-stop() activate() opens no transports on a dead manager.
coverage: exported defaultActivateChannel (the real engine) with an
injectable importer seam (optional param, default = the same non-literal
dynamic import) and covered its 3 branches — config->opts mapping (scope
carries only projectId+channelName), module-not-found friendly error, and
generic-error passthrough. Added cheap manager coverage: lazy-activate
duplicate-name reject via ready(), empty channels[] -> online + ready
resolves, non-default adapter reaches the engine config.
Call sites (no external breakage):
- CHANNEL_NAME_PATTERN/MIN/MAX: were module-private; zero references.
- parseProjectIdFromApiKey: only internal caller is
deriveChannelActivationConfig (passes the real key) + tests; <=0 throw
affects only malformed cpk-0 keys.
- defaultActivateChannel: only internal use is the ChannelManager
constructor default (called 2-arg) + tests; new 3rd param is optional
and the fn still satisfies ActivateChannelEngine.
- ChannelsIntelligenceModule: new additive export, referenced internally
+ tests only.
- activate() guard / logger bridge: internal only.
CR batch for packages/runtime managed-channels activation:
- RC11: getOrCreateChannelManager now passes a `log` adapter bridging the
ChannelManager diagnostic sink to the shared logger
(`log: (msg, meta) => logger.warn({ meta }, msg)`). Previously every
breadcrumb (setup_required, failed-to-activate, dropped-session,
teardown-stop failure) was a no-op, so a channel that failed to activate
was permanently dead with zero output.
- f2: stopEntry logs a swallowed handle.stop() error via the sink instead of
discarding it; teardown stays resilient (never rethrows).
- sync-throw guard: stopEntry wraps handle.stop() in
`Promise.resolve().then(...)` so a foreign/injected handle that throws
SYNCHRONOUSLY is caught by the same `.catch` — otherwise the throw escaped,
skipped resolveSettled(), and hung `settled` forever.
- f3: ready() short-circuits and RESOLVES when the manager is stopped. A
channel that settled to `error` before stop() had already rejected its
`settled` promise, so a later ready() threw an AggregateError even though
status().overall was "stopped" — now consistent with the after-stop case.
- RC12: parseProjectIdFromApiKey no longer slices a fixed 8 chars off an
arbitrary key (which echoed secret bytes for a `cpk-_...`-shaped key). The
failure message now echoes NONE of the key value, only the expected
`cpk-{projectId}_` format hint.
- RC13: deriveChannelActivationConfig enforces the lowercase-kebab-case
channel-name rule (/^[a-z][a-z0-9]*(?:-[a-z0-9]+)*$/, length 3-64) up front
and throws a clear ChannelConfigError, instead of passing a bad name to the
launcher where assertValidChannelRealtimeScope throws deep and the channel
is silently degraded to `error`. Regex/bounds are a literal copy of
channels-intelligence's assertValidChannelRealtimeScope (the source of
truth; not statically imported — it's an optional pure-ESM peer).
- ready() docblock reworded: async ready() REJECTS (not throws) the
ChannelConfigError.
Tests (red-green verified): RC11 (handler logger spy + manager log sink),
f3 (ready resolves post-stop after pre-stop error), sync-throw guard, RC12
(no secret-tail leak), RC13 (Slack/support_bot/cs/65-char reject; support
passes).
Call-site enumeration for changed signatures/behaviors:
- getOrCreateChannelManager (new internal `log` arg): single caller
fetch-handler.ts:402; public signature unchanged, no caller impact.
- deriveChannelActivationConfig (RC13 now throws on bad name): single caller
channel-manager.ts:323 inside activate()'s per-channel loop, already
wrapped in try/catch that converts a throw to a rejected activation ->
recorded as `error` status and surfaced via ready()'s AggregateError.
- parseProjectIdFromApiKey (RC12 message-only change): single caller
channel-activation-config.ts:131; error type/behavior identical.
- ready() (f3 early-return): only affects a stopped manager (now resolves
instead of throwing) — strictly more lenient; sole public reference is a
doc example in endpoints/node.ts:42.
Skipped: a test exercising defaultActivateChannel's module-not-found friendly
error — the fn is unexported and the dynamic import specifier is not
injectable, and channels-intelligence IS installed in the workspace so a real
MODULE_NOT_FOUND can't be forced without contorting the code. isModuleNotFound
remains unit-tested.
Centralize the stop-vs-settle race class in ChannelManager behind one guarded,
idempotent teardown path instead of per-branch patches:
- ChannelEntry gains a private `handleStopped` flag; new private `stopEntry()`
sets status="stopped" and stops the handle AT MOST once. Both settle handlers
and stop() route through it.
- RC5: a rejection arriving AFTER stop() now keeps the entry "stopped" and
resolves settled (no error/setup_required, no rejectSettled), so a late
connect failure can't resurrect a stopped channel or reject a later ready().
- RC7: stop() runs `Promise.allSettled` over per-entry stopEntry() calls; the
handleStopped guard means a handle assigned in the same tick as stop() is
stopped exactly once even when both stop() and the success handler reach it.
- RC9 (fetch-handler): getOrCreateChannelManager now calls activate() BEFORE
inserting into the WeakMap, so a synchronous throw (duplicate/missing names)
caches nothing and every retry re-throws instead of returning an inert
manager that falsely reports "online".
- RC8: reconcile class + ready() docstrings — activation throws synchronously
(ChannelConfigError) only on up-front misconfiguration; all other failures
are recorded as channel status.
- assertUniqueChannelNames checks missing/empty name FIRST so two nameless
channels get the accurate "missing name" error, not a spurious "undefined"
duplicate.
- Remove the dead ChannelEntry.promise field (unread residue of the removed
reconnect path).
- RC4 (packaging): move @copilotkit/channels-intelligence from
optionalDependencies (auto-installed, force-pulls the pure-ESM package into
every OSS consumer) to an optional peerDependency, mirroring the other
optional integrations.
- Test nits: clear the dangling stop()-hang setTimeout; drop the redundant
not.toBe("reconnecting") assertion.
Call sites of changed symbols:
- stopEntry (new private): channel-manager.ts only — success handler, reject
handler, and stop(); no external callers.
- ChannelEntry.promise (removed): grep confirms no reads anywhere in the repo
(the only .promise reads are unrelated test signals).
- getOrCreateChannelManager (reordered, no signature change): single caller at
fetch-handler.ts createCopilotRuntimeHandler.
TDD: RC5 and RC9 red-green verified against prior code (RC5 reported "error"
not "stopped"; RC9 retry returned an inert healthy manager). RC7 pins the
single-stop guarantee for the new idempotent design.
RC1 — remove manager-level re-activation reconnect (delegate to Phoenix). The
ChannelManager's supervised reconnect re-invoked the activation engine on the
SAME already-started Channel, which throws in channel.addAdapter (started=true)
— so it could never succeed on the real launcher. It was also redundant:
Phoenix's Socket auto-reconnects and auto-rejoins, re-sending the join
declaration; the gateway's join/3 re-runs record_heartbeat (re-registers the
listener) and terminate/2 releases the dead socket's leases (verified against
Intelligence #511 sdk_channel.ex). Removed: runReconnect, reconnectLoops,
onChannelClosed, the RECONNECT_BASE_DELAY_MS / RECONNECT_MAX_DELAY_MS /
RECONNECT_MAX_ATTEMPTS constants, the injectable sleep arg + defaultSleep, and
stoppedSignal/resolveStopped. onClose is now a log-only breadcrumb (no state
mutation, no re-activation). Once a channel activates it stays online; a
transient drop is invisible (Phoenix self-heals). ChannelStatus keeps
"reconnecting" in the union marked reserved to avoid churning the public type;
computeOverall no longer assigns it.
RC2 — stop() no longer aborts teardown on a throwing handle.stop(). The real
launcher's stop() rethrows after session.disconnect(), so Promise.all rejected
and skipped the status loop; with stopped already set, a retry no-oped, leaving
the manager permanently un-torn-down. Switched to Promise.allSettled so every
handle attempts teardown and every entry is marked "stopped". Red-green test
added.
RC3 — parseProjectIdFromApiKey no longer echoes the full cpk-… secret in
ChannelConfigError (it is logged and surfaced via ready()'s AggregateError).
Message now includes only a short non-sensitive prefix; test asserts the format
hint is present but the full key is not.
Reconnect unit test rewritten to the new contract (a drop makes no further
engine call, does not throw, manager stays usable/coherent); obsolete
backoff-growth / give-up-to-error / cancel-pending-backoff cases removed.
Integration test step 5 updated: a drop stays online with no extra engine call.
Call sites cleared: grep over packages/runtime/src for runReconnect,
reconnectLoops, RECONNECT_BASE_DELAY_MS, RECONNECT_MAX_DELAY_MS,
RECONNECT_MAX_ATTEMPTS, onChannelClosed, stoppedSignal, resolveStopped,
defaultSleep, and the injectable ChannelManager sleep arg returns no matches;
fetch-handler exposes no sleep/reconnect seam. Nothing external referenced the
removed symbols.
A1: ChannelManager.activate() now asserts unique channel names before any
engine call (entries keyed by name silently leaked the first session on a
duplicate). Throws ChannelConfigError naming the dup. Reworded the stale
runtime.ts comment that claimed startChannels validates uniqueness — the
managed path activates one Channel per launcher call, so uniqueness is
enforced by ChannelManager.activate().
A4: stop() no longer awaits pending activations (a hung connect that
ready({timeoutMs}) tolerates would hang teardown/SIGTERM forever). It stops
only handles that already exist; a post-settle guard on the initial-activation
path tears down any handle arriving after stop(), mirroring the reconnect
loop's guard. Idempotent.
A3: reconnect success clears the reconnectLoops marker BEFORE re-arming
onClose, so a synchronous onClose re-fire on the fresh handle starts a new loop
instead of leaving the Channel stuck reconnecting with no loop.
B1/B3: doc fixes — onClose seam is present-tense (launcher delegates to
session.onClose); parseProjectIdFromApiKey @throws no longer describes an
unreachable empty-segment case.
Call sites reviewed (behavior holds at each):
- activate() throws on dup: fetch-handler.ts getOrCreateChannelManager (l.187),
reached from createCopilotRuntimeHandler at handler-creation time → now fails
loud at startup instead of leaking; channel-manager ready() (l.424) surfaces
the throw as a rejected promise.
- stop() prompt-resolve: examples/slack/app/managed.ts:174 SIGTERM shutdown
await listener.channels?.stop() — the exact hang this fixes. Endpoint
adapters (node/express/hono) only attach .channels; no direct stop callers.
Tests: channel-manager.test.ts + channel-manager-reconnect.test.ts 17 passed
(2 new + 1 new, red→green); channel-activation-config + fetch-handler green;
@copilotkit/runtime:check-types clean.