Commit Graph

455 Commits

Author SHA1 Message Date
MikeRyanDev 69861f13df chore: release monorepo v1.63.2 2026-07-23 16:15:51 +00:00
tylerslaton a7459f4fb2 chore: release monorepo v1.63.1 2026-07-16 18:23:58 +00:00
Martha Schumann 9e9ce128dd test(runtime): verify packed managed channels dependency 2026-07-16 10:44:10 -07:00
Martha Schumann fea464de52 fix(runtime): install managed channels activation dependency 2026-07-16 10:27:31 -07:00
tylerslaton 6c354037fc chore: release monorepo v1.63.0 2026-07-15 22:18:07 +00:00
Tyler Slaton 9504e0666c refactor(channels): add batteries-included package architecture (#5948)
## Summary

- Extract the platform-neutral Channels foundation into
`@copilotkit/channels-core`.
- Make `@copilotkit/channels` the batteries-included consumer entry
point, with adapter and UI subpaths.
- Release the umbrella, core, UI, and all six adapters as one shared
`channels` version scope.
- Verify the packed consumer contract and migrate the Slack and Teams
examples to the umbrella.

## Why

Consumers should be able to install one Channels package without making
the runtime or selective integrations depend on every platform SDK.
Shipping the complete Channels family together prevents adapter/core
version drift and makes the umbrella's exact dependency set release as a
compatible unit.

## How

- Move shared bot/runtime primitives into `channels-core` and invert
adapter/runtime dependencies.
- Add exact workspace dependencies and export subpaths from the umbrella
package.
- Consolidate the existing release configuration and all three release
workflow selectors into one shared `channels` scope containing all nine
packages.
- Use `@copilotkit/channels` as the version source: the next minor
release resolves to `0.2.0` and bumps every Channels package together.
- Publish scoped packages in dependency order for both stable and canary
releases: UI, core, adapters, then the umbrella.
- For a stable Channels release, publish the UI/core/adapters first, run
the registry-backed packed-consumer verifier against those newly
published exact versions, then publish the umbrella.
- Generate Channels release notes from `channels/v*` tags rather than
the monorepo `v*` tags.
- Verify builds, type checks, tests, package artifacts, examples, and
release workflow scope synchronization locally.

### First stable release sequence

1. Bootstrap the currently unpublished `@copilotkit/channels-core`
package on npm and configure npm trusted publishing for every Channels
package against this repository's `release / publish` workflow and `npm`
environment.
2. Create and merge the `channels` minor release PR. The stable workflow
publishes the Channels family in the staged order above and validates
the packed umbrella from the registry before the umbrella is released.
3. Create and merge a subsequent `monorepo` release PR so the published
`@copilotkit/runtime` switches from the historical umbrella dependency
to `@copilotkit/channels-core`.
2026-07-15 13:12:16 -07:00
Benjamin Taylor a8407f84fd fix(channels): address Tyler CR — wire app-api URL through runtime activation + map user name (OSS-476)
- [P1] History/files were absent on the NORMAL managed path: defaultActivateChannel
  never forwarded the app-api HTTP URL, so the transport (which installs
  fetchFile/getHistory/uploadFile only when appApiBaseUrl is set) ran without
  them for Channels started by the CopilotRuntime handler — only manual
  low-level launcher callers got file/history. Thread intelligence.ɵgetApiUrl()
  through ChannelActivationConfig.apiUrl → defaultActivateChannel →
  startChannelsOverRealtimeGateway({ appApiBaseUrl }). The launcher + transport
  already accepted it.
- [P2] Managed turns exposed the provider profile under a non-public `displayName`
  field, leaving PlatformUser.name undefined. Map env.user.displayName -> name
  (parity with the direct Slack adapter, which populates `name`).

Tests: deriver returns apiUrl; defaultActivateChannel forwards appApiBaseUrl to
the launcher opts; onMessage sees message.user.name. channel-activation-config +
channels-intelligence (170) green; runtime build type-checks. (channel-manager.test
executes in CI — local vitest hits the known optional-peer-dep resolution flake.)
2026-07-15 13:36:36 -05:00
Tyler Slaton 5811cab177 fix(channels): align rebase with channel API 2026-07-15 10:29:21 -07:00
Tyler Slaton 06fb2ed34a refactor(channels): extract the platform-neutral core 2026-07-15 10:11:55 -07:00
Benjamin Taylor a758faa19d fix(channels): CR follow-ups — honest stopped status + correct .env WS port (OSS-482)
Two fixes from the pre-merge adversarial CR of this PR (both pre-existing,
flagged as in-subject):

- ChannelManager.status() reported overall "online" for a manager stopped
  BEFORE activate() (e.g. SIGTERM during startup): `entries` is empty, so the
  empty-set fold returned "online" — a torn-down manager reading healthy. Now
  short-circuits to "stopped" when `this.stopped`, matching the documented
  status() contract. New red-green test covers the stop()-before-activate() case.

- examples/slack/.env.example: COPILOTKIT_INTELLIGENCE_WS_URL example was
  ws://localhost:4401, but derivation is a scheme-only swap of the :4201 API URL
  (→ ws://localhost:4201) and 4401 is used nowhere — a user uncommenting it hit a
  dead port. Corrected to :4201 and clarified the derivation note.
2026-07-15 10:50:20 -05:00
Benjamin Taylor ed37e418ca fix(runtime): bound ChannelManager.stop(), surface ready() AggregateError on hang, forward log to transport (OSS-482)
Three ChannelManager observability/robustness hardenings from the OSS-473 CR.

stop() per-handle timeout: stopEntry awaited handle.stop() unbounded, so a
wedged stop() (e.g. a socket.disconnect that never returns) hung teardown — and
thus SIGTERM shutdown — forever. Each handle.stop() is now bounded by
stopHandleTimeoutMs (default 5000); on timeout it is logged and abandoned so
every other entry still reaches `stopped`.

ready() surfaces the real reason on hang: ready({ timeoutMs }) wrapped the whole
`allSettled` in one timeout, so when one channel settled to `error` while a
sibling hung, it rejected with only a generic timeout and DISCARDED the erroring
channel's reason. The deadline is now applied PER CHANNEL, so the AggregateError
carries both each failed channel's real reason AND a named timeout for each
still-hanging channel.

Log forwarding: the manager's `log` reached activation-level events only; the
default engine (defaultActivateChannel → startChannelsOverRealtimeGateway) never
passed it down, so transport-level drop diagnostics (e.g. a version-skew
missing-leaseToken outage) were silent in the managed path. `log` is now
forwarded down to the launcher/transport.

Note: the reconnect "gave-up → error" escalation from the same CR is already
implemented end-to-end in OSS-473 (realtime-gateway.ts reconnectGiveUpMs +
ChannelManager.registerConnectionObserver) and is not re-done here.

Tests: stop() resolves + logs a timeout when handle.stop() never settles;
ready() aggregate contains both a real activation error and the hung sibling's
named timeout; defaultActivateChannel forwards its log sink to the launcher opts.
2026-07-15 07:44:35 -05:00
Benjamin Taylor 2a72ca15e3 docs(runtime): describe channel skip as exclusive per Channel, not per platform (OSS-473) 2026-07-15 07:15:06 -05:00
Benjamin Taylor fbf35ac593 fix(runtime): defer channel activation to ready() so the Fetch handler is serverless-safe (OSS-473) 2026-07-15 07:15:06 -05:00
Benjamin Taylor 5042266824 fix(runtime): declare managed provider per-Channel instead of a hard-coded slack default (OSS-473) 2026-07-15 07:15:06 -05:00
Benjamin Taylor 6fb748c368 fix(runtime): report skipped direct-adapter channels as unmanaged, not healthy (OSS-473) 2026-07-15 07:15:05 -05:00
Benjamin Taylor 68349bc1f8 fix(runtime): non-optional handler.channels for intelligence runtimes with declared channels (OSS-473)
The PR's documented snippet — `await handler.channels.ready(...)` with no
`!`/`?.` — did not type-check under strict TS because
`createCopilotRuntimeHandler` always returned `channels?: ChannelsControl`.

Encode channel-presence at the type level:

- runtime.ts: `CopilotRuntime` is now a `const` typed as `CopilotRuntimeConstructor`
  (backed by an internal `CopilotRuntimeShim` class; behavior unchanged). A
  class constructor cannot vary its return type across overloads, so the two
  construct-signature overloads live on the constructor interface: `intelligence`
  + a non-empty `channels` tuple returns a `RuntimeWithDeclaredChannels`-branded
  runtime; every other config (SSE, intelligence-without-channels, empty
  `channels: []`, or a non-literal `Channel[]` variable) stays unbranded. The
  brand is a phantom (compile-time-only) property. `export interface CopilotRuntime`
  preserves the name as a type for existing `runtime: CopilotRuntime` / `as
  CopilotRuntime` sites.
- fetch-handler.ts: overload `createCopilotRuntimeHandler` — a branded runtime
  (unless `activateChannels: false`, constrained to `true | undefined`) returns
  the new `CopilotRuntimeFetchHandlerWithChannels` (non-optional `channels`);
  everything else keeps the optional shape. Opting out of activation honestly
  falls through to the optional overload.
- Added a compile-time type test (checked by `tsc --noEmit`, the `check-types`
  gate). It probes the optionality modifier structurally (`{} extends Pick<T,K>`)
  rather than for `undefined`, since this package compiles `strict: false`.
  Confirmed it fails pre-change on the required-channels assertion and passes
  after. Dropped the now-unnecessary `!` in handler-channels.test.ts.

Call sites: the second overload is byte-identical to the former single signature,
so every `createCopilotRuntimeHandler` caller (node/express/hono endpoints,
integration servers, examples) and every `new CopilotRuntime` site resolves
unchanged; only inline non-empty-`channels` construction gains the (strict
supertype-assignable) branded type. Verified via a clean full-package check-types.
2026-07-15 07:15:05 -05:00
Benjamin Taylor 08d14dcc8c fix(runtime): reachable setup_required + connection-health status from the real gateway (OSS-473)
P1#2 — reachable setup_required on the PRODUCTION engine path.
connectRealtimeGateway no longer flattens every join rejection into a
generic Error. A join `.receive("error", reason)` whose reason is a known
setup-required code (`channel_declaration_unavailable`, and defensively
`adapter_setup_required` / `not_configured`) now rejects with a
distinguishable `RealtimeGatewaySetupRequiredError` (`code === "SETUP_REQUIRED"`,
raw reason preserved). ChannelManager already detects that code, so an
unconfigured managed provider now degrades to `setup_required` (ready()
resolves) instead of `error`. All other reasons keep the generic error and
the socket-leak teardown is unchanged.

P1#3 — status() reflects real connection health instead of `online` forever.
ConnectedRealtimeGatewaySession exposes `onStateChange(cb)` over
`RealtimeGatewayConnectionState` (`online` | `reconnecting` | `gave_up`),
driven by the real Phoenix seams: an unexpected socket drop → `reconnecting`;
a successful (re)join (the join-push recHooks survive Phoenix `resend`, so
`"ok"` re-fires on every auto-rejoin) → `online`; and a BOUNDED give-up —
Phoenix retries forever, so a `reconnectGiveUpMs` window (default 60000, runs
from the first drop of an episode, cleared on rejoin) elapsing while still
reconnecting → `gave_up` (terminal). Our own disconnect() stays silent.
ChannelManager wires this in place of the log-only onClose breadcrumb:
`reconnecting`→status reconnecting, `online`→online, `gave_up`→error; a
stopped manager/entry ignores late events. computeOverall now ranks
`error > reconnecting > setup_required > connecting > online`. ready() keeps
its one-shot semantics (settles on the initial outcome); later health
transitions move only status(). Docs updated to state `online` means
currently-sendable.

Wording: the direct-adapter skip comment/log now states delivery is
exclusive-per-platform (managed OR direct per platform, not both — attaching
both would double-deliver) with true coexistence tracked in OSS-484. Skip
behavior unchanged.

Call-sites for the changed signatures:
- connectRealtimeGateway error shape: only caller is
  startChannelsOverRealtimeGateway (realtime-gateway-launcher.ts:216); it
  awaits and lets the rejection propagate, so the setup-required error flows
  through unchanged (no branch to update).
- new ConnectedRealtimeGatewaySession.onStateChange: passed through in
  startChannelsWithGatewaySession and startChannelsOverRealtimeGateway
  (realtime-gateway-launcher.ts); added to ChannelsHandle (runtime.ts) and the
  manager's local ChannelsHandle view (channel-manager.ts); exported from
  index.ts. RealtimeGatewaySession (base, no observer) consumers
  (realtime-gateway-transport.ts) unaffected.
- manager onClose→state transitions: registerOnClose renamed to
  registerConnectionObserver; sole caller is the online settle handler in
  activate().
2026-07-15 07:15:05 -05:00
Benjamin Taylor 6ab2d57ff3 fix(runtime): skip managed activation for channels carrying a direct adapter (OSS-473)
Honors the SoT "never infer managed intent from a direct adapter" rule.
Channel.adapters is a new additive read-only member; verified no consumer
constructs Channel literals (only createChannel does).
2026-07-15 07:15:05 -05:00
Benjamin Taylor 04a1b10e26 fix(runtime): safe-integer projectId, trim adapter, named error classes, RC7 guard test (OSS-473) 2026-07-15 07:15:04 -05:00
Benjamin Taylor 8ae942efce fix(runtime): remove drift-prone channel-name replica, log activation errors under err, cover default engine (OSS-473)
RC15: getOrCreateChannelManager bridged the manager log as
`logger.warn({ meta }, msg)`, but pino only serializes an Error's
(non-enumerable) message/stack under the `err` key — under `meta` a
failed activation rendered as `{}`, losing the cause. Route an Error to
`err` and keep `meta` for everything else.

LEVER: removed the channel-name FORMAT/length validation block plus the
replicated CHANNEL_NAME_PATTERN / MIN / MAX constants from
channel-activation-config.ts. This was a third copy of managed-specific
rules whose source of truth is channels-intelligence's
assertValidChannelRealtimeScope + assertValidChannelNames, and it kept
drifting (omitted the reserved-name rule). Now that activation failures
are logged, recorded as `error` status, and surfaced via ready(), the
up-front check is not worth cross-package rule parity. Missing/empty
name still throws (the config's own precondition). Deleted the obsolete
"Slack"/"support_bot"/"cs"/65-char rejection tests.

projectId>0: parseProjectIdFromApiKey now throws ChannelConfigError when
the parsed id is <= 0 (`cpk-0_...` matched but failed deep in the
launcher). Parser validating its own output, not a channel-name replica,
so it stays here; reuses the existing key redaction.

adapter default hardening: `adapter ?? "slack"` -> truthiness/trim check
so ""/whitespace falls back to "slack".

activate() stopped-guard: short-circuits on `this.activated || this.stopped`
so a post-stop() activate() opens no transports on a dead manager.

coverage: exported defaultActivateChannel (the real engine) with an
injectable importer seam (optional param, default = the same non-literal
dynamic import) and covered its 3 branches — config->opts mapping (scope
carries only projectId+channelName), module-not-found friendly error, and
generic-error passthrough. Added cheap manager coverage: lazy-activate
duplicate-name reject via ready(), empty channels[] -> online + ready
resolves, non-default adapter reaches the engine config.

Call sites (no external breakage):
- CHANNEL_NAME_PATTERN/MIN/MAX: were module-private; zero references.
- parseProjectIdFromApiKey: only internal caller is
  deriveChannelActivationConfig (passes the real key) + tests; <=0 throw
  affects only malformed cpk-0 keys.
- defaultActivateChannel: only internal use is the ChannelManager
  constructor default (called 2-arg) + tests; new 3rd param is optional
  and the fn still satisfies ActivateChannelEngine.
- ChannelsIntelligenceModule: new additive export, referenced internally
  + tests only.
- activate() guard / logger bridge: internal only.
2026-07-15 07:15:04 -05:00
Benjamin Taylor 040b1a848f fix(runtime): wire channel activation logging, redact key fully, validate channel-name format up-front (OSS-473)
CR batch for packages/runtime managed-channels activation:

- RC11: getOrCreateChannelManager now passes a `log` adapter bridging the
  ChannelManager diagnostic sink to the shared logger
  (`log: (msg, meta) => logger.warn({ meta }, msg)`). Previously every
  breadcrumb (setup_required, failed-to-activate, dropped-session,
  teardown-stop failure) was a no-op, so a channel that failed to activate
  was permanently dead with zero output.
- f2: stopEntry logs a swallowed handle.stop() error via the sink instead of
  discarding it; teardown stays resilient (never rethrows).
- sync-throw guard: stopEntry wraps handle.stop() in
  `Promise.resolve().then(...)` so a foreign/injected handle that throws
  SYNCHRONOUSLY is caught by the same `.catch` — otherwise the throw escaped,
  skipped resolveSettled(), and hung `settled` forever.
- f3: ready() short-circuits and RESOLVES when the manager is stopped. A
  channel that settled to `error` before stop() had already rejected its
  `settled` promise, so a later ready() threw an AggregateError even though
  status().overall was "stopped" — now consistent with the after-stop case.
- RC12: parseProjectIdFromApiKey no longer slices a fixed 8 chars off an
  arbitrary key (which echoed secret bytes for a `cpk-_...`-shaped key). The
  failure message now echoes NONE of the key value, only the expected
  `cpk-{projectId}_` format hint.
- RC13: deriveChannelActivationConfig enforces the lowercase-kebab-case
  channel-name rule (/^[a-z][a-z0-9]*(?:-[a-z0-9]+)*$/, length 3-64) up front
  and throws a clear ChannelConfigError, instead of passing a bad name to the
  launcher where assertValidChannelRealtimeScope throws deep and the channel
  is silently degraded to `error`. Regex/bounds are a literal copy of
  channels-intelligence's assertValidChannelRealtimeScope (the source of
  truth; not statically imported — it's an optional pure-ESM peer).
- ready() docblock reworded: async ready() REJECTS (not throws) the
  ChannelConfigError.

Tests (red-green verified): RC11 (handler logger spy + manager log sink),
f3 (ready resolves post-stop after pre-stop error), sync-throw guard, RC12
(no secret-tail leak), RC13 (Slack/support_bot/cs/65-char reject; support
passes).

Call-site enumeration for changed signatures/behaviors:
- getOrCreateChannelManager (new internal `log` arg): single caller
  fetch-handler.ts:402; public signature unchanged, no caller impact.
- deriveChannelActivationConfig (RC13 now throws on bad name): single caller
  channel-manager.ts:323 inside activate()'s per-channel loop, already
  wrapped in try/catch that converts a throw to a rejected activation ->
  recorded as `error` status and surfaced via ready()'s AggregateError.
- parseProjectIdFromApiKey (RC12 message-only change): single caller
  channel-activation-config.ts:131; error type/behavior identical.
- ready() (f3 early-return): only affects a stopped manager (now resolves
  instead of throwing) — strictly more lenient; sole public reference is a
  doc example in endpoints/node.ts:42.

Skipped: a test exercising defaultActivateChannel's module-not-found friendly
error — the fn is unexported and the dynamic import specifier is not
injectable, and channels-intelligence IS installed in the workspace so a real
MODULE_NOT_FOUND can't be forced without contorting the code. isModuleNotFound
remains unit-tested.
2026-07-15 07:15:04 -05:00
Benjamin Taylor 98598bdd69 fix(runtime): idempotent channel teardown + activate-before-cache; optional peer dep for channels-intelligence (OSS-473)
Centralize the stop-vs-settle race class in ChannelManager behind one guarded,
idempotent teardown path instead of per-branch patches:

- ChannelEntry gains a private `handleStopped` flag; new private `stopEntry()`
  sets status="stopped" and stops the handle AT MOST once. Both settle handlers
  and stop() route through it.
- RC5: a rejection arriving AFTER stop() now keeps the entry "stopped" and
  resolves settled (no error/setup_required, no rejectSettled), so a late
  connect failure can't resurrect a stopped channel or reject a later ready().
- RC7: stop() runs `Promise.allSettled` over per-entry stopEntry() calls; the
  handleStopped guard means a handle assigned in the same tick as stop() is
  stopped exactly once even when both stop() and the success handler reach it.
- RC9 (fetch-handler): getOrCreateChannelManager now calls activate() BEFORE
  inserting into the WeakMap, so a synchronous throw (duplicate/missing names)
  caches nothing and every retry re-throws instead of returning an inert
  manager that falsely reports "online".
- RC8: reconcile class + ready() docstrings — activation throws synchronously
  (ChannelConfigError) only on up-front misconfiguration; all other failures
  are recorded as channel status.
- assertUniqueChannelNames checks missing/empty name FIRST so two nameless
  channels get the accurate "missing name" error, not a spurious "undefined"
  duplicate.
- Remove the dead ChannelEntry.promise field (unread residue of the removed
  reconnect path).
- RC4 (packaging): move @copilotkit/channels-intelligence from
  optionalDependencies (auto-installed, force-pulls the pure-ESM package into
  every OSS consumer) to an optional peerDependency, mirroring the other
  optional integrations.
- Test nits: clear the dangling stop()-hang setTimeout; drop the redundant
  not.toBe("reconnecting") assertion.

Call sites of changed symbols:
- stopEntry (new private): channel-manager.ts only — success handler, reject
  handler, and stop(); no external callers.
- ChannelEntry.promise (removed): grep confirms no reads anywhere in the repo
  (the only .promise reads are unrelated test signals).
- getOrCreateChannelManager (reordered, no signature change): single caller at
  fetch-handler.ts createCopilotRuntimeHandler.

TDD: RC5 and RC9 red-green verified against prior code (RC5 reported "error"
not "stopped"; RC9 retry returned an inert healthy manager). RC7 pins the
single-stop guarantee for the new idempotent design.
2026-07-15 07:15:03 -05:00
Benjamin Taylor 3c718faa08 fix(runtime): delegate channel reconnect to Phoenix; resilient stop; redact key in error (OSS-473)
RC1 — remove manager-level re-activation reconnect (delegate to Phoenix). The
ChannelManager's supervised reconnect re-invoked the activation engine on the
SAME already-started Channel, which throws in channel.addAdapter (started=true)
— so it could never succeed on the real launcher. It was also redundant:
Phoenix's Socket auto-reconnects and auto-rejoins, re-sending the join
declaration; the gateway's join/3 re-runs record_heartbeat (re-registers the
listener) and terminate/2 releases the dead socket's leases (verified against
Intelligence #511 sdk_channel.ex). Removed: runReconnect, reconnectLoops,
onChannelClosed, the RECONNECT_BASE_DELAY_MS / RECONNECT_MAX_DELAY_MS /
RECONNECT_MAX_ATTEMPTS constants, the injectable sleep arg + defaultSleep, and
stoppedSignal/resolveStopped. onClose is now a log-only breadcrumb (no state
mutation, no re-activation). Once a channel activates it stays online; a
transient drop is invisible (Phoenix self-heals). ChannelStatus keeps
"reconnecting" in the union marked reserved to avoid churning the public type;
computeOverall no longer assigns it.

RC2 — stop() no longer aborts teardown on a throwing handle.stop(). The real
launcher's stop() rethrows after session.disconnect(), so Promise.all rejected
and skipped the status loop; with stopped already set, a retry no-oped, leaving
the manager permanently un-torn-down. Switched to Promise.allSettled so every
handle attempts teardown and every entry is marked "stopped". Red-green test
added.

RC3 — parseProjectIdFromApiKey no longer echoes the full cpk-… secret in
ChannelConfigError (it is logged and surfaced via ready()'s AggregateError).
Message now includes only a short non-sensitive prefix; test asserts the format
hint is present but the full key is not.

Reconnect unit test rewritten to the new contract (a drop makes no further
engine call, does not throw, manager stays usable/coherent); obsolete
backoff-growth / give-up-to-error / cancel-pending-backoff cases removed.
Integration test step 5 updated: a drop stays online with no extra engine call.

Call sites cleared: grep over packages/runtime/src for runReconnect,
reconnectLoops, RECONNECT_BASE_DELAY_MS, RECONNECT_MAX_DELAY_MS,
RECONNECT_MAX_ATTEMPTS, onChannelClosed, stoppedSignal, resolveStopped,
defaultSleep, and the injectable ChannelManager sleep arg returns no matches;
fetch-handler exposes no sleep/reconnect seam. Nothing external referenced the
removed symbols.
2026-07-15 07:15:03 -05:00
Benjamin Taylor b0cb79f83e fix(runtime): fail-loud on duplicate channel names, non-hanging stop, reconnect ordering (OSS-473)
A1: ChannelManager.activate() now asserts unique channel names before any
engine call (entries keyed by name silently leaked the first session on a
duplicate). Throws ChannelConfigError naming the dup. Reworded the stale
runtime.ts comment that claimed startChannels validates uniqueness — the
managed path activates one Channel per launcher call, so uniqueness is
enforced by ChannelManager.activate().

A4: stop() no longer awaits pending activations (a hung connect that
ready({timeoutMs}) tolerates would hang teardown/SIGTERM forever). It stops
only handles that already exist; a post-settle guard on the initial-activation
path tears down any handle arriving after stop(), mirroring the reconnect
loop's guard. Idempotent.

A3: reconnect success clears the reconnectLoops marker BEFORE re-arming
onClose, so a synchronous onClose re-fire on the fresh handle starts a new loop
instead of leaving the Channel stuck reconnecting with no loop.

B1/B3: doc fixes — onClose seam is present-tense (launcher delegates to
session.onClose); parseProjectIdFromApiKey @throws no longer describes an
unreachable empty-segment case.

Call sites reviewed (behavior holds at each):
- activate() throws on dup: fetch-handler.ts getOrCreateChannelManager (l.187),
  reached from createCopilotRuntimeHandler at handler-creation time → now fails
  loud at startup instead of leaking; channel-manager ready() (l.424) surfaces
  the throw as a rejected promise.
- stop() prompt-resolve: examples/slack/app/managed.ts:174 SIGTERM shutdown
  await listener.channels?.stop() — the exact hang this fixes. Endpoint
  adapters (node/express/hono) only attach .channels; no direct stop callers.

Tests: channel-manager.test.ts + channel-manager-reconnect.test.ts 17 passed
(2 new + 1 new, red→green); channel-activation-config + fetch-handler green;
@copilotkit/runtime:check-types clean.
2026-07-15 07:15:03 -05:00
Benjamin Taylor 7015837530 fix(runtime): declare channels-intelligence optional dep + export channels control types (OSS-473) 2026-07-15 07:15:02 -05:00
Benjamin Taylor 0e9f90cdba test(runtime): lock public handler.channels contract (OSS-473) 2026-07-15 07:15:02 -05:00
Benjamin Taylor 2e33d0cbbb test(runtime): integration proof of handler-owned channel activation (OSS-473) 2026-07-15 07:15:02 -05:00
Benjamin Taylor ea66c9328a feat(runtime): expose handler.channels through node/express/hono endpoints (OSS-473) 2026-07-15 07:15:01 -05:00
Benjamin Taylor b3081637c1 feat(runtime): handler owns managed channel activation and exposes handler.channels (OSS-473) 2026-07-15 07:15:01 -05:00
Benjamin Taylor a41da177de feat(runtime): supervised reconnect with bounded backoff for channels (OSS-473) 2026-07-15 07:15:01 -05:00
Benjamin Taylor 8f5185953f feat(runtime): add ChannelManager activation/readiness/stop lifecycle (OSS-473) 2026-07-15 07:15:01 -05:00
Benjamin Taylor b4f17e84b3 feat(runtime): derive channel activation config from intelligence config (OSS-473) 2026-07-15 07:15:00 -05:00
Benjamin Taylor 49c58bf081 feat(runtime): rename bots option to channels on the Intelligence runtime (OSS-473) 2026-07-15 07:15:00 -05:00
Benjamin Taylor 229f3f1460 refactor(channels): rename Bot/createBot public API to Channel/createChannel across the workspace (OSS-473) 2026-07-15 07:14:59 -05:00
Ben Taylor 855446e1ab fix(runtime): stop finalizeRunEvents emitting events after a terminal (#5812) (#5885)
## Summary

Pressing **Stop** while an assistant message is streaming
(CopilotRuntime + `HttpAgent` proxy) crashed the chat with:

```
Cannot send event type 'TEXT_MESSAGE_END': The run has already errored with 'RUN_ERROR'. No further events can be sent.
```

Root cause: `finalizeRunEvents` appended a trailing `TEXT_MESSAGE_END`
**after** the `RUN_ERROR` that the aborted agent had already emitted.

Fixes #5812.

## Root cause

When the upstream agent (e.g. pydantic-ai's `AGUIAdapter`) is aborted
mid-stream it emits a live `RUN_ERROR` while a text message is still
open — it does **not** close the message first. All runners
(`in-memory`, `intelligence`, `sqlite`) stream `finalizeRunEvents`'
output *after* everything the agent already emitted, so the appended
closer landed past the terminal:

| | outgoing event order |
|---|---|
| **Before** | `… TEXT_MESSAGE_CONTENT → RUN_ERROR → TEXT_MESSAGE_END` ❌
verifier throws |
| **After** | `… TEXT_MESSAGE_CONTENT → RUN_ERROR` ✅ terminal closes the
message client-side |

Per the AG-UI invariant: at most one terminal event per run, and no
sub-events after it. I confirmed against the real `@ag-ui/client`
`verifyEvents` (the verifier the browser runs) that a terminal arriving
with a message still open is valid — the terminal implicitly closes it.

## Fix

`finalizeRunEvents` (in `@copilotkit/shared`) now returns early and
appends **nothing** when the stream already contains a terminal event
(`RUN_FINISHED` or `RUN_ERROR`). The abrupt-end path (no terminal →
close open streams + synthesize a terminal, in the correct order) is
unchanged. No API/signature change; the in-memory, intelligence, and
sqlite runners all inherit the fix.

## Testing

RED→GREEN verified — each new/updated assertion was confirmed to fail
against the pre-fix code:

- **`finalize-events.test.ts`** — terminal-present appends nothing
(parametrized over `RUN_FINISHED` and `RUN_ERROR`) + a named #5812 case.
- **`in-memory-runner.test.ts`** — end-to-end mid-stream-stop
regression: a fake `HttpAgent`-style agent is stopped between
`TEXT_MESSAGE_START` and `TEXT_MESSAGE_END`; asserts no events follow
`RUN_ERROR` **and** that the collected stream passes `verifyEvents`
(before the fix this threw the exact browser error).
- **`intelligence-runner.test.ts`** — corrected a pre-existing assertion
that had encoded the buggy post-terminal `TEXT_MESSAGE_END`.

Green: full `@copilotkit/runtime` suite, `@copilotkit/sqlite-runner`,
`@copilotkit/shared`, `check-types`, `oxlint` (0 errors), and build.

## Reviewer notes

- The behavior change is a single early-return in `finalize-events.ts`;
the `terminalEventMissing` guards simplify away because they're only
reachable when no terminal exists.
- Diff is +204/−55 across 4 files, the bulk of it tests.
2026-07-14 09:03:15 -05:00
Tyler Slaton 001dda539a chore(channels-intelligence): complete channel terminology sweep 2026-07-10 14:47:06 -07:00
Tyler Slaton 68e43fefe1 refactor(channels-intelligence): rename remaining channel APIs 2026-07-10 14:39:56 -07:00
Tyler Slaton 0e787632a3 fix(runtime): stop finalizeRunEvents emitting events after a terminal (#5812)
Pressing Stop mid-stream against a CopilotRuntime + HttpAgent proxy aborts
the upstream agent, which emits a live RUN_ERROR while a text message is
still open. finalizeRunEvents then appended a trailing TEXT_MESSAGE_END
*after* that RUN_ERROR. Because the runners stream finalization events
after everything the agent already emitted, the closer landed past the
terminal and the AG-UI verifier rejected it with "the run has already
errored with 'RUN_ERROR'. No further events can be sent." — crashing the
chat.

finalizeRunEvents (in @copilotkit/shared, consumed by the in-memory,
intelligence, and sqlite runners) now returns early and appends nothing
when the stream already contains a terminal event (RUN_FINISHED or
RUN_ERROR): any message or tool call still open is closed by the terminal
on the client. The abrupt-end path (no terminal -> close open streams +
synthesize a terminal) is unchanged.

Tests:
- finalize-events.test.ts: terminal-present appends nothing (both
  RUN_FINISHED and RUN_ERROR) + a named #5812 case.
- in-memory-runner.test.ts: end-to-end mid-stream-stop regression that
  asserts no events follow RUN_ERROR and the stream passes AG-UI
  verifyEvents (the verifier the browser runs).
- intelligence-runner.test.ts: corrected an assertion that had encoded
  the buggy post-terminal TEXT_MESSAGE_END.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-08 14:47:29 -07:00
Benjamin Taylor b394f06fdc refactor(channels): rename @copilotkit/bot* packages to @copilotkit/channels* (OSS-438)
Renames the Bots SDK to the Channels SDK. Names only — no behavior change.

- 8 packages @copilotkit/bot* -> @copilotkit/channels* (git mv dirs, names,
  workspace: cross-deps). Now includes @copilotkit/bot-intelligence ->
  @copilotkit/channels-intelligence (landed on main via #5761; unpublished, so
  renamed fresh with the family).
- release.config.json scope keys + versionSource; ReleaseScope union;
  canary/stable-release/publish-release scope dropdowns; verify script
- examples/slack (Kite) + examples/teams: deps, jsxImportSource, imports
- showcase/shell-docs: content dirs docs/bots->docs/channels and
  reference/bot->reference/channels, nav registry, redirects

createBot and other API names unchanged. Old @copilotkit/bot* to be deprecated
after the new packages publish (bot-intelligence was never published).

Re-derived onto latest main (was conflicting after #5761 landed).

Refs OSS-438
2026-07-08 13:27:35 -05:00
Ben Taylor 9fa925bd95 feat(bot,runtime): managed bots SDK — run the bot SDK from Intelligence-delivered events (OSS-360/361) (#5761)
## Summary

Lets the `@copilotkit/bot` SDK run from **Intelligence-delivered
events** without a second programming model, and adds the runtime `bots`
declaration API. A managed event (delivered by Intelligence) runs the
*same* customer handlers, tools, context, commands, Bot UI, and agents
as local/custom adapters — the managed path is "just another
`PlatformAdapter`," fed by injected transports.

This is the **OSS / SDK slice** of the Hosted Managed Bots work. The
credentialed transports (Realtime Gateway, Connector Outbox) and the
frozen shared contracts live elsewhere (see *Out of scope*); this PR
ships the seams they plug into, fully runnable headless.

Relates to **OSS-360** (runtime bots API), **OSS-361** (run the SDK from
Intelligence events), **OSS-363** (Slack render/codec reuse).

## What's in here

- **`intelligenceAdapter()` bridge** (`@internal`, not publicly
documented) — implements `PlatformAdapter` over two injected transports:
`DeliverySource` (inbound) + `EgressSink` (outbound). Ingress →
`onTurn`/`onCommand`/`onInteraction`/`onThreadStarted`/`onReaction`; ack
on success / nack on throw (at-least-once). Egress emits generic
operations carrying `BotNode[]` IR with **deterministic ids**
(`turnId:seq`, reset per turn) so a redelivered turn reproduces the same
ids for the Connector Outbox to dedupe. Idempotency lives at egress, so
the managed path skips ingress dedup (`skipIngressDedup`) — a redelivery
re-runs rather than being dropped.
- **Runtime `bots` API** — `new CopilotRuntime({ intelligence, bots })`,
accepted by TypeScript **only when `intelligence` is configured**
(discriminated union). `createBot({ name })`; `startManagedBots()`
validates names (required, identifier-style, unique — fail-loud), builds
activation metadata, and wires each bot to its resolved transport.
- **`PlatformCodec` seam** + Slack egress codec (`slackCodec`) composing
the existing pure `renderSlackMessage`, so IR→native rendering is shared
(no Bolt/creds) instead of duplicated.
- **Backwards-compatible SDK foundations**: `bot.addAdapter()` +
optional `adapters`, deferred backend resolution at `start()` with
`stateStore`-provider precedence (+ multi-provider warning),
`bot.transcripts` throws pre-start, optional
`eventId`/`turnId`/`deliveryId` on ingress + handler context. Existing
`createBot` callers and every `PlatformAdapter` implementer are
unaffected.
- **In-memory transports + fixture tests** — the full dispatch path
(envelope in → handler runs → egress op out) runs with zero
Slack/Intelligence/network.

## Out of scope (external / separate tickets)

- **Realtime Gateway + Connector Outbox transports** — implemented in
the closed-source repo against the `DeliverySource`/`EgressSink`
interfaces shipped here.
- **Shared contracts freeze (OSS-377)** — consumed here via a minimal,
isolated placeholder (`managed/contracts.ts`, marked `TODO(OSS-377)`);
swaps in via one import change.
- **OSS-363 ingress normalization** — the egress codec is done;
extracting the pure Slack event→neutral mapping out of the Bolt listener
(so local + Intelligence ingress share it) is the remaining, higher-risk
half and is left to that ticket (`TODO(OSS-363)`).

## Testing

TDD throughout (RED→GREEN per behavior). New: managed adapter
dispatch/ack-nack/ids/run-renderer/exclusivity, all-kinds routing, name
validation + metadata + lifecycle, runtime `bots` option, Slack codec.
Full suites green: `bot` 147, `bot-slack` 256, `runtime` 1574. All
builds typecheck (`bot`/`bot-slack`/`bot-discord`/`runtime`);
oxlint/oxfmt clean.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-08 12:49:17 -05:00
Alem Tuzlak dff80a2a9a fix(bot,bot-slack,bot-intelligence,runtime): address managed-bots SDK review findings
Correctness:
- C1 create-bot: start() is now idempotent — a second start() no longer
  re-resolves the backend / rebuilds Transcripts+Telemetry+ActionRegistry or
  re-connects adapters (which would wipe MemoryStore state and double-bind real
  adapters). stop() clears the flag so start→stop→start is still a real restart.
- S1 bot-slack ingress: a threaded reply that @-mentions the bot is now skipped
  (app_mention handles it) so the managed path no longer double-responds. Matches
  both the plain <@U…> and labeled <@U…|handle> mention forms.
- S2 runtime: CopilotSseRuntime throws if `bots` is passed without intelligence
  instead of silently dropping them (guards a JS/as-any caller past the type).
- S3 bot-intelligence: startManagedBots rolls back — stops already-started bots —
  when a later bot fails to start, instead of leaking listeners/connections.

Lower:
- S4 ingress: stripMentions handles the labeled <@U…|handle> form; DM turns strip
  mentions too (parity with app_mention/thread_reply).
- S5 bot-intelligence: bot-name uniqueness is now case-insensitive.
- S6 runtime: fail fast at construction when a declared bot has no name (full
  shape/uniqueness validation stays at the activation seam — assertValidBotNames —
  because it can't cross into this CJS package from pure-ESM bot-intelligence).
- S7 bot-intelligence: buildActivationMetadata throws on a nameless bot instead of
  silently filtering it out of the activation set.
- S8 bot-intelligence: startManagedBots warns on an empty bots array.
- M1 intelligence-adapter: the per-turn egress seq Map entry is deleted after each
  turn so it can't grow unbounded over a long-running bot.
- M2 intelligence-adapter: an inbound file that fails to fetch degrades to a
  fail-visible text note instead of being silently dropped from model context.
- I2 contracts: dropped the now-dead `duplicate_skipped` RenderAccepted value
  (Intelligence returns duplicate_accepted or a 409 conflict).

Changelog (C2/C3, intended behavior after moving init into start()):
- bot.transcripts now throws before start() (was a concrete property).
- telemetry `oss.bot.configured` now fires at start() rather than construction, so
  a constructed-but-never-started bot no longer emits it.

Not addressed here (cross-repo, tracked on the Intelligence side):
- I1 realtime render-event kind:"file" clause on the gateway validator.
- I3 lease-token fencing on the render-accept path.
2026-07-08 19:18:19 +02:00
tylerslaton 4394f9c81d chore: release monorepo v1.62.3 2026-07-08 16:17:36 +00:00
Alem Tuzlak 2330dca267 fix(bot-slack,runtime): align tests with renderer status + AbstractAgent.run
Two pre-existing test failures on this branch, surfaced by CI's unit +
check-types jobs once main was merged:

- bot-slack event-renderer: the non-pane thread tool-call test still
  asserted the old "no composer status" behavior. Commit 13248dda0b
  deliberately drove setStatus on ANY thread anchor (not just panes), so
  the test now expects both the 🔧 row and the "is using…" status.
- runtime in-memory-runner: HangingAgent/AbortableAgent extended
  AbstractAgent but omitted the abstract run() member (@ag-ui/client
  0.0.57), failing tsc on the test tsconfig (TS2515). Add the same
  run() => EMPTY stub the sibling test agents use.
2026-07-08 18:15:29 +02:00
Benjamin Taylor 71a4ac42e4 Merge origin/main into alem/oss-360-sdk-foundations
Brings the 499-commit-stale foundations branch up to date with main so #5761
has a clean diff and no stale reverts (e.g. forwardHeaders). Conflicts:
- CopilotThreadsDrawer.tsx: took main's (main renamed CopilotDrawer -> ThreadsDrawer
  + added the collapse feature; the branch's edit was a no-op import-type split).
- pnpm-lock.yaml: regenerated with the pinned pnpm 10.33.4 (adds @copilotkit/bot-intelligence).
2026-07-08 11:01:58 -05:00
Benjamin Taylor 5994bfe482 Merge origin/main into ben1/ent-1018-stateless-suggestions
Syncs the branch with main (304 commits) to resolve CI type-check failure.
main changed extractForwardableHeaders to require a forwarding policy and
added the mergeForwardableHeaders helper (#5712); handle-suggest now uses
mergeForwardableHeaders(agent.headers, request, runtime.forwardHeadersPolicy ??
resolveForwardHeadersPolicy(undefined)) to match the run handler — fixing the
drift and adopting the server-headers-win / infra-header denylist behavior.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 17:40:50 -05:00
Benjamin Taylor 9f938e89cf refactor(suggestions): stream stateless /suggest over SSE and forward consumer state
Rework the stateless /suggest transport to reuse the AG-UI SSE pipeline
instead of a buffered JSON response, resolving the streaming + state review
feedback:

- server runs the provider agent directly and streams its events via
  createSseEventResponse (the runner's event pipeline minus GLOBAL_STORE
  persistence), gated with captureTelemetry:false so suggestions stay out of
  run telemetry
- client drives a stock HttpAgent against /agent/:id/suggest, so chips stream
  progressively via onMessagesChanged and the run never routes through the
  Intelligence websocket delegate (still no thread persistence)
- forward the consumer's deep-cloned messages + state onto the suggestion run
  (was state: {}), matching the clone fallback

Net -68 LOC of production code; the stateless and fallback paths now share one
runAgent flow.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-07 17:24:46 -05:00
Mark f452699510 Merge branch 'main' into fix/issue-5533-agentid-runtime-sync 2026-07-07 11:21:43 -07:00
Jordan Ritter 7897be4a95 feat(runtime): configurable inbound-header forwarding policy with default infra/platform denylist (#5783)
## Problem — the leak

The v2 runtime's `shouldForwardHeader` forwarded `authorization` **and
any header whose name starts with `x-`** onto the outgoing agent call.
In a real deployment the inbound request has already traversed a
browser, CDN/edge, load balancer, and hosting platform — each stamping
its own `x-*` headers — so the wide `x-*` wildcard silently forwarded:

- **Hop-by-hop / topology:** `x-forwarded-for`, `x-real-ip`,
`x-forwarded-proto/host/port`
- **Cloud / CDN tracing:** `x-amzn-trace-id`, `x-amz-cf-id`,
`x-cloud-trace-context`, `x-azure-*`, `x-fastly-*`, `x-request-id`
- **Platform-injected:** `x-vercel-*`, `x-middleware-*`
- **CopilotKit Cloud platform credential:**
`x-copilotcloud-public-api-key`

The last item is a real credential-exfiltration concern: a platform key
scoped to Copilot Cloud reaching a third-party agent URL. This is the
**breadth** half of #5712 (option 3); the **precedence** half was fixed
in #5782.

## Design — denylist default + config knob, both paths

- **Default denylist (safe default).** Keep the `authorization` + `x-*`
base eligibility, but strip a curated, greppable set of known
infra/proxy/platform headers (exact names + prefix families) before
forwarding. Legitimate custom `x-*` application headers (`x-tenant-id`,
`x-api-key`, …) keep flowing untouched. The authoritative list is a
single exported constant in `header-utils.ts`.
- **Configurable policy (`forwardHeaders` runtime option).**
- `useDefaultDenylist?: boolean` (default **true**) — `false` restores
the previous wide-open behavior.
  - `deny?` / `denyPrefixes?` — extend the default denylist.
- `allow?` — opt into strict allowlist mode (only listed headers
forward).
- **Resolve once.** The constructor resolves `forwardHeaders` into a
`forwardHeadersPolicy: ResolvedForwardHeadersPolicy` field (mirroring
the existing `debug` → `ResolvedDebugConfig` resolve-once), exposed on
`CopilotRuntimeLike` / `BaseCopilotRuntime` with a passthrough getter on
the `CopilotRuntime` shim.
- **Both paths.** The resolved policy is read at **/run**
(`configureAgentForRequest`) and **/connect** (`handleSseConnect`) via
`mergeForwardableHeaders`, so the two can never diverge. Server-wins
precedence and server-self case-dedup from #5782 are untouched.

## Semver

**Minor with an opt-out.** Removing a leak is a fix, not a contract
change, and we ship a documented escape hatch: `new CopilotRuntime({
agents, forwardHeaders: { useDefaultDenylist: false } })` restores the
prior behavior. Custom-header forwarders (the common case) are
unaffected.

## Red-green proof (real surface, both paths)

RED — with the predicate reverted to the old wide-open `authorization ||
x-*` (policy ignored), the new behavior assertions fail; the leak
reproduces (`x-forwarded-for: 203.0.113.7` forwards on both /run and
/connect):

```
 ❯ header-utils.test.ts (19 tests | 8 failed)
   × strips known infra/proxy/platform headers by exact name → expected true to be false
   × strips known infra/platform header families by prefix   → expected true to be false
   × strips denylisted headers case-insensitively            → expected true to be false
   × deny extends the default set                            → expected true to be false
   × denyPrefixes extends the default set                    → expected true to be false
   × allow switches to allowlist mode                        → expected true to be false
   × extractForwardableHeaders drops denylisted x-* infra    → expected {…4} to deeply equal {…1}
 ❯ agent-utils-header-forwarding.test.ts (/run) (10 tests | 1 failed)
   × strips denylisted infra/platform headers (#5712 breadth) → expected '203.0.113.7' to be undefined
 ❯ sse-connect-agent-id.test.ts (/connect) (5 tests | 1 failed)
   × strips denylisted infra/platform headers                → expected '203.0.113.7' to be undefined
```

GREEN — with the real policy in place:

```
 ✓ header-utils.test.ts (19 tests)
 ✓ agent-utils-header-forwarding.test.ts (10 tests)   # /run path
 ✓ sse-connect-agent-id.test.ts (5 tests)             # /connect path
 ✓ agent-header-precedence.test.ts (2 tests)
 Test Files  4 passed (4)
      Tests  36 passed (36)
```

Full `@copilotkit/runtime` suite: **113 files / 1593 tests passed.**
Typecheck, oxlint (0 errors), oxfmt, and build all green.

## Builds on #5782

This branches off #5782's head (`636bcad05`) and reuses that PR's
`mergeForwardableHeaders` (server-wins precedence + server-self
case-dedup). It should land **after #5782**. It addresses the
**forwarding-breadth half of #5712** — #5712's precedence core is fixed
by #5782; this is the breadth follow-up (not `Fixes #5712`).
2026-07-06 09:41:07 -07:00
Alem Tuzlak d8928c445a fix(runtime): server-configured agent headers take precedence over forwarded inbound headers (#5782)
## Problem

When a self-hosted v2 `CopilotRuntime` is configured with a server-side
agent (an `@ag-ui/client` `HttpAgent` with static `headers` for
service-to-service auth), the runtime forwards inbound
`authorization`/`x-*` request headers onto the agent's outgoing call
**and lets them override the headers the server configured** — silently
breaking service-to-service auth to a secured backend (e.g. a private
Cloud Run agent behind IAM).

`Fixes #5712`

## Root cause


`packages/runtime/src/v2/runtime/handlers/shared/agent-utils.ts:125-128`
merged forwarded inbound headers **last**, so they won on collision:

```ts
agent.headers = {
  ...agent.headers,                      // server-configured
  ...extractForwardableHeaders(request), // inbound — overrode the above
};
```

There are actually **two** failure modes:

1. **Same-case collision** — inbound `authorization` overwrites a server
`authorization` (last-write-wins).
2. **Case-mismatch collision** — `extractForwardableHeaders` lowercases
inbound keys (`authorization`), while the server typically configures
canonical casing (`Authorization`). A plain spread treats those as
*distinct* keys and emits **both** — which undici downstream comma-joins
into a single invalid `"Bearer A, Bearer B"` ("multiple JWTs") value.
Flipping the spread order alone does **not** fix this case.

## Fix

In `agent-utils.ts`, make server-configured `agent.headers`
authoritative on collision, matched **case-insensitively**: drop any
forwarded inbound header whose name (case-insensitively) is already set
on the agent, and let non-colliding inbound headers pass through
unchanged. This preserves the existing forward-for-auth behavior for
headers the server does *not* set, while guaranteeing a server-set token
is never overridden or duplicated.

The merge logic lives in a shared
`mergeForwardableHeaders(serverHeaders, request)` helper in
`packages/runtime/src/v2/runtime/handlers/header-utils.ts` so the
precedence semantics are defined in exactly one place.

### Scope note

This is the conservative precedence + case-insensitive-dedup fix (the
issue's suggested fix #1). I did **not** tighten the default allowlist
to drop hop-by-hop/platform `x-*` headers (`x-serverless-*`,
`x-forwarded-*`, …) or add an opt-out — those alter existing forwarding
behavior and are worth a separate, deliberate change. The precedence fix
alone resolves the reported breakage (the server-set token now wins
regardless of what the platform injects on a colliding header name).

A documented workaround already exists for users on released versions:
pass a custom `fetch` to the `HttpAgent` that builds outgoing headers
from scratch (it runs after `configureAgentForRequest` and survives the
per-request `agent.clone()`).

## Red-green proof (the real fix — `/run` path)

The load-bearing assertion: there must be exactly **one** authorization
header carrying the **server** value.

### RED (fix stashed, against unmodified `agent-utils.ts`)

```
 ❯ src/v2/runtime/__tests__/agent-header-precedence.test.ts (2 tests | 1 failed)
   × configureAgentForRequest — header precedence (#5712) > server-configured agent headers win over a colliding inbound header
AssertionError: expected [ 'Authorization', 'authorization' ] to have a length of 1 but got 2
     81|     expect(authKeys).toHaveLength(1);
 Test Files  1 failed (1)
      Tests  1 failed | 1 passed (2)
```

The pre-existing `agent-utils-header-forwarding.test.ts` also failed,
because it explicitly encoded the buggy behavior
(`expect(...["x-aimock-context"]).toBe("new-context")` — inbound
winning):

```
 FAIL  src/v2/runtime/__tests__/agent-utils-header-forwarding.test.ts > ... > request forwardable headers override matching pre-existing agent headers
AssertionError: expected 'old-context' to be 'new-context'
```

### GREEN (fix applied)

```
 ✓ src/v2/runtime/__tests__/agent-header-precedence.test.ts (2 tests) 2ms
 ✓ src/v2/runtime/__tests__/agent-utils-header-forwarding.test.ts (8 tests) 3ms
 Test Files  2 passed (2)
      Tests  10 passed (10)
```

The colliding test (`agent-utils-header-forwarding.test.ts`) was updated
from asserting the old bug to asserting corrected precedence + a new
case-insensitive-dedup guard. The non-colliding-forward test is retained
unchanged as a regression guard.

## Quality gates

```
NX  Successfully ran target check-types for project @copilotkit/runtime
NX  Successfully ran target test for project @copilotkit/runtime  — Test Files 113 passed (113), Tests 1576 passed (1576)
```

---

## `/connect`-path change — forward-looking plumbing, inert today

The original issue and a prior eval flagged the same forwarding pattern
at `handlers/sse/connect.ts`. To keep the two paths' merge semantics
consistent, the `/connect` path now builds the same server-wins merged
headers (via the shared `mergeForwardableHeaders` helper) and passes
them into `runner.connect()`.

**This is not an active auth fix, and it is not red-green-proven as one
— because there is no live bug to fix on the connect path today.** No
shipped runner consumes the `headers` field of
`AgentRunnerConnectRequest`: the in-memory, intelligence, telemetry, and
sqlite runners all destructure only `threadId` from the connect request
and ignore `headers` entirely. Connect is a thread replay/reconnect, not
a fresh outgoing agent call. So whatever headers we pass into
`runner.connect()` are dropped on the floor by every runner that ships.

What this change actually does:

- Threads the per-request agent clone through `handle-connect.ts →
handleSseConnect` so the connect path *has access to* the
server-configured `agent.headers` (it previously did not).
- Passes `mergeForwardableHeaders(agent?.headers, request)` into
`runner.connect()` — the correct, server-wins argument **shape** for a
future outbound-connecting runner that *would* consume connect-path
headers.
- Rewrites the comments/JSDoc on this path to say this plainly, rather
than implying an active auth fix. It also documents that the
connect-site `cloneAgentForRequest` call is the sole `agentId`-existence
guard (the intelligence branch never re-validates the id), and documents
`cloneAgentForRequest`'s `AbstractAgent | Response` (404) dual-return
contract that both callers depend on.

The real outbound header forwarding — the thing that fixes #5712 — is
the `/run` path's `agent.headers` mutation described above. The connect
change is staged plumbing so that if/when a runner starts honoring
connect-path headers, it inherits the same server-wins precedence
without a second fix.

### Tests on the `/connect` path

The connect tests assert the *merge shape* that reaches
`runner.connect()` (server value wins on collision, exactly one
`authorization` key, non-colliding `x-*` still forwards) and that the
agent-undefined case (no server `agent.headers`) degrades to forwarding
allowlisted inbound headers only and does not crash. These verify the
argument we construct is correctly shaped — not that any shipped runner
consumes it.

## Files

- `packages/runtime/src/v2/runtime/handlers/header-utils.ts` — shared
`mergeForwardableHeaders` helper (case-insensitive, server-wins).
- `packages/runtime/src/v2/runtime/handlers/shared/agent-utils.ts` —
`/run` path uses the helper so server headers win on collision (**the
real fix**).
- `packages/runtime/src/v2/runtime/handlers/sse/connect.ts` — `/connect`
path uses the helper; forward-looking plumbing, inert until a runner
consumes connect-path headers.
- `packages/runtime/src/v2/runtime/handlers/handle-connect.ts` — threads
the per-request agent clone into `handleSseConnect`.
-
`packages/runtime/src/v2/runtime/__tests__/agent-header-precedence.test.ts`
— `/run` regression test exercising the real `configureAgentForRequest`
surface with a real `HttpAgent`.
-
`packages/runtime/src/v2/runtime/__tests__/agent-utils-header-forwarding.test.ts`
— updated the test that encoded the old (buggy) precedence; added a
case-mismatch dedup guard.
-
`packages/runtime/src/v2/runtime/handlers/sse/__tests__/sse-connect-agent-id.test.ts`
— connect-path merge-shape + agent-undefined coverage.

### Notes

- A documented `@ag-ui/client` `HttpAgent` `fetch` workaround already
exists for attaching service-to-service auth the runtime can't override
(see the issue). This change makes the workaround unnecessary for the
`/run` precedence case.
- Conservative scope: this is the **precedence flip on `/run`** plus
forward-looking connect plumbing. Tightening the default allowlist
(dropping hop-by-hop / platform `x-serverless-*`, `x-forwarded-*`,
`x-cloud-trace-context`, …) and an opt-out switch — issue suggestions
#2/#3 — are intentionally left as a follow-up to keep the
security-policy change minimal.
2026-07-06 18:26:38 +02:00
Markus Ecker 7b72bd491a Merge branch 'main' into mme/memory-core 2026-07-03 10:20:47 +02:00