Commit Graph

61 Commits

Author SHA1 Message Date
Tyler Slaton b7d2fdcc25 fix(channels-intelligence): HTTP-transport robustness cluster (OSS-497) (#6037)
Six pre-existing `@copilotkit/channels-intelligence` HTTP-transport
robustness items surfaced by the pre-merge CR of #5983 (Linear OSS-497).
All confirmed against `main` after #5983 merged; none introduced by it.
Each is an independent, focused commit with a test.

## Fixes

1. **Heartbeat starvation mid-turn** — `HttpDeliverySource.runLoop`
heartbeated only at the top of each iteration, then blocked on
`onDelivery` for up to `turnTimeoutMs` (120s). With a 15s cadence, a
turn longer than the cadence sent no heartbeat, so app-api could mark a
healthily-working runtime stale mid-turn and withhold new deliveries.
Heartbeating now runs on a standalone recurring timer, independent of
the claim loop.
2. **`stop()` shutdown latency** — `stop()` set `running=false` then
awaited the loop, which only rechecked after its current sleep (≤15s
idle) or `onDelivery` (≤120s mid-turn); sleeps were `unref`'d but not
interruptible. Added a `stopWait` promise that the poll-sleep and
turn-wait both race, so shutdown is prompt. A mid-turn stop leaves the
lease for app-api to re-lease and does **not** nack (the turn didn't
fail); the turn's eventual settlement is always handled so a post-stop
rejection never surfaces as unhandled. *(1 & 2 share the runLoop
lifecycle, so they land in one commit.)*
3. **Empty-text `update` silently acked** — the `if (!text) return { ok:
true }` guard ran for both `post` and `update`. An empty POST is a legit
no-op, but an empty UPDATE (e.g. clearing a message body) that this
post-only fallback egress can't express was acked as success. Now
returns `{ ok: false, code: "empty_update" }`; empty posts still no-op.
4. **Static `adapter` on egress** — `emit` posted `this.cfg.adapter`
(default `"slack"`) alongside a possibly-Teams `replyTarget`. Confirmed
app-api's egress route (`sendChannelEgressMessage`) routes on
`replyTarget.adapter` + `channelName` and ignores this field, so it's
latent — but now derived from the delivery's own reply route to avoid
the contradiction. *(Note: the listener heartbeat's
`declaredChannels[].adapter` is intentionally left alone — app-api
genuinely consumes it for per-adapter health/conflict via the
`channel_adapter_configs.provider` join, and declaring the bot's full
adapter set is a separate design change, not a robustness cleanup.)*
5. **`projectId` strict-`number` only** — the realtime scope build
accepted `projectId` only as a JS `number` while org/channel were
`String()`-coerced, so a numeric-string on the untyped wire silently
fell back to the transport-default projectId, defeating per-delivery
scope authority. Extracted a `coerceWireProjectId` helper (number or
numeric-string → positive integer) with fallback.
6. **File/history client ignored injectable `fetch`** —
`fetchFile`/`getHistory`/`uploadFile` hardcoded `globalThis.fetch`, so
an injected `config.fetch` was honored everywhere except
binary/history/upload. The transport's `FetchLike` is text-only and
can't satisfy the binary client, so the file/history client got its
**own** injectable full `fetch` (`typeof fetch`), resolved via a single
helper with a global fallback.

Plus a small `docs` commit fixing a doc-comment placement displaced by
the item-5 export.

## Testing

- `nx run @copilotkit/channels-intelligence:test` — **178 passed** (170
baseline + 8 new).
- `nx run @copilotkit/channels-intelligence:check-types` — clean.
- lefthook (full package test + typecheck) ran on every commit.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-17 17:16:32 -07:00
Benjamin Taylor 15a8ddb0f7 fix(channels-intelligence): address adversarial CR findings (OSS-497)
Follow-up to the OSS-497 robustness cluster after a 3-agent adversarial review.

- Empty-text update: reclassify from a hard {ok:false} failure to a logged
  no-op. The failure path threw up the run and nack'd the WHOLE turn retryably,
  re-running valid work to dead-letter for a condition a retry can't fix (an
  agent clearing a message body the post-only fallback can't express). Now it is
  skipped with a distinct warning log — surfaced, not silent, and not turn-fatal.
- Injectable fetch: plumb it through the transports. Both HttpDeliverySource and
  RealtimeGatewayTransport now accept a binary-capable /
  and forward it to IntelligenceFileHistoryClient, so a consumer injecting a
  fetch has it honored on the download/history/upload paths (previously the
  client-level injectable was reachable only by constructing the client directly).
- coerceWireProjectId: use Number.isSafeInteger so a numeric string beyond
  MAX_SAFE_INTEGER is rejected (was silently coerced to a lossy integer).
- Malformed-present projectId: log before falling back to the transport scope
  (an absent projectId still falls back quietly), mirroring the loud drops in
  the same function instead of silently masking a wire-corruption signal.
- Restart safety: guard the recurring heartbeat with a run generation so a
  stop->start while a heartbeat POST is in flight can't leave two reschedule
  chains doubling the heartbeat rate; bump the generation on stop().
- Remove now-dead lastHeartbeatAt state (only reader was the removed in-loop
  heartbeat check).

Tests: empty-update now asserts log + no-op; added transport fileFetch
wire-through, adapter-from-route fallback, no-unhandled-rejection after a
mid-turn stop, heartbeat-stops-after-stop, elapsed-time bounds on the stop()
promptness tests, and coerceWireProjectId edge cases (leading zeros, > safe int).
nx test (green) + check-types (clean).
2026-07-17 15:22:10 -05:00
Benjamin Taylor 81562f5b24 docs(channels-intelligence): restore assertValidChannelRealtimeScope doc placement
The coerceWireProjectId export (added for OSS-497) was inserted directly above
assertValidChannelRealtimeScope, displacing that function's doc comment onto the
helper. Reorder so each function carries its own docblock.
2026-07-17 14:46:27 -05:00
Benjamin Taylor 0d412f3bb8 fix(channels-intelligence): honor an injectable fetch in the file/history client
IntelligenceFileHistoryConfig's fetchFile/getHistory/uploadFile hardcoded
globalThis.fetch, so a consumer (or test) injecting a fetch had it honored
everywhere EXCEPT the binary download, history, and upload paths. The transport's
FetchLike is text-only ({ ok, status, text() }) and cannot satisfy the binary
client (.body/.arrayBuffer()/.headers/.json()), so add the client's OWN
injectable full `fetch` (typeof fetch) on its config, resolved via a single
resolveFetch() helper that falls back to the global fetch. Each caller keeps its
existing degrade/throw contract for the no-fetch runtime.

Surfaced by the pre-merge CR of #5983 (OSS-497).
2026-07-17 14:45:34 -05:00
Benjamin Taylor ba38ee0e5f fix(channels-intelligence): heartbeat during long turns and make stop() prompt
Two coupled HttpDeliverySource runLoop lifecycle robustness fixes:

Heartbeat starvation: the loop heartbeated only at the top of each iteration,
then blocked on onDelivery for up to turnTimeoutMs (120s). With a 15s cadence,
a turn longer than the cadence sent no heartbeat for its whole duration, so
app-api could expire the activation and mark a healthily-working runtime stale
mid-turn, withholding new deliveries. Move heartbeating to a standalone
recurring timer that fires on cadence regardless of what the loop is doing
(rescheduling only after each heartbeat settles, so no overlap).

Shutdown latency: stop() set running=false then awaited the loop, but the loop
only rechecked after its current sleep (up to cadence, idle) or onDelivery (up
to turnTimeoutMs, mid-turn) — the sleeps were unref'd but not interruptible, so
shutdown could take up to 15s idle or 120s mid-turn. Add a stopWait promise that
stop() resolves; the poll-sleep and the turn-wait both race it, so shutdown is
prompt. A mid-turn stop leaves the lease for app-api to re-lease and does NOT
nack (the turn didn't fail); the turn's eventual settlement is always handled so
a post-stop rejection never surfaces as unhandled.

Surfaced by the pre-merge CR of #5983 (OSS-497).
2026-07-17 14:42:39 -05:00
Benjamin Taylor 787264fa8f fix(channels-intelligence): accept a numeric-string projectId in the realtime scope build
RealtimeGatewayTransport.toIngressEnvelope built the per-delivery scope with
`typeof delivery.projectId === 'number' ? delivery.projectId : this.scope.projectId`,
while organizationId/channelId were String()-coerced. The delivery.available
payload is untyped JSON, so a numeric-string projectId ('9' on the wire) failed
the strict check and silently substituted the transport-default projectId —
defeating the per-delivery scope authority and routing the render-accept under
the wrong project.

Extract a `coerceWireProjectId` helper (number or numeric-string -> positive
integer, else undefined) and use it with a fallback to the transport default.

Surfaced by the pre-merge CR of #5983 (OSS-497).
2026-07-17 14:35:49 -05:00
Benjamin Taylor 46517eecd9 fix(channels-intelligence): derive egress adapter from the reply route
HttpEgressSink.emit posted the static `this.cfg.adapter` (default "slack")
alongside a possibly-Teams `replyTarget`, so a Teams delivery through the
fallback egress carried a contradictory `adapter: "slack"`. app-api's egress
route (sendChannelEgressMessage) routes on `replyTarget.adapter` + channelName
and ignores this top-level field, so this is latent — but one provider-agnostic
runtime serves every adapter its bot has attached, so posting the config default
is misleading. Derive the posted adapter from the delivery's own reply route,
falling back to the config adapter when the route carries none.

Note: the listener heartbeat's declaredChannels[].adapter is NOT changed here —
app-api genuinely consumes it (the channel_adapter_configs.provider join) for
per-adapter health/conflict reporting, and the runtime only knows its single
configured adapter; declaring the bot's full adapter set is a separate design
change, not this robustness cleanup.

Surfaced by the pre-merge CR of #5983 (OSS-497).
2026-07-17 14:32:42 -05:00
Benjamin Taylor 060cebf304 fix(channels-intelligence): fail loud on empty-text egress update
HttpEgressSink.emit ran the `if (!text) return { ok: true }` no-op guard for
both post and update ops. An empty POST is a legitimate no-op (nothing to say),
but an empty UPDATE — a real intent such as clearing a message body that this
post-only fallback egress cannot express — was acked as success while nothing
was sent, inconsistent with the module's fail-loud posture. Return
`{ ok: false, code: 'empty_update' }` for that case; empty posts still no-op.

Surfaced by the pre-merge CR of #5983 (OSS-497).
2026-07-17 14:28:25 -05:00
Benjamin Taylor 0b69bb4b4d docs(channels): clarify read_thread getMessages is text-only (OSS-488)
The getMessages JSDoc cited "what was in the image" as a use case, but
image/file parts contribute no text in this mapping — read_thread is
text-only. Image content reaches the model only via conversationStore's
seeding of agent.messages. Tweak the comment so it no longer implies
read_thread can see image content.
2026-07-17 13:52:36 -05:00
AlemTuzlak 21b9a7138b chore: release channels v0.2.1 2026-07-17 09:13:35 +00:00
Alem Tuzlak a06c8cf1be feat(channels): canonical cross-platform reaction normalization for Teams & WhatsApp
Normalize inbound reaction emoji to a canonical cross-platform name at the
central channels-core ingress, so onReaction handlers match one value
regardless of provider.

- channels-ui/emoji: add "teams" and "whatsapp" to EmojiPlatform and canonical
  entries for refresh/laugh/surprised/sad/angry. Teams normalizes its
  <codepoint>_<name> emoji codes (1f504_refresh -> refresh) and classic bare
  names (like -> thumbs_up); an out-of-range codepoint degrades to passthrough
  instead of throwing. WhatsApp uses the unicode path like Discord/Telegram.
- channels-core: add teams+whatsapp to EMOJI_PLATFORMS; IncomingReaction gains
  an optional source `platform`; onReaction normalizes by
  evt.platform ?? adapter.platform.
- channels-intelligence: the managed reaction dispatch passes the delivery's
  source platform (env.platform) so managed reactions normalize too.
2026-07-16 12:55:57 -05:00
tylerslaton add6539dcd chore: release channels v0.2.0 2026-07-15 20:29:29 +00:00
Tyler Slaton 9504e0666c refactor(channels): add batteries-included package architecture (#5948)
## Summary

- Extract the platform-neutral Channels foundation into
`@copilotkit/channels-core`.
- Make `@copilotkit/channels` the batteries-included consumer entry
point, with adapter and UI subpaths.
- Release the umbrella, core, UI, and all six adapters as one shared
`channels` version scope.
- Verify the packed consumer contract and migrate the Slack and Teams
examples to the umbrella.

## Why

Consumers should be able to install one Channels package without making
the runtime or selective integrations depend on every platform SDK.
Shipping the complete Channels family together prevents adapter/core
version drift and makes the umbrella's exact dependency set release as a
compatible unit.

## How

- Move shared bot/runtime primitives into `channels-core` and invert
adapter/runtime dependencies.
- Add exact workspace dependencies and export subpaths from the umbrella
package.
- Consolidate the existing release configuration and all three release
workflow selectors into one shared `channels` scope containing all nine
packages.
- Use `@copilotkit/channels` as the version source: the next minor
release resolves to `0.2.0` and bumps every Channels package together.
- Publish scoped packages in dependency order for both stable and canary
releases: UI, core, adapters, then the umbrella.
- For a stable Channels release, publish the UI/core/adapters first, run
the registry-backed packed-consumer verifier against those newly
published exact versions, then publish the umbrella.
- Generate Channels release notes from `channels/v*` tags rather than
the monorepo `v*` tags.
- Verify builds, type checks, tests, package artifacts, examples, and
release workflow scope synchronization locally.

### First stable release sequence

1. Bootstrap the currently unpublished `@copilotkit/channels-core`
package on npm and configure npm trusted publishing for every Channels
package against this repository's `release / publish` workflow and `npm`
environment.
2. Create and merge the `channels` minor release PR. The stable workflow
publishes the Channels family in the staged order above and validates
the packed umbrella from the registry before the umbrella is released.
3. Create and merge a subsequent `monorepo` release PR so the published
`@copilotkit/runtime` switches from the historical umbrella dependency
to `@copilotkit/channels-core`.
2026-07-15 13:12:16 -07:00
Benjamin Taylor a8407f84fd fix(channels): address Tyler CR — wire app-api URL through runtime activation + map user name (OSS-476)
- [P1] History/files were absent on the NORMAL managed path: defaultActivateChannel
  never forwarded the app-api HTTP URL, so the transport (which installs
  fetchFile/getHistory/uploadFile only when appApiBaseUrl is set) ran without
  them for Channels started by the CopilotRuntime handler — only manual
  low-level launcher callers got file/history. Thread intelligence.ɵgetApiUrl()
  through ChannelActivationConfig.apiUrl → defaultActivateChannel →
  startChannelsOverRealtimeGateway({ appApiBaseUrl }). The launcher + transport
  already accepted it.
- [P2] Managed turns exposed the provider profile under a non-public `displayName`
  field, leaving PlatformUser.name undefined. Map env.user.displayName -> name
  (parity with the direct Slack adapter, which populates `name`).

Tests: deriver returns apiUrl; defaultActivateChannel forwards appApiBaseUrl to
the launcher opts; onMessage sees message.user.name. channel-activation-config +
channels-intelligence (170) green; runtime build type-checks. (channel-manager.test
executes in CI — local vitest hits the known optional-peer-dep resolution flake.)
2026-07-15 13:36:36 -05:00
Benjamin Taylor 237cbf4cf9 chore(channels): CR hygiene — real test coverage + doc/log cleanups (OSS-476)
Trivial items from the CR confirmation rounds (no behavior change to production
paths):

- intelligence-adapter.test.ts: `delete source.getHistory` was inert (getHistory
  is a prototype method; delete only removes own props), so the "no getHistory"
  branch was never actually exercised. Shadow with an own `undefined` instead so
  the `source?.getHistory?.()` short-circuit is genuinely tested.
- in-memory-transports.ts: mirror the production `limit <= 0 -> []` guard in
  InMemoryDeliverySource.getHistory (was `slice(-limit)` → returns ALL for 0).
- realtime ack(): log the empty-turn (no accepted frames) drop so the OSS-491
  redelivery pile-up is diagnosable (every other drop path here logs).
- http-transports.ts: remove two orphaned JSDoc comments dangling over
  ClaimResponse (leftovers from the file/history extraction).
- intelligence-adapter.ts: correct an inaccurate op-id comment (mintOp is the
  ${turnId}:${seq} source; the render path keys on ${turnId}:${slot}:${seq}).

channels-intelligence 170 tests + check-types green.
2026-07-15 12:58:53 -05:00
Tyler Slaton 5811cab177 fix(channels): align rebase with channel API 2026-07-15 10:29:21 -07:00
Benjamin Taylor 2718474229 fix(channels): second CR round — graceful realtime stop() + parity/hygiene (OSS-476)
From the 7-agent CR confirmation round:

- RealtimeGatewayTransport.stop() now halts intake (a `stopped` guard in
  handleDeliveryAvailable — the session exposes no `off` to detach the
  DELIVERY_AVAILABLE listener) and DRAINS the in-flight delivery (awaits the
  serial `processing` chain) before clearing state, so a turn settling at stop
  time still sends its terminal signal instead of silently no-oping and
  redelivering. Mirrors HttpDeliverySource.stop().
- Realtime nack() truncates the reason to 500 chars (parity with HTTP).
- getMessages drops empty content parts before join(" ") so a read_thread
  transcript isn't corrupted with doubled/leading/trailing spaces (a test had
  enshrined "part one  part two").
- Removed dead imports (buildContentParts, AgentContentPart, ChannelFileRef)
  left in http-transports.ts after the file/history extraction.

Deferred (delivery-contract / design; need coordination — folded into OSS-491):
timeout-nack can redeliver a still-running turnId (overlap); realtime push()
no-state fallback vs HTTP fail-loud (documented intentional — parity Q); the
empty-turn completion signal.

Tests: stop() drains in-flight + ignores post-stop deliveries; getMessages
assertion corrected. channels-intelligence 170 tests + check-types green.
2026-07-15 12:21:53 -05:00
Tyler Slaton 3dba09b7c7 feat(channels): add the batteries-included umbrella 2026-07-15 10:12:23 -07:00
Tyler Slaton 06fb2ed34a refactor(channels): extract the platform-neutral core 2026-07-15 10:11:55 -07:00
github-actions[bot] c2e2b4c0fd style: auto-fix formatting 2026-07-15 17:08:16 +00:00
Benjamin Taylor 54fa303477 fix(channels): address adversarial CR findings on the realtime transport (OSS-476)
Pre-merge 7-agent CR of #5983 surfaced several realtime-transport defects (all
in channels-intelligence). Fixes:

- Poison-payload re-lease loop: an unmappable delivery (unmodeled reply-target
  adapter / unknown input.kind) with a valid lease was logged + dropped, so
  app-api re-leased the identical payload forever. It now fails NON-retryable
  (dead-letter), mirroring the HTTP path. nack() gains a `retryable` param.
- Double-terminal-signal race: realtime ack()/nack() deleted delivery state
  AFTER the wire push, so a per-turn-timeout nack could race a late dispatch ack
  and emit BOTH fail + complete_requested. Now delete-before-push (XOR).
- Concurrent dispatch: deliveries were handled fire-and-forget; now processed
  serially (parity with the HTTP runLoop) so an in-flight redelivery can't reset
  the shared per-turn seq counter or run two turns of one conversation at once.
- actor.displayName was carried through the claim mapper then dropped in
  dispatchTo (`{ id }` only) — now forwarded to handlers (OSS-476 identity).
- fetchFile enforced MAX_INBOUND_FILE_BYTES only against declared content-length;
  the actual read was unbounded. Now streams and aborts past the cap.
- getHistory returned the FULL history for limit<=0 (slice(-0)) and hydrated file
  bytes for over-returned messages it then discarded. Now caps to the most recent
  `limit` BEFORE hydrating; limit<=0 -> [].
- stream() posted an empty text frame for an empty stream; now skips the post.
- HTTP withTimeout/defaultSleep timers now unref() (parity with realtime).
- Corrected the inaccurate thread_started comment in claim-mapping.ts.

Deferred to OSS-491 (delivery-terminal-signal contract, needs app-api
coordination): an empty turn (reaction/command that posts nothing) has no valid
completion signal (acceptedThrough requires >=1) and redelivers; and the
run_error swallow on the HTTP-fallback render path.

Tests: poison non-retryable, single-terminal XOR, serial dispatch, displayName
forwarding, getHistory limit<=0 + cap. channels-intelligence 169 tests + check-types green.
2026-07-15 12:06:50 -05:00
Benjamin Taylor bc883e89c0 feat(channels-intelligence): realtime transport parity — kinds, identity, history/files, delete (OSS-476)
The realtime-gateway transport had drifted from the HTTP transport: it built a
text-only ingress envelope that coerced commands/reactions/interactions into
empty turns, keyed conversations per-turn (breaking threaded follow-ups),
dropped the provider actor identity, and implemented none of
fetchFile/getHistory/uploadFile (so the realtime path silently ran with no
history and no file support). It also had no `delete` render kind. This brings
the realtime path to full parity with direct.

- claim-mapping: extract the claim→ingress mapping (ClaimedDelivery,
  conversationKeyFromReplyTarget, mapDeliveryToEnvelope) into a shared module
  both transports use, so they cannot drift again. Add the provider `actor` to
  the claim turn and map it to `env.user` (fixes identity on BOTH paths).
- realtime transport: build the envelope via the shared mapper — real kind
  discrimination, thread-stable conversationKey, actor→user — instead of the
  text-only inline build. Fail-closed: an unmodeled reply-target/kind is
  dropped+logged, not crashed.
- file/history: extract fetchFile/getHistory/uploadFile into a shared
  IntelligenceFileHistoryClient (HTTP-only — the gateway never relays bytes or
  history). The realtime transport gains them when configured with
  `appApiBaseUrl` + `apiKey` (threaded through the launcher); absent that, the
  methods stay undefined and the adapter degrades exactly as before.
- delete render kind: add `{ kind: "delete"; ref }` to ChannelRenderEvent
  (mirrors the frozen Intelligence contract) and route thread.delete through a
  render frame when a render sink is wired (OSS-420), like post/update/file.

Tests: shared-mapper unit tests (actor→user, kind discrimination,
conversationKey, unmodeled-adapter throw); realtime tests for non-text kind +
identity + thread-stable key and file/history capability toggling; a
delete-render-frame test. Full suite + check-types + build green.
2026-07-15 11:41:01 -05:00
Ben Taylor 183b0d9dc9 fix(channels): realtime-transport reliability hardening + residual bot-naming scrub (closes OSS-485, OSS-482) (#5972)
Combines the two OSS-473 follow-ups (originally #5972 + #5973, now
folded here) into one PR for a single review/merge. Base is `main`
(OSS-473/#5963 has merged). Reviewable **commit-by-commit** — the
mechanical rename is isolated in commit 1; the behavior changes are
commits 2–4.

## Commit 1 — `refactor(channels)`: scrub residual internal "bot" naming
(closes OSS-485)

Mechanical, naming-only, **no behavior change** — the last internal
"bot" vestiges left out of 473's atomic telemetry-surface commit.

- `bot` local variable → `channel` in `create-channel.ts` internals +
the ~15 tests that exercise it.
- `BotNode` type → `ChannelNode` in `@copilotkit/channels-ui` and
**every** importer (channels,
slack/teams/discord/telegram/whatsapp/intelligence adapters, slack/teams
examples — 256 refs).
- `botName` adapter-SPI option → `channelName` on `AdapterStartContext`
+ its `create-channel` caller + the one consumer
(`IntelligenceAdapter.start` ctx), in one change so the SPI can't drift.
(Phoenix wire contract was already `channelName` as of 473.)
- Stale `bot-ui`/`bot-slack` comment refs →
`channels-ui`/`channels-slack`; `Symbol.for("copilotkit.bot-ui.*")` →
`"copilotkit.channels-ui.*"`.
- **Kept** (out of scope): the platform `isBot?` author flag (unrelated
semantic) and example `bot` instance vars / "demo bot" prose
(user-facing).

## Commits 2–4 — `fix`: realtime-transport reliability & parity
hardening (closes OSS-482)

The self-contained observability/robustness subset of the OSS-473 CR
follow-ups (`packages/channels-intelligence` transport +
`packages/runtime` manager). Each hardening ships with a test that fails
without the fix.

- **`emit()` fail-loud** (`intelligence-adapter.ts`) — returned a
synthetic `MessageRef` on `{ ok: false }`, acking a failed
post/update/delete as success (silent egress drop). Now throws → the
delivery is nacked/retried.
- **Realtime delivery dispatch** (`realtime-gateway-transport.ts`) — the
`void handleDeliveryAvailable(...)` fire-and-forget had no error
boundary and no deadline. Replaced with a `.catch` + a bounded per-turn
timeout that nacks + logs on failure/timeout (parity with the HTTP
`runLoop`).
- **`ChannelManager.stop()` per-handle timeout** — a wedged
`handle.stop()` hung teardown/SIGTERM forever. Each is now bounded by
`stopHandleTimeoutMs` (default 5000); on timeout it's logged and
abandoned so other entries still stop.
- **`ChannelManager.ready({ timeoutMs })`** — a set-wide timeout
discarded an erroring channel's reason when a sibling hung. The deadline
is now **per channel**, so the `AggregateError` carries both the real
activation error **and** a named timeout for each hanging channel.
- **Forward `ChannelManager` `log` down** to the launcher/transport so
transport-level drop diagnostics (e.g. a version-skew
missing-`leaseToken` outage) aren't silent in the managed path.
- **`examples/slack/.env.example`** — OpenAI-only `AGENT_MODEL`
guidance; removed never-read `ANTHROPIC_API_KEY`/`GOOGLE_API_KEY`;
blanked the presence-gated `LINEAR_API_KEY`/`NOTION_*` placeholders (a
non-blank value wires a broken MCP); documented `NOTION_MCP_PORT`; added
WhatsApp to the adapter list.

### Already landed in OSS-473 (not re-done here)
The reconnect **"gave-up → error" escalation** is fully implemented on
`main`: `realtime-gateway.ts` bounds the reconnect window
(`reconnectGiveUpMs`, default 60s) and emits a terminal `gave_up`; the
launcher exposes `onStateChange`;
`ChannelManager.registerConnectionObserver` maps `gave_up → error`
(covered by `realtime-gateway.test.ts` +
`channel-manager-reconnect.test.ts`).

### Explicitly EXCLUDED — owned by OSS-474 / OSS-476
Left untouched (overlap in-flight work): no cross-turn history on the
realtime path (per-turn `conversationKey`) → OSS-474/476; empty-turn
`ack()` redelivery livelock → OSS-474/475/476; non-text parity
(command/interaction/reaction) → OSS-476; `thread.delete()` render frame
→ OSS-476 (PR #554); durable StateStore on the managed realtime path →
OSS-474/476.

## Testing
- **DoD grep** — `grep -rn "\bbot\b\|botName\|BotNode" packages/channels
packages/channels-ui` (`.ts`/`.tsx`, excl. `dist/`) → **0**; `BotNode`
gone repo-wide.
- **check-types** green across all 8 channels packages +
`@copilotkit/runtime`.
- **tests** green: channels 155, channels-ui 23, slack 277, teams 88,
discord 193, telegram 149, whatsapp 73, channels-intelligence 144,
runtime `channel-manager` 40; slack-example 63, teams-example 2. Full CI
matrix (incl. `unit` 20.x/22.x/24.x) green on this branch.
- Each 482 hardening has a test that fails pre-fix (no silent ack;
dispatch nacks on throw/timeout; `stop()` resolves + logs on a wedged
handle; `ready()` aggregate preserves the real reason;
`defaultActivateChannel` forwards `log`).

Closes OSS-485. Closes OSS-482.
2026-07-15 11:29:09 -05:00
Benjamin Taylor 2b4b19cf3b fix(channels): fail loud on egress failure + bound realtime delivery dispatch (OSS-482)
Two transport reliability hardenings on the managed realtime path, from the
OSS-473 CR.

emit() fail-loud: IntelligenceAdapter.emit() returned a synthetic MessageRef on
`{ ok: false }`, acking a failed post/update/delete as success (silent egress
drop on the HTTP-fallback path). It now throws with the failure code so the
failure propagates up the render/run path and the delivery is nacked/retried.

Realtime delivery dispatch: `session.on(delivery.available)` fired
`void this.handleDeliveryAvailable(...)` with no error boundary and no per-turn
deadline — an onDelivery rejection became an unhandled rejection (silent drop)
and a hung handler pinned the delivery forever. Replaced the `void` with a
`.catch` and wrapped the turn in a bounded per-turn timeout that nacks + logs on
failure/timeout (parity with the HTTP runLoop's turnTimeoutMs).

Tests: egress `{ ok: false }` makes thread.post throw (no silent ack); an
onDelivery throw and an onDelivery that exceeds deliveryTimeoutMs both nack
(channel.delivery.fail.v1) and log.
2026-07-15 07:41:17 -05:00
Benjamin Taylor 483880583f refactor(channels): scrub residual internal "bot" naming → "channel" (closes OSS-485)
Follow-up to the OSS-473 clean-break rename (#5963). Scrubs the last
internal "bot" vestiges deliberately left out of 473's atomic
telemetry-surface commit. Naming-only, no behavior change.

- `bot` local variable → `channel` in create-channel.ts factory internals
  and the ~15 test files that exercise it.
- `BotNode` type → `ChannelNode` in @copilotkit/channels-ui and every
  importer (channels, slack/teams/discord/telegram/whatsapp/intelligence
  adapters, and the slack/teams examples).
- `botName` adapter-SPI option → `channelName`: renamed on
  AdapterStartContext, its create-channel caller, and the one adapter that
  reads it (IntelligenceAdapter.start ctx) in the same change so the SPI
  cannot drift. The Phoenix wire contract was already `channelName` (473).
- Stale `bot-ui`/`bot-slack` package refs and "bot core"/"the bot" prose in
  channels/channels-ui comments → `channels-ui`/`channels-slack`/"channel".
- `Symbol.for("copilotkit.bot-ui.*")` → `"copilotkit.channels-ui.*"`.

Deliberately kept: the platform `isBot?` author flag (an unrelated
"is this message author a bot account" semantic), and example `bot`
instance variables / "demo bot" prose (user-facing, out of this ticket's
internal-symbol scope).

Validation: DoD grep returns zero internal bot symbols in
packages/channels + packages/channels-ui; check-types + test green across
all 8 channels packages (1093 tests) and both examples.
2026-07-15 07:39:50 -05:00
Ben Taylor 0b00f6ea37 fix(channels-intelligence): Teams thread history (getHistory + getMessages) (#5969)
## Problem
On the managed Teams path the agent always reported "no earlier
messages" / "I don't see the image in this thread."
`HttpDeliverySource.getHistory` was **Slack-only**: it keyed off
`threadTs` and returned `[]` for any route without one. A Teams route is
`{ adapter: 'teams', tenantId, conversationId }` (no `threadTs`), so it
short-circuited *before making any request* — starving **both** history
mechanisms on Teams:
- `agent.messages` seeding (`conversationStore.getOrCreate`), and
- the `read_thread` tool (via `thread.getMessages()`).

## Fix
- **`getHistory` is now adapter-aware** (mirrors
`conversationKeyFromReplyTarget`'s per-adapter switch): Slack keys off
`teamId`/`channel`/`threadTs`; Teams sends `adapter=teams` + `tenantId`
+ `conversationId`, matching app-api's
`teams:{tenantId}:{conversationId}` thread_key. app-api's
`/api/channels/history` route already accepts this shape. **Slack query
and order are unchanged.**
- **Add `getMessages` to the adapter** so `thread.getMessages()` (the
`read_thread` tool) reads reconstructed history via the transport and
maps it to `ThreadMessage[]`. Without it `Thread.getMessages()` returns
`[]` and thread-reading tools (summaries, "what was in the image") see
nothing even when history exists.

## Tests
- New: Teams-shaped `getHistory` query, and the
missing-`tenantId`/`conversationId` short-circuit (no request).
- All 115 `channels-intelligence` tests pass; `check-types` clean.

## Notes
Verified end-to-end against a live managed Teams bot as a dist patch
before porting to source (history seeding + `read_thread` both start
returning the thread's messages). The app-api counterpart (Teams
ingress: reactions, inline-media/Graph file ingest, slash commands) is a
separate PR in the Intelligence repo.
2026-07-15 07:39:17 -05:00
Benjamin Taylor 0e5dc5f97c fix(channels-intelligence): make gave_up sticky until a successful rejoin (OSS-473) 2026-07-15 07:15:06 -05:00
Benjamin Taylor decd4c20d5 fix(channels-intelligence): truthful connection-health for managed channels (OSS-473)
Classify per-channel state on channel_declaration_unavailable rejects so a
runtime_conflict is a hard error rather than being downgraded to setup_required;
make gave_up recoverable (a later rejoin restores online); and route Phoenix
channel-level close/error through the same health transition as socket drops.
2026-07-15 07:15:05 -05:00
Benjamin Taylor 496bb57e15 refactor(channels): rename telemetry surface oss.bot.* → oss.channel.* (OSS-473) 2026-07-15 07:15:05 -05:00
Benjamin Taylor 08d14dcc8c fix(runtime): reachable setup_required + connection-health status from the real gateway (OSS-473)
P1#2 — reachable setup_required on the PRODUCTION engine path.
connectRealtimeGateway no longer flattens every join rejection into a
generic Error. A join `.receive("error", reason)` whose reason is a known
setup-required code (`channel_declaration_unavailable`, and defensively
`adapter_setup_required` / `not_configured`) now rejects with a
distinguishable `RealtimeGatewaySetupRequiredError` (`code === "SETUP_REQUIRED"`,
raw reason preserved). ChannelManager already detects that code, so an
unconfigured managed provider now degrades to `setup_required` (ready()
resolves) instead of `error`. All other reasons keep the generic error and
the socket-leak teardown is unchanged.

P1#3 — status() reflects real connection health instead of `online` forever.
ConnectedRealtimeGatewaySession exposes `onStateChange(cb)` over
`RealtimeGatewayConnectionState` (`online` | `reconnecting` | `gave_up`),
driven by the real Phoenix seams: an unexpected socket drop → `reconnecting`;
a successful (re)join (the join-push recHooks survive Phoenix `resend`, so
`"ok"` re-fires on every auto-rejoin) → `online`; and a BOUNDED give-up —
Phoenix retries forever, so a `reconnectGiveUpMs` window (default 60000, runs
from the first drop of an episode, cleared on rejoin) elapsing while still
reconnecting → `gave_up` (terminal). Our own disconnect() stays silent.
ChannelManager wires this in place of the log-only onClose breadcrumb:
`reconnecting`→status reconnecting, `online`→online, `gave_up`→error; a
stopped manager/entry ignores late events. computeOverall now ranks
`error > reconnecting > setup_required > connecting > online`. ready() keeps
its one-shot semantics (settles on the initial outcome); later health
transitions move only status(). Docs updated to state `online` means
currently-sendable.

Wording: the direct-adapter skip comment/log now states delivery is
exclusive-per-platform (managed OR direct per platform, not both — attaching
both would double-deliver) with true coexistence tracked in OSS-484. Skip
behavior unchanged.

Call-sites for the changed signatures:
- connectRealtimeGateway error shape: only caller is
  startChannelsOverRealtimeGateway (realtime-gateway-launcher.ts:216); it
  awaits and lets the rejection propagate, so the setup-required error flows
  through unchanged (no branch to update).
- new ConnectedRealtimeGatewaySession.onStateChange: passed through in
  startChannelsWithGatewaySession and startChannelsOverRealtimeGateway
  (realtime-gateway-launcher.ts); added to ChannelsHandle (runtime.ts) and the
  manager's local ChannelsHandle view (channel-manager.ts); exported from
  index.ts. RealtimeGatewaySession (base, no observer) consumers
  (realtime-gateway-transport.ts) unaffected.
- manager onClose→state transitions: registerOnClose renamed to
  registerConnectionObserver; sole caller is the online settle handler in
  activate().
2026-07-15 07:15:05 -05:00
Benjamin Taylor 6703232f4f fix(channels-intelligence): isolate onClose callbacks, guard safeReason, cover join-failure teardown (OSS-473) 2026-07-15 07:15:04 -05:00
Benjamin Taylor 2c5c1138ea fix(channels-intelligence): reset onClose latch on rejoin so it fires once per drop (OSS-473) 2026-07-15 07:15:03 -05:00
Benjamin Taylor 1f6446caca fix(channels-intelligence): preserve session receiver in onClose passthrough; clarify V1 text-turn comment (OSS-473) 2026-07-15 07:15:03 -05:00
Benjamin Taylor 2a9b760c3a feat(channels-intelligence): expose onClose on managed session/handle for reconnect (OSS-473) 2026-07-15 07:15:01 -05:00
Benjamin Taylor c9f2119de3 refactor(channels-intelligence): make realtime scope org/channel ids optional (OSS-473) 2026-07-15 07:15:00 -05:00
Benjamin Taylor 229f3f1460 refactor(channels): rename Bot/createBot public API to Channel/createChannel across the workspace (OSS-473) 2026-07-15 07:14:59 -05:00
Alem Tuzlak 3adea157a4 test(channels-intelligence): cover getMessages mapping + address review
Addresses review on #5969:
- Add getMessages tests: string content, content-part array (text parts joined,
  non-text contributes ""), and role→isBot/user derivation; plus the no-getHistory
  transport case returning [].
- Log on the unexpected getMessages catch (matches conversationStore's seeding
  path) instead of degrading to [] silently — HttpDeliverySource already swallows
  its own fetch failures, so this outer catch only fires on a real throw.
- Fold the duplicated role boolean (isBot computed once, drives both fields).
2026-07-15 11:44:32 +02:00
Alem Tuzlak ecb91c8770 fix(channels-intelligence): teams thread history via getHistory + getMessages
getHistory was Slack-only: it keyed off threadTs and returned [] for any route
without one, so a Teams route ({adapter:'teams', tenantId, conversationId})
never hit /api/channels/history. That starved BOTH history paths on Teams --
agent.messages seeding (conversationStore.getOrCreate) and the read_thread tool
-- so the agent always reported "no earlier messages" / "no image in thread".

- getHistory is now adapter-aware (mirrors conversationKeyFromReplyTarget's
  per-adapter switch): Slack keys off teamId/channel/threadTs; Teams sends
  adapter=teams + tenantId + conversationId, matching app-api's
  teams:{tenantId}:{conversationId} thread_key. The /api/channels/history route
  already accepts this shape. Slack query order is unchanged.
- Add getMessages to the adapter so thread.getMessages() (the read_thread tool)
  reads reconstructed history via the transport and maps it to ThreadMessage[].
  Without it Thread.getMessages() returns [] and thread-reading tools see
  nothing even when history exists.

Adds tests for the Teams getHistory query and the missing-id short-circuit.
2026-07-14 19:35:08 +02:00
tylerslaton f4e19e2039 chore: release channels-intelligence v0.1.1 2026-07-10 22:44:50 +00:00
Tyler Slaton 001dda539a chore(channels-intelligence): complete channel terminology sweep 2026-07-10 14:47:06 -07:00
Tyler Slaton 68e43fefe1 refactor(channels-intelligence): rename remaining channel APIs 2026-07-10 14:39:56 -07:00
Tyler Slaton ea9910ae93 refactor(channels-intelligence): introduce realtime gateway abstraction 2026-07-10 14:10:30 -07:00
Tyler Slaton 4dddabd79f test(channels-intelligence): align claim test with provider-agnostic flow 2026-07-10 13:44:11 -07:00
Tyler Slaton 21d3b8c7df Merge remote-tracking branch 'origin/main' into update-intelligence-channels 2026-07-10 13:43:04 -07:00
Tyler Slaton 67fce71690 refactor(channels-intelligence): migrate HTTP contract to channels 2026-07-10 13:25:17 -07:00
Tyler Slaton 5c217538ab fix(channels-intelligence): claim deliveries provider-agnostically (#5914)
## Problem

A managed bot with **both** a Slack and a Teams adapter attached only
ever received its **Slack** deliveries. Teams deliveries stayed `queued`
forever — never claimed, never sent.

## Root cause

`channels-intelligence`'s runtime claim loop (`http-transports.ts` →
`claimOnce()`) posted a per-provider filter to
`/api/bots/listener/claim`:

```ts
{
  runtimeInstanceId: this.cfg.runtimeInstanceId,
  adapters: [this.cfg.adapter], // defaults to "slack"
}
```

app-api filters claimable deliveries by that list (`$adapters IS NULL OR
bie.provider = ANY($adapters)`), so a runtime declaring only `"slack"`
is never handed the same bot's Teams deliveries.

But the managed runtime is **provider-agnostic**: it emits abstract
render frames and Intelligence renders each reply per the delivery's own
reply target. There is no reason for the runtime to constrain claims by
provider — one `intelligenceAdapter()` should serve every channel its
bot has attached.

## Fix

Drop the `adapters` field from the claim body. `adapters` is already
optional on the app-api side (absent → `NULL` → no provider filter → all
providers), so this needs no coordinated backend change.
`this.cfg.adapter` is still used for the heartbeat's declared bots and
for egress, both unaffected.

## Testing

Verified end-to-end locally against a managed Teams bot: inbound Bot
Framework JWT → claim → agent run → render → Bot Connector egress all
`succeed` with this change. Slack continues to work unchanged.
2026-07-10 12:52:32 -07:00
Benjamin Taylor 150164a4bd fix(channels-intelligence): derive conversationKey per provider (Teams-safe)
Follow-up to the provider-agnostic claim change on this branch. Now that the
runtime claims deliveries for every provider its bot has attached, Teams
deliveries flow through the same bridge — and their reply target is a distinct
shape (serviceUrl/conversationId/tenantId, no teamId/channel/threadTs). Deriving
conversationKey from Slack-only fields collapsed every Teams conversation onto
one degenerate key, and conversationKey keys the agent/session
(getOrCreate -> makeAgent), so distinct Teams conversations would share
state/memory.

Make replyTarget a discriminated union (slack|teams) and derive conversationKey
per provider: teams:{tenantId}:{conversationId}, matching Intelligence app-api's
thread_key (OSS-441 slice 2, Intelligence #511) so client and server agree on
conversation identity. Unknown adapters fail loud (the claim loop's existing
catch nacks, not wedges).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-10 14:34:09 -05:00
Benjamin Taylor 57ddcb9532 fix(channels-intelligence): enforce runtimeInstanceId on OnChannel + leak-path tests + fail-loud managed entrypoint (OSS-406 review r3) 2026-07-10 12:31:24 -05:00
Alem Tuzlak 2a16becf61 fix(channels-intelligence): claim deliveries provider-agnostically
The managed runtime's claimOnce() declared `adapters: [this.cfg.adapter]`
(defaulting to "slack"), which Intelligence used to filter claimable
deliveries by provider. A runtime serving a bot with both Slack and Teams
adapters would therefore never receive the bot's Teams deliveries — they
stayed queued forever while Slack worked.

The managed runtime is provider-agnostic: it emits abstract render frames
and Intelligence renders per the delivery's own reply target. So the claim
must not filter by provider. Drop the adapter filter from the claim body;
one config-free `intelligenceAdapter()` now serves every channel its bot
has attached.

Verified end-to-end locally against managed Teams: inbound JWT -> claim ->
agent -> render -> Bot Connector egress all succeed with this change.
2026-07-10 19:06:44 +02:00
github-actions[bot] 2851d83a8e style: auto-fix formatting 2026-07-10 16:49:23 +00:00