## Summary
Renames the **Bots SDK → Channels SDK** (leadership decision, OSS-438).
**Names only — no behavior change.**
| old | new |
|---|---|
| `@copilotkit/bot` | `@copilotkit/channels` |
| `@copilotkit/bot-ui` | `@copilotkit/channels-ui` |
| `@copilotkit/bot-slack` | `@copilotkit/channels-slack` |
| `@copilotkit/bot-teams` | `@copilotkit/channels-teams` |
| `@copilotkit/bot-discord` | `@copilotkit/channels-discord` |
| `@copilotkit/bot-telegram` | `@copilotkit/channels-telegram` |
| `@copilotkit/bot-whatsapp` | `@copilotkit/channels-whatsapp` |
Package dirs renamed (`git mv`); versions carried over.
## What changed
- **Packages:** dir renames, `name` fields, `repository.directory`,
descriptions, and `workspace:` cross-deps rewired in lockstep.
`createBot` and other **API names are unchanged** (out of scope).
- **Release plumbing:** `release.config.json` scope keys +
`versionSource`, `ReleaseScope` union in
`scripts/release/lib/config.ts`, the
`canary`/`stable-release`/`publish-release` workflow scope dropdowns,
and `verify-release-scope-dropdowns.sh`.
- **Consumers:** `examples/slack` (Kite) + `examples/teams` — deps, the
load-bearing `jsxImportSource` pragma, build globs, imports.
- **Docs (`showcase/shell-docs`):** MDX package refs, content dirs
`docs/bots`→`docs/channels` and `reference/bot`→`reference/channels`,
nav registry, and **permanent redirects** from the old `/bots` and
`/reference/bot` URLs.
## Verification
- All 7 `@copilotkit/channels*` packages build; package test suites pass
(143+ in `channels`).
- `examples/slack` + `examples/teams` typecheck clean against the
renamed packages.
- `verify-release-scope-dropdowns.sh` green; release notification
wrapper test 30/30; `release:prepare` dry-runs for `channels` and
`channels-slack` resolve.
- `showcase/shell-docs` `frontend-options` test + full `tsc --noEmit`
clean.
- Zero stray `@copilotkit/bot` / `packages/bot` refs remain.
## ⚠️ Before merge / after merge
- **Do not squash-lose the deprecation step:** after these publish, run
`npm deprecate @copilotkit/bot@"*" "Renamed — install
@copilotkit/channels instead."` for the **5 published** old packages
(`bot`, `bot-ui`, `bot-slack`, `bot-teams`, `bot-discord`).
`bot-telegram`/`bot-whatsapp` were never published.
- New package names have **no npm version history**; the `package.json`
`version` seeds the first publish. Dry-run the OIDC publish for a
never-published scope first.
- Public-surface rename → stakeholder sign-off (kept as draft).
Refs OSS-438
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Two ChunkedEditStream tests waited an arbitrary `setTimeout(r, 10)` for the
internal `setTimeout(0)` + microtask flush chain, then asserted `callCount === 1`.
When CI scheduling jitter/GC ate that <=10ms window before the first `editAt`
ran, `callCount` was still 0 and the test failed with `expected +0 to be 1`.
It surfaced on the Node 20 shard of `test / unit`, while the near-identical
sister test passed in the same run — a classic flaky-timing race, not a Node 20
semantic difference.
Replace the arbitrary sleep with event-based waiting: resolve a promise the
instant the first `editAt` runs, so each test waits for the actual condition it
cares about rather than guessing a duration. Deterministic regardless of
scheduling. Verified with a mutation probe (a slowed flush chain fails the old
10ms-wait pattern but passes the new condition-based wait).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- ack/nack: delete the lease BEFORE the wire call so a background turn that
completes after a timeout-nack can't also ack (single terminal signal).
- targetFromRef: throw a clear, actionable error when a ref carries no delivery
routing instead of coercing deliveryId to the string "undefined".
- getOrCreate: guard getHistory with try/catch so a contract-violating source
can't nack an otherwise-fine turn; document the app-api turnId-distinctness
invariant the in-place card update depends on.
- getHistory: clamp to the requested limit (parity with the in-memory source);
note the Slack-shaped route cast and the thread_started wire asymmetry.
- fetchFile: bound the body read with a generous content-length backstop.
- state store: forward an explicit ttlMs:0 (was dropped by truthiness) and add
the missing TSDoc docblocks.
- supportsBlockingChoice: future-tense the TSDoc (no reader exists yet).
- tests: interaction update-in-place stamping (the HITL crux), content-parts
text/unknown-mime branches, default-store resolution, kv non-2xx + ttlMs:0.
## Summary
Follow-up cleanup/hardening for the two showcase entrypoint watchdogs
fixed in #5874. A 5-round code review deferred a set of non-load-bearing
bucket-(b) polish items to this PR; they are implemented here. Scope is
limited to:
- `showcase/integrations/langgraph-typescript/entrypoint.sh`
- `showcase/integrations/strands-typescript/entrypoint.sh`
Shared helper code (`_require_int`, `_agent_descendants`) is kept
byte-identical across both files. `bash -n` and `shellcheck
--severity=warning` are clean on both.
## Items implemented
### 1. `_require_int` upper bound (behavioral, both files)
The validator accepted any positive integer (`[1-9][0-9]*`), so a 20+
digit override overflowed bash's signed-64-bit arithmetic and either
wrapped to garbage or aborted the `[ -ge ]` test — which, suppressed to
false inside the guard's `if`, **silently disabled the guard** (the
exact fail-open class the validator exists to prevent). Added a 10-digit
length cap (max 9,999,999,999 — far above any real
interval/threshold/strike knob, comfortably inside int64), checked
before the digit `case` since an all-digit 23-char value would otherwise
pass. A too-long value now takes the same WARN + fall-back-to-default
fail-safe path as every other bad override.
**Cap rationale:** a single generous 10-digit cap rather than per-knob
caps — 9,999,999,999 seconds is ~317 years and 9,999,999,999 MB is ~9.3
PB, so no legitimate value is ever excluded, and it leaves 9 digits of
headroom below the 19-digit int64 ceiling.
### 2. `cleanup()` comment fix (langgraph, non-behavioral)
Corrected the note claiming `WATCHDOG_PID` "forks nothing that outlives
it" — it **does** fork the size sub-loop; the bare `kill $WATCHDOG_PID`
is safe because the watchdog's own inner EXIT trap reaps that child, not
because it forks nothing. The strands `cleanup()` comment was verified
accurate (strands has no size sub-loop) and left unchanged.
### 3. SIZE_PID trap-ordering leak window (langgraph, behavioral)
The size sub-loop was backgrounded (`( … ) &`, `SIZE_PID=$!`) **before**
its reaping `trap … EXIT` was registered, so an outer SIGTERM landing in
that window exited the watchdog subshell with no trap armed and orphaned
the sub-loop (reparented to PID 1, spinning for the container's life).
Now the reaping trap is armed **first** and reaps via a `$BASHPID`
PPID-walk that finds the child regardless of whether `SIZE_PID` is
assigned yet — no ordering-dependent leak.
### 4. Size-guard coverage during startup grace (langgraph, behavioral)
The size monitor started only **after** the up-to-180s startup-grace
loop, leaving the size ceiling unguarded during a pathological cold
boot. It now starts **before** the grace loop. Decision: starting early
is safe because `_watchdog_check_size_once` already fail-closes on every
not-yet-ready condition (agent PID not alive, `PERSIST_DIR`
missing/freshly-purged, non-numeric size or threshold), so early cycles
are harmless no-ops until the dir actually grows. No documented-gap
fallback was needed.
### 5. Cosmetic (both files, non-behavioral)
- **Done:** `_agent_descendants` per-PID `echo | awk` fork replaced with
the `read` builtin (clean drop-in; avoids forking awk once per `/proc`
entry). Byte-identical across both files.
- **Skipped — strike-window "off-by-one" log:** examined the
health-strike log `~$((HEALTH_CHECK_INTERVAL * HEALTH_STRIKE_LIMIT))s`;
with interval=30, limit=3 the third failure lands at ~90s and the
product is 90 — there is no actual off-by-one to fix, so touching it
would add churn without value.
- **Skipped — unthrottled transient-error re-warn:** throttling the
per-cycle transient `[watchdog:size]` warning requires adding per-loop
state/timestamp bookkeeping; low value and would balloon the diff into
the hot loop, so skipped per the "skip if not clean/low-risk"
instruction.
## RED → GREEN (behavioral items 1, 3, 4)
Proven on the **real committed entrypoint bytes** in `node:22-slim`
(Debian bash 5.x, real `/proc`), mirroring how #5874's fixes were
proven. RED = the merged #5874 baseline (`7912b6e51`); GREEN = this
branch's HEAD.
### Item 1 — overflow clamp (both files)
```
RED langgraph: input HEALTH_CHECK_INTERVAL='99999999999999999999999' (23 digits)
after _require_int, HEALTH_CHECK_INTERVAL='99999999999999999999999'
$(( HEALTH_CHECK_INTERVAL * 3 )) = 601129261562068989
RESULT: 23-digit value SURVIVED validation <-- BUG: no upper bound
RED strands: (identical to langgraph — byte-identical helper)
GREEN langgraph: [entrypoint] WARNING: health interval (HEALTH_CHECK_INTERVAL) is too large (got: '99999999999999999999999', 23 digits — max 10) — falling back to default 30
after _require_int, HEALTH_CHECK_INTERVAL='30'
$(( HEALTH_CHECK_INTERVAL * 3 )) = 90
RESULT: value was clamped to '30' (default) — guard safe <-- FIXED
GREEN strands: (identical — clamped to 30, arithmetic = 90)
```
### Item 3 — trap-ordering leak window (langgraph)
```
RED post-fix ordering emulated with the pre-fix spawn-then-trap sequence:
size sub-loop ticks AFTER watchdog exit = 2
RESULT: sub-loop ORPHANED — still ticking after watchdog gone <-- BUG (leak window)
GREEN size sub-loop ticks AFTER watchdog exit = 0
RESULT: sub-loop REAPED on watchdog exit — no orphan <-- FIXED
```
### Item 4 — size-guard coverage during startup grace (langgraph)
```
RED grace-loop announce at line 437; size-monitor spawn at line 459
RESULT: size monitor starts AFTER grace loop — UNGUARDED during up-to-180s grace <-- BUG
GREEN grace-loop announce at line 533; size-monitor spawn at line 507
RESULT: size monitor starts BEFORE grace loop — guarded during cold start <-- FIXED
```
Functional confirmation (real watchdog block, short grace, stubbed
size-check + never-healthy probe):
```
[watchdog] Startup grace: waiting up to 5s for first successful health probe before arming strike counter
[watchdog:size] Starting size-gated restart monitor (threshold=200MB, interval=1s, dir=/tmp/persistX)
[GRACE-WINDOW-SIZE-CHECK-RAN]
[GRACE-WINDOW-SIZE-CHECK-RAN]
--- 3s elapsed (still within 5s grace) ---
```
`--check-size-once` seam re-verified post-refactor: under threshold →
exit 0 (no kill); over threshold → tree-kill + exit 1.
## Lint
`bash -n` and `shellcheck --severity=warning` clean on both entrypoints.
## Deferred (bucket (d)) — known future work, NOT in this PR
- A dedicated Next.js frontend watchdog (health probe + tree-kill for
the `NEXTJS_PID` process-sub subshell, symmetric to the agent watchdog).
- A strands size-guard / boot-purge (strands currently has only the
health watchdog; no persistence-size ceiling or boot purge).
These are separate features, out of scope for this cleanup PR.
The exhaustiveness guard added to mapDeliveryToEnvelope threw on an unknown
wire kind, but claimOnce records the lease *before* mapping and runLoop's outer
catch only logs+sleeps — so an unmappable delivery leaked its lease and, after
the 120s lease expiry, was redelivered and re-thrown forever, permanently
wedging the single-delivery loop. Catch the map failure in claimOnce and nack
it non-retryably (deterministic failure → let app-api dead-letter it) so the
guard fails loud without blocking the queue. Adds a retryable flag to nack and
a regression test.
The launch-site comment claimed process substitution leaves $! pointing at
the real node process. That is false and contradicts the file header: $!/
AGENT_PID is the wrapping subshell, and the real npm->node server is a
descendant reached only via the tree-kill (the reason _kill_agent_tree exists).
Rewritten in both langgraph-typescript and strands-typescript entrypoints so no
maintainer reintroduces a bare kill.
Also in langgraph-typescript: clarify that the SIZE_PID kill is a retained
belt-and-suspenders backstop to the $BASHPID PPID-walk (not dead code), and
re-anchor the startup-grace rationale to the real cause (the top-level
@langchain/langgraph-api import cost), since the prod path no longer uses
langgraph-cli dev.
Comment-only; no executable code changed.
The bounded re-scan loop's sleep 0.2 was unguarded. Under set -e on a base image
whose sleep can return non-zero (e.g. a future busybox/Alpine rebase), a failed
sleep would abort the tree-kill mid-walk — root never killed, real npm→node
server left orphaned. Add || true so the walk completes regardless of sleep's
exit status. No behavior change on the current Debian base (coreutils sleep
succeeds). Helper kept byte-identical across both entrypoints.
Both langgraph-typescript and strands-typescript entrypoints inferred which
process exited via a post-hoc kill -0 if/elif after wait -n. That inference is
racy: on a near-simultaneous exit both PIDs are dead by probe time, so the first
kill -0 branch always wins and mislabels the diagnostic (naming the agent when
Next.js actually exited, attaching the wrong code to the wrong name). Use
bash's wait -n -p REAPED_PID (bash >= 5.1; node:22-slim ships 5.2) to capture
the actual reaped PID and key the message off it. Exit code (incl. 137) and the
final exit $EXIT_CODE are preserved; || EXIT_CODE=$? guard is unchanged.
Three related size-watchdog hardening changes in the langgraph entrypoint:
- Trap-order leak window: the size sub-loop was backgrounded (`( … ) &`,
SIZE_PID=$!) BEFORE its reaping `trap … EXIT` was registered, so an outer
SIGTERM landing in that window exited the watchdog subshell with no trap
armed and orphaned the sub-loop (reparented to PID 1, spinning for the
container's life). Arm the reaping trap FIRST, and reap via a $BASHPID
PPID-walk that finds the child regardless of whether SIZE_PID is assigned
yet — no ordering-dependent leak.
- Startup-grace size coverage: the size monitor started only AFTER the
up-to-180s startup-grace loop, leaving the size ceiling unguarded during a
pathological cold boot. Start it BEFORE the grace loop; safe because
_watchdog_check_size_once already fail-closes on every not-yet-ready
condition (agent PID not alive, PERSIST_DIR missing, non-numeric size/
threshold), so early cycles are harmless no-ops until the dir grows.
- cleanup() comment accuracy: corrected the note claiming WATCHDOG_PID
"forks nothing that outlives it" — it DOES fork the size sub-loop; the
bare `kill $WATCHDOG_PID` is safe because the watchdog's own inner EXIT
trap reaps that child, not because it forks nothing.
Proven RED->GREEN on the real entrypoint in node:22-slim: pre-fix the size
sub-loop keeps ticking after the watchdog exits (orphan) and the size monitor
spawns after the grace loop (unguarded); post-fix the sub-loop is reaped (0
ticks) and the monitor runs during grace (size-check fires within the grace
window). --check-size-once seam re-verified under/over threshold.
The numeric-config validator accepted any positive integer (`[1-9][0-9]*`),
so a 20+ digit override overflowed bash's signed-64-bit arithmetic and either
wrapped to a negative/garbage magnitude or aborted the `[ -ge ]` test with
"value too great for base" — which, suppressed to false inside the guard's
`if`, silently disabled the guard for the container's lifetime (the exact
fail-open class this validator exists to prevent).
Add a 10-digit length cap (max 9,999,999,999 — comfortably inside int64,
far above any real interval/threshold/strike knob) checked BEFORE the digit
`case`, since an all-digit 23-char value would otherwise pass validation.
A too-long value now takes the same WARN + fall-back-to-default fail-safe
path as every other bad override. Byte-identical across both entrypoints.
Proven RED->GREEN on the real entrypoints in node:22-slim: pre-fix a 23-digit
value survives validation and `$(( x * 3 ))` yields int64-wrapped garbage;
post-fix it WARNs, clamps to the default, and arithmetic is correct.
The /proc PPID walk forked an awk process for every entry in the process
table on every scan pass. Replace the `echo "${stat##*) }" | awk '{print $2}'`
pipeline with the `read` builtin, which word-splits the post-comm remainder
("STATE PPID PGRP …") on IFS and captures the 2nd field with no subprocess.
Byte-identical across the langgraph-typescript and strands-typescript
entrypoints. Non-behavioral; bash -n + shellcheck --severity=warning clean.
An unknown delivery kind previously fell through the dispatch switch as a
silent no-op (dispatch acks on resolve, so it would ack an unhandled delivery
as processed) and mapDeliveryToEnvelope coerced it into an empty turn. Add
matching `never` exhaustiveness guards to both so a future wire kind throws
instead of being silently swallowed.
Renames the Bots SDK to the Channels SDK. Names only — no behavior change.
- 8 packages @copilotkit/bot* -> @copilotkit/channels* (git mv dirs, names,
workspace: cross-deps). Now includes @copilotkit/bot-intelligence ->
@copilotkit/channels-intelligence (landed on main via #5761; unpublished, so
renamed fresh with the family).
- release.config.json scope keys + versionSource; ReleaseScope union;
canary/stable-release/publish-release scope dropdowns; verify script
- examples/slack (Kite) + examples/teams: deps, jsxImportSource, imports
- showcase/shell-docs: content dirs docs/bots->docs/channels and
reference/bot->reference/channels, nav registry, redirects
createBot and other API names unchanged. Old @copilotkit/bot* to be deprecated
after the new packages publish (bot-intelligence was never published).
Re-derived onto latest main (was conflicting after #5761 landed).
Refs OSS-438
## What & why
Two showcase agent containers (**langgraph-typescript**,
**strands-typescript**) could enter a *running-but-dead* state: Railway
showed the service `● Online` while `/api/health` returned **HTTP 502**.
This took down all 36 LGT dashboard cells on staging **and** prod
(prod's `.langgraph_api` had crossed the 200 MB size-watchdog
threshold).
**Root cause:** the agent is launched via process substitution (`... &>
>(awk …) &`), so `$AGENT_PID` (`=$!`) is the **wrapper subshell**, not
the real `npm`→`node` server that holds the port. Every watchdog/cleanup
did a bare `kill -9 $AGENT_PID`, which reaped only the subshell and
**orphaned the real server** (reparented to PID 1, still bound to the
port). The watchdog's "kill agent → container restart → boot-purge"
contract therefore never fired: the frontend kept proxying to a dead
agent → 502 forever.
## Fixes (each with local red-green on the real entrypoint in
`node:22-slim`)
1. **cleanup() EXIT trap** → routes through `_kill_agent_tree` (was
orphaning the agent on every SIGTERM/redeploy).
2. **`_kill_agent_tree`** → `/proc`-based tree-kill with a bounded
re-scan (root killed last) so mid-walk forks can't escape; refuses PID ≤
1 (fail-closed).
3. **size-watchdog** hardened against non-numeric `du` and transient
errors (no silent gate-disable, no permanent loop death).
4. **strands health-watchdog** → 180 s startup-grace window (parity with
langgraph); the now-effective kill would otherwise loop a slow cold
start.
5. **`wait -n` under `set -e`** → capture exit code so the restart
diagnostic isn't dead code on the primary (137) path.
6. **structural:** one `_require_int` validator over *every*
operator-overridable numeric knob (fail-safe to default), and **every**
wrapped-PID kill (incl. `NEXTJS_PID`) routed through the guarded
tree-kill; dangerous `${AGENT_PID:-0}` sentinel removed.
7. **`_require_int`** requires a positive integer (rejects `0` and
leading-zero/octal).
## Incident status
Staging **and** prod LGT were restored immediately via redeploy
(boot-purge cleared the oversized state) — both `/api/health` → 200.
This PR stops the recurrence.
## Review
Converged through a 5-round unbiased review-fix loop (1 + 4
confirmation), zero mandatory findings at close, all load-bearing guards
independently re-verified. `bash -n` + shellcheck (`-S warning`) clean;
170/170 shell bats pass.
## Follow-up (tracked, separate PR — non-load-bearing)
`_require_int` upper-bound clamp (LOW arith-overflow, needs a 20+-digit
value); a stale `cleanup()` comment; size-guard unarmed during the
startup-grace window; SIZE_PID trap-registration micro-window;
diagnostic label on near-simultaneous exit; cosmetic log nits; startup
readiness `sleep 3`+`kill -0` probes the wrapper subshell; no dedicated
Next.js frontend watchdog.
- Drop the resurrected inline buildContentParts method from the adapter; the
turn path now uses the extracted ./content-parts.js helper (#5814's refactor).
- Port main's fail-visible file-fetch behavior into content-parts.ts so a file
that can't be retrieved becomes a short note in both the live-turn and
history-seeding paths (they share the helper).
- Update the state-store conformance import to the @copilotkit/bot/testing
subpath (main moved it there to keep vitest out of consumers' runtime graph).
## Summary
Lets the `@copilotkit/bot` SDK run from **Intelligence-delivered
events** without a second programming model, and adds the runtime `bots`
declaration API. A managed event (delivered by Intelligence) runs the
*same* customer handlers, tools, context, commands, Bot UI, and agents
as local/custom adapters — the managed path is "just another
`PlatformAdapter`," fed by injected transports.
This is the **OSS / SDK slice** of the Hosted Managed Bots work. The
credentialed transports (Realtime Gateway, Connector Outbox) and the
frozen shared contracts live elsewhere (see *Out of scope*); this PR
ships the seams they plug into, fully runnable headless.
Relates to **OSS-360** (runtime bots API), **OSS-361** (run the SDK from
Intelligence events), **OSS-363** (Slack render/codec reuse).
## What's in here
- **`intelligenceAdapter()` bridge** (`@internal`, not publicly
documented) — implements `PlatformAdapter` over two injected transports:
`DeliverySource` (inbound) + `EgressSink` (outbound). Ingress →
`onTurn`/`onCommand`/`onInteraction`/`onThreadStarted`/`onReaction`; ack
on success / nack on throw (at-least-once). Egress emits generic
operations carrying `BotNode[]` IR with **deterministic ids**
(`turnId:seq`, reset per turn) so a redelivered turn reproduces the same
ids for the Connector Outbox to dedupe. Idempotency lives at egress, so
the managed path skips ingress dedup (`skipIngressDedup`) — a redelivery
re-runs rather than being dropped.
- **Runtime `bots` API** — `new CopilotRuntime({ intelligence, bots })`,
accepted by TypeScript **only when `intelligence` is configured**
(discriminated union). `createBot({ name })`; `startManagedBots()`
validates names (required, identifier-style, unique — fail-loud), builds
activation metadata, and wires each bot to its resolved transport.
- **`PlatformCodec` seam** + Slack egress codec (`slackCodec`) composing
the existing pure `renderSlackMessage`, so IR→native rendering is shared
(no Bolt/creds) instead of duplicated.
- **Backwards-compatible SDK foundations**: `bot.addAdapter()` +
optional `adapters`, deferred backend resolution at `start()` with
`stateStore`-provider precedence (+ multi-provider warning),
`bot.transcripts` throws pre-start, optional
`eventId`/`turnId`/`deliveryId` on ingress + handler context. Existing
`createBot` callers and every `PlatformAdapter` implementer are
unaffected.
- **In-memory transports + fixture tests** — the full dispatch path
(envelope in → handler runs → egress op out) runs with zero
Slack/Intelligence/network.
## Out of scope (external / separate tickets)
- **Realtime Gateway + Connector Outbox transports** — implemented in
the closed-source repo against the `DeliverySource`/`EgressSink`
interfaces shipped here.
- **Shared contracts freeze (OSS-377)** — consumed here via a minimal,
isolated placeholder (`managed/contracts.ts`, marked `TODO(OSS-377)`);
swaps in via one import change.
- **OSS-363 ingress normalization** — the egress codec is done;
extracting the pure Slack event→neutral mapping out of the Bolt listener
(so local + Intelligence ingress share it) is the remaining, higher-risk
half and is left to that ticket (`TODO(OSS-363)`).
## Testing
TDD throughout (RED→GREEN per behavior). New: managed adapter
dispatch/ack-nack/ids/run-renderer/exclusivity, all-kinds routing, name
validation + metadata + lifecycle, runtime `bots` option, Slack codec.
Full suites green: `bot` 147, `bot-slack` 256, `runtime` 1574. All
builds typecheck (`bot`/`bot-slack`/`bot-discord`/`runtime`);
oxlint/oxfmt clean.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
## What
Applies the threads-drawer grid fix to all 15 integration examples: adds
`grid-template-rows: minmax(0, 1fr)` to each example's `.layout` so the
drawer's threads list scrolls **internally** (pinned header + New
Conversation) instead of the whole page growing to content height.
Without this, `height: 100dvh` on a grid whose rows aren't bounded lets
the row size to content, so the list can't scroll within the drawer and
the delete-confirm dialog centers against a content-tall root.
## Why this is a separate PR
These are **example** changes that consume **published** `@copilotkit/*`
packages. They were pulled out of the drawer-redesign PR (#5823) so that
PR stays scoped to the packages. This one lands **after** the redesigned
drawer is released.
## ⛔ Blocked / TODO before marking ready
- [ ] Drawer redesign PR #5823 merged
- [ ] Lockstep release cut (web-components + react-core + vue + angular)
- [ ] Bump each example's `@copilotkit/*` dependency to the new
published versions **in this PR** (currently only the CSS is here)
- [ ] Re-verify one example end-to-end against the released packages
## Testing
- Grid fix verified live during the redesign work (langgraph-js against
a hosted Intelligence backend): threads list scrolls internally, header
+ New Conversation pinned, delete-confirm renders correctly.
- Example dependency bumps + a fresh end-to-end pass will be
added/redone here once the release is published.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
The 1.62.3 release publishes the CopilotThreadsDrawer redesign (web-components +
react-core wrapper) and the stateless /suggest feature. Bump the 15 integration
examples that consume the drawer from 1.62.2 -> 1.62.3 (package.json + lockfiles)
so they pick up the released packages alongside this branch's example CSS.
Validated: langgraph-js runs on the published 1.62.3 (no local links) — the
redesigned drawer renders (New Conversation, Recent Conversations, filter funnel,
desktop collapse toggle, per-row kebab), threads are licensed, and a real agent
message round-trips.
- ModeToggle: one style on both breakpoints (top-4/right-4 = 16px gutter,
46px min-height, 4px corners); symmetric p-1.5 + fixed 20px button leading
so the selected pill has an even gap on all four sides (was tight L/R vs T/B).
- Launcher: uniform 16px gutter (top + left) on both breakpoints so it mirrors
the toggle; drop the mobile-only 7px override.
- Logo: centered on the launcher/toggle middle line (pt-[23px]); wordmark
padding normalized so its height matches on both breakpoints.
- Inspector FAB: sits beneath the toggle, gap = the 16px top gutter (one rule,
no media query, since the toggle is identical across breakpoints).
Net: launcher, logo, toggle share center-y; launcher + toggle are both 46px;
the FAB tucks under the toggle with a matching gap; the selected toggle pill is
evenly inset.
- ModeToggle: move left (right-[72px]) so the top-right inspector FAB no longer
covers the App segment; grow to 46px (lg:min-h) + center on the logo line
(top-6) to match the launcher; keep the 4px corners.
- Launcher: left gutter -> 16px to match the right-side controls' inset.
- Logo: pt-7 so it centers on the same line as the launcher + toggle.
- Mobile header: max-lg:pb-0 -> pb-4 so chat content clears the fixed launcher/
toggle strip instead of butting right under it (no boundary).
- Chat/App ModeToggle: rounded-full -> rounded-[4px] container + rounded-[2px]
buttons, matching the drawer's 4px radius cap so the header controls are
visually consistent.
The 7px/16px launcher inset was tuned for the mobile off-canvas launcher; on
desktop it leaked onto the collapsed cluster. Move it into the mobile media
query so desktop-collapse uses the element's own 24px gutter default.
The floating launcher/collapsed cluster is fixed at the top-left corner. Below
1024px it always shows (already cleared via max-lg:pl-24); on desktop it appears
only when the drawer is COLLAPSED. Drive the header's left padding off
--cpk-drawer-reserved-width (0px when collapsed, 320px default otherwise) so the
logo starts at ~6rem when collapsed and pl-6 when expanded — no overlap. No-op
on current packages (var never set → stays pl-6).
Read grid-template-columns' first track from var(--cpk-drawer-reserved-width, 320px)
so when the drawer collapses on desktop (it sets the var to 0) the reserved
column collapses and the chat reclaims the space — instead of leaving an empty
placeholder column. Mobile (single-column) is unchanged.
Correctness:
- C1 create-bot: start() is now idempotent — a second start() no longer
re-resolves the backend / rebuilds Transcripts+Telemetry+ActionRegistry or
re-connects adapters (which would wipe MemoryStore state and double-bind real
adapters). stop() clears the flag so start→stop→start is still a real restart.
- S1 bot-slack ingress: a threaded reply that @-mentions the bot is now skipped
(app_mention handles it) so the managed path no longer double-responds. Matches
both the plain <@U…> and labeled <@U…|handle> mention forms.
- S2 runtime: CopilotSseRuntime throws if `bots` is passed without intelligence
instead of silently dropping them (guards a JS/as-any caller past the type).
- S3 bot-intelligence: startManagedBots rolls back — stops already-started bots —
when a later bot fails to start, instead of leaking listeners/connections.
Lower:
- S4 ingress: stripMentions handles the labeled <@U…|handle> form; DM turns strip
mentions too (parity with app_mention/thread_reply).
- S5 bot-intelligence: bot-name uniqueness is now case-insensitive.
- S6 runtime: fail fast at construction when a declared bot has no name (full
shape/uniqueness validation stays at the activation seam — assertValidBotNames —
because it can't cross into this CJS package from pure-ESM bot-intelligence).
- S7 bot-intelligence: buildActivationMetadata throws on a nameless bot instead of
silently filtering it out of the activation set.
- S8 bot-intelligence: startManagedBots warns on an empty bots array.
- M1 intelligence-adapter: the per-turn egress seq Map entry is deleted after each
turn so it can't grow unbounded over a long-running bot.
- M2 intelligence-adapter: an inbound file that fails to fetch degrades to a
fail-visible text note instead of being silently dropped from model context.
- I2 contracts: dropped the now-dead `duplicate_skipped` RenderAccepted value
(Intelligence returns duplicate_accepted or a 409 conflict).
Changelog (C2/C3, intended behavior after moving init into start()):
- bot.transcripts now throws before start() (was a concrete property).
- telemetry `oss.bot.configured` now fires at start() rather than construction, so
a constructed-but-never-started bot no longer emits it.
Not addressed here (cross-repo, tracked on the Intelligence side):
- I1 realtime render-event kind:"file" clause on the gateway validator.
- I3 lease-token fencing on the render-accept path.
The _require_int validator in the langgraph-typescript and strands-typescript
entrypoints accepted '0' and leading-zero/octal forms like '010'/'08'. Operator
typos on any numeric knob then broke a guard:
- SIZE_THRESHOLD_MB=0 kills the agent on cycle 1 (instant restart loop)
- HEALTH_STRIKE_LIMIT=0 kills on first probe miss
- SIZE_CHECK_INTERVAL=0 / HEALTH_CHECK_INTERVAL=0 busy-spin on 'while sleep 0'
- '010' is read as OCTAL (8) in arithmetic; '08'/'09' abort under set -e
Tighten the predicate to accept only a positive integer with no leading zero
([1-9][0-9]*). Invalid values keep the existing fail-safe behavior: WARN and
fall back to the documented default. Helper stays byte-identical across both
files.
CLASS 1 (guard silently disabled by a bad numeric override): add a reusable
_require_int validator and run it at startup over EVERY operator-overridable
numeric knob in both entrypoints (size threshold/interval, startup grace,
health-probe interval, strike limit). A non-integer/empty override now WARNs
and falls back to the documented default instead of breaking a sleep/loop/
arithmetic test. Closes instance #3 (LANGGRAPH_SIZE_CHECK_INTERVAL='60s'
killing the size-monitor loop on its first iteration).
CLASS 2 (wrapped-PID orphan + kill-0 footgun): route the cleanup() NEXTJS_PID
kill through _kill_agent_tree (it is process-sub-wrapped like the agent, so a
bare kill orphaned the real Next.js node server holding $PORT across redeploy).
Harden _kill_agent_tree and _agent_descendants to refuse a PID that is empty,
non-numeric, 0, or 1 (fail closed), making kill -9 0 / kill -9 1 structurally
impossible. Remove the ${AGENT_PID:-0} sentinel in the --check-size-once seam;
skip with a warning when AGENT_PID is unset instead of defaulting to 0.
Shared helper code kept byte-identical between the two entrypoints.
## Summary
- Backport the website messaging from CopilotKit/website#398 into the
shell-docs Slack and Microsoft Teams frontend pages.
- Replace the stale waitlist/managed-only framing with "get early
access" copy that presents CopilotKit Enterprise Intelligence as the
self-hosted or cloud-hosted production layer around the open source Bot
SDK.
- Frame Slack and Teams as frontends for agents built on any harness or
framework, while reserving production-layer terminology for CopilotKit
Enterprise Intelligence.
- Address browser review annotations on both pages: remove filler in the
opener, avoid setup-heavy lead copy, use "open source" without a hyphen,
add the full CopilotKit Enterprise Intelligence name to the CTA titles,
and keep CTA telemetry surfaces intact.
- Update the shell-docs nav test expectation so it matches the current
root IA, where Threads lives under Build Chat UIs rather than the
generated Intelligence Platform section.
## Validation
- `npm run lint` from `showcase/shell-docs` (passes with existing
warnings)
- `npm run typecheck` from `showcase/shell-docs`
- `npm run test` from `showcase/shell-docs`
- `npm run build` from `showcase/shell-docs`
## Notes
- Hydrated Git LFS assets locally with `git lfs pull` so the shell-docs
public asset tests could read real PNG bytes.
## What
Adds `examples/showcases/banking/stop-demo.sh` — the teardown companion
to the existing `run-demo.sh`.
## Why
`run-demo.sh` detaches everything except the Next.js dev server:
- `docker compose up -d --wait` (detached stack)
- native Metal TEI via `nohup … & disown` (Apple Silicon only)
- `exec pnpm dev` (the only foreground process)
So Ctrl-C stops *only* the dev server and silently leaves the docker
stack (`banking-memory`) and the host embedder on `:7067` running. There
was no one-command way to bring those down. This script fills that gap
and mirrors `run-demo.sh`'s conventions (same `say`/`ok` helpers, same
header-comment style, idempotent).
## What it does
Tears down, idempotently, in order:
1. Next.js dev server on `:3000` (defensive — usually already gone via
Ctrl-C)
2. docker compose stack (project `banking-memory`), **containers only**
by default so a re-run of `run-demo.sh` reuses the built composite image
+ seeded Postgres
3. native Metal TEI on `:7067` (Apple Silicon; the host process docker
doesn't manage) — SIGTERM, then SIGKILL for anything that ignores it
## Flags
- `--purge` — also delete the docker volumes (postgres/redis/minio/tei
model cache) for a full clean-slate reset
- `--keep-tei` — leave the slow-to-warm native embedder running when
only bouncing the stack
## Testing
- `bash -n stop-demo.sh` — syntax clean
- `shellcheck stop-demo.sh` — clean, no warnings
- `./stop-demo.sh --help` renders the banner correctly
🤖 Generated with [Claude Code](https://claude.com/claude-code)
## Problem
`@copilotkit/bot`'s package entry re-exports `runStateStoreConformance`
from `./testing/state-store-conformance`, which does `import { describe,
it, expect, ... } from "vitest"` at module top-level. In ESM a static
re-export **eagerly evaluates** the re-exported module, so a plain:
```ts
import { createBot } from "@copilotkit/bot";
```
drags `vitest` into the consumer's runtime module graph and throws
`ERR_MODULE_NOT_FOUND: Cannot find package 'vitest'` for any consumer
that doesn't have vitest installed (i.e. every production consumer).
`vitest` is only a devDependency.
Surfaced while smoke-testing the package rename (OSS-438) — but it's a
pre-existing bug on `main`, independent of that rename.
## Fix
- **Drop the re-export from `src/index.ts`** → the package entry is now
vitest-free.
- **Publish the helper under a `./testing` subpath**
(`@copilotkit/bot/testing`) — test tooling lives off the runtime entry,
the standard pattern.
- **Declare `vitest` as an optional `peerDependency`** so consumers of
`/testing` get the right signal.
- **Docs** updated to `import { runStateStoreConformance } from
"@copilotkit/bot/testing"`.
Only the import path of the test-only conformance helper changes; the
runtime API is untouched.
## Verification
- Static import trace: the entry graph is 12 runtime modules, **none**
import vitest; every vitest importer is a `.test.js` (not in the graph)
or `testing/state-store-conformance.js` (only reachable via `/testing`).
- `@copilotkit/bot` builds; **143/143** tests pass; `publint` + `attw`
clean (the internal conformance test imports the helper by relative
path, unaffected).
## Coordination
Touches `packages/bot` on `main`. The OSS-438 rename PR (#5849) renames
this package to `@copilotkit/channels`; that PR re-derives from `main`
before merge, so it will absorb this fix automatically. If #5849 merges
first, this rebases onto `packages/channels` mechanically.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
run-demo.sh detaches everything except the Next.js dev server (docker
compose up -d, native Metal TEI via nohup/disown, then exec pnpm dev), so
Ctrl-C on the dev server leaves the docker stack and the host embedder
running. stop-demo.sh brings those leftovers down in one command.
Tears down, idempotently:
- the Next.js dev server on :3000 (defensive; usually gone via Ctrl-C)
- the docker compose stack (project banking-memory), containers only by
default so a re-run reuses the built image + seeded data
- the native Metal TEI on :7067 (Apple Silicon; the host process docker
doesn't manage), SIGTERM then SIGKILL
Flags: --purge also drops volumes for a clean slate; --keep-tei leaves the
slow-to-warm embedder running when only bouncing the stack.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The package entry (@copilotkit/bot) re-exported runStateStoreConformance from
./testing/state-store-conformance, which imports vitest at module top-level.
An ESM re-export eagerly evaluates that module, so a bare
`import { createBot } from "@copilotkit/bot"` dragged vitest into every
consumer's runtime graph and threw ERR_MODULE_NOT_FOUND when vitest wasn't
installed (i.e. any production consumer).
- Drop the re-export from src/index.ts (entry is now vitest-free)
- Publish the conformance helper under the ./testing export subpath
- Declare vitest as an optional peerDependency (documents the /testing need)
- Update docs to import from @copilotkit/bot/testing
Names/behavior of the runtime API are unchanged; only the import path for the
test-only conformance helper moves.
## Release monorepo v1.62.3
**Scope:** `monorepo` | **Bump:** `patch`
---
### How this release process works
1. **This PR was created automatically** by the "release / create-pr"
workflow.
It bumped the `monorepo` packages to `1.62.3`
and generated AI-enhanced release notes.
2. **CI runs on this PR** — the full test suite (unit tests, lint, type
checks, build)
must pass before merging. This is the review gate.
3. **Review the release notes** in `release-notes.md` in this PR.
If a Notion draft was created, you can edit the release notes there
before merging.
4. **When this PR is merged**, the `release / publish` workflow
automatically:
- Builds all packages
- Publishes the `monorepo` packages to npm at version `1.62.3`
- Creates git tag `monorepo/v1.62.3`
- Creates a GitHub Release with the final release notes
### Before merging
- [ ] CI is green (tests, lint, types, build)
- [ ] Version bumps look correct
- [ ] Release notes are accurate (edit in Notion if a draft was created)
---
> **Do not merge until CI is fully green.** The full test suite runs
automatically on this PR.
Both entrypoints run under set -e. The tail `wait -n $AGENT_PID $NEXTJS_PID`
returns non-zero on the PRIMARY designed exit path (137 = size-gate/watchdog
SIGKILL of the agent tree, or an agent crash), so set -e aborted the script AT
that line — making EXIT_CODE=$?, the entire 'which process exited with code N'
diagnostic, and the final `exit $EXIT_CODE` dead code on exactly the
interesting exits. Capture the code with `EXIT_CODE=0; wait -n ... || EXIT_CODE=$?`
so the diagnostic and explicit exit run and preserve the exact code (incl. 137);
the container-restart path is unchanged.
Same class: langgraph's LANGGRAPH_SIZE_THRESHOLD_MB was used in
`[ "$DIR_SIZE_MB" -ge "$threshold" ]` with no numericity guard, so a
non-integer operator override made the test error and silently no-op the size
gate every cycle. Validate the threshold the same way DIR_SIZE_MB already is
(numeric case guard + 'size guard inactive' WARNING, then skip safely).
Two pre-existing test failures on this branch, surfaced by CI's unit +
check-types jobs once main was merged:
- bot-slack event-renderer: the non-pane thread tool-call test still
asserted the old "no composer status" behavior. Commit 13248dda0b
deliberately drove setStatus on ANY thread anchor (not just panes), so
the test now expects both the 🔧 row and the "is using…" status.
- runtime in-memory-runner: HangingAgent/AbortableAgent extended
AbstractAgent but omitted the abstract run() member (@ag-ui/client
0.0.57), failing tsc on the test tsconfig (TS2515). Add the same
run() => EMPTY stub the sibling test agents use.
The health-watchdog armed its 3-strike/~90s kill counter immediately with no
startup-grace window. langgraph-typescript has a 180s grace precisely to keep a
slow cold start from being killed mid-boot into a restart loop. Now that the
tree-kill makes the strands kill effective (the orphan bug previously made it
cosmetic), a slow tsx cold start (>90s) would be genuinely killed and loop.
Port langgraph's grace mechanism verbatim (GRACE=180, no env override, poll
every 5s, exit 0 on agent death during startup, arm anyway if grace elapses),
adapted to strands' :8000/health probe.
Two tightly-related defects in the size-gated restart machinery in
entrypoint.sh:
1. _watchdog_check_size_once validated the du/awk result only for
emptiness, not numericity. A non-integer value (junk du output, a
transient read error, a test-seam stub) reached the
`[ "$DIR_SIZE_MB" -ge ... ]` comparison and threw "integer expression
expected"; sitting inside an `if`, set -e was suppressed so the test
evaluated false and the size gate was SILENTLY skipped with no
warning (unlike the empty-string branch). Now match ^[0-9]+$ via a
case and emit the same "size guard inactive" WARNING, so the gate can
never silently disappear.
2. The size sub-loop used `_watchdog_check_size_once || break`, treating
ANY non-zero (including a transient check error) as a kill and
permanently ending the monitor for the container's lifetime. Now
break ONLY on the real-kill signal (rc==1) — preserving the
kill -> wait -n -> container-restart -> boot-purge contract — while a
transient non-zero keeps the monitor live and re-checks next cycle.
Verified RED->GREEN against the real entrypoint in node:22-slim with a
stubbed du seam: non-numeric du now warns and keeps the gate active; a
transient error no longer permanently disables the loop.
The tree-kill enumerated agent descendants in a single /proc snapshot then
killed. A child that forks a new child (or reparents) between the scan and the
kill escaped the walk, reparented to PID 1, and kept the agent port bound —
defeating the tree-kill's whole purpose of freeing the port before the
container restart.
Replace the single snapshot with a BOUNDED re-scan loop: keep the root alive as
the walk anchor, re-enumerate and SIGKILL live descendants deepest-first each
pass (up to 5 passes, 0.2s apart) until a scan comes back empty, then kill the
root last. Killing the root FIRST would immediately reparent every descendant to
PID 1 and make them unreachable by the root-anchored PPID walk, so root-last is
required for the re-scan to reap late/mid-walk descendants. A descendant that
fully daemonizes (double-fork to PID 1) before we reach it remains out of reach
— documented as an inherent limit of PPID-based reaping without job control; the
agent's npm->node tree does not daemonize.
Also document why the ${stat##*) } PPID parse is safe against a comm containing
") " (longest-prefix to the last ") " always lands on the true terminator).
Applied identically to langgraph-typescript (:8123) and strands-typescript
(:8000). Proven via local RED-GREEN in node:22-slim: pre-fix leaks a mid-walk
escapee (port stays bound), post-fix reaps it (0 orphans, port freed).
The cleanup() EXIT/SIGTERM trap in both langgraph-typescript and
strands-typescript entrypoints did a bare `kill $AGENT_PID`. Because
$AGENT_PID is the outer process-substitution subshell (not the real
npm->node server), this reaped only the subshell and orphaned the node
server (reparented to PID 1, still holding :8123 / :8000) on every
graceful/SIGTERM shutdown -- e.g. every Railway redeploy/rollover.
Route cleanup() through the existing _kill_agent_tree helper (as the
size-watchdog and health-strike kill sites already do), and move the
_agent_descendants/_kill_agent_tree helpers above cleanup()/the trap so
they are defined whenever the trap can first fire.
Brings the 499-commit-stale foundations branch up to date with main so #5761
has a clean diff and no stale reverts (e.g. forwardHeaders). Conflicts:
- CopilotThreadsDrawer.tsx: took main's (main renamed CopilotDrawer -> ThreadsDrawer
+ added the collapse feature; the branch's edit was a no-op import-type split).
- pnpm-lock.yaml: regenerated with the pinned pnpm 10.33.4 (adds @copilotkit/bot-intelligence).
strands-typescript/entrypoint.sh carries the identical latent trap fixed in
langgraph-typescript by this branch. The agent is launched through a process
substitution — `cd /app/src/agent && npm start &> >(awk …) &` — so
$AGENT_PID (=$!) is the outer subshell wrapping that pipeline, NOT the npm→node
tree it forks (`npm start` runs `node --import tsx server.ts`, which stays a
child of npm, not an exec-replacement). The health-strike `kill -9 $AGENT_PID`
therefore reaps only the subshell; npm and node reparent to PID 1 and KEEP
RUNNING, still bound to :8000. `wait -n` never observes the real server die →
container never restarts → the frontend proxies to a dead-but-not-restarted
agent forever (edge 502s). strands-typescript has no size-gate, so this only
fires after the 3-strike (~90s) health counter exhausts, but it is a real
latent footgun with the same root cause.
Fix: reuse the /proc-based `_kill_agent_tree` helper (node:22-slim ships
neither ps nor pgrep, and job control is off so a group kill would take out
the whole entrypoint) at the single health-strike kill site. The whole
npm→node tree now dies, :8000 is freed, `wait -n` returns and the container
restarts.
Red-green proven in a real node:22-slim container against the exact
process-sub → npm → node structure: RED (bare kill -9 $AGENT_PID) leaves node
orphaned and :8000 still LISTENing; GREEN (_kill_agent_tree) reaps the tree,
0 orphans, port freed.