Commit Graph

13191 Commits

Author SHA1 Message Date
Tyler Slaton 68dbadeab2 Merge remote-tracking branch 'origin/main' into telegram-bot-node20-test-fail 2026-07-08 12:34:11 -07:00
Ben Taylor 3b56a10807 refactor(channels): rename @copilotkit/bot* packages to @copilotkit/channels* (OSS-438) (#5849)
## Summary

Renames the **Bots SDK → Channels SDK** (leadership decision, OSS-438).
**Names only — no behavior change.**

| old | new |
|---|---|
| `@copilotkit/bot` | `@copilotkit/channels` |
| `@copilotkit/bot-ui` | `@copilotkit/channels-ui` |
| `@copilotkit/bot-slack` | `@copilotkit/channels-slack` |
| `@copilotkit/bot-teams` | `@copilotkit/channels-teams` |
| `@copilotkit/bot-discord` | `@copilotkit/channels-discord` |
| `@copilotkit/bot-telegram` | `@copilotkit/channels-telegram` |
| `@copilotkit/bot-whatsapp` | `@copilotkit/channels-whatsapp` |

Package dirs renamed (`git mv`); versions carried over.

## What changed
- **Packages:** dir renames, `name` fields, `repository.directory`,
descriptions, and `workspace:` cross-deps rewired in lockstep.
`createBot` and other **API names are unchanged** (out of scope).
- **Release plumbing:** `release.config.json` scope keys +
`versionSource`, `ReleaseScope` union in
`scripts/release/lib/config.ts`, the
`canary`/`stable-release`/`publish-release` workflow scope dropdowns,
and `verify-release-scope-dropdowns.sh`.
- **Consumers:** `examples/slack` (Kite) + `examples/teams` — deps, the
load-bearing `jsxImportSource` pragma, build globs, imports.
- **Docs (`showcase/shell-docs`):** MDX package refs, content dirs
`docs/bots`→`docs/channels` and `reference/bot`→`reference/channels`,
nav registry, and **permanent redirects** from the old `/bots` and
`/reference/bot` URLs.

## Verification
- All 7 `@copilotkit/channels*` packages build; package test suites pass
(143+ in `channels`).
- `examples/slack` + `examples/teams` typecheck clean against the
renamed packages.
- `verify-release-scope-dropdowns.sh` green; release notification
wrapper test 30/30; `release:prepare` dry-runs for `channels` and
`channels-slack` resolve.
- `showcase/shell-docs` `frontend-options` test + full `tsc --noEmit`
clean.
- Zero stray `@copilotkit/bot` / `packages/bot` refs remain.

## ⚠️ Before merge / after merge
- **Do not squash-lose the deprecation step:** after these publish, run
`npm deprecate @copilotkit/bot@"*" "Renamed — install
@copilotkit/channels instead."` for the **5 published** old packages
(`bot`, `bot-ui`, `bot-slack`, `bot-teams`, `bot-discord`).
`bot-telegram`/`bot-whatsapp` were never published.
- New package names have **no npm version history**; the `package.json`
`version` seeds the first publish. Dry-run the OIDC publish for a
never-published scope first.
- Public-surface rename → stakeholder sign-off (kept as draft).

Refs OSS-438

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-08 14:32:47 -05:00
Tyler Slaton 3fb0e3caeb test(bot-telegram): de-flake chunked-edit-stream intermediate-flush tests
Two ChunkedEditStream tests waited an arbitrary `setTimeout(r, 10)` for the
internal `setTimeout(0)` + microtask flush chain, then asserted `callCount === 1`.
When CI scheduling jitter/GC ate that <=10ms window before the first `editAt`
ran, `callCount` was still 0 and the test failed with `expected +0 to be 1`.
It surfaced on the Node 20 shard of `test / unit`, while the near-identical
sister test passed in the same run — a classic flaky-timing race, not a Node 20
semantic difference.

Replace the arbitrary sleep with event-based waiting: resolve a promise the
instant the first `editAt` runs, so each test waits for the actual condition it
cares about rather than guessing a duration. Deterministic regardless of
scheduling. Verified with a mutation probe (a slowed flush chain fails the old
10ms-wait pattern but passes the new condition-based wait).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-08 12:30:44 -07:00
github-actions[bot] 86578f8e94 style: auto-fix formatting 2026-07-08 19:28:06 +00:00
Benjamin Taylor dd67e5f401 fix(bot-intelligence): address managed-HITL review findings
- ack/nack: delete the lease BEFORE the wire call so a background turn that
  completes after a timeout-nack can't also ack (single terminal signal).
- targetFromRef: throw a clear, actionable error when a ref carries no delivery
  routing instead of coercing deliveryId to the string "undefined".
- getOrCreate: guard getHistory with try/catch so a contract-violating source
  can't nack an otherwise-fine turn; document the app-api turnId-distinctness
  invariant the in-place card update depends on.
- getHistory: clamp to the requested limit (parity with the in-memory source);
  note the Slack-shaped route cast and the thread_started wire asymmetry.
- fetchFile: bound the body read with a generous content-length backstop.
- state store: forward an explicit ttlMs:0 (was dropped by truthiness) and add
  the missing TSDoc docblocks.
- supportsBlockingChoice: future-tense the TSDoc (no reader exists yet).
- tests: interaction update-in-place stamping (the HITL crux), content-parts
  text/unknown-mime branches, default-store resolution, kv non-2xx + ttlMs:0.
2026-07-08 14:27:15 -05:00
Jordan Ritter fe2b124903 fix(showcase): watchdog follow-up hardening — _require_int overflow clamp, size-loop trap-order + grace coverage, read-builtin cleanup (#5879)
## Summary

Follow-up cleanup/hardening for the two showcase entrypoint watchdogs
fixed in #5874. A 5-round code review deferred a set of non-load-bearing
bucket-(b) polish items to this PR; they are implemented here. Scope is
limited to:

- `showcase/integrations/langgraph-typescript/entrypoint.sh`
- `showcase/integrations/strands-typescript/entrypoint.sh`

Shared helper code (`_require_int`, `_agent_descendants`) is kept
byte-identical across both files. `bash -n` and `shellcheck
--severity=warning` are clean on both.

## Items implemented

### 1. `_require_int` upper bound (behavioral, both files)
The validator accepted any positive integer (`[1-9][0-9]*`), so a 20+
digit override overflowed bash's signed-64-bit arithmetic and either
wrapped to garbage or aborted the `[ -ge ]` test — which, suppressed to
false inside the guard's `if`, **silently disabled the guard** (the
exact fail-open class the validator exists to prevent). Added a 10-digit
length cap (max 9,999,999,999 — far above any real
interval/threshold/strike knob, comfortably inside int64), checked
before the digit `case` since an all-digit 23-char value would otherwise
pass. A too-long value now takes the same WARN + fall-back-to-default
fail-safe path as every other bad override.

**Cap rationale:** a single generous 10-digit cap rather than per-knob
caps — 9,999,999,999 seconds is ~317 years and 9,999,999,999 MB is ~9.3
PB, so no legitimate value is ever excluded, and it leaves 9 digits of
headroom below the 19-digit int64 ceiling.

### 2. `cleanup()` comment fix (langgraph, non-behavioral)
Corrected the note claiming `WATCHDOG_PID` "forks nothing that outlives
it" — it **does** fork the size sub-loop; the bare `kill $WATCHDOG_PID`
is safe because the watchdog's own inner EXIT trap reaps that child, not
because it forks nothing. The strands `cleanup()` comment was verified
accurate (strands has no size sub-loop) and left unchanged.

### 3. SIZE_PID trap-ordering leak window (langgraph, behavioral)
The size sub-loop was backgrounded (`( … ) &`, `SIZE_PID=$!`) **before**
its reaping `trap … EXIT` was registered, so an outer SIGTERM landing in
that window exited the watchdog subshell with no trap armed and orphaned
the sub-loop (reparented to PID 1, spinning for the container's life).
Now the reaping trap is armed **first** and reaps via a `$BASHPID`
PPID-walk that finds the child regardless of whether `SIZE_PID` is
assigned yet — no ordering-dependent leak.

### 4. Size-guard coverage during startup grace (langgraph, behavioral)
The size monitor started only **after** the up-to-180s startup-grace
loop, leaving the size ceiling unguarded during a pathological cold
boot. It now starts **before** the grace loop. Decision: starting early
is safe because `_watchdog_check_size_once` already fail-closes on every
not-yet-ready condition (agent PID not alive, `PERSIST_DIR`
missing/freshly-purged, non-numeric size or threshold), so early cycles
are harmless no-ops until the dir actually grows. No documented-gap
fallback was needed.

### 5. Cosmetic (both files, non-behavioral)
- **Done:** `_agent_descendants` per-PID `echo | awk` fork replaced with
the `read` builtin (clean drop-in; avoids forking awk once per `/proc`
entry). Byte-identical across both files.
- **Skipped — strike-window "off-by-one" log:** examined the
health-strike log `~$((HEALTH_CHECK_INTERVAL * HEALTH_STRIKE_LIMIT))s`;
with interval=30, limit=3 the third failure lands at ~90s and the
product is 90 — there is no actual off-by-one to fix, so touching it
would add churn without value.
- **Skipped — unthrottled transient-error re-warn:** throttling the
per-cycle transient `[watchdog:size]` warning requires adding per-loop
state/timestamp bookkeeping; low value and would balloon the diff into
the hot loop, so skipped per the "skip if not clean/low-risk"
instruction.

## RED → GREEN (behavioral items 1, 3, 4)

Proven on the **real committed entrypoint bytes** in `node:22-slim`
(Debian bash 5.x, real `/proc`), mirroring how #5874's fixes were
proven. RED = the merged #5874 baseline (`7912b6e51`); GREEN = this
branch's HEAD.

### Item 1 — overflow clamp (both files)
```
RED  langgraph: input HEALTH_CHECK_INTERVAL='99999999999999999999999' (23 digits)
     after _require_int, HEALTH_CHECK_INTERVAL='99999999999999999999999'
     $(( HEALTH_CHECK_INTERVAL * 3 )) = 601129261562068989
     RESULT: 23-digit value SURVIVED validation <-- BUG: no upper bound
RED  strands:  (identical to langgraph — byte-identical helper)

GREEN langgraph: [entrypoint] WARNING: health interval (HEALTH_CHECK_INTERVAL) is too large (got: '99999999999999999999999', 23 digits — max 10) — falling back to default 30
     after _require_int, HEALTH_CHECK_INTERVAL='30'
     $(( HEALTH_CHECK_INTERVAL * 3 )) = 90
     RESULT: value was clamped to '30' (default) — guard safe <-- FIXED
GREEN strands:  (identical — clamped to 30, arithmetic = 90)
```

### Item 3 — trap-ordering leak window (langgraph)
```
RED  post-fix ordering emulated with the pre-fix spawn-then-trap sequence:
     size sub-loop ticks AFTER watchdog exit = 2
     RESULT: sub-loop ORPHANED — still ticking after watchdog gone <-- BUG (leak window)

GREEN size sub-loop ticks AFTER watchdog exit = 0
     RESULT: sub-loop REAPED on watchdog exit — no orphan <-- FIXED
```

### Item 4 — size-guard coverage during startup grace (langgraph)
```
RED  grace-loop announce at line 437; size-monitor spawn at line 459
     RESULT: size monitor starts AFTER grace loop — UNGUARDED during up-to-180s grace <-- BUG

GREEN grace-loop announce at line 533; size-monitor spawn at line 507
     RESULT: size monitor starts BEFORE grace loop — guarded during cold start <-- FIXED
```
Functional confirmation (real watchdog block, short grace, stubbed
size-check + never-healthy probe):
```
[watchdog] Startup grace: waiting up to 5s for first successful health probe before arming strike counter
[watchdog:size] Starting size-gated restart monitor (threshold=200MB, interval=1s, dir=/tmp/persistX)
[GRACE-WINDOW-SIZE-CHECK-RAN]
[GRACE-WINDOW-SIZE-CHECK-RAN]
--- 3s elapsed (still within 5s grace) ---
```
`--check-size-once` seam re-verified post-refactor: under threshold →
exit 0 (no kill); over threshold → tree-kill + exit 1.

## Lint
`bash -n` and `shellcheck --severity=warning` clean on both entrypoints.

## Deferred (bucket (d)) — known future work, NOT in this PR
- A dedicated Next.js frontend watchdog (health probe + tree-kill for
the `NEXTJS_PID` process-sub subshell, symmetric to the agent watchdog).
- A strands size-guard / boot-purge (strands currently has only the
health watchdog; no persistence-size ceiling or boot purge).

These are separate features, out of scope for this cleanup PR.
2026-07-08 12:13:19 -07:00
Benjamin Taylor 0e6a01ae4b fix(bot-intelligence): dead-letter unmappable deliveries instead of wedging the loop
The exhaustiveness guard added to mapDeliveryToEnvelope threw on an unknown
wire kind, but claimOnce records the lease *before* mapping and runLoop's outer
catch only logs+sleeps — so an unmappable delivery leaked its lease and, after
the 120s lease expiry, was redelivered and re-thrown forever, permanently
wedging the single-delivery loop. Catch the map failure in claimOnce and nack
it non-retryably (deterministic failure → let app-api dead-letter it) so the
guard fails loud without blocking the queue. Adds a retryable flag to nack and
a regression test.
2026-07-08 14:02:58 -05:00
Jordan Ritter 4b670f4e2e docs(showcase): correct misleading watchdog/PID comments in entrypoints
The launch-site comment claimed process substitution leaves $! pointing at
the real node process. That is false and contradicts the file header: $!/
AGENT_PID is the wrapping subshell, and the real npm->node server is a
descendant reached only via the tree-kill (the reason _kill_agent_tree exists).
Rewritten in both langgraph-typescript and strands-typescript entrypoints so no
maintainer reintroduces a bare kill.

Also in langgraph-typescript: clarify that the SIZE_PID kill is a retained
belt-and-suspenders backstop to the $BASHPID PPID-walk (not dead code), and
re-anchor the startup-grace rationale to the real cause (the top-level
@langchain/langgraph-api import cost), since the prod path no longer uses
langgraph-cli dev.

Comment-only; no executable code changed.
2026-07-08 12:02:05 -07:00
github-actions[bot] 6f2b66cadc style: auto-fix formatting 2026-07-08 19:00:38 +00:00
Benjamin Taylor 41aed530c3 docs(channels): fix stale bot links in package docs + Channels landing page
Addresses review (tylerslaton):
- Package README/ARCHITECTURE relative links ../bot* -> ../channels* across
  discord/slack/teams/telegram/whatsapp (404'd after the dir rename)
- Channels landing page: Card hrefs /bots/{persistence,transcripts} ->
  /channels/*, and 'Bot reference' -> 'Channels reference'
- Package-noun prose (bot engine -> channel engine, bot-ui -> channels-ui,
  bot-slack approach -> channels-slack approach)

Left unchanged: runtime/third-party 'bot' prose, /api/bots/* Intelligence
wire paths, next.config /bots redirect sources.
2026-07-08 13:59:50 -05:00
Jordan Ritter 5a16b59f1e fix(showcase): guard sleep in _kill_agent_tree re-scan loop
The bounded re-scan loop's sleep 0.2 was unguarded. Under set -e on a base image
whose sleep can return non-zero (e.g. a future busybox/Alpine rebase), a failed
sleep would abort the tree-kill mid-walk — root never killed, real npm→node
server left orphaned. Add || true so the walk completes regardless of sleep's
exit status. No behavior change on the current Debian base (coreutils sleep
succeeds). Helper kept byte-identical across both entrypoints.
2026-07-08 11:56:01 -07:00
Jordan Ritter abaf72185a fix(showcase): key exit-code diagnostic off the actual reaped PID
Both langgraph-typescript and strands-typescript entrypoints inferred which
process exited via a post-hoc kill -0 if/elif after wait -n. That inference is
racy: on a near-simultaneous exit both PIDs are dead by probe time, so the first
kill -0 branch always wins and mislabels the diagnostic (naming the agent when
Next.js actually exited, attaching the wrong code to the wrong name). Use
bash's wait -n -p REAPED_PID (bash >= 5.1; node:22-slim ships 5.2) to capture
the actual reaped PID and key the message off it. Exit code (incl. 137) and the
final exit $EXIT_CODE are preserved; || EXIT_CODE=$? guard is unchanged.
2026-07-08 11:55:50 -07:00
Jordan Ritter 8d676704c7 fix(showcase/langgraph-typescript): close size-loop trap-order leak window and guard size ceiling during startup grace
Three related size-watchdog hardening changes in the langgraph entrypoint:

- Trap-order leak window: the size sub-loop was backgrounded (`( … ) &`,
  SIZE_PID=$!) BEFORE its reaping `trap … EXIT` was registered, so an outer
  SIGTERM landing in that window exited the watchdog subshell with no trap
  armed and orphaned the sub-loop (reparented to PID 1, spinning for the
  container's life). Arm the reaping trap FIRST, and reap via a $BASHPID
  PPID-walk that finds the child regardless of whether SIZE_PID is assigned
  yet — no ordering-dependent leak.

- Startup-grace size coverage: the size monitor started only AFTER the
  up-to-180s startup-grace loop, leaving the size ceiling unguarded during a
  pathological cold boot. Start it BEFORE the grace loop; safe because
  _watchdog_check_size_once already fail-closes on every not-yet-ready
  condition (agent PID not alive, PERSIST_DIR missing, non-numeric size/
  threshold), so early cycles are harmless no-ops until the dir grows.

- cleanup() comment accuracy: corrected the note claiming WATCHDOG_PID
  "forks nothing that outlives it" — it DOES fork the size sub-loop; the
  bare `kill $WATCHDOG_PID` is safe because the watchdog's own inner EXIT
  trap reaps that child, not because it forks nothing.

Proven RED->GREEN on the real entrypoint in node:22-slim: pre-fix the size
sub-loop keeps ticking after the watchdog exits (orphan) and the size monitor
spawns after the grace loop (unguarded); post-fix the sub-loop is reaped (0
ticks) and the monitor runs during grace (size-check fires within the grace
window). --check-size-once seam re-verified under/over threshold.
2026-07-08 11:36:23 -07:00
Jordan Ritter 5c71c1d047 fix(showcase): clamp _require_int to a 10-digit upper bound to stop int64 overflow disabling a guard
The numeric-config validator accepted any positive integer (`[1-9][0-9]*`),
so a 20+ digit override overflowed bash's signed-64-bit arithmetic and either
wrapped to a negative/garbage magnitude or aborted the `[ -ge ]` test with
"value too great for base" — which, suppressed to false inside the guard's
`if`, silently disabled the guard for the container's lifetime (the exact
fail-open class this validator exists to prevent).

Add a 10-digit length cap (max 9,999,999,999 — comfortably inside int64,
far above any real interval/threshold/strike knob) checked BEFORE the digit
`case`, since an all-digit 23-char value would otherwise pass validation.
A too-long value now takes the same WARN + fall-back-to-default fail-safe
path as every other bad override. Byte-identical across both entrypoints.

Proven RED->GREEN on the real entrypoints in node:22-slim: pre-fix a 23-digit
value survives validation and `$(( x * 3 ))` yields int64-wrapped garbage;
post-fix it WARNs, clamps to the default, and arithmetic is correct.
2026-07-08 11:35:59 -07:00
Jordan Ritter 966915b847 chore(showcase): replace per-PID echo|awk fork with read builtin in _agent_descendants
The /proc PPID walk forked an awk process for every entry in the process
table on every scan pass. Replace the `echo "${stat##*) }" | awk '{print $2}'`
pipeline with the `read` builtin, which word-splits the post-comm remainder
("STATE PPID PGRP …") on IFS and captures the 2nd field with no subprocess.
Byte-identical across the langgraph-typescript and strands-typescript
entrypoints. Non-behavioral; bash -n + shellcheck --severity=warning clean.
2026-07-08 11:35:31 -07:00
Benjamin Taylor 8fa44c44e7 fix(bot-intelligence): fail loud on unhandled managed delivery kinds
An unknown delivery kind previously fell through the dispatch switch as a
silent no-op (dispatch acks on resolve, so it would ack an unhandled delivery
as processed) and mapDeliveryToEnvelope coerced it into an empty turn. Add
matching `never` exhaustiveness guards to both so a future wire kind throws
instead of being silently swallowed.
2026-07-08 13:28:54 -05:00
github-actions[bot] 1cf6eefba9 style: auto-fix formatting 2026-07-08 13:27:35 -05:00
Benjamin Taylor b394f06fdc refactor(channels): rename @copilotkit/bot* packages to @copilotkit/channels* (OSS-438)
Renames the Bots SDK to the Channels SDK. Names only — no behavior change.

- 8 packages @copilotkit/bot* -> @copilotkit/channels* (git mv dirs, names,
  workspace: cross-deps). Now includes @copilotkit/bot-intelligence ->
  @copilotkit/channels-intelligence (landed on main via #5761; unpublished, so
  renamed fresh with the family).
- release.config.json scope keys + versionSource; ReleaseScope union;
  canary/stable-release/publish-release scope dropdowns; verify script
- examples/slack (Kite) + examples/teams: deps, jsxImportSource, imports
- showcase/shell-docs: content dirs docs/bots->docs/channels and
  reference/bot->reference/channels, nav registry, redirects

createBot and other API names unchanged. Old @copilotkit/bot* to be deprecated
after the new packages publish (bot-intelligence was never published).

Re-derived onto latest main (was conflicting after #5761 landed).

Refs OSS-438
2026-07-08 13:27:35 -05:00
Jordan Ritter 7912b6e512 fix(showcase/langgraph-typescript): tree-kill agent so size-watchdog restart actually fires (#5874)
## What & why

Two showcase agent containers (**langgraph-typescript**,
**strands-typescript**) could enter a *running-but-dead* state: Railway
showed the service `● Online` while `/api/health` returned **HTTP 502**.
This took down all 36 LGT dashboard cells on staging **and** prod
(prod's `.langgraph_api` had crossed the 200 MB size-watchdog
threshold).

**Root cause:** the agent is launched via process substitution (`... &>
>(awk …) &`), so `$AGENT_PID` (`=$!`) is the **wrapper subshell**, not
the real `npm`→`node` server that holds the port. Every watchdog/cleanup
did a bare `kill -9 $AGENT_PID`, which reaped only the subshell and
**orphaned the real server** (reparented to PID 1, still bound to the
port). The watchdog's "kill agent → container restart → boot-purge"
contract therefore never fired: the frontend kept proxying to a dead
agent → 502 forever.

## Fixes (each with local red-green on the real entrypoint in
`node:22-slim`)

1. **cleanup() EXIT trap** → routes through `_kill_agent_tree` (was
orphaning the agent on every SIGTERM/redeploy).
2. **`_kill_agent_tree`** → `/proc`-based tree-kill with a bounded
re-scan (root killed last) so mid-walk forks can't escape; refuses PID ≤
1 (fail-closed).
3. **size-watchdog** hardened against non-numeric `du` and transient
errors (no silent gate-disable, no permanent loop death).
4. **strands health-watchdog** → 180 s startup-grace window (parity with
langgraph); the now-effective kill would otherwise loop a slow cold
start.
5. **`wait -n` under `set -e`** → capture exit code so the restart
diagnostic isn't dead code on the primary (137) path.
6. **structural:** one `_require_int` validator over *every*
operator-overridable numeric knob (fail-safe to default), and **every**
wrapped-PID kill (incl. `NEXTJS_PID`) routed through the guarded
tree-kill; dangerous `${AGENT_PID:-0}` sentinel removed.
7. **`_require_int`** requires a positive integer (rejects `0` and
leading-zero/octal).

## Incident status
Staging **and** prod LGT were restored immediately via redeploy
(boot-purge cleared the oversized state) — both `/api/health` → 200.
This PR stops the recurrence.

## Review
Converged through a 5-round unbiased review-fix loop (1 + 4
confirmation), zero mandatory findings at close, all load-bearing guards
independently re-verified. `bash -n` + shellcheck (`-S warning`) clean;
170/170 shell bats pass.

## Follow-up (tracked, separate PR — non-load-bearing)
`_require_int` upper-bound clamp (LOW arith-overflow, needs a 20+-digit
value); a stale `cleanup()` comment; size-guard unarmed during the
startup-grace window; SIZE_PID trap-registration micro-window;
diagnostic label on near-simultaneous exit; cosmetic log nits; startup
readiness `sleep 3`+`kill -0` probes the wrapper subshell; no dedicated
Next.js frontend watchdog.
2026-07-08 11:27:19 -07:00
Benjamin Taylor 24bb847760 fix(bot-intelligence): reconcile merge with main's #5761 changes
- Drop the resurrected inline buildContentParts method from the adapter; the
  turn path now uses the extracted ./content-parts.js helper (#5814's refactor).
- Port main's fail-visible file-fetch behavior into content-parts.ts so a file
  that can't be retrieved becomes a short note in both the live-turn and
  history-seeding paths (they share the helper).
- Update the state-store conformance import to the @copilotkit/bot/testing
  subpath (main moved it there to keep vitest out of consumers' runtime graph).
2026-07-08 13:05:48 -05:00
Benjamin Taylor 69d11a5391 Merge origin/main into codex/managed-slack-hitl
Unstale #5814 onto main after the #5761 stack landed. Conflict:
- intelligence-adapter.ts: kept #5814's interaction dispatch + supportsBlockingChoice
  alongside main's #5761 review fixes (seq.delete cleanup, file fetch-failure note).
2026-07-08 12:57:26 -05:00
Ben Taylor 9fa925bd95 feat(bot,runtime): managed bots SDK — run the bot SDK from Intelligence-delivered events (OSS-360/361) (#5761)
## Summary

Lets the `@copilotkit/bot` SDK run from **Intelligence-delivered
events** without a second programming model, and adds the runtime `bots`
declaration API. A managed event (delivered by Intelligence) runs the
*same* customer handlers, tools, context, commands, Bot UI, and agents
as local/custom adapters — the managed path is "just another
`PlatformAdapter`," fed by injected transports.

This is the **OSS / SDK slice** of the Hosted Managed Bots work. The
credentialed transports (Realtime Gateway, Connector Outbox) and the
frozen shared contracts live elsewhere (see *Out of scope*); this PR
ships the seams they plug into, fully runnable headless.

Relates to **OSS-360** (runtime bots API), **OSS-361** (run the SDK from
Intelligence events), **OSS-363** (Slack render/codec reuse).

## What's in here

- **`intelligenceAdapter()` bridge** (`@internal`, not publicly
documented) — implements `PlatformAdapter` over two injected transports:
`DeliverySource` (inbound) + `EgressSink` (outbound). Ingress →
`onTurn`/`onCommand`/`onInteraction`/`onThreadStarted`/`onReaction`; ack
on success / nack on throw (at-least-once). Egress emits generic
operations carrying `BotNode[]` IR with **deterministic ids**
(`turnId:seq`, reset per turn) so a redelivered turn reproduces the same
ids for the Connector Outbox to dedupe. Idempotency lives at egress, so
the managed path skips ingress dedup (`skipIngressDedup`) — a redelivery
re-runs rather than being dropped.
- **Runtime `bots` API** — `new CopilotRuntime({ intelligence, bots })`,
accepted by TypeScript **only when `intelligence` is configured**
(discriminated union). `createBot({ name })`; `startManagedBots()`
validates names (required, identifier-style, unique — fail-loud), builds
activation metadata, and wires each bot to its resolved transport.
- **`PlatformCodec` seam** + Slack egress codec (`slackCodec`) composing
the existing pure `renderSlackMessage`, so IR→native rendering is shared
(no Bolt/creds) instead of duplicated.
- **Backwards-compatible SDK foundations**: `bot.addAdapter()` +
optional `adapters`, deferred backend resolution at `start()` with
`stateStore`-provider precedence (+ multi-provider warning),
`bot.transcripts` throws pre-start, optional
`eventId`/`turnId`/`deliveryId` on ingress + handler context. Existing
`createBot` callers and every `PlatformAdapter` implementer are
unaffected.
- **In-memory transports + fixture tests** — the full dispatch path
(envelope in → handler runs → egress op out) runs with zero
Slack/Intelligence/network.

## Out of scope (external / separate tickets)

- **Realtime Gateway + Connector Outbox transports** — implemented in
the closed-source repo against the `DeliverySource`/`EgressSink`
interfaces shipped here.
- **Shared contracts freeze (OSS-377)** — consumed here via a minimal,
isolated placeholder (`managed/contracts.ts`, marked `TODO(OSS-377)`);
swaps in via one import change.
- **OSS-363 ingress normalization** — the egress codec is done;
extracting the pure Slack event→neutral mapping out of the Bolt listener
(so local + Intelligence ingress share it) is the remaining, higher-risk
half and is left to that ticket (`TODO(OSS-363)`).

## Testing

TDD throughout (RED→GREEN per behavior). New: managed adapter
dispatch/ack-nack/ids/run-renderer/exclusivity, all-kinds routing, name
validation + metadata + lifecycle, runtime `bots` option, Slack codec.
Full suites green: `bot` 147, `bot-slack` 256, `runtime` 1574. All
builds typecheck (`bot`/`bot-slack`/`bot-discord`/`runtime`);
oxlint/oxfmt clean.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-08 12:49:17 -05:00
Mike Ryan 0f8c5385da chore(examples): Update CPK version and bound drawer grid row so threads list scrolls internally [ENT-1051] (post-release follow-up) (#5828)
## What

Applies the threads-drawer grid fix to all 15 integration examples: adds
`grid-template-rows: minmax(0, 1fr)` to each example's `.layout` so the
drawer's threads list scrolls **internally** (pinned header + New
Conversation) instead of the whole page growing to content height.

Without this, `height: 100dvh` on a grid whose rows aren't bounded lets
the row size to content, so the list can't scroll within the drawer and
the delete-confirm dialog centers against a content-tall root.

## Why this is a separate PR

These are **example** changes that consume **published** `@copilotkit/*`
packages. They were pulled out of the drawer-redesign PR (#5823) so that
PR stays scoped to the packages. This one lands **after** the redesigned
drawer is released.

## ⛔ Blocked / TODO before marking ready

- [ ] Drawer redesign PR #5823 merged
- [ ] Lockstep release cut (web-components + react-core + vue + angular)
- [ ] Bump each example's `@copilotkit/*` dependency to the new
published versions **in this PR** (currently only the CSS is here)
- [ ] Re-verify one example end-to-end against the released packages

## Testing

- Grid fix verified live during the redesign work (langgraph-js against
a hosted Intelligence backend): threads list scrolls internally, header
+ New Conversation pinned, delete-confirm renders correctly.
- Example dependency bumps + a fresh end-to-end pass will be
added/redone here once the release is published.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-08 10:29:20 -07:00
Benjamin Taylor 654389aa9e chore(examples): bump @copilotkit deps to 1.62.3 [ENT-1051]
The 1.62.3 release publishes the CopilotThreadsDrawer redesign (web-components +
react-core wrapper) and the stateless /suggest feature. Bump the 15 integration
examples that consume the drawer from 1.62.2 -> 1.62.3 (package.json + lockfiles)
so they pick up the released packages alongside this branch's example CSS.

Validated: langgraph-js runs on the published 1.62.3 (no local links) — the
redesigned drawer renders (New Conversation, Recent Conversations, filter funnel,
desktop collapse toggle, per-row kebab), threads are licensed, and a real agent
message round-trips.
2026-07-08 12:23:35 -05:00
Benjamin Taylor 5e0727359f fix(examples): unify header controls + even toggle padding [ENT-1051]
- ModeToggle: one style on both breakpoints (top-4/right-4 = 16px gutter,
  46px min-height, 4px corners); symmetric p-1.5 + fixed 20px button leading
  so the selected pill has an even gap on all four sides (was tight L/R vs T/B).
- Launcher: uniform 16px gutter (top + left) on both breakpoints so it mirrors
  the toggle; drop the mobile-only 7px override.
- Logo: centered on the launcher/toggle middle line (pt-[23px]); wordmark
  padding normalized so its height matches on both breakpoints.
- Inspector FAB: sits beneath the toggle, gap = the 16px top gutter (one rule,
  no media query, since the toggle is identical across breakpoints).

Net: launcher, logo, toggle share center-y; launcher + toggle are both 46px;
the FAB tucks under the toggle with a matching gap; the selected toggle pill is
evenly inset.
2026-07-08 12:23:35 -05:00
Benjamin Taylor 14fa89a674 fix(examples): align header controls — toggle clears inspector, launcher left matches right gutter, logo/toggle on one line [ENT-1051]
- ModeToggle: move left (right-[72px]) so the top-right inspector FAB no longer
  covers the App segment; grow to 46px (lg:min-h) + center on the logo line
  (top-6) to match the launcher; keep the 4px corners.
- Launcher: left gutter -> 16px to match the right-side controls' inset.
- Logo: pt-7 so it centers on the same line as the launcher + toggle.
2026-07-08 12:23:35 -05:00
Benjamin Taylor b69ae43989 fix(examples): mobile header boundary + square the Chat/App toggle to match the drawer [ENT-1051]
- Mobile header: max-lg:pb-0 -> pb-4 so chat content clears the fixed launcher/
  toggle strip instead of butting right under it (no boundary).
- Chat/App ModeToggle: rounded-full -> rounded-[4px] container + rounded-[2px]
  buttons, matching the drawer's 4px radius cap so the header controls are
  visually consistent.
2026-07-08 12:23:35 -05:00
Benjamin Taylor dcec3cb8c4 fix(examples): scope the drawer launcher position override to mobile [ENT-1051]
The 7px/16px launcher inset was tuned for the mobile off-canvas launcher; on
desktop it leaked onto the collapsed cluster. Move it into the mobile media
query so desktop-collapse uses the element's own 24px gutter default.
2026-07-08 12:23:35 -05:00
Benjamin Taylor 115439afca fix(examples): add a gap between the collapsed cluster and the app logo (6rem->7rem) [ENT-1051] 2026-07-08 12:23:35 -05:00
Benjamin Taylor 43e9793cf2 fix(examples): clear the drawer's collapsed cluster from the app header on desktop [ENT-1051]
The floating launcher/collapsed cluster is fixed at the top-left corner. Below
1024px it always shows (already cleared via max-lg:pl-24); on desktop it appears
only when the drawer is COLLAPSED. Drive the header's left padding off
--cpk-drawer-reserved-width (0px when collapsed, 320px default otherwise) so the
logo starts at ~6rem when collapsed and pl-6 when expanded — no overlap. No-op
on current packages (var never set → stays pl-6).
2026-07-08 12:23:34 -05:00
Benjamin Taylor 8344a90475 feat(examples): reclaim the drawer column on desktop-collapse via --cpk-drawer-reserved-width [ENT-1051]
Read grid-template-columns' first track from var(--cpk-drawer-reserved-width, 320px)
so when the drawer collapses on desktop (it sets the var to 0) the reserved
column collapses and the chat reclaims the space — instead of leaving an empty
placeholder column. Mobile (single-column) is unchanged.
2026-07-08 12:23:34 -05:00
Alem Tuzlak dff80a2a9a fix(bot,bot-slack,bot-intelligence,runtime): address managed-bots SDK review findings
Correctness:
- C1 create-bot: start() is now idempotent — a second start() no longer
  re-resolves the backend / rebuilds Transcripts+Telemetry+ActionRegistry or
  re-connects adapters (which would wipe MemoryStore state and double-bind real
  adapters). stop() clears the flag so start→stop→start is still a real restart.
- S1 bot-slack ingress: a threaded reply that @-mentions the bot is now skipped
  (app_mention handles it) so the managed path no longer double-responds. Matches
  both the plain <@U…> and labeled <@U…|handle> mention forms.
- S2 runtime: CopilotSseRuntime throws if `bots` is passed without intelligence
  instead of silently dropping them (guards a JS/as-any caller past the type).
- S3 bot-intelligence: startManagedBots rolls back — stops already-started bots —
  when a later bot fails to start, instead of leaking listeners/connections.

Lower:
- S4 ingress: stripMentions handles the labeled <@U…|handle> form; DM turns strip
  mentions too (parity with app_mention/thread_reply).
- S5 bot-intelligence: bot-name uniqueness is now case-insensitive.
- S6 runtime: fail fast at construction when a declared bot has no name (full
  shape/uniqueness validation stays at the activation seam — assertValidBotNames —
  because it can't cross into this CJS package from pure-ESM bot-intelligence).
- S7 bot-intelligence: buildActivationMetadata throws on a nameless bot instead of
  silently filtering it out of the activation set.
- S8 bot-intelligence: startManagedBots warns on an empty bots array.
- M1 intelligence-adapter: the per-turn egress seq Map entry is deleted after each
  turn so it can't grow unbounded over a long-running bot.
- M2 intelligence-adapter: an inbound file that fails to fetch degrades to a
  fail-visible text note instead of being silently dropped from model context.
- I2 contracts: dropped the now-dead `duplicate_skipped` RenderAccepted value
  (Intelligence returns duplicate_accepted or a 409 conflict).

Changelog (C2/C3, intended behavior after moving init into start()):
- bot.transcripts now throws before start() (was a concrete property).
- telemetry `oss.bot.configured` now fires at start() rather than construction, so
  a constructed-but-never-started bot no longer emits it.

Not addressed here (cross-repo, tracked on the Intelligence side):
- I1 realtime render-event kind:"file" clause on the gateway validator.
- I3 lease-token fencing on the render-accept path.
2026-07-08 19:18:19 +02:00
Jordan Ritter fe1ad53c44 fix(showcase): require positive integer for numeric config knobs (reject 0 and leading-zero/octal)
The _require_int validator in the langgraph-typescript and strands-typescript
entrypoints accepted '0' and leading-zero/octal forms like '010'/'08'. Operator
typos on any numeric knob then broke a guard:
- SIZE_THRESHOLD_MB=0 kills the agent on cycle 1 (instant restart loop)
- HEALTH_STRIKE_LIMIT=0 kills on first probe miss
- SIZE_CHECK_INTERVAL=0 / HEALTH_CHECK_INTERVAL=0 busy-spin on 'while sleep 0'
- '010' is read as OCTAL (8) in arithmetic; '08'/'09' abort under set -e

Tighten the predicate to accept only a positive integer with no leading zero
([1-9][0-9]*). Invalid values keep the existing fail-safe behavior: WARN and
fall back to the documented default. Helper stays byte-identical across both
files.
2026-07-08 10:09:15 -07:00
Jordan Ritter 3d1d0a4240 fix(showcase): validate all numeric config overrides and route every wrapped-PID kill through guarded tree-kill
CLASS 1 (guard silently disabled by a bad numeric override): add a reusable
_require_int validator and run it at startup over EVERY operator-overridable
numeric knob in both entrypoints (size threshold/interval, startup grace,
health-probe interval, strike limit). A non-integer/empty override now WARNs
and falls back to the documented default instead of breaking a sleep/loop/
arithmetic test. Closes instance #3 (LANGGRAPH_SIZE_CHECK_INTERVAL='60s'
killing the size-monitor loop on its first iteration).

CLASS 2 (wrapped-PID orphan + kill-0 footgun): route the cleanup() NEXTJS_PID
kill through _kill_agent_tree (it is process-sub-wrapped like the agent, so a
bare kill orphaned the real Next.js node server holding $PORT across redeploy).
Harden _kill_agent_tree and _agent_descendants to refuse a PID that is empty,
non-numeric, 0, or 1 (fail closed), making kill -9 0 / kill -9 1 structurally
impossible. Remove the ${AGENT_PID:-0} sentinel in the --check-size-once seam;
skip with a warning when AGENT_PID is unset instead of defaulting to 0.

Shared helper code kept byte-identical between the two entrypoints.
2026-07-08 09:59:43 -07:00
Tyler Slaton 97058dc00f docs(shell-docs): update Slack and Teams agent framing (#5789)
## Summary

- Backport the website messaging from CopilotKit/website#398 into the
shell-docs Slack and Microsoft Teams frontend pages.
- Replace the stale waitlist/managed-only framing with "get early
access" copy that presents CopilotKit Enterprise Intelligence as the
self-hosted or cloud-hosted production layer around the open source Bot
SDK.
- Frame Slack and Teams as frontends for agents built on any harness or
framework, while reserving production-layer terminology for CopilotKit
Enterprise Intelligence.
- Address browser review annotations on both pages: remove filler in the
opener, avoid setup-heavy lead copy, use "open source" without a hyphen,
add the full CopilotKit Enterprise Intelligence name to the CTA titles,
and keep CTA telemetry surfaces intact.
- Update the shell-docs nav test expectation so it matches the current
root IA, where Threads lives under Build Chat UIs rather than the
generated Intelligence Platform section.

## Validation

- `npm run lint` from `showcase/shell-docs` (passes with existing
warnings)
- `npm run typecheck` from `showcase/shell-docs`
- `npm run test` from `showcase/shell-docs`
- `npm run build` from `showcase/shell-docs`

## Notes

- Hydrated Git LFS assets locally with `git lfs pull` so the shell-docs
public asset tests could read real PNG bytes.
2026-07-08 09:54:30 -07:00
David McKay 5245634c62 chore(banking): add stop-demo.sh teardown companion to run-demo.sh (#5877)
## What

Adds `examples/showcases/banking/stop-demo.sh` — the teardown companion
to the existing `run-demo.sh`.

## Why

`run-demo.sh` detaches everything except the Next.js dev server:

- `docker compose up -d --wait` (detached stack)
- native Metal TEI via `nohup … & disown` (Apple Silicon only)
- `exec pnpm dev` (the only foreground process)

So Ctrl-C stops *only* the dev server and silently leaves the docker
stack (`banking-memory`) and the host embedder on `:7067` running. There
was no one-command way to bring those down. This script fills that gap
and mirrors `run-demo.sh`'s conventions (same `say`/`ok` helpers, same
header-comment style, idempotent).

## What it does

Tears down, idempotently, in order:

1. Next.js dev server on `:3000` (defensive — usually already gone via
Ctrl-C)
2. docker compose stack (project `banking-memory`), **containers only**
by default so a re-run of `run-demo.sh` reuses the built composite image
+ seeded Postgres
3. native Metal TEI on `:7067` (Apple Silicon; the host process docker
doesn't manage) — SIGTERM, then SIGKILL for anything that ignores it

## Flags

- `--purge` — also delete the docker volumes (postgres/redis/minio/tei
model cache) for a full clean-slate reset
- `--keep-tei` — leave the slow-to-warm native embedder running when
only bouncing the stack

## Testing

- `bash -n stop-demo.sh` — syntax clean
- `shellcheck stop-demo.sh` — clean, no warnings
- `./stop-demo.sh --help` renders the banner correctly

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-08 09:53:44 -07:00
Ben Taylor 04461409e5 fix(bot): move runStateStoreConformance to @copilotkit/bot/testing subpath (#5875)
## Problem

`@copilotkit/bot`'s package entry re-exports `runStateStoreConformance`
from `./testing/state-store-conformance`, which does `import { describe,
it, expect, ... } from "vitest"` at module top-level. In ESM a static
re-export **eagerly evaluates** the re-exported module, so a plain:

```ts
import { createBot } from "@copilotkit/bot";
```

drags `vitest` into the consumer's runtime module graph and throws
`ERR_MODULE_NOT_FOUND: Cannot find package 'vitest'` for any consumer
that doesn't have vitest installed (i.e. every production consumer).
`vitest` is only a devDependency.

Surfaced while smoke-testing the package rename (OSS-438) — but it's a
pre-existing bug on `main`, independent of that rename.

## Fix

- **Drop the re-export from `src/index.ts`** → the package entry is now
vitest-free.
- **Publish the helper under a `./testing` subpath**
(`@copilotkit/bot/testing`) — test tooling lives off the runtime entry,
the standard pattern.
- **Declare `vitest` as an optional `peerDependency`** so consumers of
`/testing` get the right signal.
- **Docs** updated to `import { runStateStoreConformance } from
"@copilotkit/bot/testing"`.

Only the import path of the test-only conformance helper changes; the
runtime API is untouched.

## Verification
- Static import trace: the entry graph is 12 runtime modules, **none**
import vitest; every vitest importer is a `.test.js` (not in the graph)
or `testing/state-store-conformance.js` (only reachable via `/testing`).
- `@copilotkit/bot` builds; **143/143** tests pass; `publint` + `attw`
clean (the internal conformance test imports the helper by relative
path, unaffected).

## Coordination
Touches `packages/bot` on `main`. The OSS-438 rename PR (#5849) renames
this package to `@copilotkit/channels`; that PR re-derives from `main`
before merge, so it will absorb this fix automatically. If #5849 merges
first, this rebases onto `packages/channels` mechanically.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-08 11:52:49 -05:00
Maxim 95430de51a chore(banking): add stop-demo.sh teardown companion to run-demo.sh
run-demo.sh detaches everything except the Next.js dev server (docker
compose up -d, native Metal TEI via nohup/disown, then exec pnpm dev), so
Ctrl-C on the dev server leaves the docker stack and the host embedder
running. stop-demo.sh brings those leftovers down in one command.

Tears down, idempotently:
  - the Next.js dev server on :3000 (defensive; usually gone via Ctrl-C)
  - the docker compose stack (project banking-memory), containers only by
    default so a re-run reuses the built image + seeded data
  - the native Metal TEI on :7067 (Apple Silicon; the host process docker
    doesn't manage), SIGTERM then SIGKILL

Flags: --purge also drops volumes for a clean slate; --keep-tei leaves the
slow-to-warm embedder running when only bouncing the stack.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-08 18:49:34 +02:00
Benjamin Taylor 224587101f fix(bot): move runStateStoreConformance to @copilotkit/bot/testing subpath
The package entry (@copilotkit/bot) re-exported runStateStoreConformance from
./testing/state-store-conformance, which imports vitest at module top-level.
An ESM re-export eagerly evaluates that module, so a bare
`import { createBot } from "@copilotkit/bot"` dragged vitest into every
consumer's runtime graph and threw ERR_MODULE_NOT_FOUND when vitest wasn't
installed (i.e. any production consumer).

- Drop the re-export from src/index.ts (entry is now vitest-free)
- Publish the conformance helper under the ./testing export subpath
- Declare vitest as an optional peerDependency (documents the /testing need)
- Update docs to import from @copilotkit/bot/testing

Names/behavior of the runtime API are unchanged; only the import path for the
test-only conformance helper moves.
2026-07-08 11:44:01 -05:00
Mike Ryan 25339b0d07 chore: release monorepo v1.62.3 (#5876)
## Release monorepo v1.62.3

**Scope:** `monorepo` | **Bump:** `patch`

---

### How this release process works

1. **This PR was created automatically** by the "release / create-pr"
workflow.
   It bumped the `monorepo` packages to `1.62.3`
   and generated AI-enhanced release notes.

2. **CI runs on this PR** — the full test suite (unit tests, lint, type
checks, build)
   must pass before merging. This is the review gate.

3. **Review the release notes** in `release-notes.md` in this PR.
If a Notion draft was created, you can edit the release notes there
before merging.

4. **When this PR is merged**, the `release / publish` workflow
automatically:
   - Builds all packages
   - Publishes the `monorepo` packages to npm at version `1.62.3`
   - Creates git tag `monorepo/v1.62.3`
   - Creates a GitHub Release with the final release notes

### Before merging

- [ ] CI is green (tests, lint, types, build)
- [ ] Version bumps look correct
- [ ] Release notes are accurate (edit in Notion if a draft was created)

---

> **Do not merge until CI is fully green.** The full test suite runs
automatically on this PR.
v1.62.3
2026-07-08 09:35:28 -07:00
Jordan Ritter ed44611263 fix(showcase): capture wait -n exit code under set -e so restart diagnostics aren't dead code
Both entrypoints run under set -e. The tail `wait -n $AGENT_PID $NEXTJS_PID`
returns non-zero on the PRIMARY designed exit path (137 = size-gate/watchdog
SIGKILL of the agent tree, or an agent crash), so set -e aborted the script AT
that line — making EXIT_CODE=$?, the entire 'which process exited with code N'
diagnostic, and the final `exit $EXIT_CODE` dead code on exactly the
interesting exits. Capture the code with `EXIT_CODE=0; wait -n ... || EXIT_CODE=$?`
so the diagnostic and explicit exit run and preserve the exact code (incl. 137);
the container-restart path is unchanged.

Same class: langgraph's LANGGRAPH_SIZE_THRESHOLD_MB was used in
`[ "$DIR_SIZE_MB" -ge "$threshold" ]` with no numericity guard, so a
non-integer operator override made the test error and silently no-op the size
gate every cycle. Validate the threshold the same way DIR_SIZE_MB already is
(numeric case guard + 'size guard inactive' WARNING, then skip safely).
2026-07-08 09:25:04 -07:00
tylerslaton 4394f9c81d chore: release monorepo v1.62.3 2026-07-08 16:17:36 +00:00
Alem Tuzlak 2330dca267 fix(bot-slack,runtime): align tests with renderer status + AbstractAgent.run
Two pre-existing test failures on this branch, surfaced by CI's unit +
check-types jobs once main was merged:

- bot-slack event-renderer: the non-pane thread tool-call test still
  asserted the old "no composer status" behavior. Commit 13248dda0b
  deliberately drove setStatus on ANY thread anchor (not just panes), so
  the test now expects both the 🔧 row and the "is using…" status.
- runtime in-memory-runner: HangingAgent/AbortableAgent extended
  AbstractAgent but omitted the abstract run() member (@ag-ui/client
  0.0.57), failing tsc on the test tsconfig (TS2515). Add the same
  run() => EMPTY stub the sibling test agents use.
2026-07-08 18:15:29 +02:00
Jordan Ritter 5c5e139484 fix(showcase/strands-typescript): add startup-grace window to health watchdog for parity with langgraph
The health-watchdog armed its 3-strike/~90s kill counter immediately with no
startup-grace window. langgraph-typescript has a 180s grace precisely to keep a
slow cold start from being killed mid-boot into a restart loop. Now that the
tree-kill makes the strands kill effective (the orphan bug previously made it
cosmetic), a slow tsx cold start (>90s) would be genuinely killed and loop.

Port langgraph's grace mechanism verbatim (GRACE=180, no env override, poll
every 5s, exit 0 on agent death during startup, arm anyway if grace elapses),
adapted to strands' :8000/health probe.
2026-07-08 09:13:48 -07:00
Jordan Ritter 31a9d072e8 fix(showcase/langgraph-typescript): harden size-watchdog against non-numeric du and transient errors
Two tightly-related defects in the size-gated restart machinery in
entrypoint.sh:

1. _watchdog_check_size_once validated the du/awk result only for
   emptiness, not numericity. A non-integer value (junk du output, a
   transient read error, a test-seam stub) reached the
   `[ "$DIR_SIZE_MB" -ge ... ]` comparison and threw "integer expression
   expected"; sitting inside an `if`, set -e was suppressed so the test
   evaluated false and the size gate was SILENTLY skipped with no
   warning (unlike the empty-string branch). Now match ^[0-9]+$ via a
   case and emit the same "size guard inactive" WARNING, so the gate can
   never silently disappear.

2. The size sub-loop used `_watchdog_check_size_once || break`, treating
   ANY non-zero (including a transient check error) as a kill and
   permanently ending the monitor for the container's lifetime. Now
   break ONLY on the real-kill signal (rc==1) — preserving the
   kill -> wait -n -> container-restart -> boot-purge contract — while a
   transient non-zero keeps the monitor live and re-checks next cycle.

Verified RED->GREEN against the real entrypoint in node:22-slim with a
stubbed du seam: non-numeric du now warns and keeps the gate active; a
transient error no longer permanently disables the loop.
2026-07-08 09:13:43 -07:00
Jordan Ritter 31a4377853 fix(showcase): bounded re-scan in _kill_agent_tree so mid-walk forks can't escape
The tree-kill enumerated agent descendants in a single /proc snapshot then
killed. A child that forks a new child (or reparents) between the scan and the
kill escaped the walk, reparented to PID 1, and kept the agent port bound —
defeating the tree-kill's whole purpose of freeing the port before the
container restart.

Replace the single snapshot with a BOUNDED re-scan loop: keep the root alive as
the walk anchor, re-enumerate and SIGKILL live descendants deepest-first each
pass (up to 5 passes, 0.2s apart) until a scan comes back empty, then kill the
root last. Killing the root FIRST would immediately reparent every descendant to
PID 1 and make them unreachable by the root-anchored PPID walk, so root-last is
required for the re-scan to reap late/mid-walk descendants. A descendant that
fully daemonizes (double-fork to PID 1) before we reach it remains out of reach
— documented as an inherent limit of PPID-based reaping without job control; the
agent's npm->node tree does not daemonize.

Also document why the ${stat##*) } PPID parse is safe against a comm containing
") " (longest-prefix to the last ") " always lands on the true terminator).

Applied identically to langgraph-typescript (:8123) and strands-typescript
(:8000). Proven via local RED-GREEN in node:22-slim: pre-fix leaks a mid-walk
escapee (port stays bound), post-fix reaps it (0 orphans, port freed).
2026-07-08 09:13:38 -07:00
Jordan Ritter a5f082f50c fix(showcase): tree-kill agent in cleanup() EXIT trap to prevent shutdown orphan
The cleanup() EXIT/SIGTERM trap in both langgraph-typescript and
strands-typescript entrypoints did a bare `kill $AGENT_PID`. Because
$AGENT_PID is the outer process-substitution subshell (not the real
npm->node server), this reaped only the subshell and orphaned the node
server (reparented to PID 1, still holding :8123 / :8000) on every
graceful/SIGTERM shutdown -- e.g. every Railway redeploy/rollover.

Route cleanup() through the existing _kill_agent_tree helper (as the
size-watchdog and health-strike kill sites already do), and move the
_agent_descendants/_kill_agent_tree helpers above cleanup()/the trap so
they are defined whenever the trap can first fire.
2026-07-08 09:11:29 -07:00
Sam Julien 39a86447db docs(shell-docs): update Slack and Teams agent messaging 2026-07-08 09:04:47 -07:00
Benjamin Taylor 71a4ac42e4 Merge origin/main into alem/oss-360-sdk-foundations
Brings the 499-commit-stale foundations branch up to date with main so #5761
has a clean diff and no stale reverts (e.g. forwardHeaders). Conflicts:
- CopilotThreadsDrawer.tsx: took main's (main renamed CopilotDrawer -> ThreadsDrawer
  + added the collapse feature; the branch's edit was a no-op import-type split).
- pnpm-lock.yaml: regenerated with the pinned pnpm 10.33.4 (adds @copilotkit/bot-intelligence).
2026-07-08 11:01:58 -05:00
Jordan Ritter 93f2abeeb5 fix(showcase/strands-typescript): tree-kill agent so health-strike restart actually fires
strands-typescript/entrypoint.sh carries the identical latent trap fixed in
langgraph-typescript by this branch. The agent is launched through a process
substitution — `cd /app/src/agent && npm start &> >(awk …) &` — so
$AGENT_PID (=$!) is the outer subshell wrapping that pipeline, NOT the npm→node
tree it forks (`npm start` runs `node --import tsx server.ts`, which stays a
child of npm, not an exec-replacement). The health-strike `kill -9 $AGENT_PID`
therefore reaps only the subshell; npm and node reparent to PID 1 and KEEP
RUNNING, still bound to :8000. `wait -n` never observes the real server die →
container never restarts → the frontend proxies to a dead-but-not-restarted
agent forever (edge 502s). strands-typescript has no size-gate, so this only
fires after the 3-strike (~90s) health counter exhausts, but it is a real
latent footgun with the same root cause.

Fix: reuse the /proc-based `_kill_agent_tree` helper (node:22-slim ships
neither ps nor pgrep, and job control is off so a group kill would take out
the whole entrypoint) at the single health-strike kill site. The whole
npm→node tree now dies, :8000 is freed, `wait -n` returns and the container
restarts.

Red-green proven in a real node:22-slim container against the exact
process-sub → npm → node structure: RED (bare kill -9 $AGENT_PID) leaves node
orphaned and :8000 still LISTENing; GREEN (_kill_agent_tree) reaps the tree,
0 orphans, port freed.
2026-07-08 08:51:15 -07:00