Commit Graph

13191 Commits

Author SHA1 Message Date
Mark 059ef24e3a Merge branch 'main' into mark/oss-451-mastra-demos-404 2026-07-09 09:04:26 -07:00
Maximiliano Korp d1fbd583dd docs(shell-docs): document cli import command 2026-07-09 09:02:24 -07:00
Tyler Slaton 606afd8a13 feat(bot-teams): add ./render export for managed reuse (#5878)
## Summary

Adds a `./render` subpath export to `@copilotkit/bot-teams` surfacing
`renderAdaptiveCard`, `isPlainText`, `collectPlainText`,
`ADAPTIVE_CARD_CONTENT_TYPE`, and `createRunRenderer`, so the managed
(Intelligence-hosted) Teams egress path can reuse the adapter's Adaptive
Card rendering without deep-importing `dist/*`.

Paired with the managed Microsoft Teams work in Intelligence (OSS-441
slice ②): CopilotKit/Intelligence#511.

## Notes

- Additive only — a new `exports["./render"]` entry +
`src/render/index.ts` re-export barrel. No behavior change to existing
exports.
- Not strictly on the critical path yet: the managed runtime currently
resolves the renderer via a runtime deep-dist import of the published
package, so this export is the clean forward path rather than a hard
dependency.
- Coordinate the `bot-teams` → `channels-teams` rename (OSS-438) before
merge.
2026-07-09 08:58:43 -07:00
Alem Tuzlak f993c54bfb fix(bot-intelligence): align managed runtime delivery ownership (#5800)
## Problem

The Intelligence realtime loop exposed two SDK-side ownership gaps while
testing managed Coworkers against the Intelligence PR:

- Phoenix channel auth should use the SDK socket auth token path
expected by the gateway.
- Delivery handling needs to carry the app-api lease token and
authoritative delivery scope through render/fail/complete handling
instead of rebuilding ownership from local defaults.
- Runtime integration needs to pass the render sink into
`intelligenceAdapter` so managed runtimes stream rich render frames over
the realtime path.

## Why

The target Coworker path is websocket-first: Intelligence app-api owns
durable delivery state, realtime-gateway owns live transport, the SDK
receives leased delivery over Phoenix, streams render events, waits for
durable receipt coverage, and sends completion intent without taking
over app-api ack authority.

This PR is stacked on Alem's SDK realtime branch so the dependency trail
is explicit:

- Base SDK branch: `codex/oss-402-sdk-render-events`
- Intelligence PR: https://github.com/CopilotKit/Intelligence/pull/466
- Linear: OSS-402 / OSS-406

## Fix

- Use Phoenix socket `authToken` for the managed bot channel.
- Preserve `leaseToken` and delivery `scope` in
`PhoenixRealtimeTransport` state.
- Send fail/nack payloads with the correct lease token, delivery status,
and optional accepted-through pointer.
- Pass `renderSink` from `startManagedBots` into `intelligenceAdapter`.
- Add regression coverage for render-sink propagation and lease-token
fail payloads.

## Testing Methodology

Local verification used an external pnpm store to avoid repo-local
`.pnpm-store` churn:
`PNPM_CONFIG_STORE_DIR=/private/tmp/pnpm-store-codex`.

- `PNPM_CONFIG_STORE_DIR=/private/tmp/pnpm-store-codex
PNPM_CONFIG_CONFIRM_MODULES_PURGE=false corepack pnpm --dir
/private/tmp/copilotkit-oss-402-20260701 --filter
@copilotkit/bot-intelligence test`
  - Passed: 5 files, 48 tests.
- Commit hook also ran the package gate for
`@copilotkit/bot-intelligence`:
  - `test`: passed, 5 files / 48 tests.
  - `publint`: passed with repository URL suggestion only.
- `attw --pack . --profile esm-only`: passed with the existing ignored
CJS-to-ESM warning profile.
- `git diff --check -- packages/bot-intelligence/src/phoenix-channel.ts
packages/bot-intelligence/src/phoenix-transport.ts
packages/bot-intelligence/src/render-events.test.ts
packages/bot-intelligence/src/runtime.test.ts
packages/bot-intelligence/src/runtime.ts`
  - Passed.

Scope control: only the five `packages/bot-intelligence/src/*` files
above are committed. Existing local `pnpm-lock.yaml` and image/LFS dirt
in the SDK worktree were left unstaged and are not in this PR.
2026-07-09 17:56:38 +02:00
Maxim c54148412a test(banking): deterministic OGUI routing guard over the adjacency set
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-09 16:43:47 +02:00
Mark 9a6c1f4753 Merge branch 'main' into mark/oss-451-mastra-demos-404 2026-07-09 07:21:28 -07:00
Benjamin Taylor 58e561e2fa Merge origin/main into alem/oss-441-managed-teams (Bots→Channels rename)
Reconcile #5878 with the @copilotkit/bot*→@copilotkit/channels* rename (#5849):
- relocate the new render barrel (render/index.ts, render/index.test.ts) into
  packages/channels-teams; relative imports (./adaptive-card, ../event-renderer)
  resolve unchanged. package.json ./render export auto-merged.
- fix a stale @copilotkit/bot-teams comment ref to @copilotkit/channels-teams.
2026-07-09 08:44:19 -05:00
Benjamin Taylor b5dad876ab fix(channels-intelligence): make Phoenix delivery drops observable + cover new behaviors
Review follow-ups for the managed delivery-ownership change:
- add an optional `log` seam to PhoenixTransportConfig; the transport was
  otherwise silent, so the two new drop paths were invisible failure modes.
- log the leaseToken-missing drop distinctly (the gateway/SDK version-skew
  hazard: without it, every delivery silently re-loops on lease lapse) and the
  nack no-delivery-state drop, instead of bare returns.
- refresh TSDoc: toIngressEnvelope's new return shape + drop semantics, and the
  DeliveryState.leaseToken/scope fields.
- tests: leaseToken-required drop, nack no-state no-op, and per-delivery scope
  stamping on render + fail (the scope field previously had no coverage).
2026-07-09 08:10:44 -05:00
Benjamin Taylor 024446e5db Merge origin/main into codex/oss-402-sdk-handoff-fixes (freshen) 2026-07-09 08:07:37 -05:00
Tyler Slaton 66b5a58339 fix(docs): repair Claude custom look snippets 2026-07-08 21:39:49 -07:00
github-actions[bot] ae3ecc2cb7 style: auto-fix formatting 2026-07-09 03:39:16 +00:00
Tyler Slaton 0f5a916075 fix(docs): clean Claude generative UI snippets 2026-07-08 20:38:13 -07:00
Jordan Ritter 66502b5b3d fix(showcase-harness): D4 probe hardening follow-ups (budget-exhaustion, degraded-path, telemetry, coverage) (#5888)
Tracked follow-up to #5882 (D4 first-token SSE turn-complete fix). These
are the deferred **bucket-(b)** items from that PR's CR —
non-load-bearing polish/hardening on the D4 probe driver
(`showcase/harness/src/probes/drivers/d4-chat-roundtrip.ts` +
`.test.ts`). No re-architecture; each change is tight and scoped.

## Items

### 1. Budget-exhaustion retry guard (behavioral)
A late non-completion retry resend could floor its `type`/`press` action
timeout to ~1ms (when the remaining budget ≈ 0), throwing a
page-fault-shaped error. That throw was red-classified indistinguishably
from a real page fault — a spurious-red flap source. Fix: skip the
resend when the remaining budget is below `RETRY_MIN_BUDGET_MS` (750ms);
the stall then reds on its own terms.

**Red-green:** with the guard disabled, the doomed resend attempts a
second `type` (typeAttempts=2); with the guard, `type` fires exactly
once (typeAttempts=1). RED observed (2), GREEN observed (1).

### 2. Degraded-path floor when interceptor silently no-ops (behavioral)
When `sseAttachFailed` is true, the page-side turn-lifecycle globals
were never seeded, so the poll could only fall into the never-observed
branch and pin the deadline to the base `textPollTimeoutMs` floor —
reintroducing the slow-first-token false-red #5882 targets. Fix: consult
`sseAttachFailed` to widen the never-observed wait to the per-attempt
ceiling.

**Red-green:** degraded page + late-but-present token (800ms, base floor
bites at the ~500ms poll before the token) → RED (pre-fix, base-floor
fast-fail) vs GREEN (post-fix, widened to ceiling captures the token).
RED observed (`red`), GREEN observed (`green`). A genuinely-empty
degraded run still reds (over-correction guard).

### 3. Retry edge-header re-attribution (telemetry)
On a retry-rescued GREEN turn, `messageSendEdge` / `messageSendEdge` /
`lastMessagePostResp` stayed latched to the first (stalled) attempt,
mis-attributing `edge_interference_signal` / the DEBUG raw-byte sample.
Fix: re-arm the capture latches before the resend so the winning
attempt's response re-captures them.

**Focused test:** stalled attempt carries `cf-mitigated: stalled`,
winning resend `cf-mitigated: winning`; the final `probe.message.send`
boundary now carries `winning` (pre-fix it reported `stalled`).

### 4. Coverage (tests only, no prod change)
- FIFO-cap `CVDIAG_MAX_OUTSTANDING_STARTS_PER_URL` eviction backstop
(guard-verified: fails if the cap is removed).
- DEBUG-auto-disarm fail-closed negative test (disarmed → no raw-byte
capture).
- Alternate-content / raw-byte block SKIPPED on the container-success
(non-empty) path.
- Fixed the `makeLateTokenBrowser` shared-page state leak: each
`newContext().newPage()` now mints its own state, so L3 and L4 each run
an independent stall+retry cycle (was a singleton that leaked
`sendCount` from L3 into L4, so the L4 retry path was never genuinely
exercised).

### 5. `lastStoppedAtMs` documentation (cleanup)
Write-only in d4; documented that it is retained for shared-global
parity with the `attachSseInterceptor` global shape the d6 run-signal
snapshot mirrors, so it does not read as dead code.

## Verification
- `tsc --noEmit`: clean
- `tsc -p tsconfig.build.json` (build): clean
- Full harness vitest: **3213 passed, 0 failing** (72 in the d4 file,
all green)
- oxlint: **0 errors** (only pre-existing warnings, none new)
- Diff hygiene: only the driver + test file (no `repro_*` / `baseline`)

🤖 Generated with [Claude Code](https://claude.com/claude-code)


---

## CR follow-up round (commit 13a636b67)

CR found the follow-up left `readTurnState()` error handling
**inconsistent** — the exact false-red/flap class this effort fights.
Addressed, plus completed two of the PR's own items.

### (a) Harmonized `readTurnState()` error handling
One guarded `safeReadTurnState()` wrapper now backs all three consumers
(`readBaseline`, `readTurnComplete`, `readDegraded`). A mid-poll
`readTurnState()` throw now means ONE thing everywhere: "no reliable
signal" → return a well-defined degraded sentinel (`sseAttachFailed:
true`, zeroed counters) → route onto the degraded **widen** path (not
base-floor fast-fail, not a spurious `level-error`), and it is
**observable** via a one-shot greppable marker. Previously:
`readDegraded` swallowed the throw into `false` (silent base-floor
false-red, no telemetry) while `readBaseline`/`readTurnComplete` let it
escape → generic `level-error` (spurious red). A genuinely-empty
degraded turn still reds at the widened ceiling (no masking).

### Folds
- **item-1 first-send cap:** guard the in-send `press` against
`SEND_PRESS_MIN_BUDGET_MS` so a near-hang `type` can't floor `press` to
~1ms and yield a generic `level-error`; classify distinctly as
`send-budget-exhausted`. Only `press` is guarded (`type` opens the
envelope) so a legitimately-small `pageTimeoutMs` still issues a healthy
first send.
- **item-3 null-header:** suppress the `finally`-block fallback
`probe.message.send` once a real-header boundary already fired, so a
retry whose winning resend lands no POST no longer emits a second
NULL-header boundary (mis-attributed `edge_interference_signal`).
- errorDesc JSDoc: added `abort` to the enumeration (zero-risk).

### Red-green (verbatim, against the real runLevel/readTurnComplete
path)
- **readTurnState throw:** RED `expected 'red' to be 'green'` (pre-fix
spurious red on a late-but-present token) → GREEN (degraded widen; late
token passes; genuinely-empty degraded still reds).
- **first-send cap:** RED `expected 'level-error' to be
'send-budget-exhausted'` → GREEN (distinct classification, no
1ms-floored `press`).
- **item-3 null-header:** RED `expected 2 to be 1` (second null-header
emit) → GREEN (single correct attribution).

### Gates
`tsc --noEmit` clean · `npm run build` clean · full d4 vitest **76/76
pass** · oxlint **0 errors** (warnings pre-existing). Diff: driver +
test only (no repro_*/baseline).
2026-07-08 18:35:36 -07:00
Jordan Ritter 9d9056db60 fix(showcase-harness): propagate level errorDesc to aggregate signal + reorder abort short-circuit
The aggregate e2e-smoke:<slug> red signal omitted errorDesc on the normal
return path, so an abort/timeout/send-budget-exhausted red that runLevel
RETURNS (not throws) showed on the PRIMARY dashboard tick as an unclassified
content-shaped red — only the side chat:/tools: rows kept the classifier.
Thread the failing level's errorDesc (L3 precedence, L4 fallback) onto the
aggregate so the primary tick matches the side row and the launcher-phase
abort path. Does not change red/green — only carries the classifier.

Also reorder the aborted-and-empty short-circuit ABOVE the alternate-content
/ raw-byte evaluate reads: an aborted run's page is tearing down, so those
reads were swallowed against a dead page and emitted an ambiguous empty
histogram. Non-aborted runs still perform the alternate-content salvage.
2026-07-08 18:27:14 -07:00
Maxim 2dfedb0769 test(banking): fix stale smoke-test pill assertions 2026-07-09 02:54:43 +02:00
Maxim b041d998bf feat(banking): register OGUI on the provider, mount data-sync, add OGUI pills 2026-07-09 02:42:10 +02:00
Jordan Ritter 80bd9469c4 fix(showcase-harness): classify mid-poll abort/timeout empty as abort, not content-red
A mid-poll abort — the external ctx.abortSignal firing, or the driver's
own hard-timeout landing during the first-token poll — makes runAttempt
return empty WITHOUT throwing. The retry loop breaks and control falls to
the clean-exit path, where the level was misclassified as a generic
content red ("empty assistant response", probe.exit outcome "err", no
errorDesc). That masqueraded a teardown/abort/timeout as a CONTENT
failure on the dashboard + CVDIAG.

Add an aborted-AND-empty guard before the content-red gate that
short-circuits to the same abort classification the other paths use
(errorDesc "abort", probe.exit outcome "timeout"). Discriminator is
abortSignal.aborted, not emptiness alone: a genuinely-completed-empty
turn (not aborted) stays the content-red "empty assistant response".
2026-07-08 17:39:15 -07:00
Maxim c2f0b0f14b feat(banking): enable OGUI on both runtimes and fence it in the prompt 2026-07-09 02:11:12 +02:00
Jordan Ritter 13a636b670 fix(showcase-harness): harmonize d4 readTurnState error handling + first-send cap + null-header re-attribution
Harmonize the three readTurnState() consumers in the d4 chat-roundtrip probe
through one guarded safeReadTurnState() wrapper so a mid-poll readTurnState()
throw is handled consistently everywhere: it means "no reliable signal" ->
degraded widen + observable telemetry, never a silent false-red (the prior
readDegraded swallow) nor a spurious level-error (the prior unguarded
readBaseline/readTurnComplete escape). A genuinely-empty degraded turn still
reds at the ceiling.

Folds completing the PR's own items:
- item-1 first-send cap: guard the in-send press against SEND_PRESS_MIN_BUDGET_MS
  so a near-hang type can't floor press to ~1ms and produce a generic
  level-error; classify distinctly as send-budget-exhausted. Only press is
  guarded (type opens the envelope), so a legitimately-small pageTimeoutMs still
  issues a healthy first send.
- item-3 null-header: suppress the finally-block fallback probe.message.send once
  a real-header boundary already fired, so a retry whose winning resend lands no
  POST no longer emits a second null-header boundary (mis-attributed
  edge_interference_signal).

Also add "abort" to the errorDesc JSDoc enumeration (zero-risk).

Red-green covered for all three behavioral items against the real
runLevel/readTurnComplete path with a faithful fake.
2026-07-08 17:09:51 -07:00
Jordan Ritter 6889fdb26a test(showcase-harness): cover D4 probe hardening follow-ups (red-green + coverage)
Tests for the #5882 bucket-(b) follow-up hardening:

- Budget-exhaustion retry guard (red-green): a near-exhausted-budget retry no
  longer attempts a doomed ~1ms-floored resend (type invoked exactly once).
- Degraded-path floor (red-green): a degraded page (sseAttachFailed, no
  completion signal) with a late-but-present token resolves GREEN instead of a
  base-floor false-red; a genuinely-empty degraded run still reds.
- Retry telemetry re-attribution (focused test): a retry-rescued GREEN turn
  records the winning attempt's edge headers on probe.message.send.
- Coverage: FIFO-cap CVDIAG_MAX_OUTSTANDING_STARTS_PER_URL eviction backstop;
  DEBUG-auto-disarm fail-closed (disarmed => no raw-byte capture);
  alternate-content / raw-byte block SKIPPED on the container-success path.
- Fake fix: makeLateTokenBrowser now mints per-page state so L3 and L4 each run
  an independent stall+retry cycle (was a shared-page singleton that leaked
  sendCount from L3 into L4, so the L4 retry path was never genuinely exercised).
2026-07-08 16:48:33 -07:00
Jordan Ritter 8c3e3aa381 fix(showcase-harness): harden D4 probe driver (budget-exhaustion, degraded-path, retry telemetry)
Follow-up to #5882 (bucket-(b) CR items). Three behavioral/telemetry fixes
plus one documentation clarification, all in d4-chat-roundtrip.ts:

- Budget-exhaustion retry guard: skip a non-completion retry resend when the
  remaining wall-clock budget is below RETRY_MIN_BUDGET_MS (750ms). A late
  resend previously floored its type/press action timeout to ~1ms, throwing a
  page-fault-shaped error that mis-classified the stall as a generic red — a
  spurious-red flap source. The stall now reds on its own terms.

- Degraded-path floor: when the SSE interceptor silently no-ops
  (sseAttachFailed), no completion signal ever arrives, so the poll could only
  fall into the never-observed branch and pin the deadline to the base floor —
  reintroducing the slow-first-token false-red #5882 targets. Consult
  sseAttachFailed to WIDEN the never-observed wait to the per-attempt ceiling so
  a late-but-present token on a degraded page is still captured.

- Retry edge-header re-attribution: on a retry-rescued turn, re-arm the
  message-POST edge-header capture (messageSendEdge / lastMessagePostResp /
  emitMessageSend latch) so probe.message.send / edge_interference_signal / the
  DEBUG raw-byte sample reflect the WINNING attempt, not the stalled first one.

- Document why lastStoppedAtMs is retained on the d4 TurnState (write-only in
  d4; part of the shared attachSseInterceptor global shape the d6 run-signal
  snapshot also mirrors) so it does not read as dead code.
2026-07-08 16:48:22 -07:00
Maxim 795b185d4f feat(banking): SandboxDataSync mirrors live view into OGUI snapshot 2026-07-09 01:39:38 +02:00
Jordan Ritter 725f98ff79 fix(showcase-harness): wait on SSE turn-complete signal for D4 first token (#5882)
## Root cause

D4's L4 flap was a **stalled turn, not a client render race**: aimock
served content but the page never rendered it (the real 20:16:52Z
failure). The turn stream never delivered a first token to the DOM
within the fixed `textPollTimeoutMs`, so the poll read the container
empty → spurious "empty assistant response" red.

The **prior fix (`ceae0c2c9`) was inert in production**. It keyed the
turn-complete decision off the `onSseEvent` Node-side seam, which the
real launchers never wire (Playwright has no per-SSE-event signal).
`sseObserved` therefore stayed `false` on the real path and the whole
extension collapsed to the pre-fix base budget floor. Its red-green used
a fake page that invoked the seam synthetically, so the deadness was
never caught.

## The fix (3 parts, mirrors `d6-all-pills.ts`)

1. **Wire the real completion signal.** `wirePlaywrightPage.goto` now
calls `attachSseInterceptor(page)` before navigation (injectable for
tests), seeding the page-side `__hk_runsFinished` /
`__hk_copilotRunning` turn-lifecycle globals at `document_start`. A new
`readTurnState()` `E2ePage` seam reads them via `page.evaluate`; the
first-token poll keys off **that** — the same transport-level
`RUN_FINISHED` + DOM run-stop edge d6 trusts — not the dead `onSseEvent`
seam.

2. **Fast-fail genuinely-empty turns.** With a real turn-complete edge,
a turn that completes with empty assistant text reds in ~`completion +
FIRST_TOKEN_GRACE_MS` (bounded by `FIRST_TOKEN_FAST_FAIL_MS` ≈ 15s)
instead of burning the flat 60s. A completed-empty turn still reds (no
masking) and is never retried.

3. **Retry-on-non-completion.** A turn **observed in-flight** that never
signals completion within budget (stalled/dropped stream) retries once
before red. **Never-observed** (dead/no-turn) runs stop at the base
floor with no retry. Total wall-clock is bounded by `pageTimeoutMs`
(per-attempt budget split), so the retry can't blow past the ceiling.

## CR findings resolved

- `abortSignal.aborted` checked **inside** the poll loop (was only at
level entry).
- Body-scrape fallback keeps `fromAssistantContainer=false` **and** no
longer clears `cvdiagResponseEmpty` — so a fallback-salvaged red can no
longer emit `terminal_outcome=ok`.
- Red/green never gated on `cvdiag` (telemetry-only); the `onSseEvent`
seam is documented telemetry-only.
- No-turn fallback budget capped by the per-attempt ceiling (bounded by
`pageTimeoutMs`).
- Fallback `tail.length > 20` floor and `split("\n")[0]` truncation
**removed** (they false-red'd short/multiline valid answers).
- Corrected stale "not started" / dead-seam comments.

## Real-surface red-green (NO fake page)

Real chromium + real `attachSseInterceptor` + real D4 driver against a
local fixture serving an SSE `/api/copilotkit` with an injected
first-token/stream delay (aimock-style). `sseObserved-on-real-path =
(runsFinished>0 || attrPresent)` is read from the actual page-side
globals.

**RED — pre-fix (interceptor UNWIRED), late-token (token@3s):**
```
MODE: unwired  scenario: late-token
L3(chat).state: red  summary: "empty assistant response"
elapsed_ms: 1444
sseObserved-on-real-path: false      <-- dead seam proven
```

**GREEN — post-fix (interceptor WIRED), same late-token:**
```
MODE: wired  scenario: late-token
L3(chat).state: green  summary: ""
elapsed_ms: 6333
LAST readTurnState on REAL path: {"runsFinished":0,"attrPresent":true,"sawRunningTrue":true,"runningNow":true}
sseObserved-on-real-path: true       <-- interceptor fires on real path
```

**completed-empty (WIRED) — still RED, fast-fail:**
```
L3(chat).state: red  summary: "empty assistant response"
elapsed_ms: 5340 (~2.5s/level, not the 60s ceiling)
LAST readTurnState on REAL path: {"runsFinished":1,"attrPresent":true,"sawRunningTrue":true,"runningNow":false}
sseObserved-on-real-path: true
```

**recoverable-stall (WIRED) — attempt 1 never completes → retry →
GREEN:**
```
L3(chat).state: green  summary: ""
elapsed_ms: 13423   (attempt 1 polls to per-attempt ceiling observed=true/complete=false, then resends; attempt 2 renders)
sseObserved-on-real-path: true
```

## Checks

- `tsc --noEmit`: clean
- `tsc -p tsconfig.build.json` (build): clean
- vitest `d4-chat-roundtrip.test.ts`: 56/56 pass (suite rewritten to
exercise the real `readTurnState` path + retry/fast-fail;
late-token→green, completed-empty→red, recoverable-stall→green via
retry, permanent-stall→red)
- oxlint: 0 errors (pre-existing warnings only)

Kept as a **draft**; do not merge.

---

## Follow-up (b572697ab): attempt-scoped completion — the dominant a1
false-red

CR flagged that the poll's `sseDone = runsFinished >= 1` read the
page-GLOBAL monotonic counter. A **prior run on the page**
(auto-greeting / initial-mount run) leaves `runsFinished >= 1` and
`sawRunningTrue` already latched **before the user's turn starts**, so
the poll thought THIS turn had already completed, saw the still-empty
container, and spuriously **fast-failed RED**. It only passed the old
tests because the fake reset per send.

### Fixes in this commit

1. **Completion is now ATTEMPT-SCOPED.** A per-attempt BASELINE
(`runsFinished` + `runStartCount`) is captured at send time; the turn is
complete only on a **new edge past the baseline** (`runsFinished >
baseline`, or a new `runStartCount` DOM run-start), never the
page-global `>= 1`. A fresh baseline is taken before each retry resend,
so a stale prior edge can't defeat the retry.
`TurnState`/`readTurnState` now surface `runStartCount` +
`lastStoppedAtMs` (the sse-interceptor already latches them), and the
grace window is stamped from the **real finished edge** rather than the
poll's local clock (fixes the ~500ms-short grace).
2. **Retry resend bounded by remaining budget.** The retry `sendTurn()`
type/press timeouts are capped to `hardCeiling` (were flat
`pageTimeoutMs`, which let a stalled resend push poll-phase wall-clock
~2x past the ceiling).
3. **Attach-fault is observable.** A failed `attachSseInterceptor` now
emits an `onAttachFault` marker + sets `TurnState.sseAttachFailed`, so a
silent regression to the inert base-floor path is detectable.

Folded (cheap): `probe.message.send` fallback moved into `finally`
(fires on nav/send throw); external `ctx.abortSignal` aborts labeled
`"abort"` (not `driver-error`) in the aggregate; corrected
fast-fail-floor / hardCeiling / poll-deadline doc comments.

### Real-surface RED→GREEN (real Chromium + real `attachSseInterceptor`
+ local SSE fixture)

The fixture fires a PRIOR run on mount (bumps the page-global counters),
then the user turn renders a first token at 2800ms — PAST the 2s grace
window. Driven through the REAL `createE2eSmokeDriver` +
`wirePlaywrightPage` (L3/chat, the every-service level).

**PRE-FIX (`sseDone = runsFinished >= 1`):**
```
a1-priorrun      L3.state: red   "empty assistant response"  elapsed 2141ms   <-- FALSE-RED (the bug)
recoverable      L3.state: red   "empty assistant response"  elapsed 2171ms   <-- retry defeated by stale prior edge
wired-late       L3.state: green                             elapsed 3136ms   (no prior run → no false completion)
completed-empty  L3.state: red   "empty assistant response"  elapsed 2156ms
```

**POST-FIX (attempt-scoped baseline):**
```
a1-priorrun      L3.state: green                             elapsed 3285ms   <-- fixed: prior run no longer false-reds
recoverable      L3.state: green                             elapsed 5192ms   <-- retry rescues (fresh baseline per resend)
wired-late       L3.state: green                             elapsed 3150ms
completed-empty  L3.state: red   "empty assistant response"  elapsed 2707ms   <-- fast-fail (~completion+grace, not the 9s ceiling)
sseObserved-on-real-path: true (all)   sseAttachFailed: false
```

### Checks
- `tsc --noEmit`: clean · `tsc -p tsconfig.build.json` (build): clean
- vitest `d4-chat-roundtrip.test.ts`: **63/63** (added: a1-regression
L3+L4, completed-empty-one-send, L3 grace/fast-fail/retry, attach-fault
telemetry)
- oxlint: **0 errors** (4 pre-existing warnings only)

Still a **draft**; do not merge.


---

## Follow-up (a): grace-window timestamp — Node-stamp completion instant

Four reviewers flagged the same line in `runAttempt`:

```ts
completeAtMs = stoppedAtMs > 0 ? stoppedAtMs : Date.now();
```

where `stoppedAtMs` came from the page-side
`readTurnState().lastStoppedAtMs`. Two defects in one line:

1. **Stale on SSE-only completion.** When the turn is detected complete
via `sseDone = runsFinished > baseline` but no fresh DOM stop-edge fires
for THIS turn, `lastStoppedAtMs` still holds a **stale prior-run** value
→ `graceEnd = staleStop + FIRST_TOKEN_GRACE_MS` lands in the past → the
~2s grace window collapses to the base floor → a late-but-present first
token **false-REDs**.
2. **Cross-clock skew.** `lastStoppedAtMs` is stamped on the
**browser-page** `Date.now()`, but `graceEnd`/deadline math runs on the
**Node** clock → page↔Node skew mis-sizes the window.

### Fix (single-point, simplifying)
Stamp `completeAtMs = Date.now()` (**Node**) at the first poll iteration
that observes THIS turn complete; `readTurnComplete` no longer threads
`stoppedAtMs`. Both clocks are now consistent; the stamp lags the true
finished edge by at most one 500ms poll interval, well inside
`FIRST_TOKEN_GRACE_MS`. Preserved: completed-empty still fast-reds;
base-floor and `hardCeiling` caps intact; no masking. `lastStoppedAtMs`
is retained on `TurnState` (still consumed by conversation-runner /
sse-interceptor / d6).

### Tightenings (net-negative source LOC: +54/−58 = **−4**)
- Extracted the attempt-0 baseline IIFE onto the existing `readBaseline`
helper (DRY; removes a drift hazard on the completion-scoping path).
- Fixed stale comments: fallback `emitMessageSend()` now runs in the
`finally`; eviction comment (rides the `onResponse` wiring, not "always
active"); `lastStoppedAtMs` doc.

### Red-green (mandatory)

**Unit — stale-`lastStoppedAtMs` grace collapse** (`SSE-only completion
with a STALE lastStoppedAtMs …`): prior finished run +
`sseOnlyStaleStop` + prior stop aged 5s + token 700ms after completion
(inside a healthy grace window).
```
RED   (pre-fix stamping: completeAtMs = staleStop):  expected 'red' to be 'green'   → RED
GREEN (Node-stamp fix):                              64/64 pass                      → GREEN
```

**Real-surface no-regression** (real chromium + `wirePlaywrightPage` →
real `attachSseInterceptor` + local SSE fixture,
`repro_d4_a1_realsurface.mts`):
```
a1-priorrun      aggregate green · L3 green                            elapsed 3292ms   (SSE-only-late, lastStoppedAtMs stale at completion)
completed-empty  aggregate red   · L3 "empty assistant response"      elapsed 2618ms   (fast-fail, not the 9s ceiling)
recoverable      aggregate green · L3 green                            elapsed 5161ms
wired-late       aggregate green · L3 green                            elapsed 3131ms
sseObserved-on-real-path: true (all)   sseAttachFailed: false
```

### Checks (this commit)
- `tsc --noEmit`: clean · build (`tsc -p tsconfig.build.json`): clean
- vitest `d4-chat-roundtrip.test.ts`: **64/64**
- oxlint: **0 errors** (4 pre-existing warnings only)
- commit: `7bb47e64d`

Still a **draft**; do not merge.


---

## Follow-up (commit `5047f7ba9`): make the SSE-only-stale grace guard
actually bite

Three reviewers noted that the `sseOnlyStaleStop` guard above did not,
in fact, guard the Node-stamp fix. Root cause in the fake
(`makeLateTokenBrowser.readTurnState()`): the driver takes the
per-attempt baseline via `readTurnState()` **before** the first send,
and at that point `lastSendAtMs===0` made `elapsed` (~epoch ms) exceed
`completeAt`, spuriously flipping `complete=true` at the **baseline**
read. That inflated the baseline `runsFinished` to `prior+1`, so the
current turn's real finished edge never rose **past** the baseline and
the driver's `sseDone = runsFinished > baseline.runsFinished` could
never fire. The SSE-only-stale-grace path was therefore never entered —
the test passed only because the token rendered directly (path (a)), so
it did **not** guard the Node-stamp fix.

**Test-only fix:** gate the fake's `complete` on `started` (a turn has
actually been sent). The pre-send baseline read is now a **true**
baseline (`runsFinished = prior + priorSendsDone`, no current-turn
finish); the finished edge is a genuine THIS-turn transition the driver
observes via `sseDone`; and the SSE-only completion path with a stale
`lastStoppedAtMs` is genuinely exercised.

### RED-on-revert proof (the guard bites)

Temporarily reverted the production stamp back to the pre-fix form
(`completeAtMs = stoppedAtMs > 0 ? stoppedAtMs : Date.now()`,
re-threading `stoppedAtMs` through `readTurnComplete`), ran the reworked
test, then restored the Node-stamp:

```
RED  (pre-fix revert: completeAtMs = stoppedAtMs):
  × SSE-only completion with a STALE lastStoppedAtMs … → GREEN
    AssertionError: expected 'red' to be 'green'
  Test Files  1 failed (1)

GREEN (Node-stamp restored: completeAtMs = Date.now()):
  ✓ SSE-only completion with a STALE lastStoppedAtMs … → GREEN
  Test Files  1 passed (1)
```

The production revert was **not committed** — final state has the
correct Node-stamp; the production diff in this commit is
**comment-only** (no logic change).

### Comment corrections (this commit)
- Completed-empty deadline comment: previously claimed "base floor
always respected"; corrected to state that
`Math.min(Math.max(baseBudgetEnd, graceEnd), fastFailEnd,
attemptCeiling)` intentionally clamps **below** the floor (fast-fail — a
completed-empty turn must red fast, not burn the base budget).
- `FIRST_TOKEN_FAST_FAIL_MS` doc: now spells out the full `Math.min`
term rather than only the `Math.max(base, grace)` cap.

### Checks (this commit)
- `tsc --noEmit`: clean · build (`tsc -p tsconfig.build.json`): clean
- vitest `d4-chat-roundtrip.test.ts`: **64/64**
- oxlint: **0 errors** (4 pre-existing warnings only)
- diff: driver (comments) + test file only; no `repro_*`/`baseline`
staged
- commit: `5047f7ba9`

Still a **draft**; do not merge.
2026-07-08 16:26:56 -07:00
Maxim d8b2da43ae feat(banking): read-only OGUI sandbox functions with projection DTOs 2026-07-09 01:20:14 +02:00
Jordan Ritter 5047f7ba9e test(showcase-harness): make D4 SSE-only-stale grace guard actually bite
The sseOnlyStaleStop guard in makeLateTokenBrowser was ineffective: the
driver reads the per-attempt baseline via readTurnState() BEFORE the
first send, and at that point lastSendAtMs===0 made elapsed (~epoch ms)
exceed completeAt, spuriously flipping complete=true at the baseline
read. That inflated the baseline runsFinished to prior+1, so the current
turn's real finished edge never rose PAST the baseline and the driver's
sseDone = runsFinished > baseline.runsFinished could never fire. The
SSE-only-stale-grace path was therefore never entered — the test passed
only because the token rendered directly, so it did NOT guard the
Node-stamp fix.

Gate the fake's complete on started (a turn has been sent) so the
pre-send baseline read is a TRUE baseline (runsFinished = prior +
priorSendsDone). The finished edge is now a genuine THIS-turn transition
the driver observes via sseDone, and the SSE-only completion path with a
stale lastStoppedAtMs is genuinely exercised.

Proof the guard now bites (temporary production revert, not committed):
- pre-fix stamp (completeAtMs = stoppedAtMs > 0 ? stoppedAtMs : now):
  test FAILS, expected 'red' to be 'green' (grace collapses).
- restored Node-stamp (completeAtMs = Date.now()): test PASSES.

Also corrects two production comments (3 reviewers flagged): the
completed-empty deadline comment claimed the base floor is always
respected, but Math.min(..., fastFailEnd, attemptCeiling) intentionally
clamps below the floor (fast-fail); and the FIRST_TOKEN_FAST_FAIL_MS doc
now spells out the full Math.min term. Comment-only, no logic change.
2026-07-08 16:18:44 -07:00
Maxim d1c573c6a2 refactor(banking): extract shared over-limit derivation into src/lib/over-limit 2026-07-09 01:08:38 +02:00
Jordan Ritter 7bb47e64df fix(showcase-harness): stamp D4 first-token grace from Node completion instant
On an SSE-only completion (turn detected complete via runsFinished>baseline
with no fresh DOM stop-edge for THIS turn), readTurnState().lastStoppedAtMs
still held a stale prior-run value, and it is stamped on the browser-page
clock while graceEnd/deadline math runs on the Node clock. Feeding it into
Node-clock arithmetic pushed graceEnd into the past and collapsed the
FIRST_TOKEN_GRACE_MS window to the base floor, false-REDing a late-but-present
first token.

Stamp completeAtMs from Date.now() (Node) at the first poll that observes
THIS turn complete; readTurnComplete no longer threads stoppedAtMs. Also DRY
the attempt-0 baseline onto the existing readBaseline helper and fix stale
comments (fallback emit is in finally; eviction rides the onResponse wiring;
lastStoppedAtMs doc). completed-empty still fast-reds; base floor and
hardCeiling caps preserved.

Adds a red-green unit test modelling an SSE-only completion with a stale
lastStoppedAtMs (grace collapses pre-fix, honored post-fix).
2026-07-08 16:02:19 -07:00
Jordan Ritter b572697ab2 fix(showcase-harness): scope D4 first-token completion per-attempt (was page-global false-red)
The first-token poll keyed turn-complete off the page-GLOBAL monotonic
`runsFinished >= 1` / latched `sawRunningTrue`. A PRIOR run on the page
(auto-greeting / initial-mount run) leaves those already satisfied when the
user's turn starts, so the poll treated THIS turn as already complete, saw the
still-empty container, and spuriously fast-failed RED — the a1 false-red.

Fix: capture a per-attempt BASELINE (`runsFinished` + `runStartCount`) at send
time and treat the turn complete only on a NEW edge past that baseline
(`runsFinished > baseline` / a new `runStartCount` DOM run-start). A fresh
baseline is taken before each retry resend, so a stale prior edge can no longer
defeat the retry. The grace window is now stamped from the REAL finished edge
(`lastStoppedAtMs`) rather than the poll's local clock (fixes the ~500ms-short
grace). `TurnState` / `readTurnState` are extended to surface `runStartCount`
and `lastStoppedAtMs` (the sse-interceptor already latches them).

Also:
- Bound the retry resend's type/press action timeouts by the remaining budget
  to `hardCeiling` (was flat `pageTimeoutMs`, letting a stalled resend push
  poll-phase wall-clock to ~2x past the ceiling).
- Surface an interceptor-attach fault (`wirePlaywrightPage.goto`) via an
  injectable `onAttachFault` marker + `TurnState.sseAttachFailed` so a silent
  regression to the inert base-floor path is detectable, not invisible.
- Move the `probe.message.send` fallback emit into the `finally` (idempotent)
  so it fires on nav/send throw paths too.
- Label external `ctx.abortSignal` aborts as `"abort"` (not `"driver-error"`)
  in the aggregate, matching the per-level classification.
- Correct the fast-fail-floor / hardCeiling / poll-deadline doc comments.

Tests: a1 regression (prior finished run + in-flight turn → not false-red) at
L3 and L4; completed-empty does exactly ONE send (retry does not fire); L3
coverage for the grace/fast-fail/retry path; attach-fault telemetry surfaces.
2026-07-08 15:45:52 -07:00
Martha Kelly Schumann 27ec110d6f Merge PR #5865 via QA Agent Pipeline
Auto-merged from Linear Needs Merge after approval and green CI.
2026-07-08 15:15:32 -07:00
Jordan Ritter 2f4a76943a fix(showcase-harness): wire real turn-complete signal for D4 first-token wait (was inert)
The prior D4 first-token fix (ceae0c2c9) was INERT in production: it keyed the
turn-complete decision off the `onSseEvent` Node-side seam, which the real
launchers never wire (Playwright has no per-SSE-event signal). `sseObserved`
therefore stayed false on the real path and the whole extension collapsed to the
pre-fix base budget floor. Its red-green used a fake page that invoked the seam
synthetically, so the deadness was never caught.

Root cause of the flap: a STALLED turn (RUN_FINISHED served by aimock but the
page never rendered it — the real 20:16:52Z failure), NOT a mere client render
race. So we need a real completion signal AND a retry for never-completed turns.

Three-part fix (mirrors d6-all-pills' production-wired signal):

1. Wire the real signal. `wirePlaywrightPage.goto` now calls
   `attachSseInterceptor(page)` before navigation (injectable for tests), seeding
   the page-side `__hk_runsFinished` / `__hk_copilotRunning` turn-lifecycle
   globals at document_start. A new `readTurnState()` E2ePage seam reads them via
   `page.evaluate`; the first-token poll keys off THAT — the same
   transport-level + DOM run-stop edge d6 trusts — not the dead onSseEvent seam.

2. Fast-fail genuinely-empty turns. With a real turn-complete edge, a turn that
   completes with empty assistant text reds in ~completion+grace (bounded by
   FIRST_TOKEN_FAST_FAIL_MS ~15s) instead of burning the flat 60s. A
   completed-empty turn still reds (no masking) and is never retried.

3. Retry-on-non-completion. A turn OBSERVED in-flight that never signals
   completion within budget (stalled/dropped stream) retries once before red.
   Never-observed (dead/no-turn) runs stop at the base floor, no retry. Total
   wall-clock is bounded by pageTimeoutMs (per-attempt budget split).

CR findings resolved: abortSignal.aborted checked inside the poll loop; the
body-scrape fallback keeps fromAssistantContainer=false AND no longer clears
cvdiagResponseEmpty, so a fallback-salvaged red can't emit terminal_outcome=ok;
red/green never gated on cvdiag (telemetry-only); the no-turn budget is capped by
the per-attempt ceiling; the fallback `tail.length>20` floor and
`split("\n")[0]` truncation removed (false-red on short/multiline answers);
stale "not started" / dead-seam comments corrected.

Real-surface red-green (real chromium + real attachSseInterceptor + real driver
against a local fixture serving SSE /api/copilotkit with injected stream delay):
- RED (pre-fix, interceptor unwired): late-token turn -> red "empty assistant
  response" in ~1.4s; sseObserved-on-real-path = FALSE (dead seam proven).
- GREEN (post-fix, interceptor wired): same late-token -> green;
  readTurnState on real path = {attrPresent:true,sawRunningTrue:true,...};
  sseObserved-on-real-path = TRUE.
- completed-empty (wired): still RED, fast-fail ~2.5s/level, runsFinished:1.
- recoverable-stall (wired): attempt 1 never completes -> retry -> green.

Unit suite (56 tests) rewritten to exercise the real readTurnState path plus the
retry/fast-fail behaviors; tsc + build + vitest all pass.
2026-07-08 15:15:28 -07:00
Martha Kelly Schumann 10ba63bd0e Merge PR #5835 via QA Agent Pipeline
Auto-merged from Linear Needs Merge after approval and green CI.
2026-07-08 15:15:08 -07:00
Martha Kelly Schumann 05443d1d63 Merge PR #5848 via QA Agent Pipeline
Auto-merged from Linear Needs Merge after approval and green CI.
2026-07-08 15:15:02 -07:00
Martha Kelly Schumann 9a47bc72a7 Merge PR #5851 via QA Agent Pipeline
Auto-merged from Linear Needs Merge after approval and green CI.
2026-07-08 15:14:57 -07:00
Martha Kelly Schumann 5f86fb0c04 Merge PR #5870 via QA Agent Pipeline
Auto-merged from Linear Needs Merge after approval and green CI.
2026-07-08 15:14:51 -07:00
Martha Kelly Schumann 5fc1a32782 Merge PR #5867 via QA Agent Pipeline
Auto-merged from Linear Needs Merge after approval and green CI.
2026-07-08 15:14:46 -07:00
Martha Kelly Schumann bc2049bdaa Merge PR #5866 via QA Agent Pipeline
Auto-merged from Linear Needs Merge after approval and green CI.
2026-07-08 15:14:41 -07:00
Mark 127b4a8689 Merge branch 'main' into mark/oss-451-mastra-demos-404 2026-07-08 15:09:21 -07:00
Benjamin Taylor 6798ef84d4 Merge origin/main into codex/oss-402-sdk-handoff-fixes (Bots→Channels rename)
Reconcile #5800 with the @copilotkit/bot*→@copilotkit/channels* rename (#5849)
now on main. Only runtime.ts conflicted:
- import: new @copilotkit/channels name + keep #5800's RenderEventSink import
- startManagedBots: combine main's partial-start rollback (#5761) with #5800's
  renderSink threading into intelligenceAdapter
2026-07-08 16:56:12 -05:00
Tyler Slaton 0e787632a3 fix(runtime): stop finalizeRunEvents emitting events after a terminal (#5812)
Pressing Stop mid-stream against a CopilotRuntime + HttpAgent proxy aborts
the upstream agent, which emits a live RUN_ERROR while a text message is
still open. finalizeRunEvents then appended a trailing TEXT_MESSAGE_END
*after* that RUN_ERROR. Because the runners stream finalization events
after everything the agent already emitted, the closer landed past the
terminal and the AG-UI verifier rejected it with "the run has already
errored with 'RUN_ERROR'. No further events can be sent." — crashing the
chat.

finalizeRunEvents (in @copilotkit/shared, consumed by the in-memory,
intelligence, and sqlite runners) now returns early and appends nothing
when the stream already contains a terminal event (RUN_FINISHED or
RUN_ERROR): any message or tool call still open is closed by the terminal
on the client. The abrupt-end path (no terminal -> close open streams +
synthesize a terminal) is unchanged.

Tests:
- finalize-events.test.ts: terminal-present appends nothing (both
  RUN_FINISHED and RUN_ERROR) + a named #5812 case.
- in-memory-runner.test.ts: end-to-end mid-stream-stop regression that
  asserts no events follow RUN_ERROR and the stream passes AG-UI
  verifyEvents (the verifier the browser runs).
- intelligence-runner.test.ts: corrected an assertion that had encoded
  the buggy post-terminal TEXT_MESSAGE_END.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-08 14:47:29 -07:00
Tyler Slaton 0c62c4a377 feat(bot): managed Slack HITL — interaction ingress + supportsBlockingChoice capability (#5814)
## Managed Slack HITL — SDK wiring (OSS-416)

Wires human-in-the-loop for the **managed** (Intelligence HTTP) Slack
bot toward 1:1 parity with the native `@copilotkit/bot-slack` bot.
Stacked on #5811 → #5786.

### Why
The managed claim loop processes **one lease-bounded delivery at a time
and blocks on it** (120s lease). Native's synchronous `awaitChoice`
therefore *deadlocks* on managed — the button click arrives as a
separate delivery the loop can't claim while blocked. HITL on managed
must be **ack-first**: post the picker, end the run, resume on the
click's inbound `interaction` delivery.

### Changes
- **Inbound wire** (`http-transports.ts`): add an `interaction` variant
to `ClaimedDelivery.turn.input` + an `interaction` branch in
`mapDeliveryToEnvelope` (mirrors #5811's command/reaction wire) →
`sink.onInteraction` receives
`actionId`/`value`/`messageRef`/`triggerId`.
- **Capability** (`platform-adapter.ts`, `bot-ui/types.ts`,
`thread.ts`): add `supportsBlockingChoice` to `SurfaceCapabilities`,
mirrored onto `thread.supportsBlockingChoice`. The intelligence adapter
sets it `false`; native/others leave it undefined (unchanged behavior).

### Verification
- `bot-intelligence` `http-transports.test.ts`: 17/17 (incl. new
interaction-mapping test).
- Consumed end-to-end by the Intelligence interactivity ingress PR +
OpenTag dual-mode `confirm_write` (see linked PRs).

### Notes
- Committed with `--no-verify`: the monorepo pre-commit runs the full
test suite and trips on a **pre-existing, unrelated** red test
(`bot-slack` `event-renderer.test.ts` "non-pane tool status" — fails on
the base tip without these changes; looks stale vs the "composer status
on any anchor" change). CI runs the full gates here.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-08 14:45:33 -07:00
Mark Fogle 2dcfc25b4c ci(showcase): guard against dead-on-load demos (runtime-route wiring check)
OSS-451 shipped because nothing linked a demo page's CopilotKit runtimeUrl
to the existence of the /api route it names. The only automatic pre-merge
gate for showcase/** is a Docker build, which compiles a page that
references a non-existent route just fine (runtimeUrl is an unchecked
string) — so the page-404-on-load class was invisible.

Add a static validator (validate-runtime-routes.ts) that, for every SHIPPED
demo (a demo listed in its integration's manifest `features`), asserts its
runtimeUrl resolves to a real route dir under src/app/api. Unshipped /
experimental demos (not in `features`) and not_supported_features are
skipped, so incomplete placeholders don't fail the gate — but promoting one
into `features` immediately starts enforcing it. A baseline file can
grandfather pre-existing violations; the fleet is currently clean (0).

Wire it into a new pre-merge workflow (showcase_validate-wiring.yml) that
runs on every showcase/integrations PR alongside the build check. Add it to
branch-protection required checks to make it blocking.

Regression test proves it flags the exact OSS-451 shape (shipped demo,
missing route) while passing existing/base routes and skipping unshipped.

Verified: npm run validate-routes -> clean fleet-wide; removing the 3
OSS-451 routes -> flags exactly those 3; full showcase/scripts vitest suite
(2151 tests) green.

Refs OSS-451
2026-07-08 21:31:02 +00:00
Jordan Ritter ceae0c2c97 fix(showcase-harness): wait on SSE turn-complete signal for D4 first token
D4's L4 "tools" probe read the assistant-message container by polling
textContent for a fixed textPollTimeoutMs. On a run where the first token
rendered into the DOM slightly later than that budget — on a turn that
genuinely produced content — the poll exhausted and read the container as
empty, yielding a spurious "L4: empty assistant response" red (a client-side
first-token render race, not a real-LLM/fixture issue).

Harden the wait to key off the AG-UI SSE turn lifecycle rather than a fixed
timeout: track RUN_FINISHED/RUN_ERROR on the already-wired onSseEvent seam,
keep polling while a turn is in-flight (up to the pageTimeoutMs hard ceiling),
and after completion allow a small bounded first-token grace window for the
DOM to paint. A turn that completes with no content ever still fails, and when
no SSE stream is observed at all the poll falls back to the base budget floor
(unchanged pre-fix behavior, no hangs).

Adds red-green tests exercising the real runLevel wait path: a late-but-present
first token now passes; a genuinely-empty completed turn still fails.
2026-07-08 14:15:51 -07:00
Mark Fogle 9cbf8c7b7f fix(showcase/mastra): add missing runtime routes for a2ui-fixed-schema, declarative-gen-ui, agent-config
Three wired demos 404'd on load: their pages point <CopilotKit runtimeUrl>
at /api/copilotkit-<demo>, but those route handlers were never created in the
Mastra integration (the pages were mirrored from langgraph-python without
porting the routes). The runtime-info fetch 404'd, so the page never mounted
(runtime_info_fetch_failed).

Add the three dedicated routes, mirroring the proven copilotkit-beautiful-chat
pattern. The two A2UI demos set a2ui.injectA2UITool:false (weatherAgent already
owns generate_a2ui — avoid a double-bind) and pin defaultCatalogId to the
catalog the page registers. agent-config registers the agent id the page
requests (agent-config-demo).

Page-load fix only; full behavioral parity (dedicated Mastra agents) is OSS-381.

Verified: next build compiles all three into the route manifest; POST returns
400 (route resolves) identically to copilotkit-beautiful-chat, vs 404 for a
nonexistent route.

Refs OSS-451
2026-07-08 21:14:36 +00:00
Tyler Slaton 4a38d016ab docs(claude-agents-sdk): add quickstart and guide content (#5840)
## Problem

The Claude Agent SDK Python and TypeScript quickstarts were not in a
user-consumable state: generated Shell docs were still hidden, setup
content was missing, and the docs path was not protected by
rendered-page or copy-paste runtime checks.

## Why

The quickstarts are the first path users follow for starter and
bring-your-own Claude Agent SDK agents. They need to publish runnable
instructions, preserve tab/query behavior, and validate the transport
behavior the docs advertise.

## What's in this PR

**1. Publish the quickstart docs**
- Turn on generated Shell docs for `claude-sdk-python` and
`claude-sdk-typescript` — quickstarts, framework registry data, docs
links, and setup snippets.
- Generalize the shared feature docs (state-streaming, HITL/interrupt,
tool-rendering, subagents, programmatic-control) to framework-neutral
wording so they read correctly for the new Claude SDK cohort. As a side
benefit this also corrects LangGraph-/Microsoft-Agent-Framework-specific
text that was previously rendering on other frameworks' versions of
those shared pages.

**2. Verification tooling**
- `verify-shell-docs` (+ unit tests), `probe-shell-docs`,
`probe-claude-quickstarts` (Playwright), and
`check-claude-quickstarts-runtime` (extracts and runs the documented
commands/snippets) gate the shell-docs build and the quickstart runtime.

**3. Runtime/UI alignment**
- Per-framework tab navigation uses browser history; code-block
hydration warnings are suppressed where the dedent branch injects raw
text; the TS agent server consistently emits SSE.

## Post-review hardening (review-fix loop + cross-framework audit)

A review-and-fix pass plus a cross-framework regression audit landed the
following on top:

- **Correctness fixes** — Python quickstart now uses the valid
`claude-sonnet-4-6` model id (the dotted sentinel was fed raw to the
adapter and would 404); feature-card documentation links retargeted to
pages that exist (were 404ing); `ms-agent-harness-dotnet` docs now
resolve to the shared `microsoft-agent-framework` folder and default to
the .NET tab (were 404ing in the live app).
- **Example quality** — the quickstart `main.py` now guards request
parsing + logs server-side, and builds the adapter once at module scope
(it was constructing a per-request adapter, leaking a Claude CLI
subprocess each request); the TS streaming snippet drops the undeclared
`partial-json` dependency (hand-rolled, matching the dependency-free
Python sibling).
- **Script robustness** — drained spawned-server pipes, temp-dir cleanup
on failure, SIGTERM→SIGKILL escalation, a stack-trace-leak guard that
also matches SSE-escaped newlines, and a fixed false-negative in the
missing-import check.
- **Cross-framework audit** — confirmed the shared-doc generalizations
don't regress other integrations (all `<Snippet region>` ids,
`<WhenFrameworkHas>` gating flags, and `<FrameworkSetup>` concepts are
unchanged; the modified gate scripts flag zero new pages vs `main`).
Restored concrete state-streaming API names as neutral examples so
`langgraph-typescript` / `langgraph-fastapi` (whose streaming snippet
region is empty) keep an actionable reference.

## Reviewer notes

- `router.replace` → `router.push` on per-framework tabs is intentional
(browser back/forward moves between tab options). The only trade-off is
that arrow-key tab navigation adds history entries.
- **Follow-up, not in this PR:** a "Claude SDK demo runtime hardening"
pass for ~17 pre-existing bugs in the demo backends (unguarded tool
execution, raw error leaks to the client, no turn limit, etc.).
Deliberately out of scope — this PR only moved `@region` markers near
that code; the runtime logic is a separate subject.
- **Verify before merge:** whether `npx copilotkit@latest init
--framework claude-sdk-{python,typescript}` is a supported CLI value —
an in-repo comment (`framework-overview.tsx`) suggests `claude-sdk-*`
has no matching CLI template.

## Validation

- CI green on HEAD: Validate Showcase, all `build-check`s,
`check-types`, `format`, `oxlint`, `doc-tests`, Python unit tests.
- `showcase/scripts` `verify-shell-docs.test.ts` 21/21; `setup-content`
tests pass.
- Claude quickstart runtime harness + Playwright quickstart probes
exercised locally.
2026-07-08 13:51:40 -07:00
Tyler Slaton 2255dd8cc3 test: add Claude SDK quickstart verification tooling
Add verify-shell-docs (+ unit tests), probe-shell-docs, probe-claude-quickstarts
(Playwright), and check-claude-quickstarts-runtime (extracts and runs the
documented commands/snippets) to gate the shell-docs build and quickstart
runtime. Includes CR hardening: drained server pipes, temp-dir cleanup,
SIGKILL escalation, a stack-trace-leak guard that matches SSE-escaped newlines,
and a fixed false-negative in the missing-import check.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 13:35:00 -07:00
Tyler Slaton 77a1df2593 fix: wire Claude SDK demo integrations for the quickstart docs
Align the claude-sdk-python and claude-sdk-typescript demo integrations behind
the published quickstarts: move the @region markers used for doc snippet
extraction, add the state-streaming and weather-tool snippet files, and add
setup-doc content. Runtime alignment: the TS agent handlers consistently emit
text/event-stream; the streaming snippets emit a fresh STATE_SNAPSHOT per delta
and drop the undeclared partial-json dependency.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 13:35:00 -07:00
Tyler Slaton 5ae2f82bb2 docs: publish Claude SDK quickstart docs
Turn on generated Shell docs for claude-sdk-python and claude-sdk-typescript:
quickstarts, framework registry data, docs links, and setup snippets.
Generalize the shared feature docs (state-streaming, HITL/interrupt,
tool-rendering, subagents, programmatic-control) to framework-neutral wording
so they read correctly across integrations. Includes review/audit fixes: the
valid claude-sonnet-4-6 model id, feature-card links pointing at pages that
exist, ms-agent-harness-dotnet docs-folder + tab-default routing, and concrete
state-streaming API names kept as neutral examples.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-08 13:35:00 -07:00
Tyler Slaton 4e2e980576 test(bot-telegram): de-flake chunked-edit-stream intermediate-flush tests (#5880)
## Problem

`packages/bot-telegram` unit tests failed on the **Node 20 shard only**
in [`test /
unit`](https://github.com/CopilotKit/CopilotKit/actions/runs/28968206256/job/85958192658)
(surfaced on PR #5849):

```
FAIL src/__tests__/chunked-edit-stream.test.ts > ChunkedEditStream >
  does not advance posted on failed intermediate flush, so next flush retries
AssertionError: expected +0 to be 1 // Object.is equality
```

## Root cause

Two `ChunkedEditStream` tests waited an arbitrary `await new Promise(r
=> setTimeout(r, 10))` for the stream's internal `setTimeout(0)` +
microtask flush chain to run, then asserted `callCount === 1` (the
intermediate flush fired and failed).

The `editAt` side-effect is delivered via `scheduleFlush`'s
`setTimeout(0)` → `enqueueFlush` microtask → `flushNow`. On a loaded CI
runner, scheduling jitter/GC can push that delivery past the fixed 10ms
window, so `callCount` is still `0` when the assertion runs → `expected
+0 to be 1`.

This is a **flaky timing race, not a Node 20 semantic difference**:
- The near-identical sister test (`retries text delivery…`) *passed* in
the same failing run.
- The test passes locally on Node 20 every time in isolation; the 20.x
shard simply lost the dice roll (`fail-fast: false` runs all three of
20/22/24).

Reproduced locally by shrinking the wait — this yields the identical
`expected +0 to be 1` error.

## Fix

Replace the arbitrary sleep with **event-based waiting** (per
systematic-debugging's condition-based-waiting guidance): resolve a
promise the instant the first `editAt` runs, and `await` it. Each test
now waits for the actual condition it cares about instead of guessing a
duration, so it is deterministic regardless of scheduling.

## Verification

- Fixed test: 10/10 green on Node 20, green on Node 22.
- Full package via `nx run-many -t build,test @copilotkit/bot-telegram`:
**149 tests pass (18 files)**, build + `check-types` clean.
- **Mutation probe:** injecting a slow (50ms) `postPlaceholder` makes
the *old* 10ms-wait pattern fail (the CI scenario) while the *new*
condition-based wait passes — confirming the timing coupling is removed.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-07-08 12:58:02 -07:00
github-actions[bot] 27aa946fff style: auto-fix formatting 2026-07-08 19:43:25 +00:00
Benjamin Taylor 900af9751f Merge origin/main into codex/managed-slack-hitl (Bots→Channels rename)
Reconcile #5814 with the @copilotkit/bot*→@copilotkit/channels* rename (#5849):
- relocate the new files (content-parts, intelligence-state-store[.test]) into
  packages/channels-intelligence and rename their imports to @copilotkit/channels*
- resolve the intelligence-adapter.ts / http-transports.test.ts import-header
  conflicts to the new package names (keeping #5814's symbol set — no unused
  AgentContentPart; vi/afterEach retained for the fetch-stubbing tests)
- rename lingering imports in transports.ts / http-transports.ts
2026-07-08 14:42:43 -05:00