Commit Graph

2865 Commits

Author SHA1 Message Date
Alem Tuzlak b44a043304 Merge origin/main into feat/typegen
Keep the AgentId and Register typegen types. Drop the @copilotkit/cli package. That CLI belongs in the Intelligence repo.
2026-08-31 12:42:57 +02:00
Ben Taylor 69441ccfa4 fix(react-core): render every v1 tool call (#6682)
## What does this PR do?

Fixes the v1 compatibility render path so an assistant message can
render every tool call instead of only `toolCalls[0]`.

The returned lazy renderer now:

- matches each tool call with its corresponding tool result;
- renders all registered tool-call renderers in one fragment;
- removes `null` render results and returns `null` when no tool has a
renderer.

Keeping the fragment behind the existing lazy-renderer callback
preserves the exported `useLazyToolRenderer` return signature. Filtering
before returning also avoids attaching empty generative UI, so
caller-provided subcomponents are not suppressed when no renderer is
registered.

Regression tests cover multiple tool calls, per-call result matching,
the all-unhandled case, and a mixed handled/unhandled message.

## Related PRs and Issues

- Fixes #2946

## Verification

- `pnpm nx run @copilotkit/react-core:test` (133 files, 1,530 Vitest
tests plus 47 script tests)
- `pnpm nx run @copilotkit/react-core:check-types`
- `pnpm exec oxfmt --check
packages/react-core/src/v1-deprecated/hooks/use-lazy-tool-renderer.tsx
packages/react-core/src/v1-deprecated/hooks/__tests__/use-lazy-tool-renderer.test.tsx`

## Checklist

- [x] I have read the [Contribution
Guide](https://github.com/copilotkit/copilotkit/blob/master/CONTRIBUTING.md)
- [x] Documentation is unchanged because this restores existing v1
behavior without changing the public API
- [x] "Allow edits by maintainers" is checked
2026-08-30 16:59:05 -05:00
Ben Taylor 9909ca44d6 fix(react-core): make agentMetadata.nodeName match the node where the interrupt originates (#6488)
## Problem

`useAgentNodeName` must update React consumers when AG-UI node events
arrive, and `useLangGraphInterrupt.enabled()` must receive the node
where an interrupt actually occurred.

Current `main` includes the basic ref-to-state reactivity fix from
[ffd1580](https://github.com/CopilotKit/CopilotKit/commit/ffd15801d6d),
but that commit explicitly leaves #1426 open: a later `RUN_FINISHED` can
still replace the interrupting node with `"end"`, and v1 consumers can
still be hidden behind `useCoAgent`'s memoized return value.

## What remains in this PR

Rebased onto current `main` (`e9387e0`) after the v1 source migration,
this PR contains only the remaining behavior:

- Preserve the last active node when `RUN_FINISHED` reports `outcome:
"interrupt"`.
- Preserve it for the legacy `on_interrupt` custom-event flow as well.
- Continue transitioning successful and failed runs to `"end"`; reset
new runs and agent switches to `"start"`.
- Add `nodeName` to the `useCoAgent` return-value memo dependencies so
v1 consumers receive the reactive update.
- Share `INTERRUPT_EVENT_NAME` between the hook and interrupt
implementation.

The public hook signatures and AG-UI protocol remain unchanged.

## Preview workflow

- Disabled pkg-pr-new's generated all-package StackBlitz template;
package preview install URLs remain available.

## Changes

- `packages/react-core/src/v1-deprecated/hooks/use-agent-nodename.ts`
- `packages/react-core/src/v1-deprecated/hooks/use-coagent.ts`
-
`packages/react-core/src/v1-deprecated/hooks/__tests__/use-agent-nodename.test.tsx`
- `packages/react-core/src/v2/types/interrupt.ts`
- `packages/react-core/src/v2/hooks/use-interrupt.tsx`
- `.github/workflows/publish-commit.yml`

## Verification

- Full React Core Vitest suite: **131 files, 1520 tests passed**.
- Preview workflow: Nx formatting and YAML parsing passed.
- `nx run @copilotkit/react-core:check-types --skipNxCache`: passed,
including all 33 dependency tasks.
- `git diff --check origin/main...HEAD`: passed.
- The composite React Core test target then reaches the existing
Windows-only script baseline: 8 path-normalization failures plus 2
symlink-permission failures. These are outside this PR's files; the
complete Vitest suite passes before that script stage.

## Scope

This intentionally does not change the AG-UI event protocol, runtime
event ordering, HITL workflow, v1/v2 compatibility layer, or the
separate node tracking in `use-coagent-state-render-bridge.tsx`.

Fixes #1426
2026-08-30 10:05:27 -05:00
Ben Taylor 40beb51634 fix(channels): echo the originating interrupt back as command.interruptEvent (#6353)
## Problem

`thread.resume()` sends `forwardedProps.command = { resume }` and
nothing else.

LangGraph needs nothing more — it correlates a resume by `thread_id` —
so this gap has been invisible. Bridges that correlate by ids carried
**inside the interrupt** cannot work at all. `@ag-ui/mastra` gates its
resume on `command.interruptEvent` and needs `{toolCallId, runId}` out
of it:

```js
let o = input.forwardedProps?.command;
if (!o?.interruptEvent && Array.isArray(input.resume)) { /* standard path */ }
if (o?.resume === false && o?.interruptEvent) { /* cancelled */ }
if (o?.resume != null && o?.interruptEvent) { /* resume */ }
```

With no `interruptEvent`, **every branch is skipped**. The resume is not
an error — it is a silent no-op: the run proceeds as an ordinary new
turn, no `RUN_ERROR`, no tool result, and the suspended tool never
completes. That is the entire Mastra HITL story through Channels today:
the picker appears, the click lands, nothing resumes.

## Change

`Thread` retains the captured interrupt value when the interrupt fires,
and `resume` echoes it back.

- **Opaque.** Channels neither parses nor reshapes the value — it hands
back exactly what the agent sent, so no framework specifics leak in.
This mirrors what `useInterrupt` in `react-core` already does for legacy
interrupts.
- **Durable, not in-memory.** The click that resumes can arrive in a
later run, possibly after a restart — which is the entire reason this
path exists rather than `awaitChoice`.
- **One-use.** Read back with `kv.consume` (atomic take-and-delete),
matching the one-use continuation the resume already claims, so a
replayed click cannot resurrect a spent resume.
- **Retention tied to `actionRetentionMs`**, so the value and the button
that consumes it expire together. Configure a longer action retention
and the value can lapse first; the resume then omits the field, which is
the pre-existing behaviour rather than a new failure.
- **Best-effort store touches.** An agent that needs no correlation data
must not lose its interrupt to a store hiccup.

## Why this is additive

The field is **omitted** when no interrupt value was captured, so the
serialized command is byte-identical for agents that carry no
correlation data, and a normal (non-interrupt) run still writes nothing
at all. LangGraph ignores the extra field.

`interrupt-event-echo.test.tsx` covers all four properties — echo,
durability, one-use, omission. The echo test fails without the fix
(verified by reverting it).

## Verification

- `channels-core`: 268/268 tests, typecheck clean, `oxfmt` clean,
`oxlint` 0 errors.
- Verified end to end against a real Mastra agent driven through Slack:
two full cycles (approve, then decline), each producing a matching
`SUSPENDING` → `RESUMED` pair agent-side. Before the fix the same rig
produced suspends with zero resumes.
- The LangGraph interrupt path was re-checked against the same harness
and is unchanged.

## Reviewer note

One existing assertion in `hitl-continuation.test.tsx` moves from
counting `kv.set` / `kv.consume` calls to matching them by **key**. The
invariant it protects is "an interrupt turn writes exactly one action
snapshot, and one resume consumes it once" — both still asserted, now
stated directly rather than via a total that conflates action state with
unrelated bookkeeping. Flagging it because it is the one pre-existing
test this change edits.

## Out of scope

Channels ignores the standard AG-UI interrupt path (`RUN_FINISHED` with
`outcome.type === "interrupt"`, resumed via a top-level `resume` array).
Mastra emits it today alongside the legacy event, and it carries
interrupt ids and expiry the legacy event lacks. Worth a follow-up.
2026-08-29 14:29:29 -05:00
Ben Taylor 10587a8cd6 chore(channels): drop references to canceled OSS tickets in code comments (#6297)
## What

Three comments in the channels packages point at Linear tickets that
were **canceled on 2026-07-01** (superseded by OSS-401's consolidated
realtime foundation). They read as live tracking pointers when nothing
actually tracks the work.

| File | Reference | Ticket status |
|---|---|---|
| `packages/channels-core/src/codec.ts` | `TODO(OSS-363)` | Canceled |
| `packages/channels-intelligence/src/runtime.ts` | `TODO(OSS-377)` |
Canceled |
| `packages/channels-slack/src/ingress-normalize.ts` | `(OSS-362)`
provenance note | Canceled |

## How

The engineering intent is **preserved** in each case and re-marked as
`TODO (untracked)` with a note that the original ticket was canceled —
so a reader re-files rather than chasing a dead ticket. The `OSS-362`
mention was a bare provenance parenthetical carrying no intent, so it is
simply dropped.

`codec.ts` additionally records that the TODO's stated precondition is
**already met**: the pure Slack ingress mapping (mention stripping,
stable event-id derivation, real-user filtering, field extraction) now
lives in `channels-slack/src/ingress-normalize.ts`. The remaining work
is hoisting `normalizeIngress` onto the generic `PlatformCodec` — worth
knowing for whoever picks it up.

## Why this came up

While grounding a Channels x AG-UI gap analysis, several `OSS-###`
comments turned out to reference closed tickets — three canceled, two
already done. The done ones (OSS-450 Teams managed, OSS-491 delivery
terminal-signal) had already been cleaned out of `main` by subsequent
work; these three canceled ones are what remain.

## Scope / notes

- **Comment-only. No behavior change.** `oxfmt --check` and `oxlint`
both clean.
- Out of scope, flagged for a possible follow-up: `OSS-406`
(`channels-intelligence/src/index.ts`) and `OSS-599`
(`runtime/.../channel-manager.ts`, 5 references) are also canceled, but
sit outside the channels packages surveyed here. Other referenced
tickets (`OSS-96/162/476/566/568/621/623/641/670/691`) were not
status-checked.
2026-08-29 14:28:25 -05:00
Mark 17ed176736 chore: RuleTester coverage for no-single-arg-zod-record lint rule (#5032)
## Summary

Follow-up to #5025 addressing @marthakelly's [review
suggestion](https://github.com/CopilotKit/CopilotKit/pull/5025#discussion_r3307285578):
add oxlint `RuleTester` coverage for the
`copilotkit/no-single-arg-zod-record` rule. Since the underlying Zod 4
incompatibility is **type-level** (no runtime test can catch a
regression), this lint rule is the real safety net, so it's worth
testing directly.

## Cases (via `oxlint/plugins-dev` `RuleTester`)
**valid**
- two-arg `z.record(z.string(), z.unknown())` — no false positive
- two-arg with `.optional()` chain
- single-arg `.record()` on a non-`z` object (`cache.record(entry)`) —
confirms the rule is scoped to the `z` alias and won't over-fire

**invalid**
- single-arg `z.record(z.unknown())` → fires + autofix output
`z.record(z.string(), z.unknown())`
- chained `z.record(z.unknown()).optional()` → fixes the inner call
- `z.record(...spread)` → reports **without** a fix (`output: null`)

## Node version gate (important)
oxlint's `RuleTester` requires **Node ≥ 22** (it throws at parse time on
older runtimes). The CI unit matrix includes a **Node 20** job, so the
cases are gated to skip below Node 22 — verified locally: **6/6 pass on
Node 22**, **skips cleanly on Node 20**. The lint rule itself is still
exercised on every Node version through the `oxlint` job; only these
RuleTester unit tests are gated.

Also extends the `react-ui` vitest `include` to pick up co-located
`oxlint-rules/**/*.test.mjs`.

## Test plan
- [x] `vitest run` on the new file under Node 22 → 6 passed
- [x] same under Node 20 → 1 skipped (gate works; Node 20 CI job stays
green)
- [x] `oxlint` clean on the new test + config

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-08-29 07:54:58 -07:00
Ben Taylor b995233a4d test(shared): add unit tests for conditions engine (#6735)
Add comprehensive unit tests for the xecuteConditions rule engine in
packages/shared/src/utils/conditions.ts, which previously had no test
coverage. Covers all 12 comparison/existence rules, logical AND/OR/NOT
nesting, implicit AND across multiple conditions, and dot-path
resolution (including missing-path and whole-value fallback). Uses the
existing vitest setup in the shared package.
2026-08-28 23:23:22 -05:00
Aswin Kumar 99d349c6cd Merge branch 'main' into fix/angular-hitl-result-envelope 2026-08-29 00:41:13 +05:30
Benjamin Taylor 4edbbc3f28 fix(react-core): honor chatInputToolbarAddButtonLabel in the add menu tooltip
The add-menu ("+") button's tooltip hardcoded the string "Add attachments",
so `labels.chatInputToolbarAddButtonLabel` only retitled the menu item and
the tooltip stayed English. That blocked full localization of CopilotChat
without replacing the whole add-button slot.

Every other tooltip in the v2 chat surface is already label-driven, and the
Angular implementation already derives this tooltip from the same label, so
this was an oversight rather than a deliberate split.

The "/" shortcut glyph stays hardcoded — it is a key name, not prose.

Fixes #6750

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-28 13:45:54 -05:00
Ben Taylor d692f25852 feat(runtime): expose Learning Container selector (#6767)
## What does this PR do?

Adds `getLearningContainerId` to `CopilotKitIntelligence` so developers
can assign Intelligence Threads to Learning Containers without a
theta-prefixed Runtime option.

The selector receives:

- The resolved application `user`.
- The parsed AG-UI `input` for the run.
- The `agentId` and the `web` or `channel` surface.

Web runs pass the exact parsed `RunAgentInput`. Channel runs build the
same canonical input that the AgentRunner receives. Channels also carry
the resolved application user through the core and Intelligence adapter
boundaries.

The Runtime validates the selected stable ID and sends only that ID with
the existing Thread create or lock call. Intelligence stays responsible
for project scope, entitlements, Container lookup, and the one-time
Thread binding. Persisted AG-UI events remain the source for Learning
snapshots.

The old `ɵlearning` Runtime option remains as a deprecated fallback. The
Runtime rejects configurations that set both APIs.

## Why?

Learning Container assignment is an Intelligence SDK concern. Developers
also need the resolved user and complete run input to select a Container
from application data without reading raw transport details.

## Related PRs and Issues

- Refs
[ENT-1149](https://linear.app/copilotkit/issue/ENT-1149/enable-projects-to-learn-from-agent-runs-and-publish-reusable-skills)
- Related design: #6746

## Validation

- GitHub CI: 53 passed, 3 skipped
- `pnpm nx run-many -t test,check-types,build -p @copilotkit/runtime
@copilotkit/channels-core @copilotkit/channels-intelligence`
- `pnpm nx run-many -t publint,attw,check-dts -p @copilotkit/runtime
@copilotkit/channels-core @copilotkit/channels-intelligence`
- `pnpm lint` (0 errors; existing warnings remain)
- `npm run typecheck` and `npm run build` in `showcase/shell-docs`
- `npm test` in `showcase/shell-docs` has one pre-existing failure at
`inspector-docs.test.ts:141`: the tracked Threads callout contains
`Playground`.

## Checklist

- [x] I have read the [Contribution
Guide](https://github.com/copilotkit/copilotkit/blob/master/CONTRIBUTING.md)
- [x] If the PR changes or adds functionality, I have updated the
relevant documentation
- [x] Maintainer edits are available because this PR uses a branch in
the main repository
2026-08-28 13:02:10 -05:00
Ben Taylor 4ccae6fe20 fix(react-native): keep the polyfill imports in the built barrel (closes OSS-1002) (#6744)
## The bug

`src/polyfills.ts` is five side-effect-only imports plus
`installStreamingFetch()`. The published barrel was 195 bytes:

```js
// node_modules/@copilotkit/react-native/dist/polyfills.mjs — 1.69.2
import { t as installStreamingFetch } from "./streaming-fetch-BnQh3vBz.mjs";
installStreamingFetch();
export {  };
```

Every React Native app following the documented setup died on its first
runtime call:

```
E ReactNativeJS: '[CopilotKit] Error (runtime_info_fetch_failed):',
  [ReferenceError: Property 'ReadableStream' doesn't exist]
```

## Root cause

The `sideEffects` field, but not in the way it first looks. It is
correct for *consumers* and wrong for *this package's own build*:

```json
"sideEffects": ["./dist/index.*", "./dist/headless.*", "./dist/polyfills.*", "./dist/polyfills/**/*"]
```

tsdown/rolldown reads the package's own `sideEffects` while bundling and
matches it against **source** paths. `src/polyfills/streams.ts` matches
none of those `dist` globs, so it is declared side-effect-free — a hard
assertion that lets rolldown drop the import without analysing the
`globalThis` assignments inside.

Reproduced in isolation at the pinned tsdown (0.20.3):

| `sideEffects` | built barrel |
|---|---|
| `["./dist/polyfills.*", "./dist/polyfills/**/*"]` | `export { };` —
empty |
| same + `["./src/polyfills.*", "./src/polyfills/**/*"]` | `import
"./polyfills/streams.mjs";` |
| field absent | `import "./polyfills/streams.mjs";` |

**Wider than the ticket recorded:** `dist/index.mjs` and
`dist/headless.mjs` also had zero polyfill code, so the package's
advertised auto-install on first import did not happen either. Not
RN-specific in principle — but I surveyed every package at `origin/main`
and this is the only one exposed. The other `sideEffects` arrays
(`react-core`, `react-ui`, `react-textarea`) are `["**/*.css"]`, which
matches source and works.

## The fix

Add matching `./src/**` globs. Barrel goes 195B → 362B with all five
imports; `headless.mjs` now leads with `import "./polyfills.mjs"`.

## The test, and why the existing one didn't catch this

`src/__tests__/polyfills.test.ts` has ~20 assertions covering all five
groups and was green the whole time — it imports `"../polyfills"`, the
TypeScript **source**, which vitest transpiles without bundling and
therefore without tree-shaking. It exercises a graph the published
package does not contain.

So the new check runs against `dist/`. Two things it has to get right to
be honest:

- **Node ships these globals natively.** Asserting `ReadableStream` is
"defined" after import passes on an empty barrel. The probe clears all
nine first, emulating Hermes.
- **The two formats need different treatment.** CJS is executed for real
in a child realm. ESM is checked structurally — it cannot be executed
here because `encoding.mjs` takes a named import from CommonJS
`text-encoding`, which Metro rewrites to a `require()` but bare Node ESM
rejects.

It is wired into `build`, so a dead barrel fails the build rather than
reaching npm — which matters, because this shipped through a fully green
suite.

## Docs

Added the `Property 'ReadableStream' doesn't exist` symptom to
troubleshooting, which previously covered only the inverse case (a
polyfill *conflict*).

I deliberately left the reference docs' "auto-installs on first import"
claims and the crypto import-order callout alone: both become **true**
once the build is fixed, and I verified the auto-install behaviourally.

## Verification

- **Red/green proven, not assumed:** reverted the `sideEffects` change,
rebuilt → 5/5 groups FAIL in both formats. Restored → 5/5 PASS. There is
also a test for a *single* group regressing, which a whole-barrel
assertion would wave through.
- **Packed tarball** (`pnpm pack`) verified behaviourally: all nine
globals install.
- 289 vitest + 26 script tests pass; `check-types` clean; `attw` green;
`publint` clean apart from a pre-existing `repository.url` suggestion;
oxfmt/oxlint clean.
- Added `{projectRoot}/scripts/**` to the package's `test` inputs and
confirmed cache invalidation (19/19 cached → 18/19 after touching the
verifier); without it, editing the verifier alone would restore a cached
pass.

**Not verified:** the on-device round trip — no emulator in this
environment. The bare-realm equivalent passes on the packed tarball.

## Follow-up worth its own ticket

`dist/polyfills/encoding.mjs` uses a named import from CommonJS
`text-encoding`. Metro handles it; a true-ESM consumer would not.
Pre-existing and not RN-facing, so left out of this change.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-08-28 12:24:39 -05:00
Mike Ryan 9e27e10830 feat(runtime): expose Learning Container selector 2026-08-28 10:17:27 -07:00
Ben Taylor 15d4fa0e04 fix(telemetry): stamp sampling metadata and emitter markers on every event (#6749)
Closes OSS-1017, OSS-1018, OSS-1019.

Runtime volume can't currently be counted from the telemetry data alone.
Three separate reasons, all in the two `TelemetryClient`s, all fixed
here.

## What was wrong

**OSS-1017 — v2 samples but doesn't say so.**
`packages/runtime/src/v2/runtime/telemetry/telemetry-client.ts` gates
anonymous events at 5% and lets identified callers through at 100%, then
sends without stamping `sampleRate` / `sampleWeight`. The v1 client has
computed that block for a while; v2 never did. Measured in PostHog
(project 26816, `oss.runtime.copilot_request_created`), the share of
runtime volume arriving unweightable was 1.1% in May, 12.2% in June,
20.2% in July, 24.4% in August — roughly doubling each month. It has
already produced wrong GTM numbers: raw counts understate real volume
10–18× *and* overstate growth (+211% vs a true +128% Jan→Jul), because
the sampled/unsampled mix drifts as v2 adoption rises.

**OSS-1018 — identified v2 events aren't detectable downstream.** The
telemetry id travels as the `X-CopilotKit-Telemetry-Id` header, not an
event property, so whether an event was sampled at 5% or captured at
100% wasn't recoverable from the event. Consumers had to probe for
whatever identity fields happened to be present, which only ~2.7% of v2
events carry.

**OSS-1019 — v1 writes every event twice.** `capture()` sends to both
`lambdaClient.send()` and `segment.track()`, so one request becomes two
PostHog rows with nothing marking them as copies. Both prior dedupe
attempts produced wrong numbers: a "≥ 1.60 dual-emits" version cutoff is
wrong (1.59.5 dual-emits too — it added ~13M phantom requests to June),
and Segment-only drops the entire v2 runtime (19.5M in August).

## What changed

A new `packages/shared/src/telemetry/sampling.ts` holds
`computeSamplingMeta`, and both clients call it. The v1 client's inline
block is replaced by that call with identical output, so the
pre-existing v1 tests act as the regression check that the extraction is
behavior-preserving. v2 passes the result in `globalProperties`, the
slot the sink already spreads into the event and the same one v1 uses —
so both emitters are now identical on the wire.

`telemetry_identified` is carried explicitly rather than inferred from
`sampleWeight === 1`. Under `COPILOTKIT_TELEMETRY_SAMPLE_RATE=1`
anonymous events also weigh 1, and weight alone stops separating the
populations.

For the dual-write, per the issue's option 3 plus a dedupe key —
additive only, both transports keep flowing:

| property | value |
| -- | -- |
| `telemetry_emitter` | `v1-shared` \| `v2-runtime` |
| `telemetry_transport` | `segment` \| `lambda` |
| `telemetry_event_id` | one uuid per v1 `capture()`, identical on both
copies |

v2 doesn't carry an event id — it has a single transport, so there's
nothing to dedupe.

## The counting rule this enables

Drop v1's lambda copy, keep everything else, and sum the stamped weight:

```sql
sumIf(sampleWeight, NOT (telemetry_transport = 'lambda' AND telemetry_emitter = 'v1-shared'))
```

Note for whoever runs the next month-close: **this release breaks the
existing runbook query.** It isolates v2 with `NOT
JSONHas(properties,'sampleRate')`, which matches nothing once v2 starts
stamping `sampleRate` — `v2_runtime` would silently read zero and total
volume would lose ~24%. The runbook has been updated with an
era-agnostic query that gives the same answer either side of the
release; historical events don't get backfilled, so the old inference
path stays as the fallback for pre-fix data.

## Testing

Run from the worktree with `@copilotkit/shared` built so the runtime
resolves the real helper rather than a mock.

**`packages/shared` — 35 passed (35)**, including all 20 pre-existing v1
tests, unchanged. New `sampling.test.ts` (5) covers both branches, the
`sampleRate=1` collision that motivates the explicit flag, and an
empty-string telemetry id. New cases in `telemetry-client.test.ts` (4)
cover both copies sharing one `telemetry_event_id`, a fresh id per
capture, the emitter marker on both copies, and `telemetry_identified`
tracking the gate branch.

```
 ✓ src/telemetry/sampling.test.ts (5 tests)
 ✓ src/telemetry/lambda-client.test.ts (6 tests)
 ✓ src/telemetry/telemetry-client.test.ts (24 tests)
 Test Files  3 passed (3)
      Tests  35 passed (35)
```

**`packages/runtime` v2 telemetry — 37 passed (37)**, covering the new
client suite plus the four pre-existing license/telemetry integration
files:

```
 ✓ src/v2/runtime/telemetry/__tests__/telemetry-client.test.ts (6 tests)
 ✓ src/v2/runtime/telemetry/__tests__/global-properties.test.ts (6 tests)
 ✓ src/v2/runtime/telemetry/__tests__/instance-created.test.ts (5 tests)
 ✓ src/v2/runtime/__tests__/telemetry.test.ts (15 tests)
 ✓ src/v2/runtime/__tests__/sse-license-telemetry.integration.test.ts (1 test)
 ✓ src/v2/runtime/__tests__/sse-license-env-fallback.integration.test.ts (1 test)
 ✓ src/v2/runtime/__tests__/license-telemetry-endpoints.integration.test.ts (3 tests)
      Tests  37 passed (37)
```

One pre-existing assertion changed: `global-properties.test.ts` asserted
`globalProperties` equals `{}` when the caller sets none, which is wrong
by design now that sampling metadata always rides there. Rewritten to
pin the exact key set, which preserves the original intent (nothing of
the caller's is added) and is strictly stronger.

**Mutation-checked** — each new test was verified to fail when its
mechanism is broken, then the mutation reverted and green confirmed:

| mutation | result |
| -- | -- |
| `effectiveSampleRate` always `= sampleRate` (drop identified branch) |
shared 2 failed, runtime 1 failed |
| `telemetry_identified` hardcoded `false` | shared 3 failed, runtime 1
failed |
| v1 event id hoisted per-client instead of per-capture | shared 1
failed |
| both v1 copies stamped `transport: "lambda"` | shared 1 failed |
| v2 drops the sampling block from `globalProperties` | runtime 4 failed
|
| *(reverted)* | shared 35/35, runtime 17/17 |

**Build/lint** — `tsdown` clean for both `@copilotkit/shared` (115
files) and `@copilotkit/runtime` (414 files); `oxfmt` applied; `oxlint`
reports 0 errors. The one warning on new code
(`consistent-function-scoping` on a test-local `jwtWith`) is the same
pattern the existing v1 test file already uses at
`telemetry-client.test.ts:42`.

Not run locally: the pre-commit `test-and-check-packages` hook, which
builds the whole monorepo and fails in this worktree on packages that
were never installed there (`core`, `sdk-js`, `channels-ui`) —
environmental, unrelated to these files. Leaving that to CI.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-08-28 12:13:50 -05:00
BenTaylorDev bf1bb98765 chore: release angular v0.4.0 2026-08-28 15:03:28 +00:00
Alem Tuzlak 8469e72b30 feat(web-inspector): copy stored threads into Playground from Threads (#6642)
Inspector Threads now has **Try from here**. One click copies a stored
thread into a Playground scratch session. The stored thread does not
change.

If the copy fails, Inspector stays on Threads and keeps the current
Playground scratch. Example tour threads and locked Threads do not show
the button.

## What does this PR do?

Adds **Try from here** on a real stored thread in Inspector Threads. One
click copies messages and thread state into a Playground scratch
session. The stored thread does not change.

If the copy fails, Inspector stays on Threads and keeps the current
Playground scratch. Example tour threads and locked Threads do not show
the button.

## Related PRs and Issues

- Linear: OSS-873
- Playground base: https://github.com/CopilotKit/CopilotKit/pull/6580
(merged)

## Checklist

- [x] I have read the [Contribution
Guide](https://github.com/copilotkit/copilotkit/blob/master/CONTRIBUTING.md)
- [x] If the PR changes or adds functionality, I have updated the
relevant documentation
- [x] "Allow edits by maintainers" is checked

## Testing

**Commands run**

1. Rebased `feat/oss-873-try-from-here` onto `origin/main` and resolved
6 conflict files.
2. `npx nx run @copilotkit/web-inspector:test` — 626 tests passed (after
the stale-result guard).
3. `npx nx run @copilotkit/web-inspector:check-types` — passed.

**Manual test**

1. Open Inspector on localhost with Intelligence on, so a real stored
thread exists.
2. Open that thread. Confirm **Try from here** is in the thread header.
3. Click **Try from here**. Confirm Inspector opens Playground with the
copied messages and the stored thread is unchanged.
4. Open an example tour thread. Confirm **Try from here** is not shown.
5. Force a copy failure (disconnect runtime). Confirm Inspector stays on
Threads and the prior Playground scratch is unchanged.

**How this PR makes testing easy**

- `packages/web-inspector/src/__tests__/inspector-navigation.spec.ts`
covers the button, copy path, failure path, and a stale click that must
not overwrite Playground.
- `packages/web-inspector/src/lib/__tests__/telemetry.test.ts` covers
`oss.inspector.threads_try_from_here_clicked`.

## Risk / rollback

Risk is limited to Inspector Threads and Playground. A revert of this PR
removes the button and the new telemetry event. No runtime protocol
change.

## Public API change

New Inspector telemetry export and event name:

**Before**

```ts
trackThreadsTabClicked(props);
```

**After**

```ts
trackThreadsTabClicked(props);
trackThreadsTryFromHereClicked({ ...props, outcome: "success" });
```

`CpkThreadInspector` also emits a `tryFromHere` custom event when the
user clicks the button.
2026-08-28 16:10:06 +02:00
Alem Tuzlak 5a191c8e36 fix(web-inspector): put Try from here next to Expand all
Move the button onto the messages toolbar, on the right of Expand all and Collapse all, with a top-right arrow.
2026-08-28 15:57:18 +02:00
Alem Tuzlak ec146f6721 fix(web-inspector): drop stale Try from here results 2026-08-28 15:01:06 +02:00
Alem Tuzlak d6812e41c8 fix(core): report the runtime connection status from the last actual contact (#6706)
Closes OSS-904.

## The problem

`CopilotKitCore.runtimeConnectionStatus` was set only by the `/info`
handshake, which runs **once on connect**. If the runtime became
unreachable after that, the status stayed `connected` indefinitely —
measured two independent ways against `examples/v2/react/demo` and
recorded in the ticket.

The failure was not lost, it was filed in the wrong drawer: it arrived
as `agent_run_failed`, indistinguishable from an agent bug. So
everything downstream inherited the wrong answer — System Health
reported healthy, the launcher error signal could not fire for the most
common real symptom ("it worked a minute ago"), and a customer `onError`
handler written to separate wiring problems from agent problems got the
wrong classification.

## What this changes

The status now reports **the outcome of the last actual contact with the
runtime**.

A failed runtime request — or silence past a per-request watchdog —
triggers **one** bounded confirmation request. If nothing answers, the
status moves to `error` and the failure is emitted through the existing
wiring error code, so customers already handling startup wiring failures
pick up the mid-session case without changing a line. A subsequent
successful request re-syncs and clears it.

Crucially, **the conversation survives**. The transition does not
discard runtime knowledge, so the agent backing an open chat is the same
instance, its messages stay on screen, and submitting stays possible —
which matters because submitting is what restores the status.

No polling, no heartbeat, no retry loop. Every timer is bound to one
request and dies with it.

## Decisions worth knowing when reviewing

- **Reactive in both directions.** A heartbeat would put permanent
background traffic into every embedding application; a retry loop mostly
races a user who is about to retry anyway. The cost is stated rather
than hidden: while nothing is happening, nothing is detected.
- **Status change is separated from discarding knowledge.** The only
pre-existing code that set `error` also cleared `remoteAgents`. That is
right at startup and destructive mid-session, because conversation state
lives on the agent instance. Four sites now hold this invariant up
together; each carries a comment saying so.
- **The trigger is deliberately permissive, and the check is the
arbiter.** A request that received a successful response never triggers
a check; user cancellation never does; everything else may. Defining the
trigger precisely would mean maintaining a status-code list that is
complete only for the deployment topologies someone thought of.
- **Silence counts.** A server can refuse (fails fast) or hang (accepts
and never answers). A stopped dev server refuses; a container
mid-rollout, a half-switched deploy and a dropped tunnel hang. Only
bounding the check does not help, because no check starts — hence the
per-request watchdog. It observes only and never cancels the request.
- **The rule is stated by destination, not by call site**, so a runtime
route added later inherits the behaviour. Excluded: the Intelligence
realtime endpoint (a different service — reporting its outage as
"runtime unreachable" would be a false diagnosis), endpoints belonging
to the customer, and the stop request.
- **Recovery may prune, under two conditions**: the runtime must have
reported at least one agent, and the agent must carry no conversation
state. An empty list is the signature of a runtime that has not finished
registering.
- **"Answered but refused" keeps the error status and gets a different
message.** An expired token means the app cannot work, so red is right;
telling the reader "unreachable" would send them to check ports and
containers.

## Deliberately not delivered

- Detecting an outage, or a recovery, while the application is idle.
- Recovery by opening the Threads view: every binding withholds thread
requests until the status is already connected, so nothing is sent while
it is red. The thread plumbing still earns its place for *detection*.
- A signal for the Intelligence realtime endpoint failing while the
runtime is healthy — a real gap, and its own ticket.
- Memory and suggestion routes adopting the instrumented fetch.
- A new status value or a new error code.

## Costs this introduces

`error` now means two things — "never connected, no agents" and "lost
mid-session, agents intact". Documented on the enum. And because the
status can now change mid-session at all, an outage costs some churn
that did not exist before: the memory list and the Inspector's thread
list are cleared and refetched, and where the chat owns its run-activity
store it is stopped and restarted. All of it is paid on a user-caused
transition, never while idle.

## Testing

Four independent reviewers audited an earlier revision of this branch;
the ten defects they reproduced are fixed and each is pinned by a test
that was red first. A mutation audit of 110 mutants killed 90; the
surviving holes were closed in the round after.

The connection-health suites carry 72 tests. Request counting is a
first-class assertion throughout, because several decisions are
*absences* — no polling, no retry loop, one check per burst, no traffic
while red — and an absence is only testable by counting. Those tests use
fake timers advancing ten minutes; that boundary is documented where it
lives, since anything slower is invisible to them.

Verified by hand in a browser with the runtime running as its own
process, so the page outlives it: a refusing runtime, a hanging runtime,
recovery, an agent added during an outage, an agent deleted during an
outage. `performance.timeOrigin` was checked throughout to prove the
page never reloaded and the result was not an artefact of a fresh
handshake.

## Follow-ups this leaves behind

Three of these deserve their own ticket. None blocks this PR; all three
are consequences of where its scope was drawn, and they are listed here
so the boundary is explicit rather than implied.

### 1. A signal for the Intelligence realtime endpoint

In Intelligence mode the browser gets its chat events from a **second
service** at its own address; the runtime is only asked for the
credentials. If that service fails while the runtime is healthy, this
change correctly reports the runtime as reachable — and the user
experiences exactly the silence this ticket exists to remove.

It is excluded here on purpose: folding it into the runtime status would
report "runtime unreachable" about a healthy runtime, and a false
diagnosis costs more debugging time than no signal. It needs its own
signal, which is a presentation decision as much as a detection one.

### 2. Memory routes onto the instrumented fetch

The memory store still builds with the global fetch, so its
runtime-bound requests are invisible to connection health. Two costs: a
genuine failure there is a signal we discard, and a success there cannot
restore the status.

The asymmetry is what makes this worth fixing rather than leaving:
memory is the surface most disrupted by a status transition (its list is
cleared and refetched) and currently the one least able to contribute.
The change itself is small — that module already takes its request
function as an injected dependency.

### 3. Consumers should key on what they need, not on the status value

Several consumers treat "status is not connected" as "discard
everything": the memory list, the Inspector's thread list, and the
chat's run-activity store. That was harmless while the status could not
change after page load. It can now, so every outage costs churn that did
not exist before.

This is the same mistake this PR fixes three times *inside* core — a
guard bound to a state instead of to the thing it protects. The
principle was applied internally and not to these consumers. That makes
the churn listed under "Costs" above **deferred rather than inherent**,
and it is the largest of the three follow-ups: three consumers in three
packages, each with its own risk, which is why it was kept out of this
PR.

### Two smaller items

- The launcher error signal on `main` carries a comment stating the
limitation this change removes ("a runtime that dies after the page
loaded … raises nothing … closing that gap means a re-probe in the
core"). It becomes false when this lands and should be corrected then.
- `packages/web-inspector/src/styles/generated.css` is build output
under version control and re-dirties the tree on every build. Unrelated
to this PR, but the Tailwind source glob scans test files, so any prose
comment containing a utility word (`fixed`, `hidden`, `visible`,
`block`) silently changes the committed CSS. Narrowing the glob would
remove the class of problem.

Full specification, including the interview decisions and every
revision: `OSS-904-PRD.md`.
2026-08-28 14:46:15 +02:00
github-actions[bot] c71dcfbb12 style: auto-fix formatting 2026-08-28 14:24:39 +02:00
Alem Tuzlak 95285be33b feat(web-inspector): copy stored threads into Playground from Threads 2026-08-28 14:24:05 +02:00
Alem Tuzlak 1cb76928ad Turn the Home Intelligence card into an install path (#6740)
## Why

The Inspector said what Intelligence *is* and linked out to a signup
page. Of ~1,655 Inspector opens in 90 days, **under 100 clicked any
CTA**. This replaces the feature list with an argument, and the outbound
link with an install that happens in the editor the developer is already
in.

This is the unfinished half of OSS-867, whose body asks for exactly
this: *"If a capability requires Intelligence, detail why and include a
video demonstrating that feature working end-to-end."*

## What changed

**A four-slide argument, paired to the picture beside it.** Each slide
carries two sentences and the visual they describe: your users' threads
→ the pattern inside them → the skill file → what it compounds into. An
earlier draft sold Threads in prose while animating Learning; bound
together they read as one chain, and `meeting-scheduling/SKILL.md`
recurs through all four so the closing diagram is checkable rather than
decorative.

Condensed from the six-phase animation on the Intelligence home page —
not screen-recorded. A ported version is themeable, stays sharp, and
costs no asset weight; the original also runs 21.4s and opens on the
agent booking the wrong meeting, a poor first frame for a card arguing
for the product.

**A copy-prompt button instead of a link out.** It hands the CLI's own
onboarding prompt to a coding agent. Every previous Intelligence CTA
opened a new tab into a signup form, which is where developers drop out.

It carries the CLI's `onboarding_run_id`, so
`oss.inspector.home_prompt_copied` can be joined to
`cli.onboarding.completed` on the Intelligence side. `home_cta_clicked`
only ever proved that someone clicked a link — this is the first event
that can show whether an install followed.

**Section anatomy mirrors System Health** (header band, rule, content),
so the action sits in the same top-right slot the status pill and renew
link already use, and the panel keeps one section shape throughout.

## Correctness of the claims

The copy was checked against the product's own surfaces, and two claims
did not survive:

- **Skills are not applied at run time.** Candidates land at
`pending_review`, a human approves, and the published set is pulled down
with `copilotkit skills download`. Nothing reads published skills during
a run. The slide says approve → pull in → the next run starts from what
worked.
- **Insights were missing**, and with them the evidence link that makes
Learning credible: every Insight cites the Threads and messages behind
it.

Also: *Rich Threads* is the product's name for the durable ones, and the
distinction is the whole sale next to a Threads tab full of local ones
that die on reload. "Your users" means the app's end users — which is
what the platform means too (`identifyUser` resolves one user per
request; a thread carries `end_user_id`, renamed from `user_id` because
the old name *"caused repeated misdiagnosis"*).

## Behaviour worth reviewing

- The story advances **only while Home is visible and the document is
not hidden**. A debugging tool should not hold a repeating timer behind
a closed panel.
- Slide motion is horizontal and derived from each slide's index
relative to the active one, so clicking a tab backwards animates
backwards with no stored direction to fall out of sync.
- Copied state **expires after 4s** so the button invites a second
press; a failed copy **does not**, because that state is the only place
the prompt is selectable by hand.
- Three modes, not two: a lapsed plan keeps the renew link and never
sees an install prompt.
- The rotating copy is hidden from assistive tech (it would announce
four times a loop); one stable sentence is exposed in its place and is
test-covered so it cannot quietly rot.

## Deliberate omissions

**No third-party coding-agent logos on the button**, unlike the
Intelligence app. That app is a private hosted surface; this package
ships inside other people's sites, and vendoring Anthropic's and
OpenAI's marks is not a call to make quietly. The helper line names the
agents in text.

Deferred and worth discussing separately: ordering the Home sections by
state (health first when something is broken, Intelligence first when
nothing is), and putting the same button in the locked Learning and
Threads tabs, where intent is highest.

## Verification

617 tests pass, `check-types` clean, oxlint 0 errors. Verified live in
both themes: all four slides, uniform 16px padding on every slide, card
height stable across slides, copy success **and** failure paths, and the
header band unchanged at 76px when the copied hint appears. The reset
behaviour is mutation-checked — the file records which mutation each
test does and does not catch.
2026-08-28 14:14:39 +02:00
Alem Tuzlak 5686a0669e Merge branch 'main' into lukas/oss-904-runtime-connection-status 2026-08-28 14:06:40 +02:00
Alem Tuzlak 1dfc5cdafa refactor(core): remove OSS-904 design comments 2026-08-28 13:27:37 +02:00
Alem Tuzlak a7191e2a12 fix(core): bound recovery /info hang and tighten OSS-904 comments 2026-08-28 13:08:44 +02:00
Alem Tuzlak b8b35b736c fix(packages): declare the MIT SPDX license on five published packages (#6511)
Five packages publish to npm with no `license` field, so registry
metadata and automated license scanners report them as **Unknown**:

```
@copilotkit/agentcore-runner  published=1.68.1  license=<NONE>
@copilotkit/core              published=1.68.1  license=<NONE>
@copilotkit/sqlite-runner     published=1.68.1  license=<NONE>
@copilotkit/voice             published=1.68.1  license=<NONE>
@copilotkit/web-inspector     published=1.68.1  license=<NONE>
```

The repo is MIT (see `LICENSE`) and every other published
`@copilotkit/*` package already declares it — these five were simply
missed. This adds `"license": "MIT"` to each, positioned before
`"repository"` to match the sibling packages.

## Why

Reported downstream in #2860, where a corporate procurement scan refused
packages whose license it could not resolve. That class of scanner reads
the `license` field from registry metadata; a `LICENSE` file in the repo
is not enough, and these packages ship no `LICENSE` file either.

**Correcting the record on that issue while I am here:** the `@ag-ui/*`
packages named in the original report are *not* affected. Every version
the reporter’s scanner flagged already carries `"license": "MIT"`:

```
@ag-ui/client@0.0.42     MIT
@ag-ui/core@0.0.37       MIT
@ag-ui/core@0.0.42       MIT
@ag-ui/encoder@0.0.42    MIT
@ag-ui/langgraph@0.0.20  MIT
@ag-ui/proto@0.0.42      MIT
```

`@ag-ui/core` has declared MIT since at least 0.0.35. An earlier triage
note on #2860 attributed the failure to a missing SPDX field upstream;
that was wrong, and why their scanner reported `Unknown` for `@ag-ui/*`
is still unexplained. This PR fixes the part that is genuinely defective
on our side.

## Testing

Metadata-only; no source, build, or runtime change.

- Confirmed the five missing fields against the live registry with `npm
view <pkg> license` (output above), and confirmed the other published
`@copilotkit/*` packages (`runtime`, `react-core`, `react-ui`, `shared`,
`sdk-js`, `angular`, `channels`, `channels-core`) already report `MIT`.
- Enumerated every non-private `packages/*/package.json` on
`origin/main` to confirm these five are the complete set missing the
field.
- Each edited file re-parsed with `json.load` and reports `MIT`.
- The `sync-lockfile` pre-commit hook resolved all 71 workspace projects
against the edited manifests without error.

Placement matches `packages/shared/package.json`, where `"license"`
immediately precedes `"repository"`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-08-28 12:32:22 +02:00
Lukas Moschitz 669132d731 fix(web-inspector): stop lit's part marker from failing the usage-footer test
The assertion searched footer.outerHTML for "241" to prove the unclamped
thread count never reaches the user. outerHTML also carries lit's part
markers, which lit builds as `lit$` + nine digits from Math.random(),
regenerated per process. Roughly one process in a hundred rolls a marker
containing those digits and fails the assertion with no relation to what the
footer rendered — this run drew lit$924125892$.

Comments are stripped before the check. Visible text and attributes still
count, so a genuine leak in an aria-label is caught exactly as before, and the
stripping is asserted so it cannot silently stop working.

Not introduced here: the same dice roll could hit any change to this package.
The other outerHTML assertions in the suite compare two strings from the same
process and share a marker, so they were never exposed.
2026-08-28 11:44:42 +02:00
Lukas Moschitz 1efdf37a40 feat(web-inspector): report which story step a developer opens by hand (refs OSS-867)
The rail's four tabs reported nothing. Adds
oss.inspector.home_story_beat_selected, carrying the step as a property rather
than one event per label: the labels are expected to move as the story is
iterated, and per-label events would retire with them. beat_index rides along
so a reorder can be judged against where people actually click.

Only a press reports. The story also advances on its own every few seconds,
and reporting that would emit one event per idle developer per beat and bury
the handful of real interactions under a metronome — asserted, not assumed.
2026-08-28 10:36:59 +02:00
Benjamin Taylor 6ccff843c7 fix(telemetry): stamp sampling metadata and emitter markers on every event
The v2 runtime client gated anonymous events at 5% and let identified
callers through at 100%, then sent without recording which branch the
event took. A quarter of runtime volume — 24.4% in August and roughly
doubling each month — arrived carrying no record of its own sampling, so
it could not be weighted from the data alone. Downstream had to hardcode
a x20 assumption, which both understates real volume and overstates
growth as the sampled/unsampled mix drifts.

Extract the v1 client's sampling block into shared/telemetry/sampling so
the two clients cannot drift again, and call it from both. Identified
events weigh 1, anonymous ones 1/sampleRate.

Carry telemetry_identified explicitly rather than letting consumers infer
identity from sampleWeight === 1: under COPILOTKIT_TELEMETRY_SAMPLE_RATE=1
anonymous events also weigh 1 and the two populations stop being
distinguishable.

The v1 client also sends every capture to both Segment and the lambda
sink, so one request produces two rows with nothing marking them as
copies. Stamp telemetry_emitter, telemetry_transport, and a per-capture
telemetry_event_id shared by both copies, making the dedupe explicit
instead of inferred from $lib. Both transports keep flowing.

Refs OSS-1017, OSS-1018, OSS-1019

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 17:20:34 -05:00
copilotkit-qa-bot[bot] ed315ac592 fix: preserve structured system messages when exposing state 2026-08-27 15:12:55 -07:00
copilotkit-qa-bot[bot] d385b9290e test: assert structured system message content safely 2026-08-27 15:05:45 -07:00
copilotkit-qa-bot[bot] f17d0509db fix: preserve system prompt text when exposing state 2026-08-27 15:05:09 -07:00
copilotkit-qa-bot[bot] 2f1c4c9332 fix: preserve sibling middleware state for exposure 2026-08-27 14:49:56 -07:00
Benjamin Taylor 84dea7bfbb fix(react-native): keep the polyfill imports in the built barrel (closes OSS-1002)
`src/polyfills.ts` is five side-effect-only imports. The `sideEffects` globs only matched
`./dist/**`, and rolldown matches that field against SOURCE paths while bundling, so every
`src/polyfills/*.ts` was declared pure and dropped. Every published version through 1.69.2
shipped a 195-byte barrel installing nothing but streaming fetch, so an app following the
documented setup died on its first runtime call with `Property 'ReadableStream' doesn't exist`.

`dist/index.mjs` and `dist/headless.mjs` lost the same imports, so the package's advertised
auto-install on first import did not happen either.

Add matching `./src/**` globs. The barrel goes 195B to 362B with all five imports, and
`headless.mjs` now leads with `import "./polyfills.mjs"`.

`src/__tests__/polyfills.test.ts` stayed green throughout this, because it imports the source,
which is never bundled and so is never tree-shaken. Add `scripts/verify-polyfill-barrel.mjs`,
which checks `dist/` instead. It clears the nine globals first (Node ships them natively and
Hermes does not, so asserting they are merely "defined" would pass on an empty barrel), then
executes the CJS barrel in a child realm and checks the ESM barrel structurally. ESM cannot be
executed here: the encoding polyfill takes a named import from CommonJS `text-encoding`, which
Metro rewrites to a require() but bare Node ESM rejects.

The check runs from `build`, so a dead barrel fails the build rather than reaching npm.

Also document the `ReadableStream doesn't exist` symptom in troubleshooting, where only the
inverse case (a polyfill *conflict*) was covered before.

Verified: reverting the sideEffects change and rebuilding turns the check red in both formats,
5 of 5 groups; restoring it turns it green. A behavioural check on the packed tarball installs
all nine globals. 289 vitest + 26 script tests pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-27 12:22:02 -05:00
MikeRyanDev 8617f5b76b chore: release monorepo v1.69.3 2026-08-27 16:47:48 +00:00
Lukas Moschitz 9e29de156d fix(angular): stop the core mock from hiding the Inspector's exports
The spec replaced @copilotkit/core wholesale with a two-export factory. That
held until the Inspector started mounting in these tests — it is enabled by
default in browser frameworks now, and its connectedCallback calls
isInspectorThreadBridgeEnabled, one of seventeen value exports it imports from
core. A missing one throws an uncaught exception, so the run fails while all
49 test files still report passing, which is a confusing way to find out.

The factory now spreads the real module and overrides only CopilotKitCore and
the connection-status enum, which is what these tests actually drive. Listing
the seventeen would have postponed the next occurrence rather than removed it.

Surfaced by the web-inspector work on this branch: angular only runs when
affected, and it becomes affected the moment web-inspector changes — so the
first PR to touch web-inspector after the default-on change was going to hit
this regardless of what it changed.
2026-08-27 16:58:39 +02:00
Lukas Moschitz d6b01a47c4 chore(web-inspector): regenerate the stylesheet against current main
The checked-in artifact was produced mid-rebase, so its Tailwind token block
reflects the utilities in use at that point rather than on today's main. A
fresh build:css on the rebased tree differs; this is that output, so the
committed file matches what the build produces.
2026-08-27 16:16:39 +02:00
Lukas Moschitz f011e236cf fix(web-inspector): name the second route instead of promising an explainer (refs OSS-867)
"What Intelligence does" promised an explanation and pointed at
intelligence.copilotkit.ai — the product and signup page. Mis-promising a
destination is a poor trade at the moment the card is asking to be trusted,
and the label was doing no work besides.

The two actions are two routes to one outcome: let the coding agent wire it
up, or go and do it in the browser. The secondary now says so — "Set it up
yourself" — which is honest about where it lands and quietly argues for the
primary, because "yourself" implies the other route is not.
2026-08-27 16:11:32 +02:00
Lukas Moschitz 0e0f49b21c feat(web-inspector): let the copied prompt expire so the button invites a second press (refs OSS-867)
The confirmation used to stand for the rest of the session. The reasoning was
that a developer leaves for their editor and comes back and should still find
the instruction — but whoever comes back has already pasted. The likelier
reason to return is to copy again, and a button wearing a checkmark and
"Prompt copied" reads as spent even though it still works.

Copied now reverts after 4s: button label and secondary line both return, so
the "What Intelligence does" link comes back too. Longer than the 2s the
Threads setup prompt uses, because this state carries an instruction to read
and not just an acknowledgement.

A failed copy does not expire. That state is the only place the prompt text is
selectable by hand, and pulling it away mid-selection would be worse than the
clipboard failing.

Three tests, mutation-checked: disabling the reset kills the two copied-state
tests. The failed-state test is blunter and only fails when both safeguards go
(the scheduling condition and the timer's own state re-check) — recorded in the
file so nobody reads more into it than it proves.
2026-08-27 16:11:32 +02:00
Lukas Moschitz 869c8714a7 fix(web-inspector): stop the copied-prompt hint from growing the header (refs OSS-867)
The instruction was a second row under the button, which pushed the action
column past the band's 76px min-height and shoved the whole story down at the
moment the developer had just acted — the worst possible time for the layout
to move.

It now takes over the secondary slot instead of adding to it. Before the press
the useful aside is "what is this"; after it, "where to put it". Both are
single lines in the same flex slot, so the swap cannot change the height:
measured 76px and an unmoved story in both states.

Dropping the line was the alternative, but it is the one instruction that
cannot go: a prompt on the clipboard with no idea what to do with it converts
nobody. It is also vendor-neutral now — naming three editors was brittle and
the prompt identifies the agent itself.
2026-08-27 16:11:31 +02:00
Lukas Moschitz 83a626f5cd fix(web-inspector): sharpen the Intelligence slide copy (refs OSS-867)
- "Your users had thousands" carried a count a developer wiring this up
  locally does not have yet. "Your users have all the others" holds on day one
  and at scale. Whose users: the app's end users, which is what the platform
  means too — `identifyUser` resolves one user per request from the app, and a
  thread carries `end_user_id`, a column renamed away from `user_id` because
  the old name "caused repeated misdiagnosis" against control-plane users.
- The Skills line defended instead of selling. "Nothing reaches your agent
  until you approve it" answers a fear the reader has not voiced and plants
  the worry it deflects. Same fact as ownership: a SKILL.md that is yours to
  review, edit and ship.
- The last tab is "Intelligence", not "Better agents". The other three are
  the parts; this one is the whole, so the rail reads Threads · Learning ·
  Skills · Intelligence. Beat id renamed to match, selectors included.
2026-08-27 16:11:31 +02:00
Lukas Moschitz e5d406dea4 fix(web-inspector): make the Intelligence pitch match what the product does (refs OSS-867)
Checked the four slides against the product's own copy instead of against
intuition, and two claims did not survive.

- "Skills apply it — without you writing another prompt" said the platform
  applies skills at run time. It does not: candidates land at pending_review,
  a human approves, and the published set is pulled down with
  `copilotkit skills download`. Nothing reads published skills during a run.
  The slide now says approve, pull in, and the next run starts from what
  worked — which is also the stronger pitch, because a developer does not
  want a platform silently changing how their agent behaves.
- Insights were missing entirely, and with them the evidence link that makes
  Learning credible: every Insight cites the Threads and messages behind it.
  Learning's own onboarding leads with "46 evidence refs" for that reason.

Also: "Rich Threads" is the product's name for the durable ones, and the
distinction is the whole sale next to a Threads tab full of local ones that
die on reload. A skill is a directory holding a SKILL.md, so the card shows
`meeting-scheduling/SKILL.md` and is marked Pending review.

- The rail's last step was "Reuse", which named neither a product surface nor
  an outcome. It is "Better agents" — the promise, in the customer's words.
- Secondary link before the primary button: with the filled button in the
  middle it read as a block wedged between the heading and the link instead
  of the one thing to press.
- The heading gets the brand mark and its band a wash along the brand's hue
  path, reusing the account strip's existing device. Not gradient text: at
  18px it renders muddy, costs contrast, and is the most over-used premium
  tell going.
2026-08-27 16:11:30 +02:00
Lukas Moschitz d50e3d7106 feat(web-inspector): turn the Home Intelligence card into an install path (refs OSS-867)
The Inspector said what Intelligence is and linked out to a signup page. Of
~1,655 opens in 90 days, under 100 clicked any CTA. This replaces the feature
list with an argument, and the outbound link with an install that happens in
the editor the developer is already in.

- Copy: the heading names the product; a four-slide argument runs underneath,
  each slide pairing two sentences with the picture beside it (threads →
  the pattern in them → the skill file → it applies itself). An earlier draft
  sold Threads in prose while animating Learning; bound together they read as
  one chain.
- Copy prompt: hands the CLI's own onboarding prompt to a coding agent instead
  of opening a signup form. Carries the CLI's onboarding_run_id, so
  oss.inspector.home_prompt_copied can finally be joined to
  cli.onboarding.completed — home_cta_clicked only ever proved a click.
  A refused clipboard reveals the prompt instead of swallowing it.
- Three modes, not two: a lapsed plan keeps the renew link and never sees an
  install prompt.
- Section anatomy mirrors System Health (header band, rule, content), so the
  action sits in the same top-right slot the status pill and renew link use.
- The story only advances while Home is visible and the document is not
  hidden; a debugging tool should not hold a repeating timer behind a closed
  panel.

No third-party coding-agent logos on the button, unlike the Intelligence app:
this package ships inside other people's sites, and vendoring those marks is
not a call to make quietly. The helper line names the agents in text.
2026-08-27 16:11:28 +02:00
Alem Tuzlak ed28f3a908 fix(angular): export inspector development-mode token from public API 2026-08-27 13:01:01 +02:00
Alem Tuzlak 2dec983ed8 fix(angular): restore web-inspector workspace dependency 2026-08-27 12:15:03 +02:00
Alem Tuzlak 42d3c92fbd chore: merge origin/main into tyler/default-browser-inspector 2026-08-27 12:09:27 +02:00
Alem Tuzlak 0e700fa3b9 docs(react-core): correct Inspector debug-mode skill examples 2026-08-27 12:00:55 +02:00
moonturbo 1662644c29 test(shared): add unit tests for conditions engine 2026-08-27 09:53:50 +08:00
Ben Taylor d02f7e699e docs(react-core): state the blast radius of key-remount and the provisional agent (refs OSS-979) (#6710)
## Why

An Intelligence integration lost a request-to-row correlation map
partway through a user interaction — no error, no warning. It surfaced
as "our response routing is flaky". OSS-979 filed it as
`CopilotKitProvider` remounting its children.

The provider does nothing of the kind. It renders `{children}`
unconditionally at `CopilotKitProvider.tsx:952` — unkeyed, no early
return, and there is no `Suspense` boundary anywhere in `v2`. Nothing in
the SDK silently re-points the active thread either; every mutation path
(`setActiveThreadId`, `startNewThread`, the drawer row click, the
inspector override) is caller-driven.

The remount was app-side, and it was app-side because this skill told it
to be:

- `references/threads.md:98` teaches `useThreads()` → select →
`<CopilotChat key={activeId}>`, and that recipe is only reachable once
Intelligence is wired.
- `references/switching-agents.md:123` teaches "`key={activeAgent}`
forces remount so thread state doesn't leak" without saying what else
that discards.
- `examples/showcases/reskinnable-demo/src/app/[skin]/layout.tsx:223`
models `<SubagentActivityProvider key={threadId}>` above `{children}`,
commented "Remounting is deliberate".

Follow all three and you key a layout-level provider on a thread id that
changes asynchronously after mount. Everything below it dies
mid-interaction.

Two properties made it invisible:

- Durable threads exist only in Intelligence mode, so with a plain SSE
runtime `useThreads` returns nothing, the selected thread never changes,
and the remount never fires. It appears the moment Intelligence is
wired.
- Whether state survives depends on whether the user acted before the
thread list resolved.

## What changed

Docs only — no library change. Both traps now carry their blast radius,
in the four places an agent actually reads:

| File | Change |
|---|---|
| `SKILL.md` | Two invariants in the load-once section, so they land
before any reference is opened |
| `references/threads.md` | New HIGH entry on keying above app state;
note that `activeId` in the switcher recipe settles asynchronously |
| `references/switching-agents.md` | Existing HIGH entry now states the
blast radius and cross-links the threads trap |
| `references/switching-agents-recipes.md` | Key rule amended — keep it
on `<CopilotChat>`, nowhere higher |
| `references/agent-access.md` | The second route to the same symptom:
`useAgent` swaps a provisional stand-in for the real agent when `/info`
resolves, so an effect keyed on `agent` re-runs once, mid-interaction.
Adds an `isReady` pattern and a HIGH entry |

`isReady` appeared in **zero** shipped skills before this — it was
documented only in `showcase/shell-docs/.../useAgent.mdx` and in JSDoc.
Same shape as OSS-888, where the root cause was the shipped skill rather
than the library.

Also corrects a factual error: the skill claimed `useAgent` returns `{
agent }` only. It returns `{ agent, isReady }`.

The 10-file diff is 5 source files under `packages/react-core/skills/`
plus their 5 mirrors under `skills/`, regenerated with `pnpm
sync:plugin-skills`.

## Verification

- `pnpm check:plugin-skills` — mirror in sync
- `pnpm exec vitest run scripts/__tests__/sync-plugin-skills.test.ts` —
12 passed
- `oxfmt --check` — clean over both skill trees
- Full pre-commit suite green, including `test-and-check-packages`
(`test`, `publint`, `attw` across 2 projects and 20 dependent tasks)

## Not in scope

Whether the run's app keyed on `threadId` or on `agent` is not
settleable from the repo — its source is not in any checkout, and there
is no `2026-08-25` strands run report under
`tools/one-prompt-development/evaluation/runs` on any branch. Both
variants produce the reported symptom and this covers both, so a
first-hand repro is a separate task. The `reskinnable-demo` layout is
left as-is deliberately: it is a legitimate use of the pattern, and it
is now the worked example the guidance warns about.

Scoping detail in the OSS-979 comment.

refs OSS-979

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-08-26 14:38:23 -05:00
Ben Taylor 61a67e716b fix(runtime): keep thread naming task after transcript (#6722)
## Summary

- place the thread-title task after the embedded conversation transcript
- explicitly tell the reused agent not to answer the conversation
- add a regression test that locks the prompt ordering

## Why

LangGraph starter agents can interpret the final embedded `user:` line
as the active request when the transcript comes last. They answer the
conversation instead of returning title JSON, causing retries and
eventual `Untitled` thread names.

## Validation

- reproduced with the latest LangGraph starter and
`@copilotkit/runtime@1.69.2`
- original prompt: 0/12 direct calls returned title JSON; 4/4 targeted
threads fell back to `Untitled`
- reordered prompt: 12/12 direct calls returned title JSON; 4/4 targeted
threads received generated titles
- runtime thread-name unit suite: 28/28 passing
- pre-commit affected package checks passing under the repository Node
22 toolchain
2026-08-26 13:44:47 -05:00
Maximiliano Korp eeb01fc33f fix(runtime): keep thread naming task after transcript 2026-08-26 11:28:30 -07:00