Commit Graph

12110 Commits

Author SHA1 Message Date
Jordan Ritter 01079062a3 test(showcase/dashboard): cross-package starter column-slug equality
Only counts (12) were asserted on each side, so a future slug rename on
one side would silently flip a column to grey "not-supported" while both
counts stayed 12. Add a dashboard-side test that fs-reads the harness
STARTER_TO_COLUMN value set (the producer-side remap lives in a separate
pnpm workspace, so it cannot be imported) and asserts SET-EQUALITY with the
dashboard STARTER_COLUMNS. Mirrors the harness starter-mapping-drift fs-read
pattern; reds if either side renames a column slug without the other.
2026-06-04 00:24:11 -07:00
Jordan Ritter 92c2f3d06c fix(showcase/dashboard): correct starter staleness window to hourly cadence
The starter_smoke probe runs hourly (`schedule: "40 * * * *"`), but
STARTER_STALE_AFTER_MS was set to 13h with a comment claiming a 6h cadence
— ~13 missed hourly ticks before amber, defeating the intended two-miss
flip. Re-derive the window from the 1h probe period to 2.5h: strictly
> 2 periods (so two consecutive misses, last row ~3h old, flip amber) yet
< 3h (so a single missed/slow-wake tick, last row ~2h old, stays green,
absorbing a scale-to-zero cold-start). Reconcile both the staleness.ts and
starter_smoke.yml comments to the same hourly basis + 2.5h window in
lockstep (no more "6h" in the starter context). Extend the staleness tests
with explicit hourly-tick boundaries (1 miss → green, 2 misses → amber).

Also drop the inaccurate "hollow" from the not-supported ✗ comments — there
is no hollow render variant; the state renders as grey ✗ text.
2026-06-04 00:24:11 -07:00
github-actions[bot] 2602ad9b01 style: auto-fix formatting 2026-06-04 00:24:11 -07:00
Jordan Ritter f7327bd0a8 ci(showcase): publish per-starter GHCR images; reduce smoke-starter to PR gate
showcase_build.yml: add detect-starter-changes + build-starters jobs building
each of the 12 starters from examples/integrations/<slug>/Dockerfile via Depot
(--platform linux/amd64) and pushing ghcr.io/copilotkit/starter-<slug>:latest +
:<sha>. The starter- prefix is disjoint from showcase-* so harness discovery
stays clean.

test_smoke-starter.yml: drop the 6h schedule cron (live signal now comes from
the harness probing deployed Railway services + the harness alert path); keep
the examples/integrations/** PR build-sanity gate + offline aimock smoke run.

Workflows slot (S4) of the starter-row-group spec (model B). Railway
provisioning (S5) is intentionally out of scope, gated on cost approval.
2026-06-04 00:24:11 -07:00
Jordan Ritter 71de409884 feat(showcase/dashboard): add Starter row-group with 5-state smoke-health vocabulary
Builds the dashboard UI for the starter-smoke dimension: resolveStarterRow
(flat starter:<col>/<level> lookup) + buildStarterBadge implementing the full
5-state cell vocabulary (green ✓, red ✗ smoke-failed, amber ~ stale, gray ?
not-yet-run, grey ✗ not-supported keyed off the S1 mapping), the StarterSection
row-group, rollup-exclusion so starter rows don't skew column aggregates, and a
13h staleness window (> 2× the probe cadence).

UI slot (S3) of the starter-row-group spec (model B).
2026-06-04 00:24:11 -07:00
Jordan Ritter a112b2afed feat(showcase/harness): add starter_smoke probe family for live per-starter health
Adds the `starter_smoke` probe driver + config/probes/starter_smoke.yml that
fans out per-starter HTTP health/agent/chat/interaction levels, registers it
in the orchestrator, and exempts it from probe-config parity (its matrix
shape differs from the depth-dimension probes). Includes 13 driver unit tests.

Probe slot (S2) of the starter-row-group spec (model B): harness HTTP-probes
the deployed (sleepable) Railway starter services.
2026-06-04 00:24:11 -07:00
Jordan Ritter 365740a3e3 feat(showcase/harness): add starter dimension + column→starter mapping
Introduces the `starter` probe dimension in harness types and the single
source of truth for the starter-slug→dashboard-column-slug remap
(STARTER_TO_COLUMN, starterToColumnSlug, STARTER_LEVELS) plus a drift-lint
test asserting every smoke-matrix starter is mapped or explicitly excluded
and every mapped column slug resolves to a real manifest directory.

Foundation slot (S1) of the starter-row-group spec (model B).
2026-06-04 00:24:11 -07:00
Jordan Ritter 9f457dd7ac fix(showcase/harness): retire orphaned e2e_parity probe config
Commit 7ac3e59a5 ("D6 all-pills probe driver") replaced the
`e2eParityDriver` registration (kind `e2e_parity`) with `e2eFullDriver`
(kind `e2e_d6`) in BOTH branches of `registerAllProbeDrivers`. After that
commit NO driver registers kind `e2e_parity`, yet
`config/probes/e2e-parity.yml` still shipped.

The probe-loader hard-rejects any YAML whose `kind` has no registered
driver (`no driver registered for kind 'e2e_parity'`), so `e2e-parity.yml`
failed to load on every boot and was never scheduled — silently dropping
that probe family. The D6/parity dashboard dimension is now produced by the
`e2e_d6` (all-pills) driver, which emits the identical
`d6:<slug>/<featureType>` rows the dashboard reads by key prefix (plus a
`d6:<slug>` aggregate), so `e2e_parity` is genuinely superseded, not just
temporarily unregistered.

Delete the orphaned config so the loader no longer rejects it and the full
probe-config set loads clean. Add a loader test that exercises the REAL
shipped `config/probes` set against the REAL driver registry and asserts
every YAML loads with zero `probes.reload.failed` errors — the existing
orchestrator guard only checks a hardcoded kind list and never read the
on-disk YAMLs, so it missed this orphan.
2026-06-04 00:12:32 -07:00
Jordan Ritter 8578c48cf8 docs(showcase/harness): fix stale RACE3b orphan-close comment after closeContextTracked revert 2026-06-04 00:11:59 -07:00
Jordan Ritter 97aada1304 test(showcase/harness): drop tautological RACE3b assertion + correct orphan-close comments
RACE3b's post-shutdown `liveContextCount === 0` assertion is tautological —
shutdown() sets liveContextCount to 0 unconditionally — so it proves nothing.
Drop it; the load-bearing assertions remain (acquire rejects with the shutdown
sentinel, contextToBrowser.size === 0, the orphan context was closed).

Correct the RACE3a/RACE3b comments that claimed shutdown() awaits the tracked
orphan close (reverted): the straddled orphan is closed by the orphan guard's
fire-and-forget close and never lands in contextToBrowser.
2026-06-04 00:11:59 -07:00
Jordan Ritter b8cf35d83f fix(showcase/harness): revert shutdown orphan-close tracking to fire-and-forget
The empty-drain shutdown case (inFlightRecycles + pendingLaunches both
empty) exits the drain loop and resolves BEFORE a parked open settles and
registers its close via closeContextTracked(), so the "shutdown resolved =>
everything closed" guarantee was never actually delivered and had no
discriminating coverage. Revert the openContextOn orphan guard's close back
to main's accepted fire-and-forget `void this.closeContext(...)` and remove
the closeContextTracked() helper; inFlightRecycles is touched only by the
recycle / self-heal paths again.

Also unify the shutdown-condition error string: shutdown()'s queued-waiter
rejection now says "BrowserPool is shut down" to match the at-entry guard and
the straddle-reject legs (launch/relaunch-failure strings are a different
condition and unchanged).
2026-06-04 00:11:59 -07:00
Jordan Ritter 22074f0d7d test(showcase/harness): add RACE3 coverage for reachable shutdown-straddle settlement branches + relabel RACE2
Item 1 (red-green for the REACHABLE new shutdown branches): RACE1 only
covered the openContextOn orphan-guard `isShutdown` term. RACE3 adds
coverage for the two branches it did not:

  - RACE3a: a shifted waiter whose serve open straddles shutdown must be
    REJECTED by serveNextWaiter's post-open `if (this.isShutdown)` leg
    (not re-enqueued onto the cleared queue, not hung to timeout).
    RED with that guard reverted (waiter never settles), GREEN with it.
  - RACE3b: an acquire whose open straddles shutdown must REJECT via
    acquire()'s OUTER-catch `if (this.isShutdown) throw` leg. RED with
    that guard reverted (the acquire enqueues on the cleared queue and
    the test times out), GREEN with it. Both also assert shutdown()
    awaited the orphan close (item 2) — the straddled context is closed
    and the pool maps are empty post-shutdown.

Item 5: relabel RACE2's comment + title so it is framed as a regression
lock for the PRE-EXISTING generation-guard exactly-once-rollback
invariant (passes against main unchanged), NOT a shutdown-race proof.

Also extends the `internals()` test accessor with `waiters` so the new
tests can assert the queue is empty (no stranded caller).
2026-06-04 00:11:59 -07:00
Jordan Ritter b6c64dfe0b fix(showcase/harness): make shutdown() await orphan context close + tighten shutdown-straddle settlement
Item 2 (reviewer-7 MEDIUM): the openContextOn orphan guard's
`void this.closeContext(...)` under shutdown was fire-and-forget and
tracked in NO set that shutdown() drains, so shutdown() could resolve
while a context close was still in flight — violating the implied
"shutdown resolved => everything closed" contract upheld elsewhere via
the pendingLaunches / inFlightRecycles drains. Route the orphan close
through a new closeContextTracked() that registers the close promise in
inFlightRecycles (the SAME set shutdown's drain loop re-snapshots until
empty) and self-removes on settle. Reuses the existing tracked-set
mechanism — no new machinery; outside shutdown the add/delete is a
no-op since only shutdown's drain loop reads the set.

Item 1 decision (retry-leg guard): KEEP the acquire() transient-retry
leg's `if (this.isShutdown) throw`. Contrary to a reviewer note that it
is unreachable, it guards a DISTINCT, genuinely-reachable straddle
window: the outer catch already observed isShutdown===false, we
re-reserved and re-opened, and the pool can shut down WHILE the retry's
open is in flight (the orphan guard then throws into this catch).
Without it the acquire enqueues onto the already-cleared waiter queue
and hangs until timeout. Comment updated to assert that reachability
rather than imply it merely "mirrors" the outer guard.

Items 3 + 4 (serveNextWaiter straddle): correct the inaccurate comment
that claimed shutdown() "already rejected this waiter" — it cannot, the
waiter was shift()ed off this.waiters BEFORE shutdown's reject-loop ran,
so THIS branch is the sole settler of a shifted-then-straddled waiter.
Unify the shutdown error string on the at-entry guard's existing
"BrowserPool is shut down" (was "shutting down" only here).
2026-06-04 00:11:59 -07:00
Jordan Ritter dc1f0ef894 fix(showcase/harness): harden browser-pool post-shutdown context leak
A serveNextWaiter()/openContextOn() that shifted a waiter + reserved a
slot BEFORE shutdown could settle its newContext() AFTER shutdown()'s
close-pass. shutdown() drains inFlightRecycles + pendingLaunches, but a
fire-and-forget serve open is tracked by neither, so the freshly-opened
context landed in contextToBrowser/liveContexts on a torn-down pool — a
leaked context that is never closed. openContextOn's orphan guard
checked generation/recycling/isConnected but NOT isShutdown.

Fix: add an isShutdown term to openContextOn's post-await orphan guard
(treat shutdown like a recycle — close the just-opened context, roll
back the reservation, throw); bail serveNextWaiter before openContextOn
when isShutdown; and reject (not re-queue) a shifted waiter / straddling
acquire when the open throws under shutdown, so it cannot strand on the
already-cleared waiter queue. Preserves all #5185/#5221 behaviors.

Adds red-green coverage for the leak (RACE1) and a regression guard for
the acquire transient-retry -> concurrent-recycle -> orphan-guard
straddle (RACE2), which was verified to keep its reservation accounting
balanced via the existing generation guard (no overshoot).
2026-06-04 00:11:59 -07:00
Mark 81e97e66e9 Merge branch 'main' into mark/oss-162-a2ui-recovery-client-ux 2026-06-03 23:26:36 -07:00
Mark Fogle 661136084d chore(react-core): add changeset for A2UI recovery renderer (OSS-162)
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 06:23:20 +00:00
Jordan Ritter def354e36b fix(examples/integrations): sync instances to v2 north-star, restore parity-check
The integration-demo parity-check CI job was red on main: langgraph-js,
strands-python, and langgraph-fastapi diverged from the langgraph-python
north-star on two tracked verbatim files.

Root cause is the partial revert in 3721e7b36 (Revert of #5151), NOT #5222.
The v2 API migration (23af69041) moved north-star and all instances to the
v2 surface. The revert rolled the *instances'* example-layout/index.tsx and
docker-route-override.ts back to the v1 API (@copilotkit/react-core,
@copilotkit/runtime) while leaving the north-star on v2
(@copilotkit/react-core/v2, @copilotkit/runtime/v2). That left the demos
internally inconsistent: each instance's real src/app/api/copilotkit route
already uses runtime/v2 on copilotkit 1.59.3, but its Docker route override
and example layout were stuck on v1.

Re-baseline forward by syncing the three instances to the v2 north-star via
`pnpm parity:sync --all`. Only the 5 drifting verbatim files change; no
package.json keys move (all instances already pin 1.59.3) and no
agent-surface or allowed-divergence files are touched. `pnpm parity:check`
is now green (0 errors across all instances).
2026-06-03 23:01:37 -07:00
Jordan Ritter 92726458a2 fix(showcase): make Railway promote flow fault-tolerant (#5200)
## Summary

The showcase Railway promote CI flow could not promote the fleet: an
`all` promote was blocked end-to-end whenever a single service was red,
and the reported failure traced to `showcase-ag2` being chronically red
on staging.

Root causes and fixes:

- **ag2 crash-on-import (the red service).** `gen_ui_agent.py` carried
`from __future__ import annotations`, which stringified the `set_steps`
tool's `context_variables: ContextVariables` parameter into an
unresolved `ForwardRef`. AG2's tool-schema generation then raised
`PydanticUserError` at import time, so the process never came up and the
staging healthcheck failed on every deploy since 2026-05-31. Removed the
import (matching the working sibling agents) and added a regression test
that statically asserts the future-import stays absent
(version-independent) plus a live import check.

- **Promote loop was all-or-nothing.** The per-service loop ran under
`set -euo pipefail`, so the first failing service aborted the whole
`all` promote, leaving the rest unpromoted. Extracted the loop into
`showcase/scripts/promote-fleet.sh`, which attempts every service,
accumulates succeeded/failed sets, exits non-zero only after attempting
all, and exports `succeeded_csv`.

- **`verify-prod` defeated the best-effort design.** It was skipped on
any non-zero promote and verified the full requested set. It now runs
`if: !cancelled()` and scopes `--services` to the succeeded set.

- **Staging precondition blocked the fleet.**
`verify-staging-precondition` failed the whole `all` promote when any
one service was staging-red. It is now advisory — `promote` runs
regardless, and `bin/railway`'s per-service P2/P3 staging-green gates
authoritatively refuse red services while green services promote.
`notify` success keys on PROMOTE && PROD.

- **Regression tests now gate in CI.** Added a `shell-script-tests` job
(bats + shellcheck) to `showcase_validate.yml`, plus input-validation
hardening in the script (fail-loud on empty / all-empty CSV,
`RAILWAY_BIN` executability check, whitespace trim).

## Test plan

- [x] ag2 regression test passes in the 3.12 venv (`PYTHONPATH=".:src"
pytest tests/python/`) — 2/2
- [x] `promote-fleet.bats` — 12/12 (best-effort loop, succeeded_csv
export, empty/whitespace/missing-binary guards, digest forwarding)
- [x] `shellcheck promote-fleet.sh` clean; `actionlint` clean on both
workflows
- [ ] CI green on this PR
2026-06-03 22:26:20 -07:00
Jordan Ritter 8c1100e48f feat(release): automated #engr Slack notification per release (#5220)
## Summary

Adds an automated, **concise** `#engr` Slack notification when a
CopilotKit release publishes (or fails) — modeled on the
internal-skills→#engr release pings, but **one message per release**,
not one per package. The monorepo publishes ~15 npm packages + angular +
the Python SDK; naive per-package posts would be spam, so a single post
summarizes the whole release.

- **New `notify` job** in `publish-release.yml` — runs after both lanes
(`always()`), gated to real release contexts (a `workflow_dispatch` or a
*merged* PR). Posts exactly one message via the existing org-wide
`SLACK_WEBHOOK_ENGR` secret, if-guarded so an unset webhook never breaks
the step. Suppressed entirely for canaries (`mode=prerelease`) and
dry-runs.
- **Release intent is computed in the notify job from the `github.event`
payload** (dispatch inputs, merged `release/publish/*` head ref, and the
PR changed-files API for the Python SDK) — *independent* of the build
jobs. The build jobs emit their intent signals after failure-prone
steps, so a build that dies early on a genuine release would otherwise
leave the alert empty. Computing intent independently closes that
silent-swallow class. Failure arms **page on uncertainty** (fail toward
paging, never toward silence).
- **Pure, unit-tested message builder**
(`scripts/release/lib/build-release-notification.ts`) with a CLI wrapper
that serializes to `$GITHUB_OUTPUT` (random-delimiter heredoc, fail-loud
when it can't emit). Exhaustive truth-table coverage: npm/PyPI success +
failure arms, canary/dry-run suppression, cancelled-is-neutral,
empty-version/empty-scope graceful rendering, package-count
pluralization.
- **Self-watchdog:** if the notify job's own machinery fails, a
best-effort Slack post fires so the failure isn't invisible. The intent
step runs **first** (before checkout/install) so its outputs survive an
infra failure and the watchdog can still fire.

Publish/PyPI logic (the build/publish/build-python/publish-python jobs)
is untouched.

## Review

Converged through a 2-round, 7-agent unbiased code review (zero
mandatory in-scope findings remaining; the one in-scope finding — the
self-watchdog gating on a later step's output — was fixed by relocating
the intent step). Pre-existing publish-lane issues surfaced by reviewers
(prerelease canary verify-step, stable-tag retry idempotency,
canary→PyPI gating) are **outside this diff** and documented separately.

## Test plan

- [x] `vitest run` on the builder + wrapper suites — **85 tests green**
(builder 40, wrapper 27, versions 18)
- [x] `actionlint .github/workflows/publish-release.yml` — zero new
findings (4 pre-existing style findings in untouched jobs)
- [x] Both secret-using Slack steps if-guarded on `env.SLACK_WEBHOOK !=
''`
- [ ] CI green on the PR
- [ ] Observe a real `#engr` post on the next stable release
2026-06-03 22:26:17 -07:00
Jordan Ritter 4eb6cd1916 fix(examples/integrations/langgraph-js): remove stale top-level app/ dir shadowing src/app home route
The starter had two App Router dirs: a stale top-level app/ containing only
an /api/copilotkit route (with a wrong graphId "starterAgent") and the real
src/app/ with page.tsx + layout.tsx. At copilotkit 1.59.3 the build silently
dropped the / route from src/app (masked by next.config ignoreBuildErrors),
producing only /api/copilotkit/[[...slug]]. The app served GET / as 404, so the
docker healthcheck (GET / === 200) never passed and the smoke leg failed with
the app container unhealthy before tests ran.

Removing the duplicate top-level app/ lets Next build src/app unambiguously,
restoring the / route. Local docker-compose.test.yml smoke: app healthy in ~5s,
4 passed (@health, @agent, @chat, @interaction).
2026-06-03 22:03:38 -07:00
Jordan Ritter b70fd72120 fix(examples/mastra): restore REST transport + drop stale agent prop in CopilotKit provider
The v2 migration renamed the Mastra agent key to `default` and updated
page.tsx/index.ts, but layout.tsx was missed: it still pinned
`agent="weatherAgent"` (now non-existent) and lacked
`useSingleEndpoint={false}`. The single-endpoint auto-detect races the
lazily-compiled API route, producing a 404 console error on page load
that fails the @interaction smoke leg's zero-console-error assertion.
Match the passing sibling starters (pydantic-ai/adk/ms-agent-framework-python).
2026-06-03 22:03:38 -07:00
Jordan Ritter c3142d363d feat(examples/integrations): migrate starters to CopilotKit v2 @chat on 1.59.3
Migrate adk, agno, crewai-crews, crewai-flows, mastra, pydantic-ai, ms-agent-framework-dotnet, ms-agent-framework-python, and llamaindex starters to the CopilotKit v2 `@chat` API and float them to 1.59.3 (package.json, runtime route, layout/page source, lockfiles).
2026-06-03 22:03:38 -07:00
Jordan Ritter fed4fb2040 chore(examples/integrations): float strands/langgraph starters to copilotkit 1.59.3
Version-only bump for strands-python, langgraph-fastapi, langgraph-js, and langgraph-python (package.json + lockfile); no source changes required.
2026-06-03 22:03:38 -07:00
Jordan Ritter 1e1379a225 fix(showcase/harness): make the PID/nproc ulimit actually apply under bash
The runtime CMD raised the soft nproc limit via
`/bin/sh -c "ulimit -u $(ulimit -Hu) ..."`, but on node:22-bookworm-slim
`/bin/sh` is dash, whose builtin `ulimit` does NOT support the `-u`
(max-user-processes) flag — it errors `ulimit: Illegal option -u`, which the
`2>/dev/null || true` then silently swallows. So #5185's intended PID
protection never applied; the soft limit stayed at the inherited default.

Run the CMD under `/bin/bash` (present at /usr/bin/bash) instead, whose
`ulimit -u` works. Verified against the exact base image under
`--ulimit nproc=512:4096`: bash raises the soft limit 512 -> 4096 (the hard
ceiling) while dash leaves it untouched and errors. Minimal change — only
the interpreter; the command, exec-as-PID-1, and `|| true` fallback are
unchanged.
2026-06-03 21:02:12 -07:00
Jordan Ritter 4570081297 fix(showcase/harness): fix browser-pool close-during-launch teardown race
A browser that is mid-`launch()` is not yet in `this.browsers` and has no
disconnect handler attached, so it is invisible to every teardown path
(shutdown's close pass, recycle, the disconnect handler). Under the
self-heal/relaunch storm introduced by the browser-pool hardening cluster
(#5174/#5185), a concurrent teardown therefore raced an in-flight launch:
the launch escaped teardown accounting entirely — either leaked, or (in
production) the browser was closed underneath the still-resolving
`chromium.launch()`, which then rejected with `browserType.launch: Target
page, context or browser has been closed` (SIGTRAP). The self-heal loop
relaunched forever, each relaunch hitting the same race (336
self-heal-launch-failed events in 4 min), wedging the pool.

Make launches atomic w.r.t. teardown: `launchBrowser()` registers its
launch promise in a new `pendingLaunches` set BEFORE awaiting
`rawLaunchBrowser()` and removes it on settle. `shutdown()` drains
`pendingLaunches` (alongside `inFlightRecycles`) before its close pass, so
it WAITS for every in-flight launch to settle instead of closing the
browser mid-launch. The launch seam re-checks `isShutdown` the instant the
launch settles: if a shutdown intervened it closes the freshly-launched
browser cleanly THEN, exactly once, and throws a shutdown sentinel so the
caller (init / recycle relaunch / self-heal) does not register a browser
into a pool that is going away. The launch-stagger gate still chains off the
raw result, so serialization semantics are unchanged. All #5185 behaviors
(backoff, self-heal, degraded/recovered alarms, waiter draining) preserved.

Adds a red-green test (FIX#13) reproducing an in-flight launch racing a
concurrent shutdown: pre-fix the launching browser escapes teardown
(closeCount 0, leaked); post-fix it is drained + closed cleanly exactly once
with no close-during-launch rejection.
2026-06-03 21:02:12 -07:00
Jordan Ritter b11e87c33f fix(release): route release notifications to #engr instead of #oss-alerts
Release alerts belong in #engr per corrected routing. SLACK_WEBHOOK_ENGR
is the org-wide secret (visibility=all), so no provisioning is needed.
2026-06-03 20:58:57 -07:00
Jordan Ritter 4ea15b9250 feat(release): post one concise #oss-alerts message per release
Add a notify job that runs after both publish lanes via always() and
computes release intent directly from the github.event payload,
independent of the build jobs, so an infra failure in a lane can't
swallow the alert. The job posts the builder's rendered message to
#oss-alerts and includes a best-effort self-watchdog Slack post.
2026-06-03 20:31:03 -07:00
Jordan Ritter 76bbac5073 feat(release): add tested #oss-alerts release-notification builder
Add a pure builder that renders one concise #oss-alerts Slack message
per release from a unit-tested truth table: suppresses canary/dry-run
runs, gates npm/PyPI failure arms on event-derived release intent, and
pages on uncertain state. The CLI wrapper serializes env input to the
builder and writes the result to GITHUB_OUTPUT, failing loud on any
malformed or missing input.
2026-06-03 20:30:57 -07:00
Jordan Ritter 0aaad742fa fix(showcase/harness): commit detail write before advancing run-row counter
The per-target partial-rollup `runWriter.update` (which advances the durable
`probe_runs.summary.{passed,failed}` counters) ran BEFORE `writer.write`
committed the corresponding `status`/`status_history` detail row, and a
`writer.write` failure is intentionally swallowed so one hiccup can't tank
sibling targets. That ordering opened a run-row-orphan window: the run row
could report `failed: N` for a target whose detail row was never durably
written — the stale-red ingestion artifact.

Reorder so the detail/status row is committed first and the run-row counter
is stamped only after, so the persisted counter can never outrun its backing
detail row. Pure statement-order swap — no schema/migration, dashboard read
path untouched, genuine failures still counted.
2026-06-03 20:22:01 -07:00
Jordan Ritter 8f9aeb8096 test(react-core): pin the Enter-vs-button routing contract while running
CopilotChatInput intentionally diverges while a run is in flight: Enter with
sendable text routes to SEND (the consecutive-interrupt unblock from #5195),
while the send/stop button always routes to STOP. Only a comment bound these
two behaviors together. Add a regression test asserting BOTH at once so a
future refactor cannot silently re-converge them (verified red against both a
button-sends-while-running and an Enter-also-stops mutation).
2026-06-03 19:47:38 -07:00
Jordan Ritter 32551f2bdb fix(react-core): type the active-run completion contract and serialize the suggestion send path
Replace the brittle `"activeRunCompletionPromise" in agent` probe + double
`as unknown as` casts in CopilotChat with a typed `RunCompletionAware`
contract plus an `isRunCompletionAware` type guard, both exported from core.
IntelligenceAgent now declares the property and implements the interface, so
the in-flight-run await is reachable without a cast and non-Intelligence
agents still degrade safely. Cast sites in the attachments/e2e tests and the
MockStepwiseAgent helper are updated to the typed accessor.

Factor the await-then-send logic into a shared `waitForActiveRunToSettle`
helper and call it from BOTH onSubmitInput and handleSelectSuggestion. This
closes the suggestion-path in-flight gap: selecting a suggestion mid-run no
longer pre-empts/aborts the active run (e.g. an interrupt RESUME) — the same
regression PR #5195 fixed for the typed-Enter path. Adds a red-green
regression test that parks the suggestion send on the in-flight promise.
2026-06-03 19:47:38 -07:00
Ben Taylor 33dd22e987 revert: back out the ENT-679 Intelligence threads rollout (#5151, #5196, #5205, #5211) (#5217)
## Summary

Reverts the four remaining Intelligence/threads merges, completing the
back-out started by #5215 (langgraph-fastapi) and #5216 (agent-spec):

- Revert #5211 — mcp-apps
- Revert #5205 — a2a-middleware
- Revert #5196 — pydantic-ai
- Revert #5151 — north-star foundation + batch 1 (incl. crewai-flows
#5189 and llamaindex #5199, which merged into it)

## End-state verification

- `git diff be20a389cf HEAD -- examples/integrations` is **empty** —
every integration example is byte-identical to pre-#5151 main
(`be20a389cf`, the merge's first parent).
- `scripts/__tests__/integration-intelligence-migration.test.ts`
(introduced by the rollout) is deleted.
- Total scope vs main: examples/integrations + that test file. Nothing
else touched.

## Heads-up: collateral on two follow-up commits

@jpr5 — your agent-key standardization commits (4bf50c7c41, 2103d0c875)
were written against the post-#5151 trees (they update `agentId` refs on
ThreadsDrawer/CopilotChatConfigurationProvider etc.), so this revert
necessarily takes those files back to their pre-#5151 content, undoing
the rename. Keeping the rename only where it auto-applied would have
left broken half-states (e.g. adk route registering `default` while the
restored layout pins `agent="my_agent"`). If you still want the rename
on the pre-threads files, it needs a fresh pass on top of this.

Also note: #5151 carried the starter-smoke heals (v2 label keys, 1.59.1
bumps, `useSingleEndpoint`, parity sync). Reverting it returns main's
starter smoke to its pre-#5151 state (9 starters red).

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-03 20:59:21 -05:00
Benjamin Taylor 3721e7b36b Revert "feat(integrations): Intelligence threads — north-star foundation + batch 1 (7 examples on 1.59.1) [ENT-679] (#5151)"
This reverts commit f3ec5ddcec, reversing
changes made to be20a389cf.

# Conflicts:
#	examples/integrations/adk/src/app/layout.tsx
#	examples/integrations/adk/src/app/page.tsx
#	examples/integrations/agno/src/app/layout.tsx
#	examples/integrations/agno/src/app/page.tsx
#	examples/integrations/crewai-crews/src/app/layout.tsx
#	examples/integrations/llamaindex/src/app/api/copilotkit/[[...slug]]/route.ts
#	examples/integrations/llamaindex/src/app/layout.tsx
#	examples/integrations/llamaindex/src/app/page.tsx
#	examples/integrations/mastra/src/app/layout.tsx
#	examples/integrations/mastra/src/app/page.tsx
#	examples/integrations/ms-agent-framework-dotnet/src/app/layout.tsx
#	examples/integrations/ms-agent-framework-dotnet/src/app/page.tsx
#	examples/integrations/ms-agent-framework-python/src/app/layout.tsx
#	examples/integrations/ms-agent-framework-python/src/app/page.tsx
#	examples/integrations/pydantic-ai/src/app/layout.tsx
2026-06-03 20:47:46 -05:00
Benjamin Taylor 1ef370440b Revert "feat(integrations): add Intelligence threads to pydantic-ai (#5196)"
This reverts commit f4fa8e5c02, reversing
changes made to f3ec5ddcec.

# Conflicts:
#	examples/integrations/pydantic-ai/src/app/page.tsx
2026-06-03 20:46:08 -05:00
Benjamin Taylor 5dfafcbfcf Revert "feat(examples): add Intelligence threads to a2a-middleware (#5205)"
This reverts commit a5fbb5c2b5, reversing
changes made to a9d807d8e7.
2026-06-03 20:44:00 -05:00
Benjamin Taylor 797511f9e3 Revert "feat(integrations): add Intelligence threads to mcp-apps (#5211)"
This reverts commit bc8cec6092, reversing
changes made to 2103d0c875.
2026-06-03 20:43:51 -05:00
Ben Taylor ec856ab314 Revert "feat(integrations): add Intelligence threads to agent-spec" (#5216)
Reverts CopilotKit/CopilotKit#5209
2026-06-03 20:35:16 -05:00
Mike Ryan 806c35e4af Revert "feat(integrations): add Intelligence threads to agent-spec" 2026-06-03 18:33:49 -07:00
Ben Taylor 9e7e7e653d Revert "feat(integrations): add Intelligence threads to langgraph-fastapi" (#5215)
Reverts CopilotKit/CopilotKit#5197
2026-06-03 20:32:13 -05:00
Mike Ryan 1e1aadbad8 Revert "feat(integrations): add Intelligence threads to langgraph-fastapi" 2026-06-03 18:28:56 -07:00
Ben Taylor 11e18e378a feat(integrations): add Intelligence threads to agent-spec (#5209)
## Summary
- add env-gated CopilotKit Intelligence Threads support to the
`agent-spec` integration
- wire the Threads drawer + thread context around the existing A2UI
chat/canvas experience
- align `agent-spec` CopilotKit/AG-UI package pins with the north-star
set and commit its npm lockfile
- extend the Intelligence migration contract to cover `agent-spec`

## Verification
- `npm run build` from `examples/integrations/agent-spec`
- `pnpm exec vitest run
scripts/__tests__/integration-intelligence-migration.test.ts`
- `pnpm exec oxfmt --check
scripts/__tests__/integration-intelligence-migration.test.ts
examples/integrations/agent-spec/src/app/page.tsx
'examples/integrations/agent-spec/src/app/api/copilotkit/[[...slug]]/route.ts'
examples/integrations/agent-spec/next.config.ts
examples/integrations/agent-spec/src/components/threads-drawer/threads-drawer.module.css
examples/integrations/agent-spec/src/components/threads-drawer/THEME.md`
- pre-commit hook: `pnpm run test && pnpm run check:packages`

## Smoke test
With local Intelligence services and licensed env:
- `GET /api/copilotkit/info` returned `mode: intelligence`,
`licenseStatus: valid`, `a2uiEnabled: true`, and agent `my_a2ui_agent`
- `GET /api/copilotkit/threads?agentId=my_a2ui_agent&limit=20` returned
200
- opened the UI at `http://localhost:3001`, Threads drawer rendered,
sent “Check my inbox”, run endpoints returned 200, the thread persisted,
and the inbox UI rendered five messages

## Notes
- `agent-spec` is not in the parity manifest, so `pnpm parity:verify
--target=agent-spec` reports `unknown instance: agent-spec`; this PR
uses the migration contract instead.
- During smoke, the runtime logged a warning that the `agents` feature
is not licensed, but the Intelligence thread and agent run still
completed successfully.
- The smoke-rendered chat still shows the agent pre-tool sentence twice
around the inbox tool result; that appears to be existing agent/runtime
message behavior rather than Threads wiring.
2026-06-03 20:22:09 -05:00
Benjamin Taylor f26e4b1d81 Merge origin/main into codex/ent-734-agent-spec
Conflict resolution (contract test): keep main's 6-entry array + appRoots
map, add agent-spec (src/app), and merge her REST-transport relaxation
(layout+page concat — agent-spec's provider is page-level) with main's
appRoots-based paths. 43/43 passing.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-03 20:21:57 -05:00
Jordan Ritter 2cacf48ec0 fix(showcase): migrate dead useInterrupt to useHumanInTheLoop for Strategy-B gen-ui-interrupt demos
The gen-ui-interrupt demo across non-LangGraph integrations used
`useInterrupt`, which only renders in response to an AG-UI `on_interrupt`
event emitted by LangGraph's native `interrupt()` primitive. These
backends never emit that event — they expose `schedule_meeting` as a
frontend/HITL tool over the normal tool-call channel (Strategy B) — so
the picker never mounted.

Migrate the 8 clean Strategy-B integrations (ag2, agno, crewai-crews,
mastra, pydantic-ai, spring-ai, llamaindex, claude-sdk-typescript) to
`useHumanInTheLoop`, mirroring the ms-agent-python / ms-agent-dotnet
reference: same `name: "schedule_meeting"`, same zod `{ topic, attendee }`
parameters, same TimePickerCard render, resolving via `respond(...)`.
The framework-specific comment is generalized for accuracy.

The 3 LangGraph integrations keep `useInterrupt` (native interrupt).
built-in-agent and claude-sdk-python already use the equivalent working
`useFrontendTool` pattern and are left unchanged. strands and langroid
are intentionally NOT migrated — their backends declare a
`schedule_meeting(reason)` tool whose param shape conflicts with the
`topic`/`attendee` reference, which needs separate resolution.
2026-06-03 16:38:57 -07:00
Tyler Slaton 55465ded15 fix(skills): resolve Snyk W007/W011 findings in copilotkit-setup skill (#5208)
Fixes the [skills.sh Snyk
findings](https://www.skills.sh/copilotkit/copilotkit/copilotkit-setup/security/snyk)
on the `copilotkit-setup` skill.

**W007 (HIGH, insecure credential handling) — clears this.**
Removed the exact secret-shaped placeholders Snyk flagged
(`OPENAI_API_KEY=sk-...`, `licenseKey="ck_..."`) across `SKILL.md`, both
runtime assets, and `telemetry-setup.md`; replaced with `<your-...>`.
Made env vars the only path (dropped the inline license-key option), and
added `.gitignore`/secret-manager guidance.

**W011 (MEDIUM, indirect prompt injection) — improves, TBD.**
Added a Security notes section (untrusted chat input, least-privilege
`defineTool` handlers, runtime endpoint auth). CopilotKit's entire
purpose is piping user chat into an LLM, so the flagged data-flow
genuinely exists — we'll have to see whether the added guidance is
enough to satisfy the analyzer.

Also bumps the skill version 1.1.0 → 1.1.1.
2026-06-03 16:16:01 -07:00
Ben Taylor bc8cec6092 feat(integrations): add Intelligence threads to mcp-apps (#5211)
## Summary
- add env-gated CopilotKit Intelligence Threads support to the
`mcp-apps` integration
- move the runtime route to `[[...slug]]` with GET/POST/PATCH/DELETE for
thread REST routes
- wire the Threads drawer + locked gate into the existing MCP apps chat
surface
- align CopilotKit/AG-UI versions and commit npm lockfiles for the app
and Three.js MCP server
- extend the Intelligence migration contract to include `mcp-apps`,
including the app-router path differences

## Verification
- `npm run build` from `examples/integrations/mcp-apps`
- `pnpm exec vitest run
scripts/__tests__/integration-intelligence-migration.test.ts`
- `pnpm exec oxfmt --check
scripts/__tests__/integration-intelligence-migration.test.ts
examples/integrations/mcp-apps/app/page.tsx
examples/integrations/mcp-apps/app/components/threads-drawer/threads-drawer.tsx`

## Licensed smoke
With local Intelligence services from
`/Users/mothra/Projects/test-signups4` and the example `.env` exported:
- `docker compose up -d --wait` reported postgres, redis, and
intelligence healthy
- demo users `1_demo-user` and `demo-user` were seeded/no-op present
- `GET /api/copilotkit/info` returned `mode: intelligence`,
`licenseStatus: valid`, and agent `default`
- `GET /api/copilotkit/threads?agentId=default&limit=20` returned 200
and the Threads drawer rendered persisted threads
- started `npm run dev:mcp` and `next dev --turbopack -p 3024`
- created a new thread, sent “Say hello briefly for the smoke test.”,
`/agent/default/run` returned 200, and the assistant rendered “Hello!”
- after restarting the dev servers, selected the persisted “Quick
Greetings for Testing” thread; the prior user prompt and assistant
response restored, with `/agent/default/connect` returning 200

## Notes
- `3000` was occupied by another local example, so smoke used `3024` for
the Next UI.
- The local thread list contained older experimental thread names from
previous mcp-apps smoke runs; this PR now clears the client thread id
for “New thread” instead of minting a random UUID, so new thread
creation/replay works through the runtime.
2026-06-03 18:04:04 -05:00
Benjamin Taylor a42eb45997 Merge origin/main into codex/ent-734-mcp-apps
Post-merge fix: add the missing examples/integrations/mcp-apps/.env.example
(+ .gitignore negation, crewai-flows precedent) — the contract test
'mcp-apps documents the local Intelligence environment' was failing
because the file was never added. 37/37 passing on the merged tree.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-03 18:03:32 -05:00
Jordan Ritter 2103d0c875 refactor(examples): standardize mastra starter agent key to default 2026-06-03 16:01:02 -07:00
Jordan Ritter 4bf50c7c41 refactor(examples): standardize starter agent key to default
Rename the registered agent key from each starter's bespoke name
(my_agent / agno_agent / starterAgent / sample_agent) to "default" so the
7 single-agent sidebar starters match the passing langgraph/strands
starters' config exactly: drop the frontend `agent=` prop on <CopilotKit>
(falls back to "default") and update the runtime route registration plus
every agentId reference in page.tsx (ThreadsDrawer,
CopilotChatConfigurationProvider, useAgent/useCoAgent).

The functional `useSingleEndpoint={false}` fix that makes starter chat
reach the agent already landed on main; this change brings the agent-key
config in line with the passing set. mastra is intentionally left on its
framework-derived key (MastraAgent.getLocalAgents) — its keys are not
literals and cannot be forced to "default" without diverging from Mastra.

Verified locally via agno docker-compose.test.yml smoke (aimock): @chat
round-trip passes (assistant streams in ~1.0s) with the standardized
"default" key; removing useSingleEndpoint={false} reproduces the documented
404 / "no assistant response" failure (chat degrades to ~1.1m, @interaction
fails on single-endpoint 404s).
2026-06-03 16:01:02 -07:00
Austin Merrick d3a86fe217 fix(skills): resolve Snyk W007/W011 findings in copilotkit-setup
W007 (HIGH, insecure credential handling):
- Replace secret-shaped placeholders (sk-..., ck_...) with descriptive
  <your-...> placeholders in SKILL.md, both runtime assets, and
  telemetry-setup.md
- Make env vars the only path for API keys and the license key (drop the
  inline license-key option); source apiKey from process.env, never a literal
- Add .gitignore / secret-manager guidance for env files

W011 (MEDIUM, indirect prompt injection):
- Add a 'Security notes' section covering untrusted chat input,
  least-privilege defineTool handlers, and runtime endpoint auth

Bump skill version 1.1.0 -> 1.1.1.
2026-06-03 15:56:12 -07:00
Ben Taylor a5fbb5c2b5 feat(examples): add Intelligence threads to a2a-middleware (#5205)
## Summary
- add env-gated CopilotKit Intelligence runtime wiring to
`examples/integrations/a2a-middleware`
- move the app onto the v2 provider/chat thread context with the
reusable Threads drawer
- preserve the A2A research/analysis/orchestrator URL flow while
isolating runtime agent instances for thread/title generation
- pin the standalone example to the threads-capable CopilotKit/AG-UI
packages and document local Intelligence env vars
- extend the integration migration regression test coverage for
a2a-middleware

## Verification
- `pnpm exec vitest run
scripts/__tests__/integration-intelligence-migration.test.ts`
- `pnpm exec oxfmt --check
scripts/__tests__/integration-intelligence-migration.test.ts
examples/integrations/a2a-middleware/app/page.tsx
examples/integrations/a2a-middleware/components/chat.tsx
'examples/integrations/a2a-middleware/app/api/copilotkit/[[...slug]]/route.ts'
examples/integrations/a2a-middleware/next.config.ts
examples/integrations/a2a-middleware/postcss.config.mjs`
- `npm run build` from `examples/integrations/a2a-middleware`
- local Intelligence stack: `GET
/api/copilotkit/threads?agentId=a2a_chat` returned `200`
- Playwright smoke check: licensed Threads drawer renders on desktop,
mobile has no horizontal overflow after responsive layout fix

## Notes
- The full live A2A agent response path was not exercised locally
because this shell did not have a Google/Gemini key available for the
Python agents. The Next build still completes; missing local A2A
services only produce the existing non-fatal agent-card fetch noise
during static collection.
2026-06-03 17:52:10 -05:00