Commit Graph

12110 Commits

Author SHA1 Message Date
Alem Tuzlak ba159ce8e1 chore(examples): bump slack example to @copilotkit/bot* 0.0.2
@copilotkit/bot, @copilotkit/bot-slack, and @copilotkit/bot-ui are published at
0.0.2 (agent-native assistant pane + native streaming). Point the slack
example's dependency ranges at ~0.0.2 so a deployed/standalone install pulls
the new packages. Local monorepo installs already use the workspace copies via
root pnpm.overrides, so the lockfile is unchanged.
2026-06-16 14:22:50 +02:00
Alem Tuzlak 7ed4a44447 chore: release bot v0.0.2 (#5475)
## Release bot v0.0.2

**Scope:** `bot` | **Bump:** `patch`

---

### How this release process works

1. **This PR was created automatically** by the "release / create-pr"
workflow.
   It bumped the `bot` packages to `0.0.2`
   and generated AI-enhanced release notes.

2. **CI runs on this PR** — the full test suite (unit tests, lint, type
checks, build)
   must pass before merging. This is the review gate.

3. **Review the release notes** in `release-notes.md` in this PR.
If a Notion draft was created, you can edit the release notes there
before merging.

4. **When this PR is merged**, the `release / publish` workflow
automatically:
   - Builds all packages
   - Publishes the `bot` packages to npm at version `0.0.2`
   - Creates git tag `bot/v0.0.2`
   - Creates a GitHub Release with the final release notes

### Before merging

- [ ] CI is green (tests, lint, types, build)
- [ ] Version bumps look correct
- [ ] Release notes are accurate (edit in Notion if a draft was created)

---

> **Do not merge until CI is fully green.** The full test suite runs
automatically on this PR.
bot/v0.0.2
2026-06-16 13:58:23 +02:00
AlemTuzlak 035e7cf9c9 chore: release bot v0.0.2 2026-06-16 11:43:12 +00:00
Alem Tuzlak 72203fbcba chore: release bot-slack v0.0.2 (#5466)
## Release bot-slack v0.0.2

**Scope:** `bot-slack` | **Bump:** `patch`

---

### How this release process works

1. **This PR was created automatically** by the "release / create-pr"
workflow.
   It bumped the `bot-slack` packages to `0.0.2`
   and generated AI-enhanced release notes.

2. **CI runs on this PR** — the full test suite (unit tests, lint, type
checks, build)
   must pass before merging. This is the review gate.

3. **Review the release notes** in `release-notes.md` in this PR.
If a Notion draft was created, you can edit the release notes there
before merging.

4. **When this PR is merged**, the `release / publish` workflow
automatically:
   - Builds all packages
   - Publishes the `bot-slack` packages to npm at version `0.0.2`
   - Creates git tag `bot-slack/v0.0.2`
   - Creates a GitHub Release with the final release notes

### Before merging

- [ ] CI is green (tests, lint, types, build)
- [ ] Version bumps look correct
- [ ] Release notes are accurate (edit in Notion if a draft was created)

---

> **Do not merge until CI is fully green.** The full test suite runs
automatically on this PR.
bot-slack/v0.0.2
2026-06-16 13:26:27 +02:00
Jordan Ritter e941049d3c fix(harness): cascade fallback to whole-bubble-minus-toolbar for tool-only responses (Class B) (#5474)
## Summary

Adds a **Class B cascade fallback** in `readCascadeStateLast` so the
harness settle gate can resolve text for **single-bubble tool-only
responses** (recharts SVG, gen-UI cards, A2UI render-only output). These
bubbles have ALL scoped text selectors empty but carry substantial
rendered content in non-cascade children. The current cascade-pollution
guard correctly returns `null` for arbitrary-index reads
(`readCascadeState` / `findAssistantBubbleAt`) — but for the LAST
bubble, by `RUN_FINISHED` the content is mounted and stable, so we can
safely fall back to `bubble.textContent` MINUS the assistant-toolbar's
textContent (the suffix the toolbar contributes).

This is the residual fix on top of PR #5472 (Class A: multi-bubble
last-bubble-has-prose) and #5473 (probe taxonomy cleanup). The ~200-red
plateau on PocketBase after #5472 landed is consistent with Class B
affecting the long tail of tool-render-only demos
(ms-agent-python:beautiful-chat-*, plus other tool-render-only
Class-B-shaped demos).

## Class B failure mode (live-confirmed)

Probed against
`https://showcase-ms-agent-python-staging.up.railway.app/demos/beautiful-chat`
with the prompt _"Show me a bar chart of monthly sales for Q1 2026."_:

- `count = 1` (single canonical bubble)
- `bubble.querySelector('[data-message-content]')` -> empty string
- `bubble.querySelector('.cpk:prose')` -> empty string
- `bubble.querySelector('.prose')` -> empty string
- `bubble.textContent` -> **351 chars** of real content (chart title,
axis labels, animation styles, query metadata)
- Current `readCascadeStateLast` returns `text=null` -> settle gate
spins on `text-unstable` until timeout (RED).

## Red-green proof (verbatim from /tmp/cr/class-b-red-green.mjs)

The probe runs TWO in-browser readers against the same DOM snapshot —
the OLD (current production) and the NEW (with fallback). Run live
against staging, **before any code changes** in this PR:

```
[probe] navigating to https://showcase-ms-agent-python-staging.up.railway.app/demos/beautiful-chat
[probe] locating chat input
[probe] sending prompt: Show me a bar chart of monthly sales for Q1 2026.
[probe] waiting 20000ms for response to render
[probe] result: {
  "count": 1,
  "oldText": null,
  "newText": "query_dataquery:\"monthly sales for Q1 2026\"\n        @keyframes barSlideIn {\n          from { transform: translateY(40px); opacity: 0; }\n          20% { opacity: 1; }\n          to { transform: translateY(0); opacity: 1; }\n        }\n      Monthly Sales for Q1 2026Breakdown of income generated each month in Q1 2026.JanuaryFebruaryMarch0255075100January",
  "oldLen": null,
  "newLen": 351,
  "newPreview": "query_dataquery:\"monthly sales for Q1 2026\"\n        @keyframes barSlideIn {\n          from { transform: translateY(40px); opacity: 0; }\n          20% { opacity: 1; }\n          to { transform: translat"
}
[probe] verdict: {
  "red": true,
  "green": true,
  "count": 1
}
[probe] RED-GREEN confirmed: old=null, new=non-empty.
```

- **RED**: `oldText = null`, `oldLen = null` — matches current
production behavior; settle gate cannot resolve text.
- **GREEN**: `newText` is 351 chars of stable rendered content; settle
gate can lock in `text-stable`.

## The fix

In `showcase/harness/src/probes/helpers/assistant-message-count.ts`,
inside `readCascadeStateLast`'s per-tier loop — AFTER the existing
scoped-selector cascade exhausts without a non-empty match and BEFORE
returning `{count, text: null}`:

1. Read `bubble.textContent` (whole bubble).
2. Read
`bubble.querySelector('[data-testid="copilot-assistant-toolbar"]').textContent`
if present.
3. If the toolbar's text is a **trailing suffix** of the whole-bubble
text (which it canonically is — toolbar is a leaf sibling of the prose
div), strip it.
4. If what remains is non-empty after trim -> return `{count, text:
<trimmed>}`.
5. Otherwise (no toolbar, no content, or only-toolbar text) -> return
`{count, text: null}` to keep polling.

The same fallback is NOT applied to `findAssistantBubbleAt` /
`readCascadeState` because those address arbitrary indices and
intermediate bubbles in multi-step responses can transiently carry empty
scoped text while the NEXT bubble is mounting. Reading whole-bubble at
intermediate indices would re-introduce the cross-bubble text flap PR
#5462 was designed to prevent. The LAST bubble in a multi-step turn is
the agent's terminal output, mounted by `RUN_FINISHED`.

## Why the fallback is safe (toolbar-suffix invariant)

In the canonical CopilotKit assistant bubble
(`CopilotChatAssistantMessage.tsx`), the toolbar is the **last child**
of the bubble, mounted as a leaf sibling of the prose/markdown wrapper.
Its `textContent` is therefore the trailing slice of
`bubble.textContent`. Stripping it by suffix-match yields the message
content. When the suffix-match fails (defensive — e.g. the DOM shape
changes), we return `bubble.textContent` unchanged rather than mangle
it. When the toolbar isn't present (headless / custom-composer), we
return whole text. When even whole-text is empty, we return `null`
(don't lock the settle gate on a placeholder).

## Tests

- `assistant-message-count.test.ts`: 9 new test cases under a new
`describe("readCascadeStateLast")` block. Pins:
  - non-empty scoped text path unchanged (no fallback when scoped works)
- Class B fallback: scoped empty + whole-bubble non-empty + toolbar
suffix -> strip suffix, return content
  - Class B fallback works without a toolbar (returns whole text)
  - defensive: toolbar text NOT a suffix -> return whole text unchanged
  - everything empty -> return `text:null` (settle gate keeps polling)
  - only-toolbar text -> return `text:null` after strip
  - multi-bubble: reads the LAST bubble, not the first
  - no tier matches -> `{count:0, text:null}`
  - `evaluate()` throws -> swallowed -> `{count:0, text:null}`
- Full harness suite: **2769 / 2769 passing across 129 files** (`pnpm
exec vitest run` in `showcase/harness`).
- TypeScript: `pnpm exec tsc --noEmit` clean.
- Lint: `oxlint` 0 warnings, 0 errors.
- Format: `oxfmt --check` clean.

## Out-of-scope (untouched)

- `conversation-runner.ts` (the settle-gate logic stays — the cascade
does the right thing now).
- `sse-interceptor.ts`, `init-scripts.ts`, `d6-all-pills.ts`.
- `findAssistantBubbleAt` / `readCascadeState` (intentionally retain the
pollution guard — they address arbitrary indices).
- `resolveBubbleTextFromSelectors` (pure sibling of
`findAssistantBubbleAt`, intentionally retains pollution-guard
semantics).

## Verification plan after merge

1. Watch CI green on the PR.
2. Promote `showcase-ms-agent-python` and other tool-render-only
services.
3. Observe PocketBase D6 red counts: expect the ~200-red plateau to drop
further toward baseline.
4. If any demo regresses, revert `readCascadeStateLast` alone.

## Files changed

- `showcase/harness/src/probes/helpers/assistant-message-count.ts`
(production fix)
- `showcase/harness/src/probes/helpers/assistant-message-count.test.ts`
(9 new tests)
2026-06-16 04:22:21 -07:00
Jordan Ritter 72520de43d fix(harness): cascade last-bubble whole-bubble-minus-toolbar fallback for tool-only responses 2026-06-16 04:05:49 -07:00
Jordan Ritter f7779f90bf refactor(showcase): probe taxonomy cleanup — drop Smoke, rename E2E (Demo) → UI (Frontend) (#5473)
## Summary

The smoke probe was the same HTTP contract as `/health` on the same
service (200-OK JSON body), so every tick paid two HTTP calls for the
same liveness signal. This PR:

- Drops the `/smoke` GET probe and the `smoke:<slug>` primary
`ProbeResult`. The driver now emits `health:<slug>` as the primary and
`agent:<slug>` via writer side-emit (half the per-tick cost).
- Renames the dashboard's D3 / e2e row label from `E2E (Demo)` to `UI
(Frontend)` (and the short cell badge from `E2E` to `UI`) — the row was
already about "the demo page renders in a browser", not E2E in the
traditional sense.
- Drops the `Smoke` row from the cell drilldown (CellState.smoke field
is retained for back-compat so historical rows still parse).
- Adds an explicit `Health` row to the legend (was implicit).

The driver's registry `kind` stays `"smoke"` so existing YAML configs +
orchestrator family wiring keep routing to this driver — the emission
key (`health:<slug>`) is the taxonomy contract that matters. The
underlying probe key (`e2e:<slug>/<feature>`) is preserved on PocketBase
so historical rows render correctly during the rename window.

## RED proof (BEFORE)

`smoke:<slug>` primary emission in `liveness.ts` (origin/main):

```
18: *   1. `smoke:<slug>`  — the RETURN VALUE of `run()`. The invoker runs
56: * `smoke:<slug>`/`health:<slug>`/`agent:<slug>` triple the dashboard
241:        input.mode === "discovery" ? `smoke:${slug}` : input.key;
```

`E2E (Demo)` / `Smoke` strings in dashboard render code (origin/main):

```
cell-drilldown.tsx:38: * Trip)" (chat+tools round-trip), D3/e2e = "E2E (Demo)" (the demo page loads
cell-drilldown.tsx:49:  { key: "e2e", label: "E2E (Demo)" },
cell-drilldown.tsx:52:  { key: "smoke", label: "Smoke" },
cell-pieces.tsx:382:        name="E2E"
unified-cell.tsx:231:      <TestBadge name="E2E" level={model.d3} />
adaptive-legend.tsx:59:        E2E (Demo): demo page loads and round-trips in a browser
```

## GREEN proof (AFTER)

Primary emission rewritten — driver returns `health:<slug>` and no
`smoke:<slug>` ProbeResult is ever emitted:

```
33: * to `/smoke` and emitted a `smoke:<slug>` primary result, but that probe
230:  // `kind: "smoke"` is the registry/family identifier the orchestrator
236:  kind: "smoke",        // registry-key back-compat; no smoke ProbeResult is emitted
367:  if (input.mode === "discovery") return `health:${slug}`;
368:  if (input.key.startsWith("smoke:")) {
369:    return `health:${input.key.slice("smoke:".length)}`;
```

New labels in dashboard:

```
cell-drilldown.tsx:60:  { key: "e2e", label: "UI (Frontend)" },
cell-pieces.tsx:382:        name="UI"
unified-cell.tsx:231:      <TestBadge name="UI" level={model.d3} />
adaptive-legend.tsx:65:        UI (Frontend): demo page renders in browser (Playwright)
```

Confirmation that the old labels are gone from render code:

```
$ grep -rn 'E2E (Demo)\|"Smoke"' cell-drilldown.tsx cell-pieces.tsx unified-cell.tsx adaptive-legend.tsx
(no matches — labels removed)
```

## Test results

- `pnpm exec vitest run` in `showcase/shell-dashboard`: **63 files, 1089
tests passed, 1 skipped**
- `pnpm exec vitest run` in `showcase/harness`: **129 files, 2760 tests
passed**
- `pnpm exec tsc --noEmit` in both: clean

Liveness driver tests include a regression guard (`never emits a
smoke:<slug> key`) so the smoke contract cannot be re-introduced
silently.

## Test plan

- [x] Dashboard vitest green (1089/1089)
- [x] Harness vitest green (2760/2760)
- [x] tsc --noEmit green in both packages
- [x] Visual cross-check: drilldown renders 6 rows (no Smoke), e2e row
labelled `UI (Frontend)`; cell strip shows `UI` badge in place of `E2E`;
legend shows explicit `Health` row + `UI (Frontend)` D3 line
2026-06-16 01:57:15 -07:00
Jordan Ritter 53f021afe8 refactor(showcase): probe taxonomy cleanup — drop Smoke, E2E (Demo) → UI (Frontend)
The smoke probe was the same HTTP contract as /health on the same
service (200-OK JSON body), so every tick paid two HTTP calls for the
same liveness signal. Drop the /smoke GET + the smoke:<slug> primary
ProbeResult; the driver now emits health:<slug> as the primary and
agent:<slug> via writer side-emit (half the per-tick cost).

The driver's registry kind stays "smoke" so existing YAML configs and
orchestrator family wiring keep routing to this driver — the
emission key (health:<slug>) is the taxonomy contract that matters.

Dashboard:
- Drilldown: D3/e2e row labelled "UI (Frontend)" (was "E2E (Demo)");
  Smoke row dropped (CellState.smoke field retained for back-compat).
- Cell badges: short label "UI" (was "E2E") in cell-pieces + unified-cell.
- Legend: D3 = "UI (Frontend): demo page renders in browser (Playwright)"
  plus an explicit Health row at the top.

Tests updated to assert new labels:
- cell-drilldown.test.tsx: 6 dimensions (no Smoke); UI (Frontend) label.
- cell-drilldown.lazy-signal.test.tsx: drilldown-badge-ui--frontend- testid.
- cell-pieces.test.tsx + .signal-degrade.test.tsx: badge name "UI".
- unified-cell.test.tsx: mock-badge-UI.
- overlay-selector-integration.test.tsx: "UI" in place of "E2E".
- dashboard-color-matrix.test.tsx: badge: "UI" case names.
- liveness.test.ts: two-call contract (health + agent), regression guard
  asserting no smoke:<slug> ever emitted.

Underlying probe key (e2e:<slug>/<feature>) preserved on PocketBase so
historical rows render correctly during the rename window.
2026-06-16 01:48:41 -07:00
Jordan Ritter 665dd5aae6 fix(harness): waitForTurnComplete reads LAST bubble (multi-step agent regression from #5462) (#5472)
## Summary

**URGENT hotfix.** Staging is widely red for D5/D6 multi-step demos.
Root cause confirmed via live Playwright DOM inspection on staging
(langgraph-typescript : beautiful-chat : pie chart prompt).

PR #5462's defect-2 fix uses `bubbleIndex = turnIndex - 1`, which
assumes one assistant bubble per turn. Multi-step agents (LangGraph,
Mastra, CrewAI, llama-index, ag2, etc.) emit 2-3+ bubbles per turn —
tool-call bubble + tool-render bubble + final-text bubble. The
strict-index read lands on an intermediate **tool-call** bubble whose
scoped-text selectors (`.cpk:prose`, `[data-message-content]`, `.prose`,
`p`) are **empty** (tool-call content lives in a sibling card div
outside the scoped cascade). The cascade-pollution guard correctly
returns `null` → the text-stable conjunct never converges →
`reason=text-unstable` timeout across virtually every multi-step demo
(beautiful-chat-*, gen-ui-*, shared-state-*, agentic-chat on most
backends).

## Live RED-GREEN proof (staging)

Standalone Playwright probe against
`showcase-langgraph-typescript-staging.up.railway.app/demos/beautiful-chat`,
prompt = `"pie chart of revenue distribution by category from the sample
sales data"`, captured BOTH the OLD-logic read (`bubbleIndex = turnIndex
- 1 = 0`) AND the NEW-logic read (`bubbleIndex = count - 1 = 2`) in ONE
atomic page evaluate after the turn settled:

```
TOTAL BUBBLE COUNT FOR TURN 1: 3  (true multi-step — tool-call + tool-render + final-text)
FINAL TIER: [data-testid="copilot-assistant-message"]

OLD LOGIC (bubbleIndex = turnIndex - 1 = 0):
  count = 3
  text  = null
  reason = all scoped selectors empty           ← would time out: reason=text-unstable

NEW LOGIC (bubbleIndex = count - 1 = 2):
  count = 3
  text  = "Here is the pie chart showing the revenue distribution by category from the sample sales data."
  textLen = 94                                  ← gate settles cleanly
```

RED ≠ GREEN; new logic returns 94 chars of stable final-text where old
logic returns `null`. The bug and the fix are proven on real staging
DOM.

### Live DOM evidence
LangGraph TS beautiful-chat, prompt = "pie chart of revenue distribution
by category…":
```
count = 3 bubbles in [data-testid="copilot-assistant-message"]
bubble[0]: tool-call "query_data"   — `.cpk:prose` exists but EMPTY (content in sibling <div class="my-1.5">)
bubble[1]: pie-chart card           — `.cpk:prose` exists but EMPTY (content in sibling styled card div)
bubble[2]: final text "Here is the pie chart..." — `.cpk:prose` has 94 chars + toolbar mounted
```

The old pre-#5462 runner read `list[last]` (always the last bubble —
always bubble[2] above). The PR moved to `turnIndex - 1`, which is
bubble[0] (the tool-call) for any multi-step turn → broken.

## Fix

Read the LAST bubble in the matched cascade tier (`count - 1`) instead
of a fixed strict index. Defect-2 protection (don't read a leftover
bubble from a prior turn) is preserved via an explicit **pre-submit
baselineCount sentinel snapshot** in the runner — the gate now requires
`count > baselineCount` rather than `count > turnIndex - 1`. This
correctly rejects stale prior-turn bubbles **without assuming 1 bubble =
1 turn**.

### Changes
- `assistant-message-count.ts`: add `readCascadeStateLast(page)` — same
atomic single-evaluate contract as `readCascadeState`, but resolves the
bubble index internally as `count - 1`.
- `conversation-runner.ts`:
- Snapshot `baselineCount = countAssistantMessages(page)` BEFORE
`sendTurnMessage()` each turn.
- Pass through to `waitForTurnComplete` via new `baselineCount` option
(defaults to `turnIndex - 1` to preserve unit-test fake compatibility).
- `waitForTurnComplete` reads via `readCascadeStateLast` and gates on
`domOk = count > baselineCount`.
  - Cold-start retry re-snapshots `baselineCount` after `page.reload()`.
- Returned `bubbleIndex` in the success ctx = `count - 1` (matches the
bubble whose text settled the gate).
- `dom-missing` reason classifier updated to `countFinal <=
baselineCount`.

## Why this doesn't regress defect-2

The original defect-2 (un-turn-scoped bubble selection) leaked a later
turn's bubble into THIS turn's assertions because
`readLastAssistantText` read `list[last]` GLOBALLY with no per-turn
baseline. The new fix re-introduces the "read last" semantics but adds
the **per-turn pre-submit baseline count snapshot** — the gate only
advances once a NEW bubble has appeared (`count > baselineCount`), so we
cannot read a leftover from a prior turn, and the "last" we return is
the last NEW bubble (which for multi-step turns is the final-text bubble
whose scoped text actually settles).

## Test plan

- [x] **Live RED-GREEN staging probe** (see above) — confirms OLD logic
= null and NEW logic = 94-char final-text on a real 3-bubble multi-step
turn.
- [x] All 72 `conversation-runner.test.ts` +
`assistant-message-count.test.ts` unit tests pass (default
`baselineCount = turnIndex - 1` keeps existing count-progression scripts
working).
- [x] Full harness vitest suite: 2758/2758 green across 129 test files.
- [x] `tsc --noEmit` clean on `showcase/harness`.
- [x] `oxfmt --check` clean on changed files.
- [x] `oxlint` clean on changed files.
- [ ] Staging D5/D6 sweep after merge confirms multi-step demos
(beautiful-chat, gen-ui, shared-state, langgraph/mastra/crewai
agentic-chat) recover from `text-unstable` red.

## Scope

Strictly the two files above. No unrelated changes. No refactors.
2026-06-16 01:12:10 -07:00
Jordan Ritter 88e77f1403 fix(showcase): allow harness to boot locally without SHARED_SECRET (unblock local verify) (#5471)
## Problem

PR #5458 (`c81b361f1` — *fix(showcase/harness): register
/webhooks/deploy on CP path + fail-loud on missing SHARED_SECRET*) added
a fail-loud gate to `loadWebhookSecrets` in
`showcase/harness/src/orchestrator.ts` that refuses harness boot in any
deployable mode (`NODE_ENV !== "test"`) without `SHARED_SECRET` (or
`SHARED_SECRET_PREV`) set, because `POST /webhooks/deploy` is only
registered when `webhookSecrets.length > 0` (gate at
`showcase/harness/src/http/server.ts:119`).

The local `showcase/docker-compose.local.yml` harness services inherit
`NODE_ENV=production` from the image and do not set `SHARED_SECRET` (and
should not — that secret only matters for the Showcase: Verify Deploy
webhook flow on Railway). As a result, the harness crash-loops on every
local `bin/showcase test --d5/--d6` run:

```
FATAL-CONFIG: SHARED_SECRET (or SHARED_SECRET_PREV) is required — refusing to boot in any deployable mode … Set SHARED_SECRET (or SHARED_SECRET_PREV) in the env, or set NODE_ENV=test / HARNESS_ALLOW_NO_SECRET=1 for local dev. Current NODE_ENV=production.
```

…which blocks **all** local D5/D6 verify.

## Fix

Add `HARNESS_ALLOW_NO_SECRET=1` (the documented local-dev escape hatch —
explicitly named in the FATAL-CONFIG message itself, and the dedicated
env var checked by `loadWebhookSecrets` at `orchestrator.ts:190`) to
both harness services in `showcase/docker-compose.local.yml`:

- `harness-control-plane`
- `harness-pool-worker`

Inline comments explain the rationale and pin the relevant source
locations (the gate at `src/http/server.ts:119` and `loadWebhookSecrets`
in `src/orchestrator.ts`).

Diff: 17 lines added (comments + 2 env vars), no other files touched.

## Prod impact: NONE

Railway sets `SHARED_SECRET` explicitly via env on every harness
service, so `loadWebhookSecrets` sees a real secret, registers `POST
/webhooks/deploy`, and **never reads `HARNESS_ALLOW_NO_SECRET`**. The
escape hatch only takes effect when both secrets are absent. This change
only affects the local docker-compose stack.

## Local verification

Brought up the slot-1 isolated stack with the fix applied via
`SHOWCASE_ISO_SLOT=3 ./bin/showcase test langgraph-typescript --d6
--verbose --isolate` and observed clean boot:

```
showcase-iso1-harness                 Up 12 seconds (healthy)
showcase-iso1-langgraph-typescript    Up 12 seconds (healthy)
showcase-iso1-harness-pool-worker-1   Up 12 seconds (healthy)
showcase-iso1-dashboard               Up 13 seconds (healthy)
showcase-iso1-aimock                  Up 18 seconds (healthy)
showcase-iso1-pocketbase              Up 18 seconds (healthy)
```

Harness boot logs confirm:

- `fleet.role-selected` + `fleet.control-plane.started` +
`showcase-harness.fleet.control-plane.boot` (port 8080) — control-plane
up
- `webhook auth disabled — neither SHARED_SECRET nor SHARED_SECRET_PREV
is set` with `escapeHatch:true` at **warn** level — the expected log
produced by `loadWebhookSecrets` when the escape hatch fires
- `worker.registered` + `fleet.worker.boot` — worker self-registered
- `worker.claimed jobId=… probeKey=d6:langgraph-typescript` — real D6
job picked up and run end-to-end

NO `FATAL-CONFIG`, NO restart loop.

(The D6 cell may still be RED on this branch — that is a pre-existing
issue, unrelated to this infra fix. The point of this PR is "harness
boots and runs a job at all locally", which it now does.)

## References

- PR #5458 (`c81b361f1`): the fail-loud gate this PR re-enables an
escape for
- `showcase/harness/src/orchestrator.ts` `loadWebhookSecrets` (~line
180-220): the predicate that reads `HARNESS_ALLOW_NO_SECRET`
- `showcase/harness/src/http/server.ts:119`: the route-registration gate
2026-06-16 01:07:44 -07:00
Jordan Ritter d145bb1708 fix(harness): waitForTurnComplete reads LAST bubble in matched tier
PR #5462's defect-2 fix assumed 1 bubble per turn (bubbleIndex =
turnIndex - 1). Multi-step agents (LangGraph, Mastra, CrewAI) emit
2-3+ bubbles per turn — tool-call + tool-render + final-text — so
the strict-index read lands on an intermediate tool-call bubble
whose scoped-text selectors (`.cpk:prose`, `[data-message-content]`,
`p`) are EMPTY (its content lives in a sibling card div outside the
scoped cascade). The cascade-pollution guard returns null forever
and the text-stable conjunct times out with `reason=text-unstable`
across virtually every multi-step demo (beautiful-chat-*, gen-ui-*,
shared-state-*, agentic-chat on LangGraph/Mastra/CrewAI/etc.).

Live DOM evidence from staging (langgraph-typescript :
beautiful-chat : pie chart prompt):
  count = 3 bubbles
  bubble[0]: tool-call "query_data" — `.cpk:prose` empty
  bubble[1]: pie chart card — `.cpk:prose` empty
  bubble[2]: final text "Here is the pie chart..." — 90 chars

Fix: read the LAST bubble in the matched cascade tier (count - 1)
instead of a fixed strict index. Defect-2 protection (don't read a
leftover bubble from a prior turn) is preserved via an explicit
pre-submit `baselineCount` sentinel snapshot in the runner; the
gate now requires `count > baselineCount` rather than
`count > turnIndex - 1`, which correctly rejects stale bubbles
without assuming 1 bubble = 1 turn.

Changes:
  - assistant-message-count.ts: add `readCascadeStateLast(page)` —
    same atomic single-evaluate contract as `readCascadeState` but
    resolves the bubble index internally as `count - 1`.
  - conversation-runner.ts: snapshot `baselineCount` before
    sendTurnMessage; pass through to `waitForTurnComplete` via new
    `baselineCount` option (defaults to `turnIndex - 1` for unit-
    test fakes); `waitForTurnComplete` reads via
    `readCascadeStateLast` and gates on `count > baselineCount`.
    Cold-start retry re-snapshots `baselineCount` after page.reload.

Test plan:
  - All 72 conversation-runner + assistant-message-count unit tests
    pass (default baselineCount preserves count-progression scripts).
  - Full harness suite 2758/2758 green, tsc clean, oxfmt/oxlint
    clean on changed files.
2026-06-16 01:02:52 -07:00
Jordan Ritter 48ef506abf fix(showcase): allow harness to boot locally without SHARED_SECRET
PR #5458 (c81b361f1) added a fail-loud gate that refuses harness boot
in any deployable mode (NODE_ENV != "test") without SHARED_SECRET or
SHARED_SECRET_PREV: POST /webhooks/deploy is only registered when
webhookSecrets.length > 0 (src/http/server.ts:119 +
loadWebhookSecrets in src/orchestrator.ts). The local docker-compose
stack inherits NODE_ENV=production from the harness image and does
not (and should not) set SHARED_SECRET, so every local D5/D6 verify
run via bin/showcase test --d5/--d6 was crashing the harness in a
restart loop with FATAL-CONFIG.

Fix: add HARNESS_ALLOW_NO_SECRET=1 (the documented local-dev escape
hatch — explicitly referenced in the FATAL-CONFIG message itself) to
both harness services in showcase/docker-compose.local.yml. Inline
comments explain the rationale and pin the relevant source locations.

Prod impact: NONE. Railway sets SHARED_SECRET explicitly via env on
every harness service, so loadWebhookSecrets sees a real secret,
registers POST /webhooks/deploy, and never reads
HARNESS_ALLOW_NO_SECRET. This change only affects the local
docker-compose stack.

Verified locally: showcase-iso1-harness boots cleanly (Up healthy),
the expected warn-level webhook-auth-bypass log fires
(escapeHatch:true), worker registers, scheduler starts, and a real
d6:langgraph-typescript job claims successfully — confirming the
gate fires only in deployable contexts.
2026-06-16 01:00:22 -07:00
Jordan Ritter 7cc579a58d chore: drop dead .changeset/ debris and workflow path filters (#5470)
## Summary

The repo migrated off `@changesets/*` to conventional-commit-driven
releases. This chore PR removes the leftover debris:

- **11 stale `.changeset/*.md` files** (the entire `.changeset/`
directory)
- **4 dead `paths:` filter lines** in two e2e workflows

## Migration context

Releases are driven by `scripts/release/prepare-release.ts` /
`scripts/release/lib/changes.ts::getChangesSummary`, which reads `git
log <lastTag>..HEAD` commit subjects. It never touches `.changeset/`.

Verification — none of the changesets tooling is wired up anymore:

- No `.changeset/config.json`
- No `@changesets/*` in any `package.json` (root or workspace)
- No npm scripts reference `changeset`

The 11 `.changeset/*.md` files describe changes that have either already
shipped (via commit subjects in prior releases) or will ship in the next
release (via the current commit subjects in the `v1.60.1..main` window).
The files are inert with respect to releases — most are 2+ months old;
the newest (`slack-agent-native-apis.md`) was added recently by a
contributor who didn't know about the migration.

## What's removed

`.changeset/` (11 files):

- `angular-21-install-and-types.md`
- `angular-disable-license-watermark.md`
- `bump-license-verifier-ent-251.md`
- `debug-mode.md`
- `empty-mails-applaud.md`
- `ent-314-thread-connect-ux.md`
- `ent-658-generated-thread-tool-roundtrip.md`
- `five-avocados-visit.md`
- `fix-thread-switch-state-reset.md`
- `little-pears-tell.md`
- `slack-agent-native-apis.md`

Workflow `paths:` filters (4 lines across 2 files):

- `.github/workflows/test_e2e-dojo.yml` — `push.paths`,
`pull_request.paths`, and the `dorny/paths-filter` `ts:` filter
- `.github/workflows/test_e2e-legacy-v1.yml` — `push.paths`

Nothing to fire on after the directory is gone.

## Test plan

- [ ] CI green on this PR
2026-06-16 00:28:03 -07:00
Jordan Ritter 5afa55f067 chore: drop dead .changeset/ debris and workflow path filters
The repo migrated off @changesets/* to conventional-commit-driven releases.
scripts/release/lib/changes.ts::getChangesSummary reads `git log <lastTag>..HEAD`
commit subjects and never touches .changeset/. No .changeset/config.json,
no @changesets/* in any package.json, no npm scripts reference it.

The 11 .changeset/*.md files describe changes that have either already
shipped (via commit subjects in prior releases) or will ship in the next
release (via the current commit subjects in the v1.60.1..main window) —
the .changeset/ files are inert.

Also removes the dead workflow `paths:` filters in
test_e2e-dojo.yml and test_e2e-legacy-v1.yml that re-fired e2e on
.changeset/ changes — nothing to fire on after the directory is gone.
2026-06-16 00:25:05 -07:00
Jordan Ritter f8b14f1d6a fix(react-core): await runAgent in useInterrupt::resolve and propagate to demo-local hooks (#5461)
## Summary

Convert `useInterrupt::resolve` and the demo-local
`useHeadlessInterrupt::resolve` callbacks from fire-and-forget into
`async` + `return await copilotkit.runAgent(...)`, so callers receive a
Promise that settles when the resume run settles.

## Mechanism

Pre-fix: `resolve(...)` called `copilotkit.runAgent({...})` without
`await` and without `return` — the arrow returned `undefined`. The
showcase harness DOM-settle check on the assistant confirmation bubble
timed out because consumers had no handle to sequence against the resume
run's settle.

Post-fix:
- Framework `useInterrupt::resolve` is `async`, `return await`s
`runAgent`, wraps in try/catch + `setPendingEvent(null)` +
`console.error` + rethrow on rejection.
- `onRunFailed` now also `setPendingEvent(null)` — symmetric with
`onRunStartedEvent`, prevents stuck popups on run failure.
- `InterruptHandlerProps.resolve` / `InterruptRenderProps.resolve` typed
`() => Promise<RunAgentResult>` (was `() => void`).
- The same fix-pattern applied to 13 demo-local `useHeadlessInterrupt`
implementations across integrations: ag2, agno, claude-sdk-typescript,
crewai-crews, langgraph-{fastapi,python,typescript}, langroid,
llamaindex, mastra, pydantic-ai, spring-ai, strands.

## Verification

- **Unit:** new tests `resolve returns a Promise that settles when
runAgent settles (RESUME-PATH)` + `resolve rejects when runAgent
rejects, logs the failure, and clears pending (RESUME-PATH-REJECT)`.
Red-then-green on baseline. 22/22 react-core hook tests passing.
- **Backplane (end-to-end):** showcase control plane on langgraph-python
via `./bin/showcase test langgraph-python:gen-ui-interrupt --d5 --direct
--verbose` (GREEN, 70.1s, was RED-RESUME-PATH baseline) and
`:interrupt-headless --d5 --direct --verbose` (GREEN, 7.7s, was
RED-RESUME-PATH baseline). Built Docker artifact + aimock D6 fixture
replay — staging-equivalent path.

## CR

4 review rounds × 7 unbiased agents = 28 agent-runs. 2 Procedure 3
audits (3 promotions in round 1 → Fix C + D; 0 promotions in round 2).
Convergence achieved per agent-confirmed `bucket (a) empty` + Procedure
3 zero `PROMOTE_TO_A`.

## Out of scope (deferred follow-ups)

- **No manifest unquarantine.** A prior version of this branch flipped
`interrupt-headless` and `gen-ui-interrupt` from
`not_supported_features` to `features` across 12 manifests, but Phase 4
backplane validation across 8 PATCHED integrations × 2 demos surfaced
(a) `langgraph-fastapi:gen-ui-interrupt` still RED-RESUME-PATH because
the framework patch doesn't propagate from workspace to the
integration's Docker image without a published `@copilotkit/react-core`
release, and (b) a separate popup-mount RED-RUNTIME-OTHER cluster across
6 integrations that's pre-existing rot unrelated to this PR. Manifest
unquarantine + `@copilotkit/react-core` version bump are deferred to a
follow-up `copilotkit-release` PR after this fix lands.
- **No version bump.** Release follows in a separate PR.
- **Bucket (c) follow-up backlog** (audit verdicts: 11 STAY_IN_C):
  - JSON.parse guarding in showcase demos (pre-existing, all 13)
- Type-design quirks in `InterruptHandlerProps.result` (non-nullable in
type, nullable in runtime)
  - `renderInChat` dynamic toggle leak (pre-existing edge case)
- JSDoc clarity ("symmetric with onRunFailed" wording, `@typeParam
TRenderInChat`, `event.value` "any" → "unknown")
  - StrictMode test coverage gap
- 13× duplicate `useHeadlessInterrupt` → bucket (d): consolidate into a
shared module in a future PR

## Files changed

16 files, +486/-154:
- `packages/react-core/src/v2/hooks/use-interrupt.tsx`
- `packages/react-core/src/v2/hooks/__tests__/use-interrupt.test.tsx`
- `packages/react-core/src/v2/types/interrupt.ts`
- 13×
`showcase/integrations/<integ>/src/app/demos/interrupt-headless/page.tsx`

## Test plan

- [x] react-core unit tests (22/22 passing including new RESUME-PATH +
RESUME-PATH-REJECT regression tests)
- [x] Backplane D5 probe on langgraph-python (gen-ui-interrupt +
interrupt-headless GREEN end-to-end via the showcase control plane)
- [ ] CI green on this PR
2026-06-16 00:08:56 -07:00
tylerslaton 4564e6205d chore: release bot-slack v0.0.2 2026-06-16 05:12:14 +00:00
Tyler Slaton 431d5baae0 feat(bot-slack): agent-native Slack APIs — assistant pane + native streaming (default-on) (#5447)
## What

Activates Slack's agent-grade APIs as the **default** experience for
`@copilotkit/bot-slack`, with zero config and safe degradation.
Implements the Notion spec *"Slack Agent APIs in bot-slack — assistant
pane + native streaming by default"* (Rev 2).

### Assistant pane ("Agents & AI Apps")
- Opening the pane greets the user + shows tappable prompt chips; each
pane conversation is its own thread (replies stay in-thread, instead of
leaking to a merged flat DM).
- While the agent runs, native composer status
(`assistant.threads.setStatus`: "is thinking…", "is using `tool`…")
replaces placeholder/`🔧` messages.
- Pane threads are auto-titled from the first message.
- Customize via the new `assistant` option; `assistant: false` disables
pane handling. Apps without the Agents toggle behave exactly as before
(pane machinery dormant).

### Native streaming
- Replies stream via `chat.startStream` / `appendStream` / `stopStream`
(raw markdown — real tables / fenced code render natively) wherever a
thread exists.
- Flat DMs and workspaces without the streaming API fall back to the
legacy `chat.update` transport **automatically** — the first
`startStream` failure marks the workspace legacy and replays via legacy.
Opting in can never break a bot. `streaming: "legacy"` forces the old
transport.

### Portable engine surface (`@copilotkit/bot`)
- `bot.onThreadStarted` lifecycle handler + `IncomingThreadStart` sink
event.
- Capability-gated `thread.setSuggestedPrompts` / `thread.setTitle` (the
shipped `postFile` gating pattern).
- Two `SurfaceCapabilities` flags + optional `PlatformAdapter` methods.
All degrade gracefully on surfaces without support.

## Files
- **New:** `bot-slack/src/assistant.ts` (Bolt `Assistant` middleware ⇄
engine sink), `bot-slack/src/native-stream.ts` (`NativeMessageStream`,
same `append/finish` contract as `MessageStream`, legacy fallback).
- **Engine:** `platform-adapter.ts`, `thread.ts`, `create-bot.ts`,
`index.ts` (+ `bot-ui` `Thread` type).
- **Adapter:** `adapter.ts` (options/capabilities/stream
branch/methods), `slack-listener.ts` (one-line no-double-delivery
guard), `event-renderer.ts` (pane status mode + native text transport).
- **Example/docs:** `examples/slack` manifest (`assistant_view` + scope
+ events) & dev-ex, package READMEs/ARCHITECTURE, `slack.mdx`.
- **Hooks:** `chore(hooks)` makes the `lint-fix` script
double-quote-free (it was breaking `sh -c "…"` on Windows/lefthook
2.1.1).

## Tests / verification (CI-runnable)
- New unit tests: `native-stream` (cadence, 12k continuation, fence
re-open, first- **and** continuation-`startStream` fallback),
`assistant` (defaults-before-onThreadStarted ordering, thread-scoped
turn, auto-title), listener no-double-delivery, renderer pane-status
mode, engine capability gating.
- Verified locally: `@copilotkit/bot` 31 tests, `@copilotkit/bot-slack`
199 tests, `bot-ui` tests; `tsc` clean (bot + bot-slack); build clean;
oxlint/oxfmt clean; publint/attw pass.

## ⚠️ E2E not covered here
The §8 end-to-end flows need a **real Slack workspace** with the Agents
toggle (open pane → chips → streamed reply w/ live status; channel
mention → native markdown; kill `startStream` → legacy fallback;
non-agent app → shipped behavior). `examples/slack` is the vehicle.
Three items remain to confirm against a live workspace (spec §7 spikes):
the `isAssistantThread` runtime detection, the exact `assistant_view`
manifest acceptance, and `stopStream` finalize tolerance.

Spec:
https://app.notion.com/p/copilotkit/Slack-Agent-APIs-in-bot-slack-assistant-pane-native-streaming-by-default-spec-37b3aa38185281f5948dc6a665064d04
2026-06-15 22:10:17 -07:00
Jordan Ritter c88d687456 fix(showcase/harness): bubble-race elimination — 4 defects + atomic readCascadeState + cold-start retry + SSE counter (#5462)
## Summary

Fixes 4 bubble-race defects in the CopilotKit showcase e2e harness's
conversation runner that caused flaky turn-completion detection on
slow-streaming and cold-start integrations.

**Defects fixed (RED → GREEN):**
- **Defect 1** (fast-replay): turn settled on count, not text/SSE →
false-positive settle
- **Defect 2** (multi-turn flicker): un-turn-scoped bubble selection via
`list[last]` → cross-turn leak
- **Defect 3** (cascade blindness): diagnostic cascade picked wrong tier
→ wrong bubble text
- **Defect 4** (boot-time baseline staleness): pre-paint/stale bubble at
boot poisoned turn-1 settle

**Core changes:**
- `waitForTurnComplete` 3-conjunct primitive (SSE counter + DOM-at-index
+ text-quiet for settleMs)
- Atomic `readCascadeState` helper — single `page.evaluate` returns
`{count, text}` from one cascade tier (prevents within-poll cross-tier
inconsistency)
- Probe contract: `assertions(page, ctx: {bubbleIndex, text})` —
turn-scoped bubble retrieval (replaces `list[last]`)
- SSE interceptor — `__hk_runsFinished` counter via CDP + page-side
fetch wrap; idempotent attach/detach + handle-cache invalidation on stop
+ `framenavigated` reset (cold-start retry compatible)
- Cold-start retry — single bounded retry on first-attempt banner;
reload + re-resolve chat input + settleMs floor honoring shared turn
deadline
- Banner fast-fail — baseline banner snapshot + differs-from-baseline
2-poll debounce; in-poll `BannerVisibleError` translates to
`AssistantErroredError`
- Pre-paint placeholder env adapter for defect-4 repro infrastructure

**Defect repro tests:** real-browser tests in
`test/integration/bubble-race-repro-defect-{1,2,3,4}.test.ts` plus
mechanism-GREEN tests and `wait-for-turn-complete.test.ts` (3-conjunct
classification matrix), `probe-contract.test.ts` (ctx bridge),
`sse-counter.test.ts` (counter increment + framenavigated reset +
handle-cache invalidation).

## Test plan

- [x] Full harness vitest: 2742/2742 passing
- [x] Defect-1, defect-2, defect-4 RED tests verified GREEN after fix
- [x] Mechanism-GREEN tests pin internal contracts (cascade pollution
guard, SSE counter atomicity, init-script navigation re-entry, etc.)
- [x] cr-loop converged across 10 rounds (25→6→6→6→4→2→0 bucket-(a)
findings) + Procedure 3 audit clean
- [x] `oxfmt --check` clean across all 31 changed TS files
- [x] `tsc -p tsconfig.build.json` clean
- [ ] CI green on this PR
- [ ] Manual smoke: `bin/showcase test --d5
langgraph-python:agentic-chat` post-merge

## Follow-up debt (deferred from cr-loop; not subject-scope)

Detected by cr-loop reviewers but out of this PR's scope:
- `d6-all-pills.ts` deploy-churn NSF aggregation — features in
`notSupportedFeatures` lose `incapable` classification during
deploy-churn skip
- `d6-all-pills.ts` `joinAimockJournal` slug fallback under D6
concurrency — can pick wrong-feature entry when aimock doesn't echo
`x-diag-run-id`
- `d6-all-pills.ts` `BrowserDisconnectedError` sentinel is constructed
but never matched in `runFeature` catch — gets bucketed as
`driver-error`
- `d5-tool-rendering-default-catchall.ts` Path B narration
false-positive (substring matches in assistant prose)
- HITL registry side-effect tests use manual fallback registration that
masks real registration failures
- `installPrePaintFromEnv` and `attachSseInterceptor` are re-registered
on every `page.goto` — accumulate init scripts; harmless today via
in-script idempotency guards

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-15 20:14:51 -07:00
Jordan Ritter 1ef6a0960b test(harness): update d5-* + d6-all-pills unit tests to new probe + cascade contract
- d5-* unit tests: update page fakes / assertion-call shapes to pass
  the ctx bridge ({bubbleIndex, text}) instead of querying the live
  bubble list
- d6-all-pills.test.ts: update page-fake dispatch + cascade-state
  expectations to match the atomic readCascadeState return shape
  ({count, text} or null) — picks up the cross-tier consistency
  guarantee in the unit layer
2026-06-15 18:37:26 -07:00
Jordan Ritter cd2af0f9fd test(harness): conversation-runner unit coverage — banner debounce, cold-start retry, skipFill, error-banner shapes
Comprehensive unit coverage for:
- preFill + fillAndVerifySend retry behavior
- baseline banner snapshot + differs-from-baseline debounce matrix
  (stable banner = no fail; changed banner = fast-fail after 2 polls)
- cold-start retry cases (first-attempt banner → reload + retry;
  retry-then-banner → translatedErr; retry exhausts deadline)
- skipFill / skipSend short-circuit paths
- ErrorBannerReadResult union shape variants (present/absent/unknown)
2026-06-15 18:37:26 -07:00
Jordan Ritter 3d49d4c38f test(harness): bubble-race integration tests + defect reproductions
- bubble-race-repro.ts: shared driver harness for repro tests
- bubble-race-repro-defect-{1,2,3,4}.test.ts: real-browser RED tests
  that fail without the fix and pass with it
    - Defect 1: fast-replay false-positive settle on count
    - Defect 2: multi-turn flicker via list[last]
    - Defect 3: cascade-tier blindness
    - Defect 4: boot-time baseline staleness / pre-paint placeholder
- bubble-race-mechanisms.test.ts: GREEN tests pinning internal
  contracts (cascade-pollution guard, atomic readCascadeState,
  init-script idempotency)
- wait-for-turn-complete.test.ts: 3-conjunct classification matrix
- probe-contract.test.ts: ctx-bridge shape
- sse-counter.test.ts: counter increment + framenavigated reset +
  handle-cache invalidation
2026-06-15 18:37:26 -07:00
Jordan Ritter cf7efbda95 fix(harness): d6 driver — installPrePaintFromEnv + attachSseInterceptor wiring + helper-based captureDiagnostics
- d6-all-pills.ts: wire installPrePaintFromEnv and
  attachSseInterceptor PRE-goto in both defaultLauncher and pooled
  launcher paths, so first-paint state is deterministic and the SSE
  counter is armed before any page navigation
- captureDiagnostics rewired through the Node-side helper
  (readCascadeState) — single page.evaluate per call, no
  within-snapshot tier drift
- Consume BUBBLE_RACE_MESSAGES override for defect-4 repro flows
2026-06-15 18:37:25 -07:00
Jordan Ritter 38e155e141 refactor(harness): probe-contract ctx bridge — assertions(page, {bubbleIndex, text}) for d5 scripts
- d5-gen-ui-custom, d5-gen-ui-open-advanced, d5-mcp-apps,
  d5-subagents, d5-tool-rendering-default-catchall: wrap assertion
  bodies in ctx bridge so they receive {bubbleIndex, text} for the
  turn under test, replacing turn-leaky list[last] retrieval
- _gen-ui-shared: add readAssistantTextAt(page, bubbleIndex)
  adapter; remove readLastAssistantText (no longer turn-scoped)

This is the read-side counterpart to the atomic readCascadeState
on the helper side: probes ask for the bubble at the runner's
chosen index instead of guessing.
2026-06-15 18:37:25 -07:00
Jordan Ritter 3f3986b00b refactor(harness): waitForTurnComplete 3-conjunct primitive + cold-start retry + banner debounce
- waitForTurnComplete: 3-conjunct settle (SSE counter + DOM-at-index
  + text-quiet for settleMs) replaces count-only false-positive path
- Cold-start retry: single bounded retry on first-attempt banner;
  reload + re-resolve chat input + settleMs/POLL_INTERVAL_MS floor
  honoring shared turn deadline; translatedErr throw on exhaustion
- Banner fast-fail: baseline banner text snapshot +
  differs-from-baseline 2-poll debounce; in-poll BannerVisibleError
  translates to AssistantErroredError
- New error classes: BannerVisibleError, AssistantErroredError
- readErrorBanner returns 3-state union via
  ErrorBannerReadResult discriminated union
- chatInputSelector cascade for input resolution across integrations
2026-06-15 18:36:59 -07:00
Jordan Ritter 473a9539a0 feat(harness): add bubble-race helpers — assistant-message-count, sse-interceptor, init-scripts
- assistant-message-count: 4-tier cascade + atomic readCascadeState
  returning {count, text} from a single page.evaluate (prevents
  within-poll cross-tier inconsistency); null fallback for
  cascade-pollution
- sse-interceptor: __hk_runsFinished CDP counter + page-side fetch
  wrapper; idempotent attach/detach with handle-cache invalidation
  on stop and framenavigated reset (cold-start retry compatible)
- init-scripts: pre-paint placeholder env adapter
  (installPrePaintFromEnv), strip-selector helper, and
  BUBBLE_RACE_MESSAGES override consumer for defect-4 repro
  infrastructure
2026-06-15 18:36:59 -07:00
Jordan Ritter 6fbd66fa83 fix(showcase): mirror useInterrupt RESUME-PATH contract in 13 demo-local hooks
Each integration's interrupt-headless demo defines a local useHeadlessInterrupt
hook around the framework useInterrupt. Slot-2 originally identified 8
quarantined integrations (claude-sdk-typescript, langgraph-{fastapi,python,
typescript}, langroid, pydantic-ai, spring-ai, strands); review-round
follow-ups extended the sweep to llamaindex, mastra, ag2, agno, and
crewai-crews (5 more integrations sharing the same byte-identical hook).

The demo-local resolve() previously fire-and-forgot copilotkit.runAgent(...)
via `void runAgent(...).catch(() => {})`. Mirroring the framework fix:

- Make resolve async, return await copilotkit.runAgent(...).
- Use a pendingRef so resolve has stable identity (drop pending from
  useMemo deps).
- Type signature: resolve: (response: unknown) => Promise<unknown>.
- Wrap in try/catch + setPending(null) + console.error + rethrow,
  symmetric with the framework hook.
- onRunFailed also setPending(null).

13 integrations patched byte-identically.
2026-06-15 17:11:40 -07:00
Jordan Ritter 7b80590f67 fix(react-core): await runAgent in useInterrupt::resolve
resolve() previously called copilotkit.runAgent(...) without await and
without return, so callers had no handle to sequence against the resume
run's settle. The harness DOM-settle check timed out for any consumer
awaiting the assistant confirmation bubble.

Changes:
- Make resolve async, return await copilotkit.runAgent(...) so callers
  receive a Promise that settles when the resume run settles.
- Update InterruptHandlerProps / InterruptRenderProps resolve return
  type from () => void to () => Promise<RunAgentResult>.
- Wrap runAgent in try/catch + setPendingEvent(null) + rethrow, so
  rejection clears the popup AND propagates to awaiting callers
  (mirrors onRunFailed handler symmetry; closes the case where
  runAgent rejects before any run-failed event fires, e.g. network
  error pre-RUN_STARTED).
- onRunFailed now also setPendingEvent(null) symmetric with
  onRunStartedEvent.
- Regression tests: RESUME-PATH asserts resolve() returns a Promise
  that settles 1:1 with runAgent; RESUME-PATH-REJECT asserts rejection
  propagates, popup clears, console.error logs.
2026-06-15 17:11:39 -07:00
Jordan Ritter c81b361f15 fix(showcase/harness): register /webhooks/deploy on CP path + fail-loud on missing SHARED_SECRET (#5458)
## Problem

The \`Showcase: Verify Deploy\` workflow runs \`notify-harness\` after
every merge to \`main\` and POSTs an HMAC-signed payload to
\`\${SHOWCASE_HARNESS_URL}/webhooks/deploy\`. That POST has been
returning **404** on every main deploy since at least 2026-06-12 (5+
consecutive failures), so the harness dashboard has not reflected any
recent deploys.

**Root cause:**

1. \`POST /webhooks/deploy\` is only registered when
\`webhookSecrets.length > 0\` — gate at
\`showcase/harness/src/http/server.ts:119\`.
2. The CP \`buildServer\` call in \`runControlPlane\`
(\`showcase/harness/src/orchestrator.ts\` ~line 2961 pre-fix)
**omitted** both \`webhookSecrets\` and \`metrics\`. The public Railway
host running the CP role never mounted the route — every notify-harness
POST returned 404.
3. The FATAL-CONFIG guard on missing \`SHARED_SECRET\` (pre-fix
\`orchestrator.ts:782-790\`) only fired when \`NODE_ENV ===
"production"\`. Any deploy whose NODE_ENV was unset, set to something
other than the literal \`"production"\`, or set after the check,
silently booted with \`webhookSecrets=[]\` and no fatal error.

## Fix

**Part 1: Register the route on the CP path.**
- Lifted the env→secrets loader into a new exported helper
\`loadWebhookSecrets()\` (orchestrator.ts ~175-209) so both boot paths
consume the same predicate.
- Worker \`boot()\` (~line 836) calls it; CP \`runControlPlane\` (~line
2421) now also calls it.
- The CP \`buildServer\` call (~line 3047) now passes \`webhookSecrets\`
and \`metrics\` alongside the existing \`bus\`/\`probes\`/\`fleetRuns\`
wiring — same shape as the worker call.

**Part 2: Tighten the FATAL-CONFIG predicate.**
- New predicate (inside \`loadWebhookSecrets\`): throw unless EITHER a
secret is set OR \`NODE_ENV === "test"\` OR \`HARNESS_ALLOW_NO_SECRET
=== "1"\` (narrow escape hatch for local dev).
- Thrown error message names the env-vars, the gate location
(\`src/http/server.ts:119\`), the silent-404 symptom, and the escape
hatches — so an operator reading the deploy log can recover without
spelunking.

## Test coverage

Four new red-green tests in \`orchestrator.test.ts\` under \`B2:
/webhooks/deploy registered on CP + fail-loud on missing
SHARED_SECRET\`:

- **Test A** — CP boot registers \`POST /webhooks/deploy\` when
\`SHARED_SECRET\` is set. Probe with no HMAC headers: pre-fix returned
404, post-fix returns 401 (route mounted, HMAC reject path fires).
RED→GREEN.
- **Test B** — Worker boot still registers the route (regression guard).
Pre-fix and post-fix both 401.
- **Test C** — FATAL-CONFIG fires when \`SHARED_SECRET\` unset AND
\`NODE_ENV=development\`. Pre-fix resolved silently, post-fix rejects
with \`/FATAL-CONFIG.*SHARED_SECRET/\`. RED→GREEN.
- **Test D** — Boot succeeds when \`SHARED_SECRET\` unset AND
\`NODE_ENV=test\` (escape hatch). Pre-fix and post-fix both green.

Existing R5-G4 D5 test (orchestrator.test.ts:1592) updated to match the
new error message via \`/FATAL-CONFIG.*SHARED_SECRET/\` — same throw,
broader match.

## Verification

- Full harness suite: **2719/2719 green** across 128 files
- \`tsc --noEmit\`: clean
- RED capture pre-fix:
  - Test A: \`AssertionError: expected 404 to be 401\`
- Test C: \`AssertionError: promise resolved "{ port: ..., bus: { ... },
... }" instead of rejecting\`

## Deploy note

Requires a CP-role redeploy on the public Railway host for the new route
registration to take effect — once redeployed, the next
\`notify-harness\` POST will land on the harness dashboard.

## Test plan

- [x] B2 Tests A-D pass post-fix (red→green captured)
- [x] Full harness vitest suite green (2719/2719)
- [x] \`tsc --noEmit\` clean
- [ ] Post-deploy: confirm next \`Showcase: Verify Deploy\` run shows a
2xx notify-harness POST (not 404)
- [ ] Post-deploy: confirm harness dashboard reflects the next deploy
2026-06-15 16:44:07 -07:00
Jordan Ritter 3ce3f5394d fix(showcase/harness): hoist OPS_TRIGGER_TOKEN fail-loud + emit on subscribeDeployResults sync throw (R3 cleanups)
R3-F1 (OPS_TRIGGER_TOKEN hoist):
Extract loadOpsTriggerToken() and hoist its invocation to the top of both
boot() and runControlPlane(), alongside loadWebhookSecrets() and
loadPocketbaseUrl(). Pre-fix the empty/whitespace check fired AFTER pb /
bus / scheduler / writer / S3 uploader allocations, so a typo'd
`OPS_TRIGGER_TOKEN=` allocated expensive resources before throwing. Now
all three fail-loud config predicates fire at the top, before any
allocation needing teardown. Behaviour is unchanged for valid tokens
(trimmed via R3-A.5 contract) and for the unset case (router omitted with
info log).

R3-F2 (sync-throw also emits deploy.writer.failed):
subscribeDeployResults() pre-fix only emitted `deploy.writer.failed` on
writer.write() promise rejection — a synchronous throw inside
deployEventToProbeResult() (malformed event, type drift) bypassed the
catch and we lost both the log AND the bus emit. Wrap the sync mapping
in try/catch and mirror the rejection path so alert rules / metrics
subscribers observe sync throws the same way they observe async write
failures.

Tests:
- 5 unit tests for loadOpsTriggerToken: undefined / empty-string fail-loud
  / whitespace-only fail-loud / R3-A.5 trim contract / verbatim value.
- 1 test for the sync-throw path: vi.doMock deployEventToProbeResult to
  throw, assert bus emits deploy.writer.failed with the err message AND
  writer.write was NOT called. Red-green verified locally (test fails
  without the try/catch).

Suite: 2731 passing (+6 new). Typecheck clean.
2026-06-15 16:37:05 -07:00
Jordan Ritter d5bdf2c9fd fix(showcase/harness): track CP deploy.result unsubscribe + emit on write failure (R2 cleanups) 2026-06-15 16:23:29 -07:00
Jordan Ritter b0ba8e869d fix(showcase/aimock): align catchall fixture shape for 6 remaining D5 reds (#5460)
## Summary

Follow-up to #5459. That PR added the right `userMessage` matchers
("forecast for Tokyo" + "current price of AAPL") and flipped most
integrations green, but 6 cells stayed red because the FIXTURE SHAPE on
the matched fixtures was also broken.

**Failing cells (post-#5459 deploy):**

- `d5:agno/tool-rendering-custom-catchall`
- `d5:llamaindex/tool-rendering-custom-catchall`
- `d5:langroid/tool-rendering-custom-catchall`
- `d5:claude-sdk-python/tool-rendering-custom-catchall`
- `d5:strands/tool-rendering-custom-catchall`
- `d5:strands/tool-rendering-default-catchall`

## Root cause

The failing tool-emit fixtures all used `hasToolResult: false` (or
`toolName: "<tool>"`) as their gate. The custom-catchall probe drives
TWO sequential prompts (Tokyo, then AAPL). After Tokyo's tool result
lands in the thread, aimock's `hasToolResult` is permanently true
(`messages.some(m => m.role === 'tool')`), so `hasToolResult:false` can
never match the AAPL turn → no fixture → 30s timeout.

PocketBase confirms this exact mode — e.g. agno:

```
errorDesc: "timeout: assistant did not respond within 30000ms"
failure_turn: 2
turns_completed: 1 / 2
```

claude-sdk-python's `toolName` gate fails for an analogous reason on the
second turn when the backend doesn't forward the tool definition
uniformly.

The same trap is documented in
`showcase/aimock/d6/langgraph-python/tool-rendering.json`:

> Gated on toolName:get_stock_price rather than hasToolResult:false. The
D5 tool-rendering-custom-catchall probe runs 'weather in Tokyo' first,
which leaves a get_weather tool result in the thread; the aimock router
implements hasToolResult as messages.some(m=>m.role==='tool'), so
hasToolResult is permanently true on the AAPL turn and a
hasToolResult:false gate could never match (→ no_fixture_match → 503 →
30s timeout).

## Fix

Align all 6 files to the canonical pattern used by working integrations
(mastra, spring-ai, built-in-agent):

1. `toolCallId`-keyed narration fixture FIRST
2. tool-emit fixture SECOND, gated ONLY on `userMessage` + `context` (no
`hasToolResult` / `toolName`)

aimock's `toolCallId` matcher checks `messages[last].role === 'tool' &&
tool_call_id === ...`, which correctly fires only on the
post-tool-result iteration regardless of older tool results in the
thread. The tool-emit fixture below it always matches on turn-init (last
message is user).

Strands' two files additionally needed REORDERING — they had the
tool-emit fixture before the toolCallId fixture, defeating
first-match-wins.

## Scope discipline

- Pure fixture-content alignment — no harness, probe, frontend, or
backend changes.
- Did NOT touch `showcase/harness/**` (peer session territory per
`/tmp/coordinate/showcase-reds-coord/CHANNEL.md`).
- Did NOT touch the two `d5-tool-rendering-*-catchall.ts` probes.

## Caveat (strands)

During investigation, `showcase-strands-staging.up.railway.app` was
returning 502 / page-load failures (infra, not fixture). Strands'
fixture defects are still real and worth correcting, but the cells may
stay red until the backend recovers.

## Test plan

- [ ] CI green
- [ ] Wait for `showcase_build.yml -f service=aimock` rebuild + Railway
redeploy
- [ ] Re-check PocketBase `state` for the 6 cells flips to `green`
(excluding strands if backend stays down)
- [ ] Confirm no regression on the other catchall cells already green
post-#5459
2026-06-15 16:17:29 -07:00
Jordan Ritter df97ae8d2a fix(showcase/aimock): align catchall fixture shape for 6 remaining D5 reds
PR #5459 added userMessage matchers ("forecast for Tokyo" + "current price
of AAPL") on the catchall fixtures, which flipped most integrations from
red to green. Six cells stayed red on staging because the FIXTURE SHAPE
itself was broken on the matched fixtures, not just the userMessage key.

Failing cells (all custom-catchall except strands default):
  d5:agno/tool-rendering-custom-catchall
  d5:llamaindex/tool-rendering-custom-catchall
  d5:langroid/tool-rendering-custom-catchall
  d5:claude-sdk-python/tool-rendering-custom-catchall
  d5:strands/tool-rendering-custom-catchall
  d5:strands/tool-rendering-default-catchall

PocketBase confirms the failure mode: turn 1 (Tokyo) completes; turn 2
(AAPL) times out at 30s or renders the wrong content. E.g. agno:

  errorDesc: "timeout: assistant did not respond within 30000ms"
  failure_turn: 2
  turns_completed: 1 / 2

Root cause: the failing tool-emit fixtures used `hasToolResult: false`
(or `toolName: "<tool>"`) as their gate. aimock's hasToolResult check is
`messages.some(m => m.role === 'tool')` over the WHOLE thread — so once
turn 1's Tokyo tool result lands in the conversation, hasToolResult is
permanently true and `hasToolResult:false` can never match turn 2 → no
fixture → 30s timeout. `toolName:get_stock_price` likewise fails when an
integration backend doesn't forward the tool definition on turn 2.

The same trap is documented in
showcase/aimock/d6/langgraph-python/tool-rendering.json:

  "_comment": "Gated on toolName:get_stock_price rather than
   hasToolResult:false. The D5 tool-rendering-custom-catchall probe runs
   'weather in Tokyo' first, which leaves a get_weather tool result in
   the thread; the aimock router implements hasToolResult as
   messages.some(m=>m.role==='tool'), so hasToolResult is permanently
   true on the AAPL turn and a hasToolResult:false gate could never
   match (→ no_fixture_match → 503 → 30s timeout)."

Fix: align all 6 files to the canonical pattern used by mastra/spring-
ai/built-in-agent on this probe:

  1. toolCallId-keyed narration fixture FIRST
  2. tool-emit fixture SECOND with ONLY `userMessage` + `context`
     (no hasToolResult / toolName gate)

The toolCallId narration uses aimock's
`messages[last].role === 'tool' && tool_call_id === ...` check, so it
correctly wins on iteration 2 (post-tool-result) without being affected
by older turns' tool results. The tool-emit fixture matches turn 1 (last
message is user) and re-emits only when the narration above hasn't
matched.

Strands' two cells additionally needed REORDERING — they had the tool-
emit fixture before the toolCallId fixture, defeating first-match-wins.

Strands' staging backend was also returning 502 during testing; once it
recovers, the corrected fixtures should let the probe pass. The fixture
changes are necessary but may not be sufficient for strands if backend
remains down.

No probe-side, harness, or backend changes — pure fixture-content
alignment. Six fixture files modified; line totals: -87 / +68.
2026-06-15 16:09:04 -07:00
Martha Kelly Schumann fadb29550a Merge branch 'main' into docs/FAC-65-shared-v2-imports 2026-06-15 15:40:32 -07:00
Jordan Ritter 16a59a45c7 chore(showcase/harness): warn on webhook escape hatch + vi.stubEnv in B2 tests
CB-1 (Slot 2 #22): convert B2 deploy-webhook tests to vi.stubEnv with
`vi.unstubAllEnvs()` in afterEach instead of manual process.env mutation
+ try/finally restoration. Test runner now guarantees restoration and the
test bodies are shorter / less error-prone (CR Slot 2 #22).

CB-2 (Slot 2 #28): when `loadWebhookSecrets`' escape hatch fires with a
real-looking NODE_ENV (anything except "test"), log at `warn` instead of
`info` so a production typo (NODE_ENV=staging + HARNESS_ALLOW_NO_SECRET=1)
is visible in dashboards / log alerting. Pure local-dev (NODE_ENV=test)
stays at info level so a normal unit-test boot doesn't spam warnings.

CB-3 (Slot 4 #17): clarify `loadWebhookSecrets`' docstring + bypass-log
message — "set" is ambiguous (empty string would qualify pre-fix); use
"non-empty" to match the actual predicate.
2026-06-15 15:25:42 -07:00
Jordan Ritter f42180f7bf fix(showcase/harness): hoist fail-loud config checks + symmetric POCKETBASE_URL predicate
R1-F2 (bucket b, defensive ordering): hoist fail-loud config validation
to the TOP of boot() and runControlPlane() — BEFORE any pb client, queue,
bus, scheduler, writer, S3 uploader, fleet-health, or aggregator
allocations. Pre-fix `loadWebhookSecrets` lived AFTER the entire
scheduler+writer (and CP queue+aggregator) setup, so a misconfigured boot
allocated expensive resources and mounted file watchers before throwing.

R1-F3 (bucket b, predicate symmetry): broaden the POCKETBASE_URL fail-loud
predicate to match SHARED_SECRET semantics. Pre-fix the POCKETBASE_URL
guard only fired on `NODE_ENV === "production"`, while `loadWebhookSecrets`
fired unless NODE_ENV=test or an explicit escape hatch — so staging /
unset / "development" deploys silently bound to http://localhost:8090.

Extracted `loadPocketbaseUrl(logger)` with the same test-or-escape-hatch
predicate as `loadWebhookSecrets`. A new `HARNESS_ALLOW_NO_PB_URL=1` env
flag mirrors `HARNESS_ALLOW_NO_SECRET=1` for local dev. Worker
`boot()` and the CP's `resolveFleetPbConfig` both call the helper.

Tests:
  - HF13-A2 production-only assertion broadened (SHARED_SECRET set so the
    test specifically exercises the POCKETBASE_URL guard).
  - New R1-F3 tests cover NODE_ENV=development without POCKETBASE_URL
    (throws) and the HARNESS_ALLOW_NO_PB_URL=1 escape hatch (succeeds).
2026-06-15 15:25:11 -07:00
Jordan Ritter fc828bb08c fix(showcase/harness): subscribe deploy.result in CP boot to deliver dashboard event
R1-F1 (bucket a, load-bearing): the CP boot path now subscribes to
`deploy.result` events through its status writer. Pre-fix, B2 mounted POST
/webhooks/deploy on the CP role but only the worker boot path had a
`bus.on("deploy.result", ...)` listener — signed POSTs to the CP host
returned 202, the bus event fired with no subscriber, and the deploy-overall
dashboard row never landed.

Extracted the handler into a shared `subscribeDeployResults(bus, writer)`
helper exported from orchestrator.ts so both boot paths share the IDENTICAL
logic. Worker boot keeps its inline pattern (busUnsubs.push) so teardown is
unchanged; CP keeps the subscription alive for the lifetime of the bus.

Red-green tests in src/orchestrator.test.ts under "R1-F1" assert that
emitting deploy.result on the returned handle's bus drives one write keyed
"deploy:overall" through a stubbed status writer — once for runControlPlane
(the new path) and once for boot() (regression guard for the helper
extraction).
2026-06-15 15:23:24 -07:00
Jordan Ritter bb23e5dbbf fix(showcase/aimock): add forecast-Tokyo + AAPL catchall fixtures to flip remaining D5 reds (#5459)
## Summary
PR #5453 renamed stale catchall userMessage `"check Tokyo weather
forecast"` → `"forecast for Tokyo"` to match the D5 probe input. But
some integrations' catchall fixtures used **pill-aligned userMessages**
(e.g. LG-TS, LG-FastAPI used `"Chain a few tools in this single turn"`,
`"What's the weather in San Francisco?"`) — they had no `"forecast for
Tokyo"` matcher to rename. Result: D5 catchall probe still misses → live
LLM fallback → cells stay red on staging.

## Fix
Add the canonical emit+narrate fixture pairs for `"forecast for Tokyo"`
(default + custom catchall) and `"What's the current price of AAPL?"`
(custom catchall only) — adapting mastra's working template — to each
integration that was missing them.

## Scope
4 integrations × 1-2 files each (6 files total):
- LG-TS, LG-FastAPI: both default + custom catchall
- google-adk, ms-agent-dotnet: custom catchall only (their default is
already green)

Existing pill-aligned fixtures preserved (pure prepend at start of
array). First-match-wins means new fixtures match the D5 probe inputs
("forecast for Tokyo", "What's the current price of AAPL?") without
colliding with existing pill prompts. `toolCallId` values are unique per
integration to avoid cross-integration shadowing.

Note: the originally-flagged langroid, llamaindex, agno,
claude-sdk-python, and strands files already have `forecast for Tokyo` +
AAPL matchers in their catchall fixtures (likely from a prior pass) —
they did not need changes and are not in this PR. If those staging cells
are still red, the root cause is elsewhere (control-plane staleness,
deploy lag, or different probe variant).

## Verification
Post-merge, fleet-cp's e2e-deep probe will re-run within ≤6h staleness.
Cells should flip d6:<slug>/tool-rendering-{default,custom}-catchall =
green.
2026-06-15 15:10:22 -07:00
Jordan Ritter 07e4998798 fix(showcase/aimock): add forecast-Tokyo + AAPL catchall fixtures to flip remaining D5 reds
The catchall userMessage rename (#5453) only renamed STALE strings to 'forecast for Tokyo'. Integrations whose catchall fixtures used PILL-ALIGNED userMessages ('What's the weather in San Francisco?', 'Find flights from SFO to JFK.', 'Chain a few tools in this single turn', etc.) had no 'forecast for Tokyo' matcher to begin with — so the D5 catchall probe still fell through to live LLM and the cells stayed red on staging.

This PR ADDS the canonical 'forecast for Tokyo' (default+custom catchall) and 'What's the current price of AAPL?' (custom catchall only) emit+narrate fixture pairs to each integration that was missing them. Existing pill-aligned fixtures are preserved (pure prepend at the start of the fixtures array; first-match-wins means new fixtures match the D5 probe inputs without colliding with existing pill prompts).

Integrations touched:
- langgraph-typescript: both default + custom
- langgraph-fastapi: both default + custom
- google-adk: custom only (default already green)
- ms-agent-dotnet: custom only (default already green)

After merge, fleet-cp's e2e-deep probe will re-run within ≤6h staleness window and flip cells GREEN.
2026-06-15 15:00:53 -07:00
github-actions[bot] bfd1f12b00 style: auto-fix formatting 2026-06-15 21:56:46 +00:00
Jordan Ritter 2939696415 fix(showcase/harness): register /webhooks/deploy on CP path + fail-loud on missing SHARED_SECRET
The 'Showcase: Verify Deploy' workflow's notify-harness step has been
POSTing to /webhooks/deploy and getting 404 on every main deploy since
at least 2026-06-12 (5+ consecutive failures). Root cause:

  1. POST /webhooks/deploy is only registered when
     webhookSecrets.length > 0 (gate at src/http/server.ts:119).
  2. The CP buildServer call in runControlPlane omitted BOTH
     webhookSecrets and metrics — so the public Railway host running
     the CP role never mounted the route. The worker boot path had
     them, but it isn't the host receiving notify-harness POSTs.
  3. The FATAL-CONFIG guard on missing SHARED_SECRET was gated on
     NODE_ENV === 'production' — any deploy with NODE_ENV unset or set
     to something other than the literal 'production' silently shipped
     with webhookSecrets=[] and no fatal error.

Part 1 — wire webhookSecrets + metrics into the CP buildServer call.
Lifted the env→secrets loader into a new exported helper
'loadWebhookSecrets()' (~orchestrator.ts:175-209) so BOTH boot paths
consume the same predicate. The worker 'boot()' call (~line 836) and
the CP 'runControlPlane' call (~line 2421) both invoke it, and the CP
buildServer call (~line 3047) now passes webhookSecrets+metrics
alongside the existing bus/probes/fleetRuns wiring.

Part 2 — tighten the FATAL-CONFIG predicate.
The new predicate throws unless EITHER a secret is set OR NODE_ENV ===
'test' OR HARNESS_ALLOW_NO_SECRET === '1' (narrow escape hatch for
local dev). The thrown error message names the env-vars, the gate
location ('src/http/server.ts:119'), the silent-404 symptom, and the
escape hatches, so an operator reading the deploy log can recover
without spelunking.

Updated the existing R5-G4 D5 regex (orchestrator.test.ts:1592) since
the error message changed; covered by the new B2 Tests A/C/D below.

Tests: 4 new red-green tests in 'B2: /webhooks/deploy registered on CP
+ fail-loud on missing SHARED_SECRET':
  - Test A (CP /webhooks/deploy registration): RED 404 → GREEN 401.
  - Test B (worker /webhooks/deploy regression): GREEN 401.
  - Test C (FATAL-CONFIG when NODE_ENV=development): RED resolves →
    GREEN rejects with /FATAL-CONFIG.*SHARED_SECRET/.
  - Test D (NODE_ENV=test escape hatch): GREEN, no throw.

Full harness suite: 2719/2719 green. typecheck: clean.

Requires a CP-role redeploy on the public Railway host for the new
route registration to take effect.
2026-06-15 14:55:21 -07:00
Jordan Ritter fb3d5b3547 test(showcase/harness): cover CP webhook registration + fail-loud on missing SHARED_SECRET
Adds four red-green tests for the notify-harness 404 follow-up:

  A. CP boot registers POST /webhooks/deploy when SHARED_SECRET is set
     — RED pre-fix (404, route not mounted on CP); GREEN post-fix (401,
     the route's own HMAC reject path fires on an unsigned POST).
  B. Worker boot still registers POST /webhooks/deploy (regression
     guard) — already green; pins the unchanged behavior.
  C. boot throws FATAL-CONFIG when SHARED_SECRET unset AND NODE_ENV is
     non-test — RED pre-fix (current guard fires only on
     NODE_ENV='production'); GREEN post-fix once the predicate is
     tightened.
  D. boot succeeds when SHARED_SECRET unset AND NODE_ENV='test'
     (test-mode escape hatch) — already green; pins the escape hatch.

Red capture (pre-fix run):
  - Test A: AssertionError: expected 404 to be 401
  - Test C: AssertionError: promise resolved instead of rejecting
2026-06-15 14:52:31 -07:00
Martha Schumann 61471cac93 docs(showcase): fix shared randomUUID imports 2026-06-15 13:18:58 -07:00
Jordan Ritter c43ed08e7b ci(e2e-dojo): run dojo suites on 4-vCPU runner (−40% wall-clock) (#5452)
Bumps the dojo e2e matrix from `depot-ubuntu-24.04` (2 vCPU) to
`depot-ubuntu-24.04-4` (4 vCPU) and `NX_PARALLEL: 4` so the build uses
the extra cores. This is the non-serializing way to cut dojo wall-clock
(the build-once dedup tried in #5450 regressed wall-clock and was
reverted).

## Result: −40% wall-clock (measured on CI)

Dojo wall-clock = the single slowest suite (the 15 run in parallel).
Comparison vs the 2-vCPU baseline:

| metric | 2-vCPU baseline | 4-vCPU | Δ |
|---|---|---|---|
| **wall-clock** (long pole `langgraph-python`) | 623s (10.4m) | **373s
(6.2m)** | **−40%** |
| runner-minutes (wall summed, 15 suites) | 105m | 76m | −28% |
| **billed compute** (vCPU-min; 4-vCPU ≈ 2× rate) | ~210 | ~304 |
**+45%** |

Every suite got faster; the long-pole suites benefited most:

| suite | 2-vCPU | 4-vCPU |
|---|---|---|
| langgraph-python | 623s | 373s |
| langgraph-typescript | 547s | 362s |
| langgraph-fastapi | 500s | 337s |
| adk-middleware | 414s | 286s |
| (… all 15 faster …) | | |

Long-pole `langgraph-python` step breakdown:

| phase | 2-vCPU | 4-vCPU |
|---|---|---|
| Build cpk | 82s | 48s |
| Prep dojo | 94s | 52s |
| **Run tests (Playwright)** | **271s** | **117s** |
| total | 623s | 373s |

**Key finding:** the Playwright phase more than halved → the e2e suites
are **CPU/worker-bound, not LLM-latency-bound**. A bigger runner is the
right lever; test sharding is not needed to reach ~6 min.

## Trade-off
−40% wall-clock for **~+45% billed compute** (4-vCPU costs ~2×/min,
partly offset by finishing 28% sooner). If the cost bump isn't worth it
across all 15 suites, a follow-up can scope `-4` to just the slow suites
via a per-matrix `runner` field (wall ~6.5m, smaller cost increase).

Companion to #5450 (unit-test `nx affected`).
2026-06-15 12:49:19 -07:00
GeneralJerel d4fa49eed4 fix(showcase): refresh every view after a chat-driven mutation
useCreditCards() kept per-instance React state and was called independently by
the dashboard page and by copilot-context (where the chat's approve / finalize /
open-exception tools live). A mutation made through one instance refetched only
itself — so when the agent approved an over-limit charge in chat, the dashboard
pending table kept showing it as pending until a manual reload.

Add a module-level revalidation bus: each useCreditCards() instance registers a
refetch callback, and every mutation calls notifyDataChanged() to fan a refetch
out to all live instances. The dashboard now reflects agent-driven approvals
immediately (verified: recall-approve in chat drops the charge from the pending
table with no reload).
2026-06-15 12:19:52 -07:00
GeneralJerel 0b65391501 docs(showcase): clarify same-thread (OSS) vs cross-thread (Intelligence) recall
Adds a "What each mode actually recalls" subsection to the demo README: in OSS
mode the taught workflow is recalled only within the same conversation (the
saved procedure is echoed back into that thread), so a brand-new chat won't know
it — that's expected. Cross-conversation persistence is what the external
Intelligence backend provides. Names the symptom explicitly so reviewers aren't
surprised when a new conversation "doesn't know" the workflow in OSS mode.
2026-06-15 12:04:44 -07:00
GeneralJerel 7e1641ff1b feat(showcase): pending-approval table redesign + approval-gate UX fixes
Rework the dashboard's Pending approval view and fix two approval-gate UX bugs
on the banking demo (PR #5266):

- Table layout: replace the center-stacked per-row card with a scannable table
  (Merchant / Amount / Policy / Actions). Actions are check / x icon buttons plus
  a "more actions" overflow menu holding File policy exception. Status is its own
  column (Over limit / Cleared / Within limit), and Approve is gated until the
  charge is actually clearable.

- Fix the table shrinking when the more-actions menu opens: the Radix menu is
  modal by default and engaged react-remove-scroll, whose scrollbar compensation
  reflowed the table. Set modal={false} (a row menu needn't be modal) and add
  whitespace-nowrap to the status badges.

- Fix "cannot approve after filing an exception": the inline card offered all
  codes, including non-justifying ones that set activeExceptionId (flipping the
  row to Cleared) but never lift the server gate, so the approve 422'd silently.
  The card exists to clear an over-limit charge, so it now offers only justifying
  codes. The gate itself is unchanged.

Verified live in OSS dev: file (justifying) -> Cleared -> approve succeeds; the
menu opens without reflow; lint + build green.
2026-06-15 11:55:21 -07:00
Jordan Ritter 4791c86923 fix(showcase/harness): override LOCAL_SERVICES_JSON in --isolate generator to target the requested slug (#5454)
## Summary
- `showcase/docker-compose.local.yml:262` hardcodes
`LOCAL_SERVICES_JSON` to `showcase-langgraph-python` (intentional N=1
demo default).
- `showcase/bin/showcase test <slug> --d6 --isolate` was inheriting that
value verbatim into the iso1 stack, so the iso1 control-plane discovered
`showcase-langgraph-python` instead of `showcase-<requested-slug>`.
- Result: iso1 probes targeted the wrong service. Visible in iso1
harness logs as `discovery.railway-services.local-injection count:1
names:["showcase-langgraph-python"]` regardless of CLI arg.
- This PR teaches the iso1 compose generator (`apply_isolation` in
`showcase/scripts/cli/_common.sh`) to inject a per-slug
`LOCAL_SERVICES_JSON` override built from the slug's manifest.yaml demos
list. Fallback to `["agentic-chat"]` if manifest absent.

## Scope
Two files, ~55 LOC added: `showcase/scripts/cli/cmd-test.sh` (pass slug
arg), `showcase/scripts/cli/_common.sh` (inject regex sub in python
rewriter). Bash + embedded python only; no TypeScript changed.

## Verification
- **Local `--isolate` discovery confirmed correct:**
`showcase/bin/showcase test ms-agent-python --d6 --isolate` — iso1
harness log now shows `discovery.railway-services.local-injection
names:["showcase-ms-agent-python"]` (was `showcase-langgraph-python`
before this PR).
- Persistent stack default behavior (langgraph-python N=1) preserved via
fallback when no slug arg provided.
- **Heredoc hardening scope:** commit edc77f809 moves `$slug` from
bash-interpolation into the python rewriter to an env var. This is a
slug-only carve-out — `$slug` is the only value that originates from the
CLI arg path. The other bash-interpolated `$VAR`s embedded in the
heredoc (`$PORTS_FILE`, `$COMPOSE_FILE`, `$name`, `$SHOWCASE_ROOT`,
`$ISOLATE_PORT_OFFSET`) remain script-internal: each is constructed
inside `_common.sh` from validated sources (manifest reads, computed
offsets, fixed roots), not from user input, and CR Round 1 + Round 2
slot 5 both verified they are not user-tainted. A broader
env-var-pass-all-vars refactor would be a separate concern and is out of
scope for this PR.

## Out of scope
- Worker heartbeat clock-skew issue (`fleet.health.worker-unhealthy
lastHeartbeatAt N min stale`) is a separate bug; not touched.
- `buildLocalServicesJson` in `cli/control-plane-run.ts` has a similar
comment-vs-code disagreement (JSDoc says "filter", code returns env
verbatim) — flagged but not auto-fixed, as iso1 override is the correct
insertion point.
- Generalized env-var-pass for all heredoc-embedded $VARs (see
Verification) — separate refactor concern.
2026-06-15 11:38:26 -07:00
Jordan Ritter 318bd9e45f fix(showcase/aimock): align D6 tool-rendering catchall userMessage to D5 probe input (#5453)
## Summary
- D5 e2e-deep probes for `tool-rendering-{default,custom}-catchall` send
`"forecast for Tokyo"` as the test input (see
`showcase/harness/src/probes/scripts/d5-tool-rendering-{default,custom}-catchall.ts`).
- aimock uses substring match on `userMessage`.
- The catchall fixtures on main had a stale `"check Tokyo weather
forecast"` string that couldn't substring-match the probe input →
fixture miss → probe falls through to live LLM → CV ✗ red D4 on the
dashboard.
- This PR renames the userMessage to the canonical `"forecast for
Tokyo"` across 28 catchall fixture files in 16 integrations.

## Scope
28 files × ~2 userMessage occurrences each = 67 line changes. **No code,
no agent, no page.tsx changes.** Pure fixture-data alignment.

Integrations covered (default-catchall and/or custom-catchall): ag2,
agno, built-in-agent, claude-sdk-python, claude-sdk-typescript,
crewai-crews, google-adk, langgraph-python, langroid, llamaindex,
mastra, ms-agent-dotnet, ms-agent-python, pydantic-ai, spring-ai,
strands.

## Commit history note
Commit b1f19bdc8 changes the rename target from the original `"weather
in Tokyo"` (commit 388c69e68) to `"forecast for Tokyo"`. This was a CR
Round 1 catch: `"weather in Tokyo"` was a substring of the chain pill
prompt (`"weather forecast chain in Tokyo"` / similar), which would have
caused the catchall fixture to incorrectly match chain-pill traffic.
`"forecast for Tokyo"` has no such substring collision with any other
probe input.

## Verification
- **Live red-green proof** was performed on the prior `"weather in
Tokyo"` rename (commit 388c69e68) against `ms-agent-python` via
`showcase/bin/showcase test ms-agent-python --d6 --isolate`:
- Before (origin/main): catchall featureTypes red (fixture miss → live
LLM → flaky)
- After: `d6:ms-agent-python/tool-rendering-default-catchall=green`,
`d6:ms-agent-python/tool-rendering-custom-catchall=green`
- **The current HEAD's `"forecast for Tokyo"` rename (b1f19bdc8) has NOT
been re-run live.** It is verified by static analysis only: substring
math (no collision with any known D5 probe input or chain pill prompt)
and a clean CR Round 2 across all reviewing agents.
- Post-merge dashboard re-probe will be the final runtime verification.

## Out of scope (separate follow-ups)
- LG-TS / LG-FastAPI catchall fixtures don't have the stale string — use
pill-aligned userMessages; need different fix
- `tool-rendering` (non-catchall) and `tool-rendering-reasoning-chain`
featureTypes have separate failure modes
- Other red cells in the dashboard (frontend-tools-cosmic timeout,
agent-config, auth, etc) are unrelated
2026-06-15 11:33:34 -07:00
Jordan Ritter edc77f8090 fix(showcase/harness): pass slug via env var to python rewriter instead of bash interpolation
The python rewriter in apply_isolation previously interpolated $slug
directly into the inline python source via bash. A slug containing a
single quote would break the python literal. Internal-tool risk only
(slug is developer-typed), but cheap to harden.

Pass slug via SHOWCASE_ISO_SLUG env var and read os.environ.get(...)
inside the python heredoc. Defense-in-depth; no behavior change for
valid slugs.
2026-06-15 11:16:24 -07:00
Jordan Ritter b1f19bdc80 fix(showcase): rename catchall userMessage to 'forecast for Tokyo' to avoid chain pill substring collision
The previous 'weather in Tokyo' rename (388c69e68) was a substring of the
main tool-rendering chain pill prompt 'Chain a few tools in this single
turn: get the weather in Tokyo, search flights from SFO to Tokyo, and roll
a d20.' Because aimock loads fixtures alphabetically per integration dir
and uses substring match with first-match-wins, the catchall fixture
(loaded before tool-rendering.json) was intercepting chain pill matches
across 16 integrations.

Rename catchall fixture userMessage to 'forecast for Tokyo' — a phrase
not contained in any other pill prompt. Update the corresponding D5
catchall probe inputs in d5-tool-rendering-{default,custom}-catchall.ts
and the test assertions that pin those inputs.

Call-Site Enumeration: 'weather in Tokyo' remains intentionally in
page.tsx suggestions.ts pills (user-visible UX) and inside the chain
pill prompt itself — neither is in the substring-match path now.
2026-06-15 11:14:57 -07:00