mirror of
https://github.com/CopilotKit/CopilotKit.git
synced 2026-09-14 16:26:20 +08:00
codex/remove-threads-cli-path
12110 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
ba159ce8e1 |
chore(examples): bump slack example to @copilotkit/bot* 0.0.2
@copilotkit/bot, @copilotkit/bot-slack, and @copilotkit/bot-ui are published at 0.0.2 (agent-native assistant pane + native streaming). Point the slack example's dependency ranges at ~0.0.2 so a deployed/standalone install pulls the new packages. Local monorepo installs already use the workspace copies via root pnpm.overrides, so the lockfile is unchanged. |
||
|
|
7ed4a44447 |
chore: release bot v0.0.2 (#5475)
## Release bot v0.0.2 **Scope:** `bot` | **Bump:** `patch` --- ### How this release process works 1. **This PR was created automatically** by the "release / create-pr" workflow. It bumped the `bot` packages to `0.0.2` and generated AI-enhanced release notes. 2. **CI runs on this PR** — the full test suite (unit tests, lint, type checks, build) must pass before merging. This is the review gate. 3. **Review the release notes** in `release-notes.md` in this PR. If a Notion draft was created, you can edit the release notes there before merging. 4. **When this PR is merged**, the `release / publish` workflow automatically: - Builds all packages - Publishes the `bot` packages to npm at version `0.0.2` - Creates git tag `bot/v0.0.2` - Creates a GitHub Release with the final release notes ### Before merging - [ ] CI is green (tests, lint, types, build) - [ ] Version bumps look correct - [ ] Release notes are accurate (edit in Notion if a draft was created) --- > **Do not merge until CI is fully green.** The full test suite runs automatically on this PR.bot/v0.0.2 |
||
|
|
035e7cf9c9 | chore: release bot v0.0.2 | ||
|
|
72203fbcba |
chore: release bot-slack v0.0.2 (#5466)
## Release bot-slack v0.0.2 **Scope:** `bot-slack` | **Bump:** `patch` --- ### How this release process works 1. **This PR was created automatically** by the "release / create-pr" workflow. It bumped the `bot-slack` packages to `0.0.2` and generated AI-enhanced release notes. 2. **CI runs on this PR** — the full test suite (unit tests, lint, type checks, build) must pass before merging. This is the review gate. 3. **Review the release notes** in `release-notes.md` in this PR. If a Notion draft was created, you can edit the release notes there before merging. 4. **When this PR is merged**, the `release / publish` workflow automatically: - Builds all packages - Publishes the `bot-slack` packages to npm at version `0.0.2` - Creates git tag `bot-slack/v0.0.2` - Creates a GitHub Release with the final release notes ### Before merging - [ ] CI is green (tests, lint, types, build) - [ ] Version bumps look correct - [ ] Release notes are accurate (edit in Notion if a draft was created) --- > **Do not merge until CI is fully green.** The full test suite runs automatically on this PR.bot-slack/v0.0.2 |
||
|
|
e941049d3c |
fix(harness): cascade fallback to whole-bubble-minus-toolbar for tool-only responses (Class B) (#5474)
## Summary Adds a **Class B cascade fallback** in `readCascadeStateLast` so the harness settle gate can resolve text for **single-bubble tool-only responses** (recharts SVG, gen-UI cards, A2UI render-only output). These bubbles have ALL scoped text selectors empty but carry substantial rendered content in non-cascade children. The current cascade-pollution guard correctly returns `null` for arbitrary-index reads (`readCascadeState` / `findAssistantBubbleAt`) — but for the LAST bubble, by `RUN_FINISHED` the content is mounted and stable, so we can safely fall back to `bubble.textContent` MINUS the assistant-toolbar's textContent (the suffix the toolbar contributes). This is the residual fix on top of PR #5472 (Class A: multi-bubble last-bubble-has-prose) and #5473 (probe taxonomy cleanup). The ~200-red plateau on PocketBase after #5472 landed is consistent with Class B affecting the long tail of tool-render-only demos (ms-agent-python:beautiful-chat-*, plus other tool-render-only Class-B-shaped demos). ## Class B failure mode (live-confirmed) Probed against `https://showcase-ms-agent-python-staging.up.railway.app/demos/beautiful-chat` with the prompt _"Show me a bar chart of monthly sales for Q1 2026."_: - `count = 1` (single canonical bubble) - `bubble.querySelector('[data-message-content]')` -> empty string - `bubble.querySelector('.cpk:prose')` -> empty string - `bubble.querySelector('.prose')` -> empty string - `bubble.textContent` -> **351 chars** of real content (chart title, axis labels, animation styles, query metadata) - Current `readCascadeStateLast` returns `text=null` -> settle gate spins on `text-unstable` until timeout (RED). ## Red-green proof (verbatim from /tmp/cr/class-b-red-green.mjs) The probe runs TWO in-browser readers against the same DOM snapshot — the OLD (current production) and the NEW (with fallback). Run live against staging, **before any code changes** in this PR: ``` [probe] navigating to https://showcase-ms-agent-python-staging.up.railway.app/demos/beautiful-chat [probe] locating chat input [probe] sending prompt: Show me a bar chart of monthly sales for Q1 2026. [probe] waiting 20000ms for response to render [probe] result: { "count": 1, "oldText": null, "newText": "query_dataquery:\"monthly sales for Q1 2026\"\n @keyframes barSlideIn {\n from { transform: translateY(40px); opacity: 0; }\n 20% { opacity: 1; }\n to { transform: translateY(0); opacity: 1; }\n }\n Monthly Sales for Q1 2026Breakdown of income generated each month in Q1 2026.JanuaryFebruaryMarch0255075100January", "oldLen": null, "newLen": 351, "newPreview": "query_dataquery:\"monthly sales for Q1 2026\"\n @keyframes barSlideIn {\n from { transform: translateY(40px); opacity: 0; }\n 20% { opacity: 1; }\n to { transform: translat" } [probe] verdict: { "red": true, "green": true, "count": 1 } [probe] RED-GREEN confirmed: old=null, new=non-empty. ``` - **RED**: `oldText = null`, `oldLen = null` — matches current production behavior; settle gate cannot resolve text. - **GREEN**: `newText` is 351 chars of stable rendered content; settle gate can lock in `text-stable`. ## The fix In `showcase/harness/src/probes/helpers/assistant-message-count.ts`, inside `readCascadeStateLast`'s per-tier loop — AFTER the existing scoped-selector cascade exhausts without a non-empty match and BEFORE returning `{count, text: null}`: 1. Read `bubble.textContent` (whole bubble). 2. Read `bubble.querySelector('[data-testid="copilot-assistant-toolbar"]').textContent` if present. 3. If the toolbar's text is a **trailing suffix** of the whole-bubble text (which it canonically is — toolbar is a leaf sibling of the prose div), strip it. 4. If what remains is non-empty after trim -> return `{count, text: <trimmed>}`. 5. Otherwise (no toolbar, no content, or only-toolbar text) -> return `{count, text: null}` to keep polling. The same fallback is NOT applied to `findAssistantBubbleAt` / `readCascadeState` because those address arbitrary indices and intermediate bubbles in multi-step responses can transiently carry empty scoped text while the NEXT bubble is mounting. Reading whole-bubble at intermediate indices would re-introduce the cross-bubble text flap PR #5462 was designed to prevent. The LAST bubble in a multi-step turn is the agent's terminal output, mounted by `RUN_FINISHED`. ## Why the fallback is safe (toolbar-suffix invariant) In the canonical CopilotKit assistant bubble (`CopilotChatAssistantMessage.tsx`), the toolbar is the **last child** of the bubble, mounted as a leaf sibling of the prose/markdown wrapper. Its `textContent` is therefore the trailing slice of `bubble.textContent`. Stripping it by suffix-match yields the message content. When the suffix-match fails (defensive — e.g. the DOM shape changes), we return `bubble.textContent` unchanged rather than mangle it. When the toolbar isn't present (headless / custom-composer), we return whole text. When even whole-text is empty, we return `null` (don't lock the settle gate on a placeholder). ## Tests - `assistant-message-count.test.ts`: 9 new test cases under a new `describe("readCascadeStateLast")` block. Pins: - non-empty scoped text path unchanged (no fallback when scoped works) - Class B fallback: scoped empty + whole-bubble non-empty + toolbar suffix -> strip suffix, return content - Class B fallback works without a toolbar (returns whole text) - defensive: toolbar text NOT a suffix -> return whole text unchanged - everything empty -> return `text:null` (settle gate keeps polling) - only-toolbar text -> return `text:null` after strip - multi-bubble: reads the LAST bubble, not the first - no tier matches -> `{count:0, text:null}` - `evaluate()` throws -> swallowed -> `{count:0, text:null}` - Full harness suite: **2769 / 2769 passing across 129 files** (`pnpm exec vitest run` in `showcase/harness`). - TypeScript: `pnpm exec tsc --noEmit` clean. - Lint: `oxlint` 0 warnings, 0 errors. - Format: `oxfmt --check` clean. ## Out-of-scope (untouched) - `conversation-runner.ts` (the settle-gate logic stays — the cascade does the right thing now). - `sse-interceptor.ts`, `init-scripts.ts`, `d6-all-pills.ts`. - `findAssistantBubbleAt` / `readCascadeState` (intentionally retain the pollution guard — they address arbitrary indices). - `resolveBubbleTextFromSelectors` (pure sibling of `findAssistantBubbleAt`, intentionally retains pollution-guard semantics). ## Verification plan after merge 1. Watch CI green on the PR. 2. Promote `showcase-ms-agent-python` and other tool-render-only services. 3. Observe PocketBase D6 red counts: expect the ~200-red plateau to drop further toward baseline. 4. If any demo regresses, revert `readCascadeStateLast` alone. ## Files changed - `showcase/harness/src/probes/helpers/assistant-message-count.ts` (production fix) - `showcase/harness/src/probes/helpers/assistant-message-count.test.ts` (9 new tests) |
||
|
|
72520de43d | fix(harness): cascade last-bubble whole-bubble-minus-toolbar fallback for tool-only responses | ||
|
|
f7779f90bf |
refactor(showcase): probe taxonomy cleanup — drop Smoke, rename E2E (Demo) → UI (Frontend) (#5473)
## Summary
The smoke probe was the same HTTP contract as `/health` on the same
service (200-OK JSON body), so every tick paid two HTTP calls for the
same liveness signal. This PR:
- Drops the `/smoke` GET probe and the `smoke:<slug>` primary
`ProbeResult`. The driver now emits `health:<slug>` as the primary and
`agent:<slug>` via writer side-emit (half the per-tick cost).
- Renames the dashboard's D3 / e2e row label from `E2E (Demo)` to `UI
(Frontend)` (and the short cell badge from `E2E` to `UI`) — the row was
already about "the demo page renders in a browser", not E2E in the
traditional sense.
- Drops the `Smoke` row from the cell drilldown (CellState.smoke field
is retained for back-compat so historical rows still parse).
- Adds an explicit `Health` row to the legend (was implicit).
The driver's registry `kind` stays `"smoke"` so existing YAML configs +
orchestrator family wiring keep routing to this driver — the emission
key (`health:<slug>`) is the taxonomy contract that matters. The
underlying probe key (`e2e:<slug>/<feature>`) is preserved on PocketBase
so historical rows render correctly during the rename window.
## RED proof (BEFORE)
`smoke:<slug>` primary emission in `liveness.ts` (origin/main):
```
18: * 1. `smoke:<slug>` — the RETURN VALUE of `run()`. The invoker runs
56: * `smoke:<slug>`/`health:<slug>`/`agent:<slug>` triple the dashboard
241: input.mode === "discovery" ? `smoke:${slug}` : input.key;
```
`E2E (Demo)` / `Smoke` strings in dashboard render code (origin/main):
```
cell-drilldown.tsx:38: * Trip)" (chat+tools round-trip), D3/e2e = "E2E (Demo)" (the demo page loads
cell-drilldown.tsx:49: { key: "e2e", label: "E2E (Demo)" },
cell-drilldown.tsx:52: { key: "smoke", label: "Smoke" },
cell-pieces.tsx:382: name="E2E"
unified-cell.tsx:231: <TestBadge name="E2E" level={model.d3} />
adaptive-legend.tsx:59: E2E (Demo): demo page loads and round-trips in a browser
```
## GREEN proof (AFTER)
Primary emission rewritten — driver returns `health:<slug>` and no
`smoke:<slug>` ProbeResult is ever emitted:
```
33: * to `/smoke` and emitted a `smoke:<slug>` primary result, but that probe
230: // `kind: "smoke"` is the registry/family identifier the orchestrator
236: kind: "smoke", // registry-key back-compat; no smoke ProbeResult is emitted
367: if (input.mode === "discovery") return `health:${slug}`;
368: if (input.key.startsWith("smoke:")) {
369: return `health:${input.key.slice("smoke:".length)}`;
```
New labels in dashboard:
```
cell-drilldown.tsx:60: { key: "e2e", label: "UI (Frontend)" },
cell-pieces.tsx:382: name="UI"
unified-cell.tsx:231: <TestBadge name="UI" level={model.d3} />
adaptive-legend.tsx:65: UI (Frontend): demo page renders in browser (Playwright)
```
Confirmation that the old labels are gone from render code:
```
$ grep -rn 'E2E (Demo)\|"Smoke"' cell-drilldown.tsx cell-pieces.tsx unified-cell.tsx adaptive-legend.tsx
(no matches — labels removed)
```
## Test results
- `pnpm exec vitest run` in `showcase/shell-dashboard`: **63 files, 1089
tests passed, 1 skipped**
- `pnpm exec vitest run` in `showcase/harness`: **129 files, 2760 tests
passed**
- `pnpm exec tsc --noEmit` in both: clean
Liveness driver tests include a regression guard (`never emits a
smoke:<slug> key`) so the smoke contract cannot be re-introduced
silently.
## Test plan
- [x] Dashboard vitest green (1089/1089)
- [x] Harness vitest green (2760/2760)
- [x] tsc --noEmit green in both packages
- [x] Visual cross-check: drilldown renders 6 rows (no Smoke), e2e row
labelled `UI (Frontend)`; cell strip shows `UI` badge in place of `E2E`;
legend shows explicit `Health` row + `UI (Frontend)` D3 line
|
||
|
|
53f021afe8 |
refactor(showcase): probe taxonomy cleanup — drop Smoke, E2E (Demo) → UI (Frontend)
The smoke probe was the same HTTP contract as /health on the same service (200-OK JSON body), so every tick paid two HTTP calls for the same liveness signal. Drop the /smoke GET + the smoke:<slug> primary ProbeResult; the driver now emits health:<slug> as the primary and agent:<slug> via writer side-emit (half the per-tick cost). The driver's registry kind stays "smoke" so existing YAML configs and orchestrator family wiring keep routing to this driver — the emission key (health:<slug>) is the taxonomy contract that matters. Dashboard: - Drilldown: D3/e2e row labelled "UI (Frontend)" (was "E2E (Demo)"); Smoke row dropped (CellState.smoke field retained for back-compat). - Cell badges: short label "UI" (was "E2E") in cell-pieces + unified-cell. - Legend: D3 = "UI (Frontend): demo page renders in browser (Playwright)" plus an explicit Health row at the top. Tests updated to assert new labels: - cell-drilldown.test.tsx: 6 dimensions (no Smoke); UI (Frontend) label. - cell-drilldown.lazy-signal.test.tsx: drilldown-badge-ui--frontend- testid. - cell-pieces.test.tsx + .signal-degrade.test.tsx: badge name "UI". - unified-cell.test.tsx: mock-badge-UI. - overlay-selector-integration.test.tsx: "UI" in place of "E2E". - dashboard-color-matrix.test.tsx: badge: "UI" case names. - liveness.test.ts: two-call contract (health + agent), regression guard asserting no smoke:<slug> ever emitted. Underlying probe key (e2e:<slug>/<feature>) preserved on PocketBase so historical rows render correctly during the rename window. |
||
|
|
665dd5aae6 |
fix(harness): waitForTurnComplete reads LAST bubble (multi-step agent regression from #5462) (#5472)
## Summary **URGENT hotfix.** Staging is widely red for D5/D6 multi-step demos. Root cause confirmed via live Playwright DOM inspection on staging (langgraph-typescript : beautiful-chat : pie chart prompt). PR #5462's defect-2 fix uses `bubbleIndex = turnIndex - 1`, which assumes one assistant bubble per turn. Multi-step agents (LangGraph, Mastra, CrewAI, llama-index, ag2, etc.) emit 2-3+ bubbles per turn — tool-call bubble + tool-render bubble + final-text bubble. The strict-index read lands on an intermediate **tool-call** bubble whose scoped-text selectors (`.cpk:prose`, `[data-message-content]`, `.prose`, `p`) are **empty** (tool-call content lives in a sibling card div outside the scoped cascade). The cascade-pollution guard correctly returns `null` → the text-stable conjunct never converges → `reason=text-unstable` timeout across virtually every multi-step demo (beautiful-chat-*, gen-ui-*, shared-state-*, agentic-chat on most backends). ## Live RED-GREEN proof (staging) Standalone Playwright probe against `showcase-langgraph-typescript-staging.up.railway.app/demos/beautiful-chat`, prompt = `"pie chart of revenue distribution by category from the sample sales data"`, captured BOTH the OLD-logic read (`bubbleIndex = turnIndex - 1 = 0`) AND the NEW-logic read (`bubbleIndex = count - 1 = 2`) in ONE atomic page evaluate after the turn settled: ``` TOTAL BUBBLE COUNT FOR TURN 1: 3 (true multi-step — tool-call + tool-render + final-text) FINAL TIER: [data-testid="copilot-assistant-message"] OLD LOGIC (bubbleIndex = turnIndex - 1 = 0): count = 3 text = null reason = all scoped selectors empty ← would time out: reason=text-unstable NEW LOGIC (bubbleIndex = count - 1 = 2): count = 3 text = "Here is the pie chart showing the revenue distribution by category from the sample sales data." textLen = 94 ← gate settles cleanly ``` RED ≠ GREEN; new logic returns 94 chars of stable final-text where old logic returns `null`. The bug and the fix are proven on real staging DOM. ### Live DOM evidence LangGraph TS beautiful-chat, prompt = "pie chart of revenue distribution by category…": ``` count = 3 bubbles in [data-testid="copilot-assistant-message"] bubble[0]: tool-call "query_data" — `.cpk:prose` exists but EMPTY (content in sibling <div class="my-1.5">) bubble[1]: pie-chart card — `.cpk:prose` exists but EMPTY (content in sibling styled card div) bubble[2]: final text "Here is the pie chart..." — `.cpk:prose` has 94 chars + toolbar mounted ``` The old pre-#5462 runner read `list[last]` (always the last bubble — always bubble[2] above). The PR moved to `turnIndex - 1`, which is bubble[0] (the tool-call) for any multi-step turn → broken. ## Fix Read the LAST bubble in the matched cascade tier (`count - 1`) instead of a fixed strict index. Defect-2 protection (don't read a leftover bubble from a prior turn) is preserved via an explicit **pre-submit baselineCount sentinel snapshot** in the runner — the gate now requires `count > baselineCount` rather than `count > turnIndex - 1`. This correctly rejects stale prior-turn bubbles **without assuming 1 bubble = 1 turn**. ### Changes - `assistant-message-count.ts`: add `readCascadeStateLast(page)` — same atomic single-evaluate contract as `readCascadeState`, but resolves the bubble index internally as `count - 1`. - `conversation-runner.ts`: - Snapshot `baselineCount = countAssistantMessages(page)` BEFORE `sendTurnMessage()` each turn. - Pass through to `waitForTurnComplete` via new `baselineCount` option (defaults to `turnIndex - 1` to preserve unit-test fake compatibility). - `waitForTurnComplete` reads via `readCascadeStateLast` and gates on `domOk = count > baselineCount`. - Cold-start retry re-snapshots `baselineCount` after `page.reload()`. - Returned `bubbleIndex` in the success ctx = `count - 1` (matches the bubble whose text settled the gate). - `dom-missing` reason classifier updated to `countFinal <= baselineCount`. ## Why this doesn't regress defect-2 The original defect-2 (un-turn-scoped bubble selection) leaked a later turn's bubble into THIS turn's assertions because `readLastAssistantText` read `list[last]` GLOBALLY with no per-turn baseline. The new fix re-introduces the "read last" semantics but adds the **per-turn pre-submit baseline count snapshot** — the gate only advances once a NEW bubble has appeared (`count > baselineCount`), so we cannot read a leftover from a prior turn, and the "last" we return is the last NEW bubble (which for multi-step turns is the final-text bubble whose scoped text actually settles). ## Test plan - [x] **Live RED-GREEN staging probe** (see above) — confirms OLD logic = null and NEW logic = 94-char final-text on a real 3-bubble multi-step turn. - [x] All 72 `conversation-runner.test.ts` + `assistant-message-count.test.ts` unit tests pass (default `baselineCount = turnIndex - 1` keeps existing count-progression scripts working). - [x] Full harness vitest suite: 2758/2758 green across 129 test files. - [x] `tsc --noEmit` clean on `showcase/harness`. - [x] `oxfmt --check` clean on changed files. - [x] `oxlint` clean on changed files. - [ ] Staging D5/D6 sweep after merge confirms multi-step demos (beautiful-chat, gen-ui, shared-state, langgraph/mastra/crewai agentic-chat) recover from `text-unstable` red. ## Scope Strictly the two files above. No unrelated changes. No refactors. |
||
|
|
88e77f1403 |
fix(showcase): allow harness to boot locally without SHARED_SECRET (unblock local verify) (#5471)
## Problem PR #5458 (`c81b361f1` — *fix(showcase/harness): register /webhooks/deploy on CP path + fail-loud on missing SHARED_SECRET*) added a fail-loud gate to `loadWebhookSecrets` in `showcase/harness/src/orchestrator.ts` that refuses harness boot in any deployable mode (`NODE_ENV !== "test"`) without `SHARED_SECRET` (or `SHARED_SECRET_PREV`) set, because `POST /webhooks/deploy` is only registered when `webhookSecrets.length > 0` (gate at `showcase/harness/src/http/server.ts:119`). The local `showcase/docker-compose.local.yml` harness services inherit `NODE_ENV=production` from the image and do not set `SHARED_SECRET` (and should not — that secret only matters for the Showcase: Verify Deploy webhook flow on Railway). As a result, the harness crash-loops on every local `bin/showcase test --d5/--d6` run: ``` FATAL-CONFIG: SHARED_SECRET (or SHARED_SECRET_PREV) is required — refusing to boot in any deployable mode … Set SHARED_SECRET (or SHARED_SECRET_PREV) in the env, or set NODE_ENV=test / HARNESS_ALLOW_NO_SECRET=1 for local dev. Current NODE_ENV=production. ``` …which blocks **all** local D5/D6 verify. ## Fix Add `HARNESS_ALLOW_NO_SECRET=1` (the documented local-dev escape hatch — explicitly named in the FATAL-CONFIG message itself, and the dedicated env var checked by `loadWebhookSecrets` at `orchestrator.ts:190`) to both harness services in `showcase/docker-compose.local.yml`: - `harness-control-plane` - `harness-pool-worker` Inline comments explain the rationale and pin the relevant source locations (the gate at `src/http/server.ts:119` and `loadWebhookSecrets` in `src/orchestrator.ts`). Diff: 17 lines added (comments + 2 env vars), no other files touched. ## Prod impact: NONE Railway sets `SHARED_SECRET` explicitly via env on every harness service, so `loadWebhookSecrets` sees a real secret, registers `POST /webhooks/deploy`, and **never reads `HARNESS_ALLOW_NO_SECRET`**. The escape hatch only takes effect when both secrets are absent. This change only affects the local docker-compose stack. ## Local verification Brought up the slot-1 isolated stack with the fix applied via `SHOWCASE_ISO_SLOT=3 ./bin/showcase test langgraph-typescript --d6 --verbose --isolate` and observed clean boot: ``` showcase-iso1-harness Up 12 seconds (healthy) showcase-iso1-langgraph-typescript Up 12 seconds (healthy) showcase-iso1-harness-pool-worker-1 Up 12 seconds (healthy) showcase-iso1-dashboard Up 13 seconds (healthy) showcase-iso1-aimock Up 18 seconds (healthy) showcase-iso1-pocketbase Up 18 seconds (healthy) ``` Harness boot logs confirm: - `fleet.role-selected` + `fleet.control-plane.started` + `showcase-harness.fleet.control-plane.boot` (port 8080) — control-plane up - `webhook auth disabled — neither SHARED_SECRET nor SHARED_SECRET_PREV is set` with `escapeHatch:true` at **warn** level — the expected log produced by `loadWebhookSecrets` when the escape hatch fires - `worker.registered` + `fleet.worker.boot` — worker self-registered - `worker.claimed jobId=… probeKey=d6:langgraph-typescript` — real D6 job picked up and run end-to-end NO `FATAL-CONFIG`, NO restart loop. (The D6 cell may still be RED on this branch — that is a pre-existing issue, unrelated to this infra fix. The point of this PR is "harness boots and runs a job at all locally", which it now does.) ## References - PR #5458 (`c81b361f1`): the fail-loud gate this PR re-enables an escape for - `showcase/harness/src/orchestrator.ts` `loadWebhookSecrets` (~line 180-220): the predicate that reads `HARNESS_ALLOW_NO_SECRET` - `showcase/harness/src/http/server.ts:119`: the route-registration gate |
||
|
|
d145bb1708 |
fix(harness): waitForTurnComplete reads LAST bubble in matched tier
PR #5462's defect-2 fix assumed 1 bubble per turn (bubbleIndex = turnIndex - 1). Multi-step agents (LangGraph, Mastra, CrewAI) emit 2-3+ bubbles per turn — tool-call + tool-render + final-text — so the strict-index read lands on an intermediate tool-call bubble whose scoped-text selectors (`.cpk:prose`, `[data-message-content]`, `p`) are EMPTY (its content lives in a sibling card div outside the scoped cascade). The cascade-pollution guard returns null forever and the text-stable conjunct times out with `reason=text-unstable` across virtually every multi-step demo (beautiful-chat-*, gen-ui-*, shared-state-*, agentic-chat on LangGraph/Mastra/CrewAI/etc.). Live DOM evidence from staging (langgraph-typescript : beautiful-chat : pie chart prompt): count = 3 bubbles bubble[0]: tool-call "query_data" — `.cpk:prose` empty bubble[1]: pie chart card — `.cpk:prose` empty bubble[2]: final text "Here is the pie chart..." — 90 chars Fix: read the LAST bubble in the matched cascade tier (count - 1) instead of a fixed strict index. Defect-2 protection (don't read a leftover bubble from a prior turn) is preserved via an explicit pre-submit `baselineCount` sentinel snapshot in the runner; the gate now requires `count > baselineCount` rather than `count > turnIndex - 1`, which correctly rejects stale bubbles without assuming 1 bubble = 1 turn. Changes: - assistant-message-count.ts: add `readCascadeStateLast(page)` — same atomic single-evaluate contract as `readCascadeState` but resolves the bubble index internally as `count - 1`. - conversation-runner.ts: snapshot `baselineCount` before sendTurnMessage; pass through to `waitForTurnComplete` via new `baselineCount` option (defaults to `turnIndex - 1` for unit- test fakes); `waitForTurnComplete` reads via `readCascadeStateLast` and gates on `count > baselineCount`. Cold-start retry re-snapshots `baselineCount` after page.reload. Test plan: - All 72 conversation-runner + assistant-message-count unit tests pass (default baselineCount preserves count-progression scripts). - Full harness suite 2758/2758 green, tsc clean, oxfmt/oxlint clean on changed files. |
||
|
|
48ef506abf |
fix(showcase): allow harness to boot locally without SHARED_SECRET
PR #5458 (
|
||
|
|
7cc579a58d |
chore: drop dead .changeset/ debris and workflow path filters (#5470)
## Summary The repo migrated off `@changesets/*` to conventional-commit-driven releases. This chore PR removes the leftover debris: - **11 stale `.changeset/*.md` files** (the entire `.changeset/` directory) - **4 dead `paths:` filter lines** in two e2e workflows ## Migration context Releases are driven by `scripts/release/prepare-release.ts` / `scripts/release/lib/changes.ts::getChangesSummary`, which reads `git log <lastTag>..HEAD` commit subjects. It never touches `.changeset/`. Verification — none of the changesets tooling is wired up anymore: - No `.changeset/config.json` - No `@changesets/*` in any `package.json` (root or workspace) - No npm scripts reference `changeset` The 11 `.changeset/*.md` files describe changes that have either already shipped (via commit subjects in prior releases) or will ship in the next release (via the current commit subjects in the `v1.60.1..main` window). The files are inert with respect to releases — most are 2+ months old; the newest (`slack-agent-native-apis.md`) was added recently by a contributor who didn't know about the migration. ## What's removed `.changeset/` (11 files): - `angular-21-install-and-types.md` - `angular-disable-license-watermark.md` - `bump-license-verifier-ent-251.md` - `debug-mode.md` - `empty-mails-applaud.md` - `ent-314-thread-connect-ux.md` - `ent-658-generated-thread-tool-roundtrip.md` - `five-avocados-visit.md` - `fix-thread-switch-state-reset.md` - `little-pears-tell.md` - `slack-agent-native-apis.md` Workflow `paths:` filters (4 lines across 2 files): - `.github/workflows/test_e2e-dojo.yml` — `push.paths`, `pull_request.paths`, and the `dorny/paths-filter` `ts:` filter - `.github/workflows/test_e2e-legacy-v1.yml` — `push.paths` Nothing to fire on after the directory is gone. ## Test plan - [ ] CI green on this PR |
||
|
|
5afa55f067 |
chore: drop dead .changeset/ debris and workflow path filters
The repo migrated off @changesets/* to conventional-commit-driven releases. scripts/release/lib/changes.ts::getChangesSummary reads `git log <lastTag>..HEAD` commit subjects and never touches .changeset/. No .changeset/config.json, no @changesets/* in any package.json, no npm scripts reference it. The 11 .changeset/*.md files describe changes that have either already shipped (via commit subjects in prior releases) or will ship in the next release (via the current commit subjects in the v1.60.1..main window) — the .changeset/ files are inert. Also removes the dead workflow `paths:` filters in test_e2e-dojo.yml and test_e2e-legacy-v1.yml that re-fired e2e on .changeset/ changes — nothing to fire on after the directory is gone. |
||
|
|
f8b14f1d6a |
fix(react-core): await runAgent in useInterrupt::resolve and propagate to demo-local hooks (#5461)
## Summary
Convert `useInterrupt::resolve` and the demo-local
`useHeadlessInterrupt::resolve` callbacks from fire-and-forget into
`async` + `return await copilotkit.runAgent(...)`, so callers receive a
Promise that settles when the resume run settles.
## Mechanism
Pre-fix: `resolve(...)` called `copilotkit.runAgent({...})` without
`await` and without `return` — the arrow returned `undefined`. The
showcase harness DOM-settle check on the assistant confirmation bubble
timed out because consumers had no handle to sequence against the resume
run's settle.
Post-fix:
- Framework `useInterrupt::resolve` is `async`, `return await`s
`runAgent`, wraps in try/catch + `setPendingEvent(null)` +
`console.error` + rethrow on rejection.
- `onRunFailed` now also `setPendingEvent(null)` — symmetric with
`onRunStartedEvent`, prevents stuck popups on run failure.
- `InterruptHandlerProps.resolve` / `InterruptRenderProps.resolve` typed
`() => Promise<RunAgentResult>` (was `() => void`).
- The same fix-pattern applied to 13 demo-local `useHeadlessInterrupt`
implementations across integrations: ag2, agno, claude-sdk-typescript,
crewai-crews, langgraph-{fastapi,python,typescript}, langroid,
llamaindex, mastra, pydantic-ai, spring-ai, strands.
## Verification
- **Unit:** new tests `resolve returns a Promise that settles when
runAgent settles (RESUME-PATH)` + `resolve rejects when runAgent
rejects, logs the failure, and clears pending (RESUME-PATH-REJECT)`.
Red-then-green on baseline. 22/22 react-core hook tests passing.
- **Backplane (end-to-end):** showcase control plane on langgraph-python
via `./bin/showcase test langgraph-python:gen-ui-interrupt --d5 --direct
--verbose` (GREEN, 70.1s, was RED-RESUME-PATH baseline) and
`:interrupt-headless --d5 --direct --verbose` (GREEN, 7.7s, was
RED-RESUME-PATH baseline). Built Docker artifact + aimock D6 fixture
replay — staging-equivalent path.
## CR
4 review rounds × 7 unbiased agents = 28 agent-runs. 2 Procedure 3
audits (3 promotions in round 1 → Fix C + D; 0 promotions in round 2).
Convergence achieved per agent-confirmed `bucket (a) empty` + Procedure
3 zero `PROMOTE_TO_A`.
## Out of scope (deferred follow-ups)
- **No manifest unquarantine.** A prior version of this branch flipped
`interrupt-headless` and `gen-ui-interrupt` from
`not_supported_features` to `features` across 12 manifests, but Phase 4
backplane validation across 8 PATCHED integrations × 2 demos surfaced
(a) `langgraph-fastapi:gen-ui-interrupt` still RED-RESUME-PATH because
the framework patch doesn't propagate from workspace to the
integration's Docker image without a published `@copilotkit/react-core`
release, and (b) a separate popup-mount RED-RUNTIME-OTHER cluster across
6 integrations that's pre-existing rot unrelated to this PR. Manifest
unquarantine + `@copilotkit/react-core` version bump are deferred to a
follow-up `copilotkit-release` PR after this fix lands.
- **No version bump.** Release follows in a separate PR.
- **Bucket (c) follow-up backlog** (audit verdicts: 11 STAY_IN_C):
- JSON.parse guarding in showcase demos (pre-existing, all 13)
- Type-design quirks in `InterruptHandlerProps.result` (non-nullable in
type, nullable in runtime)
- `renderInChat` dynamic toggle leak (pre-existing edge case)
- JSDoc clarity ("symmetric with onRunFailed" wording, `@typeParam
TRenderInChat`, `event.value` "any" → "unknown")
- StrictMode test coverage gap
- 13× duplicate `useHeadlessInterrupt` → bucket (d): consolidate into a
shared module in a future PR
## Files changed
16 files, +486/-154:
- `packages/react-core/src/v2/hooks/use-interrupt.tsx`
- `packages/react-core/src/v2/hooks/__tests__/use-interrupt.test.tsx`
- `packages/react-core/src/v2/types/interrupt.ts`
- 13×
`showcase/integrations/<integ>/src/app/demos/interrupt-headless/page.tsx`
## Test plan
- [x] react-core unit tests (22/22 passing including new RESUME-PATH +
RESUME-PATH-REJECT regression tests)
- [x] Backplane D5 probe on langgraph-python (gen-ui-interrupt +
interrupt-headless GREEN end-to-end via the showcase control plane)
- [ ] CI green on this PR
|
||
|
|
4564e6205d | chore: release bot-slack v0.0.2 | ||
|
|
431d5baae0 |
feat(bot-slack): agent-native Slack APIs — assistant pane + native streaming (default-on) (#5447)
## What
Activates Slack's agent-grade APIs as the **default** experience for
`@copilotkit/bot-slack`, with zero config and safe degradation.
Implements the Notion spec *"Slack Agent APIs in bot-slack — assistant
pane + native streaming by default"* (Rev 2).
### Assistant pane ("Agents & AI Apps")
- Opening the pane greets the user + shows tappable prompt chips; each
pane conversation is its own thread (replies stay in-thread, instead of
leaking to a merged flat DM).
- While the agent runs, native composer status
(`assistant.threads.setStatus`: "is thinking…", "is using `tool`…")
replaces placeholder/`🔧` messages.
- Pane threads are auto-titled from the first message.
- Customize via the new `assistant` option; `assistant: false` disables
pane handling. Apps without the Agents toggle behave exactly as before
(pane machinery dormant).
### Native streaming
- Replies stream via `chat.startStream` / `appendStream` / `stopStream`
(raw markdown — real tables / fenced code render natively) wherever a
thread exists.
- Flat DMs and workspaces without the streaming API fall back to the
legacy `chat.update` transport **automatically** — the first
`startStream` failure marks the workspace legacy and replays via legacy.
Opting in can never break a bot. `streaming: "legacy"` forces the old
transport.
### Portable engine surface (`@copilotkit/bot`)
- `bot.onThreadStarted` lifecycle handler + `IncomingThreadStart` sink
event.
- Capability-gated `thread.setSuggestedPrompts` / `thread.setTitle` (the
shipped `postFile` gating pattern).
- Two `SurfaceCapabilities` flags + optional `PlatformAdapter` methods.
All degrade gracefully on surfaces without support.
## Files
- **New:** `bot-slack/src/assistant.ts` (Bolt `Assistant` middleware ⇄
engine sink), `bot-slack/src/native-stream.ts` (`NativeMessageStream`,
same `append/finish` contract as `MessageStream`, legacy fallback).
- **Engine:** `platform-adapter.ts`, `thread.ts`, `create-bot.ts`,
`index.ts` (+ `bot-ui` `Thread` type).
- **Adapter:** `adapter.ts` (options/capabilities/stream
branch/methods), `slack-listener.ts` (one-line no-double-delivery
guard), `event-renderer.ts` (pane status mode + native text transport).
- **Example/docs:** `examples/slack` manifest (`assistant_view` + scope
+ events) & dev-ex, package READMEs/ARCHITECTURE, `slack.mdx`.
- **Hooks:** `chore(hooks)` makes the `lint-fix` script
double-quote-free (it was breaking `sh -c "…"` on Windows/lefthook
2.1.1).
## Tests / verification (CI-runnable)
- New unit tests: `native-stream` (cadence, 12k continuation, fence
re-open, first- **and** continuation-`startStream` fallback),
`assistant` (defaults-before-onThreadStarted ordering, thread-scoped
turn, auto-title), listener no-double-delivery, renderer pane-status
mode, engine capability gating.
- Verified locally: `@copilotkit/bot` 31 tests, `@copilotkit/bot-slack`
199 tests, `bot-ui` tests; `tsc` clean (bot + bot-slack); build clean;
oxlint/oxfmt clean; publint/attw pass.
## ⚠️ E2E not covered here
The §8 end-to-end flows need a **real Slack workspace** with the Agents
toggle (open pane → chips → streamed reply w/ live status; channel
mention → native markdown; kill `startStream` → legacy fallback;
non-agent app → shipped behavior). `examples/slack` is the vehicle.
Three items remain to confirm against a live workspace (spec §7 spikes):
the `isAssistantThread` runtime detection, the exact `assistant_view`
manifest acceptance, and `stopStream` finalize tolerance.
Spec:
https://app.notion.com/p/copilotkit/Slack-Agent-APIs-in-bot-slack-assistant-pane-native-streaming-by-default-spec-37b3aa38185281f5948dc6a665064d04
|
||
|
|
c88d687456 |
fix(showcase/harness): bubble-race elimination — 4 defects + atomic readCascadeState + cold-start retry + SSE counter (#5462)
## Summary
Fixes 4 bubble-race defects in the CopilotKit showcase e2e harness's
conversation runner that caused flaky turn-completion detection on
slow-streaming and cold-start integrations.
**Defects fixed (RED → GREEN):**
- **Defect 1** (fast-replay): turn settled on count, not text/SSE →
false-positive settle
- **Defect 2** (multi-turn flicker): un-turn-scoped bubble selection via
`list[last]` → cross-turn leak
- **Defect 3** (cascade blindness): diagnostic cascade picked wrong tier
→ wrong bubble text
- **Defect 4** (boot-time baseline staleness): pre-paint/stale bubble at
boot poisoned turn-1 settle
**Core changes:**
- `waitForTurnComplete` 3-conjunct primitive (SSE counter + DOM-at-index
+ text-quiet for settleMs)
- Atomic `readCascadeState` helper — single `page.evaluate` returns
`{count, text}` from one cascade tier (prevents within-poll cross-tier
inconsistency)
- Probe contract: `assertions(page, ctx: {bubbleIndex, text})` —
turn-scoped bubble retrieval (replaces `list[last]`)
- SSE interceptor — `__hk_runsFinished` counter via CDP + page-side
fetch wrap; idempotent attach/detach + handle-cache invalidation on stop
+ `framenavigated` reset (cold-start retry compatible)
- Cold-start retry — single bounded retry on first-attempt banner;
reload + re-resolve chat input + settleMs floor honoring shared turn
deadline
- Banner fast-fail — baseline banner snapshot + differs-from-baseline
2-poll debounce; in-poll `BannerVisibleError` translates to
`AssistantErroredError`
- Pre-paint placeholder env adapter for defect-4 repro infrastructure
**Defect repro tests:** real-browser tests in
`test/integration/bubble-race-repro-defect-{1,2,3,4}.test.ts` plus
mechanism-GREEN tests and `wait-for-turn-complete.test.ts` (3-conjunct
classification matrix), `probe-contract.test.ts` (ctx bridge),
`sse-counter.test.ts` (counter increment + framenavigated reset +
handle-cache invalidation).
## Test plan
- [x] Full harness vitest: 2742/2742 passing
- [x] Defect-1, defect-2, defect-4 RED tests verified GREEN after fix
- [x] Mechanism-GREEN tests pin internal contracts (cascade pollution
guard, SSE counter atomicity, init-script navigation re-entry, etc.)
- [x] cr-loop converged across 10 rounds (25→6→6→6→4→2→0 bucket-(a)
findings) + Procedure 3 audit clean
- [x] `oxfmt --check` clean across all 31 changed TS files
- [x] `tsc -p tsconfig.build.json` clean
- [ ] CI green on this PR
- [ ] Manual smoke: `bin/showcase test --d5
langgraph-python:agentic-chat` post-merge
## Follow-up debt (deferred from cr-loop; not subject-scope)
Detected by cr-loop reviewers but out of this PR's scope:
- `d6-all-pills.ts` deploy-churn NSF aggregation — features in
`notSupportedFeatures` lose `incapable` classification during
deploy-churn skip
- `d6-all-pills.ts` `joinAimockJournal` slug fallback under D6
concurrency — can pick wrong-feature entry when aimock doesn't echo
`x-diag-run-id`
- `d6-all-pills.ts` `BrowserDisconnectedError` sentinel is constructed
but never matched in `runFeature` catch — gets bucketed as
`driver-error`
- `d5-tool-rendering-default-catchall.ts` Path B narration
false-positive (substring matches in assistant prose)
- HITL registry side-effect tests use manual fallback registration that
masks real registration failures
- `installPrePaintFromEnv` and `attachSseInterceptor` are re-registered
on every `page.goto` — accumulate init scripts; harmless today via
in-script idempotency guards
🤖 Generated with [Claude Code](https://claude.com/claude-code)
|
||
|
|
1ef6a0960b |
test(harness): update d5-* + d6-all-pills unit tests to new probe + cascade contract
- d5-* unit tests: update page fakes / assertion-call shapes to pass
the ctx bridge ({bubbleIndex, text}) instead of querying the live
bubble list
- d6-all-pills.test.ts: update page-fake dispatch + cascade-state
expectations to match the atomic readCascadeState return shape
({count, text} or null) — picks up the cross-tier consistency
guarantee in the unit layer
|
||
|
|
cd2af0f9fd |
test(harness): conversation-runner unit coverage — banner debounce, cold-start retry, skipFill, error-banner shapes
Comprehensive unit coverage for: - preFill + fillAndVerifySend retry behavior - baseline banner snapshot + differs-from-baseline debounce matrix (stable banner = no fail; changed banner = fast-fail after 2 polls) - cold-start retry cases (first-attempt banner → reload + retry; retry-then-banner → translatedErr; retry exhausts deadline) - skipFill / skipSend short-circuit paths - ErrorBannerReadResult union shape variants (present/absent/unknown) |
||
|
|
3d49d4c38f |
test(harness): bubble-race integration tests + defect reproductions
- bubble-race-repro.ts: shared driver harness for repro tests
- bubble-race-repro-defect-{1,2,3,4}.test.ts: real-browser RED tests
that fail without the fix and pass with it
- Defect 1: fast-replay false-positive settle on count
- Defect 2: multi-turn flicker via list[last]
- Defect 3: cascade-tier blindness
- Defect 4: boot-time baseline staleness / pre-paint placeholder
- bubble-race-mechanisms.test.ts: GREEN tests pinning internal
contracts (cascade-pollution guard, atomic readCascadeState,
init-script idempotency)
- wait-for-turn-complete.test.ts: 3-conjunct classification matrix
- probe-contract.test.ts: ctx-bridge shape
- sse-counter.test.ts: counter increment + framenavigated reset +
handle-cache invalidation
|
||
|
|
cf7efbda95 |
fix(harness): d6 driver — installPrePaintFromEnv + attachSseInterceptor wiring + helper-based captureDiagnostics
- d6-all-pills.ts: wire installPrePaintFromEnv and attachSseInterceptor PRE-goto in both defaultLauncher and pooled launcher paths, so first-paint state is deterministic and the SSE counter is armed before any page navigation - captureDiagnostics rewired through the Node-side helper (readCascadeState) — single page.evaluate per call, no within-snapshot tier drift - Consume BUBBLE_RACE_MESSAGES override for defect-4 repro flows |
||
|
|
38e155e141 |
refactor(harness): probe-contract ctx bridge — assertions(page, {bubbleIndex, text}) for d5 scripts
- d5-gen-ui-custom, d5-gen-ui-open-advanced, d5-mcp-apps,
d5-subagents, d5-tool-rendering-default-catchall: wrap assertion
bodies in ctx bridge so they receive {bubbleIndex, text} for the
turn under test, replacing turn-leaky list[last] retrieval
- _gen-ui-shared: add readAssistantTextAt(page, bubbleIndex)
adapter; remove readLastAssistantText (no longer turn-scoped)
This is the read-side counterpart to the atomic readCascadeState
on the helper side: probes ask for the bubble at the runner's
chosen index instead of guessing.
|
||
|
|
3f3986b00b |
refactor(harness): waitForTurnComplete 3-conjunct primitive + cold-start retry + banner debounce
- waitForTurnComplete: 3-conjunct settle (SSE counter + DOM-at-index + text-quiet for settleMs) replaces count-only false-positive path - Cold-start retry: single bounded retry on first-attempt banner; reload + re-resolve chat input + settleMs/POLL_INTERVAL_MS floor honoring shared turn deadline; translatedErr throw on exhaustion - Banner fast-fail: baseline banner text snapshot + differs-from-baseline 2-poll debounce; in-poll BannerVisibleError translates to AssistantErroredError - New error classes: BannerVisibleError, AssistantErroredError - readErrorBanner returns 3-state union via ErrorBannerReadResult discriminated union - chatInputSelector cascade for input resolution across integrations |
||
|
|
473a9539a0 |
feat(harness): add bubble-race helpers — assistant-message-count, sse-interceptor, init-scripts
- assistant-message-count: 4-tier cascade + atomic readCascadeState
returning {count, text} from a single page.evaluate (prevents
within-poll cross-tier inconsistency); null fallback for
cascade-pollution
- sse-interceptor: __hk_runsFinished CDP counter + page-side fetch
wrapper; idempotent attach/detach with handle-cache invalidation
on stop and framenavigated reset (cold-start retry compatible)
- init-scripts: pre-paint placeholder env adapter
(installPrePaintFromEnv), strip-selector helper, and
BUBBLE_RACE_MESSAGES override consumer for defect-4 repro
infrastructure
|
||
|
|
6fbd66fa83 |
fix(showcase): mirror useInterrupt RESUME-PATH contract in 13 demo-local hooks
Each integration's interrupt-headless demo defines a local useHeadlessInterrupt
hook around the framework useInterrupt. Slot-2 originally identified 8
quarantined integrations (claude-sdk-typescript, langgraph-{fastapi,python,
typescript}, langroid, pydantic-ai, spring-ai, strands); review-round
follow-ups extended the sweep to llamaindex, mastra, ag2, agno, and
crewai-crews (5 more integrations sharing the same byte-identical hook).
The demo-local resolve() previously fire-and-forgot copilotkit.runAgent(...)
via `void runAgent(...).catch(() => {})`. Mirroring the framework fix:
- Make resolve async, return await copilotkit.runAgent(...).
- Use a pendingRef so resolve has stable identity (drop pending from
useMemo deps).
- Type signature: resolve: (response: unknown) => Promise<unknown>.
- Wrap in try/catch + setPending(null) + console.error + rethrow,
symmetric with the framework hook.
- onRunFailed also setPending(null).
13 integrations patched byte-identically.
|
||
|
|
7b80590f67 |
fix(react-core): await runAgent in useInterrupt::resolve
resolve() previously called copilotkit.runAgent(...) without await and without return, so callers had no handle to sequence against the resume run's settle. The harness DOM-settle check timed out for any consumer awaiting the assistant confirmation bubble. Changes: - Make resolve async, return await copilotkit.runAgent(...) so callers receive a Promise that settles when the resume run settles. - Update InterruptHandlerProps / InterruptRenderProps resolve return type from () => void to () => Promise<RunAgentResult>. - Wrap runAgent in try/catch + setPendingEvent(null) + rethrow, so rejection clears the popup AND propagates to awaiting callers (mirrors onRunFailed handler symmetry; closes the case where runAgent rejects before any run-failed event fires, e.g. network error pre-RUN_STARTED). - onRunFailed now also setPendingEvent(null) symmetric with onRunStartedEvent. - Regression tests: RESUME-PATH asserts resolve() returns a Promise that settles 1:1 with runAgent; RESUME-PATH-REJECT asserts rejection propagates, popup clears, console.error logs. |
||
|
|
c81b361f15 |
fix(showcase/harness): register /webhooks/deploy on CP path + fail-loud on missing SHARED_SECRET (#5458)
## Problem
The \`Showcase: Verify Deploy\` workflow runs \`notify-harness\` after
every merge to \`main\` and POSTs an HMAC-signed payload to
\`\${SHOWCASE_HARNESS_URL}/webhooks/deploy\`. That POST has been
returning **404** on every main deploy since at least 2026-06-12 (5+
consecutive failures), so the harness dashboard has not reflected any
recent deploys.
**Root cause:**
1. \`POST /webhooks/deploy\` is only registered when
\`webhookSecrets.length > 0\` — gate at
\`showcase/harness/src/http/server.ts:119\`.
2. The CP \`buildServer\` call in \`runControlPlane\`
(\`showcase/harness/src/orchestrator.ts\` ~line 2961 pre-fix)
**omitted** both \`webhookSecrets\` and \`metrics\`. The public Railway
host running the CP role never mounted the route — every notify-harness
POST returned 404.
3. The FATAL-CONFIG guard on missing \`SHARED_SECRET\` (pre-fix
\`orchestrator.ts:782-790\`) only fired when \`NODE_ENV ===
"production"\`. Any deploy whose NODE_ENV was unset, set to something
other than the literal \`"production"\`, or set after the check,
silently booted with \`webhookSecrets=[]\` and no fatal error.
## Fix
**Part 1: Register the route on the CP path.**
- Lifted the env→secrets loader into a new exported helper
\`loadWebhookSecrets()\` (orchestrator.ts ~175-209) so both boot paths
consume the same predicate.
- Worker \`boot()\` (~line 836) calls it; CP \`runControlPlane\` (~line
2421) now also calls it.
- The CP \`buildServer\` call (~line 3047) now passes \`webhookSecrets\`
and \`metrics\` alongside the existing \`bus\`/\`probes\`/\`fleetRuns\`
wiring — same shape as the worker call.
**Part 2: Tighten the FATAL-CONFIG predicate.**
- New predicate (inside \`loadWebhookSecrets\`): throw unless EITHER a
secret is set OR \`NODE_ENV === "test"\` OR \`HARNESS_ALLOW_NO_SECRET
=== "1"\` (narrow escape hatch for local dev).
- Thrown error message names the env-vars, the gate location
(\`src/http/server.ts:119\`), the silent-404 symptom, and the escape
hatches — so an operator reading the deploy log can recover without
spelunking.
## Test coverage
Four new red-green tests in \`orchestrator.test.ts\` under \`B2:
/webhooks/deploy registered on CP + fail-loud on missing
SHARED_SECRET\`:
- **Test A** — CP boot registers \`POST /webhooks/deploy\` when
\`SHARED_SECRET\` is set. Probe with no HMAC headers: pre-fix returned
404, post-fix returns 401 (route mounted, HMAC reject path fires).
RED→GREEN.
- **Test B** — Worker boot still registers the route (regression guard).
Pre-fix and post-fix both 401.
- **Test C** — FATAL-CONFIG fires when \`SHARED_SECRET\` unset AND
\`NODE_ENV=development\`. Pre-fix resolved silently, post-fix rejects
with \`/FATAL-CONFIG.*SHARED_SECRET/\`. RED→GREEN.
- **Test D** — Boot succeeds when \`SHARED_SECRET\` unset AND
\`NODE_ENV=test\` (escape hatch). Pre-fix and post-fix both green.
Existing R5-G4 D5 test (orchestrator.test.ts:1592) updated to match the
new error message via \`/FATAL-CONFIG.*SHARED_SECRET/\` — same throw,
broader match.
## Verification
- Full harness suite: **2719/2719 green** across 128 files
- \`tsc --noEmit\`: clean
- RED capture pre-fix:
- Test A: \`AssertionError: expected 404 to be 401\`
- Test C: \`AssertionError: promise resolved "{ port: ..., bus: { ... },
... }" instead of rejecting\`
## Deploy note
Requires a CP-role redeploy on the public Railway host for the new route
registration to take effect — once redeployed, the next
\`notify-harness\` POST will land on the harness dashboard.
## Test plan
- [x] B2 Tests A-D pass post-fix (red→green captured)
- [x] Full harness vitest suite green (2719/2719)
- [x] \`tsc --noEmit\` clean
- [ ] Post-deploy: confirm next \`Showcase: Verify Deploy\` run shows a
2xx notify-harness POST (not 404)
- [ ] Post-deploy: confirm harness dashboard reflects the next deploy
|
||
|
|
3ce3f5394d |
fix(showcase/harness): hoist OPS_TRIGGER_TOKEN fail-loud + emit on subscribeDeployResults sync throw (R3 cleanups)
R3-F1 (OPS_TRIGGER_TOKEN hoist): Extract loadOpsTriggerToken() and hoist its invocation to the top of both boot() and runControlPlane(), alongside loadWebhookSecrets() and loadPocketbaseUrl(). Pre-fix the empty/whitespace check fired AFTER pb / bus / scheduler / writer / S3 uploader allocations, so a typo'd `OPS_TRIGGER_TOKEN=` allocated expensive resources before throwing. Now all three fail-loud config predicates fire at the top, before any allocation needing teardown. Behaviour is unchanged for valid tokens (trimmed via R3-A.5 contract) and for the unset case (router omitted with info log). R3-F2 (sync-throw also emits deploy.writer.failed): subscribeDeployResults() pre-fix only emitted `deploy.writer.failed` on writer.write() promise rejection — a synchronous throw inside deployEventToProbeResult() (malformed event, type drift) bypassed the catch and we lost both the log AND the bus emit. Wrap the sync mapping in try/catch and mirror the rejection path so alert rules / metrics subscribers observe sync throws the same way they observe async write failures. Tests: - 5 unit tests for loadOpsTriggerToken: undefined / empty-string fail-loud / whitespace-only fail-loud / R3-A.5 trim contract / verbatim value. - 1 test for the sync-throw path: vi.doMock deployEventToProbeResult to throw, assert bus emits deploy.writer.failed with the err message AND writer.write was NOT called. Red-green verified locally (test fails without the try/catch). Suite: 2731 passing (+6 new). Typecheck clean. |
||
|
|
d5bdf2c9fd | fix(showcase/harness): track CP deploy.result unsubscribe + emit on write failure (R2 cleanups) | ||
|
|
b0ba8e869d |
fix(showcase/aimock): align catchall fixture shape for 6 remaining D5 reds (#5460)
## Summary Follow-up to #5459. That PR added the right `userMessage` matchers ("forecast for Tokyo" + "current price of AAPL") and flipped most integrations green, but 6 cells stayed red because the FIXTURE SHAPE on the matched fixtures was also broken. **Failing cells (post-#5459 deploy):** - `d5:agno/tool-rendering-custom-catchall` - `d5:llamaindex/tool-rendering-custom-catchall` - `d5:langroid/tool-rendering-custom-catchall` - `d5:claude-sdk-python/tool-rendering-custom-catchall` - `d5:strands/tool-rendering-custom-catchall` - `d5:strands/tool-rendering-default-catchall` ## Root cause The failing tool-emit fixtures all used `hasToolResult: false` (or `toolName: "<tool>"`) as their gate. The custom-catchall probe drives TWO sequential prompts (Tokyo, then AAPL). After Tokyo's tool result lands in the thread, aimock's `hasToolResult` is permanently true (`messages.some(m => m.role === 'tool')`), so `hasToolResult:false` can never match the AAPL turn → no fixture → 30s timeout. PocketBase confirms this exact mode — e.g. agno: ``` errorDesc: "timeout: assistant did not respond within 30000ms" failure_turn: 2 turns_completed: 1 / 2 ``` claude-sdk-python's `toolName` gate fails for an analogous reason on the second turn when the backend doesn't forward the tool definition uniformly. The same trap is documented in `showcase/aimock/d6/langgraph-python/tool-rendering.json`: > Gated on toolName:get_stock_price rather than hasToolResult:false. The D5 tool-rendering-custom-catchall probe runs 'weather in Tokyo' first, which leaves a get_weather tool result in the thread; the aimock router implements hasToolResult as messages.some(m=>m.role==='tool'), so hasToolResult is permanently true on the AAPL turn and a hasToolResult:false gate could never match (→ no_fixture_match → 503 → 30s timeout). ## Fix Align all 6 files to the canonical pattern used by working integrations (mastra, spring-ai, built-in-agent): 1. `toolCallId`-keyed narration fixture FIRST 2. tool-emit fixture SECOND, gated ONLY on `userMessage` + `context` (no `hasToolResult` / `toolName`) aimock's `toolCallId` matcher checks `messages[last].role === 'tool' && tool_call_id === ...`, which correctly fires only on the post-tool-result iteration regardless of older tool results in the thread. The tool-emit fixture below it always matches on turn-init (last message is user). Strands' two files additionally needed REORDERING — they had the tool-emit fixture before the toolCallId fixture, defeating first-match-wins. ## Scope discipline - Pure fixture-content alignment — no harness, probe, frontend, or backend changes. - Did NOT touch `showcase/harness/**` (peer session territory per `/tmp/coordinate/showcase-reds-coord/CHANNEL.md`). - Did NOT touch the two `d5-tool-rendering-*-catchall.ts` probes. ## Caveat (strands) During investigation, `showcase-strands-staging.up.railway.app` was returning 502 / page-load failures (infra, not fixture). Strands' fixture defects are still real and worth correcting, but the cells may stay red until the backend recovers. ## Test plan - [ ] CI green - [ ] Wait for `showcase_build.yml -f service=aimock` rebuild + Railway redeploy - [ ] Re-check PocketBase `state` for the 6 cells flips to `green` (excluding strands if backend stays down) - [ ] Confirm no regression on the other catchall cells already green post-#5459 |
||
|
|
df97ae8d2a |
fix(showcase/aimock): align catchall fixture shape for 6 remaining D5 reds
PR #5459 added userMessage matchers ("forecast for Tokyo" + "current price of AAPL") on the catchall fixtures, which flipped most integrations from red to green. Six cells stayed red on staging because the FIXTURE SHAPE itself was broken on the matched fixtures, not just the userMessage key. Failing cells (all custom-catchall except strands default): d5:agno/tool-rendering-custom-catchall d5:llamaindex/tool-rendering-custom-catchall d5:langroid/tool-rendering-custom-catchall d5:claude-sdk-python/tool-rendering-custom-catchall d5:strands/tool-rendering-custom-catchall d5:strands/tool-rendering-default-catchall PocketBase confirms the failure mode: turn 1 (Tokyo) completes; turn 2 (AAPL) times out at 30s or renders the wrong content. E.g. agno: errorDesc: "timeout: assistant did not respond within 30000ms" failure_turn: 2 turns_completed: 1 / 2 Root cause: the failing tool-emit fixtures used `hasToolResult: false` (or `toolName: "<tool>"`) as their gate. aimock's hasToolResult check is `messages.some(m => m.role === 'tool')` over the WHOLE thread — so once turn 1's Tokyo tool result lands in the conversation, hasToolResult is permanently true and `hasToolResult:false` can never match turn 2 → no fixture → 30s timeout. `toolName:get_stock_price` likewise fails when an integration backend doesn't forward the tool definition on turn 2. The same trap is documented in showcase/aimock/d6/langgraph-python/tool-rendering.json: "_comment": "Gated on toolName:get_stock_price rather than hasToolResult:false. The D5 tool-rendering-custom-catchall probe runs 'weather in Tokyo' first, which leaves a get_weather tool result in the thread; the aimock router implements hasToolResult as messages.some(m=>m.role==='tool'), so hasToolResult is permanently true on the AAPL turn and a hasToolResult:false gate could never match (→ no_fixture_match → 503 → 30s timeout)." Fix: align all 6 files to the canonical pattern used by mastra/spring- ai/built-in-agent on this probe: 1. toolCallId-keyed narration fixture FIRST 2. tool-emit fixture SECOND with ONLY `userMessage` + `context` (no hasToolResult / toolName gate) The toolCallId narration uses aimock's `messages[last].role === 'tool' && tool_call_id === ...` check, so it correctly wins on iteration 2 (post-tool-result) without being affected by older turns' tool results. The tool-emit fixture matches turn 1 (last message is user) and re-emits only when the narration above hasn't matched. Strands' two cells additionally needed REORDERING — they had the tool- emit fixture before the toolCallId fixture, defeating first-match-wins. Strands' staging backend was also returning 502 during testing; once it recovers, the corrected fixtures should let the probe pass. The fixture changes are necessary but may not be sufficient for strands if backend remains down. No probe-side, harness, or backend changes — pure fixture-content alignment. Six fixture files modified; line totals: -87 / +68. |
||
|
|
fadb29550a | Merge branch 'main' into docs/FAC-65-shared-v2-imports | ||
|
|
16a59a45c7 |
chore(showcase/harness): warn on webhook escape hatch + vi.stubEnv in B2 tests
CB-1 (Slot 2 #22): convert B2 deploy-webhook tests to vi.stubEnv with `vi.unstubAllEnvs()` in afterEach instead of manual process.env mutation + try/finally restoration. Test runner now guarantees restoration and the test bodies are shorter / less error-prone (CR Slot 2 #22). CB-2 (Slot 2 #28): when `loadWebhookSecrets`' escape hatch fires with a real-looking NODE_ENV (anything except "test"), log at `warn` instead of `info` so a production typo (NODE_ENV=staging + HARNESS_ALLOW_NO_SECRET=1) is visible in dashboards / log alerting. Pure local-dev (NODE_ENV=test) stays at info level so a normal unit-test boot doesn't spam warnings. CB-3 (Slot 4 #17): clarify `loadWebhookSecrets`' docstring + bypass-log message — "set" is ambiguous (empty string would qualify pre-fix); use "non-empty" to match the actual predicate. |
||
|
|
f42180f7bf |
fix(showcase/harness): hoist fail-loud config checks + symmetric POCKETBASE_URL predicate
R1-F2 (bucket b, defensive ordering): hoist fail-loud config validation to the TOP of boot() and runControlPlane() — BEFORE any pb client, queue, bus, scheduler, writer, S3 uploader, fleet-health, or aggregator allocations. Pre-fix `loadWebhookSecrets` lived AFTER the entire scheduler+writer (and CP queue+aggregator) setup, so a misconfigured boot allocated expensive resources and mounted file watchers before throwing. R1-F3 (bucket b, predicate symmetry): broaden the POCKETBASE_URL fail-loud predicate to match SHARED_SECRET semantics. Pre-fix the POCKETBASE_URL guard only fired on `NODE_ENV === "production"`, while `loadWebhookSecrets` fired unless NODE_ENV=test or an explicit escape hatch — so staging / unset / "development" deploys silently bound to http://localhost:8090. Extracted `loadPocketbaseUrl(logger)` with the same test-or-escape-hatch predicate as `loadWebhookSecrets`. A new `HARNESS_ALLOW_NO_PB_URL=1` env flag mirrors `HARNESS_ALLOW_NO_SECRET=1` for local dev. Worker `boot()` and the CP's `resolveFleetPbConfig` both call the helper. Tests: - HF13-A2 production-only assertion broadened (SHARED_SECRET set so the test specifically exercises the POCKETBASE_URL guard). - New R1-F3 tests cover NODE_ENV=development without POCKETBASE_URL (throws) and the HARNESS_ALLOW_NO_PB_URL=1 escape hatch (succeeds). |
||
|
|
fc828bb08c |
fix(showcase/harness): subscribe deploy.result in CP boot to deliver dashboard event
R1-F1 (bucket a, load-bearing): the CP boot path now subscribes to
`deploy.result` events through its status writer. Pre-fix, B2 mounted POST
/webhooks/deploy on the CP role but only the worker boot path had a
`bus.on("deploy.result", ...)` listener — signed POSTs to the CP host
returned 202, the bus event fired with no subscriber, and the deploy-overall
dashboard row never landed.
Extracted the handler into a shared `subscribeDeployResults(bus, writer)`
helper exported from orchestrator.ts so both boot paths share the IDENTICAL
logic. Worker boot keeps its inline pattern (busUnsubs.push) so teardown is
unchanged; CP keeps the subscription alive for the lifetime of the bus.
Red-green tests in src/orchestrator.test.ts under "R1-F1" assert that
emitting deploy.result on the returned handle's bus drives one write keyed
"deploy:overall" through a stubbed status writer — once for runControlPlane
(the new path) and once for boot() (regression guard for the helper
extraction).
|
||
|
|
bb23e5dbbf |
fix(showcase/aimock): add forecast-Tokyo + AAPL catchall fixtures to flip remaining D5 reds (#5459)
## Summary PR #5453 renamed stale catchall userMessage `"check Tokyo weather forecast"` → `"forecast for Tokyo"` to match the D5 probe input. But some integrations' catchall fixtures used **pill-aligned userMessages** (e.g. LG-TS, LG-FastAPI used `"Chain a few tools in this single turn"`, `"What's the weather in San Francisco?"`) — they had no `"forecast for Tokyo"` matcher to rename. Result: D5 catchall probe still misses → live LLM fallback → cells stay red on staging. ## Fix Add the canonical emit+narrate fixture pairs for `"forecast for Tokyo"` (default + custom catchall) and `"What's the current price of AAPL?"` (custom catchall only) — adapting mastra's working template — to each integration that was missing them. ## Scope 4 integrations × 1-2 files each (6 files total): - LG-TS, LG-FastAPI: both default + custom catchall - google-adk, ms-agent-dotnet: custom catchall only (their default is already green) Existing pill-aligned fixtures preserved (pure prepend at start of array). First-match-wins means new fixtures match the D5 probe inputs ("forecast for Tokyo", "What's the current price of AAPL?") without colliding with existing pill prompts. `toolCallId` values are unique per integration to avoid cross-integration shadowing. Note: the originally-flagged langroid, llamaindex, agno, claude-sdk-python, and strands files already have `forecast for Tokyo` + AAPL matchers in their catchall fixtures (likely from a prior pass) — they did not need changes and are not in this PR. If those staging cells are still red, the root cause is elsewhere (control-plane staleness, deploy lag, or different probe variant). ## Verification Post-merge, fleet-cp's e2e-deep probe will re-run within ≤6h staleness. Cells should flip d6:<slug>/tool-rendering-{default,custom}-catchall = green. |
||
|
|
07e4998798 |
fix(showcase/aimock): add forecast-Tokyo + AAPL catchall fixtures to flip remaining D5 reds
The catchall userMessage rename (#5453) only renamed STALE strings to 'forecast for Tokyo'. Integrations whose catchall fixtures used PILL-ALIGNED userMessages ('What's the weather in San Francisco?', 'Find flights from SFO to JFK.', 'Chain a few tools in this single turn', etc.) had no 'forecast for Tokyo' matcher to begin with — so the D5 catchall probe still fell through to live LLM and the cells stayed red on staging. This PR ADDS the canonical 'forecast for Tokyo' (default+custom catchall) and 'What's the current price of AAPL?' (custom catchall only) emit+narrate fixture pairs to each integration that was missing them. Existing pill-aligned fixtures are preserved (pure prepend at the start of the fixtures array; first-match-wins means new fixtures match the D5 probe inputs without colliding with existing pill prompts). Integrations touched: - langgraph-typescript: both default + custom - langgraph-fastapi: both default + custom - google-adk: custom only (default already green) - ms-agent-dotnet: custom only (default already green) After merge, fleet-cp's e2e-deep probe will re-run within ≤6h staleness window and flip cells GREEN. |
||
|
|
bfd1f12b00 | style: auto-fix formatting | ||
|
|
2939696415 |
fix(showcase/harness): register /webhooks/deploy on CP path + fail-loud on missing SHARED_SECRET
The 'Showcase: Verify Deploy' workflow's notify-harness step has been
POSTing to /webhooks/deploy and getting 404 on every main deploy since
at least 2026-06-12 (5+ consecutive failures). Root cause:
1. POST /webhooks/deploy is only registered when
webhookSecrets.length > 0 (gate at src/http/server.ts:119).
2. The CP buildServer call in runControlPlane omitted BOTH
webhookSecrets and metrics — so the public Railway host running
the CP role never mounted the route. The worker boot path had
them, but it isn't the host receiving notify-harness POSTs.
3. The FATAL-CONFIG guard on missing SHARED_SECRET was gated on
NODE_ENV === 'production' — any deploy with NODE_ENV unset or set
to something other than the literal 'production' silently shipped
with webhookSecrets=[] and no fatal error.
Part 1 — wire webhookSecrets + metrics into the CP buildServer call.
Lifted the env→secrets loader into a new exported helper
'loadWebhookSecrets()' (~orchestrator.ts:175-209) so BOTH boot paths
consume the same predicate. The worker 'boot()' call (~line 836) and
the CP 'runControlPlane' call (~line 2421) both invoke it, and the CP
buildServer call (~line 3047) now passes webhookSecrets+metrics
alongside the existing bus/probes/fleetRuns wiring.
Part 2 — tighten the FATAL-CONFIG predicate.
The new predicate throws unless EITHER a secret is set OR NODE_ENV ===
'test' OR HARNESS_ALLOW_NO_SECRET === '1' (narrow escape hatch for
local dev). The thrown error message names the env-vars, the gate
location ('src/http/server.ts:119'), the silent-404 symptom, and the
escape hatches, so an operator reading the deploy log can recover
without spelunking.
Updated the existing R5-G4 D5 regex (orchestrator.test.ts:1592) since
the error message changed; covered by the new B2 Tests A/C/D below.
Tests: 4 new red-green tests in 'B2: /webhooks/deploy registered on CP
+ fail-loud on missing SHARED_SECRET':
- Test A (CP /webhooks/deploy registration): RED 404 → GREEN 401.
- Test B (worker /webhooks/deploy regression): GREEN 401.
- Test C (FATAL-CONFIG when NODE_ENV=development): RED resolves →
GREEN rejects with /FATAL-CONFIG.*SHARED_SECRET/.
- Test D (NODE_ENV=test escape hatch): GREEN, no throw.
Full harness suite: 2719/2719 green. typecheck: clean.
Requires a CP-role redeploy on the public Railway host for the new
route registration to take effect.
|
||
|
|
fb3d5b3547 |
test(showcase/harness): cover CP webhook registration + fail-loud on missing SHARED_SECRET
Adds four red-green tests for the notify-harness 404 follow-up:
A. CP boot registers POST /webhooks/deploy when SHARED_SECRET is set
— RED pre-fix (404, route not mounted on CP); GREEN post-fix (401,
the route's own HMAC reject path fires on an unsigned POST).
B. Worker boot still registers POST /webhooks/deploy (regression
guard) — already green; pins the unchanged behavior.
C. boot throws FATAL-CONFIG when SHARED_SECRET unset AND NODE_ENV is
non-test — RED pre-fix (current guard fires only on
NODE_ENV='production'); GREEN post-fix once the predicate is
tightened.
D. boot succeeds when SHARED_SECRET unset AND NODE_ENV='test'
(test-mode escape hatch) — already green; pins the escape hatch.
Red capture (pre-fix run):
- Test A: AssertionError: expected 404 to be 401
- Test C: AssertionError: promise resolved instead of rejecting
|
||
|
|
61471cac93 | docs(showcase): fix shared randomUUID imports | ||
|
|
c43ed08e7b |
ci(e2e-dojo): run dojo suites on 4-vCPU runner (−40% wall-clock) (#5452)
Bumps the dojo e2e matrix from `depot-ubuntu-24.04` (2 vCPU) to `depot-ubuntu-24.04-4` (4 vCPU) and `NX_PARALLEL: 4` so the build uses the extra cores. This is the non-serializing way to cut dojo wall-clock (the build-once dedup tried in #5450 regressed wall-clock and was reverted). ## Result: −40% wall-clock (measured on CI) Dojo wall-clock = the single slowest suite (the 15 run in parallel). Comparison vs the 2-vCPU baseline: | metric | 2-vCPU baseline | 4-vCPU | Δ | |---|---|---|---| | **wall-clock** (long pole `langgraph-python`) | 623s (10.4m) | **373s (6.2m)** | **−40%** | | runner-minutes (wall summed, 15 suites) | 105m | 76m | −28% | | **billed compute** (vCPU-min; 4-vCPU ≈ 2× rate) | ~210 | ~304 | **+45%** | Every suite got faster; the long-pole suites benefited most: | suite | 2-vCPU | 4-vCPU | |---|---|---| | langgraph-python | 623s | 373s | | langgraph-typescript | 547s | 362s | | langgraph-fastapi | 500s | 337s | | adk-middleware | 414s | 286s | | (… all 15 faster …) | | | Long-pole `langgraph-python` step breakdown: | phase | 2-vCPU | 4-vCPU | |---|---|---| | Build cpk | 82s | 48s | | Prep dojo | 94s | 52s | | **Run tests (Playwright)** | **271s** | **117s** | | total | 623s | 373s | **Key finding:** the Playwright phase more than halved → the e2e suites are **CPU/worker-bound, not LLM-latency-bound**. A bigger runner is the right lever; test sharding is not needed to reach ~6 min. ## Trade-off −40% wall-clock for **~+45% billed compute** (4-vCPU costs ~2×/min, partly offset by finishing 28% sooner). If the cost bump isn't worth it across all 15 suites, a follow-up can scope `-4` to just the slow suites via a per-matrix `runner` field (wall ~6.5m, smaller cost increase). Companion to #5450 (unit-test `nx affected`). |
||
|
|
d4fa49eed4 |
fix(showcase): refresh every view after a chat-driven mutation
useCreditCards() kept per-instance React state and was called independently by the dashboard page and by copilot-context (where the chat's approve / finalize / open-exception tools live). A mutation made through one instance refetched only itself — so when the agent approved an over-limit charge in chat, the dashboard pending table kept showing it as pending until a manual reload. Add a module-level revalidation bus: each useCreditCards() instance registers a refetch callback, and every mutation calls notifyDataChanged() to fan a refetch out to all live instances. The dashboard now reflects agent-driven approvals immediately (verified: recall-approve in chat drops the charge from the pending table with no reload). |
||
|
|
0b65391501 |
docs(showcase): clarify same-thread (OSS) vs cross-thread (Intelligence) recall
Adds a "What each mode actually recalls" subsection to the demo README: in OSS mode the taught workflow is recalled only within the same conversation (the saved procedure is echoed back into that thread), so a brand-new chat won't know it — that's expected. Cross-conversation persistence is what the external Intelligence backend provides. Names the symptom explicitly so reviewers aren't surprised when a new conversation "doesn't know" the workflow in OSS mode. |
||
|
|
7e1641ff1b |
feat(showcase): pending-approval table redesign + approval-gate UX fixes
Rework the dashboard's Pending approval view and fix two approval-gate UX bugs on the banking demo (PR #5266): - Table layout: replace the center-stacked per-row card with a scannable table (Merchant / Amount / Policy / Actions). Actions are check / x icon buttons plus a "more actions" overflow menu holding File policy exception. Status is its own column (Over limit / Cleared / Within limit), and Approve is gated until the charge is actually clearable. - Fix the table shrinking when the more-actions menu opens: the Radix menu is modal by default and engaged react-remove-scroll, whose scrollbar compensation reflowed the table. Set modal={false} (a row menu needn't be modal) and add whitespace-nowrap to the status badges. - Fix "cannot approve after filing an exception": the inline card offered all codes, including non-justifying ones that set activeExceptionId (flipping the row to Cleared) but never lift the server gate, so the approve 422'd silently. The card exists to clear an over-limit charge, so it now offers only justifying codes. The gate itself is unchanged. Verified live in OSS dev: file (justifying) -> Cleared -> approve succeeds; the menu opens without reflow; lint + build green. |
||
|
|
4791c86923 |
fix(showcase/harness): override LOCAL_SERVICES_JSON in --isolate generator to target the requested slug (#5454)
## Summary
- `showcase/docker-compose.local.yml:262` hardcodes
`LOCAL_SERVICES_JSON` to `showcase-langgraph-python` (intentional N=1
demo default).
- `showcase/bin/showcase test <slug> --d6 --isolate` was inheriting that
value verbatim into the iso1 stack, so the iso1 control-plane discovered
`showcase-langgraph-python` instead of `showcase-<requested-slug>`.
- Result: iso1 probes targeted the wrong service. Visible in iso1
harness logs as `discovery.railway-services.local-injection count:1
names:["showcase-langgraph-python"]` regardless of CLI arg.
- This PR teaches the iso1 compose generator (`apply_isolation` in
`showcase/scripts/cli/_common.sh`) to inject a per-slug
`LOCAL_SERVICES_JSON` override built from the slug's manifest.yaml demos
list. Fallback to `["agentic-chat"]` if manifest absent.
## Scope
Two files, ~55 LOC added: `showcase/scripts/cli/cmd-test.sh` (pass slug
arg), `showcase/scripts/cli/_common.sh` (inject regex sub in python
rewriter). Bash + embedded python only; no TypeScript changed.
## Verification
- **Local `--isolate` discovery confirmed correct:**
`showcase/bin/showcase test ms-agent-python --d6 --isolate` — iso1
harness log now shows `discovery.railway-services.local-injection
names:["showcase-ms-agent-python"]` (was `showcase-langgraph-python`
before this PR).
- Persistent stack default behavior (langgraph-python N=1) preserved via
fallback when no slug arg provided.
- **Heredoc hardening scope:** commit
|
||
|
|
318bd9e45f |
fix(showcase/aimock): align D6 tool-rendering catchall userMessage to D5 probe input (#5453)
## Summary
- D5 e2e-deep probes for `tool-rendering-{default,custom}-catchall` send
`"forecast for Tokyo"` as the test input (see
`showcase/harness/src/probes/scripts/d5-tool-rendering-{default,custom}-catchall.ts`).
- aimock uses substring match on `userMessage`.
- The catchall fixtures on main had a stale `"check Tokyo weather
forecast"` string that couldn't substring-match the probe input →
fixture miss → probe falls through to live LLM → CV ✗ red D4 on the
dashboard.
- This PR renames the userMessage to the canonical `"forecast for
Tokyo"` across 28 catchall fixture files in 16 integrations.
## Scope
28 files × ~2 userMessage occurrences each = 67 line changes. **No code,
no agent, no page.tsx changes.** Pure fixture-data alignment.
Integrations covered (default-catchall and/or custom-catchall): ag2,
agno, built-in-agent, claude-sdk-python, claude-sdk-typescript,
crewai-crews, google-adk, langgraph-python, langroid, llamaindex,
mastra, ms-agent-dotnet, ms-agent-python, pydantic-ai, spring-ai,
strands.
## Commit history note
Commit
|
||
|
|
edc77f8090 |
fix(showcase/harness): pass slug via env var to python rewriter instead of bash interpolation
The python rewriter in apply_isolation previously interpolated $slug directly into the inline python source via bash. A slug containing a single quote would break the python literal. Internal-tool risk only (slug is developer-typed), but cheap to harden. Pass slug via SHOWCASE_ISO_SLUG env var and read os.environ.get(...) inside the python heredoc. Defense-in-depth; no behavior change for valid slugs. |
||
|
|
b1f19bdc80 |
fix(showcase): rename catchall userMessage to 'forecast for Tokyo' to avoid chain pill substring collision
The previous 'weather in Tokyo' rename (
|