Commit Graph

4085 Commits

Author SHA1 Message Date
Jordan Ritter 9491b89340 fix(showcase): disjoint catchall userMessages + content-asserting probes (supersedes #5465)
Closed PR #5465 introduced cross-fixture leakage by stripping the
toolName discriminator on shared {userMessage, context} keys: with
aimock's alphabetical first-match-wins ordering,
tool-rendering-custom-catchall.json sorts before
tool-rendering-default-catchall.json, so default-catchall page requests
were served the custom file's content. The probe was structurally blind
because it asserted only DOM testids (copilot-tool-render +
data-tool-name=get_weather) — both fixtures emit get_weather, so the
testid signal passed regardless of which fixture won.

Real LGP-gold pattern is disjoint userMessages between default and
custom catchall fixtures (NOT shared keys discriminated by toolName).
This PR ports the LGP-gold pattern to the other 17 integrations and
adds page-text content assertions to both d5 catchall probes so this
class of regression can't recur silently:

- Probe prompts disjoint between default-catchall and custom-catchall
- Fixture userMessages updated to match the new disjoint prompts
- Page-text content assertions in both d5 catchall probes
  (default negatively asserts the custom-catchall leak phrase;
   custom positively asserts it)

Local gold-standard red-green proof captured:
- Step A: live staging Playwright baseline (RED-OF-RECORD)
- Step B: local control-plane on origin/main reproduces structural
  fragility (probes pass at testid level, custom fixture wins on
  default's userMessage path)
- Step C: local control-plane on this branch — both catchall probes
  GREEN with the new content-asserting assertions
    tool-rendering-default-catchall: pass=true (5443ms)
    tool-rendering-custom-catchall:  pass=true (8956ms)
                                     cross-tool signature passed

LGP (langgraph-python) was already disjoint; it remains untouched.
2026-06-16 09:37:51 -07:00
Jordan Ritter f316d055fb refactor(showcase): rename internal Wired counters to BeAgent
Renames the internal display-counter variables that fed the
stats-bar / coverage-bar "Wired" labels to match the new "BE (Agent)"
taxonomy: totalWired → totalBeAgent in cells-view.tsx and
parity-view.tsx, pctWired → pctBeAgent in coverage-bar.tsx. These
are local render-time aggregates, not part of any persisted
contract.

The status enum literal "wired" (the .filter((c) => c.status ===
"wired") guard, the coverage-segment-wired test ID, and the
wiredByCategory Map's local name) is intentionally preserved — those
all key directly off the persisted catalog Status enum and changing
them would expand scope into the catalog contract.
2026-06-16 09:23:00 -07:00
Jordan Ritter bfebac6565 refactor(showcase): relabel catalog status Wired display → BE (Agent)
Unifies the catalog integration status display label with the same
"BE (Agent)" taxonomy as the live-probe agent dimension. The
underlying Status enum literal ("wired"), the catalog.metadata.wired
data field, and the filter chip id ("wired" → cell-matrix.tsx:311
status === "wired" filter key) are all PRESERVED — they are
persisted catalog data + filter state contracts that must not move.
Only the user-facing display label on stats-bar, adaptive-stats-bar,
and the filter chip flips.

This collapses the two distinct dashboard "Wired" surfaces (the L1
live-probe dimension and the per-cell build-state count) under one
unified label, matching the user-confirmed Path B.
2026-06-16 09:22:41 -07:00
Jordan Ritter ef40bd7cfb refactor(showcase): relabel live-probe agent dimension Wired → BE (Agent)
Unifies the L1 "agent" live-probe display label with the taxonomy
convention established by #5473 (UI (Frontend), E2E, CV, D6 — layer
descriptor in parentheses). The dimension name stays "agent" in code
(PocketBase row keys agent:<slug>, LiveDimension union, keyFor
lookups are all unchanged stable contracts) — only the visible label
changes.

Updates the level-strip L1 badge label and the packages-section
L1-L4 header legend (W → B, "Wired" → "BE (Agent)") so the
level-strip's ToneChip first-letter abbreviation matches the legend
key. Test assertions covering the rendered letter, the legend text,
and the degraded-tone test's local variable follow suit.
2026-06-16 09:22:31 -07:00
Jordan Ritter 72520de43d fix(harness): cascade last-bubble whole-bubble-minus-toolbar fallback for tool-only responses 2026-06-16 04:05:49 -07:00
Jordan Ritter f7779f90bf refactor(showcase): probe taxonomy cleanup — drop Smoke, rename E2E (Demo) → UI (Frontend) (#5473)
## Summary

The smoke probe was the same HTTP contract as `/health` on the same
service (200-OK JSON body), so every tick paid two HTTP calls for the
same liveness signal. This PR:

- Drops the `/smoke` GET probe and the `smoke:<slug>` primary
`ProbeResult`. The driver now emits `health:<slug>` as the primary and
`agent:<slug>` via writer side-emit (half the per-tick cost).
- Renames the dashboard's D3 / e2e row label from `E2E (Demo)` to `UI
(Frontend)` (and the short cell badge from `E2E` to `UI`) — the row was
already about "the demo page renders in a browser", not E2E in the
traditional sense.
- Drops the `Smoke` row from the cell drilldown (CellState.smoke field
is retained for back-compat so historical rows still parse).
- Adds an explicit `Health` row to the legend (was implicit).

The driver's registry `kind` stays `"smoke"` so existing YAML configs +
orchestrator family wiring keep routing to this driver — the emission
key (`health:<slug>`) is the taxonomy contract that matters. The
underlying probe key (`e2e:<slug>/<feature>`) is preserved on PocketBase
so historical rows render correctly during the rename window.

## RED proof (BEFORE)

`smoke:<slug>` primary emission in `liveness.ts` (origin/main):

```
18: *   1. `smoke:<slug>`  — the RETURN VALUE of `run()`. The invoker runs
56: * `smoke:<slug>`/`health:<slug>`/`agent:<slug>` triple the dashboard
241:        input.mode === "discovery" ? `smoke:${slug}` : input.key;
```

`E2E (Demo)` / `Smoke` strings in dashboard render code (origin/main):

```
cell-drilldown.tsx:38: * Trip)" (chat+tools round-trip), D3/e2e = "E2E (Demo)" (the demo page loads
cell-drilldown.tsx:49:  { key: "e2e", label: "E2E (Demo)" },
cell-drilldown.tsx:52:  { key: "smoke", label: "Smoke" },
cell-pieces.tsx:382:        name="E2E"
unified-cell.tsx:231:      <TestBadge name="E2E" level={model.d3} />
adaptive-legend.tsx:59:        E2E (Demo): demo page loads and round-trips in a browser
```

## GREEN proof (AFTER)

Primary emission rewritten — driver returns `health:<slug>` and no
`smoke:<slug>` ProbeResult is ever emitted:

```
33: * to `/smoke` and emitted a `smoke:<slug>` primary result, but that probe
230:  // `kind: "smoke"` is the registry/family identifier the orchestrator
236:  kind: "smoke",        // registry-key back-compat; no smoke ProbeResult is emitted
367:  if (input.mode === "discovery") return `health:${slug}`;
368:  if (input.key.startsWith("smoke:")) {
369:    return `health:${input.key.slice("smoke:".length)}`;
```

New labels in dashboard:

```
cell-drilldown.tsx:60:  { key: "e2e", label: "UI (Frontend)" },
cell-pieces.tsx:382:        name="UI"
unified-cell.tsx:231:      <TestBadge name="UI" level={model.d3} />
adaptive-legend.tsx:65:        UI (Frontend): demo page renders in browser (Playwright)
```

Confirmation that the old labels are gone from render code:

```
$ grep -rn 'E2E (Demo)\|"Smoke"' cell-drilldown.tsx cell-pieces.tsx unified-cell.tsx adaptive-legend.tsx
(no matches — labels removed)
```

## Test results

- `pnpm exec vitest run` in `showcase/shell-dashboard`: **63 files, 1089
tests passed, 1 skipped**
- `pnpm exec vitest run` in `showcase/harness`: **129 files, 2760 tests
passed**
- `pnpm exec tsc --noEmit` in both: clean

Liveness driver tests include a regression guard (`never emits a
smoke:<slug> key`) so the smoke contract cannot be re-introduced
silently.

## Test plan

- [x] Dashboard vitest green (1089/1089)
- [x] Harness vitest green (2760/2760)
- [x] tsc --noEmit green in both packages
- [x] Visual cross-check: drilldown renders 6 rows (no Smoke), e2e row
labelled `UI (Frontend)`; cell strip shows `UI` badge in place of `E2E`;
legend shows explicit `Health` row + `UI (Frontend)` D3 line
2026-06-16 01:57:15 -07:00
Jordan Ritter 53f021afe8 refactor(showcase): probe taxonomy cleanup — drop Smoke, E2E (Demo) → UI (Frontend)
The smoke probe was the same HTTP contract as /health on the same
service (200-OK JSON body), so every tick paid two HTTP calls for the
same liveness signal. Drop the /smoke GET + the smoke:<slug> primary
ProbeResult; the driver now emits health:<slug> as the primary and
agent:<slug> via writer side-emit (half the per-tick cost).

The driver's registry kind stays "smoke" so existing YAML configs and
orchestrator family wiring keep routing to this driver — the
emission key (health:<slug>) is the taxonomy contract that matters.

Dashboard:
- Drilldown: D3/e2e row labelled "UI (Frontend)" (was "E2E (Demo)");
  Smoke row dropped (CellState.smoke field retained for back-compat).
- Cell badges: short label "UI" (was "E2E") in cell-pieces + unified-cell.
- Legend: D3 = "UI (Frontend): demo page renders in browser (Playwright)"
  plus an explicit Health row at the top.

Tests updated to assert new labels:
- cell-drilldown.test.tsx: 6 dimensions (no Smoke); UI (Frontend) label.
- cell-drilldown.lazy-signal.test.tsx: drilldown-badge-ui--frontend- testid.
- cell-pieces.test.tsx + .signal-degrade.test.tsx: badge name "UI".
- unified-cell.test.tsx: mock-badge-UI.
- overlay-selector-integration.test.tsx: "UI" in place of "E2E".
- dashboard-color-matrix.test.tsx: badge: "UI" case names.
- liveness.test.ts: two-call contract (health + agent), regression guard
  asserting no smoke:<slug> ever emitted.

Underlying probe key (e2e:<slug>/<feature>) preserved on PocketBase so
historical rows render correctly during the rename window.
2026-06-16 01:48:41 -07:00
Jordan Ritter 665dd5aae6 fix(harness): waitForTurnComplete reads LAST bubble (multi-step agent regression from #5462) (#5472)
## Summary

**URGENT hotfix.** Staging is widely red for D5/D6 multi-step demos.
Root cause confirmed via live Playwright DOM inspection on staging
(langgraph-typescript : beautiful-chat : pie chart prompt).

PR #5462's defect-2 fix uses `bubbleIndex = turnIndex - 1`, which
assumes one assistant bubble per turn. Multi-step agents (LangGraph,
Mastra, CrewAI, llama-index, ag2, etc.) emit 2-3+ bubbles per turn —
tool-call bubble + tool-render bubble + final-text bubble. The
strict-index read lands on an intermediate **tool-call** bubble whose
scoped-text selectors (`.cpk:prose`, `[data-message-content]`, `.prose`,
`p`) are **empty** (tool-call content lives in a sibling card div
outside the scoped cascade). The cascade-pollution guard correctly
returns `null` → the text-stable conjunct never converges →
`reason=text-unstable` timeout across virtually every multi-step demo
(beautiful-chat-*, gen-ui-*, shared-state-*, agentic-chat on most
backends).

## Live RED-GREEN proof (staging)

Standalone Playwright probe against
`showcase-langgraph-typescript-staging.up.railway.app/demos/beautiful-chat`,
prompt = `"pie chart of revenue distribution by category from the sample
sales data"`, captured BOTH the OLD-logic read (`bubbleIndex = turnIndex
- 1 = 0`) AND the NEW-logic read (`bubbleIndex = count - 1 = 2`) in ONE
atomic page evaluate after the turn settled:

```
TOTAL BUBBLE COUNT FOR TURN 1: 3  (true multi-step — tool-call + tool-render + final-text)
FINAL TIER: [data-testid="copilot-assistant-message"]

OLD LOGIC (bubbleIndex = turnIndex - 1 = 0):
  count = 3
  text  = null
  reason = all scoped selectors empty           ← would time out: reason=text-unstable

NEW LOGIC (bubbleIndex = count - 1 = 2):
  count = 3
  text  = "Here is the pie chart showing the revenue distribution by category from the sample sales data."
  textLen = 94                                  ← gate settles cleanly
```

RED ≠ GREEN; new logic returns 94 chars of stable final-text where old
logic returns `null`. The bug and the fix are proven on real staging
DOM.

### Live DOM evidence
LangGraph TS beautiful-chat, prompt = "pie chart of revenue distribution
by category…":
```
count = 3 bubbles in [data-testid="copilot-assistant-message"]
bubble[0]: tool-call "query_data"   — `.cpk:prose` exists but EMPTY (content in sibling <div class="my-1.5">)
bubble[1]: pie-chart card           — `.cpk:prose` exists but EMPTY (content in sibling styled card div)
bubble[2]: final text "Here is the pie chart..." — `.cpk:prose` has 94 chars + toolbar mounted
```

The old pre-#5462 runner read `list[last]` (always the last bubble —
always bubble[2] above). The PR moved to `turnIndex - 1`, which is
bubble[0] (the tool-call) for any multi-step turn → broken.

## Fix

Read the LAST bubble in the matched cascade tier (`count - 1`) instead
of a fixed strict index. Defect-2 protection (don't read a leftover
bubble from a prior turn) is preserved via an explicit **pre-submit
baselineCount sentinel snapshot** in the runner — the gate now requires
`count > baselineCount` rather than `count > turnIndex - 1`. This
correctly rejects stale prior-turn bubbles **without assuming 1 bubble =
1 turn**.

### Changes
- `assistant-message-count.ts`: add `readCascadeStateLast(page)` — same
atomic single-evaluate contract as `readCascadeState`, but resolves the
bubble index internally as `count - 1`.
- `conversation-runner.ts`:
- Snapshot `baselineCount = countAssistantMessages(page)` BEFORE
`sendTurnMessage()` each turn.
- Pass through to `waitForTurnComplete` via new `baselineCount` option
(defaults to `turnIndex - 1` to preserve unit-test fake compatibility).
- `waitForTurnComplete` reads via `readCascadeStateLast` and gates on
`domOk = count > baselineCount`.
  - Cold-start retry re-snapshots `baselineCount` after `page.reload()`.
- Returned `bubbleIndex` in the success ctx = `count - 1` (matches the
bubble whose text settled the gate).
- `dom-missing` reason classifier updated to `countFinal <=
baselineCount`.

## Why this doesn't regress defect-2

The original defect-2 (un-turn-scoped bubble selection) leaked a later
turn's bubble into THIS turn's assertions because
`readLastAssistantText` read `list[last]` GLOBALLY with no per-turn
baseline. The new fix re-introduces the "read last" semantics but adds
the **per-turn pre-submit baseline count snapshot** — the gate only
advances once a NEW bubble has appeared (`count > baselineCount`), so we
cannot read a leftover from a prior turn, and the "last" we return is
the last NEW bubble (which for multi-step turns is the final-text bubble
whose scoped text actually settles).

## Test plan

- [x] **Live RED-GREEN staging probe** (see above) — confirms OLD logic
= null and NEW logic = 94-char final-text on a real 3-bubble multi-step
turn.
- [x] All 72 `conversation-runner.test.ts` +
`assistant-message-count.test.ts` unit tests pass (default
`baselineCount = turnIndex - 1` keeps existing count-progression scripts
working).
- [x] Full harness vitest suite: 2758/2758 green across 129 test files.
- [x] `tsc --noEmit` clean on `showcase/harness`.
- [x] `oxfmt --check` clean on changed files.
- [x] `oxlint` clean on changed files.
- [ ] Staging D5/D6 sweep after merge confirms multi-step demos
(beautiful-chat, gen-ui, shared-state, langgraph/mastra/crewai
agentic-chat) recover from `text-unstable` red.

## Scope

Strictly the two files above. No unrelated changes. No refactors.
2026-06-16 01:12:10 -07:00
Jordan Ritter d145bb1708 fix(harness): waitForTurnComplete reads LAST bubble in matched tier
PR #5462's defect-2 fix assumed 1 bubble per turn (bubbleIndex =
turnIndex - 1). Multi-step agents (LangGraph, Mastra, CrewAI) emit
2-3+ bubbles per turn — tool-call + tool-render + final-text — so
the strict-index read lands on an intermediate tool-call bubble
whose scoped-text selectors (`.cpk:prose`, `[data-message-content]`,
`p`) are EMPTY (its content lives in a sibling card div outside the
scoped cascade). The cascade-pollution guard returns null forever
and the text-stable conjunct times out with `reason=text-unstable`
across virtually every multi-step demo (beautiful-chat-*, gen-ui-*,
shared-state-*, agentic-chat on LangGraph/Mastra/CrewAI/etc.).

Live DOM evidence from staging (langgraph-typescript :
beautiful-chat : pie chart prompt):
  count = 3 bubbles
  bubble[0]: tool-call "query_data" — `.cpk:prose` empty
  bubble[1]: pie chart card — `.cpk:prose` empty
  bubble[2]: final text "Here is the pie chart..." — 90 chars

Fix: read the LAST bubble in the matched cascade tier (count - 1)
instead of a fixed strict index. Defect-2 protection (don't read a
leftover bubble from a prior turn) is preserved via an explicit
pre-submit `baselineCount` sentinel snapshot in the runner; the
gate now requires `count > baselineCount` rather than
`count > turnIndex - 1`, which correctly rejects stale bubbles
without assuming 1 bubble = 1 turn.

Changes:
  - assistant-message-count.ts: add `readCascadeStateLast(page)` —
    same atomic single-evaluate contract as `readCascadeState` but
    resolves the bubble index internally as `count - 1`.
  - conversation-runner.ts: snapshot `baselineCount` before
    sendTurnMessage; pass through to `waitForTurnComplete` via new
    `baselineCount` option (defaults to `turnIndex - 1` for unit-
    test fakes); `waitForTurnComplete` reads via
    `readCascadeStateLast` and gates on `count > baselineCount`.
    Cold-start retry re-snapshots `baselineCount` after page.reload.

Test plan:
  - All 72 conversation-runner + assistant-message-count unit tests
    pass (default baselineCount preserves count-progression scripts).
  - Full harness suite 2758/2758 green, tsc clean, oxfmt/oxlint
    clean on changed files.
2026-06-16 01:02:52 -07:00
Jordan Ritter 48ef506abf fix(showcase): allow harness to boot locally without SHARED_SECRET
PR #5458 (c81b361f1) added a fail-loud gate that refuses harness boot
in any deployable mode (NODE_ENV != "test") without SHARED_SECRET or
SHARED_SECRET_PREV: POST /webhooks/deploy is only registered when
webhookSecrets.length > 0 (src/http/server.ts:119 +
loadWebhookSecrets in src/orchestrator.ts). The local docker-compose
stack inherits NODE_ENV=production from the harness image and does
not (and should not) set SHARED_SECRET, so every local D5/D6 verify
run via bin/showcase test --d5/--d6 was crashing the harness in a
restart loop with FATAL-CONFIG.

Fix: add HARNESS_ALLOW_NO_SECRET=1 (the documented local-dev escape
hatch — explicitly referenced in the FATAL-CONFIG message itself) to
both harness services in showcase/docker-compose.local.yml. Inline
comments explain the rationale and pin the relevant source locations.

Prod impact: NONE. Railway sets SHARED_SECRET explicitly via env on
every harness service, so loadWebhookSecrets sees a real secret,
registers POST /webhooks/deploy, and never reads
HARNESS_ALLOW_NO_SECRET. This change only affects the local
docker-compose stack.

Verified locally: showcase-iso1-harness boots cleanly (Up healthy),
the expected warn-level webhook-auth-bypass log fires
(escapeHatch:true), worker registers, scheduler starts, and a real
d6:langgraph-typescript job claims successfully — confirming the
gate fires only in deployable contexts.
2026-06-16 01:00:22 -07:00
Jordan Ritter f8b14f1d6a fix(react-core): await runAgent in useInterrupt::resolve and propagate to demo-local hooks (#5461)
## Summary

Convert `useInterrupt::resolve` and the demo-local
`useHeadlessInterrupt::resolve` callbacks from fire-and-forget into
`async` + `return await copilotkit.runAgent(...)`, so callers receive a
Promise that settles when the resume run settles.

## Mechanism

Pre-fix: `resolve(...)` called `copilotkit.runAgent({...})` without
`await` and without `return` — the arrow returned `undefined`. The
showcase harness DOM-settle check on the assistant confirmation bubble
timed out because consumers had no handle to sequence against the resume
run's settle.

Post-fix:
- Framework `useInterrupt::resolve` is `async`, `return await`s
`runAgent`, wraps in try/catch + `setPendingEvent(null)` +
`console.error` + rethrow on rejection.
- `onRunFailed` now also `setPendingEvent(null)` — symmetric with
`onRunStartedEvent`, prevents stuck popups on run failure.
- `InterruptHandlerProps.resolve` / `InterruptRenderProps.resolve` typed
`() => Promise<RunAgentResult>` (was `() => void`).
- The same fix-pattern applied to 13 demo-local `useHeadlessInterrupt`
implementations across integrations: ag2, agno, claude-sdk-typescript,
crewai-crews, langgraph-{fastapi,python,typescript}, langroid,
llamaindex, mastra, pydantic-ai, spring-ai, strands.

## Verification

- **Unit:** new tests `resolve returns a Promise that settles when
runAgent settles (RESUME-PATH)` + `resolve rejects when runAgent
rejects, logs the failure, and clears pending (RESUME-PATH-REJECT)`.
Red-then-green on baseline. 22/22 react-core hook tests passing.
- **Backplane (end-to-end):** showcase control plane on langgraph-python
via `./bin/showcase test langgraph-python:gen-ui-interrupt --d5 --direct
--verbose` (GREEN, 70.1s, was RED-RESUME-PATH baseline) and
`:interrupt-headless --d5 --direct --verbose` (GREEN, 7.7s, was
RED-RESUME-PATH baseline). Built Docker artifact + aimock D6 fixture
replay — staging-equivalent path.

## CR

4 review rounds × 7 unbiased agents = 28 agent-runs. 2 Procedure 3
audits (3 promotions in round 1 → Fix C + D; 0 promotions in round 2).
Convergence achieved per agent-confirmed `bucket (a) empty` + Procedure
3 zero `PROMOTE_TO_A`.

## Out of scope (deferred follow-ups)

- **No manifest unquarantine.** A prior version of this branch flipped
`interrupt-headless` and `gen-ui-interrupt` from
`not_supported_features` to `features` across 12 manifests, but Phase 4
backplane validation across 8 PATCHED integrations × 2 demos surfaced
(a) `langgraph-fastapi:gen-ui-interrupt` still RED-RESUME-PATH because
the framework patch doesn't propagate from workspace to the
integration's Docker image without a published `@copilotkit/react-core`
release, and (b) a separate popup-mount RED-RUNTIME-OTHER cluster across
6 integrations that's pre-existing rot unrelated to this PR. Manifest
unquarantine + `@copilotkit/react-core` version bump are deferred to a
follow-up `copilotkit-release` PR after this fix lands.
- **No version bump.** Release follows in a separate PR.
- **Bucket (c) follow-up backlog** (audit verdicts: 11 STAY_IN_C):
  - JSON.parse guarding in showcase demos (pre-existing, all 13)
- Type-design quirks in `InterruptHandlerProps.result` (non-nullable in
type, nullable in runtime)
  - `renderInChat` dynamic toggle leak (pre-existing edge case)
- JSDoc clarity ("symmetric with onRunFailed" wording, `@typeParam
TRenderInChat`, `event.value` "any" → "unknown")
  - StrictMode test coverage gap
- 13× duplicate `useHeadlessInterrupt` → bucket (d): consolidate into a
shared module in a future PR

## Files changed

16 files, +486/-154:
- `packages/react-core/src/v2/hooks/use-interrupt.tsx`
- `packages/react-core/src/v2/hooks/__tests__/use-interrupt.test.tsx`
- `packages/react-core/src/v2/types/interrupt.ts`
- 13×
`showcase/integrations/<integ>/src/app/demos/interrupt-headless/page.tsx`

## Test plan

- [x] react-core unit tests (22/22 passing including new RESUME-PATH +
RESUME-PATH-REJECT regression tests)
- [x] Backplane D5 probe on langgraph-python (gen-ui-interrupt +
interrupt-headless GREEN end-to-end via the showcase control plane)
- [ ] CI green on this PR
2026-06-16 00:08:56 -07:00
Tyler Slaton 431d5baae0 feat(bot-slack): agent-native Slack APIs — assistant pane + native streaming (default-on) (#5447)
## What

Activates Slack's agent-grade APIs as the **default** experience for
`@copilotkit/bot-slack`, with zero config and safe degradation.
Implements the Notion spec *"Slack Agent APIs in bot-slack — assistant
pane + native streaming by default"* (Rev 2).

### Assistant pane ("Agents & AI Apps")
- Opening the pane greets the user + shows tappable prompt chips; each
pane conversation is its own thread (replies stay in-thread, instead of
leaking to a merged flat DM).
- While the agent runs, native composer status
(`assistant.threads.setStatus`: "is thinking…", "is using `tool`…")
replaces placeholder/`🔧` messages.
- Pane threads are auto-titled from the first message.
- Customize via the new `assistant` option; `assistant: false` disables
pane handling. Apps without the Agents toggle behave exactly as before
(pane machinery dormant).

### Native streaming
- Replies stream via `chat.startStream` / `appendStream` / `stopStream`
(raw markdown — real tables / fenced code render natively) wherever a
thread exists.
- Flat DMs and workspaces without the streaming API fall back to the
legacy `chat.update` transport **automatically** — the first
`startStream` failure marks the workspace legacy and replays via legacy.
Opting in can never break a bot. `streaming: "legacy"` forces the old
transport.

### Portable engine surface (`@copilotkit/bot`)
- `bot.onThreadStarted` lifecycle handler + `IncomingThreadStart` sink
event.
- Capability-gated `thread.setSuggestedPrompts` / `thread.setTitle` (the
shipped `postFile` gating pattern).
- Two `SurfaceCapabilities` flags + optional `PlatformAdapter` methods.
All degrade gracefully on surfaces without support.

## Files
- **New:** `bot-slack/src/assistant.ts` (Bolt `Assistant` middleware ⇄
engine sink), `bot-slack/src/native-stream.ts` (`NativeMessageStream`,
same `append/finish` contract as `MessageStream`, legacy fallback).
- **Engine:** `platform-adapter.ts`, `thread.ts`, `create-bot.ts`,
`index.ts` (+ `bot-ui` `Thread` type).
- **Adapter:** `adapter.ts` (options/capabilities/stream
branch/methods), `slack-listener.ts` (one-line no-double-delivery
guard), `event-renderer.ts` (pane status mode + native text transport).
- **Example/docs:** `examples/slack` manifest (`assistant_view` + scope
+ events) & dev-ex, package READMEs/ARCHITECTURE, `slack.mdx`.
- **Hooks:** `chore(hooks)` makes the `lint-fix` script
double-quote-free (it was breaking `sh -c "…"` on Windows/lefthook
2.1.1).

## Tests / verification (CI-runnable)
- New unit tests: `native-stream` (cadence, 12k continuation, fence
re-open, first- **and** continuation-`startStream` fallback),
`assistant` (defaults-before-onThreadStarted ordering, thread-scoped
turn, auto-title), listener no-double-delivery, renderer pane-status
mode, engine capability gating.
- Verified locally: `@copilotkit/bot` 31 tests, `@copilotkit/bot-slack`
199 tests, `bot-ui` tests; `tsc` clean (bot + bot-slack); build clean;
oxlint/oxfmt clean; publint/attw pass.

## ⚠️ E2E not covered here
The §8 end-to-end flows need a **real Slack workspace** with the Agents
toggle (open pane → chips → streamed reply w/ live status; channel
mention → native markdown; kill `startStream` → legacy fallback;
non-agent app → shipped behavior). `examples/slack` is the vehicle.
Three items remain to confirm against a live workspace (spec §7 spikes):
the `isAssistantThread` runtime detection, the exact `assistant_view`
manifest acceptance, and `stopStream` finalize tolerance.

Spec:
https://app.notion.com/p/copilotkit/Slack-Agent-APIs-in-bot-slack-assistant-pane-native-streaming-by-default-spec-37b3aa38185281f5948dc6a665064d04
2026-06-15 22:10:17 -07:00
Jordan Ritter 1ef6a0960b test(harness): update d5-* + d6-all-pills unit tests to new probe + cascade contract
- d5-* unit tests: update page fakes / assertion-call shapes to pass
  the ctx bridge ({bubbleIndex, text}) instead of querying the live
  bubble list
- d6-all-pills.test.ts: update page-fake dispatch + cascade-state
  expectations to match the atomic readCascadeState return shape
  ({count, text} or null) — picks up the cross-tier consistency
  guarantee in the unit layer
2026-06-15 18:37:26 -07:00
Jordan Ritter cd2af0f9fd test(harness): conversation-runner unit coverage — banner debounce, cold-start retry, skipFill, error-banner shapes
Comprehensive unit coverage for:
- preFill + fillAndVerifySend retry behavior
- baseline banner snapshot + differs-from-baseline debounce matrix
  (stable banner = no fail; changed banner = fast-fail after 2 polls)
- cold-start retry cases (first-attempt banner → reload + retry;
  retry-then-banner → translatedErr; retry exhausts deadline)
- skipFill / skipSend short-circuit paths
- ErrorBannerReadResult union shape variants (present/absent/unknown)
2026-06-15 18:37:26 -07:00
Jordan Ritter 3d49d4c38f test(harness): bubble-race integration tests + defect reproductions
- bubble-race-repro.ts: shared driver harness for repro tests
- bubble-race-repro-defect-{1,2,3,4}.test.ts: real-browser RED tests
  that fail without the fix and pass with it
    - Defect 1: fast-replay false-positive settle on count
    - Defect 2: multi-turn flicker via list[last]
    - Defect 3: cascade-tier blindness
    - Defect 4: boot-time baseline staleness / pre-paint placeholder
- bubble-race-mechanisms.test.ts: GREEN tests pinning internal
  contracts (cascade-pollution guard, atomic readCascadeState,
  init-script idempotency)
- wait-for-turn-complete.test.ts: 3-conjunct classification matrix
- probe-contract.test.ts: ctx-bridge shape
- sse-counter.test.ts: counter increment + framenavigated reset +
  handle-cache invalidation
2026-06-15 18:37:26 -07:00
Jordan Ritter cf7efbda95 fix(harness): d6 driver — installPrePaintFromEnv + attachSseInterceptor wiring + helper-based captureDiagnostics
- d6-all-pills.ts: wire installPrePaintFromEnv and
  attachSseInterceptor PRE-goto in both defaultLauncher and pooled
  launcher paths, so first-paint state is deterministic and the SSE
  counter is armed before any page navigation
- captureDiagnostics rewired through the Node-side helper
  (readCascadeState) — single page.evaluate per call, no
  within-snapshot tier drift
- Consume BUBBLE_RACE_MESSAGES override for defect-4 repro flows
2026-06-15 18:37:25 -07:00
Jordan Ritter 38e155e141 refactor(harness): probe-contract ctx bridge — assertions(page, {bubbleIndex, text}) for d5 scripts
- d5-gen-ui-custom, d5-gen-ui-open-advanced, d5-mcp-apps,
  d5-subagents, d5-tool-rendering-default-catchall: wrap assertion
  bodies in ctx bridge so they receive {bubbleIndex, text} for the
  turn under test, replacing turn-leaky list[last] retrieval
- _gen-ui-shared: add readAssistantTextAt(page, bubbleIndex)
  adapter; remove readLastAssistantText (no longer turn-scoped)

This is the read-side counterpart to the atomic readCascadeState
on the helper side: probes ask for the bubble at the runner's
chosen index instead of guessing.
2026-06-15 18:37:25 -07:00
Jordan Ritter 3f3986b00b refactor(harness): waitForTurnComplete 3-conjunct primitive + cold-start retry + banner debounce
- waitForTurnComplete: 3-conjunct settle (SSE counter + DOM-at-index
  + text-quiet for settleMs) replaces count-only false-positive path
- Cold-start retry: single bounded retry on first-attempt banner;
  reload + re-resolve chat input + settleMs/POLL_INTERVAL_MS floor
  honoring shared turn deadline; translatedErr throw on exhaustion
- Banner fast-fail: baseline banner text snapshot +
  differs-from-baseline 2-poll debounce; in-poll BannerVisibleError
  translates to AssistantErroredError
- New error classes: BannerVisibleError, AssistantErroredError
- readErrorBanner returns 3-state union via
  ErrorBannerReadResult discriminated union
- chatInputSelector cascade for input resolution across integrations
2026-06-15 18:36:59 -07:00
Jordan Ritter 473a9539a0 feat(harness): add bubble-race helpers — assistant-message-count, sse-interceptor, init-scripts
- assistant-message-count: 4-tier cascade + atomic readCascadeState
  returning {count, text} from a single page.evaluate (prevents
  within-poll cross-tier inconsistency); null fallback for
  cascade-pollution
- sse-interceptor: __hk_runsFinished CDP counter + page-side fetch
  wrapper; idempotent attach/detach with handle-cache invalidation
  on stop and framenavigated reset (cold-start retry compatible)
- init-scripts: pre-paint placeholder env adapter
  (installPrePaintFromEnv), strip-selector helper, and
  BUBBLE_RACE_MESSAGES override consumer for defect-4 repro
  infrastructure
2026-06-15 18:36:59 -07:00
Jordan Ritter 6fbd66fa83 fix(showcase): mirror useInterrupt RESUME-PATH contract in 13 demo-local hooks
Each integration's interrupt-headless demo defines a local useHeadlessInterrupt
hook around the framework useInterrupt. Slot-2 originally identified 8
quarantined integrations (claude-sdk-typescript, langgraph-{fastapi,python,
typescript}, langroid, pydantic-ai, spring-ai, strands); review-round
follow-ups extended the sweep to llamaindex, mastra, ag2, agno, and
crewai-crews (5 more integrations sharing the same byte-identical hook).

The demo-local resolve() previously fire-and-forgot copilotkit.runAgent(...)
via `void runAgent(...).catch(() => {})`. Mirroring the framework fix:

- Make resolve async, return await copilotkit.runAgent(...).
- Use a pendingRef so resolve has stable identity (drop pending from
  useMemo deps).
- Type signature: resolve: (response: unknown) => Promise<unknown>.
- Wrap in try/catch + setPending(null) + console.error + rethrow,
  symmetric with the framework hook.
- onRunFailed also setPending(null).

13 integrations patched byte-identically.
2026-06-15 17:11:40 -07:00
Jordan Ritter c81b361f15 fix(showcase/harness): register /webhooks/deploy on CP path + fail-loud on missing SHARED_SECRET (#5458)
## Problem

The \`Showcase: Verify Deploy\` workflow runs \`notify-harness\` after
every merge to \`main\` and POSTs an HMAC-signed payload to
\`\${SHOWCASE_HARNESS_URL}/webhooks/deploy\`. That POST has been
returning **404** on every main deploy since at least 2026-06-12 (5+
consecutive failures), so the harness dashboard has not reflected any
recent deploys.

**Root cause:**

1. \`POST /webhooks/deploy\` is only registered when
\`webhookSecrets.length > 0\` — gate at
\`showcase/harness/src/http/server.ts:119\`.
2. The CP \`buildServer\` call in \`runControlPlane\`
(\`showcase/harness/src/orchestrator.ts\` ~line 2961 pre-fix)
**omitted** both \`webhookSecrets\` and \`metrics\`. The public Railway
host running the CP role never mounted the route — every notify-harness
POST returned 404.
3. The FATAL-CONFIG guard on missing \`SHARED_SECRET\` (pre-fix
\`orchestrator.ts:782-790\`) only fired when \`NODE_ENV ===
"production"\`. Any deploy whose NODE_ENV was unset, set to something
other than the literal \`"production"\`, or set after the check,
silently booted with \`webhookSecrets=[]\` and no fatal error.

## Fix

**Part 1: Register the route on the CP path.**
- Lifted the env→secrets loader into a new exported helper
\`loadWebhookSecrets()\` (orchestrator.ts ~175-209) so both boot paths
consume the same predicate.
- Worker \`boot()\` (~line 836) calls it; CP \`runControlPlane\` (~line
2421) now also calls it.
- The CP \`buildServer\` call (~line 3047) now passes \`webhookSecrets\`
and \`metrics\` alongside the existing \`bus\`/\`probes\`/\`fleetRuns\`
wiring — same shape as the worker call.

**Part 2: Tighten the FATAL-CONFIG predicate.**
- New predicate (inside \`loadWebhookSecrets\`): throw unless EITHER a
secret is set OR \`NODE_ENV === "test"\` OR \`HARNESS_ALLOW_NO_SECRET
=== "1"\` (narrow escape hatch for local dev).
- Thrown error message names the env-vars, the gate location
(\`src/http/server.ts:119\`), the silent-404 symptom, and the escape
hatches — so an operator reading the deploy log can recover without
spelunking.

## Test coverage

Four new red-green tests in \`orchestrator.test.ts\` under \`B2:
/webhooks/deploy registered on CP + fail-loud on missing
SHARED_SECRET\`:

- **Test A** — CP boot registers \`POST /webhooks/deploy\` when
\`SHARED_SECRET\` is set. Probe with no HMAC headers: pre-fix returned
404, post-fix returns 401 (route mounted, HMAC reject path fires).
RED→GREEN.
- **Test B** — Worker boot still registers the route (regression guard).
Pre-fix and post-fix both 401.
- **Test C** — FATAL-CONFIG fires when \`SHARED_SECRET\` unset AND
\`NODE_ENV=development\`. Pre-fix resolved silently, post-fix rejects
with \`/FATAL-CONFIG.*SHARED_SECRET/\`. RED→GREEN.
- **Test D** — Boot succeeds when \`SHARED_SECRET\` unset AND
\`NODE_ENV=test\` (escape hatch). Pre-fix and post-fix both green.

Existing R5-G4 D5 test (orchestrator.test.ts:1592) updated to match the
new error message via \`/FATAL-CONFIG.*SHARED_SECRET/\` — same throw,
broader match.

## Verification

- Full harness suite: **2719/2719 green** across 128 files
- \`tsc --noEmit\`: clean
- RED capture pre-fix:
  - Test A: \`AssertionError: expected 404 to be 401\`
- Test C: \`AssertionError: promise resolved "{ port: ..., bus: { ... },
... }" instead of rejecting\`

## Deploy note

Requires a CP-role redeploy on the public Railway host for the new route
registration to take effect — once redeployed, the next
\`notify-harness\` POST will land on the harness dashboard.

## Test plan

- [x] B2 Tests A-D pass post-fix (red→green captured)
- [x] Full harness vitest suite green (2719/2719)
- [x] \`tsc --noEmit\` clean
- [ ] Post-deploy: confirm next \`Showcase: Verify Deploy\` run shows a
2xx notify-harness POST (not 404)
- [ ] Post-deploy: confirm harness dashboard reflects the next deploy
2026-06-15 16:44:07 -07:00
Jordan Ritter 3ce3f5394d fix(showcase/harness): hoist OPS_TRIGGER_TOKEN fail-loud + emit on subscribeDeployResults sync throw (R3 cleanups)
R3-F1 (OPS_TRIGGER_TOKEN hoist):
Extract loadOpsTriggerToken() and hoist its invocation to the top of both
boot() and runControlPlane(), alongside loadWebhookSecrets() and
loadPocketbaseUrl(). Pre-fix the empty/whitespace check fired AFTER pb /
bus / scheduler / writer / S3 uploader allocations, so a typo'd
`OPS_TRIGGER_TOKEN=` allocated expensive resources before throwing. Now
all three fail-loud config predicates fire at the top, before any
allocation needing teardown. Behaviour is unchanged for valid tokens
(trimmed via R3-A.5 contract) and for the unset case (router omitted with
info log).

R3-F2 (sync-throw also emits deploy.writer.failed):
subscribeDeployResults() pre-fix only emitted `deploy.writer.failed` on
writer.write() promise rejection — a synchronous throw inside
deployEventToProbeResult() (malformed event, type drift) bypassed the
catch and we lost both the log AND the bus emit. Wrap the sync mapping
in try/catch and mirror the rejection path so alert rules / metrics
subscribers observe sync throws the same way they observe async write
failures.

Tests:
- 5 unit tests for loadOpsTriggerToken: undefined / empty-string fail-loud
  / whitespace-only fail-loud / R3-A.5 trim contract / verbatim value.
- 1 test for the sync-throw path: vi.doMock deployEventToProbeResult to
  throw, assert bus emits deploy.writer.failed with the err message AND
  writer.write was NOT called. Red-green verified locally (test fails
  without the try/catch).

Suite: 2731 passing (+6 new). Typecheck clean.
2026-06-15 16:37:05 -07:00
Jordan Ritter d5bdf2c9fd fix(showcase/harness): track CP deploy.result unsubscribe + emit on write failure (R2 cleanups) 2026-06-15 16:23:29 -07:00
Jordan Ritter df97ae8d2a fix(showcase/aimock): align catchall fixture shape for 6 remaining D5 reds
PR #5459 added userMessage matchers ("forecast for Tokyo" + "current price
of AAPL") on the catchall fixtures, which flipped most integrations from
red to green. Six cells stayed red on staging because the FIXTURE SHAPE
itself was broken on the matched fixtures, not just the userMessage key.

Failing cells (all custom-catchall except strands default):
  d5:agno/tool-rendering-custom-catchall
  d5:llamaindex/tool-rendering-custom-catchall
  d5:langroid/tool-rendering-custom-catchall
  d5:claude-sdk-python/tool-rendering-custom-catchall
  d5:strands/tool-rendering-custom-catchall
  d5:strands/tool-rendering-default-catchall

PocketBase confirms the failure mode: turn 1 (Tokyo) completes; turn 2
(AAPL) times out at 30s or renders the wrong content. E.g. agno:

  errorDesc: "timeout: assistant did not respond within 30000ms"
  failure_turn: 2
  turns_completed: 1 / 2

Root cause: the failing tool-emit fixtures used `hasToolResult: false`
(or `toolName: "<tool>"`) as their gate. aimock's hasToolResult check is
`messages.some(m => m.role === 'tool')` over the WHOLE thread — so once
turn 1's Tokyo tool result lands in the conversation, hasToolResult is
permanently true and `hasToolResult:false` can never match turn 2 → no
fixture → 30s timeout. `toolName:get_stock_price` likewise fails when an
integration backend doesn't forward the tool definition on turn 2.

The same trap is documented in
showcase/aimock/d6/langgraph-python/tool-rendering.json:

  "_comment": "Gated on toolName:get_stock_price rather than
   hasToolResult:false. The D5 tool-rendering-custom-catchall probe runs
   'weather in Tokyo' first, which leaves a get_weather tool result in
   the thread; the aimock router implements hasToolResult as
   messages.some(m=>m.role==='tool'), so hasToolResult is permanently
   true on the AAPL turn and a hasToolResult:false gate could never
   match (→ no_fixture_match → 503 → 30s timeout)."

Fix: align all 6 files to the canonical pattern used by mastra/spring-
ai/built-in-agent on this probe:

  1. toolCallId-keyed narration fixture FIRST
  2. tool-emit fixture SECOND with ONLY `userMessage` + `context`
     (no hasToolResult / toolName gate)

The toolCallId narration uses aimock's
`messages[last].role === 'tool' && tool_call_id === ...` check, so it
correctly wins on iteration 2 (post-tool-result) without being affected
by older turns' tool results. The tool-emit fixture matches turn 1 (last
message is user) and re-emits only when the narration above hasn't
matched.

Strands' two cells additionally needed REORDERING — they had the tool-
emit fixture before the toolCallId fixture, defeating first-match-wins.

Strands' staging backend was also returning 502 during testing; once it
recovers, the corrected fixtures should let the probe pass. The fixture
changes are necessary but may not be sufficient for strands if backend
remains down.

No probe-side, harness, or backend changes — pure fixture-content
alignment. Six fixture files modified; line totals: -87 / +68.
2026-06-15 16:09:04 -07:00
Jordan Ritter 16a59a45c7 chore(showcase/harness): warn on webhook escape hatch + vi.stubEnv in B2 tests
CB-1 (Slot 2 #22): convert B2 deploy-webhook tests to vi.stubEnv with
`vi.unstubAllEnvs()` in afterEach instead of manual process.env mutation
+ try/finally restoration. Test runner now guarantees restoration and the
test bodies are shorter / less error-prone (CR Slot 2 #22).

CB-2 (Slot 2 #28): when `loadWebhookSecrets`' escape hatch fires with a
real-looking NODE_ENV (anything except "test"), log at `warn` instead of
`info` so a production typo (NODE_ENV=staging + HARNESS_ALLOW_NO_SECRET=1)
is visible in dashboards / log alerting. Pure local-dev (NODE_ENV=test)
stays at info level so a normal unit-test boot doesn't spam warnings.

CB-3 (Slot 4 #17): clarify `loadWebhookSecrets`' docstring + bypass-log
message — "set" is ambiguous (empty string would qualify pre-fix); use
"non-empty" to match the actual predicate.
2026-06-15 15:25:42 -07:00
Jordan Ritter f42180f7bf fix(showcase/harness): hoist fail-loud config checks + symmetric POCKETBASE_URL predicate
R1-F2 (bucket b, defensive ordering): hoist fail-loud config validation
to the TOP of boot() and runControlPlane() — BEFORE any pb client, queue,
bus, scheduler, writer, S3 uploader, fleet-health, or aggregator
allocations. Pre-fix `loadWebhookSecrets` lived AFTER the entire
scheduler+writer (and CP queue+aggregator) setup, so a misconfigured boot
allocated expensive resources and mounted file watchers before throwing.

R1-F3 (bucket b, predicate symmetry): broaden the POCKETBASE_URL fail-loud
predicate to match SHARED_SECRET semantics. Pre-fix the POCKETBASE_URL
guard only fired on `NODE_ENV === "production"`, while `loadWebhookSecrets`
fired unless NODE_ENV=test or an explicit escape hatch — so staging /
unset / "development" deploys silently bound to http://localhost:8090.

Extracted `loadPocketbaseUrl(logger)` with the same test-or-escape-hatch
predicate as `loadWebhookSecrets`. A new `HARNESS_ALLOW_NO_PB_URL=1` env
flag mirrors `HARNESS_ALLOW_NO_SECRET=1` for local dev. Worker
`boot()` and the CP's `resolveFleetPbConfig` both call the helper.

Tests:
  - HF13-A2 production-only assertion broadened (SHARED_SECRET set so the
    test specifically exercises the POCKETBASE_URL guard).
  - New R1-F3 tests cover NODE_ENV=development without POCKETBASE_URL
    (throws) and the HARNESS_ALLOW_NO_PB_URL=1 escape hatch (succeeds).
2026-06-15 15:25:11 -07:00
Jordan Ritter fc828bb08c fix(showcase/harness): subscribe deploy.result in CP boot to deliver dashboard event
R1-F1 (bucket a, load-bearing): the CP boot path now subscribes to
`deploy.result` events through its status writer. Pre-fix, B2 mounted POST
/webhooks/deploy on the CP role but only the worker boot path had a
`bus.on("deploy.result", ...)` listener — signed POSTs to the CP host
returned 202, the bus event fired with no subscriber, and the deploy-overall
dashboard row never landed.

Extracted the handler into a shared `subscribeDeployResults(bus, writer)`
helper exported from orchestrator.ts so both boot paths share the IDENTICAL
logic. Worker boot keeps its inline pattern (busUnsubs.push) so teardown is
unchanged; CP keeps the subscription alive for the lifetime of the bus.

Red-green tests in src/orchestrator.test.ts under "R1-F1" assert that
emitting deploy.result on the returned handle's bus drives one write keyed
"deploy:overall" through a stubbed status writer — once for runControlPlane
(the new path) and once for boot() (regression guard for the helper
extraction).
2026-06-15 15:23:24 -07:00
Jordan Ritter 07e4998798 fix(showcase/aimock): add forecast-Tokyo + AAPL catchall fixtures to flip remaining D5 reds
The catchall userMessage rename (#5453) only renamed STALE strings to 'forecast for Tokyo'. Integrations whose catchall fixtures used PILL-ALIGNED userMessages ('What's the weather in San Francisco?', 'Find flights from SFO to JFK.', 'Chain a few tools in this single turn', etc.) had no 'forecast for Tokyo' matcher to begin with — so the D5 catchall probe still fell through to live LLM and the cells stayed red on staging.

This PR ADDS the canonical 'forecast for Tokyo' (default+custom catchall) and 'What's the current price of AAPL?' (custom catchall only) emit+narrate fixture pairs to each integration that was missing them. Existing pill-aligned fixtures are preserved (pure prepend at the start of the fixtures array; first-match-wins means new fixtures match the D5 probe inputs without colliding with existing pill prompts).

Integrations touched:
- langgraph-typescript: both default + custom
- langgraph-fastapi: both default + custom
- google-adk: custom only (default already green)
- ms-agent-dotnet: custom only (default already green)

After merge, fleet-cp's e2e-deep probe will re-run within ≤6h staleness window and flip cells GREEN.
2026-06-15 15:00:53 -07:00
github-actions[bot] bfd1f12b00 style: auto-fix formatting 2026-06-15 21:56:46 +00:00
Jordan Ritter 2939696415 fix(showcase/harness): register /webhooks/deploy on CP path + fail-loud on missing SHARED_SECRET
The 'Showcase: Verify Deploy' workflow's notify-harness step has been
POSTing to /webhooks/deploy and getting 404 on every main deploy since
at least 2026-06-12 (5+ consecutive failures). Root cause:

  1. POST /webhooks/deploy is only registered when
     webhookSecrets.length > 0 (gate at src/http/server.ts:119).
  2. The CP buildServer call in runControlPlane omitted BOTH
     webhookSecrets and metrics — so the public Railway host running
     the CP role never mounted the route. The worker boot path had
     them, but it isn't the host receiving notify-harness POSTs.
  3. The FATAL-CONFIG guard on missing SHARED_SECRET was gated on
     NODE_ENV === 'production' — any deploy with NODE_ENV unset or set
     to something other than the literal 'production' silently shipped
     with webhookSecrets=[] and no fatal error.

Part 1 — wire webhookSecrets + metrics into the CP buildServer call.
Lifted the env→secrets loader into a new exported helper
'loadWebhookSecrets()' (~orchestrator.ts:175-209) so BOTH boot paths
consume the same predicate. The worker 'boot()' call (~line 836) and
the CP 'runControlPlane' call (~line 2421) both invoke it, and the CP
buildServer call (~line 3047) now passes webhookSecrets+metrics
alongside the existing bus/probes/fleetRuns wiring.

Part 2 — tighten the FATAL-CONFIG predicate.
The new predicate throws unless EITHER a secret is set OR NODE_ENV ===
'test' OR HARNESS_ALLOW_NO_SECRET === '1' (narrow escape hatch for
local dev). The thrown error message names the env-vars, the gate
location ('src/http/server.ts:119'), the silent-404 symptom, and the
escape hatches, so an operator reading the deploy log can recover
without spelunking.

Updated the existing R5-G4 D5 regex (orchestrator.test.ts:1592) since
the error message changed; covered by the new B2 Tests A/C/D below.

Tests: 4 new red-green tests in 'B2: /webhooks/deploy registered on CP
+ fail-loud on missing SHARED_SECRET':
  - Test A (CP /webhooks/deploy registration): RED 404 → GREEN 401.
  - Test B (worker /webhooks/deploy regression): GREEN 401.
  - Test C (FATAL-CONFIG when NODE_ENV=development): RED resolves →
    GREEN rejects with /FATAL-CONFIG.*SHARED_SECRET/.
  - Test D (NODE_ENV=test escape hatch): GREEN, no throw.

Full harness suite: 2719/2719 green. typecheck: clean.

Requires a CP-role redeploy on the public Railway host for the new
route registration to take effect.
2026-06-15 14:55:21 -07:00
Jordan Ritter fb3d5b3547 test(showcase/harness): cover CP webhook registration + fail-loud on missing SHARED_SECRET
Adds four red-green tests for the notify-harness 404 follow-up:

  A. CP boot registers POST /webhooks/deploy when SHARED_SECRET is set
     — RED pre-fix (404, route not mounted on CP); GREEN post-fix (401,
     the route's own HMAC reject path fires on an unsigned POST).
  B. Worker boot still registers POST /webhooks/deploy (regression
     guard) — already green; pins the unchanged behavior.
  C. boot throws FATAL-CONFIG when SHARED_SECRET unset AND NODE_ENV is
     non-test — RED pre-fix (current guard fires only on
     NODE_ENV='production'); GREEN post-fix once the predicate is
     tightened.
  D. boot succeeds when SHARED_SECRET unset AND NODE_ENV='test'
     (test-mode escape hatch) — already green; pins the escape hatch.

Red capture (pre-fix run):
  - Test A: AssertionError: expected 404 to be 401
  - Test C: AssertionError: promise resolved instead of rejecting
2026-06-15 14:52:31 -07:00
Jordan Ritter 4791c86923 fix(showcase/harness): override LOCAL_SERVICES_JSON in --isolate generator to target the requested slug (#5454)
## Summary
- `showcase/docker-compose.local.yml:262` hardcodes
`LOCAL_SERVICES_JSON` to `showcase-langgraph-python` (intentional N=1
demo default).
- `showcase/bin/showcase test <slug> --d6 --isolate` was inheriting that
value verbatim into the iso1 stack, so the iso1 control-plane discovered
`showcase-langgraph-python` instead of `showcase-<requested-slug>`.
- Result: iso1 probes targeted the wrong service. Visible in iso1
harness logs as `discovery.railway-services.local-injection count:1
names:["showcase-langgraph-python"]` regardless of CLI arg.
- This PR teaches the iso1 compose generator (`apply_isolation` in
`showcase/scripts/cli/_common.sh`) to inject a per-slug
`LOCAL_SERVICES_JSON` override built from the slug's manifest.yaml demos
list. Fallback to `["agentic-chat"]` if manifest absent.

## Scope
Two files, ~55 LOC added: `showcase/scripts/cli/cmd-test.sh` (pass slug
arg), `showcase/scripts/cli/_common.sh` (inject regex sub in python
rewriter). Bash + embedded python only; no TypeScript changed.

## Verification
- **Local `--isolate` discovery confirmed correct:**
`showcase/bin/showcase test ms-agent-python --d6 --isolate` — iso1
harness log now shows `discovery.railway-services.local-injection
names:["showcase-ms-agent-python"]` (was `showcase-langgraph-python`
before this PR).
- Persistent stack default behavior (langgraph-python N=1) preserved via
fallback when no slug arg provided.
- **Heredoc hardening scope:** commit edc77f809 moves `$slug` from
bash-interpolation into the python rewriter to an env var. This is a
slug-only carve-out — `$slug` is the only value that originates from the
CLI arg path. The other bash-interpolated `$VAR`s embedded in the
heredoc (`$PORTS_FILE`, `$COMPOSE_FILE`, `$name`, `$SHOWCASE_ROOT`,
`$ISOLATE_PORT_OFFSET`) remain script-internal: each is constructed
inside `_common.sh` from validated sources (manifest reads, computed
offsets, fixed roots), not from user input, and CR Round 1 + Round 2
slot 5 both verified they are not user-tainted. A broader
env-var-pass-all-vars refactor would be a separate concern and is out of
scope for this PR.

## Out of scope
- Worker heartbeat clock-skew issue (`fleet.health.worker-unhealthy
lastHeartbeatAt N min stale`) is a separate bug; not touched.
- `buildLocalServicesJson` in `cli/control-plane-run.ts` has a similar
comment-vs-code disagreement (JSDoc says "filter", code returns env
verbatim) — flagged but not auto-fixed, as iso1 override is the correct
insertion point.
- Generalized env-var-pass for all heredoc-embedded $VARs (see
Verification) — separate refactor concern.
2026-06-15 11:38:26 -07:00
Jordan Ritter 318bd9e45f fix(showcase/aimock): align D6 tool-rendering catchall userMessage to D5 probe input (#5453)
## Summary
- D5 e2e-deep probes for `tool-rendering-{default,custom}-catchall` send
`"forecast for Tokyo"` as the test input (see
`showcase/harness/src/probes/scripts/d5-tool-rendering-{default,custom}-catchall.ts`).
- aimock uses substring match on `userMessage`.
- The catchall fixtures on main had a stale `"check Tokyo weather
forecast"` string that couldn't substring-match the probe input →
fixture miss → probe falls through to live LLM → CV ✗ red D4 on the
dashboard.
- This PR renames the userMessage to the canonical `"forecast for
Tokyo"` across 28 catchall fixture files in 16 integrations.

## Scope
28 files × ~2 userMessage occurrences each = 67 line changes. **No code,
no agent, no page.tsx changes.** Pure fixture-data alignment.

Integrations covered (default-catchall and/or custom-catchall): ag2,
agno, built-in-agent, claude-sdk-python, claude-sdk-typescript,
crewai-crews, google-adk, langgraph-python, langroid, llamaindex,
mastra, ms-agent-dotnet, ms-agent-python, pydantic-ai, spring-ai,
strands.

## Commit history note
Commit b1f19bdc8 changes the rename target from the original `"weather
in Tokyo"` (commit 388c69e68) to `"forecast for Tokyo"`. This was a CR
Round 1 catch: `"weather in Tokyo"` was a substring of the chain pill
prompt (`"weather forecast chain in Tokyo"` / similar), which would have
caused the catchall fixture to incorrectly match chain-pill traffic.
`"forecast for Tokyo"` has no such substring collision with any other
probe input.

## Verification
- **Live red-green proof** was performed on the prior `"weather in
Tokyo"` rename (commit 388c69e68) against `ms-agent-python` via
`showcase/bin/showcase test ms-agent-python --d6 --isolate`:
- Before (origin/main): catchall featureTypes red (fixture miss → live
LLM → flaky)
- After: `d6:ms-agent-python/tool-rendering-default-catchall=green`,
`d6:ms-agent-python/tool-rendering-custom-catchall=green`
- **The current HEAD's `"forecast for Tokyo"` rename (b1f19bdc8) has NOT
been re-run live.** It is verified by static analysis only: substring
math (no collision with any known D5 probe input or chain pill prompt)
and a clean CR Round 2 across all reviewing agents.
- Post-merge dashboard re-probe will be the final runtime verification.

## Out of scope (separate follow-ups)
- LG-TS / LG-FastAPI catchall fixtures don't have the stale string — use
pill-aligned userMessages; need different fix
- `tool-rendering` (non-catchall) and `tool-rendering-reasoning-chain`
featureTypes have separate failure modes
- Other red cells in the dashboard (frontend-tools-cosmic timeout,
agent-config, auth, etc) are unrelated
2026-06-15 11:33:34 -07:00
Jordan Ritter edc77f8090 fix(showcase/harness): pass slug via env var to python rewriter instead of bash interpolation
The python rewriter in apply_isolation previously interpolated $slug
directly into the inline python source via bash. A slug containing a
single quote would break the python literal. Internal-tool risk only
(slug is developer-typed), but cheap to harden.

Pass slug via SHOWCASE_ISO_SLUG env var and read os.environ.get(...)
inside the python heredoc. Defense-in-depth; no behavior change for
valid slugs.
2026-06-15 11:16:24 -07:00
Jordan Ritter b1f19bdc80 fix(showcase): rename catchall userMessage to 'forecast for Tokyo' to avoid chain pill substring collision
The previous 'weather in Tokyo' rename (388c69e68) was a substring of the
main tool-rendering chain pill prompt 'Chain a few tools in this single
turn: get the weather in Tokyo, search flights from SFO to Tokyo, and roll
a d20.' Because aimock loads fixtures alphabetically per integration dir
and uses substring match with first-match-wins, the catchall fixture
(loaded before tool-rendering.json) was intercepting chain pill matches
across 16 integrations.

Rename catchall fixture userMessage to 'forecast for Tokyo' — a phrase
not contained in any other pill prompt. Update the corresponding D5
catchall probe inputs in d5-tool-rendering-{default,custom}-catchall.ts
and the test assertions that pin those inputs.

Call-Site Enumeration: 'weather in Tokyo' remains intentionally in
page.tsx suggestions.ts pills (user-visible UX) and inside the chain
pill prompt itself — neither is in the substring-match path now.
2026-06-15 11:14:57 -07:00
Jordan Ritter 5e641d88c9 fix(showcase/harness): override LOCAL_SERVICES_JSON in --isolate generator to target the requested slug
The persistent stack's docker-compose.local.yml hardcodes LOCAL_SERVICES_JSON
to the langgraph-python sample for fast N=1 local demos. When --isolate
spawns an iso1 stack with a different slug (e.g. ms-agent-python), the
iso1 harness container inherited that hardcoded value, causing
discovery.railway-services.local-injection to enumerate the wrong service
(showcase-langgraph-python instead of showcase-<requested-slug>). The iso1
probe then targeted the wrong container, broke red-green verification, and
left D5 cells unwritten.

Inject a per-slug LOCAL_SERVICES_JSON override into the iso1 compose
generator so iso1 always probes the slug passed via --isolate.
2026-06-15 10:38:06 -07:00
Jordan Ritter 87b369fdb4 feat(showcase): declarative gen-UI demo as a sales-analyst dashboard (OSS-136) (#5396)
## Summary
- Reworks the `declarative-gen-ui` demo (LangGraph Python + Google ADK)
to feel like Beautiful Chat's sales dashboard, per
[OSS-136](https://linear.app/copilotkit/issue/OSS-136/declarative-gen-ui-demo-rework-to-feel-like-beautiful-chats-sales)
- Suggestion pills are natural business questions; chart-type steering
moved from user prompts into the agent system prompt + frontend context
(`sales-context.ts`, shared verbatim by both integrations)
- Every pill renders a dashboard-grade surface:
- **Show my sales dashboard** — bare KPI strip + regional revenue donut
+ 6-month revenue bars (no surrounding card)
  - **Team performance** — rep table + quota-attainment bar chart
- **Anything at risk?** — risk KPI strip over three severity cards (icon
badges, reason + next action)
  - **Top account details** — account fact card + product-line donut
- Renderers ported to beautiful-chat's visual language: card chrome,
metric typography, recharts donut/bars, shared palette, lucide severity
icons
- Test triad synced: D5 probe (per-pill newly-mounted testid
assertions), Playwright e2e (incl. click-dispatch guard), aimock
fixtures **captured from live model responses**, QA docs

## Verification
- Fixture validation: 738/738 · harness probe unit tests: 17/17
- Container e2e (fixture replay, both integrations rebuilt): LGP 6/6 ·
ADK 6/6
- Live-LLM iteration on LGP: all four pills produce the steered
composition consistently (11/11 captured runs + repeated interactive
verification)

## Test plan
- [ ] `pnpm exec nx run @copilotkit/showcase-harness:test --
d5-gen-ui-declarative`
- [ ] `pnpm --filter @copilotkit/showcase-scripts test aimock-fixtures`
- [ ] `BASE_URL=<container> npx playwright test
declarative-gen-ui.spec.ts` per integration
- [ ] Post-deploy:
https://dashboard.showcase.copilotkit.ai/#matrix:links,health row
"Declarative UI: Dynamic A2UI" still reaches D5 (CV badge) for
langgraph-python and google-adk
2026-06-15 09:42:45 -07:00
Jordan Ritter d5152eaa83 fix(showcase/e2e+qa): composition exclusions + KPI=4 contract alignment
- Scope clickPill locator to data-message-role='user' bubble so the pill
  button itself can no longer satisfy the dispatch guard
- Dedup clickPill retry: skip click if the user bubble already exists
- Hero pill: assert declarative-card count=0 (OSS-136 no-Card rule),
  metric count >=4 (was >=3 — KPI strip is 4 tiles per composition rule)
- At-risk pill: assert no chart and no table testids (composition rule)
- Top-account pill: assert no data-table and no status-badge testids
- Rename hero test title to 'KPI strip + pie + bar (no surrounding card)'
  so the title no longer falsifies the body
- QA docs: replace 'card + metrics + pie + bar' Expected Results with
  '4 KPI metrics + 1 PieChart + 1 BarChart, no surrounding Card per OSS-136'
- Probe responseTimeoutMs derived from FIRST_SIGNAL_TIMEOUT_MS so it
  matches the e2e 90s budget
2026-06-15 09:35:46 -07:00
Jordan Ritter 3ec2432964 fix(showcase/harness): newly-mounted gate + minCounts delta + per-pill chart asserts
- Migrate readDeclarativeTestIds from booleans to counts so leftover vs
  newly-mounted is distinguishable
- everyNewlyMounted gate uses current[k] > baseline[k] (was boolean
  !baseline[k] against a count, which falsely blocked at non-zero baseline)
- minCounts enforce newly-mounted delta, not raw current count
  (fixes the cross-pill bleed: at-risk metric:3 floor used to pass on
  hero's 3 leftover metrics with no fresh mount)
- Per-pill minCounts add the chart sibling asserts D5 was missing:
  hero=4 metric+1 pie+1 bar, team=1 table+1 bar, top-account=1 info-row+1 pie
- 31/31 harness tests green
2026-06-15 09:35:45 -07:00
Jordan Ritter 0f58f04e50 fix(showcase/renderers): stable row keys + per-card id + no-silent-zero charts
- DataTable rowKey uses first-column value + index instead of bare index,
  with JSON.stringify(row) fallback (stops re-mount on dynamic A2UI re-emits)
- Card emits data-card-id={props.title} so multi-card pills no longer
  collide on a single declarative-card testid
- PieChart/BarChart value coercion replaced 'Number(x) || 0' with
  finite-number check + console.warn on drift (no longer masks legitimate 0)
2026-06-15 09:35:45 -07:00
Jordan Ritter 8711326b5f fix(showcase/a2ui): tighten Zod schemas
- PrimaryButton.action: z.any() -> z.unknown() (forces caller narrowing)
- Row.justify/align + Column.align: z.string() -> z.enum() matching the
  renderer's CSS map
- DataTable rows accept numeric cells (z.union([string, number]))
- DataTable column-key refine documented in description (host
  CatalogComponentDefinition requires ZodObject, blocks .refine)
2026-06-15 09:35:45 -07:00
Jordan Ritter 0964823f3c fix(showcase/sales-context): honest duplication notice + extract TODO
Replace the misleading 'single source of truth' claim with an explicit
DUPLICATION NOTICE describing the per-integration parity convention and
a TODO(OSS-136) for the future shared-module extraction. Both copies
remain byte-identical.
2026-06-15 09:35:44 -07:00
Jordan Ritter df48df3587 fix(showcase/google-adk): align aimock fixture + suggestions comment
- Correct suggestions.ts file-path reference (was pointing at LP a2ui_dynamic.py)
- Tighten userMessage matchers to full pill prompts
- Strip unschema'd weight/variant fields from Metric/Card/Chart/Text payloads
2026-06-15 09:35:44 -07:00
Jordan Ritter e57ea6b864 fix(showcase/langgraph-python): align agent + aimock with ADK parity
- Replace fake gpt-5.4 with env-overridable real model (default gpt-4o)
- Register generate_a2ui tool matching SYSTEM_PROMPT + ADK structure
- Stub tool raises RuntimeError if middleware bypassed (fail-loud)
- Reorder LP fixture entries: inner render_a2ui before outer generate_a2ui
  to match ADK first-match-wins ordering
- Tighten userMessage matchers to full pill prompts (no substring hijack)
- Drop dead _design_a2ui_surface mirrors; strip unschema'd weight/variant fields
- Honest SYSTEM_PROMPT comment cross-referencing ADK _INSTRUCTION
2026-06-15 09:35:43 -07:00
Jordan Ritter 388c69e684 fix(showcase): align D6 tool-rendering catchall userMessage to D5 probe input
The D5 e2e-deep probe for tool-rendering-{default,custom}-catchall sends
"weather in Tokyo" as the test input (harness/src/probes/scripts/
d5-tool-rendering-{default,custom}-catchall.ts). The fixture userMessage
matcher uses substring match. Main's catchall fixtures had a stale
"check Tokyo weather forecast" string that could not substring-match
the probe input, causing the fixture to miss and the probe to fall
through to the live LLM — surfacing as CV x red D4 on the dashboard.

Rename to the canonical "weather in Tokyo" string across affected
integrations' tool-rendering-{default,custom}-catchall.json files.

No agent or page.tsx changes; fixture content otherwise unchanged.
2026-06-15 09:17:42 -07:00
Alem Tuzlak 21e4b87a43 docs(shell-docs): add WhatsApp bot guide 2026-06-15 17:08:27 +02:00
Alem Tuzlak a925e33be9 docs(bot-slack): document assistant pane + native streaming; enable in slack example
Update package READMEs + ARCHITECTURE for the assistant pane, native streaming,
and the new onThreadStarted / setSuggestedPrompts / setTitle surface. Reverse
the slack.mdx callout that told users to delete the assistant scopes (now
required), and enable the pane in the examples/slack manifest (assistant_view +
assistant:write + assistant_thread_* events) with a dev-ex onThreadStarted
greeting and the assistant config.
2026-06-15 15:47:46 +02:00
Jordan Ritter 97031e54e1 fix(showcase/aimock): break BIA tool-rendering cross-pill shadow on hasToolResult fallback
The custom-catchall probe sends 'weather in Tokyo' then 'AAPL'. After the
weather tool runs in turn 1, hasToolResult is true across the rest of the
thread — which fires tool-rendering.json's AAPL 'hasToolResult:true' narration
prematurely on turn 2 iteration 1, returning prose without ever emitting the
get_stock_price tool. The custom-catchall assertion (both tools rendered
through the wildcard testid) then fails with missing get_stock_price.

Replace the (hasToolResult:true narration + hasToolResult:false/turnIndex:0
emitter) layered fallbacks with a sequenceIndex:0 emitter ordered before a
bare userMessage+context narration. The per-test fixture-match counter resets
each run, so the emitter fires exactly once on iteration 1 regardless of
prior pills' tool history, then falls through to the narration on
iteration 2. Applied symmetrically to the two AAPL blocks in
tool-rendering.json (the 'What\'s the current price of AAPL?' block at the
top and the legacy 'current price of AAPL' alias block lower down). The
toolCallId-keyed narration above each block is retained for the non-BIA
fast path.

The earlier partial fix to tool-rendering-custom-catchall.json is kept (it
adds toolCallId-scoped narration legs ordered before the existing
hasToolResult:true narrations); those fixtures never match real probe
traffic (the probe sends 'weather in Tokyo' / 'current price of AAPL', not
the unique 'check Tokyo weather forecast' substring in this file) but the
reordering is consistent with the cross-file pattern and harmless.

Verified locally: built-in-agent:tool-rendering and
built-in-agent:tool-rendering-custom-catchall both green via
`./bin/showcase test ... --d6 --direct`; built-in-agent:tool-rendering-default-catchall
also green; aimock-fixtures.test.ts (collision/shadow ceilings) unchanged.
2026-06-15 00:18:53 -07:00
Jordan Ritter cb9fbe4c55 feat(showcase/built-in-agent): D6 BIA component port — tool-rendering + headless-complete + LGP-canonical naming (#5427)
## Summary

BIA D6 component port (#4 from PR #5413 followup list). Companion to PR
#5407 (claude-sdk-python), PR #5413 (initial BIA D6), PR #5421 (BIA D6
small follow-ups). Brings BIA closer to LGP gold-standard parity.

### What's in

- **tool-rendering**: 5 LGP-mirrored companion components
(`weather-card`, `flight-list-card`, `stock-card`, `d20-card`,
`custom-catchall-renderer`) + extracted `tool-renderers.tsx` wiring.
**D6 GREEN.**
- **headless-complete**: 2 new components (`stock-card`, `chart-card`) +
`get_revenue_chart` server tool + LGP-aligned 4 pill suggestions +
`data-message-role` on bubble cascade + `get_weather` tool name fix in
tool-renderers. Turns 1 (weather) + 2 (stock) GREEN.
- **roll_dice → roll_d20 rename** across BIA backend + reasoning-chain
references (LGP canonical naming).
- **aimock fixtures**: BIA-namespaced `tool-rendering.json` updated +
`tool-rendering-reasoning-chain.json` realigned to the rename.
- **PARITY_NOTES**: documents the server-tool reprompt loop
architectural gap blocking turns 3+4 of headless-complete.

### Known issue (documented in PARITY_NOTES.md, out-of-PR-scope)

headless-complete turns 3 (highlight_note) and 4 (revenue_chart) RED due
to BIA's TanStack multi-turn server-tool reprompt cycle + aimock
userMessage-keyed fixtures looping until timeout. Three remediation
options identified — all require changes outside this PR's scope (BIA
agent architecture / aimock matcher precedence / fixture matcher
gating).

### Test plan

- [x] Local `--direct` D6 on tool-rendering: GREEN (2.8s)
- [x] Local `--direct` D6 on headless-complete: turns 1+2 GREEN, turns
3+4 RED (documented)
- [x] CI green on PR
2026-06-14 20:15:30 -07:00
Tyler Slaton f655013dd2 fix(showcase): harden built-in-agent root rollout 2026-06-12 17:17:18 -07:00