Closed PR #5465 introduced cross-fixture leakage by stripping the
toolName discriminator on shared {userMessage, context} keys: with
aimock's alphabetical first-match-wins ordering,
tool-rendering-custom-catchall.json sorts before
tool-rendering-default-catchall.json, so default-catchall page requests
were served the custom file's content. The probe was structurally blind
because it asserted only DOM testids (copilot-tool-render +
data-tool-name=get_weather) — both fixtures emit get_weather, so the
testid signal passed regardless of which fixture won.
Real LGP-gold pattern is disjoint userMessages between default and
custom catchall fixtures (NOT shared keys discriminated by toolName).
This PR ports the LGP-gold pattern to the other 17 integrations and
adds page-text content assertions to both d5 catchall probes so this
class of regression can't recur silently:
- Probe prompts disjoint between default-catchall and custom-catchall
- Fixture userMessages updated to match the new disjoint prompts
- Page-text content assertions in both d5 catchall probes
(default negatively asserts the custom-catchall leak phrase;
custom positively asserts it)
Local gold-standard red-green proof captured:
- Step A: live staging Playwright baseline (RED-OF-RECORD)
- Step B: local control-plane on origin/main reproduces structural
fragility (probes pass at testid level, custom fixture wins on
default's userMessage path)
- Step C: local control-plane on this branch — both catchall probes
GREEN with the new content-asserting assertions
tool-rendering-default-catchall: pass=true (5443ms)
tool-rendering-custom-catchall: pass=true (8956ms)
cross-tool signature passed
LGP (langgraph-python) was already disjoint; it remains untouched.
Renames the internal display-counter variables that fed the
stats-bar / coverage-bar "Wired" labels to match the new "BE (Agent)"
taxonomy: totalWired → totalBeAgent in cells-view.tsx and
parity-view.tsx, pctWired → pctBeAgent in coverage-bar.tsx. These
are local render-time aggregates, not part of any persisted
contract.
The status enum literal "wired" (the .filter((c) => c.status ===
"wired") guard, the coverage-segment-wired test ID, and the
wiredByCategory Map's local name) is intentionally preserved — those
all key directly off the persisted catalog Status enum and changing
them would expand scope into the catalog contract.
Unifies the catalog integration status display label with the same
"BE (Agent)" taxonomy as the live-probe agent dimension. The
underlying Status enum literal ("wired"), the catalog.metadata.wired
data field, and the filter chip id ("wired" → cell-matrix.tsx:311
status === "wired" filter key) are all PRESERVED — they are
persisted catalog data + filter state contracts that must not move.
Only the user-facing display label on stats-bar, adaptive-stats-bar,
and the filter chip flips.
This collapses the two distinct dashboard "Wired" surfaces (the L1
live-probe dimension and the per-cell build-state count) under one
unified label, matching the user-confirmed Path B.
Unifies the L1 "agent" live-probe display label with the taxonomy
convention established by #5473 (UI (Frontend), E2E, CV, D6 — layer
descriptor in parentheses). The dimension name stays "agent" in code
(PocketBase row keys agent:<slug>, LiveDimension union, keyFor
lookups are all unchanged stable contracts) — only the visible label
changes.
Updates the level-strip L1 badge label and the packages-section
L1-L4 header legend (W → B, "Wired" → "BE (Agent)") so the
level-strip's ToneChip first-letter abbreviation matches the legend
key. Test assertions covering the rendered letter, the legend text,
and the degraded-tone test's local variable follow suit.
## Summary
The smoke probe was the same HTTP contract as `/health` on the same
service (200-OK JSON body), so every tick paid two HTTP calls for the
same liveness signal. This PR:
- Drops the `/smoke` GET probe and the `smoke:<slug>` primary
`ProbeResult`. The driver now emits `health:<slug>` as the primary and
`agent:<slug>` via writer side-emit (half the per-tick cost).
- Renames the dashboard's D3 / e2e row label from `E2E (Demo)` to `UI
(Frontend)` (and the short cell badge from `E2E` to `UI`) — the row was
already about "the demo page renders in a browser", not E2E in the
traditional sense.
- Drops the `Smoke` row from the cell drilldown (CellState.smoke field
is retained for back-compat so historical rows still parse).
- Adds an explicit `Health` row to the legend (was implicit).
The driver's registry `kind` stays `"smoke"` so existing YAML configs +
orchestrator family wiring keep routing to this driver — the emission
key (`health:<slug>`) is the taxonomy contract that matters. The
underlying probe key (`e2e:<slug>/<feature>`) is preserved on PocketBase
so historical rows render correctly during the rename window.
## RED proof (BEFORE)
`smoke:<slug>` primary emission in `liveness.ts` (origin/main):
```
18: * 1. `smoke:<slug>` — the RETURN VALUE of `run()`. The invoker runs
56: * `smoke:<slug>`/`health:<slug>`/`agent:<slug>` triple the dashboard
241: input.mode === "discovery" ? `smoke:${slug}` : input.key;
```
`E2E (Demo)` / `Smoke` strings in dashboard render code (origin/main):
```
cell-drilldown.tsx:38: * Trip)" (chat+tools round-trip), D3/e2e = "E2E (Demo)" (the demo page loads
cell-drilldown.tsx:49: { key: "e2e", label: "E2E (Demo)" },
cell-drilldown.tsx:52: { key: "smoke", label: "Smoke" },
cell-pieces.tsx:382: name="E2E"
unified-cell.tsx:231: <TestBadge name="E2E" level={model.d3} />
adaptive-legend.tsx:59: E2E (Demo): demo page loads and round-trips in a browser
```
## GREEN proof (AFTER)
Primary emission rewritten — driver returns `health:<slug>` and no
`smoke:<slug>` ProbeResult is ever emitted:
```
33: * to `/smoke` and emitted a `smoke:<slug>` primary result, but that probe
230: // `kind: "smoke"` is the registry/family identifier the orchestrator
236: kind: "smoke", // registry-key back-compat; no smoke ProbeResult is emitted
367: if (input.mode === "discovery") return `health:${slug}`;
368: if (input.key.startsWith("smoke:")) {
369: return `health:${input.key.slice("smoke:".length)}`;
```
New labels in dashboard:
```
cell-drilldown.tsx:60: { key: "e2e", label: "UI (Frontend)" },
cell-pieces.tsx:382: name="UI"
unified-cell.tsx:231: <TestBadge name="UI" level={model.d3} />
adaptive-legend.tsx:65: UI (Frontend): demo page renders in browser (Playwright)
```
Confirmation that the old labels are gone from render code:
```
$ grep -rn 'E2E (Demo)\|"Smoke"' cell-drilldown.tsx cell-pieces.tsx unified-cell.tsx adaptive-legend.tsx
(no matches — labels removed)
```
## Test results
- `pnpm exec vitest run` in `showcase/shell-dashboard`: **63 files, 1089
tests passed, 1 skipped**
- `pnpm exec vitest run` in `showcase/harness`: **129 files, 2760 tests
passed**
- `pnpm exec tsc --noEmit` in both: clean
Liveness driver tests include a regression guard (`never emits a
smoke:<slug> key`) so the smoke contract cannot be re-introduced
silently.
## Test plan
- [x] Dashboard vitest green (1089/1089)
- [x] Harness vitest green (2760/2760)
- [x] tsc --noEmit green in both packages
- [x] Visual cross-check: drilldown renders 6 rows (no Smoke), e2e row
labelled `UI (Frontend)`; cell strip shows `UI` badge in place of `E2E`;
legend shows explicit `Health` row + `UI (Frontend)` D3 line
The smoke probe was the same HTTP contract as /health on the same
service (200-OK JSON body), so every tick paid two HTTP calls for the
same liveness signal. Drop the /smoke GET + the smoke:<slug> primary
ProbeResult; the driver now emits health:<slug> as the primary and
agent:<slug> via writer side-emit (half the per-tick cost).
The driver's registry kind stays "smoke" so existing YAML configs and
orchestrator family wiring keep routing to this driver — the
emission key (health:<slug>) is the taxonomy contract that matters.
Dashboard:
- Drilldown: D3/e2e row labelled "UI (Frontend)" (was "E2E (Demo)");
Smoke row dropped (CellState.smoke field retained for back-compat).
- Cell badges: short label "UI" (was "E2E") in cell-pieces + unified-cell.
- Legend: D3 = "UI (Frontend): demo page renders in browser (Playwright)"
plus an explicit Health row at the top.
Tests updated to assert new labels:
- cell-drilldown.test.tsx: 6 dimensions (no Smoke); UI (Frontend) label.
- cell-drilldown.lazy-signal.test.tsx: drilldown-badge-ui--frontend- testid.
- cell-pieces.test.tsx + .signal-degrade.test.tsx: badge name "UI".
- unified-cell.test.tsx: mock-badge-UI.
- overlay-selector-integration.test.tsx: "UI" in place of "E2E".
- dashboard-color-matrix.test.tsx: badge: "UI" case names.
- liveness.test.ts: two-call contract (health + agent), regression guard
asserting no smoke:<slug> ever emitted.
Underlying probe key (e2e:<slug>/<feature>) preserved on PocketBase so
historical rows render correctly during the rename window.
## Summary
**URGENT hotfix.** Staging is widely red for D5/D6 multi-step demos.
Root cause confirmed via live Playwright DOM inspection on staging
(langgraph-typescript : beautiful-chat : pie chart prompt).
PR #5462's defect-2 fix uses `bubbleIndex = turnIndex - 1`, which
assumes one assistant bubble per turn. Multi-step agents (LangGraph,
Mastra, CrewAI, llama-index, ag2, etc.) emit 2-3+ bubbles per turn —
tool-call bubble + tool-render bubble + final-text bubble. The
strict-index read lands on an intermediate **tool-call** bubble whose
scoped-text selectors (`.cpk:prose`, `[data-message-content]`, `.prose`,
`p`) are **empty** (tool-call content lives in a sibling card div
outside the scoped cascade). The cascade-pollution guard correctly
returns `null` → the text-stable conjunct never converges →
`reason=text-unstable` timeout across virtually every multi-step demo
(beautiful-chat-*, gen-ui-*, shared-state-*, agentic-chat on most
backends).
## Live RED-GREEN proof (staging)
Standalone Playwright probe against
`showcase-langgraph-typescript-staging.up.railway.app/demos/beautiful-chat`,
prompt = `"pie chart of revenue distribution by category from the sample
sales data"`, captured BOTH the OLD-logic read (`bubbleIndex = turnIndex
- 1 = 0`) AND the NEW-logic read (`bubbleIndex = count - 1 = 2`) in ONE
atomic page evaluate after the turn settled:
```
TOTAL BUBBLE COUNT FOR TURN 1: 3 (true multi-step — tool-call + tool-render + final-text)
FINAL TIER: [data-testid="copilot-assistant-message"]
OLD LOGIC (bubbleIndex = turnIndex - 1 = 0):
count = 3
text = null
reason = all scoped selectors empty ← would time out: reason=text-unstable
NEW LOGIC (bubbleIndex = count - 1 = 2):
count = 3
text = "Here is the pie chart showing the revenue distribution by category from the sample sales data."
textLen = 94 ← gate settles cleanly
```
RED ≠ GREEN; new logic returns 94 chars of stable final-text where old
logic returns `null`. The bug and the fix are proven on real staging
DOM.
### Live DOM evidence
LangGraph TS beautiful-chat, prompt = "pie chart of revenue distribution
by category…":
```
count = 3 bubbles in [data-testid="copilot-assistant-message"]
bubble[0]: tool-call "query_data" — `.cpk:prose` exists but EMPTY (content in sibling <div class="my-1.5">)
bubble[1]: pie-chart card — `.cpk:prose` exists but EMPTY (content in sibling styled card div)
bubble[2]: final text "Here is the pie chart..." — `.cpk:prose` has 94 chars + toolbar mounted
```
The old pre-#5462 runner read `list[last]` (always the last bubble —
always bubble[2] above). The PR moved to `turnIndex - 1`, which is
bubble[0] (the tool-call) for any multi-step turn → broken.
## Fix
Read the LAST bubble in the matched cascade tier (`count - 1`) instead
of a fixed strict index. Defect-2 protection (don't read a leftover
bubble from a prior turn) is preserved via an explicit **pre-submit
baselineCount sentinel snapshot** in the runner — the gate now requires
`count > baselineCount` rather than `count > turnIndex - 1`. This
correctly rejects stale prior-turn bubbles **without assuming 1 bubble =
1 turn**.
### Changes
- `assistant-message-count.ts`: add `readCascadeStateLast(page)` — same
atomic single-evaluate contract as `readCascadeState`, but resolves the
bubble index internally as `count - 1`.
- `conversation-runner.ts`:
- Snapshot `baselineCount = countAssistantMessages(page)` BEFORE
`sendTurnMessage()` each turn.
- Pass through to `waitForTurnComplete` via new `baselineCount` option
(defaults to `turnIndex - 1` to preserve unit-test fake compatibility).
- `waitForTurnComplete` reads via `readCascadeStateLast` and gates on
`domOk = count > baselineCount`.
- Cold-start retry re-snapshots `baselineCount` after `page.reload()`.
- Returned `bubbleIndex` in the success ctx = `count - 1` (matches the
bubble whose text settled the gate).
- `dom-missing` reason classifier updated to `countFinal <=
baselineCount`.
## Why this doesn't regress defect-2
The original defect-2 (un-turn-scoped bubble selection) leaked a later
turn's bubble into THIS turn's assertions because
`readLastAssistantText` read `list[last]` GLOBALLY with no per-turn
baseline. The new fix re-introduces the "read last" semantics but adds
the **per-turn pre-submit baseline count snapshot** — the gate only
advances once a NEW bubble has appeared (`count > baselineCount`), so we
cannot read a leftover from a prior turn, and the "last" we return is
the last NEW bubble (which for multi-step turns is the final-text bubble
whose scoped text actually settles).
## Test plan
- [x] **Live RED-GREEN staging probe** (see above) — confirms OLD logic
= null and NEW logic = 94-char final-text on a real 3-bubble multi-step
turn.
- [x] All 72 `conversation-runner.test.ts` +
`assistant-message-count.test.ts` unit tests pass (default
`baselineCount = turnIndex - 1` keeps existing count-progression scripts
working).
- [x] Full harness vitest suite: 2758/2758 green across 129 test files.
- [x] `tsc --noEmit` clean on `showcase/harness`.
- [x] `oxfmt --check` clean on changed files.
- [x] `oxlint` clean on changed files.
- [ ] Staging D5/D6 sweep after merge confirms multi-step demos
(beautiful-chat, gen-ui, shared-state, langgraph/mastra/crewai
agentic-chat) recover from `text-unstable` red.
## Scope
Strictly the two files above. No unrelated changes. No refactors.
PR #5462's defect-2 fix assumed 1 bubble per turn (bubbleIndex =
turnIndex - 1). Multi-step agents (LangGraph, Mastra, CrewAI) emit
2-3+ bubbles per turn — tool-call + tool-render + final-text — so
the strict-index read lands on an intermediate tool-call bubble
whose scoped-text selectors (`.cpk:prose`, `[data-message-content]`,
`p`) are EMPTY (its content lives in a sibling card div outside the
scoped cascade). The cascade-pollution guard returns null forever
and the text-stable conjunct times out with `reason=text-unstable`
across virtually every multi-step demo (beautiful-chat-*, gen-ui-*,
shared-state-*, agentic-chat on LangGraph/Mastra/CrewAI/etc.).
Live DOM evidence from staging (langgraph-typescript :
beautiful-chat : pie chart prompt):
count = 3 bubbles
bubble[0]: tool-call "query_data" — `.cpk:prose` empty
bubble[1]: pie chart card — `.cpk:prose` empty
bubble[2]: final text "Here is the pie chart..." — 90 chars
Fix: read the LAST bubble in the matched cascade tier (count - 1)
instead of a fixed strict index. Defect-2 protection (don't read a
leftover bubble from a prior turn) is preserved via an explicit
pre-submit `baselineCount` sentinel snapshot in the runner; the
gate now requires `count > baselineCount` rather than
`count > turnIndex - 1`, which correctly rejects stale bubbles
without assuming 1 bubble = 1 turn.
Changes:
- assistant-message-count.ts: add `readCascadeStateLast(page)` —
same atomic single-evaluate contract as `readCascadeState` but
resolves the bubble index internally as `count - 1`.
- conversation-runner.ts: snapshot `baselineCount` before
sendTurnMessage; pass through to `waitForTurnComplete` via new
`baselineCount` option (defaults to `turnIndex - 1` for unit-
test fakes); `waitForTurnComplete` reads via
`readCascadeStateLast` and gates on `count > baselineCount`.
Cold-start retry re-snapshots `baselineCount` after page.reload.
Test plan:
- All 72 conversation-runner + assistant-message-count unit tests
pass (default baselineCount preserves count-progression scripts).
- Full harness suite 2758/2758 green, tsc clean, oxfmt/oxlint
clean on changed files.
PR #5458 (c81b361f1) added a fail-loud gate that refuses harness boot
in any deployable mode (NODE_ENV != "test") without SHARED_SECRET or
SHARED_SECRET_PREV: POST /webhooks/deploy is only registered when
webhookSecrets.length > 0 (src/http/server.ts:119 +
loadWebhookSecrets in src/orchestrator.ts). The local docker-compose
stack inherits NODE_ENV=production from the harness image and does
not (and should not) set SHARED_SECRET, so every local D5/D6 verify
run via bin/showcase test --d5/--d6 was crashing the harness in a
restart loop with FATAL-CONFIG.
Fix: add HARNESS_ALLOW_NO_SECRET=1 (the documented local-dev escape
hatch — explicitly referenced in the FATAL-CONFIG message itself) to
both harness services in showcase/docker-compose.local.yml. Inline
comments explain the rationale and pin the relevant source locations.
Prod impact: NONE. Railway sets SHARED_SECRET explicitly via env on
every harness service, so loadWebhookSecrets sees a real secret,
registers POST /webhooks/deploy, and never reads
HARNESS_ALLOW_NO_SECRET. This change only affects the local
docker-compose stack.
Verified locally: showcase-iso1-harness boots cleanly (Up healthy),
the expected warn-level webhook-auth-bypass log fires
(escapeHatch:true), worker registers, scheduler starts, and a real
d6:langgraph-typescript job claims successfully — confirming the
gate fires only in deployable contexts.
## Summary
Convert `useInterrupt::resolve` and the demo-local
`useHeadlessInterrupt::resolve` callbacks from fire-and-forget into
`async` + `return await copilotkit.runAgent(...)`, so callers receive a
Promise that settles when the resume run settles.
## Mechanism
Pre-fix: `resolve(...)` called `copilotkit.runAgent({...})` without
`await` and without `return` — the arrow returned `undefined`. The
showcase harness DOM-settle check on the assistant confirmation bubble
timed out because consumers had no handle to sequence against the resume
run's settle.
Post-fix:
- Framework `useInterrupt::resolve` is `async`, `return await`s
`runAgent`, wraps in try/catch + `setPendingEvent(null)` +
`console.error` + rethrow on rejection.
- `onRunFailed` now also `setPendingEvent(null)` — symmetric with
`onRunStartedEvent`, prevents stuck popups on run failure.
- `InterruptHandlerProps.resolve` / `InterruptRenderProps.resolve` typed
`() => Promise<RunAgentResult>` (was `() => void`).
- The same fix-pattern applied to 13 demo-local `useHeadlessInterrupt`
implementations across integrations: ag2, agno, claude-sdk-typescript,
crewai-crews, langgraph-{fastapi,python,typescript}, langroid,
llamaindex, mastra, pydantic-ai, spring-ai, strands.
## Verification
- **Unit:** new tests `resolve returns a Promise that settles when
runAgent settles (RESUME-PATH)` + `resolve rejects when runAgent
rejects, logs the failure, and clears pending (RESUME-PATH-REJECT)`.
Red-then-green on baseline. 22/22 react-core hook tests passing.
- **Backplane (end-to-end):** showcase control plane on langgraph-python
via `./bin/showcase test langgraph-python:gen-ui-interrupt --d5 --direct
--verbose` (GREEN, 70.1s, was RED-RESUME-PATH baseline) and
`:interrupt-headless --d5 --direct --verbose` (GREEN, 7.7s, was
RED-RESUME-PATH baseline). Built Docker artifact + aimock D6 fixture
replay — staging-equivalent path.
## CR
4 review rounds × 7 unbiased agents = 28 agent-runs. 2 Procedure 3
audits (3 promotions in round 1 → Fix C + D; 0 promotions in round 2).
Convergence achieved per agent-confirmed `bucket (a) empty` + Procedure
3 zero `PROMOTE_TO_A`.
## Out of scope (deferred follow-ups)
- **No manifest unquarantine.** A prior version of this branch flipped
`interrupt-headless` and `gen-ui-interrupt` from
`not_supported_features` to `features` across 12 manifests, but Phase 4
backplane validation across 8 PATCHED integrations × 2 demos surfaced
(a) `langgraph-fastapi:gen-ui-interrupt` still RED-RESUME-PATH because
the framework patch doesn't propagate from workspace to the
integration's Docker image without a published `@copilotkit/react-core`
release, and (b) a separate popup-mount RED-RUNTIME-OTHER cluster across
6 integrations that's pre-existing rot unrelated to this PR. Manifest
unquarantine + `@copilotkit/react-core` version bump are deferred to a
follow-up `copilotkit-release` PR after this fix lands.
- **No version bump.** Release follows in a separate PR.
- **Bucket (c) follow-up backlog** (audit verdicts: 11 STAY_IN_C):
- JSON.parse guarding in showcase demos (pre-existing, all 13)
- Type-design quirks in `InterruptHandlerProps.result` (non-nullable in
type, nullable in runtime)
- `renderInChat` dynamic toggle leak (pre-existing edge case)
- JSDoc clarity ("symmetric with onRunFailed" wording, `@typeParam
TRenderInChat`, `event.value` "any" → "unknown")
- StrictMode test coverage gap
- 13× duplicate `useHeadlessInterrupt` → bucket (d): consolidate into a
shared module in a future PR
## Files changed
16 files, +486/-154:
- `packages/react-core/src/v2/hooks/use-interrupt.tsx`
- `packages/react-core/src/v2/hooks/__tests__/use-interrupt.test.tsx`
- `packages/react-core/src/v2/types/interrupt.ts`
- 13×
`showcase/integrations/<integ>/src/app/demos/interrupt-headless/page.tsx`
## Test plan
- [x] react-core unit tests (22/22 passing including new RESUME-PATH +
RESUME-PATH-REJECT regression tests)
- [x] Backplane D5 probe on langgraph-python (gen-ui-interrupt +
interrupt-headless GREEN end-to-end via the showcase control plane)
- [ ] CI green on this PR
## What
Activates Slack's agent-grade APIs as the **default** experience for
`@copilotkit/bot-slack`, with zero config and safe degradation.
Implements the Notion spec *"Slack Agent APIs in bot-slack — assistant
pane + native streaming by default"* (Rev 2).
### Assistant pane ("Agents & AI Apps")
- Opening the pane greets the user + shows tappable prompt chips; each
pane conversation is its own thread (replies stay in-thread, instead of
leaking to a merged flat DM).
- While the agent runs, native composer status
(`assistant.threads.setStatus`: "is thinking…", "is using `tool`…")
replaces placeholder/`🔧` messages.
- Pane threads are auto-titled from the first message.
- Customize via the new `assistant` option; `assistant: false` disables
pane handling. Apps without the Agents toggle behave exactly as before
(pane machinery dormant).
### Native streaming
- Replies stream via `chat.startStream` / `appendStream` / `stopStream`
(raw markdown — real tables / fenced code render natively) wherever a
thread exists.
- Flat DMs and workspaces without the streaming API fall back to the
legacy `chat.update` transport **automatically** — the first
`startStream` failure marks the workspace legacy and replays via legacy.
Opting in can never break a bot. `streaming: "legacy"` forces the old
transport.
### Portable engine surface (`@copilotkit/bot`)
- `bot.onThreadStarted` lifecycle handler + `IncomingThreadStart` sink
event.
- Capability-gated `thread.setSuggestedPrompts` / `thread.setTitle` (the
shipped `postFile` gating pattern).
- Two `SurfaceCapabilities` flags + optional `PlatformAdapter` methods.
All degrade gracefully on surfaces without support.
## Files
- **New:** `bot-slack/src/assistant.ts` (Bolt `Assistant` middleware ⇄
engine sink), `bot-slack/src/native-stream.ts` (`NativeMessageStream`,
same `append/finish` contract as `MessageStream`, legacy fallback).
- **Engine:** `platform-adapter.ts`, `thread.ts`, `create-bot.ts`,
`index.ts` (+ `bot-ui` `Thread` type).
- **Adapter:** `adapter.ts` (options/capabilities/stream
branch/methods), `slack-listener.ts` (one-line no-double-delivery
guard), `event-renderer.ts` (pane status mode + native text transport).
- **Example/docs:** `examples/slack` manifest (`assistant_view` + scope
+ events) & dev-ex, package READMEs/ARCHITECTURE, `slack.mdx`.
- **Hooks:** `chore(hooks)` makes the `lint-fix` script
double-quote-free (it was breaking `sh -c "…"` on Windows/lefthook
2.1.1).
## Tests / verification (CI-runnable)
- New unit tests: `native-stream` (cadence, 12k continuation, fence
re-open, first- **and** continuation-`startStream` fallback),
`assistant` (defaults-before-onThreadStarted ordering, thread-scoped
turn, auto-title), listener no-double-delivery, renderer pane-status
mode, engine capability gating.
- Verified locally: `@copilotkit/bot` 31 tests, `@copilotkit/bot-slack`
199 tests, `bot-ui` tests; `tsc` clean (bot + bot-slack); build clean;
oxlint/oxfmt clean; publint/attw pass.
## ⚠️ E2E not covered here
The §8 end-to-end flows need a **real Slack workspace** with the Agents
toggle (open pane → chips → streamed reply w/ live status; channel
mention → native markdown; kill `startStream` → legacy fallback;
non-agent app → shipped behavior). `examples/slack` is the vehicle.
Three items remain to confirm against a live workspace (spec §7 spikes):
the `isAssistantThread` runtime detection, the exact `assistant_view`
manifest acceptance, and `stopStream` finalize tolerance.
Spec:
https://app.notion.com/p/copilotkit/Slack-Agent-APIs-in-bot-slack-assistant-pane-native-streaming-by-default-spec-37b3aa38185281f5948dc6a665064d04
- d5-* unit tests: update page fakes / assertion-call shapes to pass
the ctx bridge ({bubbleIndex, text}) instead of querying the live
bubble list
- d6-all-pills.test.ts: update page-fake dispatch + cascade-state
expectations to match the atomic readCascadeState return shape
({count, text} or null) — picks up the cross-tier consistency
guarantee in the unit layer
- d6-all-pills.ts: wire installPrePaintFromEnv and
attachSseInterceptor PRE-goto in both defaultLauncher and pooled
launcher paths, so first-paint state is deterministic and the SSE
counter is armed before any page navigation
- captureDiagnostics rewired through the Node-side helper
(readCascadeState) — single page.evaluate per call, no
within-snapshot tier drift
- Consume BUBBLE_RACE_MESSAGES override for defect-4 repro flows
- d5-gen-ui-custom, d5-gen-ui-open-advanced, d5-mcp-apps,
d5-subagents, d5-tool-rendering-default-catchall: wrap assertion
bodies in ctx bridge so they receive {bubbleIndex, text} for the
turn under test, replacing turn-leaky list[last] retrieval
- _gen-ui-shared: add readAssistantTextAt(page, bubbleIndex)
adapter; remove readLastAssistantText (no longer turn-scoped)
This is the read-side counterpart to the atomic readCascadeState
on the helper side: probes ask for the bubble at the runner's
chosen index instead of guessing.
Each integration's interrupt-headless demo defines a local useHeadlessInterrupt
hook around the framework useInterrupt. Slot-2 originally identified 8
quarantined integrations (claude-sdk-typescript, langgraph-{fastapi,python,
typescript}, langroid, pydantic-ai, spring-ai, strands); review-round
follow-ups extended the sweep to llamaindex, mastra, ag2, agno, and
crewai-crews (5 more integrations sharing the same byte-identical hook).
The demo-local resolve() previously fire-and-forgot copilotkit.runAgent(...)
via `void runAgent(...).catch(() => {})`. Mirroring the framework fix:
- Make resolve async, return await copilotkit.runAgent(...).
- Use a pendingRef so resolve has stable identity (drop pending from
useMemo deps).
- Type signature: resolve: (response: unknown) => Promise<unknown>.
- Wrap in try/catch + setPending(null) + console.error + rethrow,
symmetric with the framework hook.
- onRunFailed also setPending(null).
13 integrations patched byte-identically.
## Problem
The \`Showcase: Verify Deploy\` workflow runs \`notify-harness\` after
every merge to \`main\` and POSTs an HMAC-signed payload to
\`\${SHOWCASE_HARNESS_URL}/webhooks/deploy\`. That POST has been
returning **404** on every main deploy since at least 2026-06-12 (5+
consecutive failures), so the harness dashboard has not reflected any
recent deploys.
**Root cause:**
1. \`POST /webhooks/deploy\` is only registered when
\`webhookSecrets.length > 0\` — gate at
\`showcase/harness/src/http/server.ts:119\`.
2. The CP \`buildServer\` call in \`runControlPlane\`
(\`showcase/harness/src/orchestrator.ts\` ~line 2961 pre-fix)
**omitted** both \`webhookSecrets\` and \`metrics\`. The public Railway
host running the CP role never mounted the route — every notify-harness
POST returned 404.
3. The FATAL-CONFIG guard on missing \`SHARED_SECRET\` (pre-fix
\`orchestrator.ts:782-790\`) only fired when \`NODE_ENV ===
"production"\`. Any deploy whose NODE_ENV was unset, set to something
other than the literal \`"production"\`, or set after the check,
silently booted with \`webhookSecrets=[]\` and no fatal error.
## Fix
**Part 1: Register the route on the CP path.**
- Lifted the env→secrets loader into a new exported helper
\`loadWebhookSecrets()\` (orchestrator.ts ~175-209) so both boot paths
consume the same predicate.
- Worker \`boot()\` (~line 836) calls it; CP \`runControlPlane\` (~line
2421) now also calls it.
- The CP \`buildServer\` call (~line 3047) now passes \`webhookSecrets\`
and \`metrics\` alongside the existing \`bus\`/\`probes\`/\`fleetRuns\`
wiring — same shape as the worker call.
**Part 2: Tighten the FATAL-CONFIG predicate.**
- New predicate (inside \`loadWebhookSecrets\`): throw unless EITHER a
secret is set OR \`NODE_ENV === "test"\` OR \`HARNESS_ALLOW_NO_SECRET
=== "1"\` (narrow escape hatch for local dev).
- Thrown error message names the env-vars, the gate location
(\`src/http/server.ts:119\`), the silent-404 symptom, and the escape
hatches — so an operator reading the deploy log can recover without
spelunking.
## Test coverage
Four new red-green tests in \`orchestrator.test.ts\` under \`B2:
/webhooks/deploy registered on CP + fail-loud on missing
SHARED_SECRET\`:
- **Test A** — CP boot registers \`POST /webhooks/deploy\` when
\`SHARED_SECRET\` is set. Probe with no HMAC headers: pre-fix returned
404, post-fix returns 401 (route mounted, HMAC reject path fires).
RED→GREEN.
- **Test B** — Worker boot still registers the route (regression guard).
Pre-fix and post-fix both 401.
- **Test C** — FATAL-CONFIG fires when \`SHARED_SECRET\` unset AND
\`NODE_ENV=development\`. Pre-fix resolved silently, post-fix rejects
with \`/FATAL-CONFIG.*SHARED_SECRET/\`. RED→GREEN.
- **Test D** — Boot succeeds when \`SHARED_SECRET\` unset AND
\`NODE_ENV=test\` (escape hatch). Pre-fix and post-fix both green.
Existing R5-G4 D5 test (orchestrator.test.ts:1592) updated to match the
new error message via \`/FATAL-CONFIG.*SHARED_SECRET/\` — same throw,
broader match.
## Verification
- Full harness suite: **2719/2719 green** across 128 files
- \`tsc --noEmit\`: clean
- RED capture pre-fix:
- Test A: \`AssertionError: expected 404 to be 401\`
- Test C: \`AssertionError: promise resolved "{ port: ..., bus: { ... },
... }" instead of rejecting\`
## Deploy note
Requires a CP-role redeploy on the public Railway host for the new route
registration to take effect — once redeployed, the next
\`notify-harness\` POST will land on the harness dashboard.
## Test plan
- [x] B2 Tests A-D pass post-fix (red→green captured)
- [x] Full harness vitest suite green (2719/2719)
- [x] \`tsc --noEmit\` clean
- [ ] Post-deploy: confirm next \`Showcase: Verify Deploy\` run shows a
2xx notify-harness POST (not 404)
- [ ] Post-deploy: confirm harness dashboard reflects the next deploy
R3-F1 (OPS_TRIGGER_TOKEN hoist):
Extract loadOpsTriggerToken() and hoist its invocation to the top of both
boot() and runControlPlane(), alongside loadWebhookSecrets() and
loadPocketbaseUrl(). Pre-fix the empty/whitespace check fired AFTER pb /
bus / scheduler / writer / S3 uploader allocations, so a typo'd
`OPS_TRIGGER_TOKEN=` allocated expensive resources before throwing. Now
all three fail-loud config predicates fire at the top, before any
allocation needing teardown. Behaviour is unchanged for valid tokens
(trimmed via R3-A.5 contract) and for the unset case (router omitted with
info log).
R3-F2 (sync-throw also emits deploy.writer.failed):
subscribeDeployResults() pre-fix only emitted `deploy.writer.failed` on
writer.write() promise rejection — a synchronous throw inside
deployEventToProbeResult() (malformed event, type drift) bypassed the
catch and we lost both the log AND the bus emit. Wrap the sync mapping
in try/catch and mirror the rejection path so alert rules / metrics
subscribers observe sync throws the same way they observe async write
failures.
Tests:
- 5 unit tests for loadOpsTriggerToken: undefined / empty-string fail-loud
/ whitespace-only fail-loud / R3-A.5 trim contract / verbatim value.
- 1 test for the sync-throw path: vi.doMock deployEventToProbeResult to
throw, assert bus emits deploy.writer.failed with the err message AND
writer.write was NOT called. Red-green verified locally (test fails
without the try/catch).
Suite: 2731 passing (+6 new). Typecheck clean.
PR #5459 added userMessage matchers ("forecast for Tokyo" + "current price
of AAPL") on the catchall fixtures, which flipped most integrations from
red to green. Six cells stayed red on staging because the FIXTURE SHAPE
itself was broken on the matched fixtures, not just the userMessage key.
Failing cells (all custom-catchall except strands default):
d5:agno/tool-rendering-custom-catchall
d5:llamaindex/tool-rendering-custom-catchall
d5:langroid/tool-rendering-custom-catchall
d5:claude-sdk-python/tool-rendering-custom-catchall
d5:strands/tool-rendering-custom-catchall
d5:strands/tool-rendering-default-catchall
PocketBase confirms the failure mode: turn 1 (Tokyo) completes; turn 2
(AAPL) times out at 30s or renders the wrong content. E.g. agno:
errorDesc: "timeout: assistant did not respond within 30000ms"
failure_turn: 2
turns_completed: 1 / 2
Root cause: the failing tool-emit fixtures used `hasToolResult: false`
(or `toolName: "<tool>"`) as their gate. aimock's hasToolResult check is
`messages.some(m => m.role === 'tool')` over the WHOLE thread — so once
turn 1's Tokyo tool result lands in the conversation, hasToolResult is
permanently true and `hasToolResult:false` can never match turn 2 → no
fixture → 30s timeout. `toolName:get_stock_price` likewise fails when an
integration backend doesn't forward the tool definition on turn 2.
The same trap is documented in
showcase/aimock/d6/langgraph-python/tool-rendering.json:
"_comment": "Gated on toolName:get_stock_price rather than
hasToolResult:false. The D5 tool-rendering-custom-catchall probe runs
'weather in Tokyo' first, which leaves a get_weather tool result in
the thread; the aimock router implements hasToolResult as
messages.some(m=>m.role==='tool'), so hasToolResult is permanently
true on the AAPL turn and a hasToolResult:false gate could never
match (→ no_fixture_match → 503 → 30s timeout)."
Fix: align all 6 files to the canonical pattern used by mastra/spring-
ai/built-in-agent on this probe:
1. toolCallId-keyed narration fixture FIRST
2. tool-emit fixture SECOND with ONLY `userMessage` + `context`
(no hasToolResult / toolName gate)
The toolCallId narration uses aimock's
`messages[last].role === 'tool' && tool_call_id === ...` check, so it
correctly wins on iteration 2 (post-tool-result) without being affected
by older turns' tool results. The tool-emit fixture matches turn 1 (last
message is user) and re-emits only when the narration above hasn't
matched.
Strands' two cells additionally needed REORDERING — they had the tool-
emit fixture before the toolCallId fixture, defeating first-match-wins.
Strands' staging backend was also returning 502 during testing; once it
recovers, the corrected fixtures should let the probe pass. The fixture
changes are necessary but may not be sufficient for strands if backend
remains down.
No probe-side, harness, or backend changes — pure fixture-content
alignment. Six fixture files modified; line totals: -87 / +68.
CB-1 (Slot 2 #22): convert B2 deploy-webhook tests to vi.stubEnv with
`vi.unstubAllEnvs()` in afterEach instead of manual process.env mutation
+ try/finally restoration. Test runner now guarantees restoration and the
test bodies are shorter / less error-prone (CR Slot 2 #22).
CB-2 (Slot 2 #28): when `loadWebhookSecrets`' escape hatch fires with a
real-looking NODE_ENV (anything except "test"), log at `warn` instead of
`info` so a production typo (NODE_ENV=staging + HARNESS_ALLOW_NO_SECRET=1)
is visible in dashboards / log alerting. Pure local-dev (NODE_ENV=test)
stays at info level so a normal unit-test boot doesn't spam warnings.
CB-3 (Slot 4 #17): clarify `loadWebhookSecrets`' docstring + bypass-log
message — "set" is ambiguous (empty string would qualify pre-fix); use
"non-empty" to match the actual predicate.
R1-F2 (bucket b, defensive ordering): hoist fail-loud config validation
to the TOP of boot() and runControlPlane() — BEFORE any pb client, queue,
bus, scheduler, writer, S3 uploader, fleet-health, or aggregator
allocations. Pre-fix `loadWebhookSecrets` lived AFTER the entire
scheduler+writer (and CP queue+aggregator) setup, so a misconfigured boot
allocated expensive resources and mounted file watchers before throwing.
R1-F3 (bucket b, predicate symmetry): broaden the POCKETBASE_URL fail-loud
predicate to match SHARED_SECRET semantics. Pre-fix the POCKETBASE_URL
guard only fired on `NODE_ENV === "production"`, while `loadWebhookSecrets`
fired unless NODE_ENV=test or an explicit escape hatch — so staging /
unset / "development" deploys silently bound to http://localhost:8090.
Extracted `loadPocketbaseUrl(logger)` with the same test-or-escape-hatch
predicate as `loadWebhookSecrets`. A new `HARNESS_ALLOW_NO_PB_URL=1` env
flag mirrors `HARNESS_ALLOW_NO_SECRET=1` for local dev. Worker
`boot()` and the CP's `resolveFleetPbConfig` both call the helper.
Tests:
- HF13-A2 production-only assertion broadened (SHARED_SECRET set so the
test specifically exercises the POCKETBASE_URL guard).
- New R1-F3 tests cover NODE_ENV=development without POCKETBASE_URL
(throws) and the HARNESS_ALLOW_NO_PB_URL=1 escape hatch (succeeds).
R1-F1 (bucket a, load-bearing): the CP boot path now subscribes to
`deploy.result` events through its status writer. Pre-fix, B2 mounted POST
/webhooks/deploy on the CP role but only the worker boot path had a
`bus.on("deploy.result", ...)` listener — signed POSTs to the CP host
returned 202, the bus event fired with no subscriber, and the deploy-overall
dashboard row never landed.
Extracted the handler into a shared `subscribeDeployResults(bus, writer)`
helper exported from orchestrator.ts so both boot paths share the IDENTICAL
logic. Worker boot keeps its inline pattern (busUnsubs.push) so teardown is
unchanged; CP keeps the subscription alive for the lifetime of the bus.
Red-green tests in src/orchestrator.test.ts under "R1-F1" assert that
emitting deploy.result on the returned handle's bus drives one write keyed
"deploy:overall" through a stubbed status writer — once for runControlPlane
(the new path) and once for boot() (regression guard for the helper
extraction).
The catchall userMessage rename (#5453) only renamed STALE strings to 'forecast for Tokyo'. Integrations whose catchall fixtures used PILL-ALIGNED userMessages ('What's the weather in San Francisco?', 'Find flights from SFO to JFK.', 'Chain a few tools in this single turn', etc.) had no 'forecast for Tokyo' matcher to begin with — so the D5 catchall probe still fell through to live LLM and the cells stayed red on staging.
This PR ADDS the canonical 'forecast for Tokyo' (default+custom catchall) and 'What's the current price of AAPL?' (custom catchall only) emit+narrate fixture pairs to each integration that was missing them. Existing pill-aligned fixtures are preserved (pure prepend at the start of the fixtures array; first-match-wins means new fixtures match the D5 probe inputs without colliding with existing pill prompts).
Integrations touched:
- langgraph-typescript: both default + custom
- langgraph-fastapi: both default + custom
- google-adk: custom only (default already green)
- ms-agent-dotnet: custom only (default already green)
After merge, fleet-cp's e2e-deep probe will re-run within ≤6h staleness window and flip cells GREEN.
The 'Showcase: Verify Deploy' workflow's notify-harness step has been
POSTing to /webhooks/deploy and getting 404 on every main deploy since
at least 2026-06-12 (5+ consecutive failures). Root cause:
1. POST /webhooks/deploy is only registered when
webhookSecrets.length > 0 (gate at src/http/server.ts:119).
2. The CP buildServer call in runControlPlane omitted BOTH
webhookSecrets and metrics — so the public Railway host running
the CP role never mounted the route. The worker boot path had
them, but it isn't the host receiving notify-harness POSTs.
3. The FATAL-CONFIG guard on missing SHARED_SECRET was gated on
NODE_ENV === 'production' — any deploy with NODE_ENV unset or set
to something other than the literal 'production' silently shipped
with webhookSecrets=[] and no fatal error.
Part 1 — wire webhookSecrets + metrics into the CP buildServer call.
Lifted the env→secrets loader into a new exported helper
'loadWebhookSecrets()' (~orchestrator.ts:175-209) so BOTH boot paths
consume the same predicate. The worker 'boot()' call (~line 836) and
the CP 'runControlPlane' call (~line 2421) both invoke it, and the CP
buildServer call (~line 3047) now passes webhookSecrets+metrics
alongside the existing bus/probes/fleetRuns wiring.
Part 2 — tighten the FATAL-CONFIG predicate.
The new predicate throws unless EITHER a secret is set OR NODE_ENV ===
'test' OR HARNESS_ALLOW_NO_SECRET === '1' (narrow escape hatch for
local dev). The thrown error message names the env-vars, the gate
location ('src/http/server.ts:119'), the silent-404 symptom, and the
escape hatches, so an operator reading the deploy log can recover
without spelunking.
Updated the existing R5-G4 D5 regex (orchestrator.test.ts:1592) since
the error message changed; covered by the new B2 Tests A/C/D below.
Tests: 4 new red-green tests in 'B2: /webhooks/deploy registered on CP
+ fail-loud on missing SHARED_SECRET':
- Test A (CP /webhooks/deploy registration): RED 404 → GREEN 401.
- Test B (worker /webhooks/deploy regression): GREEN 401.
- Test C (FATAL-CONFIG when NODE_ENV=development): RED resolves →
GREEN rejects with /FATAL-CONFIG.*SHARED_SECRET/.
- Test D (NODE_ENV=test escape hatch): GREEN, no throw.
Full harness suite: 2719/2719 green. typecheck: clean.
Requires a CP-role redeploy on the public Railway host for the new
route registration to take effect.
Adds four red-green tests for the notify-harness 404 follow-up:
A. CP boot registers POST /webhooks/deploy when SHARED_SECRET is set
— RED pre-fix (404, route not mounted on CP); GREEN post-fix (401,
the route's own HMAC reject path fires on an unsigned POST).
B. Worker boot still registers POST /webhooks/deploy (regression
guard) — already green; pins the unchanged behavior.
C. boot throws FATAL-CONFIG when SHARED_SECRET unset AND NODE_ENV is
non-test — RED pre-fix (current guard fires only on
NODE_ENV='production'); GREEN post-fix once the predicate is
tightened.
D. boot succeeds when SHARED_SECRET unset AND NODE_ENV='test'
(test-mode escape hatch) — already green; pins the escape hatch.
Red capture (pre-fix run):
- Test A: AssertionError: expected 404 to be 401
- Test C: AssertionError: promise resolved instead of rejecting
## Summary
- `showcase/docker-compose.local.yml:262` hardcodes
`LOCAL_SERVICES_JSON` to `showcase-langgraph-python` (intentional N=1
demo default).
- `showcase/bin/showcase test <slug> --d6 --isolate` was inheriting that
value verbatim into the iso1 stack, so the iso1 control-plane discovered
`showcase-langgraph-python` instead of `showcase-<requested-slug>`.
- Result: iso1 probes targeted the wrong service. Visible in iso1
harness logs as `discovery.railway-services.local-injection count:1
names:["showcase-langgraph-python"]` regardless of CLI arg.
- This PR teaches the iso1 compose generator (`apply_isolation` in
`showcase/scripts/cli/_common.sh`) to inject a per-slug
`LOCAL_SERVICES_JSON` override built from the slug's manifest.yaml demos
list. Fallback to `["agentic-chat"]` if manifest absent.
## Scope
Two files, ~55 LOC added: `showcase/scripts/cli/cmd-test.sh` (pass slug
arg), `showcase/scripts/cli/_common.sh` (inject regex sub in python
rewriter). Bash + embedded python only; no TypeScript changed.
## Verification
- **Local `--isolate` discovery confirmed correct:**
`showcase/bin/showcase test ms-agent-python --d6 --isolate` — iso1
harness log now shows `discovery.railway-services.local-injection
names:["showcase-ms-agent-python"]` (was `showcase-langgraph-python`
before this PR).
- Persistent stack default behavior (langgraph-python N=1) preserved via
fallback when no slug arg provided.
- **Heredoc hardening scope:** commit edc77f809 moves `$slug` from
bash-interpolation into the python rewriter to an env var. This is a
slug-only carve-out — `$slug` is the only value that originates from the
CLI arg path. The other bash-interpolated `$VAR`s embedded in the
heredoc (`$PORTS_FILE`, `$COMPOSE_FILE`, `$name`, `$SHOWCASE_ROOT`,
`$ISOLATE_PORT_OFFSET`) remain script-internal: each is constructed
inside `_common.sh` from validated sources (manifest reads, computed
offsets, fixed roots), not from user input, and CR Round 1 + Round 2
slot 5 both verified they are not user-tainted. A broader
env-var-pass-all-vars refactor would be a separate concern and is out of
scope for this PR.
## Out of scope
- Worker heartbeat clock-skew issue (`fleet.health.worker-unhealthy
lastHeartbeatAt N min stale`) is a separate bug; not touched.
- `buildLocalServicesJson` in `cli/control-plane-run.ts` has a similar
comment-vs-code disagreement (JSDoc says "filter", code returns env
verbatim) — flagged but not auto-fixed, as iso1 override is the correct
insertion point.
- Generalized env-var-pass for all heredoc-embedded $VARs (see
Verification) — separate refactor concern.
## Summary
- D5 e2e-deep probes for `tool-rendering-{default,custom}-catchall` send
`"forecast for Tokyo"` as the test input (see
`showcase/harness/src/probes/scripts/d5-tool-rendering-{default,custom}-catchall.ts`).
- aimock uses substring match on `userMessage`.
- The catchall fixtures on main had a stale `"check Tokyo weather
forecast"` string that couldn't substring-match the probe input →
fixture miss → probe falls through to live LLM → CV ✗ red D4 on the
dashboard.
- This PR renames the userMessage to the canonical `"forecast for
Tokyo"` across 28 catchall fixture files in 16 integrations.
## Scope
28 files × ~2 userMessage occurrences each = 67 line changes. **No code,
no agent, no page.tsx changes.** Pure fixture-data alignment.
Integrations covered (default-catchall and/or custom-catchall): ag2,
agno, built-in-agent, claude-sdk-python, claude-sdk-typescript,
crewai-crews, google-adk, langgraph-python, langroid, llamaindex,
mastra, ms-agent-dotnet, ms-agent-python, pydantic-ai, spring-ai,
strands.
## Commit history note
Commit b1f19bdc8 changes the rename target from the original `"weather
in Tokyo"` (commit 388c69e68) to `"forecast for Tokyo"`. This was a CR
Round 1 catch: `"weather in Tokyo"` was a substring of the chain pill
prompt (`"weather forecast chain in Tokyo"` / similar), which would have
caused the catchall fixture to incorrectly match chain-pill traffic.
`"forecast for Tokyo"` has no such substring collision with any other
probe input.
## Verification
- **Live red-green proof** was performed on the prior `"weather in
Tokyo"` rename (commit 388c69e68) against `ms-agent-python` via
`showcase/bin/showcase test ms-agent-python --d6 --isolate`:
- Before (origin/main): catchall featureTypes red (fixture miss → live
LLM → flaky)
- After: `d6:ms-agent-python/tool-rendering-default-catchall=green`,
`d6:ms-agent-python/tool-rendering-custom-catchall=green`
- **The current HEAD's `"forecast for Tokyo"` rename (b1f19bdc8) has NOT
been re-run live.** It is verified by static analysis only: substring
math (no collision with any known D5 probe input or chain pill prompt)
and a clean CR Round 2 across all reviewing agents.
- Post-merge dashboard re-probe will be the final runtime verification.
## Out of scope (separate follow-ups)
- LG-TS / LG-FastAPI catchall fixtures don't have the stale string — use
pill-aligned userMessages; need different fix
- `tool-rendering` (non-catchall) and `tool-rendering-reasoning-chain`
featureTypes have separate failure modes
- Other red cells in the dashboard (frontend-tools-cosmic timeout,
agent-config, auth, etc) are unrelated
The python rewriter in apply_isolation previously interpolated $slug
directly into the inline python source via bash. A slug containing a
single quote would break the python literal. Internal-tool risk only
(slug is developer-typed), but cheap to harden.
Pass slug via SHOWCASE_ISO_SLUG env var and read os.environ.get(...)
inside the python heredoc. Defense-in-depth; no behavior change for
valid slugs.
The previous 'weather in Tokyo' rename (388c69e68) was a substring of the
main tool-rendering chain pill prompt 'Chain a few tools in this single
turn: get the weather in Tokyo, search flights from SFO to Tokyo, and roll
a d20.' Because aimock loads fixtures alphabetically per integration dir
and uses substring match with first-match-wins, the catchall fixture
(loaded before tool-rendering.json) was intercepting chain pill matches
across 16 integrations.
Rename catchall fixture userMessage to 'forecast for Tokyo' — a phrase
not contained in any other pill prompt. Update the corresponding D5
catchall probe inputs in d5-tool-rendering-{default,custom}-catchall.ts
and the test assertions that pin those inputs.
Call-Site Enumeration: 'weather in Tokyo' remains intentionally in
page.tsx suggestions.ts pills (user-visible UX) and inside the chain
pill prompt itself — neither is in the substring-match path now.
The persistent stack's docker-compose.local.yml hardcodes LOCAL_SERVICES_JSON
to the langgraph-python sample for fast N=1 local demos. When --isolate
spawns an iso1 stack with a different slug (e.g. ms-agent-python), the
iso1 harness container inherited that hardcoded value, causing
discovery.railway-services.local-injection to enumerate the wrong service
(showcase-langgraph-python instead of showcase-<requested-slug>). The iso1
probe then targeted the wrong container, broke red-green verification, and
left D5 cells unwritten.
Inject a per-slug LOCAL_SERVICES_JSON override into the iso1 compose
generator so iso1 always probes the slug passed via --isolate.
- Scope clickPill locator to data-message-role='user' bubble so the pill
button itself can no longer satisfy the dispatch guard
- Dedup clickPill retry: skip click if the user bubble already exists
- Hero pill: assert declarative-card count=0 (OSS-136 no-Card rule),
metric count >=4 (was >=3 — KPI strip is 4 tiles per composition rule)
- At-risk pill: assert no chart and no table testids (composition rule)
- Top-account pill: assert no data-table and no status-badge testids
- Rename hero test title to 'KPI strip + pie + bar (no surrounding card)'
so the title no longer falsifies the body
- QA docs: replace 'card + metrics + pie + bar' Expected Results with
'4 KPI metrics + 1 PieChart + 1 BarChart, no surrounding Card per OSS-136'
- Probe responseTimeoutMs derived from FIRST_SIGNAL_TIMEOUT_MS so it
matches the e2e 90s budget
- Migrate readDeclarativeTestIds from booleans to counts so leftover vs
newly-mounted is distinguishable
- everyNewlyMounted gate uses current[k] > baseline[k] (was boolean
!baseline[k] against a count, which falsely blocked at non-zero baseline)
- minCounts enforce newly-mounted delta, not raw current count
(fixes the cross-pill bleed: at-risk metric:3 floor used to pass on
hero's 3 leftover metrics with no fresh mount)
- Per-pill minCounts add the chart sibling asserts D5 was missing:
hero=4 metric+1 pie+1 bar, team=1 table+1 bar, top-account=1 info-row+1 pie
- 31/31 harness tests green
- DataTable rowKey uses first-column value + index instead of bare index,
with JSON.stringify(row) fallback (stops re-mount on dynamic A2UI re-emits)
- Card emits data-card-id={props.title} so multi-card pills no longer
collide on a single declarative-card testid
- PieChart/BarChart value coercion replaced 'Number(x) || 0' with
finite-number check + console.warn on drift (no longer masks legitimate 0)
Replace the misleading 'single source of truth' claim with an explicit
DUPLICATION NOTICE describing the per-integration parity convention and
a TODO(OSS-136) for the future shared-module extraction. Both copies
remain byte-identical.
The D5 e2e-deep probe for tool-rendering-{default,custom}-catchall sends
"weather in Tokyo" as the test input (harness/src/probes/scripts/
d5-tool-rendering-{default,custom}-catchall.ts). The fixture userMessage
matcher uses substring match. Main's catchall fixtures had a stale
"check Tokyo weather forecast" string that could not substring-match
the probe input, causing the fixture to miss and the probe to fall
through to the live LLM — surfacing as CV x red D4 on the dashboard.
Rename to the canonical "weather in Tokyo" string across affected
integrations' tool-rendering-{default,custom}-catchall.json files.
No agent or page.tsx changes; fixture content otherwise unchanged.
Update package READMEs + ARCHITECTURE for the assistant pane, native streaming,
and the new onThreadStarted / setSuggestedPrompts / setTitle surface. Reverse
the slack.mdx callout that told users to delete the assistant scopes (now
required), and enable the pane in the examples/slack manifest (assistant_view +
assistant:write + assistant_thread_* events) with a dev-ex onThreadStarted
greeting and the assistant config.
The custom-catchall probe sends 'weather in Tokyo' then 'AAPL'. After the
weather tool runs in turn 1, hasToolResult is true across the rest of the
thread — which fires tool-rendering.json's AAPL 'hasToolResult:true' narration
prematurely on turn 2 iteration 1, returning prose without ever emitting the
get_stock_price tool. The custom-catchall assertion (both tools rendered
through the wildcard testid) then fails with missing get_stock_price.
Replace the (hasToolResult:true narration + hasToolResult:false/turnIndex:0
emitter) layered fallbacks with a sequenceIndex:0 emitter ordered before a
bare userMessage+context narration. The per-test fixture-match counter resets
each run, so the emitter fires exactly once on iteration 1 regardless of
prior pills' tool history, then falls through to the narration on
iteration 2. Applied symmetrically to the two AAPL blocks in
tool-rendering.json (the 'What\'s the current price of AAPL?' block at the
top and the legacy 'current price of AAPL' alias block lower down). The
toolCallId-keyed narration above each block is retained for the non-BIA
fast path.
The earlier partial fix to tool-rendering-custom-catchall.json is kept (it
adds toolCallId-scoped narration legs ordered before the existing
hasToolResult:true narrations); those fixtures never match real probe
traffic (the probe sends 'weather in Tokyo' / 'current price of AAPL', not
the unique 'check Tokyo weather forecast' substring in this file) but the
reordering is consistent with the cross-file pattern and harmless.
Verified locally: built-in-agent:tool-rendering and
built-in-agent:tool-rendering-custom-catchall both green via
`./bin/showcase test ... --d6 --direct`; built-in-agent:tool-rendering-default-catchall
also green; aimock-fixtures.test.ts (collision/shadow ceilings) unchanged.