resolveD5 skips missing sub-rows and folds only present rows, so it does not
match isD5Green's every-key-present requirement; clarify staleness IS mirrored.
The deferred-relaunch background timer produced a new race in three review
rounds. Remove it entirely and rely on lazy, on-demand recovery: a failed
recycle parks the slot relaunchPending (never evicts), and the next acquire
re-attempts the launch via relaunchPendingSlots, serving any queued waiter
through handOff. The no-eviction + lazy-recovery design already prevents the
original drain-to-0 bug without a timer. Also clamp stats().inUse to >= 0.
In a single thread the 2nd interrupt's card never mounted. Three coordinated
issues in `useInterrupt` (v2) combined into a publish-cleanup race:
1. The `element` useMemo depended on `config.render` and `config.enabled`,
which consumers pass as inline lambdas (new identity every parent render).
Element identity churned on every render.
2. The publish effect did `setInterruptElement(element)` with a cleanup that
pushed `null`. On dep churn, the cleanup ran AFTER the previous publish —
chat subscribers reading via snapshot-style stores latched `null` between
renders, leaving the card unmounted.
3. `resolve` synchronously called `setPendingEvent(null)`, unmounting the
card before the resume run's first tokens streamed. Consumers worked
around this with a 500ms setTimeout wrapper around resolve().
Fix:
- Stabilize `render`, `enabled`, `handler` behind refs so the element memo
and handler effect depend only on `pendingEvent`/`handlerResult`/`resolve`.
Mirrors the v1 `useLangGraphInterrupt` wrapper's stabilization pattern.
- Split the publish effect into a publish-only effect (no nullify on churn)
plus a separate unmount-only cleanup with empty deps.
- Drop the synchronous `setPendingEvent(null)` from `resolve` —
`onRunStartedEvent` is the legitimate clear path when the resume run
begins. Removes the need for consumer setTimeout workarounds.
The element memo still returns null when pendingEvent is null, so the
legitimate clear paths (onRunStartedEvent / onRunFailed) continue to work.
Adds a red-green test that emits two interrupts in one thread with an
inline-render consumer, forces parent re-render after the 2nd interrupt,
and asserts no stale null follows the last non-null publish.
A stale-green D5 sub-row was masked when a fresh-green sibling won the
all-green tie, since green is the lowest rank and staleness was checked
only on the post-fold winner. Downgrade each green-but-stale sub-row to
degraded before folding so any stale-green forces amber, matching
depth-utils isD5Green per-key semantics regardless of map order.
Re-arm the deferred relaunch when a waiter parks after a recycle has
already failed with no waiter queued, so it is not stranded to timeout.
Make scheduleDeferredRelaunch idempotent via deferredRelaunchActive so
the acquire-kick and recycle-kick cannot spawn two timer loops. Guard
recycleSlot entry and the relaunchPending filter with a single
isSlotBusy predicate covering both recyclingSlots and relaunchingSlots,
so a late disconnect for a pending slot's old browser cannot double
launch and leak a process. Strengthen the double-publish test to fail
loud and pin exact-once, and add red-green tests for both fixes.
## Summary
Polish/hardening sweep over the showcase deploy + promote tooling — the
deferred follow-up items from the deploy-pipeline closeout. Five
focused, low-coupling changes:
- **asHost validator + Host brand** — reject ASCII control chars,
colon/port suffix, and any non-DNS-charset character; switch the `Host`
brand from a string-literal `__brand` to a non-exported `unique symbol`
so an out-of-module `as Host` cast is a type error. `asHost` remains the
sole runtime constructor.
- **Railway promote snapshot invariants** — a regression spec that fails
if a fleet-scoped check reads the narrowed single-service snapshot (the
historic spurious-WARN/REFUSE bug), an executable lint banning direct
`@*_snapshot` reads outside the sanctioned accessors, and PopenSpy
hardening (`UnexpectedPopen` + honest exit-code stamping).
- **aimock fixture loader fail-loud** — throw (with the offending path +
caught error as `cause`) on malformed JSON or a missing `fixtures` array
instead of silently skipping; covering tests for both throw branches.
- **deploy-to-railway cleanup** — memoize the Railway token (failures
never cached), and resolve the GHCR username with `||` so an empty
`GHCR_USERNAME` falls through to `GITHUB_ACTOR`.
- **rename** `getRuntimeConfigEdge` → `getRuntimeConfigForMiddleware`
across shell, shell-docs, shell-dashboard (definitions, middleware call
sites, tests). No behavior change.
## Test plan
- [x] `verify-deploy` vitest (99) incl. new asHost rejection cases
- [x] aimock fixture-coverage vitest (3) incl. red-green throw-branch
tests
- [x] ruby promote spec suite (111 runs / 392 assertions / 0 failures)
- [x] shell / shell-docs / shell-dashboard runtime-config vitest (46)
- [x] tsc, oxfmt, oxlint clean on changed files
Consolidate isE2eGreenAndFresh into isGreenAndFresh; both had identical
staleness semantics. Single helper now serves D3/D5/D6, removing a drift
hazard. No behavior change.
Reconcile the stale stats()/reinit()/acquire() comments: all-pending recovery
goes through relaunchPendingSlots, not the size-0 reinit backstop (which is the
truly-empty never-launched case). Rename the misleading reinit test to assert
the actual relaunchPendingSlots mechanism. Chain track()'s cleanup so a future
recovery throw cannot float as an unhandled rejection, and log a -1 sentinel
slotIndex for the never-added reinit close. Fix the self-contradicting
drainMicrotasks JSDoc.
## Why
Two sibling testid families that e2e tests need across the entire
frontend matrix, consolidated into one PR (this PR absorbs #5108). Both
are purely additive markers — no behavior, rendering, or styling
changes.
### A. Default tool-call renderer (cross-framework parity)
E2E tests for the chat surface (e.g. showcase's
`tool-rendering-default-catchall` canonical spec) need a stable selector
to count and inspect tool-call cards rendered by the framework's
built-in `DefaultToolCallRenderer` when an integration registers zero
custom render hooks. `react-core`'s renderer already emits a
`data-testid="copilot-tool-render"` wrapper with `data-tool-name`,
`data-status`, `data-args` and `data-result`; `vue`'s equivalent
renderer was missing them, so the same e2e test counted 0 cards there.
A freshly-built google-adk integration on the showcase reproduced this:
cards appeared but without the testids → assertion counted 0;
`langgraph-python` passed the same test because its build resolved a
`react-core` version that does emit them. This half fixes the framework
gap so every framework's built-in renderer emits the same stable
selectors.
### B. Error banner + loading indicator (cross-framework, all five
packages)
E2E tests for the chat surface today have no stable selector to
distinguish "errored out" vs "still loading" states, so probes time out
after 30–60s instead of failing fast and pointing to the right cause.
This half adds purely additive, framework-spanning `data-testid` markers
so a single selector works across every frontend.
## What
### A. Default tool-call renderer
Mirrors `react-core`'s contract onto the `vue`
`DefaultToolCallRenderer`:
- `packages/vue/src/v2/hooks/use-default-render-tool.ts`: wrap the card
in a `<div>` carrying `data-testid="copilot-tool-render"`,
`data-tool-name`, `data-status`, `data-args` and `data-result` (via the
same `safeStringifyForAttr` helper shape as react-core); tag the inner
name/status spans with `copilot-tool-render-name` and
`copilot-tool-render-status`.
Adds `toolCallId` parity to `react-core`'s default-renderer props so
both frameworks expose the same prop shape:
- `packages/react-core/src/v2/hooks/use-default-render-tool.tsx`:
`DefaultRenderProps` now carries `toolCallId` (vue already exposed it).
- `packages/react-core/src/v2/hooks/use-render-tool-call.tsx`: adapter
forwards `toolCallId` into the default renderer.
Locks the contract into unit tests in **both** frameworks so the markers
cannot silently disappear in a future refactor:
- `packages/vue/src/v2/hooks/__tests__/use-default-render-tool.test.ts`:
new `"default renderer emits stable copilot-tool-render testid and
metadata attrs"` test (red-green proven locally by stashing the source
change — without it, `getByTestId("copilot-tool-render")` throws).
-
`packages/react-core/src/v2/hooks/__tests__/use-default-render-tool.test.tsx`:
mirroring test asserting the same wrapper / `data-*` attrs and inner
testids (red-green proven by sentinel-swapping the testid in the
renderer source — the new test fails, others pass), plus a regression
test that asserts `toolCallId` is forwarded through the adapter.
### B. Error banner + loading indicator
Adds two stable testids on the relevant UI surfaces in every frontend
framework package — react-core, react-ui, react-native, angular, vue:
- `copilot-error-banner` on the global error banner UI:
- `packages/react-core/src/components/toast/toast-provider.tsx`
(`BannerErrorDisplay` root)
- `packages/react-core/src/components/usage-banner.tsx` (`UsageBanner`)
- `packages/react-ui/src/components/chat/messages/ErrorMessage.tsx`
(legacy in-chat ErrorMessage)
- `copilot-loading-cursor` on the flashing-dot / typing indicator:
- `packages/react-ui/src/components/chat/Messages.tsx` and
`messages/AssistantMessage.tsx` (legacy `LoadingIcon` spans)
-
`packages/angular/src/lib/components/chat/copilot-chat-message-view-cursor.ts`
- `packages/react-native/src/components/messages/TypingIndicator.tsx`
(via RN's `testID` convention)
- `packages/vue/src/v2/components/chat/CopilotChatMessageView.vue`
- The v2 react-core `CopilotChatMessageView.Cursor` already exposes
`copilot-loading-cursor`; this PR broadens that same selector to the
other frameworks.
Vue's prior `copilot-chat-cursor` testid is renamed to
`copilot-loading-cursor` for cross-framework consistency; the two e2e
tests in `packages/vue` that referenced the old name are updated in the
same commit.
Adds small static source-asserting tests in each touched package that
verify the markers stay in place (red/green proven locally by
transiently removing one marker).
## Per-framework audit (default renderer)
| Package | Has built-in default renderer? | `copilot-tool-render`
testid before this PR | Status after PR |
|---|---|---|---|
| `react-core` | YES (`DefaultToolCallRenderer`) | YES (already
shipping) | `toolCallId` added to `DefaultRenderProps` for vue parity;
regression test added |
| `vue` | YES (`DefaultToolCallRenderer` in
`use-default-render-tool.ts`) | **NO** | **added** + regression test |
| `angular` | NO (renders only when user registers `*`) | N/A | no
change needed |
| `react-native` | NO (deliberately excludes — DOM-only, see
`packages/react-native/src/index.ts` comments) | N/A | no change needed
|
| `react-ui` | inherits `react-core`'s | N/A | no change needed |
## Consolidation note
This PR supersedes and absorbs #5108. The two testid families touch
disjoint files; cherry-picking #5108's single commit onto this branch
was clean. Both halves went through the CR loop (converged).
## Scope
Purely additive — no behavior, rendering, or styling changes. The new
attributes are inert at runtime; only e2e and unit tests read them.
`toolCallId` is now exposed to both frameworks' default renderers as a
prop; emission of a `data-tool-call-id` attribute is explicitly out of
scope (see Out of scope below).
## Test plan
- [x] `pnpm --filter @copilotkit/vue run test` — 995/995 green
- [x] `pnpm --filter @copilotkit/react-core run test` — 1183/1183 vitest
green (the pre-existing `test:scripts` failure on `origin/main` is
unrelated)
- [x] `pnpm --filter @copilotkit/react-ui run test` — 36/36 green
- [x] `pnpm --filter @copilotkit/react-native run test` — 246/246 green
- [x] `pnpm --filter @copilotkitnext/angular run test` — 48/48 green
- [x] `pnpm --filter @copilotkit/vue run build` — green
- [x] `pnpm --filter @copilotkit/react-core run build` — green
- [x] `pnpm exec oxlint <touched files>` — 0 warnings, 0 errors
- [x] `pnpm exec oxfmt --check <touched files>` — clean
- [x] Typecheck (`tsc --noEmit`) on all five touched packages — no new
errors vs `origin/main` (pre-existing baselines preserved; vue actually
drops from 111 → 0 as a side effect of the toolCallId typing tightening)
- [x] Red→green proven for every new test (stash source / sentinel-swap
testid / transient marker removal)
## Out of scope (follow-up PR)
The CR loop surfaced several pre-existing issues in the
default-tool-renderer code path. They are intentionally NOT addressed
here so this PR stays a focused, additive testid + parity change. A
follow-up PR will harden them:
- `react-core`: swap the renderer's clickable `<div onClick>` for a
`<button>` for keyboard / a11y compliance.
- Exhaustive `ToolCallStatus` enum mapping with explicit logging on
unknown values (instead of silent fallback) in both frameworks.
- Emit a `data-tool-call-id` attribute on the wrapper (the prop is now
available on both sides; emitting it is a separate, deliberate
decision).
- `useDefaultRenderTool`: introduce an opt-in prop-shape adapter so
callers can pick the prop contract explicitly rather than implicitly.
- Extract `safeStringifyForAttr` into a shared util and replace
remaining unguarded `JSON.stringify` calls in both renderers with it.
- Rewrite the 5 source-grep testid unit tests (cherry-picked from #5108)
to render-based assertions instead of `readFileSync` + regex matching.
## Follow-up
- Unblocks the showcase `tool-rendering-default-catchall` D6 canonical
spec across the fleet, plus reliable error/loading state assertions in
showcase e2e.
- A `react-core` (+ peers) release will follow once this lands (gated,
separate change).
Clarify the wrapper's role (it forces noStore:false because unstable_noStore is
unavailable in middleware/Edge). Pure rename across shell, shell-docs, and
shell-dashboard: definitions, middleware call sites, and tests. No behavior
change.
Memoize getToken so the Railway config is not re-read and the deprecation
warning is not re-emitted on every GraphQL request; failures are never cached.
Resolve the GHCR username with || so an empty GHCR_USERNAME falls through to
GITHUB_ACTOR, and use a truthy cache-hit guard so an empty token is never
memoized.
loadAllFixtures silently skipped files lacking a fixtures array and let
JSON.parse throw without the path; throw a clear error naming the file in both
cases and attach the caught error as cause. Export the helper and add red-green
tests covering both throw branches.
Add a regression spec that fails if a fleet-scoped check reads the narrowed
single-service snapshot (the historic spurious-WARN/REFUSE bug), an executable
lint banning direct @*_snapshot reads outside the sanctioned accessors, and
harden PopenSpy with UnexpectedPopen + honest exit-code stamping (raise only on
spawn failure, not on a configured non-zero exit).
Reject ASCII control chars, a colon (port suffix), and any char outside the
DNS-label charset; switch the Host brand from a string-literal __brand to a
non-exported unique symbol so a stray `as Host` cast from outside the module
is a type error. asHost stays the sole runtime constructor. Adds positive +
negative test coverage incl. interior-tab rejection.
Extend the e2e staleness downgrade to resolveD5/resolveD6 in cell-model and to
depth-utils so a frozen-green pipeline no longer credits D5/D6 as healthy. Also
correct misleading comments/test-doc nits flagged in review.
relaunchPendingSlots only pushed recovered slots to `available` and never resolved
queued waiters, so the deferred relaunch scheduled by a failed recycle (scheduled
precisely because waiters were queued) left the waiter stranded for its full 30s
acquire timeout while a live browser sat idle. It also had no re-entry guard, so an
acquire()-driven call racing the deferred timer could both launch a browser for the
same pending slot — leaking one process and/or publishing the slot to `available`
twice (handing the same browser to two probes).
- Extract a single `handOff` helper (waiter-first, then `available` with an includes
guard) and use it in reinit, the recycle success path, and relaunchPendingSlots so
all three share one correct invariant.
- Add a `relaunchingSlots` re-entry guard: a slot is marked before its launch await
and skipped by concurrent invocations, cleared in finally.
- Replace the one-shot deferred relaunch with a self-rescheduling timer (bounded
backoff) that retries while pending slots and waiters coexist, so a waiter is no
longer stranded when a single deferred attempt also fails. Stops on shutdown.
- Track reinit/relaunch launch promises so shutdown() drains them, closing the leak
where a launch resolved after shutdown awaited the in-flight set.
- Route failure and close logging through the injected logger (with slot index)
instead of console.error / swallowed catches, so failures reach the harness pipeline.
- stats().size now counts only live (non-pending) capacity, making the empty-pool
backstop observable (size 0 when every slot is pending).
Tests: fix the vacuous reinit test (failAtCalls now exhausts all retries for every
slot and asserts size 0 before recovery), and add deferred-waiter recovery and
no-double-publish tests. Full harness suite green (1648 tests), typecheck green.
Matches the explicit-import convention used by sibling tests in this
package (e.g. copilot-chat-agentid.test.tsx, streaming-fetch.test.ts).
Removes 3 tsc "Cannot find name 'describe'/'it'/'expect'" errors
without changing runtime behavior — vitest globals already provided
at runtime via vitest.config.mjs (globals: true).
Adds purely additive data-testid markers to the error and loading UI
surfaces across the frontend framework packages (react-core, react-ui,
react-native, angular, vue) so e2e tests can deterministically detect
errored-out vs still-loading states. Without these, e2e probes hit
~30-60s timeouts instead of failing fast.
Testids (aligned with existing repo convention; copilot-<kebab>):
- copilot-error-banner on react-core BannerErrorDisplay (toast
provider) and UsageBanner, plus react-ui legacy in-chat ErrorMessage.
- copilot-loading-cursor on react-ui legacy LoadingIcon sites
(Messages.tsx, AssistantMessage.tsx), angular
CopilotChatMessageViewCursor, react-native TypingIndicator (via
RN testID convention), and vue CopilotChatMessageView. The v2
react-core Cursor already exposed this testid; this change broadens
it to every frontend framework so a single selector works across all.
Vue's prior copilot-chat-cursor testid is renamed to
copilot-loading-cursor for cross-framework consistency; the two e2e
tests in packages/vue that referenced the old name are updated.
No behavior, rendering, or styling changes. Adds small static
source-asserting tests in each touched package that verify the markers
stay in place.
Same pattern as ms-agent-dotnet: full d4 chat.json rewrite + d6 mirrors. Dropped stale
turnIndex gates and broad shadow matchers. Default-catchall now green via page-level
shadcn-catchall-renderer. 11 residual: declarative-gen-ui charts, multimodal conveyance,
tool-rendering, reasoning-chain.
Full d4 chat.json rewrite + d6 mirrors. Dropped stale turnIndex gates and broad shadow
matchers that no longer reflect LGP-canonical conveyance. Default-catchall now green via
page-level shadcn-catchall-renderer (not react-core-gated). 5 residual: chat-slots, hitl,
interrupt-headless, readonly-state.
When the e2e driver stops writing e2e:<slug>/<feature> status rows, the last
row freezes. A green row then reads as a healthy D3 forever, so the depth
ladder shows a false-green D3 that masks a dead probe pipeline (a wedged
browser pool, an outage) instead of surfacing it.
buildCellModel (cell-model.ts) and deriveDepth (depth-utils.ts) now treat a
green e2e row whose observed_at is older than E2E_STALE_AFTER_MS (6h, matching
the original stale-window model) as degraded/amber, so it no longer credits
D3. Only green is downgraded — a stale red/degraded row already signals a
problem and is left as-is. An unparseable timestamp is treated as not-stale so
staleness is never inferred from bad data.
Existing test fixtures that hardcoded old observed_at timestamps now default
to a recent value so green rows are not treated as stale; the new staleness
tests pass explicit timestamps.
When a chromium process crashes, recycleSlot relaunches it; if that launch
threw, the slot was permanently removed via slots.splice. A single transient
launch failure (OOM spike, fd exhaustion) thus monotonically shrank the pool
to empty with no self-heal, after which every e2e probe failed with
launcher-error / BrowserPool acquire timeout.
The recycle now retries the relaunch with bounded backoff; on exhaustion it
keeps the slot (capacity preserved) and parks it as relaunchPending for a lazy
retry on the next acquire. acquire() also re-initializes the pool as a backstop
when every slot has been lost (slots.length === 0), so the pool always
self-heals rather than wedging.
vue: replace the string-literal deps array (which was laundered through
`as unknown as any[]` because string is not a valid WatchSource) with
a getter-style deps array (`() => "compact"`), which is a valid
WatchSource<unknown>. The reference-identity assertion still holds.
react-core: rewrite the toolCallId comment to accurately describe what
this test verifies. The test calls config.render directly with
useRenderTool mocked, so it does not exercise the spread-adapter path
end-to-end — it only locks that useDefaultRenderTool passes the user's
render through untouched.
The 3 "default renderer" tests in use-default-render-tool.test.tsx were
narrowing config.render via an as-cast that omitted the now-required
toolCallId field on DefaultRenderProps, laundering the type. Switch the
casts to the real DefaultRenderProps shape and pass a realistic
toolCallId on every <DefaultRenderer/> invocation. No behavior change.
The runtime path already forwarded toolCallId to wildcard render functions
(useRenderTool spreads ReactToolCallRenderer props, which include toolCallId),
but the static DefaultRenderProps type omitted it. Vue's sibling type already
declared the field. This divergence forced an `as unknown as { toolCallId }`
cast in the react-core test.
Declare toolCallId on DefaultRenderProps (mirroring vue), thread it through
the defaultToolCallRenderAdapter so the now-required field is genuinely
populated, export the type, and drop the cast plus stale comments in the
test that claimed the field was runtime-only.
Mirror the Vue sibling test 'forwards toolCallId to custom wildcard render
function' so the react-core suite locks the same regression: the wildcard
hook must forward toolCallId to a custom render function. Closes a
symmetry gap in the cross-framework testid PR.
E2E tests for the chat surface (e.g. showcase's
tool-rendering-default-catchall canonical spec) need a stable selector
to count and inspect tool-call cards rendered by the framework's
built-in DefaultToolCallRenderer when an integration registers zero
custom render hooks. react-core's renderer already emits a
data-testid="copilot-tool-render" wrapper with data-tool-name,
data-status, data-args and data-result; vue's equivalent renderer was
missing them, so the same e2e test counted 0 cards there.
Mirrors react-core's contract onto the vue DefaultToolCallRenderer:
- packages/vue/src/v2/hooks/use-default-render-tool.ts: wrap the card
in a div carrying data-testid="copilot-tool-render", data-tool-name,
data-status, data-args and data-result (via the same
safeStringifyForAttr helper shape as react-core); tag the inner
name/status spans with copilot-tool-render-name and
copilot-tool-render-status.
Locks the contract into unit tests in both frameworks so the markers
cannot silently disappear in a future refactor:
- packages/vue/src/v2/hooks/__tests__/use-default-render-tool.test.ts:
new "default renderer emits stable copilot-tool-render testid and
metadata attrs" test (red-green proven locally by stashing the
source change).
- packages/react-core/src/v2/hooks/__tests__/use-default-render-tool.test.tsx:
mirroring test asserting the same wrapper/data-* attrs and inner
testids (red-green proven by sentinel-swapping the testid).
Scope: purely additive — no behavior, rendering, or styling changes.
The new attributes are inert at runtime; only e2e and unit tests read
them. angular has no built-in default renderer (only renders when the
user registers a wildcard) and react-native deliberately excludes the
default renderer (web DOM-only), so no changes are needed there.
This unblocks the showcase tool-rendering-default-catchall D6 spec
across frontends; a react-core release will follow once merged.
## Summary
Fixes the two failures in the `Showcase: Verify Deploy` workflow on
`main`. Both are pre-existing (since the staging-verify wiring landed)
and independent of recent PRs:
- **webhooks health-path mismatch** — the verify probe checks
`/api/health` (the standard used by every API-shaped service driver:
agent, eval, pocketbase, webhooks), but the `webhooks` (eval-webhook)
service only served `/health`, so the probe got HTTP 404. Fix: the
service now also serves `/api/health` (same `{ ok: true }` body),
matching the standard. `/health` is kept for back-compat.
- **notify-harness payload schema mismatch** — `showcase_deploy.yml`
posted `{ state, services }` to the harness deploy-webhook, but the
ingest endpoint's schema is `.strict()` and requires `{ succeeded[],
failed[], cancelled }`, rejecting the unknown `state` key → HTTP 400.
Every push-to-main notify-harness call had been 400-ing, so the harness
dashboard received no deploy-result events. Fix: the workflow now emits
a `failed_services` list from the redeploy summary and builds the
payload as `{ services: succeeded ∪ failed, succeeded, failed, cancelled
}` (optional run/build URLs omitted when empty so the schema's `.url()`
validation passes).
## Test plan
- [ ] Local: `tsc --noEmit` (eval-webhook) clean; oxfmt + oxlint clean;
`actionlint showcase_deploy.yml` clean
- [ ] Local: jq dry-run of the new payload across
success/failure/cancelled/empty cases produces a schema-valid object
with no `state` key
- [ ] Post-merge (push-to-main): `Showcase: Build & Push` redeploys
webhooks with `/api/health`, then `Showcase: Verify Deploy` goes green
(webhooks probe 200) and notify-harness POST returns 200
deploy.yml's notify-harness step was sending {state, services, ...},
but the harness ingest schema at showcase/harness/src/http/webhooks/
deploy.ts is .strict() and requires succeeded/failed/cancelled arrays
+ bool. Every push-to-main notify-harness POST has been 400ing as
invalid-payload (unknown key 'state') since the schema landed.
Build the payload from the redeploy gate's success-and-error buckets:
- Emit failed_services alongside ok_services from the redeploy-gate
step (services whose status==error), and expose both as resolve-matrix
job outputs.
- Rewrite the notify-harness payload to drop state and emit:
succeeded=ok_services, failed=failed_services, cancelled=(verify
conclusion==cancelled), services=union of succeeded and failed.
- Omit buildRunId/buildRunUrl entirely when empty (the schema runs
.url() on buildRunUrl, so an empty string would be rejected). Same
treatment for runUrl as defense-in-depth.
- Group the no-summary echoes through a single >>GITHUB_OUTPUT block
so shellcheck SC2129 stays clean once a third echo is added.
The eval-webhook server only exposed /health, but the verify-deploy
probe (showcase/scripts/verify-deploy.drivers.webhooks.ts) and every
other API-shaped showcase backend (agent, eval, pocketbase) standardize
on /api/health. Add the /api/health route so the service matches the
SSOT convention the probe expects; keep /health for back-compat with
any pre-existing callers.
Each non-LGP integration carried its own drifted/stale copy of the e2e specs, causing
inconsistent behavior and noisy diffs across the fleet. Copied langgraph-python's canonical
specs verbatim across ~15 integrations (576 spec files total, SHA-1-verified identical to
LGP) so every integration runs the same assertions.
Also removed 2 orphan specs whose underlying demo pages do not exist:
- showcase/integrations/agno/tests/e2e/hitl-in-chat-booking.spec.ts
- showcase/integrations/built-in-agent/tests/e2e/shared-state-write.spec.ts
Integration-specific variant specs were intentionally left as-is: reasoning-default-render,
byoc-*, agentic-chat-reasoning, and shared-state-write where the demo exists. google-adk and
langgraph-typescript were already in parity from earlier commits and show no new changes.
Integration specs had drifted/staled vs LGP gold standard; copied LGP's canonical
specs verbatim and removed the orphan shared-state-write spec whose demo exists
in neither LGP nor google-adk.
The whole-thread `hasToolResult: false` gate on the Chain-tools first-turn fixture caused
the chain pill to fall through to the broad "weather in Tokyo" matcher mid-thread.
Removing the gate (mirroring LGP) yields LGT 185/0/2.
Stages the canonical suggestion pill set (mirrored from langgraph-python) as new
suggestions.ts files across 13 integrations: ag2, agno, mastra, pydantic-ai,
claude-sdk-python, claude-sdk-typescript, llamaindex, langroid, strands, spring-ai,
built-in-agent, crewai-crews, langgraph-fastapi.
Also includes targeted edits to existing suggestions.ts files: open-gen-ui-advanced
rewrites + byoc-hashbrown pill[0] dashboard-prompt fix (drop the trend-card line so
it matches the canonical fixture).
NOTE: these new files are currently UNWIRED. Each integration's page.tsx still
defines its pill list inline via useConfigureSuggestions. Banking these so the
canonical source survives; a follow-up will rewire page.tsx to import from
suggestions.ts and delete the inline copies.
Mirrors the LGP first-match-wins fix on the langgraph-typescript side. Three broad
matchers in d4/langgraph-typescript/chat.json were shadowing d6 fixtures; narrowed
them so the d6 pills win.
Across the d6/langgraph-typescript suite: dropped fragile turnIndex:0 gates in favour
of hasToolResult:false for first-leg tool emissions, fixed em-dash escaping that broke
literal string matches in multi-pill threads, and added jsFunctions payloads to the
three sandboxed-ui fixtures (_from-feature-parity, headless-complete,
gen-ui-open-advanced) so the sandbox renderer has executable handlers.
LGT D6 now passes 184/1/2 locally. Residual 1 fail is the custom-catchall multi-pill
follow-up case; tracked separately.
Aimock fixture load order is shared -> d4 -> d6, and matching is first-match-wins. A
broad d4 langgraph-python chat fixture for search_flights was shadowing the d6
beautiful-chat fixture and preventing the intended response from firing. Neutered the
d4 matcher to a non-matching sentinel so the d6 fixture wins.
Also refreshed d6/tool-rendering-custom-catchall.json and d6/tool-rendering.json:
replaced fragile turnIndex:0 gates with hasToolResult:false so the AAPL first-leg
fixture fires correctly when D5 probes run a prior 'weather in Tokyo' pill in the
same thread (multi-pill turnIndex>=2).
LGP D6 now passes 185/0/2 locally (2 skips are the by-design mcp-apps iframe gap).
## Summary
Hardening pass over the showcase deploy pipeline and its operator
tooling. Five focused areas:
- **`bin/railway` rollback shell-injection fix** —
`RollbackCommitCommand` shelled out via backtick subshells interpolating
the `--sha`/`--env` flags. Replaced with `IO.popen` argument arrays (no
shell), with `--sha` validated against an anchored hex regex and `--env`
validated against the known env set before any git invocation. Both `git
ls-tree` and `git show` now gate on `$?.exitstatus` before parsing their
output (a failed `git show` previously merged stderr into stdout and fed
it to `YAML.safe_load`). New spec covers injection rejection, git-show
failure, and empty-snapshot paths.
- **Promote snapshot encapsulation** — replaced the bare `@full_*`/`@*`
snapshot ivars in `PromoteCommand` with named `fleet_*`/`target_*`
accessors so fleet-scoped vs per-service snapshot selection is
name-enforced rather than comment-enforced.
- **Branded `Host` type** — `ProbeTarget.host` is now a branded `Host`
produced only by `asHost` (rejects scheme, path, empty, whitespace,
userinfo, query, fragment), branded at the verify-pipeline ingress;
`ProbeTarget` fields are `readonly`.
- **`deploy-to-railway.ts` SSOT migration** — uses the shared
`resolveRailwayToken` and sources project/env IDs from the railway-envs
SSOT; reads the GHCR username from env (fail-loud) instead of a
hardcoded handle.
- **Workflow safety + observability** — collapsed the promote workflow
to an input-agnostic concurrency group so promotes can't race the same
Railway service; added `#oss-alerts` failure notifications to the build
and validate workflows (which previously had silent red paths).
- **Harness probe test fixes** — repointed the aimock fixture-coverage
probe to the per-integration D4/D6/shared layout and aligned
d5-multimodal assertions to the current auto-send sentinels.
## Test plan
- [ ] CI: Validate Showcase, build-checks, python tests green
- [ ] CI: static/quality (oxfmt + lint + typecheck) green
- [ ] Local: Ruby promote suite (`ruby spec/all_tests.rb`) — 103/342
green
- [ ] Local: showcase-scripts vitest green; harness probe tests green
- [ ] Local: oxfmt --check clean; actionlint clean
## Summary
- `@copilotkit/react` and `@copilotkit/agent` do not exist on npm —
every user following the setup skill fails on Step 1 with a 404
- Replaces them with the correct published packages:
`@copilotkit/react-core` and `@copilotkit/runtime`
- Updates all import paths throughout the skill to use the v2 subpath
exports (`@copilotkit/react-core/v2`, `@copilotkit/runtime/v2`)
- Removes references to the non-existent `@copilotkit/core` and
`@copilotkit/agent` packages from the Package map table
## Changes
- Step 1 install commands now use real package names that resolve on npm
- All code examples import `BuiltInAgent`, `CopilotRuntime`,
`InMemoryAgentRunner`, endpoint factories from `@copilotkit/runtime/v2`
- All frontend code examples import `CopilotKitProvider`, `CopilotChat`,
`CopilotSidebar`, styles from `@copilotkit/react-core/v2`
- Package map table updated to reflect the two real installable packages
## Test plan
- [x] `npm install @copilotkit/react-core @copilotkit/runtime hono` —
installs without 404
- [x] All import paths in examples resolve to real published subpath
exports
Adds two CI signals for keeping the published packages small and broadly compatible:
- Bundle size: size-limit file-mode config across packages plus a
CopilotChat import-size regression signal (gzip) so growth in the
headline consumer entrypoint is visible on every PR. A bundle-size
workflow comments results on the PR (Phase 1: no hard-fail).
- ES compatibility: a compat-check (es-check) script across 9 packages
with a root .browserslistrc, validating built .mjs/.cjs against the
es2022 build target.
The measure script is importable (measureBundle) and unit-tested. Dev
docs live under dev-docs/ (bundle-size.md, browser-compat.md). All
action refs are pinned to full commit SHAs for supply-chain safety.
Companion to the conveyance shim commit. The shim files were staged
without their callers; this pass wires:
- langroid / llamaindex / ms-agent-python / pydantic-ai / strands
agent_server.py: register the HeaderForwardingHTTPMiddleware
- mastra: switch every API route (copilotkit, copilotkit-auth, beautiful-chat,
byoc-hashbrown, byoc-json-render, mcp-apps, multimodal, ogui, voice) and
the mastra agents / tools / subagents modules onto the
_header_forwarding-wrapped openai provider so inbound x-* headers ride
on outbound Vercel AI SDK calls via ALS
Refresh d6 fixtures across the rollout cohort: ag2, built-in-agent,
claude-sdk-{python,typescript}, crewai-crews, google-adk, langgraph-fastapi,
langgraph-typescript, langroid, llamaindex, mastra, ms-agent-{dotnet,python},
pydantic-ai, strands. Companion d4/{langgraph-typescript,mastra,ms-agent-dotnet}
chat.json refreshes. Add the missing ms-agent-python/gen-ui-custom.json
to bring the integration up to the standard pill set.
Also narrow aimock/shared/common.json's generic 'hello' fixture to
'hello world' so it no longer shadows D6 pills whose prompts contain
'hello' as a substring (e.g. langgraph-python headless-simple sends
'Say hello in one short sentence.'). 'hello world' is unused by any
current demo pill, so the fixture remains a manual-typing fallback
without poisoning fixture matching.
This is a mid-rollout snapshot — fixture coverage is uneven across
integrations and rides alongside the conveyance shims landed earlier
in this branch.