Commit Graph

12629 Commits

Author SHA1 Message Date
Jordan Ritter f1f9dc2890 fix(showcase): narrow llamaindex d4 'summarize' fixture to 'Summarize the sales pipeline'
The bare 'summarize' userMessage in d4/llamaindex/chat.json substring-matched
the D6 gen-ui-agent pill 'Research our top competitor and summarize their
strengths and weaknesses.', returning the sales-pipeline text fixture instead
of the gen-ui-agent set_steps tool call. The competitor pill then produced
no/duplicate steps, failing d6:llamaindex. Narrow the match to the verbatim D4
toolbar probe 'Summarize the sales pipeline' (langgraph-python parity), which
no demo pill contains as a substring. D4 llamaindex stays green (the bare entry
was unused by any D4 cell).

(cherry picked from commit c23801c8b32d292cacf6fb2e7e2a68270eebaa84)
2026-06-28 10:22:53 -07:00
Jordan Ritter 088b7119dd fix(showcase/llamaindex): route frontend-tools-async to dedicated make_request_aware_router agent
The shared FixedAGUIChatWorkflow catch-all dropped request-injected query_notes,
so NotesCard never mounted. Give the cell its own agent (mirrors beautiful_chat_agent)
so request-time frontend tools forward. Verified GREEN via control-plane --direct.

(cherry picked from commit 9c1b8ce2c33afc89000355d83c995f03da208f0f)
2026-06-28 10:10:49 -07:00
Jordan Ritter fcdcc888fe fix(showcase): llamaindex a2ui-fixed-schema — emit streamed render_a2ui tool-call so a2ui-middleware mounts the surface
Same root cause as the sibling declarative-gen-ui (A2UI Dynamic Schema) fix:
the A2UI middleware mounts the surface from a STREAMED render-tool CALL whose
name is in its watched set, not from a TOOL_CALL_RESULT. The prior approach had
display_flight return an a2ui_operations container in the tool RESULT, which the
llama-index AG-UI adapter only re-emits via MESSAGES_SNAPSHOT — a shape the
middleware never inspects — so the flight-card surface stayed unmounted
(reason=surface-missing; the a2ui-fixed-card testid never appeared).

- route.ts: set a2ui.injectA2UITool: true so the middleware watches render_a2ui.
- a2ui_fixed.py: display_flight now returns the fixed-schema render_a2ui args
  (surfaceId/catalogId/components/data) as JSON; a workflow override
  (_A2UIRenderToolCallWorkflow) parses each backend tool result and re-emits it
  as a streamed render_a2ui tool-CALL (TOOL_CALL_START name=render_a2ui ->
  chunked TOOL_CALL_ARGS carrying the components+data JSON -> TOOL_CALL_END),
  mirroring how google-adk drives the middleware. These events are already in
  the upstream AG_UI_EVENTS allow-list the SSE router streams against.

Backend still produces the pre-authored flight schema (no stub). Only
display_flight (one backend tool, name unchanged) is involved, so the d6 fixture
needs no re-keying. Integration-code only; no shared/@ag-ui package touched.

(cherry picked from commit 74d61eddf7eccf22bcc96c37f0e35a70dd823a2f)
2026-06-28 08:40:28 -07:00
Jordan Ritter 485f37a8fa fix(showcase): mount llamaindex declarative-gen-ui A2UI surface + stop the OOM crash-loop (#5749)
## What was broken
`d6:llamaindex/gen-ui-declarative` (A2UI Dynamic Schema) was red, and
the whole **llamaindex D6 column** was crash-looping.

## Root cause (two stacked layers, both llamaindex-integration-level)
1. **OOM crash-loop.** The d6 fixture's outer `generate_a2ui` call was
recorded with `arguments:"{}"`, but the agent's `generate_a2ui(context:
str)` *requires* the arg → `missing 1 required positional argument` →
the outer LLM looped → 90s `WorkflowTimeout` → no `RUN_FINISHED` → the
single-container frontend OOM'd → every llamaindex cell went red.
2. **surface-missing.** Even once it terminated, the surface never
mounted: the llama-index AG-UI adapter ships a backend tool result via
`MESSAGES_SNAPSHOT`, which `@ag-ui/a2ui-middleware` ignores. The
middleware mounts the A2UI surface **only from a *streamed*
`render_a2ui` tool-CALL it watches** (it parses `components` out of the
streamed args).

## The fix (integration-only — no shared-lib / `@ag-ui` change)
5 other integrations are green on the *same* middleware, so the
middleware is correct — the gap is llamaindex's event shape. This PR:
- **Rebuilds** `aimock/d6/llamaindex/gen-ui-declarative.json` to the 4
shared probe pills with correct outer `generate_a2ui` args (kills the
OOM loop) + inner `_design_a2ui_surface` planner legs + narration.
- **Overrides `aggregate_tool_calls`** (`a2ui_dynamic.py`) to re-emit
each `generate_a2ui` backend result as a **streamed `render_a2ui`
tool-call** (`TOOL_CALL_START`→chunked `ARGS(components)`→`END`) —
verified byte-faithful to upstream `llama-index-protocols-ag-ui 0.2.2`
plus this one additive step.
- **Flips `injectA2UITool: true`** so the middleware watches
`render_a2ui`.
- Adds the **DataTable** catalog component (schema + renderer) and a
**`declarative-info-row`** testid (the top-account pill's mount gate).
- Aligns `suggestions.ts` to the shared probe pills.

## Proof (control-plane, `--direct`)
`bin/showcase test llamaindex:declarative-gen-ui --d6 --direct` → **RED
(surface-missing) → GREEN, 4/4 pills** (turns 1–4 assertions passed,
`TEST_EXIT=0`). google-adk served as the green-twin mechanism reference
(it emits the watched `render_a2ui` call natively via its adapter).

## Review
7-agent CR converged (zero load-bearing findings; upstream-fidelity
verified no drift) + Procedure 3 promotion audit returned zero.
Re-verified green after the CR doc/prompt/logging fixes.

## Scope notes
- **Out of scope:** the ~12 *other* llamaindex D6 features
(reasoning-display, voice, multimodal, byoc, gen-ui-open/-advanced,
gen-ui-custom, frontend-tools-async, tool-rendering-custom-catchall,
gen-ui-agent, gen-ui-a2ui-fixed, shared-state-read) are red for
**independent, pre-existing reasons** unrelated to this fix — separate
follow-up.
- **Follow-ups (non-blocking):** guard the inner planner `json.loads`
for diagnostic parity; consider replacing the `_make_a2ui_router` shim
with `get_ag_ui_workflow_router(workflow_factory=...)`.
- Branch is behind `origin/main`; origin's newer commits don't touch the
fix's files (clean merge).
2026-06-28 08:38:20 -07:00
Jordan Ritter 280fb747e0 fix(showcase): log llamaindex a2ui planner parse/error/empty-component failures instead of silent no-mount
The render re-emit override had three silent failure paths: a non-JSON tool
output (broad except swallowing TypeError/ValueError), the {"error": ...} dict
from generate_a2ui's no-tool-call branch, and a valid-JSON result missing
components. Each produced a blank UI with no diagnostic trail. Narrow the parse
except to json.JSONDecodeError (guarding that content is a str) and log a
contextual warning on each path. Happy path unchanged.

(cherry picked from commit 94b0a69aa4772758e3bcc05f67a6c28b1b0a503d)
2026-06-28 07:25:00 -07:00
Jordan Ritter b01dfe7ac9 fix(showcase): add DataTable to llamaindex declarative-gen-ui planner prompt catalog
The inlined planner SYSTEM_PROMPT listed every A2UI catalog component except
DataTable, even though the TS catalog and a team-performance suggestion pill
target a DataTable surface. Since the planner is a separate OpenAI call driven
solely by this hardcoded prompt (it never sees the TS Zod schema), DataTable
emission was unreliable. Add DataTable to the catalog list, mirroring the TS
definition (columns/rows shape) and the other entries' wording.

(cherry picked from commit a054e41b8d9394540e1cf7b84ccf9e8e0722c519)
2026-06-28 07:25:00 -07:00
Jordan Ritter 20a283eb69 docs(showcase): soften a2ui_dynamic byte-for-byte claim to functionally-equivalent (upstream 0.2.2)
The override docstring claimed it reproduces the upstream aggregate_tool_calls body byte-for-byte; it is functionally equivalent with two cosmetic diffs (Optional type hint, list comprehension). Reword to match reality.

(cherry picked from commit 5e91118b3a463dcefcd228f6343e472db6081c0f)
2026-06-28 07:24:59 -07:00
Jordan Ritter 196cf1dc6f docs(showcase/aimock): correct stale llamaindex gen-ui-declarative _note to streamed render_a2ui contract
The _note asserted injectA2UITool:false (unchanged) and that flipping to true
would blank-render, and that generate_a2ui returns an a2ui_operations container
for the middleware to forward. Both are now false: this PR set injectA2UITool:true,
generate_a2ui returns raw planner args, and the surface mounts from a streamed
render_a2ui tool-call (START/ARGS/END) the agent re-emits, which the middleware
watches under injectA2UITool:true. Prose-only; no match keys or payloads changed.

(cherry picked from commit 5679b001580615f2e7d988d8c7063994076ace29)
2026-06-28 07:24:59 -07:00
Jordan Ritter 12d4c9c217 docs(showcase): correct llamaindex declarative-gen-ui page comment to injectA2UITool:true
The page header still described the runtime as configured with
`injectA2UITool: false` and the backend agent as owning `generate_a2ui`,
mirroring beautiful-chat. This PR inverted the route to
`injectA2UITool: true`, so the comment was stale. Rewrite the step-3 block
to describe the current mechanism: `injectA2UITool: true` populates the A2UI
middleware's watched-names set, which mounts the surface from a STREAMED
`render_a2ui` tool-call the agent re-emits via its `aggregate_tool_calls`
override in a2ui_dynamic.py. Drops the stale generate_a2ui framing and
matches the accurate header in route.ts.

(cherry picked from commit 407d755638ebe28418f1f8ce2c558f202995284e)
2026-06-28 07:24:59 -07:00
Jordan Ritter b1b4ae6d83 fix(showcase): llamaindex declarative-gen-ui — d6 fixture, DataTable catalog, shared pills
Rebuild the per-integration d6 fixture to kill the missing-arg

generate_a2ui OOM loop; add the DataTable catalog component

(definitions + renderer); align suggestions.ts to shared probe pills.

Completes the integration-only fix: 4/4 pills mount, surface renders.
2026-06-28 07:04:56 -07:00
Jordan Ritter 0ddd3a6b7b fix(showcase): llamaindex declarative-gen-ui — emit declarative-info-row testid for top-account pill
The InfoRow renderer was the only catalog component missing a
data-testid. The top-account pill's _design_a2ui_surface leg emits
7 InfoRow facts + a PieChart, and the d5-gen-ui-declarative probe
asserts declarative-info-row (minCount 1) as top-account's
distinguishing testid. Because the renderer never painted that
testid, the completeOnMount gate (whose surfaceTestIds include
declarative-info-row) never observed a mount, the turn never
completed, and the run reported reason=surface-missing. Every other
pill passed because its distinguishing testid (metric / status-badge /
data-table) was already emitted.
2026-06-28 07:01:07 -07:00
Jordan Ritter 61a31ec701 fix(showcase): llamaindex declarative-gen-ui — emit streamed render_a2ui tool-call so a2ui-middleware mounts the surface
The A2UI middleware mounts the surface from a STREAMED render-tool CALL whose
name is in its watched set, not from a TOOL_CALL_RESULT. The prior approach
emitted a TOOL_CALL_RESULT carrying an a2ui_operations container, which the
middleware never inspects, so the surface stayed unmounted (surface-missing).

- route.ts: set a2ui.injectA2UITool: true so the middleware watches render_a2ui.
- a2ui_dynamic.py: generate_a2ui now returns the planner's render_a2ui args
  (surfaceId/catalogId/components/data) as JSON; the workflow override
  (_A2UIRenderToolCallWorkflow) parses each backend tool result and re-emits it
  as a streamed render_a2ui tool-CALL (TOOL_CALL_START name=render_a2ui →
  chunked TOOL_CALL_ARGS carrying the components JSON → TOOL_CALL_END), mirroring
  how google-adk drives the middleware. These events are already in the upstream
  AG_UI_EVENTS allow-list the SSE router streams against.

Backend still produces the components (no stub). Inner planner tool stays
_design_a2ui_surface, so the d6 fixture needs no re-keying.
2026-06-28 06:51:56 -07:00
Jordan Ritter e57f18aa50 fix(showcase): slow D5 deep sweep cadence to 30min to stop staleness banner flap (#5748)
## Problem

The Ops/coverage-tab worker-family **staleness banner flaps for D5**
("Worker family D5 e2e-deep has not completed successfully since…"). LOW
severity, self-healing — it clears on the next good sweep, but the noise
is constant.

## Root cause

D5's deep sweep runs every 15 min, but the banner fires at **2×periodMs
= 30 min**. A single slow or skipped sweep (browser-pool contention, a
deploy bounce, one long tick) pushes the last success past 30 min and
trips the banner before the next sweep lands.

## The change

Cron `5,20,35,50 * * * *` (every 15 min) → `*/30 * * * *` (every 30 min,
on :00/:30).

`periodMs` is **derived server-side** from the resolved cron via
`periodMsFromCron`
(`showcase/harness/src/fleet/control-plane/run-view.ts`), and the
dashboard banner consumes that value verbatim (no client-side cron
parsing). So changing the cron moves the cadence **and** the banner
threshold in lockstep: periodMs 900000 → 1800000, banner threshold 30
min → 60 min. A 31-min-stale sweep that used to trip the banner now sits
comfortably inside the window.

Also bumped the d5 entry in `FLEET_FAMILY_PERIODS_MS` (the
stale-pending-expiry window) from 15→30 min to track the new cadence,
and updated the d5 fixtures/assertions across the harness and dashboard
test suites.

## Local red-green proof

New test
`showcase/shell-dashboard/src/lib/d5-cadence-banner.redgreen.test.ts`
exercises the **real** banner-decision path with no mocks: the
production `periodMsFromCron` (harness) derives periodMs from the cron
string, feeding the production `isFamilySilent` (dashboard). A
~31-min-stale d5 family flips from silent→quiet purely because the cron
changed.

**RED** (assertion temporarily flipped to `.toBe(false)` for the 15-min
cron, to show the banner DOES fire at 15-min cadence):

```
 ❯ src/lib/d5-cadence-banner.redgreen.test.ts (3 tests | 1 failed) 5ms
     × a 31-min-stale d5 family on the 15-min cron IS silent (banner fires) 2ms

⎯⎯⎯⎯⎯⎯⎯ Failed Tests 1 ⎯⎯⎯⎯⎯⎯⎯

 FAIL  src/lib/d5-cadence-banner.redgreen.test.ts > D5 cadence banner threshold (real periodMsFromCron + isFamilySilent) > a 31-min-stale d5 family on the 15-min cron IS silent (banner fires)
AssertionError: expected true to be false // Object.is equality

- Expected
+ Received

- false
+ true

 ❯ src/lib/d5-cadence-banner.redgreen.test.ts:47:71

 Test Files  1 failed (1)
      Tests  1 failed | 2 passed (3)
```

**GREEN** (correct assertions: 15-min cron → silent=true, 30-min cron →
silent=false, plus the periodMs derivations 900000 / 1800000):

```
 RUN  v4.1.4 .../showcase/shell-dashboard

 Test Files  1 passed (1)
      Tests  3 passed (3)
```

The test asserts, on the real `periodMsFromCron`:
- `periodMsFromCron("5,20,35,50 * * * *") === 900_000` and
`periodMsFromCron("*/30 * * * *") === 1_800_000`
- a 31-min-stale d5 family on the **15-min** cron → `isFamilySilent ===
true` (banner fires — the old behavior)
- the **same** family on the **30-min** cron → `isFamilySilent ===
false` (banner stays clear — the fix)

## Live RED evidence (production, pre-deploy)

Production `/api/runs` d5 family, still on the old cadence:

```json
{
  "family": "d5",
  "label": "D5 e2e-deep",
  "probeKeyPrefix": "d5-single-pill-e2e",
  "schedule": "5,20,35,50 * * * *",
  "periodMs": 900000,
  "nextRunAt": "2026-06-28T07:20:00.000Z",
  "lastSuccessAt": "2026-06-28T06:25:42.804Z"
}
```

## Post-deploy live GREEN verification (not runnable pre-merge)

After deploy: `/api/runs` d5 family shows `schedule == "*/30 * * * *"`
and `periodMs == 1800000`, and the banner stays clear for ≥6 consecutive
sweeps.

## Notes / deviations

- The red-green test lives in the **dashboard** package because
`periodMsFromCron` pulls in `croner`, which only resolves inside the
pnpm workspace; placing it in the harness instead failed `tsc` (the
dashboard's `isFamilySilent` lives in a React/DOM-typed module the
harness `tsc` can't consume). To run the **real** derivation there,
`croner` was added as a dashboard **devDependency** (the dashboard is a
standalone npm app with its own lockfile; devDep, so it's excluded from
the prod `npm ci` build).
- Updated a stale JSDoc worked-example in `run-view.ts` (it used the old
d5 cron as its example; swapped to `"40 * * * *" → 3 600 000`, which is
`*/`-token-safe inside the block comment and unrelated to d5).
- In `use-worker-runs.test.ts`, the shared `familyEntry()` predicate
fixture keeps `periodMs: 900_000` (only its `schedule` string moved to
`*/30`) — the `isFamilySilent` boundary assertions hardcode the
2×=1_800_000 math against that value and would flip if bumped; added a
comment so a future reviewer doesn't "fix" it.

Spec: https://app.notion.com/p/38d3aa381852816499e1d8595ad4f7f5
2026-06-28 06:46:07 -07:00
Ben Taylor ca28191042 feat(drawer): CopilotDrawer host theming, mobile launcher, action icons + thread-management/tooltip fixes (#5707)
## CopilotDrawer — SDK polish + thread-management fixes

Polish and fixes on top of the shipped `<copilotkit-drawer>`
([#5701](https://github.com/CopilotKit/CopilotKit/pull/5701)), surfaced
by real de-fork pilots run against a locally-built SDK + the hosted
Intelligence runtime: **langgraph-js** (rich canvas / inline
`CopilotChat` / dark theme) and **mastra** (the proverbs CoAgents demo /
`CopilotSidebar` / light theme / controlled→uncontrolled provider).
Spans **`@copilotkit/web-components`** and **`@copilotkit/react-core`**
— no example/de-fork changes here (those wait on a release that
publishes web-components + a `react-core` bump that includes
`CopilotDrawer`). Captured as Phase-5 acceptance requirements in the
[design doc](https://app.notion.com/p/3883aa3818528190b5d1f9b6ba26dca0).

### `react-core` — thread management
- **Thread switching works under `<CopilotKit>`.**
`CopilotChatConfigurationProvider` treated any `threadId` prop as
controlled, so the auto-minted, non-explicit threadId that
`<CopilotKit>`'s v1 bridge seeds blocked the drawer's imperative
`setActiveThreadId` / `startNewThread` — thread switching and "+ New"
were dead in every app wrapped in `<CopilotKit>`. A non-explicit
(`hasExplicitThreadId={false}`) threadId is now overridable; a genuine
caller-supplied threadId stays controlled.
- **"+ New" resets the chat.** `CopilotChat` clears messages when
switching to a fresh non-explicit thread so the welcome screen shows
instead of the prior thread's messages — including when a still-loading
`/connect` is superseded by the switch (the aborted connect's stale
snapshot is dropped instead of repopulating the view).
- **No upsell flash/stick.** `CopilotDrawer` treats a pending (`null`)
license status as loading, not unlicensed, so the upgrade upsell no
longer flashes (or sticks) before the runtime reports the license.
- **In-header drawer launcher is mobile-only.**
`CopilotModalHeader.DrawerLauncher` toggles `drawerOpen`, which only
drives the off-canvas *mobile* drawer; on desktop the drawer is a
persistent in-flow panel that ignores `open`, so the launcher was a dead
no-op there. It now renders only at ≤767px. On desktop the chat's own
open/close (the toggle FAB + the header close, both lucide `X`) is the
consistent pair.

### `web-components` — theming, launcher, icons, tooltips
- **Host-theme inheritance (incl. dark mode).** Each token now resolves
`--cpk-drawer-* override → host theme var
(`--background`/`--card`/`--foreground`/`--border`/…) → built-in light
default`. The drawer follows the host app's light/dark theme by
inheritance, instead of baking light-only values (it showed a light
drawer on a dark app). Custom properties aren't reset by the `:host {
all: initial }` guard, so the standalone fallback still holds.
- **Row actions are icon buttons with instant, on-brand tooltips.**
Archive / Unarchive / Delete render inline lucide-style SVGs
(`currentColor`). Native `title` is replaced by an instant CSS tooltip
(matches the React components' Radix `delayDuration: 0`; `title`'s
show-delay is browser-fixed at ~1.5s), styled to match the standard
CopilotKit tooltip — a primary-colored bubble with an arrow, not a
surface/bordered box — and positioned to the side so the list's
scroll-overflow never clips it. The tooltip is suppressed while the
delete-confirm dialog is open (the clicked trash button keeps
`:focus-visible` otherwise, leaving its "Delete" tooltip floating over
the dialog).
- **Clipped thread names get a tooltip.** A truncated name exposes its
full text via the same instant primary-bubble tooltip as the row actions
(not the native `title`), shown only when actually clipped. The name
text moved to an inner span (so the outer can host the un-clipped
bubble) and the hovered row is z-lifted (each row is its own stacking
context from the entry-animation transform, which otherwise painted the
bubble under later rows).
- **Archived rows are muted**, no longer struck through.
- **Refetch preserves the list.** A refetch (e.g. Active↔All toggle)
keeps the known threads visible; the full loading state only renders on
the initial empty fetch.
- **Empty `footer` region hides** (it rendered as a stray bordered box
at the bottom). Driven by a `slotchange` listener, mirroring the
`memories` region.
- **Self-owned mobile launcher.** When the drawer is a closed off-canvas
mobile modal it renders its own floating "open threads" launcher, so
it's always openable without the host wiring a header button. The icon
is swappable via a `launcher-icon` **slot**, position is themeable via
`--cpk-drawer-launcher-top` / `--cpk-drawer-launcher-left` (so a host
can center it on its own header controls), and it exposes
`part="launcher"`. The richer in-header
`CopilotModalHeader.DrawerLauncher` remains an optional integration.

### Tests
- **react-core**: thread-control + non-explicit-override cases on
`CopilotChatConfigurationProvider`, the "+ New" message reset on
`CopilotChat`, license-pending gating on `CopilotDrawer`, and the
mobile-only / desktop-hidden launcher on `CopilotModalHeader`. New
regression tests verified red-green; chat + provider suites green, `tsc
--noEmit` clean, build green. (The superseded-connect *race* guard is
covered by live verification + its synchronous sibling test — an
async-timing unit test would be flaky, which the repo treats as a bug.)
- **web-components**: element tests for the icon tooltips +
confirm-dialog suppression, the clipped-name tooltip (`data-tooltip` +
placeholder), archived styling, refetch-preserves-list, footer
hide-when-empty, and mobile launcher render/open + desktop-absent.
Drawer suite green, `tsc --noEmit` clean, build + `es-check` green under
`CI=true`. (Truncation detection + tooltip stacking are layout/CSS —
verified live, not in jsdom.)

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-28 08:44:06 -05:00
Jordan Ritter af306b52c5 fix(showcase): slow D5 deep sweep cadence to 30min to stop staleness banner flap 2026-06-28 00:15:27 -07:00
Jordan Ritter d0758ebd3b feat(showcase): deploy rollover (overlap/draining) config-as-code + test cleanups (reclamation layer c) (#5747)
## What

Layer (c) of the worker-reclamation redesign: **deploy rollover as
Railway config-as-code** — no custom rolling-restart code. Brings the
two Railway deploy-teardown knobs under SSOT so a harness-workers
redeploy is **non-lossy and capacity-stable**, completing the a/b/c
reclamation arc.

## The two knobs (currently `0`/unset on staging — which is exactly why
deploys cause the staleness dip *and* SIGKILL-lose in-flight cells)

| Knob (`serviceInstanceUpdate` field / env var) | Effect | Target |
|---|---|---|
| **`overlapSeconds`** (`RAILWAY_DEPLOYMENT_OVERLAP_SECONDS`) — the
"Teardown" card | New deployment goes Active (healthy) **before** the
old is torn down → capacity overlap across all 6 replicas → **no
capacity-floor drop** | **45** |
| **`drainingSeconds`** (`RAILWAY_DEPLOYMENT_DRAINING_SECONDS`) |
SIGTERM→SIGKILL buffer → lets layer (b)'s 90s graceful drain
finish-and-report before the kill | **180** |

Sequence: *new Active → overlap (capacity held) → SIGTERM old → drain
(in-flight cell finishes+reports) → SIGKILL.* Composes with the merged
layer (a) (reclaim backstop) + layer (b) (graceful drain).

## Why config, not code

Investigation (with research against Railway's docs + live GraphQL)
established that **`/health` only flips Active once a worker is
registered AND claiming** (`worker-health.ts` returns 200 iff `pb() &&
loopAlive() && registered()`; the `/health` socket binds *after* the
claim loop). So `overlapSeconds` genuinely holds the capacity floor —
the load-bearing promise of layer (c) — without any custom
rolling-restart code. A single-worker drain (the original B-VAL plan) is
**infeasible** on staging's 6-replica single-deployment topology
(Railway has no per-replica primitive); the real-world operation is a
full redeploy, which is what we validate.

## Changes

- **SSOT**: `overlapSeconds=45` + `drainingSeconds=180` declared for
prod+staging in `railway-envs.ts`, regenerated
`railway-envs.generated.json`, drift-gate extended (bites on
wrong/missing values).
- **RAILWAY.md**: "Deploy rollover" section — knobs, rationale, layer
a/b/c composition, GraphQL/dashboard apply instructions.
- **`orchestrator.ts`**: comments-only — corrected stale pre-layer-(b)
prose ("drain abandons / never reports / 6s grace / ~10s stop") to the
current finish-and-report behavior.

## Pre-existing test cleanups (flushed out while validating this)

- **Stale `orchestrator.test.ts` drain test** — layer (b) (#5743)
changed drain to finish-and-report but missed this assertion (it lives
outside the `src/fleet/` subset #5743's CR ran), leaving it
**deterministically RED on `main`**
(`expect(report).not.toHaveBeenCalled()` where report *is* called — the
run finishes before grace). Flipped to the finish-and-report contract,
matching the canonical `worker-loop.test.ts:2162`. The sibling abandon
test (releases *after* grace) is unchanged and still passes.
- **`probe-invoker.test.ts` timing flake** — a brittle wall-clock
assertion (`elapsed < 100ms`, only 20ms above the driver's sleep) flaked
under load (12/15 fail). Made deterministic with fake timers + a
behavioral assertion (times-out-fast invariant preserved, now stronger).

## Validation

- Config CR'd (full reviewer round, 0 P0/P1) + test-gated confirmation
on the combined branch.
- Full harness suite **3209 / 0 failed**; probe-invoker **10/10**
deterministic; orchestrator flip red-green proven; tsc + oxlint clean;
SSOT generated.json byte-canonical.

## Pending your GO (not done in this PR)

- **Apply** `overlapSeconds=45`/`drainingSeconds=180` to staging
harness-workers.
- **Validate** via a real staging redeploy: capacity floor never 0 +
zero lost cells + no false family-silence alert (before/after dashboard
screenshots).
- **Merge**.

## Follow-ups (non-blocking)

- `registered` is a one-shot boot flag (not re-derived from heartbeats);
a PB blip during boot-verify could 503 `/health` forever — but it
**fails safe** (holds the old deployment longer, never releases early).
- `showcase/scripts` emit-railway-envs-json tests need the root-declared
`oxfmt` binary to run in isolation (scripts-only `npm ci` → `spawnSync
ENOENT`); declare it in `scripts/package.json` or skip-with-diagnostic.

Refs the reclamation redesign Notion proposal.
2026-06-27 23:19:30 -07:00
Jordan Ritter 400d90e87e fix(showcase): make probe-invoker timeout test deterministic (kill timing flake) 2026-06-27 22:54:01 -07:00
Jordan Ritter 19962ba5c7 fix(showcase): flip stale orchestrator drain test to finish-and-report contract (layer-b cleanup #5743 missed) 2026-06-27 22:40:48 -07:00
Jordan Ritter aaed2dc545 fix(showcase): correct stale drain comments in orchestrator (post-layer-b)
Three stale, pre-layer-(b) prose sites still claimed drain() ABANDONS the
in-flight run and the loop skips queue.report: the drainFleetWorker JSDoc, the
runWorker stop() inline comment, AND the case-worker SIGTERM-handler block.
Layer (b) changed this — on SIGTERM the worker stops claiming, lets its
in-flight cell FINISH within the 90s grace, and REPORTS its real terminal
result via the runAbort/abortedWithoutResult discriminator; only a run that
OVERRUNS the grace is abandoned -> reclaimed by layer (a). Preserved the still-
accurate drainReason=shutdown red-side-emit soft-wind-down + deregister-vs-crash
distinction. Also refreshed the WHY-THIS-ORDER preamble (90s finish-and-report
grace, 180s platform stop / layer-c drainingSeconds). Comments only — no
behavior change.
2026-06-27 22:22:11 -07:00
Jordan Ritter 2225a09ff9 feat(showcase): bring deploy rollover (overlap/draining) under SSOT
Layer (c) deploy-rollover config for the harness-workers fleet — pure
Railway config, no custom rolling-restart code. Declares overlapSeconds=45
(capacity floor: old deployment serves until new workers register+claim, so
no staleness dip) and drainingSeconds=180 (SIGTERM->SIGKILL window >=
PLATFORM_STOP_GRACE_MS so the layer-(b) 3s+90s composed worker-drain finishes)
for both prod and staging in the railway-envs SSOT, regenerates the JSON
snapshot, and extends the harness-workers drift gate so CI fails if either
field drifts between SSOT and snapshot. Documents both knobs, their rationale
(incl. that Railway's draining default is 0s = immediate SIGKILL), the
composition with drain layers a+b, and how to apply them (GraphQL
serviceInstanceUpdate / dashboard) in RAILWAY.md.
2026-06-27 22:22:00 -07:00
Jordan Ritter 5a78a2a0e3 fix(showcase): graceful worker drain — reclamation layer (b) (#5743)
## What
Layer (b) of the worker-reclamation redesign: **graceful worker drain**.
On a deploy/shutdown signal, a worker now stops claiming new jobs but
lets its in-flight job **finish and report** (instead of abandoning it),
keeping the lease renewed until it completes or a grace window expires.
Builds on the merged layer (a) reclaim-wins-until-cap reaper (#5738).

## How (B0–B5)
- **B2 — decouple:** a separate `runAbort` AbortController drives the
run's `ctx.abortSignal`; `drainSignal` now only stops claiming + drives
`drainReason`/heartbeat. The abandon `break` is conditional on
`abortedWithoutResult = runAbort.signal.aborted`, so a finished run
falls through to the report path.
- **B3 — lease renewal:** the heartbeat is re-keyed to the grace-expiry
signal, so a finishing job keeps renewing its lease (the reaper can't
reclaim it mid-finish); a genuinely abandoned job stops renewing and
lapses.
- **B4 — aborted-only suppression:** drain red-cell suppression fires
only for aborted/abandoned runs (gated on `errorClass==="abort"`); a
finished run reports its real terminal. (Already aborted-only by
construction post-B2; pinned with a mutation-validated test.)
- **B5 — grace budget:** `WORKER_DRAIN_GRACE_MS=90s` (env-overridable),
under the 300s lease ceiling; `PLATFORM_STOP_GRACE_MS=180s` documents
the Railway `terminationGracePeriod` requirement (3+90 < 180, ≥30s
headroom) that **layer (c) / C3** will apply.
- At grace-expiry, `runAbort.abort()` fires → the job is abandoned →
layer-(a) reclaim is the backstop.

## Review
7-agent CR + 7-agent confirmation round → **0 P0 / 0 P1**. All 5
red-green proofs empirically re-verified (each mutation flips its test
RED, restore GREEN). Fleet suite **679/679**, tsc clean, oxlint 0
errors.
- A contested grace-boundary TOCTOU (could a run finishing in the same
microtask turn the grace timer fires be spuriously abandoned?) was
**empirically confirmed SAFE** via `Promise.race` ordering —
`runAbort.abort()` lives only in the timeout leg, which loses the race
once `done` is resolvable — and regression-pinned.
- `clockHonoringSleep` is a **test-only** helper (fake-clock fix), zero
production footprint.
- Layer-(a) composition verified: a clean drain-then-finish resets
`consecutive_orphan_count` (no false poison-cap), grace-expiry abandon
follows the normal orphan→reclaim path.

## Not in this PR (gated on your GO)
- **B-VAL live-staging validation** (drain a real staging worker
mid-column, confirm zero lost cells + no staleness dip + no false
silence alert) — PENDING explicit user GO since it mutates the running
fleet.
- **Layer (c) rolling restart** — follow-on, after (b) merges + B-VAL
passes.

Refs the reclamation redesign Notion proposal.
2026-06-27 21:27:21 -07:00
Jeel Gor 9632ca901c Merge branch 'main' into fix/5601-windows-cross-platform-scripts 2026-06-27 15:48:22 +05:30
Jordan Ritter 04177348d1 style(showcase): oxfmt graceful-drain touched files 2026-06-26 23:40:14 -07:00
Jordan Ritter dfb59c3576 fix(showcase): pin same-turn drain race as safe + correct stale requestDrain JSDoc
CR P2 round on layer-(b) graceful worker drain.

P2-A (contested TOCTOU, reconciled empirically): a run that resolves with a
valid result in the same flush the grace setTimeout fires is REPORTED, not
spuriously abandoned. runAbort.abort() lives only in stop()'s Promise.race
TIMEOUT leg, which loses the race once `done` is resolvable — so a finished
run never trips the abortedWithoutResult discriminator. Reviewer crb6 was
correct; the TOCTOU is a non-bug, so no logic changed. Added a deterministic
fake-timer regression pin forcing both the run-completion timer and the grace
timer due in one advance (verified to BITE under an over-aggressive
abandon mutation), plus a clarifying comment at the discriminator.

P2-B: corrected the requestDrain() JSDoc — post-B2 the report-skip,
heartbeat-stop, and driver-cancel key on the grace-expiry signal
runAbort.signal, not stopAbort.signal (which now only stops claiming).
2026-06-26 23:29:34 -07:00
Jordan Ritter 09b78f419a fix(showcase): raise drain grace to bound one cell-job; re-derive composed SIGKILL-window budget
B5 — set the production drain grace T to 90s (was 6s). The grace stopped being
a bare TEARDOWN budget and is now the FINISH-AND-REPORT budget stop() waits for
an in-flight run to finish-and-report (layer b) before firing runAbort -> abandon
-> layer-(a) reclaim. T must BOUND a typical cell-job (so a normal in-flight job
finishes within grace) yet stay SHORTER than the platform SIGTERM->SIGKILL window
with headroom. Sized from in-repo cell-job signal: a single-service cell-job runs
~15s (light e2e-deep) up to ~200s (heavy d6-all-pills under contention); the
per-job lease ceiling is 300s. 90s covers the bulk and stays well under the lease
so a finishing job's lease never lapses. The tail (a job that cannot finish in
grace) falls back to layer (a) — grace is deliberately FINITE. 90s is a defensible
default; B-VAL confirms/retunes from staging p95. Env-overridable via
WORKER_DRAIN_GRACE_MS.

Introduce PLATFORM_STOP_GRACE_MS (180s) documenting the C3 requirement: layer-(c)
must set Railway terminationGracePeriodSeconds = 180 so the composed serial budget
DRAIN_DEREGISTER_TIMEOUT_MS (3s) + DEFAULT_WORKER_DRAIN_GRACE_MS (90s) fits with
>=30s headroom for the health-server-close + pool-shutdown remainder. The
composed-budget test now pins the relation 3s + grace < PLATFORM_STOP_GRACE_MS
(was hardcoded < 10s) and the concrete numbers.

Add a behavioral fake-clock test of the composed-budget invariant: a run that
finishes WITHIN T is reported (finish-and-report); a run still running AT
grace-expiry fires runAbort -> abandons WITHOUT a usable result -> not reported
(layer-a reclaim backstop). The raise also broke the existing wedged-driver
default-grace test (advancing a 90s fake span re-armed the 0ms-yielding heartbeat
into a runaway cascade); fixed by giving it a clock-honoring sleep so the
heartbeat stays quiet across the window and only the grace timer is crossed.
2026-06-26 23:15:30 -07:00
Jordan Ritter 2b741f39df fix(showcase): pin aborted-only drain-red suppression so a finished-on-drain run reports its real terminal result
Post-B2/B3 a run can finish-and-report after a graceful drain (drain no
longer hard-cancels the run; only grace-expiry runAbort does, which is
ctx.abortSignal). The d6 red-suppression already AND-s drainReason with
ctx.abortSignal.aborted and errorClass==="abort", which makes it
aborted-only by construction — a finished run whose legitimate red is a
genuine failure (errorClass!="abort") is NOT suppressed. This adds a
pinning test for that aborted-only contract (a finished goto-error red
under drainReason=shutdown with an un-fired abort signal must still be
reported). Verified RED via a temporary mutation to the over-broad
"drainReason alone" suppression form, GREEN on the real code.
2026-06-26 23:00:38 -07:00
Jordan Ritter f1cc1c8f08 fix(showcase): keep the lease renewing for a finishing-on-drain job so the reaper cannot double-claim it
B3 (layer-b graceful drain): the heartbeat-abort previously keyed on the
DRAIN signal, stopping lease renewal at drain-start. With B2 decoupling drain
from run-abort, a still-finishing job now runs past drain-start until grace
expiry, so killing renewal at drain-start lets the lease lapse mid-finish and
the layer-(a) reaper could reclaim the row out from under the worker
(double-run / report-after-reclaim). Gate the heartbeat-abort on the
grace-expiry signal (runAbortSignal) instead: a finishing job keeps renewing
until it reports terminal; a genuinely-abandoned job (runAbort fired at
grace-expiry) still stops renewing so its lease lapses and the reaper reclaims
it. runAbortSignal defaults to drainSignal, preserving direct-call unit-test
semantics.
2026-06-26 22:54:45 -07:00
Jordan Ritter 0aacd72d66 fix(showcase): finish-and-report a completed run on graceful drain instead of abandoning it
Decouple 'stop claiming new jobs' from 'abort the in-flight run': drain()
no longer fires the run's abortSignal. The run's hard cancel is now a
separate grace-expiry signal (runAbort), so a run seconds from done
finishes within grace and is reported instead of being abandoned to the
sweeper. The abandon break is now conditional on abortedWithoutResult
(runAbort fired at grace-expiry → no usable result); a finished run falls
through to the report path. A wedged run that overruns grace is still cut
(runAbort.abort() in stop()'s timeout leg) and abandons → layer (a)
reclaim is the backstop.

Reconcile the drain tests that encoded the old abort-on-drain coupling to
the new contract (drain no longer fires ctx.abortSignal; the grace-expiry
abort fires at grace, exercised via a short WORKER_DRAIN_GRACE_MS).
2026-06-26 22:49:38 -07:00
Jordan Ritter 5afb474c7f test(showcase): pin desired finish-and-report drain semantics (red) 2026-06-26 22:44:13 -07:00
Tyler Slaton 9fcfaf96f6 example(shadcn): rename project and add asset link in README (#5740) 2026-06-26 17:40:59 -07:00
Tyler Slaton 35abf430f2 example(shadcn): rename project and add asset link in README 2026-06-26 17:40:41 -07:00
Tyler Slaton 5ebda711f3 example(shadcn): add example using new components and shadcn primatives (#5739) 2026-06-26 17:39:55 -07:00
Tyler Slaton 0759f63aae example(shadcn): add example using new components and shadcn primatives
Signed-off-by: Tyler Slaton <tyler@copilotkit.ai>
2026-06-26 17:37:06 -07:00
Tyler Slaton 1e30edcf18 refactor(bot): drop unpublished bot-store-redis/postgres adapters (#5700)
## What

Removes the `@copilotkit/bot-store-redis` and
`@copilotkit/bot-store-postgres` adapters, added in #5613. The pluggable
`StateStore` interface and the in-memory `MemoryStore` default stay;
durable backends can be reintroduced as a follow-up when there's a
concrete need.

## Why

Both adapter packages were merged but **never published to npm**, so
there are no consumers — removal is a clean delete with no migration or
deprecation cycle. Trimming the surface keeps the bot persistence story
to one well-tested in-memory default plus a documented "bring your own
`StateStore`" path.

## Changes

- Delete `packages/bot-store-redis` and `packages/bot-store-postgres`.
- Revert the `bot` release scope and the release drift-guard test to
`bot + bot-ui` (count 2).
- Strip the Redis dependency, `demo:restart` script, restart demo,
`docker-compose.yml`, and `REDIS_URL` env from `examples/slack`.
- Rewrite the bot persistence/transcripts docs around "MemoryStore
default + implement the `StateStore` interface yourself for durability"
(unrelated enterprise/Helm Redis/Postgres docs untouched).

## Verification

- `nx run @copilotkit/bot:build` — pass
- `nx run @copilotkit/bot:test` — 121/121 pass
- Release drift guard — pass at count 2
- Repo-wide grep for
`bot-store-redis|bot-store-postgres|createRedisStore|createPostgresStore`
— zero matches
2026-06-26 16:42:38 -07:00
Jordan Ritter 31abb3eff3 fix(showcase): reclaimable leases — reclaim-wins-until-cap reaper with consecutive-orphan budget (#5738)
## What
Replaces the queue reaper's delete-wins behavior for stale long-expired
orphaned in-flight rows with **reclaim-wins-until-cap**: an orphaned row
is re-queued (not deleted) up to a bounded number of CONSECUTIVE
re-orphans before final claim-delete. This is layer (a) of the
worker-reclamation + graceful-rollover redesign — it makes a worker
bounce mid-column non-lossy (the column's in-flight work is reclaimed by
surviving workers instead of dropped).

## How
- **Reclaim-wins-until-cap** in `runSweepExpired`
(`showcase/harness/src/fleet/queue-client.ts`): stale long-expired
orphaned in-flight rows are re-queued; the prior delete-wins carve-out
is inverted.
- **Consecutive-orphan budget** (`consecutive_orphan_count`, migration
`1779990400`): the reclaim cap (`MAX_RECLAIM_ATTEMPTS=3`) is scoped to
*consecutive* sweeper re-orphans — incremented only on the sweeper
re-queue path, **reset to 0 on terminal `done`/`failed`**. Peer-worker
expired-lease *steals* bump the lifetime `reclaim_count` (retained as a
dashboard diagnostic) but do NOT consume the reclaim budget.
- **Stale-age re-anchoring** (`requeued_at ?? created`, migration
`1779990300`): a reclaimed row's staleness clock restarts so it isn't
immediately re-expired.

## Review
7-agent CR + 7-agent confirmation round → **0 P0 / 0 P1**. Highlights:
- **Red-green GENUINE (empirically re-verified):** P1-A — a long-lived
job with lifetime `reclaim_count=3` but `consecutive_orphan_count=0` is
RE-QUEUED on a fresh orphan (pre-fix: wrongly deleted); P1-B — low
boundary at `MAX-1=2` re-queues, mutation-killed. Both non-tautological
(the test fake bumps `consecutive_orphan_count` only on reclaim, never
on steal; a dedicated pin test confirms steals don't consume budget).
- **No regression / terminates / fails-safe:** 248/248 tests, tsc +
oxlint clean; reclaim loop terminates at 3 consecutive orphans →
claim-delete; queue cannot grow unbounded; SweepResult metrics +
`reclaim_count` dashboard consumers (run-view `jobs.reclaimed`,
family-silence) unbroken.
- **Migration `1779990400`** additive, nullable, default-0;
pre-migration rows evaluate to 0 → safe re-queue path.

## Follow-on (not in this PR)
- Layer (b) graceful worker drain and layer (c) rolling restart are the
remaining phases of the redesign — see the Notion proposal:
https://app.notion.com/p/38b3aa381852817bacf5c9cda1f11cc0

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-26 15:46:05 -07:00
Jordan Ritter d4d363d940 fix(showcase): tighten D4/BE probe gate to a provenance check (catch the BIA-style dead-agent false-pass) (#5737)
## What

Tightens the **D4/BE chat-roundtrip probe gate** from `text.length > 0`
to a **provenance check**. A turn now only counts green if the assistant
text came from the `[data-testid="copilot-assistant-message"]` container
— not from the `<body>` fallback scrape (which picks up static page
chrome).

## Why

This is the gate weakness that **masked the BIA prod outage**: a dead
agent (RUN_STARTED→RUN_FINISHED with zero `TEXT_MESSAGE`) left the page
chrome on screen, the `<body>` scrape returned non-empty text, and the
old `text.length > 0` gate reported **green** while the agent was
actually broken. The new gate threads a `fromAssistantContainer` flag
(set true only on the testid-container read path) so a body-scrape-only
result fails (`!fromAssistantContainer` → red). The same guard is
applied at L4 (tool-rendering).

## Review

**7-agent CR → 0 P0 / 0 P1.** Highlights:
- **Red-green genuine** (empirical): old gate false-passes BIA (1/52
RED); post-fix 52/52. 3-way guard (green /
empty-via-unchanged-length-guard / BIA-via-provenance).
- **Blast radius: no false-fail risk** — all 20 integrations'
`agentic-chat` + `tool-rendering` demos use the shared v2 `CopilotChat`
which mounts the testid unconditionally; aimock `d4/<slug>/chat.json`
fixtures are text-bearing for all 20. Verdict: **ship as a blocking
gate, no soak.**
- **No lost coverage** — the `<body>`-scrape fallback *code* is retained
(diagnostics); only the gate *semantics* narrowed. No D4-probed cell
relied on the fallback to pass.
- **No cross-family regression** — change is confined to
`d4-chat-roundtrip.ts`; other probe drivers untouched; 169/0 tests pass;
tsc clean.

## Notes / follow-ups (non-blocking, P2)
- A *future* probed demo using a `CopilotChatAssistantMessage`
children-render override (which doesn't emit the testid) would
false-fail — worth a one-line note near the fallback comment if that
pattern is ever added to a D4 route.

Part of the 2026-06-26 incident remediation (companion to #5733); this
gate would have surfaced the BIA failure that the dashboard masked.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-26 15:46:02 -07:00
Jordan Ritter ff2fca4776 fix(showcase): incident remediation — dashboard SWR + family-silence grace + llamaindex log-gate + SSOT worker-provisioning (#5733)
## What

Bundles the four code fixes from the 2026-06-26 showcase prod incident
remediation. (Prod config reconciles — fleet image promotes, worker
scaling, CVDIAG keys — were applied directly to Railway and are not
code; the post-incident debugging lesson shipped separately in #5727.)

## Incident context (why these exist)

Prod's coverage dashboard cascaded to a wall of red. Root cause was
**staleness, not a feature break**: the harness worker pool was starved
(a deploy-bounce + under-provisioning), so probe sweeps couldn't
complete within the staleness windows → cells aged out → rendered
red/BE✗ even though the apps were fine. Compounded by a `llamaindex`
container crash (per-request log flood tripping Railway's log cap) and a
2-day-stale prod image fleet. See #5727 for the debugging lesson, and
the Notion proposal below for the durable worker-reclamation redesign.

## The four fixes

1. **Dashboard stale-while-revalidate**
(`shell-dashboard/depth-chip.tsx`) — a passing (green) cell keeps its
color + a non-destructive `⟳` refreshing affordance during a re-probe
instead of flapping to grey; never-run stays grey; failure/regression
keeps its color with no spinner; staleness bound preserved. + a11y
(`role=status`, flag-gated regression label, `motion-reduce`).
2. **Family-silence grace window** (`harness/fleet/control-plane/*` +
dashboard banner) — a normal harness deploy's post-bounce worker drain
no longer fires a false "worker family silent" banner/alert (suppressed
within a `2×period` bounce grace keyed on the freshest worker
registration), while genuine silence beyond the window still fires on
both the Slack-alert and banner paths. No silent failure (PB-down →
grace disabled → real outages still alert).
3. **llamaindex per-request log gate**
(`integrations/llamaindex/.../route.ts`) — gates the chatty
`[copilotkit/route]` per-request logs behind `SHOWCASE_ROUTE_DEBUG`
(default off) so they can't flood Railway's log cap and kill the
container. Error logging untouched.
4. **SSOT worker-provisioning** (`scripts/railway-envs.ts` + drift gate)
— brings harness-workers replica provisioning under SSOT so prod/staging
can't silently drift. Models the **effective** field
(`multiRegionConfig.<region>.numReplicas`) +
`BROWSER_POOL_MAX_CONTEXTS`, declares current reconciled reality
(prod=staging=6 replicas, 40 contexts), with a CI drift-detection test.

## Review

7-agent CR round + 7-agent confirmation round → converged to **0 P0 / 0
P1**. Two post-CR fixes (SSOT effective-field model after discovering
`multiRegionConfig` is the live knob; re-anchoring a grace-edge test to
genuinely pin the 2× boundary) each re-confirmed clean. Suites:
shell-dashboard 92 + harness 63 + scripts 12 = **167 passed, 0 failed**.
No new tsc errors (pre-existing only).

## Notes / follow-ups (non-blocking)

- llamaindex log-gate is logging-only and untested (acceptable); the
chatty pattern exists in ~16 integrations incl. the gold standard — a
**fleet-wide log-flood gate** is a worthwhile follow-up.
- `freshestBounceMs` has two equivalent impls (harness `parseIso` /
dashboard `Date.parse`) — flagged for future lockstep.
- The D4/BE probe gate weakness (asserts `text.length>0`, which masked
the BIA outage) is being addressed in a **separate** PR
(`fix/d4-gate-tighten`).

## Refs
- Post-incident debugging lesson: #5727
- Worker reclamation + graceful-rollover redesign (Notion proposal):
https://app.notion.com/p/38b3aa381852817bacf5c9cda1f11cc0

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-26 15:45:59 -07:00
Jordan Ritter c3a13ac988 fix(showcase): scope reclaim cap to consecutive orphans, not lifetime steals (P1-A + P1-B)
P1-A (cap semantics): the MAX_RECLAIM_ATTEMPTS cap was keyed on
`reclaim_count`, a LIFETIME tally bumped by BOTH the sweeper re-queue
path AND the peer-worker expired-lease steal (claim CAS). A long-lived
job that accrues benign peer steals could exhaust its 3-budget and then
get claim-DELETED on its first real orphan rather than re-queued.

Fix: introduce a dedicated `consecutive_orphan_count` column (migration
1779990400) that is bumped ONLY by the sweeper re-queue path in the
fleet-claim release CAS, and reset to 0 on every terminal done|failed
release. The peer-worker steal (claim CAS wasExpiredSteal branch) does
NOT touch this counter. The reaper's cap check now uses
`consecutive_orphan_count` instead of `reclaim_count`. `reclaim_count`
is left intact as the lifetime dashboard diagnostic (jobs.reclaimed).

P1-B (low boundary): adds a test at consecutive_orphan_count = MAX-1
(= 2) asserting the row is RE-QUEUED, not deleted. The off-by-one
mutation `>= MAX` -> `>= MAX-1` causes this test to go RED.

Test-fake honesty: `makeReclaimClaim`'s claimJob now explicitly models
the steal-bump on `reclaim_count` (matching the real hook) while
intentionally NOT bumping `consecutive_orphan_count`, and adds an
explicit pin test confirming steals do not consume the reclaim budget.

JSDoc on MAX_RECLAIM_ATTEMPTS updated to describe the correct semantics:
consecutive re-orphans scoped by sweeper re-queue, reset on terminal.

Red-green proof:
- P1-A RED: revert cap to reclaim_count → "P1-A CAP SCOPE" fails with
  `expect(undefined).toBeDefined()` (job deleted instead of re-queued)
- P1-A GREEN: consecutive_orphan_count cap → test passes (re-queued)
- P1-B RED: mutate `>= MAX` to `>= MAX-1` → low-boundary test fails
- P1-B GREEN: revert mutation → low-boundary test passes

Suite: 145 queue-client + 103 producer = 248 total, all green.
2026-06-26 15:21:41 -07:00
Benjamin Taylor 9223fc476a fix(web-components): stamp name-clipped on the row so the tooltip z-lift fires
_syncNameClipping toggled name-clipped on .row-name, but the stacking fix in
styles.ts targets .row.name-clipped:hover — a different element — so the z-index
lift never matched and clipped-name tooltips still painted under later rows
(each row is its own transform stacking context). Stamp the flag on the owning
.row too; the tooltip bubble stays scoped to .row-name:hover so a row-action
hover never surfaces it. Adds a red-green contract test asserting the class
lands on both the row-name and the row.

Addresses MikeRyanDev review on #5707.
2026-06-26 17:11:25 -05:00
Jordan Ritter 159de7b1ae feat(showcase): reclaimable leases — invert long-expired carve-out to reclaim-wins until cap
Layer (a) of the worker reclamation+rollover redesign: make a worker bounce
non-lossy. The reaper's long-expired carve-out (G1d) used to claim-DELETE an
orphaned in-flight (claimed/running) row whose lease expired beyond its
family's stale window AND whose created-age was past that window —
`reclaimed=0, expiredPending++` — silently dropping work an abrupt bounce
(SIGKILL past grace / OOM / crash) left mid-flight.

Invert it to RECLAIM-WINS-UNTIL-CAP: re-queue the orphan to pending (it
re-runs; idempotent probes make at-least-once safe) until its durable
`reclaim_count` reaches MAX_RECLAIM_ATTEMPTS (3), only then claim-deleting a
row that keeps re-orphaning so a poison job cannot loop forever.

The carve-out existed to dodge an honesty bind: re-queueing a `created`-stale
row emitted a "back in flight" gray the next sweep falsified by claim-deleting
it off the renewal-immune `created` age. Dissolve the bind with a new
`requeued_at` column (migration 1779990300) the release CAS stamps on every
pending re-queue; both stale phases now age off `staleAgeAnchorMs`
(`requeued_at ?? created`), so a reclaimed row is genuinely young again and the
next sweep does not delete it. `reclaim_count` (migration 1779990200) is reused
as the attempt counter — no second tally to drift.

Red-green proven on the real reaper (queue-client.test.ts): a stale-aged
long-expired orphan below the cap goes RED (deleted, reclaimed=0,
expiredPending=1) on delete-wins and GREEN (re-queued, reclaimed=1,
expiredPending=0, requeued_at stamped) on reclaim-wins; plus an attempt-cap
deletion test and a next-sweep no-falsification test.
2026-06-26 15:02:32 -07:00
Jordan Ritter 7364bdfe43 fix(showcase): tighten D4/BE chat gate to require a real assistant turn
The L3 chat-roundtrip gate only asserted `text.length > 0`, which the
`<body>` fallback scrape can satisfy with static page text (nav links,
footer copy, demo blurb) trailing the sent message. When the agent run
finished with ZERO assistant content (RUN_STARTED -> RUN_FINISHED, no
TEXT_MESSAGE) the assistant-message bubble never rendered, so the probe
fell back to scraping <body> and false-PASSED on incidental page chrome —
the mechanism that masked the BIA outage (dead agent reported green).

Track the provenance of the captured response and require it to come from
the [data-testid="copilot-assistant-message"] container (a genuine
assistant turn — the DOM-layer equivalent of RUN_FINISHED + a non-empty
TEXT_MESSAGE), not the body fallback. Apply the same guard to L4 so weather
content must also come from a real assistant turn. The instrumented probe
sends the aimock fixture headers and renders real content into the
container, so this does not affect a working agent; it only rejects the
empty-turn false-pass.

Adds a red-green regression test reproducing the BIA false-pass.
2026-06-26 14:59:55 -07:00
Benjamin Taylor 31f31eeebd fix(web-components): tooltip for clipped CopilotDrawer thread names
A long thread name is clipped with an ellipsis; expose the full text via the
same instant primary-bubble tooltip as the row actions (not the native `title`),
shown only when the name is actually truncated. The name text moved to an inner
span so the outer `.row-name` can host the un-clipped bubble, and the hovered
row is z-lifted so the bubble isn't painted under later rows (each row is its
own stacking context from the entry-animation transform).
2026-06-26 16:54:55 -05:00
Benjamin Taylor a4d129d8fa fix(react-core): render the CopilotModalHeader drawer launcher only on mobile
The in-header thread-drawer launcher toggles `drawerOpen`, which only drives the
off-canvas MOBILE drawer; on desktop the drawer is a persistent in-flow panel
that ignores `open`, so the launcher was a dead no-op there (it appeared to do
nothing on click). Gate it on a reactive mobile viewport check so it renders
only at ≤767px. On desktop the chat's own open/close — the toggle FAB + the
header close (both lucide X) — is the consistent pair; nothing masquerades as a
chat control.
2026-06-26 16:54:55 -05:00
Benjamin Taylor 4fbb5d1f65 fix(web-components): match CopilotDrawer row-action tooltips to the standard CPK tooltip
Row-action tooltips used a surface background + border, which read as a button.
Restyle to the standard CopilotKit tooltip: primary bubble, primary-foreground
text, an arrow, no border. Also suppress the tooltip while the delete-confirm
dialog is open — the clicked trash button keeps :focus-visible, which otherwise
left its "Delete" tooltip floating over the dialog.
2026-06-26 16:54:55 -05:00
Benjamin Taylor af854aa5ff fix(react-core): make CopilotDrawer thread switching work under <CopilotKit>
CopilotChatConfigurationProvider treated any `threadId` prop as controlled, so
the auto-minted non-explicit threadId that <CopilotKit>'s v1 bridge seeds
blocked the drawer's imperative setActiveThreadId/startNewThread — thread
switching and "+ New" were dead in every app wrapped in <CopilotKit>. A
non-explicit (`hasExplicitThreadId={false}`) threadId is now overridable; a
genuine caller-supplied threadId stays controlled.

Also, surfaced during the langgraph-js de-fork validation:
- CopilotChat clears messages when switching to a fresh non-explicit thread
  ("+ New") so the welcome screen shows instead of the prior thread — including
  when a still-loading /connect is superseded by the switch.
- CopilotDrawer treats a pending (null) license status as loading, not
  unlicensed, so the upgrade upsell no longer flashes or sticks before the
  runtime reports the license.
2026-06-26 16:54:55 -05:00
Benjamin Taylor 2664df54a9 feat(web-components): CopilotDrawer host-theme inheritance, action icons, self-owned mobile launcher, and UX fixes
Polish on top of the shipped <copilotkit-drawer> (#5701), surfaced by the first
real de-fork pilot (langgraph-js):

- Theming follows the host app's light/dark theme: each token resolves
  override -> host theme var (--background/--card/...) -> built-in light default,
  instead of baking light-only values (drawer was light on a dark app).
- Row actions are icon buttons (inline lucide-style SVG, currentColor) with
  instant tooltips. Native title is replaced by a CSS tooltip (matches the
  react components' delayDuration:0; title's show-delay is browser-fixed) and is
  positioned to the side so the list's scroll-overflow never clips it.
- Archived rows are muted, no longer struck through.
- A refetch (e.g. Active<->All filter toggle) keeps the known list visible;
  the full loading state shows only on the initial empty fetch.
- The reserved footer region hides when nothing is slotted into it (it rendered
  as an empty box at the bottom); mirrors the memories region.
- Mobile gets a self-owned floating launcher so the drawer is always openable
  with no host header wiring. Its icon is swappable via a launcher-icon slot and
  its position via --cpk-drawer-launcher-top/left (host can center it on its own
  header controls); exposes part="launcher".

Adds element tests for icons+tooltips, archived styling, refetch-preserves-list,
footer hide-when-empty, and the mobile launcher (render + open + desktop-absent).
2026-06-26 16:54:55 -05:00
Jordan Ritter 3d1e2eafa5 fix(showcase): isolation CLI correctness — port-guard propagation + --isolate arg-parsing (#5731)
## Summary
Two pre-existing isolation-CLI bugs that #5730 (slot-reaping) surfaced
and correctly deferred as out-of-subject — both real, now fixed with
real-surface red-green.

1. **Port-conflict guard silent-failure** — `_slot_ports_free` consumed
`_slot_offset_ports` via process substitution (`done < <(...)`), so a
`die` on an out-of-range/non-numeric slot exited only the subshell; the
loop read zero ports and the function returned 0 ("all free"), silently
defeating the port-conflict guard for a bad slot. Now captures into a
variable with `|| die` so the failure propagates to the caller.
2. **`--isolate` arg-parsing** — `--isolate <name> <slug>` (name before
slug) mis-parsed the name as the slug then died "Unexpected argument"; a
bare `--isolate=` silently auto-picked. Now order-independent (deferred
`pending_iso_name` resolution after the parse loop), rejects an empty
value loudly, and `--isolate=<name>` (non-numeric) binds an explicit
name. `--isolate=<N>` numeric pin unchanged. **Intended behavior
change:** `--isolate=<name>` previously died in the picker; it now binds
a name.

## Tests
`bats showcase/scripts/__tests__/` = 170/0 — real slot dirs, real
bad-slot input, real parser end-to-end (no mocks). CR converged round 1
(zero bucket-a), Procedure 3 promotion audit clean.

## Deferred to a potential follow-up (pre-existing "port-enumeration
completeness")
CR surfaced, and correctly kept out of scope: `_slot_offset_ports`
hardcoded/incomplete infra-port list (phantom 8081, misses other
published ports); apply_isolation port-offset regex matches quoted
short-form only; `.last-test-ts` not repointed per-isolate (concurrent
clobber); `--repeat` treats a following flag as its count.

🤖 Generated with [Claude Code](https://claude.com/claude-code)
2026-06-26 14:49:23 -07:00
github-actions[bot] 0e05d94f18 style: auto-fix formatting 2026-06-26 14:40:11 -07:00
Alem Tuzlak e0307facc8 refactor(bot): drop bot-store-redis/postgres adapters (never published)
The StateStore interface and the in-memory MemoryStore default remain;
durable backends can be reintroduced as a follow-up. Both adapter packages
were merged in #5613 but never published to npm, so removal is a clean
delete with no consumer impact.

- Delete packages/bot-store-redis and packages/bot-store-postgres.
- Revert the bot release scope and drift guard to bot + bot-ui.
- Strip the Redis dep, demo:restart script, restart demo, docker-compose,
  and REDIS_URL env from examples/slack.
- Rewrite the bot persistence/transcripts docs around "MemoryStore default
  + implement the StateStore interface yourself for durability".
2026-06-26 14:40:10 -07:00
Sam Julien a40731ffd3 docs(shell-docs): ungate Slack and Teams docs (#5735)
## Summary
- remove Slack and Teams early-access gates from frontend and bot docs
- remove the bot reference early-access wrapper
- add inline CopilotKit Intelligence waitlist CTAs for managed Slack and
Teams agents

## Verification
- `npm exec -- vitest run
src/components/__tests__/early-access-gate.test.tsx
src/lib/__tests__/docs-render.test.ts -t "Slack and Teams|early-access
frontmatter"`
- `npm run typecheck`
- `npm run build`
- `curl -I --max-time 10 https://www.copilotkit.ai/opentag-managed`
2026-06-26 14:38:08 -07:00