## Release monorepo v1.62.3
**Scope:** `monorepo` | **Bump:** `patch`
---
### How this release process works
1. **This PR was created automatically** by the "release / create-pr"
workflow.
It bumped the `monorepo` packages to `1.62.3`
and generated AI-enhanced release notes.
2. **CI runs on this PR** — the full test suite (unit tests, lint, type
checks, build)
must pass before merging. This is the review gate.
3. **Review the release notes** in `release-notes.md` in this PR.
If a Notion draft was created, you can edit the release notes there
before merging.
4. **When this PR is merged**, the `release / publish` workflow
automatically:
- Builds all packages
- Publishes the `monorepo` packages to npm at version `1.62.3`
- Creates git tag `monorepo/v1.62.3`
- Creates a GitHub Release with the final release notes
### Before merging
- [ ] CI is green (tests, lint, types, build)
- [ ] Version bumps look correct
- [ ] Release notes are accurate (edit in Notion if a draft was created)
---
> **Do not merge until CI is fully green.** The full test suite runs
automatically on this PR.
Documents the new **CopilotDrawer** prebuilt threads drawer (component
PR #5746).
## Changes
- **Guide** — `docs/prebuilt-components/copilot-drawer.mdx` (new) +
registered in `prebuilt-components/meta.json`. React drop-in
(`<CopilotDrawer />` beside `<CopilotChat />`, zero active-thread
wiring), props table, customization (slots / `::part`s /
`--cpk-drawer-*` tokens), Angular usage, license.
- **Threads how-to** — `snippets/shared/threads/threads.mdx` (shared →
propagates to ~10 integration threads pages + `/threads`): a callout
recommending `CopilotDrawer` as the prebuilt path and framing the
`useThreads` steps as the **headless** alternative (kept, not removed) +
Next-steps cross-links.
- **Reference** — `reference/components/CopilotDrawer.mdx` (React) and
`reference/angular/components/CopilotDrawer.mdx` (Angular),
hand-authored with `<PropertyReference>`.
## Validation
`npm run build` in `showcase/shell-docs` is green locally — compiled
successfully, 215/215 static pages generated, no MDX errors. All
internal links verified to resolve on `main`.
## Why draft / release-gated
Merges **with or after** the drawer release
(`@copilotkit/web-components` published +
`react-core`/`@copilotkit/angular` > 1.61.2 containing `CopilotDrawer`).
Until then the documented imports don't exist in a published package.
## Follow-ups (out of scope)
- Live `<InlineDemo>` of the drawer (needs a demo-registry entry) — uses
plain code blocks for now.
- Angular `injectThreads` / `CopilotChatConfiguration` reference pages
(named without links here; they don't exist on `main` yet).
Spec: https://app.notion.com/p/38f3aa381852814e8884e3153d3d4688 ·
Tracking: ENT-1021
🤖 Generated with [Claude Code](https://claude.com/claude-code)
> **Update (SSE rework — addresses review):** `/suggest` now **streams
AG-UI SSE** instead of buffering a JSON `{ messages }` response, and the
run **forwards the consumer's messages + state**. Server-side it reuses
`createSseEventResponse` (the runner's event pipeline minus
`GLOBAL_STORE` persistence, `captureTelemetry:false`); client-side it
drives a stock `HttpAgent` at `/agent/:id/suggest`, so chips fill in
progressively via `onMessagesChanged` and the run never touches the
Intelligence websocket delegate. The stateless and fallback paths now
share one `runAgent` flow (net −68 LOC of production code). The `## How`
/ behavioral notes below are updated to match.
## What & why
Dynamic suggestions were implemented as a **full shadow-thread agent
run** (`SuggestionEngine.generateSuggestions` clones the provider agent
onto a throwaway `threadId` and calls `runAgent`). Against a
thread-persisting backend (CopilotKit Intelligence) every reload
materialized a **listed, auto-named thread** (plus gateway events + a
run lock), flooding the thread drawer. Root cause is the design: a
suggestion is a **stateless structured completion**, not a persisted
conversation.
This makes it one. A suggestion now runs the provider agent **directly**
behind a dedicated `POST /agent/:agentId/suggest` handler — no thread,
lock, gateway, runner-store, or name-gen — for **all**
CopilotKit-runtime backends. The debris is structurally impossible
rather than hidden.
## How
- **`get-runtime-info`** advertises a `suggestions` capability on `GET
/info` (`RuntimeInfo.suggestions`).
- **`handleSuggestAgent`** (new) runs the resolved provider agent
directly and **streams its AG-UI events as SSE** via
`createSseEventResponse` — the runner's event pipeline
(`agent.runAgent({ onEvent })` + `finalizeRunEvents`) **minus** the
`GLOBAL_STORE` persistence that backs SSE-mode thread listings.
`captureTelemetry:false` and no debug bus keep suggestions out of run
telemetry/the inspector. Header-forwarding only; no middleware.
- **Router**: `POST /agent/:agentId/suggest` wired through
`fetch-router`/`fetch-handler` (multi- and single-route), with
`assertNever` exhaustiveness guards.
- **core `SuggestionEngine`**: when the runtime advertises `suggestions`
(and transport isn't single-route), it seeds a stock `HttpAgent`
(pointed at `/agent/:id/suggest`, credentials via a fetch wrapper) with
the consumer's deep-cloned messages + state and `runAgent`s it — a plain
`HttpAgent` only speaks REST SSE, so it never routes through the
Intelligence websocket delegate. Otherwise it falls back to the
clone+`runAgent` path. Both paths share one flow; abort-aware; failures
logged.
## Testing
Unit + integration across both packages (**504 core + 1587 runtime
green**, `check-types` clean): the **no-thread-leak** proof (real
`InMemoryAgentRunner`, asserts the suggest thread never appears in the
live listing), progressive-streaming + state-forwarding assertions,
cross-mode SSE handler behavior, error/abort robustness (Safari
`DOMException` + undici non-`AbortError`), header/credential forwarding,
multi-route `200` + `GET→405`, and single-route/`suggestions:false`
fallback.
Validated end-to-end against a running runtime + real `CopilotKitCore`
client over HTTP (local built packages):
- **In-memory managed threads:** a normal `/run` creates exactly 1
listed thread; **4 dynamic-suggestion reloads add 0**. Chips stream
progressively (`onSuggestionsChanged` 0→1→2); the `/suggest` request
carries the consumer's state.
- **Live hosted CopilotKit Intelligence** (`mode: "intelligence"`, real
gateway): `/info` advertises the capability, suggestions stream via the
stateless path, and the **hosted thread listing is byte-identical
(50→50) after 6 reloads** — the drawer's actual data source is
untouched.
## Compatibility
**Migrations:** none. No schema/DB change; self-hosted deployments need
no coordinated update — the change is entirely in `@copilotkit/core` +
`@copilotkit/runtime`. No release flag.
**Additive + capability-gated — version skew is safe both directions:**
- **New core + old runtime** (no `/suggest`, capability absent):
`core.suggestions` is `undefined` → client uses today's clone+`runAgent`
path.
- **Old core + new runtime**: the new optional `RuntimeInfo.suggestions`
field is ignored; the `/suggest` endpoint sits unused.
- **New + new**: stateless `/suggest`.
`RuntimeInfo.suggestions` is a new **optional** field and the
`/agent/:agentId/suggest` route (+ `RouteInfo`/`METHOD_NAMES` entries)
are purely additive. No exported symbol was removed or renamed.
**Behavioral notes (matched new versions):**
- Dynamic suggestions **stream progressively** over SSE (chips fill in
as the provider emits them), matching the pre-`/suggest` UX.
- Suggestion runs no longer flow through the run handler, so they create
**no thread** (the fix) and won't appear in `oss.runtime.*` run
telemetry.
- Single-route transport and non-CopilotKit AG-UI backends keep the
clone+`runAgent` fallback.
**⚠️ Deploy note:** a reverse proxy that path-allowlists CopilotKit
routes must add `POST /agent/:agentId/suggest`. If the capability is
advertised but that path is blocked, suggestions silently no-op — the
client does **not** fall back on a `/suggest` error (by design, to avoid
re-introducing the thread flood).
## Review
Landed through a full CR loop (3 rounds / 21 agent-reviews + a
bucket-(c) promotion audit) — bucket (a) converged to zero. Known
pre-existing follow-ups tracked separately (out of scope):
`resolveLicenseStatus` grace-period ordering and
`AgentRegistry.getAgent` using `in` vs `hasOwnProperty` for reserved
ids.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
The prior preview paired the real <copilotkit-threads-drawer> with a hand-coded
mock chat panel (from the screenshot harness), so the chat half wasn't an actual
CopilotKit component — which is what the reviewer flagged. Re-captured with the
real default <CopilotThreadsDrawer> + <CopilotChat> together, stock (untheme d)
light styling, showing a thread's replayed conversation.
Syncs the branch with main (304 commits) to resolve CI type-check failure.
main changed extractForwardableHeaders to require a forwarding policy and
added the mergeForwardableHeaders helper (#5712); handle-suggest now uses
mergeForwardableHeaders(agent.headers, request, runtime.forwardHeadersPolicy ??
resolveForwardHeadersPolicy(undefined)) to match the run handler — fixing the
drift and adopting the server-headers-win / infra-header denylist behavior.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Rework the stateless /suggest transport to reuse the AG-UI SSE pipeline
instead of a buffered JSON response, resolving the streaming + state review
feedback:
- server runs the provider agent directly and streams its events via
createSseEventResponse (the runner's event pipeline minus GLOBAL_STORE
persistence), gated with captureTelemetry:false so suggestions stay out of
run telemetry
- client drives a stock HttpAgent against /agent/:id/suggest, so chips stream
progressively via onMessagesChanged and the run never routes through the
Intelligence websocket delegate (still no thread persistence)
- forward the consumer's deep-cloned messages + state onto the suggestion run
(was state: {}), matching the clone fallback
Net -68 LOC of production code; the stateless and fallback paths now share one
runAgent flow.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- /threads page title -> 'Headless Threads' to distinguish the headless
useThreads path from the prebuilt drawer (slug kept, so no inbound links
break; sidebar label follows the frontmatter title). (samjulien #9)
- prebuilt-components index: the 'saved conversations' line now leads with the
drop-in CopilotThreadsDrawer and offers Headless Threads as the DIY path, plus
a companion-sidebar mention in 'Pick a surface' — the drawer was absent from
the prebuilt landing page.
- chat page: same, point at the prebuilt drawer first, headless second.
- Remove 'rename' from the prebuilt CopilotThreadsDrawer capability claims
(guide, React reference, shared Threads callout) — the row kebab only does
archive/unarchive + delete. Add an explicit note that rename is available via
the headless useThreads path. (MikeRyanDev)
- Reference CSS parts list now matches the shipped element: adds row/row-active,
collapse-toggle, close-toggle, backdrop, launcher-cluster, launcher-new-thread,
load-more, fetching-more, fetch-more-error, fetch-more-retry, licensed,
licensed-cta; grouped by area. (MikeRyanDev)
- Rewrite the 'no threadId state / no onSelect plumbing' line to stand on its own
by contrasting with a hand-rolled sidebar. (samjulien)
Stand-in for a live showcase example (out of scope for this PR): a real
screenshot of <CopilotThreadsDrawer> beside <CopilotChat>, rendered from the
v2 react demo against the Intelligence platform. Embedded as a <Frame> preview
right under the intro.
Match the final component name. Renames the guide and reference pages
(copilot-drawer.mdx -> copilot-threads-drawer.mdx, CopilotDrawer.mdx ->
CopilotThreadsDrawer.mdx), their slugs/URLs, the nav meta entry, the
data-testid default, the <copilotkit-threads-drawer> element mention, and
all prose/import references. The --cpk-drawer-* CSS tokens and ::part names
are unchanged.
Per review: remove the onUnlicensed prop, the unlicensed slot, and the
unlicensed/unlicensed-cta parts from the guide + reference so neither
humans nor agents surface them. Replace with a single neutral line:
threads require Intelligence; a locked view shows without a license key.
- Guide: lead with the user benefit (less reference-y opening), reframe the
headless useThreads alternative to stand alone, add the OpsPlatformCTA
sign-up callout (per review).
- Add threads.mdx (shared-snippet include) for a2a, adk, agent-spec,
deepagents + register each under Intelligence Platform in meta.json.
Per review: defer the Angular drawer docs. Removes the Angular reference
page, the guide's Angular section, and Angular cross-links; keeps the
React guide + reference + the Threads how-to callout.
Redesigns the shared `<copilotkit-threads-drawer>` element
(`@copilotkit/web-components`) to the new Figma UX, keeping the React,
Vue, and Angular wrappers in lockstep. Pure-VIEW change — no
`@copilotkit/core`, runtime, or `useThreads` changes.
**Ticket:** [ENT-1051](https://linear.app/copilotkit/issue/ENT-1051) ·
**Figma:** [Thread
Drawer](https://www.figma.com/design/feSsBJw1qCfLp0JNnOurrJ/CopilotKit-Intelligence?node-id=723-78)
· **Spec:**
[Notion](https://app.notion.com/p/3953aa381852819ab464dae3894e7f18)
## What changed
- **Header** → right-aligned icon row. On desktop it holds the
**collapse** toggle (sidebar glyph); on mobile the **close** toggle. No
title text, no "+ New" pill. Optional `slot="header"` preserved (empty
by default; the toggle right-aligns after it).
- **New Conversation** row (plus-square + label) below the header —
keeps `part="new-thread-button"` + the `new-thread` event.
- **Recent Conversations** heading + **funnel** filter icon → Active/All
popover. Preserves `_filter` + `filter-change` and
`part="filter-active"`/`filter-all"`.
- **Per-row kebab menu** (⋮) holding Archive/Unarchive + Delete —
preserves those events + parts. An open kebab now shields the rest of
the list from hover so it reads as a single surface (see Review fixes).
- **Delete confirm** is a native `<dialog>` opened with `showModal()`
(browser top layer), centered over the drawer's visible box — it can't
paint under other UI or drop below the fold. jsdom falls back to the
`open` attribute.
- **Archived rows** render italic/muted inline in the "All" view.
- **Desktop collapse** → `collapsed` / `collapsible` props (default
**expanded**) + a `collapse-change` event / `CollapseChangeDetail`.
Collapsing sets `--cpk-drawer-reserved-width: 0` on the document root so
the host grid reclaims the column with no hydration flicker.
- **Unified floating cluster** (Figma "closed" mockup) =
`[sidebar-toggle] [+ New Conversation]`, shown in both the mobile-closed
and desktop-collapsed states. Mobile stays an off-canvas modal (backdrop
/ Escape / focus-trap).
## Compatibility
- All existing `::part()` names and events are preserved; only
**additive** parts are introduced: `collapse-toggle`, `close-toggle`,
`section-heading`, `filter-toggle`, `row-menu`, `row-menu-popover`,
`launcher-cluster`, `launcher`, `launcher-new-thread`, plus
`confirm-dialog`/`confirm-cancel`/`confirm-delete`/`backdrop`. One
additive event: `collapse-change`.
- Additive wrapper props: `recentLabel` (all frameworks);
`collapsible`/`collapsed` + `onCollapseChange` (React) / equivalents in
Vue & Angular.
- **Usage note (now in the docs):** the drawer and `<CopilotChat>` must
share a chat-configuration provider so the drawer drives the chat —
`CopilotChatConfigurationProvider` (React/Vue) /
`provideCopilotChatConfiguration()` (Angular). The v2
`CopilotKitProvider` does not provide that context on its own.
- Verified: **no example `::part()` theme changes required** — every
example themes the drawer via inherited `--cpk-drawer-*` custom
properties.
## Descoped / changed during development (re: earlier review)
- **Client-side search was removed at the designer's request.** There is
**no** search UI, `search` event, `search-toggle`/`search-input` part,
or `onSearch` wrapper prop in the shipped element. Any remaining
"search" mention in older comments is stale.
- **Desktop collapse was briefly backed out, then re-restored** per the
designer (commit `8bfd245305`). The shipped element **has** collapse
(`collapsed` is a live public property — it was not removed).
## Review fixes (commit `0b4f6ca392`)
Addressing @MikeRyanDev and @marthakelly:
- **`core/threads.ts`** — a full-list refetch (filter-change / retry)
now clears `fetchMoreError` on both `listRequested` and `listSucceeded`,
so the inline "couldn't load more — retry" banner no longer survives
onto a fresh list.
- **Escape while confirming delete** — the host keydown handler now
consumes Escape while a confirmation is open; previously the bubbled
keydown fell through and closed the whole mobile drawer along with the
confirmation.
- **Open kebab menu shields the list** — `.list.menu-open
.row:not(.menu-open)` gets `pointer-events: none`, so other rows no
longer reveal their kebab / paint a host `::part(row):hover` background
around or behind the open popover. Click-away dismissal is preserved via
the existing document pointerdown handler. (Verified live in the
langgraph-js example.)
- **Docs token** — dropped the removed `--cpk-drawer-rail-width` from
the web-components README.
## Testing
All suites run via `nx`, green through each package's lefthook
pre-commit gate:
- `@copilotkit/web-components` — **92** drawer element tests + `:build`
green. Covers header collapse/close toggles, New Conversation, funnel
filter switch, `recentLabel`, kebab open + archive/delete routing,
confirm-dialog gating + native cancel + **backdrop-click dismiss**,
**Escape-while-confirming (no drawer close)**, **open-menu row shield**,
collapse/cluster/column-reclaim, archived-italic, preserved
parts/events, `header` slot + `label` aria-labels.
- `@copilotkit/core` — **553** tests incl. the new `clears a lingering
fetchMoreError when a full list refetch succeeds`.
- `@copilotkit/react-core` — CopilotThreadsDrawer suite + full package
**1419** green.
- `@copilotkit/vue` — **32** incl. SSR + the `collapsible`
boolean-prop-default regression test.
- `@copilotkit/angular` — CopilotThreadsDrawer spec **35** (incl.
**scoped-chat-input focus**: prefers the ancestor `copilot-chat-view`
over the document-global fallback); full package **178**.
## Follow-on
- Docs (screenshot + reference/guide) on the release-gated docs PR
**#5780**.
- Release (`web-components` + `react-core` + `vue` + `angular`,
lockstep) + CLI scaffolding bump.
- Example grid/theme updates ride the release in **#5828**.
- **[ENT-1080](https://linear.app/copilotkit/issue/ENT-1080)** — dedup
the per-wrapper `findChatInput` / open-state fallback (marthakelly #6,
deliberately deferred as a cross-package refactor).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
## Problem
The **Shell script tests (bats + shellcheck)** job in
`showcase_validate.yml` installs `bats` with:
```yaml
sudo apt-get update
sudo apt-get install -y bats
```
GitHub's `ubuntu-latest` runner image preconfigures third-party apt
repos (Microsoft / `azure-cli`, pointing at `packages.microsoft.com`)
for preinstalled tooling. `apt-get update` refreshes **every**
configured repo, not just the ones a job needs. When one of those
Microsoft repos serves invalid release metadata:
```
E: Failed to fetch https://packages.microsoft.com/.../InRelease Clearsigned file isn't valid, got 'NOSPLIT'
##[error]Process completed with exit code 100.
```
`apt-get update` exits non-zero and — because the step runs under `bash
-e` — the whole step aborts before `bats` installs. This fails the job
even though `bats` comes from Ubuntu's own `universe` repo, which is
unaffected. It's a runner-image / external-repo outage, not anything in
the shell tests.
## Fix
This job only needs Ubuntu packages, so remove the unused third-party
repos before updating:
```yaml
sudo rm -f /etc/apt/sources.list.d/*microsoft* /etc/apt/sources.list.d/*azure-cli*
sudo apt-get update
sudo apt-get install -y bats
```
Ubuntu's archive (where `bats` lives) is in the base `sources.list` and
is untouched, so `bats` still installs; `apt-get update` no longer
refreshes the broken, unused Microsoft repos, so it stops failing the
job.
## Testing
- Change is confined to the `Install bats` step of the
`shell-script-tests` job. The rest of the job (shellcheck + the bats
suite) is unchanged.
- CI on this PR exercises the modified step directly — a green
`shell-script-tests` run here confirms the install path.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
The shell-script-tests job installs bats via apt. GitHub's ubuntu-latest runner
image preconfigures third-party apt repos (Microsoft / azure-cli) for
preinstalled tooling this job never uses. When one of those repos serves invalid
release metadata, `apt-get update` exits non-zero and `bash -e` aborts the step
before bats installs — even though bats comes from Ubuntu's own `universe` repo,
which is unaffected.
This job only needs Ubuntu packages, so remove those unused third-party repos
before `apt-get update`.
Addresses PR #5823 review (MikeRyanDev + marthakelly):
- core/threads.ts: a full-list refetch (filter-change / retry) now clears
fetchMoreError on both listRequested and listSucceeded, so the inline
'couldn't load more - retry' banner no longer survives onto a fresh list.
- web-components: Escape while the confirm-delete <dialog> is open is now
consumed by a confirm guard in the host keydown handler; previously the
bubbled keydown fell through to the mobile branch and closed the whole drawer
along with the confirmation.
- web-components: an open kebab popover now shields the rest of the list -
.list.menu-open .row:not(.menu-open) gets pointer-events:none so other rows
no longer reveal their kebab or paint a host ::part(row):hover background
around/behind the menu. Click-away dismissal is preserved via the existing
document pointerdown handler.
- README: drop the removed --cpk-drawer-rail-width from the documented tokens.
Tests: core clears-fetchMoreError-on-refetch; web-components Escape-while-
confirming, confirm-dialog backdrop-click, menu-open row shield; Angular
scoped-chat-input focus (ancestor copilot-chat-view over the global fallback).
Menu-shield verified live in the langgraph-js example (:3002).
## Summary
Fixes#5533
When a runtime registers an agent under a **non-default** name (e.g.
`agents: { TravelBookingAgent }`) and the frontend renders `<CopilotChat
agentId="TravelBookingAgent" />` without an `agent` prop on
`<CopilotKit>`, the app throws after runtime sync:
> useAgent: Agent 'default' not found after runtime sync (runtimeUrl=…).
Known agents: [TravelBookingAgent]
## Root cause
`useAgent()` (`packages/react-core/src/v2/hooks/use-agent.tsx`) resolved
its `agentId` only from its own prop, falling back straight to
`DEFAULT_AGENT_ID`. It never consulted the surrounding
`CopilotChatConfigurationProvider` — even though `CopilotChat` installs
that provider around its subtree with the resolved (non-default)
agentId.
`CopilotChat` resolves its *own* `useAgent` call correctly, so the chat
works. But any **descendant** that calls `useAgent()` without re-passing
`agentId` (a custom message/tool-render component, a sibling hook)
silently resolves to `'default'`. Once `/info` sync lands and the
registry holds only the non-default agent, that consumer throws — which
is why the thrown id is `'default'`, not `'TravelBookingAgent'`.
## Fix
Resolve `agentId` in `useAgent` with the same precedence `CopilotChat`
already uses:
```ts
const resolvedAgentId = agentId ?? chatConfig?.agentId ?? DEFAULT_AGENT_ID;
```
The hook already imported and called `useCopilotChatConfiguration` (for
`threadId`); the call is hoisted and reused — no duplicate hook call. An
explicit `agentId` prop still wins; with no chat config it still falls
back to `DEFAULT_AGENT_ID`. No changes to core or the providers.
## Tests added
`packages/react-core/src/v2/hooks/__tests__/use-agent-nondefault-agentid.test.tsx`
— a `useAgent()` consumer inside a chat configured for
`TravelBookingAgent` (runtime synced to `agents:{TravelBookingAgent}`)
must not throw `Agent 'default' not found`, and must inherit
`TravelBookingAgent`. Fails before the fix, passes after.
## Checklist
- [x] Failing test written and confirmed failing before the fix
- [x] Fix applied, test passes
- [x] Full `@copilotkit/react-core` suite passes (1291 passed)
- [x] Build succeeds (`nx build @copilotkit/react-core`)
- [x] Formatter passes (`pnpm format`)
The collapse toggle keeps the header bar visible on desktop, so the row's 12px
top margin doubled up with the header's padding (extra gap vs mobile). Make the
top margin conditional on the header being HIDDEN (collapsible=false + no header
slot) via .header[hidden] + .new-conversation; otherwise the header supplies the
top spacing on both breakpoints.
- Floating cluster/launcher default gutter → 24px on both top and left (was 12px).
- Selected row: drop the border-color; the background change alone marks it
(base row keeps its 1px transparent border for layout stability).
- Cap --_radius at 4px via min(theme, 4px) and lower the hardcoded 6px button
radii to 4px, so no bordered element exceeds a 4px corner radius.
Intersect the .root rect with the viewport before centering the confirm dialog.
A host grid that doesn't bound the drawer's row lets .root grow to content
height, so centering over the raw rect dropped the modal far down the page (seen
in the langgraph-js example, whose grid has no row bound). Clamping to the
on-screen band keeps it centered in the visible drawer regardless of host sizing.
## What
`examples/showcases/banking/run-demo.sh` launches a native Metal
`text-embeddings-router` (TEI) on `:7067` for the self-hosted
durable-memory path. This passes `--max-batch-tokens 512` so TEI's
warmup uses a small forward pass.
## Why
On some Apple Silicon machines, TEI's default `--max-batch-tokens`
(16384) **faults the Metal backend during its warmup forward pass**.
Observed two failure modes at the exact same step (`Warming up model`):
- **Deadlock** — every thread, including the main thread, parked in
`__psynch_cvwait` at 0% CPU. Never binds `:7067`.
- **Silent death** — process exits mid-warmup with no panic / no OOM
line (the signature of a GPU-level abort).
Either way `:7067` never comes up, the script's `wait_http … 300` times
out, and the demo appears to "crash" with only:
```
ERROR: native Metal TEI did not come up at http://localhost:7067/health within 300s
```
The 300s timeout looks like a slow model download (the weights are ~1.1
GB), but that's a red herring — with weights cached the process still
hangs at warmup. Two different `--dtype` values (fp16, float32) both
failed identically, ruling out dtype; the variable is the warmup batch
size.
## Fix
`--max-batch-tokens 512` shrinks the warmup forward pass, which clears
reliably (`Ready` in ~3s, health `200`, verified 1024-dim `/embed`). It
bounds only per-request tokens — memory texts are short — **not** the
embedding vectors, so recall stays byte-identical to the docker/CI
embedder (the runbook's byte-identical guarantee holds).
## Scope
One-line flag change + explanatory comment. Only affects the Apple
Silicon native-TEI branch of the self-hosted demo path; amd64/CI (docker
`tei`) is untouched.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
The self-hosted `run-demo.sh` path launches a native Metal
`text-embeddings-router` on :7067 for the durable-memory demo. TEI's
default `--max-batch-tokens` (16384) can fault the Metal backend during
its warmup forward pass on some Apple Silicon machines. The process then
either deadlocks (every thread parked in a pthread cond wait at 0% CPU)
or dies silently with no panic — a GPU-level abort — so it never binds
:7067 and the 300s health wait times out. The demo appears to "crash"
with no actionable error.
Pass `--max-batch-tokens 512` so warmup uses a small forward pass, which
clears reliably. This only bounds per-request tokens (memory texts are
short), not the embedding vectors, so recall stays byte-identical to the
docker/CI embedder.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Move the collapse toggle after the header slot so it right-aligns (the slot has
flex:1 and pushes it over), matching the mobile close button — it was reading as
left-aligned above New Conversation, which the designer flagged.
Design iteration (Ben's designer):
- RESTORE desktop collapse. Re-add collapsed/collapsible + collapse-change
(element + all three wrappers, lockstep), the desktop header collapse toggle,
and the CollapseChangeDetail type/re-export. Default is EXPANDED.
- UNIFY the closed affordance into one floating cluster (Figma 'closed' mockup):
a sidebar-glyph toggle + a New Conversation (+) icon button, shown in BOTH the
mobile-closed state (adds New Conversation to the old single launcher) and the
desktop-collapsed state. Parts: launcher-cluster, launcher, launcher-new-thread.
- COLUMN RECLAIM (no empty reserved gap): on desktop-collapse the element sets
--cpk-drawer-reserved-width: 0px on the document root (reaches the grid past
the wrapper host via :root inheritance); hosts read it in grid-template-columns.
Default expanded never sets it, so no hydration flicker.
- DELETE MODAL centered over the DRAWER PANEL, not the viewport: keeps the
top-layer showModal() robustness (never clipped) but drives --confirm-cx/cy
from the visible .root rect and caps width to the drawer band.
- Vue fix: default collapsible to true in the wrapper. Vue coerces an omitted
Boolean prop to false, which was silently forcing collapsible=false (collapse
toggle vanished) — React/Angular pass undefined and keep the element default.
Validated live in the Nuxt (Vue) and Angular demos against managed Intelligence:
collapse/expand, cluster + New Conversation, column reclaim, drawer-centered
top-layer modal, mobile cluster. Tests: web-components 89, react-core 1419,
vue 32, angular 34 — all green.
## Summary
- Fixes langroid `a2ui-fixed-schema` D6 cell
(`d6:langroid/gen-ui-a2ui-fixed`): the a2ui container was emitting via
`TextMessageContentEvent` so the A2UI middleware never detected it and
the flight card never mounted — raw JSON rendered as plain text in chat
- Fixed by emitting `ToolCallResultEvent` (matching the
claude-sdk-python peer)
- Also fixed the operation shape: upgraded from legacy flat form
(`{"type": "create_surface", ...}`) to v0.9 nested form (`{"version":
"v0.9", "createSurface": {...}}`) — the renderer silently ignores the
flat form
- Adds the missing langroid aimock D6 fixture for `gen-ui-a2ui-fixed`
(without it, aimock returned 'no fixture match', causing a fetch error)
## Changes
**`showcase/integrations/langroid/src/agents/a2ui_fixed_agent.py`**
- Added `ToolCallResultEvent` import
- Replaced `TextMessageStart` + `TextMessageContent` + `TextMessageEnd`
block with a single `ToolCallResultEvent(tool_call_id=call_id,
content=json.dumps(operations))`
- Upgraded `_build_a2ui_operations()` to emit v0.9 nested op shape
**`showcase/aimock/d6/langroid/gen-ui-a2ui-fixed.json`** (new file)
- Adds aimock fixtures: turn 1 (no-tool-result → call `display_flight`)
and turn 2 (has-tool-result → confirmation text)
**Peer matched:** claude-sdk-python (event path), strands / google-adk
(v0.9 op shape)
## RED → GREEN Proof
### RED (before fix)
```
✗ d6:langroid/gen-ui-a2ui-fixed red (0.0s)
state=red
0 passed, 1 failed
```
Flap-diagnostics: `expected [data-testid="a2ui-fixed-card"] to mount
within 60000ms` — bodyTextSnippet showed raw `{"a2ui_operations":
[{"type": "create_surface"...` text rendered inline in chat.
### GREEN (after fix — rebuilt container + v0.9 ops +
ToolCallResultEvent + aimock fixture)
```
langroid [d6]
[conversation-runner] turn 1/1 — assertions passed
[conversation-runner] conversation completed successfully { turnsCompleted: 1, totalDurationMs: 2121 }
✓ d6:langroid green (2.6s)
1 passed (2.6s)
```
## Aimock Ceiling Check
`__tests__/aimock-fixtures.test.ts` passes as-is (825/825) — the new
langroid `gen-ui-a2ui-fixed.json` fixture uses `context: "langroid"`
scoping with unique `hasToolResult: false/true` discriminators, so no
new duplicate-ceiling collisions introduced.
## Note
This is a re-home of #5839. The original PR's head was on a
`worktree-agent-*` branch which never triggers CI workflows (0 runs,
even after close+reopen). Branch renamed to
`fix/langroid-a2ui-tool-call-result` to trigger normal CI.
Two bugs fixed:
1. Tool result event path: the `a2ui_operations` container was emitted
inside a `TextMessageContentEvent` block. The A2UI middleware only scans
`TOOL_CALL_RESULT` events for the container, so the card never mounted
and the raw JSON appeared as plain text in the chat. Fixed by emitting a
`ToolCallResultEvent` (matching the claude-sdk-python peer).
2. Operation shape: the ops used the legacy flat form
(`{"type": "create_surface", ...}`) which the renderer silently ignores.
Updated to the v0.9 nested form (`{"version": "v0.9", "createSurface":
{...}}`) used by every other working peer (claude-sdk-python, strands,
google-adk).
Also adds the missing langroid aimock D6 fixture for `gen-ui-a2ui-fixed`
(`display_flight` → tool result → confirmation text) so the D6 probe has
a mock response to drive the full surface-render assertion.
D6 cell: d6:langroid/gen-ui-a2ui-fixed red → green
## Summary
Spring-ai's `DisplayFlightTool` was emitting legacy flat A2UI operations
(`{"type":"create_surface",...}`) but the A2UI middleware expects v0.9
nested operations (`{"version":"v0.9","createSurface":{...}}`). This is
the same flat→nested migration done for Python/TS in #5832 and langroid
in #5839. The flat shape was silently ignored by the middleware, so the
flight card never mounted and the `a2ui-fixed-schema` D6 cell was
permanently red.
**Root cause** (confirmed by prior local red-green on the disproven TS
fix):
`DisplayFlightTool.apply()` emitted `a2ui_operations` in the legacy flat
format. The middleware's `tryParseA2UIOperations` parses the container
correctly, but the op dispatchers inside require the v0.9 shape. No
surface
was ever created → card never mounted.
## Fix
Updated the three operations in `DisplayFlightTool.java` to v0.9 nested
format:
| Before (flat, ignored) | After (v0.9 nested, works) |
|---|---|
| `{"type":"create_surface","surfaceId":...,"catalogId":...}` |
`{"version":"v0.9","createSurface":{"surfaceId":...,"catalogId":...}}` |
| `{"type":"update_components","surfaceId":...,"components":...}` |
`{"version":"v0.9","updateComponents":{"surfaceId":...,"components":...}}`
|
| `{"type":"update_data_model","surfaceId":...,"data":{...}}` |
`{"version":"v0.9","updateDataModel":{"surfaceId":...,"path":"/","value":{...}}}`
|
Shape matches `sdk-python/copilotkit/a2ui.py` and the shared Python
tools.
## Local Red-Green Proof (real control-plane probe, `--rebuild` both
runs)
**RED — original flat ops (`{"type":"create_surface",...}`):**
```
$ showcase test spring-ai:a2ui-fixed-schema --d6 --rebuild --keep
✗ d6:spring-ai/gen-ui-a2ui-fixed red (0.0s)
state=red
⚠ Tests failed for spring-ai:a2ui-fixed-schema (exit 1)
```
**GREEN — v0.9 nested ops
(`{"version":"v0.9","createSurface":{...}}`):**
```
$ showcase test spring-ai:a2ui-fixed-schema --d6 --rebuild --keep
✓ d6:spring-ai/gen-ui-a2ui-fixed green (0.0s)
1 passed
✓ Tests passed for spring-ai:a2ui-fixed-schema
```
## Java Build & Tests
All 64 existing spring-ai JUnit tests pass after the change (`mvn test`:
64 run, 0 failures, 0 errors). Code compiles cleanly (`mvn compile -q`).
## Files Changed
-
`showcase/integrations/spring-ai/src/main/java/com/copilotkit/showcase/springai/tools/DisplayFlightTool.java`
— v0.9 nested op format
The `route.ts` file is unchanged from main (the prior no-op `a2ui: {
injectA2UITool: true }` addition has been reverted — it was disproven as
a fix by a real local red-green).
## Summary
- PR #5426 added `showcase/aimock/d6/ag2/multimodal.json` with two
fixtures keyed on `userMessage + turnIndex:0 + context:ag2` (the image
and PDF multimodal probes).
- Those exact same match keys were already present in
`showcase/aimock/d6/ag2/agentic-chat.json` (placed there at the same
time #5426 updated the agentic-chat file).
- Result: 2 exact duplicates within the `ag2` context scope → collision
count 297→299, tripping the ceiling (297).
## Fix
Dedupe: remove the two multimodal-probe entries from
`agentic-chat.json`. The dedicated `multimodal.json` (added by #5426) is
the authoritative home. The `agentic-chat` probe never sends image/PDF
turns; aimock routes those to `multimodal.json` via `context: ag2`.
## RED → GREEN
**RED** (origin/main at `81c577f067`): CI run #28840506865 (`Showcase:
Validate main`):
```
AssertionError: Exact duplicate count (299) exceeds ceiling (297).
```
**GREEN** (this branch, `fe96b3f254`): local run against worktree
fixtures:
```
✓ fixture collision detection > no exact duplicate match keys within the same context scope 2ms
Tests 824 passed (824)
```
## Test plan
- [x] `__tests__/aimock-fixtures.test.ts > fixture collision detection >
no exact duplicate match keys` passes (count ≤ 297)
- [x] No new fixtures added or ceiling bumped — pure dedupe
- [ ] CI `Showcase: Validate main` flips green on this branch
🤖 Generated with [Claude Code](https://claude.com/claude-code)
PR #5426 (ag2 multimodal unquarantine) added showcase/aimock/d6/ag2/multimodal.json
with two fixtures keyed on:
- userMessage: "can you tell me what is in this demo image I just attached", turnIndex: 0, context: ag2
- userMessage: "can you tell me what is in this demo pdf I just attached", turnIndex: 0, context: ag2
Those same keys already existed in showcase/aimock/d6/ag2/agentic-chat.json,
creating 2 exact duplicates within the ag2 context scope and pushing the
collision count from 297 → 299 (ceiling = 297), breaking the validate CI job.
Fix: remove the two multimodal-probe entries from agentic-chat.json since
the dedicated multimodal.json is now the authoritative home. The agentic-chat
probe does not send image/PDF turns; the multimodal probe matches via context
"ag2" against multimodal.json directly.
RED: AssertionError: Exact duplicate count (299) exceeds ceiling (297)
→ confirmed in CI run #28840506865 (Showcase: Validate main)
GREEN: all 824 tests pass after removing the duplicate entries
## Summary
- **Restores ag2's `multimodal` D6 pill** from `skipped-incapable` (NSF)
to a working feature by adding a showcase-local ASGI middleware that
normalises AG-UI image/document/binary content parts to OpenAI Chat
Completions `image_url` parts before they hit AG2's `ConversableAgent`.
- **Surgical scope**: middleware mounted only on the multimodal sub-app
— other ag2 routes never see image parts and pay no body-buffer cost.
- **No upstream wait**: option (A) showcase shim, not an autogen PR.
autogen still lacks AG-UI image-part support; the moment they add it the
normalizer is a no-op and the RED-half regression pin flips to alert us.
- **Reverses commit d8a0a25db** for the multimodal half: removes
`multimodal` from `not_supported_features`, adds it back to `features`,
and restores the D6 aimock fixture pair.
`tool-rendering-reasoning-chain` stays quarantined (a different upstream
gap — no `REASONING_MESSAGE_*` events emitted by AGUIStream).
## What was failing
AG2's `autogen.code_utils.content_str` only accepts content-part types
`{"text", "input_text", "image_url", "input_image", "function",
"tool_call", "tool_calls"}`. The harness sends user messages whose
`content` carries:
- modern AG-UI: `{"type": "image" \| "document", "source": {"type":
"data" \| "url", "value": ..., "mime_type": ...}}`
- legacy mirror (appended by `legacy-converter-shim.tsx` for LangChain
integrations): `{"type": "binary", mimeType, data \| url}`
Both trip the gate with `ValueError("Wrong content format: unknown type
image within the content")` BEFORE the request reaches the vision model
— observed live on staging in the D6 multimodal probe. That's why the
feature was quarantined NSF in d8a0a25db.
## How the fix works
`agents/_multimodal_normalize.py` adds a raw-ASGI middleware (mirrors
the existing `RequestUserMessageMiddleware` pattern) that:
1. Buffers each inbound POST body.
2. Walks `messages[*].content` on user-role messages only.
3. Rewrites each AG-UI image/document/binary part to `{"type":
"image_url", "image_url": {"url": ...}}` — data sources become
`data:<mime>;base64,<value>` URLs; URL sources pass through unchanged.
4. Updates the request's `content-length` header.
5. Replays the rewritten body to the downstream AGUIStream endpoint.
Non-user messages, plain-text content, already-normalised parts, and
unknown shapes pass through untouched (identity-preserved on no-op
turns). Any body-parse failure logs at WARNING and replays the ORIGINAL
body so autogen's verbatim error surface stays intact — visibility, not
silent rewrite.
## RED → GREEN evidence
`tests/python/test_multimodal_normalize.py` — 14 unit tests, all pass:
| # | Test | What it pins |
|---|------|------|
| 1 | `test_autogen_rejects_raw_agui_image_part` | RED: `content_str`
raises the verbatim `ValueError` text the D6 probe surfaced |
| 2 | `test_normalized_content_is_accepted_by_autogen` | GREEN: after
normalize, `content_str` returns the rendered string with `<image>`
placeholder |
| 3-7 | shape coverage | image data/url, document data, binary data/url,
mimeType camelCase alias |
| 8-10 | passthrough | text-only, plain-string content, assistant/tool
messages |
| 11 | idempotency | re-running on already-normalised content is a no-op
|
| 12 | error path | unrecognised source → text placeholder (not a hard
fail) |
| 13 | tripwire | middleware class exposes `__init__(app)` + `__call__`
|
RED was independently verified by monkey-patching
`_normalize_content_part` to passthrough — that reproduces the exact
`ValueError("Wrong content format: unknown type image within the
content")` from the staging probe. Restoring the normalizer flips it
back to GREEN.
End-to-end ASGI smoke (run inline during development): a synthetic AGUI
POST body with a modern image part is sent through
`MultimodalContentNormalizerMiddleware` → inner ASGI app sees rewritten
body with correct `content-length`. PASS.
## Out of scope / follow-ups
- **PDF rendering**: PDFs ride through as
`data:application/pdf;base64,...` inside an `image_url` part — they
survive autogen's gate but the vision model can't read them natively.
Flattening PDFs to inline text (the pattern langgraph-python uses via
pypdf) is a separate enhancement; this PR's scope is unblocking the
image path that the D6 `multimodal` pill assertion checks.
- **Upstream**: autogen could fix this in `content_str` by accepting
AG-UI's `image`/`document`/`binary` content types directly. When/if that
lands, the normalizer becomes a no-op and the RED-half test will start
failing (which is the signal to delete the shim).
## Test plan
- [x] `cd showcase/integrations/ag2 && python -m pytest tests/python/` —
16 passed (2 pre-existing gen_ui guard tests + 14 new
multimodal_normalize tests)
- [x] `ruff format --check` on touched python files — clean
- [x] `ruff check` on touched python files — clean
- [x] `cd showcase/scripts && pnpm validate-manifests` — ag2 manifest
validates
- [x] `oxfmt --check showcase/aimock/d6/ag2/multimodal.json` — clean
- [x] Verified `multimodal_app.user_middleware` includes
`MultimodalContentNormalizerMiddleware` after import
- [x] End-to-end ASGI smoke: middleware rewrites body + updates
content-length, downstream app sees normalised payload
- [ ] Staging deploy: D6 `multimodal` pill flips from
`skipped-incapable` to GREEN with image fixture (1×1 PNG → "image
attachment shows a small abstract test pattern..."). Validated
post-merge via the staging deploy.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
- Remove unused imports: Iterable, ConversableAgent (from autogen),
AGStreamInput (from autogen.ag_ui.adapter) — none appear in executable
code, only in docstring prose.
- Fix raw_msgs possibly-unbound at dispatch guard: initialize to None
before the try block so the identity check at line 302 is always
safe even if model_dump raises before raw_msgs is assigned. Also
tighten the guard to `raw_msgs is not None` to make the no-normalization
fallback explicit.
autogen.ag_ui import unresolved and LLMConfig(dict) "Expected 0 positional
arguments" are ENVIRONMENT findings — autogen.ag_ui ships only in the
ag2[ag-ui] extra (present in the container, not in local Pyright's venv),
and LLMConfig({...}) is the codebase-wide pattern that works at runtime
with ag2>=0.9 as installed in the container.
AG2's ConversableAgent runs every user message through
``autogen.code_utils.content_str``, which only accepts content-part
types in {"text", "input_text", "image_url", "input_image", "function",
"tool_call", "tool_calls"}. CopilotChat / the AG-UI runtime emits image
and document attachments as the modern shape
{"type": "image" | "document", "source": {...}}
and the demo page's legacy-converter-shim.tsx ALSO appends a legacy
{"type": "binary", mimeType, data | url}
mirror alongside it (to keep the @ag-ui/langgraph converter happy on
LangChain-based integrations — it rides through on the ag2 path too).
Both shapes trip autogen's allowed-types gate with
ValueError("Wrong content format: unknown type image within the
content")
…BEFORE the request reaches the vision model — observed live in the
D6 multimodal probe (commit d8a0a25db, which originally quarantined
the feature as NSF).
Fix
---
Add ``agents/_multimodal_normalize.py``: a ``NormalizingAGUIStream``
subclass of ``AGUIStream`` that overrides ``dispatch()`` to normalize
AG-UI image/document/binary content parts to OpenAI Chat Completions
``image_url`` parts AFTER ``RunAgentInput`` Pydantic parsing and BEFORE
``AgentService`` serialises the messages for autogen.
This is the only correct interception point:
- Too early (ASGI body rewrite before Pydantic): ``RunAgentInput``
rejects ``image_url`` because it is not an AG-UI standard type —
the discriminated union only accepts image/document/binary/text.
- Too late (inside ConversableAgent): requires patching autogen
internals.
The override works by calling ``normalize_messages_for_autogen()`` on
the dict-serialised messages (same form as ``run_stream`` produces via
``model_dump()``) and re-injecting them via a ``_PatchedRunAgentInput``
wrapper that overrides only ``.messages``, delegating all other
attribute access to the original ``RunAgentInput``.
Conversions:
- {"type": "image", "source": {"type": "data", value, mime_type}} →
{"type": "image_url", "image_url": {"url": "data:<mime>;base64,<value>"}}
- {"type": "image", "source": {"type": "url", value}} →
{"type": "image_url", "image_url": {"url": value}}
- {"type": "document", "source": ...} → image_url with the document's
mime preserved (data:application/pdf;base64,...). The vision model
still can't natively read PDFs, but the request reaches the model
instead of being rejected upstream, which is the failure mode this
fix targets.
- {"type": "binary", mimeType, data | url} → image_url (the
legacy-shim parts ride through cleanly).
- {"type": "text", ...} and already-normalised image_url parts pass
through unchanged (identity-preserved on no-op turns).
Failure path: any normalization error is logged at WARNING and the
original messages are forwarded unchanged — autogen's own ValueError
fires verbatim with its error surface intact.
Manifest + fixture
------------------
- showcase/integrations/ag2/manifest.yaml: remove multimodal from
not_supported_features (with its now-stale comment) and add it back
to the features list next to voice.
- showcase/aimock/d6/ag2/multimodal.json: add the D6 fixture pair
using the actual autoPrompt strings from sample-attachment-buttons.tsx
("can you tell me what is in this demo image I just attached" /
"can you tell me what is in this demo pdf I just attached").
TDD evidence (red-green)
------------------------
showcase/integrations/ag2/tests/python/test_multimodal_normalize.py
contains 14 unit tests, pinned at three layers:
1. RED/GREEN against autogen's actual content gate:
* test_autogen_rejects_raw_agui_image_part — confirms
content_str([{type: image, source: ...}]) raises the verbatim
ValueError the D6 probe surfaced. This is the regression pin: if
autogen ever relaxes the gate, this test fails and we know to
revisit the normalizer.
* test_normalized_content_is_accepted_by_autogen — after
normalize_messages_for_autogen(...), content_str accepts every
part and renders "<image>" for the image_url part.
2. Shape coverage: modern image data/url, modern document, legacy
binary data/url, mimeType camelCase alias, plain-text passthrough,
plain-string content, assistant/tool messages untouched,
unrecognised source → text placeholder, idempotency.
3. NormalizingAGUIStream class surface tripwire.
Control-plane D6 RED→GREEN:
RED (no normalizer, pre-fix container): d6:ag2/multimodal → red
(HTTP 500 agent_run_error_event from content_str ValueError)
GREEN (NormalizingAGUIStream applied): d6:ag2/multimodal → green
## What
Adds a **Session-stack discipline / Cleanup after isolated runs**
subsection to `showcase/TESTING.md`, governing how `--isolate`/`--keep`
is used across a debugging/testing session.
## Why
`--keep` correctly lets an `--isolate <name>` stack survive a run so it
can be reused for a session-long test set. The leak was **agent
discipline**, not the flag:
1. Agents minted a **new** named kept stack per individual cell instead
of reusing ONE stack for the whole session — which is how Docker
accumulated `cvtest2`, `greenproof`, `conformred`, `conformgreen`,
`gp1`..`gp10`, `showcase-iso2/4`, etc.
2. When the session's work was done, the stacks it created were never
torn down — each one holds a slot and offset ports until the host fills
up.
## The discipline encoded
1. **One stack per session, reused** — choose ONE stable `--isolate
<session-name> --keep` and reuse it for ALL tests in the session (derive
the name from the primary slug, e.g. `--isolate <slug>-session`). Never
mint a new named stack per cell/feature/pill.
2. **`--keep` is for intra-session reuse only, never a license to leak**
— if you pass `--keep`, you OWN teardown at session end.
3. **Tear down at session end** — use the survival-notice command
`docker compose -p <name> down --remove-orphans --volumes && rm -rf
<run-dir> <slot-dir>`. `bin/showcase down` does NOT tear down isolated
stacks (it only stops the default `showcase-*` project). Bare
`--isolate` (no `--keep`) auto-cleans and is preferred for one-off
tests.
Teardown mechanics live once in `DEBUGGING.md → Cleanup` (cross-linked);
this section owns the discipline.
## Verification
This is a docs-only behavior change — no probe/Playwright/code red-green
surface applies. Doc quality gates run instead:
- `oxfmt --check showcase/TESTING.md` → passes (the file was correctly
formatted on `main`; the only formatter touch was `*new*` → `_new_`).
- Diff is purely additive (+49 lines, one file).
- Cross-link anchor `DEBUGGING.md#cleanup` verified against the `###
Cleanup` heading.
The three harness facts the guidance relies on were verified by reading
`scripts/cli/_common.sh`, `cmd-test.sh`, and `bin/showcase`: each
`--keep` run claims a fresh slot + idempotent pre-down + brings the
stack up (no attach); a same-name re-run against a still-live kept stack
fails loudly on the duplicate-name guard; the exact teardown command
matches the survival notice at `_common.sh:~1074`.
## Follow-up (not in this PR)
The harness could make this self-enforcing — e.g. warn when a session
uses >1 distinct kept `--isolate` name, or add a `bin/showcase slots
--reap-mine` convenience to tear down all stacks this user created.
Noted for later; no harness changes here.
🤖 Generated with [Claude Code](https://claude.com/claude-code)