## Release bot v0.1.0
**Scope:** `bot` | **Bump:** `minor`
---
### How this release process works
1. **This PR was created automatically** by the "release / create-pr"
workflow.
It bumped the `bot` packages to `0.1.0`
and generated AI-enhanced release notes.
2. **CI runs on this PR** — the full test suite (unit tests, lint, type
checks, build)
must pass before merging. This is the review gate.
3. **Review the release notes** in `release-notes.md` in this PR.
If a Notion draft was created, you can edit the release notes there
before merging.
4. **When this PR is merged**, the `release / publish` workflow
automatically:
- Builds all packages
- Publishes the `bot` packages to npm at version `0.1.0`
- Creates git tag `bot/v0.1.0`
- Creates a GitHub Release with the final release notes
### Before merging
- [ ] CI is green (tests, lint, types, build)
- [ ] Version bumps look correct
- [ ] Release notes are accurate (edit in Notion if a draft was created)
---
> **Do not merge until CI is fully green.** The full test suite runs
automatically on this PR.
## What
Adds anonymous **`oss.bot.*`** usage telemetry to the `@copilotkit/bot`
SDK — a **configured → started → agent_run** adoption funnel that
answers "what are people doing with the bot packages?" (which adapters,
which store backends, how they iterate on config, where they fail) —
without knowing *who* the customer is.
Design spec: [Bot SDK Telemetry — oss.bot.* Adoption
Funnel](https://app.notion.com/p/38b3aa3818528177a3abfda9ce35bdbc) ·
Plan: [Implementation
Plan](https://app.notion.com/p/38b3aa38185281d5a09cfea114df8897)
## Events (all 100%, anonymous, metadata-only)
| Event | When | Notable props |
|-------|------|---------------|
| `oss.bot.configured` | `createBot()` returns (repeats per restart →
iteration signal) | `platforms`, `store`, `toolsCount`, `hasComponents`,
… |
| `oss.bot.started` | `bot.start()` connected ≥1 adapter | `platforms`,
`startedCount`, `failedCount`, handler flags |
| `oss.bot.start_failed` | an adapter threw during start | `platform`,
`errorClass` |
| `oss.bot.agent_run` | a successful agent run | `platform`,
`durationMs`, `toolCallCount`, `iterations`, `interrupted` |
| `oss.bot.agent_run_failed` | an agent run errored | `platform`,
`errorClass`, `stage` |
Catalog + property reference: `packages/bot/telemetry-events.json`.
## Design
- **Anonymous.** Bot deployments carry no `telemetry_id`/license/API
key, so events are stitched by a persisted `anonymous_id` (durable
`StateStore` → project-local cache file → per-process UUID) plus a
per-`createBot` `bot_session_id`. Reuses `lambdaClient.send` from
`@copilotkit/shared` (no new deps).
- **Zero-config.** Works out of the box; **no new env vars**. Reads only
the pre-existing optional `COPILOTKIT_TELEMETRY_URL` (override) and
`COPILOTKIT_TELEMETRY_DISABLED` / `DO_NOT_TRACK` (opt-out), and
`NODE_ENV`/`VITEST` for the `environment` tag + test suppression.
- **Fire-and-forget.** `capture()` never throws into or blocks the host
app; all dispatch/fetch/fs failures are swallowed.
- **PII guardrails.** Snapshots are scalars/counts only (never
adapter/store option objects → no token/connection-string leakage);
errors are mapped to a bounded `errorClass` category
(`auth`/`network`/`timeout`/`validation`/`unknown`) — never a raw
message or stack; no message text, user ids, or channel names anywhere.
**Sink side** (`oss.bot.` allow-list prefix + `anonymous_id` → PostHog
`distinct_id`) shipped separately in oss-path-to-production PR #175
(already merged).
## Testing
**Unit (TDD, red→green per module):**
- `sanitize-error.test.ts` — category mapping + a "never leaks the
message" secret-redaction assertion (2)
- `install-id.test.ts` — durable-store persistence, file persistence,
unwritable-dir fallback (3)
- `bot-telemetry.test.ts` — global-props shape, disabled no-op,
never-throws-on-send-reject, event-name set (4)
- `events-catalog.test.ts` — drift guard: catalog keys == emitted event
names (1)
- `run-loop.test.ts` — new `{iterations, interrupted}` return value
(interrupt + normal paths)
- `create-bot-telemetry.test.ts` — mocked-telemetry wiring: configured
snapshot, started/start_failed (with `xoxb-SECRET` redaction assertion),
agent_run
**End-to-end:**
- `e2e-telemetry.test.ts` — drives the **real `BotTelemetry`** through
`createBot → start → runAgent` (only the network boundary is stubbed),
asserts the `configured → started → agent_run` payloads carry
`anonymous_id`/`bot_session_id` and **runs with an empty telemetry env**
(proves zero-config; asserts no license token attached).
**Full suite (authoritative, on a real `pnpm install` in the integration
worktree):**
- `tsc --noEmit` → 0 errors · `@copilotkit/bot` vitest → **26 files /
141 tests pass**
**In-session manual smoke (real wire bytes):** stubbed
`globalThis.fetch`, drove a real bot through one turn, captured the
actual serialized POSTs:
```
sink URL (no env var set → default): https://telemetry.copilotkit.ai/ingest
oss.bot.configured {"properties":{"platforms":["fake"],"adapterCount":1,"store":"memory",...},
"global_properties":{"anonymous_id":"74b3c598-…","bot_session_id":"7cddd757-…","environment":"production"},
"package":{"name":"@copilotkit/bot","version":"0.0.3"},"ts":1782502088}
oss.bot.started {"properties":{"platforms":["fake"],"startedCount":1,"failedCount":0,"hasMentionHandler":true,...}, "global_properties":{ same anonymous_id + bot_session_id }}
oss.bot.agent_run {"properties":{"platform":"fake","durationMs":0,"toolCallCount":0,"iterations":1,"interrupted":false}, "global_properties":{ same anonymous_id + bot_session_id }}
```
All three share one `anonymous_id` + `bot_session_id` (funnel
stitching), default sink URL with no env var set, no license/key
anywhere.
**No new env vars (verified):** `git diff origin/main --
packages/bot/src/create-bot.ts packages/bot/src/thread.ts | grep
process.env` → none; wiring code reads no env var directly.
**Code review:** one focused reviewer pass on the full diff — no
blocker/high findings; PII, fire-and-forget, zero-config, and
no-regression axes all confirmed clean. Two low/nit findings applied
(dropped `TypeError`→`validation` mis-categorization; documented the
interrupt→resume `agent_run` double-count in the catalog).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
The harness-workers SSOT modeled only the top-level numReplicas, whose drift
gate watched a field that does not drive the live replica count. harness-workers
is single-region (us-west2); Railway derives the live count from
multiRegionConfig.us-west2.numReplicas.
- WorkerProvisioning gains effectiveReplicas (= multiRegionConfig.us-west2.
numReplicas), the authoritative field the drift gate now asserts. Top-level
numReplicas is retained as a documented mirror.
- Declared values reflect current reconciled reality (verified live via the
Railway GraphQL environment.config staged-config read): prod and staging both
at effectiveReplicas=6 (parity achieved by B-reconcile scaling prod 3 -> 6 in
both the top-level field and multiRegionConfig). BROWSER_POOL_MAX_CONTEXTS is
40 on both envs (verified live).
- Regenerated railway-envs.generated.json; drift-gate test asserts
effectiveReplicas (RED-GREEN proven). RAILWAY.md documents multiRegionConfig
as the effective knob and the achieved parity.
Adds a monorepo-invariant test that scans sibling bot-* adapter packages for their declared `platform` literal and fails if any is missing from normalizePlatform's allow-list. It immediately caught the new bot-teams adapter (platform "teams") that landed on main — now added to the allow-list.
Add WorkerProvisioning interface and workerProvisioning field to the
harness-workers ServiceEntry in railway-envs.ts. Declares current live
reality: prod=3 replicas, staging=6 replicas, BROWSER_POOL_MAX_CONTEXTS=40
per worker (both envs).
Worker model: 1-worker-per-replica (Railway runs one process per
container, keyed on HOSTNAME). HARNESS_POOL_COUNT is informational
only — not a fork factor. Authoritative concurrency knob per worker is
BROWSER_POOL_MAX_CONTEXTS.
Comments flag: staging config-field drift (Railway field=2, live=6) as
a follow-up item; and the prod/staging parity decision (prod=3 vs
staging=6) as deliberately deferred.
Extends emit-railway-envs-json.ts to emit workerProvisioning into the
generated JSON snapshot. Adds drift-gate test
(harness-workers-provisioning.test.ts) that fails if SSOT numReplicas
diverges from the committed JSON snapshot — no live Railway API calls.
Red-green-red-green verified locally.
Updates RAILWAY.md with the 1-worker-per-replica model, declared
values, manual apply procedure, and drift-gate reference.
The existing tooling is verify-only for numReplicas; applying a replica
count change to Railway remains a manual operation (Railway Dashboard or
GraphQL API).
The llamaindex copilotkit API route logged '[copilotkit/route] POST ...'
and '[copilotkit/route] Response status: 200' unconditionally on EVERY
request. Under d6 probe fan-out this exceeded Railway's 500-logs/sec cap
('Messages dropped' -> 'Stopping Container'), killing the replica.
Gate both per-request console.log lines behind a SHOWCASE_ROUTE_DEBUG env
flag (default off). Module-load logs and error logging are unchanged.
This chatty pattern is shared/copied across ~16 integrations (including
the langgraph-python gold standard); this commit scopes the fix to
llamaindex. The others are flagged as follow-up.
A normal harness deploy rebuilds the shared showcase-harness image and
bounces the pool workers (PR #5715). Immediately after the bounce the
workers re-register, the producers re-arm, and every family is mid-sweep:
lastSuccessAt still points at the pre-bounce success, so it reads stale
against the silence thresholds (banner 2x period, Slack alert 3x period +
3 consecutive ticks). The result was a FALSE "worker family X has not
completed successfully" banner AND Slack family-silence alert during the
expected post-bounce drain window.
Fix: a bounce-keyed grace window. The freshest worker registered_at across
the /api/runs workers strip is the fleet's most-recent bounce instant
(independent of CP boot — a worker can bounce while the CP stays up). While
now - bounce < 2 x period, a family with no success yet is DRAINING, not
silent, so neither the §7.4 banner, the §7.3 cell glyph, nor the §9 Slack
alert flags it. Beyond the window with still-no-success, genuine silence
fires exactly as before.
The determination lives in two surfaces (server monitor for the Slack
alert; client isFamilySilent for the banner + glyph), so both now consume
the same new SSOT field (WorkerView.registeredAt) and the same 2x-period
grace constant, keeping them consistent.
- run-view.ts: project registered_at -> WorkerView.registeredAt (server)
- family-silence-monitor.ts: BOUNCE_GRACE_PERIOD_MULTIPLIER + freshest-bounce
grace gate, keyed off body.workers
- worker-runs-context.tsx: freshestBounceMs + bounceAtMs grace arg on
isFamilySilent; banner + cell glyph pass it
- ops-api.ts: WorkerView.registeredAt on the client DTO
A GREEN coverage cell flapped green -> grey -> green every time its probe
job's worker lease lapsed and the control-plane sweeper re-queued the job
(worker-reclaimed-pending). The data layer already preserves the
last-known-good colour in chipColor while surfaceState flips to "pending",
but depth-chip.tsx's `if (pending)` branch did a destructive grey
early-return that never read chipColor.
Split the pending branch:
- prior-good (chipColor is a real colour, or depth > 0): render the normal
coloured chip via the default branch's exact ternary (threading regression
through) PLUS a non-destructive refreshing affordance -- a corner ⟳ glyph
and a subtle pulsing ring in the chip's own colour, with
data-refreshing="true" / data-has-prior="true". Colour preserved.
- no-prior (never-run / first load: gray + depth 0): keep today's honest grey
⟳ chip, now tagged data-has-prior="false".
The unreachable red ⚡ overlay and the "failure never masked" gate are
untouched; red/regression cells pass through with no spinner. The refreshing
cue is conveyed by shape (⟳) + motion (ring), never by colour alone, and the
pulse/spin respect prefers-reduced-motion (motion-reduce:animate-none), so the
static ⟳ carries the meaning for colour-blind and reduced-motion operators.
Renderer-only change; unified-cell.tsx is the sole caller passing `pending`.
Address PR review: bound the free-form adapter platform label to slack|discord|telegram|whatsapp|custom (no tenant/project leakage, capped cardinality); emit oss.bot.agent_run only after the transcript-append + renderer.finish() steps succeed, with a finalize stage on oss.bot.agent_run_failed for late failures.
--isolate <name> <slug> (name before slug) mis-parsed the name as the
slug, then died "Unexpected argument"; a bare --isolate= silently
fell through to auto-pick. Defer the ambiguous post---isolate token to
pending_iso_name and resolve it after the parse loop (slug present =>
token was the name; no slug => token was the slug); reject an empty
--isolate= loudly; --isolate=<name> (non-numeric) now binds an explicit
isolate name. The --isolate=<N> numeric pin is unchanged.
_slot_ports_free consumed _slot_offset_ports via process substitution
(done < <(...)), so a die on an out-of-range/non-numeric slot exited only
the subshell — the loop read zero ports, any_held stayed 0, and the
function returned 0 ("all free"), silently defeating the port-conflict
guard for a bad slot. Capture into a variable with || die so the failure
propagates to the caller. Real-surface bats prove a bad slot now fails
loudly while valid-slot free/held behavior is unchanged.
createBot emits configured + started/start_failed; Thread emits agent_run/agent_run_failed around runAgentLoop. Zero new env vars. Includes mocked-wiring unit tests and a real-BotTelemetry e2e test.
BotTelemetry posts 5 oss.bot.* events at 100% via @copilotkit/shared's lambdaClient. Anonymous (no telemetry_id/license/key), opt-out via COPILOTKIT_TELEMETRY_DISABLED/DO_NOT_TRACK, suppressed under test. 3-tier anonymous_id (durable store -> project cache file -> per-process UUID). errorClass() maps errors to a bounded category, never a raw message.
## Summary
Hardens the showcase `--isolate` slot lifecycle to stop leaked Docker
stacks (we had 16 leaked `--keep` stacks accumulate: 89 containers, 16
volumes).
1. **Slot-liveness false-positive fix** — a `--keep`'d stack whose
owning process exited (but whose containers kept running) was classified
`live` forever and never reaped. New start-time-verified
`_owner_liveness` probe + a `kept` state; `slots` now renders
`<pid>(dead)`/`(reused)` and adds a `--reapable` filter.
2. **Kept-stack TTL** — `ISOLATE_KEEP_TTL` (4h,
`SHOWCASE_ISOLATE_KEEP_TTL`-overridable) flips an over-age `kept` slot
to `stale` so the claim-time sweep reclaims it (with a loud warning).
3. **`showcase reap` subcommand** — dry-run by default;
`--force`/`--all`/`--include-live`/`<name|slot>`; identifies
harness-owned stacks via slot-record ∪ run-dir ∪ `showcase-iso<N>` ∪ a
new `com.copilotkit.showcase.isolate` self-id label stamped by
`apply_isolation`; never touches the base `showcase` stack or BuildKit
resources.
Spec: https://app.notion.com/p/38b3aa3818528137a399fafee3750463
## Tests
Real-surface bats (real slot dirs, real dead PIDs via spawn+wait, real
running compose projects — never mocks): `bats
showcase/scripts/__tests__/` = 161/0. CR converged in 2 rounds (4
fixes), Procedure 3 promotion audit clean.
## Deferred to a follow-up PR (pre-existing, out of this PR's subject)
- `_slot_ports_free` die-in-subshell defeats the port-conflict guard for
a bad slot (byte-identical in `main`).
- `cmd-test.sh` `--isolate <name> <slug>` (name-before-slug) mis-parse +
bare `--isolate=` validation gap.
- `slots` table shows slot 0's PORTS at offset +200 while displaying
OFFSET +0.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Dry-run by default (lists the plan, changes nothing); --force executes,
--all ignores TTL/keep, --include-live opts into reaping a live-owner
target, <name|slot> targets one. Identifies harness-owned projects via
the slot-record / run-dir / showcase-iso<N> / self-id-label union, and
never touches the base 'showcase' stack or BuildKit resources. Real
docker bats prove dry-run/--force/--all + the base/buildkit guards.
A --keep'd isolated stack whose owning process had exited (but whose
containers kept running) was classified 'live' forever and never reaped,
leaking Docker stacks indefinitely. Introduce a start-time-verified
_owner_liveness probe and a new 'kept' state, an ISOLATE_KEEP_TTL (4h,
SHOWCASE_ISOLATE_KEEP_TTL-overridable) that flips an over-age kept slot
to 'stale' so the sweep reclaims it, a com.copilotkit.showcase.isolate
self-id label stamped by apply_isolation, a 'slots --reapable' filter,
and a macOS lsof COMMAND-truncation fix in the own-project port filter.
Real-surface bats cover the liveness false-positive and TTL reaping.
## What
Adds **Strategy 10** to `showcase/DEBUGGING.md` capturing a hard-won
production-debugging lesson from the 2026-06-26 incident: **a red / BE✗
dashboard cell does NOT mean the feature is broken — it's often
staleness.**
## The lesson
- The coverage dashboard's per-cell **BE (D4) flag = `resolveD4` =
worst-of(`chat:<slug>`, `tools:<slug>`)**, then folded by a **staleness
window** (`staleness.ts`: `D4_STALE_AFTER_MS = 60m`; D3/D5/D6 + family
aggregates use `E2E_STALE_AFTER_MS = 6h`). A green row older than its
window folds to stale → renders red / BE✗.
- **Reading a single PocketBase collection row (e.g. `chat:<slug>`) is
NOT the dashboard's flag** — it ignores `tools:` and ignores staleness,
and will falsely report "BE green." Reproduce `resolveD4`'s worst-of +
staleness logic.
- **Root failure mode:** if a probe sweep takes longer than the
staleness window, cells the sweep hasn't re-touched go stale and render
red even when the app is fine. Evidence (2026-06-26 prod): sweep
durations vs periods — d5 41m/15m, e2e-smoke 45m/15m, e2e-demos 97m/60m,
d6 127m/60m; with the worker pool starved (concurrency = `numReplicas ×
HARNESS_POOL_COUNT`), the D4 sweep blew past 60m → ~13 integration
columns showed BE✗ while the apps were healthy. Scaling worker
concurrency so a sweep completes in-window restored them.
- **Prod-vs-staging disparity** is frequently this — same code, but one
env's harness can't complete sweeps within the staleness windows.
Includes a 5-step diagnostic checklist (check freshness + worker
throughput *before* blaming the app) and an Anti-Patterns entry.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
A red/BE✗ dashboard cell is often staleness, not a broken feature. The
per-cell BE flag is resolveD4 = worst-of(chat,tools) folded by a staleness
window; a green row older than its window folds to stale-red even when the
app is healthy. Add D5 Strategy 10 with a diagnostic checklist (harness
/api/runs sweep-duration vs window, observed_at age vs *_STALE_AFTER_MS,
numReplicas × HARNESS_POOL_COUNT concurrency) plus an anti-pattern entry.
Evidence from the 2026-06-26 prod incident: D4 sweep (127m d6 / 97m e2e-demos)
blew past the 60m window with a starved worker pool, stale-reddening ~13
integration columns while apps were fine.
## What & why
Brings showcase **strands-typescript** to full **D6 parity** and fixes
the production/staging aws-strands column flap.
**Root cause of the flap:** strands-typescript is a *two-process*
integration — the Next route is a bare `@ag-ui/client` `HttpAgent`
proxy; the actual model call happens in the separate Express agent.
`@ag-ui/aws-strands@0.2.3` drops inbound headers before `agent.run()`,
so the probe's `X-AIMock-Strict` never reached the outbound aimock call.
On a fixture miss, instead of a strict **503**, the request **silently
proxied to real OpenAI** — non-deterministic, intermittently red.
## Changes (by concern)
1. **`feat`: forward `X-AIMock-Strict` end-to-end through the
two-process hop** — per-request header forwarding via
`AsyncLocalStorage` + a custom `fetch` shim on both the Next `HttpAgent`
(`forwardingProxyFetch`, null-guarded) and the Express `OpenAIModel`
(`forwardingFetch`), including the sub-agent `openaiClient`.
Never-clobber merge keeps the static `x-aimock-context` slug
authoritative; byte-identical to plain `fetch` when no `x-*` are in
scope (demo traffic unaffected).
2. **`feat`: agent-side CVDIAG backend instrumentation** — emitter
middleware mounted before the aws-strands handler (the real backend
boundary lives in the Express process), staged by the `cvdiag-stage-ts`
generator; `sseChunkByteLength` counts `byteLength` for any
`ArrayBufferView` (was zeroing typed-array SSE chunks).
3. **`fix`: guard `crypto.randomUUID`** in the 2 headless chat shells
(undefined on insecure-origin harness).
4. **`fix`: `COPY src/cvdiag`** into the Docker runner so the
two-process agent boots; **`docs`**: RAILWAY.md +
INTEGRATION-CHECKLIST.md now require CVDIAG + strict-forwarding + the
Dockerfile COPY for new/promoted integrations.
## Red→green proof (local, real failure surface)
- **X-AIMock-Strict e2e**: pre-fix the outbound aimock call carried only
the static context slug → fixture miss fell through; post-fix the
inbound strict header is forwarded across all 5 hops (traced file:line)
→ miss fails loud.
- **route null-guard** (`forwarding-proxy-fetch-nullguard.test.mts`,
5/5): pre-fix `new Headers(requestInit.headers)` throws `TypeError` on
undefined init; post-fix `requestInit?.headers` safe.
- **sub-agent forwarding** (`tools.test.ts`): pre-fix outbound
`x-aimock-strict` absent (spy asserts `null`); post-fix present.
Re-proven by revert→fail, restore→pass against `openai@6.44.0`.
- **cvdiag byteLength** (`cvdiag-backend-strands.test.ts`): pre-fix
Uint8Array chunk size `0`; post-fix `byteLength` (6). Covers
string/Buffer/Uint8Array/unknown.
## Deploy verification
Built + pushed to GHCR (`sha256:da25c776…385bba2`), pinned + deployed to
**staging** (production untouched): `/api/health` 200, agent healthy,
emitter import verified in-container, and the **full aws-strands D6
column is green, visually verified** via the dashboard.
## Review
- 7-agent CR confirmation round **converged: 0 blocking (bucket-a)**
findings after adjudication (the one flagged item — a pre-existing
`runSubagent` catch — is outside this PR's diff and the PR strictly
improves on it).
- Procedure-3 promotion audit: **PROMOTE: none** (the cvdiag classifier
is failure-*diagnosis*, not the D6 grade computation, so no
backend-outcome label can flip a cell's grade).
- Pre-push: oxlint clean, `next build` 58/58, all red-green tests pass.
## Caveats / follow-ups
- The new `vitest` tests are **local/CR red-green guards and are not run
by any CI workflow** (CI ignores `showcase/**` integration `src/agent`
vitest; same as sibling integrations built-in-agent /
langgraph-typescript). They guard against local regressions, not in CI.
- **Deferred to a separate security pass** (pre-existing, out of scope
here): `route.ts` 500-handler leaks `err.stack` (153-156) and the GET
health endpoint reports `OPENAI_API_KEY` presence (160-182).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Emit CVDIAG backend boundary markers from the agent process for strands-typescript (byteLength fix on sseChunkByteLength), enable the emitter in docker-compose.local.yml, vendor src/cvdiag, and exclude tests from tsconfig.
Forward inbound X-AIMock-Strict header through the two-process strands-typescript hop (Next route -> agent -> sub-agent fetch), with null-guard on the forwarding proxy fetch and supporting unit tests.
Bring the agnostic root A2UI docs up to the catalog-on-provider model and
make every generated framework serve them consistently.
- Root /generative-ui/a2ui (index, fixed-schema, dynamic-schema): lead with
passing a catalog on the provider (auto-enables A2UI and auto-injects the
generate_a2ui tool), add a manual opt-out section explaining the two pieces
you wire yourself (the generate_a2ui agent tool and the A2UIMiddleware), and
set fixed-schema to injectA2UITool: false since the agent owns the tool.
- Flip langgraph-fastapi, strands, strands-typescript to docs_mode: generated
so they serve the shared root A2UI docs 1:1 with langgraph-python.
Generated frameworks covered: langgraph-python/fastapi/typescript, google-adk,
strands, strands-typescript. deepagents (authored) is handled separately.
## Root cause
The shared `static / quality` workflow's **`format`** job auto-formats
PR files and pushes a fixup commit. Two pieces were misaligned:
- **"Run formatter (fix on PR)"** set `format_fixed=true` from a
**whole-tree** `git diff --name-only` (old line ~126).
- **"Commit formatting fixes"** (gated on `format_fixed == 'true'`)
staged **only the scoped** PR files (`xargs -a
.pr-format-files.existing.txt git add --`, old line ~163) and ran a bare
`git commit` (old line ~164).
When a PR's own files are already formatter-clean **but the runner's
working tree is dirty for an unrelated reason** — e.g. an LFS smudge on
`examples/teams/appPackage/*.png` (declared `*.png filter=lfs`) — the
whole-tree diff falsely set `format_fixed=true`, the **scoped `git add`
staged nothing**, and `git commit` exited **1** ("nothing to commit") →
the job **failed**.
This intermittently red-flagged any PR depending on per-runner
LFS-smudge state (cf. #5715, where the format check failed and then
passed on re-run with no code change).
## The fix (minimal, `format` job only)
1. **Scope the trigger.** Set `format_fixed` from a **scoped** diff
(`git diff --quiet -- <scoped files>`) guarded by `[ -s
.pr-format-files.existing.txt ]` so unrelated working-tree drift no
longer triggers the commit path, and the empty case never degrades to a
whole-tree diff.
2. **Guard the commit.** After the scoped `git add`, treat an empty
index as a no-op (`git diff --cached --quiet && exit 0`) instead of
letting `git commit` exit 1.
The real behavior is preserved: when a scoped PR file genuinely needs
formatting, the job still stages, commits, and pushes the fix. No other
job is touched.
```diff
@@ Run formatter (fix on PR) @@
- if [ -n "$(git diff --name-only)" ]; then
+ # shellcheck disable=SC2046 # intentional split: each path is a
+ # separate `git diff` pathspec arg; the `-s` guard rules out the
+ # empty-arg (whole-tree) case, and PR paths never contain spaces.
+ if [ -s .pr-format-files.existing.txt ] && \
+ ! git diff --quiet -- $(cat .pr-format-files.existing.txt); then
echo "format_fixed=true" >> "$GITHUB_ENV"
fi
@@ Commit formatting fixes @@
xargs -a .pr-format-files.existing.txt git add --
+ if git diff --cached --quiet; then
+ echo "No scoped formatting changes to commit"
+ exit 0
+ fi
git commit -m "style: auto-fix formatting"
git push
```
## Local RED → GREEN proof
GitHub Actions can't run locally, so the job's **exact shell** was
reproduced in a throwaway `/tmp` git repo, using real `oxfmt@0.36` and
GNU `gxargs` for Linux-runner fidelity (`xargs -a` is GNU-only). Scoped
PR file = `app.js`; unrelated tracked file = `unrelated.bin`, left
**dirty** to simulate the LFS smudge.
**RED — current logic (whole-tree trigger + scoped add + bare commit):**
```
[RED] trigger env='format_fixed=true' # whole-tree diff saw unrelated.bin
[staged after scoped add]: '' # app.js already clean → nothing staged
nothing to commit
>>> RED git commit exit code: 1 (JOB FAILS)
```
**GREEN — fixed logic, same repo state:**
```
[GREEN] trigger env='' # scoped app.js clean; unrelated.bin ignored
[gate] format_fixed NOT set -> Commit step SKIPPED
>>> GREEN exit code: 0 (job passes, no false trigger)
```
**POSITIVE — scoped file genuinely needs formatting
(committed-unformatted `app.js`, `unrelated.bin` still dirty):**
```
after oxfmt — app.js: const z = 3;
[POS] trigger env='format_fixed=true'
[staged]: 'app.js' # only the scoped file
[POS] commit exit=0
[POS] git log: 46d4849 style: auto-fix formatting
[POS] worktree: ' M unrelated.bin' # unrelated drift left uncommitted
```
So: false-trigger failure is eliminated (RED→GREEN), and the real
auto-format path still commits & pushes exactly the scoped fix
(positive).
## Validation
- `python3 yaml.safe_load(...)` → YAML OK
- `actionlint .github/workflows/static_quality.yml` → exit 0 (baseline
on `main` is also clean; the SC2046 word-split warning introduced by the
scoped diff is suppressed with a narrowly-scoped, commented `shellcheck
disable` for the intentional split).
🤖 Generated with [Claude Code](https://claude.com/claude-code)
The format job set format_fixed=true from a whole-tree git diff and then
git-committed only the scoped PR files. When a PR's own files are already
formatter-clean but the runner's working tree is dirty for an unrelated
reason (e.g. an LFS smudge on examples/teams/appPackage/*.png, which are
*.png filter=lfs), the whole-tree diff falsely triggered the commit path
while the scoped git add staged nothing, so git commit exited 1 and failed
the job. This intermittently red-flagged any PR depending on per-runner
LFS-smudge state (cf. #5715).
- Trigger format_fixed only when a SCOPED file actually changed.
- Guard the commit so an empty staged set is a no-op (exit 0) instead of a
hard failure.
## Root cause
Prod's `harness-workers` fleet worker runs a **stale `showcase-harness`
image** because it had **no `prod` env entry** in the railway-envs SSOT.
- The worker (`serviceId c2aa8a0b-350e-4b76-8541-3012dfac41d0`, prod
instance `7c48ee43-6df4-457b-b977-10f1f1ac1680`, `HARNESS_ROLE=worker`)
consumes the shared `showcase-harness` image via `imageOf: "harness"`.
- `expandImageConsumers(names, env)` is **env-aware**: a consumer only
enters an env's redeploy scope if it declares that env
(`redeploy-env.ts:278` — `if (!Object.hasOwn(entry.environments, env))
continue;`).
- Because the worker modeled **staging only**, a rebuilt
`showcase-harness:latest` bounced the prod control-plane but **silently
skipped the prod worker**, which kept its stale **2026-06-19** image.
- That stale image bakes a **1-demo `registry.json`** for
`ms-agent-harness-dotnet` (only `beautiful-chat`). The hourly
`e2e_demos` driver runs on the worker → resolves 1 demo → writes only
`e2e:ms-agent-harness-dotnet/beautiful-chat`. The other 38 feature rows
never exist in prod PocketBase → `resolveD3.exists === false` → `UI`
badge omitted → broken D3 rung collapses the ladder → **D0**.
(d5/d6 populate fully in prod because the D5/D6 drivers enumerate from a
**compiled-in** script registry, not the `registry.json` data file —
only `e2e_demos` is data-driven, which is why only `UI` was affected.)
## Fix
Backfill the live prod worker as a real `prod` env entry in
`scripts/railway-envs.ts` (real serviceInstance ID `7c48ee43…`), flip
`gateIgnore` off, and set `gateValidated: true`. The env-aware `imageOf`
expansion now pulls the prod worker into the **prod** redeploy scope on
every `showcase-harness` rebuild, so it can no longer drift onto a stale
image.
Also regenerates `railway-envs.generated.json` (Ruby/jq boundary
artifact) and the golden behavior-preservation fixture, and updates the
two gate-count assertions (`gateValidated` services 40→41;
`harness-workers` removed from the gateIgnore set).
## Local RED → GREEN proof
Failure surface: the real `expandImageConsumers("harness", "prod")`
against the real SSOT must include `harness-workers`.
**RED** (prod env entry absent from SSOT — the bug):
```
× includes harness-workers in the PROD redeploy scope when showcase-harness rebuilds
AssertionError: expected [ 'harness' ] to include 'harness-workers'
at __tests__/redeploy-env.harness-worker-prod-scope.test.ts:31:19
Tests 1 failed | 1 passed (2)
```
**GREEN** (after adding the prod `harness-workers` env entry):
```
✓ includes harness-workers in the PROD redeploy scope when showcase-harness rebuilds
✓ still includes harness-workers in the STAGING redeploy scope (no regression)
Tests 2 passed (2)
```
Full SSOT-dependent suite (golden snapshot, emit-json, image-ref gate,
promote closure, verify-matrix, redeploy-env): **140 passed**.
## Note (out of scope for this PR)
This SSOT change ensures the prod worker is bounced on **future**
rebuilds. The currently-live prod worker still needs a one-time
redeploy/restart onto the current `showcase-harness:latest` (39-demo
registry) to immediately backfill the 38 missing rows; that is an
operational step, not a code change.
## What
Brings the **Authentication demo** of three integrations into 1:1
conformance with the `langgraph-python` (LGP) gold standard, completing
the work started in #5713. The showcase Iron Law: LGP is the reference;
every integration must have (1) identical tests, (2) near-identical
frontends, (3) minimal backends, (4) per-integration fixtures.
A conformance audit against LGP found 3 violators (the other 17
integrations already conform):
| Integration | Violation | Fix |
|---|---|---|
| **claude-sdk-python** | Legacy auth-*first* shape: class
`ChatErrorBoundary` + `lastError`, no `handleAuthError`, missing
`sign-in-card.tsx`, divergent banner/hook | Ported
`page.tsx`/`use-demo-auth.ts`/`auth-banner.tsx` **byte-identical** to
LGP + new `sign-in-card.tsx`; added the shared shadcn primitives it
lacked (`lib/utils.ts`, `components/ui/{button,card}.tsx`) +
`radix-ui@^1.4.3` (matching the claude-sdk-typescript peer) |
| **built-in-agent** | Distinct legacy variant:
`ChatErrorBoundary`→`auth-demo-chat-boundary`, local 401-regex
`onError`, auth-first hook | Normalized error-handling shape + hook to
LGP; **preserved** the forced `<CopilotKitProvider>` (default-agent) +
raw-Tailwind divergences (documented in a new `README.md`) |
| **ms-agent-harness-dotnet** | Missing `tests/e2e/auth.spec.ts` (rule
1) | Added LGP's spec **byte-identical** (sha256 `603a68e5…`) |
After this PR, all auth `page.tsx`/hook files are byte-identical to LGP
except documented, forced per-integration wiring; all `auth.spec.ts`
share LGP's sha256.
## Red–green proof (per integration, on the real probe surface)
The shared `d5-auth.ts` probe accepts *either* `auth-demo-error` *or*
`auth-demo-chat-boundary`, so it passes leniently on the legacy shape —
the **discriminating gate is the byte-identical `auth.spec.ts`**
(asserts unauth-first `SignInCard` + `auth-authenticate-button` +
post-sign-out `auth-demo-error`):
- **claude-sdk-python:** legacy frontend → `auth.spec.ts` **6/6 FAIL**
(timeout on `auth-sign-in-button`); conformed → **6/6 PASS** (`next
build` clean).
- **built-in-agent:** legacy → 6/6 FAIL; conformed → 4 conformance
assertions flip FAIL→PASS incl. unauthenticated-send surfaces
`auth-demo-error` (`next build` clean).
- **ms-agent-harness-dotnet:** spec absent (coverage gap) → added →
`--d5 --isolate` green, full real-browser auth flow passes.
## Review
7-agent CR round + mandatory 7-agent confirmation round → **converged to
zero findings** (correctness, conformance, types/build, deps/lockfile,
tests, silent-failures, cross-integration regressions). 2 P2 conformance
nits found and fixed (import-style alignment; restored `DEMO_TOKEN` so
built-in-agent's hook is byte-identical to LGP).
## Known limitation (non-blocking, pre-existing infra)
The GHA workflow `test_e2e-showcase-on-demand.yml` runs Playwright only
for slugs with a Python agent, so the **built-in-agent /
ms-agent-harness-dotnet auth specs are not executed in PR CI**. This is
a pre-existing infra gap (those integrations have no Python agent), not
introduced here. Coverage **does** exist post-merge: the Railway staging
**d6 harness** enumerates services language-agnostically and runs the
auth probe against live `/demos/auth` for both — verified, and it's what
drives their dashboard cells green at D6. A follow-up to add a
non-Python e2e execution path is warranted.
## Notes (pre-existing, not introduced)
- `npm ci`/`npm install` in `showcase/integrations/claude-sdk-python`
shows a micromark/unified desync and a zod/openai ERESOLVE peer conflict
— both reproduce identically at the base commit `ab85b939ac`
(independent of the `radix-ui` add); handled by the existing
`--legacy-peer-deps` path.
Ref: #5713 (original post-sign-out auth rejection fix).
Add a d5-a2ui-recovery probe so the A2UI error-recovery demo runs on every
PR via the d5/d6 fleet harness, not only the manual on-demand workflow.
- New probe d5-a2ui-recovery.ts drives both pills in one session: HEAL
asserts >=2 newly-mounted declarative-metric tiles and no hard-failure
card; EXHAUST asserts the "Couldn't generate the UI" card appears and
no surface paints. Deltas (vs a pre-send baseline) keep the two
mutually-exclusive negatives correct across the shared session. The
transient "Retrying..." label is not asserted (timing-flaky).
- Prompts are sent as typed input, keyed per integration slug, mirroring
each slug's suggestions.ts message verbatim. The recovery prompts are
unique per slug because the inner render_a2ui calls carry no
x-aimock-context; a typed message is byte-identical to the pill
dispatch, so it matches the same fixture. Sending via input (not a
preFill pill click) lets the runner snapshot its run-lifecycle baseline
first, avoiding a false done-signal-missing failure.
- Register a2ui-recovery in d5-registry, map it in d5-feature-mapping,
add its representative fixture, and mirror the mapping in the dashboard
CATALOG_TO_D5_KEY (kept in lock-step via the drift test).
Verified green locally on both recovery paths: langgraph-python
(backend-owned get_a2ui_tools) and strands (auto-inject middleware).
## What
Ports the google-adk **A2UI Error Recovery** demo to five more
frameworks:
- LangGraph (Python)
- LangGraph (FastAPI)
- LangGraph (TypeScript)
- AWS Strands (Python)
- AWS Strands (TypeScript)
Each integration gets a dedicated recovery agent, a scoped runtime
route, the demo page/chat/suggestions, a `manifest.yaml` feature + demo
entry, aimock D6 fixtures, an e2e spec, and a QA doc. The demo reuses
each integration's existing `declarative-gen-ui` catalog
(`declarative-gen-ui-catalog`), so no new components are introduced.
## How recovery is wired (two paths)
- **LangGraph (py/fastapi/ts):** backend-owned. The graph owns
`generate_a2ui` via `ag_ui_langgraph.get_a2ui_tools` /
`@ag-ui/langgraph` `getA2UITools` with `recovery`, and the route sets
`injectA2UITool: false` so the runtime does not double-inject.
(langgraph-python adds `ag-ui-langgraph==0.0.41` to requirements.)
- **Strands (py/ts):** the adapter runs the toolkit validate-retry loop
on its auto-inject path, so the recovery agent is a dedicated clone of
the dynamic agent with no explicit tool.
Two pills per demo:
- **Recover a bad render** (heal): first render is structurally invalid,
the loop retries, the second render is valid and paints.
- **Show an unrecoverable failure** (exhaust): every attempt is invalid,
the loop hits the cap and returns `a2ui_recovery_exhausted`, rendered as
a graceful failure.
## Fixture design notes
- The toolkit's `validate_a2ui_components` rejects the whole surface on
any invalid entry (no single-pass sanitize), so the heal is a genuine
invalid-then-valid sequence staged via aimock `sequenceIndex` (0
invalid, 1 valid). ADK's stringified `parse_and_fix` single-pass heal is
ADK-middleware-specific and does not apply on the langgraph/strands
toolkit loop.
- The inner `render_a2ui` sub-agent call carries no `x-aimock-context`
header (only the harness sets it), and aimock loads every framework's D6
dir into one process. To avoid cross-framework fixture collisions AND to
make the demos fire for real browser (dojo) traffic, each framework uses
unique recovery prompts and the fixtures carry no `context` match field
(userMessage alone disambiguates).
- Also hardens the strands `declarative-gen-ui` composition guide to
name the exact catalog component (`Metric`, not `MetricTile`), which a
real LLM was mis-naming.
## Verification
- langgraph-python recovery e2e: 3/3 (page load, heal paints, exhaust
shows the failure UI).
- Both backend paths confirmed in a real browser via the local dojo
(langgraph-python + strands heal and paint).
- Strands declarative + recovery confirmed grounded under a real LLM
(correct `declarative-gen-ui-catalog` + real components).
- `showcase/scripts` suite green (836 tests), incl. updated
`generate-catalog` (langgraph-python wired 36 to 37) and
`aimock-fixtures` collision checks. Manifest validation 20/20.
## Known limitation
The heal stages invalid-then-valid via aimock `sequenceIndex`, whose
match counter only resets with a fresh `x-test-id` (which the browser
does not send). On a long-lived/shared aimock a repeat heal click
advances past the staged pair. Exhaust is fully repeatable. A follow-up
can send a per-session `x-test-id` so the heal is repeatable and
multi-user safe.
ADK is intentionally left as-is (its recovery fixtures remain
context-scoped).
getA2UITools changed signature: 0.0.39 is getA2UITools(model, options) (positional),
0.0.42 is getA2UITools(params) (single object). The agent code (recovery-agent.ts
and graph.ts) calls the single-object form, but the override pinned 0.0.39, so the
whole params object was treated as the model -> e.bindTools undefined -> the tool
returned {"error":"Provided model does not support bindTools"} and the render
sub-agent never ran. Bumping the override to 0.0.42 aligns the dep with the API the
code uses; verified the recovery graph now emits a healed a2ui_operations surface
(invalid seq0 -> valid seq1) and fires the render_a2ui sub-agent.
The lg-ts agent serves graphs from a hardcoded graphSpec in src/agent/server.mjs
(mirrors langgraph.json). The a2ui_recovery graph was added to langgraph.json but
not graphSpec, so the langgraph server returned 404 on its runs and the demo
never dispatched. Add a2ui_recovery to graphSpec.
NOTE: this fixes graph REGISTRATION. The lg-ts recovery render does not yet fire
(getA2UITools 0.0.39 returns from generate_a2ui without invoking the render
sub-agent); tracked separately, likely needs @ag-ui/langgraph >= 0.0.42.
Port the google-adk a2ui-recovery demo to langgraph (python, fastapi,
typescript) and aws-strands (python, typescript). Each ships a dedicated
recovery agent, route, demo page/chat/suggestions, manifest entry, aimock
d6 fixtures, e2e spec, and QA doc.
Backend-owned recovery on langgraph via get_a2ui_tools / getA2UITools
(injectA2UITool=false); auto-inject recovery on the strands adapter path.
Heal stages an invalid-then-valid render via aimock sequenceIndex (the
toolkit validate->retry loop rejects the whole surface, so a single-pass
parse_and_fix heal is ADK-specific and does not apply here). Recovery
prompts are unique per framework and the fixtures carry no context match
field, so they fire for real browser (dojo) traffic, not just the harness.
Also harden the strands declarative-gen-ui composition guide to name the
exact catalog component (Metric, not MetricTile) and update the
generate-catalog + aimock-fixtures test expectations.
Add an optional onAction interceptor to createA2UIMessageRenderer so apps
can handle A2UI actions client-side (e.g. navigate) instead of forwarding
every action to the agent.
- onAction(action, forward) runs before forwarding. Return null to handle
client-side and stop forwarding, return a modified action to forward it,
or return undefined to forward unchanged.
- Threaded from createA2UIMessageRenderer options through ReactSurfaceHost.
- Extracted runA2UIAction helper for deterministic unit testing.
- Default behavior is unchanged when onAction is not supplied.
## What
Types the bot UI surface and broadens the cross-platform component
vocabulary, adding **only** capabilities more than one adapter can
express.
### 1. `ui` typed as `Renderable` (was `unknown`)
The `Thread` interface's `post` / `update` / `awaitChoice` /
`postEphemeral` now accept `Renderable` instead of `unknown`. JSX is
type-checked at the call site, and the `{ raw }` escape hatch is the
explicit way out when the JSX vocabulary doesn't fit. (The concrete
`Thread` class already used `Renderable` — only the interface leaked
`unknown`.)
### 2. `<Message onReaction>` — per-message reaction callback
```tsx
<Message onReaction={(emoji, r) => (r.added && emoji === "bug" ? triage() : ack())}>
Deploy finished — react 🐛 to file a bug
</Message>
```
- First arg is the emoji (matches the common `r => r === "bug"` shape);
second carries `{ added, user, rawEmoji, messageId }`. Fires on add
**and** remove.
- The handler is stripped from the IR before it reaches the adapter,
then associated with the posted message's id; inbound reactions route to
it.
- **Durable on the same terms as a component `onClick`:** a `{
component, props }` snapshot is persisted and the component is
re-rendered to re-derive the handler after a restart (when the
`<Message>` comes from a registered component + a durable `store`).
Inline handlers route in-process only — identical degradation to an
inline `onClick`.
### 3. `Button.url` link buttons
`<Button url="…">` → native link buttons on **Slack, Discord, Teams,
Telegram** (Telegram already supported it; now typed + wired
everywhere).
### 4. `Field.label`
`<Field label="Status">Online</Field>` — **Discord and Telegram already
read this prop untyped** (latent type gap); now typed and additionally
rendered on Slack.
### 5. `Select.multi` multi-select
`<Select multi onSelect={(vals) => …}>` — **Slack**
(`multi_static_select` in an input block, since Slack forbids it in
`actions` blocks; decodes `selected_options` → `string[]`), **Discord**
(min/max values; decodes via `interaction.component` bounds), **Teams**
(`isMultiSelect`). Telegram/WhatsApp degrade to single-select.
`onSelect` widened to `ClickHandler<string | string[]>`.
### Intentionally dropped (Slack-only)
`Button.confirm`, `Image.title`,
`Input.label`/`initialValue`/`required`, `Select.initialValue` —
single-platform, excluded by design.
## Testing
- All 7 affected packages (`bot-ui`, `bot`, `bot-slack`, `bot-discord`,
`bot-teams`, `bot-telegram`, `bot-whatsapp`) build clean and pass their
full suites.
- New coverage: reaction routing + durability cold-resolve (`bot`), link
button / field label / multi-select render + source-ordering
(`bot-slack`), link button + multi-select bounds + multi decode
(`bot-discord`/`bot-slack`), `Action.OpenUrl` + `isMultiSelect`
(`bot-teams`).
- Two code-review passes (incl. an adversarial one on the durability
change); all findings addressed.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
The reaction handler now receives the conversation `thread` and an
update-capable `messageRef` for the reacted message — the same surface an
`onClick` gets via `ctx.thread`/`ctx.message.ref`. A reaction can now post new
UI (`thread.post`), swap the message in place (`thread.update(messageRef, …)`),
run the agent, or block on a human choice (`thread.awaitChoice`, HITL).
- bot-ui: `MessageReaction` gains `thread` and `messageRef`.
- platform-adapter: `IncomingReaction` gains an optional adapter-provided
`messageRef` (engine falls back to `{ id: messageId }`).
- create-bot: threads `thread` + `messageRef` into both the global
`ReactionEvent` and the per-message handler.
- Slack/Discord/Telegram reaction decoders emit a platform-specific,
update-capable `messageRef` (channel+ts / channelId+id / chatId+messageId).
Type the bot UI surface and broaden the cross-platform component vocabulary,
adding only capabilities that more than one adapter can express.
- bot-ui: type the `Thread` interface's `ui` params (`post`/`update`/
`awaitChoice`/`postEphemeral`) as `Renderable` instead of `unknown`, so JSX
is checked at the call site and the `{ raw }` escape hatch is explicit.
- `<Message onReaction>`: per-message reaction callback `(emoji, { added, user,
rawEmoji, messageId })`. Stripped from the IR before it reaches the adapter,
routed from reaction ingress by message id. Durable on the same terms as a
component `onClick` — a `{ component, props }` snapshot is persisted and the
component is re-rendered to re-derive the handler after a restart; inline
handlers route in-process only.
- `Button.url` link buttons — Slack, Discord, Teams, Telegram.
- `Field.label` — typed (Discord and Telegram already consumed it untyped);
newly rendered on Slack.
- `Select.multi` multi-select — Slack (`multi_static_select` in an input block
+ `selected_options` decode), Discord (min/max values + `component` bounds
decode), Teams (`isMultiSelect`); Telegram/WhatsApp degrade to single-select.
Slack-only capabilities (Button.confirm, Image.title, Input label/initialValue/
required, Select.initialValue) were intentionally left out.
The prod harness-workers backfill (e88d01a) inverts the old
"harness-workers is staging-only" invariant. Update the 8 stale
assertions across 3 test files that still encoded staging-only,
deriving the new expected values from the SSOT (railway-envs.ts) and
the regenerated railway-envs.generated.json:
- healthcheckPathFor/emit healthcheckPath: prod now /health (was undefined/omitted)
- repoNameFor(prod): now resolves showcase-harness (was throw)
- envsFor: now [prod, staging] (was [staging])
- the worker-shape test: dual-env, domainless+probe-disabled in BOTH
envs, gateValidated:true / gateIgnore dropped (per SSOT)
- computePromoteClosure: harness-workers now Tier-1 promoted, not
skipped; the always-Tier-1 set no longer filters it out
- expandImageConsumers(prod) / default prod redeploy scope (39->40):
the dual-env worker now joins the prod showcase-harness redeploy scope
claude-sdk-python was the last integration still on the legacy auth-first
shape: an authenticated-on-load page guarded by a class-based
`ChatErrorBoundary`, a `useDemoAuth` exposing `authenticate`/`authenticated`,
an `auth-banner` with an `onAuthenticate` prop and bespoke buttons, and NO
`sign-in-card`. The byte-identical `auth.spec.ts` (which asserts an
unauthenticated-first `SignInCard` with `auth-sign-in-button` /
`auth-demo-token`) therefore failed all six cases against it.
Port the four auth files verbatim from the langgraph-python gold standard
(adapting nothing — the per-integration wiring, `agent="auth-demo"` and
`runtimeUrl="/api/copilotkit-auth"`, was already identical):
- use-demo-auth.ts: unauth-first, localStorage-backed, exposes
`isAuthenticated`/`hasEverSignedIn`/`signIn`/`signOut`.
- page.tsx: render `SignInCard` until first sign-in, then keep `<CopilotKit>`
mounted across the sign-out cycle; shared `handleAuthError` on BOTH the
provider and agent-scoped `<CopilotChat onError>`; clear-on-auth effect;
amber `auth-demo-error` surface.
- auth-banner.tsx: shared `<Button>`, `onSignIn`/`onSignOut` props.
- sign-in-card.tsx: new, ported from the gold standard.
Add the shared shadcn primitives the gold-standard frontend depends on and
which claude-sdk-python was missing (`src/lib/utils.ts`,
`src/components/ui/button.tsx`, `src/components/ui/card.tsx`) plus the
`radix-ui` dependency they require, matching the claude-sdk-typescript peer.
Red/green on the real surfaces: against the legacy frontend `auth.spec.ts`
fails 6/6 (every test times out waiting for `auth-sign-in-button`); against
the rebuilt frontend it passes 6/6 and the `--d5 --isolate` auth probe is
green.