mirror of
https://github.com/CopilotKit/CopilotKit.git
synced 2026-09-14 16:26:20 +08:00
codex/cloudplot-showcase-migration
15424 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
7fd4c5ee78 |
fix(showcase/harness): reap chromium zombies, alarm on PID saturation, stop saturated workers claiming (#6333)
> **This PR carries three changes to the same prod incident.** Change 1
fixes
> the PID leak itself; changes 2 and 3 fix the two reasons the leak was
able to
> starve the fleet unannounced. Each has its own RED → GREEN below.
>
> | # | Commit | What it fixes |
> |---|---|---|
> | 1 | `e66fb1af13` | **The leak.** No PID-1 reaper, so crashed
chromium grandchildren become permanent zombies holding cgroup PID
slots. |
> | 2 | `7f28e9bd39` | **Nothing alarmed.** The worker heartbeats its
cgroup PID gauges and no non-test code read them. |
> | 3 | `63ef4081c2` | **Saturated workers kept claiming.** The claim
gate only checked free browser-context slots, so a worker at 1000/1000
PIDs won jobs it could not run and burned each to a 600000ms abort. |
---
# CHANGE 1 — the leak (tini as PID 1)
## The bug
`harness-workers` has **no PID-1 zombie reaper**, so it leaks cgroup PID
slots until it cannot launch a browser.
libuv's `SIGCHLD` handler only `waitpid()`s the pids **Node itself
spawned**. Playwright spawns the chromium *browser* process (reaped
fine), but that browser's **zygote/renderer/GPU children are
grandchildren**. When a browser process dies — prod's `Target crashed`
path, or a pool self-heal relaunch — those grandchildren are re-parented
onto PID 1 and exit. Node-as-PID-1 never waits on them, so each becomes
a permanent `<defunct>` still holding a cgroup PID slot.
Prod, on a single container up 6d22h: `zombieCount` 70 → **757**,
`cgroupPidsCurrent` 286 → **1000/1000**. Once saturated, probes could
not launch browsers (`Target crashed`, `browserContext.newPage: … has
been closed`, `feature exceeded 300000ms wall-clock`) and **295 of 414
red prod d6 cells went `abort`, fleet-wide across ~20 slugs.**
This Dockerfile already documented the outage mode in a 40-line `FIX #3`
block — but **both of its stated mitigations are demand-side**
(`BROWSER_POOL_MAX_CONTEXTS` 40→24, resource-gauge warning logs). Those
cap the *peak* of `pids.current`. They cannot stop a *monotonic* leak,
and prod died of a monotonic leak. `resource-gauges.ts` only ever
**measured** zombies (`if (s.state === "Z") zombieCount += 1`); nothing
reaped them.
## The change
One file, +59/-3. `tini` as PID 1.
```dockerfile
RUN node ./node_modules/playwright/cli.js install --with-deps chromium \
&& apt-get install -y --no-install-recommends tini \
&& /usr/bin/tini --version \
&& rm -rf /var/lib/apt/lists/*
…
ENTRYPOINT ["/usr/bin/tini", "--"]
CMD ["/bin/bash", "-c", "ulimit -u $(ulimit -Hu) 2>/dev/null || true; exec node dist/orchestrator.js"]
```
Why this form:
- **`tini` baked into the image, not Docker `--init`** — `--init` is a
*runtime* flag and Railway does not expose it, the same reason
`pids.max` isn't settable from here (already documented in this file).
- **Not an in-process reaper** — Node has no `waitpid` binding, so an
application `SIGCHLD` handler would need a native addon. This is the
platform's job.
- **Rides the existing apt layer** — playwright's `--with-deps` just ran
`apt-get update`, so this costs one ~267KB package download instead of a
second index refresh. The `tini --version` call is a build-time
assertion on the path baked into `ENTRYPOINT`: if a future Debian moves
the binary, the **build** fails rather than the container failing to
start in prod.
- **No `-g`** — tini's process-group broadcast would signal the live
chromium pool alongside node and pre-empt `orchestrator.ts`'s
`process.once("SIGTERM", drainAndExit)`. Without it, delivery is
byte-identical to today: exactly one process gets the signal.
- **Exec form, `exec` kept in CMD** — tini really is PID 1 and really
receives Railway's SIGTERM; node is tini's direct child with no bash
lingering in the tree.
No existing convention to match: `grep` for `tini`/`dumb-init`/`--init`
across every Dockerfile in the repo returns zero hits. Nothing in the
repo asserts on this image's `CMD`/`ENTRYPOINT`, and nothing outside the
CI build step `docker run`s it, so adding an `ENTRYPOINT` breaks no
consumer.
---
# RED → GREEN
Real container, real chromium, real zombies. **Identical argv for
both**; only the image differs, and the image supplies PID 1:
```
docker run --rm --pids-limit 1000 -e REPRO_CYCLES=230 \
-v <repro>:/repro:ro <image> /bin/bash -c "exec node /repro/zombie-repro.mjs"
```
`--pids-limit 1000` mirrors Railway's platform-fixed ceiling. Each cycle
launches chromium with the exact args `browser-pool.ts` uses
(`headless`, `--no-sandbox`, `--disable-dev-shm-usage`), opens a
context+page, then **SIGKILLs the browser process** — prod's `Target
crashed` path, orphaning its children onto PID 1.
**Measurement is container-wide on purpose.** The harness's own
`sampleResourceGauges()` walks the tree from `process.pid`; under a real
init the orphans re-parent to the *init*, not to node, so a node-rooted
walk would report zero zombies **whether or not they were reaped** — a
vacuous GREEN. Counting every state-`Z` process in `/proc` plus the
cgroup counter is blind to which process is PID 1, so it is honest in
both directions.
## RED (unmodified `main`, `pid1=node`)
20 cycles — exactly **+5 permanent zombies per crashed browser**:
```
baseline zombieCount=0 procCount=1 threadCount=7 cgroupPidsCurrent=7 cgroupPidsMax=1000 selfPid=1 pid1=node
PID PPID S THR COMMAND
1 0 R 7 node
cycle=1 pre-kill zombieCount=0 procCount=7 threadCount=76 cgroupPidsCurrent=76 cgroupPidsMax=1000 selfPid=1 pid1=node
cycle=1 zombieCount=5 procCount=6 threadCount=16 cgroupPidsCurrent=16 cgroupPidsMax=1000 selfPid=1 pid1=node
cycle=2 zombieCount=10 procCount=11 threadCount=21 cgroupPidsCurrent=21 cgroupPidsMax=1000 selfPid=1 pid1=node
cycle=3 zombieCount=15 procCount=16 threadCount=26 cgroupPidsCurrent=26 cgroupPidsMax=1000 selfPid=1 pid1=node
…
cycle=20 zombieCount=100 procCount=101 threadCount=111 cgroupPidsCurrent=111 cgroupPidsMax=1000 selfPid=1 pid1=node
final zombieCount=100 procCount=101 threadCount=111 cgroupPidsCurrent=111 cgroupPidsMax=1000 selfPid=1 pid1=node
SUMMARY cycles=20 zombiesBefore=0 zombiesAfter=100 pidsBefore=7 pidsAfter=111 launchFailure=none
```
The leaked processes, verbatim from the final `/proc` table — all
`ppid=1`, all `Z`:
```
21 1 Z 1 headless_shell <defunct>
22 1 Z 1 headless_shell <defunct>
36 1 Z 1 headless_shell <defunct>
61 1 Z 1 headless_shell <defunct>
77 1 Z 1 headless_shell <defunct>
87 1 Z 1 headless_shell <defunct>
88 1 Z 1 headless_shell <defunct>
104 1 Z 1 headless_shell <defunct>
```
`cgroupPidsCurrent` after each cycle never returns to 7 — it is `7 +
5×cycles`. A monotonic ratchet, the same curve prod walked over 7 days.
### RED, driven to saturation (started before the Dockerfile was
touched)
```
baseline zombieCount=0 cgroupPidsCurrent=7 /1000
cycle=25 zombieCount=125 cgroupPidsCurrent=136 /1000
cycle=50 zombieCount=250 cgroupPidsCurrent=261 /1000
cycle=75 zombieCount=375 cgroupPidsCurrent=386 /1000
cycle=100 zombieCount=500 cgroupPidsCurrent=511 /1000
cycle=125 zombieCount=625 cgroupPidsCurrent=636 /1000
cycle=150 zombieCount=750 cgroupPidsCurrent=761 /1000 ← prod's ~757 reproduced
cycle=175 zombieCount=875 cgroupPidsCurrent=886 /1000
cycle=184 pre-kill zombieCount=915 cgroupPidsCurrent=993 /1000
cycle=185 pre-kill zombieCount=920 cgroupPidsCurrent=996 /1000
cycle=185 zombieCount=925 cgroupPidsCurrent=936 /1000
← no further output. Cycle 186 never completed.
```
The run **wedged** at cycle 185 with no output for >12 minutes:
`chromium.launch()` could not get PIDs and never returned — it did not
even surface Playwright's 30s launch timeout, which is exactly why prod
shows 300s feature timeouts and 1200s run aborts instead of a clean
launch error. `docker ps`: `Up 17 minutes (unhealthy)`.
```
$ docker exec red-sat cat /sys/fs/cgroup/pids.current /sys/fs/cgroup/pids.max
1000
1000
```
And the errno-11 signature this Dockerfile's own comment names — the
container could no longer `fork()` at all:
```
$ docker exec red-sat /bin/bash -c 'for i in 1 2 3; do cat /proc/uptime; done'
/bin/bash: fork: retry: Resource temporarily unavailable
/bin/bash: fork: retry: Resource temporarily unavailable
/bin/bash: fork: retry: Resource temporarily unavailable
/bin/bash: fork: retry: Resource temporarily unavailable
/bin/bash: fork: Resource temporarily unavailable
```
That is the prod outage, end to end, in a local container.
## GREEN (this PR's image, `pid1=tini`) — same repro, same argv, 230
cycles
```
baseline zombieCount=0 procCount=2 threadCount=8 cgroupPidsCurrent=8 cgroupPidsMax=1000 selfPid=7 pid1=tini
PID PPID S THR COMMAND
1 0 S 1 tini
7 1 R 7 node
cycle=1 pre-kill zombieCount=0 procCount=8 threadCount=76 cgroupPidsCurrent=76 cgroupPidsMax=1000 selfPid=7 pid1=tini
cycle=1 zombieCount=0 procCount=2 threadCount=12 cgroupPidsCurrent=12 cgroupPidsMax=1000 selfPid=7 pid1=tini
cycle=50 zombieCount=0 procCount=2 threadCount=12 cgroupPidsCurrent=12 cgroupPidsMax=1000 selfPid=7 pid1=tini
cycle=100 zombieCount=0 procCount=2 threadCount=12 cgroupPidsCurrent=12 cgroupPidsMax=1000 selfPid=7 pid1=tini
cycle=150 zombieCount=0 procCount=2 threadCount=12 cgroupPidsCurrent=12 cgroupPidsMax=1000 selfPid=7 pid1=tini
cycle=200 zombieCount=0 procCount=2 threadCount=12 cgroupPidsCurrent=12 cgroupPidsMax=1000 selfPid=7 pid1=tini
cycle=230 pre-kill zombieCount=0 procCount=8 threadCount=77 cgroupPidsCurrent=77 cgroupPidsMax=1000 selfPid=7 pid1=tini
cycle=230 zombieCount=0 procCount=2 threadCount=12 cgroupPidsCurrent=12 cgroupPidsMax=1000 selfPid=7 pid1=tini
final zombieCount=0 procCount=2 threadCount=12 cgroupPidsCurrent=12 cgroupPidsMax=1000 selfPid=7 pid1=tini
SUMMARY cycles=230 zombiesBefore=0 zombiesAfter=0 pidsBefore=8 pidsAfter=12 launchFailure=none
```
`procCount` returning to **2** (tini + node) after every single cycle is
the reap actually happening: the browser's 6 processes appear pre-kill
(`procCount=8`) and are fully gone post-kill.
| | RED (`pid1=node`) | GREEN (`pid1=tini`) |
| --- | --- | --- |
| zombies after 20 cycles | 100 | **0** |
| zombies at cycle 185 | 925 | **0** |
| `pids.current` floor at cycle 185 | 936, peaking 996/1000 | **12** |
| reached cycle 230? | **no — wedged at 185** | yes |
| `chromium.launch()` | hangs; `fork: Resource temporarily unavailable`
| fine, `launchFailure=none` |
| leak per crashed browser | **+5** | **0** |
---
## Signals and non-zero exit still work
The prober's container must still self-restart on crash, so this was
proven, not assumed. `signal-child.mjs` installs the same
`process.once("SIGTERM", …)` shape `orchestrator.ts` uses; run on
**both** images:
```
--- GREEN (ENTRYPOINT tini, pid1=tini) ---
[showcase-harness:green MODE=term-graceful] exitCode=0
CHILD mode=term-graceful pid=7 ppid=1
CHILD ready — idling until signalled
CHILD received SIGTERM at pid=7 — draining ← tini FORWARDED it to node
[showcase-harness:green MODE=exit-42] exitCode=42 ← non-zero preserved
[showcase-harness:green MODE=abort] exitCode=134 ← 128+6 (SIGABRT), non-zero
--- RED (no ENTRYPOINT, pid1=node) — baseline ---
[showcase-harness:red MODE=term-graceful] exitCode=0
CHILD received SIGTERM at pid=1 — draining
[showcase-harness:red MODE=exit-42] exitCode=42
[showcase-harness:red MODE=abort] exitCode=133
```
- **SIGTERM reaches node, not just tini** — the handler fires at
`pid=7`, so `docker stop`'s SIGTERM was forwarded into the process that
owns the graceful drain. `orchestrator.ts`'s
`drainControlPlaneAndExit("SIGTERM")` path is intact.
- **Non-zero exits survive** — `exit(42)` → container `42`; a crash
still exits non-zero, so Railway still restarts. The HEALTHCHECK
restart-on-sustained-503 behaviour this Dockerfile calls "THE INTENDED
OUTCOME" is untouched.
- **tini is strictly *more* correct here**: note RED's abort → **133**
vs GREEN's → **134**. A process running as PID 1 in a namespace has
default signal actions *ignored* by the kernel, so node-as-PID-1 was
distorting its own fatal-signal exit code (128+SIGTRAP instead of
128+SIGABRT). Both non-zero, so no regression — but the fixed image
reports crashes accurately.
## Cross-arch check
Local images are arm64; prod builds amd64. `tini` availability was
verified on amd64 explicitly rather than assumed — same path the
build-time assertion checks:
```
Package: tini
Architecture: amd64
Version: 0.19.0-1+b3
-rwxr-xr-x 1 root root 27792 Jun 6 2025 /usr/bin/tini
tini version 0.19.0
```
## Pre-push checks
| Check | Result |
| --- | --- |
| `docker buildx build -f showcase/harness/Dockerfile` (== CI's build
check) | **PASS**, both images |
| `nx run @copilotkit/showcase-harness:typecheck` | **PASS** |
| `nx run @copilotkit/showcase-harness:test:ci` | 3677 passed, 18
skipped, 1 failed — **environmental**, see below |
| commitlint | **PASS** |
| lefthook `pre-commit` + `commit-msg` | **PASS** |
| `git status` | clean |
The single failure is `d0-gone-predicate.test.ts > resolves the real
generated registry.json…`, failing on `ENOENT …
showcase/shell/src/data/registry.json`. That file is **gitignored and
generated at build time** — this Dockerfile's own comment says so, which
is why the image generates it via `generate-registry.ts`. Proven
environmental rather than assumed: generated the file, re-ran that exact
spec, **11/11 passed**. A Dockerfile-only diff cannot affect a vitest
run.
No formatter/linter covers this file: lefthook's `lint-fix` glob is
`*.{js,jsx,ts,tsx,mjs,cjs,md,css,yml,yaml,html,vue,py}`, and there is no
hadolint anywhere in CI.
## Not in this PR
- **This does not fix prod by itself.** Prod is digest-pinned and still
running the leaking container; it needs the live mitigation restart
(owned by a separate worker — **prod was not touched here**) and then a
promote to pick this image up.
- **The ≥3-cell value test is not done and cannot be done from this PR**
— confirming red cells flip requires the prod restart plus a fleet
sweep. Note the ~20 `conversation-error` and ~16 `goto-error` cells are
almost certainly unrelated real bugs this will **not** fix; only the
`abort`/`feature-timeout`/`driver-error` population is in scope.
- **Alerting / scheduled recycle is deliberately out of scope** (one
concern per change). Still worth doing: the documented
`pool-unrecoverable` alarm demonstrably never fired at 1000/1000, and
`cgroupPidsCurrent/cgroupPidsMax > 0.75` would have given ~36h of
warning before the flip.
- Leak rate on the **graceful** `browser.close()` path was not measured
— irrelevant to the fix (tini reaps either way), but it means prod's
exact 757 can't be attributed to crashes alone.
---
# RE-VERIFICATION 2026-08-12 — rebased onto `origin/main`, red-green
re-run from scratch
The proof above was produced 2026-08-03 against a base that predates
`ced993447f` ("copy patches/ into harness build context"), which touches
THIS
Dockerfile. The branch has been rebased onto current `origin/main` and
the
whole red-green was re-run on freshly built images so the numbers
describe the
actual merge candidate.
**Both images built locally from this repo, same machine, same context,
~3 min apart:**
```
showcase-harness:red-main sha256:b86686e9a06d51f884d5e59e9d15d667fd298284efe7a15622ff3bf2575a3ffd created=2026-08-12T20:15:50Z (origin/main Dockerfile, /usr/bin/tini ABSENT, Entrypoint=["docker-entrypoint.sh"])
showcase-harness:green-tini sha256:d666b823f16bda07bf62d35f1d72ef099d13406a61e2f0cd51c99bf51eb14df4 created=2026-08-12T20:18:54Z (this branch, /usr/bin/tini present, Entrypoint=["/usr/bin/tini","--"])
```
Identical driver, identical argv, 40 cycles each, runs 18 seconds apart:
```
docker run --rm --pids-limit 1000 -e REPRO_CYCLES=40 -v <repro>:/repro:ro <image> \
/bin/bash -c "exec node /repro/zombie-repro.mjs"
```
## RED — `showcase-harness:red-main`, `pid1=node` (2026-08-12T20:19:28Z)
```
baseline zombieCount=0 procCount=1 threadCount=7 cgroupPidsCurrent=7 cgroupPidsMax=1000 selfPid=1 pid1=node
cycle=1 pre-kill zombieCount=0 procCount=7 threadCount=76 cgroupPidsCurrent=76 cgroupPidsMax=1000 selfPid=1 pid1=node
cycle=1 zombieCount=5 procCount=6 threadCount=16 cgroupPidsCurrent=16 cgroupPidsMax=1000 selfPid=1 pid1=node
cycle=2 pre-kill zombieCount=5 procCount=12 threadCount=83 cgroupPidsCurrent=83 cgroupPidsMax=1000 selfPid=1 pid1=node
cycle=2 zombieCount=10 procCount=11 threadCount=21 cgroupPidsCurrent=21 cgroupPidsMax=1000 selfPid=1 pid1=node
cycle=10 pre-kill zombieCount=45 procCount=52 threadCount=120 cgroupPidsCurrent=120 cgroupPidsMax=1000 selfPid=1 pid1=node
cycle=10 zombieCount=50 procCount=51 threadCount=61 cgroupPidsCurrent=61 cgroupPidsMax=1000 selfPid=1 pid1=node
cycle=20 pre-kill zombieCount=95 procCount=102 threadCount=174 cgroupPidsCurrent=174 cgroupPidsMax=1000 selfPid=1 pid1=node
cycle=20 zombieCount=100 procCount=101 threadCount=111 cgroupPidsCurrent=111 cgroupPidsMax=1000 selfPid=1 pid1=node
cycle=30 pre-kill zombieCount=145 procCount=152 threadCount=219 cgroupPidsCurrent=219 cgroupPidsMax=1000 selfPid=1 pid1=node
cycle=30 zombieCount=150 procCount=151 threadCount=161 cgroupPidsCurrent=161 cgroupPidsMax=1000 selfPid=1 pid1=node
cycle=40 pre-kill zombieCount=195 procCount=202 threadCount=271 cgroupPidsCurrent=271 cgroupPidsMax=1000 selfPid=1 pid1=node
cycle=40 zombieCount=200 procCount=201 threadCount=211 cgroupPidsCurrent=211 cgroupPidsMax=1000 selfPid=1 pid1=node
final zombieCount=200 procCount=201 threadCount=211 cgroupPidsCurrent=211 cgroupPidsMax=1000 selfPid=1 pid1=node
SUMMARY cycles=40 zombiesBefore=0 zombiesAfter=200 pidsBefore=7 pidsAfter=211 launchFailure=none
```
All 200 leaked processes are `ppid=1`, state `Z` (tail of the final
`/proc` table):
```
2610 1 Z 1 headless_shell <defunct>
2612 1 Z 1 headless_shell <defunct>
2646 1 Z 1 headless_shell <defunct>
2661 1 Z 1 headless_shell <defunct>
2662 1 Z 1 headless_shell <defunct>
2678 1 Z 1 headless_shell <defunct>
2680 1 Z 1 headless_shell <defunct>
2717 1 Z 1 headless_shell <defunct>
```
## GREEN — `showcase-harness:green-tini`, `pid1=tini`
(2026-08-12T20:19:46Z)
```
baseline zombieCount=0 procCount=2 threadCount=8 cgroupPidsCurrent=8 cgroupPidsMax=1000 selfPid=7 pid1=tini
cycle=1 pre-kill zombieCount=0 procCount=8 threadCount=76 cgroupPidsCurrent=76 cgroupPidsMax=1000 selfPid=7 pid1=tini
cycle=1 zombieCount=0 procCount=2 threadCount=12 cgroupPidsCurrent=12 cgroupPidsMax=1000 selfPid=7 pid1=tini
cycle=2 pre-kill zombieCount=0 procCount=8 threadCount=78 cgroupPidsCurrent=78 cgroupPidsMax=1000 selfPid=7 pid1=tini
cycle=2 zombieCount=0 procCount=2 threadCount=12 cgroupPidsCurrent=12 cgroupPidsMax=1000 selfPid=7 pid1=tini
cycle=10 pre-kill zombieCount=0 procCount=8 threadCount=79 cgroupPidsCurrent=79 cgroupPidsMax=1000 selfPid=7 pid1=tini
cycle=10 zombieCount=0 procCount=2 threadCount=12 cgroupPidsCurrent=12 cgroupPidsMax=1000 selfPid=7 pid1=tini
cycle=20 pre-kill zombieCount=0 procCount=8 threadCount=76 cgroupPidsCurrent=76 cgroupPidsMax=1000 selfPid=7 pid1=tini
cycle=20 zombieCount=0 procCount=2 threadCount=12 cgroupPidsCurrent=12 cgroupPidsMax=1000 selfPid=7 pid1=tini
cycle=30 pre-kill zombieCount=0 procCount=8 threadCount=77 cgroupPidsCurrent=77 cgroupPidsMax=1000 selfPid=7 pid1=tini
cycle=30 zombieCount=0 procCount=2 threadCount=12 cgroupPidsCurrent=12 cgroupPidsMax=1000 selfPid=7 pid1=tini
cycle=40 pre-kill zombieCount=0 procCount=8 threadCount=79 cgroupPidsCurrent=79 cgroupPidsMax=1000 selfPid=7 pid1=tini
cycle=40 zombieCount=0 procCount=2 threadCount=12 cgroupPidsCurrent=12 cgroupPidsMax=1000 selfPid=7 pid1=tini
final zombieCount=0 procCount=2 threadCount=12 cgroupPidsCurrent=12 cgroupPidsMax=1000 selfPid=7 pid1=tini
SUMMARY cycles=40 zombiesBefore=0 zombiesAfter=0 pidsBefore=8 pidsAfter=12 launchFailure=none
```
`grep -c defunct` over the GREEN run: **0**. `procCount` returns to
**2**
(tini + node) after every cycle — the browser's 6 processes appear
pre-kill
(`procCount=8`, `threadCount` 76–79) and are fully gone post-kill.
| | RED (`pid1=node`) | GREEN (`pid1=tini`) |
| --- | --- | --- |
| image id | `b86686e9a06d` | `d666b823f16b` |
| cycles | 40 | 40 |
| zombies after | **200** | **0** |
| zombies per crashed browser | **+5** | **0** |
| `pids.current` before → after | 7 → **211** | 8 → **12** |
| `<defunct>` rows in final table | **200** | **0** |
| `launchFailure` | none | none |
## Signals / non-zero exit, both images, this run
```
[showcase-harness:RED MODE=term-graceful] exitCode=0
CHILD mode=term-graceful pid=1 ppid=0
CHILD received SIGTERM at pid=1 — draining
[showcase-harness:RED MODE=exit-42] exitCode=42
[showcase-harness:RED MODE=abort] exitCode=133
RED /proc/1/comm = bash
[showcase-harness:GREEN MODE=term-graceful] exitCode=0
CHILD mode=term-graceful pid=7 ppid=1
CHILD received SIGTERM at pid=7 — draining <- tini FORWARDED it to node
[showcase-harness:GREEN MODE=exit-42] exitCode=42
[showcase-harness:GREEN MODE=abort] exitCode=134
GREEN /proc/1/comm = tini, child self=7
```
SIGTERM reaches node (handler fires at `pid=7`), non-zero exits survive
(42 → 42; abort → 134 = 128+6), so Railway's restart-on-crash is intact.
## The prod leak, measured live the same day
`resource_snapshots` in prod PocketBase, immediately before the
mitigation
restart on the unfixed digest `sha256:80a0c5c2…`:
```
2026-08-12 20:10:00.674Z worker-7abf6 heartbeat pids=1000/1000 zombies=758 procs=778 threads=1000
2026-08-12 20:09:59.783Z worker-5a213 heartbeat pids=1000/1000 zombies=761 procs=781 threads=1000
2026-08-12 20:09:54.868Z worker-de14c heartbeat pids=837/1000 zombies=744 procs=753 threads=837
```
and immediately after it (`init` rows): `pids=164–169/1000 zombies=0
procs=16`.
A new zombie was already recorded at `2026-08-12 20:12:27.182Z
worker-de14c … zombies=1`.
## Cross-arch
Local images are arm64 (`/usr/bin/tini` 68592 bytes); prod builds amd64.
Measured on amd64 directly: `Package: tini / Architecture: amd64 /
Version: 0.19.0-1+b3`, `/usr/bin/tini` 27792 bytes, `tini version
0.19.0`;
package download 267,204 bytes, 749 kB on disk.
## Deploy note
`computePromoteClosure(["showcase-harness"])` = `pocketbase(0),
harness(1),
dashboard(1), harness-workers(1)` — the workers are in the promote
scope.
Prod `startCommand` is `null` and no repo consumer overrides the
entrypoint,
so Railway uses the image's ENTRYPOINT+CMD. **PR #5750 is CLOSED
unmerged**, so
promoting `harness-workers` resets
`multiRegionConfig.us-west2.numReplicas` to
1 and it must be re-asserted to 6 and read back after the promote.
---
# CHANGE 2 — nothing alarmed (worker PID-saturation alarm)
Commit `7f28e9bd39`.
## The gap
The measured prod condition on 2026-08-12 20:09–20:10Z was
`pids=1000/1000`,
`1000/1000`, `837/1000` across three workers, 320 `timeout after
600000ms`
abort rows, an `e2e-smoke` run 42 minutes into a 15-minute cadence with
3 jobs
still pending, prod d6 534/784 vs staging 744/788. **No alarm fired**,
and this
was the second occurrence (2026-08-03, 2026-08-10).
The documented `pool-unrecoverable` alarm did not fire *by
construction*: it is
raised only when `BrowserPool`'s self-heal circuit-breaker exhausts
`selfHealMaxHardRecoveries`. A worker that is PID-starved but still
heartbeating
never reaches that state — it launches browsers fine, it just cannot
fork the
process tree a probe needs.
The signal was already on the wire and read by nobody:
- `worker/registration.ts:192` writes `capacity_pids_current` /
`capacity_pids_max` onto the worker's `workers` row on its ~75s
heartbeat
(`DEFAULT_WORKER_HEARTBEAT_MS = 75_000`).
- `control-plane/fleet-health.ts` lists that entire roster every 15s
(`DEFAULT_FLEET_HEALTH_INTERVAL_MS`) — and its own comment on the row
type
said *"the capacity gauges are ignored here."*
- Nothing in non-test source compared `pidsCurrent` to `pidsMax`.
## Where the alarm goes, and why
**Control-plane, not worker.** `writers/status-writer.ts` documents
`fleet-cp`
as *"the only authoritative fleet writer; workers never write status
directly"*,
so a worker cannot raise a `system:` health key. `fleet-health` already
holds
the roster read this needs, on a 15s cadence, so the alarm costs zero
extra PB
traffic.
**Routed to `#oss-alerts`, not to the browser-pool webhook.**
`SLACK_WEBHOOK_BROWSER_POOL_UNRECOVERABLE` appears **nowhere in this
repo
outside its own definition** — `grep` finds it only in `orchestrator.ts`
— so it
is unset everywhere and that alarm has no receiver even when it does
fire.
`SLACK_WEBHOOK_OSS_ALERTS` is a real wired secret that the
family-silence
monitor and the D0-gone monitor already post through. Routing a new
alarm to a
known-unwired webhook would reproduce the exact gap this closes.
## Thresholds
- **Alarm at `pids.current / pids.max >= 0.75`.** At the observed leak
rate
(~1000 PIDs over ~7 days) that is roughly 36 hours of lead time — more
than a
working day. A busy worker's context budget is nowhere near 750 PIDs.
- **Latched per worker, clearing at 0.65.** The monitor ticks every 15s;
an
unlatched alarm would page ~8600 times across that 36-hour window. An
alarm
that fires constantly is worse than none, so `createFleetHealthMonitor`
now
**fails loud at construction** on a ratio pair with no hysteresis gap.
- **Evaluated before the online/stale branch**, because the failing
worker is
fully `online` — it heartbeats on time throughout.
- **Unmeasured gauges never alarm.** `registration.ts` maps the pool's
`-1`
sentinel to `null`; `pidUsageRatio` returns `null` for
null/absent/negative
current or non-positive max, and null is never read as zero.
---
## RED → GREEN (change 2)
Driver exercises the **real chain end-to-end**, faking only the OS
cgroup read
(`BrowserPoolOptions.cgroupPidsReader`, an existing injectable seam) and
the PB
transport. The column names are whatever the real registration writer
emits and
whatever the real fleet-health reader consumes, so a column drift fails
this
driver:
```
BrowserPool.budget() [real, probes/helpers/browser-pool.ts]
-> workerCapacityFromBudget [real, fleet/contracts.ts]
-> registerWorker() [real, fleet/worker/registration.ts]
-> workers row
-> createFleetHealthMonitor().checkOnce()
[real, fleet/control-plane/fleet-health.ts]
-> alarm sink
```
Identical command both sides; only the source differs:
```
./node_modules/.bin/tsx <scratch>/drive-alarm.mts
```
### RED — unmodified source
```
=== A1 SATURATED 900/1000 (ratio 0.90) ===
pool.budget() pids : 900/1000
workers row gauges : capacity_pids_current=900 capacity_pids_max=1000
fleet-health cycles run : 1
ALARM HOOK CALLS : 0
result.pidSaturated : <field does not exist>
error-level log events : <none>
=== A2 CEILING 1000/1000 (ratio 1.00) ===
pool.budget() pids : 1000/1000
workers row gauges : capacity_pids_current=1000 capacity_pids_max=1000
fleet-health cycles run : 1
ALARM HOOK CALLS : 0
result.pidSaturated : <field does not exist>
error-level log events : <none>
=== A6 LATCH 900/1000 over 5 cycles ===
fleet-health cycles run : 5
ALARM HOOK CALLS : 0
result.pidSaturated : <field does not exist>
error-level log events : <none>
```
The exact prod condition — a worker at the ceiling — produces **zero**
alarms.
### GREEN — with this change
```
=== A1 SATURATED 900/1000 (ratio 0.90) ===
pool.budget() pids : 900/1000
workers row gauges : capacity_pids_current=900 capacity_pids_max=1000
fleet-health cycles run : 1
ALARM HOOK CALLS : 1
result.pidSaturated : 1
error-level log events : ["fleet.health.worker-pid-saturated"]
ALARM PAYLOAD : {"workerId":"worker-drv","pidsCurrent":900,"pidsMax":1000,"ratio":0.9,"threshold":0.75,"lastHeartbeatAt":"2026-08-12T20:44:54.244Z","observedAt":"2026-08-12T20:44:54.244Z"}
=== A2 CEILING 1000/1000 (ratio 1.00) ===
ALARM HOOK CALLS : 1
result.pidSaturated : 1
error-level log events : ["fleet.health.worker-pid-saturated"]
ALARM PAYLOAD : {"workerId":"worker-drv","pidsCurrent":1000,"pidsMax":1000,"ratio":1,"threshold":0.75,"lastHeartbeatAt":"2026-08-12T20:44:54.253Z","observedAt":"2026-08-12T20:44:54.253Z"}
```
### The alarm must also stay QUIET — same run, same command
```
=== A3 QUIET 500/1000 (ratio 0.50) ===
ALARM HOOK CALLS : 0
result.pidSaturated : 0
error-level log events : <none>
=== A4 QUIET 740/1000 (ratio 0.74, just under) ===
ALARM HOOK CALLS : 0
result.pidSaturated : 0
error-level log events : <none>
=== A5 QUIET -1/-1 (cgroup unreadable) ===
pool.budget() pids : -1/-1
workers row gauges : capacity_pids_current=null capacity_pids_max=null
ALARM HOOK CALLS : 0
result.pidSaturated : 0
error-level log events : <none>
=== A6 LATCH 900/1000 over 5 cycles ===
fleet-health cycles run : 5
ALARM HOOK CALLS : 1 <- once, not five times
```
`A4` pins the boundary at 0.74. `A5` pins the off-Linux case: the `-1`
sentinel
lands as `null` on the row and stays silent rather than reading as `0`.
`A6`
pins the latch — 5 saturated cycles, 1 alarm.
---
# CHANGE 3 — saturated workers kept claiming (PID headroom claim gate)
Commit `63ef4081c2`.
## The defect
`fleet/worker/worker-loop.ts` had exactly one conditional in the claim
loop:
```ts
if (budget.available <= 0) { … idle … }
```
and `available` is free Playwright **browser-context** slots and nothing
else —
`browser-pool.ts`: `available: Math.max(0, this.maxContexts -
this.liveContextCount)`.
Context slots have no relationship to cgroup PIDs. A worker whose
container has
leaked to `pids=1000/1000` therefore still advertises **full** capacity,
keeps
winning claims, cannot fork the browser tree, and burns each job's
entire
600000ms lease before aborting. That is the mechanism behind the 320
abort rows.
## Blast radius — what happens when ALL workers are saturated
The gate deliberately reuses the **same decline-and-idle path** as the
existing
no-budget branch, which is what makes total saturation safe:
- the job is **never claimed**, so it is never dropped and never
requeued — it
stays `pending` for a healthy worker, or for this one after a redeploy;
- the loop then sleeps a full `pollIntervalMs`, so a fully-saturated
fleet
**idles at the poll cadence** instead of spinning;
- the queue **stalls visibly**: `probe_jobs` rows pile up pending, a
rising-edge
`fleet.worker.pid-headroom-exhausted` warn says why (warn → stderr →
Sentry;
steady state drops to debug so a multi-hour stall does not flood), and
change 2's `system:worker-pid-saturation` alarm fired at 0.75 **before**
this
gate engaged at 0.90.
A stalled, alarmed queue is recoverable by redeploy. Claim-fail-repeat
silently
burned the whole cadence.
## Why the gate is 0.90 and the alarm is 0.75
They do different jobs and must not be equal. 0.75 pages an operator
while the
worker is still perfectly able to run jobs; gating dispatch that early
would
convert a warning into ~36 hours of withheld fleet capacity. 0.90 is
where we
stop trusting the worker to fork at all — 100 free PIDs at the prod
ceiling.
**That reserve is chosen, not derived: the per-job PID cost of a
chromium
context tree is not measured**, hence the `WORKER_PID_CLAIM_GATE_RATIO`
override. The invariant that matters is ordering — the alarm always
fires before
capacity is withdrawn.
The gate is **inert** when the cgroup gauges are unreadable
(`pidUsageRatio` →
`null`: off-Linux, every macOS dev box), so local workers claim exactly
as
before.
---
## RED → GREEN (change 3)
Driver runs the **real `startWorkerLoop`** against the **real
`BrowserPool`**
budget path, faking only the cgroup read and the queue transport. The
queue fake
holds a real `pending` list, so *"was the job dropped / does it stay
pending"* is
directly observed, not inferred. `budget()` calls are counted as a
loop-iteration proxy for the spin check.
Identical command both sides:
```
./node_modules/.bin/tsx <scratch>/drive-gate.mts
```
### RED — unmodified source
```
=== B1 SATURATED 1 worker @ 1000/1000, 3 jobs pending ===
workers=1 cgroup pids=1000/1000 contexts available=24
window : 600ms @ poll 50ms
queue.claimNext CALLS : 15
jobs CLAIMED : 3 ["job-1","job-2","job-3"]
jobs STILL PENDING : 0 []
jobs REPORTED : 3
loop iterations : 15
=== B2 ALL-SATURATED 3 workers @ 1000/1000, 3 jobs pending ===
workers=3 cgroup pids=1000/1000 contexts available=24
queue.claimNext CALLS : 39
jobs CLAIMED : 3 ["job-1","job-2","job-3"]
jobs STILL PENDING : 0 []
jobs REPORTED : 3
loop iterations : 39
```
A worker at the PID ceiling claims **all three jobs**. In prod each of
those
becomes a 600000ms abort.
### GREEN — with this change
```
=== B1 SATURATED 1 worker @ 1000/1000, 3 jobs pending ===
workers=1 cgroup pids=1000/1000 contexts available=24
window : 600ms @ poll 50ms
queue.claimNext CALLS : 0
jobs CLAIMED : 0 []
jobs STILL PENDING : 3 ["job-1","job-2","job-3"]
jobs REPORTED : 0
loop iterations : 12 (spin check: bounded by window/poll)
```
`claimNext` is never called; all three jobs are **still pending** — not
dropped,
not requeued.
### ALL workers saturated — the blast-radius case
```
=== B2 ALL-SATURATED 3 workers @ 1000/1000, 3 jobs pending ===
workers=3 cgroup pids=1000/1000 contexts available=24
window : 600ms @ poll 50ms
queue.claimNext CALLS : 0
jobs CLAIMED : 0 []
jobs STILL PENDING : 3 ["job-1","job-2","job-3"]
jobs REPORTED : 0
loop iterations : 36 (spin check: bounded by window/poll)
```
Zero claims, **3 of 3 jobs still pending**, and 36 iterations across 3
workers
over a 600ms window at a 50ms poll — i.e. 12 per worker, the poll
cadence. No
spin, no loss.
### The gate must NOT engage below threshold — same run, same command
```
=== B3 HEALTHY 1 worker @ 100/1000, 3 jobs pending ===
queue.claimNext CALLS : 15
jobs CLAIMED : 3 jobs STILL PENDING : 0 jobs REPORTED : 3
=== B4 UNDER-GATE 1 worker @ 890/1000, 3 jobs pending ===
queue.claimNext CALLS : 15
jobs CLAIMED : 3 jobs STILL PENDING : 0 jobs REPORTED : 3
=== B5 UNMEASURED 1 worker @ -1/-1, 3 jobs pending ===
queue.claimNext CALLS : 15
jobs CLAIMED : 3 jobs STILL PENDING : 0 jobs REPORTED : 3
```
`B4` is the load-bearing one: 0.89 is **above** change 2's 0.75 alarm
and below
this gate. The operator has been paged and the worker keeps working —
which is
the whole reason the two thresholds differ. `B5` pins the off-Linux
case.
---
# Tests and gates (changes 2 and 3)
`showcase/harness`, on the pushed head `63ef4081c2`:
| Gate | Command | Result |
|---|---|---|
| Typecheck | `npx tsc --noEmit -p tsconfig.json` | exit 0 |
| Lint | `npx oxlint src/fleet src/orchestrator.ts` | 0 errors; warning
count unchanged vs base (2 on `fleet-health.ts` before and after) |
| Format | `npx oxfmt --write` | clean |
| Unit suite | `vitest run` | **3696 passed, 18 skipped, 1 failed** —
175/178 files pass |
The single failure is `src/probes/frontend-matrix.test.ts` —
**pre-existing**.
Verified by stashing this branch's changes and re-running it on the
unmodified
tree, where it fails identically (`1 failed | 5 passed`).
14 tests are new (`3682 → 3696`).
**Mutation-tested for vacuity.** With the four source files reverted to
`e66fb1af13` and the new test files kept, **11 of the new tests fail**:
```
Test Files 2 failed (2)
Tests 11 failed | 80 passed (91)
FAIL fleet-health.test.ts > PID saturation > alarms on an ONLINE worker whose PID gauges cross the threshold
FAIL fleet-health.test.ts > PID saturation > stays silent below the threshold, including just under it
FAIL fleet-health.test.ts > PID saturation > stays silent when the gauges are UNMEASURED (null / unbounded max)
FAIL fleet-health.test.ts > PID saturation > is EDGE-triggered: a sustained saturation alarms once, not every cycle
FAIL fleet-health.test.ts > PID saturation > re-arms only after the worker drops below the CLEAR ratio
FAIL fleet-health.test.ts > PID saturation > never lets a throwing alarm hook abort the cycle
FAIL fleet-health.test.ts > PID saturation > fails loud on a ratio config that would alarm every cycle
FAIL worker-loop.test.ts > PID headroom gate > declines to claim at the PID ceiling — the job is never claimed, so it stays pending
FAIL worker-loop.test.ts > PID headroom gate > declines at the gate ratio boundary (0.90)
FAIL worker-loop.test.ts > PID headroom gate > honours an injected gate ratio
FAIL fleet-health.test.ts > checkOnce > never throws when the roster read fails — returns an empty cycle
```
The three that pass on both sides are the sub-threshold controls (`still
claims
just BELOW the gate`, `claims normally on a healthy worker`, `stays
INERT when
the cgroup gauges are unreadable`) — they are supposed to be unchanged
by the
fix, and they are.
# New env knobs (all optional, all defaulted)
| Var | Default | Effect |
|---|---|---|
| `WORKER_PID_SATURATION_RATIO` | `0.75` | Control-plane alarm
threshold. |
| `WORKER_PID_SATURATION_CLEAR_RATIO` | `0.65` | Alarm latch hysteresis
clear. |
| `WORKER_PID_CLAIM_GATE_RATIO` | `0.90` | Worker claim-gate threshold.
|
An out-of-range override on the two alarm ratios falls back to the
default
rather than tripping the fail-loud guard and taking the control-plane
down over
a typo'd env var.
# Not verified
- Neither change 2 nor change 3 has been exercised against **live prod
or
staging PocketBase**; both red-greens are local, on the real modules
with the
cgroup read and PB/queue transport injected.
- **No Slack post was sent.** `SLACK_WEBHOOK_OSS_ALERTS` was not set in
the
driver, so the send leg was not exercised end-to-end — only that its
failure
is swallowed without costing the durable status row.
- The per-job PID cost of a chromium context tree is **not measured**,
so the
0.90 gate's 100-PID reserve is a chosen value.
|
||
|
|
528dea6483 |
fix(react-core): ship a single v2 context instance (#6440)
## Problem
`@copilotkit/react-core` ships **two independent copies** of the v2
context module, so `useLicenseContext` imported from
`@copilotkit/react-core/v2/context` returns the default forever —
`status: null` even when `/info` reports `licenseStatus: "valid"`.
Reported downstream as a chat-history sidebar that never loads, because
`useThreads` is gated on license status.
`src/v2/context.ts` is compiled by two separate tsdown builds:
| Build | Output | Contains |
|---|---|---|
| `entry: ["src/index.tsx", "src/v2/index.ts"]` | `dist/` shared chunk |
inlined copy **A** |
| `entry: {context: "src/v2/context.ts"}` | `dist/v2/context.*` |
standalone copy **B** |
There is no import edge between them, so `createContext()` runs twice.
`CopilotKitProvider` lives in the shared chunk and publishes to **A**;
`@copilotkit/react-core/v2/context` exports **B**, which nothing ever
provides.
Verified against the published 1.66.4 artifact:
```
$ grep -n "createContext" dist/v2/context.mjs
104:const CopilotKitContext = createContext(null);
124:const LicenseContext = createContext({
$ grep -n "createContext" dist/copilotkit-nRjRp2_5.mjs # inside //#region src/v2/context.ts
1522:const CopilotKitContext = createContext(null);
1544:const LicenseContext = createContext({
$ grep -E '^import .*from "[^"]*context[^"]*"' dist/copilotkit-nRjRp2_5.mjs
# (empty — no import edge)
```
`CopilotKitContext` is duplicated identically, so `useCopilotKit`
imported from that subpath throws `"useCopilotKit must be used within
CopilotKitProvider"`. The subpath was effectively unusable for web
consumers; license was just the *silent* failure mode.
**Compounding defect:** `src/v2/providers/index.ts` enumerates its
exports by name and omits `useLicenseContext` (even though
`CopilotKitProvider.tsx:19` re-exports it). So the live copy had **no
public import path at all**, leaving consumers with no correct
alternative.
Not a 1.66.x regression — broken since
|
||
|
|
63ef4081c2 |
fix(showcase/harness): stop a PID-saturated worker from claiming jobs
worker-loop's claim gate had exactly one conditional -- `if (budget.available <= 0)` -- and `available` is free Playwright browser-CONTEXT slots (maxContexts - liveContextCount), nothing else. A worker whose container has leaked PIDs to the cgroup ceiling therefore still advertises full capacity and keeps winning claims it cannot possibly run: the driver cannot fork, so each job burns its entire 600000ms lease and lands as an abort. Measured 2026-08-10: workers at pids=1000/1000 with 758-761 zombies, 320 `timeout after 600000ms` rows. Decline the claim above 0.90 of pids.max. Deliberately far above the control-plane's 0.75 saturation ALARM: the two thresholds do different jobs. 0.75 pages an operator while the worker is still perfectly able to run jobs; gating dispatch that early would convert a warning into ~36h of withheld fleet capacity. 0.90 leaves 100 free PIDs at the prod ceiling -- a chosen reserve, not a derived one (the per-job PID cost of a chromium context tree is not measured), hence the env override. The ordering that matters is that the alarm always fires before capacity is withdrawn. BLAST RADIUS -- reuses the SAME decline-and-idle path as the existing no-budget branch, which is what makes an all-workers-saturated fleet safe: the job is never claimed, so it is never dropped and never requeued; it simply stays `pending`. The loop then sleeps a full pollIntervalMs, so a fully-saturated fleet idles at the poll cadence instead of spinning. The queue stalls VISIBLY -- pending rows pile up, the rising-edge warn says why, and the control-plane's system:worker-pid-saturation alarm fired at 0.75 before this gate engaged at 0.90. A stalled, alarmed queue is recoverable by redeploy; claim-fail-repeat silently burned the whole cadence. The gate is INERT when the cgroup gauges are unreadable (pidUsageRatio -> null: off-Linux, every macOS dev box), so local workers claim exactly as before. |
||
|
|
7f28e9bd39 |
fix(showcase/harness): alarm when a worker's cgroup PIDs saturate
Prod harness-workers replicas reach the platform-fixed pids.max=1000
ceiling roughly every 7 days and NOTHING alarmed. Measured 2026-08-10
20:09-20:10Z: pids=1000/1000 zombies=758, 1000/1000 zombies=761,
837/1000 zombies=744, with 320 `timeout after 600000ms` abort rows and
an e2e-smoke run 42 minutes into a 15-minute cadence with 3 jobs still
pending. The documented pool-unrecoverable alarm did not fire: it only
trips when the self-heal breaker gives up, which a PID-starved but
still-heartbeating worker never reaches.
The signal was already on the wire and read by nobody. The worker's
~75s heartbeat writes capacity_pids_current / capacity_pids_max onto
its `workers` row (worker/registration.ts), and fleet-health lists that
whole roster every 15s — its own comment said "the capacity gauges are
ignored here". Nothing in non-test source compared them.
Alarm control-plane-side rather than worker-side, because status-writer
documents `fleet-cp` as the only authoritative fleet writer ("workers
never write status directly"), and because fleet-health already holds
the roster read this needs.
- fleet-health raises a rising-edge alarm at pids.current/pids.max >=
0.75 (~36h of lead time at the observed leak rate), evaluated BEFORE
the online/stale branch since the failing worker is fully ONLINE.
- Latched per worker with a 0.65 hysteresis clear: the monitor ticks
every 15s, so an unlatched alarm would page ~8600 times across the
lead-time window. Construction fails loud on a ratio pair with no gap.
- Unmeasured gauges (null off-Linux, unbounded pids.max) never alarm --
pidUsageRatio returns null and null is never read as zero.
- Routed to a system:worker-pid-saturation status row plus #oss-alerts
via SLACK_WEBHOOK_OSS_ALERTS, the target family-silence and the
D0-gone monitor already post to. Deliberately NOT
SLACK_WEBHOOK_BROWSER_POOL_UNRECOVERABLE, which appears nowhere in
this repo outside its own definition and is unset everywhere -- an
alarm nobody receives is the gap being closed.
|
||
|
|
e66fb1af13 |
fix(showcase/harness): reap orphaned chromium children with tini as PID 1
harness-workers leaked one cgroup PID slot per orphaned chromium grandchild. Node as PID 1 only waitpid()s processes it spawned, so every browser crash stranded ~5 <defunct> renderers permanently. Prod climbed to pids.current=1000/1000 with zombieCount=757 over 6d22h uptime, after which no browser could launch and ~295 d6 cells went abort fleet-wide. Install tini in the existing playwright apt layer and run it as PID 1 via exec-form ENTRYPOINT. No -g, so signal delivery to node is unchanged and orchestrator.ts's SIGTERM drain still runs; tini propagates the child's exit status so a crash still exits non-zero and Railway still restarts. |
||
|
|
be3485d8f3 | fix(ci): narrow AEO synthetic rollout | ||
|
|
a01b9f1034 | ci: monitor public AEO surfaces | ||
|
|
25ce46ba5f | fix(docs): simplify AEO surface contract | ||
|
|
d771e54993 | fix(docs): include shared contract in Turbopack root | ||
|
|
9f2fb8e43b | docs: define public AEO surface contract | ||
|
|
fb2aedb0e0 |
docs(reskinnable-demo): de-narrate the reskin skill and its guard comments
Second pass of the history sweep. The first cleaned CLAUDE.md and README.md; this finishes the reskin skill and the in-code comments that still recounted who hit a defect, when it was found, and how long it survived. Every rule, gate, command and checklist item is kept. What went is the narration around them — "it named only the first four skins for two releases", "caught by `eslint --print-config`, by hand, once", "drifted out of true three review rounds running", "it shipped that way once", "one CR pass found sixteen of them live", "measured in logistics", "each raised after the fact". Where a cut would have left a rule reading as arbitrary, the mechanism is restated in one present-tense clause instead: a hand-copied list rots silently and nothing fails when it is stale; flat-config `rules` are REPLACED, not merged, so a block silently drops every selector it does not restate; a schema leak is routinely line-wrapped, so a source-text guard never matches. Two stale cross-references fixed while in there: failure-modes.md quoted a CLAUDE.md sentence that the first pass removed, and claimed the roster-docs test header lists "two" known instances outside its doc set (it lists one). Skill impact, per the standing rule in CLAUDE.md: this change IS the skill, and it is prose-only — no contract field, link builder, lint rule, gate, beat mechanism, skin identity or file path changed, so no template or verification step needed a matching edit. The two doc properties `skin-roster-docs.test.ts` depends on were preserved deliberately: templates.md keeps "the six shipped skins" ahead of its brace glob, and SKILL.md keeps its "Six are registered —" id list, since both are what arm the brace-glob and valid-id-list rules. Gates: `pnpm lint`, `pnpm exec tsc --noEmit`, `pnpm test:unit` (197 files / 2227 tests) and `pnpm build` all green. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
6740613fb8 |
docs(reskinnable-demo): de-narrate in-code comments that recounted build history
Same cutting rule, applied only where a comment narrated what a past slot or
agent did rather than explaining the code: "a later slot owns that file", "as of
the beat-parity work", "two parallel agents each hand-edited this paragraph",
"this very paragraph did it once", "learned the hard way in the banking skin".
Every WHY stays; several are restated as present-tense properties (do not
reintroduce a client ticker, do not add a second seed of AV1423).
Four of these were also FACTUALLY STALE and are now correct:
- airline/attach-hotel-confirmation.ts claimed no pill carries
HOTEL_CONFIRMATION_MESSAGE; suggestions.ts has carried it since beat 3d landed.
- keel/attach-bulletin.ts said the same of BULLETIN_MESSAGE.
- keel/tools-replay-safety.test.ts said keel was not yet in the
statusKeyedTerminalRender glob; it is.
- airline/data/{store,trip-types,types}.ts described use-data.ts / useAirlineData
as "still live and still driving the trip, loyalty and disruption pages"; the
hook is deleted and the ledger is the only substrate.
skin-roster-docs.test.ts: comments only. No fixture entry, exemption or rule was
touched — the "legitimate phrasings" list still pins the numeral+adjective
discriminator, and the header still documents both false-positive shapes.
Reskin skill impact: checked — no rule, path or symbol the skill references
changed, so no skill edit is required beyond the prose pass in
|
||
|
|
e9759825b6 | test(skills): simplify public contract guard | ||
|
|
7bba4446ff | test(skills): narrow public API contract checks | ||
|
|
4f258e409a | test(skills): add manifest-backed public evals | ||
|
|
224dd10f07 |
feat(reskinnable-demo): add bookstore b2b skin (#6467)
## What does this PR do?
Adds `bookstore` ("Bookstore"), the reskinnable-demo's seventh skin, and
implements
demo beat 5 — **the book club run** — on it.
**The skin.** A customer-facing storefront on an in-memory substrate: a
25-book
catalog, generated covers, browse/book/cart pages, per-shopper cart and
orders
mirrored to `localStorage`, and per-shopper Intelligence identity. It is
the only
skin facing an end customer rather than an internal operator, which is
what makes
it a different proof from `commerce` (a merchant ops console) rather
than a second
copy of it.
**Beat 5.** The agent is *seeded* with an operational procedure it must
recall and
replay, rather than being told the steps in its prompt:
1. `addToCart` the club pick's hardcover
2. `swapEdition` to the paperback (same `workId`, different edition)
3. `applyPromoCode` with the club's code
4. `setDeliveryBy` the club's next meeting date
Three sibling tools — `addToWishlist`, `setReminder`, `applyStoreCredit`
— are
registered and genuinely work. Declining them is the beat's actual test:
a
distractor that errors proves nothing.
The promo code appears nowhere in `tools.tsx` (0 occurrences), so the
agent can
only obtain it from the club readable. If the code gets applied, the
readable was
genuinely read.
### Details worth a reviewer's attention
- **`localCalendarDay()`** (`data/club.ts`) re-anchors the caller's
*local*
calendar day onto UTC midnight. `nextMeetingISO` is UTC-only by design;
feeding
it a wall-clock `new Date()` gives a presenter west of UTC the wrong
meeting
day — at `2026-12-31T23:00-08:00` it skips a **full week**. Both the
club
readable and `setDeliveryBy`'s past-check use it, so the date the
readable
advertises can never be refused as past (proven over 23,107 weekday×date
combinations).
- **The discount renders as two rows, not one.** `discountCents` is a
single
scalar, so one row labelled with the club would render club-plus-credit
under
the club's name and misattribute the credit — reachable exactly when the
store-credit distractor misfires. `splitCartDiscount` recomputes the
club-only
part and takes credit as the remainder; `clubPart + creditPart ===
discountCents`
holds in every case, including a credit exceeding the subtotal.
- **All four totals agree to the cent.** The agent's cart readable, the
cart page,
the checkout form and the persisted order record all compute the same
figure.
Two of those were priced *without* the discount before this PR, so the
agent
would have announced the pre-discount total at the moment the procedure
tells it
to "confirm the new total in bold".
- **Every write tool uses `[]` deps and `dataRef.current`.**
`useFrontendTool`
keys its registration effect on `JSON.stringify(deps)`, so a callback in
a dep
array stringifies to a constant and permanently pins the pre-hydration
store.
This bug shipped once in this skin already.
- **Renders key off `result`, never `status`**, so a reopened thread
replays
correctly.
- **The empty-recall branch in `agent.ts` clause 7 came from a live
failure.** With
the memory unseeded, the agent correctly reported the miss and then
offered to
*learn* the procedure — which the clause already forbade, and which is
beat 6's
moment. It now says so plainly and stops, without guessing the pick,
edition,
code or date from the catalog or cart. An invented answer that looks
right is
worse than an honest failure: on stage the two are indistinguishable.
### Does this make anything in `.claude/skills/reskin/` wrong or
misleading?
**Yes — and it's updated in this PR.** Both `templates.md` (`SAVED
PROCEDURE`) and
`demo-beats.md` (beat 5's spec list) documented only the happy path, so
a new skin
author would ship the empty-recall gap above. Both now require the
branch and cite
the worked example.
`CLAUDE.md` had also drifted in five places, each phrased differently
enough that
keyword searches kept missing one; an exhaustive audit of all 36
bookstore claims
in the file closed it. `§ Commands` now also records that `pnpm lint` is
ESLint
only while lefthook's `pre-commit` runs `oxlint --fix` + `oxfmt --write`
and
re-stages — the two disagree on type-import style, so a contributor can
satisfy
the documented gate and be silently rewritten.
### Known gaps, deliberately not addressed here
- **`banking`'s beat-5 clause has the same empty-recall gap**, and its
only such
instruction affirmatively calls `offerWorkflowRecording`. Correctly
scoped to its
teach path, but it leaves a pattern pointing the wrong way for beat 5.
Left for a
separate change, since banking is the reference demo.
- Both empty-recall docs distinguish the beats by *beat number*, which
the agent
never sees. Banking makes it runtime-detectable via `overLimit: true`.
Bookstore
has no beat 6 so nothing is broken today; the first skin shipping both
will hit it.
- `beats 2, 4 and 5` exist only in Intelligence mode. Without it the
skin still
runs, but three of its four headline claims disappear — the runtime
warning at the
top of `suggestions.ts` says so.
## Related PRs and Issues
- (none)
## Checklist
- [x] I have read the [Contribution
Guide](https://github.com/copilotkit/copilotkit/blob/master/CONTRIBUTING.md)
- [x] If the PR changes or adds functionality, I have updated the
relevant documentation
- [x] "Allow edits by maintainers" is checked
|
||
|
|
085c92e6cc |
docs(reskinnable-demo): cut historical prose from the reskin skill
Same rule as the previous commit, applied to SKILL.md, demo-beats.md, failure-modes.md and templates.md: keep the rule and the mechanism that makes it a rule, drop who hit it, when, how it was found and how long it survived. Largest removals: the "there is no longer a partial skin to warn you off" retrospective closing demo-beats.md, the CR-pass provenance header on failure-modes.md, "this paragraph has now been wrong twice" under the pill count, the three-copies-of-the-staging-chain incident report, and the count-of-selectors paragraph that recorded its own rot. Past-tense incident illustrations were restated in the present tense rather than deleted, so every worked example still names its file. Two stale claims fixed while passing through: templates.md § tools.tsx said only banking, people and commerce key renders off `result` (every skin does), and SKILL.md described `--nw-nav-inset-*` as recently retired rather than simply absent. Reskin skill impact: this IS the skill; the app docs move in the commit before this one and the two are consistent. |
||
|
|
050814b433 |
docs(reskinnable-demo): cut historical prose from CLAUDE.md and README
Record the current state and the forward-looking instruction; drop the retrospective narration around it. Removed: the "MIGRATION not a split" substrate-history block, the three worked examples of past skill staleness under the standing rule, "was the FIRST skin"/"the second skin built"/"the retrofits" framing, "it did rot for two releases", the glass-engine replacement note, and the TS2352-in-a-green-slot anecdote. Every rule, gate, command, derivation and mechanism is kept; where a cut would have left a rule reading as arbitrary the reason is restated in the present tense (a nested thread rail compounds the assistant's floor; a client ticker is a second clock; a template teaching a removed pattern still compiles). Reskin skill impact: checked — the skill is edited in the following commit for the same reason, so the two stay in step. |
||
|
|
2ab26d66d5 |
docs: beat-5 coverage and reskin skill
Bring the app's own documentation back in line with what the skin now does, and record in the authoring skill the failure that a live run exposed. CLAUDE.md had drifted in five separate places, each phrased differently enough that keyword searches kept missing one: the seed catalog size, the count of SKIPPED beat-map rows (twice), the beat-matrix cell for stored-procedure replay, the intro paragraph's list of skipped beats, and a claim that the seed file seeds "no procedure at all". An exhaustive audit of every bookstore claim in the file — 36 of them — is what finally closed it. Also documents a repo-level trap in § Commands: pnpm lint is ESLint only, while lefthook's pre-commit additionally runs oxlint --fix and oxfmt --write over staged files and re-stages the result. The two disagree (oxlint enforces prefer-top-level type imports; ESLint does not), so a contributor can satisfy the documented gate and still be silently rewritten at commit time. This already misled a reviewer into filing a finding asking for the exact thing the hook auto-reverts. The reskin skill gains an empty-recall requirement for the stored-procedure beat, in both templates.md and demo-beats.md, which previously documented only the happy path. A skin author following either would ship the gap this skin shipped: with the memory empty, the agent reported the miss correctly and then offered to learn the procedure — beat 6's moment, arriving as an improvised fallback. Both now require saying so and stopping, with no guessing and no teach-offer, and cite the worked example. Note the reference skin has the same gap: banking's beat-5 clause has no empty-recall branch, and its only such instruction affirmatively calls offerWorkflowRecording. Correctly scoped to its teach path, but it leaves a pattern pointing the wrong way for beat 5. Left for a separate change. |
||
|
|
79cbc7ef84 |
feat(bookstore): seeded procedure, prompt and pill
The demo half of beat 5: a procedure the agent already knows, an instruction to recall rather than improvise it, and a pill so the presenter never types. - intelligence/seed-memories.ts: a kind "operational", scope "user" memory naming addToCart -> swapEdition -> applyPromoCode -> setDeliveryBy in order, and explicitly excluding the three distractors. The procedure is SEEDED, not taught — it is recalled. Scope is "user" and never "project", which would return the memory for every user of a shared instance. - agent.ts: clause 7 calls recall_memory FIRST, runs all four steps in order without confirmation, and states that finding the club is not running the procedure — reporting the pick, code or date and stopping is the failure mode, not a partial success. It scopes openCheckout out, so the run ends with a filled but unpaid cart, and refuses the teach-offer: this is a recall, not a teaching moment. - The empty-recall branch exists because a live run went off-script the moment the store was empty. With nothing recalled the model said so correctly and then offered to LEARN the procedure — which the clause already forbade, and which is beat 6's moment. It now says plainly that nothing was found and stops, without guessing the pick, edition, code or date from the catalog or cart: an invented answer that looks right is worse than an honest failure, because on stage the two are indistinguishable. The four tool names are frozen string literals shared by the prompt and the seed, and no test reads either, so renaming one breaks the beat with a green suite. A drift guard is the next commit's concern, not this one's. |
||
|
|
f1913bf485 |
feat(bookstore): cart discount and delivery UI
Price the cart through the three-argument cartTotals and show what the club run actually did to it. The discount is rendered as up to TWO rows, not one. discountCents is a single scalar, so a single row labelled with the club would render club-plus-credit under the club's name and silently misattribute the credit — and that combined case is reachable exactly when the applyStoreCredit distractor misfires, the most scrutinised second of the demo. splitCartDiscount recomputes the club-only discount and takes the credit as the remainder, so both parts are attributed honestly and clubPart + creditPart === discountCents holds for every case, including a credit that exceeds the subtotal (the club keeps its full percentage; credit takes only the applied remainder). Also adds the delivery-by badge, wishlist and reminder counts, and the same figures on the page readable so "what's on my screen" agrees with what the agent says. card_last4 remains the only card datum that leaves the checkout. |
||
|
|
4ee549a616 |
feat(bookstore): the book club mechanism
Everything the saved book-club procedure needs in order to run: the club constant and its computed next-meeting date, the edition pair the swap moves between, discount-aware pricing, the six store writes, and the twelve registered frontend tools. - data/club.ts: BOOKSTORE_CLUB (pick, promo code, 15%, meeting weekday), nextMeetingDate/nextMeetingISO (UTC-only by design) and localCalendarDay, which re-anchors the caller's LOCAL calendar day onto UTC midnight. Without it a presenter west of UTC demoing on a Thursday evening gets next Thursday: at 2026-12-31T23:00-08:00 the naive path skips a full week. - data/seed.ts: a 25th book, the club pick's paperback, sharing workId "trust" with the hardcover so swapEdition has a real work to move within. - data/query.ts: cartTotals gains an optional pricing object and returns subtotalCents/discountCents alongside totalCents, which stays the POST-discount amount charged. Inputs are sanitised so 0 <= discountCents <= subtotalCents holds for any input, including a non-finite credit or discountPercent. - data/use-data.ts: promoCode, deliverBy, storeCreditCents, wishlist and reminders persist under one extras key with a field-by-field validator; six writes returning WriteResult; placeOrder prices through cartTotals and consumes all three sticky fields. swapEdition merges into an existing target line rather than duplicating a bookId, and setDeliveryBy's past-check reads the local calendar day so the club's own date is never refused. - tools.tsx: the club readable (the only agent-reachable source of the promo code), the three procedure writes, the three distractors that genuinely work, and discount-aware pricing in both the cart readable and openCheckout's render so the total the agent speaks matches the cart page, the checkout form and the order record. Every registration uses [] deps and reaches the store through dataRef: useFrontendTool keys its effect on JSON.stringify(deps), so a callback in a dep array stringifies to a constant and pins the pre-hydration store. Reskin skill: checked, no impact — skin-internal data, store and tool wiring; no Skin contract field, registration, routing or gate changed. |
||
|
|
1c4003d2a3 | fix(bookstore): seed the default memory bucket and stop claiming per-shopper isolation | ||
|
|
80727a47f7 |
docs(reskinnable-demo): document the bookstore skin and correct the reskin skill
Answers the standing question in CLAUDE.md — this work found the skill wrong, so the fixes ship with it. Rule 1 on tool deps told authors to 'pass the data each closure reads' without noting that useFrontendTool keys its effect on JSON.stringify(deps). A Map, a Set or a function stringifies to a constant, so the registration is inert and the closure never refreshes — the skill's own words for the bug it warns about described the fix it recommended. The useData template taught a bare useState(SEED) and said nothing about a storage-mirrored variant, so an author needing one writes a hydration effect and trips react-hooks/set-state-in-effect immediately. Roster prose across CLAUDE.md, README.md, .env.example and the skill now covers seven skins. Most count claims were rephrased without a numeral rather than renumbered, so the next skin cannot re-falsify them — skin-roster-docs.test.ts is what caught them, and its roster fixtures are updated to match. |
||
|
|
3bde5f4443 |
feat(bookstore): assemble the Skin and register it across the shell
resolvePage uses a Map, never a plain object: segments[0] is untrusted URL input, and an object lookup walks the prototype chain, so /bookstore/constructor would resolve a Function where a ComponentType is declared and crash React instead of 404ing. skin.test.tsx pins that with the prototype-chain keys. An unknown book slug resolves the detail page and renders a not-found body rather than 404ing — the agent hands out these links, and 404ing a renamed book would break a deep link. Registration is four files, not two: both registries plus skins-config (whose test asserts skinIds and skinIdentities match the live registry, and which LOCK_SKIN is validated against) and eslint.config.mjs, where the id joins LINTED_SKIN_IDS — the array the URL-contract selectors interpolate, so without it lint is blind to this skin. |
||
|
|
02fdf443ea |
feat(bookstore): the agent prompt, its six tools, catalog and demo pills
The prompt is where the beats are enforced: recall memory before recommending and name the recalled preference in the note, never ask for or repeat card digits, never emit a markdown table where a gen-UI component exists. No temperature is set. gpt-5.4 rejects the parameter and logs that it is unsupported on every run, so pinning it alongside a comment claiming determinism would assert a guarantee the model discards. Tool registrations read live store data through a ref and close with empty deps where a dep cannot re-register them: useFrontendTool keys its effect on JSON.stringify(deps), so a Map or a function stringifies to a constant and the closure keeps its first values forever. openCheckout additionally must not re-register mid-call — placeOrder mutates the cart, and a teardown would lose respond() and fail the thread. Every render keys off the recorded result rather than status: a reopened thread replays with a stored result and no status transition, so a status-keyed render looks correct live and blanks on reload. |
||
|
|
6855ad9cca |
feat(bookstore): layout chrome and the browse, book and cart pages
The route readable in the layout plus one readable per page is what makes the screen-awareness beat work: asking on two pages must give two different correct answers. All four payloads are deliberately disjoint. The active segment comes from useSkinSegments, not a pathname slice — the shell hook strips a leading skin id rather than a fixed offset, so it stays correct under a LOCK_SKIN deploy where the segment is absent entirely. The presenter reset is a full-page assign, not a router.push: it clears storage with removeItem, bypassing the store, and the store has no storage listener, so only a document load re-reads it. A client navigation would leave the cart visibly full right after a successful reset. The cart page has no checkout button by design — checkout is the agent's beat. |
||
|
|
7a4f8dde79 |
feat(bookstore): generated covers, cards and the in-chat surfaces
Covers are typographic and generated rather than sourced images: 24 scans would be a licensing problem, would not reskin with the theme, and would read as stock photography in a demo whose argument is that the UI belongs to the product. checkout-card carries the security boundary — onSubmit receives only the last four digits, the other digits are cleared from state at that boundary, and both sensitive inputs are type=password because this card appears on a projector. Its receipt mode re-derives from a replayed result so a reopened thread shows a receipt rather than a blank form. filter-bar keeps the ebook lever even though no seed book has that format: the agent can set format=ebook via browseWithFilters, and a missing lever would make an agent-applied filter invisible, which is the one thing the component exists to prevent. |
||
|
|
22c8492922 |
feat(bookstore): cart and orders store with a per-shopper storage mirror
Reads through useSyncExternalStore rather than useState plus a hydration effect: the effect form fails react-hooks/set-state-in-effect, and a useState lazy initialiser that reads storage makes the server and client markup disagree. layout-preferences.tsx is the shell's sanctioned pattern. getSnapshot caches the parsed value and getServerSnapshot returns a frozen module-level constant — cart and orders are arrays, and a fresh array per call fails Object.is and infinite-loops during hydration. The storage mirror exists for the durable-thread beat: its proof is a hard reload, and a useState-only cart empties at exactly that moment. |
||
|
|
f9b649c59e |
feat(bookstore): per-shopper Intelligence identity, seeded memory and presenter reset
Memory is scoped per shopper, so the same suggestion pill answers differently for Maya (one seeded taste preference) and Guest (none). That contrast is the demo's headline claim, so identifyUser never derives a scope from userRole — both shoppers share the role and a role-derived scope would merge them. forgetAllMemories deliberately skips scope:'project' rows: project scope is global to the Intelligence backend instance rather than partitioned per product, all skins share one instance locally, and banking seeds a project-scoped procedure memory a bookstore reset must not destroy. The reset route maps raw shopper ids through resolveBookstoreUserId before clearing. Passing the raw ids would clear scopes nothing writes to and no-op while reporting success. |
||
|
|
860063ee48 |
feat(bookstore): brand identity, theme tokens, nav and lock-safe link builders
Theme values are space-separated HSL channels, not hex: globals.css wraps every token in hsl(), so a hex value yields invalid CSS and the whole skin silently falls back to the gray :root defaults. href.ts and nav-target.ts route every URL through useSkinHref. A hardcoded /bookstore/... path puts the tenant segment back in the address bar on a LOCK_SKIN deploy, and concatenating onto the builder's base emits the protocol-relative //book/x because that base is '/' under a lock. |
||
|
|
e3c6b75293 |
feat(bookstore): data types, 24-book seed and pure query functions
The seed test encodes the demo's falsifiability rule: the literary and
translated shelves must carry hardcovers and over-$20 titles, or the
recalled 'paperback only, under $20' preference has no visible effect.
cartTotals returns { itemCount, totalCents } — no tax and no shipping, so a
separate subtotal would duplicate the total and could drift.
|
||
|
|
d9667bc5f1 | test(sdk-python): cover intercepted event delivery boundaries | ||
|
|
ca2818f6dc | fix(sdk-python): deliver intercepted action events through the adapter | ||
|
|
9f77123df4 | fix(sdk-python): route intercepted action events through adapter filtering | ||
|
|
fbaa645bad |
docs(reskinnable-demo): finish the airline and keel beat-map cleanup
Follow-up to the previous commit, which fixed the two beat-maps' headers and their flagged risks but left the body still speaking in the future tense about work that has landed. Comment/doc only. src/skins/airline/data/beat-map.md - § "It is ADDITIVE" claimed `use-data.ts` (`useAirlineData`) "is untouched and still drives the trip, loyalty and disruption pages". Both are DELETED; every component reads `useAirlineLedger()` through `components/concierge-view.ts`. - Risk #5, "Two substrates, one passenger — they must not disagree on stage", told a later slot to "migrate BOTH readings" and not to touch either seed's AV1423 without the other. Marked RESOLVED, with the answer it actually got: the duplicate reading was DELETED rather than kept in sync, because a hand-synced pair of seeds was never going to survive a reseed and is a pair of figures that can contradict each other on a projector with nothing checking. - The beat table's four "(later slot)" cells, the beat-4 seed note, and the route list's "memory re-seed: later slot" — all shipped. src/skins/keel/data/beat-map.md - Beat 4's "the seeded memory (a later slot writes the file)" — written, and the scope it was written at (`user`, never `project`) named, since that is the choice the next author most needs and the one CLAUDE.md and demo-beats.md now flag. Skill impact: none beyond the previous commit. These two files are per-skin design records, not part of `.claude/skills/reskin/`, and the generalisable lessons in them (the two-substrates seed, the second clock) were already lifted into demo-beats.md § "Seeding memories" and SKILL.md § step 3 in that commit. Verified: pnpm lint, pnpm exec tsc --noEmit, pnpm test:unit (197 files / 2227 tests), pnpm build — all green. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
acac2b7f27 |
docs(reskinnable-demo): correct the in-code comments the beat-parity wave slots could not reach
Comment-only, no behaviour change. Each of these sat OUTSIDE the boundary of the slot whose work falsified it, so each described the tree as it was mid-migration. All verified against the current tree before editing. src/shell/skin-contract.ts - `RuntimeProviders`: "airline needs neither, so omits it". Airline omits `RuntimeProviders` and DOES supply `useRuntimeProperties` — the two are SEPARABLE, and airline is now the worked example: one account holder, no switcher, so its hook reads no context and returns a frozen module constant. Do not mount an empty provider for symmetry. - `useData`: "Omit when a skin has no shell-managed data (banking …)". Every skin omits it now; documented with the grep, plus WHY the field is kept (the shape is legitimate, it just has no worked example left). - `CanvasSurface`: the "omit if the skin has no report canvas" branch is currently unexercised — every shipped skin has one. src/shell/agent-registry.ts - The logistics entry said it ships "neither `intelligence/seed-memories.ts` nor `intelligence/forget-memories.ts`" and is "identity plumbing only — do NOT read it as a durable-memory demo". It ships both (`ls src/skins/*/intelligence/`), and its `dev/reset` sweeps and re-seeds through them. Corrected, with the same properties-forwarding caveat the other five entries carry, and a note that "expensive half built, cheap half skipped" was its state for two releases. src/skins/keel/data/types.ts - The `THE REST SUBSTRATE` banner said the two substrates are "deliberately not merged yet", that `useKeelData` "holds runs in `useState` and ticks them on a 900 ms interval", and that "the pages still read it through `useSkinData`". All three are false: one substrate, one clock, `useSkinData` returns undefined. Rewritten to name the server-settled read (`settle-runs.ts`, called by both `GET /ledger` and `GET /runs/[runId]`) and to record WHY the deleted client ticker was a defect rather than a design choice — it was a second clock that painted progress the server never heard of, which the next re-read after any write silently rewound. - The `KeelData` interface header claimed to be "the interface every page, component, and tool codes against". It is not referenced by any code at all (`grep -rn KeelData src` returns only comments). Marked HISTORICAL, with a do-not-add-a-consumer note. NOT deleted: it is a doc pass, several comments across the skin describe the migration in terms of this shape, and removing an exported type is a code change for a separate commit. Flagged as a follow-up. src/skins/keel/data/store.ts - "it does not advance them on a timer, because the ticker lives in `useKeelData` on the client. Whichever slot migrates that hook has to decide where the ticker ends up" — decided: the server is the only clock now. src/skins/keel/data/beat-map.md - Header: "Keel today is `useKeelData`, an in-memory `useState` store, and it hits about one beat." Marked BUILT and reframed as the design record. - Risk #2 (the two-substrates/ticker question, correctly called "the biggest single risk in the migration") marked RESOLVED, with the answer (move the clock, do not relocate the ticker) and the generalised lesson. src/skins/airline/data/beat-map.md - Header: the tools/prompt/pages/pills were "later slots". All landed. - § "What this slot did NOT build": every row of the deferral table has shipped. Kept as the retrofit record — which is the most useful thing about it — with a third column saying where each landed, and the three flagged traps marked resolved (including "the reset route says memoryBeats: unarmed on purpose", which was removed in exactly the change that added the seed module, as instructed). src/skins/airline/data/fare-waiver-codes.ts - "⚠️ THE LINT GUARD DOES NOT COVER THIS SKIN YET." It does: both `src/skins/airline/tools.tsx` and `agent.ts` are in `withheldGateVocabulary`'s `files` glob. Also dropped its "COUNT the selectors" instruction (that count has rotted twice) in favour of the resolved-selector table in `skins-config.test.ts`, and spelled out that a green lint still leaves the three prose channels AND `waiverGround` — which matches no `*_CODES` pattern, so the rule cannot see it — as hand-review items. src/skins/airline/tools.test.ts - Header said `statusKeyedTerminalRender` "covers logistics only; airline's glob entry is a later slot's, so until it lands this file is the whole guard" and that `withheldGateVocabulary`'s glob "does not list airline yet either". Both globs list airline now. Also fixed "Three defect classes" over a list of five. src/skins/keel/skin.tsx - "exactly as it does for the four other REST-backed skins" → every skin; nothing sets `useData`. src/proxy.ts - "matching how the other three skins behave" → numeral-free. This was one of the two known-stale instances named in `skin-roster-docs.test.ts`'s header; that header is updated in the app-docs commit, and the remaining one (`e2e/inset-layout.spec.ts`'s hardcoded four-skin loop) is deliberately left — fixing it means adding assertions against skins the spec has never visited, which is a coverage change rather than a prose fix. docs/teach-mode/README.md - It correctly refuses to write the roster into prose, but its verified-by-role paragraph named only banking/commerce/logistics/people and its "so copy commerce or logistics" line named the only two skins with pinned replay behaviour. The `offerWorkflowRecording` grep now returns every registered skin, and `ls src/skins/*/teach-mode-directives.ts` — added as the mechanical discriminator for role #3 — returns four. Also: "logistics and commerce both" skip project-scoped rows in `forget-memories.ts` is now every skin but banking, replaced with the grep that proves it. Does this make anything in .claude/skills/reskin/ stale? No — the reverse. The preceding commit updated the skill for exactly these facts, and these comments were brought into line with it. Checked: `grep -rn "useKeelData\|use-data\|in-memory" .claude/skills/reskin/` names no path or symbol that no longer exists. Verified: pnpm lint, pnpm exec tsc --noEmit, pnpm test:unit (197 files / 2227 tests), pnpm build — all green. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a58ea35b19 |
docs(reskinnable-demo): bring the reskin skill up to six demo-complete REST-backed skins
This IS the answer to CLAUDE.md's standing question for the beat-parity work:
"does this change make anything in .claude/skills/reskin/ wrong, incomplete, or
misleading for the next person authoring a skin?" It did, in all four files.
Every claim replaced was checked against the tree first, and replacements state a
derivation rather than a roster wherever one exists.
SKILL.md
- "`banking`, `people`, `commerce` and `logistics` are the four demo-complete
skins ... `airline` is the minimal contract surface" — false. Rewritten to "every
registered skin is demo-complete", so the choice is which one is CLEANEST for a
given problem, and a new "do not model a new skin on the ABSENCE of a field"
paragraph with the grep that derives what each skin actually sets. Airline's
restraint was the single most quoted licence to skip fields in this skill.
- The `useData?` contract row and the "4-2 split ... airline and keel are
in-memory and SET it" paragraph: zero implementors
(`ls src/skins/*/data/use-data.ts` is empty).
- The optional-slots list ("airline omits all of these EXCEPT toolLabels and
useData"): derived per field — the only real omissions in the tree are airline's
`sandboxFunctions` and `RuntimeProviders`.
- "banking, logistics, keel, people and commerce all five ship them, airline
none" (identity plumbing): all six ship `useRuntimeProperties` + `identifyUser`.
Added the point that the three parts are SEPARABLE, with airline as the worked
example of `useRuntimeProperties` without `RuntimeProviders` — do not mount an
empty provider for symmetry.
- "Only banking, people and commerce" register route + page readables: all six do.
- The `statusKeyedTerminalRender` glob comment ("logistics today; keel and airline
still carry the defect") — both are in the glob now; replaced with the grep, so
the sentence cannot rot again.
- The `agentRegistry` snippet handed out `airline: { createAgent }` as the
no-identity example. Rewritten so the identity-bearing form is the default.
- Authoring step 3 and the file tree taught `data/` + `useXData` as the substrate.
Now REST + `ledger-context.tsx`, with a new warning that time-dependent data must
be settled SERVER-side — a client ticker beside a server store is a second clock,
which is the defect keel had to be migrated off.
- § Verification step 1 and the authoring-order footer both claimed `pnpm build`
type-checks the whole app. It does not: `next build` only visits what the module
graph reaches, so it never opens a test file, and Vitest transpiles without
type-checking. Both now name `pnpm exec tsc --noEmit` as the only full
type-check, with four gates in cheapest-first order and the reason it matters
here specifically (several guards this skill asks for are TYPE-ONLY, so they are
decoration until tsc runs).
demo-beats.md
- Beat 3b: "this beat is impossible in airline and keel today" — false; both
register route + page readables. Replaced with the two greps, both of which now
return every skin, so a MISSING entry is the signal.
- Beat 3d: "In-memory skins can fake half of this" — no in-memory skin exists.
- Beat 6: the FIVE-channel leak list is now explicitly non-exhaustive, because
airline has a sixth nobody would look for: `Booking.waiverGround` is a
code-shaped token on a record the ledger publishes, so the LEDGER READABLE is a
channel. The transferable question is "what does my GET /ledger answer with", not
only "what did I put in a prompt".
- § "Seeding memories": the scope rule is rewritten and promoted to its own
flagged subsection. It described banking's `project` scope as the pattern. It is
now the minority and it is a trap for anyone copying a modern skin:
`forget-memories.ts` SKIPS project-scoped rows in every skin but banking
(`grep -ln 'scope !== "project"' src/skins/*/intelligence/forget-memories.ts`),
because project scope is global to the one shared Intelligence instance and
sweeping it would delete a sibling skin's seeds — so in such a skin a
project-scoped learned procedure CANNOT be cleared by the reset, and beat 6 opens
already-taught on the second run of the day, proving nothing while looking
perfect. `user` is documented as the default; banking is documented as
self-consistent the other way (project scope + a sweep that deletes everything)
and therefore not copyable. The beat-6 walkthrough's `scope: project` mention now
says so inline.
- § "Which skin to copy for what": every row re-derived. "the four at 9/9 beats",
"In-memory `useData` substrate → airline, keel" and "Minimal contract surface →
airline" were all false. Added rows for the three genuinely new references
(runtime identity with no context to read; a server-settled clock; an
entitlement-rather-than-authority gate) and for "raising an EXISTING skin", which
is what three of the six now demonstrate.
- The closing "do not use airline or keel as demo-completeness references" warning
is gone. What replaces it is the lesson the retrofits actually taught and which
generalises past any roster: keel shipped the full per-user identity plumbing and
then no seed file, so it got ZERO demo value from the hardest part of what it
built — build the seed file in the same phase as the plumbing.
- Pill-count prose: was wrong twice for the same reason. Now derivation-only.
templates.md
- Checked the `resolvePage` scaffolds first, because the `PAGES[key] ?? null`
object-literal pattern is a live bug (an object literal inherits
Object.prototype, so `PAGES["constructor"]` is truthy, `?? null` never fires,
and the shell's `if (!Page) notFound()` is bypassed — a 500 where a 404 belongs).
Both scaffolds ALREADY use `new Map()` and already explain why, so nothing to
fix; recorded here so the next reader does not re-check.
- `data/use-data.ts` pointed at `src/skins/airline/data/use-data.ts`, a DELETED
file. Rewritten to say plainly that nothing sets `useData`, that this template is
the only reference left, and to give the two reasons both skins migrated (beat
3d's artifact cannot outlive the tab in client state; a client ticker is a second
clock).
- The page and tools scaffolds defaulted to `useSkinData<<Id>Data>()` — i.e. they
taught the one shape no skin uses. Now default to the skin's own ledger context,
with `useSkinData` as the documented alternative.
- The `skin.tsx` scaffold's "airline omits every one below EXCEPT toolLabels +
useData — and airline hits one beat of nine" is false on both halves.
- "Mirror src/skins/airline/agent.ts (minimal)": airline's agent is 247 lines and
none of the six is minimal, because agent.ts is where most beats are enforced.
- The pill-count paragraph.
failure-modes.md — two new classes, in the file's voice
- § 13 "Nothing in this app type-checks a test file unless you run tsc yourself".
This is a lies-shaped defect, not a commands note: a green gate is a claim, and
three of them were green over a live TS2352 this run. It also silently voids a
class of guard this skill asks for, since several of the strongest assertions in
the tree are type-only. Cross-linked from § 7's "what would have to break for
this to go red", one level up.
- § 14 "A lookup keyed by URL input must be a Map, or the 404 branch never fires":
the prototype-chain bug above, why `Record<string, ComponentType>` cannot catch
it (the annotation is a lie about a plain object), and the § 11 class sweep — the
second instance was the operator→identity map in `user-id.ts`, keyed by a
CLIENT-forwarded `properties.userId`, where the same defect hands Intelligence an
`undefined` memory bucket and silently misroutes beats 4/5/6.
- § 10 now says the five-channel list is not exhaustive and documents airline's
`waiverGround` sixth channel plus why no identifier selector can see it, and its
"COUNT the selectors (covered files: four; any other in-skin file: three)"
instruction is replaced — that count had already rotted (`actions.ts` resolves to
two, and the number moved again when `statusKeyedTerminalRender` joined the
block). The check is the resolved-selector table in `skins-config.test.ts`, which
asserts the list BY NAME.
- The header's "everything here came out of one review of commerce" is now
"most of this", with §§ 13-14 attributed to the beat-parity run.
Verified: pnpm lint, pnpm exec tsc --noEmit, pnpm test:unit (197 files / 2227
tests), pnpm build — all green. skin-roster-docs.test.ts passes, including over
this skill's three files in its DOC_SET.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
d7675421a4 |
docs(reskinnable-demo): retire the two-substrate framing; all six skins are REST-backed and demo-complete
CLAUDE.md, README.md and .env.example described the app as it stood two waves
ago. Every claim below was verified stale against the tree before being changed,
and each replacement carries the command that derives it rather than a number or
a name list.
WHAT WAS FALSE, AND THE DERIVATION THAT PROVED IT
- "two different data substrates — banking/logistics/people/commerce REST-backed,
airline and keel in-memory ... to prove the contract is substrate-agnostic".
`ls src/skins/*/data/use-data.ts` returns NOTHING; both in-memory skins now read
REST ledgers (`ls -d src/app/api/*/v1` names six). Rather than delete the app's
central architectural claim, it is restated on the stronger footing the
migration actually earned: the proof is now that TWO SKINS CHANGED SUBSTRATE
with no change to the `Skin` contract and no change to the shell. Stated
honestly that no in-memory skin remains, so that shape has no worked example
left in the tree.
- The `useData` contract row ("the standard mechanism for the two in-memory
skins ... splits exactly along the substrate line"). It has ZERO implementors:
`grep -rn useData src/skins/*/skin.tsx` returns only six comments recording the
omission. Now stated plainly as an optional escape hatch nothing uses — live
rather than vestigial (the shell still runs the hook), with templates.md named
as the only remaining reference.
- The beat matrix. `airline` and `keel` showed ❌ on nine of ten rows and hit all
ten. Every cell updated. The per-skin GEN-UI COUNTS were removed rather than
corrected, and the derivation printed instead — the old command
(`grep -A3 useComponent .../tools.tsx | grep -c name:`) under-reports banking
by one, because banking registers a component in `pages/cards.tsx`; the
replacement counts the whole skin folder. Same for pill counts.
- Per-skin bullets for airline ("in-memory", "the minimal end of the contract",
omits eight optional fields) and keel ("in-memory", `useData: useKeelData`).
Derived per field:
`grep -nE '^\s+(Providers|CanvasSurface|...)[,:]' src/skins/*/skin.tsx`. Airline
omits exactly `sandboxFunctions` and `RuntimeProviders`; keel omits nothing.
Keel's parameterized-routes claim is TRUE and kept.
- "That is five of the six skins; airline is the only one that omits [identifyUser]".
`ls src/skins/*/intelligence/user-id.ts` returns all six, as do
`seed-memories.ts` and `forget-memories.ts`. Replaced with the three `ls`
commands. The "dev/reset is the wider set" note is kept as a WARNING while
recording that the two sets now coincide.
- "there is no `typecheck` script — `pnpm build` type-checks as part of
`next build`". Wrong, and it cost this run real time. `next build` type-checks
only what the app's module graph REACHES, so it never opens any of the 197 test
files; Vitest transpiles without type-checking; `tsconfig.json` DOES include
them. `pnpm exec tsc --noEmit` is the only full type-check in the tree and it
found a live TS2352 in a slot that had reported three green gates. § Commands
now names four gates in cheapest-first order; README's Testing block too.
- .env.example: "The other presenter-ready skins (commerce, logistics, people)".
All six ship `v1/dev/reset`; replaced with the route + button greps, and the
banking-gates-unconditionally difference recorded.
- New: the memory-SCOPE rule. `forget-memories.ts` skips project-scoped rows in
every skin but banking (`grep -ln 'scope !== "project"'`), so a project-scoped
learned procedure survives every presenter reset and beat 6 opens already-taught
on the second run. Banking is self-consistent the other way and documented as
the historical exception.
FIXTURE COUPLING — src/shell/skin-roster-docs.test.ts
Its "legitimate phrasings" list holds strings labelled "verbatim from the current
docs", three of which this commit made false in the docs ("the two in-memory
skins", "the four REST-backed skins", "five of the six skins"). Worth recording
precisely: the list is a SHAPE fixture passed straight to `findStaleCountClaims`,
not a doc mirror — it reads no file — so correcting CLAUDE.md alone would NOT have
turned the guard red. What it would have done is make the header comment a lie
and invite the next author to prune the entries, silently dropping the
discriminator they pin (an adjective between the numeral and "skins" is what
separates a subset claim from a total claim, and nothing else asserts it). So the
entries were moved to the PAST tense rather than deleted, the header now says
they are shapes and must not be pruned to match the docs, and three present-tense
phrasings from this commit were added.
The guard did its job on the way through: it failed on three sentences this
commit introduced ("the two skins" x2, "two of them"), all reworded numeral-free.
The header's "two known stale instances outside the doc set" note is updated —
`src/proxy.ts` is fixed in the following commit; `e2e/inset-layout.spec.ts`
remains, and is left alone deliberately because fixing it is a coverage change,
not a prose fix.
DOES THIS MAKE THE RESKIN SKILL STALE?
Yes, extensively — that is the point of this slot, and it is answered by the two
commits that follow (the skill, then the stale in-code comments). Per CLAUDE.md's
standing rule the skill is updated in the same PR.
Verified: pnpm lint, pnpm exec tsc --noEmit, pnpm test:unit (197 files / 2227
tests), pnpm build — all green.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
20da9fae41 |
feat(runtime): report managed Channel drops and recoveries as telemetry (refs OSS-825)
A managed Channel that loses its gateway link was invisible outside the host process: the only trace was the injected `log` seam, which for a self-hosted or Railway-hosted runtime reaches nobody who can act on it. One 2026-08-12 outage ran 2.7 days on `kite-community` before a customer reported it, and reconstructing it needed pod forensics. Emit `channel_session_dropped` (carrying the cause we already compute for the log line) and `channel_session_recovered` (carrying how long the outage lasted). A session may report `online` without a preceding drop, which is not a recovery and is not reported as one. Events deliberately omit the Channel name — it is a customer-chosen identifier — and carry no message content. Capture is fire-and-forget with failures swallowed, the same contract fireInstanceCreatedTelemetry uses: telemetry must never break a live session. Also replay the drop cause on each "still down" reminder. In prod those lines read `still down after 233134s; Phoenix is retrying` with no cause at all, so an operator had to find the first line, 15 minutes earlier, to learn it was an HTTP 502. |
||
|
|
1ea8901fe2 |
chore(deps): update astral-sh/setup-uv action to v10 (#6463)
This PR contains the following updates: | Package | Type | Update | Change | |---|---|---|---| | [astral-sh/setup-uv](https://redirect.github.com/astral-sh/setup-uv) | action | major | `v9.0.0` → `v10.0.0` | --- ### Release Notes <details> <summary>astral-sh/setup-uv (astral-sh/setup-uv)</summary> ### [`v10.0.0`](https://redirect.github.com/astral-sh/setup-uv/releases/tag/v10.0.0): 🌈 Disable automatic caching for sensitive events and new QOL features [Compare Source](https://redirect.github.com/astral-sh/setup-uv/compare/v9.0.0...v10.0.0) ##### Changes Another breaking release, directly after v9.0.0 but we think the added security justifies that. ##### Extra security by default If you use the default `enable-cache: auto` this will now **DISABLE THE CACHE** to protect against cache poisoning for the following events: - `pull_request_target` - `workflow_run` - `release` You can read the full reasoning in [#​984](https://redirect.github.com/astral-sh/setup-uv/issues/984) ##### `version: latest-known` ```yaml - name: Install the latest version of uv known to setup-uv uses: astral-sh/setup-uv@v10.0.0 with: version: "latest-known" ``` This will now install the latest version with a checksum that is known by this action. The [known `uv` checksums](https://redirect.github.com/astral-sh/setup-uv/blob/4f6036f71cec78afb113b323f220c9185d983c12/src/download/checksum/known-checksums.ts) are automatically updated but will take a release of this action to take effect. You won't be always using the latest & greatest but you will have an extra level of security. ##### Read python version from `.tool-versions` ```yaml - name: Install uv based on the version defined in .tool-versions and also set python uses: astral-sh/setup-uv@v10.0.0 with: version-file: "pyproject.toml" ``` Will now also set the python version if it is defined in `.tool-versions`. You can read the details [in the docs](https://redirect.github.com/astral-sh/setup-uv/blob/main/docs/advanced-version-configuration.md#install-a-version-defined-in-a-requirements-or-config-file) ##### 🚨 Breaking changes - Disable automatic caching for sensitive events [@​eifinger](https://redirect.github.com/eifinger) ([#​992](https://redirect.github.com/astral-sh/setup-uv/issues/992)) ##### 🐛 Bug fixes - Reject paths in .tool-versions [@​eifinger](https://redirect.github.com/eifinger) ([#​1007](https://redirect.github.com/astral-sh/setup-uv/issues/1007)) ##### 🚀 Enhancements - Read Python version from .tool-versions [@​eifinger](https://redirect.github.com/eifinger) ([#​996](https://redirect.github.com/astral-sh/setup-uv/issues/996)) - Add latest-known version selector [@​eifinger](https://redirect.github.com/eifinger) ([#​993](https://redirect.github.com/astral-sh/setup-uv/issues/993)) ##### 🧰 Maintenance - Require pull requests for Dependabot rollups [@​eifinger](https://redirect.github.com/eifinger) ([#​1005](https://redirect.github.com/astral-sh/setup-uv/issues/1005)) - ci: pin Alpine container image [@​eifinger](https://redirect.github.com/eifinger) ([#​995](https://redirect.github.com/astral-sh/setup-uv/issues/995)) - chore: update known checksums for 0.12.3 @​[github-actions\[bot\]](https://redirect.github.com/apps/github-actions) ([#​991](https://redirect.github.com/astral-sh/setup-uv/issues/991)) - chore: update known checksums for 0.12.2 @​[github-actions\[bot\]](https://redirect.github.com/apps/github-actions) ([#​985](https://redirect.github.com/astral-sh/setup-uv/issues/985)) - chore: update known checksums for 0.12.1 @​[github-actions\[bot\]](https://redirect.github.com/apps/github-actions) ([#​982](https://redirect.github.com/astral-sh/setup-uv/issues/982)) - chore: update known checksums for 0.12.0 @​[github-actions\[bot\]](https://redirect.github.com/apps/github-actions) ([#​981](https://redirect.github.com/astral-sh/setup-uv/issues/981)) - chore: update known checksums for 0.11.31/0.11.32 @​[github-actions\[bot\]](https://redirect.github.com/apps/github-actions) ([#​972](https://redirect.github.com/astral-sh/setup-uv/issues/972)) ##### 📚 Documentation - docs: update version references to v9.0.0 @​[github-actions\[bot\]](https://redirect.github.com/apps/github-actions) ([#​971](https://redirect.github.com/astral-sh/setup-uv/issues/971)) ##### ⬆️ Dependency updates - chore(deps): roll up Dependabot updates [@​eifinger](https://redirect.github.com/eifinger) ([#​1013](https://redirect.github.com/astral-sh/setup-uv/issues/1013)) - chore(deps): roll up Dependabot updates [@​eifinger](https://redirect.github.com/eifinger) ([#​1004](https://redirect.github.com/astral-sh/setup-uv/issues/1004)) - chore(deps): roll up Dependabot updates [@​eifinger](https://redirect.github.com/eifinger) ([#​994](https://redirect.github.com/astral-sh/setup-uv/issues/994)) - chore(deps): bump zizmorcore/zizmor-action from 0.5.7 to 0.6.0 @​[dependabot\[bot\]](https://redirect.github.com/apps/dependabot) ([#​976](https://redirect.github.com/astral-sh/setup-uv/issues/976)) - chore(deps): bump actions/checkout from 7.0.0 to 7.0.1 @​[dependabot\[bot\]](https://redirect.github.com/apps/dependabot) ([#​980](https://redirect.github.com/astral-sh/setup-uv/issues/980)) </details> --- ### Configuration 📅 **Schedule**: (in timezone America/Los_Angeles) - Branch creation - "before 9am every weekday" - Automerge - At any time (no schedule defined) 🚦 **Automerge**: Enabled. ♻ **Rebasing**: Whenever PR is behind base branch, or you tick the rebase/retry checkbox. 🔕 **Ignore**: Close this PR and you won't be reminded about this update again. --- - [ ] <!-- rebase-check -->If you want to rebase/retry this PR, check this box --- This PR was generated by [Mend Renovate](https://mend.io/renovate/). View the [repository job log](https://developer.mend.io/github/CopilotKit/CopilotKit). <!--renovate-debug:eyJjcmVhdGVkSW5WZXIiOiI0NC4yNC4wIiwidXBkYXRlZEluVmVyIjoiNDQuMjQuMCIsInRhcmdldEJyYW5jaCI6Im1haW4iLCJsYWJlbHMiOltdfQ==--> |
||
|
|
a255d9bcf3 |
chore(reskinnable-demo): put airline and keel under the beat-2 and beat-6 lint guards
Both skins now ship a withheld gate vocabulary and replay-safe terminal
renders, so both belong in the two guards that check those properties. Held
back until now on purpose: a glob covering an unfixed skin turns the tree red
for a whole phase, and a phase that cannot end green is a phase nobody can
bisect.
Verified clean BEFORE widening, not after:
- beat 2 (statusKeyedTerminalRender). Every remaining ToolCallStatus
reference in either skin is a comment or a
`=== ToolCallStatus.Executing && respond` HITL branch -- the interactive
affordance drawn while a response is awaited. Neither has a `.Complete`
terminal render, which is the shape the selector catches.
- beat 6 (withheldGateVocabulary). Each skin contributes exactly its two
agent-facing files. The three human filing FORMS stay OUT: each
legitimately imports its label map, because the operator reads it and the
agent learns the code by watching them choose one. A withheld catalogue
with no form is an unlearnable gate.
skins-config.test.ts's resolved-selector table is updated in the same commit,
which is the point of that table: it asserts, per file, the selector LIST that
`ESLint#calculateConfigForFile` actually resolves. Its existing keel row
expected three selectors and would have failed. Six rows added or changed,
including both filing forms, so the deliberate exclusion is pinned rather than
merely intended.
Also resolves a bug this commit introduced and then fixed, kept because the
comment is the fix: a glob star followed by a slash inside a block comment
CLOSES the comment, and the rest of the file becomes a syntax error. `pnpm lint`
caught it as `Parsing error: ',' expected` five lines below the real cause.
|
||
|
|
00162f5e1b |
feat(reskinnable-demo): give keel memory, a stored procedure and a teach loop
Merges blitz slot keel-teach. Conflict resolution note. Both teach slots hand-edited the same paragraph in agent-registry.ts describing which skins supply identifyUser, and BOTH sides were wrong by the time they merged -- air-teach's said "logistics and keel scope threads only" (keel-teach had just given keel durable memory), and keel-teach's said "skins without it (e.g. airline)" (air-teach had just given airline a resolver). Each was made false by the other's slot, in the same hour. Resolved by deriving instead of picking a side: ls src/skins/*/intelligence/user-id.ts -> all six ls src/skins/*/intelligence/seed-memories.ts -> all six ls src/skins/*/intelligence/forget-memories.ts -> all six All six skins now supply identifyUser AND use it for durable memory, so the generic-identity fallback is unreachable from the registry and is kept only for skins that do not exist yet. The comment now says that and points at the derivation, with the near-miss recorded so the next person does not re-add a hand-maintained list. |
||
|
|
6cbbd0757e | feat(reskinnable-demo): give airline memory, a stored procedure and a teach loop | ||
|
|
2c698c3724 |
fix(reskinnable-demo): make keel's presenter reset re-arm the memory beats
The route restored the DATA STORE ONLY and said so loudly in its own header, because keel had no seed/forget pair. Now it has one, so the reset does the other half — and, critically, VERIFIES it rather than assuming. - Wipes learned memory in every bucket `memoryScopeUserIds()` reports, ASKED FOR rather than hardcoded. Bellwether shipped a hardcoded list and it could not possibly be right: `playwright.config.ts` pins `INTELLIGENCE_USER_ID`, which collapses the whole set onto one bucket the list did not contain, so the reset scrubbed buckets nothing was reading while a taught procedure survived. - Re-seeds every target bucket, then COMPARES the count against `seedTargets.length * SEED_MEMORIES.length`. `seedMemories` never throws — it counts stored rows and logs the rest — so without the comparison the route would answer `reset: ["store","memory"]` against a backend that had rejected every POST, and the presenter would walk on stage believing beats 4/5 were armed. - `reset: ["store","memory"]` is claimed ONLY when the wipe proved itself complete AND every expected seed landed. Partial and total shortfalls both 502, because a shortfall does not say WHICH memory is missing and the only caller branches on `res.ok`. The interrupted path reports MEASURED counters (`bucketsSwept`, not `forgot` — an empty bucket forgets zero rows, which is the normal state of a second reset in a row). - Every free-text field goes through `redactSecrets`: this route's gate is a demo convenience, not an authorization boundary, and `memoryError`'s cause can quote the backend address verbatim while a 401 body can echo the key it rejected. The address stays in the LOG, which is where a human debugging a reset reads it. - The OSS path (no Intelligence env) still answers `reset: ["store"]` and never the word "memory" — the most misleading string this route could return, because a presenter reading it stops looking for the reason beat 6 opened already taught. The test that pinned the old body is updated to that case rather than deleted. Does this make anything in `.claude/skills/reskin/` wrong? Checked: no. The skill already prescribes the seed-then-verify reset shape; keel was the outlier and is no longer one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
e381d58d64 |
feat(reskinnable-demo): give keel one suggestion pill per beat
The pills are the demo's script — the presenter should never have to type — and keel's four were the whole list, covering one beat. Now twelve, in demo order. Eight beat-carrying additions (1, 3a, 3b, 3c, 3d, 4, 5, 6) plus keel's original four, which survive VERBATIM: spec §11's walk-through is scripted against their copy, and each is the only pill reaching its tool — grounded citation, the run plan-preview HITL, the persona-scoped approval queue, the a2ui canvas report. The beat map targeted eight-to-nine, and that figure is incompatible with its own other two instructions once the pills are counted: eight beat pills plus four survivors is twelve, and nine is only reachable by dropping an identity pill or leaving a beat without one. Both were checked against the tools and neither is available. The arithmetic is written out in the file header so the next reader does not "fix" the count by deleting a beat. Two things that would break silently and are now pinned by tests: - The beat-3d pill's message IS `BULLETIN_MESSAGE`, imported. `onSuggestionSelect` keys on that exact string; a retyped sentence takes the default send path, which DROPS attachments — the model then invents the bulletin's contents and files a durable brief that reads perfectly and proves the opposite of the beat. Asserted through the real `onSuggestionSelect` rather than by string comparison, so it fails if either side drifts, and every OTHER pill is asserted to return false. - Beats 3a, 5 and 6 target three DIFFERENT documents (STD-045 endorsed, POL-121 stale, POL-114 gated). Beat 6's unaided replay on POL-208 deliberately has no pill — a scripted sentence would let the room suspect it was rehearsed. Also asserted: no pill says "Knowledge" (the nav label is Register while the segment is still `knowledge`), and no pill names a variance code — a pill's message reaches the model, so it is the prompt's leak channel by another route. Does this make anything in `.claude/skills/reskin/` wrong? Checked: no. The skill's "one pill per beat, in demo order" rule is what this implements; the stale number is in keel's own `data/beat-map.md`, and the reconciliation is documented at the call site rather than by editing that record of the original design. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8b8b08603b |
feat(reskinnable-demo): wire keel's beats 3c, 4, 5 and 6
The tools, the prompt clauses and the mount points. Keel went from one beat to ten.
BEAT 3c — `showRegister`, a HITL maneuver rather than a link. The Register's four
controls existed and nothing drove them. Space x attention x sort x top-N arrive
from the query string through the one shared `data/register-levers.ts`, the confirm
card draws its chips from the SAME normalized record the URL is built from, and the
controls tint. Every lever is REQUIRED with an explicit not-pulled member ("all", or
0): logistics needed a fix commit for exactly this, because a model facing an
optional enum fills it anyway — omission is not a choice it can state — and it put
an empty board on screen under four confidently tinted controls.
BEAT 4 — `showRegisterSummary(note)`, plus a prompt clause forcing `recall_memory`
BEFORE any question about the library's shape and requiring the recalled preference
in `note`. Without the visible why the beat is invisible on stage.
BEAT 5 — `raiseReviewFlag` -> `sendOwnerNotice` -> `addDocumentNote`, all three
`useFrontendTool` and NOT HITL: banking's equivalent once opened a confirmation card
mid-procedure, a presenter moved on, and the next message failed the whole thread
with "Tool result is missing for tool call ...". Their vocabularies are ENUMERATED on
the schemas — the exact opposite of beat 6 — because the claim is that it already
knows the procedure. Nine distractor tools sit alongside, so "it picked the right
three" means something. The prompt adds FINDING IS NOT HANDLING and states that this
is a DIFFERENT procedure from beat 6's with no offer to record.
BEAT 6 — `fileReleaseVariance` plus the HITL chain `offerWorkflowRecording` ->
`awaitDemonstration` -> `saveLearnedProcedure` -> `save_memory`. All five leak
channels are closed: no readable, no `z.enum` (free `z.string()` whose `.describe()`
states the withholding), no code in any description, none in the prompt, and the
route refusals are relayed verbatim without enumerating the catalogue. The prompt's
ACTION DISCIPLINE clause also shuts the two doors a run gate would open — a persona
switch and the e-signature card — because neither can clear a gate about the
REVISION. The replay lands on POL-208 Rev C, a different record from the POL-114
Rev D taught on stage.
Mounts: `KeelProviders` (new) carries the shell's `RecordingProvider` +
`RecordingVignette` BELOW CopilotKitProvider — the only point enclosing both the app
card where the operator demonstrates and the chat card that reads the feed; a
narrower mount makes every `logStep` a silent no-op. The variance filing form goes
on the Register page and is deliberately ABSENT from that page's readable. The
presenter Reset button goes in the header behind the same gate the route enforces.
Does this make anything in `.claude/skills/reskin/` wrong? Checked: no — but it does
leave `eslint.config.mjs` and `src/shell/skins-config.test.ts` needing keel's globs
(withheldGateVocabulary for `keel/{tools.tsx,agent.ts}`, and keel is now free of
`ToolCallStatus` so it can join statusKeyedTerminalRender). Both files are the
orchestrator's; the exact additions are in the slot report and in
`data/variance-codes.ts`'s header. Until they land, keel's own
`tools-replay-safety.test.ts` and `agent.test.ts` are deliberately stronger than the
rule.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
8aa3f62fa9 |
feat(reskinnable-demo): add keel's teach-mode surfaces and beat-4 card
The client half of beats 4 and 6. The server gate and unlock already existed; what was missing was every surface a human or the recorder touches. - `teach-mode-directives.ts` — the four strings the teach chain settles with, each BUILDER beside its READER so they cannot drift. `classifySaveProcedureResult` exists because both save buttons settle with a string: branching on presence prints "Saved — I'll use this next time" after the presenter clicked "Don't save", asserting a durable write that never happened, live and identically on every replay. `SAVE_PROCEDURE_CONFIRMED` names scope 'user' — project scope survives the forget sweep, so a project-scoped procedure would leave beat 6 opening already taught. - `components/variance-form.tsx` — the operator's filing form, and the ONE sanctioned consumer of `VARIANCE_CODE_LABELS`. It is the SIXTH channel and the one that must be OPEN: the agent learns which code lifts the gate by WATCHING the operator pick one. The menu lists justifying codes and decoys together, unmarked, in catalogue order — a form that flagged the working ones turns the demonstration into a guided tour. The filing step logs the code the operator ACTUALLY chose, decoy included, because a recorder that quietly corrected them would report a procedure nobody demonstrated. - `components/demonstration-card.tsx` — owns the OUTER recording bracket, held from "show me" to "I'm done" so the two clicks a demonstration takes (file, then release) read as one recording. If the ref count reaches zero between them the shell clears the feed and STRANDS the code, and `getDemonstratedCode()` then reports null on a demonstration that plainly happened. - `components/register-summary-card.tsx` — beat 4, with the `note` slot that makes the recall visible. Without it a grouped list is not evidence of memory; a model with none could produce one. Note renders ABOVE the groups, because it is the claim they are evidence for. - `components/presenter-reset-button.tsx` — hard-navigates on success (the thread, canvas, levers and recorder feed are all state the reset threw away) and stays put on a 502, saying the register WAS restored but memory was not. - `data/variance-codes.ts` header updated: the form it reserved `VARIANCE_CODE_LABELS` for now exists, and the header names the eslint glob keel still needs plus the tests standing in until it lands. Does this make anything in `.claude/skills/reskin/` wrong? Checked: no. It follows the skill's existing teach-mode guidance (shell recorder, never a private copy; the sixth channel open) rather than changing it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
a5732a235b |
feat(reskinnable-demo): ship airline's eight pills and arm its presenter reset
The pills were the pre-beat-map set — check in, pick a seat, loyalty, delay, bags — which demonstrated none of the nine beats and made a presenter type every one. Replaced with ONE PILL PER BEAT in demo order, each carrying a comment recording what it was VERIFIED to produce against the seeded ledger and where verification stops (anything downstream of the model's tool choice is labelled runtime-conditional rather than claimed). Two orderings are load-bearing, not cosmetic. Beat 4 sits BEFORE beat 5, because beat 5 rebooks the cancelled return that beat 4's seeded preference says to lead with. Beat 6 is LAST, because the room has to watch the concierge succeed at everything else before it is shown failing. The beat-3d pill's message is `HOTEL_CONFIRMATION_MESSAGE` imported from `./attach-hotel-confirmation`, never a retyped sentence: `onSuggestionSelect` matches on that exact value, and a drifted string takes the default send path with the attachment DROPPED — which fails beat 3d while looking like a model problem. `suggestions.test.ts` asserts the identity, that the interceptor claims that pill AND ONLY that pill, and that claiming it actually sends. Presenter reset — `dev/reset` now sweeps and re-seeds durable memory (commerce's route, verbatim in structure), so a cold reset arms beats 4/5 with no warm-up run and leaves beat 6 unlearned. `memoryBeats: "unarmed"` and its `memoryNote` are DELETED in this same change, which is the ordering `data/beat-map.md` trap 3 asked for: the field existed to stop a store-only reset being a silent trap, and keeping it now would be the same lie pointing the other way. `route.test.ts` pins its absence alongside the shortfall/wipe-incomplete/interrupted paths, and keeps the REAL store for the "puts back everything the beats wrote" case. Airline had NO reset control at all, so the sidebar gains one behind `PRESENTER_RESET_ENABLED` — gated exactly as the route is, so a production booth never shows a button that 403s. A non-ok response alerts and does NOT reload: a 502 means the wipe could not prove it finished or the seed fell short, and both break a beat silently, so reloading to a clean-looking app is the wrong answer. Reskin-skill review: checked, no skill impact. The pills, the reset route shape and the sidebar control all follow patterns `demo-beats.md` § "Presentation requirements" and § "Seeding memories" already prescribe; nothing about the contract, the registration sites or the verification commands changed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
02be94c4a7 |
feat(reskinnable-demo): add keel's intelligence seed/forget pair
Keel had `intelligence/user-id.ts` and nothing else, so its presenter reset restored the DATA STORE ONLY: it could not wipe a procedure the agent had been taught, nor re-arm the memories beats 4 and 5 depend on. Run beat 6 twice and the second room watched an agent that already knew the answer. - `seed-memories.ts` — beat 4's TOPICAL reading preference (group by knowledge space, overdue first, coverage as a whole percent, owner beside every ref, and "not measurable" rather than 0%) and beat 5's OPERATIONAL procedure (raiseReviewFlag -> sendOwnerNotice -> addDocumentNote, immediately, no confirmation). Both `scope: "user"`, never `project`: the forget sweep skips project rows because that scope is global to the shared Intelligence instance, so a project-scoped procedure would survive every reset. Beat 6's unlock is DELIBERATELY not seeded — that is what must be taught on stage — and the file says so where the next author will read it. - The memory text is addressed to the DESK, not to a named persona. It lands in every persona's bucket and keel's role switcher sits in the header, so "when Sam asks…" would be recalled while Ana is on screen. - `forget-memories.ts` — mirrors commerce's: bare-list enumeration, verified list->delete passes rather than a guessed page size, per-row failures stepped over rather than abandoning the bucket, and project-scoped rows left alone so a keel reset cannot destroy banking's seeded memories. - `user-id.ts` gains `memoryScopeUserIds()` and `memorySeedTargetUserIds()`, DERIVED from the persona roster so the reset asks rather than restates. Seeds the DEFAULT bucket as well as every mapped persona's, because runs frequently resolve to the default and a single-bucket seed recalls nothing while looking perfectly stored one id over. Does this make anything in `.claude/skills/reskin/` wrong? Checked: no. The skill describes seed/forget as the pair a memory-claiming skin ships, which is now true of keel; `demo-beats.md`'s "Seeding memories" section still describes banking's `project` scope, which was already out of date before this change and is called out in `seed-memories.ts`'s own comment rather than silently followed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |