stalePendingFilters silently falls back to the 1h default period for
any family missing from FLEET_FAMILY_PERIODS_MS, so a typo'd key (e.g.
"d5" vs "d5-single-pill-e2e") would never throw — it would just quietly
mis-size that family's stale-pending drain window. Lock the map's keys
to the probe-key families derived by RUNNING the four real enumerator
factories against a fake discovery source, so either side drifting
breaks the test. Also document the known d6 FLEET_PRODUCER_CRON
override drift limitation on the map (an env override changes d6's real
cadence without updating the nominal period).
The field was dead on the only consumer path: queue-client's enqueue()
destructures only `payload` and never reads leaseSeconds (the claim
lease comes from the WORKER side — claimNext(workerId, leaseSeconds)
with the worker-loop's DEFAULT_LEASE_SECONDS). No production call site
ever set it, so wiring it up would add a knob nothing needs; delete is
chosen over wire-it.
Call-site enumeration (all removed):
- contracts.ts EnqueueJobInput.leaseSeconds (declaration; never read by
queue-client.ts enqueue, the sole FleetQueueClient.enqueue impl)
- job-producer.ts ServiceJobSpec.leaseSeconds (only producer thereof)
- job-producer.ts toEnqueueInput() spec.leaseSeconds -> input.leaseSeconds
threading (only writer of the field)
- job-producer.test.ts spec fixture leaseSeconds: 600 + the
`expect(input.leaseSeconds).toBe(600)` assertion that legitimized the
dead plumbing
Worker-side lease plumbing (claimNext/renewLease/worker-loop
leaseSeconds) is unrelated and untouched.
maybeSweep's catch arm returned the same shape as a clean zero-reclaim
sweep (sweptExpired: true, reclaimed: 0), so a thrown sweepExpired call
was indistinguishable from success in the TickResult and the
tick-complete log. Add sweepFailed to the sweep outcome, TickResult,
and the tick-complete log; the cadence latch is unchanged (a failed
sweep still consumes its window so a persistently-failing sweep cannot
fire on every tick).
The sweep no longer synthesizes worker-crashed-mid-job (it cannot tell a
crash from a platform teardown); it re-queues the job and emits the neutral
worker-reclaimed-pending kind. Update the queue-client module header, the
contracts kind/heartbeat/SweepResult/sweepExpired docs, the job-producer
sink/tick docs, the sweep test fixtures and titles, and the dashboard's
mirrored kind description. worker-crashed-mid-job is now documented as the
worker self-observed in-driver crash only. No runtime behavior changes.
Only the d6 producer was built with onSweepCommErrors, but all four family
producers run the same GLOBAL queue.sweepExpired on their own crons — and the
sweep's S0 CAS means whichever producer ticks first wins each expired job's
reclaim, along with its synthesized comm error. With smoke/demos/deep sweeping
far more often than d6's hourly :40, the worker-reclaimed-pending dashboard
overlay (and stale-pending telemetry) was dropped ~11 of 12 sweeps, since
job-producer's maybeSweep forwards comm errors only when the sink is wired.
Share the ONE control-plane sink (surfaceSweepCommErrors -> aggregator) across
all four producers and correct the now-false "preserves the current behavior"
comment. Sweeps remain CAS-safe across producers: the S0 CAS guarantees exactly
one producer reclaims (and forwards) each expired job, and the surfacing leg is
best-effort per error, so the shared sink introduces no double-write.
Red-green: new runControlPlane REQ-B test drives the SMOKE producer's tick and
asserts its swept overlay reaches the status row (RED against d6-only wiring,
GREEN after). The test doUnmocks/re-mocks everything it touches so it passes in
isolation despite the file's leaked doMock factories.
filterBackloggedFamilies ran BEFORE maybeSweep, so the tick whose own
stale-pending drain cleared a family's backlog still counted the
about-to-be-expired rows and skipped that family — production resumed a
full cron period late. Reordered tick() to sweep first; the cadence gate
and fail-open semantics (maybeSweep swallows sweep failures) are
unchanged, only the order moved.
A single 50-row page per sweep was far slower than the incident the
stale-pending drain exists for: against the 3,734-row staging backlog at
~10 sweeps/hour that is ~7.5 hours of drain. The sweep now loops candidate
pages (re-listing page 1 — deletes shift pagination) up to a cap of 10
pages / 500 rows per sweep, draining the same backlog in well under an
hour while bounding a single sweep's PB load. CAS-claim-then-delete per
row is unchanged; a pass that expires nothing terminates the loop.
The sweepExpired lease phase listed claimed/running rows with perPage 50
and NO sort: with >50 such rows (mass worker crash), PB's unspecified
default order could return the same 50 live-lease rows every sweep,
leaving expired leases beyond the page permanently orphaned with zero
signal. Sort by lease_expires_at ascending (indexed) so the most-expired
rows always head the page, and WARN when the page is full so truncation
is observable. Single page per sweep is kept deliberately — the sort
guarantees progressive forward drain.
When the stale-pending sweep CAS-claims a row under stale-pending-sweeper
and the delete fails, the next lease sweep treated the expired sweeper
lease like a crashed worker's: it re-queued the row AND synthesized a
worker-reclaimed-pending comm error — a gray "re-queued / back in flight"
dashboard overlay for stale garbage mid-deletion, attributed to a
non-existent worker. The lease sweep now special-cases rows held by the
stale-pending sweeper: still re-queued (the self-healing delete-retry
contract is unchanged) but silently — no comm error, no reclaimed count,
just a stale-sweeper-retry-requeue debug line. The lease holder is
snapshotted before the release CAS so attribution reflects who held the
expired lease, not the post-release row.
sweepExpired's lease phase re-queues an expired-lease row to pending and
emits worker-reclaimed-pending ("back in flight"), but the stale-pending
phase of the SAME call lists pending fresh and ages rows off PB's system
`created` (the ORIGINAL enqueue time) — so a long-claimed job was
re-queued then immediately claimed-and-deleted, falsifying the comm
error and nulling downstream aggregate-key resolution on the deleted
row. Track the ids re-queued in this sweep and exclude them from this
call's stale phase; a truly stale job ages out on the next sweep. The
`created` anchor is kept (re-anchoring needs a column; out of scope).
sweepExpired only reclaimed claimed/running leases — a pending row had
no terminal path, so an accumulated backlog (staging: 3,734 pending,
oldest 22h) could only drain through 2 serial workers and effectively
never did.
The sweep now also expires pending jobs older than expiryPeriods x
their family's production period (default 3 periods; the family has
enqueued fresher batches since, so the stale job's result would be
ancient data). Each stale row is first CLAIMED via the S0 CAS under a
synthetic stale-pending-sweeper id — so the delete can never race a
worker (exactly-one-winner) — then deleted; a failed delete self-heals
via normal lease expiry + re-queue. Unparseable created timestamps are
conservatively skipped (delete is destructive). Policy is configurable
via FleetQueueClientConfig.stalePending; the control-plane wires the
real per-family cadences (FLEET_FAMILY_PERIODS_MS: d4/d5 15min, d6/
e2e-demos hourly) so 15min families expire on a 45min window. No comm
error is synthesized for expired-pending rows (they never ran); the
count surfaces as SweepResult.expiredPending and in the producer's
sweep log.
Producers enqueued a fresh batch every scheduled tick regardless of
whether the family's previous batch had even been claimed — against 2
serial browser workers the queue compounded without bound (staging:
3,734 pending), feeding the claim-page starvation.
A scheduled tick now skips its family's batch when that family already
has pending (unclaimed) jobs, bounding the per-family backlog to one
batch, with a structured fleet.producer.skipped-for-backlog log
(family, pendingCount, skippedJobs) and a skippedForBacklog count on
TickResult. The check is per family, fails OPEN on a count blip (a PB
read failure must never stop production), and is BYPASSED by
operator-triggered ticks (explicit intent wins; the trigger CLI treats
0 enqueued as failure). Backed by a new
FleetQueueClient.countPendingForFamily (server-side totalItems count of
a family's pending rows).
claimNext listed ONE global oldest-50 pending page; with the d4+d5
producers ticking every 15min against 2 serial browser workers, a
persistent backlog permanently saturated that page and e2e-demos jobs
never entered the candidate set (prod: all 18 e2e-demos jobs pending
forever; staging: 3,734 pending, oldest 22h).
claimNext now discovers the distinct families present in pending
(oldest first, one perPage=1 query per family) and tries them in
round-robin rotation across calls, listing a per-family candidate page
for each. Every discovered family is attempted before giving up, so no
family starves while any of its jobs are claimable. The S0 CAS
exactly-one-winner semantics and the per-page anti-herd shuffle are
unchanged; only candidate SELECTION changed.
## Summary
End-to-end hardening of the showcase shell's URL plane:
- Carries `backendHostPattern` + `docsHost` in the shell runtime config
(no longer baked from `registry.json` at Docker build time). Derives
demo backend URLs at runtime from the pattern, so a new pattern via env
reconfigures every integration on the next deploy without a registry
rebuild.
- Issues docs-host 301/308 redirects from middleware with a runtime
`DOCS_HOST`; misconfigured values no longer 500 every docs route — they
fall back to a sentinel that disables the docs-redirect step.
- Hardens the redirect table builder + matcher: first-match-wins for
duplicate exact sources, deduped wildcard prefixes with warn, malformed
entries rejected at lookup-build time, case-insensitive matching parity,
trailing-slash normalization, structural `/integrations` namespace guard
above the docs-host redirect (closes R15/R17 hijack class).
- Hardens runtime-config + backend-url env readers:
scheme/whitespace/control-char normalization, query/fragment/userinfo
rejection on `DOCS_HOST` / `POSTHOG_HOST` / backend-host pattern /
local-override URLs, present-but-empty `posthogKey` rejection, prod
loopback `BASE_URL`/`DOCS_HOST` rejection (no silent `http://` prepend),
once-guarded FATAL logging (no per-request spam).
- Brings the build-time twin in `showcase/scripts/generate-registry.ts`
to parity with the runtime normalizer (scheme/trailing-slash strip,
degenerate fallback, `NEXT_PUBLIC` fallback, slug validation).
- Surfaces PostHog capture failures once per failure class; keeps
capture alive across the redirect via `event.waitUntil`; missing
`POSTHOG_KEY` is `console.error` in production and surfaces at
config-resolution time (not first redirect).
- Open-redirect hardening on `/shared//evil.com` (SU-18); `//` rejected
at both source and destination; root `/` exact-source and `/:path*`
wildcard sources rejected; non-printable-ASCII source/destination
rejected.
Stream: SU (shell-runtime-urls). Subject groups
SU-2/8/11/13/14/15/16/17/18/19/20 (initial), SU2-A/B (env-hardening),
CR2-C (test infra), SU5-A1..A7 (registry/builder/lint/matcher boundary),
SU6-A1..A6 (request-time normalization parity), SU6-B1..B7
(parsed-normalized return forms, SHOWCASE_LOCAL states, generator
`{slug}` validation), SU7-F1..F3 (final-round backend-pattern + POSTHOG
+ table + script-side parity).
## Test plan
- [x] `pnpm test` in `showcase/shell` — 7 files, 235/235 pass
- [x] `pnpm test` in `showcase/scripts` — 51 files, 1857/1857 pass (one
pre-existing flake in `generate-registry.test.ts` unrelated to this
diff; passes on subsequent runs, classic vitest fork-reuse cross-file
pollution)
- [x] `pnpm exec tsc --noEmit` in `showcase/shell` — clean
- [x] `oxlint` on every changed file — 0 warnings, 0 errors
- [x] `oxfmt --check` on every changed file — clean
- [x] `pnpm build` in `showcase/shell` (full Next 15.x build with
registry+demo-content+starter-content+search-index generation) — clean
- [ ] CI green on the PR — confirm via `gh pr checks` after push
## Summary
Anti-dual-writer defense for the showcase `status` collection — the
flap-comb incident class where the legacy monolith scheduler and the
fleet writer fight over the same rows.
- **Writer identity columns**: new `written_by` / `state_written_at`
columns on `status` (two PB migrations, idempotency-symmetric up/down
paths). Every write stamps a stable host-derived writer identity; the
legacy orchestrator wires `writtenBy: "legacy"` explicitly.
- **Flip / foreign-write detection**: the status writer detects
cross-writer state flips inside a validated window and warns
(TTL-deduped re-warns so sustained fighting stays visible; one-time warn
when the self-write memory cap starts evicting; same-identity replica
hint).
- **PB-safe date normalization**: `observedAt` is normalized to RFC-3339
PB-safe shapes before any date-field write (zone-less shapes, colonless
offsets rejected) — no PB 400 unlatch-retry hot loops.
- **Aggregator honesty**: padded/blank projected keys normalized or
skipped loudly; explicit outcome discriminators replace ambiguous
asserts (no consumer hot-loop); `droppedCommError` surfaced whenever a
comm error misses the aggregate row; trusted-negative duplicate handling
scoped to cell-vs-cell so no cell can impersonate an aggregate row.
- **CLI persistence honesty**: thrown driver errors are persisted to PB
in the `runDriverInputs` catch; the summary carries `WriteOutcome`
discriminators (dropped-write counts, correct pluralization); writer
error classification documented honestly (401 vs 403 split, new
`pb_not_found` reason).
- **Alert/fixture truthfulness**: the alert engine's synthesized cron
outcome is stamped `persisted: false`; alert/orchestrator/probe test
fixtures align with the real `StatusWriter` / `OverlayWriteOutcome`
contracts, with fail-loud fake-PB fixtures hardened against silent
divergence from real PocketBase.
Hardened over 10 CR rounds (~60 reviewer reports), converged to zero
findings.
## Test plan
- [x] Full harness vitest suite: 124 test files / 2378 tests passing
- [x] `tsc --noEmit` and build typecheck (`tsconfig.build.json`) clean
- [x] oxfmt clean on all touched files; oxlint 0 errors
- [x] Known flake note: `probe-invoker.test.ts` wall-clock assertion
(elapsed 101ms vs <100ms bound) fired once under full-suite load and
passes in isolation (67/67) — pre-existing timing sensitivity, unrelated
to this diff
SU7-F1 — backend host pattern hardening:
- F1.1 Reject bare trailing ?/# in the backend host pattern
- F1.2 Strip internal tab/CR/LF from the backend host pattern
- F1.3 Warn when ignoring an empty-string local backend override
- F1.4 Reject empty-userinfo @ in the backend host pattern authority
- F1.5 Keep __proto__ keys as data in local-backend maps
- F1.6 Commit the local-backends memo key only after the value computes
- F1.7 Trim local backend overrides before validation and name the real
rejection
- F1.8 Honest FATAL when the pattern host is a stray scheme fragment
- F1.9 Canonicalize the pattern authority for parity with the override
path
- F1.10 Acknowledge the staging-to-prod fail-open in the pattern fallback
- F1.11 Harden backend-url/local-backends-env test hygiene
SU7-F2 — runtime-config & client-config edge cases:
- F2.1 Branch POSTHOG_HOST rejection reasons (scheme/degenerate/parse-
failure) instead of the catch-all mislabel
- F2.2 Reject loopback BASE_URL/DOCS_HOST in production instead of the
silent http:// prepend
- F2.3 Key the DOCS_HOST fallback once-guard on (mode, shellHost, value)
and mode-prefix all value-only guard keys
- F2.4 Reject a present-but-empty posthogKey in the client config reader
- F2.5 Drop the trailing slash from SSR_PLACEHOLDER_URL for structural
parity with server values
- F2.6 Attribute the DOCS_HOST slash-strip to readDocsHost itself
- F2.7 Normalize trailing-dot FQDN spellings in the docs self-host loop
guard (both compare sides)
- F2.8 Harden console spies to capture all log args; pin the full all-env
config shape; converge SSR simulation on vi.stubGlobal
SU7-F3 — script-side parity, table classification & test isolation:
- F3 #1 Handle a missing reference integration per the error contract
- F3 #2 Port the runtime backend-host-pattern normalization into the
generator — scheme/trailing-slash strip, degenerate fallback,
NEXT_PUBLIC fallback
- F3 #3 Treat non-mapping manifest parses (empty/null/scalar/array YAML)
as validation errors, not TypeErrors
- F3 #4 Label a missing/unreadable constraints.yaml per the stderr+exit(1)
error contract
- F3 #5 Align atomic-write tmp naming with the test harness straggler-
sweep convention; guard main() on direct invocation
- F3 #6 Correct the determineCellStatus unshipped docstring; replace
stale hardcoded cell counts with formulas
- F3 #7 Isolate the pattern suite on a per-suite tmpdir harness; snapshot
the generator's full write set
- F3 #8 Classify discarded duplicate wildcards as duplicates — hoist the
owner check above the destination warns
- F3 #9 Reject a root ("/") EXACT seo-redirect source — homepage-hijack
twin of the root-wildcard guard
- F3 #10 Reject seo-redirect entries with non-printable-ASCII source/
destination — close the silent-dead-entry class
- F3 #11 Strip trailing slashes in normalizePosthogHost before the scheme
test
- F3 #12 Message-filter the empty-slug-set error count; pin the single
matcher entry
Round-by-round CR convergence covering the redirect builder, the middleware
matcher, the docs-host self-loop guard, and the runtime-config env readers.
Highlights:
- Clear module-load warns after fresh middleware import
- Validate SET BASE_URL values (scheme-less/degenerate/garbage) with
sentinel fallback + once-guarded FATAL log
- Normalize path/query/fragment-bearing DOCS_HOST to origin; reject
non-http(s) schemes; branch rejection reasons
- Harden POSTHOG_HOST (degenerate-host/scheme rejection); expose
posthogKey via readEnvPair semantics
- Reject a DOCS_HOST equal to the shell's own host (redirect-loop guard,
authority compare)
- Warn on missing local-ports.json under SHOWCASE_LOCAL=1 and validate
TCP port range; extract helper for tests
- backend-url hardening — slug charset guard, frozen local-backends memo,
pattern path-segment warn, local-override URL validation
- Client config fail-loud covers all four URL fields with type checks
- Make RuntimeConfig.posthogKey optional — absence is a valid state, not
a wiring bug
- Drop R15/R17 and guard /integrations from SEO redirects
- Dedup duplicate wildcard prefixes with first-match-wins warn
- Validate malformed SEO entries at lookup-build time
- Restore case-insensitive redirect matching parity
- Normalize trailing slashes before redirect matching
- Keep the framework segment on F13, pin MG3 case fix
- Read posthogKey from runtime config in middleware, not raw process.env
- Fall back to the default backend host pattern for degenerate values
- Disable docs redirects when the default fallback collides with the
shell host
- Bring validateBaseUrl to parity with its sibling readers
- Strip query/fragment from POSTHOG_HOST while keeping reverse-proxy paths
- Restrict local backend overrides to http(s) URLs
- Add server-only guard to runtime-config
- Harden localBackendsEnv failure posture
- Hoist /integrations namespace guard above the docs-host redirect
- Validate seo-redirect sources and cross-kind shadowing in
buildRedirectLookup
- Unify slash normalization for middleware matching
- Lowercase-normalize REGISTRY_FRAMEWORK_SLUGS at construction
- Escalate missing POSTHOG_KEY to console.error in production
- Skip all redirect steps when docs redirects are disabled (sentinel
consumer)
- Reject userinfo credentials in DOCS_HOST, POSTHOG_HOST, and the backend
host pattern
- Branch dev-vs-prod logging in readDocsHost and fatalPatternOnce
- Prepend http:// (not https://) to scheme-less loopback hosts
- Round-5 micro-finding batch across the URL config libs
SU5-A1..A7 — registry safety, // reject, builder lint batch (case-
insensitive :path*, same-destination twin allowlist, original-case
divergence remainder), matcher api boundary, generator+vitest infra, test
hygiene + empty docs-host guard, comment batch.
SU6-A1..A6 — reject miscased :path* tokens, warn on tokenless wildcards,
normalize redirect-destination comparisons like request time, reject
destinations containing "//", surface missing POSTHOG_KEY at config-
resolution time, compile matcher harness like Next's runtime, type
parse/tokensToRegexp in the path-to-regexp shim, keep buildRedirectLookup
JSDoc attached.
SU6-B1..B7 — reject query/fragment/userinfo in pattern and local-override
URL gates, return parsed-normalized URL form from validation success
paths, distinguish unset/blank/padded SHOWCASE_LOCAL states, warn when
SHOWCASE_LOCAL is set to a value other than 1, validate {slug} placeholder
in generate-registry, mirror middleware drop semantics in the wiring
test's registry re-derivation, pin the noStore spy and calls to one fresh
module instance in the Edge-path test.
SU2-B series — runtime-config / backend-url env robustness:
- Correct the Edge-safety story in runtime-config (SU2-B1)
- Stop per-request FATAL-CONFIG spam for unset BASE_URL (SU2-B2)
- Prepend https:// to a scheme-less POSTHOG_HOST (SU2-B3)
- Trim whitespace paste artifacts in env values and host patterns (SU2-B4)
- Memoize parseLocalBackends and warn once per value (SU2-B5)
- Make {slug} substitution immune to $-patterns (SU2-B6)
- Harden the client runtime-config reader (SU2-B7)
- runtime-config hardening batch (SU2-B8)
- Validate local-ports.json before baking NEXT_PUBLIC_LOCAL_BACKENDS (SU2-B9)
- test: warn-once assertions retry-safe; stop console leaks (SU2-B10)
CR2-C series — test infrastructure:
- Generate registry.json in a vitest globalSetup (CR2-C1)
- Stop ambient POSTHOG_KEY firing real fetches in middleware tests (CR2-C2)
- Assert the production slug set, not a re-derivation (CR2-C3)
- Make the registry generator subprocess robust (CR2-C4)
- Middleware/wiring test hygiene batch (CR2-C5)
SU2-A series — redirect-layer & PostHog capture:
- Stop $-pattern expansion in wildcard redirect substitution (SU2-A1)
- Surface PostHog capture failures once per failure class (SU2-A2)
- Duplicate exact redirect sources are first-match-wins (SU2-A3)
- Resolve runtime config once per redirected request (SU2-A4)
- Include destination host in seo_redirect capture (SU2-A5)
- Normalize scheme-less POSTHOG_HOST at the capture use site (SU2-A6)
- Correct redirect-layer comments and guard wildcard prefix boundary (SU2-A7)
- Cover docs-host hardening branches, compile matcher via path-to-regexp (SU2-A8)
Resolve SEO redirect destinations against the docs host (SU-17); forward
the query string on SEO redirects (SU-16); match bare paths on wildcard
SEO sources (SU-19). Collapse duplicate slashes in docs-host redirect
destinations (SU-13). Regression test for /shared//evil.com open redirect
(SU-18). Emit 308 for docs-host redirects to match next.config parity
(SU-2). Add a path boundary to the api matcher exclusion (SU-15). Loud
guard when registry yields zero framework slugs (SU-20). Keep PostHog
capture alive via event.waitUntil (SU-14). Note docs-host redirects are
untracked by design (SU-8). Cover docs-host redirects at the middleware
level (SU-11).
Squash of the SEO-table + matcher-hardening cluster.
Carry backendHostPattern + docsHost in the shell runtime config (no longer
baked from registry.json at Docker build time). Derive demo backend URLs at
runtime from the pattern; issue docs-host 301s from middleware with a runtime
DOCS_HOST so a misconfigured value can no longer 500 every docs route.
Validate NEXT_PUBLIC_LOCAL_BACKENDS and empty overrides; guard the backend
host pattern against silent env misconfigs. Reword the stale demo-page
comment about backend URL derivation. Pin the registry slug set and SSR
placeholder URL composition; fix env/spy/global leaks in runtime-config
test cleanup.
Squash of the initial runtime-URL refactor cluster:
- feat(showcase): carry backendHostPattern + docsHost in shell runtime config
- fix(showcase): derive demo backend URLs at runtime instead of baked registry values
- fix(showcase): issue docs-host 301s from middleware with runtime DOCS_HOST
- fix(showcase): never let a misconfigured DOCS_HOST 500 every docs route
- fix(showcase): guard backend host pattern against silent env misconfigs
- fix(showcase): validate NEXT_PUBLIC_LOCAL_BACKENDS values and empty overrides
- docs(showcase): reword stale demo-page comment about backend URL derivation
- test(showcase): fix env/spy/global leaks in runtime-config test cleanup
- test(showcase): pin registry slug set and SSR placeholder URL composition
Blur the Slack guide (/slack and its framework-scoped variants) and the
entire bot reference section behind a client-side unlock card. The card
follows the scroll in a sticky scrollport-height frame, persists unlock
state in localStorage, and leaves the sidebar and top nav usable. Doc
pages opt in via earlyAccess frontmatter; the bot reference gates on
version === "bot". Visitors without the password get a product shot
(light/dark variants in git LFS, shown beside the form on wide cards
via container query) and a link to the beyond-the-web early-access
form.
Wire the legacy monolith scheduler's status writer with an explicit
writtenBy:"legacy" identity, correct the dual-writer dedupe comments
to describe the real upsert-collapse (and why it is not safe), stamp
the alert engine's synthesized cron outcome persisted:false honestly,
and align alert/orchestrator/probe test fixtures with the real
StatusWriter and OverlayWriteOutcome contracts.
Persist thrown driver errors to PB in the runDriverInputs catch,
carry WriteOutcome discriminators through the CLI summary (dropped
write counts with correct pluralization, no duplicate cause
clauses), and pin the runner/results behavior with StatusWriter
contract-typed stubs so writer contract drift is compile-checked
at the stub site.
Normalize padded projected keys to trimmed canonical form, skip
blank keys loudly, and replace ambiguous outcomes with explicit
discriminators: honest outage-skip/empty-projection semantics,
droppedCommError surfaced whenever the comm error misses the
aggregate row, trusted-negative duplicate preference scoped to
cell-vs-cell (no aggregate-row impersonation), and comm-error
identity asserts converted to discriminators so the consumer
cannot hot-loop.
Pin the writer-identity stamping, flip/foreign-write warn paths
(TTL re-warns, self-write memory cap eviction), date normalization
shapes, overlay write outcomes, idempotency latches, and error
classification against fail-loud fake-PB fixtures hardened against
silent divergence from real PocketBase behavior.
Add written_by/state_written_at columns (PB migrations) and stamp every
status write with a stable host-derived writer identity. The status
writer now detects cross-writer state flips and foreign writes (the
anti-dual-writer flap-comb defense), normalizes observedAt to PB-safe
RFC-3339 shapes before date-field writes, and classifies writer errors
honestly (401 auth vs 403 permission split, new pb_not_found reason).
## Summary
The beautiful-chat demo's "Calculator App (Open Generative UI)" pill
rendered a calculator whose "=" key was inert: the fixture's
`jsFunctions` called `Websandbox.connection.remote.evaluateExpression`,
a host-bridge function only the open-gen-ui-advanced demo registers —
beautiful-chat never does, so the call could never resolve. While fixing
it, review and live verification surfaced four more behavioral defects
in the same fixture family, all fixed here across all 18 integrations:
- **Self-contained evaluation**: `jsFunctions` now evaluates in-sandbox
via an allowlist-gated strict-mode `Function()` (digits/operators/`eE`
only; failures → `err`). No host bridge required.
- **Fixture shadowing**: the specific "with standard buttons" pair is
ordered before the generic "build a modern calculator" pair (aimock is
first-match-wins by load order), so the intended fixture actually serves
the pill.
- **Chained evaluation**: results in exponential notation (e.g.
`1.728e+18`) re-evaluate instead of wiping to `err` (regex allows `eE`).
- **Interleaved pills**: removed the thread-global `hasToolResult` gate
that made the calculator fall through to the live proxy (502) when
clicked after any tool-producing pill; toolCallId-anchored follow-ups
(ordered first) disambiguate legs instead.
- **Repeat clicks**: distinct `tool_call_id`s per click via
`sequenceIndex` variants + a non-sequenced fallback, fixing the
second-widget collapse (duplicate id collapsed both renders into one
slot, wiping state). Scoping caveats (per-X-Test-Id counters,
cross-integration co-increment, DEFAULT_TEST_ID degradation floor =
pre-fix behavior) are documented in the fixture comments and GOTCHAS.
Also: SUPERSEDED annotations on the unreachable recorded.json calculator
entries, GOTCHAS corrections (sequenceIndex scoping, hasToolResult
semantics, statelessness claim), and a routing-invariant unit test
pinning the entry ordering structurally and behaviorally for all 18
integrations (55 tests).
## Test plan
- [x] Local Playwright red-green on the built stack (langgraph-python):
broken "=" reproduced pre-fix; post-fix visual proof of render,
`7+8=15`, chained exponent math, calculator-after-dashboard interleave,
and two independent working widgets across repeat clicks
- [x] aimock version bisect (1.28.0/1.29.0/1.30.0) ruling out an aimock
regression before the fixture root-cause
- [x] `showcase/scripts` vitest suite: 51 files / 1899 tests green
(incl. new `calculator-fixture-routing.test.ts` 55/55, red-proofed via
mutation)
- [x] validate-parity 19/19; harness aimock-fixture-coverage 3/3;
oxfmt/oxlint/commitlint clean
- [ ] CI green on PR HEAD
Follow-ups (readonly-state matcher shadowing parity across 15
integrations, sibling-pill interleave gates, GOTCHAS accuracy pass,
aimock sequenceIndex scoping) are tracked in the showcase follow-up
ledger.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
- structural + behavioral pins for all 18 integrations (four ordering invariants; click walk
_001->_002->_003->_003 with follow-ups mirroring the server match/increment flow)
- fail-loud guards against loader error-swallowing and vacuous ordering passes
- full production pill text
- oxfmt applied
- SUPERSEDED annotations on unreachable recorded calculator entries (load-order shadowing is
the only guarantee; model gate is not a safety net)
- GOTCHAS sequenceIndex rewritten to per-X-Test-Id semantics with co-increment/eviction caveats
- hasToolResult paragraph corrected (omission = no gate, thread-global predicate)
- statelessness claim reconciled with sequence counters
- replace Websandbox host-bridge evaluateExpression (never registered in beautiful-chat) with
in-sandbox allowlist-gated Function() eval
- reorder specific calculator pair before generic (first-match-wins)
- allow exponential notation (eE) so chained evaluation of large results works
- drop thread-global hasToolResult gate that broke the pill after other tool-producing pills,
reorder toolCallId follow-ups before leg-1
- mint distinct tool_call_ids per repeat click via sequenceIndex variants + non-sequenced
fallback (fixes second-widget collapse)
- restore google-adk trailing newline
- document ordering invariants, gate tradeoffs, and sequence-counter scoping caveats in
fixture comments
Verified via local Playwright red-green (render, 7+8=15, chained exponent, interleaved pills,
repeat clicks).
- New Platforms entry: /platform/slack quickstart — manifest-based app
creation, Socket Mode tokens, minimal createBot bot run with tsx,
interactive JSX with inline onClick, slash commands, production split
- New "Bots" SDK tab in the reference picker with per-symbol pages for
@copilotkit/bot, @copilotkit/bot-ui, and @copilotkit/bot-slack
(Components / Functions / Classes / Types)
- Rename reference picker labels to React (V2) / React (V1)
- Remove the retired /reference/sdk pages (LangGraph/CrewAI SDK,
Remote Endpoints); search/sitemap/llms indexes derive from the
content tree, so they de-index with the deletion
- Retarget the one inbound link to its /reference/v1 copy
Co-Authored-By: Claude <noreply@anthropic.com>
Closes [OSS-299](https://linear.app/copilotkit/issue/OSS-299).
Follow-up to #5248 (already merged).
## Problem
The `hero_command_copied` PostHog event added in #5248
(`showcase/shell-docs/src/components/hero-start-commands.tsx`) carries
no surface discriminator. `HeroStartActions` renders on **both** the
home hero and **every** framework landing hero:
- The **create** card embeds the framework in `command` (`--framework
langgraph-js`), so it's recoverable.
- The **onboard** card's command (`npx copilotkit@latest skills
onboard`) is byte-identical on every page — so onboard copies **cannot**
be attributed to a surface from the event alone.
Every sibling event in shell-docs already carries a "where" property —
`cli_command_copied` → `location: window.location.pathname`, the nav
events → `location`, `markdown_copied`/`open_in_llm_clicked` → `path`.
`hero_command_copied` was the only one without one.
## Fix
Add `location: window.location.pathname` to the `hero_command_copied`
payload, mirroring the `cli_command_copied` event the global
`<CopyTracker>` already emits for the same copy (verified: it
monkeypatches `navigator.clipboard.writeText`, which the hero calls).
The two paired events now join cleanly on the same dimension. Guarded
for SSR (`typeof window !== "undefined"`) to match the sibling.
## Test
Adds a colocated source-assertion guard test. shell-docs vitest runs in
the `node` environment (no jsdom/RTL), so this follows the suite's
existing convention (`readFileSync` + assertions, like
`brand-nav.test.tsx`) rather than introducing a behavioral render
harness.
```
✓ src/components/__tests__/hero-start-commands.test.tsx (3 tests)
```
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Pin alignment fixes 9 validate-pins FAILs; ratchet the drift baseline
count and hash accordingly. Also tighten the _comment: document the exact
hash recipe (SHA-256 of the stderr-only [FAIL] lines, LC_ALL=C sort -u)
and correct baselineDemoCount semantics (exact expected demo count per
package; deviation either direction warns).
Align showcase integration requirements.txt files (strands,
langgraph-fastapi, langgraph-python, pydantic-ai, google-adk,
crewai-crews) to the fleet pin standard, including an accurate
typing_extensions comment in crewai-crews and a trailing newline in
langgraph-python.
Replace floating "beta" dist-tags with exact versions for @ag-ui/mastra,
@mastra/{client-js,core,libsql,memory}, and mastra in both the examples
and showcase mastra packages. Showcase mastra also raises its zod floor
^3.24.0 -> ^3.25.0. The examples mastra package additionally carries the
fleet-wide @ag-ui/client 0.0.55 bump and single-tree overrides here, since
its manifest mixes both changes.
The hero_command_copied event fired by the landing-hero command cards carried
no surface discriminator. HeroStartActions renders on both the home hero and
every framework landing hero; the "onboard" card's command is byte-identical
on every page, so onboard copies could not be attributed to a surface from the
event alone (only the "create" card embeds the framework in `command`).
Add `location: window.location.pathname` to the payload, mirroring the
`cli_command_copied` event the global <CopyTracker> already emits for the same
copy so the two paired events join on the same dimension. Guarded for SSR to
match the sibling.
Adds a source-assertion guard test in the shell-docs node-env convention.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
## What
Replaces the landing-page CTA with **three entry points**, framed by
situation, and renders the **identical action block on the home hero and
every framework landing hero**:
| | action |
|---|---|
| **New project** | `npx copilotkit create` |
| **Existing project** | `npx copilotkit skills onboard` |
| **Guided walkthrough** | **Quickstart** button (preserved from the
previous hero) |
- **Unified `<HeroStartActions>` block**: two equal-weight command cards
plus a quickstart row beneath, shared verbatim by the home hero and the
framework landing heroes (per review: the two surfaces previously
diverged).
- **Quickstart preserved** in its original accent treatment. On the home
hero it is the framework-picker dropdown (`<HeroQuickstartDropdown>`,
restored); on framework pages it links straight to that framework's
quickstart guide. The home hero also keeps the "Learn more about
building with agents" link in the same row.
- **Framework landing heroes** (e.g. `/langgraph-typescript`): the
create command **pre-fills the framework** via the CLI's `--framework`
flag (e.g. `--framework langgraph-js`).
**Framework-flag mapping**: docs slug to CLI `--framework` value,
verified against the CLI's `AGENT_FRAMEWORKS` enum
(`langgraph-typescript`→`langgraph-js`,
`langgraph-python`→`langgraph-py`, `google-adk`→`adk`,
`strands`→`aws-strands-py`,
`ms-agent-dotnet`→`microsoft-agent-framework-dotnet`, identical for
`mastra`/`pydantic-ai`/`llamaindex`/`agno`/`ag2`). Slugs with **no** 1:1
CLI template fall back to a bare `npx copilotkit create`, notably
`crewai-crews` (the CLI ships *CrewAI Flows*, not Crews), plus
`langgraph-fastapi`, `claude-sdk-*`, `langroid`, `spring-ai`,
`agent-spec`, `deepagents`. `skills onboard` has no framework flag, so
it is identical everywhere. Frameworks with bespoke setup (`a2a` `git
clone`, `ms-agent-dotnet`) keep the pre-cards layout: quickstart button
plus their own copy-command chip.
**Responsive, with all text always visible.** Commands **wrap, never
truncate**:
- Wraps happen at spaces only; every token is non-breaking, so
`--framework` can never split into a dangling `-` at a line edge.
- `text-wrap: balance` splits multi-line commands evenly, typically
right at the flag boundary (`npx copilotkit@latest create` /
`--framework langgraph-js`).
- The block caps at 740px with 12px mono, the narrowest cap where both
home commands fit one line with enough headroom to survive platform
mono-font width differences.
- Cards sit two-up from `sm` and stack below it; the grid (`min-w-0`,
`items-stretch`) keeps long commands inside their track and the card
pair equal-height.
## Screenshots
**Home**: two cards, quickstart dropdown, learn-more link

**Home, quickstart dropdown open** (framework picker preserved)

**Framework landing (LangGraph)**: same block, framework pre-filled,
create command balanced across two lines, quickstart links to the guide

**Worst case (Microsoft Agent Framework, Python)**: longest CLI flag
value, three balanced lines, fully readable

**Bespoke setup (A2A)**: quickstart button plus own command chip
(pre-cards layout preserved)

**Mobile (375px)**: cards stack, quickstart goes full-width
| home | framework |
|---|---|
| 
| 
|
## Telemetry
Both hero copy buttons are now explicitly instrumented: each click
captures **`hero_command_copied`** (`command_id`: `create` | `onboard`,
full `command` string, `clipboard_blocked`), so create-vs-onboard
funnels are queryable per landing page. The pre-existing global
`cli_command_copied` (fired by `CopyTracker` on any clipboard copy)
still fires for volume metrics; the new event uses a different name so
that funnel is not double-counted. Validated locally against a live
PostHog client: each click POSTs both events (plus `$autocapture`) to
`/ingest/e` with HTTP 200.
## Notes
- Both cards equal weight; accent only on hover. Copy rows copy on click
with `aria-live` feedback plus a clipboard-blocked fallback; cursor is
`pointer`.
- Removes `agent-start-prompt.tsx` and `hero-command-copy.tsx`.
`hero-quickstart-dropdown.tsx` is back (restored unchanged after review
feedback).
## Summary
- Live staging redeploy evidence (2026-06-10 16:37Z): 2/6 workers
completed the full SIGTERM → abandon → deregister sequence in **under 1
second**, while 4/6 were SIGKILLed mid-browser-teardown because
Railway's ~10s stop grace is shorter than the old 25s drain budget —
leaving 4 stale roster rows and a reclaim splash on every deploy.
- This PR makes abandon + deregister the **guarded, sub-second critical
path** and demotes teardown to best-effort within a composed <10s
budget, so a platform kill mid-teardown is harmless.
## Design
- `drainFleetWorker` ordering: drain → `registration.stop` → bounded
deregister → graced `worker.stop` → always-run pool shutdown, with
stop-error precedence (a pool-shutdown failure can never mask the stop
error).
- `DRAIN_DEREGISTER_TIMEOUT_MS` (3s) bounds the **whole registration
write chain**, so a hung—not failing—PocketBase cannot consume the kill
window; timeout degrades to the documented crash-path reclaim.
- `safeLog` guards every loop/stop/drain-path log: a throwing logger can
neither reject the worker loop's done-promise nor skip the roster delete
or teardown (abort-before-log in `requestDrain`; structural-caller
guards in `drainFleetWorker`).
- Drain-aware lease renewal (an abandoned job's lease lapses instead of
being re-extended), mid-drain claim skip (a claim won after the drain
decision is never run), and mid-report precision (a run that began
reporting is never logged as abandoned).
- Never-throws loop closure: loop-crash logging via `done.catch` +
`/health` 503, heartbeat and idle-poll sleep hardening with a non-busy
pacing floor, aggregate-key protocol-violation wrap.
- `WORKER_DRAIN_GRACE_MS` default 25s → 6s; the composed 3+6 < 10s
budget is **pinned by a test**; present-but-invalid overrides warn;
overrides at/above ~7s are documented as forfeiting the composed budget.
- Boot-failure teardown catches now log (no silent chromium stranding).
## Review
- 6 unbiased 7-agent CR rounds + 5 fix rounds; every behavioral change
red-green or mutation-proven; `Promise.race` loser semantics empirically
pinned by test.
- ~30 pre-existing harness findings deferred to the flap-fix follow-up
backlog (top of the next fleet-robustness PR: lease-renew
retry-on-throw, empty-registry guard/dispatch mismatch,
`registered`-flag refresh, worker `/health` async bind race, queue fetch
timeouts).
## Test plan
- [x] 2176/2176 vitest (32 new tests)
- [x] `tsc --noEmit` both configs
- [x] oxfmt clean
- [ ] CI green on this PR
🤖 Generated with [Claude Code](https://claude.com/claude-code)
## Summary
Hardens the `--isolate` showcase verification flow across three areas:
**1. XDG state migration.** Isolate slot registry and per-run
rewritten-compose scratch dirs move off `/tmp` (wiped on reboot,
world-writable) to
`${XDG_STATE_HOME:-$HOME/.local/state}/copilotkit/showcase/` (`slots/` +
`runs/<name>/`). `/tmp` clearing silently destroyed a kept stack's
compose file and slot, making `--keep` unreliable. Run dirs are keyed by
the finalized project name (not PID) so a kept run is locatable for
manual teardown.
**2. Slot reaping + registry concurrency.** Since the state dir is now
persistent, slots are reaped by compose-project liveness (`docker ps
--filter label=com.docker.compose.project=<name>`), with PID/age
heuristics as fallback. The registry is made safe under concurrent
claimers: a sweep lock with heartbeat updates, own-pid lock release, and
tombstones; a claim-then-verify duplicate-name guard closing the TOCTOU
window; crash-safe reap ordering with compose-down of reap remnants and
a path-traversal guard. Failed `--isolate` setup no longer tears down
the default stack; half-initialized state is cleaned up on the way out.
Teardown uses `--volumes` everywhere, and a failed compose-down
preserves state for diagnosis. `--isolate` names are validated (must
start with lowercase letter/digit; `showcase` is reserved — it aliases
the default stack), and a fail-loud warning precedes pre-down of an
existing stack.
**3. `--keep` now actually persists an isolated stack.** Previously the
unconditional `trap restore_isolation EXIT` tore the stack down
regardless of `--keep`. Teardown is now gated on the keep flag
(`ISOLATE_KEEP` promoted to a global so it survives `cmd_test` return
into the trap scope): the slot + run dir are retained and a survival
notice prints the project name, the three offset host ports, and the
exact `docker compose -p <name> down` command — no silent port/slot
leak. A kept stack's live containers keep its slot from being reaped.
Shell-only — confined to `showcase/scripts/cli/_common.sh` +
`cmd-test.sh`; the harness TS only reads the env vars the shell exports
(unchanged). Follows up the `--keep` caveat documented in #5346.
## Review hardening
The branch went through an 8-round, 7-agent code-review loop with
red-green-verified fixes — that loop produced the state-machine
hardening commit (trap-scope fix, default-stack guards, registry
concurrency/teardown robustness, name validation) and grew the test
suite to pin every fix. A live end-to-end `--keep` verification run is
what surfaced the trap-scope bug (`--keep` silently not honored),
driving the `ISOLATE_KEEP` global fix.
## Test plan
- [x] `showcase/scripts/__tests__/isolate.bats` — 41 isolate tests
(red→green): XDG path resolution (+`XDG_STATE_HOME` override,
`~/.local/state` fallback, `runs/<name>`), liveness-based reaping (dead
project reaped/reclaimed, live project preserved), real-trap-path
`--keep` tests (no simulated-trap shortcuts), sweep/lock/tombstone race
pins (heartbeat resurrection, lock takeover, duplicate-name TOCTOU),
reap-order probe pinning live-slot protection, root/PID-reuse/DST
guards, and sentinel anti-vacuity discipline so trap tests cannot pass
vacuously.
- [x] Full `bats showcase/scripts/__tests__/` green, matching CI's Shell
script tests invocation.
- [x] shellcheck: no new warnings.
- [x] Live end-to-end: `bin/showcase test <slug> --d6 --isolate <name>
--keep` persists the stack under `~/.local/state/copilotkit/showcase`,
survival notice + manual teardown work, follow-up run reaps the stale
slot.
## Summary
- PR #5352's worker-side flap fixes never reached staging automatically:
`harness-workers` runs the same `showcase-harness` image as the
`harness` scheduler, but the SSOT's `ciBuilt: false` conflated "owns a
build slot" with "should be redeployed when its image is rebuilt" — so
main merges redeployed only the scheduler and the workers silently kept
running a stale image (a manual redeploy was required to ship the
fixes).
- This adds an `imageOf` field to the Railway SSOT so a rebuilt image
redeploys **all** of its consumers: the CI redeploy scope is now built
slots ∪ their `imageOf` consumers that declare the target env.
- Staging default scope becomes 27 (26 ciBuilt + `harness-workers` via
expansion); prod is unchanged at 26 (the worker is staging-only and the
expansion is env-aware).
## Design
- `imageOf: "<ssot-key>"` on consumer entries (`harness-workers` →
`harness`), enforced by a module-load invariant
`assertImageConsumersValid`: dangling targets, non-ciBuilt producers,
consumer chains, and consumer envs not a subset of the producer's all
fail loud at import; lookups are prototype-safe (`Object.hasOwn`).
- `expandImageConsumers` in `redeploy-env.ts` performs the env-aware,
single-level expansion and fails loud on unnormalized env names
(synonyms like `production` must go through `resolveEnv`) — the first
real consumer of `ENV_ID_BY_NAME`.
- Service-name resolution (`resolveTargetServices`/`runRedeploy`) now
rejects inherited `Object.prototype` keys with the proper
Unknown-service operator error.
- The explicit `--services` passthrough (a named service is attempted
even in an env it does not declare) is documented and contract-pinned by
a test.
## Review
- 5 unbiased 7-agent CR rounds plus a diff-attribution triage; every
diff-authored finding fixed with red-green proofs.
- ~30 pre-existing script-hygiene findings (env-registry consolidation,
accessor leniency, fetch timeout, parseArgs edges, coverage gaps in
`makeLiveRedeploy`/summary-JSON, etc.) deferred to the flap-fix
follow-up backlog.
## Test plan
- [x] 82/82 vitest (13 new tests: expansion, env-awareness, invariants
incl. prototype keys and env-subset, contract pins)
- [x] `tsc --noEmit -p showcase/scripts/tsconfig.json` clean
- [x] oxfmt clean on all changed files
- [ ] CI green on this PR
🤖 Generated with [Claude Code](https://claude.com/claude-code)
The webhooks SSOT comment claimed the push-driven default scope is
"guaranteed" to leave webhooks untouched — false: a push touching the
build workflow files trips the workflow_config paths-filter disjunct,
which selects every matrix slot (webhooks included; its skip_build slot
still reports success and enters the redeploy CSV). Reworded to state
the actual behavior. imageOf doc now states the enforced NON-EMPTY
subset constraint; serviceEnvPairs doc now truthfully says it has no
consumers yet; file header notes the probe flag default and
bin/railway's Ruby-only "stage" synonym.
Test hygiene: drop the stale bin/railway line-number citation from a
test name, the _envConfigTypeAnchor (EnvironmentConfig is genuinely
referenced by the shape-compile test), a dead eslint-disable, and a
dead `as never` cast (env is an open string); align the webhooks
dispatch-name pin regex with the extraction regex's whitespace
tolerance.