Both resolvers collapsed a family with a missing sub-row via
`worstState !== "red"` literal equality, silently swallowing
out-of-vocabulary states (e.g. "error") that the A2 worstStateRank
machinery deliberately ranks ABOVE red. Compare by rank instead so a
red-or-worse fold dominates no-data; worstState is typed State but can
hold raw strings at runtime, so the rank form is the honest comparison.
The harness FleetSurfaceState union had no "pending" member and the
derivation mapped EVERY comm error to the red "unreachable" overlay,
contradicting the worker-reclaimed-pending neutral-surface contract that
POOL_COMM_ERROR_KINDS mandates and the dashboard already implements.
Mirror the dashboard cell-model derivation exactly: reclaimed-pending on
a non-red row -> "pending"; a red row passes through unmasked; every
other kind -> "unreachable". Fix the stale binary-derivation docstring
and extend the cross-package drift test to pin the union overlay
members AND the derivation shape on both sides.
TickOptions.filter / EnumerateContext.filter were documented trigger-only
but tick() forwarded the filter unconditionally — a scheduled tick could
be scoped by a stray operator filter. The filter is now forwarded only
when triggered (dropped with a warn otherwise). ServiceJobSpec.priority's
'higher pulls first' doc described behavior that does not exist (claimNext
never reads it) — rewritten as reserved/not-consulted; no priority
behavior added.
nowMs was captured before the potentially-seconds-long enumerate() await,
back-dating lease-expiry decisions and lastSweepAt, and stamping
meta.enqueuedAt (documented 'ISO timestamp the control-plane enqueued the
job') with the tick-START time. The clock is now re-read immediately
before maybeSweep and enqueuedAt is stamped per job at enqueue time. The
default runId factory also reads the injected now() for
injection-discipline consistency (behavior otherwise identical).
Each producer gets an independent default runIdFactory with its counter
starting at 0, so two producers ticking in the same ms with equal tick
counts minted the SAME runId — and the aggregator groups results by
meta.runId. The factory now bakes in a per-factory random discriminator
segment (generated once at creation); ids stay sortable-prefixed by
timestamp.
The enumerate-throw path returned the same shape as a legitimately empty
run (enqueued: 0), so a discovery outage was indistinguishable from an
empty catalog in the tick outcome — the exact ambiguity class sweepFailed
was added to remove. TickResult now carries enumerateFailed, set on the
enumerate-throw path. (control-plane.test.ts fake TickResult gains the
new required field.)
A synchronously-throwing injected fetchImpl escaped the warm loop and
aborted the whole tick before any job was enqueued, violating the 'never
block or fail job production' contract — the per-spec dispatch is now
try/caught. The success handler also discarded the Response without
consuming it (an unread body pins the socket under undici) — the body is
now cancelled best-effort — and the abort timers are unref()'d (guarded)
so a pending warm timer can't hold the process open.
sweepExpired only synthesizes comm errors for rows reclaimed in that call,
so a transient onSweepCommErrors sink failure permanently dropped the
reclaimed jobs' dashboard signal (REQ-B violation) — the old comment's
'next sweep retries' claim was false. The producer now buffers undelivered
comm errors (capped at 500, oldest dropped with a warn) and prepends them
to the next sweep's sink delivery.
- decodePayload validated only half the payload: now asserts
meta.triggered (boolean), meta.enqueuedAt (string), cellIds (string[]
when present) and driverInputs (plain record when present) via a
shared assertServiceJobPayload, failing loud at the boundary.
- enqueue validates the payload BEFORE pb.create (it used to deref
payload.meta.runId after creating the row, persisting a poison row on
a malformed caller payload).
- report()'s refused-release message no longer claims 'nothing was
lost' — the computed result IS discarded and the job re-runs; the
message and comment now say so.
- emptyPayloadForLease carries an explicit HEARTBEAT-ONLY warning: the
placeholder must never feed aggregation (empty runId/serviceSlug
sentinels would corrupt grouping).
- Family discovery warns (queue-client.family-discovery-truncated) when
the 16-family bound trips with families still hidden — previously a
silent starvation; 17-family test.
- escapeLikeLiteral backslash-escapes %/_ for the ~/!~ legs (verified:
PB 0.22.21 builds LIKE ... ESCAPE '\' and skips auto-wrap when the
operand has an unescaped %, and fexpr passes backslashes verbatim), so
a family like 'd%' can no longer occlude d6:/d4: from discovery or
over-count in the backlog gate; = legs keep plain literal escaping.
The gate reads totalItems off a perPage=1 list but never passed
skipTotal. If totals are skipped PB returns totalItems: -1, which is
never above the producer's backlog threshold — the per-tick dedupe gate
would silently FAIL OPEN and enqueue fresh batches on top of an existing
backlog (the exact compounding the gate exists to stop). Pass
skipTotal: false explicitly; test pins the param.
payloadCache entries were only evicted in report()'s finally — a worker
whose renew loses the CAS (lease stolen/swept/terminal) never reports
that job, so its claim-time cache entry stranded forever and the
per-client map grew with every abandoned job. renewLease now evicts on
the lost-CAS return. Tests prove eviction by observing the later renew
take the convenience re-read path (cache miss → pb.getOne) for both the
renew-lost and the claim→report→renew sequences.
The decode-failure path released a won-but-poisoned job as 'failed' with
NO result — and per the file's own contract a terminal resultless row is
synthesized by the result consumer as worker-crashed-mid-job, painting a
FALSE red 'crashed' overlay for what is a payload/protocol problem. The
taxonomy already has worker-protocol-violation for exactly this. After a
successful decode-fail release the client now best-effort writes a
synthetic ServiceJobResult (error state, probe_key aggregate-key
fallback mirroring the worker's comm-error builder) carrying that kind;
a lost write falls back to the consumer's crash synthesis (logged,
swallowed). The result-write retry loop is extracted into a shared
writeResult helper reused by report(). Also covers the two untested
arms: a releaseJob THROW during decode-fail cleanup (warn + continue)
and all-candidates-decode-failing returning { claimed: false }.
sweep-lease-page-truncated fired on ANY exactly-full page — including 50
healthy in-flight jobs at steady state, a guaranteed false positive
every sweep. Under the ascending lease_expires_at sort, truncation only
hides expirable rows when the page TAIL is itself expired (everything
beyond has a later expiry), so the warn now requires a full page AND an
expired tail. Boundary-tested at exactly 50 all-live (no warn) and 50
all-expired (warn).
The stale-pending drain broke out of its page loop whenever a pass
expired nothing, even on a FULL page — but expiry is per-family while
the sort is absolute created, so younger expirable fast-family rows
sitting on later pages behind a full page of not-yet-expirable
slow-family rows were stranded for the whole sweep (the cross-family
occlusion class this drain exists to fix). Now: a non-full page is the
only early exit (the tail was seen); a full page with zero claim
attempts advances the page cursor (nothing left pending, pagination did
not shift); a page with attempts re-lists the same index (rows left
pending — deleted or claim-CAS'd either way — shifted pagination back).
Also corrects the wrong 're-listing returns the same rows forever'
justification: claim-CAS-lost rows LEAVE pending. The per-sweep page
cap still bounds the work.
sweepExpired (and fleet-health's reclaim) decide 'expired' from a listed
SNAPSHOT, then releaseJob(jobId, holder, 'pending') authorizes on
claimed_by alone — a worker renewing between the list and the release
still matches claimed_by, so a live just-renewed job was yanked back to
pending (duplicate execution + a false worker-reclaimed-pending comm
error). The /api/fleet/release hook now refuses a pending-target release
while the row's CURRENT lease is still live, re-checked inside the same
transaction with the leaseExpired helper that stays byte-equivalent to
the client's anchored parse. The client needs no change: released:false
already maps to the sweep's skip path. Pinned by a hook-source parity
test plus a renewed-after-list race test against a hook-faithful fake.
A thrown (not refused) releaseJob in the lease phase escaped the loop,
and the stale phase's pb.list/claimJob were uncaught — either throw
discarded the commErrors already synthesized for rows ALREADY released
to pending (the producer swallows sweepExpired throws), so their gray
're-queued' dashboard surfaces were never rendered and never
regenerated. Now: per-row try/catch around the lease-phase release
(queue-client.sweep-release-threw, continue), and the whole stale drain
is wrapped (queue-client.sweep-stale-phase-threw) with expiredPending
counted per row so a mid-pass throw still returns partial progress.
- Replace the shared ../logger.js module logger with a per-file silent
vi.fn-backed logger rebuilt in afterEach (the shared-logger spy-leak
class: a failed assertion skips mockRestore and poisons sibling files
under fork-reuse).
- makeFakePb now honors status AND probe_key clauses via a shared
LIKE-faithful matcher (incl. PocketBase's ESCAPE '\' semantics) and
THROWS on any clause it can't honor, so multi-family tests can never
pass vacuously; makePagingPb reuses the same matcher.
- Fix sampleResult fixture leaking the driver kind into the probe-key
position (e2e_d6:<slug> -> d6:<slug>; contracts.ts: no e2e_d6 rows).
- Replace the V8-leniency-dependent leading-space leaseExpired test with
STRING-level pins of the exported PB_DATE_SEP_RE anchor (engine
independent; residual V8/goja divergence documented), add a boundary
test at the exact expiry millisecond, and pin the hook source's
anchored regex + t <= Date.now() operator parity (3 handlers).
Composition fix across the cherry-picked sweep branches: the lease phase's
silent re-queue of a STALE_PENDING_SWEEPER row must join requeuedThisSweep,
so the same sweep's multi-page stale drain cannot claim-and-delete a row
the lease phase just re-queued — the retry contract is a LATER sweep. Adds
a composition test pinning that the grace set is honored on EVERY page of
the drain loop (a graced row surfacing only on pass 2 must survive).
- orchestrator.test.ts: file-level afterEach doUnmocks queue-client,
status-writer, and result-consumer (vi.doMock factories persist across
the file; resetModules clears the module cache, not the mock
registry) so a leaked stub can't poison later tests. Full file
re-run: 100/100 pass with the leaks closed — no test was depending on
a leaked factory.
- orchestrator.test.ts: the R5-G4 webhook-secret tests now save/restore
POCKETBASE_URL like the HF13-A2 pattern instead of unconditionally
deleting it in finally.
- job-producer.test.ts: the no-warm test stubbed a local fetch spy it
never wired in (vacuously zero calls); stub GLOBAL fetch via
vi.stubGlobal (+ vi.unstubAllGlobals in afterEach) and assert the
unconfigured producer never falls back to it.
- queue-client.test.ts: famOf re-implemented probeKeyFamily; import the
production helper from contracts so the tests can't drift from the
real family rule.
- control-plane.test.ts: the invalid-cron latch test now also retries
start() on the FAILED instance and asserts it throws again (a
stuck-true latch would make the retry a silent no-op).
stalePendingFilters silently falls back to the 1h default period for
any family missing from FLEET_FAMILY_PERIODS_MS, so a typo'd key (e.g.
"d5" vs "d5-single-pill-e2e") would never throw — it would just quietly
mis-size that family's stale-pending drain window. Lock the map's keys
to the probe-key families derived by RUNNING the four real enumerator
factories against a fake discovery source, so either side drifting
breaks the test. Also document the known d6 FLEET_PRODUCER_CRON
override drift limitation on the map (an env override changes d6's real
cadence without updating the nominal period).
The field was dead on the only consumer path: queue-client's enqueue()
destructures only `payload` and never reads leaseSeconds (the claim
lease comes from the WORKER side — claimNext(workerId, leaseSeconds)
with the worker-loop's DEFAULT_LEASE_SECONDS). No production call site
ever set it, so wiring it up would add a knob nothing needs; delete is
chosen over wire-it.
Call-site enumeration (all removed):
- contracts.ts EnqueueJobInput.leaseSeconds (declaration; never read by
queue-client.ts enqueue, the sole FleetQueueClient.enqueue impl)
- job-producer.ts ServiceJobSpec.leaseSeconds (only producer thereof)
- job-producer.ts toEnqueueInput() spec.leaseSeconds -> input.leaseSeconds
threading (only writer of the field)
- job-producer.test.ts spec fixture leaseSeconds: 600 + the
`expect(input.leaseSeconds).toBe(600)` assertion that legitimized the
dead plumbing
Worker-side lease plumbing (claimNext/renewLease/worker-loop
leaseSeconds) is unrelated and untouched.
maybeSweep's catch arm returned the same shape as a clean zero-reclaim
sweep (sweptExpired: true, reclaimed: 0), so a thrown sweepExpired call
was indistinguishable from success in the TickResult and the
tick-complete log. Add sweepFailed to the sweep outcome, TickResult,
and the tick-complete log; the cadence latch is unchanged (a failed
sweep still consumes its window so a persistently-failing sweep cannot
fire on every tick).
The sweep no longer synthesizes worker-crashed-mid-job (it cannot tell a
crash from a platform teardown); it re-queues the job and emits the neutral
worker-reclaimed-pending kind. Update the queue-client module header, the
contracts kind/heartbeat/SweepResult/sweepExpired docs, the job-producer
sink/tick docs, the sweep test fixtures and titles, and the dashboard's
mirrored kind description. worker-crashed-mid-job is now documented as the
worker self-observed in-driver crash only. No runtime behavior changes.
Only the d6 producer was built with onSweepCommErrors, but all four family
producers run the same GLOBAL queue.sweepExpired on their own crons — and the
sweep's S0 CAS means whichever producer ticks first wins each expired job's
reclaim, along with its synthesized comm error. With smoke/demos/deep sweeping
far more often than d6's hourly :40, the worker-reclaimed-pending dashboard
overlay (and stale-pending telemetry) was dropped ~11 of 12 sweeps, since
job-producer's maybeSweep forwards comm errors only when the sink is wired.
Share the ONE control-plane sink (surfaceSweepCommErrors -> aggregator) across
all four producers and correct the now-false "preserves the current behavior"
comment. Sweeps remain CAS-safe across producers: the S0 CAS guarantees exactly
one producer reclaims (and forwards) each expired job, and the surfacing leg is
best-effort per error, so the shared sink introduces no double-write.
Red-green: new runControlPlane REQ-B test drives the SMOKE producer's tick and
asserts its swept overlay reaches the status row (RED against d6-only wiring,
GREEN after). The test doUnmocks/re-mocks everything it touches so it passes in
isolation despite the file's leaked doMock factories.
filterBackloggedFamilies ran BEFORE maybeSweep, so the tick whose own
stale-pending drain cleared a family's backlog still counted the
about-to-be-expired rows and skipped that family — production resumed a
full cron period late. Reordered tick() to sweep first; the cadence gate
and fail-open semantics (maybeSweep swallows sweep failures) are
unchanged, only the order moved.
A single 50-row page per sweep was far slower than the incident the
stale-pending drain exists for: against the 3,734-row staging backlog at
~10 sweeps/hour that is ~7.5 hours of drain. The sweep now loops candidate
pages (re-listing page 1 — deletes shift pagination) up to a cap of 10
pages / 500 rows per sweep, draining the same backlog in well under an
hour while bounding a single sweep's PB load. CAS-claim-then-delete per
row is unchanged; a pass that expires nothing terminates the loop.
The sweepExpired lease phase listed claimed/running rows with perPage 50
and NO sort: with >50 such rows (mass worker crash), PB's unspecified
default order could return the same 50 live-lease rows every sweep,
leaving expired leases beyond the page permanently orphaned with zero
signal. Sort by lease_expires_at ascending (indexed) so the most-expired
rows always head the page, and WARN when the page is full so truncation
is observable. Single page per sweep is kept deliberately — the sort
guarantees progressive forward drain.
When the stale-pending sweep CAS-claims a row under stale-pending-sweeper
and the delete fails, the next lease sweep treated the expired sweeper
lease like a crashed worker's: it re-queued the row AND synthesized a
worker-reclaimed-pending comm error — a gray "re-queued / back in flight"
dashboard overlay for stale garbage mid-deletion, attributed to a
non-existent worker. The lease sweep now special-cases rows held by the
stale-pending sweeper: still re-queued (the self-healing delete-retry
contract is unchanged) but silently — no comm error, no reclaimed count,
just a stale-sweeper-retry-requeue debug line. The lease holder is
snapshotted before the release CAS so attribution reflects who held the
expired lease, not the post-release row.
sweepExpired's lease phase re-queues an expired-lease row to pending and
emits worker-reclaimed-pending ("back in flight"), but the stale-pending
phase of the SAME call lists pending fresh and ages rows off PB's system
`created` (the ORIGINAL enqueue time) — so a long-claimed job was
re-queued then immediately claimed-and-deleted, falsifying the comm
error and nulling downstream aggregate-key resolution on the deleted
row. Track the ids re-queued in this sweep and exclude them from this
call's stale phase; a truly stale job ages out on the next sweep. The
`created` anchor is kept (re-anchoring needs a column; out of scope).
sweepExpired only reclaimed claimed/running leases — a pending row had
no terminal path, so an accumulated backlog (staging: 3,734 pending,
oldest 22h) could only drain through 2 serial workers and effectively
never did.
The sweep now also expires pending jobs older than expiryPeriods x
their family's production period (default 3 periods; the family has
enqueued fresher batches since, so the stale job's result would be
ancient data). Each stale row is first CLAIMED via the S0 CAS under a
synthetic stale-pending-sweeper id — so the delete can never race a
worker (exactly-one-winner) — then deleted; a failed delete self-heals
via normal lease expiry + re-queue. Unparseable created timestamps are
conservatively skipped (delete is destructive). Policy is configurable
via FleetQueueClientConfig.stalePending; the control-plane wires the
real per-family cadences (FLEET_FAMILY_PERIODS_MS: d4/d5 15min, d6/
e2e-demos hourly) so 15min families expire on a 45min window. No comm
error is synthesized for expired-pending rows (they never ran); the
count surfaces as SweepResult.expiredPending and in the producer's
sweep log.
Producers enqueued a fresh batch every scheduled tick regardless of
whether the family's previous batch had even been claimed — against 2
serial browser workers the queue compounded without bound (staging:
3,734 pending), feeding the claim-page starvation.
A scheduled tick now skips its family's batch when that family already
has pending (unclaimed) jobs, bounding the per-family backlog to one
batch, with a structured fleet.producer.skipped-for-backlog log
(family, pendingCount, skippedJobs) and a skippedForBacklog count on
TickResult. The check is per family, fails OPEN on a count blip (a PB
read failure must never stop production), and is BYPASSED by
operator-triggered ticks (explicit intent wins; the trigger CLI treats
0 enqueued as failure). Backed by a new
FleetQueueClient.countPendingForFamily (server-side totalItems count of
a family's pending rows).
claimNext listed ONE global oldest-50 pending page; with the d4+d5
producers ticking every 15min against 2 serial browser workers, a
persistent backlog permanently saturated that page and e2e-demos jobs
never entered the candidate set (prod: all 18 e2e-demos jobs pending
forever; staging: 3,734 pending, oldest 22h).
claimNext now discovers the distinct families present in pending
(oldest first, one perPage=1 query per family) and tries them in
round-robin rotation across calls, listing a per-family candidate page
for each. Every discovered family is attempted before giving up, so no
family starves while any of its jobs are claimable. The S0 CAS
exactly-one-winner semantics and the per-page anti-herd shuffle are
unchanged; only candidate SELECTION changed.
## Summary
End-to-end hardening of the showcase shell's URL plane:
- Carries `backendHostPattern` + `docsHost` in the shell runtime config
(no longer baked from `registry.json` at Docker build time). Derives
demo backend URLs at runtime from the pattern, so a new pattern via env
reconfigures every integration on the next deploy without a registry
rebuild.
- Issues docs-host 301/308 redirects from middleware with a runtime
`DOCS_HOST`; misconfigured values no longer 500 every docs route — they
fall back to a sentinel that disables the docs-redirect step.
- Hardens the redirect table builder + matcher: first-match-wins for
duplicate exact sources, deduped wildcard prefixes with warn, malformed
entries rejected at lookup-build time, case-insensitive matching parity,
trailing-slash normalization, structural `/integrations` namespace guard
above the docs-host redirect (closes R15/R17 hijack class).
- Hardens runtime-config + backend-url env readers:
scheme/whitespace/control-char normalization, query/fragment/userinfo
rejection on `DOCS_HOST` / `POSTHOG_HOST` / backend-host pattern /
local-override URLs, present-but-empty `posthogKey` rejection, prod
loopback `BASE_URL`/`DOCS_HOST` rejection (no silent `http://` prepend),
once-guarded FATAL logging (no per-request spam).
- Brings the build-time twin in `showcase/scripts/generate-registry.ts`
to parity with the runtime normalizer (scheme/trailing-slash strip,
degenerate fallback, `NEXT_PUBLIC` fallback, slug validation).
- Surfaces PostHog capture failures once per failure class; keeps
capture alive across the redirect via `event.waitUntil`; missing
`POSTHOG_KEY` is `console.error` in production and surfaces at
config-resolution time (not first redirect).
- Open-redirect hardening on `/shared//evil.com` (SU-18); `//` rejected
at both source and destination; root `/` exact-source and `/:path*`
wildcard sources rejected; non-printable-ASCII source/destination
rejected.
Stream: SU (shell-runtime-urls). Subject groups
SU-2/8/11/13/14/15/16/17/18/19/20 (initial), SU2-A/B (env-hardening),
CR2-C (test infra), SU5-A1..A7 (registry/builder/lint/matcher boundary),
SU6-A1..A6 (request-time normalization parity), SU6-B1..B7
(parsed-normalized return forms, SHOWCASE_LOCAL states, generator
`{slug}` validation), SU7-F1..F3 (final-round backend-pattern + POSTHOG
+ table + script-side parity).
## Test plan
- [x] `pnpm test` in `showcase/shell` — 7 files, 235/235 pass
- [x] `pnpm test` in `showcase/scripts` — 51 files, 1857/1857 pass (one
pre-existing flake in `generate-registry.test.ts` unrelated to this
diff; passes on subsequent runs, classic vitest fork-reuse cross-file
pollution)
- [x] `pnpm exec tsc --noEmit` in `showcase/shell` — clean
- [x] `oxlint` on every changed file — 0 warnings, 0 errors
- [x] `oxfmt --check` on every changed file — clean
- [x] `pnpm build` in `showcase/shell` (full Next 15.x build with
registry+demo-content+starter-content+search-index generation) — clean
- [ ] CI green on the PR — confirm via `gh pr checks` after push
## Summary
Anti-dual-writer defense for the showcase `status` collection — the
flap-comb incident class where the legacy monolith scheduler and the
fleet writer fight over the same rows.
- **Writer identity columns**: new `written_by` / `state_written_at`
columns on `status` (two PB migrations, idempotency-symmetric up/down
paths). Every write stamps a stable host-derived writer identity; the
legacy orchestrator wires `writtenBy: "legacy"` explicitly.
- **Flip / foreign-write detection**: the status writer detects
cross-writer state flips inside a validated window and warns
(TTL-deduped re-warns so sustained fighting stays visible; one-time warn
when the self-write memory cap starts evicting; same-identity replica
hint).
- **PB-safe date normalization**: `observedAt` is normalized to RFC-3339
PB-safe shapes before any date-field write (zone-less shapes, colonless
offsets rejected) — no PB 400 unlatch-retry hot loops.
- **Aggregator honesty**: padded/blank projected keys normalized or
skipped loudly; explicit outcome discriminators replace ambiguous
asserts (no consumer hot-loop); `droppedCommError` surfaced whenever a
comm error misses the aggregate row; trusted-negative duplicate handling
scoped to cell-vs-cell so no cell can impersonate an aggregate row.
- **CLI persistence honesty**: thrown driver errors are persisted to PB
in the `runDriverInputs` catch; the summary carries `WriteOutcome`
discriminators (dropped-write counts, correct pluralization); writer
error classification documented honestly (401 vs 403 split, new
`pb_not_found` reason).
- **Alert/fixture truthfulness**: the alert engine's synthesized cron
outcome is stamped `persisted: false`; alert/orchestrator/probe test
fixtures align with the real `StatusWriter` / `OverlayWriteOutcome`
contracts, with fail-loud fake-PB fixtures hardened against silent
divergence from real PocketBase.
Hardened over 10 CR rounds (~60 reviewer reports), converged to zero
findings.
## Test plan
- [x] Full harness vitest suite: 124 test files / 2378 tests passing
- [x] `tsc --noEmit` and build typecheck (`tsconfig.build.json`) clean
- [x] oxfmt clean on all touched files; oxlint 0 errors
- [x] Known flake note: `probe-invoker.test.ts` wall-clock assertion
(elapsed 101ms vs <100ms bound) fired once under full-suite load and
passes in isolation (67/67) — pre-existing timing sensitivity, unrelated
to this diff
Rewire the before-first-message suggestion pills to drive the FOR-137
self-learning story: (1) the teachable over-limit ask, (2) surface the pending
charges so the officer can demonstrate the unlock, (3) recall on a different
over-limit charge on a fresh thread. Titles stay symptom-only so they do not
hint at the exception path the agent is meant to learn on its own.
The banking demo's human-in-the-loop tools registered their render in a
mount-keyed effect (useFrontendTool), so without a deps array the render
closure froze on the EMPTY initial cards/policies/transactions. Those arrays
load async after mount, so the registered render kept filtering empty data:
showAndApproveTransactions painted a card with no rows (the agent-driven
approve flow appeared to do nothing), and assignPolicyToCard / setCardPin /
addNoteToTransaction showed raw ids instead of the resolved card/transaction.
Pass the data each render reads as the useHumanInTheLoop deps so it
re-registers when that data loads, mirroring the existing selectCard and
showTransactions (useComponent) deps. addNewCard / openPolicyException /
finalizePolicyException render their args only, so they are left as-is.
Verified against unmodified workspace react-core: the agent-driven approval
card now renders the Google Ads charge with its over-limit badge and the
file-exception action instead of a blank card.
## What
Completes the v2 API migration across the CopilotKit agent skills,
extending #5345 (which started it for `copilotkit-setup`). Docs only, no
runtime code changed.
## Why
The journey skills a new user reaches for were still teaching pre-v2
APIs, so copy-pasting their examples failed before anything ran:
- `copilotkit-develop` and `copilotkit-integrations` imported from
`@copilotkit/react`, which does not exist (the package is
`@copilotkit/react-core`, v2 surface at `/v2`).
- `copilotkit-setup` assets imported the phantom `@copilotkit/agent`
package and a non-existent `@copilotkit/runtime/express` subpath.
- `copilotkit-integrations` used the v1 route and an uncompilable
`useAgent` shape.
- The recommended multi-route provider snippet omitted
`useSingleEndpoint={false}`, so the chat rendered but never connected
(404).
## Changes
| Skill | Migration |
|-------|-----------|
| `copilotkit-setup` | assets + refs to `@copilotkit/runtime/v2` (and
`/v2/express`), `createCopilotHonoHandler` /
`createCopilotExpressHandler`, `@copilotkit/react-core/v2`; transport
fix (`useSingleEndpoint={false}`); PATCH/DELETE route exports |
| `copilotkit-develop` | hooks and components to `react-core/v2`,
runtime to `/v2`, `useThreads` signature, complete route example |
| `copilotkit-integrations` | `CopilotKit` provider, v2 catch-all Hono
route, per-framework agent classes matched to the shipped
`examples/integrations/*` |
| `runtime` skill | LangGraph to `@copilotkit/runtime/langgraph`, Agno
to `HttpAgent` from `@ag-ui/client` |
| `react-core` skill | `defaultOpen`, `onSubmitMessage`, headless
`CopilotChatView` render prop |
Standardized on the non-deprecated `createCopilotHonoHandler` /
`createCopilotExpressHandler` factories (matching #5345).
## Validation
- Reviewed across 5 unbiased review rounds (7 agents per round); the
final round found zero actionable findings in the changed code.
- Functional build test against the published `@copilotkit/*@1.60.0`
packages: `npm install` plus `tsc --noEmit` pass with zero errors, and
every `/v2` subpath and imported symbol resolves.
## Out of scope (tracked separately)
Pre-existing issues in files this PR did not change: integration example
Python bodies, `runtime`-skill citation paths into the retired `docs/`
tree, reference completeness gaps, `showDevConsole` doc accuracy, the
`copilotkit-upgrade` skill, the `publicApiKey` vs `publicLicenseKey`
source-of-truth decision, and skill eval modernization.
Bring the agent skills in line with the shipped v2 API so their examples
install, compile, and connect. Extends #5345 (which migrated the
copilotkit-setup SKILL.md body) to the rest of the skills.
- Imports: drop the nonexistent @copilotkit/react and @copilotkit/agent
packages and the bare @copilotkit/runtime/express subpath; use
@copilotkit/react-core/v2 and @copilotkit/runtime/v2 (+ /v2/express),
and createCopilotHonoHandler / createCopilotExpressHandler rather than
the deprecated createCopilotEndpoint aliases.
- Provider: CopilotKit from @copilotkit/react-core/v2 with
useSingleEndpoint={false} on multi-route setups (the v1-compat bridge
defaults to single transport and would 404 a multi-route backend).
- Routes: v2 catch-all Hono handler exporting GET/POST/PATCH/DELETE via
handle() from hono/vercel, replacing the v1
copilotRuntimeNextJSAppRouterEndpoint + ExperimentalEmptyAdapter.
- Integrations: per-framework agent classes matched to the shipped
examples (LangGraphAgent/LangGraphHttpAgent from @copilotkit/runtime/
langgraph, CrewAIAgent, MastraAgent, LlamaIndexAgent, HttpAgent from
@ag-ui/client; Agno via HttpAgent, not @ag-ui/agno).
- Hooks/props: correct useAgent, useThreads, useRenderTool, identifyUser,
and the chat-component props (defaultOpen, onSubmitMessage, the headless
CopilotChatView render prop).
Validated across review rounds and a build test against the published
@copilotkit/*@1.60.0 packages (tsc passes, every /v2 subpath resolves).
Regenerated the skills/runtime and skills/react-core mirrors.
SU7-F1 — backend host pattern hardening:
- F1.1 Reject bare trailing ?/# in the backend host pattern
- F1.2 Strip internal tab/CR/LF from the backend host pattern
- F1.3 Warn when ignoring an empty-string local backend override
- F1.4 Reject empty-userinfo @ in the backend host pattern authority
- F1.5 Keep __proto__ keys as data in local-backend maps
- F1.6 Commit the local-backends memo key only after the value computes
- F1.7 Trim local backend overrides before validation and name the real
rejection
- F1.8 Honest FATAL when the pattern host is a stray scheme fragment
- F1.9 Canonicalize the pattern authority for parity with the override
path
- F1.10 Acknowledge the staging-to-prod fail-open in the pattern fallback
- F1.11 Harden backend-url/local-backends-env test hygiene
SU7-F2 — runtime-config & client-config edge cases:
- F2.1 Branch POSTHOG_HOST rejection reasons (scheme/degenerate/parse-
failure) instead of the catch-all mislabel
- F2.2 Reject loopback BASE_URL/DOCS_HOST in production instead of the
silent http:// prepend
- F2.3 Key the DOCS_HOST fallback once-guard on (mode, shellHost, value)
and mode-prefix all value-only guard keys
- F2.4 Reject a present-but-empty posthogKey in the client config reader
- F2.5 Drop the trailing slash from SSR_PLACEHOLDER_URL for structural
parity with server values
- F2.6 Attribute the DOCS_HOST slash-strip to readDocsHost itself
- F2.7 Normalize trailing-dot FQDN spellings in the docs self-host loop
guard (both compare sides)
- F2.8 Harden console spies to capture all log args; pin the full all-env
config shape; converge SSR simulation on vi.stubGlobal
SU7-F3 — script-side parity, table classification & test isolation:
- F3 #1 Handle a missing reference integration per the error contract
- F3 #2 Port the runtime backend-host-pattern normalization into the
generator — scheme/trailing-slash strip, degenerate fallback,
NEXT_PUBLIC fallback
- F3 #3 Treat non-mapping manifest parses (empty/null/scalar/array YAML)
as validation errors, not TypeErrors
- F3 #4 Label a missing/unreadable constraints.yaml per the stderr+exit(1)
error contract
- F3 #5 Align atomic-write tmp naming with the test harness straggler-
sweep convention; guard main() on direct invocation
- F3 #6 Correct the determineCellStatus unshipped docstring; replace
stale hardcoded cell counts with formulas
- F3 #7 Isolate the pattern suite on a per-suite tmpdir harness; snapshot
the generator's full write set
- F3 #8 Classify discarded duplicate wildcards as duplicates — hoist the
owner check above the destination warns
- F3 #9 Reject a root ("/") EXACT seo-redirect source — homepage-hijack
twin of the root-wildcard guard
- F3 #10 Reject seo-redirect entries with non-printable-ASCII source/
destination — close the silent-dead-entry class
- F3 #11 Strip trailing slashes in normalizePosthogHost before the scheme
test
- F3 #12 Message-filter the empty-slug-set error count; pin the single
matcher entry
Round-by-round CR convergence covering the redirect builder, the middleware
matcher, the docs-host self-loop guard, and the runtime-config env readers.
Highlights:
- Clear module-load warns after fresh middleware import
- Validate SET BASE_URL values (scheme-less/degenerate/garbage) with
sentinel fallback + once-guarded FATAL log
- Normalize path/query/fragment-bearing DOCS_HOST to origin; reject
non-http(s) schemes; branch rejection reasons
- Harden POSTHOG_HOST (degenerate-host/scheme rejection); expose
posthogKey via readEnvPair semantics
- Reject a DOCS_HOST equal to the shell's own host (redirect-loop guard,
authority compare)
- Warn on missing local-ports.json under SHOWCASE_LOCAL=1 and validate
TCP port range; extract helper for tests
- backend-url hardening — slug charset guard, frozen local-backends memo,
pattern path-segment warn, local-override URL validation
- Client config fail-loud covers all four URL fields with type checks
- Make RuntimeConfig.posthogKey optional — absence is a valid state, not
a wiring bug
- Drop R15/R17 and guard /integrations from SEO redirects
- Dedup duplicate wildcard prefixes with first-match-wins warn
- Validate malformed SEO entries at lookup-build time
- Restore case-insensitive redirect matching parity
- Normalize trailing slashes before redirect matching
- Keep the framework segment on F13, pin MG3 case fix
- Read posthogKey from runtime config in middleware, not raw process.env
- Fall back to the default backend host pattern for degenerate values
- Disable docs redirects when the default fallback collides with the
shell host
- Bring validateBaseUrl to parity with its sibling readers
- Strip query/fragment from POSTHOG_HOST while keeping reverse-proxy paths
- Restrict local backend overrides to http(s) URLs
- Add server-only guard to runtime-config
- Harden localBackendsEnv failure posture
- Hoist /integrations namespace guard above the docs-host redirect
- Validate seo-redirect sources and cross-kind shadowing in
buildRedirectLookup
- Unify slash normalization for middleware matching
- Lowercase-normalize REGISTRY_FRAMEWORK_SLUGS at construction
- Escalate missing POSTHOG_KEY to console.error in production
- Skip all redirect steps when docs redirects are disabled (sentinel
consumer)
- Reject userinfo credentials in DOCS_HOST, POSTHOG_HOST, and the backend
host pattern
- Branch dev-vs-prod logging in readDocsHost and fatalPatternOnce
- Prepend http:// (not https://) to scheme-less loopback hosts
- Round-5 micro-finding batch across the URL config libs
SU5-A1..A7 — registry safety, // reject, builder lint batch (case-
insensitive :path*, same-destination twin allowlist, original-case
divergence remainder), matcher api boundary, generator+vitest infra, test
hygiene + empty docs-host guard, comment batch.
SU6-A1..A6 — reject miscased :path* tokens, warn on tokenless wildcards,
normalize redirect-destination comparisons like request time, reject
destinations containing "//", surface missing POSTHOG_KEY at config-
resolution time, compile matcher harness like Next's runtime, type
parse/tokensToRegexp in the path-to-regexp shim, keep buildRedirectLookup
JSDoc attached.
SU6-B1..B7 — reject query/fragment/userinfo in pattern and local-override
URL gates, return parsed-normalized URL form from validation success
paths, distinguish unset/blank/padded SHOWCASE_LOCAL states, warn when
SHOWCASE_LOCAL is set to a value other than 1, validate {slug} placeholder
in generate-registry, mirror middleware drop semantics in the wiring
test's registry re-derivation, pin the noStore spy and calls to one fresh
module instance in the Edge-path test.
SU2-B series — runtime-config / backend-url env robustness:
- Correct the Edge-safety story in runtime-config (SU2-B1)
- Stop per-request FATAL-CONFIG spam for unset BASE_URL (SU2-B2)
- Prepend https:// to a scheme-less POSTHOG_HOST (SU2-B3)
- Trim whitespace paste artifacts in env values and host patterns (SU2-B4)
- Memoize parseLocalBackends and warn once per value (SU2-B5)
- Make {slug} substitution immune to $-patterns (SU2-B6)
- Harden the client runtime-config reader (SU2-B7)
- runtime-config hardening batch (SU2-B8)
- Validate local-ports.json before baking NEXT_PUBLIC_LOCAL_BACKENDS (SU2-B9)
- test: warn-once assertions retry-safe; stop console leaks (SU2-B10)
CR2-C series — test infrastructure:
- Generate registry.json in a vitest globalSetup (CR2-C1)
- Stop ambient POSTHOG_KEY firing real fetches in middleware tests (CR2-C2)
- Assert the production slug set, not a re-derivation (CR2-C3)
- Make the registry generator subprocess robust (CR2-C4)
- Middleware/wiring test hygiene batch (CR2-C5)
SU2-A series — redirect-layer & PostHog capture:
- Stop $-pattern expansion in wildcard redirect substitution (SU2-A1)
- Surface PostHog capture failures once per failure class (SU2-A2)
- Duplicate exact redirect sources are first-match-wins (SU2-A3)
- Resolve runtime config once per redirected request (SU2-A4)
- Include destination host in seo_redirect capture (SU2-A5)
- Normalize scheme-less POSTHOG_HOST at the capture use site (SU2-A6)
- Correct redirect-layer comments and guard wildcard prefix boundary (SU2-A7)
- Cover docs-host hardening branches, compile matcher via path-to-regexp (SU2-A8)
Resolve SEO redirect destinations against the docs host (SU-17); forward
the query string on SEO redirects (SU-16); match bare paths on wildcard
SEO sources (SU-19). Collapse duplicate slashes in docs-host redirect
destinations (SU-13). Regression test for /shared//evil.com open redirect
(SU-18). Emit 308 for docs-host redirects to match next.config parity
(SU-2). Add a path boundary to the api matcher exclusion (SU-15). Loud
guard when registry yields zero framework slugs (SU-20). Keep PostHog
capture alive via event.waitUntil (SU-14). Note docs-host redirects are
untracked by design (SU-8). Cover docs-host redirects at the middleware
level (SU-11).
Squash of the SEO-table + matcher-hardening cluster.
Carry backendHostPattern + docsHost in the shell runtime config (no longer
baked from registry.json at Docker build time). Derive demo backend URLs at
runtime from the pattern; issue docs-host 301s from middleware with a runtime
DOCS_HOST so a misconfigured value can no longer 500 every docs route.
Validate NEXT_PUBLIC_LOCAL_BACKENDS and empty overrides; guard the backend
host pattern against silent env misconfigs. Reword the stale demo-page
comment about backend URL derivation. Pin the registry slug set and SSR
placeholder URL composition; fix env/spy/global leaks in runtime-config
test cleanup.
Squash of the initial runtime-URL refactor cluster:
- feat(showcase): carry backendHostPattern + docsHost in shell runtime config
- fix(showcase): derive demo backend URLs at runtime instead of baked registry values
- fix(showcase): issue docs-host 301s from middleware with runtime DOCS_HOST
- fix(showcase): never let a misconfigured DOCS_HOST 500 every docs route
- fix(showcase): guard backend host pattern against silent env misconfigs
- fix(showcase): validate NEXT_PUBLIC_LOCAL_BACKENDS values and empty overrides
- docs(showcase): reword stale demo-page comment about backend URL derivation
- test(showcase): fix env/spy/global leaks in runtime-config test cleanup
- test(showcase): pin registry slug set and SSR placeholder URL composition
## What
Across the **threads-enabled integration examples**
(`examples/integrations/*`, excluding the vestigial
`langgraph-python-threads`):
1. **Drop demo-user provisioning** — removes the `provision-user`
one-shot service from the shared `_intelligence/docker-compose.yml`.
2. **Bump Intelligence → 0.5.0** —
`ghcr.io/copilotkit/intelligence/composite` `0.2.0` → `0.5.0`.
3. **Bump CopilotKit → 1.60.0** — every `@copilotkit/*` pin
(`react-core`, `react-ui`, `runtime`, `a2ui-renderer`, `sdk-js`)
`1.59.5` → `1.60.0`, with regenerated `package-lock.json` files.
## Why
The `provision-user` service existed to seed `cpki.users` (a bare id
plus a per-project `<projectId>_<userId>` alias) so the runtime's
`demo-user` identity satisfied the `threads_user_id_fkey` constraint.
The **0.5.0 composite provisions thread users on demand**, so the manual
seed is obsolete and can be dropped from the local-dev overlay.
## Scope
18 threads-enabled examples (each renders the `ThreadsDrawer` panel):
a2a-a2ui, a2a-middleware, adk, agent-spec, agentcore, agno,
crewai-crews, crewai-flows, langgraph-fastapi, langgraph-js,
langgraph-python, llamaindex, mastra, mcp-apps,
ms-agent-framework-dotnet, ms-agent-framework-python, pydantic-ai,
strands-python.
`langgraph-python-threads` is intentionally untouched (vestigial, slated
for removal).
## Notes
- The `identifyUser` demo stub in each route handler is **kept** — it
selects the demo identity; provisioning that identity is now the
composite's job.
- `_intelligence/docker-compose.yml` is not parity-tracked;
`@copilotkit` pins are — `pnpm parity:check` passes (0 errors).
- The Intelligence `0.5.0` image + on-demand provisioning is validated
at runtime (needs the image + license token + Docker), not in this PR's
static checks.