- decode-failure REFUSED release now warns with the hook's refusal reason
- probeKeySlug empty-slug ('d6:') falls back to the whole probe_key; a
fully-empty key falls back to unknown-<jobId> (never the forbidden
empty serviceSlug sentinel)
- countPendingForFamily throws on non-numeric/negative totalItems instead
of failing the producer backlog gate open with -1
- stale-pending drain contains a thrown claim CAS per row (a 4xx no
longer aborts the whole drain) and documents the 5xx->won:false
bounded page-stall on passAttempted
- writeResult paces retries (RESULT_WRITE_RETRY_DELAY_MS, injectable
sleep) instead of burning attempts back-to-back
- protocol-violation warns for won&&!job / renewed&&!job before falling
through
- ensureAuth memoizes the in-flight auth promise (no concurrent stampede)
and wraps the auth-body JSON.parse with status context
- releaseJob interface doc corrected: the hook admits claimed AND running
rows (the decode-failure cleanup depends on claimed->failed); the doc's
'running state only' was the wrong side
- documented the LIKE-vs-= ASCII case-sensitivity divergence for
mixed-case families (unreachable: keys are lowercase slugs)
The conservative thrown-release path treated EVERY throw as
may-have-committed, so a wedge row (e.g. empty claimed_by -> hook 400 on
the missing workerId) produced a permanent per-sweep false
worker-reclaimed-pending overlay. job-claim now threads the HTTP status
onto thrown endpoint errors (JobClaimEndpointError); the sweep maps a
4xx (hook rejected, nothing committed) to an error-level log with no
grace, no comm error, no reclaimed++, while 5xx/transport/2xx-unreadable
keep the conservative at-least-once handling.
On the refused_terminal_same_holder retry path, writeResult's
result_processed:false seed could overwrite a row the consumer had
already aggregated (and latched), un-latching it for a second
aggregation — a double-count. The retry now reads the row first: a
present+processed result is skipped (info 'result already aggregated'),
a present+unprocessed result is skipped too (idempotent), only a
missing result falls through to writeResult. A failed pre-write read
throws instead of blind-writing.
With job-claim now throwing on unreadable-2xx (and renew already throwing
on 5xx), a thrown renew escaped queue-client.renewLease into the worker
heartbeat, which treats any renewLease throw as fatal — the heartbeat
died and the sweeper reclaimed a live job (false worker-crashed-mid-job).
renewLease now catches throws from claim.renewLease, warns, and returns
the last-known lease unchanged (no payloadCache eviction, never null) so
the heartbeat retries next beat; only a definitive renewed:false stops
it. Documented at-most-one-lease-duration duplicate-execution risk; the
release CAS still arbitrates terminal writes.
A 2xx response whose body fails to read, is empty, or fails to parse was
mapped to {} — i.e. claimed/renewed/released: false — fabricating a CAS
LOSS for a transition the server had already COMMITTED (the renew flavor
recreated the false worker-crashed-mid-job class; the release flavor made
report() falsely declare the result discarded). Now any unreadable 2xx
body THROWS with path context; callers contain the throw.
probeKeyFamily treats a leading-colon probe_key as its own whole-key
family, so the family value itself contains colons — expanding it with
the <family>:% LIKE leg folded the unrelated family ":foo:bar" under
":foo" in the inclusion/count legs and hid it from discovery via the
exclusion leg. Special-case colon-bearing families to the equality leg
only, implementing the invariant pinned on probeKeyFamily (G3e).
countPendingForFamily filtered on status = "pending" only, so a family
whose whole batch was claimed/running stopped gating and a scheduled
tick could enqueue a fresh batch on top of the in-flight one, doubling
the family's concurrency. Broaden the filter to non-terminal
(pending || claimed || running) and align the gate docs in the contract
and job-producer.
probeKeyFamily(':foo:bar') returns the WHOLE key as its own family (the
round-2 leading-colon fix), but the queue-client clause builders
(familyInclusionClause/familyExclusionClause) expand any family with a
prefix-LIKE <family>:% leg, which would include/exclude ':foo:bar' under
family ':foo' — disagreeing with the fairness partition. The clause
builders live in queue-client.ts (out of scope here), so this documents the
invariant on the contracts side — whole-key (colon-bearing) families must
match by equality only — and pins the partition behavior in the test suite.
The queue-client special-case is reported to the integrator separately.
The taxonomy doc defined the kind as exclusively worker-self-monitor-
reported, but the control-plane result consumer ALSO synthesizes it for
terminal-but-resultless rows past the grace window (the queue-client's
bounded result-write retry exhausted — see RESULT_WRITE_MAX_ATTEMPTS).
Both the contracts.ts kind doc and the dashboard mirror now enumerate
the two sources.
typeof raw !== "object" admits arrays: an array carrying comm-error fields
as expando properties decoded as a well-formed PoolCommError. Both copies
(harness contracts.ts and the dashboard live-status.ts mirror) now reject
arrays explicitly and stay byte-identical — the drift suite gains a
byte-identity pin on the full function source so a one-sided fix can never
silently re-open the gap.
The rollup fold used literal includes("red")/includes("degraded") checks,
so an out-of-vocabulary contributor state (e.g. "error", which the harness
can persist at runtime) matched neither bucket and rolled the cell up GRAY
(benign no-data) while its own badge rendered the loud error tone —
violating the documented precedence red > degraded > green > error >
unknown and the A2 never-swallow rule. Contributor states now route through
worstStateRank, so an unknown state rolls up at least red-severity.
The pass-through gate in fleetSurfaceState was a literal === "red" check, so
a row whose last-known state was degraded/error/out-of-vocab got masked by
the neutral pending overlay — violating the function's own never-mask-a-
genuine-failure invariant (A2: error ranks ABOVE red). Only green now becomes
pending; the dashboard cell-model derivation mirrors it over its ChipColor
vocabulary (red AND amber pass through). Fixes the self-contradicting JSDoc
closing sentence, the stale live-status.ts mirror doc (union was missing
pending), and extends the drift test to pin the pass-through semantics on
both sides.
- count a warm GET as fired only after successful dispatch (sync-throwing
fetchImpl no longer inflates the warmed log)
- route an enumerator resolving to a non-array through the same
enumerate-failure handling as a throw
- mint the runId AFTER the running check (stopped ticks no longer burn
the factory counter or log phantom runIds; empty-runId sentinel)
- log tick-start BEFORE the enumerate await so a hung discovery still
leaves a tick trace (services count moved to tick-complete)
- warn once, with jobIds, when swept comm errors exist but no sink is
configured (previously dropped with only a count)
- terminal .catch on the fire-and-forget warm chain (a throwing logger
no longer raises an unhandled rejection)
- document the SweepCommErrorSink at-least-once contract (aggregator
must be idempotent per jobId+observedAt)
- tests: fix stale "Four producers" comment; replace the fragile
two-shot Math.random spy with a never-exhausting deterministic
sequence plus a call-count guard
A slow tick overlapping the next cron tick let both read the backlog
gate's pending count before either enqueued, so both produced a batch
for the same family — the motivating concurrent same-service-run
incident class. An overlapping tick is now skipped (warned) via the
tracked in-flight tick promise; it returns the empty-runId no-op result
without minting a run.
NOTE for the integrator: the gate also under-counts — queue-client's
countPendingForFamily filters status="pending" only, so a fully-claimed
but-still-running batch doesn't gate. Broadening the filter to
non-terminal rows (pending || claimed || running) is a one-line edit in
queue-client.ts (plus the contracts.ts doc), both sibling-owned —
deferred to the integration pass.
stop() used to flip flags and resolve while an in-flight tick kept
enumerating/sweeping/enqueueing, and buffered sweep comm errors were
silently dropped at shutdown. Now stop() awaits the tracked in-flight
tick promise, the enqueue loop re-checks `running` and truncates loudly
when stopped mid-batch, and stop() makes one final sink delivery of the
buffered comm errors — logging dropped count + jobIds at error level if
that last attempt fails too.
The undelivered comm-error buffer was drained only on maybeSweep's try
SUCCESS path, so a persistently-throwing sweepExpired — the exact
failure mode the buffer exists to ride out — never handed the buffered
batch to a now-healthy sink. Extract the delivery into
deliverSweepCommErrors and run it on every sweep attempt, including the
catch arm.
rowMatchesFilter now THROWS on operators it can't honor (<,>,<=,>=,?=,
?~) and on status clauses with !=/~/!~ instead of silently matching all
rows, with the quoted-literal false-positive limitation documented;
both fakes return totalItems:-1 (and -1 totalPages) under skipTotal —
faithful to PB, closing the fail-open count class beyond the single
countPendingForFamily pin; makePagingPb honors the '-' sort-direction
prefix and computes totalPages honestly; makeOrderedPb's
ignore-all-filters behavior is documented as deliberate coupling to the
duplicate-family defensive break; the result-write retry bound is
pinned with the exported RESULT_WRITE_MAX_ATTEMPTS constant.
(i) per-candidate try/catch in claimNext's CAS race — a thrown transport
claim no longer aborts the whole rotation (warn + next candidate);
(ii) discoverPendingFamilies' duplicate-family defensive break now warns
before breaking; (iii) renew re-read triage split — a decodePayload
failure logs as a protocol violation, not a read blip; (iv) the
decode-failure synthetic result uses an injected clock (new config.now)
instead of new Date(); (v) backslash charset guard for family clause
building (the equality/LIKE escape contracts contradict for backslash —
skip such families with a warn) and the over-claiming VERIFIED comment
scoped to %/_ only; (vi) single-sweeper assumption documented as
load-bearing at the grace-set declaration; (vii) hook comment misname
fixed (worker-crashed-mid-job -> worker-reclaimed-pending); (viii)
claimJob maps a 5xx claim response to a lost CAS (won:false, warn) —
a WAL serialization error escaping runInTransaction surfaces as 500 —
while 4xx and renew/release 5xx still throw loud.
After a release-CAS success + result-write exhaustion, report() throws;
a natural retry got REFUSED (row already terminal) and emitted an error
claiming the result is discarded and the job re-runs — both false (the
result is still writable by this holder; terminal rows never re-run).
The hook's release response now carries a refusal reason
(refused_terminal_same_holder / refused_not_holder /
refused_lease_live), threaded through job-claim's ReleaseResult.
report() treats refused-terminal-under-my-workerId as the second leg of
a timeout-after-commit retry and proceeds to writeResult; the
not-holder error is reworded to 're-runs only if reclaimed to pending'.
A reason-less refusal still fails closed.
Stale-pending age is anchored on PB created and never re-anchored on
re-queue: a job running longer than its family expiry window got
lease-reclaimed with a 'back in flight' comm error, then claim-deleted
by the NEXT sweep before any plausible re-run — the dashboard
permanently showed 're-queued' for silently-discarded work. Schema-free
fix: the release hook now RETAINS the expired lease_expires_at on a
pending re-queue (claim admits pending rows regardless of lease), and
the stale phase skips rows whose retained lease is recent (parseable
and within now - familyExpiryWindow) — recently in flight means the
created-based age is stale evidence. Comments cover the heuristic, the
requeued_at-column alternative, and the sweeper-garbage lingering
tradeoff; hook parity test pins the retention.
The decode-failure synthesis persisted serviceSlug: "" and runId: "" —
the exact empty sentinels this file's own emptyPayloadForLease warning
forbids feeding aggregation (an empty runId groups into nothing, an
empty serviceSlug corrupts the per-service rollup). Recover the slug
from the probe_key's slug segment (new probeKeySlug, the complement of
contracts.ts probeKeyFamily) and mint a non-colliding pviol_<jobId>
runId; reconcile the contradictory comments on both sides. Verified
downstream: result-consumer/aggregator key on aggregateKey + dedupe on
jobId, with serviceSlug/runId riding into logs and batch grouping.
A release that THROWS in the lease sweep may have COMMITTED server-side
(timeout-after-commit): the row is then pending but absent from the
grace set — the same call's stale phase could claim-and-delete it — and
its worker-reclaimed-pending comm error was never synthesized, losing
the gray 're-queued' surface forever (no later sweep re-emits pending
rows). Conservatively add the row to the grace set and synthesize the
comm error on a throw (at-least-once: a duplicate gray overlay is
harmless, a missing one is not); sweeper-held rows stay silent,
mirroring the committed path.
The three /api/fleet/* routerAdd handlers carried no auth middleware —
a middleware-less PB 0.22 routerAdd handler is PUBLIC, so any
unauthenticated caller could claim/renew/release arbitrary jobs despite
the header claiming superuser auth was required. Append
$apis.requireAdminAuth() (verified against PB 0.22.21 JSVM types) to
all three routes; the client already authenticates as superuser with a
401-reauth retry, so enforcement is compat-safe. Also clamp
leaseSeconds in claim+renew (numeric only, 3600s ceiling, 30s default
on garbage) and pin both contracts in the hook-source parity tests.
A probe key beginning with ":" yielded the empty-string family, which
then flowed into countPendingForFamily and the claimNext fairness
partition as a phantom real bucket. Treat such a key as having no family
prefix (the whole key is its own family), matching the no-colon case.
The truthiness guard (`if (featureId && ...)`) let an empty-string
featureId bypass the delimiter validation and fall through to the
integration-aggregate key shape, silently fabricating `<dim>:<slug>` for
what the caller meant as a per-feature lookup. Throw loudly instead,
preserving the function's defensive posture.
The JSDoc said data-bearing states are delegated to buildBadge "under
the `health` dimension label", but the code passes "starter" — and must:
the starter ✓/✗/~ glyph vocabulary requires formatLabel's non-health
branch (the health branch renders up/down/stale word labels instead).
rowsAreNoop ignored fail_count, first_failure_at, and id, so an SSE
delta moving only one of them was swallowed as a no-op: first_failure_at
is load-bearing in formatTooltip ("red since ..."), fail_count feeds the
drilldown/alerting surfaces, and a deleted-and-recreated PB row (same
key, fresh id) kept the stale id in the map. Add the three fields to the
comparison and correct the comparator's field-list doc claims.
Both resolvers collapsed a family with a missing sub-row via
`worstState !== "red"` literal equality, silently swallowing
out-of-vocabulary states (e.g. "error") that the A2 worstStateRank
machinery deliberately ranks ABOVE red. Compare by rank instead so a
red-or-worse fold dominates no-data; worstState is typed State but can
hold raw strings at runtime, so the rank form is the honest comparison.
The harness FleetSurfaceState union had no "pending" member and the
derivation mapped EVERY comm error to the red "unreachable" overlay,
contradicting the worker-reclaimed-pending neutral-surface contract that
POOL_COMM_ERROR_KINDS mandates and the dashboard already implements.
Mirror the dashboard cell-model derivation exactly: reclaimed-pending on
a non-red row -> "pending"; a red row passes through unmasked; every
other kind -> "unreachable". Fix the stale binary-derivation docstring
and extend the cross-package drift test to pin the union overlay
members AND the derivation shape on both sides.
TickOptions.filter / EnumerateContext.filter were documented trigger-only
but tick() forwarded the filter unconditionally — a scheduled tick could
be scoped by a stray operator filter. The filter is now forwarded only
when triggered (dropped with a warn otherwise). ServiceJobSpec.priority's
'higher pulls first' doc described behavior that does not exist (claimNext
never reads it) — rewritten as reserved/not-consulted; no priority
behavior added.
nowMs was captured before the potentially-seconds-long enumerate() await,
back-dating lease-expiry decisions and lastSweepAt, and stamping
meta.enqueuedAt (documented 'ISO timestamp the control-plane enqueued the
job') with the tick-START time. The clock is now re-read immediately
before maybeSweep and enqueuedAt is stamped per job at enqueue time. The
default runId factory also reads the injected now() for
injection-discipline consistency (behavior otherwise identical).
Each producer gets an independent default runIdFactory with its counter
starting at 0, so two producers ticking in the same ms with equal tick
counts minted the SAME runId — and the aggregator groups results by
meta.runId. The factory now bakes in a per-factory random discriminator
segment (generated once at creation); ids stay sortable-prefixed by
timestamp.
The enumerate-throw path returned the same shape as a legitimately empty
run (enqueued: 0), so a discovery outage was indistinguishable from an
empty catalog in the tick outcome — the exact ambiguity class sweepFailed
was added to remove. TickResult now carries enumerateFailed, set on the
enumerate-throw path. (control-plane.test.ts fake TickResult gains the
new required field.)
A synchronously-throwing injected fetchImpl escaped the warm loop and
aborted the whole tick before any job was enqueued, violating the 'never
block or fail job production' contract — the per-spec dispatch is now
try/caught. The success handler also discarded the Response without
consuming it (an unread body pins the socket under undici) — the body is
now cancelled best-effort — and the abort timers are unref()'d (guarded)
so a pending warm timer can't hold the process open.
sweepExpired only synthesizes comm errors for rows reclaimed in that call,
so a transient onSweepCommErrors sink failure permanently dropped the
reclaimed jobs' dashboard signal (REQ-B violation) — the old comment's
'next sweep retries' claim was false. The producer now buffers undelivered
comm errors (capped at 500, oldest dropped with a warn) and prepends them
to the next sweep's sink delivery.
- decodePayload validated only half the payload: now asserts
meta.triggered (boolean), meta.enqueuedAt (string), cellIds (string[]
when present) and driverInputs (plain record when present) via a
shared assertServiceJobPayload, failing loud at the boundary.
- enqueue validates the payload BEFORE pb.create (it used to deref
payload.meta.runId after creating the row, persisting a poison row on
a malformed caller payload).
- report()'s refused-release message no longer claims 'nothing was
lost' — the computed result IS discarded and the job re-runs; the
message and comment now say so.
- emptyPayloadForLease carries an explicit HEARTBEAT-ONLY warning: the
placeholder must never feed aggregation (empty runId/serviceSlug
sentinels would corrupt grouping).
- Family discovery warns (queue-client.family-discovery-truncated) when
the 16-family bound trips with families still hidden — previously a
silent starvation; 17-family test.
- escapeLikeLiteral backslash-escapes %/_ for the ~/!~ legs (verified:
PB 0.22.21 builds LIKE ... ESCAPE '\' and skips auto-wrap when the
operand has an unescaped %, and fexpr passes backslashes verbatim), so
a family like 'd%' can no longer occlude d6:/d4: from discovery or
over-count in the backlog gate; = legs keep plain literal escaping.
The gate reads totalItems off a perPage=1 list but never passed
skipTotal. If totals are skipped PB returns totalItems: -1, which is
never above the producer's backlog threshold — the per-tick dedupe gate
would silently FAIL OPEN and enqueue fresh batches on top of an existing
backlog (the exact compounding the gate exists to stop). Pass
skipTotal: false explicitly; test pins the param.
payloadCache entries were only evicted in report()'s finally — a worker
whose renew loses the CAS (lease stolen/swept/terminal) never reports
that job, so its claim-time cache entry stranded forever and the
per-client map grew with every abandoned job. renewLease now evicts on
the lost-CAS return. Tests prove eviction by observing the later renew
take the convenience re-read path (cache miss → pb.getOne) for both the
renew-lost and the claim→report→renew sequences.
The decode-failure path released a won-but-poisoned job as 'failed' with
NO result — and per the file's own contract a terminal resultless row is
synthesized by the result consumer as worker-crashed-mid-job, painting a
FALSE red 'crashed' overlay for what is a payload/protocol problem. The
taxonomy already has worker-protocol-violation for exactly this. After a
successful decode-fail release the client now best-effort writes a
synthetic ServiceJobResult (error state, probe_key aggregate-key
fallback mirroring the worker's comm-error builder) carrying that kind;
a lost write falls back to the consumer's crash synthesis (logged,
swallowed). The result-write retry loop is extracted into a shared
writeResult helper reused by report(). Also covers the two untested
arms: a releaseJob THROW during decode-fail cleanup (warn + continue)
and all-candidates-decode-failing returning { claimed: false }.
sweep-lease-page-truncated fired on ANY exactly-full page — including 50
healthy in-flight jobs at steady state, a guaranteed false positive
every sweep. Under the ascending lease_expires_at sort, truncation only
hides expirable rows when the page TAIL is itself expired (everything
beyond has a later expiry), so the warn now requires a full page AND an
expired tail. Boundary-tested at exactly 50 all-live (no warn) and 50
all-expired (warn).
The stale-pending drain broke out of its page loop whenever a pass
expired nothing, even on a FULL page — but expiry is per-family while
the sort is absolute created, so younger expirable fast-family rows
sitting on later pages behind a full page of not-yet-expirable
slow-family rows were stranded for the whole sweep (the cross-family
occlusion class this drain exists to fix). Now: a non-full page is the
only early exit (the tail was seen); a full page with zero claim
attempts advances the page cursor (nothing left pending, pagination did
not shift); a page with attempts re-lists the same index (rows left
pending — deleted or claim-CAS'd either way — shifted pagination back).
Also corrects the wrong 're-listing returns the same rows forever'
justification: claim-CAS-lost rows LEAVE pending. The per-sweep page
cap still bounds the work.
sweepExpired (and fleet-health's reclaim) decide 'expired' from a listed
SNAPSHOT, then releaseJob(jobId, holder, 'pending') authorizes on
claimed_by alone — a worker renewing between the list and the release
still matches claimed_by, so a live just-renewed job was yanked back to
pending (duplicate execution + a false worker-reclaimed-pending comm
error). The /api/fleet/release hook now refuses a pending-target release
while the row's CURRENT lease is still live, re-checked inside the same
transaction with the leaseExpired helper that stays byte-equivalent to
the client's anchored parse. The client needs no change: released:false
already maps to the sweep's skip path. Pinned by a hook-source parity
test plus a renewed-after-list race test against a hook-faithful fake.
A thrown (not refused) releaseJob in the lease phase escaped the loop,
and the stale phase's pb.list/claimJob were uncaught — either throw
discarded the commErrors already synthesized for rows ALREADY released
to pending (the producer swallows sweepExpired throws), so their gray
're-queued' dashboard surfaces were never rendered and never
regenerated. Now: per-row try/catch around the lease-phase release
(queue-client.sweep-release-threw, continue), and the whole stale drain
is wrapped (queue-client.sweep-stale-phase-threw) with expiredPending
counted per row so a mid-pass throw still returns partial progress.
- Replace the shared ../logger.js module logger with a per-file silent
vi.fn-backed logger rebuilt in afterEach (the shared-logger spy-leak
class: a failed assertion skips mockRestore and poisons sibling files
under fork-reuse).
- makeFakePb now honors status AND probe_key clauses via a shared
LIKE-faithful matcher (incl. PocketBase's ESCAPE '\' semantics) and
THROWS on any clause it can't honor, so multi-family tests can never
pass vacuously; makePagingPb reuses the same matcher.
- Fix sampleResult fixture leaking the driver kind into the probe-key
position (e2e_d6:<slug> -> d6:<slug>; contracts.ts: no e2e_d6 rows).
- Replace the V8-leniency-dependent leading-space leaseExpired test with
STRING-level pins of the exported PB_DATE_SEP_RE anchor (engine
independent; residual V8/goja divergence documented), add a boundary
test at the exact expiry millisecond, and pin the hook source's
anchored regex + t <= Date.now() operator parity (3 handlers).
Composition fix across the cherry-picked sweep branches: the lease phase's
silent re-queue of a STALE_PENDING_SWEEPER row must join requeuedThisSweep,
so the same sweep's multi-page stale drain cannot claim-and-delete a row
the lease phase just re-queued — the retry contract is a LATER sweep. Adds
a composition test pinning that the grace set is honored on EVERY page of
the drain loop (a graced row surfacing only on pass 2 must survive).
- orchestrator.test.ts: file-level afterEach doUnmocks queue-client,
status-writer, and result-consumer (vi.doMock factories persist across
the file; resetModules clears the module cache, not the mock
registry) so a leaked stub can't poison later tests. Full file
re-run: 100/100 pass with the leaks closed — no test was depending on
a leaked factory.
- orchestrator.test.ts: the R5-G4 webhook-secret tests now save/restore
POCKETBASE_URL like the HF13-A2 pattern instead of unconditionally
deleting it in finally.
- job-producer.test.ts: the no-warm test stubbed a local fetch spy it
never wired in (vacuously zero calls); stub GLOBAL fetch via
vi.stubGlobal (+ vi.unstubAllGlobals in afterEach) and assert the
unconfigured producer never falls back to it.
- queue-client.test.ts: famOf re-implemented probeKeyFamily; import the
production helper from contracts so the tests can't drift from the
real family rule.
- control-plane.test.ts: the invalid-cron latch test now also retries
start() on the FAILED instance and asserts it throws again (a
stuck-true latch would make the retry a silent no-op).