Commit Graph

4085 Commits

Author SHA1 Message Date
Jordan Ritter 5dcbe31202 fix(showcase): queue-client/job-claim hardening batch — silent paths get logs, poisoned values fail loud
- decode-failure REFUSED release now warns with the hook's refusal reason
- probeKeySlug empty-slug ('d6:') falls back to the whole probe_key; a
  fully-empty key falls back to unknown-<jobId> (never the forbidden
  empty serviceSlug sentinel)
- countPendingForFamily throws on non-numeric/negative totalItems instead
  of failing the producer backlog gate open with -1
- stale-pending drain contains a thrown claim CAS per row (a 4xx no
  longer aborts the whole drain) and documents the 5xx->won:false
  bounded page-stall on passAttempted
- writeResult paces retries (RESULT_WRITE_RETRY_DELAY_MS, injectable
  sleep) instead of burning attempts back-to-back
- protocol-violation warns for won&&!job / renewed&&!job before falling
  through
- ensureAuth memoizes the in-flight auth promise (no concurrent stampede)
  and wraps the auth-body JSON.parse with status context
- releaseJob interface doc corrected: the hook admits claimed AND running
  rows (the decode-failure cleanup depends on claimed->failed); the doc's
  'running state only' was the wrong side
- documented the LIKE-vs-= ASCII case-sensitivity divergence for
  mixed-case families (unreachable: keys are lowercase slugs)
2026-06-11 20:38:29 -07:00
Jordan Ritter 3c14e0ae15 fix(showcase): sweep discriminates deterministic 4xx release throws from indeterminate failures
The conservative thrown-release path treated EVERY throw as
may-have-committed, so a wedge row (e.g. empty claimed_by -> hook 400 on
the missing workerId) produced a permanent per-sweep false
worker-reclaimed-pending overlay. job-claim now threads the HTTP status
onto thrown endpoint errors (JobClaimEndpointError); the sweep maps a
4xx (hook rejected, nothing committed) to an error-level log with no
grace, no comm error, no reclaimed++, while 5xx/transport/2xx-unreadable
keep the conservative at-least-once handling.
2026-06-11 20:38:28 -07:00
Jordan Ritter 33bfa92039 fix(showcase): report() retry reads the row before rewriting — never un-latch an aggregated result
On the refused_terminal_same_holder retry path, writeResult's
result_processed:false seed could overwrite a row the consumer had
already aggregated (and latched), un-latching it for a second
aggregation — a double-count. The retry now reads the row first: a
present+processed result is skipped (info 'result already aggregated'),
a present+unprocessed result is skipped too (idempotent), only a
missing result falls through to writeResult. A failed pre-write read
throws instead of blind-writing.
2026-06-11 20:38:28 -07:00
Jordan Ritter 208817e7b0 fix(showcase): contain indeterminate renews in queue-client.renewLease — keep the lease assumed-live
With job-claim now throwing on unreadable-2xx (and renew already throwing
on 5xx), a thrown renew escaped queue-client.renewLease into the worker
heartbeat, which treats any renewLease throw as fatal — the heartbeat
died and the sweeper reclaimed a live job (false worker-crashed-mid-job).
renewLease now catches throws from claim.renewLease, warns, and returns
the last-known lease unchanged (no payloadCache eviction, never null) so
the heartbeat retries next beat; only a definitive renewed:false stops
it. Documented at-most-one-lease-duration duplicate-execution risk; the
release CAS still arbitrates terminal writes.
2026-06-11 20:38:27 -07:00
Jordan Ritter 0b84849bb0 fix(showcase): job-claim 2xx-unreadable body throws indeterminate instead of fabricating a CAS loss
A 2xx response whose body fails to read, is empty, or fails to parse was
mapped to {} — i.e. claimed/renewed/released: false — fabricating a CAS
LOSS for a transition the server had already COMMITTED (the renew flavor
recreated the false worker-crashed-mid-job class; the release flavor made
report() falsely declare the result discarded). Now any unreadable 2xx
body THROWS with path context; callers contain the throw.
2026-06-11 20:38:27 -07:00
Jordan Ritter ee1269b96e fix(showcase): whole-key (colon-bearing) families match by equality only in queue-client clauses
probeKeyFamily treats a leading-colon probe_key as its own whole-key
family, so the family value itself contains colons — expanding it with
the <family>:% LIKE leg folded the unrelated family ":foo:bar" under
":foo" in the inclusion/count legs and hid it from discovery via the
exclusion leg. Special-case colon-bearing families to the equality leg
only, implementing the invariant pinned on probeKeyFamily (G3e).
2026-06-11 20:38:26 -07:00
Jordan Ritter 3f898ee0d5 fix(showcase): backlog gate counts non-terminal rows — bound concurrent runs per family
countPendingForFamily filtered on status = "pending" only, so a family
whose whole batch was claimed/running stopped gating and a scheduled
tick could enqueue a fresh batch on top of the in-flight one, doubling
the family's concurrency. Broaden the filter to non-terminal
(pending || claimed || running) and align the gate docs in the contract
and job-producer.
2026-06-11 20:38:26 -07:00
Jordan Ritter 022674ac87 test(harness): harden fleet worker port-closed assertion against port reclaim by a parallel process 2026-06-11 20:38:26 -07:00
Jordan Ritter 3fd3592609 refactor(dashboard): hoist the byte-identity drift parser to module scope (G3c follow-up)
oxlint flagged the inline extract closure (recreated per call); the parser
now sits alongside the file's other parse helpers.
2026-06-11 20:38:25 -07:00
Jordan Ritter ed5eedebfb docs(fleet): pin the equality-only matching invariant for whole-key probe families (G3e)
probeKeyFamily(':foo:bar') returns the WHOLE key as its own family (the
round-2 leading-colon fix), but the queue-client clause builders
(familyInclusionClause/familyExclusionClause) expand any family with a
prefix-LIKE <family>:% leg, which would include/exclude ':foo:bar' under
family ':foo' — disagreeing with the fairness partition. The clause
builders live in queue-client.ts (out of scope here), so this documents the
invariant on the contracts side — whole-key (colon-bearing) families must
match by equality only — and pins the partition behavior in the test suite.
The queue-client special-case is reported to the integrator separately.
2026-06-11 20:38:25 -07:00
Jordan Ritter bdef17c781 docs(fleet): worker-crashed-mid-job enumerates both reporting sources (G3d)
The taxonomy doc defined the kind as exclusively worker-self-monitor-
reported, but the control-plane result consumer ALSO synthesizes it for
terminal-but-resultless rows past the grace window (the queue-client's
bounded result-write retry exhausted — see RESULT_WRITE_MAX_ATTEMPTS).
Both the contracts.ts kind doc and the dashboard mirror now enumerate
the two sources.
2026-06-11 20:38:24 -07:00
Jordan Ritter 08e8827a16 fix(fleet): commErrorFromStatusSignal rejects arrays in both contract copies (G3c)
typeof raw !== "object" admits arrays: an array carrying comm-error fields
as expando properties decoded as a well-formed PoolCommError. Both copies
(harness contracts.ts and the dashboard live-status.ts mirror) now reject
arrays explicitly and stay byte-identical — the drift suite gains a
byte-identity pin on the full function source so a one-sided fix can never
silently re-open the gap.
2026-06-11 20:38:24 -07:00
Jordan Ritter f535f7756f fix(dashboard): resolveCell rollup folds contributor states through worstStateRank (G3b)
The rollup fold used literal includes("red")/includes("degraded") checks,
so an out-of-vocabulary contributor state (e.g. "error", which the harness
can persist at runtime) matched neither bucket and rolled the cell up GRAY
(benign no-data) while its own badge rendered the loud error tone —
violating the documented precedence red > degraded > green > error >
unknown and the A2 never-swallow rule. Contributor states now route through
worstStateRank, so an unknown state rolls up at least red-severity.
2026-06-11 20:38:23 -07:00
Jordan Ritter 830355d3b8 fix(fleet): worker-reclaimed-pending overlay passes through every non-green failure state (G3a)
The pass-through gate in fleetSurfaceState was a literal === "red" check, so
a row whose last-known state was degraded/error/out-of-vocab got masked by
the neutral pending overlay — violating the function's own never-mask-a-
genuine-failure invariant (A2: error ranks ABOVE red). Only green now becomes
pending; the dashboard cell-model derivation mirrors it over its ChipColor
vocabulary (red AND amber pass through). Fixes the self-contradicting JSDoc
closing sentence, the stale live-status.ts mirror doc (union was missing
pending), and extends the drift test to pin the pass-through semantics on
both sides.
2026-06-11 20:38:23 -07:00
Jordan Ritter ffc25bb03a fix(showcase/harness): producer tick/warm-up hygiene batch (CR round 3)
- count a warm GET as fired only after successful dispatch (sync-throwing
  fetchImpl no longer inflates the warmed log)
- route an enumerator resolving to a non-array through the same
  enumerate-failure handling as a throw
- mint the runId AFTER the running check (stopped ticks no longer burn
  the factory counter or log phantom runIds; empty-runId sentinel)
- log tick-start BEFORE the enumerate await so a hung discovery still
  leaves a tick trace (services count moved to tick-complete)
- warn once, with jobIds, when swept comm errors exist but no sink is
  configured (previously dropped with only a count)
- terminal .catch on the fire-and-forget warm chain (a throwing logger
  no longer raises an unhandled rejection)
- document the SweepCommErrorSink at-least-once contract (aggregator
  must be idempotent per jobId+observedAt)
- tests: fix stale "Four producers" comment; replace the fragile
  two-shot Math.random spy with a never-exhausting deterministic
  sequence plus a call-count guard
2026-06-11 20:38:22 -07:00
Jordan Ritter 7ec2890b40 fix(showcase/harness): guard producer tick re-entrancy (backlog-gate TOCTOU double-enqueue)
A slow tick overlapping the next cron tick let both read the backlog
gate's pending count before either enqueued, so both produced a batch
for the same family — the motivating concurrent same-service-run
incident class. An overlapping tick is now skipped (warned) via the
tracked in-flight tick promise; it returns the empty-runId no-op result
without minting a run.

NOTE for the integrator: the gate also under-counts — queue-client's
countPendingForFamily filters status="pending" only, so a fully-claimed
but-still-running batch doesn't gate. Broadening the filter to
non-terminal rows (pending || claimed || running) is a one-line edit in
queue-client.ts (plus the contracts.ts doc), both sibling-owned —
deferred to the integration pass.
2026-06-11 20:38:22 -07:00
Jordan Ritter 4b482c37cd fix(showcase/harness): make producer stop() quiesce — await in-flight tick, truncate mid-batch, final comm-error drain
stop() used to flip flags and resolve while an in-flight tick kept
enumerating/sweeping/enqueueing, and buffered sweep comm errors were
silently dropped at shutdown. Now stop() awaits the tracked in-flight
tick promise, the enqueue loop re-checks `running` and truncates loudly
when stopped mid-batch, and stop() makes one final sink delivery of the
buffered comm errors — logging dropped count + jobIds at error level if
that last attempt fails too.
2026-06-11 20:38:21 -07:00
Jordan Ritter c2f8383b8a fix(showcase/harness): drain buffered sweep comm errors even when the sweep itself throws
The undelivered comm-error buffer was drained only on maybeSweep's try
SUCCESS path, so a persistently-throwing sweepExpired — the exact
failure mode the buffer exists to ride out — never handed the buffered
batch to a now-healthy sink. Extract the delivery into
deliverSweepCommErrors and run it on every sweep attempt, including the
catch arm.
2026-06-11 20:38:21 -07:00
Jordan Ritter bd7c02f966 test(showcase): close the vacuous-pass holes in the queue-client fakes (CR G1g)
rowMatchesFilter now THROWS on operators it can't honor (<,>,<=,>=,?=,
?~) and on status clauses with !=/~/!~ instead of silently matching all
rows, with the quoted-literal false-positive limitation documented;
both fakes return totalItems:-1 (and -1 totalPages) under skipTotal —
faithful to PB, closing the fail-open count class beyond the single
countPendingForFamily pin; makePagingPb honors the '-' sort-direction
prefix and computes totalPages honestly; makeOrderedPb's
ignore-all-filters behavior is documented as deliberate coupling to the
duplicate-family defensive break; the result-write retry bound is
pinned with the exported RESULT_WRITE_MAX_ATTEMPTS constant.
2026-06-11 20:38:20 -07:00
Jordan Ritter dc3acb8427 fix(showcase): queue-client robustness batch (CR G1f)
(i) per-candidate try/catch in claimNext's CAS race — a thrown transport
claim no longer aborts the whole rotation (warn + next candidate);
(ii) discoverPendingFamilies' duplicate-family defensive break now warns
before breaking; (iii) renew re-read triage split — a decodePayload
failure logs as a protocol violation, not a read blip; (iv) the
decode-failure synthetic result uses an injected clock (new config.now)
instead of new Date(); (v) backslash charset guard for family clause
building (the equality/LIKE escape contracts contradict for backslash —
skip such families with a warn) and the over-claiming VERIFIED comment
scoped to %/_ only; (vi) single-sweeper assumption documented as
load-bearing at the grace-set declaration; (vii) hook comment misname
fixed (worker-crashed-mid-job -> worker-reclaimed-pending); (viii)
claimJob maps a 5xx claim response to a lost CAS (won:false, warn) —
a WAL serialization error escaping runInTransaction surfaces as 500 —
while 4xx and renew/release 5xx still throw loud.
2026-06-11 20:38:19 -07:00
Jordan Ritter 5730b86f94 fix(showcase): make report() retryable via release refusal reasons (CR G1e)
After a release-CAS success + result-write exhaustion, report() throws;
a natural retry got REFUSED (row already terminal) and emitted an error
claiming the result is discarded and the job re-runs — both false (the
result is still writable by this holder; terminal rows never re-run).
The hook's release response now carries a refusal reason
(refused_terminal_same_holder / refused_not_holder /
refused_lease_live), threaded through job-claim's ReleaseResult.
report() treats refused-terminal-under-my-workerId as the second leg of
a timeout-after-commit retry and proceeds to writeResult; the
not-holder error is reworded to 're-runs only if reclaimed to pending'.
A reason-less refusal still fails closed.
2026-06-11 20:38:19 -07:00
Jordan Ritter 937ec924de fix(showcase): stop the next sweep from claim-deleting a re-queued long-runner (CR G1d)
Stale-pending age is anchored on PB created and never re-anchored on
re-queue: a job running longer than its family expiry window got
lease-reclaimed with a 'back in flight' comm error, then claim-deleted
by the NEXT sweep before any plausible re-run — the dashboard
permanently showed 're-queued' for silently-discarded work. Schema-free
fix: the release hook now RETAINS the expired lease_expires_at on a
pending re-queue (claim admits pending rows regardless of lease), and
the stale phase skips rows whose retained lease is recent (parseable
and within now - familyExpiryWindow) — recently in flight means the
created-based age is stale evidence. Comments cover the heuristic, the
requeued_at-column alternative, and the sweeper-garbage lingering
tradeoff; hook parity test pins the retention.
2026-06-11 20:38:19 -07:00
Jordan Ritter abe27b38d5 fix(showcase): populate serviceSlug/runId on the synthetic protocol-violation result (CR G1c)
The decode-failure synthesis persisted serviceSlug: "" and runId: "" —
the exact empty sentinels this file's own emptyPayloadForLease warning
forbids feeding aggregation (an empty runId groups into nothing, an
empty serviceSlug corrupts the per-service rollup). Recover the slug
from the probe_key's slug segment (new probeKeySlug, the complement of
contracts.ts probeKeyFamily) and mint a non-colliding pviol_<jobId>
runId; reconcile the contradictory comments on both sides. Verified
downstream: result-consumer/aggregator key on aggregateKey + dedupe on
jobId, with serviceSlug/runId riding into logs and batch grouping.
2026-06-11 20:38:18 -07:00
Jordan Ritter 39b154ee6c fix(showcase): treat thrown sweep release as committed — grace + synthesize comm error (CR G1b)
A release that THROWS in the lease sweep may have COMMITTED server-side
(timeout-after-commit): the row is then pending but absent from the
grace set — the same call's stale phase could claim-and-delete it — and
its worker-reclaimed-pending comm error was never synthesized, losing
the gray 're-queued' surface forever (no later sweep re-emits pending
rows). Conservatively add the row to the grace set and synthesize the
comm error on a throw (at-least-once: a duplicate gray overlay is
harmless, a missing one is not); sweeper-held rows stay silent,
mirroring the committed path.
2026-06-11 20:38:18 -07:00
Jordan Ritter 547e6f2841 fix(showcase): enforce superuser auth + leaseSeconds clamp on fleet claim endpoints (CR G1a)
The three /api/fleet/* routerAdd handlers carried no auth middleware —
a middleware-less PB 0.22 routerAdd handler is PUBLIC, so any
unauthenticated caller could claim/renew/release arbitrary jobs despite
the header claiming superuser auth was required. Append
$apis.requireAdminAuth() (verified against PB 0.22.21 JSVM types) to
all three routes; the client already authenticates as superuser with a
401-reauth retry, so enforcement is compat-safe. Also clamp
leaseSeconds in claim+renew (numeric only, 3600s ceiling, 30s default
on garbage) and pin both contracts in the hook-source parity tests.
2026-06-11 20:38:17 -07:00
Jordan Ritter a8a44a9cd9 docs(harness): align meta.priority JSDoc with claimNext reality — reserved, not consulted 2026-06-11 20:38:17 -07:00
Jordan Ritter 3686a82987 fix(harness): probeKeyFamily treats a leading-colon key as its own family
A probe key beginning with ":" yielded the empty-string family, which
then flowed into countPendingForFamily and the claimNext fairness
partition as a phantom real bucket. Treat such a key as having no family
prefix (the whole key is its own family), matching the no-colon case.
2026-06-11 20:37:35 -07:00
Jordan Ritter e5eceb2e76 fix(shell-dashboard): keyFor throws on empty-string featureId
The truthiness guard (`if (featureId && ...)`) let an empty-string
featureId bypass the delimiter validation and fall through to the
integration-aggregate key shape, silently fabricating `<dim>:<slug>` for
what the caller meant as a per-feature lookup. Throw loudly instead,
preserving the function's defensive posture.
2026-06-11 20:37:34 -07:00
Jordan Ritter a2f666376d docs(shell-dashboard): correct buildStarterBadge JSDoc dimension claim
The JSDoc said data-bearing states are delegated to buildBadge "under
the `health` dimension label", but the code passes "starter" — and must:
the starter ✓/✗/~ glyph vocabulary requires formatLabel's non-health
branch (the health branch renders up/down/stale word labels instead).
2026-06-11 20:37:34 -07:00
Jordan Ritter d368fd986c fix(shell-dashboard): compare id, fail_count, first_failure_at in rowsAreNoop
rowsAreNoop ignored fail_count, first_failure_at, and id, so an SSE
delta moving only one of them was swallowed as a no-op: first_failure_at
is load-bearing in formatTooltip ("red since ..."), fail_count feeds the
drilldown/alerting surfaces, and a deleted-and-recreated PB row (same
key, fresh id) kept the stale id in the map. Add the three fields to the
comparison and correct the comparator's field-list doc claims.
2026-06-11 20:37:34 -07:00
Jordan Ritter bb2484d1ed fix(shell-dashboard): rank-based anyMissing collapse in resolveD5Row/resolveD6Row
Both resolvers collapsed a family with a missing sub-row via
`worstState !== "red"` literal equality, silently swallowing
out-of-vocabulary states (e.g. "error") that the A2 worstStateRank
machinery deliberately ranks ABOVE red. Compare by rank instead so a
red-or-worse fold dominates no-data; worstState is typed State but can
hold raw strings at runtime, so the rank form is the honest comparison.
2026-06-11 20:37:34 -07:00
Jordan Ritter f83d1ea2ca fix(harness): route worker-reclaimed-pending to the neutral 'pending' surface in fleetSurfaceState
The harness FleetSurfaceState union had no "pending" member and the
derivation mapped EVERY comm error to the red "unreachable" overlay,
contradicting the worker-reclaimed-pending neutral-surface contract that
POOL_COMM_ERROR_KINDS mandates and the dashboard already implements.
Mirror the dashboard cell-model derivation exactly: reclaimed-pending on
a non-red row -> "pending"; a red row passes through unmasked; every
other kind -> "unreachable". Fix the stale binary-derivation docstring
and extend the cross-package drift test to pin the union overlay
members AND the derivation shape on both sides.
2026-06-11 20:37:33 -07:00
Jordan Ritter 1f9ae1e43f fix(harness): enforce trigger-only filter; correct priority doc (G2f)
TickOptions.filter / EnumerateContext.filter were documented trigger-only
but tick() forwarded the filter unconditionally — a scheduled tick could
be scoped by a stray operator filter. The filter is now forwarded only
when triggered (dropped with a warn otherwise). ServiceJobSpec.priority's
'higher pulls first' doc described behavior that does not exist (claimNext
never reads it) — rewritten as reserved/not-consulted; no priority
behavior added.
2026-06-11 20:37:33 -07:00
Jordan Ritter bfe60adecf fix(harness): read the clock at sweep/enqueue time, not tick start (G2e)
nowMs was captured before the potentially-seconds-long enumerate() await,
back-dating lease-expiry decisions and lastSweepAt, and stamping
meta.enqueuedAt (documented 'ISO timestamp the control-plane enqueued the
job') with the tick-START time. The clock is now re-read immediately
before maybeSweep and enqueuedAt is stamped per job at enqueue time. The
default runId factory also reads the injected now() for
injection-discipline consistency (behavior otherwise identical).
2026-06-11 20:37:32 -07:00
Jordan Ritter dfde37af5b fix(harness): de-collide default runId factory across producer instances (G2d)
Each producer gets an independent default runIdFactory with its counter
starting at 0, so two producers ticking in the same ms with equal tick
counts minted the SAME runId — and the aggregator groups results by
meta.runId. The factory now bakes in a per-factory random discriminator
segment (generated once at creation); ids stay sortable-prefixed by
timestamp.
2026-06-11 20:37:32 -07:00
Jordan Ritter f24045d829 fix(harness): report enumerate failure distinctly in TickResult (G2c)
The enumerate-throw path returned the same shape as a legitimately empty
run (enqueued: 0), so a discovery outage was indistinguishable from an
empty catalog in the tick outcome — the exact ambiguity class sweepFailed
was added to remove. TickResult now carries enumerateFailed, set on the
enumerate-throw path. (control-plane.test.ts fake TickResult gains the
new required field.)
2026-06-11 20:37:31 -07:00
Jordan Ritter da46297ba5 fix(harness): harden warm-up loop — sync-throw isolation, body cancel, timer unref (G2b)
A synchronously-throwing injected fetchImpl escaped the warm loop and
aborted the whole tick before any job was enqueued, violating the 'never
block or fail job production' contract — the per-spec dispatch is now
try/caught. The success handler also discarded the Response without
consuming it (an unread body pins the socket under undici) — the body is
now cancelled best-effort — and the abort timers are unref()'d (guarded)
so a pending warm timer can't hold the process open.
2026-06-11 20:37:30 -07:00
Jordan Ritter 18e12b198f fix(harness): buffer undelivered sweep comm errors and redeliver next sweep (G2a)
sweepExpired only synthesizes comm errors for rows reclaimed in that call,
so a transient onSweepCommErrors sink failure permanently dropped the
reclaimed jobs' dashboard signal (REQ-B violation) — the old comment's
'next sweep retries' claim was false. The producer now buffers undelivered
comm errors (capped at 500, oldest dropped with a warn) and prepends them
to the next sweep's sink delivery.
2026-06-11 20:37:30 -07:00
Jordan Ritter 4d65fdb741 fix(harness): queue-client small batch — full payload validation, pre-create enqueue check, honest messages, discovery observability, LIKE-wildcard escaping
- decodePayload validated only half the payload: now asserts
  meta.triggered (boolean), meta.enqueuedAt (string), cellIds (string[]
  when present) and driverInputs (plain record when present) via a
  shared assertServiceJobPayload, failing loud at the boundary.
- enqueue validates the payload BEFORE pb.create (it used to deref
  payload.meta.runId after creating the row, persisting a poison row on
  a malformed caller payload).
- report()'s refused-release message no longer claims 'nothing was
  lost' — the computed result IS discarded and the job re-runs; the
  message and comment now say so.
- emptyPayloadForLease carries an explicit HEARTBEAT-ONLY warning: the
  placeholder must never feed aggregation (empty runId/serviceSlug
  sentinels would corrupt grouping).
- Family discovery warns (queue-client.family-discovery-truncated) when
  the 16-family bound trips with families still hidden — previously a
  silent starvation; 17-family test.
- escapeLikeLiteral backslash-escapes %/_ for the ~/!~ legs (verified:
  PB 0.22.21 builds LIKE ... ESCAPE '\' and skips auto-wrap when the
  operand has an unescaped %, and fexpr passes backslashes verbatim), so
  a family like 'd%' can no longer occlude d6:/d4: from discovery or
  over-count in the backlog gate; = legs keep plain literal escaping.
2026-06-11 20:37:29 -07:00
Jordan Ritter 714178ae60 fix(harness): request totals explicitly in countPendingForFamily — backlog gate must not fail open
The gate reads totalItems off a perPage=1 list but never passed
skipTotal. If totals are skipped PB returns totalItems: -1, which is
never above the producer's backlog threshold — the per-tick dedupe gate
would silently FAIL OPEN and enqueue fresh batches on top of an existing
backlog (the exact compounding the gate exists to stop). Pass
skipTotal: false explicitly; test pins the param.
2026-06-11 20:37:29 -07:00
Jordan Ritter 9064fb4e87 fix(harness): evict the payload cache when a renew CAS is lost (leak on abandoned/stolen jobs)
payloadCache entries were only evicted in report()'s finally — a worker
whose renew loses the CAS (lease stolen/swept/terminal) never reports
that job, so its claim-time cache entry stranded forever and the
per-client map grew with every abandoned job. renewLease now evicts on
the lost-CAS return. Tests prove eviction by observing the later renew
take the convenience re-read path (cache miss → pb.getOne) for both the
renew-lost and the claim→report→renew sequences.
2026-06-11 20:37:28 -07:00
Jordan Ritter ad17e8542a fix(harness): attribute decode-failed jobs as worker-protocol-violation, not a synthetic crash
The decode-failure path released a won-but-poisoned job as 'failed' with
NO result — and per the file's own contract a terminal resultless row is
synthesized by the result consumer as worker-crashed-mid-job, painting a
FALSE red 'crashed' overlay for what is a payload/protocol problem. The
taxonomy already has worker-protocol-violation for exactly this. After a
successful decode-fail release the client now best-effort writes a
synthetic ServiceJobResult (error state, probe_key aggregate-key
fallback mirroring the worker's comm-error builder) carrying that kind;
a lost write falls back to the consumer's crash synthesis (logged,
swallowed). The result-write retry loop is extracted into a shared
writeResult helper reused by report(). Also covers the two untested
arms: a releaseJob THROW during decode-fail cleanup (warn + continue)
and all-candidates-decode-failing returning { claimed: false }.
2026-06-11 20:37:28 -07:00
Jordan Ritter 1619b9e6a0 fix(harness): gate the lease-sweep truncation warn on an expired page tail
sweep-lease-page-truncated fired on ANY exactly-full page — including 50
healthy in-flight jobs at steady state, a guaranteed false positive
every sweep. Under the ascending lease_expires_at sort, truncation only
hides expirable rows when the page TAIL is itself expired (everything
beyond has a later expiry), so the warn now requires a full page AND an
expired tail. Boundary-tested at exactly 50 all-live (no warn) and 50
all-expired (warn).
2026-06-11 20:37:27 -07:00
Jordan Ritter db78a5c68b fix(harness): page past full no-attempt pages in the stale drain — close cross-family occlusion
The stale-pending drain broke out of its page loop whenever a pass
expired nothing, even on a FULL page — but expiry is per-family while
the sort is absolute created, so younger expirable fast-family rows
sitting on later pages behind a full page of not-yet-expirable
slow-family rows were stranded for the whole sweep (the cross-family
occlusion class this drain exists to fix). Now: a non-full page is the
only early exit (the tail was seen); a full page with zero claim
attempts advances the page cursor (nothing left pending, pagination did
not shift); a page with attempts re-lists the same index (rows left
pending — deleted or claim-CAS'd either way — shifted pagination back).
Also corrects the wrong 're-listing returns the same rows forever'
justification: claim-CAS-lost rows LEAVE pending. The per-sweep page
cap still bounds the work.
2026-06-11 20:37:27 -07:00
Jordan Ritter 5362f93993 fix(showcase): re-check lease expiry on pending-target releases — close the sweeper TOCTOU
sweepExpired (and fleet-health's reclaim) decide 'expired' from a listed
SNAPSHOT, then releaseJob(jobId, holder, 'pending') authorizes on
claimed_by alone — a worker renewing between the list and the release
still matches claimed_by, so a live just-renewed job was yanked back to
pending (duplicate execution + a false worker-reclaimed-pending comm
error). The /api/fleet/release hook now refuses a pending-target release
while the row's CURRENT lease is still live, re-checked inside the same
transaction with the leaseExpired helper that stays byte-equivalent to
the client's anchored parse. The client needs no change: released:false
already maps to the sweep's skip path. Pinned by a hook-source parity
test plus a renewed-after-list race test against a hook-faithful fake.
2026-06-11 20:37:27 -07:00
Jordan Ritter 116083cbb8 fix(harness): contain sweepExpired throws so partial reclaim results survive (REQ-B)
A thrown (not refused) releaseJob in the lease phase escaped the loop,
and the stale phase's pb.list/claimJob were uncaught — either throw
discarded the commErrors already synthesized for rows ALREADY released
to pending (the producer swallows sweepExpired throws), so their gray
're-queued' dashboard surfaces were never rendered and never
regenerated. Now: per-row try/catch around the lease-phase release
(queue-client.sweep-release-threw, continue), and the whole stale drain
is wrapped (queue-client.sweep-stale-phase-threw) with expiredPending
counted per row so a mid-pass throw still returns partial progress.
2026-06-11 20:37:26 -07:00
Jordan Ritter 17b6ac9439 test(harness): queue-client test hygiene — silent per-file logger, honest fakes, hook-parity pins
- Replace the shared ../logger.js module logger with a per-file silent
  vi.fn-backed logger rebuilt in afterEach (the shared-logger spy-leak
  class: a failed assertion skips mockRestore and poisons sibling files
  under fork-reuse).
- makeFakePb now honors status AND probe_key clauses via a shared
  LIKE-faithful matcher (incl. PocketBase's ESCAPE '\' semantics) and
  THROWS on any clause it can't honor, so multi-family tests can never
  pass vacuously; makePagingPb reuses the same matcher.
- Fix sampleResult fixture leaking the driver kind into the probe-key
  position (e2e_d6:<slug> -> d6:<slug>; contracts.ts: no e2e_d6 rows).
- Replace the V8-leniency-dependent leading-space leaseExpired test with
  STRING-level pins of the exported PB_DATE_SEP_RE anchor (engine
  independent; residual V8/goja divergence documented), add a boundary
  test at the exact expiry millisecond, and pin the hook source's
  anchored regex + t <= Date.now() operator parity (3 handlers).
2026-06-11 20:37:25 -07:00
Jordan Ritter 4e00b56e6a style(harness): oxfmt formatting on cherry-picked fleet tests 2026-06-11 20:37:25 -07:00
Jordan Ritter 7cca03e99e fix(harness): feed the stale-sweeper silent re-queue into the sweep grace set
Composition fix across the cherry-picked sweep branches: the lease phase's
silent re-queue of a STALE_PENDING_SWEEPER row must join requeuedThisSweep,
so the same sweep's multi-page stale drain cannot claim-and-delete a row
the lease phase just re-queued — the retry contract is a LATER sweep. Adds
a composition test pinning that the grace set is honored on EVERY page of
the drain loop (a graced row surfacing only on pass 2 must survive).
2026-06-11 20:37:24 -07:00
Jordan Ritter f0b4c328e2 test(harness): close mock/env leaks and tighten fleet test hygiene
- orchestrator.test.ts: file-level afterEach doUnmocks queue-client,
  status-writer, and result-consumer (vi.doMock factories persist across
  the file; resetModules clears the module cache, not the mock
  registry) so a leaked stub can't poison later tests. Full file
  re-run: 100/100 pass with the leaks closed — no test was depending on
  a leaked factory.
- orchestrator.test.ts: the R5-G4 webhook-secret tests now save/restore
  POCKETBASE_URL like the HF13-A2 pattern instead of unconditionally
  deleting it in finally.
- job-producer.test.ts: the no-warm test stubbed a local fetch spy it
  never wired in (vacuously zero calls); stub GLOBAL fetch via
  vi.stubGlobal (+ vi.unstubAllGlobals in afterEach) and assert the
  unconfigured producer never falls back to it.
- queue-client.test.ts: famOf re-implemented probeKeyFamily; import the
  production helper from contracts so the tests can't drift from the
  real family rule.
- control-plane.test.ts: the invalid-cron latch test now also retries
  start() on the FAILED instance and asserts it throws again (a
  stuck-true latch would make the retry a silent no-op).
2026-06-11 20:37:24 -07:00