The fail-closed gate matched a message substring duplicated in 3 places;
any rewording silently flipped it fail-closed -> fail-open. The refusal is
now a dedicated exported class, the producer gates via instanceof, the test
literal copies use the real class, and a drift test drives the REAL
queue-client refusal through the producer gate.
A resolved sink call clears the producer's undelivered comm-error buffer
(deliverSweepCommErrors' at-least-once contract), so the late-bound sink
silently 'succeeding' before controlPlaneRef was assigned dropped the
batch permanently. Extract buildSweepCommErrorSink (exported for tests)
and reject on the unbound window so the producer re-buffers and
redelivers after assembly; pin with red-green unit tests.
- assertServiceJobPayload rejects an empty meta.enqueuedAt — the
emptyPayloadForLease never-aggregate sentinel, same forbidden class as
the other empties (the minted fallback never crosses this boundary)
- 408/429 are excluded from the deterministic-4xx class (shared
deterministicEndpointRejection helper): transient by meaning, they keep
the assumed-live renew and at-least-once sweep containments
- the family-discovery truncation warn threads nextFamilyClauseSafe so a
clause-unsafe garbage head can't masquerade as a hidden claimable family
- fix 'CONTROl' comment typo in queue-client.test.ts
Enforce the documented jobId/workerId equality between the top-level
input and the echoed result (a mismatched caller would release one row
while filing the result under another), and move the terminalJobStatus
computation inside the try so a malformed result can no longer throw
past the finally and leak the claim-time cache entry.
NOTE: contracts.ts's INTEGRATOR NOTE on ReportJobInput still says report()
does not validate this — contracts.ts is sibling-owned this round and is
NOT touched here; the note needs a follow-up edit.
The long-expired carve-out required a finite parsed lease, so an expired
claimed/running row with an unparseable lease_expires_at was re-queued
with a 'back in flight' worker-reclaimed-pending signal — which the next
sweep falsified by claim-deleting the row (the recent-lease protection
also requires a finite lease). An unparseable lease carries no
recent-flight evidence either phase can honor: treat it as long-expired
and delete directly so the emitted signal matches the outcome. Pinned
with a two-sweep test.
An array is typeof 'object' and can carry expando fields that satisfy
every per-field check — the same hole the nested meta check's
Array.isArray guard already closes one level down. Reject it at the
shared enqueue/decode boundary.
authenticate() threw plain Errors for every non-2xx, so the renew and
sweep deterministic-4xx carve-outs (instanceof JobClaimEndpointError +
4xx) never fired for auth failures — rotated creds rode the indeterminate
path forever: a phantom assumed-live lease per renew beat and a false
worker-reclaimed-pending comm error per sweep. 4xx auth responses now
carry the discriminable class (auth path + status); 5xx/network stay
plain (indeterminate). End-to-end pins go through the real job-claim
client for both the renew and sweep halves.
The pending-target release could never succeed: the hook refuses every
pending-target release on a live lease (refused_lease_live, no holder
exemption) and refused_not_holder otherwise, so the containment was inert
— the row wedged a lease window and got a false worker-reclaimed-pending
overlay. Mirror the decode-failure containment (terminal target passes
the live-lease gate) and share its synthetic-result builder. Tests now
use hook-faithful release fakes instead of unconditional released:true.
- report ordering assertion made non-vacuous: capture rows[0].result INSIDE
the releaseJob fake and assert undefined at release time.
- empty-slug-segment fixture carries probe_key 'd6:' on the claim fake's
returned view too (no pinning of the internal key source).
- job-claim.test.ts: file-local silent logger replacing the shared logger
import (spy-leak class under fork-reuse).
- auth-order tests: auth + endpoint URLs recorded in one call log, order
asserted by index; fallback test pins _superusers-first order and the
Authorization header carried to the endpoint.
- makeFakePb.list honesty: models created/lease_expires_at sorts (throws on
unmodeled keys) and honors perPage/page truncation; self-tested in the
fake-honesty suite.
- documented the interleaved-|| connective-guard limitation in
rowMatchesFilter.
- renamed the renew 'convenience re-read returns null' test to the cache-hit
pin it actually covers, with expect(getOneSpy).not.toHaveBeenCalled().
- report() retry: a null getOne resolution is a FAILED read (throw, no blind
write) and a "" result is PB's unset-JSON shape (absent → write proceeds).
- sweepExpired: thrown-release conservative maybes now counted on a separate
reclaimedIndeterminate (SweepResultWithIndeterminate); reclaimed counts only
CAS-confirmed re-queues. Producer one-liner documented for when the
sibling-owned TickResult gains the field.
- job-claim 401 retry: snapshot the token the failed request used; only null
authToken if unchanged (no clobbering a concurrently refreshed token).
- decode-failure synthetic result write: single attempt, no 250ms retry pacing
inside the claim race (consumer crash-synthesis is the documented backstop).
- fleet-claim.pb.js: typeof jobId !== "string" → 400 in all three handlers.
- docs: recent-lease bound is expiryPeriods × period (not one window); claim
5xx→won:false bounded false-overlay source; report retryability deploy-skew
note; drainStalePending page-advance indeterminacy note.
The won:true/no-job protocol breach may still have committed the claim, so
falling through abandoned a row this worker owns — wedging a full lease window
and producing a false worker-reclaimed-pending overlay on the next sweep.
Mirror the decode-failure containment: best-effort releaseJob(id, worker,
"pending") (no work happened) before continuing, with refusals and throws
swallowed+warned.
A row whose lease expired beyond its family's stale window is past the
recent-lease heuristic's cross-sweep protection: re-queueing it emitted a
'back in flight' worker-reclaimed-pending signal that the NEXT sweep falsified
by stale-deleting the row. When a lease-phase row is already stale-expirable
(parseable created past maxAgeMs AND lease expired longer than the window),
claim it under the sweeper id and delete it — no comm error, counted in
expiredPending. Recently-expired rows keep today's re-queue path; header's
cross-call-protection claim corrected.
A backslash (or empty) family used to BREAK family discovery, starving every
younger family behind the offending row for up to 3h. The offending ROW needs
no family-charset semantics to exclude — push an id != "<rowId>" clause and
continue, keeping the warn. Younger families are still discovered; the unsafe
row stays unclaimable.
probeKeyFamily("") takes the prefix-LIKE clause leg, so the empty family's
inclusion clause matched every leading-colon key and its exclusion clause hid
ALL leading-colon families from discovery. Treat family "" as clause-unsafe
in familyClauseSafe (skip+warn, like the backslash class), and require
non-empty probeKey/serviceSlug/driverKind/meta.runId in assertServiceJobPayload
so the documented forbidden empty sentinels fail loud at both the enqueue and
claim-decode boundaries.
A renewed:true response missing its job view is a protocol violation on a
SUCCESSFUL renew — the worker still holds the live lease. Evicting + returning
null stopped the heartbeat and let the sweeper falsely reclaim the live job.
Return the cached lease assumed-live (like the indeterminate path); only
evict+null when no cached lease exists. Also corrects the stale 'We ONLY
return null when the CAS itself failed' comment.
A JobClaimEndpointError 4xx from the renew endpoint fails identically every
beat, so the indeterminate assumed-live containment never converged — the
worker held a phantom lease forever while the server lease lapsed and another
worker double-ran the job. Mirror the sweep's 4xx carve-out: error-log
renew-rejected, evict payload+lease caches, return null. 5xx/transport/
2xx-unreadable keep the assumed-live path.
The global lease sweep mirrors comm errors onto the status row keyed by the
reclaimed job's probe_key (resolveSweepAggregateKey → aggregateCommError) for
ALL four fleet families, but the dashboard's decodeCellCommError candidate
scan only covered the d6 family's d6:<slug> aggregate. The smoke/demos/deep
families land on d4:<slug>, e2e-demos:<slug>, and d5-single-pill-e2e:<slug> —
rows the dashboard reads nowhere else — so reclaim/crash overlays on those
families were invisible. Add the three candidate keys (e2e window).
- Object.freeze the UNSUPPORTED CellModel + NOT_WIRED_LEVEL singletons —
they are returned by reference to every caller, so one consumer mutating
them would corrupt every unsupported/not-wired cell.
- Remove the unfirable !d5.exists chip arm (d5.exists and d6.exists derive
from the SAME CATALOG_TO_D5_KEY entry, so !d5.exists && d6.exists is
impossible); fix the decision-table and d6Effective docs that described
the unfirable behavior, keeping a note for a future map split.
- Constrain upsertByKey's generic to StatusRow so rowsAreNoop can never
vacuously match a non-StatusRow type whose fields are all undefined.
- keyFor: validate the dimension segment for ':'/'/' like slug/featureId.
- decodeCellCommError: scope the comm-error staleness window per row family
(e2e 6h for d5/d6/e2e, D4 1h for chat/tools, liveness 45m for health)
instead of applying the 6h E2E window to liveness-cadence rows.
The three-way contradiction: harness fleetSurfaceState routes green-only →
pending, cell-model routes green OR gray (non-regression) → pending while its
own comment claimed 'only green', and both files claimed 'exact mirror'.
gray→pending IS intended dashboard-side (gray is the dashboard-only no-data
colour ProbeState cannot represent). Fix the comments on both sides to state
the asymmetry precisely, extend the drift test to pin the derivation
difference itself (harness green-equality + no gray; dashboard
failure-passthrough + isRegression), add behavioral gray→pending and
red-passthrough tests, and fix chipColorToSurface's stale union doc
(now ChipColor | unreachable | pending).
The stale-green branch returned the RAW row (status "amber" with row.state
still "green"), violating the .row.state ↔ .status invariant resolveD4/D5/D6
maintain. Return { ...row, state: "degraded" } like the other resolvers.
stateToTestStatus mapped unknown runtime states (e.g. "error") to null,
swallowing the A2 rank-fold winner one step after the fold surfaced it — the
D5/D6 chip went benign gray no-data while live-status rendered the loud
"error" tone for the same row. Add foldStateToTestStatus (out-of-vocab →
"red") for the D5/D6 resolvers; D3/D4 keep the base mapping since the
chip's D1-D4 gate check already rescues their null (pinned by tests).
The strict missing-sub-row collapse used the literal `worstState !== "red"`,
which silently swallowed out-of-vocabulary runtime states (e.g. "error") that
the A2 rank machinery deliberately ranks ABOVE red — collapsing exactly the
state the rank fold exists to surface into benign gray no-data. Use
`rankOfState(worstState) < STATE_RANK.red`, mirroring the already-fixed
resolveD5Row/resolveD6Row in live-status.ts.
Follow-up to the TickResult.truncatedByStop accounting field — the
control-plane wiring test's producer fake builds a literal TickResult
and needed the new required field for tsc to stay clean.
A failed assertion in the loop-crash test leaked the worker's bound
/health server (and the default-boot pool) across tests because stop()
was only reached via the in-body rejects.toThrow. Mirror the sibling
default-boot test's try/finally pattern with a rejection-swallowing
safety-net stop.
- derive the buffer-cap overflow boundary from
MAX_BUFFERED_SWEEP_COMM_ERRORS instead of hardcoding 'old100'
- race the re-entrancy-guard's overlapping tick against a short timer so
a broken guard fails diagnostically instead of deadlocking
- pin the warm-up abort test's AbortController+setTimeout coupling
(AbortSignal.timeout would evade fake timers)
- document makeFakeQueue.enqueued as recording enqueue ATTEMPTS
- align queue-client samplePayload driverKind with the probeKey family
so future payload-coherence cross-validation doesn't mass-fail
Stop-truncated specs vanished from the tick outcome: tick-complete's
'services' could not be reconciled against enqueued + enqueueFailures +
skippedForBacklog. Add truncatedByStop to TickResult and the
tick-complete log (invariant: services == enqueued + enqueueFailures +
skippedForBacklog + truncatedByStop) and fix the 'enqueued' doc to
mention stop truncation.
The one-shot no-sink warn kept a sink-less deployment from burying its
logs, but made every subsequent drop completely traceless. Later drops
now log at debug with the dropped jobIds and a running dropped-total
counter; the counter also rides the first full warn.
The first stop() flips 'running' synchronously, so a second concurrent
stop() hit the !running early-return and resolved immediately while the
tick the first stop() was quiescing on kept enqueueing. Both stop()
paths now loop-await the in-flight tick AND the queued trigger before
resolving.
The producer's re-entrancy guard skipped TRIGGERED ticks too, silently
losing an explicit operator 'run it NOW' whenever a slow scheduled tick
was in flight — contradicting the operator-intent-wins rationale that
already lets triggers bypass the backlog gate. A triggered tick that
hits the guard now waits for the in-flight tick and then runs; scheduled
ticks keep the skip. Capped at one queued trigger: a second concurrent
trigger gets the skip + warn.
- rowMatchesFilter throws when multiple positive clauses for one field
are ANDed (the OR-model would evaluate them wrong → vacuous pass)
- makePagingPb throws on unmodeled sort keys instead of silently
returning insertion order
- documented the escaped-quote limitation of the clause value-extraction
regex and likeToRegExp's dangling-backslash fallback
(buffer-cap boundary in job-producer.test.ts deliberately untouched —
sibling slot owns that file)
- release: status is REQUIRED (the old '|| done' fallback silently
finished a job whose caller omitted/emptied status) — 400 like
jobId/workerId
- claim+renew: leaseSeconds floored at 1s (0.001 previously yielded a
1ms lease — instantly stealable, renew thrash)
- all handlers: reject non-string workerId (a JSON number coerces into
the text claimed_by column and the holder can never renew/release)
- claim: same-holder live-lease re-claim answers claimed:true with an
alreadyHeld:true marker (timeout-after-commit retry no longer abandons
a row the worker actually holds); client treats it as a plain win
(ClaimEndpointBody documents the no-op marker)
- hook-parity pins updated/added for every change