Commit Graph

12110 Commits

Author SHA1 Message Date
Jordan Ritter fc14f2e39c fix(harness): claim-delete long-expired stale-aged leases in the lease phase instead of re-queueing (G1d)
A row whose lease expired beyond its family's stale window is past the
recent-lease heuristic's cross-sweep protection: re-queueing it emitted a
'back in flight' worker-reclaimed-pending signal that the NEXT sweep falsified
by stale-deleting the row. When a lease-phase row is already stale-expirable
(parseable created past maxAgeMs AND lease expired longer than the window),
claim it under the sweeper id and delete it — no comm error, counted in
expiredPending. Recently-expired rows keep today's re-queue path; header's
cross-call-protection claim corrected.
2026-06-11 20:38:39 -07:00
Jordan Ritter ef5ec1bcf5 fix(harness): exclude clause-unsafe family rows by id and continue discovery (G1f)
A backslash (or empty) family used to BREAK family discovery, starving every
younger family behind the offending row for up to 3h. The offending ROW needs
no family-charset semantics to exclude — push an id != "<rowId>" clause and
continue, keeping the warn. Younger families are still discovered; the unsafe
row stays unclaimable.
2026-06-11 20:38:38 -07:00
Jordan Ritter d7505d2644 fix(harness): close the empty-string family partition gap and forbid empty payload sentinels (G1c)
probeKeyFamily("") takes the prefix-LIKE clause leg, so the empty family's
inclusion clause matched every leading-colon key and its exclusion clause hid
ALL leading-colon families from discovery. Treat family "" as clause-unsafe
in familyClauseSafe (skip+warn, like the backslash class), and require
non-empty probeKey/serviceSlug/driverKind/meta.runId in assertServiceJobPayload
so the documented forbidden empty sentinels fail loud at both the enqueue and
claim-decode boundaries.
2026-06-11 20:38:38 -07:00
Jordan Ritter 620a24b75c fix(harness): keep a renewed-without-job lease assumed-live instead of false-reclaiming a live job (G1b)
A renewed:true response missing its job view is a protocol violation on a
SUCCESSFUL renew — the worker still holds the live lease. Evicting + returning
null stopped the heartbeat and let the sweeper falsely reclaim the live job.
Return the cached lease assumed-live (like the indeterminate path); only
evict+null when no cached lease exists. Also corrects the stale 'We ONLY
return null when the CAS itself failed' comment.
2026-06-11 20:38:37 -07:00
Jordan Ritter 8a8b1bd004 fix(harness): treat deterministic 4xx renew rejections as definitive lease loss (G1a)
A JobClaimEndpointError 4xx from the renew endpoint fails identically every
beat, so the indeterminate assumed-live containment never converged — the
worker held a phantom lease forever while the server lease lapsed and another
worker double-ran the job. Mirror the sweep's 4xx carve-out: error-log
renew-rejected, evict payload+lease caches, return null. 5xx/transport/
2xx-unreadable keep the assumed-live path.
2026-06-11 20:38:37 -07:00
Jordan Ritter d16b878923 fix(showcase): scan non-d6 fleet-family sweep aggregate keys for comm errors (G3f)
The global lease sweep mirrors comm errors onto the status row keyed by the
reclaimed job's probe_key (resolveSweepAggregateKey → aggregateCommError) for
ALL four fleet families, but the dashboard's decodeCellCommError candidate
scan only covered the d6 family's d6:<slug> aggregate. The smoke/demos/deep
families land on d4:<slug>, e2e-demos:<slug>, and d5-single-pill-e2e:<slug> —
rows the dashboard reads nowhere else — so reclaim/crash overlays on those
families were invisible. Add the three candidate keys (e2e window).
2026-06-11 20:38:36 -07:00
Jordan Ritter 2fdd6132b6 fix(showcase): dashboard status-model hardening batch (G3e)
- Object.freeze the UNSUPPORTED CellModel + NOT_WIRED_LEVEL singletons —
  they are returned by reference to every caller, so one consumer mutating
  them would corrupt every unsupported/not-wired cell.
- Remove the unfirable !d5.exists chip arm (d5.exists and d6.exists derive
  from the SAME CATALOG_TO_D5_KEY entry, so !d5.exists && d6.exists is
  impossible); fix the decision-table and d6Effective docs that described
  the unfirable behavior, keeping a note for a future map split.
- Constrain upsertByKey's generic to StatusRow so rowsAreNoop can never
  vacuously match a non-StatusRow type whose fields are all undefined.
- keyFor: validate the dimension segment for ':'/'/' like slug/featureId.
- decodeCellCommError: scope the comm-error staleness window per row family
  (e2e 6h for d5/d6/e2e, D4 1h for chat/tools, liveness 45m for health)
  instead of applying the 6h E2E window to liveness-cadence rows.
2026-06-11 20:38:36 -07:00
Jordan Ritter 13138208e8 docs(showcase): pin the deliberate pending-gate asymmetry harness↔dashboard (G3d)
The three-way contradiction: harness fleetSurfaceState routes green-only →
pending, cell-model routes green OR gray (non-regression) → pending while its
own comment claimed 'only green', and both files claimed 'exact mirror'.
gray→pending IS intended dashboard-side (gray is the dashboard-only no-data
colour ProbeState cannot represent). Fix the comments on both sides to state
the asymmetry precisely, extend the drift test to pin the derivation
difference itself (harness green-equality + no gray; dashboard
failure-passthrough + isRegression), add behavioral gray→pending and
red-passthrough tests, and fix chipColorToSurface's stale union doc
(now ChipColor | unreachable | pending).
2026-06-11 20:38:35 -07:00
Jordan Ritter 4c2cb309c1 fix(showcase): resolveD3 stale-green downgrade returns the effective row (G3c)
The stale-green branch returned the RAW row (status "amber" with row.state
still "green"), violating the .row.state ↔ .status invariant resolveD4/D5/D6
maintain. Return { ...row, state: "degraded" } like the other resolvers.
2026-06-11 20:38:35 -07:00
Jordan Ritter 3b5e9987a7 fix(showcase): map out-of-vocab D5/D6 fold states to a failing status (G3b)
stateToTestStatus mapped unknown runtime states (e.g. "error") to null,
swallowing the A2 rank-fold winner one step after the fold surfaced it — the
D5/D6 chip went benign gray no-data while live-status rendered the loud
"error" tone for the same row. Add foldStateToTestStatus (out-of-vocab →
"red") for the D5/D6 resolvers; D3/D4 keep the base mapping since the
chip's D1-D4 gate check already rescues their null (pinned by tests).
2026-06-11 20:38:34 -07:00
Jordan Ritter 8868047018 fix(showcase): rank-based anyMissing collapse in cell-model resolveD5/resolveD6 (G3a)
The strict missing-sub-row collapse used the literal `worstState !== "red"`,
which silently swallowed out-of-vocabulary runtime states (e.g. "error") that
the A2 rank machinery deliberately ranks ABOVE red — collapsing exactly the
state the rank fold exists to surface into benign gray no-data. Use
`rankOfState(worstState) < STATE_RANK.red`, mirroring the already-fixed
resolveD5Row/resolveD6Row in live-status.ts.
2026-06-11 20:38:34 -07:00
Jordan Ritter d2f56c7359 test(showcase): add truncatedByStop to the control-plane test's TickResult fake
Follow-up to the TickResult.truncatedByStop accounting field — the
control-plane wiring test's producer fake builds a literal TickResult
and needed the new required field for tsc to stay clean.
2026-06-11 20:38:33 -07:00
Jordan Ritter eed508da23 test(showcase): loop-crash worker test guarantees teardown via try/finally stop()
A failed assertion in the loop-crash test leaked the worker's bound
/health server (and the default-boot pool) across tests because stop()
was only reached via the in-body rejects.toThrow. Mirror the sibling
default-boot test's try/finally pattern with a rejection-swallowing
safety-net stop.
2026-06-11 20:38:33 -07:00
Jordan Ritter dca5731c42 test(showcase): harden job-producer + queue-client fixtures (CR G2e batch)
- derive the buffer-cap overflow boundary from
  MAX_BUFFERED_SWEEP_COMM_ERRORS instead of hardcoding 'old100'
- race the re-entrancy-guard's overlapping tick against a short timer so
  a broken guard fails diagnostically instead of deadlocking
- pin the warm-up abort test's AbortController+setTimeout coupling
  (AbortSignal.timeout would evade fake timers)
- document makeFakeQueue.enqueued as recording enqueue ATTEMPTS
- align queue-client samplePayload driverKind with the probeKey family
  so future payload-coherence cross-validation doesn't mass-fail
2026-06-11 20:38:32 -07:00
Jordan Ritter 31efa201b9 fix(showcase): account stop-truncated specs in TickResult.truncatedByStop so the tick outcome partitions exactly
Stop-truncated specs vanished from the tick outcome: tick-complete's
'services' could not be reconciled against enqueued + enqueueFailures +
skippedForBacklog. Add truncatedByStop to TickResult and the
tick-complete log (invariant: services == enqueued + enqueueFailures +
skippedForBacklog + truncatedByStop) and fix the 'enqueued' doc to
mention stop truncation.
2026-06-11 20:38:32 -07:00
Jordan Ritter 0f92764aaa fix(showcase): no-sink comm-error drops after the first warn are never fully silent — debug log with jobIds + running dropped-total
The one-shot no-sink warn kept a sink-less deployment from burying its
logs, but made every subsequent drop completely traceless. Later drops
now log at debug with the dropped jobIds and a running dropped-total
counter; the counter also rides the first full warn.
2026-06-11 20:38:31 -07:00
Jordan Ritter 394736b8ed fix(showcase): a second concurrent stop() must quiesce too — await the in-flight tick and any queued trigger
The first stop() flips 'running' synchronously, so a second concurrent
stop() hit the !running early-return and resolved immediately while the
tick the first stop() was quiescing on kept enqueueing. Both stop()
paths now loop-await the in-flight tick AND the queued trigger before
resolving.
2026-06-11 20:38:31 -07:00
Jordan Ritter 1bca4decf8 fix(showcase): queue an operator-triggered tick behind an overlapping in-flight tick instead of dropping it
The producer's re-entrancy guard skipped TRIGGERED ticks too, silently
losing an explicit operator 'run it NOW' whenever a slow scheduled tick
was in flight — contradicting the operator-intent-wins rationale that
already lets triggers bypass the backlog gate. A triggered tick that
hits the guard now waits for the in-flight tick and then runs; scheduled
ticks keep the skip. Capped at one queued trigger: a second concurrent
trigger gets the skip + warn.
2026-06-11 20:38:30 -07:00
Jordan Ritter 3a3634ee4b test(showcase): queue-client fakes throw on shapes they cannot honor
- rowMatchesFilter throws when multiple positive clauses for one field
  are ANDed (the OR-model would evaluate them wrong → vacuous pass)
- makePagingPb throws on unmodeled sort keys instead of silently
  returning insertion order
- documented the escaped-quote limitation of the clause value-extraction
  regex and likeToRegExp's dangling-backslash fallback
(buffer-cap boundary in job-producer.test.ts deliberately untouched —
sibling slot owns that file)
2026-06-11 20:38:30 -07:00
Jordan Ritter 9ce9b2d947 fix(showcase): fleet-claim hook hardening — explicit status, lease floor, workerId type, claim idempotency
- release: status is REQUIRED (the old '|| done' fallback silently
  finished a job whose caller omitted/emptied status) — 400 like
  jobId/workerId
- claim+renew: leaseSeconds floored at 1s (0.001 previously yielded a
  1ms lease — instantly stealable, renew thrash)
- all handlers: reject non-string workerId (a JSON number coerces into
  the text claimed_by column and the holder can never renew/release)
- claim: same-holder live-lease re-claim answers claimed:true with an
  alreadyHeld:true marker (timeout-after-commit retry no longer abandons
  a row the worker actually holds); client treats it as a plain win
  (ClaimEndpointBody documents the no-op marker)
- hook-parity pins updated/added for every change
2026-06-11 20:38:29 -07:00
Jordan Ritter 5dcbe31202 fix(showcase): queue-client/job-claim hardening batch — silent paths get logs, poisoned values fail loud
- decode-failure REFUSED release now warns with the hook's refusal reason
- probeKeySlug empty-slug ('d6:') falls back to the whole probe_key; a
  fully-empty key falls back to unknown-<jobId> (never the forbidden
  empty serviceSlug sentinel)
- countPendingForFamily throws on non-numeric/negative totalItems instead
  of failing the producer backlog gate open with -1
- stale-pending drain contains a thrown claim CAS per row (a 4xx no
  longer aborts the whole drain) and documents the 5xx->won:false
  bounded page-stall on passAttempted
- writeResult paces retries (RESULT_WRITE_RETRY_DELAY_MS, injectable
  sleep) instead of burning attempts back-to-back
- protocol-violation warns for won&&!job / renewed&&!job before falling
  through
- ensureAuth memoizes the in-flight auth promise (no concurrent stampede)
  and wraps the auth-body JSON.parse with status context
- releaseJob interface doc corrected: the hook admits claimed AND running
  rows (the decode-failure cleanup depends on claimed->failed); the doc's
  'running state only' was the wrong side
- documented the LIKE-vs-= ASCII case-sensitivity divergence for
  mixed-case families (unreachable: keys are lowercase slugs)
2026-06-11 20:38:29 -07:00
Jordan Ritter 3c14e0ae15 fix(showcase): sweep discriminates deterministic 4xx release throws from indeterminate failures
The conservative thrown-release path treated EVERY throw as
may-have-committed, so a wedge row (e.g. empty claimed_by -> hook 400 on
the missing workerId) produced a permanent per-sweep false
worker-reclaimed-pending overlay. job-claim now threads the HTTP status
onto thrown endpoint errors (JobClaimEndpointError); the sweep maps a
4xx (hook rejected, nothing committed) to an error-level log with no
grace, no comm error, no reclaimed++, while 5xx/transport/2xx-unreadable
keep the conservative at-least-once handling.
2026-06-11 20:38:28 -07:00
Jordan Ritter 33bfa92039 fix(showcase): report() retry reads the row before rewriting — never un-latch an aggregated result
On the refused_terminal_same_holder retry path, writeResult's
result_processed:false seed could overwrite a row the consumer had
already aggregated (and latched), un-latching it for a second
aggregation — a double-count. The retry now reads the row first: a
present+processed result is skipped (info 'result already aggregated'),
a present+unprocessed result is skipped too (idempotent), only a
missing result falls through to writeResult. A failed pre-write read
throws instead of blind-writing.
2026-06-11 20:38:28 -07:00
Jordan Ritter 208817e7b0 fix(showcase): contain indeterminate renews in queue-client.renewLease — keep the lease assumed-live
With job-claim now throwing on unreadable-2xx (and renew already throwing
on 5xx), a thrown renew escaped queue-client.renewLease into the worker
heartbeat, which treats any renewLease throw as fatal — the heartbeat
died and the sweeper reclaimed a live job (false worker-crashed-mid-job).
renewLease now catches throws from claim.renewLease, warns, and returns
the last-known lease unchanged (no payloadCache eviction, never null) so
the heartbeat retries next beat; only a definitive renewed:false stops
it. Documented at-most-one-lease-duration duplicate-execution risk; the
release CAS still arbitrates terminal writes.
2026-06-11 20:38:27 -07:00
Jordan Ritter 0b84849bb0 fix(showcase): job-claim 2xx-unreadable body throws indeterminate instead of fabricating a CAS loss
A 2xx response whose body fails to read, is empty, or fails to parse was
mapped to {} — i.e. claimed/renewed/released: false — fabricating a CAS
LOSS for a transition the server had already COMMITTED (the renew flavor
recreated the false worker-crashed-mid-job class; the release flavor made
report() falsely declare the result discarded). Now any unreadable 2xx
body THROWS with path context; callers contain the throw.
2026-06-11 20:38:27 -07:00
Jordan Ritter ee1269b96e fix(showcase): whole-key (colon-bearing) families match by equality only in queue-client clauses
probeKeyFamily treats a leading-colon probe_key as its own whole-key
family, so the family value itself contains colons — expanding it with
the <family>:% LIKE leg folded the unrelated family ":foo:bar" under
":foo" in the inclusion/count legs and hid it from discovery via the
exclusion leg. Special-case colon-bearing families to the equality leg
only, implementing the invariant pinned on probeKeyFamily (G3e).
2026-06-11 20:38:26 -07:00
Jordan Ritter 3f898ee0d5 fix(showcase): backlog gate counts non-terminal rows — bound concurrent runs per family
countPendingForFamily filtered on status = "pending" only, so a family
whose whole batch was claimed/running stopped gating and a scheduled
tick could enqueue a fresh batch on top of the in-flight one, doubling
the family's concurrency. Broaden the filter to non-terminal
(pending || claimed || running) and align the gate docs in the contract
and job-producer.
2026-06-11 20:38:26 -07:00
Jordan Ritter 022674ac87 test(harness): harden fleet worker port-closed assertion against port reclaim by a parallel process 2026-06-11 20:38:26 -07:00
Jordan Ritter 3fd3592609 refactor(dashboard): hoist the byte-identity drift parser to module scope (G3c follow-up)
oxlint flagged the inline extract closure (recreated per call); the parser
now sits alongside the file's other parse helpers.
2026-06-11 20:38:25 -07:00
Jordan Ritter ed5eedebfb docs(fleet): pin the equality-only matching invariant for whole-key probe families (G3e)
probeKeyFamily(':foo:bar') returns the WHOLE key as its own family (the
round-2 leading-colon fix), but the queue-client clause builders
(familyInclusionClause/familyExclusionClause) expand any family with a
prefix-LIKE <family>:% leg, which would include/exclude ':foo:bar' under
family ':foo' — disagreeing with the fairness partition. The clause
builders live in queue-client.ts (out of scope here), so this documents the
invariant on the contracts side — whole-key (colon-bearing) families must
match by equality only — and pins the partition behavior in the test suite.
The queue-client special-case is reported to the integrator separately.
2026-06-11 20:38:25 -07:00
Jordan Ritter bdef17c781 docs(fleet): worker-crashed-mid-job enumerates both reporting sources (G3d)
The taxonomy doc defined the kind as exclusively worker-self-monitor-
reported, but the control-plane result consumer ALSO synthesizes it for
terminal-but-resultless rows past the grace window (the queue-client's
bounded result-write retry exhausted — see RESULT_WRITE_MAX_ATTEMPTS).
Both the contracts.ts kind doc and the dashboard mirror now enumerate
the two sources.
2026-06-11 20:38:24 -07:00
Jordan Ritter 08e8827a16 fix(fleet): commErrorFromStatusSignal rejects arrays in both contract copies (G3c)
typeof raw !== "object" admits arrays: an array carrying comm-error fields
as expando properties decoded as a well-formed PoolCommError. Both copies
(harness contracts.ts and the dashboard live-status.ts mirror) now reject
arrays explicitly and stay byte-identical — the drift suite gains a
byte-identity pin on the full function source so a one-sided fix can never
silently re-open the gap.
2026-06-11 20:38:24 -07:00
Jordan Ritter f535f7756f fix(dashboard): resolveCell rollup folds contributor states through worstStateRank (G3b)
The rollup fold used literal includes("red")/includes("degraded") checks,
so an out-of-vocabulary contributor state (e.g. "error", which the harness
can persist at runtime) matched neither bucket and rolled the cell up GRAY
(benign no-data) while its own badge rendered the loud error tone —
violating the documented precedence red > degraded > green > error >
unknown and the A2 never-swallow rule. Contributor states now route through
worstStateRank, so an unknown state rolls up at least red-severity.
2026-06-11 20:38:23 -07:00
Jordan Ritter 830355d3b8 fix(fleet): worker-reclaimed-pending overlay passes through every non-green failure state (G3a)
The pass-through gate in fleetSurfaceState was a literal === "red" check, so
a row whose last-known state was degraded/error/out-of-vocab got masked by
the neutral pending overlay — violating the function's own never-mask-a-
genuine-failure invariant (A2: error ranks ABOVE red). Only green now becomes
pending; the dashboard cell-model derivation mirrors it over its ChipColor
vocabulary (red AND amber pass through). Fixes the self-contradicting JSDoc
closing sentence, the stale live-status.ts mirror doc (union was missing
pending), and extends the drift test to pin the pass-through semantics on
both sides.
2026-06-11 20:38:23 -07:00
Jordan Ritter ffc25bb03a fix(showcase/harness): producer tick/warm-up hygiene batch (CR round 3)
- count a warm GET as fired only after successful dispatch (sync-throwing
  fetchImpl no longer inflates the warmed log)
- route an enumerator resolving to a non-array through the same
  enumerate-failure handling as a throw
- mint the runId AFTER the running check (stopped ticks no longer burn
  the factory counter or log phantom runIds; empty-runId sentinel)
- log tick-start BEFORE the enumerate await so a hung discovery still
  leaves a tick trace (services count moved to tick-complete)
- warn once, with jobIds, when swept comm errors exist but no sink is
  configured (previously dropped with only a count)
- terminal .catch on the fire-and-forget warm chain (a throwing logger
  no longer raises an unhandled rejection)
- document the SweepCommErrorSink at-least-once contract (aggregator
  must be idempotent per jobId+observedAt)
- tests: fix stale "Four producers" comment; replace the fragile
  two-shot Math.random spy with a never-exhausting deterministic
  sequence plus a call-count guard
2026-06-11 20:38:22 -07:00
Jordan Ritter 7ec2890b40 fix(showcase/harness): guard producer tick re-entrancy (backlog-gate TOCTOU double-enqueue)
A slow tick overlapping the next cron tick let both read the backlog
gate's pending count before either enqueued, so both produced a batch
for the same family — the motivating concurrent same-service-run
incident class. An overlapping tick is now skipped (warned) via the
tracked in-flight tick promise; it returns the empty-runId no-op result
without minting a run.

NOTE for the integrator: the gate also under-counts — queue-client's
countPendingForFamily filters status="pending" only, so a fully-claimed
but-still-running batch doesn't gate. Broadening the filter to
non-terminal rows (pending || claimed || running) is a one-line edit in
queue-client.ts (plus the contracts.ts doc), both sibling-owned —
deferred to the integration pass.
2026-06-11 20:38:22 -07:00
Jordan Ritter 4b482c37cd fix(showcase/harness): make producer stop() quiesce — await in-flight tick, truncate mid-batch, final comm-error drain
stop() used to flip flags and resolve while an in-flight tick kept
enumerating/sweeping/enqueueing, and buffered sweep comm errors were
silently dropped at shutdown. Now stop() awaits the tracked in-flight
tick promise, the enqueue loop re-checks `running` and truncates loudly
when stopped mid-batch, and stop() makes one final sink delivery of the
buffered comm errors — logging dropped count + jobIds at error level if
that last attempt fails too.
2026-06-11 20:38:21 -07:00
Jordan Ritter c2f8383b8a fix(showcase/harness): drain buffered sweep comm errors even when the sweep itself throws
The undelivered comm-error buffer was drained only on maybeSweep's try
SUCCESS path, so a persistently-throwing sweepExpired — the exact
failure mode the buffer exists to ride out — never handed the buffered
batch to a now-healthy sink. Extract the delivery into
deliverSweepCommErrors and run it on every sweep attempt, including the
catch arm.
2026-06-11 20:38:21 -07:00
Jordan Ritter bd7c02f966 test(showcase): close the vacuous-pass holes in the queue-client fakes (CR G1g)
rowMatchesFilter now THROWS on operators it can't honor (<,>,<=,>=,?=,
?~) and on status clauses with !=/~/!~ instead of silently matching all
rows, with the quoted-literal false-positive limitation documented;
both fakes return totalItems:-1 (and -1 totalPages) under skipTotal —
faithful to PB, closing the fail-open count class beyond the single
countPendingForFamily pin; makePagingPb honors the '-' sort-direction
prefix and computes totalPages honestly; makeOrderedPb's
ignore-all-filters behavior is documented as deliberate coupling to the
duplicate-family defensive break; the result-write retry bound is
pinned with the exported RESULT_WRITE_MAX_ATTEMPTS constant.
2026-06-11 20:38:20 -07:00
Jordan Ritter dc3acb8427 fix(showcase): queue-client robustness batch (CR G1f)
(i) per-candidate try/catch in claimNext's CAS race — a thrown transport
claim no longer aborts the whole rotation (warn + next candidate);
(ii) discoverPendingFamilies' duplicate-family defensive break now warns
before breaking; (iii) renew re-read triage split — a decodePayload
failure logs as a protocol violation, not a read blip; (iv) the
decode-failure synthetic result uses an injected clock (new config.now)
instead of new Date(); (v) backslash charset guard for family clause
building (the equality/LIKE escape contracts contradict for backslash —
skip such families with a warn) and the over-claiming VERIFIED comment
scoped to %/_ only; (vi) single-sweeper assumption documented as
load-bearing at the grace-set declaration; (vii) hook comment misname
fixed (worker-crashed-mid-job -> worker-reclaimed-pending); (viii)
claimJob maps a 5xx claim response to a lost CAS (won:false, warn) —
a WAL serialization error escaping runInTransaction surfaces as 500 —
while 4xx and renew/release 5xx still throw loud.
2026-06-11 20:38:19 -07:00
Jordan Ritter 5730b86f94 fix(showcase): make report() retryable via release refusal reasons (CR G1e)
After a release-CAS success + result-write exhaustion, report() throws;
a natural retry got REFUSED (row already terminal) and emitted an error
claiming the result is discarded and the job re-runs — both false (the
result is still writable by this holder; terminal rows never re-run).
The hook's release response now carries a refusal reason
(refused_terminal_same_holder / refused_not_holder /
refused_lease_live), threaded through job-claim's ReleaseResult.
report() treats refused-terminal-under-my-workerId as the second leg of
a timeout-after-commit retry and proceeds to writeResult; the
not-holder error is reworded to 're-runs only if reclaimed to pending'.
A reason-less refusal still fails closed.
2026-06-11 20:38:19 -07:00
Jordan Ritter 937ec924de fix(showcase): stop the next sweep from claim-deleting a re-queued long-runner (CR G1d)
Stale-pending age is anchored on PB created and never re-anchored on
re-queue: a job running longer than its family expiry window got
lease-reclaimed with a 'back in flight' comm error, then claim-deleted
by the NEXT sweep before any plausible re-run — the dashboard
permanently showed 're-queued' for silently-discarded work. Schema-free
fix: the release hook now RETAINS the expired lease_expires_at on a
pending re-queue (claim admits pending rows regardless of lease), and
the stale phase skips rows whose retained lease is recent (parseable
and within now - familyExpiryWindow) — recently in flight means the
created-based age is stale evidence. Comments cover the heuristic, the
requeued_at-column alternative, and the sweeper-garbage lingering
tradeoff; hook parity test pins the retention.
2026-06-11 20:38:19 -07:00
Jordan Ritter abe27b38d5 fix(showcase): populate serviceSlug/runId on the synthetic protocol-violation result (CR G1c)
The decode-failure synthesis persisted serviceSlug: "" and runId: "" —
the exact empty sentinels this file's own emptyPayloadForLease warning
forbids feeding aggregation (an empty runId groups into nothing, an
empty serviceSlug corrupts the per-service rollup). Recover the slug
from the probe_key's slug segment (new probeKeySlug, the complement of
contracts.ts probeKeyFamily) and mint a non-colliding pviol_<jobId>
runId; reconcile the contradictory comments on both sides. Verified
downstream: result-consumer/aggregator key on aggregateKey + dedupe on
jobId, with serviceSlug/runId riding into logs and batch grouping.
2026-06-11 20:38:18 -07:00
Jordan Ritter 39b154ee6c fix(showcase): treat thrown sweep release as committed — grace + synthesize comm error (CR G1b)
A release that THROWS in the lease sweep may have COMMITTED server-side
(timeout-after-commit): the row is then pending but absent from the
grace set — the same call's stale phase could claim-and-delete it — and
its worker-reclaimed-pending comm error was never synthesized, losing
the gray 're-queued' surface forever (no later sweep re-emits pending
rows). Conservatively add the row to the grace set and synthesize the
comm error on a throw (at-least-once: a duplicate gray overlay is
harmless, a missing one is not); sweeper-held rows stay silent,
mirroring the committed path.
2026-06-11 20:38:18 -07:00
Jordan Ritter 547e6f2841 fix(showcase): enforce superuser auth + leaseSeconds clamp on fleet claim endpoints (CR G1a)
The three /api/fleet/* routerAdd handlers carried no auth middleware —
a middleware-less PB 0.22 routerAdd handler is PUBLIC, so any
unauthenticated caller could claim/renew/release arbitrary jobs despite
the header claiming superuser auth was required. Append
$apis.requireAdminAuth() (verified against PB 0.22.21 JSVM types) to
all three routes; the client already authenticates as superuser with a
401-reauth retry, so enforcement is compat-safe. Also clamp
leaseSeconds in claim+renew (numeric only, 3600s ceiling, 30s default
on garbage) and pin both contracts in the hook-source parity tests.
2026-06-11 20:38:17 -07:00
Jordan Ritter a8a44a9cd9 docs(harness): align meta.priority JSDoc with claimNext reality — reserved, not consulted 2026-06-11 20:38:17 -07:00
Jordan Ritter 3686a82987 fix(harness): probeKeyFamily treats a leading-colon key as its own family
A probe key beginning with ":" yielded the empty-string family, which
then flowed into countPendingForFamily and the claimNext fairness
partition as a phantom real bucket. Treat such a key as having no family
prefix (the whole key is its own family), matching the no-colon case.
2026-06-11 20:37:35 -07:00
Jordan Ritter e5eceb2e76 fix(shell-dashboard): keyFor throws on empty-string featureId
The truthiness guard (`if (featureId && ...)`) let an empty-string
featureId bypass the delimiter validation and fall through to the
integration-aggregate key shape, silently fabricating `<dim>:<slug>` for
what the caller meant as a per-feature lookup. Throw loudly instead,
preserving the function's defensive posture.
2026-06-11 20:37:34 -07:00
Jordan Ritter a2f666376d docs(shell-dashboard): correct buildStarterBadge JSDoc dimension claim
The JSDoc said data-bearing states are delegated to buildBadge "under
the `health` dimension label", but the code passes "starter" — and must:
the starter ✓/✗/~ glyph vocabulary requires formatLabel's non-health
branch (the health branch renders up/down/stale word labels instead).
2026-06-11 20:37:34 -07:00
Jordan Ritter d368fd986c fix(shell-dashboard): compare id, fail_count, first_failure_at in rowsAreNoop
rowsAreNoop ignored fail_count, first_failure_at, and id, so an SSE
delta moving only one of them was swallowed as a no-op: first_failure_at
is load-bearing in formatTooltip ("red since ..."), fail_count feeds the
drilldown/alerting surfaces, and a deleted-and-recreated PB row (same
key, fresh id) kept the stale id in the map. Add the three fields to the
comparison and correct the comparator's field-list doc claims.
2026-06-11 20:37:34 -07:00