Commit Graph

4085 Commits

Author SHA1 Message Date
Jordan Ritter c600222cf3 fix(harness): poisoned-count backlog gate discriminates via exported PoisonedBacklogCountError, not message text (CF7 #1)
The fail-closed gate matched a message substring duplicated in 3 places;
any rewording silently flipped it fail-closed -> fail-open. The refusal is
now a dedicated exported class, the producer gates via instanceof, the test
literal copies use the real class, and a drift test drives the REAL
queue-client refusal through the producer gate.
2026-06-11 20:42:35 -07:00
Jordan Ritter 8119635ade docs(harness): ReportJobInput INTEGRATOR NOTE reflects report()'s equality enforcement (integration) 2026-06-11 20:42:35 -07:00
Jordan Ritter 5e471aab0e test(dashboard): drop the [0]! non-null assertion in the G3f aggregate-key helper — destructure with fallback survives noUncheckedIndexedAccess (CF6-G5 #9) 2026-06-11 20:42:34 -07:00
Jordan Ritter aa294dd6c3 fix(dashboard): drift-test surface-derivation parser guards against silent truncation at an internal ';' blank-line boundary (CF6-G5 #7) 2026-06-11 20:42:34 -07:00
Jordan Ritter c9663e258c fix(dashboard): drift-test kind parser strips comments before matching; fix the misleading line-position comment (CF6-G5 #6) 2026-06-11 20:42:33 -07:00
Jordan Ritter e742155e7f fix(dashboard): health label honors the staleness split — fresh producer-emitted degraded reads 'degraded', not 'stale' (CF6-G5 #2) 2026-06-11 20:42:33 -07:00
Jordan Ritter bbaf58a392 fix(dashboard): D4 missing-chat collapse renders the gray no-data chip, not a red gate failure (CF6-G5 #1) 2026-06-11 20:42:32 -07:00
Jordan Ritter 57054b9830 fix(harness): REQ-B sweep sink throws while the control-plane is unbound
A resolved sink call clears the producer's undelivered comm-error buffer
(deliverSweepCommErrors' at-least-once contract), so the late-bound sink
silently 'succeeding' before controlPlaneRef was assigned dropped the
batch permanently. Extract buildSweepCommErrorSink (exported for tests)
and reject on the unbound window so the producer re-buffers and
redelivers after assembly; pin with red-green unit tests.
2026-06-11 20:42:31 -07:00
Jordan Ritter 4bfeb121c4 fix(harness): drop enumerator specs with an empty probeKey at the tick boundary — no phantom '' family in the backlog gate, no unjoinable claim row; counted as enqueueFailures 2026-06-11 20:38:46 -07:00
Jordan Ritter 13a33bf97a fix(harness): warm-up branches on res.ok — a settled 404/5xx logs warm-failed (with status) instead of masquerading as warm-ok 2026-06-11 20:38:46 -07:00
Jordan Ritter ce22ff454d fix(harness): producer tick/stop survive a throwing logger transport — safeLogger wrapper keeps the 'a tick promise never rejects' invariant true and stopPromise unpoisoned 2026-06-11 20:38:45 -07:00
Jordan Ritter b3b3dbdc52 fix(harness): minor queue-client honesty fixes (empty enqueuedAt, 408/429 transient class, truncation warn, typo)
- assertServiceJobPayload rejects an empty meta.enqueuedAt — the
  emptyPayloadForLease never-aggregate sentinel, same forbidden class as
  the other empties (the minted fallback never crosses this boundary)
- 408/429 are excluded from the deterministic-4xx class (shared
  deterministicEndpointRejection helper): transient by meaning, they keep
  the assumed-live renew and at-least-once sweep containments
- the family-discovery truncation warn threads nextFamilyClauseSafe so a
  clause-unsafe garbage head can't masquerade as a hidden claimable family
- fix 'CONTROl' comment typo in queue-client.test.ts
2026-06-11 20:38:45 -07:00
Jordan Ritter 7c467f6231 fix(harness): report() enforces the ReportJobInput equality invariant and never skips cache eviction
Enforce the documented jobId/workerId equality between the top-level
input and the echoed result (a mismatched caller would release one row
while filing the result under another), and move the terminalJobStatus
computation inside the try so a malformed result can no longer throw
past the finally and leak the claim-time cache entry.

NOTE: contracts.ts's INTEGRATOR NOTE on ReportJobInput still says report()
does not validate this — contracts.ts is sibling-owned this round and is
NOT touched here; the note needs a follow-up edit.
2026-06-11 20:38:44 -07:00
Jordan Ritter 2982db35ef fix(harness): sweep deletes stale-aged expired rows with UNPARSEABLE leases in the lease phase
The long-expired carve-out required a finite parsed lease, so an expired
claimed/running row with an unparseable lease_expires_at was re-queued
with a 'back in flight' worker-reclaimed-pending signal — which the next
sweep falsified by claim-deleting the row (the recent-lease protection
also requires a finite lease). An unparseable lease carries no
recent-flight evidence either phase can honor: treat it as long-expired
and delete directly so the emitted signal matches the outcome. Pinned
with a two-sweep test.
2026-06-11 20:38:44 -07:00
Jordan Ritter ac1b5a2415 fix(harness): assertServiceJobPayload rejects a top-level array payload
An array is typeof 'object' and can carry expando fields that satisfy
every per-field check — the same hole the nested meta check's
Array.isArray guard already closes one level down. Reject it at the
shared enqueue/decode boundary.
2026-06-11 20:38:44 -07:00
Jordan Ritter 4108bc39f4 fix(harness): auth 4xx throws JobClaimEndpointError so the deterministic carve-outs fire for rotated creds
authenticate() threw plain Errors for every non-2xx, so the renew and
sweep deterministic-4xx carve-outs (instanceof JobClaimEndpointError +
4xx) never fired for auth failures — rotated creds rode the indeterminate
path forever: a phantom assumed-live lease per renew beat and a false
worker-reclaimed-pending comm error per sweep. 4xx auth responses now
carry the discriminable class (auth path + status); 5xx/network stay
plain (indeterminate). End-to-end pins go through the real job-claim
client for both the renew and sweep halves.
2026-06-11 20:38:44 -07:00
Jordan Ritter cedd7d481a fix(harness): won-without-job containment releases terminal 'failed' with a synthetic protocol-violation result
The pending-target release could never succeed: the hook refuses every
pending-target release on a live lease (refused_lease_live, no holder
exemption) and refused_not_holder otherwise, so the containment was inert
— the row wedged a lease window and got a false worker-reclaimed-pending
overlay. Mirror the decode-failure containment (terminal target passes
the live-lease gate) and share its synthetic-result builder. Tests now
use hook-faithful release fakes instead of unconditional released:true.
2026-06-11 20:38:44 -07:00
Jordan Ritter 46775478a6 docs(harness): SweepResult.commErrors doc covers the indeterminate (thrown-release) synthesis too — length = reclaimed + reclaimedIndeterminate (CR G2 #3) 2026-06-11 20:38:44 -07:00
Jordan Ritter d57a1320f8 fix(harness): commErrorFromStatusSignal rejects an array SIGNAL blob, not only an array nested value; dashboard mirror updated in lockstep (drift test pins byte-identity) (CR G2 #2) 2026-06-11 20:38:43 -07:00
Jordan Ritter cb5e4e4ccf fix(harness): land reclaimedIndeterminate on the shared SweepResult contract; producer reads it directly (integration) 2026-06-11 20:38:43 -07:00
Jordan Ritter 60c692d824 fix(dashboard): D5/D6 resolvers return the effective fold winner, fresh-degraded tooltip copy split from stale, document CellState.d6 as ungated (G2f ii+iii+v) 2026-06-11 20:38:43 -07:00
Jordan Ritter cb35d7fb4b fix(dashboard): D4 collapses a green fold when the unconditional chat row is missing; tie-break doc drops unreachable 'absent' case (G2f i+iv) 2026-06-11 20:38:43 -07:00
Jordan Ritter e7eca9174f fix(harness): fleet contracts polish — clamp WorkerCapacity.available at 0, ReportJobInput equality-invariant doc, enumerate both fleetSurfaceState/dashboard divergences (G2e) 2026-06-11 20:38:43 -07:00
Jordan Ritter c4854d40cc fix(harness): producer tick-outcome polish — queued-trigger-after-stop debug event, reclaimed at-least-once doc + reclaimedIndeterminate surfacing (G2d) 2026-06-11 20:38:42 -07:00
Jordan Ritter 86e717a370 fix(harness): share one stop completion so every producer stop() resolution implies the final comm-error drain finished (G2c) 2026-06-11 20:38:42 -07:00
Jordan Ritter 31d119683d fix(harness): make producer stop() before start() a no-op instead of a permanent brick (G2b) 2026-06-11 20:38:41 -07:00
Jordan Ritter bbd42449d1 fix(harness): fail the producer backlog gate CLOSED on the queue-client's poisoned-count refusal (G2a) 2026-06-11 20:38:41 -07:00
Jordan Ritter f8033fe6bf test(harness): G1h test-quality batch for the fleet queue/claim suites
- report ordering assertion made non-vacuous: capture rows[0].result INSIDE
  the releaseJob fake and assert undefined at release time.
- empty-slug-segment fixture carries probe_key 'd6:' on the claim fake's
  returned view too (no pinning of the internal key source).
- job-claim.test.ts: file-local silent logger replacing the shared logger
  import (spy-leak class under fork-reuse).
- auth-order tests: auth + endpoint URLs recorded in one call log, order
  asserted by index; fallback test pins _superusers-first order and the
  Authorization header carried to the endpoint.
- makeFakePb.list honesty: models created/lease_expires_at sorts (throws on
  unmodeled keys) and honors perPage/page truncation; self-tested in the
  fake-honesty suite.
- documented the interleaved-|| connective-guard limitation in
  rowMatchesFilter.
- renamed the renew 'convenience re-read returns null' test to the cache-hit
  pin it actually covers, with expect(getOneSpy).not.toHaveBeenCalled().
2026-06-11 20:38:40 -07:00
Jordan Ritter 34e67f7b4a fix(harness): G1g batch — report retry guard, at-least-once sweep split, 401 race, single-attempt decode write, hook jobId guard, doc corrections
- report() retry: a null getOne resolution is a FAILED read (throw, no blind
  write) and a "" result is PB's unset-JSON shape (absent → write proceeds).
- sweepExpired: thrown-release conservative maybes now counted on a separate
  reclaimedIndeterminate (SweepResultWithIndeterminate); reclaimed counts only
  CAS-confirmed re-queues. Producer one-liner documented for when the
  sibling-owned TickResult gains the field.
- job-claim 401 retry: snapshot the token the failed request used; only null
  authToken if unchanged (no clobbering a concurrently refreshed token).
- decode-failure synthetic result write: single attempt, no 250ms retry pacing
  inside the claim race (consumer crash-synthesis is the documented backstop).
- fleet-claim.pb.js: typeof jobId !== "string" → 400 in all three handlers.
- docs: recent-lease bound is expiryPeriods × period (not one window); claim
  5xx→won:false bounded false-overlay source; report retryability deploy-skew
  note; drainStalePending page-advance indeterminacy note.
2026-06-11 20:38:40 -07:00
Jordan Ritter 8a45823906 fix(harness): release a won-without-job claim back to pending instead of abandoning a possibly-owned row (G1e)
The won:true/no-job protocol breach may still have committed the claim, so
falling through abandoned a row this worker owns — wedging a full lease window
and producing a false worker-reclaimed-pending overlay on the next sweep.
Mirror the decode-failure containment: best-effort releaseJob(id, worker,
"pending") (no work happened) before continuing, with refusals and throws
swallowed+warned.
2026-06-11 20:38:39 -07:00
Jordan Ritter fc14f2e39c fix(harness): claim-delete long-expired stale-aged leases in the lease phase instead of re-queueing (G1d)
A row whose lease expired beyond its family's stale window is past the
recent-lease heuristic's cross-sweep protection: re-queueing it emitted a
'back in flight' worker-reclaimed-pending signal that the NEXT sweep falsified
by stale-deleting the row. When a lease-phase row is already stale-expirable
(parseable created past maxAgeMs AND lease expired longer than the window),
claim it under the sweeper id and delete it — no comm error, counted in
expiredPending. Recently-expired rows keep today's re-queue path; header's
cross-call-protection claim corrected.
2026-06-11 20:38:39 -07:00
Jordan Ritter ef5ec1bcf5 fix(harness): exclude clause-unsafe family rows by id and continue discovery (G1f)
A backslash (or empty) family used to BREAK family discovery, starving every
younger family behind the offending row for up to 3h. The offending ROW needs
no family-charset semantics to exclude — push an id != "<rowId>" clause and
continue, keeping the warn. Younger families are still discovered; the unsafe
row stays unclaimable.
2026-06-11 20:38:38 -07:00
Jordan Ritter d7505d2644 fix(harness): close the empty-string family partition gap and forbid empty payload sentinels (G1c)
probeKeyFamily("") takes the prefix-LIKE clause leg, so the empty family's
inclusion clause matched every leading-colon key and its exclusion clause hid
ALL leading-colon families from discovery. Treat family "" as clause-unsafe
in familyClauseSafe (skip+warn, like the backslash class), and require
non-empty probeKey/serviceSlug/driverKind/meta.runId in assertServiceJobPayload
so the documented forbidden empty sentinels fail loud at both the enqueue and
claim-decode boundaries.
2026-06-11 20:38:38 -07:00
Jordan Ritter 620a24b75c fix(harness): keep a renewed-without-job lease assumed-live instead of false-reclaiming a live job (G1b)
A renewed:true response missing its job view is a protocol violation on a
SUCCESSFUL renew — the worker still holds the live lease. Evicting + returning
null stopped the heartbeat and let the sweeper falsely reclaim the live job.
Return the cached lease assumed-live (like the indeterminate path); only
evict+null when no cached lease exists. Also corrects the stale 'We ONLY
return null when the CAS itself failed' comment.
2026-06-11 20:38:37 -07:00
Jordan Ritter 8a8b1bd004 fix(harness): treat deterministic 4xx renew rejections as definitive lease loss (G1a)
A JobClaimEndpointError 4xx from the renew endpoint fails identically every
beat, so the indeterminate assumed-live containment never converged — the
worker held a phantom lease forever while the server lease lapsed and another
worker double-ran the job. Mirror the sweep's 4xx carve-out: error-log
renew-rejected, evict payload+lease caches, return null. 5xx/transport/
2xx-unreadable keep the assumed-live path.
2026-06-11 20:38:37 -07:00
Jordan Ritter d16b878923 fix(showcase): scan non-d6 fleet-family sweep aggregate keys for comm errors (G3f)
The global lease sweep mirrors comm errors onto the status row keyed by the
reclaimed job's probe_key (resolveSweepAggregateKey → aggregateCommError) for
ALL four fleet families, but the dashboard's decodeCellCommError candidate
scan only covered the d6 family's d6:<slug> aggregate. The smoke/demos/deep
families land on d4:<slug>, e2e-demos:<slug>, and d5-single-pill-e2e:<slug> —
rows the dashboard reads nowhere else — so reclaim/crash overlays on those
families were invisible. Add the three candidate keys (e2e window).
2026-06-11 20:38:36 -07:00
Jordan Ritter 2fdd6132b6 fix(showcase): dashboard status-model hardening batch (G3e)
- Object.freeze the UNSUPPORTED CellModel + NOT_WIRED_LEVEL singletons —
  they are returned by reference to every caller, so one consumer mutating
  them would corrupt every unsupported/not-wired cell.
- Remove the unfirable !d5.exists chip arm (d5.exists and d6.exists derive
  from the SAME CATALOG_TO_D5_KEY entry, so !d5.exists && d6.exists is
  impossible); fix the decision-table and d6Effective docs that described
  the unfirable behavior, keeping a note for a future map split.
- Constrain upsertByKey's generic to StatusRow so rowsAreNoop can never
  vacuously match a non-StatusRow type whose fields are all undefined.
- keyFor: validate the dimension segment for ':'/'/' like slug/featureId.
- decodeCellCommError: scope the comm-error staleness window per row family
  (e2e 6h for d5/d6/e2e, D4 1h for chat/tools, liveness 45m for health)
  instead of applying the 6h E2E window to liveness-cadence rows.
2026-06-11 20:38:36 -07:00
Jordan Ritter 13138208e8 docs(showcase): pin the deliberate pending-gate asymmetry harness↔dashboard (G3d)
The three-way contradiction: harness fleetSurfaceState routes green-only →
pending, cell-model routes green OR gray (non-regression) → pending while its
own comment claimed 'only green', and both files claimed 'exact mirror'.
gray→pending IS intended dashboard-side (gray is the dashboard-only no-data
colour ProbeState cannot represent). Fix the comments on both sides to state
the asymmetry precisely, extend the drift test to pin the derivation
difference itself (harness green-equality + no gray; dashboard
failure-passthrough + isRegression), add behavioral gray→pending and
red-passthrough tests, and fix chipColorToSurface's stale union doc
(now ChipColor | unreachable | pending).
2026-06-11 20:38:35 -07:00
Jordan Ritter 4c2cb309c1 fix(showcase): resolveD3 stale-green downgrade returns the effective row (G3c)
The stale-green branch returned the RAW row (status "amber" with row.state
still "green"), violating the .row.state ↔ .status invariant resolveD4/D5/D6
maintain. Return { ...row, state: "degraded" } like the other resolvers.
2026-06-11 20:38:35 -07:00
Jordan Ritter 3b5e9987a7 fix(showcase): map out-of-vocab D5/D6 fold states to a failing status (G3b)
stateToTestStatus mapped unknown runtime states (e.g. "error") to null,
swallowing the A2 rank-fold winner one step after the fold surfaced it — the
D5/D6 chip went benign gray no-data while live-status rendered the loud
"error" tone for the same row. Add foldStateToTestStatus (out-of-vocab →
"red") for the D5/D6 resolvers; D3/D4 keep the base mapping since the
chip's D1-D4 gate check already rescues their null (pinned by tests).
2026-06-11 20:38:34 -07:00
Jordan Ritter 8868047018 fix(showcase): rank-based anyMissing collapse in cell-model resolveD5/resolveD6 (G3a)
The strict missing-sub-row collapse used the literal `worstState !== "red"`,
which silently swallowed out-of-vocabulary runtime states (e.g. "error") that
the A2 rank machinery deliberately ranks ABOVE red — collapsing exactly the
state the rank fold exists to surface into benign gray no-data. Use
`rankOfState(worstState) < STATE_RANK.red`, mirroring the already-fixed
resolveD5Row/resolveD6Row in live-status.ts.
2026-06-11 20:38:34 -07:00
Jordan Ritter d2f56c7359 test(showcase): add truncatedByStop to the control-plane test's TickResult fake
Follow-up to the TickResult.truncatedByStop accounting field — the
control-plane wiring test's producer fake builds a literal TickResult
and needed the new required field for tsc to stay clean.
2026-06-11 20:38:33 -07:00
Jordan Ritter eed508da23 test(showcase): loop-crash worker test guarantees teardown via try/finally stop()
A failed assertion in the loop-crash test leaked the worker's bound
/health server (and the default-boot pool) across tests because stop()
was only reached via the in-body rejects.toThrow. Mirror the sibling
default-boot test's try/finally pattern with a rejection-swallowing
safety-net stop.
2026-06-11 20:38:33 -07:00
Jordan Ritter dca5731c42 test(showcase): harden job-producer + queue-client fixtures (CR G2e batch)
- derive the buffer-cap overflow boundary from
  MAX_BUFFERED_SWEEP_COMM_ERRORS instead of hardcoding 'old100'
- race the re-entrancy-guard's overlapping tick against a short timer so
  a broken guard fails diagnostically instead of deadlocking
- pin the warm-up abort test's AbortController+setTimeout coupling
  (AbortSignal.timeout would evade fake timers)
- document makeFakeQueue.enqueued as recording enqueue ATTEMPTS
- align queue-client samplePayload driverKind with the probeKey family
  so future payload-coherence cross-validation doesn't mass-fail
2026-06-11 20:38:32 -07:00
Jordan Ritter 31efa201b9 fix(showcase): account stop-truncated specs in TickResult.truncatedByStop so the tick outcome partitions exactly
Stop-truncated specs vanished from the tick outcome: tick-complete's
'services' could not be reconciled against enqueued + enqueueFailures +
skippedForBacklog. Add truncatedByStop to TickResult and the
tick-complete log (invariant: services == enqueued + enqueueFailures +
skippedForBacklog + truncatedByStop) and fix the 'enqueued' doc to
mention stop truncation.
2026-06-11 20:38:32 -07:00
Jordan Ritter 0f92764aaa fix(showcase): no-sink comm-error drops after the first warn are never fully silent — debug log with jobIds + running dropped-total
The one-shot no-sink warn kept a sink-less deployment from burying its
logs, but made every subsequent drop completely traceless. Later drops
now log at debug with the dropped jobIds and a running dropped-total
counter; the counter also rides the first full warn.
2026-06-11 20:38:31 -07:00
Jordan Ritter 394736b8ed fix(showcase): a second concurrent stop() must quiesce too — await the in-flight tick and any queued trigger
The first stop() flips 'running' synchronously, so a second concurrent
stop() hit the !running early-return and resolved immediately while the
tick the first stop() was quiescing on kept enqueueing. Both stop()
paths now loop-await the in-flight tick AND the queued trigger before
resolving.
2026-06-11 20:38:31 -07:00
Jordan Ritter 1bca4decf8 fix(showcase): queue an operator-triggered tick behind an overlapping in-flight tick instead of dropping it
The producer's re-entrancy guard skipped TRIGGERED ticks too, silently
losing an explicit operator 'run it NOW' whenever a slow scheduled tick
was in flight — contradicting the operator-intent-wins rationale that
already lets triggers bypass the backlog gate. A triggered tick that
hits the guard now waits for the in-flight tick and then runs; scheduled
ticks keep the skip. Capped at one queued trigger: a second concurrent
trigger gets the skip + warn.
2026-06-11 20:38:30 -07:00
Jordan Ritter 3a3634ee4b test(showcase): queue-client fakes throw on shapes they cannot honor
- rowMatchesFilter throws when multiple positive clauses for one field
  are ANDed (the OR-model would evaluate them wrong → vacuous pass)
- makePagingPb throws on unmodeled sort keys instead of silently
  returning insertion order
- documented the escaped-quote limitation of the clause value-extraction
  regex and likeToRegExp's dangling-backslash fallback
(buffer-cap boundary in job-producer.test.ts deliberately untouched —
sibling slot owns that file)
2026-06-11 20:38:30 -07:00
Jordan Ritter 9ce9b2d947 fix(showcase): fleet-claim hook hardening — explicit status, lease floor, workerId type, claim idempotency
- release: status is REQUIRED (the old '|| done' fallback silently
  finished a job whose caller omitted/emptied status) — 400 like
  jobId/workerId
- claim+renew: leaseSeconds floored at 1s (0.001 previously yielded a
  1ms lease — instantly stealable, renew thrash)
- all handlers: reject non-string workerId (a JSON number coerces into
  the text claimed_by column and the holder can never renew/release)
- claim: same-holder live-lease re-claim answers claimed:true with an
  alreadyHeld:true marker (timeout-after-commit retry no longer abandons
  a row the worker actually holds); client treats it as a plain win
  (ClaimEndpointBody documents the no-op marker)
- hook-parity pins updated/added for every change
2026-06-11 20:38:29 -07:00