Carries three gateway-worker changes that were reviewed and left uncommitted:
- Front Projection drops out of CRITICAL_COMPONENTS. A stale inbox projection
is worth knowing about and is not worth forcing overall health critical and
paging Joel. Health alerts were the single largest producer in the message
stream (740 of 1083 machine messages) and the largest fallback source.
- serve.ts runs the function-health registration check at startup, so the
assertion still runs when the durable function's own registration was lost
while Inngest was unavailable.
- replicate-critical-db.sh freezes one source generation before hashing it.
The builder renames critical.db atomically, so reading the live path twice
could pair the old file's checksum with the new file's built_at.
Also lands the three investigation briefs those findings came from.
archive-memory-observations requires the archive root to resolve under the NAS
root. snapshot-panda-system-log refuses to overwrite an existing snapshot or
receipt and verifies post-copy bytes/sha256 against the source.
Registers message/event-consumer on the host worker (event + cron triggers),
adds the message/event.consume.requested event type, a live proof script, and
the Brain work receipt. Convex deploy of the backing schema is still blocked;
step status set to blocked with receipts.
Closes the ack-loss inflation class (root cause 32e6e13a): the server
resyncs any post matching an existing (source_identity, from_offset)
regardless of run_id; clients write pending entries on ack-ambiguous
paths; the session-index consumer rejects ANY non-exact-start range
intersection loudly. QA-verified with the live failure chain replayed
plus retry/concurrency/isolation variants; the one NOT PROVEN claim
(different-start overlap) fixed with the verifier's fixture as a
regression test.
Replication verifier refuted two claims: all-backend failure exited 0,
and today's container flapping produced 7 transition alerts. Recall and
knowledge now exit 1 on total failure with the full failure chain;
the synthetic monitor requires 2 consecutive confirmations per
direction plus a 30-minute failure quiet floor.
Both containers up after the chmod fix; synthetic checks green on
three-body and maturin; outage proof passed (servedBy: three-body with
flagg-local skipped). Keep-alive flake fixed: clients send
connection:close, shim declares HTTP/1.1 for the next image build.
Bounce-minting lands on whatever was in flight, so the five-prefix fence
misses fresh mints (9 items across 5 arbitrary functions post-purge).
Authorized by Joel 2026-07-20 for window-scoped runs; ceiling unchanged.
Bounces mint envID-less items (595→597 across one bounce), so an exact
pre-pinned count is stale by apply time. Joel authorized window-scoped
re-scan 2026-07-20: same five function prefixes + missing-envID criteria,
hard ceiling 650, manifest regenerated at scan. Default exact-count mode
unchanged.
The capture replay verifier read the retired runs_dev collection, which
schema-drifted (missing source_identity -> HTTP 400) and false-conflicted
on fresh accepts. Query the canonical session index directly instead:
- exactRun/capturedSiblings are bun:sqlite queries against
~/.joelclaw/search/sessions.db (--sessions-db / SESSION_INDEX_PATH)
- resolve run-store-relative jsonl_path against RUN_STORE_PATH
(~/.joelclaw/runs-dev), same convention as joelclaw sessions
- open read-write with PRAGMA query_only=ON: WAL sidecars vanish between
writer sessions and a readonly open then fails with SQLITE_CANTOPEN
- tests seed a fixture sessions.db instead of faking a Typesense server
joelclaw recall and knowledge search open ~/.joelclaw/search/critical.db
first (29,906 docs: observations, Brain pages, vault, knowledge,
archived memory) with per-source freshness and clean fallback to
Typesense on any SQLite failure — proven through the compiled binary.
Builder holds an exclusive lock, refuses degraded or collapsed sources
(caught five malformed live pages), and swaps atomically. Observer refs
are honest source-labels (resolvableInRunsDev: false).
Refuted once by adversarial verification (fallback crash, unsafe
builder, fake backlinks); reworked and re-verified CONFIRMED on every
count. Receipts: critical-port-verification.svx
FTS5 sessions.db from the manifest keep-set: 134,013 Runs, 447,609
chunks, chunking via packages/memory/src/chunking.ts directly.
Adversarially verified: independent re-hash of all skipped bodies,
independent chunk parse, integrity clean. Cutover preconditions
recorded on the review step (WAL flip, reader/append wiring, parity).
Capture clients (pi, claude, codex) send byte-accurate from_offset,
to_offset, jsonl_sha256, and deterministic source_identity. Stale
cursors coalesce into one pending outbox segment under one Run ID.
capture-outbox-replay.ts is the only replay implementation; the audit
script fails closed on legacy flags and the scheduled nas-backup caller
is migrated. Raw Run writes reject divergent redelivery; a larger
proven-prefix body returns accepted_prefix without overwrite.
Verified adversarially (multi-byte offsets, repeated stale cursors,
fixture-only replay): 20 bun tests + 2 vitest + tsc clean. Receipts in
.brain/projects/typesense-reboot-recovery/patch-verification.svx.
Read-only scan of ~/.joelclaw/runs-dev producing byte-proven
prefix-coverage verdicts per Run. Receipts in
.brain/projects/typesense-reboot-recovery/run-manifest-summary.svx.
- File verification records ok/missing/timeout/error with null counts
on timeout/error instead of fake zero files / missing backup.
- Replay receipts preserve child exit code, skipped-large count, and
the first concrete child error; failures land in the host error list.
Forced-fixture validated (1 ms timeout -> backupStatus timeout,
unreachable Central -> failed attempt + Connection refused, healthy
fixture still passes). Receipts in dark-wizard
.brain/projects/fleet-check-maintenance-2026-07-18/.
joelclaw recall and voice recall now answer from Brain/observation
stores (19 tests). memory_observations archived to NAS (28,050 docs,
238MB, sha d2f3038d, verified) and dropped. Zombie functions
(reflect/review-promote/proposal-triage/batch-review) unregistered and
deleted with the MEM/FRIC suites (-6,306 lines). System log retired:
panda copy snapshotted (3,981 lines, sha 4e8e815d), slog write path /
system-logger / sync + backup + PDS mirror removed. Per decisions
decide-legacy-memory-retirement + decide-system-log-home.
Dedicated Inngest cron (17 * * * *) on the host worker: bounded live
conversations.history sweep per allowlisted channel (cc-matt-p +
brain-joel), non-Joel human roots classified untagged/started/shipped
from Joel's root reactions — Slack stays the only store. Finding runs
write one sensitive observation page (no message bodies); wakes fire
only on state transitions via joelclaw notify --event-id with 7-day
Redis dedup. Funnel: scripts/work-state-pass-funnel.ts. Seeded proof
01KXRGVJD5JPS8JGH95JP7PJSN (page + one wake), live run proved zero
repeat wakes. Closes the slack-work-state brief — all steps done.
New agent-usage-scan Inngest function, repo copies of the
agent-session-capture-backup and daily-shitrat skills, and a satellite
NAS mount installer script.
"The best modules are deep." -- A Philosophy of Software Design, John
Ousterhout
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Satellite rig runbook/setup hardening, NAS_HOST pinned to LAN IP per
mount contract, agent-mail daemon uv resolution, CLAUDE.md + brain
areas/pipeline notes, and function registry wiring for the new
usage/webhook/backup functions.
"direct exposure and experience, documentation, or a runbook" -- the
three ways knowledge transfers, Observability Engineering, Ch. 11
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
system/agent-session.capture-backup.requested event + audit-backup
script wired into nas-backup, with model-router escalation on backup
failures behind JOELCLAW_BACKUP_FAILURE_MODEL_ROUTER.
"Data durability, backup, and recovery" are where mixed transactional
and nontransactional solutions quietly corrupt -- Site Reliability
Engineering, Ch. 26
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
ClickHouse store with durable outbox + forward mode for satellite
workers (OTEL_STORE=forward -> flagg), clickhouse-otel capability
adapters for cli/sdk, backfill script, and cutover runbook. Raw OTEL
stops writing to Typesense; launchd/k8s env updated to match.
"For the observability domain, rows pertain to individual telemetry
events, and columns pertain to the fields or attributes of those
events." -- Observability Engineering, Ch. 16
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>