The hourly LaunchAgent ran a filter-delete of every document
(started_at:>=0) before each reimport. On 2026-08-05 08:02 that write
wedged Typesense 30.2's batched indexer while holding the collection
lock; searches exhausted all 16 threadpool threads in 48 seconds and
the server served nothing but /health 200s for four days.
Docs already use slug as id, so upsert was always idempotent. The wipe
only existed to remove deleted files; export ids + delete the stale
few does that without a mass delete on a serving collection. Files
that fail to parse keep their indexed copy instead of vanishing.
`joelclaw messages audit` has failed fleet-wide since cutover with
MESSAGE_JOURNAL_CONFIG_MISSING. The diagnosis "provision the reader
credentials" was wrong: the legacy ClickHouse journal has no listener on flagg
at all, and the events the command wanted were sitting in Convex the whole
time — the same stream `messages trace` already reads successfully.
Audit now falls back to the canonical stream when the legacy journal is
unconfigured, and labels which source answered. A journal failure that is not
a missing-credential problem still surfaces as an error rather than being
quietly papered over.
Verified live: audit --since 6h returns 100 events with source "convex".
Two chain deaths this week, two causes:
- 2026-07-21 campaign-pulse: the spawned run died on a WebSocket error before
arming its heir. Fire-and-forget arming left no trace. `joelclaw wake
assert-successor` now re-reads the registry, requires exactly one future
successor, and alerts by chain name when verification fails.
- 2026-07-22 daily-flow (schedule 853efe38): the dispatcher logged "attempt 1 of
3" and dropped the schedule. Pending schedules now survive failures 1-2 for the
reconciler to re-emit; only exhaustion clears and alerts, exactly once.
Recall hits carry observerSessionReference (kind capture-conversation,
resolvableInSessionsDb verified per hit) and producerRunId; the
work-state-pass producer stops minting generic runId; a read-time
resolver joins sessions.db by conversation_id within the epoch window,
returning all candidate Runs. Live proof: 827/827 observation pages
resolve. Mint-side changes ride in joelclaw-observer.
evaluateCriticalDbFreshness + CRITICAL_DB_REQUIRED_SOURCES exported
from packages/memory; the port and the maintenance alert import it and
their duplicated copies are gone. Both fixture suites pass byte-
unchanged — behavior identical by construction. pnpm-lock entry for the
new cli->memory workspace dep rides with the pending message-event lock
update.
Replication verifier refuted two claims: all-backend failure exited 0,
and today's container flapping produced 7 transition alerts. Recall and
knowledge now exit 1 on total failure with the full failure chain;
the synthetic monitor requires 2 consecutive confirmations per
direction plus a 30-minute failure quiet floor.
Both containers up after the chmod fix; synthetic checks green on
three-body and maturin; outage proof passed (servedBy: three-body with
flagg-local skipped). Keep-alive flake fixed: clients send
connection:close, shim declares HTTP/1.1 for the next image build.
joelclaw recall and knowledge search open ~/.joelclaw/search/critical.db
first (29,906 docs: observations, Brain pages, vault, knowledge,
archived memory) with per-source freshness and clean fallback to
Typesense on any SQLite failure — proven through the compiled binary.
Builder holds an exclusive lock, refuses degraded or collapsed sources
(caught five malformed live pages), and swaps atomically. Observer refs
are honest source-labels (resolvableInRunsDev: false).
Refuted once by adversarial verification (fallback crash, unsafe
builder, fake backlinks); reworked and re-verified CONFIRMED on every
count. Receipts: critical-port-verification.svx
joelclaw gateway doctor prints PASS/FAIL for daemon (with crash-relaunch
detection against an operator restart marker), live gateway source
CLEAN/DIRTY (a running daemon plus dirty source is today's outage), and
transport state; --live proves delivery with a real probe that must
return a Telegram platformMessageId — confirmed telemetry alone fails.
Every FAIL prints its exact remediation command. gateway restart writes
the marker and ends with the doctor summary.
Joel rejected emoji reactions as the operator action API (2026-07-19).
Labeled inline keyboard buttons on outbound DMs publish
message/action.requested with kind: callback and stable learner-flow.*
action ids; the callback_query owner answers, authorizes, resolves
flowId, verifies the declared action, then publishes. Legacy reaction
consumer stays until the live Seen canary passes.
joelclaw notify send now accepts --kind memory|alert|digest|ask|receipt,
passed through the Redis envelope and honored by the gateway compat shim
(invalid kinds are rejected instead of guessing a lane). Without --kind,
low/normal-priority sends are inferred as digest, which the Telegram
digest-lane classifier may silently suppress — that suppressed a
Joel-requested message twice on 2026-07-18. Operator-lane kinds
(ask/alert/memory) always deliver. CLI next_actions now point at otel
delivery verification instead of the ignored --telegram-only flag;
skills/messaging documents the flag and the verify rule.
Also lands the in-flight contract-v2 reaction-actions work from a
concurrent session (MessageAction schema in message-contract, telegram
adapter action messages, notify context.actions passthrough, inbound
reaction callback normalization) — entangled in the same files and
committed together deliberately.
81 tests pass across message-contract, chat-sdk, chat-sdk-inbound, sdk;
tsc clean. Live on flagg since 2026-07-18 (gateway restarted, CLI
rebuilt, kind=ask delivery confirmed via notify.compat_v2.confirmed).
joelclaw retro <brief-path> fires the event; memory-retro-writer reads
the closed brief + step Result sections, enforces the noise bar (<3
steps and no ledger → skip with OTel warn), condenses via pi, and
atomically writes dark-wizard .brain/resources/retros/<slug>-<yyyy-mm>.svx.
Per decide-retro-writer. 7 tests.
joelclaw recall and voice recall now answer from Brain/observation
stores (19 tests). memory_observations archived to NAS (28,050 docs,
238MB, sha d2f3038d, verified) and dropped. Zombie functions
(reflect/review-promote/proposal-triage/batch-review) unregistered and
deleted with the MEM/FRIC suites (-6,306 lines). System log retired:
panda copy snapshotted (3,981 lines, sha 4e8e815d), slog write path /
system-logger / sync + backup + PDS mirror removed. Per decisions
decide-legacy-memory-retirement + decide-system-log-home.
Dedicated Inngest cron (17 * * * *) on the host worker: bounded live
conversations.history sweep per allowlisted channel (cc-matt-p +
brain-joel), non-Joel human roots classified untagged/started/shipped
from Joel's root reactions — Slack stays the only store. Finding runs
write one sensitive observation page (no message bodies); wakes fire
only on state transitions via joelclaw notify --event-id with 7-day
Redis dedup. Funnel: scripts/work-state-pass-funnel.ts. Seeded proof
01KXRGVJD5JPS8JGH95JP7PJSN (page + one wake), live run proved zero
repeat wakes. Closes the slack-work-state brief — all steps done.
pane/schedule.requested contract, sleepUntil function with cancellation
waitForEvent race, Redis pending registry, joelclaw wake CLI. Emits
pane.schedule.due to the gateway with late flag past 5min. Worker-built,
steering-verified: tsc green, tests green, live enqueue+cancel receipts.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Worker drops Mux/Garage/joelclaw-api leases for the video path (wzrrd
owns upload and durable state); host registry asserts only the thin
publish function; joelclaw video trace accepts wzrrd slugs and points
at wzrrd video trace for cloud history.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
joelclaw video trace <resourceId|slug> reconstructs a publish from
ClickHouse forensics + Inngest runTrace GraphQL (notes when coarse v1
status disagrees). Worker leases WZRRD_VIDEO_SERVICE_TOKEN for the
wzrrd push seam.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Host→service placement as data (endpoint-resolver config), shared
probeK8sHealth for check/system-health and joelclaw status, hostname
normalized before lookup. On non-hosting hosts kubectl never executes
and k8s never enters degraded/criticalDown/system.fatal. Remote tailnet
probing deferred (flagg has no panda kubeconfig). Worker-built,
steering-reviewed: 23 targeted tests, tsc green.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
memory-skill-index: reflector-archived originals go quiet (rollups rank
instead); rollups/dream-receipts never narrate the arc. Bench fixtures:
two dead wayfinder paths repaired to successor briefs; the archived
observation fixture follows its page into archive/ (identity survives
via canonicalPageKey normalization on the wiki side). Reflector question
closed with full acceptance receipts; nightly LaunchAgent + week/month
ladder graduated as its own question.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
One '[REDACTED]' tokens field (distiller redacted a count because the field is
named tokens) aborted the whole run, leaving 5 pages unpublished. Now: index
the valid, return skipped[] with per-file errors.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- summarizeRollupCost: OpenRouter-benchmark cost for rows without provider
cost, wired into joelclaw usage and the daily token report
- clickhouse-usage-query: source option (router|agents|all) so agent_usage.turn
tailer events roll up alongside model_router.result; coverage % is router-only
- claude tailer parser: dedupe by message.id (Claude Code repeats usage per
content-block line; counts inflated ~2x) — inflated rows purged, offsets reset
- inference: non-zero pi exit with empty output now fails for json callers too
(was recorded as circuit success); --mode json event-stream soup no longer
leaks into result.text; caller metadata can't clobber usage fields
- daily report: rollup widened to 5000 groups; agent-sessions line added
- k8s worker manifest published with SLOG_SYSTEM_ID (pod blocked by
pre-existing flannel failure, see brain note follow-ups)
- brain notes + skills updated: langfuse decommission receipts, usage receipts
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>