Replace current 5.3 model rules/defaults with gpt-5.4 across AGENTS, skills,
gateway fallback docs/config, system-bus model constants, inference-router,
pi config, and Langfuse normalization.
Add packages/restate/src/trigger-prd.ts to compile markdown PRDs into a DAG
using pi with openai-codex/gpt-5.4, then submit Restate runs whose story nodes
dispatch host agents with a mandatory joelclaw mail contract.
Add host-only internal endpoints in packages/system-bus/src/serve.ts for:
- /internal/agent-dispatch
- /internal/agent-result/:id
- /internal/agent-await/:id
Allow per-node Restate handler timeouts so long-running PRD stories can wait on
agent completion without faking success.
emitEvent passes the full serialized message transcript as a curl -d argument.
Long sessions produce multi-MB payloads that exceed macOS ~256KB argv limit,
causing 'spawn E2BIG' during auto-compaction.
Switch from -d <payload> to -d @- with stdin pipe. Same curl behavior,
no argument size limit.
V3: Thread fade lifecycle with memory archival
- runFadeCycle() checks every 5min (debounced in turn_end hook)
- Threads older than 72h get archived via joelclaw memory write
- OTEL emitted for each archival (thread.archived event)
- readThreadSnapshot() helper reads daemon-persisted state
V4: Budget-efficient per-turn thread injection
- Daemon injects thread index (formatThreadIndexForPrompt) into
every prompt — active threads show full summary, cool show label only
- Current thread tag shows label + summary + message count
- No extension-level injection (avoids context accumulation from
hidden sendMessage calls on every turn)
Combined with V1 (classification + bidirectional reply-to) and V2
(compaction survival), the full ADR-0209 is now wired:
- Classify → Tag → Inject → Reply → Persist → Fade → Archive
Verify: restart gateway, observe thread lifecycle over 72h+,
check memory observations for archived threads.
Daemon persists thread snapshot to ~/.joelclaw/state/thread-snapshot.json
after each classification (haiku, reply-to, or fallback). Gateway
extension reads snapshot on session_compact and injects active/warm
threads into recovery context with emoji lifecycle, labels, age, and
message counts.
Thread state now survives compaction cycles. On recovery, the model
sees which threads were active and can continue them coherently.
Snapshot is best-effort (fire-and-forget write, silent catch on read).
Thread anchors are NOT persisted across gateway restarts — that would
need Redis storage (V3).
Verify: restart gateway, have multi-topic conversation, trigger
compaction, verify threads appear in recovery injection.
Root cause: pi.sendMessage() inside session_compact handlers queues messages
that trigger agent.continue() → model response → _checkCompaction() → second
compaction. The compaction summary (~16K) + system prompt (~35K) + kept messages
(~20K) already consumes 36-55% of context window, so recovery pointer injection
pushes past threshold again.
Fix: Track lastCompactionTs in both session-lifecycle and gateway extensions.
If session_compact fires within 60s of the previous compaction, skip recovery
pointer injection entirely. First compaction gets full pointers; cascading
second is a no-op.
OTEL events added for tuning:
- compaction.before: contextPercent, contextTokens, contextLimit, timeSinceLastCompactionMs
- compaction.inject: contentChars, estimatedTokens (enriched)
- compaction.inject.skipped: reason=rapid-recompaction
- gateway.compaction.inject: contentChars, estimatedTokens (enriched)
- gateway.compaction.inject.skipped: reason=rapid-recompaction
To verify: joelclaw otel search "compaction.inject.skipped" --hours 24
To tune: adjust COMPACTION_COOLDOWN_MS (currently 60s) or GW_COMPACTION_COOLDOWN_MS
ADR-0203 updated with full root cause trace through pi-mono source.
Gateway periodic context refresh and compaction recovery were dead code —
spawn (child_process) was used by gwRunRecall but never imported, and
emitGatewayOtel was called but never defined/imported.
Added spawn import and inline emitGatewayOtel helper that shells to
joelclaw otel (same pattern as memory-enforcer) to avoid depending on
@joelclaw/telemetry resolution in pi extension runtime.
The once-per-session injection guards were declared inside memoryEnforcer()
closure but referenced by seedRecall() and seedSystemKnowledge() at module
level. Bun transpiles without type-checking so it compiled, but threw
ReferenceError at runtime — silently discarding every recall result.
Move to module scope, keep reset in session_start handler.
ADR-0204 Change 1 completion.
seedSystemKnowledge runs on before_agent_start which fires EVERY turn.
Without a guard, each turn added a new hidden sendMessage to context.
Over N turns this accumulated ~500*N tokens of duplicate injections,
causing compaction cascades (compact → first response → compact again).
Fix: recallInjected and knowledgeInjected flags, reset on session_start.
Each injection fires at most once per session lifecycle. After compaction,
the recovery pipeline (ADR-0203) provides context — no need to re-inject.
Add three new capabilities to the gateway extension:
1. Periodic context refresh (every 30min):
- Runs joelclaw recall with topics from recent conversation
- Injects top 5 observations as hidden context-refresh message
- Gateway now knows what's happening in the system over time
- 60s check interval, 30min refresh interval, async non-blocking
2. turn_end compaction awareness:
- Tracks recent conversation topics for targeted recall
- Warm zone (40%): fire recall to cache relevant memories
- Hot zone (60%): force refresh before compaction hits
3. session_compact recovery:
- Builds pointer message from cached recall + recent topics
- Injects as hidden gateway-recovery message (display: false)
- Resets zone flags for next compaction cycle
- OTEL event: gateway.compaction.inject
The gateway session is long-running (hours/days). Previously it had zero
context refresh and zero compaction recovery. Now it maintains rolling
awareness via Typesense recall and survives compaction with real context.
ADR-0204 Changes 2 + 5.
When the agent first edits a file in a package (system-bus, gateway, cli,
web, etc.), automatically run a targeted Typesense recall for that domain
and inject relevant memories as hidden context.
Scope detection maps file paths to recall queries:
packages/system-bus/ → 'system-bus inngest worker functions deployment'
packages/gateway/ → 'gateway daemon telegram redis event bridge'
apps/web/ → 'joelclaw.com next.js web RSC content'
pi/extensions/ → 'pi extension hooks lifecycle session'
(etc.)
Each scope fires once per session (tracked in injectedScopes Set).
Uses existing runRecall() from ADR-0203 (lean budget, ~370ms).
Injects via pi.sendMessage({ customType: 'project-context', display: false }).
OTEL event: project.context.injected with scope, query, hitCount.
ADR-0204 Change 4.
seedRecall() and seedSystemKnowledge() were running Typesense queries,
emitting OTEL about the results, and DISCARDING the actual data. The
only thing memory-enforcer injected was the MEMORY_NUDGE write-pressure
prompt.
Now:
- seedRecall() parses recall output, formats top observations, injects
via pi.sendMessage({ customType: 'memory-recall', display: false })
- seedSystemKnowledge() injects formatted knowledge hits via
pi.sendMessage({ customType: 'system-knowledge', display: false })
- New OTEL events: memory.recall.injected, system_knowledge.injected
- Added formatRecallHits() and formatKnowledgeHits() parsers
Every session (interactive + gateway) now gets actual Typesense recall
results as hidden context. ADR-0204 Change 1.
Rip out all regex-based signal extraction (DECISION_SIGNAL_PATTERNS,
FAILURE_SIGNAL_PATTERNS, matchSignalPatterns, extractDecisionSignals,
extractFailureSignals, deriveRecallQueries, extractAssistantText).
Replace with actual Typesense hybrid search via joelclaw recall CLI:
- triggerRecall() fires async recall queries (lean budget, ~370ms) in
warm zone (40%+), caching results in recallCache
- buildCheckpoint() pulls validated queries + observation hits from
cache, with fallback to simple task-derived query
- session_compact pointer now includes 'Related memories' section with
real observations from Typesense, not regex scraps
- writeTaskContextToMemory() writes session task summary to memory so
future sessions can find it via recall
The async Inngest observe pipeline still handles deep semantic extraction
from full session transcripts. This pipeline focuses on RETRIEVAL of
existing knowledge for post-compaction context, not extraction.
- Move memory-enforcer canonical source to pi/extensions/memory-enforcer/
(was a loose file in ~/.pi/agent/extensions/, now git-tracked + symlinked)
- Remove duplicate memory-enforcer from pi-tools git/ (old HTTP-OTEL version)
- Remove duplicate identity-inject symlink from pi-tools git/
(extensions/ symlink to repo is the canonical one)
- Remove MEMORY_NUDGE section from AGENTS.md (keep only in extension hook
to avoid duplicating instructions in the system prompt)
- Update context budget analysis with specialist agent plan and
dynamic skill retrieval assessment
Fixes: double memory-enforcer hooks, double identity-inject loading,
~320 tokens of duplicated memory instructions in AGENTS.md.