Tool-heavy turns (Opus doing 120s+ of tool calls before text) falsely
triggered fallback. New onActivity() method restarts the timeout watch
on tool_call, tool_result, message_start, and message_update events.
Also fixed fallback model in Redis from gpt-5.3-codex back to
claude-sonnet-4-6.
Canonical TelemetryEmitter interface + createGatewayEmitter factory.
~120 call sites migrated from gateway/observability.ts.
model-fallback and message-store now import TelemetryEmitter from here.
Old observability.ts deleted. 3 tests.
Three fixes for gateway stability:
1. Proactive compaction now blocks the drain loop (awaited, not fire-and-forget).
Pauses fallback timeout watch during compaction and resumes after with a
fresh timeout window. Prevents false fallback triggers when compaction
takes >90s.
2. iMessage connectAndRun() now rejects (not resolves) when socket closes
before connecting. This preserves the exponential backoff counter —
previously every ENOENT resolved the promise, resetting backoff to 0
and causing 5s spam forever.
3. ModelFallbackController gains pauseTimeoutWatch()/resumeTimeoutWatch()
for known slow operations (compaction, session migration) where no
streaming tokens are expected.
- New @joelclaw/model-fallback with TelemetryEmitter port (DI)
- Gateway imports from package, passes emitGatewayOtel adapter
- FallbackConfig type owned by model-fallback package
- 6/6 unit tests passing (fake timers, telemetry, recovery)
- Removed gateway/src/model-fallback.ts