2.9 KiB
Observability
Canonical notes for joelclaw telemetry, OTEL events, and triage semantics.
Health summary semantics
packages/system-bus/src/inngest/functions/check-system-health.ts emits system.health.checked after each health pass.
levelreflects whether the run observed any degradation at all.info= nothing degradedwarn= one or more health surfaces degraded
successis narrower thanlevel.success: falsemeans a critical health surface degraded (Redis,Inngest,Worker,Gateway,Typesense,Agent Secrets) or the agent-dispatch canary is unhealthy.success: truewithlevel: warnis valid when degradation is non-critical only, such asNFS Mounts.
This split is deliberate. O11y triage escalates failed operations, so non-critical degradations must stay visible without looking like a failed health-check operation.
CLI emission
Use --metadata for JSON context on manual OTEL events. The CLI does not have an --attributes flag.
joelclaw otel emit "task.completed" \
--source system \
--component skills \
--success true \
--metadata '{"session":"NimbleBadger","task":"install wzrrd-publish skill"}'
Metadata contract
system.health.checked metadata should include:
degradedCountcriticalDegradedCountnonCriticalDegradedCountcriticalDegradedServicesnonCriticalDegradedServices- full
servicesinventory with per-servicecriticalflag
That keeps dashboards honest while preventing noisy tier-3 escalations from warn-only, non-critical drift.
Talon worker supervision
Talon must not supervise the host system-bus worker when the canonical com.joel.system-bus-worker LaunchDaemon is loaded. The worker LaunchDaemon lives in the system bootstrap domain, so Talon checks both the current launchctl list <label> path and launchctl print system/<label> before deciding to start its internal worker supervisor. Dynamic launchd.* service probes also check launchctl list, launchctl print system/<label>, and launchctl print gui/$(id -u)/<label> so system LaunchDaemons like com.joel.gateway do not look dead from Talon's user LaunchAgent domain. If this detection regresses, Talon and worker-supervisor will fight over localhost:3111, producing repeated EADDRINUSE, SIGTERM/SIGKILL churn, and Inngest runs stuck at Unable to reach SDK URL or RUNNING.
Talon's node_schedulable probe treats explicit cordon (spec.unschedulable=true) and non-control-plane NoSchedule taints as critical, but allows the default node-role.kubernetes.io/control-plane:NoSchedule taint on Panda's single-node Talos/Colima cluster. That taint is normal; paging SOS on it is watchdog theatre, not reliability. Built-in critical probes now debounce through probes.critical_after_consecutive_failures (default 2) before healing/escalating, and repeated SOS alerts use a 4h default cooldown so one persistent outage does not page every 30 minutes.