* docs+ux: make device ownership discoverable end to end Complete the #1320 agent experience so 'busy? -> inspect -> choose or release' is discoverable from every surface an agent actually reads: - devices now projects the blocking claim owner per row (claimedBy with session and workspace, observe-policy projection; provably dead owners are excluded because the next open replaces them automatically), so an agent told a device is busy can pick a free one from the same listing. - help debugging gains a 'Device busy and ownership' section separating the two DEVICE_IN_USE flavors and their exact recoveries. - AGENTS.md documents both flavors; docs/agents/device-verification.md retires the last ps/kill recovery guidance in favor of device status, daemon stop --state-dir, and device release --stale (Stage 5 of #1320). - ADR-0010 no longer calls DEVICE_IN_USE 'the only retriable code' without naming the claim path's non-retriable override. - The rendered cross-worktree claim error gains a help-conformance quiz case binding (sample-output-device-claim-inspects-owner). - README points at device status / device release --stale. Part of #1320. * fix: key ownership projection by canonical device identity end to end Review findings on #2165: - blockingClaimOwnersByDevice keyed claims and inventory rows by bare device.id, so a live Android claim could project claimedBy onto an unrelated same-id Apple/Harmony/Vega row, with scan order picking the displayed owner. Both sides now use the canonical local device key (claim.deviceKey against canonicalLocalDeviceKey of the row's claim identity). The cross-family same-id regression was observed red against the bare-id keying. - The projection is now asserted across every hop the PR promises: client normalization preserves well-formed claimedBy and drops malformed ones, and the devices CLI formatter carries it through JSON data and renders the text line (MCP shares the same serialization).
4.4 KiB
ADR 0010: Error system conventions
- Status: accepted
- Date: 2026-07-03
Amended by ADR 0019. Existing normalization and cause-preservation rules remain accepted. When an operation and binding/resource cleanup both fail, the operation remains primary and cleanup is structured secondary diagnostic evidence; cleanup-only failure surfaces normally.
Context
The error system is centralized in src/kernel/errors.ts (AppError, normalizeError,
defaultHintForCode, retriableForErrorCode) but has ~1,200 construction sites across the CLI,
daemon, and platform layers. A July 2026 audit (code sweep + iOS/Android dogfooding) found the
kernel sound while call-site quality drifted: internal validation thrown as bare Error
(surfacing as UNKNOWN with a misleading hint), semantically wrong codes (COMMAND_FAILED for
busy devices, INVALID_ARGS for expired server-side resources), missing hints on the most common
agent failures (selector/ref misses), and consumers that drop fields (MCP tool errors carried only
message). Nothing documented the conventions, so each new site re-decided them.
Decision
- Every thrown error a user or agent can reach is an
AppError. Barethrow new Error(...)is reserved for provably unreachable invariants. Input validation — including daemon-side decoding of command input — throwsAppError('INVALID_ARGS', ...). Coercion of unknown caught errors usesasAppError(err, fallbackCode), which preserves the cause chain; do not hand-rollerr instanceof AppError ? err : new AppError(...). - Code selection. Use the most specific
KnownAppErrorCode;COMMAND_FAILEDis for genuine runtime failures of a well-formed request, never a catch-all for capability gaps (UNSUPPORTED_OPERATION), contention (DEVICE_IN_USE, the only code retriable by default; the cross-worktree device-claim path overrides it toretriable: false), or ambiguity (AMBIGUOUS_MATCH). New codes are added to the union deliberately; machine-dispatchable sub-classification rides indetails.reason(the lease registry is the model). - Hints answer "what should the agent run next". A hint is required wherever the per-code
default from
defaultHintForCodewould mislead; it is omitted where the default suffices — mass-adding boilerplate hints is worse than the default. Shared failure modes get shared hint constants next to the code that detects them (selectorFailureHint,STALE_REF_HINTinpackages/selectors/src/internal/resolve.ts;resolveIosDevicectlHint;bootFailureHint), not copy-pasted strings. Re-wraps preserve an existing hint rather than clobbering it. - Wrapping external tool failures. Prefer
exec.tserrors as-is. A hand-rolled wrap of anallowFailureresult must carry{ stdout, stderr, exitCode, processExitError: true }sonormalizeErrorcan surface the first meaningful stderr line instead of "tool exited with code N". - The wire error contract is
code,message,hint,details(redacted),diagnosticId,logPath, plus optional typed signalsretriable/supportedOn— the fields are absent unless confidently known. Every consumer surface (CLI human, CLI--json, SDK, MCP, events timeline) must carry code + message + hint at minimum; rehydration helpers (throwDaemonError,toDaemonHttpRpcError) copy all fields, andNormalizedErrorhoists the typed signals so--jsonconsumers see them. - Observability. Failed requests always flush diagnostics:
diagnosticId+ ndjsonlogPathaccompany every error (daemon-side under the session state dir, client-side under~/.agent-device/logs). Messages must not embed raw stderr dumps or secrets — structured context belongs indetails, which is redacted at normalize/write time.
Consequences
- Agents can branch on
code, retry onretriable, and followhintwithout parsing prose; wording can improve without breaking consumers (except strings pinned by help-text guarantees — update those tests in lockstep). - The known-code union stays small and meaningful; the widened
(string & {})type still lets runner/daemon codes pass through, so SDK consumers keep adefaultbranch. - New call sites have a documented bar: right code, contextual message naming the target, hint only when the default misleads, cause preserved.