* refactor(snapshot): give the Wave 4 policies neutral host seams (#1983)
#2005 established the presentation ownership boundary and moved the iOS
presentation policies out of `src/daemon/`. It left the three remaining Wave 4
policies behind their existing daemon adapters. This closes that gap, so
`src/snapshot/` owns host-side snapshot policy generally rather than
presentation alone.
Freshness recovery: the window vocabulary, the Android staleness classification
and its thresholds, and the retry loop move to `src/snapshot/snapshot-freshness/`.
The loop is parameterized by a classifier and a retry schedule, so how long a
backend may lag behind a real transition is a policy input rather than a
constant the loop owns. `src/daemon/session-snapshot-freshness.ts` keeps only
what needs a session — reading and retiring the window on store-owned
`SessionState`, and choosing the comparison baseline from snapshot lineage — and
remains the declared R7 owner of `androidSnapshotFreshness`. The two call sites
#1739 named as the Wave 5 blockers, `selector-capture-runtime.ts` and
`deferred-interaction-outcome.ts`, now reach freshness through the seam.
Timeout evidence: whether a failure is the accessibility-timeout shape becomes a
policy in `src/snapshot/snapshot-timeout-policy.ts`. The published
`details.androidSnapshotTimeoutScreenshot` payload becomes vocabulary in
`@agent-device/contracts/snapshot-timeout-evidence`, built through constructors
so an assembly site cannot publish a fifth, undeclared arm. It gets its own
subpath rather than riding the shared capture facade, which keeps it out of the
CLI cold-start closure. Typed details, diagnostics and screenshot evidence are
unchanged.
Screenshot-overlay policy: which Android nodes earn an overlay ref, and what
rectangle an overlay covers, move to `src/snapshot/screenshot-overlay/`. The
daemon keeps approved artifact and ref assembly only — ranking, projection to
screenshot pixels, drawing and PNG IO.
The boundary test generalizes from the presentation subtree to the whole facet:
nothing under `src/snapshot/` may import `src/daemon/`. It gains a positive
control, because a filter that stopped matching would look identical to a
boundary being obeyed.
The residual call sites #1983 also named are audited and deliberately left in
place. `direct-ios-selector.ts` carries no presentation policy; its two pure
exports are selector derivation and ADR 0011 delegation-on-error, whose owner
would be the selector pipeline governed by R19, not this facet. ADR 0004 records
the finding so it does not have to be re-derived.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GLYhmt5ZNHQATG8T8ZFo7R
* refactor(snapshot): address adversarial review of the Wave 4 seams
Three findings from an adversarial pass over bc95d7f, all in the new seams.
`SnapshotFreshnessRetrySchedule.deadlineMs` was an absolute epoch instant named
almost identically to the duration constant `ANDROID_FRESHNESS_RETRY_DEADLINE_MS`
that feeds it. A backend binding the loop with the duration instead of
`markedAt + duration` type-checked, drove `remainingMs` hugely negative, and
silently ran zero retries with no annotation. Renamed to `retryUntilMs` — the
pre-refactor local's name — and the doc now says which one it is. The recovery
loop also gains direct tests it never had: the trustworthy, recovered and
still-suspicious paths, plus an already-expired deadline that pins the budget to
the action rather than to whenever the first capture returned, which is the shape
the mis-binding would have taken.
Two stale doc references from earlier drafts of the same commit: the timeout
assembly claimed its evidence shape lives in `@agent-device/contracts/capture`,
which is where it deliberately does NOT live — following that comment would
re-home the type into the shared facade and reintroduce the cold-start closure
cost the dedicated subpath exists to avoid. And the freshness window doc cited
`SnapshotFreshnessPolicy`, a type removed before commit for being unused; the
real seam is the loop's `classify` callback.
No production behavior change.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GLYhmt5ZNHQATG8T8ZFo7R
* refactor(snapshot): key timeout evidence on a typed reason, budget retries by duration
Addresses the three review findings on the Wave 4 facet work.
1. The recovery loop accepted an absolute `retryUntilMs`, and its own comment
admitted that passing a duration type-checks and silently disables retries.
Documenting a footgun is not removing one. The schedule is now a duration
budget and the loop derives the deadline from the window's `markedAt`
itself, so there is no absolute instant a caller can get wrong. Two tests
pin the invariant: a budget already spent before the loop starts runs the
capture once, and the same budget retries or not depending only on how old
the window is — a loop measuring from its own start would return the same
count for both.
2. The timeout policy recognized failures from hint prose and helper message
text. That is a message shape standing in for a decision, and extracting it
into a named facet made it worse by promoting the sniffing to declared
policy. The Android platform boundary now decides once and publishes the
typed reason `accessibility-timeout`, joining the existing
`ANDROID_CONTENT_RECOVERY_REASONS` taxonomy in the contract that already
exists to stop producers and consumers growing separate ones. The facet
reads that reason. The hint is derived from it rather than decided
alongside it, so rewording prose can no longer change what a reader
concludes. Coverage now runs producer to consumer: the platform tests assert
that both timeout shapes publish the reason, that an ordinary helper failure
does not, and that the real policy recognizes exactly what the real producer
emits — the message-sniffing approvals are gone.
3. `SnapshotTimeoutEvidence` still permitted `annotated: true` with zero refs.
The annotated arm now carries a non-empty tuple, so the contradiction is
unconstructible rather than merely unconstructed, with a `@ts-expect-error`
guard that fails the build if it ever becomes valid again.
The timeout tests moved out of `snapshot.test.ts` into a cohesive
`snapshot-capture-failure-reason.test.ts` rather than growing a file already
over the size tripwire; its pin ratchets down 1495 -> 1445.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GLYhmt5ZNHQATG8T8ZFo7R
* refactor(snapshot): decide the capture-failure reason from machine values only
Addresses the two remaining typed-policy blockers on #2014.
P1. The previous commit moved the message sniff rather than removing it:
`androidCaptureFailureReasonOf` still ran `/timed out/i` over helper and
wrapper prose, and a regex over the wrapper message for exit 137. A producer
sniffing prose is the same defect as a consumer sniffing prose, one layer down.
The decision now happens at the deepest boundary that holds the evidence, from
machine-defined values only. `snapshot-capture-failure-reason.ts` maps the
helper's structured `errorType` field by exact equality against
`java.util.concurrent.TimeoutException` — the same constant
`isUiAutomationConnectionTimeoutResponse` already compares — and the SIGKILL
exit code 137, which the fallback constructor knows structurally instead of
re-deriving from the message it just wrote. The helper-result,
session-protocol, and killed-instrumentation constructors attach the reason;
every layer above rewraps it. Both regexes are deleted, and the only
`TimeoutException` string left on the path is that constant.
This tightens behavior deliberately: a helper reporting ok=false with
timeout-looking prose but some other `errorType` is no longer classified as a
timeout. Both directions are proved end to end against the real producer —
four rewordings of the helper message (including empty) keep the typed value,
and three timeout-looking messages under non-timeout error types produce no
value and are not recognized by the real policy.
P2. The evidence union stored `overlayRefCount` beside the refs, so
`{annotated: true, count: 0, refs: [ref]}` and arbitrary mismatches stayed
assignable. No arm stores a count now — it is derived from `overlayRefs`, the
one source of truth — and the arms that carry no refs have nothing to count,
which `overlayRefsAnnotated: false` already states. Two type regressions guard
it: the empty-annotated contradiction, and the reintroduction of a stored
count, both as `@ts-expect-error` so the build fails if either becomes valid.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GLYhmt5ZNHQATG8T8ZFo7R
* refactor(snapshot): retire the duplicate timeout classifier on the session path
`isUiAutomationConnectionTimeoutResponse` compared `helper.errorType` to
`java.util.concurrent.TimeoutException` on its own, so the session fallback
diagnostic decided "was this a UiAutomation timeout" a second time. I cited it
as precedent for the constant in the previous round without noticing that
leaving it standing is the drift it was cited against: one taxonomy, two
deciders. The session protocol already publishes the typed reason on exactly
these errors, so the diagnostic now reads it.
The regression is proved rather than assumed: with the protocol's
`androidCaptureFailureReason` attachment removed, the new session-path test
fails; with it restored, it passes. It rides the existing
`ui-automation-timeout` fixture, so it exercises the real socket response
shape rather than a hand-built error.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GLYhmt5ZNHQATG8T8ZFo7R
---------
Co-authored-by: Claude <noreply@anthropic.com>
12 KiB
Agent Device Domain Language
Canonical vocabulary for the automation domain. Use these names in code, tests, issues, and architecture notes; implementation decisions and procedures belong in ADRs and task guidance.
Language
Sessions, targets, and devices
Platform family: An internal ownership group for related automation platforms: Apple, Android, HarmonyOS, Vega, Linux, or web.
Platform leaf: A concrete OS and device shape within a platform family whose support is classified independently, such as iOS simulator, physical iOS, tvOS, or macOS.
Platform module: A private package that owns one platform family's device mechanics and exposes its metadata and runtime bindings.
Device inventory gateway: The platform-neutral composition of local-family and provider inventory sources.
Device runtime gateway: The platform-neutral boundary that reports runtime facts and binds an admitted device to its runtime owner.
Runtime owner: The one local platform module or provider runtime selected to execute behavior for an ownership-qualified device.
Request binding: A request-lived attachment of cancellation, diagnostics, progress, and admitted context to a runtime owner.
Bound device runtime: The behavior-bearing view returned after a request binding proves the required runtime operations.
Runtime facet: A capability-cohesive interface on a bound device runtime with semantic inputs and typed outcomes.
Runtime fact: A typed claim about behavior available for one exact platform leaf, device or backend, and provider mode.
Narrowed bound runtime: A projection that exposes required facets, optional preferred facets, and no undeclared facets.
Host capability: Narrow authority supplied to a platform module for host execution, diagnostics, progress, or native assets.
Target: The selected automation destination, such as mobile, TV, or desktop.
Session: Daemon-owned state for one selected target and its opened app or surface.
Device key: A stable provider-scoped identity used for device ownership and contention.
Device lease: Logical remote ownership of a selected device for a tenant, run, or client.
Lease provider: The remote connection source that routes and owns a device lease.
Runner lease: A mutual-exclusion guard for a platform helper process. It is not remote client ownership. Avoid: Device lease, process lease
Device claim: Host-global exclusive ownership of one local device by an open session or a sessionless mutating command.
Device-claim policy: A command's declared relationship to local device ownership, including observation, acquisition, release, and exclusive mutation.
Commands and routing
Command surface: The catalog of public command identity, interface exposure, adapter policy, and shared metadata across CLI, Node.js, MCP, and batch entrypoints.
Runtime use: A command's platform-neutral declaration of operations required for admission and preferred fast paths whose absence does not reject the command.
Inventory use: An inventory command's platform-neutral declaration for composing device sources without binding a selected device.
Daemon command registry: The daemon-side source of truth for route ownership and request-policy traits.
Runner command traits: Per-command classifications that control Apple runner lifecycle and recovery behavior independently of the public command surface.
Daemon RPC protocol version: The integer used to detect breaking compatibility across the remote daemon boundary.
Version-skew invariant: Local client and daemon versions must match; compatibility handling is reserved for remote daemons, separately versioned helpers, persisted artifacts, and released API consumers.
Interactions, selectors, and refs
Interactor: The legacy monolithic interface between dispatch and platform behavior, retained only for commands not yet migrated to request-bound runtimes. Avoid: New or migrated command behavior
Interaction dispatch path: One concrete route an interaction command takes from a resolved target to device execution.
Coordinate-first resolved element activation: An Apple interaction that resolves a semantic element and then activates its resolved center point, avoiding a second element lookup after navigation.
Parent-owned touch point: A point that preserves the selected parent's identity while avoiding independently interactive descendants at its center.
Guarantee cell: One dispatch-path-by-guarantee classification stating where an interaction guarantee is enforced, delegated, inapplicable, or waived.
Owned waiver: A guarantee gap with a tracking issue and explicit owner.
Delegation-on-error: A fast path returning semantic failures to the shared path. It establishes failure-side handling, not success-path parity.
Parity table: A golden cross-language rule table consumed by both TypeScript and native tests.
Coverage manifest: A contract test's declaration of the guarantee cells it proves.
Ref frame:
The session's authorization namespace for mutating @ref targets, containing a frozen observation
epoch and issuance scope separate from the latest operational snapshot.
Frame expiry seam: The point immediately before a mutating device operation where the active ref frame becomes invalid.
Mutation admission: The decision that an active ref frame's epoch and issuance scope authorize a requested ref mutation.
Ref generation pin:
An optional ~s<n> suffix that carries the snapshot generation from which an @ref was minted.
Deferred interaction outcome: Post-response state that records whether a mutation may still need outcome retry, stabilization, or snapshot freshness recovery.
Settled observation: An optional post-action observation that waits for a quiet UI and reports the difference from the pre-action tree.
Resolution disclosure: Bounded response evidence describing how an interaction target was resolved without issuing new actionable refs.
Gestures and touch
Gesture plan: A typed, platform-neutral normalization of one- or two-contact gesture intent into bounded pointer trajectories.
Android planned-touch executor: The Android boundary that selects a provider-native or instrumentation-backed executor for a normalized touch plan.
Multi-touch geometry: The centroid, span, angle, translation, scale, and rotation used to construct two-contact motion.
Snapshots and capture
Raw AX node: A backend-owned accessibility value before snapshot presentation.
Snapshot acquisition: One backend attempt's raw accessibility nodes and attempt-level capture facts.
Snapshot producer: The acquisition component that produced a snapshot's raw tree — the third identity axis beside the platform channel and the in-plan capture strategy. Producers on one channel carry different guarantees, so presentation, scope, and geometry logic keys on the producer, never the channel alone.
Presentation options: The policy input controlling how one snapshot acquisition becomes a public projection.
Snapshot policy facet: The host-side owner of neutral snapshot policy: presentation, freshness, timeout and overlay. Platform acquisition supplies raw facts and a fold policy; runner-side Swift presentation remains separate across the process boundary.
Capture hint: The acquisition-facing view of a snapshot request, derived once from presentation options. It names the projection a backend must serve, keeps raw traversal depth separate from regular presented depth, and may narrow acquisition only where that backend can prove the narrowing complete.
Regular presented-depth frontier: The acquisition boundary for an unscoped regular snapshot, measured against regular presented depth after structural wrappers collapse. It is distinct from raw traversal depth.
Snapshot eligibility: Membership in a presented snapshot projection, independent of whether a node is currently hittable.
Clip fold: The regular projection's single visibility interpreter, run inside presentation for every backend: viewport and scroll-container clipping, ancestor projection, scroll hints, and collapsed depth. Platform differences enter as a fold policy, never as a backend exception.
Presented node: A wire-facing snapshot value produced at the presentation boundary.
Snapshot capture plan: An ordered set of capture backends executed under one shared wall-clock budget.
Snapshot quality verdict: A structured statement of capture state, backend, degradation reason, effective depth, and collapsed content.
Snapshot projection: A view of one acquired tree. Interactive is a subset of regular, and regular is a subset of raw.
Declared capture residue: A fidelity limit in acquired evidence that presentation cannot repair and must disclose.
AX-unavailable target invalidation: The Apple behavior that discards a suspect cached application target after a root accessibility failure so the next command reacquires it.
Recording and replay
Script recording:
Session mode that captures portable actions and target evidence into a .ad script.
Avoid: Screen recording
Recorded input parameterization:
An explicit fill contract that sends literal text to the live app while storing a caller-chosen
${VAR} placeholder in durable records.
Open-to-destination script:
A self-contained .ad script that opens an app, reaches and verifies a destination, and leaves the
session active.
Destination guard: A selector-targeted wait near the end of an open-to-destination script that verifies its ready state.
Replay script source bundle: The complete caller-resolved set of script paths and contents needed for one replay or test run.
Screen-recording facet: A runtime facet that starts video capture and returns a live handle plus a durable descriptor.
Live resource handle: Process-local authority to finish or forcibly dispose active logging, recording, or profiling work.
Durable resource descriptor: Bounded, versioned identity and recovery state from which the same runtime owner can reattach to a resource.
Reattachment: A fenced recovery attempt by the descriptor's exact runtime owner that returns a live handle, completed result, missing state, or typed refusal.
Maestro compatibility
Maestro program: A source-preserving typed representation of the Maestro Flow syntax and behavior supported by agent-device, interpreted through the compatibility runtime.
Maestro observation generation: Compatibility-engine evidence captured since the most recent mutation; mutation invalidates it before dispatch.
Providers and tests
Provider: An external adapter that owns a device runtime or contributes transport to a platform module.
Provider-backed integration scenario: A device-free test through the real daemon request path that replaces only external device or host tool execution.
Cloud WebDriver runtime: A provider runtime that maps a cloud-owned Appium or WebDriver session into agent-device inventory, leases, runtime behavior, artifacts, and release.
Cloud artifact: Provider-hosted session output such as video, automation logs, device logs, or dashboard links.
Daemon artifact type: An optional semantic category supplied by the owner of a daemon-managed downloadable artifact.
Provider transcript: An exact record of external provider calls used to verify command translation.
Scenario transcript: A command-level integration flow describing user-visible behavior through daemon commands.
In-process provider scenario harness: An integration runner that invokes the daemon request handler without opening an HTTP listener.
HTTP contract test: A narrow test of JSON-RPC transport, authentication, and response finalization.