Commit Graph

91 Commits

Author SHA1 Message Date
agent 4778a27512 feat(daemon): land the selector capture seam with get as its first consumer
Takes ownership of the request-bound selector capture seam from #1876, which
cannot ship standalone: with find's cutover deferred it had no consuming
command (ADR 0019 §10) and was not dead-code clean (check:production-exports
19 -> 20). `get` is its first consumer, so it lands here.

Adopts find's handoff as given. The one shape change, approved by the
coordinator: the selector family gets its own capture uses carrying a PREFERRED
`readTextAtPoint`, declared ALONGSIDE the snapshot uses so `snapshot`/`diff`
keep binding exactly what they bind today. The read is surfaced through the
existing arms of `bindSnapshotCaptureRuntime`, reusing the same
selectActiveAppSnapshot / selectSnapshotWithoutActiveApp selectors — no second
plan-to-operation dispatch.

`get` now runs through `createBoundSelectorRuntime`; `resolveBoundGetRuntime`
and its test are deleted as superseded, and `'get'` leaves the
`createSelectorRuntime` capability union.

The legacy read adapter survives for `find <q> get text` and is selected by
which command constructed the runtime — never by failure, family, environment,
or flag — so `get` cannot reach it. It retires in find's cutover, where the
last consumer moves.
2026-08-19 14:37:30 +02:00
agent 9056a4c790 fix(get): admit before the direct-iOS fast path; close the element-read outcome
Review blockers on #1877.

1. `dispatchGetViaRuntime` could complete the direct-iOS selector query before
   `resolveBoundGetRuntime`. Once `get` declares `device-runtime`, ADR 0019
   requires resolve -> admit -> bind before anything in the request path
   operates, so admission now runs first for every target shape and the fast
   path is a fast path *within* an admitted request. Regression: an eligible
   direct selector cannot operate when facts refuse admission.

2. `readTextAtPoint` returned `Promise<string>` and `readTextForNode` caught
   any throw and fell back, assigning a typed diagnostic after an untyped
   failure. It now returns a closed `ElementTextReadOutcome`; fallback happens
   only for the contract's classified reasons; unexpected errors propagate.
   The reason union is derived from its runtime list so the two cannot drift,
   and an unhandled reason is a compile error at the consumer.

This retires the generic catch the start record promised.
2026-08-19 13:58:54 +02:00
agent 8f7d55001d refactor: migrate get to the request-bound device runtime
`get` declares `elementReadRuntimeUse` (required `captureSnapshot`, preferred
`readTextAtPoint`), admits once from exact owner facts, refuses before binding,
and binds exactly once. Its capability bucket, the static HarmonyOS/Web command
sets that augmented it, and `requireCommandSupported` admission for `get` are
gone; `'get'` leaves the `createSelectorRuntime` capability union.

The neutral `readTextAtPoint` operation replaces the branch-per-family legacy
`read` dispatch on the `get` path. Every local family and both providers now
classify it exhaustively — Web, HarmonyOS, Vega and every provider row report it
unavailable, which is behaviour-preserving because the legacy dispatch had no arm
for them and threw on every call before falling back.

R36 is the new parametrized cutover row.
2026-08-19 13:58:54 +02:00
agent 30435df1b6 refactor(daemon): one capture-input builder and one admit-then-bind step
Behaviour-neutral. No descriptor changes platform execution and the cutover
table is untouched.

- buildRuntimeCaptureInput moves to its own module so every request-bound
  capture consumer builds CaptureSnapshotInput one way.
- The admit-then-bind sequence in the snapshot/diff resolver becomes one named
  step, ready for the selector units' second caller.
- CaptureSnapshotInput gains an optional per-capture signal, composed through
  captureSnapshotSignal by every snapshot runtime owner, so a polling consumer
  can enforce a poll deadline rather than inheriting only the bind-time signal.
- handlers/find.ts splits into focused target-capture and match-resolution
  concepts (600 -> 346 lines); behaviour unchanged.
2026-08-19 13:44:37 +02:00
Michał Pierzchała 79dddaf781 refactor: tighten viewport runtime facts (#1873) 2026-08-19 11:06:18 +02:00
Michał Pierzchała 37b1bc8cbd refactor: migrate viewport to request runtime (#1864)
* refactor: migrate viewport to request runtime

* fix: preserve viewport cutover evidence
2026-08-19 10:38:27 +02:00
Michał Pierzchała 294654a3a5 fix(android): resolve snapshot scope once and disclose the API 23 occlusion-scan gap (#1832 C1/C2) (#1846)
* fix(android): resolve snapshot scope once and disclose the API 23 occlusion-scan gap (#1832 C1/C2)

- Android resolves --scope inside its projection only, under the shared scope specification
  (matchesSnapshotScope in @agent-device/contracts/snapshot: first document-order match over
  label/value/identifier, empty on no match). The daemon post-wire scopeSnapshotNodes pass skips
  the android backend, so scope has one owner and one no-match semantics instead of BFS+fallback
  followed by document-order+empty.
- Golden table contracts/fixtures/snapshot-scope-policy.json is asserted against the predicate,
  the Android projection, and the daemon pass; the Swift runner twin (#1797) consumes the same table.
- androidSnapshot.occlusionScanUnavailable discloses helper trees without drawing-order (API 23),
  where the covered-sibling pruner cannot run. Disclosure only; C1 stays open until occlusion moves
  to the daemon annotator.

* fix(android): resolve scope over the presented tree and stop dropping it on interaction captures

Adversarial review findings on the first commit:

- BLOCKER: captureSnapshotData spread `snapshotScope: undefined` over flags, so an interaction
  capture (press/click/fill/longpress/hover --scope, --settle observation) reached the Android
  platform unscoped while buildSnapshotState still saw the scope. The post-wire pass used to rescue
  it; after skipping android it returned the unscoped tree. One effective scope now feeds both.
- Scope resolves over the PRESENTED nodes of the requested projection, not the acquired tree, so an
  acquired match that membership drops no longer empties the snapshot, and Android matches the
  domain iOS's pass uses.
- Slicing after the walk keeps ancestor context (hittable/collection/chrome) above the scope root,
  so scoped -i is a subset of unscoped -i; --depth stays scope-relative.
- Shared findSnapshotScopeRange/reindexSnapshotNodes so the daemon pass and the Android projection
  run one implementation; scope slice extracted to ui-hierarchy-scope.ts (mirrors its test).
- parseUiHierarchy moved to a test fixture module (it had no production caller left).
- Golden rows sharpened (value row no longer matches via label on Android); the Android leg runs
  raw AND regular. CHANGELOG entry; docs wording corrected for iOS/@ref.

* fix(layering): keep the contracts snapshot façade exhaustive over snapshot-scope

* fix(android): scope to the first match whose subtree still has presented content

Review P1 on #1846: with scope resolved strictly over presented nodes, `snapshot -i --scope panel`
answered "no nodes" whenever the matched container was a structural view membership drops — even
though the button inside it was exactly what was asked for — and `--depth 0` then hid a node the
response prints at depth 0.

The scope root is now the first document-order match whose subtree contributes at least one node to
the requested projection, and the result is that subtree's presented nodes re-rooted at depth 0.
Both failure modes die: a decorative match membership drops no longer empties the snapshot, and a
dropped container still scopes to its content. `--depth` under scope filters the depths the
response emits, so a node shown at depth 0 survives `--depth 0`.

Tests: the golden legs stay raw+regular (bare TextViews cannot survive -i, so an -i leg would
measure membership, not scope) with the projection interplay pinned by two dedicated tests on
actionable shapes; the 'not re-scoped after the wire' case now runs a real parsed scoped tree
instead of fabricated depth-0 siblings.
2026-08-18 18:44:13 +02:00
Michał Pierzchała d76e0f94e9 refactor: migrate snapshot to device runtime (#1779)
* refactor: migrate snapshot to device runtime

* refactor: complete snapshot runtime policy cutover

* test: enforce snapshot owner-facts admission

* refactor: consolidate desktop snapshot capture

* fix: scroll to visible iOS smoke targets

* fix: close snapshot cutover alias bypasses

* fix: constrain snapshot admission identity flow

* fix: enforce snapshot admission through owner facts

* fix: adapt replay source tests to snapshot runtime
2026-08-18 15:49:24 +02:00
Michał Pierzchała a853734f0c fix(webdriver): give cloud session creation its own budget and stop leaking billed sessions (#1782)
* fix(webdriver): give cloud session creation its own budget and stop leaking billed sessions

Cloud lease allocation ran under the generic 30s/1-retry request policy, so
BrowserStack iOS real-device session creation (45-90s) aborted client-side at
~60s on most runs. Each timed-out POST /session still completed server-side and,
being non-idempotent, was retried — leaving two billed provider sessions per
failed open with no id to release them.

- POST /session is its own phase: a 180s create budget (default), zero retries,
  and no request-bound abort, so the daemon always learns the session id.
- lease_allocate carries a 300s allocation budget surfaced to providers as
  LeaseLifecycleContext.deadline, and a matching 330s client envelope that
  preserves the daemon on timeout (a reset would SIGKILL mid-create and orphan
  every billed session the daemon held).
- The request's cancellation signal is ownership evidence: a session that
  completes after the requester left is released, not registered; a create that
  the transport gives up on surfaces typed evidence (provider + lease) so an
  operator can find and stop the maybe-orphaned session.

Closes #1774

* refactor: one canceled-request error, and tighten the #1774 shapes

Review pass over the session-create fix:

- The canceled-request error had nine hand-rolled copies (src/request/cancel,
  maestro shared, exec, retry, install-source x2, and the new provider one).
  It now has one definition in @agent-device/kernel/errors:
  createRequestCanceledError(details?, cause?) + isRequestCanceledError +
  REQUEST_CANCELED_REASON. Callers add evidence or a sharper hint; the reason
  itself is not overridable, so nothing can build one the predicate misses.
- lease_allocate's timeout bundle moves beside INSTALL_TIMEOUT_POLICY in the
  registry (same {...DEFAULT, envelopeMs, onTimeout} shape); the request timeout
  constant stays exported from timeout-policy like its siblings.
- Transport: fetch helper returns Response's own ok/status; the timeout reason
  const is private behind isWebDriverRequestTimeout.
- Client: one-use options type inlined; the two deadline helpers share one floor.
- Session-manager tests: shared makeRuntime/jsonResponse/afterEach restore.

Net -29 lines with the feature in.

* chore: keep the canceled-request reason private to the kernel

* fix: typed cancellation everywhere + own the AWS remote-access ARN through startup

Second-order follow-ups the #1774 refactor made cheap:

- markRequestCanceled aborts the request signal WITH the kernel's typed
  canceled error as its reason. Every signal.throwIfAborted(), aborted fetch,
  and 'throw signal.reason' in the daemon (20+ sites) now surfaces a canceled
  request as such instead of a bare DOMException that normalized to UNKNOWN —
  and no site has to know the factory exists.
- AWS Device Farm prepareSession owns the remote-access ARN from the moment
  create-remote-access-session answers: a startup timeout, the allocation
  deadline, or a canceled request now stops it before the failure surfaces
  (previously a timed-out startup left a RUNNING billed session behind — the
  same leak class as the WebDriver session, one phase earlier). The startup
  wait is capped by LeaseLifecycleContext.deadline and wakes on cancellation.
- BrowserStack's pre-session local app upload honors the request signal (an
  upload is not billed, so plain abort is right there).
- lease_heartbeat/lease_release share lease_allocate's preserve-daemon policy:
  the rationale — the daemon owns billed sessions; a reset orphans them all —
  applies verbatim.

Each AWS ownership test proven red without the guard (3/3).

* refactor: dedupe billed-resource cleanup and lease-signal wiring

Shrink pass — same behavior, less duplication:

- releaseOnFailure(primaryError, release) in webdriver-utils replaces the two
  identical 'best-effort stop the billed resource, attach cleanupError to the
  primary AppError' helpers (WebDriver session + AWS remote-access ARN); shared
  errorMessage too.
- The lease handler pulls the request signal from getRequestSignal(requestId)
  like every sibling handler, instead of threading a requestSignal arg through
  LeaseHandlerArgs and the request-handler chain. Drops the field, the wiring,
  and five mechanical test edits; the handler test now proves the request-bound
  signal (abort it, watch the provider's signal flip) rather than arg identity.
- Inlined the one-use requestHeaders back into fetchWebDriver.

Handler-signal test proven red without the wiring.

* fix(lease): the daemon releases a lease allocated for a gone requester; honest release evidence

Review follow-up. The provider was doing the daemon's job: it treated the request
signal as 'ownership evidence, not an interrupt' and needed three paragraphs to
say so. The daemon owns the request, so it now decides — generically, for every
provider — what happens to a lease that finished allocating after its requester
left: release it (provider + registry) and answer with the canceled error.

- lease.ts: after allocate returns, isRequestCanceled(requestId) →
  releaseAllocationForGoneRequester(). Release evidence is claimed ONLY on a
  clean release (no warnings, no throw); a WEBDRIVER_SESSION_DELETE_FAILED
  release is reported released:false with providerSessionId + a stop-by-hand
  hint (thymikee's finding: the previous evidence was success-shaped even when
  DELETE failed).
- WebDriverSessionManager: the createOwnedSession/releaseCanceledSession trio is
  gone; allocate is plain 'create with a budget; on failure clean up' again.
- LeaseLifecycleContext.signal is just cancellation, like everywhere else; the
  ownership-semantics comments on the contract, client, registry, AWS prepare and
  utils shrink to what the code no longer says itself.
- Tests: the two provider-level cancellation tests move to the daemon handler
  (where the logic now lives), plus the failing-DELETE regression; both proven
  red without the post-allocate check.

* fix(aws): the allocation deadline bounds remote-access startup, not the 120s default

Live iOS real-device run: startup needed ~128s and hit the standalone 120s
default while the daemon's 300s allocation budget still had room — the new
ownership guard correctly stopped the ARN, but the open failed for no reason.
When the daemon supplies a deadline it is the bound; the default only applies
standalone. Rerun: open in 112s, snapshot, clean close, session STOPPING.

* test(aws): pin that the allocation deadline outlives the 120s startup default; drop empty import

Review follow-ups on 7f9d1481a: a virtual-clock test (Date.now advanced 10s per
poll, RUNNING at 150s, deadline 300s) that fails on the old min(default,
deadline) logic and passes now; and the empty 'import {} from kernel/errors'
left in maestro/shared.ts is removed.

* refactor: finish the dedupe — one release path, kernel errorMessage, AWS on releaseOnFailure

Code-quality review at 7f9d1481a:
1. aws-device-farm.ts still carried its own copy of releaseOnFailure (the dedupe
   commit's script aborted before reaching it and I mis-verified). Now uses the
   shared helper; private copy deleted.
2. Empty 'import {} from kernel/errors' in maestro/shared.ts removed (273870099).
3. errorMessage() lives in @agent-device/kernel/errors; the two copies this PR
   had added (lease.ts, webdriver-utils.ts) import it. Sweeping the pre-existing
   copies is a follow-up.
4. lease.ts has ONE release path: releaseLease(registry, provider, lease,
   request, ctx) → { released (registry), provider } used by both the
   lease_release case (wire shape unchanged) and the gone-requester branch, which
   folds a throwing provider release into releaseError. 'released' now means the
   same thing in both; the provider verdict is a separate 'providerReleased'
   (warnings-free, no throw) that drives the stop-by-hand hint. -~35 lines.
5. sessionCreateTimeoutMs is Omit-ed at the WebDriverTransportOptions boundary
   instead of Pick-ed back out internally.

* fix(lease): 'released' on a canceled allocation means the billed session is confirmed gone

Re-review at 3665ea06: unifying the release path had made the cancellation
error report released:true from the daemon's registry record while the provider
DELETE had failed — success-shaped again, with the operator verdict demoted to
a second key. Fixed at the source of the ambiguity:

- LeaseReleaseOutcome names its bookkeeping field registryReleased.
- On the canceled error, 'released' is true only when registryReleased AND the
  provider released without warnings AND without throwing; the registry record
  is exposed as 'registryReleased'. The stop-by-hand hint keys on 'released'.
- lease_release keeps its existing wire field ('released' = registry; provider
  cleanup rides in 'provider'), unchanged.
- Regressions: failed DELETE and throwing release both pin released:false /
  registryReleased:true (+ providerSessionId, warnings|releaseError, hint);
  both proven red on registry-only semantics.

* ci: retrigger default-setup CodeQL

Run 32051017472 is wedged on GitHub's side: status=completed with
Analyze (python) still queued and Analyze (java-kotlin) failed only at SARIF
upload (503, 'No server is currently available'). It can be neither cancelled
nor rerun, and default-setup CodeQL has no dispatchable workflow, so a new push
is the only way to get a fresh run. No source change.

* test(webdriver): assert the typed timeout contract on the shared-budget probe

main's #1790 tightened this test to expect the raw TimeoutError DOMException,
which this PR intentionally normalizes into AppError{reason:
webdriver_request_timeout}. On the merge ref the two met and Coverage went red.
The regression now asserts the structured contract and that the second request's
budget is the shared remainder (~118ms of 200 after an 80ms first call).
2026-08-18 15:48:58 +02:00
Michał Pierzchała d0d5c8594c fix: serve remote daemon request diagnostics to the caller (#1801) (#1814) 2026-08-18 15:36:09 +02:00
Michał Pierzchała 60f6356b04 fix: read replay scripts on the caller and ship them with the request (#1810)
* fix: read replay scripts on the caller and ship them with the request

Closes #1802

* test: assert the caller-side replay path as a substring, not a hand-escaped regex

* perf(cli): load the Maestro engine only when a replay entry is a flow

The command registry evaluates every command family on CLI startup, so the replay script-source builder's static @agent-device/maestro import put the YAML parser on the --help path. It now loads on demand behind the format check, and the startup import-closure guard covers the engine the way it already covers node:http.

* refactor: share the replay request field vocabulary across the CLI and client views

The new replay script-source flags appear in both CliFlags and CommandExecutionOptions, which fallow flagged as a clone; ReplayRequestFields declares them once. The test-suite handler's missing-sources rejection now travels the typed-error path its sibling rejections already use, so the fix adds no branch to handleSessionReplayCommands.
2026-08-18 14:56:34 +02:00
Michał Pierzchała f843dc2df1 fix(scroll): keep saturated scroll gestures out of the status bar; gate Android replays from android/emulator (#1781 A1) (#1820)
* fix(scroll): keep saturated scroll gestures out of the status bar; gate Android replays from android/emulator (#1781 A1)

`pnpm gate replay-android` failed 4/8 whenever it ran after the full-tier Android E2E
(replays-nightly run 32107665052, job 95620294899): 05-app-lifecycle, 06-swipe-gestures and
both fixture replays diverged under "A system surface covers the app". The E2E was not the
cause. Reproduced on a pixel_7 / API 36 AVD with the same cutout geometry CI's
`avdmanager --device pixel_7` produces (status bar 136px, not the 63px of a plain 1080x2400
skin):

- `03-scroll-discovery.ad` runs `scroll up 3`. The scroll planner clamps travel to the viewport
  minus a 5% band, so the touch-down landed at y=120 — inside the 136px status bar — and
  pulled the notification shade instead of scrolling. On API 36 the app window is
  edge-to-edge, so the reported viewport starts at y=0 and includes that bar.
- The shade then covered every replay until `04`'s `back` closed it. Native readdir order on
  the runner (03, 05, 06, fixture/02, fixture/01, 04, 01, 02) put four files in that window;
  the last green run (2026-07-30) had 04 right after 03, so the pull was masked.

Fix in the product, not the lane: DEFAULT_EDGE_PADDING_FRACTION 0.05 -> 0.1 in the TS scroll
planner and its Swift port. Every real Pixel has a cutout (5.7% of a Pixel 7's height) and an
iPhone's Dynamic Island status bar is 6.9%, so any saturated `scroll up` opened the shade /
Notification Center for real agents too. Parity vectors updated in both suites plus a Pixel 7
regression vector (1080x2400, amount 3 -> touch-down y=240 > 136).

Second contamination the same order exposed once the shade was gone:
`fixture/02-selector-routes-covered-diagnosis.ad` is a #1715 reproduction recipe that FAILS
BY DESIGN at step 9 (covered-target refusal) and leaves the device in landscape, yet the
gate enumerated `test/integration/replays/android` recursively. iOS keeps gate replays in
`replays/ios/simulator` and fixture recipes in `replays/ios/fixture`; Android now mirrors that:
the six Settings replays move to `replays/android/emulator`, `test:replay:android` points there,
and `fixture/` stays E2E-owned (`full:fixture-replays` already runs 01 by path). android.yml
and the workflow-evidence fixture follow the path; the replay-compat manifest keeps the
historical paths it pins at released tags.

Verified live (Pixel 7 geometry, API 36, --retries 0): control run at main head in CI order
reproduces exactly CI's 4/8; with the fix, `pnpm gate replay-android` 6/6 in both native and
CI order, and `03` leaves Settings on screen (scroll up 3 now touches down at y=240).

* test(scroll): drive the TS and Swift scroll-plan parity vectors from one table (#1820 review)

The two suites hand-mirrored the same vectors and #1820 had to edit both by hand — the drift
class the repo already closes for the tap-point rule via contracts/fixtures/tap-point-policy.json.
The scroll vectors (plus both planner constants, pinned behaviourally on a 1000px axis) now live in
contracts/fixtures/scroll-gesture.json; scroll-gesture.test.ts and RunnerTests+ScrollGesture.swift
iterate it. Verified: vitest 10/10; the four XCTests run on an iOS 26.2 simulator with the unit
flag on (Executed 4 tests, 0 failures).

Also: test/ci/android-workflow-evidence.json says what it guards.

Follow-up for content-safe viewport bounds + discovery order: #1821.
2026-08-18 14:31:54 +02:00
Michał Pierzchała 8db36299e4 feat(web): add hover command for hover-gated UI (#1783) (#1786)
* feat(web): add hover command for hover-gated UI (#1783)

Add a first-class `hover <x y|@ref|selector> [--settle]` verb, admitted on
web only, that moves the pointer without pressing via the agent-browser
backend (mouse move). It rides the existing targeted-touch pipeline
(ref/selector/coordinate resolution, occlusion/off-screen guards, settle
observation, response builder, recording) through a new optional
Interactor/backend `hover` op that only the web provider implements.

Touch platforms have no hover state: capabilities advertise it on web
only and iOS/Android/Linux reject it at admission with a --platform web
hint; longpress stays the mobile hold-gesture verb.

Closes #1783

* fix(hover): native hoverRef route for web @ref, android coverage pin, revert skill edit

Review follow-ups on #1786:
- hover @ref on web now dispatches through the provider's own element handle
  (agent-browser `hover <ref>`) via a new backend hoverTarget, mirroring
  click/fill's ADR 0011 native-ref path — web ref frames carry no rects, so
  the coordinate route could never resolve them. The shared preflight +
  exact-ref dispatch is extracted into dispatchNativeRefInteraction and used
  by tap/fill/hover; the guarantee matrix native-ref row now lists hover.
- Daemon regression test is production-faithful: rect-less web ref frame,
  scoped provider, asserts no coordinate dispatch. Selector→coordinate and
  provider hoverRef tests added.
- Android emulator coverage summary pin 2/53 → 3/54.
- skills/agent-device/SKILL.md reverted (out of scope, AGENTS.md rule).
- Docs/help disclose that --settle with @ref on web shares click's existing
  limitation; use a selector or coordinates for the settled diff.

* test: drive hover through the apple output guard; cover direct hover dispatch

The provider-integration apple-leak guard partitions every public command
into driven/skipped; hover was neither, which failed Integration Tests and
took Coverage down with it. Drive it (it reaches the Apple capability
refusal, which is scanned like any other error response). Also cover the
direct-dispatch handleHoverCommand seam.

* test(web): drive hover @ref in the provider-backed web scenario

The integration-progress gate requires every public command to be referenced
by a provider-backed scenario. Add hover @ref to the web desktop flow: it
must reach the provider's hoverRef handle (never a coordinate) and be
recorded on the session without fabricated x/y, like click @ref.
2026-08-18 11:32:48 +02:00
Michał Pierzchała 681cad2222 test: assert the specific error code instead of any failure (#1781 B4) (#1790)
Converts the 20 test assertions across the repo that accepted ANY
failure (bare `expect(...).toThrow()`, bare `assert.throws(fn)`, bare
`assert.rejects(p)`) into assertions on the specific AppError `code`
each test is actually about, or — where the propagated error is
genuinely opaque (a mocked upstream failure whose identity, not its
shape, is the point) — identity assertions with a comment explaining
why.

Added a synchronous `assertThrowsAppError(fn, {code, message?})`
sibling to the existing `assertRejectsAppError` helper in
src/__tests__/test-utils/app-error.ts, exported via the test-utils
index, for the two src/ sites that needed it.
packages/provider-limrun and packages/provider-webdriver have no
test-utils dir and cannot import from src/, so those sites use
vitest's `expect(...).toThrow(expect.objectContaining({ code }))` or
an inline `assert.rejects(p, matcherFn)` instead.

Sites converted:
- packages/provider-limrun/src/app-log-runtime.test.ts:153-155
  (bare `.toThrow()` x3 -> `UNSUPPORTED_OPERATION`)
- src/daemon/__tests__/app-log.test.ts:39 (bare `.toThrow()` ->
  message match; plain Error, not AppError, from verified-file's
  identity check)
- src/daemon/__tests__/resumable-upload-range.test.ts:13 (bare
  `assert.throws(fn)` -> `INVALID_ARGS`); also fixed line 19's
  `assert.throws(fn, value)`, a documented Node.js gotcha where a
  string second argument is the failure message, not a matcher, so
  it was equally bare in effect
- packages/provider-webdriver/src/webdriver-client.test.ts:229 (bare
  `assert.rejects(p)` -> asserts the raw AbortSignal.timeout()
  rejection's `name`, since the transport re-throws it unwrapped)
- src/daemon/handlers/__tests__/session-device-claims.test.ts:129,
  151, 174 (bare `assert.rejects(p)` x3 -> identity assertions; each
  test's point is device-claim rollback/retention around an opaque
  mocked upstream failure)
- src/platforms/android/__tests__/settings.test.ts:109 (bare
  `assert.rejects(p)` -> `UNSUPPORTED_OPERATION`)
- src/platforms/android/__tests__/snapshot.test.ts:1071, 1342 (bare
  `assert.rejects(p)` x2 -> `COMMAND_FAILED` + message)
- src/platforms/android/__tests__/touch-helper-session.test.ts:526
  (bare `assert.rejects(p)` -> `COMMAND_FAILED`, wrong-protocol
  message)
- src/platforms/apple/core/__tests__/runner-command-retry.test.ts:472,
  527, 550, 762, 881, 1016 (bare `assert.rejects(p)` x6 ->
  `COMMAND_FAILED` with the recovery-path-specific details/message)
- src/platforms/apple/core/__tests__/runner-transport.test.ts:61
  (bare `assert.rejects(p)` -> identity assertion; fetchWithTimeout
  does not wrap fetch() failures into an AppError)

No repo-wide scanner/lint rule added (explicitly out of scope per
#1781); no test loosened.
2026-08-18 10:13:28 +02:00
Michał Pierzchała 0d3b7413c5 fix: prevent private AX subtree leaks at source (#1807)
* fix: prevent private AX subtree leaks at source

* fix: preserve values in settle signals

* fix: normalize settle signal semantics
2026-08-18 09:23:39 +02:00
Michał Pierzchała ed1d44fa62 fix: unify private AX scroll visibility (#1798)
* fix: unify iOS scroll snapshot visibility

* fix: preserve composed scroll hints
2026-08-18 00:47:25 +02:00
Michał Pierzchała f378050586 feat(snapshot): attach a fallback screenshot to sparse captures (#1764)
A sparse verdict already tells the caller to use a screenshot as visual truth,
which made that screenshot the guaranteed next command on every unreadable
screen — a second round trip to obey advice we authored. The user-facing
`snapshot` dispatch now takes the shot itself and links the path in its
warnings.

The fallback is deliberately hung off `dispatchSnapshotViaRuntime` and skipped
for internal observations: selector resolution, settle, and wait polling reach
`captureSnapshot` directly, so a wait polling an unreadable screen cannot turn
into a screenshot per poll. A failed shot is swallowed — the verdict's own
warning still carries the manual remedy, so the fallback can never fail the
snapshot that was asked for.

Sparse captures also say when the screen is the app's problem. Only the
`sparse-tree` reason code is evidence about the app: every backend reached the
screen and it published no semantic content, which is the same emptiness
assistive tech gets. `ax-rejected`, `budget`, `no-nodes` and `capture-failed`
are limits of this tool and stay unattributed, so readers are not sent to file
bugs against code that is not broken.
2026-08-16 16:06:10 +02:00
Michał Pierzchała d8a7d03faf refactor: route application lifecycle through runtime facts (#1759)
* refactor: route application lifecycle through runtime facts

Moves the canonical `open`, `prepare`, `close` and internal `runtime` descriptors
behind package-owned lifecycle bindings admitted from device runtime facts, while
daemon request/session policy and public response construction stay put.

Based on main, which already carries the boot unit, the parametrized cutover gate
and the apps unit. Readiness is package-owned there, so the Apple and Android
bindings call ensureAppleReady/ensureAndroidReady rather than a root readiness
bag; ensureAppleReady gained an onColdBootStart hook so open keeps warming the
runner cache in parallel with a cold boot, and a narrow markBooted port publishes
readiness' fresh observation so a flow still makes one simctl listing.

Cutover rows take R24-R27, clear of the accepted catalog and the sibling install
stack, and cutoverTableDefects rejects a duplicate rule id.

Two defects this unit introduced are fixed here rather than shipped:
`open <app> <url>` dropped the URL on a first open, and test-IME activation was
first fatal on an unobtainable helper and then over-caught. Helper unavailability
is a typed non-activation outcome now; fence, lock and post-record failures
propagate.

The duplication the unit had accumulated is gone: one runtime-admission module
instead of five per-command copies, one direct-lifecycle binding factory instead
of six hand-rolled packages, one transport-hint predicate, one session
finalization path, and no identity-wrapper module.

* fix: allocate lifecycle cutover rows after deployment

* chore: preserve lifecycle union reconstruction

* fix: reconcile lifecycle runtime stack

* refactor: tighten lifecycle runtime topology

* refactor: remove superseded runtime adapters

* fix: preserve stacked runtime cutovers

* test: preserve migrated runtime ownership

* test: move Android deployment retry ownership

* test: extract runtime hint fixtures

* fix: preserve lifecycle stack invariants

* fix: complete lifecycle runtime cutover

* fix: remove lifecycle cutover residue
2026-08-16 15:13:10 +02:00
Michał Pierzchała 66cca1a5b8 refactor: route install commands through platform runtime (#1758)
* refactor: route install commands through platform runtime

* fix: preserve stacked runtime facts

* fix: align deployment facts with shutdown runtime

* refactor: simplify capability facts projection

* style: format capability facts projection

* fix: preserve migrated capability ownership

* fix: propagate deployment artifact cancellation

* refactor: move Harmony deployment mechanics into package

* refactor: move Apple deployment tools into package

* refactor: move Android deployment tools into package

* refactor: inject deployment temporary storage

* refactor: remove superseded deployment helpers

* fix: preserve provider deployment transport
2026-08-16 15:13:09 +02:00
Michał Pierzchała 5855dfc2e0 refactor: route shutdown through device runtime (#1757)
* refactor: route shutdown through device runtime

* fix: cover shutdown cutover review gaps

* fix: propagate shutdown cancellation

* fix: move shutdown mechanics to platform owners

* fix: pass device to shutdown fact fixture

* test: cover shutdown facts in session state fixtures

* test: simplify Android shutdown assertions

* fix: preserve Apple shutdown cancellation
2026-08-16 15:13:08 +02:00
Michał Pierzchała 39cd4d346a refactor: route appstate through platform runtime (#1755)
* refactor: route appstate through platform runtime

* test: keep appstate capability fixture below complexity limit

* test: cover appstate required readiness fact

* fix: align appstate facts with boot readiness

* fix: keep appstate use declaration minimal

* fix: close appstate parity and ownership gaps

* docs: record final appstate size accounting

* fix: merge neutral runtime imports

* docs: align final appstate size totals

* fix: remove stale runtime dependency edges

* docs: correct appstate size accounting

* refactor: keep runtime-use factory internal

* docs: itemize runtime-use relocation

* fix: move appstate queries into runtime packages

* refactor: retire root foreground query paths

* refactor: share Android foreground parser ownership

* fix: preserve Android appstate parser precedence

* docs: keep appstate evidence in review artifacts

* fix: keep appstate runtime loading lazy

* fix: fail closed for stale limrun appstate

* fix: preserve limrun recovery and abort appstate

* fix: narrow limrun exact-owner recovery

* fix: allocate appstate cutover rule

* fix: reconcile appstate with merged main

* style: format harmony runtime test

* fix: allocate appstate rule id

* fix: allocate appstate layering rule

* fix: remove stale app command admissions

* fix: close appstate layering regressions

* fix: align Harmony capability parity with runtime facts

* test: cover limrun recovery-only readiness

* fix: keep Limrun recovery binding app-log only

* fix: parse Android app state in linear time
2026-08-12 16:55:50 +02:00
Michał Pierzchała c7565cb1f8 refactor(snapshot): clean snapshot ownership (#1754)
* refactor(snapshot): clean snapshot ownership

* fix(snapshot): address ownership review feedback
2026-08-12 13:44:21 +02:00
Michał Pierzchała 15133b587f fix: validate macOS recording finalization (#1732)
* fix: validate macOS recording finalization

* fix: preserve local macOS recording ownership

* fix: resolve recording teardown session keys
2026-08-12 12:51:13 +02:00
Michał Pierzchała f98814ff10 fix: preserve recording evidence across fault boundaries (#1734)
* fix(recording): preserve native evidence across fault boundaries

* fix(android): retire evidence after active publish rollback
2026-08-12 12:11:37 +02:00
Michał Pierzchała eabc936a0f refactor: route apps through request runtime (#1756)
* refactor: route apps through request runtime

* test: remove stale apps adapter mock

* fix: clean apps runtime replay artifacts

* fix: remove stale runtime test exports

* refactor: simplify apps runtime admission

* test: exercise runtime use through facade

* fix: close apps runtime admission gaps

* fix: disambiguate doctor app inventory callback

* fix: keep capability fixture below complexity limit

* fix: keep HarmonyOS app inventory fail-closed

* test: type HarmonyOS app admission fixture

* fix: restore HarmonyOS app inventory parity

* test: align HarmonyOS readiness fixture

* fix: preserve HarmonyOS doctor app parity

* test: type HarmonyOS doctor fixture

* fix: isolate HarmonyOS doctor policy

* fix: align apps cutover with shared rule catalog
2026-08-12 11:55:09 +02:00
Michał Pierzchała b13c06338e refactor: route boot through readiness runtime (#1747)
* refactor: route boot through readiness runtime

* fix: separate boot admission from readiness

* fix: register boot cutover policy

* refactor(runtime): move readiness into platform owners

* fix(test): tolerate provider temp cleanup races
2026-08-12 11:55:09 +02:00
Michał Pierzchała 97c87eec4d refactor(registry): exhaustive platformExecution discriminator (ADR 0019 §6) (#1740)
* refactor(registry): make the platform-execution discriminator exhaustive

ADR 0019 §6 (amended): every command descriptor declares its platform-execution
mode explicitly. Adds the `none` mode to `CommandPlatformExecution`, removes the
silent `{ kind: 'legacy' }` default at registry entry, and annotates all 76
descriptors so the migration denominator is machine-readable.

Part of #1739 (wave 0)

* fix(registry): react-devtools executes delegated platform behavior

`react-devtools start` on a Limrun Android instance dispatches internal
`runtime port-reverse`, which reaches a provider device runtime, so ADR 0019 §6
`none` is false for it. Reclassify as `legacy` and add the derived coherence gate
that catches delegated platform execution: if a CLI route for command R
dispatches command D, R may declare `none` only when D is `none`.

Part of #1739 (wave 0)

* fix(registry): attribute CLI dispatches by occurrence, not command name

Subtracting attributed command NAMES let a stray dispatch hide behind a routed
one that names the same command, so the gate's totality claim did not hold.
Dispatch sites now carry their source offset and attribution subtracts
occurrences.

Part of #1739 (wave 0)

* fix(registry): unresolvable CLI daemon-send targets fail the gate

An unknown literal or computed command target resolved to undefined and never
entered the scan, so a dispatch could evade attribution by naming a target the
gate could not read. Daemon-send envelopes are now located by their send call and
an unresolvable target is reported instead of skipped.

Part of #1739 (wave 0)

* refactor(cli): own injected daemon dispatches at a typed construction seam

The syntactic scan recognized only a direct sendToDaemon call whose first
argument was an inline object literal, so a variable envelope or a computed
callee was omitted from every result. Rather than teach the scanner more shapes,
the CLI's injected dispatches now flow through one typed construction point whose
route/command pairs are declared, and the gate reads that declaration instead of
recovering it from syntax.

Part of #1739 (wave 0)

* fix: constrain injected daemon transport handoffs
2026-08-12 10:40:42 +02:00
Michał Pierzchała 74eab2a554 refactor: route selector-resolution structural stages into typed policy (#1744)
* refactor: route selector structural stages into typed policy

#1649 landed the per-caller ambiguity matrix and deliberately left four
structural columns out: occlusion, off-screen, hittable-ancestor promotion,
and the poll budget were per-caller pipeline code, so declaring them would
have been an unverifiable claim (nothing consumed them; flipping one left the
suite green).

This adds the missing half as a table with runners. `SELECTOR_PIPELINE_POLICIES`
(src/core/selector-pipeline-policy.ts) gives each caller ONE row naming its
ambiguity contract plus its four stages, and every stage is reached only
through a runner that reads the row:

- occlusion -> selectorPipelineCandidates (candidacy) and
  resolveSelectorPipelineTarget (refusal). Acting rows exclude covered nodes
  and refuse covered targets; `find` and the diagnosis probe keep them as
  candidates and refuse at the target; reads and `wait` ignore them.
- promotion -> resolveSelectorPipelineTarget. The per-call-site
  `promoteToHittableAncestor: boolean` is gone: click/press/longpress name
  `promotedTarget`, fill/focus/scroll/drag endpoints and the native-ref
  preflight name `resolvedTarget`. `find`'s below-the-root variant is a
  declared value rather than a second local helper.
- off-screen -> throwIfOffscreenInteractionTarget, which now takes the row and
  returns the node untouched (no iOS rescue round trip) for observation rows.
- poll -> selectorPollBudget, which createWaitPolling derives its deadline and
  inter-poll delay from; the two wait loops carry a budget, every other row
  carries none and cannot be polled.

Behavior is byte-identical. The acting refusal keeps its exact node, label and
details in every branch (promotion declines to retarget away from a covered
node, so the "both covered" case names the same node it always did), and
`find` carries the occlusion verdict to the focus/type seam rather than
raising it early, because find click/fill still delegate that refusal to the
interaction leaf's own error shape.

selector-pipeline-policy.test.ts drives EVERY row through EVERY runner,
including the rows whose answer is "skip" — the half that used to be an
absence of code, and an absence cannot fail. Each stage was proven red by
flipping its cell (occlusion, promotion, off-screen, poll, plus the
declare-only-what-is-enforced guard). The ADR 0011 occlusion/nonHittable
`via` pointers for the runtime tree paths now name the runner that makes the
decision, not the predicate it applies.

Closes #1656; prework for #1739 (waves 4-5).

* docs: state constraints instead of narrating the refactor

Comment pass over #1656: drop the "used to be per-caller code" /
"not module constants" / "rather than an omission" narration — a comment
should say what a future edit must respect, not what the previous shape was —
and compress the find occlusion-verdict and poll-budget notes to the
constraint they actually carry.

* refactor: make the selector pipeline the only door to the engine

Review of #1744: the structural rows were declared but bypassable. Read and
wait routes composed `selectorPipelineCandidates(row, nodes)` with the raw
`resolveSelectorChainWithPolicy(..., row.resolution)` and never entered the
promotion or off-screen stages, so flipping a read row's `promotion` or
`offscreen` changed only the policy unit tests — production `get`/`is`/`wait`
were unaffected, which is the unverifiable-column failure #1656 exists to
remove. Callers could also pair one row's candidate set with another row's
ambiguity contract, and `find list` reached the engine directly.

The owning interface (src/core/selector-pipeline.ts) now runs every stage a row
declares, skips included, and the stage functions are private to it:

- `resolveSelectorPipeline` — single-target rows: candidacy, ambiguity, the
  replay-guard hook, promotion, occlusion, off-screen.
- `listSelectorPipelineMatches` — `reject-candidates` rows, returning the
  candidate set AND the tree the row sees, so ranking and equivalence
  classification judge the same nodes candidacy produced.
- `runNodePipelineStages` — the node stages for a target from a non-chain
  matcher (`@ref`, find's fuzzy locator) or a narrowed candidate set.

A row whose off-screen stage refuses must supply a refusal shape, so flipping
an observation row to `refuse` fails on its real route instead of silently
observing. `find list` now names a `readList` row (the new
`reject-candidates`/no-rect ambiguity row) instead of calling the engine.

R17 selector-pipeline-ownership (scripts/layering/) makes the bypass
structurally inexpressible: only the owner may import the engine entry points.
Proven against a planted import in selector-read.ts, which the repo-wide scan
rejects with the entry points that replace it.

Flips now fail through REAL command routes, verified one at a time:
readUnique.occlusion/offscreen/promotion and wait.occlusion via get attrs / is
/ wait; readAny.offscreen via is exists and find; readList.occlusion via find
list; promotedTarget.promotion via runtime click. The wait route test needed an
advancing clock first — with the frozen one a refused wait spun instead of
failing, so the flip hung rather than asserting.

* refactor: drop find's dead candidate binding

The selector branch bound the row's candidate set and never read it: only the
acting classification needs that tree, and find's locator branch brings its own
matcher. Names what actually governs the locator target — the shared node
stages below, not a candidate set it never had.

* refactor: reserve the selector engine behind the pipeline owner

Review of #1744 (three blockers).

**Listing rows no longer claim stages they cannot run.** `find <q> list`
resolves to a candidate SET, so promotion, the off-screen guard and a poll
budget have nothing to apply to — a listing has no single element to retarget,
keep on screen, or wait for. `readList` now declares only the two stages a
listing executes (`SelectorListPolicy`: resolution + occlusion), and the
narrower shape is load-bearing: `runNodePipelineStages` and `selectorPollBudget`
take the full row, so handing them a listing row is a compile error rather than
a silently skipped stage. Pinned with `@ts-expect-error` — widening `readList`
makes the directives unused and fails the typecheck.

**The engine door is a specifier, not a symbol.** R17's regex could not see a
namespace import, a re-export, or a deferred `import()`, none of which mention
the symbol it matched. The two engine entries moved to
`@agent-device/selectors/engine`, and R19 enforces over the resolved import
graph, where every one of those forms is the same edge. Proven on the
repo-wide scan by planting each form into a shipped route: namespace import,
dynamic import, and `export *` laundering all come back red.

`resolveImportEdges` drops an edge whose specifier resolves to nothing, so a
specifier rule goes quiet — not red — if the subpath is ever retired. The gate
now says that out loud instead of scanning clean.

**R19, not R17.** #1750 allocates R17/R18. Verified free against origin/main
and that PR's diff, then validated by real merges in both directions: the
uniqueness gate passes either way and the three ids stay distinct.

The gate itself is new (`scripts/layering/rule-ids.ts`): two branches taking one
free number do not conflict in git, so nothing caught R17 twice. Matching whole
string literals is what separates a declaration from prose that names a rule,
and it is what let the gate see #1750's `const RULE = '…'` shape — the first
version missed it and would have been vacuous. `main`'s two pre-existing
collisions (R11, R13) are listed as known, not pinned by equality, so #1750
lands in either order without breaking this.

Also: the root façade now exposes no resolver at all, and its surface test
pins both doors.

* fix(layering): make each rule-id allowance expire with its collision

Review of #1744: `KNOWN_RULE_ID_COLLISIONS` filtered the exact R11/R13
collision strings, so once #1750 renames those rules apart the entries would
keep waving those very collisions through if anyone reintroduced them. "Inert"
was wrong — a stale allowance fails open, permanently.

`ruleIdCollisionFailures` now checks the transition from both sides: a
collision nobody allowed fails, AND an allowance whose collision is absent
from the scan fails as a stale allowance. The entry therefore has to be deleted
in the same change that removes the collision, and the list burns down to
empty, which admits nothing.

#1750 is still open, so the transitional entries stay for now (option (b)).
Verified against a scratch tree carrying that PR's rename: leaving the list
untouched reports both entries as stale; deleting them is clean; and
reintroducing `R11 names contracts-implementation-authority and
package-boundaries` afterwards is rejected. The last of those is also a unit
regression, so the post-transition guarantee is pinned rather than argued.
2026-08-12 07:57:08 +02:00
Michał Pierzchała f18f8b076f refactor(contracts): consolidate per-domain defineUse wrappers into one neutral defineUse (#1741)
* refactor(contracts): consolidate per-domain defineUse wrappers into one neutral defineUse

ADR 0019 §9: use declarations share one neutral defineUse; per-domain
currying wrappers around runtimeUse<PlatformRuntimeOperations>() add a
module per domain for no information. network-runtime-plan.ts,
logs-runtime-plan.ts, screen-recording-runtime-plan.ts, and
app-log-resource-recovery.ts each re-derived their own curried alias
(defineNetworkUse, appLogUse, defineScreenRecordingUse, and an inline
instantiation) from the same generic factory with the same type
parameter.

Export defineUse = runtimeUse<PlatformRuntimeOperations>() once from
platform-runtime.ts (where runtimeUse lives) and re-export it from the
platform facade. Every runtime-use declaration (networkDumpUse,
networkAdmissionUse, the app-log uses, the screen-recording uses, and
appLogRecoveryUse) now builds through that single export; the
per-domain wrappers are deleted.

Type-level only: no required/preferred keys changed, and no use
declaration was added, removed, or altered. Existing deepEqual
assertions in network-runtime-plan.test.ts, logs-runtime-plan.test.ts,
and screen-recording-runtime-plan.test.ts already pin every produced
use object's exact {required, preferred} shape, so they double as the
before/after regression proof that this refactor is behavior-neutral.

Part of #1739 (wave 0).

* fix(contracts): move defineUse into platform-runtime-operations.ts

Review: defining defineUse in platform-runtime.ts required importing the
concrete PlatformRuntimeOperations catalog into the generic runtimeUse
primitive module, while platform-runtime-operations.ts already imports
generic runtime types from platform-runtime.ts. That's an avoidable reverse
type dependency — the lower generic primitive depended on its concrete
aggregate catalog. The unchanged SCC file count didn't prove this harmless;
it counts cycle members, not newly introduced back-edges.

Move defineUse = runtimeUse<PlatformRuntimeOperations>() into
platform-runtime-operations.ts, alongside PlatformRuntimeOperations.
platform-runtime.ts no longer imports the concrete catalog. Re-export
defineUse through the platform facade from its new source module; every
call site keeps importing it from @agent-device/contracts/platform
unchanged, and the three contracts-internal call sites now import it
directly from platform-runtime-operations.ts.

Validation: tsc (full workspace + examples/sdk), check:layering (136/136,
type-cycle count unchanged at 46), and check:affected --run (473 files /
3939 tests) all clean at the new head.
2026-08-11 18:05:32 +02:00
Michał Pierzchała f5d9789764 feat: enforce local device claims and reconcile stale owners (#1735)
* feat: enforce local device claims

* fix: address device claim review feedback

* fix: persist canonical daemon claim state directory
2026-08-11 16:18:45 +02:00
Michał Pierzchała 602b7a2995 refactor: narrow perf API to actionable evidence (#1731)
* refactor: narrow perf API to actionable evidence

* fix: address perf API review feedback

* fix: preserve deprecated Android CPU metrics
2026-08-11 15:28:58 +02:00
Michał Pierzchała b8dd6a5854 refactor: tighten capture ownership boundaries (#1736) 2026-08-11 13:55:28 +02:00
Michał Pierzchała 1b2e786128 refactor: move screen recording onto platform runtime (#1724) 2026-08-11 10:24:57 +02:00
Michał Pierzchała 1b75e102d7 perf: speed up device inventory and status (#1723) 2026-08-11 07:37:50 +02:00
Michał Pierzchała 338aa2a0d5 refactor: route every native selector resolution through the policy interface (#1715)
* refactor: route every native selector resolution through the policy interface

#1649 declared the per-caller ambiguity matrix; four native call sites still
bypassed it, spreading `selectorResolutionKnobs(row)` into a raw
`resolveSelectorChain` instead of naming the row. That left the "one
interface" claim aspirational: a caller could restate its contract as engine
knobs and nothing would notice.

- `is` non-exists, `get text`/`get attrs`, find's read actions, and the
  covered-selector diagnosis probe now call `resolveSelectorChainWithPolicy`
  with their existing row. Semantics are byte-identical: the knob-backed
  branch of that interface forwards to the same engine call the call sites
  built by hand.
- The façade drops `resolveSelectorChain` and `selectorResolutionKnobs`, so
  no knob-taking resolver is reachable from outside the package and a call
  site cannot re-acquire the knobs even by accident.
  `requireUnique`/`disambiguateAmbiguous` are now named in exactly one
  function, which `resolve-with-policy.ts` and the replay resolver both
  derive through.
- `get` names the two rows it may consume as a type, so pointing it at any
  other ambiguity contract is a compile error.

Tests: selector-read-policy.test.ts pins which row each read command
consumes, end to end, on one ambiguous fixture — the only tree the rows
disagree on. Each assertion was proven red by re-pointing its caller at a
neighbouring row. The knob-consistency check moves into the package beside
the now-private helper. Test call sites that used the raw resolver move to
`resolveRecordedTarget`, the same knobs and the path that actually replays a
recorded chain.

Extracting the failure branch drops `resolveSelectorInteractionTarget` below
the complexity threshold; its `fallow-ignore` waiver is removed (verified
load-bearing before the extraction, unnecessary after).

Closes #1630. Structural stages (occlusion, off-screen, promotion, poll
budget) stay per-caller pipeline code, tracked in #1656.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD

* test: observe which node find's row selected, not just that one existed

#1715 review, P2: the find row assertion was only half a pin. `find exists`
returns `found: true` for any resolved node, and the `list` call it leaned on
goes through listFindMatches — a path that consumes no policy row at all. So
repointing findFirstLocatorMatch at `readText` left both assertions green
while selection silently moved from the document-order head to the tiebreak
winner.

Assert through `find get_attrs`, which returns the ref of the node the row
actually selected. Both neighbouring rows are now red: `readText` fails
'@e3' !== '@e2' (the move the old test missed), `readUnique` fails by
refusing the ambiguous screen. `exists` stays as a second, weaker assertion
on the same resolution.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD

* refactor: route is exists through the matrix, collapse the double match pass

Follow-up tightening on the same seam.

`is exists` reached findSelectorChainMatch directly while the `readAny` row's
own doc claimed to serve "`exists` and find's read-only actions" — true of the
docs, not of the code, which is the unverifiable-claim shape #1656's review
called out. It now names `readAny`, the row it always described. Equivalent by
construction: both take the first alternative with any match under
requireRect: false, and disclose that alternative's count.

That leaves the root façade with no consumer for findSelectorChainMatch, so it
goes the way of resolveSelectorChain — dropped from the string-only façade,
kept on the published ./ast surface. Its façade-twin type SelectorChainMatch
dies with it (fallow caught it).

resolveSelectorChainWithPolicy matched twice on the uniqueness path: once via
resolveSelectorChain, then again to fill matchedNodes. Hoisting the single
list call above the row switch removes that second pass, collapses two
duplicated ambiguous literals into one helper, and drops a `?? [resolution.node]`
fallback that was unreachable — a resolution implies its alternative matched,
so the list is never null there.

While hoisting: the resolved arm's matchedNodes can describe a different
alternative than resolution.selector, because uniqueness skips an ambiguous
alternative to try the next one. Unreachable today (only first-match callers
read it, where both come from one list), and left as-is rather than silently
changed — but the doc claimed "the alternative it came from", so it now says
what is actually true.

Tests: is exists gets a caller-level pin on the shared ambiguous fixture —
passes with matches: 2 where its fail-closed siblings refuse — proven red by
pointing it at readUnique.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD

* test: discriminate is exists's row by alternative, guard the façade structurally

#1715 review, second regression-validity gap. The `is exists` pin observed
only `pass: true` and `matches: 2` on a fixture whose first alternative was
merely TIEBREAKABLE — so disambiguation succeeded there and reported the same
count first-match would. `readAny`, `readText`, and the pre-migration raw
lookup all produced that, and only the readUnique swap I had checked went
red. One mutation proven is not the same as the row being pinned.

`exists` exposes no node ref, so the row has to be read off WHICH alternative
answered. New fixture: alternative one matches two nodes that are genuinely
indistinguishable (same depth, same area, both on screen) so the tiebreak
declines; alternative two matches exactly one. First-match answers from
alternative one; every uniqueness row skips the undecidable alternative and
answers from alternative two. Asserting the selector now separates them —
readText and readUnique both fail with `id="save-unique"` where
`label="Save"` is expected.

Restoring the raw lookup stays behaviourally invisible, though:
findSelectorChainMatch is equivalent to the readAny row it migrated to, which
is precisely why that migration preserved semantics. No fixture assertion can
catch that revert, so the guard is structural — the façade's export list must
not carry resolveSelectorChain, findSelectorChainMatch, or
selectorResolutionKnobs. Follows the packages/maestro index.test.ts
absence-assertion precedent. Verified red by re-exporting the lookup.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD

* fix: cover selector routes in device replays

* test: simplify selector replay regression

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-11 07:34:11 +02:00
Michał Pierzchała cdc754e6ed perf: speed up iOS agent recovery and streamline CLI guidance (#1700)
* Avoid interactive children in parent taps

* docs: streamline no-skill CLI help

* perf: recover faster from sparse iOS trees

* fix: preserve selector context for blocked parent taps

* fix: preserve coordinate text-entry focus

* fix: preserve thin parent touch targets

* fix: fail closed for unscoped iOS typing

* test: isolate replay lock fixture

* test: share node integration process
2026-08-10 20:43:01 +02:00
Michał Pierzchała b15c502318 refactor: extract platform network runtime (#1702)
* refactor: extract platform network runtime

* fix: preserve platform network recovery routes

* test: guard network parser placement
2026-08-10 17:58:42 +02:00
Michał Pierzchała b1ed5353d1 refactor: extract platform log runtime (#1701)
* refactor: extract platform log runtime

* fix: clear terminal app log recovery markers

* fix: preserve scoped app log tooling

* fix: preserve app log cancellation

* fix: handle large changed coverage diffs

* fix: harden Limrun runtime identity

* refactor: tighten platform log runtime

* fix: close app log trust gaps

* fix: accept canonical session path aliases

* refactor: extract durable capture kit

* fix: refresh retained log marker admission

* fix: rotate app logs after process relaunch
2026-08-10 17:58:42 +02:00
Michał Pierzchała 13cc90ffc6 fix: harden Android snapshot and fill reliability (#1708)
* fix: harden Android automation reliability

* test: isolate CLI flush integration

* test: close Android review gaps

* test: register CLI transport fixture

* test: consolidate CLI subprocess fixture
2026-08-10 13:10:18 +02:00
Michał Pierzchała c06bed9f77 refactor: extract platform device inventory runtime (#1699)
* refactor: extract platform inventory runtime

* fix: preserve scoped Apple inventory tooling

* fix: preserve Apple tool cancellation

* refactor: tighten platform inventory boundaries
2026-08-10 12:51:59 +02:00
vw2x 4b432fb59b feat: add HarmonyOS support (#1683)
* feat: add HarmonyOS device automation foundation

Add HDC-backed discovery, snapshots, application lifecycle, and core mobile interactions.

Route HarmonyOS through the platform registry and client contracts.

Cover parsing and capability parity with focused tests.

* feat: support HarmonyOS HAP deployment

Install and reinstall signed HAP archives through HDC.

Resolve bundle identities from module metadata and relaunch after package replacement.

Extend deploy routing and capability coverage for HarmonyOS.

* feat: add HarmonyOS single-pointer gestures

Execute pan, fling, and swipe plans through HDC uiInput primitives.

Derive scroll coordinates from the live ArkUI viewport.

Keep unsupported multi-touch gestures explicitly rejected.

* refactor: split session inventory command handling

Separate session, device, capability, and app inventory response paths.

Preserve the public inventory response contract while reducing handler complexity.

* feat: support HarmonyOS keyboard actions

Route HarmonyOS enter, return, and dismiss through HDC key events.

Expose supported keyboard actions through the system command metadata.

Keep keyboard visibility inspection explicitly unsupported.

* fix: reject unsupported HarmonyOS drag gestures

Keep drag unavailable until HDC can preserve source and destination hold semantics.

* feat: add HarmonyOS app log streaming

1. Stream HarmonyOS app logs through PID-scoped hilog sessions.\n2. Record HarmonyOS app identity during bundle-id opens for app-scoped commands.\n3. Cover backend routing and bundle identity resolution.

* feat: report HarmonyOS foreground app state

1. Read the foreground HarmonyOS mission through aa dump.\n2. Expose HarmonyOS appstate with package and ability metadata.\n3. Add parser coverage for foreground and missing-state cases.

* fix: advertise appstate through capabilities

1. Classify appstate in the command descriptor capability matrix.\n2. Surface supported appstate commands in capability inventory.\n3. Cover the advertised Android capability contract.

* feat: sample HarmonyOS process performance

1. Sample HarmonyOS process CPU and resident memory through HDC.\n2. Expose the verified metrics through the shared perf command.\n3. Keep frame and memory snapshot collection explicitly unavailable.

* feat: clear HarmonyOS app state

1. Add HarmonyOS settings clear-app-state through bundle cleanup.\n2. Force stop the app before clearing data and cache.\n3. Reject all unverified HarmonyOS settings explicitly.

* docs: document HarmonyOS support

1. Describe HarmonyOS HDC prerequisites and HAP installation.\n2. Add HarmonyOS to platform discovery and product documentation.\n3. Document verified performance limits for the public HDC surface.

* fix: preserve HarmonyOS deploy session identity

1. Bind a resolved HarmonyOS bundle after install or reinstall.\n2. Keep app-scoped logs and observability available after deployment.\n3. Cover session identity preservation for HarmonyOS reinstall.

* test: lock HarmonyOS capability boundary

1. Add an independent HarmonyOS capability-matrix oracle and exact advertised-command regression test.
2. Document current HDC-backed support and evidence-based unsupported command boundaries.

* refactor: simplify HarmonyOS shared platform boundaries

1. Split device selection and settings dispatch into focused helpers without changing behavior.
2. Keep HarmonyOS serial selection and lock-policy classification covered by regression tests.
3. Remove Fallow complexity findings from the HarmonyOS diff against upstream main.

* fix: bound default HarmonyOS HDC commands

1. Apply a 15 second timeout to ordinary HDC operations.
2. Preserve operation-specific timeout budgets for installation and capture paths.
3. Add regression coverage for default and overridden HDC timeouts.

* feat: add HarmonyOS screen recording

Implement physical-device whole-screen recording through the system recorder and HDC media transfer.

Reject unsupported HarmonyOS recording scopes and export flags.

Cover capability routing, media retrieval, cleanup, and simulator rejection.

* feat: report HarmonyOS HDC readiness

Add an HDC version check to the HarmonyOS doctor flow.

Document HarmonyOS as a supported doctor platform and cover the result.

* refactor: simplify HarmonyOS recording checks

Reduce recording validation and test complexity without changing behavior.

* test: cover HarmonyOS platform contracts

Synchronize public platform expectations across CLI, MCP, replay, and inventory tests.

Mock HarmonyOS inventory probes to preserve concurrent test behavior.

* test: model HarmonyOS recording capability

Require a physical HarmonyOS device in the independent capability parity oracle.

* test: cover HarmonyOS input and lifecycle paths

Exercise HDC input, lifecycle, installation, and relaunch command sequences.

* test: cover HarmonyOS device observability paths

Exercise discovery, screenshot validation, and process performance sampling.

* docs: define HarmonyOS CI hardware policy

Keep HDC hardware validation local and require mocked CI contract tests.

* fix: honor HarmonyOS app inventory filters

* fix: bound HarmonyOS app inventory classification

1. 限制应用元数据分类并发并为默认清单设置整体时限.
2. 将请求取消信号传递给 HarmonyOS 应用清单读取.
3. 补充失败时中止在飞读取且不继续排队的回归测试.

* fix: preserve HarmonyOS inventory failure causes

1. 保留触发应用元数据分类失败的原始错误, 避免被取消同级任务覆盖.
2. 补充总时限中止在飞读取且不启动排队任务的回归测试.
3. 验证后序任务失败时保留默认筛选的恢复提示.
2026-08-09 10:29:20 +02:00
Michał Pierzchała 18291ba8e2 perf: collapse app-driving startup turns (#1693)
* perf: collapse app-driving startup turns

* fix: align foreground open guidance
2026-08-09 10:10:11 +02:00
Michał Pierzchała 6c0fcb64a1 fix: reject distinct ambiguous mutation targets (#1667)
* fix: reject distinct ambiguous mutation targets

* fix(ios): scope the raw-match rejection to mutating dispatches

`RunnerTests+Interaction.findElement` applied the new fail-closed
classification to `querySelector` as well as press/type, because the read
call site takes the default `allowNonHittableFallback: false`. With one
visible/hittable match and one non-hittable same-selector duplicate the
query started returning AMBIGUOUS_MATCH where it previously selected the
hittable element, and `queryDirectIosSelectorOrFallback` preserves that
error for read callers — so `get`, `is`, and `wait` surfaced an error
instead of their prior answer.

`classifyDirectSelectorCandidates` now takes a `rawMatchPolicy`. Mutations
keep `.rejectDistinctMatches` (the default, so no mutation call site
changes); `queryElement` passes `.preferHittableMatch`, restoring the
prior read rule: prefer the single hittable match, ambiguous only when
hittable matches compete, and never adopt the Maestro coordinate fallback.
The Maestro expected-point path is untouched.

Covers the one-hittable + one-non-hittable read, competing hittable reads,
and the non-hittable-only read. ADR 0011's amendment now states the scope.

* test(ios): execute selector read ambiguity regression

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-08 08:57:32 +02:00
Michał Pierzchała 4b89482a06 fix(daemon): open --foreground P1 hotfix — selector rejection, interactive snapshot, capture-failure masking (#1671)
* fix(daemon): open --foreground P1 hotfix — selector rejection, interactive snapshot, capture-failure masking

Three P1s from the post-merge review of #1670, all at the
session-open-foreground dispatch seam:

1. Explicit device selectors were silently overwritten: the resolved-device
   rewrite pinned --udid/--platform over whatever the caller passed, so
   `open --foreground --udid B` with sim A sole-booted silently opened A.
   Now fails fast with INVALID_ARGS (matching the existing app-positional
   rejection) on --udid/--device, and on --platform other than ios; an
   explicit --platform ios passes through.

2. The promised interactive snapshot was never requested: the composed
   snapshot dispatch forwarded the open request's flags untouched, without
   snapshotInteractiveOnly — so the capture was NOT the `snapshot -i` path
   the doc comment promised and returned no interactive presentation. The
   composed request now sets snapshotInteractiveOnly: true (the exact key
   the CLI maps -i to and the snapshot runtime reads as interactiveOnly).

3. A capture failure masked the successful open: returning the snapshot
   error discarded openResponse even though the session exists, so a retry
   of `open --foreground` failed with "close the current session first".
   Open success + snapshot failure now returns ok with an explicit
   initialSnapshotError {code, message} detail and a rendered warning that
   the session IS open and how to capture manually (snapshot -i).

Regressions added for all three: explicit-selector rejection
(udid/device/both/non-iOS platform + ios pass-through), the composed
dispatch carrying snapshotInteractiveOnly, and the snapshot-failure path
returning ok + warning + usable session.

* fix(cli): render the composed open --foreground snapshot on default stdout and project initialSnapshotError through the public surfaces

Post-merge review on #1671 found the daemon fixes never reached the public
boundaries: openCliOutput ignored the nested snapshot (the one-call promise
held only under --json), and initialSnapshotError was daemon-only — absent
from AppOpenResult, Node normalization, and serializeOpenResult, with the
normalized shape truncated to code+message.

- default open output now renders the composed interactive tree through the
  same snapshotCliOutput path snapshot -i uses
- AppOpenResult carries initialSnapshotError as the FULL daemon error
  (hint/details/diagnosticId/logPath preserved) through normalization and
  serialization

Worker-authored; committed by the coordinating session after the worker
stalled twice mid-push. Tests: output.test.ts + session-open-foreground
(26 pass), typecheck, oxfmt.

* refactor(client): one daemon-error normalizer + client-route regressions for initialSnapshotError

Review follow-ups on #1671: normalizeInitialSnapshotError duplicated the
target-shutdown error normalization and pushed the module over the fallow
complexity threshold; consolidated into a single internal normalizeDaemonError
(table-driven, full shape incl. retriable/supportedOn) used by both result
paths, projected through normalizeOpenForegroundComposition.

Client-route regressions: createAgentDeviceClient().apps.open now proves the
full initialSnapshotError shape (hint/details/diagnosticId/logPath/retriable)
survives normalization, and that a malformed one is dropped — deleting the
boundary normalization fails these tests.

Also rebased onto current main.

* fix(daemon): a thrown initial-snapshot capture failure gets the same successful-open contract

Review P1 on #1671: dispatchSnapshotViaRuntime rethrows ordinary
capture/runner exceptions; the composition only handled a returned
{ ok: false }, so a thrown failure escaped to the router and failed the
whole open after the session was created — retrying then wedged on the
existing session. The catch normalizes the rejection (kernel
normalizeError, same conversion the router applies) into the shared
openWithInitialSnapshotFailure path: ok response, full-shape
initialSnapshotError, session-usable warning. Rejecting-mock regression
added alongside the returned-failure case.
2026-08-08 07:54:33 +02:00
Michał Pierzchała 13bc70f24f refactor(find): resolve a mutating find's target once, not twice (#1654) (#1669)
* fix(find): resolve a mutating find's target once, not twice (#1654)

A mutating `find click`/`find fill` captured the screen, matched by
locator under the `findAct` policy, promoted to a hittable ancestor, minted
`@eN` off the node it chose — and then re-entered the interaction leaf by
bare `@ref`, which looked that ref up AGAIN via resolveSnapshotForRef. The
second lookup reads the SESSION frame tree, not the fresh capture find
matched against, so the two could disagree: it could hand the action a
different node than find picked, or refuse as unresolvable a ref find had
observed a moment earlier.

find now passes the node itself. `internal.findPreresolvedTarget` carries
the resolved node and its tree over the in-process invoke hop, and the ref
branch adopts them instead of resolving the ref a second time.

What is NOT skipped: the guarantees. Occlusion, hittable-ancestor
promotion, and the off-screen guard all still run, on that node, at the
same symbols the ADR 0011 `runtime-ref` cells name. Only the LOOKUP is
replaced. Recording, ref-frame effects, settle/observation, and
deferred-outcome marking are untouched — they live in the dispatch wrapper,
not in resolution, which is why the invoke hop is kept rather than bypassed
the way find focus/type bypass it.

The channel is a second field rather than widening `findResolvedTarget`
because they answer different questions: that flag governs ref-frame
admission and staleness (policy), this governs resolution (lookup), and
focus/type set neither. It is `internal`-only, so it never crosses the
wire and carrying live node references is sound.

ADR 0011 re-check: the `runtime-ref` disclosure cell gains a second
producer, adoptPreresolvedRefTarget, reporting `exact`. Truthful rather
than borrowed — the ref is minted off the node handed over. It cannot
report label-fallback, which is right: label recovery is a property of
looking a stale ref up, and this path performs no lookup.

Tests pin the behavior by tap coordinates, so they name which tree the leaf
resolved against, plus a control proving the ordinary @ref path is
unchanged and a case where the session tree cannot resolve the ref at all —
that one can only pass if no second lookup runs. Verified revert-sensitive:
removing the short-circuit fails 3 of the 4, and the control stays green.

Out of scope, unchanged: the ADR 0011 `native-ref` path (web provider
clickRef/fillRef only), where the ref is the provider's own element handle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117wfrvC6MRDJUWdRBErNEF

* test(find): prove the production route, and correct the divergence claim (#1654 review P2)

The regression tests built `internal.findPreresolvedTarget` by hand and
called handleInteractionCommands directly, so they proved the leaf consumes
the channel but never that find ATTACHES it. Either producer could have
stopped doing so and they would all have stayed green — the exact gap
#1649's review named, reappearing one layer up.

Four tests now drive the real handleFindCommands with invoke wired to the
real handleInteractionCommands. Two assert click and fill each carry the
selected node (separate producers, separate forwarding hops, so asserted
separately); two advance the session tree between find's match and the
leaf's resolution and assert the dispatched point is still find's node.

That last pair also corrects the record. The claim that the leaf resolves
against a different tree than find matched is too strong for the common
path: `omitRefFrameSnapshot` (interaction-runtime.ts) makes find's internal
dispatch skip the authorized frame tree and resolve against
`session.snapshot`, which find's own capture just wrote — so the second
lookup normally AGREES, and the two non-diverged tests above pass with or
without the fix.

The divergence is real but narrower: it needs something to advance
`session.snapshot` between find's match and the leaf's resolution, which
`refreshAndroidRefSnapshotIfFreshnessActive` does on the ref path. Measured
at this head, the old code taps (60,720) "Delete" where find matched
"Save" at (310,510) — a wrong-element mutation, now pinned by both click
and fill.

So the change's value is what #1654 asked for — one resolution end to end —
plus closing that window, not a fix for a divergence occurring on every
find.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117wfrvC6MRDJUWdRBErNEF

* fix(find): tighten resolved target provenance

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-07 17:45:43 +02:00
Michał Pierzchała 04c33f9e1e feat(ios): expose AX custom actions on merged accessibility elements (#1665)
* feat(ios): expose AX custom actions on merged accessibility elements

Apps that merge a card into one accessibility element for VoiceOver (React
Native's `accessible` prop) publish the card's real affordances as
UIAccessibilityCustomActions rather than as child elements. Our snapshot showed
only the merged node, so an agent looking for a feed card's options control had
nothing to aim at and fell back to coordinate guessing.

`snapshot --actions` now names them:

    @e8 [link] "feedItem-by-whiskers.test" actions: ["Reply", "Repost", "Open post options menu"]

Opt-in, because the AX server cannot serve custom actions in a bulk tree
request: adding the attribute makes testmanagerd's reply decoder reject the
nested arrays a custom action serializes into, drop the reply, and time the
request out (~65s vs ~110ms). Only per-element reads answer, at ~100ms each, so
the runner reads at most 12 labelled childless nodes and stops at the
capture-plan deadline. The request pins the private-AX backend, since no other
backend can read the attribute, and reports that as its own `requested-backend`
verdict so a deliberate pin never renders as a degradation warning.

Invocation is not shipped: the actions are readable but not invocable from the
runner. RunnerAXSnapshotBridge.h records the five APIs that were tried.

* fix(ios): disclose a capped custom-action pass, and read on-screen elements first

Two gaps in the first cut.

An element the bounded pass never reached rendered identically to one with no
custom actions, so a capped capture silently taught the reader that later feed
cards have no affordances — the exact mis-inference this feature exists to
prevent. The runner already counted reads against candidates; it now carries
both to the response, and the verdict renders one response-level line when the
pass was incomplete. A complete pass stays silent, and "never asked" stays
distinguishable from "read none" (the key is absent, not (0, 0)).

The obvious remedy for a capped pass would be a scoped re-run, but scope is
applied when the Swift walk builds nodes, long after the read pass, so it does
not redirect the budget at all — measured: reads=12 candidates=18 with and
without --scope. Rather than print a remedy that does nothing, the read pass now
orders candidates on-screen first. That makes the budget land on elements an
agent can act on, and makes the disclosed remedy true: scrolling changes the
on-screen set, so a re-run reads elements the previous pass could not.

Also states plainly, in the flag help and the tool/SDK field description, that
the names are for planning: nothing invokes them, so the affordance is reached
through the element's detail screen, the same control exposed elsewhere, or
coordinates.

* test(snapshot): pin the custom-action coverage pair in the verdict shape assertion

* chore(scripts): classify the --actions flag in the integration progress model

The completeness gate flagged snapshotCustomActions as unclassified, which is
what it is for. It gets its own bucket rather than joining the provider-scenario
table: the values come from the private AX client inside the runner process, and
the fake runner derives its behavior from fixture tables that cannot fabricate
custom actions, so there is no provider-backed scenario to claim. The owning
coverage is named instead — runner XCTest unit, snapshot-lines, snapshot-quality.

* fix(ios): fail closed on unserviceable --actions, bound each read, cap output, and count actions in identity

Four review findings.

1. `--actions` with `--raw`, or on any target that is not an iOS simulator, used
to succeed and return nodes with no actions — a requested capability silently
no-opped, indistinguishable from "this screen has none". Both now fail closed.
The raw pairing is rejected at the shared request seam (INVALID_ARGS) so CLI,
Node client and MCP answer alike before any device work; the platform case is
rejected once the session device is resolved (UNSUPPORTED_OPERATION), naming the
resolved target. `diff --actions` was already rejected as an unsupported flag.

2. The per-element AX read had no timeout, so one wedged element could consume
the whole capture budget. Each read now runs off-thread behind a 1s wait. A
timed-out element counts as unread, never as "read, and it has no actions", so
the existing partial-pass disclosure already covers it.

3. The element budget bounded element count only; one element could still return
an unbounded list of unbounded names. Capped at 8 names of 80 characters, and
clipped elements are counted into the coverage so a truncated list is disclosed
rather than silently presented as complete.

4. Action names were rendered unescaped, and no comparison key read them. Names
now get the same escaping as text previews plus control-character folding, so an
app-authored name cannot split or corrupt a line. `actions` joins the diff
comparable key, the unchanged-comparison projection, and — the sharper bug —
the snapshot presentation key, without which `snapshot` followed by
`snapshot --actions` on a still screen answered "unchanged" and never delivered
the actions that were explicitly requested.

* fix(ios): contain a hung custom-action read instead of accumulating orphans

The 1s read deadline frees the caller, but the underlying AX call is a
synchronous XPC round trip that cannot be cancelled — it keeps running. On a
global concurrent queue that meant repeated `snapshot --actions` against a
wedged element piled up orphaned reads, all using the shared XCAXClient
concurrently. The deadline was containment for the capture, not for the runner.

Since the call cannot be cancelled, contain it instead:

- every read runs on one dedicated serial queue, so a wedged call can never be
  joined by a second concurrent user of the shared client;
- a single-flight guard refuses to dispatch at all while an abandoned read is
  still outstanding, so a repeated capture adds no work — the dispatch counter
  stands still;
- the read pass stops at that point rather than paying a deadline per element
  on reads that would all be refused, and reports `blocked` so the capture stays
  honest. That gets its own line, because the partial-pass remedy (scroll and
  re-run) cannot clear a hang and would send the reader in circles.

Recovery needs no reset: when the hung call finally returns, in-flight drops to
zero and reads resume.

The regression drives a fake AX client that never returns, and asserts the three
things the fix exists for — exactly one in-flight read with no further
dispatches across repeated captures, immediate returns with the skip disclosed
instead of the scroll remedy, and reads working again once the wedge clears.

* ci(ios): execute the custom-action runner regressions instead of only compiling them

The iOS workflow runs a targeted -only-testing list, so a runner test that is
not named there is compiled by the build step and then never executed. All seven
custom-action tests were in that gap — including the containment regression,
which is the only executable proof that a hung AX read cannot accumulate
orphaned in-flight reads.

Red/green against the containment regression, with the fix reverted to its
pre-fix concurrency behavior (global concurrent queue, no single-flight guard,
no blocked exit):

  RED   in-flight 6 (want 1), dispatches 6 (want 1), each repeat paid the full
        1.004s deadline (want <0.2s), blocked=false (want true), and the
        in-flight drain never completed — "Exceeded timeout of 5 seconds".
  GREEN 7/7 pass, containment regression in 1.02s.
2026-08-07 17:35:48 +02:00
Michał Pierzchała e14c9d8d7b fix(cloud): unblock iOS snapshot and gate cloud fill on text-entry focus (#1658) (#1666)
* fix(cloud): unblock iOS snapshot and gate cloud fill on text-entry focus (#1658)

Two bugs isolated to the cloud-webdriver iOS path.

`snapshot`/`diff` refused every capture on a live BrowserStack session with
SESSION_NOT_FOUND, instantly and without a driver round trip. The app-session
guard they ran belongs to the local XCUITest runner, which must attach to a
target app; a cloud capture reads the provider's own driver session and needs
no app identity, so it now applies to local Apple targets only. The session
was empty in the first place because the provider open path skips local app
resolution wholesale — no simctl/devicectl reaches a hosted device — and
dropped an explicitly spelled bundle id along with it. A dotted, non-deep-link
target is the bundle id under the same convention resolveIosApp applies
locally, so a cloud `open com.example.app` now records it.

`fill` tapped and sent its keys in back-to-back requests. A WebView input —
an OAuth page in a Safari view controller — takes first responder
asynchronously, so the keys landed with nothing focused while the command
still answered "Filled N chars"; tapping and filling as two separate commands
worked only because the round trip between them gave the field time to focus.
The cloud interactor now waits on the same signal the Apple runner uses, the
software keyboard going from hidden to shown after its tap, and discloses what
it observed as `textEntryReadiness` so a fill with no witness cannot pass for
a filled field. Where keyboard visibility cannot witness the focus move —
back-to-back fills into one form, the shape that failed most often — it spends
the runner's full readiness budget rather than racing the app with a short
settle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Zjbzf7HdziX9SzFpNP7WW

* fix(cloud): let a new bundle-id open replace the tracked cloud iOS app

Adopting an explicitly spelled bundle id on a provider-backed open (the
fix that makes snapshot/diff work at all) also made a previously dead
precedence rule live: the provider branch returned currentAppBundleId
first, so once a first open had populated it, `open com.a` followed by
`open com.b` left the session still reporting com.a to every
appBundleId-gated command.

The local path does the opposite, and is the convention this branch is
meant to mirror: resolveIosApp returns a dotted target unchanged and
never consults the session's current app. Only its deep-link branches
prefer the tracked id. Flip the provider branch to match — an explicit
bundle-id target wins, and everything the branch cannot name (deep
links, display names, bare open) still falls back to the tracked id.

* fix(cloud): fail a witness-less cloud fill instead of reporting it filled

Review follow-ups on #1658.

`not-observed` was still a success: it sent the keys and answered "Filled N
chars", and nothing renders `textEntryReadiness` in default CLI output — so the
exact silent success this branch exists to remove survived whenever focus never
happened. A tap that raises no keyboard now fails with
`text_entry_focus_not_observed` and sends no keys, leaving the field untouched
rather than half-written, and the readiness vocabulary keeps only outcomes that
describe a fill that did type.

The readiness budget was advertised but not enforced at the request boundary:
each keyboard probe inherited the client's 30s default, so one hung probe could
hold a 2s wait for far longer. Probes now carry their own bound, threaded
through the client as a per-request timeout override.

The probe also swallowed every error as "this driver cannot answer", which
degraded a dead session, an auth rejection, or a grid outage into a blind text
entry. Only a positively classified unimplemented route counts as unsupported
now — classified on the W3C error code rather than the status, since `unknown
command` and `invalid session id` share HTTP 404 — and everything else
propagates.

The provider scenario proved request ordering against a stub that always
accepted keys. Its fake now models the device: focus lands a beat after the
tap, and keys arriving while the keyboard is down are accepted and dropped,
exactly as an unfocused field does. The tests assert the field's own value, and
both go red against the pre-fix `fill`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Zjbzf7HdziX9SzFpNP7WW

* fix(cloud): witness the tapped field's focus before a cloud fill types

Review of cc23f2b found two ways a fill could still report success
without evidence that OUR tap focused the field it was aimed at.

P1. `settled-keyboard-up` and `settled-unknown` both typed and returned
normal success. Keyboard visibility can only witness that *a* field took
focus, never *which*: filling a second field in an already-open form
reads the same before and after, so a missed tap left the first field
focused and `POST /keys` — which the driver routes to whatever holds
first responder — appended to it while every request returned 200.

Failing those closed outright would have broken ordinary multi-field
form fills, which do work: a live AWS Device Farm run types both fields
of a WebView login correctly. So witness focus properly instead. W3C
`GET /element/active` answers the question keyboard visibility cannot —
is the thing focused now the thing I tapped — and answers it whether or
not the keyboard was already up. That becomes the primary signal
(`focused-element`); the keyboard transition stays as the fallback for
drivers without the route, and a keyboard already up on such a driver
now refuses rather than typing.

The test is identity, not geometry. Containment of the tap point looks
like the obvious rule and is wrong: focusing a field can re-lay it out.
On a live iPhone 16, tapping Safari's collapsed address bar expands it
into a taller field that no longer covers the tapped point, and a
containment-only rule refused a fill that plainly worked. So a tap that
MOVES focus counts, with containment as the second half of the test —
re-filling the already-focused field moves nothing, and only geometry
tells that from a tap that missed. Both readings are taken before the
tap, since each is evidence only as a change.

P2. The 2s budget bounded the loop but not the calls inside it: every
probe got a fixed 1500ms, so one begun near the deadline finished well
past it. Both the probe timeout and the sleep are now capped by the
remaining budget.

Also fixes a related escape the review did not name: the poll loop had
no catch, so one transient grid error aborted a fill the next poll would
have satisfied. Probe failures are now tolerated within the budget, but
a budget that expires without a single answered probe rethrows, so a
dead session surfaces as itself rather than as "the tap missed".

The provider scenario gains the two-field case the review asked for: it
begins keyboard-up with the email field focused, misses the password
tap, and asserts no keys reach the email field.

Verified on AWS Device Farm iPhone 16 / iOS 18.0 at this exact tree:
address bar (the re-layout case) and both WebView login fields all
report `focused-element`, the second with the keyboard already up, and
the device reads back `tomsmith` and a 20-character password.

* fix(cloud): refuse a cloud fill no focus probe can witness, and bound the composite probe

Two blockers from the review of 3a9aceb9.

P1. `settled-unknown` was the last path that typed without evidence: when
both the active-element and keyboard routes are positively unsupported,
`fill` settled 350ms, typed, and returned ordinary success. Nothing
renders `textEntryReadiness`, so that reached a caller looking exactly
like a fill that worked — the same silent false success #1658 is about,
just narrowed to one branch. It now refuses with a distinct reason,
`text_entry_focus_unobservable`: nothing is wrong with the target, the
driver simply cannot answer, so the caller's next move differs from a
missed tap and the hint names it — `press` then `type` stays the
deliberate way to enter text unwitnessed.

`CLOUD_TEXT_ENTRY_READINESS` is now `focused-element` and
`keyboard-shown` only. Every value describes a fill that witnessed focus
before sending a key; there is deliberately no value for typing blind.

P2. `activeElement(timeoutMs)` bounded each of its two sequential
requests by the full timeout rather than bounding the operation, so a
probe handed the 1.5s left of a 2s readiness deadline could spend ~3s
across `/element/active` and `/element/{id}/rect` and overrun the
deadline it was derived from. It now derives one deadline at entry and
gives the second request only what the first left, floored at zero so an
already-spent budget aborts immediately instead of falling back to the
client default.

The regression pins elapsed transport time across both calls, which is
what the defect is made of: the rect request answers only its own abort,
so the time it was allowed to run IS the budget it was handed. It
measures ~202ms of a shared 200ms budget before the fix and ~120ms
after.

Also updates the generic Cloud WebDriver facade scenario, whose stub
answered `{value: null}` to everything and so read as a driver with
neither route. It now answers the two focus probes, since that scenario
exercises facade wiring rather than text-entry semantics — those live in
cloud-webdriver-ios-text-entry.test.ts, which models focus properly.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-07 15:52:24 +02:00
Michał Pierzchała cc943400a9 feat(daemon): [RFC] prototype foreground-attach convenience (#1670)
Prototype `open --foreground`: on a fresh session with no app argument,
auto-resolves the target from the sole booted iOS simulator's sole
foreground app (reusing the exact same ambiguity-detection probe that
enriches the SESSION_NOT_FOUND hint), then attaches the initial
interactive snapshot to the response by composing the existing
snapshot-runtime dispatch. Collapses the documented 3-call
snapshot-fails -> read-hint -> open -> snapshot-succeeds dance into a
single call for the unambiguous case, while failing closed
(AMBIGUOUS_MATCH) with no guessing otherwise.

First-pass RFC, not reviewed — see PR body for the design tradeoff
writeup, live before/after evidence, and scoped-out follow-ups.
2026-08-07 13:17:13 +02:00
Michał Pierzchała 10ff339d14 refactor: declare selector resolution policy as data (#1649)
* refactor: declare selector resolution policy as data (#1630)

Five native consumers of "resolve a selector against the screen" each
hand-declared their ambiguity contract as inline requireUnique/
disambiguateAmbiguous literals, so the repo's real policy matrix was only
discoverable by reading four files. SELECTOR_RESOLUTION_POLICIES
(packages/selectors) now declares one row per caller — ambiguity kind plus
the structural columns (rect, occlusion, off-screen guard, promotion, poll)
— and selectorResolutionKnobs turns a row into the engine knobs it stands
for. Callers consume rows; zero ambiguity literals remain in src.

Semantics are unchanged by construction: each row was read off its call
site. The matrix names what was previously implicit — act and get text
disambiguate, is/get attrs fail closed, exists/find-reads and wait take the
first match, mutating find rejects candidates unless narrowed (#1625).
`reject-candidates` is declaration-only and rejected by
selectorResolutionKnobs at the type level, because find enforces it through
its own narrowing rather than engine knobs.

resolution-policy-parity.test.ts gate-tests the matrix against the callers
(ADR 0011's declared-plus-gate-tested pattern): knobs must match the named
ambiguity contract, every claimed structural column must appear in the
caller's source, the read/wait pipelines must genuinely lack the machinery
they disclaim, and no caller may reintroduce an inline literal. Verified
revert-sensitive: flipping readUnique to disambiguate and faking wait's
occlusion column each fail it.

Out of scope, unchanged, per the issue: the Maestro engine (ADR 0015) and
the open click-implicit-wait product decision.

* refactor: route wait and mutating find through the policy interface (#1649 review)

P1 was right: the first head declared seven rows but genuinely routed five.
selector-wait.ts never imported its row (it called listSelectorChainMatches
directly), findAct consumed only requireRect while its ambiguity contract
stayed bespoke, and the parity test sniffed marker strings in source files —
so it stayed green across exactly that gap. Asserting about the layer I had
edited instead of the behavior it produces.

resolveSelectorChainWithPolicy is now the one policy-driven entry: it
returns a discriminated outcome (none / resolved / ambiguous) because the
rows genuinely disagree about what several matches mean, which is what
previously forced each caller to re-derive its contract inline. wait and
find's selector branch both route through it; find additionally asserts its
row still says reject-candidates rather than assuming.

The parity test is rebuilt on fixture trees driven through that interface —
no source sniffing. Wiring verified revert-sensitive: flipping the wait row
fails the policy tests, and flipping findAct fails REAL find handler tests
(ambiguous-candidate listing), which is the proof the previous version
could not produce.

One behavior nuance the fixture work surfaced and now pins: disambiguation
declines on genuinely indistinguishable candidates (the tiebreak is
evidence, not a coin flip), so an acting row surfaces ambiguity there rather
than binding one silently.

* fix(test): let fallow see the host-process mock helper's real consumers

Rebase onto main brought #1642's host-process-mock.ts into this PR's
fallow scope, where its export reports as unused. It is not: three suites
consume it, but only through `(await import(...)).pinOwnProcessStartTime`
inside vi.mock factories — vitest hoists those above static imports, so the
dynamic form is required and fallow cannot trace it statically. Documented
suppression rather than a restructure that would break the hoisting
contract.

Latent on main rather than introduced here: the audit gate is
changed-files-only, so main sees the file in scope only from a PR whose
diff contains it.

* fix: keep every candidate when a policy resolves one winner (#1649 review P1)

A real regression I introduced, not a test gap: routing wait through the
policy interface collapsed the candidate set to the winner, and the #1349
landmark check is satisfied when SOME match carries the recorded identity.
A first same-selector impostor therefore hid a later genuine landmark and
timed the wait out.

The resolved outcome now carries `matchedNodes` — the full candidate set of
the alternative the winner came from — so a policy that picks one node no
longer throws the rest away. wait passes that straight to the landmark
check, restoring the original semantics.

Regression test added at the within-one-poll shape the existing suite did
not cover (both candidates in the SAME capture, impostor first); verified
it goes red against the singleton reconstruction it replaces.

* refactor: declare only the policy fields the matrix enforces (#1649 review)

The occlusion / offscreenGuard / promotion / poll columns were never
consumed by resolveSelectorChainWithPolicy or selectorResolutionKnobs:
changing any of them left behavior and the suite green, so they were
unverifiable claims that read as truth. (My earlier source-sniffing test
"verified" them by grepping caller files for marker strings — which is why
it also stayed green when a row was disconnected entirely.)

The matrix now declares exactly what it enforces: the ambiguity contract and
the rect requirement, both consumed by the resolution interface and pinned
behaviorally. A new test asserts every row's field set, so an unenforceable
column cannot reappear without coverage — verified by re-adding one and
watching it fail. Routing the structural stages into typed behavior is
tracked in #1656 with the constraint that each field must be consumed, not
merely declared.

* fix(selectors): flatten the policy outcome at the package boundary

`PolicyResolutionOutcome.resolution` was typed as `AstSelectorResolution` and
the root façade returned it unchanged, so the parser AST #1589 confined to
`@agent-device/selectors/ast` came back through a nested field.
`selector-wait.ts` reading `outcome.resolution.selector.raw` was the runtime
proof. The existing boundary gate reads exported *names*, so it could not see
this.

The public outcome now lives beside `SelectorResolution` in
public-resolution-types.ts with its selector as text; the parser-side shape is
renamed `AstPolicyResolutionOutcome` and stays package-private, and the façade
wrapper flattens on the way out — the same treatment `resolveSelectorChain`
already gave `AstSelectorResolution`.

Two new pins, both verified red against the shape they replace: a behavioral
one asserting the façade returns selector text under every policy row, and a
structural one asserting resolution shapes are re-exported from
public-resolution-types.ts rather than from a parser-side module — which is
what distinguishes the leak from a correct re-export in a name list.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Rva4YGtSCAKJqH5PbpcCU

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-06 21:10:30 +02:00
Michał Pierzchała 3937036e5e feat: support --settle on scroll and back (#1638) (#1650)
* feat: support --settle on scroll and back (#1638)

Scroll-then-observe and back-then-observe are legitimate agent pairs, but
the post-action observation registry never grew past the touch commands, so
`--settle` on either was rejected with INVALID_ARGS — burning a tool call
each in AppControlBench's bsky-16.

Both commands now carry the `settle` descriptor trait, and every surface
derives from it rather than a hand list: CLI allowed flags, MCP/SDK input
fields, the flag-sourced timeout envelope, and MCP ref-pinning. The CLI
flag/metadata helpers moved out of the interaction family into
post-action-observation-grammar.ts (back is a system command), and
SETTLE_REF_ISSUING_TOOLS became a derivation — a hand list would have
silently stopped pinning the new commands' refs.

settleAfterInteraction and the new settleObservationCommand are two entry
points over one engine: same loop, storage, hints, and diff bounds, with the
target-less path supplying its own baseline and no proximity point. The
daemon reaches that command through the runtime surface, never by importing
`commands/` (R2) — the same seam the touch handlers use for press/fill —
and generic-settle.ts is loaded through a lazy `await import` returning a
closure, so the interaction runtime subgraph stays out of this dispatcher's
static graph (a static edge folded ~18 files into the daemon-server type
cycle; R10 caught it).

Both of generic-settle's orderings are load-bearing and tested: the baseline
is frozen before dispatch (and before the Android dialog preflight), and the
observation runs after markDeferredInteractionOutcome so settle's first
capture folds in the #1542 stabilization rather than racing it. The ADR 0014
"a settled diff publishes refs" rule moved to settle-ref-issuance.ts, shared
by both routes.

One divergence is deliberate: scroll/back resolve no element, so the diff
baseline is the session's STORED pre-action tree — "settled tree vs the
last tree you observed" — not press's freshly resolved pre-action capture.

Both commands also switch to preserve-daemon on timeout, which changes the
non-settle path too: with --settle their dominant hang mode is now a wedged
accessibility bridge, and a timed-out capture must not reset the daemon and
lose every session (#1105). The reviewed-set gate records it.

Live-validated on an iOS 26.2 simulator (Settings): scroll --settle settled
in 1786ms with a +6/-6 diff carrying fresh refs; back --settle in 771ms with
+15/-6. Alternating cost runs, one call vs the pair it replaces:
scroll 2.9-3.0s vs 5.3-5.6s, back 3.1-3.2s vs 4.7-5.1s. Those include the
#1627 deep-capture extension.

* fix: render settled-diff refs paste-ready in CLI output

A settled diff activates a PARTIAL ref frame (ADR 0014), which admits only
the pinned `@eN~s<gen>` form of the refs it issued. The unchanged-interactive
tail already rendered that way, but the diff's own added lines rendered the
bare `@eN` embedded in the snapshot line — so a CLI caller who copied the
ref the diff just handed them got `plain_ref_requires_complete_frame` and had
to append the generation by hand.

Added lines now render pinned when the response carries `refsGeneration`,
exactly like the tail. Removed lines render verbatim: they name elements that
just left the screen, and `SettleDiffLine` never gives them a ref.

This is not new to scroll/back — press/click/fill/longpress had the same gap
since #1101. MCP was never affected: its ref-pin store rewrites plain refs on
the way in, which is why the model never sees a suffix.

Live: `scroll down --settle` now emits `+ @e14~s218078 [cell] "Game Center"`,
and `press @e14~s218078` copied straight out of that line taps successfully.

* test: record the pinned-diff-ref bytes in the output-economy baseline

Rendering added diff-line refs pinned costs 8 bytes in the two settle CLI
text samples (two `~s<gen>` suffixes). The output-economy baseline is the
tripwire for exactly this, so the increase takes an explicit reviewed waiver
rather than a silent baseline bump — the same one the settled TAIL's pins
already carry, for the same ADR 0014 reason.

Only `bytes` moves: lines, refs, hints, and shape are unchanged, which is the
evidence that this is a suffix on existing refs and not a new payload.

Caught by CI, not locally: `pnpm test:unit` runs unit-core and
subprocess-stub only, while the Coverage lane runs every vitest project.

* test: prove the generic settle degrades when its runtime cannot be built

`createGenericSettleRuntime` catches and returns undefined so an observation
that cannot even start does not fail an action that already succeeded. That
was a claim in a docstring with nothing behind it — the one changed line the
coverage gate reported uncovered (95/96).

The test puts the session in the state the catch exists for: the router
handed us a session that is no longer in the store, so building the settle
runtime throws SESSION_NOT_FOUND. The response keeps its scroll result and
simply carries no settle payload. Removing the try/catch fails it.

* build: teach fallow that vi.mock reaches pinOwnProcessStartTime dynamically

Not from this PR: #1642 added `pinOwnProcessStartTime` on main, and its three
consumers reach it the only way a Vitest module mock can —
`vi.mock(path, async (importOriginal) => (await import('...')).pinOwnProcessStartTime(...))`.
Dependency analysis cannot follow that dynamic import to a consumer, so the
export reads as dead the moment any PR pulls that file into its audit scope.
This PR is the one that did.

The entry records the consumers by path and the reason, matching the
daemon route-handler entry directly above it, which exists for the same
dynamic-`import()` limitation.

* refactor: adopt the best of the parallel #1653 implementation

Two sessions independently built #1638 (PR #1650 and PR #1653) and converged
on the same architecture — trait in the registry, one engine with two entry
points, runtime-command seam, lazy import, preserve-daemon, stored-baseline
honesty. #1650 continues; this folds in what #1653 did better:

- The agent-facing help core loop (cli-help.ts) now names scroll and back as
  settle-capable. Without this, the benchmarked closed-grammar help line kept
  instructing agents that --settle is only for press/click/fill/longpress —
  actively steering the AppControlBench models away from what #1638 shipped.
- issueSettleRefs moves into session-snapshot.ts, beside the partial-frame
  primitive it wraps, deleting the single-function settle-ref-issuance module.
- Their seam tests: back reader→writer settle plumbing, back CLI settle
  rendering, and a trait-less generic command (home) ignoring a stray settle
  flag rather than observing or rejecting.

What #1650 had that #1653 lacked, for the record: the SETTLE_REF_ISSUING_TOOLS
registry derivation (without it, MCP never pins a scroll/back settle diff's
refs and the partial frame rejects every follow-up), BackCommandResult.settle
in contracts, back's MCP output schema, paste-ready pinned diff refs, and the
docs/changelog/baseline surfaces.

* bench: help-conformance case for settled scroll-to-find planning

The #1638 extension of the closed --settle grammar to scroll/back is the
feature's entire payoff — collapsing scroll-then-observe into one call — and
the closed command list is an enumerated N whose enumerator is this bench.
The regex over the help text proves the sentence exists; this case checks
whether a model plans differently because of it.

One focused case, deliberately not coached: a pinned visible-first snapshot
(rendered by formatSnapshotText, pinned by the sample-producers gate) whose
wanted row is summarized off-screen with no ref anywhere in the output. The
tempting pre-#1638 plan is `scroll` plus a separate `snapshot -i`; acceptance
is the single settled call. Scoring was verified against eight plan shapes in
both directions before recording.

Model-backed record (claude-haiku-4-5, 3 trials, current help): 0/3 — but the
decomposition is the finding. Settle eligibility GENERALIZED (3/3 trials put
--settle on scroll unprompted; the mutation-suffix framing concern did not
materialize) and the two-call habit is residual (1/3). All three trials failed
on `scroll @e3 down --settle` — the pre-existing #1366 scroll-takes-no-target
confusion, which the live CLI recovers with a dedicated hint but a single-shot
bench cannot. The recorded gap is therefore a first-30 doc gap (nothing
teaches that scroll takes no target), not a settle-eligibility gap; tuning the
case until it passes would just delete the evidence.
2026-08-06 20:04:16 +02:00