Commit Graph

68 Commits

Author SHA1 Message Date
Michał Pierzchała 5bc3354113 refactor(runtime): move readiness into platform owners 2026-08-12 11:42:19 +02:00
Michał Pierzchała 123e2c607e fix: separate boot admission from readiness 2026-08-12 11:42:19 +02:00
Michał Pierzchała cd3b4af782 refactor: route boot through readiness runtime 2026-08-12 11:42:19 +02:00
Michał Pierzchała 97c87eec4d refactor(registry): exhaustive platformExecution discriminator (ADR 0019 §6) (#1740)
* refactor(registry): make the platform-execution discriminator exhaustive

ADR 0019 §6 (amended): every command descriptor declares its platform-execution
mode explicitly. Adds the `none` mode to `CommandPlatformExecution`, removes the
silent `{ kind: 'legacy' }` default at registry entry, and annotates all 76
descriptors so the migration denominator is machine-readable.

Part of #1739 (wave 0)

* fix(registry): react-devtools executes delegated platform behavior

`react-devtools start` on a Limrun Android instance dispatches internal
`runtime port-reverse`, which reaches a provider device runtime, so ADR 0019 §6
`none` is false for it. Reclassify as `legacy` and add the derived coherence gate
that catches delegated platform execution: if a CLI route for command R
dispatches command D, R may declare `none` only when D is `none`.

Part of #1739 (wave 0)

* fix(registry): attribute CLI dispatches by occurrence, not command name

Subtracting attributed command NAMES let a stray dispatch hide behind a routed
one that names the same command, so the gate's totality claim did not hold.
Dispatch sites now carry their source offset and attribution subtracts
occurrences.

Part of #1739 (wave 0)

* fix(registry): unresolvable CLI daemon-send targets fail the gate

An unknown literal or computed command target resolved to undefined and never
entered the scan, so a dispatch could evade attribution by naming a target the
gate could not read. Daemon-send envelopes are now located by their send call and
an unresolvable target is reported instead of skipped.

Part of #1739 (wave 0)

* refactor(cli): own injected daemon dispatches at a typed construction seam

The syntactic scan recognized only a direct sendToDaemon call whose first
argument was an inline object literal, so a variable envelope or a computed
callee was omitted from every result. Rather than teach the scanner more shapes,
the CLI's injected dispatches now flow through one typed construction point whose
route/command pairs are declared, and the gate reads that declaration instead of
recovering it from syntax.

Part of #1739 (wave 0)

* fix: constrain injected daemon transport handoffs
2026-08-12 10:40:42 +02:00
Michał Pierzchała 74eab2a554 refactor: route selector-resolution structural stages into typed policy (#1744)
* refactor: route selector structural stages into typed policy

#1649 landed the per-caller ambiguity matrix and deliberately left four
structural columns out: occlusion, off-screen, hittable-ancestor promotion,
and the poll budget were per-caller pipeline code, so declaring them would
have been an unverifiable claim (nothing consumed them; flipping one left the
suite green).

This adds the missing half as a table with runners. `SELECTOR_PIPELINE_POLICIES`
(src/core/selector-pipeline-policy.ts) gives each caller ONE row naming its
ambiguity contract plus its four stages, and every stage is reached only
through a runner that reads the row:

- occlusion -> selectorPipelineCandidates (candidacy) and
  resolveSelectorPipelineTarget (refusal). Acting rows exclude covered nodes
  and refuse covered targets; `find` and the diagnosis probe keep them as
  candidates and refuse at the target; reads and `wait` ignore them.
- promotion -> resolveSelectorPipelineTarget. The per-call-site
  `promoteToHittableAncestor: boolean` is gone: click/press/longpress name
  `promotedTarget`, fill/focus/scroll/drag endpoints and the native-ref
  preflight name `resolvedTarget`. `find`'s below-the-root variant is a
  declared value rather than a second local helper.
- off-screen -> throwIfOffscreenInteractionTarget, which now takes the row and
  returns the node untouched (no iOS rescue round trip) for observation rows.
- poll -> selectorPollBudget, which createWaitPolling derives its deadline and
  inter-poll delay from; the two wait loops carry a budget, every other row
  carries none and cannot be polled.

Behavior is byte-identical. The acting refusal keeps its exact node, label and
details in every branch (promotion declines to retarget away from a covered
node, so the "both covered" case names the same node it always did), and
`find` carries the occlusion verdict to the focus/type seam rather than
raising it early, because find click/fill still delegate that refusal to the
interaction leaf's own error shape.

selector-pipeline-policy.test.ts drives EVERY row through EVERY runner,
including the rows whose answer is "skip" — the half that used to be an
absence of code, and an absence cannot fail. Each stage was proven red by
flipping its cell (occlusion, promotion, off-screen, poll, plus the
declare-only-what-is-enforced guard). The ADR 0011 occlusion/nonHittable
`via` pointers for the runtime tree paths now name the runner that makes the
decision, not the predicate it applies.

Closes #1656; prework for #1739 (waves 4-5).

* docs: state constraints instead of narrating the refactor

Comment pass over #1656: drop the "used to be per-caller code" /
"not module constants" / "rather than an omission" narration — a comment
should say what a future edit must respect, not what the previous shape was —
and compress the find occlusion-verdict and poll-budget notes to the
constraint they actually carry.

* refactor: make the selector pipeline the only door to the engine

Review of #1744: the structural rows were declared but bypassable. Read and
wait routes composed `selectorPipelineCandidates(row, nodes)` with the raw
`resolveSelectorChainWithPolicy(..., row.resolution)` and never entered the
promotion or off-screen stages, so flipping a read row's `promotion` or
`offscreen` changed only the policy unit tests — production `get`/`is`/`wait`
were unaffected, which is the unverifiable-column failure #1656 exists to
remove. Callers could also pair one row's candidate set with another row's
ambiguity contract, and `find list` reached the engine directly.

The owning interface (src/core/selector-pipeline.ts) now runs every stage a row
declares, skips included, and the stage functions are private to it:

- `resolveSelectorPipeline` — single-target rows: candidacy, ambiguity, the
  replay-guard hook, promotion, occlusion, off-screen.
- `listSelectorPipelineMatches` — `reject-candidates` rows, returning the
  candidate set AND the tree the row sees, so ranking and equivalence
  classification judge the same nodes candidacy produced.
- `runNodePipelineStages` — the node stages for a target from a non-chain
  matcher (`@ref`, find's fuzzy locator) or a narrowed candidate set.

A row whose off-screen stage refuses must supply a refusal shape, so flipping
an observation row to `refuse` fails on its real route instead of silently
observing. `find list` now names a `readList` row (the new
`reject-candidates`/no-rect ambiguity row) instead of calling the engine.

R17 selector-pipeline-ownership (scripts/layering/) makes the bypass
structurally inexpressible: only the owner may import the engine entry points.
Proven against a planted import in selector-read.ts, which the repo-wide scan
rejects with the entry points that replace it.

Flips now fail through REAL command routes, verified one at a time:
readUnique.occlusion/offscreen/promotion and wait.occlusion via get attrs / is
/ wait; readAny.offscreen via is exists and find; readList.occlusion via find
list; promotedTarget.promotion via runtime click. The wait route test needed an
advancing clock first — with the frozen one a refused wait spun instead of
failing, so the flip hung rather than asserting.

* refactor: drop find's dead candidate binding

The selector branch bound the row's candidate set and never read it: only the
acting classification needs that tree, and find's locator branch brings its own
matcher. Names what actually governs the locator target — the shared node
stages below, not a candidate set it never had.

* refactor: reserve the selector engine behind the pipeline owner

Review of #1744 (three blockers).

**Listing rows no longer claim stages they cannot run.** `find <q> list`
resolves to a candidate SET, so promotion, the off-screen guard and a poll
budget have nothing to apply to — a listing has no single element to retarget,
keep on screen, or wait for. `readList` now declares only the two stages a
listing executes (`SelectorListPolicy`: resolution + occlusion), and the
narrower shape is load-bearing: `runNodePipelineStages` and `selectorPollBudget`
take the full row, so handing them a listing row is a compile error rather than
a silently skipped stage. Pinned with `@ts-expect-error` — widening `readList`
makes the directives unused and fails the typecheck.

**The engine door is a specifier, not a symbol.** R17's regex could not see a
namespace import, a re-export, or a deferred `import()`, none of which mention
the symbol it matched. The two engine entries moved to
`@agent-device/selectors/engine`, and R19 enforces over the resolved import
graph, where every one of those forms is the same edge. Proven on the
repo-wide scan by planting each form into a shipped route: namespace import,
dynamic import, and `export *` laundering all come back red.

`resolveImportEdges` drops an edge whose specifier resolves to nothing, so a
specifier rule goes quiet — not red — if the subpath is ever retired. The gate
now says that out loud instead of scanning clean.

**R19, not R17.** #1750 allocates R17/R18. Verified free against origin/main
and that PR's diff, then validated by real merges in both directions: the
uniqueness gate passes either way and the three ids stay distinct.

The gate itself is new (`scripts/layering/rule-ids.ts`): two branches taking one
free number do not conflict in git, so nothing caught R17 twice. Matching whole
string literals is what separates a declaration from prose that names a rule,
and it is what let the gate see #1750's `const RULE = '…'` shape — the first
version missed it and would have been vacuous. `main`'s two pre-existing
collisions (R11, R13) are listed as known, not pinned by equality, so #1750
lands in either order without breaking this.

Also: the root façade now exposes no resolver at all, and its surface test
pins both doors.

* fix(layering): make each rule-id allowance expire with its collision

Review of #1744: `KNOWN_RULE_ID_COLLISIONS` filtered the exact R11/R13
collision strings, so once #1750 renames those rules apart the entries would
keep waving those very collisions through if anyone reintroduced them. "Inert"
was wrong — a stale allowance fails open, permanently.

`ruleIdCollisionFailures` now checks the transition from both sides: a
collision nobody allowed fails, AND an allowance whose collision is absent
from the scan fails as a stale allowance. The entry therefore has to be deleted
in the same change that removes the collision, and the list burns down to
empty, which admits nothing.

#1750 is still open, so the transitional entries stay for now (option (b)).
Verified against a scratch tree carrying that PR's rename: leaving the list
untouched reports both entries as stale; deleting them is clean; and
reintroducing `R11 names contracts-implementation-authority and
package-boundaries` afterwards is rejected. The last of those is also a unit
regression, so the post-transition guarantee is pinned rather than argued.
2026-08-12 07:57:08 +02:00
Michał Pierzchała f18f8b076f refactor(contracts): consolidate per-domain defineUse wrappers into one neutral defineUse (#1741)
* refactor(contracts): consolidate per-domain defineUse wrappers into one neutral defineUse

ADR 0019 §9: use declarations share one neutral defineUse; per-domain
currying wrappers around runtimeUse<PlatformRuntimeOperations>() add a
module per domain for no information. network-runtime-plan.ts,
logs-runtime-plan.ts, screen-recording-runtime-plan.ts, and
app-log-resource-recovery.ts each re-derived their own curried alias
(defineNetworkUse, appLogUse, defineScreenRecordingUse, and an inline
instantiation) from the same generic factory with the same type
parameter.

Export defineUse = runtimeUse<PlatformRuntimeOperations>() once from
platform-runtime.ts (where runtimeUse lives) and re-export it from the
platform facade. Every runtime-use declaration (networkDumpUse,
networkAdmissionUse, the app-log uses, the screen-recording uses, and
appLogRecoveryUse) now builds through that single export; the
per-domain wrappers are deleted.

Type-level only: no required/preferred keys changed, and no use
declaration was added, removed, or altered. Existing deepEqual
assertions in network-runtime-plan.test.ts, logs-runtime-plan.test.ts,
and screen-recording-runtime-plan.test.ts already pin every produced
use object's exact {required, preferred} shape, so they double as the
before/after regression proof that this refactor is behavior-neutral.

Part of #1739 (wave 0).

* fix(contracts): move defineUse into platform-runtime-operations.ts

Review: defining defineUse in platform-runtime.ts required importing the
concrete PlatformRuntimeOperations catalog into the generic runtimeUse
primitive module, while platform-runtime-operations.ts already imports
generic runtime types from platform-runtime.ts. That's an avoidable reverse
type dependency — the lower generic primitive depended on its concrete
aggregate catalog. The unchanged SCC file count didn't prove this harmless;
it counts cycle members, not newly introduced back-edges.

Move defineUse = runtimeUse<PlatformRuntimeOperations>() into
platform-runtime-operations.ts, alongside PlatformRuntimeOperations.
platform-runtime.ts no longer imports the concrete catalog. Re-export
defineUse through the platform facade from its new source module; every
call site keeps importing it from @agent-device/contracts/platform
unchanged, and the three contracts-internal call sites now import it
directly from platform-runtime-operations.ts.

Validation: tsc (full workspace + examples/sdk), check:layering (136/136,
type-cycle count unchanged at 46), and check:affected --run (473 files /
3939 tests) all clean at the new head.
2026-08-11 18:05:32 +02:00
Michał Pierzchała f5d9789764 feat: enforce local device claims and reconcile stale owners (#1735)
* feat: enforce local device claims

* fix: address device claim review feedback

* fix: persist canonical daemon claim state directory
2026-08-11 16:18:45 +02:00
Michał Pierzchała 602b7a2995 refactor: narrow perf API to actionable evidence (#1731)
* refactor: narrow perf API to actionable evidence

* fix: address perf API review feedback

* fix: preserve deprecated Android CPU metrics
2026-08-11 15:28:58 +02:00
Michał Pierzchała b8dd6a5854 refactor: tighten capture ownership boundaries (#1736) 2026-08-11 13:55:28 +02:00
Michał Pierzchała 1b2e786128 refactor: move screen recording onto platform runtime (#1724) 2026-08-11 10:24:57 +02:00
Michał Pierzchała 1b75e102d7 perf: speed up device inventory and status (#1723) 2026-08-11 07:37:50 +02:00
Michał Pierzchała 338aa2a0d5 refactor: route every native selector resolution through the policy interface (#1715)
* refactor: route every native selector resolution through the policy interface

#1649 declared the per-caller ambiguity matrix; four native call sites still
bypassed it, spreading `selectorResolutionKnobs(row)` into a raw
`resolveSelectorChain` instead of naming the row. That left the "one
interface" claim aspirational: a caller could restate its contract as engine
knobs and nothing would notice.

- `is` non-exists, `get text`/`get attrs`, find's read actions, and the
  covered-selector diagnosis probe now call `resolveSelectorChainWithPolicy`
  with their existing row. Semantics are byte-identical: the knob-backed
  branch of that interface forwards to the same engine call the call sites
  built by hand.
- The façade drops `resolveSelectorChain` and `selectorResolutionKnobs`, so
  no knob-taking resolver is reachable from outside the package and a call
  site cannot re-acquire the knobs even by accident.
  `requireUnique`/`disambiguateAmbiguous` are now named in exactly one
  function, which `resolve-with-policy.ts` and the replay resolver both
  derive through.
- `get` names the two rows it may consume as a type, so pointing it at any
  other ambiguity contract is a compile error.

Tests: selector-read-policy.test.ts pins which row each read command
consumes, end to end, on one ambiguous fixture — the only tree the rows
disagree on. Each assertion was proven red by re-pointing its caller at a
neighbouring row. The knob-consistency check moves into the package beside
the now-private helper. Test call sites that used the raw resolver move to
`resolveRecordedTarget`, the same knobs and the path that actually replays a
recorded chain.

Extracting the failure branch drops `resolveSelectorInteractionTarget` below
the complexity threshold; its `fallow-ignore` waiver is removed (verified
load-bearing before the extraction, unnecessary after).

Closes #1630. Structural stages (occlusion, off-screen, promotion, poll
budget) stay per-caller pipeline code, tracked in #1656.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD

* test: observe which node find's row selected, not just that one existed

#1715 review, P2: the find row assertion was only half a pin. `find exists`
returns `found: true` for any resolved node, and the `list` call it leaned on
goes through listFindMatches — a path that consumes no policy row at all. So
repointing findFirstLocatorMatch at `readText` left both assertions green
while selection silently moved from the document-order head to the tiebreak
winner.

Assert through `find get_attrs`, which returns the ref of the node the row
actually selected. Both neighbouring rows are now red: `readText` fails
'@e3' !== '@e2' (the move the old test missed), `readUnique` fails by
refusing the ambiguous screen. `exists` stays as a second, weaker assertion
on the same resolution.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD

* refactor: route is exists through the matrix, collapse the double match pass

Follow-up tightening on the same seam.

`is exists` reached findSelectorChainMatch directly while the `readAny` row's
own doc claimed to serve "`exists` and find's read-only actions" — true of the
docs, not of the code, which is the unverifiable-claim shape #1656's review
called out. It now names `readAny`, the row it always described. Equivalent by
construction: both take the first alternative with any match under
requireRect: false, and disclose that alternative's count.

That leaves the root façade with no consumer for findSelectorChainMatch, so it
goes the way of resolveSelectorChain — dropped from the string-only façade,
kept on the published ./ast surface. Its façade-twin type SelectorChainMatch
dies with it (fallow caught it).

resolveSelectorChainWithPolicy matched twice on the uniqueness path: once via
resolveSelectorChain, then again to fill matchedNodes. Hoisting the single
list call above the row switch removes that second pass, collapses two
duplicated ambiguous literals into one helper, and drops a `?? [resolution.node]`
fallback that was unreachable — a resolution implies its alternative matched,
so the list is never null there.

While hoisting: the resolved arm's matchedNodes can describe a different
alternative than resolution.selector, because uniqueness skips an ambiguous
alternative to try the next one. Unreachable today (only first-match callers
read it, where both come from one list), and left as-is rather than silently
changed — but the doc claimed "the alternative it came from", so it now says
what is actually true.

Tests: is exists gets a caller-level pin on the shared ambiguous fixture —
passes with matches: 2 where its fail-closed siblings refuse — proven red by
pointing it at readUnique.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD

* test: discriminate is exists's row by alternative, guard the façade structurally

#1715 review, second regression-validity gap. The `is exists` pin observed
only `pass: true` and `matches: 2` on a fixture whose first alternative was
merely TIEBREAKABLE — so disambiguation succeeded there and reported the same
count first-match would. `readAny`, `readText`, and the pre-migration raw
lookup all produced that, and only the readUnique swap I had checked went
red. One mutation proven is not the same as the row being pinned.

`exists` exposes no node ref, so the row has to be read off WHICH alternative
answered. New fixture: alternative one matches two nodes that are genuinely
indistinguishable (same depth, same area, both on screen) so the tiebreak
declines; alternative two matches exactly one. First-match answers from
alternative one; every uniqueness row skips the undecidable alternative and
answers from alternative two. Asserting the selector now separates them —
readText and readUnique both fail with `id="save-unique"` where
`label="Save"` is expected.

Restoring the raw lookup stays behaviourally invisible, though:
findSelectorChainMatch is equivalent to the readAny row it migrated to, which
is precisely why that migration preserved semantics. No fixture assertion can
catch that revert, so the guard is structural — the façade's export list must
not carry resolveSelectorChain, findSelectorChainMatch, or
selectorResolutionKnobs. Follows the packages/maestro index.test.ts
absence-assertion precedent. Verified red by re-exporting the lookup.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD

* fix: cover selector routes in device replays

* test: simplify selector replay regression

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-11 07:34:11 +02:00
Michał Pierzchała cdc754e6ed perf: speed up iOS agent recovery and streamline CLI guidance (#1700)
* Avoid interactive children in parent taps

* docs: streamline no-skill CLI help

* perf: recover faster from sparse iOS trees

* fix: preserve selector context for blocked parent taps

* fix: preserve coordinate text-entry focus

* fix: preserve thin parent touch targets

* fix: fail closed for unscoped iOS typing

* test: isolate replay lock fixture

* test: share node integration process
2026-08-10 20:43:01 +02:00
Michał Pierzchała b15c502318 refactor: extract platform network runtime (#1702)
* refactor: extract platform network runtime

* fix: preserve platform network recovery routes

* test: guard network parser placement
2026-08-10 17:58:42 +02:00
Michał Pierzchała b1ed5353d1 refactor: extract platform log runtime (#1701)
* refactor: extract platform log runtime

* fix: clear terminal app log recovery markers

* fix: preserve scoped app log tooling

* fix: preserve app log cancellation

* fix: handle large changed coverage diffs

* fix: harden Limrun runtime identity

* refactor: tighten platform log runtime

* fix: close app log trust gaps

* fix: accept canonical session path aliases

* refactor: extract durable capture kit

* fix: refresh retained log marker admission

* fix: rotate app logs after process relaunch
2026-08-10 17:58:42 +02:00
Michał Pierzchała 13cc90ffc6 fix: harden Android snapshot and fill reliability (#1708)
* fix: harden Android automation reliability

* test: isolate CLI flush integration

* test: close Android review gaps

* test: register CLI transport fixture

* test: consolidate CLI subprocess fixture
2026-08-10 13:10:18 +02:00
Michał Pierzchała c06bed9f77 refactor: extract platform device inventory runtime (#1699)
* refactor: extract platform inventory runtime

* fix: preserve scoped Apple inventory tooling

* fix: preserve Apple tool cancellation

* refactor: tighten platform inventory boundaries
2026-08-10 12:51:59 +02:00
vw2x 4b432fb59b feat: add HarmonyOS support (#1683)
* feat: add HarmonyOS device automation foundation

Add HDC-backed discovery, snapshots, application lifecycle, and core mobile interactions.

Route HarmonyOS through the platform registry and client contracts.

Cover parsing and capability parity with focused tests.

* feat: support HarmonyOS HAP deployment

Install and reinstall signed HAP archives through HDC.

Resolve bundle identities from module metadata and relaunch after package replacement.

Extend deploy routing and capability coverage for HarmonyOS.

* feat: add HarmonyOS single-pointer gestures

Execute pan, fling, and swipe plans through HDC uiInput primitives.

Derive scroll coordinates from the live ArkUI viewport.

Keep unsupported multi-touch gestures explicitly rejected.

* refactor: split session inventory command handling

Separate session, device, capability, and app inventory response paths.

Preserve the public inventory response contract while reducing handler complexity.

* feat: support HarmonyOS keyboard actions

Route HarmonyOS enter, return, and dismiss through HDC key events.

Expose supported keyboard actions through the system command metadata.

Keep keyboard visibility inspection explicitly unsupported.

* fix: reject unsupported HarmonyOS drag gestures

Keep drag unavailable until HDC can preserve source and destination hold semantics.

* feat: add HarmonyOS app log streaming

1. Stream HarmonyOS app logs through PID-scoped hilog sessions.\n2. Record HarmonyOS app identity during bundle-id opens for app-scoped commands.\n3. Cover backend routing and bundle identity resolution.

* feat: report HarmonyOS foreground app state

1. Read the foreground HarmonyOS mission through aa dump.\n2. Expose HarmonyOS appstate with package and ability metadata.\n3. Add parser coverage for foreground and missing-state cases.

* fix: advertise appstate through capabilities

1. Classify appstate in the command descriptor capability matrix.\n2. Surface supported appstate commands in capability inventory.\n3. Cover the advertised Android capability contract.

* feat: sample HarmonyOS process performance

1. Sample HarmonyOS process CPU and resident memory through HDC.\n2. Expose the verified metrics through the shared perf command.\n3. Keep frame and memory snapshot collection explicitly unavailable.

* feat: clear HarmonyOS app state

1. Add HarmonyOS settings clear-app-state through bundle cleanup.\n2. Force stop the app before clearing data and cache.\n3. Reject all unverified HarmonyOS settings explicitly.

* docs: document HarmonyOS support

1. Describe HarmonyOS HDC prerequisites and HAP installation.\n2. Add HarmonyOS to platform discovery and product documentation.\n3. Document verified performance limits for the public HDC surface.

* fix: preserve HarmonyOS deploy session identity

1. Bind a resolved HarmonyOS bundle after install or reinstall.\n2. Keep app-scoped logs and observability available after deployment.\n3. Cover session identity preservation for HarmonyOS reinstall.

* test: lock HarmonyOS capability boundary

1. Add an independent HarmonyOS capability-matrix oracle and exact advertised-command regression test.
2. Document current HDC-backed support and evidence-based unsupported command boundaries.

* refactor: simplify HarmonyOS shared platform boundaries

1. Split device selection and settings dispatch into focused helpers without changing behavior.
2. Keep HarmonyOS serial selection and lock-policy classification covered by regression tests.
3. Remove Fallow complexity findings from the HarmonyOS diff against upstream main.

* fix: bound default HarmonyOS HDC commands

1. Apply a 15 second timeout to ordinary HDC operations.
2. Preserve operation-specific timeout budgets for installation and capture paths.
3. Add regression coverage for default and overridden HDC timeouts.

* feat: add HarmonyOS screen recording

Implement physical-device whole-screen recording through the system recorder and HDC media transfer.

Reject unsupported HarmonyOS recording scopes and export flags.

Cover capability routing, media retrieval, cleanup, and simulator rejection.

* feat: report HarmonyOS HDC readiness

Add an HDC version check to the HarmonyOS doctor flow.

Document HarmonyOS as a supported doctor platform and cover the result.

* refactor: simplify HarmonyOS recording checks

Reduce recording validation and test complexity without changing behavior.

* test: cover HarmonyOS platform contracts

Synchronize public platform expectations across CLI, MCP, replay, and inventory tests.

Mock HarmonyOS inventory probes to preserve concurrent test behavior.

* test: model HarmonyOS recording capability

Require a physical HarmonyOS device in the independent capability parity oracle.

* test: cover HarmonyOS input and lifecycle paths

Exercise HDC input, lifecycle, installation, and relaunch command sequences.

* test: cover HarmonyOS device observability paths

Exercise discovery, screenshot validation, and process performance sampling.

* docs: define HarmonyOS CI hardware policy

Keep HDC hardware validation local and require mocked CI contract tests.

* fix: honor HarmonyOS app inventory filters

* fix: bound HarmonyOS app inventory classification

1. 限制应用元数据分类并发并为默认清单设置整体时限.
2. 将请求取消信号传递给 HarmonyOS 应用清单读取.
3. 补充失败时中止在飞读取且不继续排队的回归测试.

* fix: preserve HarmonyOS inventory failure causes

1. 保留触发应用元数据分类失败的原始错误, 避免被取消同级任务覆盖.
2. 补充总时限中止在飞读取且不启动排队任务的回归测试.
3. 验证后序任务失败时保留默认筛选的恢复提示.
2026-08-09 10:29:20 +02:00
Michał Pierzchała 18291ba8e2 perf: collapse app-driving startup turns (#1693)
* perf: collapse app-driving startup turns

* fix: align foreground open guidance
2026-08-09 10:10:11 +02:00
Michał Pierzchała 6c0fcb64a1 fix: reject distinct ambiguous mutation targets (#1667)
* fix: reject distinct ambiguous mutation targets

* fix(ios): scope the raw-match rejection to mutating dispatches

`RunnerTests+Interaction.findElement` applied the new fail-closed
classification to `querySelector` as well as press/type, because the read
call site takes the default `allowNonHittableFallback: false`. With one
visible/hittable match and one non-hittable same-selector duplicate the
query started returning AMBIGUOUS_MATCH where it previously selected the
hittable element, and `queryDirectIosSelectorOrFallback` preserves that
error for read callers — so `get`, `is`, and `wait` surfaced an error
instead of their prior answer.

`classifyDirectSelectorCandidates` now takes a `rawMatchPolicy`. Mutations
keep `.rejectDistinctMatches` (the default, so no mutation call site
changes); `queryElement` passes `.preferHittableMatch`, restoring the
prior read rule: prefer the single hittable match, ambiguous only when
hittable matches compete, and never adopt the Maestro coordinate fallback.
The Maestro expected-point path is untouched.

Covers the one-hittable + one-non-hittable read, competing hittable reads,
and the non-hittable-only read. ADR 0011's amendment now states the scope.

* test(ios): execute selector read ambiguity regression

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-08 08:57:32 +02:00
Michał Pierzchała 4b89482a06 fix(daemon): open --foreground P1 hotfix — selector rejection, interactive snapshot, capture-failure masking (#1671)
* fix(daemon): open --foreground P1 hotfix — selector rejection, interactive snapshot, capture-failure masking

Three P1s from the post-merge review of #1670, all at the
session-open-foreground dispatch seam:

1. Explicit device selectors were silently overwritten: the resolved-device
   rewrite pinned --udid/--platform over whatever the caller passed, so
   `open --foreground --udid B` with sim A sole-booted silently opened A.
   Now fails fast with INVALID_ARGS (matching the existing app-positional
   rejection) on --udid/--device, and on --platform other than ios; an
   explicit --platform ios passes through.

2. The promised interactive snapshot was never requested: the composed
   snapshot dispatch forwarded the open request's flags untouched, without
   snapshotInteractiveOnly — so the capture was NOT the `snapshot -i` path
   the doc comment promised and returned no interactive presentation. The
   composed request now sets snapshotInteractiveOnly: true (the exact key
   the CLI maps -i to and the snapshot runtime reads as interactiveOnly).

3. A capture failure masked the successful open: returning the snapshot
   error discarded openResponse even though the session exists, so a retry
   of `open --foreground` failed with "close the current session first".
   Open success + snapshot failure now returns ok with an explicit
   initialSnapshotError {code, message} detail and a rendered warning that
   the session IS open and how to capture manually (snapshot -i).

Regressions added for all three: explicit-selector rejection
(udid/device/both/non-iOS platform + ios pass-through), the composed
dispatch carrying snapshotInteractiveOnly, and the snapshot-failure path
returning ok + warning + usable session.

* fix(cli): render the composed open --foreground snapshot on default stdout and project initialSnapshotError through the public surfaces

Post-merge review on #1671 found the daemon fixes never reached the public
boundaries: openCliOutput ignored the nested snapshot (the one-call promise
held only under --json), and initialSnapshotError was daemon-only — absent
from AppOpenResult, Node normalization, and serializeOpenResult, with the
normalized shape truncated to code+message.

- default open output now renders the composed interactive tree through the
  same snapshotCliOutput path snapshot -i uses
- AppOpenResult carries initialSnapshotError as the FULL daemon error
  (hint/details/diagnosticId/logPath preserved) through normalization and
  serialization

Worker-authored; committed by the coordinating session after the worker
stalled twice mid-push. Tests: output.test.ts + session-open-foreground
(26 pass), typecheck, oxfmt.

* refactor(client): one daemon-error normalizer + client-route regressions for initialSnapshotError

Review follow-ups on #1671: normalizeInitialSnapshotError duplicated the
target-shutdown error normalization and pushed the module over the fallow
complexity threshold; consolidated into a single internal normalizeDaemonError
(table-driven, full shape incl. retriable/supportedOn) used by both result
paths, projected through normalizeOpenForegroundComposition.

Client-route regressions: createAgentDeviceClient().apps.open now proves the
full initialSnapshotError shape (hint/details/diagnosticId/logPath/retriable)
survives normalization, and that a malformed one is dropped — deleting the
boundary normalization fails these tests.

Also rebased onto current main.

* fix(daemon): a thrown initial-snapshot capture failure gets the same successful-open contract

Review P1 on #1671: dispatchSnapshotViaRuntime rethrows ordinary
capture/runner exceptions; the composition only handled a returned
{ ok: false }, so a thrown failure escaped to the router and failed the
whole open after the session was created — retrying then wedged on the
existing session. The catch normalizes the rejection (kernel
normalizeError, same conversion the router applies) into the shared
openWithInitialSnapshotFailure path: ok response, full-shape
initialSnapshotError, session-usable warning. Rejecting-mock regression
added alongside the returned-failure case.
2026-08-08 07:54:33 +02:00
Michał Pierzchała 13bc70f24f refactor(find): resolve a mutating find's target once, not twice (#1654) (#1669)
* fix(find): resolve a mutating find's target once, not twice (#1654)

A mutating `find click`/`find fill` captured the screen, matched by
locator under the `findAct` policy, promoted to a hittable ancestor, minted
`@eN` off the node it chose — and then re-entered the interaction leaf by
bare `@ref`, which looked that ref up AGAIN via resolveSnapshotForRef. The
second lookup reads the SESSION frame tree, not the fresh capture find
matched against, so the two could disagree: it could hand the action a
different node than find picked, or refuse as unresolvable a ref find had
observed a moment earlier.

find now passes the node itself. `internal.findPreresolvedTarget` carries
the resolved node and its tree over the in-process invoke hop, and the ref
branch adopts them instead of resolving the ref a second time.

What is NOT skipped: the guarantees. Occlusion, hittable-ancestor
promotion, and the off-screen guard all still run, on that node, at the
same symbols the ADR 0011 `runtime-ref` cells name. Only the LOOKUP is
replaced. Recording, ref-frame effects, settle/observation, and
deferred-outcome marking are untouched — they live in the dispatch wrapper,
not in resolution, which is why the invoke hop is kept rather than bypassed
the way find focus/type bypass it.

The channel is a second field rather than widening `findResolvedTarget`
because they answer different questions: that flag governs ref-frame
admission and staleness (policy), this governs resolution (lookup), and
focus/type set neither. It is `internal`-only, so it never crosses the
wire and carrying live node references is sound.

ADR 0011 re-check: the `runtime-ref` disclosure cell gains a second
producer, adoptPreresolvedRefTarget, reporting `exact`. Truthful rather
than borrowed — the ref is minted off the node handed over. It cannot
report label-fallback, which is right: label recovery is a property of
looking a stale ref up, and this path performs no lookup.

Tests pin the behavior by tap coordinates, so they name which tree the leaf
resolved against, plus a control proving the ordinary @ref path is
unchanged and a case where the session tree cannot resolve the ref at all —
that one can only pass if no second lookup runs. Verified revert-sensitive:
removing the short-circuit fails 3 of the 4, and the control stays green.

Out of scope, unchanged: the ADR 0011 `native-ref` path (web provider
clickRef/fillRef only), where the ref is the provider's own element handle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117wfrvC6MRDJUWdRBErNEF

* test(find): prove the production route, and correct the divergence claim (#1654 review P2)

The regression tests built `internal.findPreresolvedTarget` by hand and
called handleInteractionCommands directly, so they proved the leaf consumes
the channel but never that find ATTACHES it. Either producer could have
stopped doing so and they would all have stayed green — the exact gap
#1649's review named, reappearing one layer up.

Four tests now drive the real handleFindCommands with invoke wired to the
real handleInteractionCommands. Two assert click and fill each carry the
selected node (separate producers, separate forwarding hops, so asserted
separately); two advance the session tree between find's match and the
leaf's resolution and assert the dispatched point is still find's node.

That last pair also corrects the record. The claim that the leaf resolves
against a different tree than find matched is too strong for the common
path: `omitRefFrameSnapshot` (interaction-runtime.ts) makes find's internal
dispatch skip the authorized frame tree and resolve against
`session.snapshot`, which find's own capture just wrote — so the second
lookup normally AGREES, and the two non-diverged tests above pass with or
without the fix.

The divergence is real but narrower: it needs something to advance
`session.snapshot` between find's match and the leaf's resolution, which
`refreshAndroidRefSnapshotIfFreshnessActive` does on the ref path. Measured
at this head, the old code taps (60,720) "Delete" where find matched
"Save" at (310,510) — a wrong-element mutation, now pinned by both click
and fill.

So the change's value is what #1654 asked for — one resolution end to end —
plus closing that window, not a fix for a divergence occurring on every
find.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117wfrvC6MRDJUWdRBErNEF

* fix(find): tighten resolved target provenance

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-07 17:45:43 +02:00
Michał Pierzchała 04c33f9e1e feat(ios): expose AX custom actions on merged accessibility elements (#1665)
* feat(ios): expose AX custom actions on merged accessibility elements

Apps that merge a card into one accessibility element for VoiceOver (React
Native's `accessible` prop) publish the card's real affordances as
UIAccessibilityCustomActions rather than as child elements. Our snapshot showed
only the merged node, so an agent looking for a feed card's options control had
nothing to aim at and fell back to coordinate guessing.

`snapshot --actions` now names them:

    @e8 [link] "feedItem-by-whiskers.test" actions: ["Reply", "Repost", "Open post options menu"]

Opt-in, because the AX server cannot serve custom actions in a bulk tree
request: adding the attribute makes testmanagerd's reply decoder reject the
nested arrays a custom action serializes into, drop the reply, and time the
request out (~65s vs ~110ms). Only per-element reads answer, at ~100ms each, so
the runner reads at most 12 labelled childless nodes and stops at the
capture-plan deadline. The request pins the private-AX backend, since no other
backend can read the attribute, and reports that as its own `requested-backend`
verdict so a deliberate pin never renders as a degradation warning.

Invocation is not shipped: the actions are readable but not invocable from the
runner. RunnerAXSnapshotBridge.h records the five APIs that were tried.

* fix(ios): disclose a capped custom-action pass, and read on-screen elements first

Two gaps in the first cut.

An element the bounded pass never reached rendered identically to one with no
custom actions, so a capped capture silently taught the reader that later feed
cards have no affordances — the exact mis-inference this feature exists to
prevent. The runner already counted reads against candidates; it now carries
both to the response, and the verdict renders one response-level line when the
pass was incomplete. A complete pass stays silent, and "never asked" stays
distinguishable from "read none" (the key is absent, not (0, 0)).

The obvious remedy for a capped pass would be a scoped re-run, but scope is
applied when the Swift walk builds nodes, long after the read pass, so it does
not redirect the budget at all — measured: reads=12 candidates=18 with and
without --scope. Rather than print a remedy that does nothing, the read pass now
orders candidates on-screen first. That makes the budget land on elements an
agent can act on, and makes the disclosed remedy true: scrolling changes the
on-screen set, so a re-run reads elements the previous pass could not.

Also states plainly, in the flag help and the tool/SDK field description, that
the names are for planning: nothing invokes them, so the affordance is reached
through the element's detail screen, the same control exposed elsewhere, or
coordinates.

* test(snapshot): pin the custom-action coverage pair in the verdict shape assertion

* chore(scripts): classify the --actions flag in the integration progress model

The completeness gate flagged snapshotCustomActions as unclassified, which is
what it is for. It gets its own bucket rather than joining the provider-scenario
table: the values come from the private AX client inside the runner process, and
the fake runner derives its behavior from fixture tables that cannot fabricate
custom actions, so there is no provider-backed scenario to claim. The owning
coverage is named instead — runner XCTest unit, snapshot-lines, snapshot-quality.

* fix(ios): fail closed on unserviceable --actions, bound each read, cap output, and count actions in identity

Four review findings.

1. `--actions` with `--raw`, or on any target that is not an iOS simulator, used
to succeed and return nodes with no actions — a requested capability silently
no-opped, indistinguishable from "this screen has none". Both now fail closed.
The raw pairing is rejected at the shared request seam (INVALID_ARGS) so CLI,
Node client and MCP answer alike before any device work; the platform case is
rejected once the session device is resolved (UNSUPPORTED_OPERATION), naming the
resolved target. `diff --actions` was already rejected as an unsupported flag.

2. The per-element AX read had no timeout, so one wedged element could consume
the whole capture budget. Each read now runs off-thread behind a 1s wait. A
timed-out element counts as unread, never as "read, and it has no actions", so
the existing partial-pass disclosure already covers it.

3. The element budget bounded element count only; one element could still return
an unbounded list of unbounded names. Capped at 8 names of 80 characters, and
clipped elements are counted into the coverage so a truncated list is disclosed
rather than silently presented as complete.

4. Action names were rendered unescaped, and no comparison key read them. Names
now get the same escaping as text previews plus control-character folding, so an
app-authored name cannot split or corrupt a line. `actions` joins the diff
comparable key, the unchanged-comparison projection, and — the sharper bug —
the snapshot presentation key, without which `snapshot` followed by
`snapshot --actions` on a still screen answered "unchanged" and never delivered
the actions that were explicitly requested.

* fix(ios): contain a hung custom-action read instead of accumulating orphans

The 1s read deadline frees the caller, but the underlying AX call is a
synchronous XPC round trip that cannot be cancelled — it keeps running. On a
global concurrent queue that meant repeated `snapshot --actions` against a
wedged element piled up orphaned reads, all using the shared XCAXClient
concurrently. The deadline was containment for the capture, not for the runner.

Since the call cannot be cancelled, contain it instead:

- every read runs on one dedicated serial queue, so a wedged call can never be
  joined by a second concurrent user of the shared client;
- a single-flight guard refuses to dispatch at all while an abandoned read is
  still outstanding, so a repeated capture adds no work — the dispatch counter
  stands still;
- the read pass stops at that point rather than paying a deadline per element
  on reads that would all be refused, and reports `blocked` so the capture stays
  honest. That gets its own line, because the partial-pass remedy (scroll and
  re-run) cannot clear a hang and would send the reader in circles.

Recovery needs no reset: when the hung call finally returns, in-flight drops to
zero and reads resume.

The regression drives a fake AX client that never returns, and asserts the three
things the fix exists for — exactly one in-flight read with no further
dispatches across repeated captures, immediate returns with the skip disclosed
instead of the scroll remedy, and reads working again once the wedge clears.

* ci(ios): execute the custom-action runner regressions instead of only compiling them

The iOS workflow runs a targeted -only-testing list, so a runner test that is
not named there is compiled by the build step and then never executed. All seven
custom-action tests were in that gap — including the containment regression,
which is the only executable proof that a hung AX read cannot accumulate
orphaned in-flight reads.

Red/green against the containment regression, with the fix reverted to its
pre-fix concurrency behavior (global concurrent queue, no single-flight guard,
no blocked exit):

  RED   in-flight 6 (want 1), dispatches 6 (want 1), each repeat paid the full
        1.004s deadline (want <0.2s), blocked=false (want true), and the
        in-flight drain never completed — "Exceeded timeout of 5 seconds".
  GREEN 7/7 pass, containment regression in 1.02s.
2026-08-07 17:35:48 +02:00
Michał Pierzchała e14c9d8d7b fix(cloud): unblock iOS snapshot and gate cloud fill on text-entry focus (#1658) (#1666)
* fix(cloud): unblock iOS snapshot and gate cloud fill on text-entry focus (#1658)

Two bugs isolated to the cloud-webdriver iOS path.

`snapshot`/`diff` refused every capture on a live BrowserStack session with
SESSION_NOT_FOUND, instantly and without a driver round trip. The app-session
guard they ran belongs to the local XCUITest runner, which must attach to a
target app; a cloud capture reads the provider's own driver session and needs
no app identity, so it now applies to local Apple targets only. The session
was empty in the first place because the provider open path skips local app
resolution wholesale — no simctl/devicectl reaches a hosted device — and
dropped an explicitly spelled bundle id along with it. A dotted, non-deep-link
target is the bundle id under the same convention resolveIosApp applies
locally, so a cloud `open com.example.app` now records it.

`fill` tapped and sent its keys in back-to-back requests. A WebView input —
an OAuth page in a Safari view controller — takes first responder
asynchronously, so the keys landed with nothing focused while the command
still answered "Filled N chars"; tapping and filling as two separate commands
worked only because the round trip between them gave the field time to focus.
The cloud interactor now waits on the same signal the Apple runner uses, the
software keyboard going from hidden to shown after its tap, and discloses what
it observed as `textEntryReadiness` so a fill with no witness cannot pass for
a filled field. Where keyboard visibility cannot witness the focus move —
back-to-back fills into one form, the shape that failed most often — it spends
the runner's full readiness budget rather than racing the app with a short
settle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Zjbzf7HdziX9SzFpNP7WW

* fix(cloud): let a new bundle-id open replace the tracked cloud iOS app

Adopting an explicitly spelled bundle id on a provider-backed open (the
fix that makes snapshot/diff work at all) also made a previously dead
precedence rule live: the provider branch returned currentAppBundleId
first, so once a first open had populated it, `open com.a` followed by
`open com.b` left the session still reporting com.a to every
appBundleId-gated command.

The local path does the opposite, and is the convention this branch is
meant to mirror: resolveIosApp returns a dotted target unchanged and
never consults the session's current app. Only its deep-link branches
prefer the tracked id. Flip the provider branch to match — an explicit
bundle-id target wins, and everything the branch cannot name (deep
links, display names, bare open) still falls back to the tracked id.

* fix(cloud): fail a witness-less cloud fill instead of reporting it filled

Review follow-ups on #1658.

`not-observed` was still a success: it sent the keys and answered "Filled N
chars", and nothing renders `textEntryReadiness` in default CLI output — so the
exact silent success this branch exists to remove survived whenever focus never
happened. A tap that raises no keyboard now fails with
`text_entry_focus_not_observed` and sends no keys, leaving the field untouched
rather than half-written, and the readiness vocabulary keeps only outcomes that
describe a fill that did type.

The readiness budget was advertised but not enforced at the request boundary:
each keyboard probe inherited the client's 30s default, so one hung probe could
hold a 2s wait for far longer. Probes now carry their own bound, threaded
through the client as a per-request timeout override.

The probe also swallowed every error as "this driver cannot answer", which
degraded a dead session, an auth rejection, or a grid outage into a blind text
entry. Only a positively classified unimplemented route counts as unsupported
now — classified on the W3C error code rather than the status, since `unknown
command` and `invalid session id` share HTTP 404 — and everything else
propagates.

The provider scenario proved request ordering against a stub that always
accepted keys. Its fake now models the device: focus lands a beat after the
tap, and keys arriving while the keyboard is down are accepted and dropped,
exactly as an unfocused field does. The tests assert the field's own value, and
both go red against the pre-fix `fill`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Zjbzf7HdziX9SzFpNP7WW

* fix(cloud): witness the tapped field's focus before a cloud fill types

Review of cc23f2b found two ways a fill could still report success
without evidence that OUR tap focused the field it was aimed at.

P1. `settled-keyboard-up` and `settled-unknown` both typed and returned
normal success. Keyboard visibility can only witness that *a* field took
focus, never *which*: filling a second field in an already-open form
reads the same before and after, so a missed tap left the first field
focused and `POST /keys` — which the driver routes to whatever holds
first responder — appended to it while every request returned 200.

Failing those closed outright would have broken ordinary multi-field
form fills, which do work: a live AWS Device Farm run types both fields
of a WebView login correctly. So witness focus properly instead. W3C
`GET /element/active` answers the question keyboard visibility cannot —
is the thing focused now the thing I tapped — and answers it whether or
not the keyboard was already up. That becomes the primary signal
(`focused-element`); the keyboard transition stays as the fallback for
drivers without the route, and a keyboard already up on such a driver
now refuses rather than typing.

The test is identity, not geometry. Containment of the tap point looks
like the obvious rule and is wrong: focusing a field can re-lay it out.
On a live iPhone 16, tapping Safari's collapsed address bar expands it
into a taller field that no longer covers the tapped point, and a
containment-only rule refused a fill that plainly worked. So a tap that
MOVES focus counts, with containment as the second half of the test —
re-filling the already-focused field moves nothing, and only geometry
tells that from a tap that missed. Both readings are taken before the
tap, since each is evidence only as a change.

P2. The 2s budget bounded the loop but not the calls inside it: every
probe got a fixed 1500ms, so one begun near the deadline finished well
past it. Both the probe timeout and the sleep are now capped by the
remaining budget.

Also fixes a related escape the review did not name: the poll loop had
no catch, so one transient grid error aborted a fill the next poll would
have satisfied. Probe failures are now tolerated within the budget, but
a budget that expires without a single answered probe rethrows, so a
dead session surfaces as itself rather than as "the tap missed".

The provider scenario gains the two-field case the review asked for: it
begins keyboard-up with the email field focused, misses the password
tap, and asserts no keys reach the email field.

Verified on AWS Device Farm iPhone 16 / iOS 18.0 at this exact tree:
address bar (the re-layout case) and both WebView login fields all
report `focused-element`, the second with the keyboard already up, and
the device reads back `tomsmith` and a 20-character password.

* fix(cloud): refuse a cloud fill no focus probe can witness, and bound the composite probe

Two blockers from the review of 3a9aceb9.

P1. `settled-unknown` was the last path that typed without evidence: when
both the active-element and keyboard routes are positively unsupported,
`fill` settled 350ms, typed, and returned ordinary success. Nothing
renders `textEntryReadiness`, so that reached a caller looking exactly
like a fill that worked — the same silent false success #1658 is about,
just narrowed to one branch. It now refuses with a distinct reason,
`text_entry_focus_unobservable`: nothing is wrong with the target, the
driver simply cannot answer, so the caller's next move differs from a
missed tap and the hint names it — `press` then `type` stays the
deliberate way to enter text unwitnessed.

`CLOUD_TEXT_ENTRY_READINESS` is now `focused-element` and
`keyboard-shown` only. Every value describes a fill that witnessed focus
before sending a key; there is deliberately no value for typing blind.

P2. `activeElement(timeoutMs)` bounded each of its two sequential
requests by the full timeout rather than bounding the operation, so a
probe handed the 1.5s left of a 2s readiness deadline could spend ~3s
across `/element/active` and `/element/{id}/rect` and overrun the
deadline it was derived from. It now derives one deadline at entry and
gives the second request only what the first left, floored at zero so an
already-spent budget aborts immediately instead of falling back to the
client default.

The regression pins elapsed transport time across both calls, which is
what the defect is made of: the rect request answers only its own abort,
so the time it was allowed to run IS the budget it was handed. It
measures ~202ms of a shared 200ms budget before the fix and ~120ms
after.

Also updates the generic Cloud WebDriver facade scenario, whose stub
answered `{value: null}` to everything and so read as a driver with
neither route. It now answers the two focus probes, since that scenario
exercises facade wiring rather than text-entry semantics — those live in
cloud-webdriver-ios-text-entry.test.ts, which models focus properly.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-07 15:52:24 +02:00
Michał Pierzchała cc943400a9 feat(daemon): [RFC] prototype foreground-attach convenience (#1670)
Prototype `open --foreground`: on a fresh session with no app argument,
auto-resolves the target from the sole booted iOS simulator's sole
foreground app (reusing the exact same ambiguity-detection probe that
enriches the SESSION_NOT_FOUND hint), then attaches the initial
interactive snapshot to the response by composing the existing
snapshot-runtime dispatch. Collapses the documented 3-call
snapshot-fails -> read-hint -> open -> snapshot-succeeds dance into a
single call for the unambiguous case, while failing closed
(AMBIGUOUS_MATCH) with no guessing otherwise.

First-pass RFC, not reviewed — see PR body for the design tradeoff
writeup, live before/after evidence, and scoped-out follow-ups.
2026-08-07 13:17:13 +02:00
Michał Pierzchała 10ff339d14 refactor: declare selector resolution policy as data (#1649)
* refactor: declare selector resolution policy as data (#1630)

Five native consumers of "resolve a selector against the screen" each
hand-declared their ambiguity contract as inline requireUnique/
disambiguateAmbiguous literals, so the repo's real policy matrix was only
discoverable by reading four files. SELECTOR_RESOLUTION_POLICIES
(packages/selectors) now declares one row per caller — ambiguity kind plus
the structural columns (rect, occlusion, off-screen guard, promotion, poll)
— and selectorResolutionKnobs turns a row into the engine knobs it stands
for. Callers consume rows; zero ambiguity literals remain in src.

Semantics are unchanged by construction: each row was read off its call
site. The matrix names what was previously implicit — act and get text
disambiguate, is/get attrs fail closed, exists/find-reads and wait take the
first match, mutating find rejects candidates unless narrowed (#1625).
`reject-candidates` is declaration-only and rejected by
selectorResolutionKnobs at the type level, because find enforces it through
its own narrowing rather than engine knobs.

resolution-policy-parity.test.ts gate-tests the matrix against the callers
(ADR 0011's declared-plus-gate-tested pattern): knobs must match the named
ambiguity contract, every claimed structural column must appear in the
caller's source, the read/wait pipelines must genuinely lack the machinery
they disclaim, and no caller may reintroduce an inline literal. Verified
revert-sensitive: flipping readUnique to disambiguate and faking wait's
occlusion column each fail it.

Out of scope, unchanged, per the issue: the Maestro engine (ADR 0015) and
the open click-implicit-wait product decision.

* refactor: route wait and mutating find through the policy interface (#1649 review)

P1 was right: the first head declared seven rows but genuinely routed five.
selector-wait.ts never imported its row (it called listSelectorChainMatches
directly), findAct consumed only requireRect while its ambiguity contract
stayed bespoke, and the parity test sniffed marker strings in source files —
so it stayed green across exactly that gap. Asserting about the layer I had
edited instead of the behavior it produces.

resolveSelectorChainWithPolicy is now the one policy-driven entry: it
returns a discriminated outcome (none / resolved / ambiguous) because the
rows genuinely disagree about what several matches mean, which is what
previously forced each caller to re-derive its contract inline. wait and
find's selector branch both route through it; find additionally asserts its
row still says reject-candidates rather than assuming.

The parity test is rebuilt on fixture trees driven through that interface —
no source sniffing. Wiring verified revert-sensitive: flipping the wait row
fails the policy tests, and flipping findAct fails REAL find handler tests
(ambiguous-candidate listing), which is the proof the previous version
could not produce.

One behavior nuance the fixture work surfaced and now pins: disambiguation
declines on genuinely indistinguishable candidates (the tiebreak is
evidence, not a coin flip), so an acting row surfaces ambiguity there rather
than binding one silently.

* fix(test): let fallow see the host-process mock helper's real consumers

Rebase onto main brought #1642's host-process-mock.ts into this PR's
fallow scope, where its export reports as unused. It is not: three suites
consume it, but only through `(await import(...)).pinOwnProcessStartTime`
inside vi.mock factories — vitest hoists those above static imports, so the
dynamic form is required and fallow cannot trace it statically. Documented
suppression rather than a restructure that would break the hoisting
contract.

Latent on main rather than introduced here: the audit gate is
changed-files-only, so main sees the file in scope only from a PR whose
diff contains it.

* fix: keep every candidate when a policy resolves one winner (#1649 review P1)

A real regression I introduced, not a test gap: routing wait through the
policy interface collapsed the candidate set to the winner, and the #1349
landmark check is satisfied when SOME match carries the recorded identity.
A first same-selector impostor therefore hid a later genuine landmark and
timed the wait out.

The resolved outcome now carries `matchedNodes` — the full candidate set of
the alternative the winner came from — so a policy that picks one node no
longer throws the rest away. wait passes that straight to the landmark
check, restoring the original semantics.

Regression test added at the within-one-poll shape the existing suite did
not cover (both candidates in the SAME capture, impostor first); verified
it goes red against the singleton reconstruction it replaces.

* refactor: declare only the policy fields the matrix enforces (#1649 review)

The occlusion / offscreenGuard / promotion / poll columns were never
consumed by resolveSelectorChainWithPolicy or selectorResolutionKnobs:
changing any of them left behavior and the suite green, so they were
unverifiable claims that read as truth. (My earlier source-sniffing test
"verified" them by grepping caller files for marker strings — which is why
it also stayed green when a row was disconnected entirely.)

The matrix now declares exactly what it enforces: the ambiguity contract and
the rect requirement, both consumed by the resolution interface and pinned
behaviorally. A new test asserts every row's field set, so an unenforceable
column cannot reappear without coverage — verified by re-adding one and
watching it fail. Routing the structural stages into typed behavior is
tracked in #1656 with the constraint that each field must be consumed, not
merely declared.

* fix(selectors): flatten the policy outcome at the package boundary

`PolicyResolutionOutcome.resolution` was typed as `AstSelectorResolution` and
the root façade returned it unchanged, so the parser AST #1589 confined to
`@agent-device/selectors/ast` came back through a nested field.
`selector-wait.ts` reading `outcome.resolution.selector.raw` was the runtime
proof. The existing boundary gate reads exported *names*, so it could not see
this.

The public outcome now lives beside `SelectorResolution` in
public-resolution-types.ts with its selector as text; the parser-side shape is
renamed `AstPolicyResolutionOutcome` and stays package-private, and the façade
wrapper flattens on the way out — the same treatment `resolveSelectorChain`
already gave `AstSelectorResolution`.

Two new pins, both verified red against the shape they replace: a behavioral
one asserting the façade returns selector text under every policy row, and a
structural one asserting resolution shapes are re-exported from
public-resolution-types.ts rather than from a parser-side module — which is
what distinguishes the leak from a correct re-export in a name list.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Rva4YGtSCAKJqH5PbpcCU

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-06 21:10:30 +02:00
Michał Pierzchała 3937036e5e feat: support --settle on scroll and back (#1638) (#1650)
* feat: support --settle on scroll and back (#1638)

Scroll-then-observe and back-then-observe are legitimate agent pairs, but
the post-action observation registry never grew past the touch commands, so
`--settle` on either was rejected with INVALID_ARGS — burning a tool call
each in AppControlBench's bsky-16.

Both commands now carry the `settle` descriptor trait, and every surface
derives from it rather than a hand list: CLI allowed flags, MCP/SDK input
fields, the flag-sourced timeout envelope, and MCP ref-pinning. The CLI
flag/metadata helpers moved out of the interaction family into
post-action-observation-grammar.ts (back is a system command), and
SETTLE_REF_ISSUING_TOOLS became a derivation — a hand list would have
silently stopped pinning the new commands' refs.

settleAfterInteraction and the new settleObservationCommand are two entry
points over one engine: same loop, storage, hints, and diff bounds, with the
target-less path supplying its own baseline and no proximity point. The
daemon reaches that command through the runtime surface, never by importing
`commands/` (R2) — the same seam the touch handlers use for press/fill —
and generic-settle.ts is loaded through a lazy `await import` returning a
closure, so the interaction runtime subgraph stays out of this dispatcher's
static graph (a static edge folded ~18 files into the daemon-server type
cycle; R10 caught it).

Both of generic-settle's orderings are load-bearing and tested: the baseline
is frozen before dispatch (and before the Android dialog preflight), and the
observation runs after markDeferredInteractionOutcome so settle's first
capture folds in the #1542 stabilization rather than racing it. The ADR 0014
"a settled diff publishes refs" rule moved to settle-ref-issuance.ts, shared
by both routes.

One divergence is deliberate: scroll/back resolve no element, so the diff
baseline is the session's STORED pre-action tree — "settled tree vs the
last tree you observed" — not press's freshly resolved pre-action capture.

Both commands also switch to preserve-daemon on timeout, which changes the
non-settle path too: with --settle their dominant hang mode is now a wedged
accessibility bridge, and a timed-out capture must not reset the daemon and
lose every session (#1105). The reviewed-set gate records it.

Live-validated on an iOS 26.2 simulator (Settings): scroll --settle settled
in 1786ms with a +6/-6 diff carrying fresh refs; back --settle in 771ms with
+15/-6. Alternating cost runs, one call vs the pair it replaces:
scroll 2.9-3.0s vs 5.3-5.6s, back 3.1-3.2s vs 4.7-5.1s. Those include the
#1627 deep-capture extension.

* fix: render settled-diff refs paste-ready in CLI output

A settled diff activates a PARTIAL ref frame (ADR 0014), which admits only
the pinned `@eN~s<gen>` form of the refs it issued. The unchanged-interactive
tail already rendered that way, but the diff's own added lines rendered the
bare `@eN` embedded in the snapshot line — so a CLI caller who copied the
ref the diff just handed them got `plain_ref_requires_complete_frame` and had
to append the generation by hand.

Added lines now render pinned when the response carries `refsGeneration`,
exactly like the tail. Removed lines render verbatim: they name elements that
just left the screen, and `SettleDiffLine` never gives them a ref.

This is not new to scroll/back — press/click/fill/longpress had the same gap
since #1101. MCP was never affected: its ref-pin store rewrites plain refs on
the way in, which is why the model never sees a suffix.

Live: `scroll down --settle` now emits `+ @e14~s218078 [cell] "Game Center"`,
and `press @e14~s218078` copied straight out of that line taps successfully.

* test: record the pinned-diff-ref bytes in the output-economy baseline

Rendering added diff-line refs pinned costs 8 bytes in the two settle CLI
text samples (two `~s<gen>` suffixes). The output-economy baseline is the
tripwire for exactly this, so the increase takes an explicit reviewed waiver
rather than a silent baseline bump — the same one the settled TAIL's pins
already carry, for the same ADR 0014 reason.

Only `bytes` moves: lines, refs, hints, and shape are unchanged, which is the
evidence that this is a suffix on existing refs and not a new payload.

Caught by CI, not locally: `pnpm test:unit` runs unit-core and
subprocess-stub only, while the Coverage lane runs every vitest project.

* test: prove the generic settle degrades when its runtime cannot be built

`createGenericSettleRuntime` catches and returns undefined so an observation
that cannot even start does not fail an action that already succeeded. That
was a claim in a docstring with nothing behind it — the one changed line the
coverage gate reported uncovered (95/96).

The test puts the session in the state the catch exists for: the router
handed us a session that is no longer in the store, so building the settle
runtime throws SESSION_NOT_FOUND. The response keeps its scroll result and
simply carries no settle payload. Removing the try/catch fails it.

* build: teach fallow that vi.mock reaches pinOwnProcessStartTime dynamically

Not from this PR: #1642 added `pinOwnProcessStartTime` on main, and its three
consumers reach it the only way a Vitest module mock can —
`vi.mock(path, async (importOriginal) => (await import('...')).pinOwnProcessStartTime(...))`.
Dependency analysis cannot follow that dynamic import to a consumer, so the
export reads as dead the moment any PR pulls that file into its audit scope.
This PR is the one that did.

The entry records the consumers by path and the reason, matching the
daemon route-handler entry directly above it, which exists for the same
dynamic-`import()` limitation.

* refactor: adopt the best of the parallel #1653 implementation

Two sessions independently built #1638 (PR #1650 and PR #1653) and converged
on the same architecture — trait in the registry, one engine with two entry
points, runtime-command seam, lazy import, preserve-daemon, stored-baseline
honesty. #1650 continues; this folds in what #1653 did better:

- The agent-facing help core loop (cli-help.ts) now names scroll and back as
  settle-capable. Without this, the benchmarked closed-grammar help line kept
  instructing agents that --settle is only for press/click/fill/longpress —
  actively steering the AppControlBench models away from what #1638 shipped.
- issueSettleRefs moves into session-snapshot.ts, beside the partial-frame
  primitive it wraps, deleting the single-function settle-ref-issuance module.
- Their seam tests: back reader→writer settle plumbing, back CLI settle
  rendering, and a trait-less generic command (home) ignoring a stray settle
  flag rather than observing or rejecting.

What #1650 had that #1653 lacked, for the record: the SETTLE_REF_ISSUING_TOOLS
registry derivation (without it, MCP never pins a scroll/back settle diff's
refs and the partial frame rejects every follow-up), BackCommandResult.settle
in contracts, back's MCP output schema, paste-ready pinned diff refs, and the
docs/changelog/baseline surfaces.

* bench: help-conformance case for settled scroll-to-find planning

The #1638 extension of the closed --settle grammar to scroll/back is the
feature's entire payoff — collapsing scroll-then-observe into one call — and
the closed command list is an enumerated N whose enumerator is this bench.
The regex over the help text proves the sentence exists; this case checks
whether a model plans differently because of it.

One focused case, deliberately not coached: a pinned visible-first snapshot
(rendered by formatSnapshotText, pinned by the sample-producers gate) whose
wanted row is summarized off-screen with no ref anywhere in the output. The
tempting pre-#1638 plan is `scroll` plus a separate `snapshot -i`; acceptance
is the single settled call. Scoring was verified against eight plan shapes in
both directions before recording.

Model-backed record (claude-haiku-4-5, 3 trials, current help): 0/3 — but the
decomposition is the finding. Settle eligibility GENERALIZED (3/3 trials put
--settle on scroll unprompted; the mutation-suffix framing concern did not
materialize) and the two-call habit is residual (1/3). All three trials failed
on `scroll @e3 down --settle` — the pre-existing #1366 scroll-takes-no-target
confusion, which the live CLI recovers with a dedicated hint but a single-shot
bench cannot. The recorded gap is therefore a first-30 doc gap (nothing
teaches that scroll takes no target), not a settle-eligibility gap; tuning the
case until it passes would just delete the evidence.
2026-08-06 20:04:16 +02:00
Michał Pierzchała 58a9d4bf71 refactor(daemon): consolidate surface-evidence helpers, delete the wording heuristic (#1615)
Re-derived onto #1633's split (`post-gesture-stabilization.ts` became
`post-gesture-stability.ts` + `deferred-interaction-outcome.ts` +
`gesture-no-effect.ts`). None of this had been subsumed by that refactor —
it relocated the code and carried every one of these forward untouched.

- Three copies of the four-field rect comparison in
  `interaction-outcome-policy.ts` become one `rectsWithinTolerance`.
- `identifiedContent` returns the entry instead of `{ entry }`, dropping
  the `.entry` indirection at every call site.
- `haveIdenticalDiscriminatingSurfaces` records why it keys on `key` while
  `classifyBaselineSurfaceEvidence` keys on `identity` — opposite choices,
  once, at the function that makes the stricter one.
- `SnapshotCaptureBackend` names the capture-strategy union in
  kernel/snapshot beside `SnapshotBackend`, and
  `PostGestureStabilization.baselineBackend` uses it instead of `string`,
  closing the silent-typo gap on the comparison the field exists for.
  Deliberately NOT applied to `post-gesture-stability.ts`: #1633 made that
  module generic over the surface type, and `backend: string` is right for
  an interface that must not know about iOS capture strategies.
- `formatGestureNoEffectWarning` echoes positionals verbatim. The
  `/^[\d.-]+$/` filter it replaces ate all four coordinates of
  `swipe <x1> <y1> <x2> <y2>` and emitted a contentless bare "swipe"; the
  warning names the gesture the agent issued, and `scroll down 1` is what
  they issued.
- `PostGestureStabilization.positionals` is required — its only writer
  always sets it, so the `?? []` at the read site guarded an impossible
  state. Tightening it caught six test fixtures building the state
  directly.

Red evidence: restoring the numeric filter fails the wording test
("scroll down 1 produced no visible change"); 95 files / 757 tests green
with it deleted.
2026-08-06 17:27:18 +02:00
Michał Pierzchała 870d12c406 fix: honest find contract — press/tap aliases, read-only list action, selector uniqueness (#1637)
* fix: honest find contract — press/tap aliases, read-only list, selector uniqueness (#1625)

Three defects in find's contract, fixed together because they are one
vocabulary:

press/tap are the same action as click everywhere else in this CLI, yet
find rejected them — agents using the vocabulary the tool itself
established burned a tool call per attempt (four in one bench run).
Both parsers now normalize press/tap to click; longpress/swipe stay
real exclusions.

The #1602 recovery hint told agents to run bare find to 'list matches',
but bare find CLICKS a unique match — inspection guidance pointing at a
mutation (the #1625 report: 'find Dictionary' navigated into
Dictionary). find <q> list is the read-only surface that guidance
needed: every match with its @ref, unique match included, never a tap.
Captured UNSCOPED (the label-scope optimization narrows to the first
match, exactly wrong for listing), published as an ADR 0014 partial
frame authorizing every listed ref.

Selector-shaped queries skipped the ambiguity check and took the first
match silently — the mis-binding path the AMBIGUOUS_MATCH recovery
advice itself pointed agents at, while --first/--last were documented
as explicit opt-ins. Selector and text queries now share one contract:
multiple matches reject with the #1597 candidates listing unless
--first/--last narrows explicitly.

The hint is rewritten around the new contract; docs and the MCP find
output schema follow. Regressions at every layer: both parsers (alias,
list token, unsupported-action hint shape), the daemon handler
(selector ambiguity with candidates, --first opt-out, list returns all
matches with zero action dispatches, unique-match list does not tap).

* refactor: single-home the find read result and flatten parseFindArgs (fallow)

The daemon's DaemonFindResult had drifted into an identical structural
twin of the engine's FindReadCommandResult — the two grew the list
variant in parallel and crossed the clone threshold. The shape now
lives in contracts as FindReadResult (below both zones, per R2's own
remedy) with the engine and daemon both aliasing it.

parseFindArgs collapses the four bare single-token actions into one
membership check and extracts the get sub-action parser, bringing it
back under the complexity threshold instead of waiving it.

* style: merge duplicate contracts import (lint)

* fix: accept list on the MCP input surface and pin every listed ref (review)

- FIND_ACTION_VALUES gains 'list' so field-metadata/MCP input no longer
  rejects the action the CLI parser accepts
- FindCommandResponseData types 'matches' (public client response)
- MCP mergeFindRefPins learns every matches[] ref, so a plain @eN press
  after find-list forwards pinned and the partial frame admits it
- CLI/MCP text renders every listed match as its own pinned line via the
  snapshot-line role/label normalizers
- regressions: daemon partial-frame scope, pin store, CLI output, MCP
  schema + find-list->press chain, typed client list response
2026-08-06 15:11:12 +02:00
Szymon Dziedzic 3e4828d68d feat: add scale-only screenshot sizing (#1617)
* feat: add scale-only screenshot sizing

* fix: refuse retired --max-size inputs on every released surface

Released sizing inputs must fail closed with migration guidance instead of
silently producing native-size artifacts:

- contracts: RETIRED_SCREENSHOT_MAX_SIZE declaration + SCREENSHOT_SCALE_LIMITS
  as the single source for the scale bounds and migration messages
- .ad parser: released 'screenshot ... --max-size N' and 'record start ...
  --max-size N' lines now refuse at parse time (frozen replay-compat witnesses)
- daemon: screenshot rejects old-client screenshotMaxSize like recording does;
  the recording guard now shares the same contract data
- Node client: screenshot/record daemon writers refuse the removed { maxSize }
  option before transport
- CLI: --max-size unknown-flag error carries the migration guidance
- config/env: stale screenshotMaxSize config keys and the retired
  AGENT_DEVICE_SCREENSHOT_MAX_SIZE env var are refused for sizing commands
  (other commands keep working)

Quality: numberField now reuses the canonical readOptionalNumber contract
helper (AppError bounds instead of plain Error); png-resize inlines one-use
wrappers and restores the worker-thread rationale; docs typo fixed.

* test: drop retired maxSize entries from the MCP undocumented-input allowlist

* fix: refuse retired maxSize at the MCP field-projection seam + release-provenance corpus witnesses

- readFieldInput silently dropped undeclared keys before the daemon writers
  could refuse them, so an MCP call carrying { maxSize } reached transport and
  returned native-size success. New retiredField() combinator declares the
  removed key in the field map: the projection seam refuses it with the
  canonical migration message and the JSON schema no longer advertises it.
  Real-route MCP executor regressions cover screenshot and record.
- replay-compat corpus: derived v0.20.5 witnesses for the released screenshot
  and record --max-size forms (SHA-256 pinned, new retired-capture-size
  coverage surface) so check:replay-compat proves the shipped syntax refuses
  with migration guidance instead of degrading silently.

---------

Co-authored-by: Michał Pierzchała <thymikee@gmail.com>
2026-08-06 15:11:00 +02:00
Michał Pierzchała f93b259d15 fix: stop the slow-snapshot warning from firing on a single cold start (#1628)
* fix: stop the slow-snapshot warning from firing on a single cold start

The session's first capture folds one-time startup (runner launch,
helper install) into its duration, and nearest-rank p95 over a small
sample set equals its largest one or two values — so one 12s cold start
produced 'snapshots are slow in this run: p95 12417ms over 1 captures'
with hints blaming device load or a stale daemon, inviting exactly the
restart spirals the hints exist to prevent (observed on every
AppControlBench run).

The warning now judges only warm captures (first sample excluded) and
only once at least three exist; the displayed stats still cover every
sample, so the cold start remains visible as maxMs.

* refactor: single-home the slow-snapshot warning policy (review)

Push the warm-judging rule down into summarizeSnapshotTimingSamples so
the interactive session path and all three replay handler paths share
one policy, and summarizeSnapshotDiagnostics returns to a one-line
delegate. Merge no longer judges slowness from lossy order-less
reconstructed samples (a run's cold start comes back as both its p95
and max): it aggregates display stats and carries a warning only when a
constituent run judged one itself. MIN_WARNING_SAMPLE_COUNT renamed to
MIN_WARM_SAMPLE_COUNT — it gates warm samples, not total captures.

Cold-start regression tests move to the shared layer; suite-aggregation
tests now pin that individually-silent runs merge silent.

* style: oxfmt

* fix: quorum-gate the warm warning and make the merged message speak about warned runs (review)

A single warm outlier could still fire the warning (nearest-rank p95 is
the maximum through nineteen samples): chronic now additionally requires
at least two slow warm captures. And the merged warning formatted its
number from the reconstructed aggregate, so one slow run among many fast
ones produced 'slow: p95 <fast number>' — the merged message now reports
how many runs warned and the worst warned run's own p95, never the
aggregate. Regressions for both: one-warm-outlier stays silent; a slow
run merged with many fast ones warns with the slow run's number while
the aggregate p95 sits below the threshold.
2026-08-06 13:52:20 +02:00
Michał Pierzchała a67c72c211 fix(ios): pin tap-outcome corroboration probes to the baseline's backend (#1634)
* fix(ios): pin tap-outcome corroboration probes to the baseline's backend

The recorded-failure screens are exactly where the capture plan flips
between XCTest and private-AX (the penalty boundary), so #1605's
same-backend requirement failed closed right where XCTest tap false
negatives actually happen: the baseline was captured via private-AX
under penalty, the probe came back via tree, and a landed tap surfaced
as XCTEST_RECORDED_FAILURE. In the AppControlBench bsky-16 run this
fired four times, each sending the model into a re-observe/retry spiral.

The comparison stays same-backend by design (backends are not comparable
views of a screen); instead the probe is now CAPTURED the way its
baseline was: a new internal preferredBackend option (never CLI-exposed)
threads daemon -> runner, and a private-AX-preferred capture takes the
exact penalized route — privateAX-first plan, 'deferred' verdict, no
degradation warning, no settle budget reset.

Live-verified on the deterministic repro (Bluesky drawer-menu press
under penalty, seeded bench feed): errored with the backend-mismatch
diagnostic before, corroborates as landed after, with no mismatch phase
in the request diagnostics. Daemon tests cover pinned and unpinned
baselines end to end through the dispatch context; the Swift plan gate
is a pure function with an executed in-bundle test (added to the ios.yml
regression list).

* style: oxfmt

* fix: exclude raw baselines from corroboration and prove the pin end to end (review)

Raw baselines could not be pinned: the raw diagnostic plan keeps
tree-first error propagation by contract and is never rerouted by the
penalty or the preferred backend, so preserving 'raw: true' on the probe
recreated exactly the backend-mismatch false failure this PR removes.
Corroboration now declines raw baselines up front (they are diagnostics,
not evidence baselines) with a regression pinning that no probe capture
is dispatched at all.

The wire is now regression-proven at every hop: a dispatch-level test
drives dispatchCommand with the context flag and asserts the emitted
RunnerCommand carries preferredBackend (red if handleSnapshotCommand or
the interactor stops forwarding); the injected-transport test asserts
the interactor's snapshot payload both ways; and a runner unit test
decodes the wire JSON, projects it through the extracted
snapshotOptions(from:), and composes it with the plan rule — pinned
regular plan defers to privateAX-first, RAW plan stays untouched.
Executed on-simulator; added to the ios.yml regression list.
2026-08-06 13:27:30 +02:00
Michał Pierzchała d5f99bab1c refactor: sink backend.ts's cycle-closing types below both zones (#1632) (#1636)
backend.ts imported RepeatedInput up from commands/command-input.ts and
ScreenshotResultData up from utils/screenshot-result.ts — the interface hub
typed in terms of the zones that depend on it, R6's textbook inversion shape.

- RepeatedInput now lives in @agent-device/contracts/interaction;
  command-input.ts re-exports it for its existing importers.
- ScreenshotResultData already had a byte-identical canonical declaration in
  contracts/snapshot-types.ts (exported via contracts/capture); the utils
  copy is now a re-export of it, deleting the duplicate outright.

Measured member-by-member: the R9 type cycle collapses 76 -> 49 files.
backend.ts, runtime-contract.ts, commands/runtime-types.ts, and
commands/runtime-common.ts all leave the component (27 files stranded out at
once); zone ceilings lowered to the measured values (commands 33 -> 14,
platforms 7 -> 2, root 5 -> 3, daemon-server 20 -> 19) and CONTEXT.md's hub
list recomputed (core/dispatch.ts 8, command-catalog.ts 7, resolution.ts 6,
command-descriptor/registry.ts 6). No TYPE_INVERSION_BASELINE additions.
2026-08-06 13:25:07 +02:00
Michał Pierzchała 8c800ae53f refactor(contracts): one viewport-root predicate for the whole repo (#1613)
* refactor(contracts): one viewport-root predicate for the whole repo

"Is this the Application/Window root" was written nine times: three
spellings normalizing `type|role|subrole`, five lowercasing `type` alone,
and one comparing the normalized type for EQUALITY. Two of the nine sat
in `contracts/snapshot-visibility.ts` itself, disagreeing with each other.

Measured before collapsing, using #1592's method — ground the comparison
in what each backend ACTUALLY emits, not in fixture strings. Over the 31
names iOS's `elementTypeName` can return, the 18 fully-qualified class
names Android emits, and the 24 mapped/raw forms the macOS helper
produces, the nine agreed on 71 of 73. The two exceptions are macOS
window subroles, and the only spelling that disagreed is maestro's `===`,
whose platform union is `android | ios` — so it can never see them. The
duplication was textual, not behavioral, which is what made the collapse
safe.

`isViewportRootNode` reads role and subrole because the macOS helper is
the only backend populating them and the only one able to emit a window
whose `type` does not say so: `normalizedSnapshotType` returns the raw
subrole for a non-standard window, so an `AXWindow` with subrole
`AXSystemDialog` or `AXUnknown` reads as neither from `type` alone. Those
two shapes are the whole behavioral delta of this change, at the six
call sites that were type-only, and they are windows by role.

`snapshot-viewport-root.test.ts` pins the predicate over those three
emitted vocabularies. Red evidence: reverting the canonical definition to
the type-only spelling fails 2 of 5 cells, to the equality spelling 4 of 5.

Also drops two kernel re-declarations this made visible: maestro's local
`containsPoint` and `rectsOverlap` were character-identical to
`@agent-device/kernel/rect`'s `containsPoint` and `isRectVisibleInViewport`,
in a file that already imports from that module. And `resolveViewportRect`
loses three `as Rect` casts that only existed because `.filter()` cannot
narrow `node.rect` — one `flatMap` states the same thing honestly.

Deliberately NOT in this change: the three viewport RESOLVERS still
diverge, and on Android that is a live defect rather than duplication.
Filed separately with the measurement.

* test(contracts): enumerate the macOS emitter's real vocabulary

Review found the table claimed to pin "the vocabulary each backend actually
emits" while omitting most of it. `normalizedSnapshotType` has three output
classes and only two were represented:

  1. thirteen roles mapped to fixed short names — six were missing
     (StaticText, TextField, TextArea, MenuBarItem, Menu, MenuItem);
  2. AXWindow, whose output is the SUBROLE unless it is AXStandardWindow;
  3. the `default:` arm, `subrole ?? role`, emitting the raw AX-prefixed
     value for every unmapped role.

All three are now enumerated, and the table asserts its own completeness
against the emitter's fixed-output set — a role added to that switch without
being added here fails, which is the emitter-drift protection the docblock
was promising but not delivering.

Re-measuring over the complete tables also corrected the header's own
numbers. The claim was "71 of 73 agree, 2 disagree"; over 75 names it is 71
agree and FOUR disagree, because AXSystemDialog and AXUnknown were absent
from the old table. Those two are the behavioral delta of this PR — an
AXWindow whose subrole is emitted as the type, invisible to the six
type-only spellings and named exactly by `role` — so the incomplete table
had been hiding the very rows that justify reading role/subrole. The other
two (AXFloatingWindow, AXSystemFloatingWindow) remain inert: only the `===`
spelling misses them and its platform union is `android | ios`.

* test(contracts): derive the macOS fixed-output set from the emitter

Two test-validity defects from review, both real.

The raw-fallback row `{ type: 'AXSearchField', role: 'AXTextField', subrole:
'AXSearchField' }` was unreachable: the `AXTextField` arm returns `TextField`
whatever the subrole, so no emitter run can produce it. Replaced with
`{ type: 'AXSortButton', role: 'AXCell', subrole: 'AXSortButton' }` — a
subrole on a genuinely unmapped role, which is what the `subrole ?? role`
default arm actually emits.

`MACOS_FIXED_OUTPUTS` was a hand-kept twin compared against a hand-kept
table, which is circular: a new mapped Swift role is absent from BOTH, so
they agree and the gate stays green. The "emitter-drift protection" the
docblock promised did not exist. The set is now parsed out of
`normalizedSnapshotType` in SnapshotTraversal.swift, so the comparison is
against the emitter rather than against a copy of the table's own
assumptions. `case "AXWindow"` returns a subrole expression rather than a
literal and is deliberately outside the literal-return set.

Red evidence: adding `case "AXDisclosureTriangle": return "DisclosureTriangle"`
to the Swift switch fails with `expected [ 'DisclosureTriangle' ] to deeply
equal []`; 6 pass once reverted. The parser throws rather than silently
matching nothing if the function is renamed or moved.

* chore: restore maestro conformance corpus to main

45 corpus YAMLs carried an unrelated quote-style churn ("Button" ->
'Button'). They were already modified in the worktree when this branch
started and a `git add -A` swept them into the predicate commit. Nothing
in this PR reads them. Restored verbatim to main.
2026-08-05 18:16:58 +02:00
Michał Pierzchała d8b309c6db refactor(contracts): name façade exports explicitly and retire the pin table (#1614)
* refactor(contracts): name façade exports explicitly and retire the pin table

Thirteen of the fourteen `@agent-device/contracts` façades were bare
`export *` barrels. `facades/snapshot.ts`, added by #1582, was the one
exception — explicit named re-exports — and that is now the rule.

Everything #1574 built to cope with `export *` goes with them:

  scripts/layering/facade-symbols.ts          -980   (816 pinned names)
  scripts/layering/facade-exports.ts          -192   (readFacadeExports)
  scripts/layering/facade-exports.test.ts     -234   (star semantics)
  scripts/layering/package-boundaries.test.ts  -55

`readFacadeExports` re-implemented ESM `GetExportedNames`/`ResolveExport`
— star-chain resolution, ambiguity rejection, diamond binding identity,
cycle guards, spec-accurate `default` filtering at the star rather than
the source. All of it existed to enumerate what `export *` hides. 523 of
the 816 pinned names belonged to contracts, i.e. to those thirteen files.
Once a façade names its exports, the façade file IS the pin, and it is
visible in the diff of the file that widened rather than in a separate
table a reviewer has to cross-check.

`readNamedExports` (20 lines) stays and is enough: it already throws on
bare `export *` and on `export default`. The pin is replaced by one
structural gate — no façade may contain a bare star — which reuses that
rejection rather than adding a regex.

Surface equivalence verified independently, not asserted: main's own
`readFacadeExports` run over the new façades, compared against main's own
`FACADE_SYMBOLS` table — 31 subpaths, 0 added, 0 removed.

Red evidence for the new gate: planting `export * from '../request-progress.ts'`
back into facades/progress.ts fails it with the file named and the reason
quoted; 12 pass / 0 fail once reverted.

Not included: the `lowerAndroidTouchPlan` tuple-assertion drive-by. It
needs `sampleGestureOffsets` to carry a min-arity tuple through `.map()`,
which TypeScript will not infer without a typed helper — a real change to
the gesture-plan contract rather than a drive-by, so it stays out.

* test(layering): assert façades stay exhaustive over their sources

Review on #1614 caught this conversion silently narrowing the public
surface. The explicit lists were generated against the surface at fork
time; #1567 landed 13 exports meanwhile — `DragOptions`, the drag-gesture
vocabulary (`COORDINATE_GESTURE_KINDS`, `CoordinateGesturePayload`, the
three `DEFAULT_DRAG_*` constants, `DragGestureInput`, `DragGesturePayload`,
`GestureCommandInput`, `buildDragGesturePlan`,
`dragGesturePayloadFromPositionals`, `normalizeGestureCommandInput`) and
`MultiTargetAnnotationV1`. The `export *` barrels had been forwarding all
13 automatically; the rebase dropped every one, and only a human diff
caught it.

The star-rejection gate could not: it only proves a façade does not WIDEN
invisibly. Narrowing is the failure an explicit list newly makes possible,
because `export *` could not narrow by construction. So the property the
stars gave for free is now asserted directly — every name a re-exported
source declares must appear in the façade.

Scoped to `packages/*/src/facades/`, the barrels this PR converted. A
hand-curated package `index.ts` is a different thing: `ad-replay`
deliberately publishes two values out of a much larger `internal/`, and
forcing exhaustiveness there would widen a surface its owner narrowed on
purpose (#1555). A source that itself carries a bare `export *` is skipped
— unknowable from that file alone, and reachable because the façade
re-exports the starred module directly too, which IS checked.

Red evidence: dropping `MultiTargetAnnotationV1` from facades/replay.ts —
one of the 13 the old gate was blind to — fails with the file, the source
and the symbol named. 13 pass / 0 fail once restored.

* fix(layering): close the exhaustiveness gate's starred-source hole

Two review findings, plus a third the gate caught on itself.

P1 — the three `DEFAULT_DRAG_*` constants join the existing public-façade
suppression, alongside `COORDINATE_GESTURE_KINDS` and
`normalizePublicGesture` which the same conversion surfaced. All five are
#1567's drag vocabulary, made individually visible to `--production`
analysis for the first time because a bare star used to hide them from
that exact check. Kept rather than narrowed, for the reason the existing
entry already states: the façade's surface stays byte-identical to what
the retired pin table asserted, and narrowing is a follow-up with its own
review.

P2 — the exhaustiveness gate skipped any source carrying a bare
`export *`, which dropped that module's DIRECT exports from the check too.
`gesture-plan.ts` stars `gesture-plan-types.ts`, so removing
`buildDragGesturePlan` from the façade narrowed the public surface and
still passed. `readDirectNamedExports` now reads exactly the names a module
declares or re-exports BY NAME and ignores the star, so direct exports are
checked while the starred set stays covered by the façade's own direct
re-export of that module.

Red evidence: removing `buildDragGesturePlan` from facades/interaction.ts
now fails naming file, source and symbol; 13 pass / 0 fail restored.

Third, and the reason the gate is worth having: rebasing onto main after
#1612 merged silently dropped `TEXT_ENTRY_ROUTES`, `TextEntryRoute` and
`TypeTextBackendResult` from the interaction façade — the same narrowing
class as the #1567 one review caught by hand, one merge later. The gate
failed on it before CI did. Restored.
2026-08-05 15:58:02 +02:00
Michał Pierzchała ee473b6adc refactor(daemon): give the Maestro fallback and ambiguous-match details real types (#1612)
Three places smuggled structured data through untyped bags and re-read it
with runtime guards. Each gets an explicit typed boundary.

A. The resolution-suppression rule was encoded twice in
   interaction-touch-response.ts — a spread ternary in the runner-payload
   branch and an unconditional destructure used conditionally in the runtime
   branch, with the ADR 0012 rationale living on only one source variant.
   Both branches now read one `suppressesResolutionDisclosure(source)`
   predicate through one `applyResolutionDisclosurePolicy` helper, where the
   reason is stated once. The union field is renamed
   `maestroCoordinateFallbackDispatched` (the dispatch path that ran) and
   hoisted into a shared base. handleFillCommand's two-arm interactor.fill
   call collapses to one.

B. `Interactor.type` narrows from `Record<string, unknown> | void` to
   `TypeTextBackendResult | void`; the Apple runner boundary is the single
   place the wire payload becomes that type. `maestroFallbackDetails` returns
   a typed `{ used, extra }` instead of a bag both call sites re-read.

C. `details.candidates` meant two incompatible things. The device-domain
   resolvers now key their list `devices`, so the shared renderer drops its
   shape-disambiguation guards and the device list actually renders.
2026-08-05 14:32:25 +02:00
Thiago Brezinski a13a6832ee feat: add selector-targeted drag gestures (#1567)
* feat: add selector-targeted drag gestures

* fix: address drag gesture review feedback

* fix: satisfy drag review quality gates

* fix(android): lower drag trajectories piecewise

* test(replay): validate drag fixture selectors

* fix(ios): ignore full-viewport chrome containers

* test(drag): prove destination on live devices
2026-08-05 12:37:02 +02:00
Michał Pierzchała 611858103e fix(ios): harden Bluesky-class interaction reliability (#1588)
* fix: type into focused iOS inputs without AX

* fix: fill AX-hostile iOS text inputs

* fix: keep scrolling containers from stealing taps

* fix: stop agents after explicit task success

* chore: format benchmark guidance

* fix(ios): preserve fill semantics across fast paths

* test: retire direct selector fill expectations

* test: assert runtime selector fill evidence

* fix(ios): preserve verified and Maestro fill paths

* refactor(ios): isolate synthesized text entry

* fix(client): preserve open diagnostic paths

* fix(ios): expose structured text entry route

* fix(packaging): strip text entry policy tests
2026-08-05 08:02:44 +02:00
Michał Pierzchała 20e903c117 fix(maestro): unify the scrollable-ancestor walks and fix Android scroll-container selection (#1592)
* refactor(maestro): collapse the duplicate scrollable-ancestor walk

fallow reported three structurally identical "walk up the parent chain to
the nearest scrollable ancestor" implementations as clone groups
(dup:1b401a24, dup:ce01e1de). Two of the three predicates classify
identically, one does not.

snapshot-policy.ts's isScrollableNode is logically identical to contracts'
isScrollableNodeLike -- same six type patterns, same `=== 'table'`
equality, same role/subrole fallback, and neither normalizes the type
first. The walks match too, so findScrollableAncestorRect collapses onto
findNearestScrollableAncestor with `(n) => Boolean(n.rect)`.

runtime-port-geometry.ts's isScrollableSnapshotType does NOT agree. It
equality-matches the NORMALIZED type, so over 227 node-type strings
harvested from the repo's fixtures and tests it disagrees in both
directions: Android ListView/GridView/RecyclerView, HorizontalScrollView,
AXScrollBar and role-only scrollables clip but are not swipe containers,
while XCUIElementTypeTable and AXTable are swipe containers but do not
clip (the clip predicate compares 'table' against the unnormalized type,
so prefixed forms miss). Only bare `table` satisfies both. It stays
separate, with the divergence and the reason each call site needs its own
answer written down where it can be read.

The new test is load-bearing rather than decorative: replacing
isScrollableSnapshotType with the contracts predicate leaves
`pnpm maestro:conformance` at 46/46 and the pre-existing maestro suite at
206/206 green. The oracle does not cover scroll-container selection, so
nothing else in the repo fails on that collapse.

resolveRootViewport is deliberately left alone -- it resembles contracts'
resolveViewportRect but lacks its third "largest containing rect of any
node" fallback, so it is a real divergence and not the next dedup.

* fix(maestro): recognize Android scroll containers when aiming scrollUntilVisible

The divergence note added in the previous commit was wrong, and it was
covering for a bug rather than describing a design.

Grounding the comparison at the call site instead of in fixture text
changes the answer. `node.type` is never normalized on the way in -- it
carries the raw platform string -- so the domain of each predicate is
exactly what each platform emits:

  iOS    `elementTypeName` returns 31 fixed short names ("Table",
         "ScrollView", "CollectionView", ...), never "XCUIElementType*".
  Android `attrs.className`, fully qualified.
  macOS   role-mapped short names; outside Maestro's platform union.

Over all 31 iOS names the two predicates agree on every single one. The
claimed `XCUIElementTypeTable` / `AXTable` divergence was measured on
strings the runner cannot produce; the real emission is "Table", which
both predicates accept. The role/subrole arm is macOS-helper-only, so it
is inert for Maestro entirely.

What remains is Android, one-directional, and a defect: matching a
normalized type for EQUALITY recognizes bare `android.widget.ScrollView`
and silently misses HorizontalScrollView, NestedScrollView, RecyclerView,
ListView and GridView. `scrollUntilVisible` therefore selected no
container and fell back to a screen-centred swipe inside essentially
every RecyclerView-backed list -- contradicting the function's own
documented intent, and contradicting the existing Android test that
expects `android.widget.ScrollView` to be selected.

So the third walk collapses onto the shared helper too: the substring
predicate is also the better fit for Android's open class-name space,
where an allow-list would keep missing NestedScrollView and every custom
subclass. All three walks now share
`@agent-device/contracts/snapshot`, and the explanatory comment is gone
because there is nothing left to explain.

The test is rewritten to pin the classification over the vocabulary each
platform actually emits, with the Android rows as the regression guard.
2026-08-05 07:58:15 +02:00
Michał Pierzchała 543e9f8c05 fix(ios): give keyboard dismiss a safe-area-tap fallback (#1598) (#1606)
* fix(ios): give keyboard dismiss a safe-area-tap fallback (#1598)

The runner already tapped a keyboard's own Hide/Dismiss/Done key when the
AX tree exposed one, but iPhone's default software keyboard has no such
key, so `keyboard dismiss` returned UNSUPPORTED_OPERATION on the common
case and agents proceeded with the keyboard (and any live QuickType
predictive-text bar) still up.

Live-validated on throwaway simulators before choosing a design: hardware
escape key (no effect without a connected hardware keyboard), swipe-down
starting on the keyboard (does not trigger UIKit's interactive dismissal
on Settings/Safari/Contacts), and a private
`performAction:onElement:value:error:` AX call (hung the runner for 90s on
a guessed action name, force-killed by the daemon timeout) were all ruled
out. The dismiss-key tap remains the primary mechanism (iPad, or any app
with an inputAccessoryView Done/Cancel button); a new snapshot-derived
safe-area tap is added as the disclosed last resort, computed to land
outside both the keyboard and every currently-hittable element so it is a
safe no-op even when it fails to dismiss.

The response now discloses which mechanism actually fired
(`mechanism: 'dismissKey' | 'safeAreaTap'`) across the CLI/daemon dispatch
path, the SDK runtime.backend surface, and session-event summaries, so
callers can tell a real dismiss-key press apart from a best-effort tap.
UNSUPPORTED_OPERATION now says both mechanisms were tried.

* fix: satisfy CI formatting and complexity gates

oxfmt on three touched files; buildKeyboardActionSummary split so the
dismiss wording (incl. the safeAreaTap mechanism disclosure) lives in its
own helper below the complexity threshold.

* fix: any-element obstacle rule for the safe-area dismiss tap (#1606 review P1)

A role allowlist cannot prove a point is AX-empty: an unlabeled RN
Pressable surfaces as a hittable Other, and a tappable parent can cover a
point its static-text child does not. Every known element frame now counts
as an obstacle regardless of role or hittability, with only ~window-sized
structural frames exempt (isStructuralRootFrame, 95% coverage) — exempting
those is what keeps the rule satisfiable, and a genuinely tappable
full-screen backdrop staying exempt is the disclosed, accepted behavior of
this fallback. One .any resolution replaces ten typed queries (single tree
snapshot, no per-element isHittable round trips), so the stricter rule is
also cheaper.

* fix: drop the safe-area tap — background-tap dismissal is unsupported (#1606 review P1, round 2)

No geometry or role query can prove a coordinate is side-effect-free: after
the any-element rule, the structural-root exemption still deliberately
removed full-screen actionable elements (RN Pressable backdrops) from the
obstacle set, so the tap could navigate or submit — and report success
because the mutation hid the keyboard. Per review, generic background-tap
dismissal is now explicitly unsupported: the dismiss key is the only
mechanism the runner vouches for, UNSUPPORTED_OPERATION says so and steers
callers to press-the-next-target / keyboard enter, and the mechanism field
narrows to 'dismissKey'. Unrecognized wire mechanisms degrade to the bare
message and are dropped from event details.
2026-08-04 22:30:04 +02:00
Michał Pierzchała 8ba5f9b8de fix: surface AMBIGUOUS_MATCH candidates and name find's supported actions (#1602)
* fix: surface AMBIGUOUS_MATCH candidates and name find's supported actions (#1597)

AMBIGUOUS_MATCH errors now list the matching candidates (ref, role,
label/identifier) rendered the same way as snapshot -i lines, capped at
5 with a "+N more" marker. buildAmbiguousMatchError (the single
producer, src/daemon/handlers/find.ts) reuses formatSnapshotLine to
build the list; formatAmbiguousMatchCandidateLines (src/utils/output.ts)
renders it unconditionally on both text surfaces an agent actually
reads (CLI printHumanError and MCP formatToolErrorText) — previously
the candidates lived only in details, which neither surface printed.

find's "Unsupported find action: X" (e.g. from `find <text> press`)
now attaches a hint naming every action find actually supports and the
two-step recovery shape: run find "<text>" to resolve the ref, then
dispatch the gesture as its own command (press @eNN). The hint is a
single exported constant (UNSUPPORTED_FIND_ACTION_HINT) shared by both
throw sites — packages/selectors' raw-token parser and the CLI's typed
reader (src/commands/interaction/selectors.ts) — so they can't drift.

Matching semantics are unchanged; ambiguous rejection stays by-design.
The help-conformance corpus's AMBIGUOUS_MATCH quiz is updated: its
premise ("candidate refs were not shown") no longer holds, but with 3
identically-labeled candidates the lesson (don't guess a specific ref)
still holds.

* fix: guard the AMBIGUOUS_MATCH candidate renderer against device-domain shapes

Review on #1602 (P2): formatAmbiguousMatchCandidateLines ran for every
normalized error and stringified details.candidates unconditionally,
but device-domain AMBIGUOUS_MATCH/APP_NOT_INSTALLED errors
(findBootedAppleSimulatorWithApp, src/core/dispatch-resolve.ts) reuse
that key for { id, name } device objects with no `matches` field —
CLI and MCP would have printed "Candidates: [object Object]" for
those. The renderer now requires numeric details.matches AND every
candidate to be a string before rendering anything, restricting it to
buildAmbiguousMatchError's element-match shape; unrecognized shapes
render nothing, same as before this feature existed. Added regression
tests against the exact device-error shape on both text surfaces.

Also unexports AMBIGUOUS_MATCH_CANDIDATE_LIMIT (fallow flagged it as
an unused production export) — it has no consumer outside find.ts.
2026-08-04 21:10:05 +02:00
Michał Pierzchała 4f8dc3f31e refactor: move selector engine into workspace package (#1589)
* refactor: move selector engine into workspace package

* refactor(selectors): trim the package façade to its real consumers

Follow-up to the selector-package cutover, from a structural review of it.

- Drop 15 façade symbols with no consumer anywhere in the repo:
  selectorUsesKey (added by the cutover, never called), isNodeVisible /
  isNodeEditable (the real helpers are contracts/snapshot's), normalizeText,
  splitIsSelectorArgs, IS_PREDICATE_REQUIRED_MESSAGE, four nested Replay
  types, SelectorDisambiguationDisclosure, and the four kernel type
  re-exports every consumer already imports from kernel directly.
- Delete SelectorCapturePolicyInput.selectorExpression, which
  deriveSelectorCapturePolicy never read; the policy varies only by
  predicate, so it takes one now. Two of the four tests asserted that the
  unread parameter had no effect and could not fail; they go with it.
- Return the Maestro export vocabulary to the maestro package. The cutover
  inlined MAESTRO_TEXT/STATE_SELECTOR_KEYS' values into the CLI call site,
  leaving both constants dead in the package that owns the concept and no
  gate over the two copies. MAESTRO_SELECTOR_PROJECTION is now the one
  statement of it.
- Dedupe SelectorDiagnostics and SelectorDisambiguationDisclosure, declared
  character-for-character twice across the AST/string seam, and name the two
  shared option shapes once instead of five inline copies. The parser-side
  resolution types take an Ast prefix so the twins read as twins.
- Delete three identity wrappers: parsePrivateSelector,
  selectorExpressionToMaestro, and the formatSelectorFailure forwarder —
  nothing passes it a chain any more, so the SelectorChain | string union
  and its branch go too.
- Delete internal/index.ts, an AST barrel whose only consumer was one test
  in the same directory (renamed to engine.test.ts), and the match.ts
  pass-through that existed to feed it.
- ReplaySelectorGrammar had three variants for two behaviors; 'wait' and
  'ordinary' were the same path. It is 'is' | 'positional' now.
- Drop the deleted src/sdk/selectors.ts from .fallowrc.json's entry list.

Behavior unchanged. pnpm check green: 598 unit files / 5278 tests, smoke
35 passed / 3 live skipped, layering 71/71, depgraph 22/22, mutation config
45/45, fallow clean, package smoke sound. Counterfactual: pointing
MAESTRO_SELECTOR_PROJECTION.textKeys at the state keys turns three
replay-maestro-export cells red; restored before commit.

* test(selectors): split the engine aggregation test by source concept

`internal/index.test.ts` (renamed `engine.test.ts` when its barrel went away)
was a 708-line aggregation over the whole engine — past the 500-line tripwire
and mirroring no source module, so it also ran as one serial unit.

It becomes five files that each mirror what they test, plus the parser cells
folded into the existing parse test:

  resolve.test.ts                 alternative fallback, strict uniqueness,
                                  first-match existence
  resolve-disambiguation.test.ts  ADR 0012 ranking: deepest, smallest-area,
                                  winner-vs-challenger disclosure, tie fallback
  resolve-viewport.test.ts        the visibility half: on-screen beats
                                  off-screen, including inside an off-screen
                                  scroll container
  match.test.ts                   per-key matching semantics (text, role,
                                  focused, appname/windowtitle, decoded
                                  newline labels)
  arguments.test.ts               where the selector ends and the command's
                                  positionals begin, both grammars
  parse.test.ts                   +6 grammar/escape cells beside the existing
                                  property tests

The login-form tree shared by resolve.test.ts and match.test.ts moves to
`__tests__/login-form-nodes.ts` rather than being copied into both.

All 27 cells are carried over unchanged and still pass; no file now exceeds
224 lines. pnpm check green: 602 unit files / 5278 tests, layering 71/71,
depgraph 22/22, mutation config 45/45, fallow clean over 127 changed files.

* revert(selectors): keep agent-device/selectors public, behind one AST subpath

The cutover removed the `agent-device/selectors` public subpath as part of
tightening the API. It is in use, so the removal is reverted: the subpath ships
the same ten symbols v0.20.5 shipped, with the same signatures.

That has to coexist with the reason the package façade is string-only, so the
AST leaves through one named door instead of the main one:

  @agent-device/selectors        string-in/string-out; every in-repo consumer
  @agent-device/selectors/ast    the published parser surface; one consumer,
                                 src/sdk/selectors.ts

`packages/selectors/src/ast.ts` re-exports parseSelectorChain,
tryParseSelectorChain, isSelectorToken, the AST-taking findSelectorChainMatch
and resolveSelectorChain, isNodeVisible, isNodeEditable, and types
SelectorChain / SelectorDiagnostics. `formatSelectorFailure` keeps its
published `SelectorChain | string` first parameter as a shim here rather than
widening internal/resolve.ts back to a union — the compatibility obligation
sits at the boundary that owes it.

This is strictly narrower than main, where the AST was reachable from anywhere
in src/ via src/selectors/*. Two gates hold it there: facade-symbols.ts pins
./ast to exactly the v0.20.5 list, and package-boundaries.test.ts asserts
src/sdk/selectors.ts is the only file outside the package that imports it.

Restored alongside: the ./selectors export and tsdown entry/chunk group, the
.fallowrc.json entry, the package-exports supported-subpath list, and both
client-api.md sections. No CHANGELOG entry — nothing is removed any more.

pnpm check green: 602 unit files / 5278 tests, smoke 35 passed / 3 live
skipped, layering 71/71 (10 packages, 32 subpaths), depgraph 22/22, mutation
config 45/45, fallow clean over 129 changed files, package smoke imported all
12 published entry points with publint and attw passing. Verified functionally
against the built dist: the doc's parse -> findSelectorChainMatch example
returns the same shapes as before, resolveSelectorChain still returns an AST
`selector`, and formatSelectorFailure still accepts a chain.

* fix(selectors): correct the two expectations that still assume the removal

Review P1s on a792415a: restoring the public subpath left two gates asserting
it was gone.

- installed-package-metro.test.ts moved `agent-device/selectors` into the
  blocked-specifier list. It goes back to the subpath smoke set, running the
  same `isSelectorToken('||')` + `parseSelectorChain` check it ran before the
  removal, so the file's only remaining delta from main is a formatter reflow.
- owner-files-no-leak.test.ts asserted `dist/src/sdk-selectors.js` was absent.
  It requires the stable named chunk again, and still rejects an auto-numbered
  `selectors2.js` fallback — the pair is what proves the restored tsdown chunk
  group is doing its job, verified against a clean build.

PR body corrected: the removal is no longer described as intentional API
tightening.

* refactor(selectors): satisfy the widened fallow scope after rebase

main's #1591 (the follow-up filed from this review) removed `packages/**` from
.fallowrc.json's ignorePatterns, so the new package is audited for the first
time. Everything below is a finding fallow could not previously see.

Dead surface, all confirmed consumer-free:

- 12 type re-exports from the `.` façade whose shapes consumers only ever
  reach structurally.
- MAESTRO_TEXT_SELECTOR_KEYS / MAESTRO_STATE_SELECTOR_KEYS, orphaned by this
  branch's own MAESTRO_SELECTOR_PROJECTION change, and the test-util
  SELECTOR_VALUE_HAZARDS. All three are module-local now.
- IS_PREDICATE_USAGE_HINT fails --production because its only consumer is the
  is-argument-surface parity test. It gets a commented `ignoreExports` entry
  rather than deletion: the constant is what makes the daemon and CLI raise
  ONE hint instead of two copied strings (ADR 0010), so the test asserting
  that is the point, not an accident.

`fast-check` is now declared by the package that imports it.

Duplication, split by what could be proven:

- `isUsefulVisibilityAnchor` existed character-for-character in both
  packages/selectors and packages/maestro. Moved to
  @agent-device/contracts/snapshot, which both already depend on and which
  already owns this vocabulary. Safe because the `normalizeType` each copy
  called is itself character-identical to the contracts one — checked before
  moving, since a different normalizer would have silently changed which
  nodes anchor.
- maestro additionally reimplemented `normalizeType`, `buildSnapshotNodeMap`
  (as `buildSnapshotNodeByIndex`) and `findSnapshotAncestor`, all
  character-identical to contracts'. Deleted in favour of the shared ones.
- The three scroll-ancestor walks are NOT deduped. They are structurally the
  same walk but each uses a different scrollable predicate, and I have no
  evidence the three agree; collapsing them would be a Maestro-conformance
  change, not a cleanup. Both maestro sites now say so, and the work is filed
  separately.

`projectSelectorExpression` (15 cyclomatic / 22 cognitive, written by the
cutover) splits into a dispatcher plus `readAgreedTextValue` and
`projectSelectorTerms`; all three are under threshold.

Rebase note: the one conflict, in package-boundaries.test.ts, resolved to
NEITHER side — #1591 had already deleted `AdReplayVerifiedTargetGuard` as an
unused export, and this branch deletes the seven ReplaySelectorPort names, so
the conflicting block is empty.

* build: record fast-check for packages/selectors in the lockfile

Declaring the dependency in packages/selectors/package.json without
regenerating pnpm-lock.yaml made every CI job fail in its install step with
ERR_PNPM_OUTDATED_LOCKFILE. My local `pnpm install --frozen-lockfile` printed
"+ 1 dependencies were added: fast-check@^4.9.0" and exited 0, which read as
success but was the same mismatch CI refuses.

Regenerated with the pinned pnpm 11.17.0, not the 11.5.3 on this machine:
11.5.3 rewrites peer-dependency resolution keys repo-wide (dropping
`(supports-color@7.2.0)` suffixes) and produced a 222-line diff. With the
pinned version the diff is the 4 lines this change actually needs, plus
pnpm's alphabetical re-sort of the root selectors entry.
2026-08-04 19:06:39 +02:00
Michał Pierzchała eb3fc5b28d chore: scan packages/** with fallow instead of ignoring it (#1591)
`ignorePatterns: ["packages/**"]` landed in #1494 W0 with the recorded
reason "its resolver cannot follow workspace specifiers". That was either
wrong at the time or never re-checked: the fallow version has not moved
(^2.95.0 then and now) and it resolves @agent-device/* through each
package's exports map today. packages/kernel alone exposes 8 subpaths and
~110 exports reachable only via workspace specifiers, and scanning it
reports zero findings — a resolver that could not follow the specifier
would report all of them.

The cost of the ignore is that every package extraction silently removes
its code from dead-code analysis. #1589 moved the selector engine into
packages/selectors/ and shipped a façade with 15 zero-consumer exports,
including `selectorUsesKey`, written in that PR and never called. A
follow-up commit removed them by hand; nothing would have caught them.

Removing the pattern surfaced 43 findings, driven to zero by deleting the
dead code rather than by baselining or excluding it (fallow-baselines/*.json
are empty on purpose — the posture is fix-or-document-the-exemption, so a
first baseline entry would be a policy change):

- 38 are deleted. 24 façade type re-exports whose only claim was that a
  consumer might one day want to name them — typecheck is green without
  every one, so the claim was theoretical; 5 façade value re-exports; 9
  `export` keywords on symbols used only inside their own file. Every
  deleted façade symbol comes off scripts/layering/facade-symbols.ts (and
  ad-replay's inline pin in package-boundaries.test.ts) in the same change,
  so R11 is narrowed with the façade, never weakened around it.
- 4 stale suppressions in src/provider-limrun-runtime.ts existed only
  because packages/ was invisible.
- 5 have consumers analysis genuinely cannot see, and get an
  `ignoreExports` entry naming the consumer per the existing `comment`
  convention: four test-tree importers that --production does not walk, and
  `LimrunIosCommandExecution`, which src/sdk/limrun.ts republishes as
  agent-device/limrun — its only importer compiles in a temp checkout, so
  no static edge reaches it. test/integration/limrun-public-types.test.ts
  is the standing proof that one is real API.

Three doc comments named types their façade no longer exports and are
corrected rather than left asserting something false — including #1555's
claim in session-replay-target-verification.ts that the daemon imports
`AdReplayVerifiedTargetGuard` directly. It does not; it reaches that shape
through `AdReplayTargetClassification`/`AdReplayDispatchGuard`, which is
why the name read as dead.

`scripts/maestro-conformance/**` was ignored wholesale to cover its corpus
data. Narrowed to `corpus/**`, which un-hides the tooling beside it and
turned up one more file-local export (`buildManifest`); regenerate.mjs's
importer of `fixtureContentHash` becomes visible, so that needs no
exemption at all.

scripts/check-affected/model.ts deliberately did not select the `fallow`
check for packages/*/src/**, carrying the same stale rationale as a
comment. Without that selection the new scope would never run in the
affected-driven lane, so the ignore removal would have bought nothing.
model.test.ts now pins the selection.

Verified: check:fallow and check:production-exports green with packages in
scope; full-repo `fallow dead-code` back to its one pre-existing finding;
typecheck, layering (R11), lint, format, build, check:package, and the
limrun published-types integration test all pass. Probed by adding a fresh
zero-consumer export to the xml façade — check:production-exports reports
it, so the #1589 case now fails the gate.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 18:26:03 +02:00
Michał Pierzchała 8d526f2402 perf(ios): halve hostile-screen capture cost under the XCTest-channel penalty (#1587)
* perf(ios): stop re-paying known-failing work on penalized private-AX captures

Live-measured on the Bluesky bench feed (139-171 nodes), each private-AX
capture wasted ~1.35s of its ~1.65s runner-side cost re-doing work a prior
capture already proved futile:

- ~1000ms: the viewport read is XCTest main-thread work; under the channel
  penalty it reliably burns its full timeout and falls back to the root
  frame anyway. Honor the penalty in privateAXSnapshotViewport the same way
  capture plans do.
- ~310ms: the depth ladder re-paid the kAXErrorIllegalArgument rejection of
  the default depth on every capture. Remember the accepted rung per bundle
  (penalty-duration TTL, cleared on target process change); expiry re-probes
  the full depth so screens that recover are not capped forever. Explicit
  --depth requests bypass the memory in both directions.

Steady-state hostile-screen captures drop 2.65s -> ~0.55s CLI wall, and
press --settle round-trips drop ~5.4s -> ~2s (settle needs two captures).

* fix(snapshot): distinguish penalty-deferred captures from genuine recoveries

A capture whose backend was PRE-selected by the XCTest-channel penalty was
stamped with the same recovered verdict as one that ground through a live
failure. Two costs followed on hostile screens (Bluesky bench: 130 repeats
per 30-task run):

- the daemon repeated the full fell-back warning on every capture, long
  after the arming capture had already said it once, and
- settle's one-shot private-AX budget reset fired on every loop even though
  the capture paid no grind to give the budget back for.

The deferred plan now stamps reasonCode 'deferred' (a new code; older
daemons drop unknown codes and keep today's behavior). The daemon keeps the
verdict recovered but suppresses the repeated warning, keeps the depth-cap
line, and skips the settle budget reset for deferred captures.

* style: wrap reasonCode union to satisfy oxfmt

* fix(ios): bind accepted-depth memory to the process, pin deferred through the wire parser

Review follow-up (#1587):

- The depth memory was cleared only in refreshCachedTargetIfProcessChanged;
  resetTargetAfterExternalRelaunch -> invalidateCachedTarget drops the cached
  PID without clearing it, so the next activation had no old PID to compare
  and could reuse a stale shallow rung for up to 120s (also A->B->A when A
  restarted while inactive). The memory now stores the PID it was learned
  under and only matches the same live process; recording without a PID is
  refused. Every invalidation path is covered automatically because they all
  drop currentAppProcessIdentifier.

- The deferred settle/warning tests constructed typed verdicts directly, so
  removing 'deferred' from the accepted reason-code set would silently
  restore the repeated warning and budget reset while tests stayed green.
  They now parse a raw runner-wire object through readSnapshotQualityVerdict
  (red on base: parser strips the code -> warning re-appears, reset fires).

- Extracted shouldReadPrivateAXViewportViaXCTest() and pinned the penalized
  viewport skip with an in-bundle regression test.
2026-08-04 17:12:30 +02:00
Michał Pierzchała 351ef7a14f refactor(android): enforce transport lowering in the type system (#1583)
`AndroidLoweredTouchPlan` widened the canonical two-sample trajectory to a
plain sample array, so a plan that skipped `lowerAndroidTouchPlan` still
satisfied the transport types. That is the mistake the lowering exists to
prevent: an un-lowered plan injects a two-sample gesture, which is the sparse
delivery #1572 removed from the shared plan in the first place.

Transport samples are now a minimum-arity tuple. `sampleGestureOffsets` floors
the frame count at three, so lowering always yields at least four samples,
which makes "denser than the canonical endpoint pair" a true statement about
the data rather than a comment. A canonical plan is no longer assignable, so
skipping the lowering fails typecheck at every injection seam.

Tightening the type caught three call sites that were passing un-lowered plans
straight to the helper transport, which is the evidence the previous signature
enforced nothing. `longPressPlan` now returns `AndroidLongPressTouchPlan`
instead of the wide union it never produced, and the dual-pointer normalize
test routes through the lowering like every other transport call.

Also drops the unused `= 'default'` on `sampleGestureOffsets` so every caller
states which platform sampling convention it wants, which is the point of
having centralized the policy.

Sample values are unchanged by construction, so Android injection stays
bit-identical to #1572.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 15:09:27 +02:00
Michał Pierzchała bcaa106845 refactor: extract snapshot and replay identity semantics (#1582)
* refactor: extract snapshot and replay identity semantics

* refactor: move pure rect primitives from contracts to kernel/rect

containsPoint, pickLargestRect, and isRectVisibleInViewport are raw
rectangle arithmetic with no snapshot awareness, so they belong beside
rectContains/rectArea in @agent-device/kernel/rect rather than in the
snapshot-semantics vocabulary. The node-aware resolveViewportRect folds
into contracts/snapshot-visibility.ts, retiring snapshot-geometry.ts;
after this split, everything behavioral in @agent-device/contracts/snapshot
is policy that interprets the snapshot model.

* refactor: restore ADR-0012 rationale docs and dedupe replay identity shapes

The #1478/#1581 extraction moved the identity/structural helpers but
compressed their invariant documentation to one-liners; the deliberate
no-ancestry-exclusion rule on idMatchCountInTree, the fail-closed guard
comparison, and the who-throws/who-detects contracts on the two
divergence reason markers now travel with their definitions again.

LocalIdentity and NodeStructuralDenotation move to
contracts/target-annotation.ts (beside TargetAncestryEntry, which
WaitLandmarkMismatchEvidence now references directly), so the guard
shapes in contracts/replay.ts are nominal instead of hand-rolled
structural twins; ad-script re-exports the vocabulary beside the
readers that produce it. Also inlines the demoteNonUniqueId
pass-through wrapper in session-target-evidence.ts.

* style: fix oxfmt formatting in snapshot-visibility

* refactor: move findSnapshotAncestor into contracts/snapshot-tree

The last root value import from src/selectors: predicates.ts reached
src/snapshot/snapshot-processing.ts for the ancestor walker. The walker
is index-based tree traversal with no presentation policy, so it joins
buildSnapshotNodeMap in contracts/snapshot-tree.ts; both consumers
repoint to the façade and the non-contiguous-index/cycle coverage moves
to the package test. src/selectors now has zero value imports from root
src in production files.
2026-08-04 15:08:51 +02:00
Michał Pierzchała 56d9ee605c fix(ios): preserve timed pan duration (#1572)
* fix(ios): preserve timed pan gesture execution

* fix(gestures): encode linear pans as endpoint plans

* fix(ci): pin wait contract exports

* fix(android): lower endpoint gesture plans for touch transport

* fix(gestures): preserve timed pan duration across adapters
2026-08-04 14:07:53 +02:00
Michał Pierzchała 111b85fc5a refactor(replay): make the daemon's artifact set the run's one ledger (#1575)
Artifact-path accumulation was double-written after the P5 extraction: the
engine's step loop kept its own `Set` while `runReplayScriptFile` kept an
outer `Set` that only its exception handler read, and `dispatchStep` wrote
both — two mutable collections with no single owner, kept in sync by hand.

The daemon's `Set` is now the run's only ledger. `dispatchStep` remains its
sole writer and returns its CONTENTS (cumulative for the run, not just the
step's own entries); the engine drops its `Set` for a plain `readonly
string[]` re-bound to whatever the capability last returned.
`AdReplayRunOutcome.artifactPaths` stays a façade field — it is wire-relevant,
the daemon's success response reports it — but is now a projection of what the
capability handed back rather than an independent accumulation.

The exception path is preserved byte-identically. The two old sets differed in
exactly one way: the engine's also absorbed a divergence build's own fresh
capture, which the daemon's never saw. Writing those into the ledger would
change what the catch block reports when `handleActionFailure` itself throws,
so they stay out of it and reach `handleActionFailure` through a derived union
(`mergeArtifactPaths`) instead — a value, not a write. A failing step always
ends the run, so nothing downstream observes that union.

Adds a counterfactual test at the ledger's one independent observation point:
a mid-loop throw (an unresolved `${VAR}` on step 3) after two artifact-producing
steps must report exactly those two artifacts. Dropping the `dispatchStep`
write fails it; returning per-step entries instead of the ledger fails its
companion completed-run assertion — verified in both directions, then restored.

The P5 `declaredScriptPlatform` duplicate this follow-up was also meant to
unwind landed inside #1555 itself (`resolveDeclaredScriptPlatform`, owned by
packages/ad-script, consumed by both the engine's inspect/digest path and the
daemon's `readScriptReplaySelection`), so there is nothing left to dedupe.

Gates: typecheck / lint / format:check / check:layering (56) /
check:replay-compat (12 digest-pinned entries) / vitest packages src/daemon
(255 files, 2179 tests) — all green.


Claude-Session: https://claude.ai/code/session_01JrwynLjFHMfz42PFBzoEwK

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-04 10:37:05 +02:00
Michał Pierzchała 6baa5d97e1 fix: make wait verdicts evidence-based (#1570) 2026-08-04 10:32:58 +02:00
Michał Pierzchała 761317deb7 refactor(daemon): extract native .ad replay to packages/ad-replay (#1478 P5) (#1555)
* refactor(replay): move the dependency-free engine leaves into packages/ad-replay

Stage A of the #1478 P5 extraction: vars, plan-digest (+canonical-json,
sole consumer), the target-identity classification core, report-action,
and suggestion-ranking move verbatim; imports updated. The package facade
temporarily re-exports the moved symbols so root consumers keep compiling;
a later stage narrows it to inspectAdReplay/runAdReplay only.

* chore(layering): register packages/ad-replay in the workspace and DAG

* refactor(replay): define the three-operation replay selector port with dual adapters (#1478 P5)

* refactor(daemon): route replay handlers through the selector port (#1478 P5)

* refactor(replay): split target verification into engine policy and daemon authority (#1478 P5)

* refactor(replay): move the .ad step loop behind inspectAdReplay/runAdReplay (#1478 P5)

* refactor(replay): lock the ad-replay façade to its real consumers (#1478 P5)

* test(replay): prove shared-id demotion on both selector-port adapters (#1555 review)

* fix(replay): restore invalid replayBackend rejection on the native path (#1555 review)

* refactor(replay): move shared .ad vocabulary to its owner, packages/ad-script (#1555 review)

* refactor(replay): neutral step/run outcomes and digest/resume behind inspectAdReplay (#1555 review)

P1 "do not smuggle daemon wire failures through a generic": drop the
TResponse generic from AdReplayStepRuntime/runAdReplay. executeStep and
handleActionFailure now return neutral tagged AdReplayStepOutcome/
AdReplayStepFailure values (kind/message/artifactPaths only); runAdReplay
returns a neutral completed/failed AdReplayRunOutcome. The engine never
holds or returns a DaemonResponse. The daemon adapter
(createAdReplayStepRuntime, session-replay-runtime.ts) keeps its real wire
response in a local side-map as it builds each neutral outcome, and
runReplayScriptFile reads it back once runAdReplay reports which step
failed, so the final response is byte-identical to before this split.

P1 "parsing/planning/digest/resume must also occur behind runAdReplay":
relocate computeReplayPlanDigest's call site and the --from/--plan-digest
resume-point math (resolveReplayEntryIndex) behind inspectAdReplay's
manifest as planDigest and a resolveEntryIndex closure. Neither is a new
top-level export -- inspectAdReplay/runAdReplay stay the only two. Timing
is preserved exactly (still called eagerly in prepareReplayPlan, before
prepareReplaySession's coordinator-mutating side effects) since moving
resume validation to run inside runAdReplay itself would let a rejected
--from request mutate coordinator/session state first -- a real ordering
hazard, not just a cosmetic one.

computeReplayPlanDigest/ReplayPlanDigestMetadata/resolveReplayEntryIndex
leave the ad-replay façade; request-router-repair-expired.test.ts and
prepareReplayPlan read the digest/resume result off the manifest instead.

* refactor(replay): relocate classifyTargetBindingMatch and pin the ad-replay façade (#1555 review)

P1 "complete the binding façade instead of documenting deviations":
classifyTargetBindingMatch never had a real consumer reachable through
inspectAdReplay/runAdReplay -- both its callers (the daemon's record-time
self-check in session-target-evidence.ts and its replay-time
classification wrapper in session-replay-target-classification.ts) are
daemon files that imported it directly. It interprets TargetAnnotationV1
evidence semantics shared beyond the engine, so it moves to
packages/ad-script alongside target-annotation-identity.ts (new
target-annotation-classification.ts + its test), and both daemon call
sites now import it from there instead of @agent-device/ad-replay.

One deviation remains and is reported rather than papered over per the
review's own instruction: the four target-verification policy functions
(planPreDispatchTargetVerification, planPostResolutionTargetVerification,
deriveReplayTargetGuardMismatchEvidence,
deriveWaitLandmarkMismatchEvidence) and the ReplaySelectorPort type
family stay exported. Their sole caller,
session-replay-target-verification.ts, interleaves these pure decisions
with daemon-only async work (capture, SessionStore, coordinator/resume
stamping, wire shaping) that must stay outside the engine by design;
moving their call sites to live only behind runAdReplay would require
restructuring that whole orchestration into new fine-grained
AdReplayStepRuntime capabilities, which is out of scope for this pass.
See packages/ad-replay/src/index.ts's header comment for the full
reasoning.

P1 "add the reviewer-required exact exported-symbol gate": adds
readNamedExports (scripts/layering/package-boundaries.ts), a small
parser over a façade's `export { .. } from`, `export type { .. } from`,
and direct-declaration forms, and pins @agent-device/ad-replay's exact
21-symbol export list in package-boundaries.test.ts. Plant-verified: a
stray `export const` addition failed the assertion; removed it and the
gate went green again.

* refactor(replay): drive target verification from the engine step loop (#1555 review)

Moves the verify-then-dispatch decision flow into packages/ad-replay's
step loop so the four target-verification policy functions
(plan{PostResolution,PreDispatch}TargetVerification,
derive{ReplayTargetGuardMismatch,WaitLandmark}MismatchEvidence) become
engine-private and leave the ad-replay façade. The daemon
(session-replay-target-verification.ts) shrinks to the narrow
AdReplayStepRuntime capabilities the engine drives: routing
(beginTargetVerification), capture (captureObservation), classification
(classifyTarget), dispatch (dispatchStep), and wire-building
(buildRecordedUnverifiableFailure, buildTargetBindingFailure,
buildPostDispatchTargetBindingFailure). Wire output and replay-compat
stay byte-identical; the exact-symbol façade gate is updated to the
shrunken export list.

* refactor(daemon): decompose the replay adapter's two over-threshold functions (#1555)

* refactor(replay): fold #1554's keep-session terminal-lifecycle policy into the ad-replay engine

Rebasing p5/extract-ad-replay onto main pulled in #1554's --keep-session
feature, which had grown its own daemon-side terminal-close-suppression
predicate (session-replay-terminal-lifecycle.ts's
resolveSuppressedTerminalCloseIndex/countExecutedReplayActions) independently
of this branch's own engine-side one (step-loop.ts's
isRepairArmedTerminalCloseAction). Both are the same decision family — replay
--keep-session and an active --save-script repair now share ONE structural
resolution (resolveSuppressedTerminalCloseIndex, generalized to "terminal
among EXECUTABLE actions" rather than the old physical-last-index check) and
one suppression check inside runAdReplay, gated on keepSession OR
runtime.isRepairArmed(). AdReplayRunRequest grew a keepSession field; the
neutral 'replayed' count in AdReplayRunOutcome is now computed inline in the
loop instead of the daemon's old actions.length - entryIndex approximation.

requireLiveSessionForKeepSession (the --keep-session live-session
postcondition) stays daemon-side, inlined into session-replay-runtime.ts,
since it inspects SessionStore state the engine never sees. The daemon-only
session-replay-terminal-lifecycle.ts this arrived with is deleted entirely —
its isExecutableReplayAction was a duplicate of the engine's own.

runReplayScriptFile's Maestro-format routing (including the new --keep-session
Maestro rejection) was extracted into routeMaestroReplay to keep the function
under fallow's complexity threshold after re-threading keepSession through it.

Added packages/ad-replay/src/internal/__tests__/step-loop.test.ts covering the
unified suppression decision (both keepSession and repair-armed) directly
against runAdReplay, including the terminal-among-executable-actions case with
a trailing nested replay marker. The daemon-level integration tests (6 tests
in session-replay-terminal-lifecycle.test.ts, exercising the same behavior
through runReplayScriptFile) and the SDK provider-scenario test
(active-session-script-publication.test.ts) needed no changes and pass
unmodified.

* refactor(daemon): decompose session-replay-runtime.ts into three modules (#1555)

Splits the ~1096-line replay runtime into cohesive pieces, keeping
session-replay-runtime.ts as thin orchestration (~240 LOC):

- session-replay-runtime-engine-adapter.ts: the AdReplayStepRuntime
  adapter (createAdReplayStepRuntime, the build*Failure capability
  implementations, and the lastResponse/lastObservation side-map
  mechanics), extracted verbatim.
- session-replay-runtime-plan.ts: extended with the plan-side helpers
  (validateReplayBackendFlag, inspectReplayPlanManifest,
  resolveReplayPlanEntryIndex, prepareReplayPlan, routeMaestroReplay)
  alongside the buildReplayMetadataFlags helper already there —
  buildReplayMetadataFlags is now module-private since its one caller
  moved into the same file. Also introduces ReplayScriptFileParams,
  named here (instead of derived via Parameters<typeof
  runReplayScriptFile>) so routeMaestroReplay can reference the shape
  without importing back from session-replay-runtime.ts.
- session-replay-runtime-session.ts (new): session preparation
  (prepareReplaySession and its coordinator arming/repair-preflight
  helpers), extracted verbatim.

Coordinator ownership is unchanged: createReplayCoordinator is still
constructed only in session-replay-runtime.ts, matching
replay-coordinator-ownership.test.ts's allowlist as-is — every
extracted module receives the already-constructed ReplayCoordinator as
a parameter. Pure move; no behavior change.

* test(replay): cover pre-step artifact ordering and resume-before-mutation (#1555)

Two invariants found during the P5 decomposition pass now have direct
counterfactual-verified coverage:

- packages/ad-replay/src/internal/__tests__/step-loop.test.ts: a
  post-dispatch target-binding mismatch (dispatchWithGuard) must report
  the accumulated PRE-step artifact snapshot it was called with, never
  the artifacts the failed dispatch itself produced. Verified red by
  swapping the buildPostDispatchTargetBindingFailure call to
  outcome.artifactPaths.

- src/daemon/handlers/__tests__/session-replay-runtime-plan.test.ts: a
  rejected --from/--plan-digest resume must never reach
  prepareReplaySession's coordinator-mutating writes (the R2 ordering
  invariant) — a pre-armed repair transaction and corrective-resume
  watermark are asserted byte-for-byte unchanged after rejection.
  Verified red by calling prepareReplaySession before honoring the
  plan-validation rejection.

* fix(ad-replay): enforce the exact two-entrypoint facade (#1555 review P1)

packages/ad-replay/src/index.ts now exports exactly two value symbols,
inspectAdReplay and runAdReplay, and zero types — formatReplaySuccessMessage
(presentation) moves beside its one caller in session-replay-runtime.ts, and
every type a root daemon file needs is derived structurally off the two
entrypoints in the one new src/daemon/ad-replay-facade-types.ts module
instead of being named off the façade.

scripts/layering/package-boundaries.ts's readNamedExports is rewritten on
oxc-parser's own static-export table instead of a regex, so it can no longer
silently miss a widening export form: a bare `export *` re-export or an
`export default` now throws (an un-enumerable, and therefore un-pinnable,
export), while `export * as ns` and every other enumerable form is still
counted. The pinned exact-symbol assertion in package-boundaries.test.ts is
narrowed to ['inspectAdReplay', 'runAdReplay'].

* fix(ad-replay): translate wire failures before the engine boundary (#1555 review P1)

AdReplayDispatchOutcome's guard-mismatch/landmark-mismatch variants carried
a generic `details: Record<string, unknown> | undefined` bag straight off
the wire response — a daemon wire projection crossing into the engine even
though the outcome itself was already a neutral type. The daemon adapter
(session-replay-runtime-engine-adapter.ts) now narrows that bag into the
typed AdReplayGuardMismatchEvidence/AdReplayLandmarkMismatchEvidence shapes
(observed identity, expected/observed structural denotation, ancestry
entries, match count) before returning the outcome; the unknown-parsing
readers move there with the wire-reading responsibility they always were.
target-verification.ts's deriveReplayTargetGuardMismatchEvidence/
deriveWaitLandmarkMismatchEvidence now consume only the typed values — no
`unknown`-valued record type remains on any engine-crossing signature.

* fix(ad-replay): move variable semantics/planning behind runAdReplay (#1555 review P1)

The daemon assembled the `${VAR}` scope (buildPreparedReplayScope) and
interpolated actions at two independent call sites: dispatch's own
(invokeReplayAction) and target verification's separate one
(resolveTargetVerificationEntry) — duplicated orchestration the P5 design
assigns to the engine.

runAdReplay's request now carries the raw scope INPUTS (varSources: plain
builtins/file/shell/cli-env data, plus actionLines/actionSourcePaths/
resolvedPath for interpolation-error location) instead of a built scope; the
engine builds the scope and resolves each action exactly once per step,
handing the RESOLVED action to dispatchStep/beginTargetVerification while
every other capability still receives the ORIGINAL recorded action (a
target-binding divergence reports the recorded selector, never an expanded
${VAR}). This is the one resolution site now — session-replay-action-runtime.ts's
invokeReplayAction and session-replay-target-verification.ts's
resolveTargetVerificationEntry no longer hold a scope or call
resolveReplayAction themselves.

Scrub-value collection (collectReplayScrubbableVarValues, for divergence-report
redaction) is kept single-sourced in the engine too: it's computed from the
engine's own live scope and threaded to each build-failure/handleActionFailure
capability as an explicit scrubVars argument, rather than the daemon
recomputing it from a second scope object (which would have gone stale,
since expandedBuiltinNames tracking now only happens engine-side).

The Maestro replay path's own daemon-side vars usage is unrelated (a
different engine) and is out of scope here.

* fix(ad-script): make ${VAR} interpolation a linear scanner

CodeQL flagged the interpolation regex's fallback group as js/polynomial-redos
once vars.ts moved into packages/ (library-input classification): every
${NAME:- prefix of an unclosed input rescanned to end-of-string, quadratic
overall — 1,857 ms measured on 20k repetitions of '${A:-['. Replaced with a
single-pass scanner; failed fallback scans emit their span verbatim and resume
after it (escape-pair alignment is identical from every candidate start inside
the span, so no later candidate can terminate where the failed scan could not).
Equivalence: 200k-trial differential fuzz against the retired regex over the
adversarial alphabet, zero mismatches; both adversarial shapes now resolve in
1-2 ms.

* refactor(ad-replay): typed façade replaces the zero-type rule (#1555 structural-quality review)

Reverses the exact-two-value zero-type export shape #1555's second review
pass established: it forced every root type derivation through one shim
(src/daemon/ad-replay-facade-types.ts) and left four daemon-side twin types
(TargetVerificationEntry, TargetClassificationOutcome,
TargetBindingFailureEvidence, ReplayVerifiedTargetGuard) plus a
toDaemonEvidence copy translator shadowing the engine's own shapes.

packages/ad-replay/src/index.ts now exports inspectAdReplay/runAdReplay
(unchanged, still the only two values) plus the neutral vocabulary their
signatures are built from, by name — following packages/maestro's façade
precedent. The exact-symbol gate in scripts/layering/package-boundaries.test.ts
is widened to pin the full sorted list (values + types).

The four daemon twins are deleted; session-replay-target-verification.ts and
session-replay-runtime-engine-adapter.ts now use the engine's own
AdReplayVerificationEntry/AdReplayTargetClassification/
AdReplayTargetBindingEvidence/AdReplayVerifiedTargetGuard directly.
TargetBindingDivergenceBuilt's array fields are now readonly-compatible, so
toDaemonEvidence's copy is gone — evidence flows through unchanged.

* fix(ad-replay): honor the selector port's own contract in the parse gate

target-verification.ts's planPreDispatchTargetVerification used
resolveRecordedTarget (operation 2, resolve) over an empty node tree purely
to read its parse-invalid reason — a resolve call standing in for a parse
call, even though readSelectorExpression (operation 1, parse) exists to
answer exactly that question and was already unused inside the engine.

Replaced with port.readSelectorExpression('ordinary', [token]). The mapping
is not 'invalid' -> skip: production's 'ordinary'/'wait' grammars only ever
record a boundary once it has already parsed, so a single malformed token
can only come back 'not-applicable' there ('invalid' is unreachable from
this call site on the production adapter). Both non-'expression' outcomes
map to skip, matching the historical behavior (a single parse-invalid reason
covered both cases). platform dropped from the function's params — it was
only ever threaded to the resolve call this replaces.

Added a contract-suite cell pinning the exact (diverging) discriminant each
adapter reports for a selector-shaped-but-malformed bare token, and why the
divergence is harmless for the one real consumer.

* refactor(ad-replay): split step-loop.ts and shrink the daemon adapter (#1555 structural-quality review)

step-loop.ts (810 LOC) splits three ways, following packages/maestro's own
precedent:
- internal/runtime-port-types.ts: the AdReplayStepRuntime boundary
  vocabulary (all the neutral types the engine/daemon exchange).
- internal/verify-dispatch.ts: verifyAndDispatchStep + its dispatchNoGuard/
  dispatchWithGuard helpers.
- internal/step-loop.ts: runAdReplay itself plus the terminal-close/
  executable-action structural logic (isExecutableReplayAction,
  resolveSuppressedTerminalCloseIndex).

packages/ad-replay/src/index.ts's type exports now source from
runtime-port-types.ts. step-loop.test.ts's AdReplayStepRuntime import moves
to the new path (no assertion changes).

src/daemon/handlers/session-replay-runtime-engine-adapter.ts (553 LOC after
item 1's twin removal) shrinks to 294 via two further extractions:
- session-replay-dispatch-narrowing.ts: the wire `details` bag -> typed
  evidence narrowing and dispatch-failure classification.
- session-replay-runtime-step-support.ts: ReplayStepContext (moved here to
  avoid a cycle with the adapter, which re-exports it by name) plus the
  failure-wrapping/diagnostics-support helpers.

Final LOC: adapter 294, dispatch-narrowing 148, step-support 153,
step-loop 225, verify-dispatch 246, runtime-port-types 374.

* test(ad-replay): package-local tests for resume.ts/target-verification.ts + terminal-lifecycle test rename

resume.test.ts covers resolveReplayEntryIndex directly (previously only
exercised transitively through the daemon's session-replay-runtime-plan
tests): no --from/--plan-digest, the paired-flags requirement, in-range
--from, out-of-range rejection, stale-digest rejection, the authorized
empty-tail boundary (actionCount + 1) gated on a matching watermark, and the
unperformed-record-and-heal growth check. Counterfactual run and restored:
widening describeOutOfRangeResumeFrom's bound turns the out-of-range/
empty-tail-without-watermark assertions red (2 failures observed).

target-verification.test.ts covers all four engine policy functions
directly: the two plan* pre-capture gates and the two derive* post-dispatch
evidence builders, including item 2's own new decision surface (a fake
ReplaySelectorPort proving both non-'expression' readSelectorExpression
outcomes map to skip). Counterfactual run and restored: narrowing the check
to the literal `'invalid' -> skip` reading turns the 'not-applicable' case
red (reports recorded-unverifiable instead of skip).

session-replay-terminal-lifecycle.test.ts renamed to
session-replay-runtime-keep-session.test.ts: its production module
(session-replay-terminal-lifecycle.ts) was already deleted by the #1554
fold-in, and its six cases drive the full runReplayScriptFile round trip
against a real SessionStore (including daemon-only postconditions the
engine's step loop never reaches) rather than testing engine policy through
the façade in isolation — the engine's own terminal-close-suppression
decision already has direct, cheaper coverage in step-loop.test.ts. No
assertion changes; both files' header comments cross-reference the split.

* refactor(ad-replay): compute scrub values once per step, one name end to end

collectReplayScrubbableVarValues(scope) was called fresh at 5 separate
return points inside one verifyAndDispatchStep invocation plus once more in
handleActionFailure — always the same result, since nothing between them
mutates scope. step-loop.ts's runAdReplay now computes scrubVars ONCE per
step, right after resolveReplayAction (the one call that can grow the
scope's expanded-builtins set), and threads it as a plain
readonly AdReplayScrubValue[] value; verify-dispatch.ts no longer imports
ReplayVarScope or collectReplayScrubbableVarValues at all.

"One name" end to end: the daemon's TargetBindingDivergenceContext.scrubVars
and withReplayFailureDiagnostics's scrubVars param used a separately-derived
ReturnType<typeof collectReplayScrubbableVarValues> (mutable array) instead
of the engine's own AdReplayScrubValue, requiring a [...scrubVars] copy at
every daemon call site to satisfy the mutable-array type. Both now use
readonly AdReplayScrubValue[]/readonly ReplayVarScrubEntry[] (structurally
identical, already readonly-safe downstream — scrubReplayVarValues and
createReplayDivergenceSanitizer already accepted readonly arrays), so the
four [...scrubVars] copies in session-replay-runtime-engine-adapter.ts are
gone.

* fix(daemon): make lastObservation genuinely per-step, not per-run

createAdReplayStepRuntime's lastObservation closure lives for the whole
replay run (one factory call covers every step), but was never reset
between steps. Every current buildTargetBindingFailure call site happens to
be preceded by this same step's own captureObservation, so the
`lastObservation ?? { reason: 'observation-missing' }` fallback could never
actually fire — but if it ever did (a future call path reaching
buildTargetBindingFailure without capturing first), it would silently
attach the PREVIOUS step's screen instead of reporting the missing-capture
condition the fallback message claims.

armStep runs exactly once per step, before any of that step's other
capabilities (verified against step-loop.ts's runAdReplay loop order) — the
natural per-step boundary. It now clears lastObservation first. No behavior
change on any reachable path today (full daemon + ad-replay suite: 1766/1766
green); an unrelated device-claim-prune contention flake was observed once
and did not reproduce on isolated or full-suite reruns.

* docs(ad-replay): fix decayed review-changelog comments naming defunct symbols

Four comments named symbols/paths that no longer exist, left behind by
earlier review passes describing PR history rather than the current
constraint:
- session-replay-runtime-step-support.ts / session-replay-runtime.ts (2
  sites): referenced a function called executeStep, which was never
  reintroduced under that name after the P5 split — the actual mechanism is
  the runtime's dispatch/build-failure capabilities recording into the
  lastResponse side-map.
- session-replay-runtime.ts: referenced an engine collectArtifactPaths
  capability that does not exist — artifactPaths is a daemon-side Set the
  adapter mutates via collectReplayActionArtifactPaths.
- packages/ad-replay/src/internal/selector-port.ts: pointed at
  ./testing/in-memory-selector-port.ts, the in-memory adapter's pre-stage-D
  location — it has lived at
  src/__tests__/test-utils/in-memory-replay-selector-port.ts since.
- session-replay-repair-hint.ts / session-replay-runtime-step-support.ts (2
  sites): named target-identity.ts, which does not exist (the real file is
  target-identity-node.ts); the second site additionally mislabeled
  classifyReplayTarget as engine-side when it is daemon-side
  (session-replay-target-classification.ts).

Comment-only; no behavior change.

* refactor(ad-script): move declaredScriptPlatform to its natural shared owner

packages/ad-replay/src/internal/inspect.ts's declaredScriptPlatform and
src/daemon/replay-device-selection.ts's readScriptReplaySelection each kept
their own copy of the same "platform declared before the first open" scan
over runtime/open actions — .ad script semantics, not engine or daemon
policy, needed independently by ad-replay's plan-digest precedence and the
daemon's device-selection platform resolution.

Verified this was a genuine duplicate (not the single-sourced state I
initially reported): readScriptReplaySelection's platform-tracking loop
computes the identical result via a differently-shaped traversal fused with
its own app-target scan.

resolveDeclaredScriptPlatform now lives in packages/ad-script (its natural
owner: the one package both ad-replay and the daemon already depend on,
avoiding the R11 issue that justified the original duplication). The
daemon's app-target scan stays its own separate pass; fusing it back into
the shared function would smuggle a daemon-only concern into ad-script for
no measurable cost (the actions array is small, and the shared function
already stops at the same point the app-target scan needs to look).

* docs(ad-replay): fix package.json description to match the current façade

Described "target-identity, variable substitution, plan-digest, and report
primitives" — the wide pre-#1555-review façade shape. Vars/identity/report
vocabulary moved to ad-script/daemon across the P5 and #1555 review passes;
the package now exports exactly inspectAdReplay/runAdReplay plus the
neutral AdReplayStepRuntime vocabulary. Description updated to match.

* refactor(daemon): fold the step-support fragment back into the engine adapter

A simplicity audit judged session-replay-runtime-step-support.ts a
size-target fragment, not a concern boundary: four unrelated concerns,
one consumer, and a header comment admitting it existed to satisfy the
<300 LOC metric. Folded back; the previously-exported helpers are
module-private again; the adapter's honest size is renegotiated from the
plan metric (dispatch-narrowing stays extracted — it has one nameable
job).
2026-08-03 17:25:13 +02:00