28 Commits

Author SHA1 Message Date
Michał Pierzchała 0feb4e26a0 feat(ios): drive ASWebAuthenticationSession sign-in sheets in place (#2438) (#2448)
* 0.21.1

* feat(ios): drive ASWebAuthenticationSession sign-in sheets in place (#2438)

iOS apps that sign in via ASWebAuthenticationSession present the identity
provider in com.apple.SafariViewService, out of the app's process. Two facts,
both verified live on the iOS 26.2 Simulator, made these flows unautomatable:
activating or launching the host cancels the auth session, and the host AX
bridge cannot see the sheet because the app stays the AX primaryApp.

Serve and drive the sheet in place. A closed registry names the host (shared by
the TypeScript and Swift sides under a parity test); the runner reads and drives
it without activation and never adopts it as the session target; and the
Simulator route detects a running host with a cheap device-scoped ps probe and
takes the runner path, since the bridge would serve the occluded app tree as if
healthy. open refuses to launch a registered host, and captures carry a
system-surface disclosure.

Presence is foreground state, not tree content: a torn-down host serves a richer
tree than a live one, so content heuristics cannot tell them apart. The
never-activate guard is what keeps the foreground predicate sound, which also
makes the stale-tree failure mode unrepresentable for this flow.

Closes #2438

* chore(gates): register contracts/ios-system-surface in the export snapshot

* fix(ios): close the system-surface correctness gaps from review

Presence probe: absence and probe failure are no longer reported as "no
surface". The probe returns present/absent/unknown and the route takes the
runner for anything but a proven absent, so a sheet opened between two captures,
or a probe that cannot answer, can no longer fall through to a bridge capture
that would answer confidently from the occluded app tree. Only a positive
observation is memoized. The probe now matches with pgrep and reads only a
matched pid's environment, which is ~3x cheaper than the previous full
process-environment dump and stops copying every process's environment.

Open guard: the refusal moved to every resolved-host launch and terminate, so
the URL, deep-link and launch-args branches that returned before the old check
can no longer launch the host. Terminating a host is refused too, since that
cancels the presented session just as launching it does.

Comparison: the surface identity now reaches SnapshotState, and tap-failure
corroboration refuses outright when a baseline and a post-action capture
disagree about it, instead of letting app and sheet captures meet in legacy
same-presentation matching. Selector routes disclose an iOS system surface
through the shared disclosure seam rather than reading only the Android field.

The contracts import in the launch path is deferred so the app-lifecycle
facade's eager closure stays flat, and the runner's comment prose is trimmed
because apple/runner ships to npm as uncompiled source.

* fix(ios): route the system-surface probe through the Apple tool provider

The probe shelled out with runCmd, so every eligible capture spawned a real
process even in provider-backed tests that stub the Apple tool seam — 17 real
spawns in one scenario file, which is both wasted work and added latency on
timing-sensitive settle paths. It now goes through runAppleToolCommand like the
sibling ps probe, so a stubbed provider answers instead of spawning.

* fix(ios): disclose a skipped bridge when the surface probe cannot answer

Routing an unprovable probe to the runner is right, but the early return also
skipped runFallback, so the response lost its warning and kept an identity that
could still be compared against a bridge publication. An unknown probe now falls
back through the same disclosed path as a bridge failure, with its own reason.

* fix(ios): keep surface identity through comparison, find, and probe scope

A ps read that carries no SIMULATOR_UDID at all was reported as absence, so an
unreadable or truncated environment could route a live sheet to the occluded app
tree. Only a scope naming a different device is a real negative now; a missing
one stays unknown.

The shared post-gesture comparison token used comparisonKey or the backend
alone, so an app capture and a sheet capture — both XCTest — compared equal and
a sheet appearing or dismissing read as a stable surface. The token now carries
the surface, which covers stabilization, verify and settle through the one path
they share.

Mutating find rebuilt its capture without iosSystemSurfaceBundleId, so the
shared disclosure helper could not report the sheet on either outcome. It is
preserved now.

Each fix has a regression that fails without it.

* fix(ios): keep surface identity in verify and settle comparisons

`--verify` compared node digests and `--settle` diffed node-only baselines, so an app
baseline and an in-place system-surface capture (a web sign-in sheet) were treated as one
presentation: a meaningless changed verdict, and a whole-surface replacement presented as an
in-surface diff with refs.

The pre-action baseline now travels with the surface its capture described, from the resolution
and the session frame through to the settled capture, and one module owns the comparison for
both routes. Across a surface change no same-surface claim is made: evidence reports the
transition instead of a digest comparison, the settled diff and its refs are withheld, and both
payloads disclose the transition.

* refactor(test): move the cross-surface settle tests onto their source mirror

The #2438 cross-surface cases were appended to `settle.test.ts`, taking it over the
test-file size ratchet (2528 lines, 2359 at the merge-base). They assert the
comparison `post-action-surface.ts` owns, so they move to that module's mirror test
file, and the device double plus the trees both files drive move to a sibling
fixtures module under `__tests__/` rather than being duplicated.

Pure move: every test and every assertion is unchanged, and `settle.test.ts` is back
under its merge-base length.

* test(daemon): cover the cross-surface settle refusal on the generic route

`scroll --settle` and `back --settle` plumb the baseline's surface identity
through `baselineSurfaceBundleId`, but nothing asserted it: the generic route
had zero coverage of the #2438 refusal, so a regression there would have been
silent while the element-targeted route stayed green.

Assert the same contract the targeted route guarantees, in both directions and
for both commands: no diff is attached across an app/sheet boundary — therefore
no tail and no `refsGeneration` — the transition is disclosed, and the settle
observation still reports its own verdict alongside that disclosure.

Each direction falsifies a different half of the plumbing, so both are needed:
dropping the baseline's surface identity fails only the sheet-to-app tests (an
app baseline has no surface id to lose), and dropping the settled capture's
fails only the app-to-sheet tests. No production change: the plumbing was
correct, only untested.

* refactor(ios): inline the single-caller surface disclosure wrapper

iosSystemSurfaceDisclosure() only mapped provenance-or-nothing onto the shared
constant for one caller, so the caller now reads the constant directly and the
wrapper is gone. Its test becomes a test of the transition disclosure, which is
the function that still earns its place (the "sheet is gone" sentence).

readAppleSnapshotResult also called readSystemSurfaceProvenance twice inside one
spread; it is bound to a local and read once.

* docs(adr): state that a presented surface outranks a requested bundle id

prepareActiveCommandContext checks for a presented system surface before it
resolves or activates command.appBundleId, so a command naming a different app is
still served the sheet. That is intended, but the code does not read that way;
the amendment now says it plainly.

* refactor(ios): carry the system surface in the capture's comparison lineage

A capture of an in-place system surface (a web sign-in sheet) describes a
different presentation than a capture of the app, so it must never compare
equal to one. The `present` branch of the iOS snapshot route returned a bare
fallback, so that capture carried no comparison identity at all, and two
comparison sites hand-rolled the distinction from `iosSystemSurfaceBundleId`
instead.

The probe now reports which host it matched, and the `present` branch goes
through `runFallback` like the `unknown` branch beside it, lineaged to
`<device>:<host bundle>`. The comparison key then differs from an app
capture's by construction, so the surface branch in `hasMatchingPresentation`
and the surface concatenation in `snapshotComparisonKey` are gone: both sites
are plain key equality again, and neither knows that system surfaces exist.
Two captures of the same surface still share a lineage, so they stay
comparable with each other.

A presented surface is not a bridge failure, so it gets its own warning
wording: the bridge is inapplicable here, not unavailable.

* refactor(interaction): carry the pre-action baseline as one surface-scoped value

The same pre-action tree travelled as a flattened nodes/surface pair at every
boundary, and each boundary rebuilt it with a conditional spread. Carry
SurfaceScopedNodes itself instead:

- ResolvedInteractionTarget gets preAction?: SurfaceScopedNodes, replacing the
  preActionNodes/preActionSurfaceBundleId pair and the PreActionBaselineFields
  intersection on all three arms of the union.
- SettleObservationCommandOptions gets baseline: SurfaceScopedNodes, replacing
  baselineNodes/baselineSurfaceBundleId.
- RefResolution carries tree: SurfaceScopedNodes instead of nodes plus a loose
  surfaceBundleId.

That retires preActionBaselineFields(), preActionBaseline(), evidenceBaseline(),
the local SettleBaseline type, the split-then-reassemble in
settleObservationCommand, and the 'preActionNodes' in resolved narrowing tests.
SurfaceScopedNodes moves to contracts, where ResolvedInteractionTarget can name
it; only two sites now mint one from a SnapshotState.

Behaviour is unchanged: the cross-surface guarantees keep their existing tests.

* fix(ios): identify a surface capture by what the runner served

The `present` path stamped the capture's comparison lineage from the host-side
presence probe. That probe answers about a host PROCESS and deliberately stays
positive while a dismissed host lingers, so during that window the runner
truthfully returned APP content while the route lineaged it to the HOST: the
sheet capture before the dismissal and the app capture after it compared equal,
and a post-gesture poll could read the transition as a stable surface.

Derive the identity from the returned capture's `systemSurface` instead - the
runner stamps the surface it actually served - and say which of the two the
capture holds in the warning. The probe's host is now evidence only: it names
the matched host in a route diagnostic so a lingering window is legible in the
daemon log. Other reasons keep their lineage and wording byte for byte.

Captures that bypass the route's planning (a pinned backend, a custom-actions
read) also reach the runner, and the runner serves the sheet there too. They
carried no comparison key at all, so a sheet and app content fell through to
legacy presentation matching as one presentation and could corroborate a tap
across the two. The capture owner now gives those a surface-scoped identity as
well, with no fallback-source residue: nothing fell back. An app capture off
the route is untouched.

* fix(ios): derive a served surface identity at the one stamping point

A runner fallback's comparison identity was decided per call site. The
`present` path and the off-route path read the runner's `systemSurface`
stamp, but the plain `runFallback` path did not: it stamped the app
lineage the route had planned, whatever the runner returned.

The probe and the capture are separate observations, so a sheet can
appear in the gap between them. With the bridge circuit already disabled
for the generation, an app capture and a later sheet capture both
received the same app-generation key, so tap corroboration could treat
two different surfaces as comparable.

`stampFallback` now owns the decision for every runner fallback: the
surface the runner served outranks the app lineage the route planned.
The reason the bridge was skipped survives either way, and
app-generation evidence leaves with the app lineage it describes, so two
captures of the same sheet still compare equal. `runSurfaceFallback`
keeps only the reason, which is the one thing that path decides.
2026-09-11 17:37:35 +02:00
Michał Pierzchała fef0b12cc5 chore: hoist shared snapshot/selector test fixtures into a single canonical location (#2419)
* chore: hoist shared snapshot/selector test fixtures into @agent-device/selectors

PR #2397 left two copies of the snapshot-state builder and duplicated
geometry/touch-point arbitraries (root's src/__tests__/test-utils/ and the
package's internal/__tests__/), because packages cannot import root src/.
Move the canonical versions into a new @agent-device/selectors/test-fixtures
subpath and have both root and the selectors package import from it, leaving
buildNodes and the root-only replay/gesture arbitraries in place.

Fixes #2402

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01USabYbQjD16A2UkkvMpf5x

* chore: exempt test-fixtures.ts's test-only arbitraries from dead-code check

PROPERTY_RUNS, scrollingContainerTypeArb, distinctRectPairArb, and
interactionTouchPointScenarioArb are consumed only by *.test.ts files, which
Fallow's --production analysis does not see, matching the existing pattern
for other workspace-package symbols reached only from the test tree.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01USabYbQjD16A2UkkvMpf5x

* chore: consolidate makeSnapshotState into capture-kit, rename fixtures file

An adversarial review of the #2402 fixture-hoisting change found a third
copy of makeSnapshotState in packages/capture-kit/src/snapshot-state.fixtures.ts,
predating PR #2397. Since @agent-device/selectors already depends on
capture-kit, make capture-kit's copy canonical (exported as
./snapshot-state-fixtures) and have the selectors package's fixtures module
re-export it instead of duplicating it a third time.

Also rename packages/selectors/src/test-fixtures.ts to
snapshot-geometry.fixtures.ts (subpath ./snapshot-geometry-fixtures) to match
every other test-fixture module's *.fixtures.ts convention in this repo,
which lets it fall under .fallowrc.json's existing blanket **/*.fixtures.ts
dead-code exemption instead of needing a bespoke per-symbol entry.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01USabYbQjD16A2UkkvMpf5x

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-09-09 13:56:40 +02:00
Michał Pierzchała 947582a3cc refactor(daemon): move interaction and find routes behind facade (#2178) (#2228) 2026-09-02 07:59:06 +02:00
Michał Pierzchała 1826b2e68b refactor(ios): integrate runner with snapshot engine (#2214)
* refactor(ios): integrate runner with snapshot engine

* fix(ios): preserve macOS runner snapshots

* refactor(ios): keep runner presentation device-aware

* fix(ios): validate runner scroll presentation

* fix(ios): close presenter package boundaries

* fix(ios): preserve snapshot source lineage

* test(ios): colocate snapshot engine coverage

* fix(ios): settle post-merge audit checks

* test(ios): fix manifest parity lint

* refactor(ios): simplify runner source walk

* test(ios): cover shared package source fixture

* fix(ios): close post-merge audit gaps

* perf(ios): avoid bundling acquired snapshot path
2026-09-01 18:36:09 +02:00
Michał Pierzchała aeb2ff8402 refactor(contracts): retire the platform/interaction compatibility façades (#2048)
#1959 granularized the contracts entry surface but left the wide
platform/interaction façades in place as a compatibility surface for
~490 type-only importers. This mechanically moves every importer
(~350 files) onto the granular subpath each symbol actually lives on,
deletes the two façade files, and removes the fallow ignoreExports
entries and eager-closure-budget rows that existed only to cover them.

Eight previously-unexported source files needed new package.json
subpaths (clipboard, keyboard, network-traffic, platform-plugin,
platform-providers, runner-lease-context, screen-recording-runtime-host),
and ./interaction now points directly at src/interaction.ts instead of
the deleted barrel. .oxlintrc.json's no-restricted-imports rules for
the two façades are removed since there's no wide facade left to warn
against value-importing.
2026-08-26 13:13:25 +02:00
Michał Pierzchała 759f175332 refactor: move Wave 5 touch commands to platform runtime (#1987)
* refactor: move touch commands to platform runtime

* fix: require direct selector touch binding

* fix: address touch runtime review

* fix: classify maestro direct click guarantee

* fix: distinguish Maestro direct selector dispatch
2026-08-24 14:54:44 +02:00
Michał Pierzchała cb65d6ca1f refactor(tests): replace the test-utils barrel with direct module imports (#1956)
* refactor(tests): replace the test-utils barrel with direct module imports

The barrel re-exported 13 modules, so every importer evaluated all of them
(store-factory alone drags 16 daemon session-store files; property-arbitraries
drags fast-check). Importing the backing modules directly cuts the unit
suite's aggregate eager module evaluations from 153,401 to 144,344 (-5.9%),
measured with the eager-import-closure walker. Deleting the barrel makes the
tax unrepresentable instead of pinning it with a guard test.

* docs(testing): point fixture guidance at the test-utils modules, not the deleted barrel

* test: extract replay session fixture
2026-08-22 12:21:35 +02:00
Michał Pierzchała 81409f1a7c refactor: migrate type to the request-bound device runtime (#1935)
* refactor: migrate type to the request-bound device runtime

Wave 5 unit 2 for #1739 (ADR 0019), find's last blocker. `type "text"` and
`find <q> type "text"` now reach the device through one admitted, request-bound
`typeText` operation instead of the `handleTypeCommand` interactor leaf and its
dispatch-table arm.

- New `TypeTextRuntimeOperations` contract riding the same `Interactor` seam as
  focus/screenshot/element-text; the operation returns the interactor's own
  closed `TypeTextBackendResult`, so Apple route evidence passes through and
  every other owner types blind, exactly as before. The iOS synthesized-type
  commit wait (#1676) is Apple-interactor-internal and moves nowhere.
- The interaction backend's `typeText` member exists only when the `type`
  handler admitted and bound a runtime — no caller can fall back to legacy
  dispatch, so the command keeps exactly one execution path (R41).
- `executeBoundTypeText` reproduces the retired leaf byte-for-byte: leading-ref
  rejection with the same hint, space-joined positionals, the 0-10000 delay
  bound, and only textEntryRoute surviving from the owner's result. Its parse
  pins moved from the dispatch-level tests into the daemon runtime test.
- Exact-owner facts replace the capability bucket (apple sim+device, android
  all-but-simulator-row, harmonyos emulator+device, linux device, web device,
  vega unavailable, providers wherever their interactor is reachable) and
  `type` leaves HARMONYOS_SUPPORTED_COMMANDS / WEB_INTERACTION_COMMANDS.
- The Linux desktop replay types a digit on real hardware; the coverage
  manifest promotes `type` contract -> live with the two-sided count pins.
- Android/webdriver facts helpers extracted (androidTouchFact, interactorCell)
  to keep inspectFacts under the complexity gate.

`find` stays legacy: both of its direct execution legs now share bound
runtimes, so the atomic R35 cutover is next.

* refactor(type): single-pass daemon routing, shared binder source, owner cell tests

Review follow-ups on #1935.

- The type handler now calls the bound executor directly: the interaction-
  runtime hop validated and formatted what executeBoundTypeText validates and
  formats again, so it is gone — no boundTypeText backend member, no second
  result rebuild. The ADR 0014 frame expiry moves to the handler.
- New contracts/interactor-operation-binding.ts: one local resolver and one
  fail-closed provider resolver shared by the screenshot, focus, and type
  binders — three private copies retired, provider error text preserved.
- provider-limrun/interaction-operations.ts: the interactor-backed interaction
  cells move out of the app-log owner (586 -> 563 lines, below its pre-unit
  size); text interaction is composed by that owner, not defined in it.
- Every owner runtime test now pins the focusPoint/typeText fact cells and
  bound-operation presence for its exact kinds: apple, android (incl. the
  synthetic-simulator refusal), harmonyos, linux, web, vega (refusal + hint),
  webdriver (reachability-gated, incl. inactive session), limrun (live +
  recovery). The webdriver unsupported-capability row documents that
  interaction gates on interactor reachability, not capture declarations.
- The Linux replay assertion is now change-sensitive: type "555" then wait for
  a 555 node — no calculator button carries that label, so the wait passes only
  if the keystrokes landed in the display; deleting the type step turns it red.

* fix(replay): give the calculator focus before the Linux type assertion

The change-sensitive wait exposed what the review predicted: the typed digits
never landed, because `focus 100 100` clicks the DESKTOP and takes keyboard
focus away from the calculator. The old broad assertion masked exactly this.

The retries then wedged on a latent quirk: attempt-1 leaves the pointer at
(100,100), and the next attempt's `xdotool mousemove --sync` to the same point
waits for a motion event that never comes, so every retry dies at the focus
step with a 10s timeout — which is why the lane reported step 7, not the
failing wait.

New tail: `focus 100 100` (R40 evidence + survival assert), then
`click "label=1"` — a resolved press inside the window that restores keyboard
focus, proves pointer input lands in the app, and moves the pointer off
(100,100) so retries cannot trip the mousemove no-op hang — then `type "55"`
and `wait "label=155 || text=155 || value=155"`. No button is labelled 155, so
the wait passes only if the typed keystrokes reached the display.

* refactor(type): delete the fallow SDK typeText surface, drop dead surface fields

Thermo-nuclear review follow-ups (reviewed at 82fb8c2dc; the three-hop relay
it names was already deleted in 3d44b6dea — these are the residuals).

- typeTextCommand had zero production callers after the direct-executor
  routing: the daemon was its only consumer, and the released SDK types over
  the wire (executeCommand('type')), not through the embedded runtime catalog.
  Deleted: the command, its Options/Result types, the catalog registrations,
  the AgentDeviceBackend.typeText member, and every fixture/pin that kept the
  dead surface alive. R41's retirement claim now names typeTextCommand, so a
  revival fails the gate. One parse/compose owner remains: executeBoundTypeText.
- TypeTextInput.options.surface and FocusPointInput.options.surface were dead
  clones — never set by a projector, never read by a binder. Removed both.
- The Linux replay click disambiguates with role=button: the calculator tree
  carries [text] digit nodes beside the buttons, so a bare label=1 was an
  AMBIGUOUS_MATCH rejection on attempt-1 — which parked the pointer at
  (100,100) and made every retry hang in the xdotool no-op move, reporting as
  the step-7 focus timeout.

Live-verified on iPhone 17 Pro at this head: coordinate focus -> bare type ->
route synthesized-first-responder -> text observable in the captured tree.

* fix(ci): lower the settle.test.ts pin, add pixel evidence to the Linux type wait

- settle.test.ts shrank to 2359 lines when its dead typeText fixture member
  left; the ratchet pin follows it down (the history-backed gate caught the
  gap on CI while the local affected run had not re-selected the ratchet).
- The Linux replay now screenshots the calculator right after the type step.
  The lane uploads test/screenshots/replays/*.png on pass and fail, so a
  wait-155 timeout becomes diagnosable from the artifact: display showing 155
  means a tree-exposure gap; an empty display means the input never landed.
  The 20-ref divergence dump cannot show the entry node either way.

* fix(replay): assert the Linux type through the computed result

Run 32485981780's typed-state artifact settled both open questions with one
image: the display shows "55" — the bound typeText keystrokes LAND on Linux
CI — while the wait for that value timed out, so the calculator entry does not
expose its text to selectors; and the display shows no leading "1", so the
resolved click on button 1 missed its target entirely.

Both discoveries leave the replay: the click dependency goes (a pre-existing
click-coordinate defect is not this unit's evidence chain), and the assertion
moves to where the tree can answer — the typed string is now a full
calculation ("100+55=") whose `=` creates a history row, and the wait matches
the computed 155. No button carries that label and the typed string never
contains it, so the wait passes only if the keystrokes executed; deleting the
type step turns it red. The typed-state screenshot stays as per-run pixel
evidence either way.

* fix(linux): guard the no-op --sync mousemove; assert typed digits via value=

Two artifact PNGs decided this. Run 32485981780 shows "55" in the entry while
the wait for it timed out on a mismatched target; run 32487868346 shows
"100+55" — digits and + land, the = keystroke does not, and the retries died
in the focus step again because the previous commit removed the pointer-moving
click.

- moveTo now probes `xdotool getmouselocation --shell` and skips the --sync
  move when the pointer already sits on the target: a no-op move emits no
  motion event and hangs until the action timeout, which is what reported
  every failed replay retry as its first coordinate step. A failed probe never
  blocks the move. The provider test pins both sequences, including the
  skip case.
- The replay types digits only ("155") and waits on `value=155`: the AT-SPI
  dumper reads the entry's Text interface into the node's value and the
  selector engine matches it — the earlier "exposure gap" conclusion came
  from a pair that never tested the matching value. No button carries 155, so
  the wait passes only if the keystrokes executed.

* fix(replay): keep Linux type at contract tier — GTK4 exposes no entry text

Run 32490373693 closed the investigation: the typed-state artifact shows
"155" in the calculator entry while the wait for value=155 timed out with no
interactive filtering in play and a matcher that does compare node.value. The
AT-SPI dumper's get_text_iface() route returns nothing for GTK4
gnome-calculator, so no tree-level assertion on typed text can hold on this
lane today.

The replay keeps the type step and uploads the typed-entry screenshot every
run — live pixel evidence that the migrated typeText path lands keystrokes on
real Linux hardware — and closes with the survival assertion. The manifest
claim returns to the contract tier with the reason written at the entry, and
the count pins follow. The GTK4 exposure defect joins the Linux input-defect
chip; fixing PyGObject-vs-GTK4 blind through CI rounds is not a sane loop.

The moveTo no-op guard from the previous commit stays: it is why this run
finally reported the true failing step on every attempt instead of the
step-7 hang.
2026-08-21 17:15:35 +02:00
Michał Pierzchała b8dd6a5854 refactor: tighten capture ownership boundaries (#1736) 2026-08-11 13:55:28 +02:00
Michał Pierzchała cdc754e6ed perf: speed up iOS agent recovery and streamline CLI guidance (#1700)
* Avoid interactive children in parent taps

* docs: streamline no-skill CLI help

* perf: recover faster from sparse iOS trees

* fix: preserve selector context for blocked parent taps

* fix: preserve coordinate text-entry focus

* fix: preserve thin parent touch targets

* fix: fail closed for unscoped iOS typing

* test: isolate replay lock fixture

* test: share node integration process
2026-08-10 20:43:01 +02:00
Michał Pierzchała 6c0fcb64a1 fix: reject distinct ambiguous mutation targets (#1667)
* fix: reject distinct ambiguous mutation targets

* fix(ios): scope the raw-match rejection to mutating dispatches

`RunnerTests+Interaction.findElement` applied the new fail-closed
classification to `querySelector` as well as press/type, because the read
call site takes the default `allowNonHittableFallback: false`. With one
visible/hittable match and one non-hittable same-selector duplicate the
query started returning AMBIGUOUS_MATCH where it previously selected the
hittable element, and `queryDirectIosSelectorOrFallback` preserves that
error for read callers — so `get`, `is`, and `wait` surfaced an error
instead of their prior answer.

`classifyDirectSelectorCandidates` now takes a `rawMatchPolicy`. Mutations
keep `.rejectDistinctMatches` (the default, so no mutation call site
changes); `queryElement` passes `.preferHittableMatch`, restoring the
prior read rule: prefer the single hittable match, ambiguous only when
hittable matches compete, and never adopt the Maestro coordinate fallback.
The Maestro expected-point path is untouched.

Covers the one-hittable + one-non-hittable read, competing hittable reads,
and the non-hittable-only read. ADR 0011's amendment now states the scope.

* test(ios): execute selector read ambiguity regression

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-08 08:57:32 +02:00
Michał Pierzchała ee473b6adc refactor(daemon): give the Maestro fallback and ambiguous-match details real types (#1612)
Three places smuggled structured data through untyped bags and re-read it
with runtime guards. Each gets an explicit typed boundary.

A. The resolution-suppression rule was encoded twice in
   interaction-touch-response.ts — a spread ternary in the runner-payload
   branch and an unconditional destructure used conditionally in the runtime
   branch, with the ADR 0012 rationale living on only one source variant.
   Both branches now read one `suppressesResolutionDisclosure(source)`
   predicate through one `applyResolutionDisclosurePolicy` helper, where the
   reason is stated once. The union field is renamed
   `maestroCoordinateFallbackDispatched` (the dispatch path that ran) and
   hoisted into a shared base. handleFillCommand's two-arm interactor.fill
   call collapses to one.

B. `Interactor.type` narrows from `Record<string, unknown> | void` to
   `TypeTextBackendResult | void`; the Apple runner boundary is the single
   place the wire payload becomes that type. `maestroFallbackDetails` returns
   a typed `{ used, extra }` instead of a bag both call sites re-read.

C. `details.candidates` meant two incompatible things. The device-domain
   resolvers now key their list `devices`, so the shared renderer drops its
   shape-disambiguation guards and the device list actually renders.
2026-08-05 14:32:25 +02:00
Thiago Brezinski a13a6832ee feat: add selector-targeted drag gestures (#1567)
* feat: add selector-targeted drag gestures

* fix: address drag gesture review feedback

* fix: satisfy drag review quality gates

* fix(android): lower drag trajectories piecewise

* test(replay): validate drag fixture selectors

* fix(ios): ignore full-viewport chrome containers

* test(drag): prove destination on live devices
2026-08-05 12:37:02 +02:00
Michał Pierzchała 611858103e fix(ios): harden Bluesky-class interaction reliability (#1588)
* fix: type into focused iOS inputs without AX

* fix: fill AX-hostile iOS text inputs

* fix: keep scrolling containers from stealing taps

* fix: stop agents after explicit task success

* chore: format benchmark guidance

* fix(ios): preserve fill semantics across fast paths

* test: retire direct selector fill expectations

* test: assert runtime selector fill evidence

* fix(ios): preserve verified and Maestro fill paths

* refactor(ios): isolate synthesized text entry

* fix(client): preserve open diagnostic paths

* fix(ios): expose structured text entry route

* fix(packaging): strip text entry policy tests
2026-08-05 08:02:44 +02:00
Michał Pierzchała 0ee2a86129 refactor: extract contracts workspace package (#1499)
* refactor: extract contracts workspace package

* fix: preserve screenshot diff result contract

* test: stabilize Android keyboard smoke
2026-07-30 17:07:46 +02:00
Michał Pierzchała cd9a7ce41b test(android): add comprehensive emulator E2E coverage (#1482)
* test(android): add catalog emulator smoke coverage

* test(android): use stable snapshot diff mutation

* test(android): assert actual back destination

* test(android): separate keyboard and fill IMEs

* fix(ci): keep Android timing report in one shell

* refactor(test): simplify simulator e2e coverage

* test(android): assert stable diff landmarks

* fix(android): release snapshot helper gracefully

* fix(android): fully release snapshot helper runtime

* test(android): report coverage classifications

* fix(android): stabilize accessibility root capture

* fix(android): bound UiAutomation connection

* ci: upload worktree daemon diagnostics

* fix(android): cancel stalled wait captures

* fix(android): bound helper fallback lifecycle

* fix(android): harden emulator e2e lifecycle

* fix: align e2e changes with kernel package

* test(android): prove alert helper reuse directly

* fix(android): cancel stalled settle captures

* fix(android): separate helper retirement budgets
2026-07-30 15:10:21 +02:00
Michał Pierzchała 2e6849ddb6 test: stabilize parallel provider scenarios (#1497) 2026-07-30 14:41:37 +02:00
Michał Pierzchała 76453add71 refactor: pnpm workspace + @agent-device/kernel pilot (#1490 W0) (#1494)
* refactor: pnpm workspace + @agent-device/kernel pilot (#1490 W0)

Extend the workspace with packages/* and move the kernel behind an
enforced public API: packages/kernel with nine consumer-earned subpath
exports (errors, device, snapshot, contracts, collections, rect,
redaction, daemon-error, bounds — the last absorbed from utils as Rect
vocabulary). Every kernel import repo-wide becomes the
@agent-device/kernel/<sub> specifier; kernel tests move to
src/__tests__/kernel/ and exercise the package surface. The root
declares the package in devDependencies (workspace:*), tsdown bundles
it (noExternal) so the published artifact and its runtime dependency
manifest are unchanged.

Gate rewiring in the same change, per the W0 brief:
- R1 kernel-sink retires (physically subsumed); new R11
  package-boundaries guards no-root-back-imports, relative tunnelling
  past exports maps, undeclared workspace deps, and non-exported
  subpaths, with runtime resolution pins via import.meta.resolve.
- resolveImportEdges and mutation ownership follow workspace
  specifiers through exports maps, keeping R4 cycle checks, depgraph,
  and derived test ownership connected across the seam (kernel-errors
  still owns 495 tests). listSourceFiles includes packages/*/src.
- kernel becomes an unranked zone; mutation registry, stryker mutate
  globs, and the mutation-affected workflow path filter move to
  packages/kernel/src/errors.ts.
- check:affected gains packages/ ownership (manifests fail open);
  vitest and coverage include packages/*/src; fallow ignores
  packages/** (its resolver cannot follow workspace specifiers).
- The affected-selector CI job installs dependencies: its closure now
  crosses workspace specifiers, and the R8 relative exception is
  unsafe for production src files (Node ESM does not realpath, so dual
  specifier/relative loads would instantiate modules twice). The R8
  zero-dep set is pinned empty with that rationale.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUv7bvbWNryuXgSBuqTtep

* fix: address W0 review — mutation sandbox, exports-map resolution, tsc -b

Review findings on #1494, all five:

1. contracts-schema-public.test.ts reads the kernel source at its
   packages/ path (fs access invisible to the codemod and typecheck).
2. Mutation lane: Stryker sandboxes the tree but pnpm's node_modules
   symlink resolves @agent-device/* back to the real repo, so mutants
   in the sandbox never load and vitest.related finds no tests.
   vitest.mutation.config.ts now aliases each EXPORTED specifier to
   its source (derived from exports maps, never a wildcard), keeping
   resolution inside the mutated tree. Validated: kernel-errors module
   runs end to end (dry run 3,984 tests, mutants killed, exit 0).
3. Layering/depgraph resolve workspace specifiers through the
   exports-derived map (workspaceSpecifierTargets) instead of
   reconstructing paths, so '.'-facade packages resolve; the
   positional fallback remains only for map-less fixtures (P0 pin).
4. Per-package project references implemented: packages/kernel is
   composite (emitDeclarationOnly -> dist-types, gitignored), the root
   references it, and typecheck becomes tsc -b — probed to catch type
   errors on both sides under TypeScript 7 native.
5. R11's relative-route exception now requires membership in an actual
   R8 zero-dep job closure (zeroDepClosureFiles walks entries), not
   mere scripts/ placement — closing the dual-instantiation bypass.

Also from review discussion: daemon-error moves out of the kernel
package to src/client/ — its consumers (cli, client facade) rehydrate
wire DaemonErrors client-side; the daemon only produces them. Kernel
drops to 8 exported subpaths before any of them ship.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUv7bvbWNryuXgSBuqTtep

* refactor: one exports-map reader for mutation alias and ownership

Fallow flagged workspaceExportAliases (cognitive 15, CRAP 90). The
manifest-reading logic already exists as workspaceSpecifierTargets in
scripts/layering/package-boundaries.ts, so both the Stryker sandbox
alias table and the mutation ownership walker now consume it instead
of carrying near-clones. Behavior unchanged; mutation suite 45/45 and
changed-code fallow green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUv7bvbWNryuXgSBuqTtep

* fix: composite kernel without a root references edge

FreeRange runs plain `tsc -p tsconfig.json`, and a root `references`
entry makes non-build-mode TypeScript demand the referenced project's
built declarations (TS6305) — a standing "build first" tax on every
plain -p consumer (fr, editors). Keep the per-package composite
project and build it in typecheck (`tsc -b packages/kernel` before the
root and examples/sdk passes), but drop the root references edge: root
consumption resolves through exports to source, identical to runtime
and to the bundler. Probed: plain -p green with no prebuilt output;
kernel-side type errors still caught by its own build.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUv7bvbWNryuXgSBuqTtep

* fix: R11 uses the layering parser; mutation config is a fallow entry

Review blockers on #1494:

- R11's private single-quote regex could miss a double-quoted or
  re-export route into packages/*/src. specifierSites now delegates to
  the layering model's parseImports (both quote styles, side-effect
  imports, re-exports, dynamic imports), with direct regressions for
  each formerly-invisible form.
- vitest.mutation.config.ts becomes a declared fallow entry instead of
  a tolerated unused-file finding: the full-repo audit now reports it
  reachable (unused files 2 -> 1; the remainder predates this PR).

FreeRange clean-checkout evidence: with packages/kernel/dist-types and
every *.tsbuildinfo deleted, `pnpm check:freerange` reports 0 findings
on this head — the TS6305 topology died with the root references edge
in the previous commit; check:freerange has no build precondition.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUv7bvbWNryuXgSBuqTtep

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-30 12:12:46 +02:00
Michał Pierzchała 95f838f514 test: stop pinning settle capture counts against a wall-clock loop (#1307)
* test: model the UI in settle-observation fixtures instead of a capture count

settle-observation's "contention flakes" were a zero-margin comparison
meeting a 1ms clock skew, not generic load.

runStableCaptureLoop derives pollMs = min(300, max(25, quietMs)), so the
test's settleQuietMs: 25 made pollMs === quietMs, and the settle check
(one sleep(25) plus capture time, against >= 25) a 0ms margin. Node's
setTimeout(25) advances Date.now() by only 24ms in 0.13% of calls idle
and 0.63% under load, because libuv's timers and Date.now() read
different clocks. On a 24 the loop takes a third capture that the
transcript never scripted, and settle's best-effort catch reports the
resulting throw as settled: false.

The fixtures now model the surface rather than the runner's speed: a
quiet UI serves the settled tree to every capture, a busy one a fresh
tree per capture, via new transcript `repeat` entries and result
factories. The snapshot-floor economy guard survives as a bound (2-3
captures), and the follow-up's "no fresh capture" cost — previously
implied by the exact transcript — is now asserted directly.

Production is untouched: the same zero margin only costs a wasted extra
capture and poll there, filed separately as #1306.

* test: stop pinning the settle capture count in the iOS contract scenario too

A sweep for the same bug class found direct-ios-selector's settleObservation
scenario carrying the identical 0ms margin: settleQuietMs: 25 (so pollMs ===
quietMs) against a consume-once transcript scripting exactly two settle
captures, with no injected clock. It has not lost the coin flip in CI yet,
but it fails the same way when it does — a third capture finds no entry and
settle's best-effort catch reports settled: false.

Same fix: the fixture models a quiet UI (every settle capture sees the same
tree) instead of scripting how many captures fit in a wall-clock window.

The rest of the sweep was clean. The other contract scenarios and the
interaction runtime tests already use clamped mocks that repeat the last
snapshot, so any capture count is tolerated; the fake-clock tests are correct
to pin exact counts.

* style: oxfmt quietRunnerSnapshotEntry signature

* test: make one-shot-outranks-repeat a real transcript rule (P2 review)

The review is right: `one-shot entries still outrank a repeat entry` asserted
a guarantee the lookup did not provide. It passed only because the one-shot
happened to be declared first — unordered lookup took the first match, so a
repeat declared ahead of a matching one-shot shadowed it forever and left it
permanently unconsumed. Both failure modes reproduce; the reverse-order test
added here fails on the previous implementation.

Unordered lookup now searches matching one-shots before repeats, so
outranking holds whatever the declaration order. A repeat is documented as
its command's fallback.

Ordered transcripts now reject repeats at construction: ordered lookup only
ever reads the head, so a repeat there never advances and strands every entry
behind it. Refusing the combination beats failing later as a confusing
"Provider command mismatch".

Coverage added for both: reverse declaration order, and ordered + repeat.
2026-07-16 20:13:07 +02:00
devin-ai-integration[bot] 8246362999 chore: baseline-free production-exports cleanup (#1276) (#1282)
* chore: baseline-free production-exports cleanup (#1276)

Classify and burn down the 32 baseline-tolerated unused production exports.

- Live seams: annotate with @internal JSDoc visibility tags (test hooks,
  introspection helpers, public install-source constant) so fallow no longer
  treats them as dead production exports.
- Wrappers: collapse re-export wrappers in commands/index.ts (ref/selector)
  and daemon/lease-context.ts (buildLeaseDiagnosticsContext); update all
  importers to pull directly from the source module.
- Stale baseline entry: remove the non-existent
  resetAndroidMultiTouchHelperInstallCache entry.
- Empty fallow-baselines/production-unused-exports.json so
  check:production-exports now fails loudly on any new dead export.

Fixes #1276

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: address review feedback on production-exports cleanup (#1276)

- CONTRIBUTING.md: document that intentional non-production exports should use
  JSDoc @internal with a short justification, treated as a reviewed baseline entry.
- isPlatform: fix JSDoc tag to "@internal" and remove conflicting "public" wording.
- ARCHIVE_EXTENSIONS: re-export from src/sdk/install-source.ts so the public
  install-source subpath has a real consumer story for the constant.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* chore: make production-exports check truly baseline-free (#1276)

- Drop --baseline from pnpm check:production-exports and remove the
check:production-exports:baseline generation script.
- Delete fallow-baselines/production-unused-exports.json.
- Update CONTRIBUTING.md to describe the baseline-free behavior and remove
references to reviewed baseline entries for production unused exports.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Michał Pierzchała <thymikee@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-16 11:44:07 +02:00
Michał Pierzchała 37895caf99 refactor: replace Maestro compat with typed direct engine (#1217)
* test: add pinned Maestro conformance harness

* feat: add typed Maestro program IR parser

* docs: define direct Maestro engine architecture

* test: compare Maestro oracle with typed IR

* feat: add direct Maestro program engine

* refactor: narrow Maestro execution context

* refactor: tighten Maestro program parsing

* fix: verify iOS Maestro visibility waits

* refactor: isolate retained Maestro runtimes

* refactor: type Maestro target resolution

* refactor: harden typed Maestro execution

* refactor: share in-page swipe planning

* feat: add typed Maestro runtime port

* refactor: parse Maestro suite metadata from typed IR

* refactor: centralize Maestro include loading

* feat: execute Maestro files through typed engine

* refactor: share replay built-in variables

* fix: make Maestro target intent explicit

* fix: refresh Maestro targets before input

* refactor: format Maestro progress from typed IR

* feat: compile typed Maestro replay plans

* feat: bind typed Maestro runtime to public commands

* feat: route Maestro YAML through typed runtime

* refactor: remove legacy Maestro runtime

* refactor: remove obsolete replay control model

* refactor: split typed Maestro plan modules

* fix: harden typed Maestro runtime semantics

* docs: update direct Maestro architecture

* fix: reconcile Maestro runtime with merged contracts

* fix: harden typed Maestro execution boundaries

* fix: harden typed Maestro runtime evidence

* perf: avoid eager Maestro device resolution

* refactor: finalize typed Maestro execution

* fix: reject Android system-only helper snapshots

* fix: preserve Android system dialog snapshots

* fix: make helper-backed CI deterministic

* refactor: invalidate Maestro observations before dispatch

* fix: make Maestro selector policy explicit

* refactor: remove Maestro ranking sentinels

* refactor: make Maestro own observation stabilization

* refactor: source Maestro compatibility presets

* refactor: keep Maestro failure reports typed

* refactor: simplify Maestro runtime policy

* fix: isolate Maestro engine failures

* refactor: consolidate Maestro swipe presets

* fix: align Maestro selector and observation semantics

* fix: preserve atomic iOS Maestro taps

* fix: require semantic uniqueness for Maestro taps

* fix: preserve Maestro parse provenance

* docs: pin Maestro compatibility presets

* docs: reconcile Maestro gesture viewport contract

* perf: resolve Maestro gesture viewport directly

* test: align Maestro replay regressions

* fix: order Android gesture lift after endpoint

* fix: settle Maestro gestures before continuation

* fixup! fix: order Android gesture lift after endpoint

* refactor: normalize Maestro swipes once

* refactor: fail impossible Maestro observations

* refactor: normalize Maestro defaults alias

* test: reconcile Android provider scenarios

* fix(android): synchronize single-pointer move events

* test: align repair digest parsing

* refactor: type Maestro runtime operations

* refactor: keep Maestro controls compact

* refactor: name Maestro diagnostic limit

* fix: align Maestro parser and settling semantics

* fix: complete Maestro compatibility semantics

* docs: define Maestro compatibility boundaries

* fix: refresh iOS runner target after relaunch

* fix: reset prewarmed iOS runner after URL open

* fix: preserve iOS Maestro target and swipe intent

* fix: harden direct Maestro runtime semantics

* fix: preserve ranked Maestro replay suggestions

* fix: align maestro tap runtime semantics

* fix: stabilize maestro ci contracts

* fix: tighten maestro runtime architecture

* fix: reconcile maestro replay with latest main

* perf: tighten Maestro iOS stabilization

* fix: preserve Maestro app lifecycle sessions

* fix: restore Maestro CI coverage

* fix: address Maestro engine review findings

* refactor: consolidate Maestro compatibility internals

* fix: scope Maestro target evidence to childOf
2026-07-15 21:26:42 +02:00
Michał Pierzchała c93dcdbc90 feat: disclose selector resolution in interaction responses (#1193)
* feat: disclose selector resolution in interaction responses

Implements ADR-0012 migration step 1 (decision 2). Adds an additive
`resolution` field to press/click/fill/longpress responses: runtime-selector
carries the full pre-action diagnostic shape (unique or disambiguated with
matchCount/winnerDiagnostic/tiebreak/bounded alternatives), runtime-ref and
native-ref carry the exact ref-provenance shape, direct-ios-selector carries
the explicit not-observed marker, and coordinate/maestro-non-hittable-fallback
stay inapplicable (no field). The comparator in selectors-resolve.ts now
records which criterion (visible/deepest/smallest-area) decided each
disambiguation without changing resolveSelectorChain's winner.

Extends the ADR-0011 guarantee matrix with the resolutionDisclosure guarantee
across all six dispatch paths, wires the shared response builder and MCP
output schema, adds digest-level trimming (drops alternatives, keeps the
verdict/counts), and proves via contract tests that resolution diagnostics
are never ref-issued or MCP-pinned and cannot be reused as @ref targets.

* fix: address resolution disclosure review findings

* refactor: make resolution-disclosure choices self-evident

Replace the direct-iOS/maestro message-sniffing (and its justification
paragraph) with an explicit maestroFallback flag passed from the dispatch
site that already owns the path decision, and shrink every why-this-is-OK
paragraph to one-line constraint statements per the maintainer directive.

* fix: usage-based maestro fallback disclosure + spec label-fallback

Blocker 1: the runner-payload source now carries maestroFallbackUsed derived
from the runner's actual execution outcome (the usedNonHittableFallback
message bit RunnerTests+CommandExecution.swift reports, the same signal
directIosSelectorFallbackDetails already keys on) instead of the permission
flag. A fallback-allowed dispatch that hit its element normally discloses
direct-ios/not-observed; only an actually-executed coordinate fallback is the
inapplicable maestro cell. Contract tests cover both sides.

Blocker 2: ADR-0012 decision 2 now defines the ref/label-fallback disclosure
(runtime-ref trailing-label recovery via tryResolveRefNode's fallbackLabel;
native-ref stays exact because the backend receives only the ref handle),
amends the matrix-cell enumeration, layer-3 coverage list, and validation
bullet, and the runtime-ref contract suite proves the label-fallback shape.

* fix: honest runtime-ref registry cells for label recovery

The disambiguation cell no longer claims refs identify exactly one node by
construction — trailing-label recovery is a first-match lookup without the
ranking, now an intentional waiver whose outcome the label-fallback
disclosure surfaces per-response. resolutionDisclosure.via points at
tryResolveRefNode (now exported), the resolver producing both exact and
label-fallback, with direct unit coverage of both outcomes.

* docs: correct native-ref exactness rationale and tiebreak doc

Native-ref forwards fallbackLabel to the backend; exact is justified by
non-observability of any backend-side label recovery, not by non-forwarding.
The tiebreak doc now states the derived winner-vs-runner-up decisive margin.

* fix: disclose Maestro fill fallback usage
2026-07-11 09:09:44 +02:00
Michał Pierzchała 58995bf51b feat: parse and preserve .ad target-v1 evidence (ADR 0012 migration step 3) (#1196)
* feat: parse and preserve .ad target-v1 evidence (ADR 0012 migration step 3)

Recording now emits a `# agent-device:target-v1 {...}` comment immediately
before every click/press/longpress/fill action that resolves through the
tree (ref or selector), carrying the identity/ancestry/sibling/scrollRegion/
viewportOrder tuple from decision 3's record-time write algorithm plus a
record-time self-check (verified/unverifiable). The parser accepts known
fields in any order, NFC-normalizes, rejects malformed/oversized annotations
with INVALID_ARGS, binds an annotation only to the immediately-next action
line, and treats unknown future target-vN comments as ordinary comments.
writeReplayScript's read-then-rewrite (heal) path preserves v1 annotations
in canonical form. Nothing consumes the parsed evidence yet beyond
preservation — replay-time enforcement is a later migration step.

The identity tuple's own node/tree only ever lived on the internal
visualization/session-history payload (never the public response); it is
converted to the compact target-v1 evidence and stripped before anything is
persisted or returned.

* fix: bound candidate identity before self-check comparison

The identity-set scan in computeTargetEvidence compared an untruncated
candidate identity against the 256-byte-bounded recorded identity, so a
node whose own id/label exceeds the field cap could fail to match itself,
corrupting the record-time self-check. Compare through the same bounding
on both sides instead.

* fix: address PR #1196 review — contract boundary, formatting, path-6 reasons, worst-case sizing

- native-ref contract test: deliberately updated to assert the new ADR 0012
  boundary — the preflight's node/preActionNodes ride the INTERNAL runtime
  result and visualization payload only; the public responseData never
  carries node/preActionNodes/targetEvidence (asserted directly against
  buildInteractionResponseData).
- oxfmt formatting on all touched files.
- classifyTargetBindingMatch path 6 now distinguishes decision 3's two
  spec-distinct outcomes: a signal isolating a member that differs from the
  winner ('signal-isolated-wrong', the paths-4/5 comparison class, future
  identity-mismatch) vs true fall-through ('no-signal-isolation', future
  identity-unverifiable), so step 4 can consume the reasons directly.
- writer reduction loop sizes every candidate against the worst-case
  verification value ("unverifiable" is 4 bytes longer than "verified"), so
  a fail-closed self-check downgrade can never push an accepted payload over
  the 4 KiB cap; pinned by a 1-byte-granularity sweep across the boundary
  that fails against the old placeholder-sized check.

* fix: reject missing role keys in target-v1 annotations, correct data-flow comment

Second-review nits: the writer emits `role` unconditionally (top level,
ancestry entries, scrollRegion) — possibly as the empty string for typeless
nodes, which stays accepted — so a MISSING role key can only come from a
hand-edited/adversarial annotation and is now rejected with INVALID_ARGS
instead of silently parsing as an implicit empty role, which step-4
enforcement could otherwise match against anonymous wrapper nodes. Also
corrects the interaction-touch-response comment that overstated where the
raw node/tree flows (finalizeTouchInteraction strips it before both session
history and touch overlay telemetry).

* fix: record iOS selector target evidence

* refactor: make the record-time evidence channel structural, drop argument-comments

Maintainer directive: comments that argue a workaround is acceptable mark
code to redesign. Applied to the whole PR diff:

- The record-time node/tree now travel on a typed `recordedTarget` side
  channel of InteractionResponsePayloads instead of being smuggled through
  the visualization Record and stripped later. The construction site routes
  them there exclusively, finalizeTouchInteraction consumes the channel
  directly, and extractTargetEvidenceForRecording (the strip helper and its
  justifying doc) is deleted — the public/internal split is now enforced by
  the type shape, so the contract test asserts it in two lines instead of a
  paragraph.
- findNearestScrollableContainer/findNearestAncestor made generic over the
  node type, removing a cast plus its safety-argument comment.
- computeLocalIdentity/boundedLocalIdentity collapsed into one always-capped
  identity reader — the raw/bounded split was the root of the earlier
  self-match bug and existed only to be explained.
- parseReplayScriptDetailed's rejectUnbound takes the pending annotation as
  a parameter, removing a non-null assertion and its comment.
- Remaining paragraph-length argument-comments trimmed to one-line
  constraint statements (module docs, sizing/floor comments, regex/role
  rationales, test comments); spec-mapping docs stay.

* chore: oxfmt the selector-evidence test files

* fix: record get target evidence, fail closed on broken parent linkage

Maintainer blockers on ADR 0012 step 3:

- get text/attrs now records target-v1 evidence: GetCommandResult carries
  the resolution tree (preActionNodes, internal — neither the recorded
  result nor the public payload copies it), recordIfSession gains the typed
  recordedTarget channel, and dispatchGetViaRuntime threads the capture
  through. The direct-iOS get query is gated during recording so the
  snapshot path supplies the evidence tree, mirroring the tap/fill gating.
  find/is stay uncovered: their results are match-set shaped (found flags,
  predicate booleans over possibly-many matches), not a single resolved
  winner — noted in the PR.
- buildAncestryChain now reports a broken parent walk (dangling parentIndex
  or cycle) instead of silently producing a root-like chain; the writer
  fails the annotation closed to 'unverifiable' per decision 3's
  capture-anomaly rule, and broken candidates cannot prove an ancestry
  prefix. Regression tests for both anomaly shapes.
- New ADR 0012 recording tests live in focused files
  (interaction-target-evidence.test.ts) instead of growing the
  interaction.test.ts aggregation; the earlier additions moved out.

* refactor: extract direct-iOS get selector guards below the complexity threshold
2026-07-10 18:43:17 +02:00
Michał Pierzchała 9dfebbe3be fix: omit interaction uptime from wire responses (#1142)
* fix: omit interaction uptime from wire responses

* fix: simplify interaction wire sanitizer
2026-07-07 16:53:44 +02:00
Michał Pierzchała 0dcc1aa553 fix: normalize interaction response wire shapes (#1114)
* fix: normalize interaction response wire shapes

* fix: reduce interaction response complexity
2026-07-07 11:29:01 +02:00
Michał Pierzchała 5c5fa012f7 feat: --settle returns the settled diff in the interaction response (#1101) (#1106)
* feat: --settle returns the settled diff in the interaction response (#1101)

press/click/fill/longpress --settle executes the action, waits for the UI
to go quiet (wait stable's loop, shared via stable-capture.ts), and returns
the settled diff vs the pre-action tree in the same response — one round
trip instead of the interact -> observe pair.

- payload: changed lines only (bounded), summary counts, added-line refs,
  refsGeneration; best-effort (settled:false + hint on never-quiet content,
  never an action failure); --verify shares the settle captures
- ref issuance: the settled tree becomes the session snapshot; a
  diff-carrying settle response clears snapshotRefsStale and the MCP layer
  merge-only re-pins added-line refs at the settle generation
- grammar: --settle + --settle-quiet <ms> + --timeout <ms> (flag-sourced
  descriptor budget with new envelope:'widen' semantics mirroring wait)
- ADR 0011: new settleObservation guarantee classified on every path with
  contract scenarios per enforced/delegated cell

* test: give the two contention-flaky doctor scenarios explicit budgets

The doctor provider scenarios sit at ~5s of real daemon-harness work on a
loaded host and flake at vitest's 5s default during full-suite runs (the
known contention flake AGENTS.md documents). Same in-file precedent as the
Metro-probe scenario's 10s budget.

* fix: move SettleParams to contracts to satisfy the layering DAG

daemon/handlers/interaction-flags.ts imported the type across the
daemon -> commands boundary (R2 commands-floor). The tuning params are
part of the interaction contract like SettleObservation, so they live
in contracts/interaction.ts and both layers import from there.

* feat: keep settle diffs content-first — drop Key nodes, added lines win the cap

Bluesky dogfood: a fill that summons the iOS keyboard spent 49 of the 80
capped diff lines spelling out QWERTY keys, and a screen transition with
269 removals could starve out the added lines entirely. Key-type nodes
are now filtered from both diff sides (the [keyboard] container line
still signals presence), and under truncation added lines — the ones
carrying fresh refs — win slots over removals.

* docs: state the core loop in the top-level help starting point

Benchmarked with headless haiku/sonnet agents given only --help: both
models skipped the help-workflow pointer and started with plain
snapshot (38KB payloads they then had to re-read from files). One
core-loop line at the starting point is what teaches snapshot -i and
--settle to models that never read a second help page.

* fix: preserve settle digest refs for mcp

* fix: reduce settle fallow complexity

* fix: surface settle output in CLI text

* fix: complete settle handling for longpress

* refactor: localize daemon timeout envelopes

* refactor: deepen post-action observation

* refactor: centralize post-action observation planning

* refactor: derive settle capability from descriptors

* refactor: trim settle descriptor helpers
2026-07-06 20:18:44 +02:00
Michał Pierzchała ad754ac0a4 feat: direct iOS selector falls back to tree resolution on semantic failures (ADR 0011) (#1091)
* feat: direct iOS selector falls back to tree resolution on semantic failures (ADR 0011)

Runner ELEMENT_NOT_FOUND/AMBIGUOUS_MATCH on the direct iOS selector
dispatch now delegate to the tree-based runtime path, which supplies
runtime disambiguation, occlusion refusal, non-hittable
promotion/annotation, and rich selector diagnostics/hints.

Maestro replay dispatches (allowNonHittableCoordinateFallback) keep the
runner-native error shapes: the new delegateSemanticFailures option is
threaded from the dispatch flag, and the query path
(allowElementNotFound at selector-runtime) keeps its exact semantics.

Registry: direct-ios-selector errorTaxonomy/nonHittable/occlusion flip
from gap waivers to delegated cells; the pinned gap list shrinks 6 -> 3.
Success-path parity cells (disambiguation, responseIdentity)
intentionally remain gaps per ADR 0011.

Refs #1081

* test: contract coverage for the three newly-delegated direct-iOS cells

The coverage gate demanded scenarios the moment the rebase flipped
errorTaxonomy/nonHittable/occlusion to delegated — as designed. Three
transcript scenarios prove the delegation end to end: runner
ELEMENT_NOT_FOUND falls back to runtime no-match diagnostics, to the
covered refusal, and to an annotated coordinate tap (the transcript
asserts the fallback tap is coordinate-keyed).
2026-07-04 16:41:03 +02:00
Michał Pierzchała f3e07ff236 test: interaction contract suite with registry-driven coverage gate (ADR 0011 Layer 3) (#1092)
* fix: preserve the runner's non-hittable fallback marker through direct selector press

The direct iOS selector press handler spread successText('Tapped <selector>')
after the runner payload, clobbering the 'tapped via non-hittable coordinate
fallback' message that directIosSelectorFallbackDetails keys on — so
maestroNonHittableCoordinateFallbackUsed could never be true end-to-end (the
existing unit test passes because it mocks dispatchCommand above this layer).
Found by the ADR 0011 Layer-3 maestro-fallback contract scenario.

Share the marker string as MAESTRO_NON_HITTABLE_FALLBACK_MESSAGE and keep it
as the success message when the runner reports fallback usage.

* test: interaction contract suite with registry-driven coverage gate (ADR 0011 Layer 3)

test/integration/interaction-contract/ holds one scenario file per dispatch
path, each with a sibling .coverage.ts manifest declaring which guarantee
matrix cells it proves (scenario strings double as the vitest test titles).
index.ts aggregates the manifests statically, and a new Layer-3 gate in
src/contracts/__tests__/interaction-contract-coverage.test.ts fails when any
enforced (runtime/runner/delegated) cell lacks a scenario or when a scenario
claims a waived/inapplicable cell — coverage of the matrix is by
construction in both directions.

Path forcing is natural (no test-only env switch needed): selector/ref
targets take the runtime path, a tapTarget backend takes the native-ref fast
path, x/y takes the coordinate path, and simple-selector clicks on an iOS
provider transcript take the direct runner path (the transcript itself
proves which path ran via assertComplete).

Fixtures are the permanent Bluesky shapes: closed drawer, drawer + visible
twin, edge-grazing container, covered button, non-hittable cell.

Refs #1081
2026-07-04 16:02:41 +02:00