Files
Michał Pierzchała 338aa2a0d5 refactor: route every native selector resolution through the policy interface (#1715)
* refactor: route every native selector resolution through the policy interface

#1649 declared the per-caller ambiguity matrix; four native call sites still
bypassed it, spreading `selectorResolutionKnobs(row)` into a raw
`resolveSelectorChain` instead of naming the row. That left the "one
interface" claim aspirational: a caller could restate its contract as engine
knobs and nothing would notice.

- `is` non-exists, `get text`/`get attrs`, find's read actions, and the
  covered-selector diagnosis probe now call `resolveSelectorChainWithPolicy`
  with their existing row. Semantics are byte-identical: the knob-backed
  branch of that interface forwards to the same engine call the call sites
  built by hand.
- The façade drops `resolveSelectorChain` and `selectorResolutionKnobs`, so
  no knob-taking resolver is reachable from outside the package and a call
  site cannot re-acquire the knobs even by accident.
  `requireUnique`/`disambiguateAmbiguous` are now named in exactly one
  function, which `resolve-with-policy.ts` and the replay resolver both
  derive through.
- `get` names the two rows it may consume as a type, so pointing it at any
  other ambiguity contract is a compile error.

Tests: selector-read-policy.test.ts pins which row each read command
consumes, end to end, on one ambiguous fixture — the only tree the rows
disagree on. Each assertion was proven red by re-pointing its caller at a
neighbouring row. The knob-consistency check moves into the package beside
the now-private helper. Test call sites that used the raw resolver move to
`resolveRecordedTarget`, the same knobs and the path that actually replays a
recorded chain.

Extracting the failure branch drops `resolveSelectorInteractionTarget` below
the complexity threshold; its `fallow-ignore` waiver is removed (verified
load-bearing before the extraction, unnecessary after).

Closes #1630. Structural stages (occlusion, off-screen, promotion, poll
budget) stay per-caller pipeline code, tracked in #1656.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD

* test: observe which node find's row selected, not just that one existed

#1715 review, P2: the find row assertion was only half a pin. `find exists`
returns `found: true` for any resolved node, and the `list` call it leaned on
goes through listFindMatches — a path that consumes no policy row at all. So
repointing findFirstLocatorMatch at `readText` left both assertions green
while selection silently moved from the document-order head to the tiebreak
winner.

Assert through `find get_attrs`, which returns the ref of the node the row
actually selected. Both neighbouring rows are now red: `readText` fails
'@e3' !== '@e2' (the move the old test missed), `readUnique` fails by
refusing the ambiguous screen. `exists` stays as a second, weaker assertion
on the same resolution.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD

* refactor: route is exists through the matrix, collapse the double match pass

Follow-up tightening on the same seam.

`is exists` reached findSelectorChainMatch directly while the `readAny` row's
own doc claimed to serve "`exists` and find's read-only actions" — true of the
docs, not of the code, which is the unverifiable-claim shape #1656's review
called out. It now names `readAny`, the row it always described. Equivalent by
construction: both take the first alternative with any match under
requireRect: false, and disclose that alternative's count.

That leaves the root façade with no consumer for findSelectorChainMatch, so it
goes the way of resolveSelectorChain — dropped from the string-only façade,
kept on the published ./ast surface. Its façade-twin type SelectorChainMatch
dies with it (fallow caught it).

resolveSelectorChainWithPolicy matched twice on the uniqueness path: once via
resolveSelectorChain, then again to fill matchedNodes. Hoisting the single
list call above the row switch removes that second pass, collapses two
duplicated ambiguous literals into one helper, and drops a `?? [resolution.node]`
fallback that was unreachable — a resolution implies its alternative matched,
so the list is never null there.

While hoisting: the resolved arm's matchedNodes can describe a different
alternative than resolution.selector, because uniqueness skips an ambiguous
alternative to try the next one. Unreachable today (only first-match callers
read it, where both come from one list), and left as-is rather than silently
changed — but the doc claimed "the alternative it came from", so it now says
what is actually true.

Tests: is exists gets a caller-level pin on the shared ambiguous fixture —
passes with matches: 2 where its fail-closed siblings refuse — proven red by
pointing it at readUnique.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD

* test: discriminate is exists's row by alternative, guard the façade structurally

#1715 review, second regression-validity gap. The `is exists` pin observed
only `pass: true` and `matches: 2` on a fixture whose first alternative was
merely TIEBREAKABLE — so disambiguation succeeded there and reported the same
count first-match would. `readAny`, `readText`, and the pre-migration raw
lookup all produced that, and only the readUnique swap I had checked went
red. One mutation proven is not the same as the row being pinned.

`exists` exposes no node ref, so the row has to be read off WHICH alternative
answered. New fixture: alternative one matches two nodes that are genuinely
indistinguishable (same depth, same area, both on screen) so the tiebreak
declines; alternative two matches exactly one. First-match answers from
alternative one; every uniqueness row skips the undecidable alternative and
answers from alternative two. Asserting the selector now separates them —
readText and readUnique both fail with `id="save-unique"` where
`label="Save"` is expected.

Restoring the raw lookup stays behaviourally invisible, though:
findSelectorChainMatch is equivalent to the readAny row it migrated to, which
is precisely why that migration preserved semantics. No fixture assertion can
catch that revert, so the guard is structural — the façade's export list must
not carry resolveSelectorChain, findSelectorChainMatch, or
selectorResolutionKnobs. Follows the packages/maestro index.test.ts
absence-assertion precedent. Verified red by re-exporting the lookup.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD

* fix: cover selector routes in device replays

* test: simplify selector replay regression

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-11 07:34:11 +02:00
..
2026-04-26 20:49:59 -04:00

Agent Device Tester

Agent Device Tester is a minimal Expo Router fixture app for agent-device experiments.

It is intentionally small, but each surface is dense with durable accessibility targets so a few screens cover a large share of the workflows we care about.

Why this app exists

  • It gives agent-device a stable React Native target that we control.
  • It keeps the number of screens low while still covering roughly 50 practical interaction and verification cases.

Screens

  • Home: visible-text checks, dismissible banner, modal open/close, async loading, status badge, switch state
  • Catalog: search debounce, filter chips, direction-aware scroll canary, favorite toggles, cart updates, drill-in navigation
  • Product detail: back navigation, quantity stepper, multiline notes, save action
  • Checkout form: required-field validation, fill vs type, checkbox state, choice groups, keyboard dismiss, success summary
  • Settings: switch rows, accordion content, loading and error states, retry flow, destructive-confirm modal
  • Automation lab: long-press, alert-result, app-event, app-state, appearance, orientation, permission-recovery, and log canaries
  • WebView accessibility: a deterministic semantic fixture plus live websites with varied HTML for native accessibility snapshot verification

Navigation uses Expo Router native bottom tabs, so the tab bar itself is also part of the test surface.

The deterministic WebView fixture is the stable accessibility oracle. On iOS, an interactive snapshot should expose its root as webview, both titles as heading, paragraph and label content as text, and the form controls as text-field, switch, and button. The live-site buttons are exploratory smoke coverage for real-world WebKit trees, not stable assertion targets. Use an unscoped snapshot for this oracle: XCTest can detach a scoped WebKit document subtree from its WebView ancestor, leaving insufficient evidence for safe semantic projection.

Coverage map

These are the main case families this app can support without adding more screens:

  • app open and close
  • visible text verification with plain snapshot
  • interactive discovery with snapshot -i
  • press on stable buttons, pills, and rows
  • fill on single-line and multiline fields
  • type after focus for append flows
  • get text on headings, badges, summaries, and accordion content
  • is visible and is exists assertions
  • wait for async loading and success states
  • diff snapshot after dismissals and submits
  • long-list scrolling and scrollintoview
  • selector-based navigation across repeated cards
  • modal open, cancel, and confirm flows
  • switch and checkbox state changes
  • validation-error and recovery loops
  • retryable error banners
  • cart counters and quantity changes
  • screenshot and recording proof capture

Run locally

This fixture uses an Expo development build, not Expo Go. Expo's development-build workflow installs expo-dev-client, builds the native app with expo run:ios or expo run:android, and then serves JavaScript from Metro with expo start. The app declares @expo/dom-webview directly to keep Expo's development runtime on the SDK 56 native module; Android verification failed when the dev client resolved an older transitive copy.

Build cache

Local pnpm test-app:ios / test-app:android cache the native build on disk via the expo-build-disk-cache provider (configured in app.config.js), keyed by the Expo fingerprint. A second run with no native change reuses the first build instead of recompiling; editing screens never needs a rebuild, because Metro serves JavaScript at runtime. A fresh checkout still pays for the first native build — the disk cache only spares you the repeats.

That fingerprint is why /ios and /android are gitignored: ignoring the prebuild output is what makes @expo/fingerprint treat this app as CNG and skip hashing it. Un-ignore them and the fingerprint starts describing your machine rather than the project.

CI does not use the disk cache. .github/workflows/test-app-build-cache.yml builds a Release binary (JS embedded, no Metro) per platform when the fingerprint has no artifact yet, and publishes it as a GitHub Actions artifact named fingerprint.<hash>.<platform>. Jobs that drive the app install it through .github/actions/setup-fixture-app, which downloads that artifact and refreshes its JS with @expo/repack-app (~seconds) — so a JS-only change reuses the same native binary. A consuming job needs permissions: actions: read.

The /automation route is intentionally JavaScript-only and can be opened from Settings → Open automation lab or with the agent-device-test-app:///automation scheme. Its stable automation-* ids expose durable outcomes for long press, native alert actions, app-event name/payload, app state, appearance, window orientation, and microphone permission recovery. CI repacks JavaScript-only changes into the cached Release app without starting Metro; native configuration changes intentionally produce one new fingerprinted build that all simulator consumers share.

iOS simulator

From the repo root, install dependencies and run the development build on the target simulator:

pnpm test-app:install
pnpm test-app:ios -- --device "iPhone 17 Pro"

expo run:* keeps Metro in the foreground after launching the app. Leave that terminal running, then use a separate terminal for agent-device or Maestro commands.

iOS physical device

Use the physical device name from agent-device devices --platform ios or xcrun devicectl list devices. Keep the expo run:ios terminal running so Metro stays visible to the development build:

pnpm test-app:install
pnpm test-app:ios -- --device "<physical device name>"

Then verify the installed development build from another terminal with the same physical device identifier:

agent-device open com.callstack.agentdevicelab --platform ios --udid "<physical udid>" --session test-app-physical
agent-device snapshot -i --platform ios --udid "<physical udid>" --session test-app-physical

The snapshot should show the Agent Device Tester home screen, for example the Agent Device Tester heading and tab bar. An already installed com.callstack.agentdevicelab is not enough evidence by itself: confirm Metro is running for the development build and verify the visible app surface before using the session for manual logs, network, replay, or interaction checks. Close the same session when verification is complete:

agent-device close --platform ios --udid "<physical udid>" --session test-app-physical

AccessorySetupKit picker fixture

The Settings tab links to a dedicated Accessory setup lab backed by a local Expo module. The development client uses this fixed test service UUID, so no build-time environment variables are required:

FFF0

Advertise that service from the test accessory, build with the normal physical-device command above, then open Settings → Open accessory setup lab. The picker requires physical iOS 18+ hardware; use the normal session hygiene above when validating its snapshot, wait, and selector paths.

Android emulator or device

Install dependencies and run the development build on the target Android emulator or device:

pnpm test-app:install
pnpm test-app:android -- --device "$ANDROID_DEVICE"

For Android app/package launches connected to local Metro, run adb reverse for the Metro port when needed before opening the app with agent-device.

Running from the app folder

If you prefer to work from inside the app folder:

cd examples/test-app
pnpm install --ignore-workspace
pnpm ios

Or on Android:

cd examples/test-app
pnpm install --ignore-workspace
pnpm android

After the first native build is installed, use pnpm test-app:start when you only need to restart Metro for JavaScript or TypeScript changes. test-app:start starts Metro only; it does not build, install, or prove a physical device is running the development build. Once the app is running and verified with snapshot -i, use agent-device against Agent Device Tester like any other target app.

Non-default Metro ports

If the default Metro port is already in use, start Metro on another port. Do not reinstall the native development build just to change the JavaScript server port:

pnpm test-app:start -- --port 8082

If you are building and installing for the first time in that terminal, Expo's run:ios and run:android commands also accept --port:

pnpm test-app:ios -- --device "<device name>" --port 8082
pnpm test-app:android -- --device "$ANDROID_DEVICE" --port 8082

After the development build is installed, keep using the same native app. The current agent-device open CLI does not accept --metro-host or --metro-port; open the app normally, then use the Metro command surface for Metro-specific actions:

agent-device metro prepare --project-root examples/test-app --kind expo --port 8082 --public-base-url http://127.0.0.1:8082
agent-device metro reload --metro-host 127.0.0.1 --metro-port 8082

Use metro prepare when you want agent-device to start or reuse Metro and print the runtime URLs. Use metro reload when Metro is already running and the installed development build is connected to that server. For Android local device/emulator runs, also run adb reverse tcp:8082 tcp:8082 when the device needs host port forwarding.

Local Agent Device suites

The repo includes two local suites for iterating on the fixture app:

pnpm test-app:replay:ios
pnpm test-app:replay:android

These run the .ad replay suite in examples/test-app/replays.

The Android gesture replay pins coordinates to the CI emulator profile — pixel_7, 1080x2400 @ 420 dpi (gh workflow uses exactly this AVD). Run it on a matching emulator; on a different size or density the gesture card moves and the canary waits fail with a wait timeout naming the missed state, which is fixture geometry, not a product regression. The checkout replay is selector-driven and runs on any emulator.

The iOS gesture-lab.ad and Android gesture-lab-android.ad replays verify gesture pan, gesture fling, gesture pinch, and gesture rotate against the gesture metrics rendered by the Home screen. They also prove that the default pan does not activate an exactly-two-pointer recognizer, while gesture pan ... --pointer-count 2 does without changing pinch or rotation state.

Each gesture replay relaunches the app before its combined gesture transform canary, verifies the clean pan/pinch/rotate state, then checks that one atomic two-pointer gesture changes all three semantic states. On Android, these checks are intentionally qualitative because recognizers can report non-exact centroid, scale, and rotation values for one simultaneous two-finger gesture.

To target a specific iOS simulator or an installed Expo development build, run the underlying command directly so global flags stay before replay inputs:

node bin/agent-device.mjs test examples/test-app/replays \
  --platform ios \
  --device "iPhone 17 Pro" \
  --env APP_TARGET=com.callstack.agentdevicelab \
  --env APP_URL=<project-url> \
  --artifacts-dir .tmp/test-app-replay/ios

Omit APP_URL when the installed development build can discover the local Metro server from its launcher.

The Maestro prototype suite lives in examples/test-app/maestro and runs through agent-device replay --maestro:

pnpm test-app:maestro:ios
pnpm test-app:maestro:android

The Maestro flow includes launchApp, so the suite launches the app inside each test attempt. Start Metro first when the installed development build needs the local bundle.

The suite intentionally covers the compat layer syntax used by public Maestro suites: runFlow file/inline blocks, when.platform, config hooks, deterministic repeat.times, flow env, selectors, input, assertions, and swipe.