* feat: add scale-only screenshot sizing
* fix: refuse retired --max-size inputs on every released surface
Released sizing inputs must fail closed with migration guidance instead of
silently producing native-size artifacts:
- contracts: RETIRED_SCREENSHOT_MAX_SIZE declaration + SCREENSHOT_SCALE_LIMITS
as the single source for the scale bounds and migration messages
- .ad parser: released 'screenshot ... --max-size N' and 'record start ...
--max-size N' lines now refuse at parse time (frozen replay-compat witnesses)
- daemon: screenshot rejects old-client screenshotMaxSize like recording does;
the recording guard now shares the same contract data
- Node client: screenshot/record daemon writers refuse the removed { maxSize }
option before transport
- CLI: --max-size unknown-flag error carries the migration guidance
- config/env: stale screenshotMaxSize config keys and the retired
AGENT_DEVICE_SCREENSHOT_MAX_SIZE env var are refused for sizing commands
(other commands keep working)
Quality: numberField now reuses the canonical readOptionalNumber contract
helper (AppError bounds instead of plain Error); png-resize inlines one-use
wrappers and restores the worker-thread rationale; docs typo fixed.
* test: drop retired maxSize entries from the MCP undocumented-input allowlist
* fix: refuse retired maxSize at the MCP field-projection seam + release-provenance corpus witnesses
- readFieldInput silently dropped undeclared keys before the daemon writers
could refuse them, so an MCP call carrying { maxSize } reached transport and
returned native-size success. New retiredField() combinator declares the
removed key in the field map: the projection seam refuses it with the
canonical migration message and the JSON schema no longer advertises it.
Real-route MCP executor regressions cover screenshot and record.
- replay-compat corpus: derived v0.20.5 witnesses for the released screenshot
and record --max-size forms (SHA-256 pinned, new retired-capture-size
coverage surface) so check:replay-compat proves the shipped syntax refuses
with migration guidance instead of degrading silently.
---------
Co-authored-by: Michał Pierzchała <thymikee@gmail.com>
* fix(ios): pin tap-outcome corroboration probes to the baseline's backend
The recorded-failure screens are exactly where the capture plan flips
between XCTest and private-AX (the penalty boundary), so #1605's
same-backend requirement failed closed right where XCTest tap false
negatives actually happen: the baseline was captured via private-AX
under penalty, the probe came back via tree, and a landed tap surfaced
as XCTEST_RECORDED_FAILURE. In the AppControlBench bsky-16 run this
fired four times, each sending the model into a re-observe/retry spiral.
The comparison stays same-backend by design (backends are not comparable
views of a screen); instead the probe is now CAPTURED the way its
baseline was: a new internal preferredBackend option (never CLI-exposed)
threads daemon -> runner, and a private-AX-preferred capture takes the
exact penalized route — privateAX-first plan, 'deferred' verdict, no
degradation warning, no settle budget reset.
Live-verified on the deterministic repro (Bluesky drawer-menu press
under penalty, seeded bench feed): errored with the backend-mismatch
diagnostic before, corroborates as landed after, with no mismatch phase
in the request diagnostics. Daemon tests cover pinned and unpinned
baselines end to end through the dispatch context; the Swift plan gate
is a pure function with an executed in-bundle test (added to the ios.yml
regression list).
* style: oxfmt
* fix: exclude raw baselines from corroboration and prove the pin end to end (review)
Raw baselines could not be pinned: the raw diagnostic plan keeps
tree-first error propagation by contract and is never rerouted by the
penalty or the preferred backend, so preserving 'raw: true' on the probe
recreated exactly the backend-mismatch false failure this PR removes.
Corroboration now declines raw baselines up front (they are diagnostics,
not evidence baselines) with a regression pinning that no probe capture
is dispatched at all.
The wire is now regression-proven at every hop: a dispatch-level test
drives dispatchCommand with the context flag and asserts the emitted
RunnerCommand carries preferredBackend (red if handleSnapshotCommand or
the interactor stops forwarding); the injected-transport test asserts
the interactor's snapshot payload both ways; and a runner unit test
decodes the wire JSON, projects it through the extracted
snapshotOptions(from:), and composes it with the plan rule — pinned
regular plan defers to privateAX-first, RAW plan stays untouched.
Executed on-simulator; added to the ios.yml regression list.
* feat(ios): extend depth-capped private-AX captures via element-rooted requests
The AX server's depth limit is per-request (kAXErrorIllegalArgument above
a tree-size-dependent value), so a capture capped at depth 56 on
Bluesky-class React Native trees returned chains of unlabeled [other]
containers and hid every actionable control below the cap — agents fell
back to screenshot-and-coordinate guessing.
After a capped serialization, the bridge now re-issues the same snapshot
request rooted at each deepest-level childless node's live accessibility
element, splicing the returned children in and chaining further while
capped subtrees remain — bounded by a call limit, the shared node
budget, and the capture-plan deadline. The AX server counts node levels,
not edges: a maxDepth=56 request emits nodes to depth 55, so the
frontier is the deepest observed level, kept only when it sits at the
cap; a tree that ends naturally above the cap yields no frontiers and
costs nothing.
A capture whose frontier extension drained every capped node no longer
reports itself depth-limited — the re-run hint it used to trigger could
not add anything. Explicit --depth requests stay exact captures with no
extension.
Live on the seeded Bluesky bench feed: 8 extension calls (~106ms each)
turn the 117-node all-[other] tree into a 182-node tree carrying post
text, testID links, and every feed control (Reply/Repost/Like/options);
steady-state capture 0.52s -> 1.37s. Observation-only captures paying
the extension needlessly is #1626.
* fix: count missed frontiers and tighten deep-extension shape (review)
The completeness verdict inverted in the failure paths: a frontier whose
live element vanished or whose re-rooted request failed was silently
dropped, so an all-miss extension reported pendingFrontiers=0 and the
capture presented as complete while whole subtrees were missing. Missed
frontiers are now counted, logged, and keep the depth-limited verdict —
the pure decision lives in privateAXDepthLimited with in-bundle tests.
The truncation hint no longer advertises --depth (an explicit --depth
capture disables extension, so following it returned strictly less than
the capture that produced the hint); the honest remedy is a plain
re-run with a fresh extension budget.
Shape: candidates collected only at the cap boundary and not at all
when extension is disabled (exact --depth captures pay zero
bookkeeping); the zero-caller 3-arg overload is gone; the response
shape's keys are shared constants; the unsupported-selector diagnostic
is restored; the exact-depth policy is hoisted to one named local.
* test: execute the deep-extension miss-path contract in CI (review)
The depth-limited regression compiled but never ran (absent from
ios.yml's -only-testing list), and the pure Swift consumer test was
vacuous against the Objective-C producer — a literal missedFrontiers
proves nothing about the increments. RunnerAXSnapshotFrontier and
extendSnapshotFrontiers move to the header as the executed contract
seam, and testDeepExtensionCountsMissedFrontiers drives both real miss
paths with fabricated snapshots: a nil accessibilityElement (explicit
nil property — bare NSObject resolves the key through a UIKit category
and takes the call path instead) counts missed without consuming a
call; a resolving element whose client cannot serve the re-rooted
request consumes a call AND counts missed. Both tests join the
executed ios.yml list.
Red-before verified on-simulator: with both increments stripped the
producer test fails at the counter assertion (0 != 2); restored, both
tests pass.
* fix(ios): corroborate recorded tap outcomes
* fix(ios): preserve corroborated tap target identity
* fix(ios): suppress corroborated tap retries
* fix: require comparable iOS tap evidence
* fix: bound iOS tap corroboration baseline
* test(ios): deterministic injection seam for recorded-tap-failure corroboration (#1605 merge gate)
The field failure cannot be reproduced on this head: the tap false-failures
were a downstream symptom of XCTest-channel saturation, which the #1587
capture fixes removed. The seam records a real XCTIssue AFTER the real
gesture inside the per-command failure-count window, so
xctestRecordedFailureResponse and target invalidation fire byte-for-byte
like the field failure. Armed via a decrementing /tmp flag file (the daemon
regenerates tampered xctestrun templates, so env plumbing cannot reach a
daemon-spawned runner); compiled only under AGENT_DEVICE_RUNNER_UNIT_TESTS.
Live evidence on a daemon-spawned runner (Bluesky, ad-bsky-repro sim):
- landed case: injected failure on a real Search-tab tap -> success with
the corroboration warning, screen verifiably on Search, no redispatch,
runner serving next commands; flag consumed exactly once.
- unchanged case: injected failure on a dead-coordinate tap -> capture
unchanged -> XCTEST_RECORDED_FAILURE preserved with the new honest hint;
runner still usable.
- field-shape race (relaunch -> full snapshot -> immediate press, 5
attempts): no natural recorded failure occurs on this head — the hostile
tree needed for channel saturation is gone, corroborating the causal
story.
* fix: reconcile tap corroboration with current interaction semantics
RunnerSynthesizedGesture.m and RunnerSynthesizedTextEntry.m each independently
reflected into XCSynthesizedEventRecord/XCPointerEventPath. Extract the
common resolution (classes, addPointerEventPath:/setTargetProcessID:/
synthesizeWithError:/processID, the RunnerRequireClass/Selector/
ApplicationSelector checks, the shared objc_msgSend typedefs, and the
"name: reason" exception formatter) into RunnerXCTestEventBridge.h/.m.
Each caller resolves its own extras on top of the shared core: gesture
resolves the 2-arg initWithName:interfaceOrientation: plus touch-path
selectors, text entry resolves the 1-arg initWithName: plus text-input
selectors.
* fix: type into focused iOS inputs without AX
* fix: fill AX-hostile iOS text inputs
* fix: keep scrolling containers from stealing taps
* fix: stop agents after explicit task success
* chore: format benchmark guidance
* fix(ios): preserve fill semantics across fast paths
* test: retire direct selector fill expectations
* test: assert runtime selector fill evidence
* fix(ios): preserve verified and Maestro fill paths
* refactor(ios): isolate synthesized text entry
* fix(client): preserve open diagnostic paths
* fix(ios): expose structured text entry route
* fix(packaging): strip text entry policy tests
* fix(ios): give keyboard dismiss a safe-area-tap fallback (#1598)
The runner already tapped a keyboard's own Hide/Dismiss/Done key when the
AX tree exposed one, but iPhone's default software keyboard has no such
key, so `keyboard dismiss` returned UNSUPPORTED_OPERATION on the common
case and agents proceeded with the keyboard (and any live QuickType
predictive-text bar) still up.
Live-validated on throwaway simulators before choosing a design: hardware
escape key (no effect without a connected hardware keyboard), swipe-down
starting on the keyboard (does not trigger UIKit's interactive dismissal
on Settings/Safari/Contacts), and a private
`performAction:onElement:value:error:` AX call (hung the runner for 90s on
a guessed action name, force-killed by the daemon timeout) were all ruled
out. The dismiss-key tap remains the primary mechanism (iPad, or any app
with an inputAccessoryView Done/Cancel button); a new snapshot-derived
safe-area tap is added as the disclosed last resort, computed to land
outside both the keyboard and every currently-hittable element so it is a
safe no-op even when it fails to dismiss.
The response now discloses which mechanism actually fired
(`mechanism: 'dismissKey' | 'safeAreaTap'`) across the CLI/daemon dispatch
path, the SDK runtime.backend surface, and session-event summaries, so
callers can tell a real dismiss-key press apart from a best-effort tap.
UNSUPPORTED_OPERATION now says both mechanisms were tried.
* fix: satisfy CI formatting and complexity gates
oxfmt on three touched files; buildKeyboardActionSummary split so the
dismiss wording (incl. the safeAreaTap mechanism disclosure) lives in its
own helper below the complexity threshold.
* fix: any-element obstacle rule for the safe-area dismiss tap (#1606 review P1)
A role allowlist cannot prove a point is AX-empty: an unlabeled RN
Pressable surfaces as a hittable Other, and a tappable parent can cover a
point its static-text child does not. Every known element frame now counts
as an obstacle regardless of role or hittability, with only ~window-sized
structural frames exempt (isStructuralRootFrame, 95% coverage) — exempting
those is what keeps the rule satisfiable, and a genuinely tappable
full-screen backdrop staying exempt is the disclosed, accepted behavior of
this fallback. One .any resolution replaces ten typed queries (single tree
snapshot, no per-element isHittable round trips), so the stricter rule is
also cheaper.
* fix: drop the safe-area tap — background-tap dismissal is unsupported (#1606 review P1, round 2)
No geometry or role query can prove a coordinate is side-effect-free: after
the any-element rule, the structural-root exemption still deliberately
removed full-screen actionable elements (RN Pressable backdrops) from the
obstacle set, so the tap could navigate or submit — and report success
because the mutation hid the keyboard. Per review, generic background-tap
dismissal is now explicitly unsupported: the dismiss key is the only
mechanism the runner vouches for, UNSUPPORTED_OPERATION says so and steers
callers to press-the-next-target / keyboard enter, and the mechanism field
narrows to 'dismissKey'. Unrecognized wire mechanisms degrade to the bare
message and are dropped from event details.
The Apple runner ships as source in dist/apple/runner, and packaging strips
#if AGENT_DEVICE_RUNNER_UNIT_TESTS blocks — but tests outside such blocks
shipped whole and compiled on every user's machine. Two files leaked six
tests this way (RunnerTests+LifecycleCacheTests, RunnerTests+SnapshotTraversalIdentityTests).
Wrap the strays and close the class: packaging now fails if any XCTest-shaped
method (func test*) survives stripping, with testCommand in RunnerTests.swift
as the only allowlisted entrypoint.
* perf(ios): stop re-paying known-failing work on penalized private-AX captures
Live-measured on the Bluesky bench feed (139-171 nodes), each private-AX
capture wasted ~1.35s of its ~1.65s runner-side cost re-doing work a prior
capture already proved futile:
- ~1000ms: the viewport read is XCTest main-thread work; under the channel
penalty it reliably burns its full timeout and falls back to the root
frame anyway. Honor the penalty in privateAXSnapshotViewport the same way
capture plans do.
- ~310ms: the depth ladder re-paid the kAXErrorIllegalArgument rejection of
the default depth on every capture. Remember the accepted rung per bundle
(penalty-duration TTL, cleared on target process change); expiry re-probes
the full depth so screens that recover are not capped forever. Explicit
--depth requests bypass the memory in both directions.
Steady-state hostile-screen captures drop 2.65s -> ~0.55s CLI wall, and
press --settle round-trips drop ~5.4s -> ~2s (settle needs two captures).
* fix(snapshot): distinguish penalty-deferred captures from genuine recoveries
A capture whose backend was PRE-selected by the XCTest-channel penalty was
stamped with the same recovered verdict as one that ground through a live
failure. Two costs followed on hostile screens (Bluesky bench: 130 repeats
per 30-task run):
- the daemon repeated the full fell-back warning on every capture, long
after the arming capture had already said it once, and
- settle's one-shot private-AX budget reset fired on every loop even though
the capture paid no grind to give the budget back for.
The deferred plan now stamps reasonCode 'deferred' (a new code; older
daemons drop unknown codes and keep today's behavior). The daemon keeps the
verdict recovered but suppresses the repeated warning, keeps the depth-cap
line, and skips the settle budget reset for deferred captures.
* style: wrap reasonCode union to satisfy oxfmt
* fix(ios): bind accepted-depth memory to the process, pin deferred through the wire parser
Review follow-up (#1587):
- The depth memory was cleared only in refreshCachedTargetIfProcessChanged;
resetTargetAfterExternalRelaunch -> invalidateCachedTarget drops the cached
PID without clearing it, so the next activation had no old PID to compare
and could reuse a stale shallow rung for up to 120s (also A->B->A when A
restarted while inactive). The memory now stores the PID it was learned
under and only matches the same live process; recording without a PID is
refused. Every invalidation path is covered automatically because they all
drop currentAppProcessIdentifier.
- The deferred settle/warning tests constructed typed verdicts directly, so
removing 'deferred' from the accepted reason-code set would silently
restore the repeated warning and budget reset while tests stayed green.
They now parse a raw runner-wire object through readSnapshotQualityVerdict
(red on base: parser strips the code -> warning re-appears, reset fires).
- Extracted shouldReadPrivateAXViewportViaXCTest() and pinned the penalized
viewport skip with an in-bundle regression test.
* fix(ios): wait for post-dismiss content settle before the next gesture (#1542)
Dismissing the keyboard can trigger the app's own ScrollView content-offset
correction (e.g. releasing the inset it grew to keep a focused field above
the keyboard). That correction is a separate, unsynchronized animation that
`keyboard.waitForNonExistence` knows nothing about — the keyboard AX element
can disappear well before the app visually settles. The very next command is
frequently a synthesized, AX-free drag (scroll/gesture, kept AX-free so it
still works under #1105-family AX degradation), which has no XCTest
quiescence wait of its own, so it can land mid-animation and net to zero —
the "scroll does nothing" symptom on the Form screen's checkout-form.ad leg.
Add a bounded, AX-free screenshot-stability wait to dismissKeyboard() so the
runner only returns once the screen has actually stopped changing (or a
generous cap elapses). The stopping decision is a pure function
(runnerScreenshotStabilitySettled) covered by unit tests under
AGENT_DEVICE_RUNNER_UNIT_TESTS; the surrounding capture/sleep loop is the
thin, untestable I/O shell around it.
Live-verified on iPhone 17 Pro / iOS 26.2: the scroll now visually lands at
the correct position (confirmed via screen-recording frame correlation)
instead of leaving content at its pre-scroll offset.
Not a full fix for #1542: the checkout-form.ad corpus leg still fails at the
same step, now because the daemon's shared post-gesture snapshot
stabilization (src/daemon/post-gesture-stabilization.ts) can read a
stale-but-internally-consistent AX tree after the AX-free scroll and
mistake "unchanged across polls" for "settled", so the following click's
off-screen guard sees pre-scroll node positions. That is a cross-platform,
cross-command stabilization semantics change and needs a design decision,
not a unilateral fix here — see the PR description.
* ci(ios): execute the screenshot-stability runner tests (#1559 review)
* chore: remove verified dead code and migration scaffolding
Multi-agent audit of accumulated waste, every finding adversarially
verified against call sites, git history, and the published surface
before removal. Net -710 lines.
- delete src/core/platform-descriptor/ (superseded ADR-0009 migration
scaffold; parity tests now assert an inline table)
- remove test-only seams: registry introspection exports,
CommandFacet.extraDaemonWriters, MaestroEngineOptions.timing
- remove dead flexibility: backend capability allow-list,
screenshot-diff maxRegions, CloudWebDriverSupportLevel 'partial',
clearFirst on the TS+Swift runner wire contract
- remove dead deprecated surface: --session-locked /
--session-lock-conflicts aliases (hard migration error now points at
--session-lock), replay export --format single-value enum,
unused Lease*Payload contract types, runtime-layer rotate duplicate
- collapse pass-throughs/duplication: withRetry adapter,
default-cloud-artifact-provider, connect-profile client-id hashing
(3x sha256 impls -> one helper, byte-identical output), shared
scripts walker, cloneValue -> structuredClone, fill-diagnostics
moved into android/
BREAKING CHANGE: --session-locked and --session-lock-conflicts now fail
with a migration error pointing at --session-lock; replay export
--format is removed (Maestro was the only value); Lease*Payload types
are dropped from the ./contracts subpath.
* chore: satisfy fallow gates tightened by #1363/#1364 after rebase
- drop the consumer-less AndroidFillVerificationNode re-export
- reuse requireSnapshotSession in resolveSnapshotForRef instead of
inlining the same authorized-frame resolution (fallow clone group);
the helper's return type now guarantees the session it already
throws for
* chore: address review — keep cloud-webdriver partial capability metadata
The partial/supported/unsupported levels and their notes are part of the
lease-response capability contract for genuinely limited operations
(Appium page-source snapshots, upload-then-install), not dead
scaffolding. Restore them and the asserting tests unchanged from main.
Also add the missing CHANGELOG entry for the Lease*Payload type removal
from agent-device/contracts.
tvOS has no SpringBoard (it uses PineBoard/HeadBoard), so probing
com.apple.springboard for blocking system modals raises an XCTest failure
("Application com.apple.springboard is not running") that safely() cannot
trap, terminating the whole runner test. Every snapshot and alert
resolution then failed, and the CLI misreported it as the screen
"overwhelming the accessibility capture".
resolveBlockingSystemModal — the single chokepoint every snapshot/alert
path funnels through — assumed a SpringBoard host always exists. Model that
assumption explicitly with `hasSpringBoardSystemModalHost` and bail to
`.absent` when no host exists, so app-owned alert queries still run. The
fourth probe site (shouldRouteToSpringboardBlockingSystemModal) is already
`#if os(iOS)` and needs no change.
Add tvOS regression tests covering the snapshot and alert-resolution paths.
Move the scattered root-level native projects into per-platform folders and drop
the now-redundant platform prefix:
- android-ime-helper/ -> android/ime-helper/
- android-multitouch-helper/ -> android/multitouch-helper/
- android-snapshot-helper/ -> android/snapshot-helper/
- apple-runner/ -> apple/runner/
- macos-helper/ -> apple/macos-helper/
- src/platforms/linux/atspi-dump.py -> linux/atspi-dump.py
Only repo source paths move. Identity surfaces stay frozen so no user's runner
cache is invalidated on upgrade: the derived-cache key hashes source paths
relative to AgentDeviceRunner and excludes packageVersion, and the
~/.agent-device/{apple-runner,macos-helper} namespaces, the
agent-device-android-*-helper artifact/manifest/protocol names, the
AgentDeviceRunner Xcode project, and the `prepare ios-runner` CLI command are
unchanged. Updates build/package scripts, CI, package.json files+scripts,
ignore/attr/fallow configs, runtime path resolvers, and test fixtures.
Also: re-base repo-root-relative refs inside the moved apple/runner for the
added nesting level (gated XCUITest fixture walk + two doc links), and clean the
legacy dist/apple-runner packaged output so the relocated runner can't
double-ship into the wholesale-included dist (with a regression test).