Commit Graph

63 Commits

Author SHA1 Message Date
Michał Pierzchała 7f3e355426 fix(ios): preserve regular snapshot depth through structural wrappers (#1947)
* fix(ios): complete regular snapshot depth frontier

* fix(ios): align depth frontier with visibility fold

* fix(ios): exercise regular depth frontier in CI

* fix(ios): cover visible-depth frontier through public snapshot

* fix(ios): tolerate absent deep-link confirmation

* test(ios): expose visible-depth fixture hierarchy

* test(ios): wait for visible-depth fixture subtree

* fix(ios): keep visible-depth fixture minimal

* fix(ios): update snapshot hint fixtures

* test(ios): avoid fixture label aggregation

* test(ios): match fixture raw hierarchy

* test(ios): prove visible-depth raw ancestry

* test(ios): align depth smoke with AX hierarchy
2026-08-22 13:53:39 +02:00
Michał Pierzchała e5bfde3d13 diagnose(1874): instrument the synthesized commit wait and add a dispatchable stall loop (#1941)
* diagnose(1874): instrument synthesized commit wait and add stall loop workflow

* diagnose(1874): fix empty-array expansion under set -u; raise default iterations

* diagnose(1874): add arm64 matrix leg to isolate the Rosetta factor

* ci: build the iOS runner for the native arm64 slice

A generic simulator destination leaves the active arch undefined; Xcode 26.6
defaults it to x86_64, running the whole runner under Rosetta on arm64 hosts.
Pin ARCHS=arm64 across every lane that builds the iOS runner and bump the
derived-data cache suffixes. Measured ~30% faster commits on identical CI
hardware; delivery-throttle episodes still occur but start from a lower base.

* diagnose(1874): keep commit-wait cadence evidence value-free

The per-poll trace logged the observed field's contents (prefix(40)) on the
shipped type path; that value is user content and runner.log persists. Log
lengths and the expected-prefix walk instead, allowlist every
string-interpolating NSLog format in the module behind a source-scan guard,
and pin commonPrefixLength in the host-lane policy tests.

* diagnose(1874): narrow the log-format match for typecheck

* diagnose(1874): route cadence evidence through a typed value-free boundary

logCommitCadence accepts Int lengths and a timestamp only, so observed field
contents are unrepresentable at the poll call site; its emitted line is pinned
by a sentinel-secret test in the host-lane policy tests. The source guard
becomes structural — boundary present, poll path logs through it, no raw NSLog
in the observe closure — instead of parsing Swift format strings. #1874 is
reopened as the removal-tracking thread for this temporary instrumentation.
2026-08-22 13:39:09 +02:00
Michał Pierzchała 991c08561b fix(ios): enforce regular snapshot clip invariant (#1946)
* fix(ios): enforce regular snapshot clip invariant

* fix(ios): restore typed snapshot failure construction

* fix(ios): linearize snapshot clip validation

* fix(ios): propagate snapshot presentation errors

* fix(snapshot): clarify presentation failure recovery
2026-08-22 12:15:40 +02:00
Michał Pierzchała 17da776350 feat(ios): add snapshot backend conformance (#1930)
* feat(ios): add snapshot backend conformance

* fix(ios): load built SDK at live runtime

* test(client): isolate snapshot forwarding regression

* refactor(snapshot): keep backend capability metadata internal

* fix(test): merge backend conformance imports

* fix(snapshot): keep backend forcing internal

* refactor(snapshot): isolate backend capability fixtures

* refactor(snapshot): keep capability governance internal

* fix(ios): align snapshot actionability contract
2026-08-21 15:01:10 +02:00
Michał Pierzchała 07023eb202 fix(ios): separate snapshot actionability from occlusion (#1933) 2026-08-21 12:47:25 +02:00
Michał Pierzchała d57aa69777 test: add macOS platform command coverage manifest (#1922)
* test: add macOS platform command coverage manifest

* fix: remove unused macOS coverage type exports

* test: route macOS coverage away from iOS lane

* fix: account for host-dependent macOS audio capability

* fix: run macOS coverage manifest in CI
2026-08-21 12:39:35 +02:00
Michał Pierzchała af96c6608d feat(ios): publish effective snapshot geometry (#1931) 2026-08-21 11:27:04 +02:00
Michał Pierzchała 73db7be2ff feat(ios): move the regular-projection clip fold into snapshot presentation (#1797) (#1929)
* feat(ios): move the regular-projection clip fold into snapshot presentation

Both iOS snapshot backends carried their own copy of the visibility fold: the
tree walker and the private-AX serializer each computed viewport-and-scroll-clip
intersection, ancestor projection, hidden-content hints, and collapsed depth
during acquisition. Hand-synchronized copies of that interpretation are what
produced the scroll-overflow leak class (#1784), and C1 (fact-availability
neutrality) could not hold while acquisition decided what a screen shows.

Acquisition backends are now fact serializers: every traversed node is emitted
at raw traversal depth with its reported frame, and SnapshotAcquisition carries
the viewport. presentRegular runs the one clip fold for every backend --
viewport ∩ scroll clip, the ancestor cursor (an out-of-clip Cell or scroll
container hides its clamped descendants), the sub-pixel decoration rule,
scroll hints booked onto anchors, reparenting with collapsed depth -- and
narrows the emitted hittable to the clip: nothing outside its clip, and nothing
without geometry, is ever hittable, whatever the backend reported. Platform
differences are a SnapshotFoldPolicy input to the shared algorithm (iOS
cursor-projected; macOS/tvOS plain viewport), never a backend exception.

The private-AX backend collapses to ONE serializer for both projections, and
the flat filter-decision family dies with the acquisition gates it fed.

Three intentional edge deltas, each toward one backend-neutral rule: sub-pixel
content-free decorations now drop on every backend (was private-AX only);
labeled offscreen Application/Window carriers survive on every backend (was
tree only), never hittable; query-sweep regular without -i is viewport-folded.
Declared acquisition residues: the traversal-depth budget cut, the sweep's
frameless-element drop, the private-AX bridge's device-side cap.

Refs #1797 (migration step 3, clip-fold delta).

* refactor(ios): isolate snapshot visibility fold
2026-08-21 11:27:04 +02:00
Michał Pierzchała 2a9a4ee80c test: add Linux platform command coverage manifest (#1921)
* test: add Linux platform command coverage manifest

* test: address platform coverage review feedback
2026-08-20 21:48:05 +02:00
Michał Pierzchała 4137e4275e feat(ios): split raw and regular snapshot projections behind one capture hint (#1926)
The private-AX backend interpreted `--raw` as the regular projection: it folded
the viewport and scroll clips and dropped sub-pixel decorations before returning,
so a raw capture that recovered onto it answered with viewport-pruned nodes
labeled raw (#1797 D4). Nothing related the two copies of that decision.

Presentation now exposes the two projections it always implied. `presentRegular`
folds visibility, eligibility, scope and scroll hints; `presentRaw` is the
acquired tree, normalized, narrowed only by a scope or depth the request asked
for. Acquisition reads one derived `CaptureHint` instead of the request itself,
so what a capture may skip is stated once, beside the reason skipping it keeps
the projection complete: scope and its relative depth never narrow, raw depth
does (raw depth is traversal depth), and the raw projection never carries
`interactiveOnly` — `--raw -i` is the acquired tree.

Two structural rules replace the hand-synchronized ones. The raw plan is derived
from `SnapshotBackendKind.supportsRawProjection`, so the query sweep — an
interactive element query with no hierarchy to return — cannot be planned for a
raw request. And presentation compares the requested projection with the hint
the acquisition was captured under, dropping that tier with a structured
`IOS_SNAPSHOT_PROJECTION_MISMATCH` failure rather than presenting it under the
requested label.

Declared residue: a regular `--depth` request still cuts the traversal at that
depth while regular presentation emits collapsed depth, so a node that would
present within the limit can be dropped. The cut is what keeps `--depth 1`
probes cheap; making it complete is the open visible-depth frontier obligation.

Refs #1797 (migration step 3, raw-projection delta).
2026-08-20 21:44:48 +02:00
Michał Pierzchała 06d27de4d0 test(gesture): assert pan duration in the iOS gesture-lab replay (#1901)
* test(gesture): assert pan duration in the iOS gesture-lab replay (#1584)

The only replay exercising the `gesture pan` command class that regressed
in #1562 asserted a counter, which stays green even if the requested
duration collapses — nothing in CI could catch the regression coming back.

Record an observed-duration bucket from a single-pointer Gesture.Pan's
begin/end timestamps in GestureLab.tsx (iOS-only, so Android's raw-touch
transform handling in the same shared component is untouched), render it
as plain text, and assert it with a one-line wait in gesture-lab.ad. No
runner protocol changes needed.

* style: fix oxfmt line-wrap in GestureLab.tsx

* ci(ios): run the pan-duration canary automatically on every PR

gesture-lab.ad (and its new duration assertion) only runs under full:fixture-replays,
which is currently dispatch-only in replays-manual.yml — the PR-triggered ios.yml lane
runs the smoke tier, and replays-nightly.yml no longer carries device replays at all
(#1781 A1). So the #1584 guard could not actually catch a regression automatically.

Split the duration check into its own minimal, isolated replay
(gesture-pan-duration.ad) and run it as a smoke-tier step in ios.yml, so it's cheap
and doesn't depend on gesture-lab.ad's multi-touch commands, which stay full-tier only.

* test: require pan recognition in duration canary
2026-08-20 19:08:33 +02:00
Michał Pierzchała b4331815e2 test: add web platform command coverage manifest (#1902)
* test: add web platform command coverage manifest

* fix: preserve web coverage report on cleanup failure
2026-08-20 15:59:50 +02:00
Michał Pierzchała 9ce1ef7587 feat(snapshot): move scope into presentation (#1855)
* feat(snapshot): move scope into presentation

Use one preorder label/identifier/value policy across Swift and TypeScript, keep scoped iOS acquisition conservative, and remove the daemon's second scope pass.

Non-vacuity: label-only matching failed identifier/value parity fixtures; Android pass-through failed its boundary test; disconnecting Swift applyScope produced eight scope/depth/projection failures.

* fix(snapshot): select scopes with presented content

* docs(snapshot): describe presentation-owned scope

* docs(snapshot): record contribution-aware scope

* refactor(snapshot): drop unrelated provider churn

* fix(snapshot): accept healthy empty scoped capture

* refactor(snapshot): isolate empty-scope admission

* fix(snapshot): align scope ownership across runtimes

* test(snapshot): pin post-wire scope owner

* test(snapshot): retain find test shrink
2026-08-20 10:30:18 +02:00
Michał Pierzchała 393eb30a28 ci: give check:affected real Apple ownership rules and route ios.yml on them (#1781 A9-2) (#1857)
* ci: give check:affected real Apple ownership rules and route ios.yml on them (#1781 A9-2)

Device-lane ownership by platform family in the affected selector
(scripts/check-affected/device-lanes.ts): a TypeScript-only Apple change now
carries replay-ios/replay-ios-device/replay-macos in a narrow plan, other
families own only their own lanes, shared runtime surface owns every lane,
unit tests own none. Golden tables (contracts/fixtures) own the parity unit
test and both runner builds instead of failing open.

ios.yml pull_request paths-ignore is routed on that ownership; the gate
manifest asserts the list against the selector over every tracked path both
ways (scripts/gate/routing.ts, ROUTED_LANES). push to main is unfiltered.
Path coverage exempts declared manual-only checks the way owned does.

* ci: tighten routing assertion shape (fallow: unused exports, complexity)

* ci: name parked checks in check:affected --run skips

* ci: bound the routed-lane exemption to sibling workflows (review of #1857)

The exact-name .github exemption was unbounded: naming the lane's own
setup-apple-runner-build or boot-ios-test-simulator action skipped the lane
that runs them and the manifest stayed green. Lane now carries the transitive
composite-action closure plus its own workflow file (Lane.uses, same walk
declaredGates does), and the exemption refuses anything in it.

Also: an unowned path under an ignored root (a non-TS fixture under a family
root) asked for the ignore entry to be removed, which would un-route every
sibling in that tree; it now asks for a selector owner. Both cases pinned,
both proven red against the pre-fix code. Documents GitHub's 300-changed-file
path-filter limit in docs/agents/testing.md.

* ci: close the routed-lane exemption over composite-action support files

Lane.uses recorded only each composite action's action.yml, so a support file
the descriptor executes was exemptible as if it were an unrelated sibling
workflow: ios.yml uses setup-fixture-app, whose action.yml runs
"$GITHUB_ACTION_PATH/fetch-artifact.sh", and that script runs its siblings
resolve-artifact-name.sh and trusted-artifact.mjs — references that exist only
inside shell, one level past anything YAML parsing sees.

The closure unit is the action's directory now. It needs no shell model and
cannot miss a file however deep the reference chain runs; the coarseness is
harmless because a file in an action's own directory belongs to that action.
All three files pinned, red against the descriptor-only closure.
2026-08-19 17:35:23 +02:00
Michał Pierzchała ee13203a16 feat(ios): unify snapshot eligibility (#1850)
Make iOS regular snapshot eligibility one backend-neutral presentation rule.

Acquire tree nodes conservatively, preserve interactive scroll containers, normalize surviving hierarchy, and keep raw membership plus daemon publication policy unchanged. Part of #1797.

- iOS and macOS unit-enabled runner builds
- 2 focused XCTest cases
- 3 production-path publication tests
- live Settings snapshots: 73 regular nodes and 167 raw nodes, both healthy tree captures
2026-08-18 18:57:29 +02:00
Michał Pierzchała f03c0309a1 fix: derive iOS transition snapshots from visible presentation (#1831)
* fix: project iOS transition semantics

* fix: derive iOS transition semantics from visible state

* fix: preserve iOS presentation context for scoped snapshots

* fix: confirm broad iOS transition settlement

* ci: run coordinate input regression on pull requests

* test: mock migrated snapshot capture seam

* fix: confirm transitions across snapshot backends

* fix: arm transition confirmation after first capture

* fix: settle against immutable action baseline
2026-08-18 17:53:23 +02:00
Michał Pierzchała 8300fa131e refactor(ios): establish snapshot presentation seam (#1845)
Introduce RawAXNode and PresentedNode so acquisition backends can no longer construct the wire-facing snapshot shape directly. Preserve current output while #1797 moves semantics behind the seam.

Non-vacuity: setting PresentedNode.label to nil made testSnapshotPresentationPreservesCurrentWireShape execute once and fail on the missing label field; restoring the production mapping made the same focused XCTest pass.
2026-08-18 17:39:26 +02:00
Michał Pierzchała 8b0560a51d fix: hide clamped descendants of offscreen iOS rows (#1811) 2026-08-18 11:00:19 +02:00
Michał Pierzchała f45228ae71 ci: skip device lanes for root-level docs-only changes (#1781 A9) (#1791)
* ci: skip device lanes for root-level docs-only changes (#1781 A9)

Add AGENTS.md, CHANGELOG.md, CONTEXT.md, CONTRIBUTING.md, LICENSE, and
SECURITY.md to the pull_request paths-ignore block in ios.yml,
android.yml, linux.yml, macos.yml, ci.yml, and size.yml. These
root-level docs files were the only gap left after docs/**,
website/**, and README.md — PRs #1568 (SECURITY.md only), #1697
(CONTEXT.md + docs/adr only), and #1722 (AGENTS.md + docs/) each still
triggered a full 9-15 min macOS iOS run despite touching only prose.

Why each file is safe to ignore for every one of these six workflows:

- None of the four device workflows (ios/android/linux/macos) or their
  composite actions read any of these six files at runtime; the only
  hits from `grep -rln` across scripts/, src/, test/, and
  .github/actions/ are prose comments pointing humans at CONTEXT.md or
  AGENTS.md sections (e.g. scripts/layering/check.ts,
  scripts/wire-compat/run.ts, src/mcp/tool-ref-pins.ts) — never an
  `fs.readFileSync`/`readFile` of the file itself.
- The check-affected selector (scripts/check-affected/model.ts)
  already classifies all six as pure docs: `isDocs()` matches any
  `.md` file plus the literal `LICENSE`, and `docsOwnership()` only
  special-cases `website/docs/docs/commands.md` (unrelated). So these
  files already select zero checks — they only ever produced
  `docsOnlyPaths` entries, never `SelectionReason`s.
- Because they select zero checks, the gate-manifest's path-coverage
  category derivation (`scripts/gate/model.ts` `categories()`, which
  iterates `plan.reasons`) never records a category for them, so
  ci.yml has nothing check-manifest-only that these six files would
  need to keep reachable. `pnpm check:gate-manifest` and
  `pnpm check:gate-manifest:test` both stay green after the change
  (48 checks / 33 lanes, 28/28 gate tests passing).
- size.yml's bundle-size job (scripts/size-report.mjs) measures the
  `pnpm build` dist output and startup timing only — no reference to
  any of these six files. (npm packs LICENSE/README.md into the
  publishable tarball, but that's a `pnpm check:package` node-22.12
  concern in ci.yml's packaged-cli job, which is driven by `dist`
  contents and `package.json`, not by LICENSE/README prose — already
  evidenced by README.md being ignored here since before this change.)

Scope disclosure: `mutation-affected.yml` uses a `paths:` allowlist
(not paths-ignore) so it's structurally unaffected; `test-app-build-cache.yml`
has no path filter at all. Neither was touched.

actionlint and `pnpm check:gate-manifest`/`:test` pass on the changed
workflows.

* test: pin root-doc paths-ignore entries with a regression test

Addresses review feedback on #1791 from thymikee: the docs-only
classifier for AGENTS.md/CHANGELOG.md/CONTEXT.md/CONTRIBUTING.md/
LICENSE/SECURITY.md across ios.yml/android.yml/linux.yml/macos.yml/
ci.yml/size.yml had no regression pin. Neither check:gate-manifest
(only proves a *registered check* is reachable) nor actionlint (only
validates YAML shape) nor generic Markdown coverage would catch a
single dropped entry — e.g. LICENSE reappearing in one workflow's
paths-ignore list but not another's would silently put a full 9-15 min
device run back on prose-only PRs.

test/ci/root-docs-paths-ignore.test.ts parses the six real workflow
files and asserts, using the same matchesGlob the gate-manifest model
uses to decide lane triggering, that each of the six root docs is
ignored by each workflow's pull_request paths-ignore. Registered in
vitest.config.ts's unit-core project next to its sibling
upload-agent-device-artifacts.test.ts (parse-only, no device/subprocess
lane needed).

Verified red on main (all 36 file x doc assertions fail — confirmed via
a throwaway script reading `git show main:.github/workflows/*.yml`)
and green on this branch (6/6). Full unit-core project (873 files /
6641 tests) still passes; check:gate-manifest and
check:gate-manifest:test unchanged (48 checks / 33 lanes, 28/28).
2026-08-18 09:57:09 +02:00
Michał Pierzchała 0d3b7413c5 fix: prevent private AX subtree leaks at source (#1807)
* fix: prevent private AX subtree leaks at source

* fix: preserve values in settle signals

* fix: normalize settle signal semantics
2026-08-18 09:23:39 +02:00
Michał Pierzchała 8b698e8efc fix: stabilize private AX settle snapshots (#1784)
* fix: stabilize private AX settle snapshots

* fix: harden private AX settling
2026-08-17 18:58:12 +02:00
Michał Pierzchała 9c22467832 refactor(ci): make gate ownership structural (#1429) (#1753)
* test(ci): prove every registered gate is owned and reachable (#1429)

A check that silently stops running looks exactly like a green build. Two
suites had already stopped: `check:tmpdir-leaks` (with its model tests) and
`test:fixture-cache` are real package scripts that no workflow ran, reachable
only through the `check:unit` aggregate CI never invokes.

`CHECK_CATALOG` becomes the registry of every check and `pnpm gate <id>` the
only way CI runs one, so finding what a lane runs is a scan for `pnpm gate`
rather than an attempt to interpret shell. `pnpm check:gate-manifest` then
asserts against the real workflows that every registered check is run by some
qualifying lane (per unit, not per script name), that every check the real
selector activates for a path is run by a lane that path would start (#1420's
class), and that every Vitest project and suite script belongs to a check.

The wiring that keeps those honest is asserted too: a gate id must name a
registered check, an `if:` must be ruled on in GATE_CONDITIONS so `if: false`
unowns what it guards, an action declared to run a gate is proven to, and a
job whose steps the loader cannot open fails closed.

It deliberately does not try to prove CI runs project code only through
`pnpm gate`. Whether a shell block executes project code is not decidable from
its text, so shell this model does not recognise earns no ownership credit —
the failure direction is a check reported unowned, never one waved through.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* test(ci): update the two suites that assert on rewired workflow text

`scripts/mutation/workflow.test.ts` and `test/ci/trusted-fixture-artifact.test.mjs`
read the workflow and action files and assert on their command text, so routing
those steps through `pnpm gate <id>` moved what they were matching.

They are the two suites the manifest cannot help with: it proves a gate is still
run, not that a test asserting on how CI spells a command was updated with it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* fix(ci): credit gates by execution shape, and keep every guard

Three ways the manifest could report a gate as owned when it does not run.

1. Crediting was a substring scan over `run:`, which #1429 explicitly rules
   out — "do not infer reachability from a command name merely appearing in
   workflow text". `false && pnpm gate x`, a gate inside `if false; then … fi`,
   one named in a heredoc, and `echo pnpm gate x` all credited it. There is a
   live instance: conformance-regenerate.yml's "Fail if regeneration changed
   anything" step names `pnpm gate maestro-regenerate` inside an error message
   telling a human to run it, and that credited the gate.

   A gate now counts only as the first command segment of a line, and a body
   carrying shell structure earns nothing. Reachability inside a script is not
   decidable, so this does not try: unrecognised shape means no credit and the
   check reports unowned. `VAR=$(pnpm gate x …)` is read, since the assignment
   form is unambiguous and the gate runs.

2. Job-level `if:` was not modelled at all, though six live jobs carry one, so
   a job that cannot run still credited every gate inside it. Two conditions on
   the mutation lanes are now declared.

3. A caller's `if:` REPLACED the guard on a nested composite-action step
   (`guard[0] ?? step.condition`), so an outer `always()` erased an inner
   `if: false`. Steps carry every guard between the lane and the step.

Also corrects two source comments that still claimed project code run outside
the runner fails the manifest. It does not: such a step earns no credit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* ci: add the run-gate action that names a gate structurally

The seam the ownership proof will read instead of shell. A lane says which
gate it runs in `with.gate`, a typed input the manifest reads straight out of
the YAML and validates against CHECK_CATALOG.

Nothing here is wired yet — the ~60 call sites and the model change follow.
Added first so the target of that conversion is reviewable on its own.

`args` cannot select which gate runs; it is appended after the id, so the
worst a wrong value does is fail the gate it already named. There is no
`|| true` and no output capture: the gate's exit code is the step's exit code,
so a gate cannot run without being able to fail its lane.

Part of #1429.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* merge: main (#1770) and route its three new steps through the runner

#1770 landed the orphan-check fix on main, wiring `check:tmpdir-leaks`,
`check:tmpdir-leaks:test` and `test:fixture-cache` into Coverage, Layering
Guard and Integration Tests. This branch had wired the same three through
`pnpm gate`, so the merge produced two steps per check rather than a conflict
— each check ran twice.

Kept main's steps, with the placement and reasoning reviewed on #1770, and
changed only their `run:` line to the canonical runner. Dropped this branch's
duplicates. Net effect on CI is unchanged: the same three checks, in the same
three lanes, once each.

Gate manifest green after the merge: 47 checks wired across 33 lanes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* fix(ci): address review — suite detection, freerange, glob, vacuous skip-list

Six review findings plus the mutation blocker.

[bug] `registered` was shape-only, so a `test:*` script running
`node src/bin.ts test <dir>` resolved to a `script:` leaf and was invisible.
Four `test:replay:*` scripts were owned only because someone hand-registered
them; `test:replay:android` was neither registered nor reported while the
nightly ran the same six .ad files by inlining them. A `test:*` script is now
a suite by name. `replay-android` is registered, and the nightly runs the
script instead of re-listing its files so the two cannot drift.

  The nightly invokes it inside `reactivecircus/android-emulator-runner`'s
  `script:` input — shell handed to a third-party action this loader does not
  read — so the suite executes but cannot be credited. Recorded in
  UNPROVABLE_OWNERS with that exact reason rather than assumed.

  The fixed detector also found a second orphan the review did not name:
  `test:integration:progress`. That one is a reporter whose `--check` sibling
  is the registered gate, so it is declared in REPORTING_SCRIPTS — a
  declaration that itself fails when inert.

[bug] `freerange` defaulted to localRunnable, so fail-open ran `fr` (a Bun
binary) on the pre-push path. Now false.

[suggestion] The `--run` skip-list asserted `build:android-snapshot-helper`,
a name `android-helpers` no longer uses, so it could not fail. Derived from
the catalog instead.

[suggestion] `matchesGlob` joined `**` splits with `.*`, making the adjacent
slash mandatory — GitHub's `**` matches zero directories, so
`src/**/*.test.ts` did not match `src/a.test.ts`. Pinned against
`packages/*/src/**/*.test.ts`.

[suggestion] Deleted the unwired `run-gate` action. It had no callers, was
absent from GATE_ACTIONS, and its comment described a system that had not
shipped. It returns with the rewiring, not before.

[suggestion] Collapsed the module headers that narrated discarded designs.

Mutation: `daemon entrypoint publishes HTTP metadata and cleans up on
shutdown` is the only test here that spawns a real daemon process. It takes
~1.1s alone but exceeds Vitest's 5s default inside Stryker's dry run, which
aborts the sweep before a single mutant runs. Given 30s.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* fix(mutation): order sandbox aliases longest-first so subpaths resolve

Every shard of the mutation sweep aborted in Stryker's dry run with:

  Cannot find package '@agent-device/selectors/engine' imported from
    .tmp/stryker/sandbox-*/src/core/selector-pipeline.ts

The alias was generated correctly; it just never won. Vite matches a STRING
alias by prefix and takes the first hit, and `workspaceSpecifierTargets`
emitted the bare `@agent-device/selectors` ahead of the subpath entries. The
bare entry therefore captured `@agent-device/selectors/engine` and rewrote it
to `…/src/index.ts/engine`, which does not exist; Node fell back to real
package resolution, could not find the subpath inside the sandbox, and the dry
run failed before a single mutant ran — so the shard uploaded an empty
envelope instead of a report and the ratchet failed for want of one.

Sorting longest specifier first makes the most specific alias win:

  @agent-device/selectors/engine -> packages/selectors/src/engine.ts
  @agent-device/selectors/ast    -> packages/selectors/src/ast.ts
  @agent-device/selectors        -> packages/selectors/src/index.ts

`/ast` never tripped this because nothing in a related test set imported it;
`selector-pipeline.ts` introduced the first subpath import that mattered
(#1744), so the mutation lane has been unable to run since that landed. Any
PR touching `scripts/mutation/**` — which fails open into the full sweep —
would have hit it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* refactor: derive gate ownership from workflow structure

* fix: run gates without optional arguments

* fix: resolve mutation workspace subpaths exactly

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-12 16:02:18 +02:00
Michał Pierzchała cdc754e6ed perf: speed up iOS agent recovery and streamline CLI guidance (#1700)
* Avoid interactive children in parent taps

* docs: streamline no-skill CLI help

* perf: recover faster from sparse iOS trees

* fix: preserve selector context for blocked parent taps

* fix: preserve coordinate text-entry focus

* fix: preserve thin parent touch targets

* fix: fail closed for unscoped iOS typing

* test: isolate replay lock fixture

* test: share node integration process
2026-08-10 20:43:01 +02:00
Michał Pierzchała 6c0fcb64a1 fix: reject distinct ambiguous mutation targets (#1667)
* fix: reject distinct ambiguous mutation targets

* fix(ios): scope the raw-match rejection to mutating dispatches

`RunnerTests+Interaction.findElement` applied the new fail-closed
classification to `querySelector` as well as press/type, because the read
call site takes the default `allowNonHittableFallback: false`. With one
visible/hittable match and one non-hittable same-selector duplicate the
query started returning AMBIGUOUS_MATCH where it previously selected the
hittable element, and `queryDirectIosSelectorOrFallback` preserves that
error for read callers — so `get`, `is`, and `wait` surfaced an error
instead of their prior answer.

`classifyDirectSelectorCandidates` now takes a `rawMatchPolicy`. Mutations
keep `.rejectDistinctMatches` (the default, so no mutation call site
changes); `queryElement` passes `.preferHittableMatch`, restoring the
prior read rule: prefer the single hittable match, ambiguous only when
hittable matches compete, and never adopt the Maestro coordinate fallback.
The Maestro expected-point path is untouched.

Covers the one-hittable + one-non-hittable read, competing hittable reads,
and the non-hittable-only read. ADR 0011's amendment now states the scope.

* test(ios): execute selector read ambiguity regression

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-08 08:57:32 +02:00
Michał Pierzchała 04c33f9e1e feat(ios): expose AX custom actions on merged accessibility elements (#1665)
* feat(ios): expose AX custom actions on merged accessibility elements

Apps that merge a card into one accessibility element for VoiceOver (React
Native's `accessible` prop) publish the card's real affordances as
UIAccessibilityCustomActions rather than as child elements. Our snapshot showed
only the merged node, so an agent looking for a feed card's options control had
nothing to aim at and fell back to coordinate guessing.

`snapshot --actions` now names them:

    @e8 [link] "feedItem-by-whiskers.test" actions: ["Reply", "Repost", "Open post options menu"]

Opt-in, because the AX server cannot serve custom actions in a bulk tree
request: adding the attribute makes testmanagerd's reply decoder reject the
nested arrays a custom action serializes into, drop the reply, and time the
request out (~65s vs ~110ms). Only per-element reads answer, at ~100ms each, so
the runner reads at most 12 labelled childless nodes and stops at the
capture-plan deadline. The request pins the private-AX backend, since no other
backend can read the attribute, and reports that as its own `requested-backend`
verdict so a deliberate pin never renders as a degradation warning.

Invocation is not shipped: the actions are readable but not invocable from the
runner. RunnerAXSnapshotBridge.h records the five APIs that were tried.

* fix(ios): disclose a capped custom-action pass, and read on-screen elements first

Two gaps in the first cut.

An element the bounded pass never reached rendered identically to one with no
custom actions, so a capped capture silently taught the reader that later feed
cards have no affordances — the exact mis-inference this feature exists to
prevent. The runner already counted reads against candidates; it now carries
both to the response, and the verdict renders one response-level line when the
pass was incomplete. A complete pass stays silent, and "never asked" stays
distinguishable from "read none" (the key is absent, not (0, 0)).

The obvious remedy for a capped pass would be a scoped re-run, but scope is
applied when the Swift walk builds nodes, long after the read pass, so it does
not redirect the budget at all — measured: reads=12 candidates=18 with and
without --scope. Rather than print a remedy that does nothing, the read pass now
orders candidates on-screen first. That makes the budget land on elements an
agent can act on, and makes the disclosed remedy true: scrolling changes the
on-screen set, so a re-run reads elements the previous pass could not.

Also states plainly, in the flag help and the tool/SDK field description, that
the names are for planning: nothing invokes them, so the affordance is reached
through the element's detail screen, the same control exposed elsewhere, or
coordinates.

* test(snapshot): pin the custom-action coverage pair in the verdict shape assertion

* chore(scripts): classify the --actions flag in the integration progress model

The completeness gate flagged snapshotCustomActions as unclassified, which is
what it is for. It gets its own bucket rather than joining the provider-scenario
table: the values come from the private AX client inside the runner process, and
the fake runner derives its behavior from fixture tables that cannot fabricate
custom actions, so there is no provider-backed scenario to claim. The owning
coverage is named instead — runner XCTest unit, snapshot-lines, snapshot-quality.

* fix(ios): fail closed on unserviceable --actions, bound each read, cap output, and count actions in identity

Four review findings.

1. `--actions` with `--raw`, or on any target that is not an iOS simulator, used
to succeed and return nodes with no actions — a requested capability silently
no-opped, indistinguishable from "this screen has none". Both now fail closed.
The raw pairing is rejected at the shared request seam (INVALID_ARGS) so CLI,
Node client and MCP answer alike before any device work; the platform case is
rejected once the session device is resolved (UNSUPPORTED_OPERATION), naming the
resolved target. `diff --actions` was already rejected as an unsupported flag.

2. The per-element AX read had no timeout, so one wedged element could consume
the whole capture budget. Each read now runs off-thread behind a 1s wait. A
timed-out element counts as unread, never as "read, and it has no actions", so
the existing partial-pass disclosure already covers it.

3. The element budget bounded element count only; one element could still return
an unbounded list of unbounded names. Capped at 8 names of 80 characters, and
clipped elements are counted into the coverage so a truncated list is disclosed
rather than silently presented as complete.

4. Action names were rendered unescaped, and no comparison key read them. Names
now get the same escaping as text previews plus control-character folding, so an
app-authored name cannot split or corrupt a line. `actions` joins the diff
comparable key, the unchanged-comparison projection, and — the sharper bug —
the snapshot presentation key, without which `snapshot` followed by
`snapshot --actions` on a still screen answered "unchanged" and never delivered
the actions that were explicitly requested.

* fix(ios): contain a hung custom-action read instead of accumulating orphans

The 1s read deadline frees the caller, but the underlying AX call is a
synchronous XPC round trip that cannot be cancelled — it keeps running. On a
global concurrent queue that meant repeated `snapshot --actions` against a
wedged element piled up orphaned reads, all using the shared XCAXClient
concurrently. The deadline was containment for the capture, not for the runner.

Since the call cannot be cancelled, contain it instead:

- every read runs on one dedicated serial queue, so a wedged call can never be
  joined by a second concurrent user of the shared client;
- a single-flight guard refuses to dispatch at all while an abandoned read is
  still outstanding, so a repeated capture adds no work — the dispatch counter
  stands still;
- the read pass stops at that point rather than paying a deadline per element
  on reads that would all be refused, and reports `blocked` so the capture stays
  honest. That gets its own line, because the partial-pass remedy (scroll and
  re-run) cannot clear a hang and would send the reader in circles.

Recovery needs no reset: when the hung call finally returns, in-flight drops to
zero and reads resume.

The regression drives a fake AX client that never returns, and asserts the three
things the fix exists for — exactly one in-flight read with no further
dispatches across repeated captures, immediate returns with the skip disclosed
instead of the scroll remedy, and reads working again once the wedge clears.

* ci(ios): execute the custom-action runner regressions instead of only compiling them

The iOS workflow runs a targeted -only-testing list, so a runner test that is
not named there is compiled by the build step and then never executed. All seven
custom-action tests were in that gap — including the containment regression,
which is the only executable proof that a hung AX read cannot accumulate
orphaned in-flight reads.

Red/green against the containment regression, with the fix reverted to its
pre-fix concurrency behavior (global concurrent queue, no single-flight guard,
no blocked exit):

  RED   in-flight 6 (want 1), dispatches 6 (want 1), each repeat paid the full
        1.004s deadline (want <0.2s), blocked=false (want true), and the
        in-flight drain never completed — "Exceeded timeout of 5 seconds".
  GREEN 7/7 pass, containment regression in 1.02s.
2026-08-07 17:35:48 +02:00
Michał Pierzchała 5f90d908e2 fix(ios): wait for hidden-keyboard synthesized text to commit before responding (#1676)
* fix(test): wait for typed text to settle in the hidden-keyboard runner test

testBareTypeUsesTappedInputWhenSoftwareKeyboardIsHidden read textField.value
in one shot right after executeTypeCommand returned. The simulator commits
synthesized keystrokes after the command responds, so on a loaded CI machine
the read landed mid-word — observed failures reported ("h") and
("hardware-ke"). It failed on 3 of 5 runs of a branch carrying zero Swift
changes and passed on re-run.

Poll the value until it holds the expected text (10s ceiling) and assert on
the last value read, so a real regression still fails with what the field
actually held. Both reads in the test use the same helper; the assertions are
unchanged.

* Revert "fix(test): wait for typed text to settle in the hidden-keyboard runner test"

This reverts commit caecdc6d5b.

* fix(ios): wait for hidden-keyboard synthesized text to commit before responding

The synthesized-first-responder bare-type route returned ok as soon as the
private XCTest event record was posted, while the target app was still
committing characters. On slow CI simulators the trailing characters landed
after the response, so agents (and the smoke test) observed a truncated
field value through the public type path.

After dispatch, poll the tapped element until its value reaches
textBefore + typedText, exit immediately when the app transforms the input
(formatter, mid-text caret, autocomplete), and if progress stalls as a
strict prefix, re-synthesize the missing tail once. Submit-suffixed text
keeps the old immediate return so a repair can never double-submit.

Validated on a booted iPhone 17 Pro simulator: 5/5 passes of
testBareTypeUsesTappedInputWhenSoftwareKeyboardIsHidden under full-core CPU
load, with the commit wait absorbing up to ~390ms of post-dispatch lag that
the previous code ignored (uniform ~494ms dispatch-only before); the tail
repair never had to fire, consistent with commit lag rather than true drops.

* fix(ios): make the synthesized commit wait observation-only

Addresses the P1 on #1673: a 600 ms quiet prefix cannot distinguish delayed
app-side commit from a genuinely dropped suffix, so re-synthesizing the tail
could post it while the original was still queued and commit the text twice
after the command had already reported ok.

Drop the repair path entirely — the tail builder, the quiet-window constant,
the stall tracking, and the synthesizer dependency the wait only needed in
order to repair. What remains is a bounded observation: poll the tapped
element until the value commits, the app transforms it, the value becomes
unreadable, or the 3s ceiling expires. A dropped suffix still reports ok,
exactly as before this change; only the false truncation from commit lag is
removed.

The submit-key skip stays, for its own reason: the app may clear or rewrite
the field on submit, so textBefore + typedText is not the value to wait for.

Also wire testSynthesizedTextCommitProgressWalksExpectedPrefixOnly into the
iOS smoke lane — that workflow enumerates its tests by hand with
-only-testing, so a new test that is not listed never runs.
2026-08-07 17:34:09 +02:00
Michał Pierzchała ac52281448 fix(test): deterministic temp-dir cleanup across node --test lanes (#1661)
* fix(test): deterministic temp-dir cleanup across node --test lanes

node --test has no global setup/teardown hook, so unlike Vitest (#1593) every
node --test package.json script (maestro:conformance, mutation:test,
check:affected:test, check:coverage-changed:test, check:layering,
depgraph:test, check:tmpdir-leaks:test, check:contention-retry,
test:fixture-cache, test:smoke(:web), test:integration:node,
test:concurrency-torture) still created scratch directories against the
real, unredirected os.tmpdir(), with cleanup only as reliable as each call
site's own try/finally — which a crash, OOM, or timeout kill bypasses
entirely.

Add scripts/node-test-tmpdir.ts: it wraps the whole `node --test`
invocation as a child process, redirecting TMPDIR to one disposable,
pid-tagged directory (shared root/prefix with the Vitest lane) and removing
it from the process 'exit' event, which fires on normal completion, a
thrown error, or a forwarded SIGINT/SIGTERM alike. Every node --test script
now runs through it. check-tmpdir-leaks.ts already scans by root/prefix, so
it covers both mechanisms with no changes to its detection logic.

Verified: a node --test process that mkdtemp's then gets SIGKILL'd leaves a
directory behind unwrapped; wrapped and SIGTERM'd, TMPDIR is redirected and
the directory is gone with no orphaned processes. All 13 wrapped lanes and
the full Vitest suite (5,591 tests) pass with zero residual
agent-device-test-run-* directories after the run.

Fixes #1595

* test(tmpdir): ratchet every node --test script through the wrapper

The 13 lanes wrapped in package.json were a one-time hand sweep with
nothing enforcing the pattern going forward — a 14th node --test script
added later without scripts/node-test-tmpdir.ts would silently reopen
#1595 for that one lane.

Add a structural check to scripts/node-test-tmpdir.test.ts (now part of
check:tmpdir-leaks:test) that reads package.json and fails if any script
invokes `node ... --test` without routing through the wrapper. Dumb
string matching over the scripts map, no shell parsing, with an explicit
(currently empty) NODE_TEST_WRAPPER_BYPASS_ALLOWLIST for any lane that
must legitimately bypass it. Verified it both passes on the current
package.json and fails when a synthetic unwrapped `node --test` script is
added.

* fix(test): preserve the Swift cache and close the raw node --test bypasses

Review on #1661 found two gaps:

1. The wrapper only overrode TMPDIR, so it discarded and forced a
   recompile of the durable Swift compiler cache every run instead of
   mirroring vitest-tmpdir-global-setup.ts's carve-out for it. Read
   os.tmpdir() before the child's TMPDIR redirect takes effect and set
   AGENT_DEVICE_SWIFT_CACHE_DIR from that (only when unset), same as the
   Vitest lane — the two now share one durable cache instead of each
   discarding and recompiling their own. Added a probe assertion
   (scripts/node-test-tmpdir.test.ts) that fails without the fix and
   passes with it (verified both ways).

2. docs/agents/testing.md documented raw `node --test` commands for the
   iOS smoke files, and the android/ios/conformance-regenerate/nightly
   workflows invoked `node --test` directly outside package.json. Routed
   all of them through scripts/node-test-tmpdir.ts so the documented
   local commands and CI lanes get the same crash/timeout-safe cleanup
   the package.json scripts already have.
2026-08-07 13:26:20 +02:00
Michał Pierzchała a67c72c211 fix(ios): pin tap-outcome corroboration probes to the baseline's backend (#1634)
* fix(ios): pin tap-outcome corroboration probes to the baseline's backend

The recorded-failure screens are exactly where the capture plan flips
between XCTest and private-AX (the penalty boundary), so #1605's
same-backend requirement failed closed right where XCTest tap false
negatives actually happen: the baseline was captured via private-AX
under penalty, the probe came back via tree, and a landed tap surfaced
as XCTEST_RECORDED_FAILURE. In the AppControlBench bsky-16 run this
fired four times, each sending the model into a re-observe/retry spiral.

The comparison stays same-backend by design (backends are not comparable
views of a screen); instead the probe is now CAPTURED the way its
baseline was: a new internal preferredBackend option (never CLI-exposed)
threads daemon -> runner, and a private-AX-preferred capture takes the
exact penalized route — privateAX-first plan, 'deferred' verdict, no
degradation warning, no settle budget reset.

Live-verified on the deterministic repro (Bluesky drawer-menu press
under penalty, seeded bench feed): errored with the backend-mismatch
diagnostic before, corroborates as landed after, with no mismatch phase
in the request diagnostics. Daemon tests cover pinned and unpinned
baselines end to end through the dispatch context; the Swift plan gate
is a pure function with an executed in-bundle test (added to the ios.yml
regression list).

* style: oxfmt

* fix: exclude raw baselines from corroboration and prove the pin end to end (review)

Raw baselines could not be pinned: the raw diagnostic plan keeps
tree-first error propagation by contract and is never rerouted by the
penalty or the preferred backend, so preserving 'raw: true' on the probe
recreated exactly the backend-mismatch false failure this PR removes.
Corroboration now declines raw baselines up front (they are diagnostics,
not evidence baselines) with a regression pinning that no probe capture
is dispatched at all.

The wire is now regression-proven at every hop: a dispatch-level test
drives dispatchCommand with the context flag and asserts the emitted
RunnerCommand carries preferredBackend (red if handleSnapshotCommand or
the interactor stops forwarding); the injected-transport test asserts
the interactor's snapshot payload both ways; and a runner unit test
decodes the wire JSON, projects it through the extracted
snapshotOptions(from:), and composes it with the plan rule — pinned
regular plan defers to privateAX-first, RAW plan stays untouched.
Executed on-simulator; added to the ios.yml regression list.
2026-08-06 13:27:30 +02:00
Michał Pierzchała 4a3ed1e8c3 feat(ios): extend depth-capped private-AX captures via element-rooted requests (#1627)
* feat(ios): extend depth-capped private-AX captures via element-rooted requests

The AX server's depth limit is per-request (kAXErrorIllegalArgument above
a tree-size-dependent value), so a capture capped at depth 56 on
Bluesky-class React Native trees returned chains of unlabeled [other]
containers and hid every actionable control below the cap — agents fell
back to screenshot-and-coordinate guessing.

After a capped serialization, the bridge now re-issues the same snapshot
request rooted at each deepest-level childless node's live accessibility
element, splicing the returned children in and chaining further while
capped subtrees remain — bounded by a call limit, the shared node
budget, and the capture-plan deadline. The AX server counts node levels,
not edges: a maxDepth=56 request emits nodes to depth 55, so the
frontier is the deepest observed level, kept only when it sits at the
cap; a tree that ends naturally above the cap yields no frontiers and
costs nothing.

A capture whose frontier extension drained every capped node no longer
reports itself depth-limited — the re-run hint it used to trigger could
not add anything. Explicit --depth requests stay exact captures with no
extension.

Live on the seeded Bluesky bench feed: 8 extension calls (~106ms each)
turn the 117-node all-[other] tree into a 182-node tree carrying post
text, testID links, and every feed control (Reply/Repost/Like/options);
steady-state capture 0.52s -> 1.37s. Observation-only captures paying
the extension needlessly is #1626.

* fix: count missed frontiers and tighten deep-extension shape (review)

The completeness verdict inverted in the failure paths: a frontier whose
live element vanished or whose re-rooted request failed was silently
dropped, so an all-miss extension reported pendingFrontiers=0 and the
capture presented as complete while whole subtrees were missing. Missed
frontiers are now counted, logged, and keep the depth-limited verdict —
the pure decision lives in privateAXDepthLimited with in-bundle tests.

The truncation hint no longer advertises --depth (an explicit --depth
capture disables extension, so following it returned strictly less than
the capture that produced the hint); the honest remedy is a plain
re-run with a fresh extension budget.

Shape: candidates collected only at the cap boundary and not at all
when extension is disabled (exact --depth captures pay zero
bookkeeping); the zero-caller 3-arg overload is gone; the response
shape's keys are shared constants; the unsupported-selector diagnostic
is restored; the exact-depth policy is hoisted to one named local.

* test: execute the deep-extension miss-path contract in CI (review)

The depth-limited regression compiled but never ran (absent from
ios.yml's -only-testing list), and the pure Swift consumer test was
vacuous against the Objective-C producer — a literal missedFrontiers
proves nothing about the increments. RunnerAXSnapshotFrontier and
extendSnapshotFrontiers move to the header as the executed contract
seam, and testDeepExtensionCountsMissedFrontiers drives both real miss
paths with fabricated snapshots: a nil accessibilityElement (explicit
nil property — bare NSObject resolves the key through a UIKit category
and takes the call path instead) counts missed without consuming a
call; a resolving element whose client cannot serve the re-rooted
request consumes a call AND counts missed. Both tests join the
executed ios.yml list.

Red-before verified on-simulator: with both increments stripped the
producer test fails at the counter assertion (0 != 2); restored, both
tests pass.
2026-08-06 12:09:42 +02:00
Michał Pierzchała 4c7a899a05 fix: prevent iOS text entry runner wedge (#1604)
* fix: prevent iOS text entry runner wedge

* fix: preserve iOS hardware-keyboard text entry

* fix: expire iOS text-entry tap witnesses

* fix: fail interrupted iOS text entry

* test: run interrupted iOS typing regression
2026-08-05 10:24:04 +02:00
Michał Pierzchała c18636315a fix(ios): keyboard-dismiss content settle race (#1542) — partial, defect 2 needs a decision (#1559)
* fix(ios): wait for post-dismiss content settle before the next gesture (#1542)

Dismissing the keyboard can trigger the app's own ScrollView content-offset
correction (e.g. releasing the inset it grew to keep a focused field above
the keyboard). That correction is a separate, unsynchronized animation that
`keyboard.waitForNonExistence` knows nothing about — the keyboard AX element
can disappear well before the app visually settles. The very next command is
frequently a synthesized, AX-free drag (scroll/gesture, kept AX-free so it
still works under #1105-family AX degradation), which has no XCTest
quiescence wait of its own, so it can land mid-animation and net to zero —
the "scroll does nothing" symptom on the Form screen's checkout-form.ad leg.

Add a bounded, AX-free screenshot-stability wait to dismissKeyboard() so the
runner only returns once the screen has actually stopped changing (or a
generous cap elapses). The stopping decision is a pure function
(runnerScreenshotStabilitySettled) covered by unit tests under
AGENT_DEVICE_RUNNER_UNIT_TESTS; the surrounding capture/sleep loop is the
thin, untestable I/O shell around it.

Live-verified on iPhone 17 Pro / iOS 26.2: the scroll now visually lands at
the correct position (confirmed via screen-recording frame correlation)
instead of leaving content at its pre-scroll offset.

Not a full fix for #1542: the checkout-form.ad corpus leg still fails at the
same step, now because the daemon's shared post-gesture snapshot
stabilization (src/daemon/post-gesture-stabilization.ts) can read a
stale-but-internally-consistent AX tree after the AX-free scroll and
mistake "unchanged across polls" for "settled", so the following click's
off-screen guard sees pre-scroll node positions. That is a cross-platform,
cross-command stabilization semantics change and needs a design decision,
not a unilateral fix here — see the PR description.

* ci(ios): execute the screenshot-stability runner tests (#1559 review)
2026-08-03 09:11:05 +02:00
Michał Pierzchała cd9a7ce41b test(android): add comprehensive emulator E2E coverage (#1482)
* test(android): add catalog emulator smoke coverage

* test(android): use stable snapshot diff mutation

* test(android): assert actual back destination

* test(android): separate keyboard and fill IMEs

* fix(ci): keep Android timing report in one shell

* refactor(test): simplify simulator e2e coverage

* test(android): assert stable diff landmarks

* fix(android): release snapshot helper gracefully

* fix(android): fully release snapshot helper runtime

* test(android): report coverage classifications

* fix(android): stabilize accessibility root capture

* fix(android): bound UiAutomation connection

* ci: upload worktree daemon diagnostics

* fix(android): cancel stalled wait captures

* fix(android): bound helper fallback lifecycle

* fix(android): harden emulator e2e lifecycle

* fix: align e2e changes with kernel package

* test(android): prove alert helper reuse directly

* fix(android): cancel stalled settle captures

* fix(android): separate helper retirement budgets
2026-07-30 15:10:21 +02:00
Michał Pierzchała 4e4ecdea0d test(ios): expand simulator e2e coverage (#1408)
* test(ios): expand simulator e2e coverage

* test(ios): make coverage checks host portable

* test(ios): handle deep link confirmation

* test(ios): fix deep link prompt selector

* ci: stabilize full simulator nightly

* test(ios): stabilize permission prompt lifecycle

* test(e2e): wait for route-specific landmarks

* test(e2e): reset permissions from inactive app

* test(e2e): redeliver trusted cold deep links

* test(ios): verify orientation native readback

* test(ios): stabilize simulator permission coverage

* test(ios): wait for tab target after deep link

* test(ios): paginate full event timeline

* test(ios): simplify event pagination coverage

* test(ios): stabilize simulator e2e coverage

* test(ios): model simulator recorder lifetime

* ci(test-app): cache fixture dependencies

* fix(ci): isolate test app cache by node

* fix(ios): settle fixture route navigation

* test(ci): waive unbenchmarked ios system UI help

* fix(ios): tolerate delayed simulator scale lookup

* 0.20.1

* test(ci): remove superseded system UI waiver

* fix(ios): harden simulator e2e reliability

* chore: clarify Apple runner CI steps

* fix(ios): wait before fixture home snapshot

* fix(ios): require exact catalog navigation

* chore(ci): format rebased workflows

* fix(ios): retry unobserved fixture navigation

* refactor(test): remove iOS e2e workarounds

* fix(ci): verify fixture artifact provenance

* fix(ci): align fixture artifact fingerprints

* fix(ci): use unified Android helper packager

* perf(ci): scope fixture build concurrency

* chore: format fixture artifact tests

* test(ios): update split Apple coverage owner

* fix(ios): accept deep-link confirmation alerts
2026-07-28 20:08:45 +02:00
Michał Pierzchała 37895caf99 refactor: replace Maestro compat with typed direct engine (#1217)
* test: add pinned Maestro conformance harness

* feat: add typed Maestro program IR parser

* docs: define direct Maestro engine architecture

* test: compare Maestro oracle with typed IR

* feat: add direct Maestro program engine

* refactor: narrow Maestro execution context

* refactor: tighten Maestro program parsing

* fix: verify iOS Maestro visibility waits

* refactor: isolate retained Maestro runtimes

* refactor: type Maestro target resolution

* refactor: harden typed Maestro execution

* refactor: share in-page swipe planning

* feat: add typed Maestro runtime port

* refactor: parse Maestro suite metadata from typed IR

* refactor: centralize Maestro include loading

* feat: execute Maestro files through typed engine

* refactor: share replay built-in variables

* fix: make Maestro target intent explicit

* fix: refresh Maestro targets before input

* refactor: format Maestro progress from typed IR

* feat: compile typed Maestro replay plans

* feat: bind typed Maestro runtime to public commands

* feat: route Maestro YAML through typed runtime

* refactor: remove legacy Maestro runtime

* refactor: remove obsolete replay control model

* refactor: split typed Maestro plan modules

* fix: harden typed Maestro runtime semantics

* docs: update direct Maestro architecture

* fix: reconcile Maestro runtime with merged contracts

* fix: harden typed Maestro execution boundaries

* fix: harden typed Maestro runtime evidence

* perf: avoid eager Maestro device resolution

* refactor: finalize typed Maestro execution

* fix: reject Android system-only helper snapshots

* fix: preserve Android system dialog snapshots

* fix: make helper-backed CI deterministic

* refactor: invalidate Maestro observations before dispatch

* fix: make Maestro selector policy explicit

* refactor: remove Maestro ranking sentinels

* refactor: make Maestro own observation stabilization

* refactor: source Maestro compatibility presets

* refactor: keep Maestro failure reports typed

* refactor: simplify Maestro runtime policy

* fix: isolate Maestro engine failures

* refactor: consolidate Maestro swipe presets

* fix: align Maestro selector and observation semantics

* fix: preserve atomic iOS Maestro taps

* fix: require semantic uniqueness for Maestro taps

* fix: preserve Maestro parse provenance

* docs: pin Maestro compatibility presets

* docs: reconcile Maestro gesture viewport contract

* perf: resolve Maestro gesture viewport directly

* test: align Maestro replay regressions

* fix: order Android gesture lift after endpoint

* fix: settle Maestro gestures before continuation

* fixup! fix: order Android gesture lift after endpoint

* refactor: normalize Maestro swipes once

* refactor: fail impossible Maestro observations

* refactor: normalize Maestro defaults alias

* test: reconcile Android provider scenarios

* fix(android): synchronize single-pointer move events

* test: align repair digest parsing

* refactor: type Maestro runtime operations

* refactor: keep Maestro controls compact

* refactor: name Maestro diagnostic limit

* fix: align Maestro parser and settling semantics

* fix: complete Maestro compatibility semantics

* docs: define Maestro compatibility boundaries

* fix: refresh iOS runner target after relaunch

* fix: reset prewarmed iOS runner after URL open

* fix: preserve iOS Maestro target and swipe intent

* fix: harden direct Maestro runtime semantics

* fix: preserve ranked Maestro replay suggestions

* fix: align maestro tap runtime semantics

* fix: stabilize maestro ci contracts

* fix: tighten maestro runtime architecture

* fix: reconcile maestro replay with latest main

* perf: tighten Maestro iOS stabilization

* fix: preserve Maestro app lifecycle sessions

* fix: restore Maestro CI coverage

* fix: address Maestro engine review findings

* refactor: consolidate Maestro compatibility internals

* fix: scope Maestro target evidence to childOf
2026-07-15 21:26:42 +02:00
Kenichi Saito 236016ed8a fix(ios): support remote-hosted alerts on physical devices (#1232)
* fix(ios): probe remote-hosted system modals (AccessorySetupKit picker) when the springboard mirror yields no hittable actions

* fix(ios): fail closed on host state, guard dismissal re-query, unit-test probe routing

Addresses review on #1232:
- Gate the remote-host probe to a foreground host
  (RemoteHostedSystemModalPolicy.isEligibleHostState); background/unknown hosts
  fail closed instead of substituting an unrelated action tree.
- Wrap the alert-resolution fallback query in safeElementsQuery so a dismissed
  remote host raising kAXErrorServerNotFound is absorbed.
- Extract routing/gating into RemoteHostedSystemModalPolicy and add
  simulator-free unit tests under AGENT_DEVICE_RUNNER_UNIT_TESTS.

* refactor(ios): centralize blocking system modal resolution

* fix(ios): bound alert dismissal rechecks

* feat(ios): enable alerts on physical devices

* fix(ios): bound alert system modal resolution

* test(ios): add AccessorySetupKit picker fixture

* fix(ios): validate remote-hosted system modal interactions

* chore: keep pnpm checks non-interactive

* fix(ios): share alert command deadline

---------

Co-authored-by: Michał Pierzchała <thymikee@gmail.com>
2026-07-14 17:19:17 +02:00
Michał Pierzchała 139153ce64 fix: bound the iOS system-modal snapshot probe to the capture deadline (#1244) (#1248)
* fix: bound the iOS system-modal snapshot probe to the capture deadline (#1244)

The pre-plan SpringBoard system-modal probe (`blockingSystemAlertSnapshot`)
ran before `runSnapshotCapturePlan`, outside the 20s plan budget, the bounded
main-thread watchdog, and the XCTest recovery envelope. A slow
`springboard.alerts`/`sheets`/descendant enumeration (seen on
ASWebAuthenticationSession consent and notification-permission dialogs) could
therefore stall `snapshot -i` for 30-39s.

Run the probe as a bounded capture tier instead: it shares the snapshot plan
deadline and executes on the same `runMainThreadWork` watchdog as the tree and
query backends (`systemModalProbeBudget`, clamped by the remaining deadline).
On timeout the abandoned probe is tracked like an abandoned tree capture, so the
plan recovers through the independent (private-AX on simulator) backend, later
commands fail fast as busy instead of queueing behind it, and the runner is
never wedged. Diagnostics identify the probe and its elapsed time.

The shared abandonment bookkeeping used by both the tree capture and the modal
probe is extracted into retain/release helpers.

* fix: don't re-enter main for bookkeeping behind an abandoned snapshot probe (#1244)

After the bounded system-modal probe times out, runMainThreadWork abandons its
query but that query keeps grinding on the main thread. The capture plan recovers
through the independent private-AX backend and returns — but executeSnapshotDispatched
then ran its didRecordXCTestFailure/retry bookkeeping through another runMainThreadWork
hop, which queues behind the same abandoned query and re-stalls the command for up to
the 30s execution watchdog (or throws), reintroducing the very stall the recovery
avoided. Skip that bookkeeping while abandoned XCTest work is outstanding — the policy
setNeedsPostSnapshotInteractionDelay already uses — so the recovered response returns
immediately and a subsequent command still reports RUNNER_BUSY until the work drains.

Adds an in-bundle regression that fails if the guard is removed. The snapshot recovery
loop is factored into executeDispatchedWithRecovery so the guard is exercisable without
a live capture.

The alert command's SpringBoard detection is intentionally left on its existing 30s
command watchdog here; giving alert its own physical-iOS capability is tracked as a
focused follow-up (#1231).

* test: force the bounded system-modal probe timeout through snapshotFast

Adds a minimal probeWork seam to boundedBlockingSystemAlertSnapshot (threaded
through snapshotFast/snapshotRaw, defaulting to today's exact production
closure) so a test can substitute a blocking probe body while still running
the real runMainThreadWork wrap and the real onAbandoned/onDrained hooks.

The new regression test drives snapshotFast (the real entry point, not
boundedBlockingSystemAlertSnapshot directly) with an injected probeWork that
blocks past the probe's slice, forcing a genuine timeout, and asserts in
order: busy/penalty accounting once the timeout fires, a recovered payload
returned while the probe is still abandoned (before drain), and release
(hasAbandonedTreeCapture() false / idle) once the probe drains. Verified
revert-sensitive: bypassing the runMainThreadWork wrap, or dropping the
onAbandoned/onDrained hooks, each turn this test red.

* ci(ios): run the #1244 system-modal probe regressions so a wrapper/hook revert is caught

* fix: gate the #1244 probe-timeout test seam behind the unit-test flag

Addresses PR #1248 re-review blockers:

1. `probeWork` was a test-only DI parameter on the production signatures of
   `snapshotFast`, `snapshotRaw`, and `boundedBlockingSystemAlertSnapshot`.
   Gate it behind `#if AGENT_DEVICE_RUNNER_UNIT_TESTS` (the same active
   compilation condition the build script already sets via
   AGENT_DEVICE_XCUITEST_INCLUDE_UNIT_TESTS, and the same one this file
   already uses to gate its unit-test methods): production callers now
   compile the original, unparameterized signatures byte-identical to
   pre-`probeWork`. Both overloads of `boundedBlockingSystemAlertSnapshot`
   delegate to a new private `boundedBlockingSystemAlertSnapshotBody`, which
   is the only place the real `runMainThreadWork` wrap and the real
   `onAbandoned`/`onDrained` hooks are defined, so a revert there fails
   through both the production and the test-seam overload.

2. The regression test previously asserted `.idle`/`hasAbandonedTreeCapture()
   == false` right after signaling the probe's release semaphore, racing the
   drain instead of synchronizing on it. It now polls
   `hasAbandonedTreeCapture()` (bounded) on the same background queue after
   signaling release, fulfilling a dedicated `drained` expectation that the
   test `wait(for:timeout:)`s on before the release assertions, so a slow or
   missing drain fails deterministically instead of racing.

3. The regression only drove `snapshotFast`, leaving a `snapshotRaw`-only
   wrapper regression uncaught. The test body is now a private helper
   parameterized over the entry point, called once for `snapshotFast`
   (existing test, unchanged name) and once for `snapshotRaw` (new
   `...ForSnapshotRaw` test), keeping ios.yml's `-only-testing` list in sync.

Verified revert-sensitive for both entry points: temporarily bypassing the
bounding wrap in `snapshotFast` or `snapshotRaw` turns each entry point's own
test red (and only that one); dropping the onAbandoned/onDrained hooks in
the shared body turns both tests red. Restoring the code turns all green
again.

* fix: collapse snapshotFast/snapshotRaw to one production entry point

Reviewer blocker on #1248: the previous #if/#else split compiled a unit-test
overload with a probeWork parameter that the regression tests called, while
shipping builds compiled a separate #else implementation that no test ever
exercised. Reverting the shipped snapshotFast/snapshotRaw/
boundedBlockingSystemAlertSnapshot to bypass the bounded probe would have left
the tests green.

Collapse each command to a single, always-compiled production implementation
(no probeWork parameter anywhere), and move the only injectable seam to a
tiny systemModalProbeOverrideForTesting property (stored on RunnerTests since
extensions can't hold stored properties) consulted from inside
boundedBlockingSystemAlertSnapshot's probe closure. Tests now call the real
snapshotFast/snapshotRaw entry points and set the override instead of passing
probeWork.
2026-07-14 09:29:30 +02:00
Michał Pierzchała f474f0784e feat: unify gesture planning and multi-touch execution (#1212)
* feat: unify gesture planning and multi-touch execution

* fix: correct unified gesture helper behavior

* refactor: tighten unified gesture architecture

* fix: preserve gesture routing contracts

* test: account for fresh gesture viewport

* refactor: remove retired gesture series

* fix: preserve example app navigation targets

* test: reconcile unified gestures with helper ownership

* docs: update Android helper gesture protocol

* fix: refresh Maestro percentage swipe frames

* refactor: remove stale Maestro frame cache

* fix: harden unified gesture execution

* fix: model gesture viewport in providers

* refactor: remove legacy gesture paths

* fix: remove unused swipe preset parser

* refactor: tighten unified gesture boundaries

* fix: close gesture review gaps

* fix: preserve gesture compatibility contracts

* fix: preserve multi-touch recording semantics

* fix: refresh Apple runner state after app relaunch

* test: lock Apple fling fallback route

* fix: close Apple runner review gaps

* refactor: tighten unified gesture seams

* refactor: consolidate gesture planning policy

* fix: preserve swipe response compatibility

* fix: keep gesture lab aligned with replay coordinates
2026-07-13 13:16:38 +02:00
Michał Pierzchała 26ac865c63 refactor: consolidate Apple platform internals (#968) 2026-06-30 21:30:46 +02:00
Michał Pierzchała edc8dd059b ci: automate iOS runner request-count gate for the Apple runner unwind (Phase 3 step c prep) (#966)
Replaces the manual "run with --debug, hand-count the runner phases" check with
an automated, committed assertion so the Phase 3 step (c) runner relocation (and
future runner refactors) can prove byte-identical runner request behavior.

- src/daemon/runner-request-count.ts: pure, unit-testable counter. Parses the
  daemon --debug diagnostics ndjson and counts the iOS-runner round-trip phases,
  plus baseline parse/compare logic. Owns RUNNER_ROUND_TRIP_PHASES as the single
  source of truth, now imported by request-router.ts (was a local const) so the
  in-process cost graft and the external counter never drift.
- src/daemon/__tests__/runner-request-count.test.ts: 13 unit tests over synthetic
  ndjson fixtures (tolerant parse, counting, baseline parse/compare). Run in the
  normal unit suite; no hardware.
- scripts/runner-request-count/: assertion harness (run.ts) + committed baseline
  (expected-counts.json). Drives the existing smoke-ios replay scenario with
  --debug in an isolated --state-dir, counts runner round-trips from daemon.log,
  and asserts against the baseline. --update regenerates the baseline. Infra
  hiccups are inconclusive (don't fail); only a real count drift fails.
- .github/workflows/ios.yml: new "Assert iOS runner request count" step in the
  smoke-ios job, reusing the booted simulator.
- package.json: `validate:runner-count` script. .fallowrc.json: harness entry.

The baseline ships unarmed (established=false); the harness records observed
counts (printed + uploaded as a test/artifacts artifact) without failing, so the
maintainer arms it once from a real CI run.
2026-06-30 17:26:01 +02:00
Michał Pierzchała bf8d952e21 fix: speed up web snapshots (#842)
* fix: speed up web snapshots

* ci: stabilize iOS simulator smoke boot
2026-06-23 14:13:30 +02:00
Michał Pierzchała d29b86d7b8 fix: report maestro ios runner setup failures (#809) 2026-06-16 16:28:28 +01:00
Michał Pierzchała 0f7187f543 fix: scope source daemon state by worktree (#719)
* fix: scope source daemon state by worktree

* docs: clarify worktree daemon state tradeoffs

* ci: harden Apple runner cache

* chore: keep daemon state helper internal

* ci: validate Apple runner cache restores

* ci: simplify Apple runner cache setup
2026-06-10 10:32:58 +02:00
Michał Pierzchała c2b29d5600 fix: stabilize Maestro replay on iOS (#713)
* fix: stabilize Maestro replay on iOS

* fix: scope iOS runner cleanup to daemon owner

* fix: lease iOS runner ownership per device

* fix: release prepared iOS runner daemon in CI

* fix: inline runner lease release cleanup
2026-06-09 14:12:57 +02:00
Michał Pierzchała 2014cb6873 fix: harden covered snapshot targets (#708)
* fix: block covered snapshot targets

* fix: harden covered snapshot targets
2026-06-08 14:08:52 +02:00
Michał Pierzchała 5c083eacc6 fix: harden iOS replay runner prewarm (#705)
* fix: harden iOS replay runner prewarm

* fix: avoid stale iOS runner during relaunch

* fix: stop stale iOS runner processes

* fix: clean stale iOS runners before startup
2026-06-07 20:37:23 +02:00
Michał Pierzchała ad7b386444 feat: cache iOS runner artifacts during prepare (#688)
* feat: cache ios runner artifacts during prepare

* refactor: deepen ios runner lifecycle

* refactor: generalize apple runner prepare

* refactor: simplify apple runner lifecycle

* fix: clarify runner prepare recovery diagnostics

* refactor: trim apple runner provider surface

* test: simplify runner recovery diagnostics assertion
2026-06-04 15:37:05 -07:00
Michał Pierzchała 7400701857 perf: speed up iOS swipes and harden runner cache (#676)
* perf: speed up ios swipes and harden runner cache

* fix: harden maestro replay smoke tests
2026-06-03 04:59:13 -07:00
Michał Pierzchała 233df070f1 ci: skip platform smoke for docs-only PRs (#604) 2026-05-28 17:06:51 +02:00
Michał Pierzchała 840bef56ca fix: tighten env var surface (#560) 2026-05-19 17:11:31 +02:00
Michał Pierzchała 094c290703 perf: speed up iOS replay runner (#557)
* perf: speed up iOS replay runner

* fix: harden ios replay fast paths

* fix: address ci validation failures

* refactor: trim unused ios replay surface
2026-05-19 11:36:48 +02:00