1372 Commits

Author SHA1 Message Date
Michał Pierzchała 700d85ec54 0.20.7 v0.20.7 2026-08-10 21:24:34 +02:00
Michał Pierzchała c2c81549d9 feat: add simulator verification skills (#1716)
* feat: add simulator verification skills

* chore: simplify simulator skills

* docs: refine simulator skill guidance

* test: guard simulator skill workflows

* style: format simulator skill contract test
2026-08-10 21:15:53 +02:00
Michał Pierzchała fa9a350361 docs: prefer design fixes over regression-only guards (#1722)
* ci: require simplicity review for large tooling changes

* docs: prefer design constraints over regression-only fixes

* docs: simplify design-first guidance
2026-08-10 21:15:41 +02:00
Michał Pierzchała 05a1d76f2e test: add daemon RPC wire-surface compatibility gate (#1717)
* test: gate daemon RPC wire compatibility against the last released tag (#1432)

ADR 0006 fixes exactly when DAEMON_RPC_PROTOCOL_VERSION must be bumped, and
nothing checked that it was. The runtime guard (readRemoteDaemonHealth) refuses
a mismatched peer, but only fires when someone remembered the bump — a wire
change that skipped it left both sides advertising protocol 2 while parsing
different payloads, which is the failure ADR 0006 exists to prevent.

Local daemons cannot skew (isReusableDaemonInfo takes over on any package
version mismatch). Cross-machine is skewed by design — proxy, cloud/limrun, a
remote macOS host — and ADR 0006 explicitly rules package version out as the
compatibility gate there, so the one boundary where skew is intended was the
one boundary with no gate.

test/wire-compat/surface.ts declares the wire surface grouped by the ADR bullet
each group serves, quoting it, with an `uncovered` note where a bullet is only
partly digestible (the /health and /rpc literals inside http-server.ts stay
reviewer-owned: a moved route 404s at connect time rather than misparsing).
ledger.json records what each declaration hashes to, at which protocol version.

Two gates, split for the same reason the replay-compat corpus splits:
- unit-core holds the ledger to its source and prints the digest to paste;
- Released-Surface Compatibility reads the ledger at the last RELEASED tag and
  requires the drift since then to carry a bump or a compatibleChanges ack.

From one commit a bumped ledger and an unbumped one are both just an edited
file, so only a released baseline can tell them apart. Acks are keyed by the
digest they cover, so one "added an optional field" cannot launder later
changes. Digests ignore comments and formatting; the manifest's closure is
derived from the AST, so a field typed by an unlisted sibling fails rather than
sitting outside the gate.

CI cost: one added job (checkout + toolchain + two node scripts, ~1 min),
mirroring the existing full-history replay-compat job.

* test: close wire-surface overclaim and make the closure fail closed (#1432)

Addresses both review P1s on #1717.

P1 — the manifest materially overclaimed ADR 0006 coverage. It quoted all four
bullets while digesting only the payload TYPES, so the producer and consumer
seams could break a skewed peer without moving a listed digest. Now listed on
both sides of every boundary: JSON-RPC method sets and the projections that
turn each method's params into a DaemonRequest, createRpcError/sendJson/
writeRpcResponseEnvelope, resolveToken and the auth-hook types, upload
preflight/finalize/308 handlers and the resumable ticket shape, artifact route
and download/inventory framing, REST error mapping, and the client's own
payload builder, lease-method mapping, response parser and error projection.
57 -> 117 declarations.

What stays out is now named rather than implied: createDaemonHttpServer's
dispatch wiring and the /health and /rpc literals inside it. Everything it
dispatches WITH is digested individually, and a moved route 404s at connect
time rather than misparsing — the loud failure, not the silent one.

P1 — imported and re-exported payload shapes escaped the closure.
declarationHomes() scanned only the manifest's own files and the walk
continued silently when a name could not be placed, so a listed type could
gain foo?: ImportedShape from a new module and stay green. Resolution is now
explicit and fails closed: relative imports, workspace specifiers (through the
owning package's own exports map, so a re-pointed export cannot drop a type),
and facade re-export chains. Every referenced name must land on a listed
declaration, a waiver with a written reason, a declared external module, or the
TS/Node global set. Fixed two extractor blind spots the walk exposed: a
declaration's own generic parameters and `as const` were being reported as
references.

Planted-red proofs (wire-mutations.test.ts): 13 cases independently mutate
method naming, response serialization, response parsing, auth projection,
upload ticket shape, 308 framing, artifact framing, REST error mapping, and
progress framing, each asserting the digest moves; 3 probes prove the closure
really reaches across a package boundary, a facade re-export, and a plain
relative import. Mutations apply inside the declaration's own span — a
whole-file replace silently hit a sibling sharing the substring, which is how
the first draft of one case passed vacuously.

The largest waiver pair (InternalRequestOptions, CommandFlags) rests on ADR
0006's own additive rule: they reach the peer inside DaemonRequest's untyped
flags/input bags, and the decision says a new flag needs no bump. Digesting
them would fire the gate on every new CLI flag and train reviewers to
rubber-stamp acks.

* test: list the consumer half of the auxiliary HTTP boundaries (#1432)

Addresses the remaining review P1 on #1717. The manifest claimed both sides of
response/upload/artifact framing while listing nothing from upload-client.ts,
daemon-artifacts.ts, or the health consumer in daemon-client-transport.ts, so
those parsers could narrow without moving a listed digest or protocol 2.

Now listed (117 -> 141 declarations):

- /health consumer: RemoteDaemonHealth, readHealthPayload, readDaemonHttpHealth,
  readRemoteDaemonHealth. This is the sharpest of the three — narrowing the
  reader or the comparison disables the very refusal ADR 0006 exists to
  guarantee, and nothing else in the repo would notice.
- /upload consumer: UploadResponse, UploadPreflightResponse, UploadPreflightResult,
  parseUploadPreflightResult, requestUploadPreflight, uploadDirectArtifact,
  tryDirectUploadWithResume, shouldRetryDirectUpload, finalizeDirectUpload,
  uploadLegacyArtifact, ARTIFACT_HASH_ALGORITHM, isStringRecord, and
  PreparedUploadArtifact — whose sha256/sizeBytes/fileName/artifactType/
  contentType fields ARE the preflight body the daemon parses.
- /artifacts/* consumer: DaemonArtifactEndpoint, buildDaemonArtifactUrl,
  isRemoteDaemon, DownloadRemoteArtifactParams, downloadRemoteArtifact,
  materializeRemoteArtifacts, resolveMaterializedArtifactPath.

Running the closure fail-closed over the new files surfaced three more stops,
each decided rather than skipped: PreparedUploadArtifact listed (it is payload),
UploadProgressSink waived (client-local rendering, never leaves the process),
and src/daemon/types.ts#DaemonArtifact waived as a re-export alias of the listed
kernel type, matching its DaemonRequest/DaemonResponse siblings.

10 more planted-red mutations cover the new seams: health version-read and
mismatch-refusal defeated, RemoteDaemonHealth field dropped, preflight parser
narrowed, preflight/legacy response shapes narrowed, finalize body key renamed,
ticket field renamed, artifact tenant header dropped, artifact URL moved. A
fourth closure probe proves the upload-consumer files are genuinely reached by
the walk rather than merely listed. 22 -> 33 tests.

The README now states the coverage as a producer/consumer table per boundary,
so the claim is checkable at a glance instead of asserted in prose.

* test: list the client half of the resumable 308 contract (#1432)

Addresses the third review P1 on #1717. Listing the daemon's
handleResumableUpload proved it still PRODUCES 308; nothing proved the client
still CONSUMES the released one. src/remote/upload-stream.ts owns that half and
was entirely outside the manifest, so a newer client could stop accepting
`upload-offset`, change how it reads `Range: bytes=0-N`, or emit a different
resumed `Content-Range` without moving one of the 141 listed digests.

Now listed (141 -> 151): UploadStreamResponse, streamFileToHttpRequest,
streamFileToHttpRequestAttempt, buildUploadRequestHeaders, isUploadResumeStatus,
isUploadRedirectStatus, parseUploadResumeOffset, parseNonNegativeIntegerHeader,
firstHeaderValue, MAX_UPLOAD_REDIRECTS.

streamFileToHttpRequestAttempt is listed despite its size, unlike
createDaemonHttpServer which stays in `uncovered`. The distinction is stated at
the declaration: the HTTP server only dispatches to handlers that are each
digested, while the attempt loop IS the resume state machine — it decides
whether a 308 continues the upload and what the next request carries, so its
sequencing alone can break a released daemon while every helper keeps its digest.

6 new planted-red mutations prove the client half moves the ledger: a dropped
`upload-offset` fallback, narrowed Range parsing, a changed resumed
Content-Range, 308 no longer treated as continue, a narrowed UploadStreamResponse,
and dropped header-value coercion. 33 -> 39 tests.

Closure fail-closed surfaced two more stops: UploadStreamProgressOptions waived
(local byte-progress rendering) and URL/URLSearchParams added to the global set.

README now carries a `/upload` resume row in the producer/consumer table, and
names the pattern behind three rounds of review: the coverage sentence kept
getting written ahead of the coverage, so the table and the `uncovered` notes
are the claims to trust — they are checkable against surface.ts, prose is not.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-10 20:52:29 +02:00
Michał Pierzchała cdc754e6ed perf: speed up iOS agent recovery and streamline CLI guidance (#1700)
* Avoid interactive children in parent taps

* docs: streamline no-skill CLI help

* perf: recover faster from sparse iOS trees

* fix: preserve selector context for blocked parent taps

* fix: preserve coordinate text-entry focus

* fix: preserve thin parent touch targets

* fix: fail closed for unscoped iOS typing

* test: isolate replay lock fixture

* test: share node integration process
2026-08-10 20:43:01 +02:00
Michał Pierzchała 4f9aded0b8 fix(record): make an aborted authoring recording terminal by construction (#1712)
* fix(record): make an aborted authoring recording terminal by construction

A second successful `open` on an `open --save-script` session aborts the
recording: the aggregate goes to `authoring{aborted}`, `recordSession` is
cleared, and the caller is warned. `close --save-script` then refuses it with
"Retry with plain close; it will tear down the session without writing."

That promise was not kept. When the second `open` itself carried
`--save-script`, the recorder's shared flag ingress re-armed `recordSession`
while leaving the status terminal, and a bare `close` published the full
session log — the writer gated only on `recordSession` and the repair variant,
so nothing on the ordinary authoring path refused an aborted lifecycle.

The abort is now terminal by construction rather than inert by ordering:

- `isAuthoringAborted` gives the pure aggregate one home for the question.
- `applyRecordedSaveScriptFlags` takes no branch for an aborted lifecycle:
  it neither re-arms recording nor retargets the output path.
- `SessionScriptWriter` asks one `isPublicationWriteBlocked` question covering
  all three reasons to publish nothing, so every path reaching the writer
  (bare `close`, teardown, idle-reap, active publication) refuses it.

Armed recordings, published recordings, and every repair transaction are
unaffected; the control tests for those stay green against the pre-fix code
while the five new regressions go red.

Closes #1533

* docs: record the #1533 resolution in ADR 0016 and fix a stale symbol reference

The ADR 0016 close-time amendment still described #1533 as unresolved, and
the session-close.ts note named `isAuthoringAbortedWriteBlocked` — a private
helper that was folded into `isPublicationWriteBlocked` when the writer's
three sequential guards collapsed into one predicate, so the symbol names
nothing in the tree.

* fix(record): arm recording through the publication lifecycle, not around it

`recordSession` is an evidence-capture flag, but three surfaces set it
directly without consulting the publication aggregate, so it could
contradict a terminal ABORTED authoring status. The #1533 fix closed the
recorded-action ingress and made the writer refuse an ABORTED lifecycle,
then documented the remaining contradiction as acceptable — the writer's
own comment noted that "something can re-arm that boolean behind the
terminal status".

That something was live: `buildNextOpenSession` re-armed recording for any
`open --save-script`, and `applyOrdinaryScriptRecordingOpenOutcome` only
aborts a lifecycle that is still ARMED. A third `open --save-script` on an
already-ABORTED session therefore left `recordSession` true behind the
terminal status. The writer gate hid the publication symptom, but the
session kept paying recording-time costs for a recording that can never
publish: `recordSession` disables the direct iOS selector fast paths for
click and get, forcing every interaction onto the snapshot route.

Route the flag through one rule owned by the publication projection
(`recordSessionAfterSaveScriptFlag`), which answers "not recording" for an
ABORTED lifecycle on every surface that handles it — the re-open builder,
the close finalizer, and the recorded-action ingress. The writer's gate is
unchanged and still correct; it now stands on the aggregate alone rather
than as a net under a known drift, so the comments defending the drift are
replaced by statements of the rule.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFW9gJqz1wEHoowkdd1nFW

* docs: describe the full #1533 surface in the changelog entry

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFW9gJqz1wEHoowkdd1nFW

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-10 18:01:21 +02:00
Michał Pierzchała f569b91d65 docs: record platform runtime adoption checkpoint (#1703)
* docs: record platform runtime checkpoint revision

* docs: correct platform runtime checkpoint evidence

* docs: finalize platform runtime checkpoint
2026-08-10 17:58:43 +02:00
Michał Pierzchała b15c502318 refactor: extract platform network runtime (#1702)
* refactor: extract platform network runtime

* fix: preserve platform network recovery routes

* test: guard network parser placement
2026-08-10 17:58:42 +02:00
Michał Pierzchała b1ed5353d1 refactor: extract platform log runtime (#1701)
* refactor: extract platform log runtime

* fix: clear terminal app log recovery markers

* fix: preserve scoped app log tooling

* fix: preserve app log cancellation

* fix: handle large changed coverage diffs

* fix: harden Limrun runtime identity

* refactor: tighten platform log runtime

* fix: close app log trust gaps

* fix: accept canonical session path aliases

* refactor: extract durable capture kit

* fix: refresh retained log marker admission

* fix: rotate app logs after process relaunch
2026-08-10 17:58:42 +02:00
Michał Pierzchała 4279d4c580 docs: align snapshot fallback and actions guidance (#1713)
* docs: align snapshot fallback and actions guidance

The snapshot guide claimed a zero-node XCTest result fails without ever
switching to AX, but regular iOS capture has an explicit recursive-tree →
query-sweep → private-AX recovery plan (ADR 0004,
RunnerTests+SnapshotCapturePlan.swift). The public CLI also exposes
`--actions`, `--force-full`, and `--timeout`, while both website reference
pages published a three-flag snapshot usage line.

- Add a schema-derived gate: `commands.md` must publish the exact usage
  `buildCommandUsage('snapshot', getCliCommandSchema('snapshot'))` produces,
  so the canonical invocation cannot drift from the command schema again. The
  flag list is never restated in the test. Proven red against the pre-fix
  `commands.md`.
- Extract the fence walker both doc checks now share, and prove the new gate
  fails on a planted usage drift.
- Publish the canonical snapshot usage in the command reference and describe
  `--actions` as iOS-simulator-only and planning-only.
- Replace the "Backends (iOS)" list with an iOS capture behavior section
  written from ADR 0004 and the live capture plan: regular visible strategy
  with a bounded recovery ladder, raw diagnostic strategy preserving strict
  capture failures, and recovered/sparse/degraded output staying observable
  through quality warnings. Capture tiers are documented as internal, not as
  user-selectable backends.

Custom-action discovery stays separate from invocation: the runner can read
names but cannot trigger them (RunnerAXSnapshotBridge.h), so neither page
implies otherwise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014VzyMim5q4jbp1xDMm3jVo

* docs: state that --actions and --raw are rejected as a pair

Both pages described the combination as a silent no-op ("returns a raw
tree without them"), which cannot happen: `customActionFlagsResponse` in
src/daemon/request-router.ts rejects `snapshotCustomActions` + `snapshotRaw`
with INVALID_ARGS at the shared request seam, before any session or device
work, so CLI, Node client, and MCP all get the same answer. Pinned by
src/daemon/__tests__/request-router-custom-action-flags.test.ts.

The underlying reason was right and is kept — custom actions are only
readable through the private-AX capture path, which the raw diagnostic
strategy does not take — but the user-visible outcome is a rejection, not a
degraded capture, so both pages now say to choose one flag or the other.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014VzyMim5q4jbp1xDMm3jVo

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-10 16:45:34 +02:00
Michał Pierzchała b7470698a0 fix(test): include workspace packages in changed-line coverage (#1711)
The full Vitest coverage run instruments both root `src/` and
`packages/*/src/**`, but the changed-line gate's prefilter rejected every
package path before consulting LCOV. A PR could add uncovered lines to
contracts, selectors, kernel, or provider packages while the 70%
changed-line gate reported no package denominator.

Make the source-root predicate package-generic so it accepts both
coverage roots, while package tests (`.test.ts`, `__tests__/`),
package-level `test/`, and `.tsx` stay excluded as before. Scoring,
waivers, the excluded-line tally, and branch reporting are unchanged.

Pinned at both levels: the classification matrix and a pure-model
scoring regression in model.test.ts, plus a temporary-repository
regression in run.test.ts proving the executable gate fails and names the
uncovered package path and line.


Claude-Session: https://claude.ai/code/session_01HgngjVMdk2eSKQ9d1gYLGo

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-10 15:58:51 +02:00
Michał Pierzchała 4af1307024 test: cap Vitest workers for parallel worktrees (#1710)
* test: cap vitest workers for parallel worktrees

* perf: leave Vitest workers uncapped in CI
2026-08-10 15:19:54 +02:00
Michał Pierzchała 057f0e6eb1 fix(android): shell-quote free-text arguments reaching the device shell (#1645)
Text entry (input text) and clipboard write (cmd clipboard set text) now
quote their free-form text argument with the same shellQuoteIfNeeded
helper app-lifecycle.ts already uses for deep-link URLs and launch
arguments, and app-lifecycle.ts's local duplicate of that helper is
retired in favor of the shared one. Multi-word clipboard writes also now
arrive at the device as a single argument instead of being re-tokenized
into separate ones.

Updates the provider-scenario test harness's scripted clipboard-state
simulator to unwrap shell quoting the same way a device shell does, so
it keeps modelling what the device actually receives.
2026-08-10 13:47:30 +02:00
Michał Pierzchała 13cc90ffc6 fix: harden Android snapshot and fill reliability (#1708)
* fix: harden Android automation reliability

* test: isolate CLI flush integration

* test: close Android review gaps

* test: register CLI transport fixture

* test: consolidate CLI subprocess fixture
2026-08-10 13:10:18 +02:00
Michał Pierzchała c06bed9f77 refactor: extract platform device inventory runtime (#1699)
* refactor: extract platform inventory runtime

* fix: preserve scoped Apple inventory tooling

* fix: preserve Apple tool cancellation

* refactor: tighten platform inventory boundaries
2026-08-10 12:51:59 +02:00
Michał Pierzchała 44c298d7f3 docs: adopt request-bound platform runtime (#1697)
* docs: adopt request-bound platform runtime

* docs(adr): record the rejected process-local live-handle ledger alternative
2026-08-09 17:50:23 +02:00
vw2x 4b432fb59b feat: add HarmonyOS support (#1683)
* feat: add HarmonyOS device automation foundation

Add HDC-backed discovery, snapshots, application lifecycle, and core mobile interactions.

Route HarmonyOS through the platform registry and client contracts.

Cover parsing and capability parity with focused tests.

* feat: support HarmonyOS HAP deployment

Install and reinstall signed HAP archives through HDC.

Resolve bundle identities from module metadata and relaunch after package replacement.

Extend deploy routing and capability coverage for HarmonyOS.

* feat: add HarmonyOS single-pointer gestures

Execute pan, fling, and swipe plans through HDC uiInput primitives.

Derive scroll coordinates from the live ArkUI viewport.

Keep unsupported multi-touch gestures explicitly rejected.

* refactor: split session inventory command handling

Separate session, device, capability, and app inventory response paths.

Preserve the public inventory response contract while reducing handler complexity.

* feat: support HarmonyOS keyboard actions

Route HarmonyOS enter, return, and dismiss through HDC key events.

Expose supported keyboard actions through the system command metadata.

Keep keyboard visibility inspection explicitly unsupported.

* fix: reject unsupported HarmonyOS drag gestures

Keep drag unavailable until HDC can preserve source and destination hold semantics.

* feat: add HarmonyOS app log streaming

1. Stream HarmonyOS app logs through PID-scoped hilog sessions.\n2. Record HarmonyOS app identity during bundle-id opens for app-scoped commands.\n3. Cover backend routing and bundle identity resolution.

* feat: report HarmonyOS foreground app state

1. Read the foreground HarmonyOS mission through aa dump.\n2. Expose HarmonyOS appstate with package and ability metadata.\n3. Add parser coverage for foreground and missing-state cases.

* fix: advertise appstate through capabilities

1. Classify appstate in the command descriptor capability matrix.\n2. Surface supported appstate commands in capability inventory.\n3. Cover the advertised Android capability contract.

* feat: sample HarmonyOS process performance

1. Sample HarmonyOS process CPU and resident memory through HDC.\n2. Expose the verified metrics through the shared perf command.\n3. Keep frame and memory snapshot collection explicitly unavailable.

* feat: clear HarmonyOS app state

1. Add HarmonyOS settings clear-app-state through bundle cleanup.\n2. Force stop the app before clearing data and cache.\n3. Reject all unverified HarmonyOS settings explicitly.

* docs: document HarmonyOS support

1. Describe HarmonyOS HDC prerequisites and HAP installation.\n2. Add HarmonyOS to platform discovery and product documentation.\n3. Document verified performance limits for the public HDC surface.

* fix: preserve HarmonyOS deploy session identity

1. Bind a resolved HarmonyOS bundle after install or reinstall.\n2. Keep app-scoped logs and observability available after deployment.\n3. Cover session identity preservation for HarmonyOS reinstall.

* test: lock HarmonyOS capability boundary

1. Add an independent HarmonyOS capability-matrix oracle and exact advertised-command regression test.
2. Document current HDC-backed support and evidence-based unsupported command boundaries.

* refactor: simplify HarmonyOS shared platform boundaries

1. Split device selection and settings dispatch into focused helpers without changing behavior.
2. Keep HarmonyOS serial selection and lock-policy classification covered by regression tests.
3. Remove Fallow complexity findings from the HarmonyOS diff against upstream main.

* fix: bound default HarmonyOS HDC commands

1. Apply a 15 second timeout to ordinary HDC operations.
2. Preserve operation-specific timeout budgets for installation and capture paths.
3. Add regression coverage for default and overridden HDC timeouts.

* feat: add HarmonyOS screen recording

Implement physical-device whole-screen recording through the system recorder and HDC media transfer.

Reject unsupported HarmonyOS recording scopes and export flags.

Cover capability routing, media retrieval, cleanup, and simulator rejection.

* feat: report HarmonyOS HDC readiness

Add an HDC version check to the HarmonyOS doctor flow.

Document HarmonyOS as a supported doctor platform and cover the result.

* refactor: simplify HarmonyOS recording checks

Reduce recording validation and test complexity without changing behavior.

* test: cover HarmonyOS platform contracts

Synchronize public platform expectations across CLI, MCP, replay, and inventory tests.

Mock HarmonyOS inventory probes to preserve concurrent test behavior.

* test: model HarmonyOS recording capability

Require a physical HarmonyOS device in the independent capability parity oracle.

* test: cover HarmonyOS input and lifecycle paths

Exercise HDC input, lifecycle, installation, and relaunch command sequences.

* test: cover HarmonyOS device observability paths

Exercise discovery, screenshot validation, and process performance sampling.

* docs: define HarmonyOS CI hardware policy

Keep HDC hardware validation local and require mocked CI contract tests.

* fix: honor HarmonyOS app inventory filters

* fix: bound HarmonyOS app inventory classification

1. 限制应用元数据分类并发并为默认清单设置整体时限.
2. 将请求取消信号传递给 HarmonyOS 应用清单读取.
3. 补充失败时中止在飞读取且不继续排队的回归测试.

* fix: preserve HarmonyOS inventory failure causes

1. 保留触发应用元数据分类失败的原始错误, 避免被取消同级任务覆盖.
2. 补充总时限中止在飞读取且不启动排队任务的回归测试.
3. 验证后序任务失败时保留默认筛选的恢复提示.
2026-08-09 10:29:20 +02:00
Michał Pierzchała fcf429e0bd docs: split device cloud integration guides (#1695)
* docs: split device cloud integration guides

* docs: refine device cloud integration copy
2026-08-09 10:16:58 +02:00
Michał Pierzchała 18291ba8e2 perf: collapse app-driving startup turns (#1693)
* perf: collapse app-driving startup turns

* fix: align foreground open guidance
2026-08-09 10:10:11 +02:00
Michał Pierzchała e18a183ac9 fix: harden artifact ingestion boundaries (#1692)
* fix: harden artifact ingestion boundaries

* fix: bound archive inspection and upload expiry

* fix: preserve upload preflight expiry
2026-08-09 09:56:34 +02:00
Michał Pierzchała 588a419524 feat(mcp): serve the stateless 2026-07-28 revision alongside the legacy handshake (#1678)
* feat(mcp): serve the stateless 2026-07-28 revision alongside the legacy handshake

MCP 2026-07-28 drops the initialize handshake: each request carries its
protocol version and client capabilities in `_meta`, and clients probe
`server/discover` to tell a modern server from a handshake-only one.
agent-device answered that probe with -32601, so a dual-era client fell back
to `initialize` and a modern-only client had no way to connect at all.

Serve both eras from the one stdio process, which is what the spec calls a
dual-era server:

- `server/discover` advertises the supported revisions, the tools
  capability, and server identity.
- A request declaring a protocol version in `_meta` is served modern: its
  result carries `resultType: "complete"` and
  `_meta["io.modelcontextprotocol/serverInfo"]`.
- `tools/list` and `server/discover` return `ttlMs`/`cacheScope`, so clients
  can cache the 55-tool ~223KB list instead of re-fetching it every start.
  The list was already emitted sorted, which is the other half of what makes
  it cacheable.
- A declared revision we do not implement is rejected with
  `UnsupportedProtocolVersionError` (-32022) naming the ones we do.

Also fixes legacy version negotiation, which the era split surfaced:
`initialize` returned 2025-11-25 whatever the client asked for, so a client
pinned to 2025-06-18 was answered with a revision it had not requested — the
lifecycle contract's cue to disconnect. It now echoes the requested revision
when we implement it, and otherwise names the newest legacy one we do.

Legacy responses are otherwise byte-identical: `initialize` and `ping` are
still served, and no cache, `resultType`, or `_meta` field is added to them.
The stdio transport, the tool set, and every tool's schema are untouched, so
the CLI, Node, and daemon surfaces are unaffected.

Era handling lives in its own module so the router stays a dispatcher.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TKWswqMZ6Z9UAQDAda83Xz

* fix(mcp): keep 2025 revisions on the legacy wire contract and require modern metadata

Review found the era model was too loose in two ways.

Membership was one flat set, so a request declaring 2025-11-25 or
2025-06-18 through modern `_meta` was served the 2026-only envelope
(`resultType`, `serverInfo`, cache hints) — fields absent from those
revisions' schemas — and `initialize` would echo 2026-07-28, agreeing to a
revision whose handshake the modern era removed. Split modern and legacy
membership so the declared revision picks the wire contract: 2025 declared
through modern framing is answered legacy-shaped, `server/discover` requires
a modern revision because it exists in no legacy one, and `initialize`
negotiates only within the legacy set.

Modern request metadata is now required rather than guessed. `_meta`
carries `protocolVersion` and `clientCapabilities` as required fields, so
`server/discover` without them is malformed instead of being promoted to
modern, and a half-declared `_meta` is rejected as invalid params (-32602)
rather than having its lenient handling locked in by tests.

Adds black-box router cases for declared-2025 requests, initialize(2026),
`server/discover` with missing and with legacy metadata, and a declared
revision without client capabilities.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TKWswqMZ6Z9UAQDAda83Xz

* fix(mcp): validate supplied client identity and gate the methods 2026-07-28 removed

Two protocol-boundary gaps from review.

`clientInfo` was read but never checked. The field is optional in 2026-07-28,
so its absence is fine, but a supplied one must be an `Implementation` —
`clientInfo: 42` was accepted and the request served. A present value now has
to carry string `name` and `version`, matching how `clientCapabilities` is
already validated; omitting it stays legal.

`initialize` and `ping` were served regardless of era, so a request carrying
valid modern `_meta` could call methods its own revision deleted and get a
`resultType: "complete"` envelope back — with `initialize` reporting a legacy
`protocolVersion` inside a modern result. Both are now gated by the resolved
era and answer -32601 to modern-framed callers, while metadata-free legacy
calls keep working unchanged.

Adds black-box router cases for modern-framed `initialize`/`ping` (each
paired with its still-working legacy call) and for malformed `clientInfo`,
including the omitted-is-legal case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TKWswqMZ6Z9UAQDAda83Xz

* fix(mcp): validate every recognized clientInfo field, not just the required ones

The previous round checked `name` and `version` but let malformed recognized
optional fields through, so `{name:'c', version:'1', websiteUrl:42}` — and the
same for `title`, `description`, and `icons` — was accepted and served.

`Implementation` validation now type-checks each recognized field when
present: `title`, `description`, and `websiteUrl` as strings, and `icons` as
an array of `Icon`, where `src` is required and `mimeType`, `sizes`, and
`theme` are typed when supplied (`theme` against its `light`/`dark` union).
Unrecognized keys still pass — `_meta` payloads carry extension fields, and
rejecting those would reject the future.

Adds regressions across both layers: a wrong scalar per optional field, a
wrong icons container, an icon entry missing `src`, and each malformed typed
icon member. The positive cases pin the other direction — a fully populated
clientInfo carrying an extension key must still be served, so the validator
cannot harden into rejecting what the spec allows.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TKWswqMZ6Z9UAQDAda83Xz

* fix(mcp): accept schema-valid empty strings in clientInfo identity fields

`isImplementation` reused `stringField`, which requires a non-empty string, for
the required `name` and `version`. The 2026-07-28 schema declares both as plain
`string` with no minimum length, so `{name: "", version: ""}` is a conforming
`Implementation` and was being answered -32602. Required now means present and
a string.

`Icon.src` gets the same treatment. Its `format: uri` annotation is not
something this server enforces — any other non-URI string is accepted — so
rejecting the empty one alone was arbitrary rather than stricter.

Adds positive regressions at both layers for empty `name`/`version` and an
empty `Icon.src`, alongside the existing malformed cases, so the validator is
pinned against over-rejection as well as under-rejection.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01TKWswqMZ6Z9UAQDAda83Xz

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-08 12:48:38 +02:00
Michał Pierzchała 9c25bc66f4 docs(cli): advertise open --foreground and snapshot --actions in the workflow card (#1682)
* docs(cli): advertise open --foreground and snapshot --actions in the workflow card

open --foreground (#1670/#1671) and snapshot -i --actions (#1665) shipped
with no mention in the compact `help workflow` card, so a planning model
never discovers either. Add one terse line each: the foreground fast-path
in Bootstrap, and the merged-element custom-action guidance in Validation
and evidence. Stays under the 9,000-byte compact-card budget (8493 -> 8908
bytes).

Adds two help-conformance bench cases per the repo's changed-guidance rule:
foreground-attach-single-sim (correct plan starts with `open --foreground`
in an unambiguous single-sim scenario, fail-closed alternative forbidden)
and merged-card-actions-not-directly-invokable (a merged Bluesky-style feed
card's actions list is evidence, not a selector). Both use a real pinned
sample rebuilt through the production snapshot renderer.

* fix(scripts): accept flag order in the foreground-attach conformance matcher

Flag order after `open` isn't semantically meaningful (`open --platform ios
--foreground` is exactly as correct as `open --foreground --platform ios`),
but startsWithForegroundOpen required --foreground to be the literal next
token after `open`. Rescoring the completed repeat=3 bench report shows this
docked codex:gpt-5.4-mini on all 3 trials even though its plan was
config-order noise, not a real deviation -- the no-positional/no-device
guarantee already comes from the forbidden checks. Loosened to require
--foreground anywhere on the open line; foreground-attach-single-sim now
scores 54/54 across both runners.

* fix: close workflow help conformance gaps
2026-08-08 10:35:12 +02:00
Michał Pierzchała ac9e4d0f04 test: measure oracle liveness suite-wide; pin the one dead-path oracle (#1679)
* test: measure oracle liveness suite-wide; pin the one dead-path oracle

- docs/agents/oracle-negation-spike.md: assertion-negation sweep over 677
  test files (5,687 verdicts). Zero vacuous tests: all 150 negation
  survivors decompose into assert.rejects-validator artifacts (112),
  helper-oracles (31), in-file-fake breakage (6), and one conditional
  oracle. Records the companion mock-coupled coverage-uniqueness numbers
  and the follow-ups they motivate (diff-scoped mutation gate, provider
  seam closures, transcript provenance).
- watchos-sentinel: the non-watchOS test's only assertion sat in a catch
  block that never fires (tvOS interactor creation succeeds), so no
  assertion executed on the observed path. Pin creation success instead.
  Red-run proof: the old shape survived the negation sweep; the new shape
  fails under it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(android): inject fake adb through the provider scope, not PATH stubs

Adds withFakeAdb to test-utils: a scripted in-process AndroidAdbProvider
installed through the production withAndroidAdbProvider seam — the same
scope the daemon installs per request — replacing PATH-stub shell scripts
that spawn a real subprocess per adb call. No PATH mutation, no spawns,
no real subprocess waits.

Converts settings.test.ts (15 tests, 23ms; waiver said "waits real
settings-apply poll time") and notifications.test.ts (2 tests, 9ms).
Assertions move from args-log regex greps to structural checks on the
recorded call list; the fake receives device-scoped args with the
-s serial pair stripped, so serial routing is enforced by the scoped
provider matching device.id instead of asserted per call.

Remaining PATH-stub files convert next; their contention-retry waiver
entries lift together with the conversions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(android): convert device-input-state to fake adb provider injection

10 PATH-stub cases move to withFakeAdb through the production provider
scope; the 2 tests that already inject an executor directly are
unchanged. Cross-invocation shell STATE_FILE state becomes a closure
boolean; args-log regex asserts become structural checks on recorded
calls. 12/12 green at 386ms — the residue is dismissAndroidKeyboard's
two fixed 120ms retry sleeps, not stub subprocess waits.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(android): convert app-lifecycle-install adb stubbing to fake provider

The adb half of every case moves to withFakeAdb through the production
provider scope; installs take the documented exec-shaped fallback
(exec(['install','-r',...])), matching what the PATH stub saw minus the
serial pair. bundletool/zip/unzip stay real or PATH-stubbed — they run
via runCmd outside the adb seam, so this file remains in the serialized
subprocess-stub lane with its waiver reason to be corrected from adb to
bundletool. 13/13 green at ~130ms; no case enters a retry/poll loop.

Conversion note: manifest identity's `unzip -p` failure is silently
swallowed (readZipEntry catch -> undefined, aapt fallback) — an
invisible degradation path worth a future explicit diagnostic.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(android): convert input-actions adb stubbing to fake provider

9 PATH-stub cases move to withFakeAdb; the 3 tests already injecting
providers directly are unchanged. Chunked shell-input assertions become
ordered deepEqual on the recorded calls; never-called negatives and
call-count checks preserved 1:1. 12/12 green.

File time drops to 2.2s, all of it production sleeps:
verifyAndroidFilledText unconditionally waits its [0,150,350]ms
verification cadence even when the first inspection matches, so each
fill verification pass costs ~500ms with an instant fake. A
budget-derived cadence there (testing.md pattern 1) would put this file
near 25ms; flagged as follow-up rather than changed here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(android): extract shared oracles; make fake adb failure-faithful

Three test-utils extractions applied across the six converted files:

- assertRejectsAppError collapses the hand-rolled AppError code+message
  rejection validator (10 sites here; ~30 more repo-wide can adopt it
  incrementally). Validators asserting details or multiple differently-
  flagged regexes stay explicit on purpose.
- withFakeAdb gains a `provider` option for extra capabilities
  (snapshotHelperArtifact, reverse, ...), replacing input-actions'
  nested re-scoping bridge.
- withFakeAdb now mirrors the local executor's contract: a scripted
  nonzero exit throws androidAdbResultError unless the call site passed
  allowFailure. Provider-scoped exec bypasses exec.ts's throw-on-close-
  failure, so returning {exitCode:1} took a different production path
  than the PATH-stub `exit 1` these fakes replaced. All 75 tests hold
  under the corrected semantics.

Also swaps settings' inline emulator DeviceInfo literals for the shared
ANDROID_EMULATOR fixture.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test: lift five converted Android files from the contention-retry waiver

settings, notifications, device-input-state, input-actions, and
app-lifecycle-open no longer stub binaries on PATH or spawn
subprocesses, so their contention mechanism is gone: they leave
CONTENTION_RETRY_FILES and, through the derived SUBPROCESS_STUB_TESTS
constant, the serialized subprocess-stub project (17 -> 12 files).
app-lifecycle-install stays with its reason corrected: adb is now
in-process, but bundletool stays PATH-stubbed and zip/unzip spawn for
.aab packaging paths.

Full unit suite green at the new membership: 638 files, 5,724 tests,
with the five files running at unit-core's default parallelism.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test: apply review findings to the fake-adb conversion batch

- app-lifecycle-open: the missing-package launch failure returns
  {stderr, exitCode: 1} and lets withFakeAdb's throw path produce the
  production-shaped androidAdbResultError instead of hand-modeling the
  thrown AppError — the drift the helper exists to eliminate.
- withScriptedAdb deleted: the six converted files were its only
  callers, and a live PATH-stub export invites new tests back into the
  serialized lane this batch shrank. withMockedAdb stays (dispatch and
  runtime-hints tests still stub other binaries).
- android-snapshot-helper gains androidSnapshotHelperScriptResponse so
  the version-probe detection and versionCode reply have one source of
  truth; input-actions' local copy delegates to it.
- withFakeAdb's provider option becomes a distributed Omit over the
  AndroidAdbProvider union, so touch without gestureViewport is a
  compile error at the fake's boundary (planted and verified) instead
  of a TypeError inside production gesture planning.
- spike-doc re-run checklist restores wider than the codemod globs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test: drop the consumer-less FakeAdbScript barrel re-export

Fallow's dead-code gate flagged it: scripts are always passed as inline
lambdas, so only FakeAdbResponse needs a name at the barrel. The type
stays exported from fake-adb.ts where the withFakeAdb signature uses it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(apple): inject fake xcrun through the tool-provider scope, not PATH stubs

withFakeAppleTool mirrors withFakeAdb for the Apple seam: a scripted
provider installed via the production withAppleToolProvider scope, flat
simctl/devicectl invocations recorded exactly as the PATH-stub shell
scripts saw them, throw-on-nonzero fidelity matching exec.ts unless the
call site passed allowFailure, and the canned `simctl privacy help`
listing served by default (the block withMockedXcrun injected into
every script). screenshot-status-bar.test.ts converts as the exemplar:
3/3 green at 9ms with deepEqual call-sequence assertions replacing the
args-log regexes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(apple): convert apps.test.ts xcrun stubbing to fake tool provider

All withMockedXcrun scripts and hand-rolled PATH stubs move to
withFakeAppleTool; args-log regexes become structural call assertions
(exact deepEqual where order is deterministic, presence checks where
the 5s simulatorBootedMemo TTL makes boot-probe order test-dependent).
12 hand-rolled AppError validators collapse into assertRejectsAppError.
54/54 green; file test time 1172ms -> ~400ms with no test over 201ms.

Five .ipa install tests keep a minimal PATH stub for unzip only:
install-artifact.ts:112 and install-source.ts:438 call runCmd('unzip')
directly, outside the Apple tool provider seam — the file therefore
stays in the serialized subprocess-stub lane with its waiver reason
corrected from xcrun to unzip.

Also observed: getSimctlPrivacyServices caches per PATH+simulatorSetPath
and simulatorBootedMemo keys on deviceId|setPath, so neither cache
accounts for the provider scope — worked around per test, follow-up
worthy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test: lift six Apple waivers; fix format and fallow findings from CI

- interactions, simulator, screenshot, physical-device-screenshot,
  devicectl, and screenshot-status-bar leave CONTENTION_RETRY_FILES:
  the first five stopped stubbing PATH binaries in earlier refactors
  (measured 3-64ms per file, no subprocess activity), and
  screenshot-status-bar now injects through the fake tool provider.
  apps.test.ts stays with its reason corrected to the unzip PATH stub
  (xcrun is in-process; install-artifact.ts:112 / install-source.ts:438
  call runCmd('unzip') outside the Apple seam). Serialized lane 12 -> 6.
- oxfmt: fake-apple-tool.ts and contention-retry.ts were pushed
  unformatted (local check piped through tail masked the failure).
- fallow complexity: the three fake-script arrows in apps.test.ts drop
  under threshold via shared predicates (isSimctlMainScreenScale,
  isSimctlScreenshot, isDevicectlDevice), which also deduplicate the
  screenshot pair.

Full unit suite green at the new membership: 638 files, 5,724 tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-08 09:01:05 +02:00
Michał Pierzchała 6c0fcb64a1 fix: reject distinct ambiguous mutation targets (#1667)
* fix: reject distinct ambiguous mutation targets

* fix(ios): scope the raw-match rejection to mutating dispatches

`RunnerTests+Interaction.findElement` applied the new fail-closed
classification to `querySelector` as well as press/type, because the read
call site takes the default `allowNonHittableFallback: false`. With one
visible/hittable match and one non-hittable same-selector duplicate the
query started returning AMBIGUOUS_MATCH where it previously selected the
hittable element, and `queryDirectIosSelectorOrFallback` preserves that
error for read callers — so `get`, `is`, and `wait` surfaced an error
instead of their prior answer.

`classifyDirectSelectorCandidates` now takes a `rawMatchPolicy`. Mutations
keep `.rejectDistinctMatches` (the default, so no mutation call site
changes); `queryElement` passes `.preferHittableMatch`, restoring the
prior read rule: prefer the single hittable match, ambiguous only when
hittable matches compete, and never adopt the Maestro coordinate fallback.
The Maestro expected-point path is untouched.

Covers the one-hittable + one-non-hittable read, competing hittable reads,
and the non-hittable-only read. ADR 0011's amendment now states the scope.

* test(ios): execute selector read ambiguity regression

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-08 08:57:32 +02:00
Michał Pierzchała 919700fd42 perf(cli): keep node:http and ICU off the CLI startup path (-54% warm call) (#1681)
* perf(cli): keep node:http and ICU off the CLI startup path (-54% warm call)

A warm `appstate` cost 146ms; 79ms of that was startup work no local
command needs.

`node:http` was value-imported by four modules inside the eager import
closure of `src/cli.ts`, so every invocation initialized undici -- and,
under `NODE_USE_SYSTEM_CA=1`, the macOS trust store behind it. That is
~70ms of mostly off-CPU time (a blocking trustd lookup, invisible to
--cpu-prof), paid by commands that never speak HTTP: the default daemon
transport is a plain `node:net` socket. Cheap daemon-only commands paid
it in full; only commands slow enough to absorb it, like `snapshot`,
did not.

Separately, `option-schema.ts` sorted its specs with `localeCompare` at
module scope. The first `localeCompare` in a process costs ~5.4ms of ICU
collator initialization (every later one is ~0.002ms), so the CLI paid
ICU init purely to alphabetize an internal list. A case-insensitive
comparator reproduces the exact same order for the ASCII flag-key
vocabulary -- verified against all 159 keys.

Both HTTP modules are now loaded on demand, via `.default` so they stay
the same mutable module object a static import yields (the transport
tests stub `http.request` on it). Behavior is unchanged: the HTTP
transport, remote daemons and artifact upload/download all still work.

Interleaved A/B, alternating prebuilt dists in one loop, 5-min load avg
< 10, n=30 per arm:

  appstate      146.4ms -> 67.6ms   -53.8%
  session list  137.9ms -> 57.7ms   -58.2%
  snapshot -i   195.5ms -> 188.8ms   -3.4%  (device-work bound)

Without NODE_USE_SYSTEM_CA the appstate win is 73.9ms -> 57.9ms (-22%),
so this is not purely an artifact of that setting.

Measured and rejected: splitting command metadata from runtime handlers
(the facet restructure deferred by #1660). registry.js is 142kB of the
643kB eager closure but only ~1.6ms of module evaluation, and ALL module
compilation is ~9.5ms -- the whole refactor could not clear the 20ms bar
set for attempting it.

A unit test walks the eager closure of src/cli.ts and fails if any file
in it value-imports node:http/node:https again; type-only imports stay
allowed.

* test(cli): classify startup-closure imports with the ES module record

The regex-based guard only recognized `import ... from` / `export ... from`
forms, so a bare side-effect `import 'node:http'` anywhere in the eager
closure would reintroduce the full undici + system-CA initialization cost
while the guard stayed green -- the one form that binds nothing yet still
evaluates the module.

Classification now comes from oxc-parser's ES module record, the parser
this repo already uses for import analysis in scripts/lib/shipped-imports.
A side-effect import is exactly an entry-less static import, so it falls
out of the record rather than needing another regex branch; type-only
imports and re-exports are identified by their own `isType` flag instead
of being pattern-matched.

A fixture table now pins all eleven import forms -- side-effect, default,
namespace, named, value re-export, star re-export, mixed value+type,
type-only default/named/re-export, and dynamic -- so the classifier's
contract is asserted directly rather than only through the closure walk.
Verified red end to end: a side-effect import inserted into a real closure
module fails the guard and names the file.

* test(cli): follow workspace subpath imports in the startup-closure guard

The walk stopped at the package boundary, so a `node:http` import inside
any @agent-device/* module the CLI evaluates would have reintroduced the
cost invisibly -- the same blind spot the module-record rewrite closed for
side-effect imports, just relocated. Workspace specifiers now resolve
through each package's exports map.

The anti-vacuity assertion checks the closure actually contains package
files, so a resolver that silently returned null for every workspace
specifier would fail rather than leave the guard green over a src-only
closure. Verified red both ways: a side-effect import inserted into
packages/kernel/src/errors.ts, and one in a src/ module, each fail the
guard and name the offending file.

* test(cli): drop a redundant cast in the workspace exports lookup

* test(cli): classify a top-level dynamic import as eager, and share one HTTP loader

"On demand" is a claim about SCOPE, not syntax, and the guard only knew
syntax. `import('node:http')` at MODULE TOP LEVEL runs during module
evaluation -- so a top-level `await import(...)`, an `import(...).then(...)`,
or an immediately-invoked top-level function reintroduced the whole undici
plus trust-store cost while the guard stayed green. A dynamic import is lazy
only when its nearest enclosing function scope is not the module itself.

The guard now walks each closure file's OXC AST carrying an `eager` flag that
drops on entry to a FunctionDeclaration/FunctionExpression/ArrowFunction body,
and reports `import()` calls still reached with it set. Immediately-invoked
top-level functions are detected by unwrapping the callee's
ParenthesizedExpression, so `(async () => { ... })()` is eager too. The one
shape that still reads as lazy -- a top-level function invoked indirectly at
load time -- is stated as a known limitation in the test rather than left
implied.

Verified red by inserting each shape into a REACHABLE closure module
(daemon-client-transport.ts): top-level await, `.then(...)`, and the
top-level IIFE each fail the guard and name the file. The lazy case is
proven by the suite itself rather than a fixture alone -- the shared loader
below lives in the closure, so its function-local `await import('node:http')`
would turn the guard red if scope tracking were wrong.

The three lazy-load sites had drifted into three copies of the same ternary.
They now share `loadNodeHttpRequester` in src/utils/node-http.ts, next to the
other node:http helpers, which is where the reason for the indirection is
documented once. This also clears the Fallow complexity finding the previous
push introduced.

Fixture table extended to 19 cases covering both axes: static form
(side-effect, default, namespace, named, value/star re-export, mixed,
type-only x3) and dynamic scope (top-level x3, function/arrow/method-local,
and the exact ternary the shipped loader uses).
2026-08-08 08:57:19 +02:00
Michał Pierzchała e6b4fa2810 fix: isolate concurrent remote connections (#1675)
* fix: isolate concurrent remote connections

* refactor: harden remote connection state

* fix(cli): scope every emitted connect command to its own session

`scopeNextSteps` only reached `ConnectReadiness.nextSteps`, so two
command-bearing outputs still shipped unscoped:

- `providerArtifactNotes()` emitted `agent-device artifacts --json` as a
  prose note, which never passes through that helper.
- `buildDeferredRuntimeNotice()` emitted `agent-device metro prepare
  --remote-config <path>` independently in connection.ts.

On the shared-host concurrency path this branch fixes, following either
one resolves against the host-global active connection, so the artifacts
instruction can return another job's provider video and log URLs (#1659).

Both producers now take the connection state and format through one
exported `scopeCommand` helper, which is the single place a suggested
command is bound to its originating session. The metro config path is
shell-quoted alongside the session name.

Coverage: the BrowserStack route asserts human and JSON shapes carry one
`--session` per suggested command and that each emitted session resolves
back through `readRemoteConnectionState` to the connection that printed
it; `connection status` pins the scoped deferred-metro `nextStep`.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-08 07:54:43 +02:00
Michał Pierzchała 4b89482a06 fix(daemon): open --foreground P1 hotfix — selector rejection, interactive snapshot, capture-failure masking (#1671)
* fix(daemon): open --foreground P1 hotfix — selector rejection, interactive snapshot, capture-failure masking

Three P1s from the post-merge review of #1670, all at the
session-open-foreground dispatch seam:

1. Explicit device selectors were silently overwritten: the resolved-device
   rewrite pinned --udid/--platform over whatever the caller passed, so
   `open --foreground --udid B` with sim A sole-booted silently opened A.
   Now fails fast with INVALID_ARGS (matching the existing app-positional
   rejection) on --udid/--device, and on --platform other than ios; an
   explicit --platform ios passes through.

2. The promised interactive snapshot was never requested: the composed
   snapshot dispatch forwarded the open request's flags untouched, without
   snapshotInteractiveOnly — so the capture was NOT the `snapshot -i` path
   the doc comment promised and returned no interactive presentation. The
   composed request now sets snapshotInteractiveOnly: true (the exact key
   the CLI maps -i to and the snapshot runtime reads as interactiveOnly).

3. A capture failure masked the successful open: returning the snapshot
   error discarded openResponse even though the session exists, so a retry
   of `open --foreground` failed with "close the current session first".
   Open success + snapshot failure now returns ok with an explicit
   initialSnapshotError {code, message} detail and a rendered warning that
   the session IS open and how to capture manually (snapshot -i).

Regressions added for all three: explicit-selector rejection
(udid/device/both/non-iOS platform + ios pass-through), the composed
dispatch carrying snapshotInteractiveOnly, and the snapshot-failure path
returning ok + warning + usable session.

* fix(cli): render the composed open --foreground snapshot on default stdout and project initialSnapshotError through the public surfaces

Post-merge review on #1671 found the daemon fixes never reached the public
boundaries: openCliOutput ignored the nested snapshot (the one-call promise
held only under --json), and initialSnapshotError was daemon-only — absent
from AppOpenResult, Node normalization, and serializeOpenResult, with the
normalized shape truncated to code+message.

- default open output now renders the composed interactive tree through the
  same snapshotCliOutput path snapshot -i uses
- AppOpenResult carries initialSnapshotError as the FULL daemon error
  (hint/details/diagnosticId/logPath preserved) through normalization and
  serialization

Worker-authored; committed by the coordinating session after the worker
stalled twice mid-push. Tests: output.test.ts + session-open-foreground
(26 pass), typecheck, oxfmt.

* refactor(client): one daemon-error normalizer + client-route regressions for initialSnapshotError

Review follow-ups on #1671: normalizeInitialSnapshotError duplicated the
target-shutdown error normalization and pushed the module over the fallow
complexity threshold; consolidated into a single internal normalizeDaemonError
(table-driven, full shape incl. retriable/supportedOn) used by both result
paths, projected through normalizeOpenForegroundComposition.

Client-route regressions: createAgentDeviceClient().apps.open now proves the
full initialSnapshotError shape (hint/details/diagnosticId/logPath/retriable)
survives normalization, and that a malformed one is dropped — deleting the
boundary normalization fails these tests.

Also rebased onto current main.

* fix(daemon): a thrown initial-snapshot capture failure gets the same successful-open contract

Review P1 on #1671: dispatchSnapshotViaRuntime rethrows ordinary
capture/runner exceptions; the composition only handled a returned
{ ok: false }, so a thrown failure escaped to the router and failed the
whole open after the session was created — retrying then wedged on the
existing session. The catch normalizes the rejection (kernel
normalizeError, same conversion the router applies) into the shared
openWithInitialSnapshotFailure path: ok response, full-shape
initialSnapshotError, session-usable warning. Rejecting-mock regression
added alongside the returned-failure case.
2026-08-08 07:54:33 +02:00
Michał Pierzchała 13bc70f24f refactor(find): resolve a mutating find's target once, not twice (#1654) (#1669)
* fix(find): resolve a mutating find's target once, not twice (#1654)

A mutating `find click`/`find fill` captured the screen, matched by
locator under the `findAct` policy, promoted to a hittable ancestor, minted
`@eN` off the node it chose — and then re-entered the interaction leaf by
bare `@ref`, which looked that ref up AGAIN via resolveSnapshotForRef. The
second lookup reads the SESSION frame tree, not the fresh capture find
matched against, so the two could disagree: it could hand the action a
different node than find picked, or refuse as unresolvable a ref find had
observed a moment earlier.

find now passes the node itself. `internal.findPreresolvedTarget` carries
the resolved node and its tree over the in-process invoke hop, and the ref
branch adopts them instead of resolving the ref a second time.

What is NOT skipped: the guarantees. Occlusion, hittable-ancestor
promotion, and the off-screen guard all still run, on that node, at the
same symbols the ADR 0011 `runtime-ref` cells name. Only the LOOKUP is
replaced. Recording, ref-frame effects, settle/observation, and
deferred-outcome marking are untouched — they live in the dispatch wrapper,
not in resolution, which is why the invoke hop is kept rather than bypassed
the way find focus/type bypass it.

The channel is a second field rather than widening `findResolvedTarget`
because they answer different questions: that flag governs ref-frame
admission and staleness (policy), this governs resolution (lookup), and
focus/type set neither. It is `internal`-only, so it never crosses the
wire and carrying live node references is sound.

ADR 0011 re-check: the `runtime-ref` disclosure cell gains a second
producer, adoptPreresolvedRefTarget, reporting `exact`. Truthful rather
than borrowed — the ref is minted off the node handed over. It cannot
report label-fallback, which is right: label recovery is a property of
looking a stale ref up, and this path performs no lookup.

Tests pin the behavior by tap coordinates, so they name which tree the leaf
resolved against, plus a control proving the ordinary @ref path is
unchanged and a case where the session tree cannot resolve the ref at all —
that one can only pass if no second lookup runs. Verified revert-sensitive:
removing the short-circuit fails 3 of the 4, and the control stays green.

Out of scope, unchanged: the ADR 0011 `native-ref` path (web provider
clickRef/fillRef only), where the ref is the provider's own element handle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117wfrvC6MRDJUWdRBErNEF

* test(find): prove the production route, and correct the divergence claim (#1654 review P2)

The regression tests built `internal.findPreresolvedTarget` by hand and
called handleInteractionCommands directly, so they proved the leaf consumes
the channel but never that find ATTACHES it. Either producer could have
stopped doing so and they would all have stayed green — the exact gap
#1649's review named, reappearing one layer up.

Four tests now drive the real handleFindCommands with invoke wired to the
real handleInteractionCommands. Two assert click and fill each carry the
selected node (separate producers, separate forwarding hops, so asserted
separately); two advance the session tree between find's match and the
leaf's resolution and assert the dispatched point is still find's node.

That last pair also corrects the record. The claim that the leaf resolves
against a different tree than find matched is too strong for the common
path: `omitRefFrameSnapshot` (interaction-runtime.ts) makes find's internal
dispatch skip the authorized frame tree and resolve against
`session.snapshot`, which find's own capture just wrote — so the second
lookup normally AGREES, and the two non-diverged tests above pass with or
without the fix.

The divergence is real but narrower: it needs something to advance
`session.snapshot` between find's match and the leaf's resolution, which
`refreshAndroidRefSnapshotIfFreshnessActive` does on the ref path. Measured
at this head, the old code taps (60,720) "Delete" where find matched
"Save" at (310,510) — a wrong-element mutation, now pinned by both click
and fill.

So the change's value is what #1654 asked for — one resolution end to end —
plus closing that window, not a fix for a divergence occurring on every
find.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0117wfrvC6MRDJUWdRBErNEF

* fix(find): tighten resolved target provenance

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-07 17:45:43 +02:00
Michał Pierzchała 04c33f9e1e feat(ios): expose AX custom actions on merged accessibility elements (#1665)
* feat(ios): expose AX custom actions on merged accessibility elements

Apps that merge a card into one accessibility element for VoiceOver (React
Native's `accessible` prop) publish the card's real affordances as
UIAccessibilityCustomActions rather than as child elements. Our snapshot showed
only the merged node, so an agent looking for a feed card's options control had
nothing to aim at and fell back to coordinate guessing.

`snapshot --actions` now names them:

    @e8 [link] "feedItem-by-whiskers.test" actions: ["Reply", "Repost", "Open post options menu"]

Opt-in, because the AX server cannot serve custom actions in a bulk tree
request: adding the attribute makes testmanagerd's reply decoder reject the
nested arrays a custom action serializes into, drop the reply, and time the
request out (~65s vs ~110ms). Only per-element reads answer, at ~100ms each, so
the runner reads at most 12 labelled childless nodes and stops at the
capture-plan deadline. The request pins the private-AX backend, since no other
backend can read the attribute, and reports that as its own `requested-backend`
verdict so a deliberate pin never renders as a degradation warning.

Invocation is not shipped: the actions are readable but not invocable from the
runner. RunnerAXSnapshotBridge.h records the five APIs that were tried.

* fix(ios): disclose a capped custom-action pass, and read on-screen elements first

Two gaps in the first cut.

An element the bounded pass never reached rendered identically to one with no
custom actions, so a capped capture silently taught the reader that later feed
cards have no affordances — the exact mis-inference this feature exists to
prevent. The runner already counted reads against candidates; it now carries
both to the response, and the verdict renders one response-level line when the
pass was incomplete. A complete pass stays silent, and "never asked" stays
distinguishable from "read none" (the key is absent, not (0, 0)).

The obvious remedy for a capped pass would be a scoped re-run, but scope is
applied when the Swift walk builds nodes, long after the read pass, so it does
not redirect the budget at all — measured: reads=12 candidates=18 with and
without --scope. Rather than print a remedy that does nothing, the read pass now
orders candidates on-screen first. That makes the budget land on elements an
agent can act on, and makes the disclosed remedy true: scrolling changes the
on-screen set, so a re-run reads elements the previous pass could not.

Also states plainly, in the flag help and the tool/SDK field description, that
the names are for planning: nothing invokes them, so the affordance is reached
through the element's detail screen, the same control exposed elsewhere, or
coordinates.

* test(snapshot): pin the custom-action coverage pair in the verdict shape assertion

* chore(scripts): classify the --actions flag in the integration progress model

The completeness gate flagged snapshotCustomActions as unclassified, which is
what it is for. It gets its own bucket rather than joining the provider-scenario
table: the values come from the private AX client inside the runner process, and
the fake runner derives its behavior from fixture tables that cannot fabricate
custom actions, so there is no provider-backed scenario to claim. The owning
coverage is named instead — runner XCTest unit, snapshot-lines, snapshot-quality.

* fix(ios): fail closed on unserviceable --actions, bound each read, cap output, and count actions in identity

Four review findings.

1. `--actions` with `--raw`, or on any target that is not an iOS simulator, used
to succeed and return nodes with no actions — a requested capability silently
no-opped, indistinguishable from "this screen has none". Both now fail closed.
The raw pairing is rejected at the shared request seam (INVALID_ARGS) so CLI,
Node client and MCP answer alike before any device work; the platform case is
rejected once the session device is resolved (UNSUPPORTED_OPERATION), naming the
resolved target. `diff --actions` was already rejected as an unsupported flag.

2. The per-element AX read had no timeout, so one wedged element could consume
the whole capture budget. Each read now runs off-thread behind a 1s wait. A
timed-out element counts as unread, never as "read, and it has no actions", so
the existing partial-pass disclosure already covers it.

3. The element budget bounded element count only; one element could still return
an unbounded list of unbounded names. Capped at 8 names of 80 characters, and
clipped elements are counted into the coverage so a truncated list is disclosed
rather than silently presented as complete.

4. Action names were rendered unescaped, and no comparison key read them. Names
now get the same escaping as text previews plus control-character folding, so an
app-authored name cannot split or corrupt a line. `actions` joins the diff
comparable key, the unchanged-comparison projection, and — the sharper bug —
the snapshot presentation key, without which `snapshot` followed by
`snapshot --actions` on a still screen answered "unchanged" and never delivered
the actions that were explicitly requested.

* fix(ios): contain a hung custom-action read instead of accumulating orphans

The 1s read deadline frees the caller, but the underlying AX call is a
synchronous XPC round trip that cannot be cancelled — it keeps running. On a
global concurrent queue that meant repeated `snapshot --actions` against a
wedged element piled up orphaned reads, all using the shared XCAXClient
concurrently. The deadline was containment for the capture, not for the runner.

Since the call cannot be cancelled, contain it instead:

- every read runs on one dedicated serial queue, so a wedged call can never be
  joined by a second concurrent user of the shared client;
- a single-flight guard refuses to dispatch at all while an abandoned read is
  still outstanding, so a repeated capture adds no work — the dispatch counter
  stands still;
- the read pass stops at that point rather than paying a deadline per element
  on reads that would all be refused, and reports `blocked` so the capture stays
  honest. That gets its own line, because the partial-pass remedy (scroll and
  re-run) cannot clear a hang and would send the reader in circles.

Recovery needs no reset: when the hung call finally returns, in-flight drops to
zero and reads resume.

The regression drives a fake AX client that never returns, and asserts the three
things the fix exists for — exactly one in-flight read with no further
dispatches across repeated captures, immediate returns with the skip disclosed
instead of the scroll remedy, and reads working again once the wedge clears.

* ci(ios): execute the custom-action runner regressions instead of only compiling them

The iOS workflow runs a targeted -only-testing list, so a runner test that is
not named there is compiled by the build step and then never executed. All seven
custom-action tests were in that gap — including the containment regression,
which is the only executable proof that a hung AX read cannot accumulate
orphaned in-flight reads.

Red/green against the containment regression, with the fix reverted to its
pre-fix concurrency behavior (global concurrent queue, no single-flight guard,
no blocked exit):

  RED   in-flight 6 (want 1), dispatches 6 (want 1), each repeat paid the full
        1.004s deadline (want <0.2s), blocked=false (want true), and the
        in-flight drain never completed — "Exceeded timeout of 5 seconds".
  GREEN 7/7 pass, containment regression in 1.02s.
2026-08-07 17:35:48 +02:00
Michał Pierzchała 5f90d908e2 fix(ios): wait for hidden-keyboard synthesized text to commit before responding (#1676)
* fix(test): wait for typed text to settle in the hidden-keyboard runner test

testBareTypeUsesTappedInputWhenSoftwareKeyboardIsHidden read textField.value
in one shot right after executeTypeCommand returned. The simulator commits
synthesized keystrokes after the command responds, so on a loaded CI machine
the read landed mid-word — observed failures reported ("h") and
("hardware-ke"). It failed on 3 of 5 runs of a branch carrying zero Swift
changes and passed on re-run.

Poll the value until it holds the expected text (10s ceiling) and assert on
the last value read, so a real regression still fails with what the field
actually held. Both reads in the test use the same helper; the assertions are
unchanged.

* Revert "fix(test): wait for typed text to settle in the hidden-keyboard runner test"

This reverts commit caecdc6d5b.

* fix(ios): wait for hidden-keyboard synthesized text to commit before responding

The synthesized-first-responder bare-type route returned ok as soon as the
private XCTest event record was posted, while the target app was still
committing characters. On slow CI simulators the trailing characters landed
after the response, so agents (and the smoke test) observed a truncated
field value through the public type path.

After dispatch, poll the tapped element until its value reaches
textBefore + typedText, exit immediately when the app transforms the input
(formatter, mid-text caret, autocomplete), and if progress stalls as a
strict prefix, re-synthesize the missing tail once. Submit-suffixed text
keeps the old immediate return so a repair can never double-submit.

Validated on a booted iPhone 17 Pro simulator: 5/5 passes of
testBareTypeUsesTappedInputWhenSoftwareKeyboardIsHidden under full-core CPU
load, with the commit wait absorbing up to ~390ms of post-dispatch lag that
the previous code ignored (uniform ~494ms dispatch-only before); the tail
repair never had to fire, consistent with commit lag rather than true drops.

* fix(ios): make the synthesized commit wait observation-only

Addresses the P1 on #1673: a 600 ms quiet prefix cannot distinguish delayed
app-side commit from a genuinely dropped suffix, so re-synthesizing the tail
could post it while the original was still queued and commit the text twice
after the command had already reported ok.

Drop the repair path entirely — the tail builder, the quiet-window constant,
the stall tracking, and the synthesizer dependency the wait only needed in
order to repair. What remains is a bounded observation: poll the tapped
element until the value commits, the app transforms it, the value becomes
unreadable, or the 3s ceiling expires. A dropped suffix still reports ok,
exactly as before this change; only the false truncation from commit lag is
removed.

The submit-key skip stays, for its own reason: the app may clear or rewrite
the field on submit, so textBefore + typedText is not the value to wait for.

Also wire testSynthesizedTextCommitProgressWalksExpectedPrefixOnly into the
iOS smoke lane — that workflow enumerates its tests by hand with
-only-testing, so a new test that is not listed never runs.
2026-08-07 17:34:09 +02:00
Michał Pierzchała f5b6e8f28c fix: enforce replay timeout envelope (#1674)
* fix: enforce replay timeout envelope

* fix: preserve replay timeout metadata
2026-08-07 17:33:53 +02:00
Michał Pierzchała 7b4e461220 fix(test): clean Swift toolchain temporary directories (#1664)
* fix(test): clean Swift toolchain temporary directories

* fix(test): wait for Swift toolchain shutdown

* fix(test): terminate Swift toolchain process groups
2026-08-07 16:14:37 +02:00
Michał Pierzchała e14c9d8d7b fix(cloud): unblock iOS snapshot and gate cloud fill on text-entry focus (#1658) (#1666)
* fix(cloud): unblock iOS snapshot and gate cloud fill on text-entry focus (#1658)

Two bugs isolated to the cloud-webdriver iOS path.

`snapshot`/`diff` refused every capture on a live BrowserStack session with
SESSION_NOT_FOUND, instantly and without a driver round trip. The app-session
guard they ran belongs to the local XCUITest runner, which must attach to a
target app; a cloud capture reads the provider's own driver session and needs
no app identity, so it now applies to local Apple targets only. The session
was empty in the first place because the provider open path skips local app
resolution wholesale — no simctl/devicectl reaches a hosted device — and
dropped an explicitly spelled bundle id along with it. A dotted, non-deep-link
target is the bundle id under the same convention resolveIosApp applies
locally, so a cloud `open com.example.app` now records it.

`fill` tapped and sent its keys in back-to-back requests. A WebView input —
an OAuth page in a Safari view controller — takes first responder
asynchronously, so the keys landed with nothing focused while the command
still answered "Filled N chars"; tapping and filling as two separate commands
worked only because the round trip between them gave the field time to focus.
The cloud interactor now waits on the same signal the Apple runner uses, the
software keyboard going from hidden to shown after its tap, and discloses what
it observed as `textEntryReadiness` so a fill with no witness cannot pass for
a filled field. Where keyboard visibility cannot witness the focus move —
back-to-back fills into one form, the shape that failed most often — it spends
the runner's full readiness budget rather than racing the app with a short
settle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Zjbzf7HdziX9SzFpNP7WW

* fix(cloud): let a new bundle-id open replace the tracked cloud iOS app

Adopting an explicitly spelled bundle id on a provider-backed open (the
fix that makes snapshot/diff work at all) also made a previously dead
precedence rule live: the provider branch returned currentAppBundleId
first, so once a first open had populated it, `open com.a` followed by
`open com.b` left the session still reporting com.a to every
appBundleId-gated command.

The local path does the opposite, and is the convention this branch is
meant to mirror: resolveIosApp returns a dotted target unchanged and
never consults the session's current app. Only its deep-link branches
prefer the tracked id. Flip the provider branch to match — an explicit
bundle-id target wins, and everything the branch cannot name (deep
links, display names, bare open) still falls back to the tracked id.

* fix(cloud): fail a witness-less cloud fill instead of reporting it filled

Review follow-ups on #1658.

`not-observed` was still a success: it sent the keys and answered "Filled N
chars", and nothing renders `textEntryReadiness` in default CLI output — so the
exact silent success this branch exists to remove survived whenever focus never
happened. A tap that raises no keyboard now fails with
`text_entry_focus_not_observed` and sends no keys, leaving the field untouched
rather than half-written, and the readiness vocabulary keeps only outcomes that
describe a fill that did type.

The readiness budget was advertised but not enforced at the request boundary:
each keyboard probe inherited the client's 30s default, so one hung probe could
hold a 2s wait for far longer. Probes now carry their own bound, threaded
through the client as a per-request timeout override.

The probe also swallowed every error as "this driver cannot answer", which
degraded a dead session, an auth rejection, or a grid outage into a blind text
entry. Only a positively classified unimplemented route counts as unsupported
now — classified on the W3C error code rather than the status, since `unknown
command` and `invalid session id` share HTTP 404 — and everything else
propagates.

The provider scenario proved request ordering against a stub that always
accepted keys. Its fake now models the device: focus lands a beat after the
tap, and keys arriving while the keyboard is down are accepted and dropped,
exactly as an unfocused field does. The tests assert the field's own value, and
both go red against the pre-fix `fill`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Zjbzf7HdziX9SzFpNP7WW

* fix(cloud): witness the tapped field's focus before a cloud fill types

Review of cc23f2b found two ways a fill could still report success
without evidence that OUR tap focused the field it was aimed at.

P1. `settled-keyboard-up` and `settled-unknown` both typed and returned
normal success. Keyboard visibility can only witness that *a* field took
focus, never *which*: filling a second field in an already-open form
reads the same before and after, so a missed tap left the first field
focused and `POST /keys` — which the driver routes to whatever holds
first responder — appended to it while every request returned 200.

Failing those closed outright would have broken ordinary multi-field
form fills, which do work: a live AWS Device Farm run types both fields
of a WebView login correctly. So witness focus properly instead. W3C
`GET /element/active` answers the question keyboard visibility cannot —
is the thing focused now the thing I tapped — and answers it whether or
not the keyboard was already up. That becomes the primary signal
(`focused-element`); the keyboard transition stays as the fallback for
drivers without the route, and a keyboard already up on such a driver
now refuses rather than typing.

The test is identity, not geometry. Containment of the tap point looks
like the obvious rule and is wrong: focusing a field can re-lay it out.
On a live iPhone 16, tapping Safari's collapsed address bar expands it
into a taller field that no longer covers the tapped point, and a
containment-only rule refused a fill that plainly worked. So a tap that
MOVES focus counts, with containment as the second half of the test —
re-filling the already-focused field moves nothing, and only geometry
tells that from a tap that missed. Both readings are taken before the
tap, since each is evidence only as a change.

P2. The 2s budget bounded the loop but not the calls inside it: every
probe got a fixed 1500ms, so one begun near the deadline finished well
past it. Both the probe timeout and the sleep are now capped by the
remaining budget.

Also fixes a related escape the review did not name: the poll loop had
no catch, so one transient grid error aborted a fill the next poll would
have satisfied. Probe failures are now tolerated within the budget, but
a budget that expires without a single answered probe rethrows, so a
dead session surfaces as itself rather than as "the tap missed".

The provider scenario gains the two-field case the review asked for: it
begins keyboard-up with the email field focused, misses the password
tap, and asserts no keys reach the email field.

Verified on AWS Device Farm iPhone 16 / iOS 18.0 at this exact tree:
address bar (the re-layout case) and both WebView login fields all
report `focused-element`, the second with the keyboard already up, and
the device reads back `tomsmith` and a 20-character password.

* fix(cloud): refuse a cloud fill no focus probe can witness, and bound the composite probe

Two blockers from the review of 3a9aceb9.

P1. `settled-unknown` was the last path that typed without evidence: when
both the active-element and keyboard routes are positively unsupported,
`fill` settled 350ms, typed, and returned ordinary success. Nothing
renders `textEntryReadiness`, so that reached a caller looking exactly
like a fill that worked — the same silent false success #1658 is about,
just narrowed to one branch. It now refuses with a distinct reason,
`text_entry_focus_unobservable`: nothing is wrong with the target, the
driver simply cannot answer, so the caller's next move differs from a
missed tap and the hint names it — `press` then `type` stays the
deliberate way to enter text unwitnessed.

`CLOUD_TEXT_ENTRY_READINESS` is now `focused-element` and
`keyboard-shown` only. Every value describes a fill that witnessed focus
before sending a key; there is deliberately no value for typing blind.

P2. `activeElement(timeoutMs)` bounded each of its two sequential
requests by the full timeout rather than bounding the operation, so a
probe handed the 1.5s left of a 2s readiness deadline could spend ~3s
across `/element/active` and `/element/{id}/rect` and overrun the
deadline it was derived from. It now derives one deadline at entry and
gives the second request only what the first left, floored at zero so an
already-spent budget aborts immediately instead of falling back to the
client default.

The regression pins elapsed transport time across both calls, which is
what the defect is made of: the rect request answers only its own abort,
so the time it was allowed to run IS the budget it was handed. It
measures ~202ms of a shared 200ms budget before the fix and ~120ms
after.

Also updates the generic Cloud WebDriver facade scenario, whose stub
answered `{value: null}` to everything and so read as a driver with
neither route. It now answers the two focus probes, since that scenario
exercises facade wiring rather than text-entry semantics — those live in
cloud-webdriver-ios-text-entry.test.ts, which models focus properly.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-07 15:52:24 +02:00
Michał Pierzchała 9b4239b2f3 fix: heal dangling XCTestDevices redirect and prove lock-owner death before reclaim (#1672)
* fix(ios): heal dangling XCTestDevices symlink during device-set reconcile

fs.existsSync follows symlinks, so a redirect symlink whose target set was
deleted read as absent: reconcile skipped the unlink and the backup rename
failed with ENOTDIR, wedging every XCTest command on the machine until manual
cleanup. Detect symlinks with lstat (throwIfNoEntry) in both the reconcile
check and unlinkIfSymlink so the stale link is removed and the backup restored.

* fix: prove owner death before reclaiming locks, and detect zombies

A zombified lock owner passes kill(pid, 0) and still reports its original
lstart, so a killed-but-unreaped daemon held the XCTest device-set lock
forever. Conversely a ps read lost to CPU contention condemned a live owner
and let waiters steal a held lock (the flake pinOwnProcessStartTime papers
over in tests).

Unify both liveness surfaces on classifyOwnerLiveness: condemn zombies via
ps state, treat failed ps reads as no-proof rather than death, and reclaim
null-start-time owners whose pid provably started after the lock was acquired
(ps etime bound). Lock timeout errors now report ownerLiveness.

* fix: keep null-start-time lock owners fail-closed

Review P1 on #1672: the started-after-acquisition bound compared a wall-clock
birth estimate (Date.now() - ps etime) against the persisted wall-clock
acquiredAtMs, so a clock step larger than the fixed 30s slack could condemn a
live null-start owner and let a waiter steal the held lock — the failure class
this PR eliminates. There is no same-clock-domain proof of birth order to be
had here (macOS ps etime is itself wall-derived), so drop the bound: an alive
pid with no recorded start time is never reclaimed. Regression covers the
steal at both the classifier and the lock level.
2026-08-07 14:33:06 +02:00
Michał Pierzchała a158434a9c feat(cli): compact workflow help card + version header (#1663)
* feat(cli): compact workflow help card + version header

Shrinks the per-task agent protocol tax of the help/skill surface.
`agent-device help workflow` drops from 41025 to 8466 bytes (-79%) by
moving depth into new `help scripting` (save-script, secret-safe
fills, batch JSON, replay divergence/repair) and `help gestures`
(multi-touch shapes/quirks) topics, and folding a few paragraphs into
topics that already owned the subject (help debugging,
help physical-device, help validate). Content is moved, not deleted.

Every `help <topic>` first line is now `agent-device <version> —
<topic>`, so the skill router reads the CLI version off the mandatory
first help read instead of a separate `agent-device --version` call.
SKILL.md is updated to do that and stays a thin router otherwise.

The compact card also gains two terse behavioral rules: chain
confident consecutive steps with `&&` (falling back to one command at
a time when uncertain), and confirm the requested end state is
actually visible on screen before declaring a task done.

help-conformance-bench (22 cases x 2 runners) improves after the
change: 29/44 -> 32/44 passing checks.

* fix(cli): review follow-ups on the compact workflow card (#1663)

Three fixes from PR review:

- Extend the help-conformance plan validator to split a command line
  on unquoted && and validate each chained segment independently, so
  a plan that follows the workflow card's "chain confident consecutive
  steps with &&" guidance is accepted instead of rejected as one
  shell-projection violation. && inside a quoted selector value (e.g.
  label="A && B") is not a chain boundary and does not split. Adds
  unit tests for the splitter and a chains-confident-consecutive-
  settle-steps conformance case. batch stays out of this: it is
  deliberately stop-only.

- Replace the literal @ref placeholder the compact card used in its
  own "snapshot -s @ref" example with a concrete ref
  (snapshot -s @e12 (the current concrete ref)), matching the same
  card's rule against placeholder targets. Reverts the test to demand
  the concrete shape.

- Give help scripting and help gestures real conformance cases
  instead of waivers: a secret-safe recorded-fill + publish case, and
  an Android transform-then-verify case whose exact verification text
  only appears in the gestures topic. Removes both waivers.

help-conformance-bench (25 cases x 2 runners, repeat=1) after these
fixes: two full runs landed at 32/50 and 33/50. That is on par with
the pre-change baseline (29/44) once the topic-untouched cases'
run-to-run swings are accounted for (confirmed noise: one case with
zero exposure to any change here flipped 10/10 -> 1/10 on a runner API
error, and another swung across all three post-fix runs). The new
scripting case now passes 8/8 for both runners; the new chaining case
correctly reports the model's choice not to chain as a soft signal,
not a validator failure.

* fix(cli): update session.test.ts help pointer for moved script-authoring content

* fix(cli): reject empty && chain operands in the plan validator (#1663)

splitOnUnquotedAnd() previously trimmed and filtered out empty
segments, so a plan with a leading (`&& press ...`), trailing
(`press ... &&`), or doubled (`a && && b`) operator passed
validPlanCommands even though a real shell rejects all three as a
syntax error. The validator would bless a plan that fails at
execution.

Empty segments are now surfaced as an `empty-chain-operand` issue
instead of being silently dropped. The quoted-&& non-split behavior
(label="A && B") is unchanged, and a normal single command with no
chain still parses identically to before.

Adds regression tests for all three empty-operand shapes plus the
quoted-&& case.
2026-08-07 13:26:48 +02:00
Michał Pierzchała ac52281448 fix(test): deterministic temp-dir cleanup across node --test lanes (#1661)
* fix(test): deterministic temp-dir cleanup across node --test lanes

node --test has no global setup/teardown hook, so unlike Vitest (#1593) every
node --test package.json script (maestro:conformance, mutation:test,
check:affected:test, check:coverage-changed:test, check:layering,
depgraph:test, check:tmpdir-leaks:test, check:contention-retry,
test:fixture-cache, test:smoke(:web), test:integration:node,
test:concurrency-torture) still created scratch directories against the
real, unredirected os.tmpdir(), with cleanup only as reliable as each call
site's own try/finally — which a crash, OOM, or timeout kill bypasses
entirely.

Add scripts/node-test-tmpdir.ts: it wraps the whole `node --test`
invocation as a child process, redirecting TMPDIR to one disposable,
pid-tagged directory (shared root/prefix with the Vitest lane) and removing
it from the process 'exit' event, which fires on normal completion, a
thrown error, or a forwarded SIGINT/SIGTERM alike. Every node --test script
now runs through it. check-tmpdir-leaks.ts already scans by root/prefix, so
it covers both mechanisms with no changes to its detection logic.

Verified: a node --test process that mkdtemp's then gets SIGKILL'd leaves a
directory behind unwrapped; wrapped and SIGTERM'd, TMPDIR is redirected and
the directory is gone with no orphaned processes. All 13 wrapped lanes and
the full Vitest suite (5,591 tests) pass with zero residual
agent-device-test-run-* directories after the run.

Fixes #1595

* test(tmpdir): ratchet every node --test script through the wrapper

The 13 lanes wrapped in package.json were a one-time hand sweep with
nothing enforcing the pattern going forward — a 14th node --test script
added later without scripts/node-test-tmpdir.ts would silently reopen
#1595 for that one lane.

Add a structural check to scripts/node-test-tmpdir.test.ts (now part of
check:tmpdir-leaks:test) that reads package.json and fails if any script
invokes `node ... --test` without routing through the wrapper. Dumb
string matching over the scripts map, no shell parsing, with an explicit
(currently empty) NODE_TEST_WRAPPER_BYPASS_ALLOWLIST for any lane that
must legitimately bypass it. Verified it both passes on the current
package.json and fails when a synthetic unwrapped `node --test` script is
added.

* fix(test): preserve the Swift cache and close the raw node --test bypasses

Review on #1661 found two gaps:

1. The wrapper only overrode TMPDIR, so it discarded and forced a
   recompile of the durable Swift compiler cache every run instead of
   mirroring vitest-tmpdir-global-setup.ts's carve-out for it. Read
   os.tmpdir() before the child's TMPDIR redirect takes effect and set
   AGENT_DEVICE_SWIFT_CACHE_DIR from that (only when unset), same as the
   Vitest lane — the two now share one durable cache instead of each
   discarding and recompiling their own. Added a probe assertion
   (scripts/node-test-tmpdir.test.ts) that fails without the fix and
   passes with it (verified both ways).

2. docs/agents/testing.md documented raw `node --test` commands for the
   iOS smoke files, and the android/ios/conformance-regenerate/nightly
   workflows invoked `node --test` directly outside package.json. Routed
   all of them through scripts/node-test-tmpdir.ts so the documented
   local commands and CI lanes get the same crash/timeout-safe cleanup
   the package.json scripts already have.
2026-08-07 13:26:20 +02:00
Michał Pierzchała cc943400a9 feat(daemon): [RFC] prototype foreground-attach convenience (#1670)
Prototype `open --foreground`: on a fresh session with no app argument,
auto-resolves the target from the sole booted iOS simulator's sole
foreground app (reusing the exact same ambiguity-detection probe that
enriches the SESSION_NOT_FOUND hint), then attaches the initial
interactive snapshot to the response by composing the existing
snapshot-runtime dispatch. Collapses the documented 3-call
snapshot-fails -> read-hint -> open -> snapshot-succeeds dance into a
single call for the unambiguous case, while failing closed
(AMBIGUOUS_MATCH) with no guessing otherwise.

First-pass RFC, not reviewed — see PR body for the design tradeoff
writeup, live before/after evidence, and scoped-out follow-ups.
2026-08-07 13:17:13 +02:00
Michał Pierzchała 4f510ebd18 fix(daemon): SESSION_NOT_FOUND hint names the detected foreground app (#1662)
* fix(daemon): SESSION_NOT_FOUND hint names the detected foreground app

iOS snapshot/diff without an app session always emitted the generic
"Run open first" hint, even when the environment was unambiguous — costing
an agent 2-3 extra turns to discover the bundle id itself (27/30 benchmark
tasks hit this). When exactly one iOS simulator is booted and exactly one
installed app is currently running on it, the hint now names the exact
runnable `open` command instead. Any ambiguity (0/2+ booted simulators, 0/2+
running apps, or a probe failure) keeps the existing generic hint — the
detection never guesses.

- listBootedIosSimulators (devices.ts): cheap, bounded simctl probe for every
  currently booted simulator, used to detect true device-level ambiguity
  independent of what device resolution silently picked.
- detectSoleRunningIosSimulatorApp (app-resolution.ts): cross-references
  `launchctl list` running UIKitApplication jobs against `simctl listapps`
  to find the sole real installed app running, filtering out simulator-
  internal services without a version-fragile bundle-id denylist.
- ios-app-session-hint.ts composes both into the hint text, wired into
  snapshot-runtime.ts's existing SESSION_NOT_FOUND guard.

* fix(daemon): say the detected app is running, not in the foreground

launchctl only confirms the app's process is alive, not that it's the
frontmost scene — a resident-but-backgrounded app (launched, then sent
home) would match the same way. The actionable open command is correct
either way, but "in the foreground" overclaimed what the probe actually
verified. Say "running" instead.

* fix(daemon): bound and catch the hint probe, pin the printed command

Two precision fixes from review:

- The probe was best-effort in spirit but not in fact: allowFailure only
  suppresses a non-zero exit code, not a timeout or spawn-failure rejection
  (exec.ts always rejects on those). listSimulatorApps' listapps call had no
  timeout at all, so it could hang the error path indefinitely; the
  launchctl probe had no catch, so a rejection would propagate and replace
  the deterministic SESSION_NOT_FOUND with a probe failure. Threaded an
  optional timeoutMs through listSimulatorApps/listSimulatorAppMetadata
  (including its plutil fallback) and wrapped both
  detectSoleRunningIosSimulatorApp and buildIosOpenCommandHint in try/catch
  so any throw falls back to the generic hint, never propagates.

- The printed command now always pins --udid (the sole-booted-device check
  only proves uniqueness within the probed device's own simulator set, not
  that a fresh CLI invocation without --ios-simulator-device-set would
  resolve to the same device) and echoes --ios-simulator-device-set when the
  detected device carries one, so the hint is the exact command that was
  live-validated, not a shorter one that happens to work by coincidence.
  Added a length guard: wire-level error details are redacted before send
  and silently truncate any string over 400 chars, so a hint that would
  exceed a safe budget (a long device-set path pushes it there) now falls
  back instead of printing a truncated, non-copy-pasteable command.

Also: "in the foreground" overclaimed what the probe verifies (a resident
but backgrounded app would match the same running-process check) — reworded
to "running".
2026-08-07 10:43:38 +02:00
Michał Pierzchała 35f4cd63d7 perf(cli): lazy-load dedicated client command handlers (#1660)
router.ts statically imported all 10 "dedicated" command handlers
(connect, disconnect, connection, auth, daemon, device, proxy, replay,
screenshot, diff) at module top level, so every real CLI invocation paid
for loading all of them -- including replay.ts's @agent-device/maestro
subtree -- even when dispatching to an unrelated command like snapshot
or press. The generic client-backed commands (snapshot, press, tap,
fill, ...) already lazy-load their handler via a dynamic import('./generic.ts');
this brings the 10 dedicated handlers in line with that existing pattern.

ClientCommandHandlerMap (router-types.ts) is repointed to describe the
lazy-loader shape (Record<CliCommandName, () => Promise<ClientCommandHandler>>)
and the loader map is checked against it via `satisfies`, so a typo'd or
stale command key still fails to compile, matching the safety the old
eager map had.

Measured via git-stash-built before/after dist trees, interleaved 1:1 to
cancel drift, run against a live iOS session (own throwaway simulator):
the CLI-side "pre-daemon" phase (node boot + import graph + arg parse,
before the process contacts the daemon) drops from a ~121-125ms median
to a ~76-80ms median -- about a 37% cut -- consistently across snapshot -i
and press --settle. Total command wall time drops 10-21% depending on
background system load (daemon/device RPC time is unaffected by this
change and dominates more under contention).
2026-08-07 09:29:12 +02:00
Michał Pierzchała f4fc089099 0.20.6 v0.20.6 2026-08-07 08:34:50 +02:00
Michał Pierzchała 628406390b fix(daemon): report the device claim retained by a failed close (#1647)
* fix(daemon): report the device claim retained by a failed close

When close cannot confirm the device was released, it deliberately keeps the
advisory claim (handing an unconfirmed device to the next session would be
worse) but deletes the session record on the next line regardless, leaving a
claim naming a session the daemon no longer tracks with no trace. Emit a warn
diagnostic naming the device key and session, mirroring the open path's
existing rollbackNewSessionClaim handling. Retention policy is unchanged.

* test(daemon): pin claim retention on the cleanup-failure branch too

The retention decision reads `platformCloseError ?? cleanupAggregate`, and the
existing pair only drove the first input. A best-effort cleanup failure — a
wedged perfetto stop, a dead helper — is the branch operators hit more often
and reaches the same retention through a different value, so narrowing the
diagnostic to the platform-close branch left every existing test green.

Verified red against exactly that: gating the emit on `platformCloseError`
fails this test alone, 28 others unaffected.

Live evidence for the device-facing path is still outstanding; this closes the
untested residual the review named, not that requirement.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Rva4YGtSCAKJqH5PbpcCU

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-07 07:50:37 +02:00
Michał Pierzchała f4ebd6f14f fix(ios): synthesize hidden-keyboard text through responder (#1657)
* fix(ios): synthesize hidden-keyboard text through responder

* docs: document iOS text synthesis failure
2026-08-06 22:16:55 +02:00
Michał Pierzchała fd1ed578d5 fix(ios): keep the baseline presentation when a tap corroboration has no request flags (#1646)
* fix(ios): keep the baseline presentation when a tap corroboration has no request flags

matchingCaptureFlags dropped the baseline snapshot's scope/depth/raw whenever
the incoming request carried no flags, so the post-action corroboration
capture ran at the default presentation and could never match a non-default
baseline's presentationKey. The corroboration then silently declined to
engage, leaking the raw XCTEST_RECORDED_FAILURE it exists to eliminate.

* docs(test): correct tap corroboration reachability
2026-08-06 22:16:42 +02:00
Michał Pierzchała b19ee116b5 fix(remote): stop persisting the daemon bearer token, and authenticate forced-reconnect release correctly (#1648)
* fix(remote): stop persisting the daemon bearer token in connection state

ADR 0007 requires generated connection profiles to strip daemon and Metro
bearer tokens; only the Metro half was honored. `connect` was writing the
daemon bearer token into the 0600 connection-state file, and every later
command read it back out.

Stop writing `authToken` into `RemoteConnectionState['daemon']` and resolve
it at each reader from the existing flag -> environment
(AGENT_DEVICE_DAEMON_AUTH_TOKEN) -> remote-config-profile chain instead,
matching src/cli/auth-session.ts's precedence.

Behavior change: a user who ran `connect --daemon-auth-token <value>` and
relied on later commands picking the token back up from the state file will
now get an auth failure. They must export AGENT_DEVICE_DAEMON_AUTH_TOKEN,
set daemonAuthToken in their remote config, or pass --daemon-auth-token on
each command. website/docs/docs/remote-proxy.md is updated to show the
supported env-var workflow.

* fix(remote): authenticate forced-reconnect lease release with the previous endpoint's own credential

connect --force released the previous connection's lease using the new
connection's ambient daemonAuthToken instead of the previous endpoint's own
credential, and swallowed the resulting auth failure — silently orphaning the
old lease when replacing a connection with a differently-authenticated one.

Resolve the release token from the previous connection's own remote-config
profile first, fall back to the ambient token only when the two connections
share the same daemon endpoint, and otherwise skip the release and surface an
actionable notice (tenant, run id, lease id, endpoint) through the existing
connect notice channel instead of hiding the failure.

* fix(remote): stop merging ambient env defaults into the previous lease's own token

resolvePreviousOwnDaemonAuthToken read the previous connection's profile
through resolveRemoteConfigProfile, which folds AGENT_DEVICE_DAEMON_AUTH_TOKEN
(and other env defaults) into the result. When the previous config file
declared no token and the new connection's credential came from that same
global env var, it was misclassified as belonging to the previous endpoint
and sent there on forced-reconnect release — recreating the credential leak
the prior fix was meant to close, just via env instead of --daemon-auth-token.

Read the previous profile with the new readRemoteConfigFile (a provenance-
preserving, file-only load with no ambient env/CLI merging), so only a token
the previous config file itself declares can satisfy rule 1. Rules 2 and 3
are unchanged.

* fix(remote): verify the previous config file still speaks for its endpoint

Rule 1 reads the previous connection's own config file to recover a credential
that provably belongs to the previous endpoint. It re-read
`previous.remoteConfigPath` and trusted whatever token that file holds *now* —
but a config path is routinely reused, so "connect to A from ./remote.json,
re-point ./remote.json at B, connect --force" classified B's token as A's own
and sent it to A during lease release. Same cross-endpoint leak the env-merge
fix closed, arriving through the file instead of the environment.

The file must now still vouch for the previous endpoint, by either of two
independent facts: its bytes still hash to the `remoteConfigHash` recorded at
connect time (so it is literally the declaration that stood up the previous
connection), or — if it changed — it still declares the same daemon base URL.
The second is what keeps an ordinary credential rotation releasing its lease
instead of orphaning one; endpoint equality, not the fact of an edit, is what
separates rotation from re-pointing.

Endpoint comparison runs both sides through `buildRemoteConnectionDaemonState`,
the same normalizer that produced the stored `daemon.baseUrl`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Rva4YGtSCAKJqH5PbpcCU

* fix(remote): bind previous config token to its endpoint

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-06 21:32:02 +02:00
Michał Pierzchała 09a0881e6c refactor(ios): runner error classification as data; recovery tested at the transport seam (#1644)
* refactor(ios): runner error classification as data; recovery tested at the transport seam (#1631)

Error-retry classification lived in four places consulted from one catch
dispatch: two overlapping substring chains in runner-contract.ts, a
code-keyed fatality chain in runner-session.ts, and an inline composite in
runner-lifecycle.ts. RUNNER_ERROR_RULES is now the one declaration
(mirroring RUNNER_COMMAND_TRAIT_MANIFEST): each rule names an error shape
and its verdicts across the four axes (retryable, connect-retry,
session-fatal reason, restart-before-send); the predicates keep their
signatures and derive from the table. One deliberate widening, noted at the
declaration: restart-before-send matching is now case-insensitive like the
other axes (the message is our own transport literal).

runner-command-recovery.ts gets its first direct coverage — through the
real stack: a scripted fake iOS runner (local HTTP server) stands in for
the XCTest runner process, executeRunnerCommandWithSession and the
transport fetch run for real, and invalidation is observed through the
module's existing invalidateSession parameter. Seven scenarios (retained
response, runner-reported failure, in-flight, unknown lifecycle, probe
failure, missing command id, read-only completed-without-response) plus an
11-case classification suite. Zero vi.mocks of runner internals.

* test(ios): prove the recovery wiring and the recycled-pid guarantee (#1644 review)

P1 was right, and I verified it before fixing: disabling the shipped
recovery callsite in runner-lifecycle.ts left all seven recovery tests
green. They call handleRunnerTransportErrorAfterCommandSend directly, so
they prove the module, not the wiring.

runner-recovery-wiring.test.ts enters at runAppleRunnerCommand — the
production facade and provider seam — and fakes only session creation (the
xcodebuild spawn). Real: command-id assignment, provider resolution,
executeRunnerCommand's catch/classification, the recovery callsite, the
recovery module, executeRunnerCommandWithSession, the transport fetch, and
response parsing. Disabling the callsite now turns all three red. Entering
one layer lower silently defeats it: the id is assigned by the facade and
recovery declines to probe a command without one.

P2: runner-lease-recycled-pid.test.ts covers #1621 with a REAL spawned
process and real readProcessStartTime — no mock choreography. It asserts
through the injected cleanup adapter rather than by observing a kill, so a
regression reports a pid instead of signalling a live process; removing the
identity guard turns it red. Only the refusal direction, which is
contention-proof: the recorded start time never matches, so a timed-out
verification read yields the same verdict. The opposite direction needs a
second live `ps` read inside production code and flakes under load, so it
stays with the mocked cases in runner-session.test.ts where pinning is the
point.

The fake runner now scripts per command rather than as one queue, because
production sends a readiness `uptime` probe before a mutating command and
`status` during recovery; a rigid queue coupled every test to that order.

* test(ios): route the recycled-pid test through exec.ts and off the real lease tree (#1644 review)

Two fixes to the same test:

- Process execution goes through utils/exec.ts's runCmdBackground, per
  AGENTS.md's hard rule, instead of node:child_process spawn. The
  kill-induced `wait` rejection is owned deliberately.
- The lease root defaults to the developer's REAL ~/.agent-device tree, so
  the test was writing leases there. It now redirects
  AGENT_DEVICE_IOS_RUNNER_LEASE_DIR to a temp dir per test and restores it,
  writing nothing outside the sandbox. (Verified by inspecting and cleaning
  the leases the earlier revision left behind.)
2026-08-06 21:19:48 +02:00
Michał Pierzchała 10ff339d14 refactor: declare selector resolution policy as data (#1649)
* refactor: declare selector resolution policy as data (#1630)

Five native consumers of "resolve a selector against the screen" each
hand-declared their ambiguity contract as inline requireUnique/
disambiguateAmbiguous literals, so the repo's real policy matrix was only
discoverable by reading four files. SELECTOR_RESOLUTION_POLICIES
(packages/selectors) now declares one row per caller — ambiguity kind plus
the structural columns (rect, occlusion, off-screen guard, promotion, poll)
— and selectorResolutionKnobs turns a row into the engine knobs it stands
for. Callers consume rows; zero ambiguity literals remain in src.

Semantics are unchanged by construction: each row was read off its call
site. The matrix names what was previously implicit — act and get text
disambiguate, is/get attrs fail closed, exists/find-reads and wait take the
first match, mutating find rejects candidates unless narrowed (#1625).
`reject-candidates` is declaration-only and rejected by
selectorResolutionKnobs at the type level, because find enforces it through
its own narrowing rather than engine knobs.

resolution-policy-parity.test.ts gate-tests the matrix against the callers
(ADR 0011's declared-plus-gate-tested pattern): knobs must match the named
ambiguity contract, every claimed structural column must appear in the
caller's source, the read/wait pipelines must genuinely lack the machinery
they disclaim, and no caller may reintroduce an inline literal. Verified
revert-sensitive: flipping readUnique to disambiguate and faking wait's
occlusion column each fail it.

Out of scope, unchanged, per the issue: the Maestro engine (ADR 0015) and
the open click-implicit-wait product decision.

* refactor: route wait and mutating find through the policy interface (#1649 review)

P1 was right: the first head declared seven rows but genuinely routed five.
selector-wait.ts never imported its row (it called listSelectorChainMatches
directly), findAct consumed only requireRect while its ambiguity contract
stayed bespoke, and the parity test sniffed marker strings in source files —
so it stayed green across exactly that gap. Asserting about the layer I had
edited instead of the behavior it produces.

resolveSelectorChainWithPolicy is now the one policy-driven entry: it
returns a discriminated outcome (none / resolved / ambiguous) because the
rows genuinely disagree about what several matches mean, which is what
previously forced each caller to re-derive its contract inline. wait and
find's selector branch both route through it; find additionally asserts its
row still says reject-candidates rather than assuming.

The parity test is rebuilt on fixture trees driven through that interface —
no source sniffing. Wiring verified revert-sensitive: flipping the wait row
fails the policy tests, and flipping findAct fails REAL find handler tests
(ambiguous-candidate listing), which is the proof the previous version
could not produce.

One behavior nuance the fixture work surfaced and now pins: disambiguation
declines on genuinely indistinguishable candidates (the tiebreak is
evidence, not a coin flip), so an acting row surfaces ambiguity there rather
than binding one silently.

* fix(test): let fallow see the host-process mock helper's real consumers

Rebase onto main brought #1642's host-process-mock.ts into this PR's
fallow scope, where its export reports as unused. It is not: three suites
consume it, but only through `(await import(...)).pinOwnProcessStartTime`
inside vi.mock factories — vitest hoists those above static imports, so the
dynamic form is required and fallow cannot trace it statically. Documented
suppression rather than a restructure that would break the hoisting
contract.

Latent on main rather than introduced here: the audit gate is
changed-files-only, so main sees the file in scope only from a PR whose
diff contains it.

* fix: keep every candidate when a policy resolves one winner (#1649 review P1)

A real regression I introduced, not a test gap: routing wait through the
policy interface collapsed the candidate set to the winner, and the #1349
landmark check is satisfied when SOME match carries the recorded identity.
A first same-selector impostor therefore hid a later genuine landmark and
timed the wait out.

The resolved outcome now carries `matchedNodes` — the full candidate set of
the alternative the winner came from — so a policy that picks one node no
longer throws the rest away. wait passes that straight to the landmark
check, restoring the original semantics.

Regression test added at the within-one-poll shape the existing suite did
not cover (both candidates in the SAME capture, impostor first); verified
it goes red against the singleton reconstruction it replaces.

* refactor: declare only the policy fields the matrix enforces (#1649 review)

The occlusion / offscreenGuard / promotion / poll columns were never
consumed by resolveSelectorChainWithPolicy or selectorResolutionKnobs:
changing any of them left behavior and the suite green, so they were
unverifiable claims that read as truth. (My earlier source-sniffing test
"verified" them by grepping caller files for marker strings — which is why
it also stayed green when a row was disconnected entirely.)

The matrix now declares exactly what it enforces: the ambiguity contract and
the rect requirement, both consumed by the resolution interface and pinned
behaviorally. A new test asserts every row's field set, so an unenforceable
column cannot reappear without coverage — verified by re-adding one and
watching it fail. Routing the structural stages into typed behavior is
tracked in #1656 with the constraint that each field must be consumed, not
merely declared.

* fix(selectors): flatten the policy outcome at the package boundary

`PolicyResolutionOutcome.resolution` was typed as `AstSelectorResolution` and
the root façade returned it unchanged, so the parser AST #1589 confined to
`@agent-device/selectors/ast` came back through a nested field.
`selector-wait.ts` reading `outcome.resolution.selector.raw` was the runtime
proof. The existing boundary gate reads exported *names*, so it could not see
this.

The public outcome now lives beside `SelectorResolution` in
public-resolution-types.ts with its selector as text; the parser-side shape is
renamed `AstPolicyResolutionOutcome` and stays package-private, and the façade
wrapper flattens on the way out — the same treatment `resolveSelectorChain`
already gave `AstSelectorResolution`.

Two new pins, both verified red against the shape they replace: a behavioral
one asserting the façade returns selector text under every policy row, and a
structural one asserting resolution shapes are re-exported from
public-resolution-types.ts rather than from a parser-side module — which is
what distinguishes the leak from a correct re-export in a name list.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Rva4YGtSCAKJqH5PbpcCU

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-06 21:10:30 +02:00
Michał Pierzchała 0033002ed5 fix(daemon): say why a no-effect claim was withheld (#1655)
* fix(daemon): say why a no-effect claim was withheld

A vetoed gestureNoEffect claim is invisible from outside: the response
looks exactly like a gesture that worked. #1620 spent weeks unable to
tell a cross-backend pair from capture drift from real movement for
precisely this reason, and its own suggested next step was to log the
corroboration inputs and re-run.

Every veto on an accept-stale verdict now emits
post_gesture_no_effect_vetoed with a reason: baseline_rebased (the pair
is cross-backend, #1569) or surface_divergence carrying
onlyInBaseline/onlyInCurrent/rectMismatched/shared from a new
summarizeDiscriminatingSurfaceDivergence — one-sided keys are membership
drift, rectMismatched is movement.

Measured live against the seeded Bluesky fixture (167 nodes, depth-56,
private-ax): inert scrolls corroborate and the warning fires 2/2,
successful gestures never claim 2/2, zero rebases, and the one veto
observed was rectMismatched:4 shared:158 onlyIn*:0 from ambient feed
motion over an artificial 105s gap. #1620's truncation-drift half does
not reproduce; membership is stable under the remembered depth cap.

Also: marking no longer stores an empty baseline signature. '[]' is
truthy, so an invented baseline was being rebased onto a post-gesture
capture; 'no usable baseline' now has one representation, matching
markPendingInteractionOutcome.

* test(daemon): pin the veto diagnostic's payload, not just its phase

The claim tests counted `post_gesture_no_effect_vetoed` events by phase. The
operator-facing behavior this PR adds is the *reason* and the divergence
counts, and both survive an event tally: emitting the wrong reason, or dropping
the counts entirely, keeps every count at 1.

The four veto tests now assert the whole emitted payload, read back out of a
diagnostics trace file so the assertion pins the serialized NDJSON line rather
than an in-memory event. Backend rebase pins `reason: baseline_rebased` with no
counts riding along (the pair is cross-backend; there is nothing comparable to
count over). Both divergence cases pin `reason: surface_divergence` with exact
onlyInBaseline/onlyInCurrent/rectMismatched/shared — the numbers that separate
scope drift from movement, and the two membership-drift directions from each
other.

Verified red both ways the reviewer named: dropping the counts fails 2 tests,
forcing the reason to surface_divergence fails 2. The old count assertions
stayed green under both.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Rva4YGtSCAKJqH5PbpcCU

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-06 21:08:33 +02:00
Michał Pierzchała d85072d935 perf(cli): route command aliases through the help fast path (#1641)
* perf(cli): route command aliases through the help fast path

bin.ts's `--help` fast path resolved aliases through a hand-written
two-entry table that had drifted out of sync with the real
CLI_COMMAND_ALIASES registry (five entries). `tap`, `launch`, and
`relaunch` missed the table and silently fell through to a full
runCli() bootstrap just to print static help text (~150-165ms vs
~45-50ms for aliases already in the table).

Delegate to the shared normalizeCliCommandAlias registry instead of
the stale local table, so every alias the registry knows about gets
the fast path automatically.

* test(cli): add R12 layering guard for bin.ts's alias delegation

The unit test added for the alias fast-path fix (cli-help-alias-fast-path.test.ts)
calls normalizeCliCommandAlias directly, so it stays green even if bin.ts
itself reverts to a hand-rolled table — it pins the registry composition,
not bin.ts's own wiring, and bin.ts cannot be safely unit-imported (it runs
unguarded top-level dispatch on import and is deliberately excluded from
coverage).

Add an AST-based structural guard instead, in the style already established
by scripts/layering/session-state.ts, facade-exports.ts, and zero-dep-jobs.ts
(oxc-parser's module/program records, not a line scan, so a fixture's string
literal can't produce a false hit). R12 asserts two facts about src/bin.ts:
it holds a value import of normalizeCliCommandAlias from
commands/cli-command-aliases.ts, and it contains none of the registry's own
alias tokens as string literals. The token list is read out of the
registry's own source (CLI_COMMAND_ALIASES's `alias:` property values), not
hard-coded, so a future sixth alias is covered automatically. Both facts
were false on the pre-fix bin.ts, verified by reverting locally and
capturing the failure before restoring the fix.

Wired into the existing check:layering chain (already part of
check:tooling), next to R7's session-state ownership rule, which pins the
same "delegate to your single owner" shape.

* test(cli): pin the alias-resolver call into buildCommandUsageText (R12 P2)

Maintainer review of R12 (PR #1641): import-presence and literal-absence
alone let bin.ts regress to buildCommandUsageText(helpTarget) while the
normalizeCliCommandAlias import stays in place, used harmlessly elsewhere
(or not at all) — the real-tree gate stayed green through that exact
regression.

Add a third fact: bin.ts's call to buildCommandUsageText must receive, as
its argument, a call to the LOCAL binding the resolver was imported as
(aliasResolverLocalName + usageTextCallsResolver, both AST-based). Binding
by local name rather than the literal export name means a renamed import
(`as resolveAlias`) still verifies, and an unrelated same-named local
cannot be mistaken for it.

Verified by reverting locally to exactly the missed regression — import
left in place, call reverted to buildCommandUsageText(helpTarget) — and
confirming R12 now fails where the two-fact version passed; restored after.
Two negative fixtures pin the scenario going forward: import present but
unused, and import present but used only unrelated to the call.

* test(cli): make R12's delegation fact universal and value-bound

The previous fact 3 asked whether *any* `buildCommandUsageText(resolver(...))`
existed in bin.ts. That quantifier is satisfied by a decoy call while the line
that actually ships resolves nothing:

    void buildCommandUsageText(normalizeCliCommandAlias('open'));
    const commandHelp = buildCommandUsageText(helpTarget);

Fact 3 now requires EVERY `buildCommandUsageText` call to receive the imported
resolver applied to the fast path's own help-target binding, which rejects both
lines above independently. The help-target name is read from bin.ts (the
variable initialized by `resolveSimpleHelpTarget`), so renaming it re-points
the guard instead of disarming it.

Because fact 3 claims binding identity by name, it also now rejects a local
shadow of the resolver and an ambiguous second help-target declaration — a
same-named local would otherwise let the composition read as delegation while
calling something that resolves nothing.

The predicate returns the reason rather than a boolean, so the gate names which
of the several distinct failures happened.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Rva4YGtSCAKJqH5PbpcCU

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-06 21:00:22 +02:00
Michał Pierzchała 3937036e5e feat: support --settle on scroll and back (#1638) (#1650)
* feat: support --settle on scroll and back (#1638)

Scroll-then-observe and back-then-observe are legitimate agent pairs, but
the post-action observation registry never grew past the touch commands, so
`--settle` on either was rejected with INVALID_ARGS — burning a tool call
each in AppControlBench's bsky-16.

Both commands now carry the `settle` descriptor trait, and every surface
derives from it rather than a hand list: CLI allowed flags, MCP/SDK input
fields, the flag-sourced timeout envelope, and MCP ref-pinning. The CLI
flag/metadata helpers moved out of the interaction family into
post-action-observation-grammar.ts (back is a system command), and
SETTLE_REF_ISSUING_TOOLS became a derivation — a hand list would have
silently stopped pinning the new commands' refs.

settleAfterInteraction and the new settleObservationCommand are two entry
points over one engine: same loop, storage, hints, and diff bounds, with the
target-less path supplying its own baseline and no proximity point. The
daemon reaches that command through the runtime surface, never by importing
`commands/` (R2) — the same seam the touch handlers use for press/fill —
and generic-settle.ts is loaded through a lazy `await import` returning a
closure, so the interaction runtime subgraph stays out of this dispatcher's
static graph (a static edge folded ~18 files into the daemon-server type
cycle; R10 caught it).

Both of generic-settle's orderings are load-bearing and tested: the baseline
is frozen before dispatch (and before the Android dialog preflight), and the
observation runs after markDeferredInteractionOutcome so settle's first
capture folds in the #1542 stabilization rather than racing it. The ADR 0014
"a settled diff publishes refs" rule moved to settle-ref-issuance.ts, shared
by both routes.

One divergence is deliberate: scroll/back resolve no element, so the diff
baseline is the session's STORED pre-action tree — "settled tree vs the
last tree you observed" — not press's freshly resolved pre-action capture.

Both commands also switch to preserve-daemon on timeout, which changes the
non-settle path too: with --settle their dominant hang mode is now a wedged
accessibility bridge, and a timed-out capture must not reset the daemon and
lose every session (#1105). The reviewed-set gate records it.

Live-validated on an iOS 26.2 simulator (Settings): scroll --settle settled
in 1786ms with a +6/-6 diff carrying fresh refs; back --settle in 771ms with
+15/-6. Alternating cost runs, one call vs the pair it replaces:
scroll 2.9-3.0s vs 5.3-5.6s, back 3.1-3.2s vs 4.7-5.1s. Those include the
#1627 deep-capture extension.

* fix: render settled-diff refs paste-ready in CLI output

A settled diff activates a PARTIAL ref frame (ADR 0014), which admits only
the pinned `@eN~s<gen>` form of the refs it issued. The unchanged-interactive
tail already rendered that way, but the diff's own added lines rendered the
bare `@eN` embedded in the snapshot line — so a CLI caller who copied the
ref the diff just handed them got `plain_ref_requires_complete_frame` and had
to append the generation by hand.

Added lines now render pinned when the response carries `refsGeneration`,
exactly like the tail. Removed lines render verbatim: they name elements that
just left the screen, and `SettleDiffLine` never gives them a ref.

This is not new to scroll/back — press/click/fill/longpress had the same gap
since #1101. MCP was never affected: its ref-pin store rewrites plain refs on
the way in, which is why the model never sees a suffix.

Live: `scroll down --settle` now emits `+ @e14~s218078 [cell] "Game Center"`,
and `press @e14~s218078` copied straight out of that line taps successfully.

* test: record the pinned-diff-ref bytes in the output-economy baseline

Rendering added diff-line refs pinned costs 8 bytes in the two settle CLI
text samples (two `~s<gen>` suffixes). The output-economy baseline is the
tripwire for exactly this, so the increase takes an explicit reviewed waiver
rather than a silent baseline bump — the same one the settled TAIL's pins
already carry, for the same ADR 0014 reason.

Only `bytes` moves: lines, refs, hints, and shape are unchanged, which is the
evidence that this is a suffix on existing refs and not a new payload.

Caught by CI, not locally: `pnpm test:unit` runs unit-core and
subprocess-stub only, while the Coverage lane runs every vitest project.

* test: prove the generic settle degrades when its runtime cannot be built

`createGenericSettleRuntime` catches and returns undefined so an observation
that cannot even start does not fail an action that already succeeded. That
was a claim in a docstring with nothing behind it — the one changed line the
coverage gate reported uncovered (95/96).

The test puts the session in the state the catch exists for: the router
handed us a session that is no longer in the store, so building the settle
runtime throws SESSION_NOT_FOUND. The response keeps its scroll result and
simply carries no settle payload. Removing the try/catch fails it.

* build: teach fallow that vi.mock reaches pinOwnProcessStartTime dynamically

Not from this PR: #1642 added `pinOwnProcessStartTime` on main, and its three
consumers reach it the only way a Vitest module mock can —
`vi.mock(path, async (importOriginal) => (await import('...')).pinOwnProcessStartTime(...))`.
Dependency analysis cannot follow that dynamic import to a consumer, so the
export reads as dead the moment any PR pulls that file into its audit scope.
This PR is the one that did.

The entry records the consumers by path and the reason, matching the
daemon route-handler entry directly above it, which exists for the same
dynamic-`import()` limitation.

* refactor: adopt the best of the parallel #1653 implementation

Two sessions independently built #1638 (PR #1650 and PR #1653) and converged
on the same architecture — trait in the registry, one engine with two entry
points, runtime-command seam, lazy import, preserve-daemon, stored-baseline
honesty. #1650 continues; this folds in what #1653 did better:

- The agent-facing help core loop (cli-help.ts) now names scroll and back as
  settle-capable. Without this, the benchmarked closed-grammar help line kept
  instructing agents that --settle is only for press/click/fill/longpress —
  actively steering the AppControlBench models away from what #1638 shipped.
- issueSettleRefs moves into session-snapshot.ts, beside the partial-frame
  primitive it wraps, deleting the single-function settle-ref-issuance module.
- Their seam tests: back reader→writer settle plumbing, back CLI settle
  rendering, and a trait-less generic command (home) ignoring a stray settle
  flag rather than observing or rejecting.

What #1650 had that #1653 lacked, for the record: the SETTLE_REF_ISSUING_TOOLS
registry derivation (without it, MCP never pins a scroll/back settle diff's
refs and the partial frame rejects every follow-up), BackCommandResult.settle
in contracts, back's MCP output schema, paste-ready pinned diff refs, and the
docs/changelog/baseline surfaces.

* bench: help-conformance case for settled scroll-to-find planning

The #1638 extension of the closed --settle grammar to scroll/back is the
feature's entire payoff — collapsing scroll-then-observe into one call — and
the closed command list is an enumerated N whose enumerator is this bench.
The regex over the help text proves the sentence exists; this case checks
whether a model plans differently because of it.

One focused case, deliberately not coached: a pinned visible-first snapshot
(rendered by formatSnapshotText, pinned by the sample-producers gate) whose
wanted row is summarized off-screen with no ref anywhere in the output. The
tempting pre-#1638 plan is `scroll` plus a separate `snapshot -i`; acceptance
is the single settled call. Scoring was verified against eight plan shapes in
both directions before recording.

Model-backed record (claude-haiku-4-5, 3 trials, current help): 0/3 — but the
decomposition is the finding. Settle eligibility GENERALIZED (3/3 trials put
--settle on scroll unprompted; the mutation-suffix framing concern did not
materialize) and the two-call habit is residual (1/3). All three trials failed
on `scroll @e3 down --settle` — the pre-existing #1366 scroll-takes-no-target
confusion, which the live CLI recovers with a dedicated hint but a single-shot
bench cannot. The recorded gap is therefore a first-30 doc gap (nothing
teaches that scroll takes no target), not a settle-eligibility gap; tuning the
case until it passes would just delete the evidence.
2026-08-06 20:04:16 +02:00
Michał Pierzchała 8030fc10d6 fix: make Android record stop survive static screens, slow moov finalization, and dead-pid recovery (#1651)
* fix: make Android record stop survive static screens, slow moov finalization, and dead-pid recovery

Three live-reproduced defects on a loaded Pixel_7_CI emulator shared the
"pulled file is not a playable MP4" / "manifest could not be verified"
symptom family:

- The Swift video validator required duration > 0, permanently rejecting
  valid single-frame recordings of fully static screens (AVFoundation
  reports their duration as 0). Whether a run passed depended on whether
  anything — even the status-bar clock — changed during the window.
- screenrecord finalizes by patching a front-reserved moov in place, so
  the remote file size never changes; the copy path now re-pulls with
  escalating delays (750/1500/3000ms) and detects finalization from the
  pulled bytes via a mandatory ftyp+moov container sniff, which also
  preserves the truncation detection the duration check provided by
  accident. The local waitForStableFile call is gone: a pull is complete
  when adb returns.
- toybox `ps -p <missing-pid>` exits 1 with empty output — the normal
  pid-gone signature — but the recovery liveness probe read every
  non-zero exit as an uncertain adb failure, making a finished recording
  behind a live-status manifest unrecoverable forever. Empty-output
  failures now corroborate via the full process list: healthy listing
  with the pid absent recovers the finished recording; listing failure
  stays conservatively uncertain (a stale verdict deletes the manifest,
  so transport health is proven first). Transport failures keep stderr
  and exec-layer timeouts throw, which is what makes the empty-output
  signature safe to trust.

Provider-scenario coverage: in-place same-size finalization landing past
the first retry, and dead-pid recovery on a responsive device — both
fail on the previous implementation with the live-observed errors.

* refactor: extract Android liveness probe and split recording scenarios per file-shape rule

record-trace-android-recovery.ts (763 lines) and android-recording.test.ts
(1,764 lines) were both past the 500-LOC extraction tripwire. The
screenrecord liveness/probe concept this PR modified now lives in
record-trace-android-liveness.ts, and the new scenarios moved to test
files mirroring the source modules they cover
(record-trace-android-copy.test.ts, record-trace-android-liveness.test.ts)
with shared scenario plumbing in the android-recording-fixtures.ts
sibling. android-recording.test.ts shrinks to 1,476 lines — below its
pre-PR size.
2026-08-06 17:38:10 +02:00