288 Commits

Author SHA1 Message Date
Michał Pierzchała 57fc0f99fa test: fail closed on unknown recording provider commands (#1728) 2026-08-11 11:30:20 +02:00
Michał Pierzchała 1b2e786128 refactor: move screen recording onto platform runtime (#1724) 2026-08-11 10:24:57 +02:00
Michał Pierzchała 338aa2a0d5 refactor: route every native selector resolution through the policy interface (#1715)
* refactor: route every native selector resolution through the policy interface

#1649 declared the per-caller ambiguity matrix; four native call sites still
bypassed it, spreading `selectorResolutionKnobs(row)` into a raw
`resolveSelectorChain` instead of naming the row. That left the "one
interface" claim aspirational: a caller could restate its contract as engine
knobs and nothing would notice.

- `is` non-exists, `get text`/`get attrs`, find's read actions, and the
  covered-selector diagnosis probe now call `resolveSelectorChainWithPolicy`
  with their existing row. Semantics are byte-identical: the knob-backed
  branch of that interface forwards to the same engine call the call sites
  built by hand.
- The façade drops `resolveSelectorChain` and `selectorResolutionKnobs`, so
  no knob-taking resolver is reachable from outside the package and a call
  site cannot re-acquire the knobs even by accident.
  `requireUnique`/`disambiguateAmbiguous` are now named in exactly one
  function, which `resolve-with-policy.ts` and the replay resolver both
  derive through.
- `get` names the two rows it may consume as a type, so pointing it at any
  other ambiguity contract is a compile error.

Tests: selector-read-policy.test.ts pins which row each read command
consumes, end to end, on one ambiguous fixture — the only tree the rows
disagree on. Each assertion was proven red by re-pointing its caller at a
neighbouring row. The knob-consistency check moves into the package beside
the now-private helper. Test call sites that used the raw resolver move to
`resolveRecordedTarget`, the same knobs and the path that actually replays a
recorded chain.

Extracting the failure branch drops `resolveSelectorInteractionTarget` below
the complexity threshold; its `fallow-ignore` waiver is removed (verified
load-bearing before the extraction, unnecessary after).

Closes #1630. Structural stages (occlusion, off-screen, promotion, poll
budget) stay per-caller pipeline code, tracked in #1656.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD

* test: observe which node find's row selected, not just that one existed

#1715 review, P2: the find row assertion was only half a pin. `find exists`
returns `found: true` for any resolved node, and the `list` call it leaned on
goes through listFindMatches — a path that consumes no policy row at all. So
repointing findFirstLocatorMatch at `readText` left both assertions green
while selection silently moved from the document-order head to the tiebreak
winner.

Assert through `find get_attrs`, which returns the ref of the node the row
actually selected. Both neighbouring rows are now red: `readText` fails
'@e3' !== '@e2' (the move the old test missed), `readUnique` fails by
refusing the ambiguous screen. `exists` stays as a second, weaker assertion
on the same resolution.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD

* refactor: route is exists through the matrix, collapse the double match pass

Follow-up tightening on the same seam.

`is exists` reached findSelectorChainMatch directly while the `readAny` row's
own doc claimed to serve "`exists` and find's read-only actions" — true of the
docs, not of the code, which is the unverifiable-claim shape #1656's review
called out. It now names `readAny`, the row it always described. Equivalent by
construction: both take the first alternative with any match under
requireRect: false, and disclose that alternative's count.

That leaves the root façade with no consumer for findSelectorChainMatch, so it
goes the way of resolveSelectorChain — dropped from the string-only façade,
kept on the published ./ast surface. Its façade-twin type SelectorChainMatch
dies with it (fallow caught it).

resolveSelectorChainWithPolicy matched twice on the uniqueness path: once via
resolveSelectorChain, then again to fill matchedNodes. Hoisting the single
list call above the row switch removes that second pass, collapses two
duplicated ambiguous literals into one helper, and drops a `?? [resolution.node]`
fallback that was unreachable — a resolution implies its alternative matched,
so the list is never null there.

While hoisting: the resolved arm's matchedNodes can describe a different
alternative than resolution.selector, because uniqueness skips an ambiguous
alternative to try the next one. Unreachable today (only first-match callers
read it, where both come from one list), and left as-is rather than silently
changed — but the doc claimed "the alternative it came from", so it now says
what is actually true.

Tests: is exists gets a caller-level pin on the shared ambiguous fixture —
passes with matches: 2 where its fail-closed siblings refuse — proven red by
pointing it at readUnique.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD

* test: discriminate is exists's row by alternative, guard the façade structurally

#1715 review, second regression-validity gap. The `is exists` pin observed
only `pass: true` and `matches: 2` on a fixture whose first alternative was
merely TIEBREAKABLE — so disambiguation succeeded there and reported the same
count first-match would. `readAny`, `readText`, and the pre-migration raw
lookup all produced that, and only the readUnique swap I had checked went
red. One mutation proven is not the same as the row being pinned.

`exists` exposes no node ref, so the row has to be read off WHICH alternative
answered. New fixture: alternative one matches two nodes that are genuinely
indistinguishable (same depth, same area, both on screen) so the tiebreak
declines; alternative two matches exactly one. First-match answers from
alternative one; every uniqueness row skips the undecidable alternative and
answers from alternative two. Asserting the selector now separates them —
readText and readUnique both fail with `id="save-unique"` where
`label="Save"` is expected.

Restoring the raw lookup stays behaviourally invisible, though:
findSelectorChainMatch is equivalent to the readAny row it migrated to, which
is precisely why that migration preserved semantics. No fixture assertion can
catch that revert, so the guard is structural — the façade's export list must
not carry resolveSelectorChain, findSelectorChainMatch, or
selectorResolutionKnobs. Follows the packages/maestro index.test.ts
absence-assertion precedent. Verified red by re-exporting the lookup.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD

* fix: cover selector routes in device replays

* test: simplify selector replay regression

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-11 07:34:11 +02:00
Michał Pierzchała 05a1d76f2e test: add daemon RPC wire-surface compatibility gate (#1717)
* test: gate daemon RPC wire compatibility against the last released tag (#1432)

ADR 0006 fixes exactly when DAEMON_RPC_PROTOCOL_VERSION must be bumped, and
nothing checked that it was. The runtime guard (readRemoteDaemonHealth) refuses
a mismatched peer, but only fires when someone remembered the bump — a wire
change that skipped it left both sides advertising protocol 2 while parsing
different payloads, which is the failure ADR 0006 exists to prevent.

Local daemons cannot skew (isReusableDaemonInfo takes over on any package
version mismatch). Cross-machine is skewed by design — proxy, cloud/limrun, a
remote macOS host — and ADR 0006 explicitly rules package version out as the
compatibility gate there, so the one boundary where skew is intended was the
one boundary with no gate.

test/wire-compat/surface.ts declares the wire surface grouped by the ADR bullet
each group serves, quoting it, with an `uncovered` note where a bullet is only
partly digestible (the /health and /rpc literals inside http-server.ts stay
reviewer-owned: a moved route 404s at connect time rather than misparsing).
ledger.json records what each declaration hashes to, at which protocol version.

Two gates, split for the same reason the replay-compat corpus splits:
- unit-core holds the ledger to its source and prints the digest to paste;
- Released-Surface Compatibility reads the ledger at the last RELEASED tag and
  requires the drift since then to carry a bump or a compatibleChanges ack.

From one commit a bumped ledger and an unbumped one are both just an edited
file, so only a released baseline can tell them apart. Acks are keyed by the
digest they cover, so one "added an optional field" cannot launder later
changes. Digests ignore comments and formatting; the manifest's closure is
derived from the AST, so a field typed by an unlisted sibling fails rather than
sitting outside the gate.

CI cost: one added job (checkout + toolchain + two node scripts, ~1 min),
mirroring the existing full-history replay-compat job.

* test: close wire-surface overclaim and make the closure fail closed (#1432)

Addresses both review P1s on #1717.

P1 — the manifest materially overclaimed ADR 0006 coverage. It quoted all four
bullets while digesting only the payload TYPES, so the producer and consumer
seams could break a skewed peer without moving a listed digest. Now listed on
both sides of every boundary: JSON-RPC method sets and the projections that
turn each method's params into a DaemonRequest, createRpcError/sendJson/
writeRpcResponseEnvelope, resolveToken and the auth-hook types, upload
preflight/finalize/308 handlers and the resumable ticket shape, artifact route
and download/inventory framing, REST error mapping, and the client's own
payload builder, lease-method mapping, response parser and error projection.
57 -> 117 declarations.

What stays out is now named rather than implied: createDaemonHttpServer's
dispatch wiring and the /health and /rpc literals inside it. Everything it
dispatches WITH is digested individually, and a moved route 404s at connect
time rather than misparsing — the loud failure, not the silent one.

P1 — imported and re-exported payload shapes escaped the closure.
declarationHomes() scanned only the manifest's own files and the walk
continued silently when a name could not be placed, so a listed type could
gain foo?: ImportedShape from a new module and stay green. Resolution is now
explicit and fails closed: relative imports, workspace specifiers (through the
owning package's own exports map, so a re-pointed export cannot drop a type),
and facade re-export chains. Every referenced name must land on a listed
declaration, a waiver with a written reason, a declared external module, or the
TS/Node global set. Fixed two extractor blind spots the walk exposed: a
declaration's own generic parameters and `as const` were being reported as
references.

Planted-red proofs (wire-mutations.test.ts): 13 cases independently mutate
method naming, response serialization, response parsing, auth projection,
upload ticket shape, 308 framing, artifact framing, REST error mapping, and
progress framing, each asserting the digest moves; 3 probes prove the closure
really reaches across a package boundary, a facade re-export, and a plain
relative import. Mutations apply inside the declaration's own span — a
whole-file replace silently hit a sibling sharing the substring, which is how
the first draft of one case passed vacuously.

The largest waiver pair (InternalRequestOptions, CommandFlags) rests on ADR
0006's own additive rule: they reach the peer inside DaemonRequest's untyped
flags/input bags, and the decision says a new flag needs no bump. Digesting
them would fire the gate on every new CLI flag and train reviewers to
rubber-stamp acks.

* test: list the consumer half of the auxiliary HTTP boundaries (#1432)

Addresses the remaining review P1 on #1717. The manifest claimed both sides of
response/upload/artifact framing while listing nothing from upload-client.ts,
daemon-artifacts.ts, or the health consumer in daemon-client-transport.ts, so
those parsers could narrow without moving a listed digest or protocol 2.

Now listed (117 -> 141 declarations):

- /health consumer: RemoteDaemonHealth, readHealthPayload, readDaemonHttpHealth,
  readRemoteDaemonHealth. This is the sharpest of the three — narrowing the
  reader or the comparison disables the very refusal ADR 0006 exists to
  guarantee, and nothing else in the repo would notice.
- /upload consumer: UploadResponse, UploadPreflightResponse, UploadPreflightResult,
  parseUploadPreflightResult, requestUploadPreflight, uploadDirectArtifact,
  tryDirectUploadWithResume, shouldRetryDirectUpload, finalizeDirectUpload,
  uploadLegacyArtifact, ARTIFACT_HASH_ALGORITHM, isStringRecord, and
  PreparedUploadArtifact — whose sha256/sizeBytes/fileName/artifactType/
  contentType fields ARE the preflight body the daemon parses.
- /artifacts/* consumer: DaemonArtifactEndpoint, buildDaemonArtifactUrl,
  isRemoteDaemon, DownloadRemoteArtifactParams, downloadRemoteArtifact,
  materializeRemoteArtifacts, resolveMaterializedArtifactPath.

Running the closure fail-closed over the new files surfaced three more stops,
each decided rather than skipped: PreparedUploadArtifact listed (it is payload),
UploadProgressSink waived (client-local rendering, never leaves the process),
and src/daemon/types.ts#DaemonArtifact waived as a re-export alias of the listed
kernel type, matching its DaemonRequest/DaemonResponse siblings.

10 more planted-red mutations cover the new seams: health version-read and
mismatch-refusal defeated, RemoteDaemonHealth field dropped, preflight parser
narrowed, preflight/legacy response shapes narrowed, finalize body key renamed,
ticket field renamed, artifact tenant header dropped, artifact URL moved. A
fourth closure probe proves the upload-consumer files are genuinely reached by
the walk rather than merely listed. 22 -> 33 tests.

The README now states the coverage as a producer/consumer table per boundary,
so the claim is checkable at a glance instead of asserted in prose.

* test: list the client half of the resumable 308 contract (#1432)

Addresses the third review P1 on #1717. Listing the daemon's
handleResumableUpload proved it still PRODUCES 308; nothing proved the client
still CONSUMES the released one. src/remote/upload-stream.ts owns that half and
was entirely outside the manifest, so a newer client could stop accepting
`upload-offset`, change how it reads `Range: bytes=0-N`, or emit a different
resumed `Content-Range` without moving one of the 141 listed digests.

Now listed (141 -> 151): UploadStreamResponse, streamFileToHttpRequest,
streamFileToHttpRequestAttempt, buildUploadRequestHeaders, isUploadResumeStatus,
isUploadRedirectStatus, parseUploadResumeOffset, parseNonNegativeIntegerHeader,
firstHeaderValue, MAX_UPLOAD_REDIRECTS.

streamFileToHttpRequestAttempt is listed despite its size, unlike
createDaemonHttpServer which stays in `uncovered`. The distinction is stated at
the declaration: the HTTP server only dispatches to handlers that are each
digested, while the attempt loop IS the resume state machine — it decides
whether a 308 continues the upload and what the next request carries, so its
sequencing alone can break a released daemon while every helper keeps its digest.

6 new planted-red mutations prove the client half moves the ledger: a dropped
`upload-offset` fallback, narrowed Range parsing, a changed resumed
Content-Range, 308 no longer treated as continue, a narrowed UploadStreamResponse,
and dropped header-value coercion. 33 -> 39 tests.

Closure fail-closed surfaced two more stops: UploadStreamProgressOptions waived
(local byte-progress rendering) and URL/URLSearchParams added to the global set.

README now carries a `/upload` resume row in the producer/consumer table, and
names the pattern behind three rounds of review: the coverage sentence kept
getting written ahead of the coverage, so the table and the `uncovered` notes
are the claims to trust — they are checkable against surface.ts, prose is not.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-10 20:52:29 +02:00
Michał Pierzchała cdc754e6ed perf: speed up iOS agent recovery and streamline CLI guidance (#1700)
* Avoid interactive children in parent taps

* docs: streamline no-skill CLI help

* perf: recover faster from sparse iOS trees

* fix: preserve selector context for blocked parent taps

* fix: preserve coordinate text-entry focus

* fix: preserve thin parent touch targets

* fix: fail closed for unscoped iOS typing

* test: isolate replay lock fixture

* test: share node integration process
2026-08-10 20:43:01 +02:00
Michał Pierzchała b15c502318 refactor: extract platform network runtime (#1702)
* refactor: extract platform network runtime

* fix: preserve platform network recovery routes

* test: guard network parser placement
2026-08-10 17:58:42 +02:00
Michał Pierzchała b1ed5353d1 refactor: extract platform log runtime (#1701)
* refactor: extract platform log runtime

* fix: clear terminal app log recovery markers

* fix: preserve scoped app log tooling

* fix: preserve app log cancellation

* fix: handle large changed coverage diffs

* fix: harden Limrun runtime identity

* refactor: tighten platform log runtime

* fix: close app log trust gaps

* fix: accept canonical session path aliases

* refactor: extract durable capture kit

* fix: refresh retained log marker admission

* fix: rotate app logs after process relaunch
2026-08-10 17:58:42 +02:00
Michał Pierzchała 057f0e6eb1 fix(android): shell-quote free-text arguments reaching the device shell (#1645)
Text entry (input text) and clipboard write (cmd clipboard set text) now
quote their free-form text argument with the same shellQuoteIfNeeded
helper app-lifecycle.ts already uses for deep-link URLs and launch
arguments, and app-lifecycle.ts's local duplicate of that helper is
retired in favor of the shared one. Multi-word clipboard writes also now
arrive at the device as a single argument instead of being re-tokenized
into separate ones.

Updates the provider-scenario test harness's scripted clipboard-state
simulator to unwrap shell quoting the same way a device shell does, so
it keeps modelling what the device actually receives.
2026-08-10 13:47:30 +02:00
Michał Pierzchała 13cc90ffc6 fix: harden Android snapshot and fill reliability (#1708)
* fix: harden Android automation reliability

* test: isolate CLI flush integration

* test: close Android review gaps

* test: register CLI transport fixture

* test: consolidate CLI subprocess fixture
2026-08-10 13:10:18 +02:00
Michał Pierzchała c06bed9f77 refactor: extract platform device inventory runtime (#1699)
* refactor: extract platform inventory runtime

* fix: preserve scoped Apple inventory tooling

* fix: preserve Apple tool cancellation

* refactor: tighten platform inventory boundaries
2026-08-10 12:51:59 +02:00
Michał Pierzchała e18a183ac9 fix: harden artifact ingestion boundaries (#1692)
* fix: harden artifact ingestion boundaries

* fix: bound archive inspection and upload expiry

* fix: preserve upload preflight expiry
2026-08-09 09:56:34 +02:00
Michał Pierzchała 6c0fcb64a1 fix: reject distinct ambiguous mutation targets (#1667)
* fix: reject distinct ambiguous mutation targets

* fix(ios): scope the raw-match rejection to mutating dispatches

`RunnerTests+Interaction.findElement` applied the new fail-closed
classification to `querySelector` as well as press/type, because the read
call site takes the default `allowNonHittableFallback: false`. With one
visible/hittable match and one non-hittable same-selector duplicate the
query started returning AMBIGUOUS_MATCH where it previously selected the
hittable element, and `queryDirectIosSelectorOrFallback` preserves that
error for read callers — so `get`, `is`, and `wait` surfaced an error
instead of their prior answer.

`classifyDirectSelectorCandidates` now takes a `rawMatchPolicy`. Mutations
keep `.rejectDistinctMatches` (the default, so no mutation call site
changes); `queryElement` passes `.preferHittableMatch`, restoring the
prior read rule: prefer the single hittable match, ambiguous only when
hittable matches compete, and never adopt the Maestro coordinate fallback.
The Maestro expected-point path is untouched.

Covers the one-hittable + one-non-hittable read, competing hittable reads,
and the non-hittable-only read. ADR 0011's amendment now states the scope.

* test(ios): execute selector read ambiguity regression

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-08 08:57:32 +02:00
Michał Pierzchała e6b4fa2810 fix: isolate concurrent remote connections (#1675)
* fix: isolate concurrent remote connections

* refactor: harden remote connection state

* fix(cli): scope every emitted connect command to its own session

`scopeNextSteps` only reached `ConnectReadiness.nextSteps`, so two
command-bearing outputs still shipped unscoped:

- `providerArtifactNotes()` emitted `agent-device artifacts --json` as a
  prose note, which never passes through that helper.
- `buildDeferredRuntimeNotice()` emitted `agent-device metro prepare
  --remote-config <path>` independently in connection.ts.

On the shared-host concurrency path this branch fixes, following either
one resolves against the host-global active connection, so the artifacts
instruction can return another job's provider video and log URLs (#1659).

Both producers now take the connection state and format through one
exported `scopeCommand` helper, which is the single place a suggested
command is bound to its originating session. The metro config path is
shell-quoted alongside the session name.

Coverage: the BrowserStack route asserts human and JSON shapes carry one
`--session` per suggested command and that each emitted session resolves
back through `readRemoteConnectionState` to the connection that printed
it; `connection status` pins the scoped deferred-metro `nextStep`.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-08 07:54:43 +02:00
Michał Pierzchała e14c9d8d7b fix(cloud): unblock iOS snapshot and gate cloud fill on text-entry focus (#1658) (#1666)
* fix(cloud): unblock iOS snapshot and gate cloud fill on text-entry focus (#1658)

Two bugs isolated to the cloud-webdriver iOS path.

`snapshot`/`diff` refused every capture on a live BrowserStack session with
SESSION_NOT_FOUND, instantly and without a driver round trip. The app-session
guard they ran belongs to the local XCUITest runner, which must attach to a
target app; a cloud capture reads the provider's own driver session and needs
no app identity, so it now applies to local Apple targets only. The session
was empty in the first place because the provider open path skips local app
resolution wholesale — no simctl/devicectl reaches a hosted device — and
dropped an explicitly spelled bundle id along with it. A dotted, non-deep-link
target is the bundle id under the same convention resolveIosApp applies
locally, so a cloud `open com.example.app` now records it.

`fill` tapped and sent its keys in back-to-back requests. A WebView input —
an OAuth page in a Safari view controller — takes first responder
asynchronously, so the keys landed with nothing focused while the command
still answered "Filled N chars"; tapping and filling as two separate commands
worked only because the round trip between them gave the field time to focus.
The cloud interactor now waits on the same signal the Apple runner uses, the
software keyboard going from hidden to shown after its tap, and discloses what
it observed as `textEntryReadiness` so a fill with no witness cannot pass for
a filled field. Where keyboard visibility cannot witness the focus move —
back-to-back fills into one form, the shape that failed most often — it spends
the runner's full readiness budget rather than racing the app with a short
settle.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Zjbzf7HdziX9SzFpNP7WW

* fix(cloud): let a new bundle-id open replace the tracked cloud iOS app

Adopting an explicitly spelled bundle id on a provider-backed open (the
fix that makes snapshot/diff work at all) also made a previously dead
precedence rule live: the provider branch returned currentAppBundleId
first, so once a first open had populated it, `open com.a` followed by
`open com.b` left the session still reporting com.a to every
appBundleId-gated command.

The local path does the opposite, and is the convention this branch is
meant to mirror: resolveIosApp returns a dotted target unchanged and
never consults the session's current app. Only its deep-link branches
prefer the tracked id. Flip the provider branch to match — an explicit
bundle-id target wins, and everything the branch cannot name (deep
links, display names, bare open) still falls back to the tracked id.

* fix(cloud): fail a witness-less cloud fill instead of reporting it filled

Review follow-ups on #1658.

`not-observed` was still a success: it sent the keys and answered "Filled N
chars", and nothing renders `textEntryReadiness` in default CLI output — so the
exact silent success this branch exists to remove survived whenever focus never
happened. A tap that raises no keyboard now fails with
`text_entry_focus_not_observed` and sends no keys, leaving the field untouched
rather than half-written, and the readiness vocabulary keeps only outcomes that
describe a fill that did type.

The readiness budget was advertised but not enforced at the request boundary:
each keyboard probe inherited the client's 30s default, so one hung probe could
hold a 2s wait for far longer. Probes now carry their own bound, threaded
through the client as a per-request timeout override.

The probe also swallowed every error as "this driver cannot answer", which
degraded a dead session, an auth rejection, or a grid outage into a blind text
entry. Only a positively classified unimplemented route counts as unsupported
now — classified on the W3C error code rather than the status, since `unknown
command` and `invalid session id` share HTTP 404 — and everything else
propagates.

The provider scenario proved request ordering against a stub that always
accepted keys. Its fake now models the device: focus lands a beat after the
tap, and keys arriving while the keyboard is down are accepted and dropped,
exactly as an unfocused field does. The tests assert the field's own value, and
both go red against the pre-fix `fill`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Zjbzf7HdziX9SzFpNP7WW

* fix(cloud): witness the tapped field's focus before a cloud fill types

Review of cc23f2b found two ways a fill could still report success
without evidence that OUR tap focused the field it was aimed at.

P1. `settled-keyboard-up` and `settled-unknown` both typed and returned
normal success. Keyboard visibility can only witness that *a* field took
focus, never *which*: filling a second field in an already-open form
reads the same before and after, so a missed tap left the first field
focused and `POST /keys` — which the driver routes to whatever holds
first responder — appended to it while every request returned 200.

Failing those closed outright would have broken ordinary multi-field
form fills, which do work: a live AWS Device Farm run types both fields
of a WebView login correctly. So witness focus properly instead. W3C
`GET /element/active` answers the question keyboard visibility cannot —
is the thing focused now the thing I tapped — and answers it whether or
not the keyboard was already up. That becomes the primary signal
(`focused-element`); the keyboard transition stays as the fallback for
drivers without the route, and a keyboard already up on such a driver
now refuses rather than typing.

The test is identity, not geometry. Containment of the tap point looks
like the obvious rule and is wrong: focusing a field can re-lay it out.
On a live iPhone 16, tapping Safari's collapsed address bar expands it
into a taller field that no longer covers the tapped point, and a
containment-only rule refused a fill that plainly worked. So a tap that
MOVES focus counts, with containment as the second half of the test —
re-filling the already-focused field moves nothing, and only geometry
tells that from a tap that missed. Both readings are taken before the
tap, since each is evidence only as a change.

P2. The 2s budget bounded the loop but not the calls inside it: every
probe got a fixed 1500ms, so one begun near the deadline finished well
past it. Both the probe timeout and the sleep are now capped by the
remaining budget.

Also fixes a related escape the review did not name: the poll loop had
no catch, so one transient grid error aborted a fill the next poll would
have satisfied. Probe failures are now tolerated within the budget, but
a budget that expires without a single answered probe rethrows, so a
dead session surfaces as itself rather than as "the tap missed".

The provider scenario gains the two-field case the review asked for: it
begins keyboard-up with the email field focused, misses the password
tap, and asserts no keys reach the email field.

Verified on AWS Device Farm iPhone 16 / iOS 18.0 at this exact tree:
address bar (the re-layout case) and both WebView login fields all
report `focused-element`, the second with the keyboard already up, and
the device reads back `tomsmith` and a 20-character password.

* fix(cloud): refuse a cloud fill no focus probe can witness, and bound the composite probe

Two blockers from the review of 3a9aceb9.

P1. `settled-unknown` was the last path that typed without evidence: when
both the active-element and keyboard routes are positively unsupported,
`fill` settled 350ms, typed, and returned ordinary success. Nothing
renders `textEntryReadiness`, so that reached a caller looking exactly
like a fill that worked — the same silent false success #1658 is about,
just narrowed to one branch. It now refuses with a distinct reason,
`text_entry_focus_unobservable`: nothing is wrong with the target, the
driver simply cannot answer, so the caller's next move differs from a
missed tap and the hint names it — `press` then `type` stays the
deliberate way to enter text unwitnessed.

`CLOUD_TEXT_ENTRY_READINESS` is now `focused-element` and
`keyboard-shown` only. Every value describes a fill that witnessed focus
before sending a key; there is deliberately no value for typing blind.

P2. `activeElement(timeoutMs)` bounded each of its two sequential
requests by the full timeout rather than bounding the operation, so a
probe handed the 1.5s left of a 2s readiness deadline could spend ~3s
across `/element/active` and `/element/{id}/rect` and overrun the
deadline it was derived from. It now derives one deadline at entry and
gives the second request only what the first left, floored at zero so an
already-spent budget aborts immediately instead of falling back to the
client default.

The regression pins elapsed transport time across both calls, which is
what the defect is made of: the rect request answers only its own abort,
so the time it was allowed to run IS the budget it was handed. It
measures ~202ms of a shared 200ms budget before the fix and ~120ms
after.

Also updates the generic Cloud WebDriver facade scenario, whose stub
answered `{value: null}` to everything and so read as a driver with
neither route. It now answers the two focus probes, since that scenario
exercises facade wiring rather than text-entry semantics — those live in
cloud-webdriver-ios-text-entry.test.ts, which models focus properly.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-07 15:52:24 +02:00
Michał Pierzchała b19ee116b5 fix(remote): stop persisting the daemon bearer token, and authenticate forced-reconnect release correctly (#1648)
* fix(remote): stop persisting the daemon bearer token in connection state

ADR 0007 requires generated connection profiles to strip daemon and Metro
bearer tokens; only the Metro half was honored. `connect` was writing the
daemon bearer token into the 0600 connection-state file, and every later
command read it back out.

Stop writing `authToken` into `RemoteConnectionState['daemon']` and resolve
it at each reader from the existing flag -> environment
(AGENT_DEVICE_DAEMON_AUTH_TOKEN) -> remote-config-profile chain instead,
matching src/cli/auth-session.ts's precedence.

Behavior change: a user who ran `connect --daemon-auth-token <value>` and
relied on later commands picking the token back up from the state file will
now get an auth failure. They must export AGENT_DEVICE_DAEMON_AUTH_TOKEN,
set daemonAuthToken in their remote config, or pass --daemon-auth-token on
each command. website/docs/docs/remote-proxy.md is updated to show the
supported env-var workflow.

* fix(remote): authenticate forced-reconnect lease release with the previous endpoint's own credential

connect --force released the previous connection's lease using the new
connection's ambient daemonAuthToken instead of the previous endpoint's own
credential, and swallowed the resulting auth failure — silently orphaning the
old lease when replacing a connection with a differently-authenticated one.

Resolve the release token from the previous connection's own remote-config
profile first, fall back to the ambient token only when the two connections
share the same daemon endpoint, and otherwise skip the release and surface an
actionable notice (tenant, run id, lease id, endpoint) through the existing
connect notice channel instead of hiding the failure.

* fix(remote): stop merging ambient env defaults into the previous lease's own token

resolvePreviousOwnDaemonAuthToken read the previous connection's profile
through resolveRemoteConfigProfile, which folds AGENT_DEVICE_DAEMON_AUTH_TOKEN
(and other env defaults) into the result. When the previous config file
declared no token and the new connection's credential came from that same
global env var, it was misclassified as belonging to the previous endpoint
and sent there on forced-reconnect release — recreating the credential leak
the prior fix was meant to close, just via env instead of --daemon-auth-token.

Read the previous profile with the new readRemoteConfigFile (a provenance-
preserving, file-only load with no ambient env/CLI merging), so only a token
the previous config file itself declares can satisfy rule 1. Rules 2 and 3
are unchanged.

* fix(remote): verify the previous config file still speaks for its endpoint

Rule 1 reads the previous connection's own config file to recover a credential
that provably belongs to the previous endpoint. It re-read
`previous.remoteConfigPath` and trusted whatever token that file holds *now* —
but a config path is routinely reused, so "connect to A from ./remote.json,
re-point ./remote.json at B, connect --force" classified B's token as A's own
and sent it to A during lease release. Same cross-endpoint leak the env-merge
fix closed, arriving through the file instead of the environment.

The file must now still vouch for the previous endpoint, by either of two
independent facts: its bytes still hash to the `remoteConfigHash` recorded at
connect time (so it is literally the declaration that stood up the previous
connection), or — if it changed — it still declares the same daemon base URL.
The second is what keeps an ordinary credential rotation releasing its lease
instead of orphaning one; endpoint equality, not the fact of an edit, is what
separates rotation from re-pointing.

Endpoint comparison runs both sides through `buildRemoteConnectionDaemonState`,
the same normalizer that produced the stored `daemon.baseUrl`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017Rva4YGtSCAKJqH5PbpcCU

* fix(remote): bind previous config token to its endpoint

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-06 21:32:02 +02:00
Michał Pierzchała 3937036e5e feat: support --settle on scroll and back (#1638) (#1650)
* feat: support --settle on scroll and back (#1638)

Scroll-then-observe and back-then-observe are legitimate agent pairs, but
the post-action observation registry never grew past the touch commands, so
`--settle` on either was rejected with INVALID_ARGS — burning a tool call
each in AppControlBench's bsky-16.

Both commands now carry the `settle` descriptor trait, and every surface
derives from it rather than a hand list: CLI allowed flags, MCP/SDK input
fields, the flag-sourced timeout envelope, and MCP ref-pinning. The CLI
flag/metadata helpers moved out of the interaction family into
post-action-observation-grammar.ts (back is a system command), and
SETTLE_REF_ISSUING_TOOLS became a derivation — a hand list would have
silently stopped pinning the new commands' refs.

settleAfterInteraction and the new settleObservationCommand are two entry
points over one engine: same loop, storage, hints, and diff bounds, with the
target-less path supplying its own baseline and no proximity point. The
daemon reaches that command through the runtime surface, never by importing
`commands/` (R2) — the same seam the touch handlers use for press/fill —
and generic-settle.ts is loaded through a lazy `await import` returning a
closure, so the interaction runtime subgraph stays out of this dispatcher's
static graph (a static edge folded ~18 files into the daemon-server type
cycle; R10 caught it).

Both of generic-settle's orderings are load-bearing and tested: the baseline
is frozen before dispatch (and before the Android dialog preflight), and the
observation runs after markDeferredInteractionOutcome so settle's first
capture folds in the #1542 stabilization rather than racing it. The ADR 0014
"a settled diff publishes refs" rule moved to settle-ref-issuance.ts, shared
by both routes.

One divergence is deliberate: scroll/back resolve no element, so the diff
baseline is the session's STORED pre-action tree — "settled tree vs the
last tree you observed" — not press's freshly resolved pre-action capture.

Both commands also switch to preserve-daemon on timeout, which changes the
non-settle path too: with --settle their dominant hang mode is now a wedged
accessibility bridge, and a timed-out capture must not reset the daemon and
lose every session (#1105). The reviewed-set gate records it.

Live-validated on an iOS 26.2 simulator (Settings): scroll --settle settled
in 1786ms with a +6/-6 diff carrying fresh refs; back --settle in 771ms with
+15/-6. Alternating cost runs, one call vs the pair it replaces:
scroll 2.9-3.0s vs 5.3-5.6s, back 3.1-3.2s vs 4.7-5.1s. Those include the
#1627 deep-capture extension.

* fix: render settled-diff refs paste-ready in CLI output

A settled diff activates a PARTIAL ref frame (ADR 0014), which admits only
the pinned `@eN~s<gen>` form of the refs it issued. The unchanged-interactive
tail already rendered that way, but the diff's own added lines rendered the
bare `@eN` embedded in the snapshot line — so a CLI caller who copied the
ref the diff just handed them got `plain_ref_requires_complete_frame` and had
to append the generation by hand.

Added lines now render pinned when the response carries `refsGeneration`,
exactly like the tail. Removed lines render verbatim: they name elements that
just left the screen, and `SettleDiffLine` never gives them a ref.

This is not new to scroll/back — press/click/fill/longpress had the same gap
since #1101. MCP was never affected: its ref-pin store rewrites plain refs on
the way in, which is why the model never sees a suffix.

Live: `scroll down --settle` now emits `+ @e14~s218078 [cell] "Game Center"`,
and `press @e14~s218078` copied straight out of that line taps successfully.

* test: record the pinned-diff-ref bytes in the output-economy baseline

Rendering added diff-line refs pinned costs 8 bytes in the two settle CLI
text samples (two `~s<gen>` suffixes). The output-economy baseline is the
tripwire for exactly this, so the increase takes an explicit reviewed waiver
rather than a silent baseline bump — the same one the settled TAIL's pins
already carry, for the same ADR 0014 reason.

Only `bytes` moves: lines, refs, hints, and shape are unchanged, which is the
evidence that this is a suffix on existing refs and not a new payload.

Caught by CI, not locally: `pnpm test:unit` runs unit-core and
subprocess-stub only, while the Coverage lane runs every vitest project.

* test: prove the generic settle degrades when its runtime cannot be built

`createGenericSettleRuntime` catches and returns undefined so an observation
that cannot even start does not fail an action that already succeeded. That
was a claim in a docstring with nothing behind it — the one changed line the
coverage gate reported uncovered (95/96).

The test puts the session in the state the catch exists for: the router
handed us a session that is no longer in the store, so building the settle
runtime throws SESSION_NOT_FOUND. The response keeps its scroll result and
simply carries no settle payload. Removing the try/catch fails it.

* build: teach fallow that vi.mock reaches pinOwnProcessStartTime dynamically

Not from this PR: #1642 added `pinOwnProcessStartTime` on main, and its three
consumers reach it the only way a Vitest module mock can —
`vi.mock(path, async (importOriginal) => (await import('...')).pinOwnProcessStartTime(...))`.
Dependency analysis cannot follow that dynamic import to a consumer, so the
export reads as dead the moment any PR pulls that file into its audit scope.
This PR is the one that did.

The entry records the consumers by path and the reason, matching the
daemon route-handler entry directly above it, which exists for the same
dynamic-`import()` limitation.

* refactor: adopt the best of the parallel #1653 implementation

Two sessions independently built #1638 (PR #1650 and PR #1653) and converged
on the same architecture — trait in the registry, one engine with two entry
points, runtime-command seam, lazy import, preserve-daemon, stored-baseline
honesty. #1650 continues; this folds in what #1653 did better:

- The agent-facing help core loop (cli-help.ts) now names scroll and back as
  settle-capable. Without this, the benchmarked closed-grammar help line kept
  instructing agents that --settle is only for press/click/fill/longpress —
  actively steering the AppControlBench models away from what #1638 shipped.
- issueSettleRefs moves into session-snapshot.ts, beside the partial-frame
  primitive it wraps, deleting the single-function settle-ref-issuance module.
- Their seam tests: back reader→writer settle plumbing, back CLI settle
  rendering, and a trait-less generic command (home) ignoring a stray settle
  flag rather than observing or rejecting.

What #1650 had that #1653 lacked, for the record: the SETTLE_REF_ISSUING_TOOLS
registry derivation (without it, MCP never pins a scroll/back settle diff's
refs and the partial frame rejects every follow-up), BackCommandResult.settle
in contracts, back's MCP output schema, paste-ready pinned diff refs, and the
docs/changelog/baseline surfaces.

* bench: help-conformance case for settled scroll-to-find planning

The #1638 extension of the closed --settle grammar to scroll/back is the
feature's entire payoff — collapsing scroll-then-observe into one call — and
the closed command list is an enumerated N whose enumerator is this bench.
The regex over the help text proves the sentence exists; this case checks
whether a model plans differently because of it.

One focused case, deliberately not coached: a pinned visible-first snapshot
(rendered by formatSnapshotText, pinned by the sample-producers gate) whose
wanted row is summarized off-screen with no ref anywhere in the output. The
tempting pre-#1638 plan is `scroll` plus a separate `snapshot -i`; acceptance
is the single settled call. Scoring was verified against eight plan shapes in
both directions before recording.

Model-backed record (claude-haiku-4-5, 3 trials, current help): 0/3 — but the
decomposition is the finding. Settle eligibility GENERALIZED (3/3 trials put
--settle on scroll unprompted; the mutation-suffix framing concern did not
materialize) and the two-call habit is residual (1/3). All three trials failed
on `scroll @e3 down --settle` — the pre-existing #1366 scroll-takes-no-target
confusion, which the live CLI recovers with a dedicated hint but a single-shot
bench cannot. The recorded gap is therefore a first-30 doc gap (nothing
teaches that scroll takes no target), not a settle-eligibility gap; tuning the
case until it passes would just delete the evidence.
2026-08-06 20:04:16 +02:00
Michał Pierzchała 8030fc10d6 fix: make Android record stop survive static screens, slow moov finalization, and dead-pid recovery (#1651)
* fix: make Android record stop survive static screens, slow moov finalization, and dead-pid recovery

Three live-reproduced defects on a loaded Pixel_7_CI emulator shared the
"pulled file is not a playable MP4" / "manifest could not be verified"
symptom family:

- The Swift video validator required duration > 0, permanently rejecting
  valid single-frame recordings of fully static screens (AVFoundation
  reports their duration as 0). Whether a run passed depended on whether
  anything — even the status-bar clock — changed during the window.
- screenrecord finalizes by patching a front-reserved moov in place, so
  the remote file size never changes; the copy path now re-pulls with
  escalating delays (750/1500/3000ms) and detects finalization from the
  pulled bytes via a mandatory ftyp+moov container sniff, which also
  preserves the truncation detection the duration check provided by
  accident. The local waitForStableFile call is gone: a pull is complete
  when adb returns.
- toybox `ps -p <missing-pid>` exits 1 with empty output — the normal
  pid-gone signature — but the recovery liveness probe read every
  non-zero exit as an uncertain adb failure, making a finished recording
  behind a live-status manifest unrecoverable forever. Empty-output
  failures now corroborate via the full process list: healthy listing
  with the pid absent recovers the finished recording; listing failure
  stays conservatively uncertain (a stale verdict deletes the manifest,
  so transport health is proven first). Transport failures keep stderr
  and exec-layer timeouts throw, which is what makes the empty-output
  signature safe to trust.

Provider-scenario coverage: in-place same-size finalization landing past
the first retry, and dead-pid recovery on a responsive device — both
fail on the previous implementation with the live-observed errors.

* refactor: extract Android liveness probe and split recording scenarios per file-shape rule

record-trace-android-recovery.ts (763 lines) and android-recording.test.ts
(1,764 lines) were both past the 500-LOC extraction tripwire. The
screenrecord liveness/probe concept this PR modified now lives in
record-trace-android-liveness.ts, and the new scenarios moved to test
files mirroring the source modules they cover
(record-trace-android-copy.test.ts, record-trace-android-liveness.test.ts)
with shared scenario plumbing in the android-recording-fixtures.ts
sibling. android-recording.test.ts shrinks to 1,476 lines — below its
pre-PR size.
2026-08-06 17:38:10 +02:00
Michał Pierzchała d5f11f6e2f refactor: import package types directly — no internal re-export laundering (#1640)
* refactor: import package types directly instead of re-exporting from internal modules

Post-#1636 review feedback: internal src modules were re-exporting package
types (export type { X } from '@agent-device/...'), giving one declaration
several import paths and hiding its provenance. New rule applied repo-wide:
internal modules import directly from the owning package; only published
entry surfaces (src/sdk/* entries, client-types, finders, metro composition,
remote-config-schema) may re-export.

Eleven internal re-exports removed and ~110 import sites redirected to the
packages, the big two being CommandFlags (core/dispatch chain, 39 sites) and
SessionAction (daemon/types.ts, 19 sites). Two were already dead
(RefFrameEffect via daemon-command-registry, DiffSnapshotCommandResult via
capture/runtime/snapshot). Entry-surface chains now re-export from the
package rather than laundering through a second internal module
(client-types/client-metro MetroBridgeScope).

Side effect: the R9 type cycle shrinks again, 49 -> 47 (daemon-server
19 -> 17); ceilings lowered to match.

* refactor: drop command-schema's CliFlags re-export (#1640 review P2)

The one consumer (cli/parser/args.ts, a multi-line import the sweep's
single-line scan missed) now imports CliFlags from contracts/command;
FlagDefinition/FlagKey stay — they are src-declared types, not package
laundering.
2026-08-06 15:22:51 +02:00
Szymon Dziedzic 3e4828d68d feat: add scale-only screenshot sizing (#1617)
* feat: add scale-only screenshot sizing

* fix: refuse retired --max-size inputs on every released surface

Released sizing inputs must fail closed with migration guidance instead of
silently producing native-size artifacts:

- contracts: RETIRED_SCREENSHOT_MAX_SIZE declaration + SCREENSHOT_SCALE_LIMITS
  as the single source for the scale bounds and migration messages
- .ad parser: released 'screenshot ... --max-size N' and 'record start ...
  --max-size N' lines now refuse at parse time (frozen replay-compat witnesses)
- daemon: screenshot rejects old-client screenshotMaxSize like recording does;
  the recording guard now shares the same contract data
- Node client: screenshot/record daemon writers refuse the removed { maxSize }
  option before transport
- CLI: --max-size unknown-flag error carries the migration guidance
- config/env: stale screenshotMaxSize config keys and the retired
  AGENT_DEVICE_SCREENSHOT_MAX_SIZE env var are refused for sizing commands
  (other commands keep working)

Quality: numberField now reuses the canonical readOptionalNumber contract
helper (AppError bounds instead of plain Error); png-resize inlines one-use
wrappers and restores the worker-thread rationale; docs typo fixed.

* test: drop retired maxSize entries from the MCP undocumented-input allowlist

* fix: refuse retired maxSize at the MCP field-projection seam + release-provenance corpus witnesses

- readFieldInput silently dropped undeclared keys before the daemon writers
  could refuse them, so an MCP call carrying { maxSize } reached transport and
  returned native-size success. New retiredField() combinator declares the
  removed key in the field map: the projection seam refuses it with the
  canonical migration message and the JSON schema no longer advertises it.
  Real-route MCP executor regressions cover screenshot and record.
- replay-compat corpus: derived v0.20.5 witnesses for the released screenshot
  and record --max-size forms (SHA-256 pinned, new retired-capture-size
  coverage surface) so check:replay-compat proves the shipped syntax refuses
  with migration guidance instead of degrading silently.

---------

Co-authored-by: Michał Pierzchała <thymikee@gmail.com>
2026-08-06 15:11:00 +02:00
Michał Pierzchała ee473b6adc refactor(daemon): give the Maestro fallback and ambiguous-match details real types (#1612)
Three places smuggled structured data through untyped bags and re-read it
with runtime guards. Each gets an explicit typed boundary.

A. The resolution-suppression rule was encoded twice in
   interaction-touch-response.ts — a spread ternary in the runner-payload
   branch and an unconditional destructure used conditionally in the runtime
   branch, with the ADR 0012 rationale living on only one source variant.
   Both branches now read one `suppressesResolutionDisclosure(source)`
   predicate through one `applyResolutionDisclosurePolicy` helper, where the
   reason is stated once. The union field is renamed
   `maestroCoordinateFallbackDispatched` (the dispatch path that ran) and
   hoisted into a shared base. handleFillCommand's two-arm interactor.fill
   call collapses to one.

B. `Interactor.type` narrows from `Record<string, unknown> | void` to
   `TypeTextBackendResult | void`; the Apple runner boundary is the single
   place the wire payload becomes that type. `maestroFallbackDetails` returns
   a typed `{ used, extra }` instead of a bag both call sites re-read.

C. `details.candidates` meant two incompatible things. The device-domain
   resolvers now key their list `devices`, so the shared renderer drops its
   shape-disambiguation guards and the device list actually renders.
2026-08-05 14:32:25 +02:00
Michał Pierzchała 9fc266351f fix: stabilize Replay Nightly fixture boundaries (#1610)
* fix: stabilize replay nightly fixture boundaries

* test: stabilize exit flush integration coverage

* chore: drop the deleted exit-naive fixture from fallow's entry list (#1610 review P3)
2026-08-05 12:43:55 +02:00
Thiago Brezinski a13a6832ee feat: add selector-targeted drag gestures (#1567)
* feat: add selector-targeted drag gestures

* fix: address drag gesture review feedback

* fix: satisfy drag review quality gates

* fix(android): lower drag trajectories piecewise

* test(replay): validate drag fixture selectors

* fix(ios): ignore full-viewport chrome containers

* test(drag): prove destination on live devices
2026-08-05 12:37:02 +02:00
Michał Pierzchała 611858103e fix(ios): harden Bluesky-class interaction reliability (#1588)
* fix: type into focused iOS inputs without AX

* fix: fill AX-hostile iOS text inputs

* fix: keep scrolling containers from stealing taps

* fix: stop agents after explicit task success

* chore: format benchmark guidance

* fix(ios): preserve fill semantics across fast paths

* test: retire direct selector fill expectations

* test: assert runtime selector fill evidence

* fix(ios): preserve verified and Maestro fill paths

* refactor(ios): isolate synthesized text entry

* fix(client): preserve open diagnostic paths

* fix(ios): expose structured text entry route

* fix(packaging): strip text entry policy tests
2026-08-05 08:02:44 +02:00
Michał Pierzchała 65450dd3a7 fix: flush stdout/stderr before every CLI process.exit() (#1596) (#1603)
* fix: flush stdout/stderr before every CLI process.exit() (#1596)

Node only flushes process.stdout/stderr synchronously to a file or TTY;
on a pipe (the normal condition for this CLI when driven as a
subprocess) a write queued right before process.exit() can be silently
dropped. handleRunCliFailure's --debug daemon-log-tail dump made this
reachable from the exact path that renders a SESSION_NOT_FOUND error
right after a daemon replace, matching field reports of the driving
process going silent immediately after "Replacing daemon ... unreachable"
plus the SESSION_NOT_FOUND error.

Add exitAfterFlush() and route every process.exit() in src/cli.ts and
src/bin.ts through it, so a piped caller always receives the full
structured error (with its "run open first" hint) before the process
terminates. Also bound the --debug log-tail dump to a byte cap instead
of an unbounded 200 lines.

Verified directly against the real CLI (piped subprocess, pre-fix vs
post-fix): a live daemon with a seeded >64KB log truncates its --debug
error output before the fix and delivers it in full after.

* fix: satisfy CI gates on #1596 (format, fallow, coverage)

- oxfmt formatting on the new integration test file.
- Register the two exit-flush regression fixtures (support/exit-naive.ts,
  support/exit-after-flush.ts) as fallow entry points: they're run as real
  subprocesses via a string path (runCmdSync), which fallow's static
  dependency analysis can't follow, same as the existing
  test/contention-retry-fixtures/* entries. exit-payload.ts becomes
  reachable transitively through their static imports. Also switched the
  integration test's local PAYLOAD_MARKER duplicate to import the one
  fallow flagged as unused from exit-payload.ts.
- Added real unit coverage for the new exitAfterFlush code paths, since
  node --test integration files aren't measured by the vitest coverage
  gate: src/utils/__tests__/process-exit.test.ts exercises the
  already-drained, backlogged-then-drains, and never-drains/timeout
  branches directly against a fake stream; src/__tests__/cli-exit-paths.test.ts
  drives runCli() for --version, bare help, no-command, and web to cover
  their exitAfterFlush call sites, plus a --debug case with a >64KB seeded
  daemon.log proving printDaemonLogTailOnError's new byte cap actually
  trims the oldest lines.

Changed-line coverage gate now passes at 92.59% (was 59.26%); the two
remaining uncovered lines are the bottom-of-file `isDirectRun` catch
handler, which only runs when cli.ts is executed as the literal entry
script and is not reachable by importing it as a module in a test (the
same shape as bin.ts's already-excluded top-level fast paths).
2026-08-04 21:57:00 +02:00
Michał Pierzchała 4f8dc3f31e refactor: move selector engine into workspace package (#1589)
* refactor: move selector engine into workspace package

* refactor(selectors): trim the package façade to its real consumers

Follow-up to the selector-package cutover, from a structural review of it.

- Drop 15 façade symbols with no consumer anywhere in the repo:
  selectorUsesKey (added by the cutover, never called), isNodeVisible /
  isNodeEditable (the real helpers are contracts/snapshot's), normalizeText,
  splitIsSelectorArgs, IS_PREDICATE_REQUIRED_MESSAGE, four nested Replay
  types, SelectorDisambiguationDisclosure, and the four kernel type
  re-exports every consumer already imports from kernel directly.
- Delete SelectorCapturePolicyInput.selectorExpression, which
  deriveSelectorCapturePolicy never read; the policy varies only by
  predicate, so it takes one now. Two of the four tests asserted that the
  unread parameter had no effect and could not fail; they go with it.
- Return the Maestro export vocabulary to the maestro package. The cutover
  inlined MAESTRO_TEXT/STATE_SELECTOR_KEYS' values into the CLI call site,
  leaving both constants dead in the package that owns the concept and no
  gate over the two copies. MAESTRO_SELECTOR_PROJECTION is now the one
  statement of it.
- Dedupe SelectorDiagnostics and SelectorDisambiguationDisclosure, declared
  character-for-character twice across the AST/string seam, and name the two
  shared option shapes once instead of five inline copies. The parser-side
  resolution types take an Ast prefix so the twins read as twins.
- Delete three identity wrappers: parsePrivateSelector,
  selectorExpressionToMaestro, and the formatSelectorFailure forwarder —
  nothing passes it a chain any more, so the SelectorChain | string union
  and its branch go too.
- Delete internal/index.ts, an AST barrel whose only consumer was one test
  in the same directory (renamed to engine.test.ts), and the match.ts
  pass-through that existed to feed it.
- ReplaySelectorGrammar had three variants for two behaviors; 'wait' and
  'ordinary' were the same path. It is 'is' | 'positional' now.
- Drop the deleted src/sdk/selectors.ts from .fallowrc.json's entry list.

Behavior unchanged. pnpm check green: 598 unit files / 5278 tests, smoke
35 passed / 3 live skipped, layering 71/71, depgraph 22/22, mutation config
45/45, fallow clean, package smoke sound. Counterfactual: pointing
MAESTRO_SELECTOR_PROJECTION.textKeys at the state keys turns three
replay-maestro-export cells red; restored before commit.

* test(selectors): split the engine aggregation test by source concept

`internal/index.test.ts` (renamed `engine.test.ts` when its barrel went away)
was a 708-line aggregation over the whole engine — past the 500-line tripwire
and mirroring no source module, so it also ran as one serial unit.

It becomes five files that each mirror what they test, plus the parser cells
folded into the existing parse test:

  resolve.test.ts                 alternative fallback, strict uniqueness,
                                  first-match existence
  resolve-disambiguation.test.ts  ADR 0012 ranking: deepest, smallest-area,
                                  winner-vs-challenger disclosure, tie fallback
  resolve-viewport.test.ts        the visibility half: on-screen beats
                                  off-screen, including inside an off-screen
                                  scroll container
  match.test.ts                   per-key matching semantics (text, role,
                                  focused, appname/windowtitle, decoded
                                  newline labels)
  arguments.test.ts               where the selector ends and the command's
                                  positionals begin, both grammars
  parse.test.ts                   +6 grammar/escape cells beside the existing
                                  property tests

The login-form tree shared by resolve.test.ts and match.test.ts moves to
`__tests__/login-form-nodes.ts` rather than being copied into both.

All 27 cells are carried over unchanged and still pass; no file now exceeds
224 lines. pnpm check green: 602 unit files / 5278 tests, layering 71/71,
depgraph 22/22, mutation config 45/45, fallow clean over 127 changed files.

* revert(selectors): keep agent-device/selectors public, behind one AST subpath

The cutover removed the `agent-device/selectors` public subpath as part of
tightening the API. It is in use, so the removal is reverted: the subpath ships
the same ten symbols v0.20.5 shipped, with the same signatures.

That has to coexist with the reason the package façade is string-only, so the
AST leaves through one named door instead of the main one:

  @agent-device/selectors        string-in/string-out; every in-repo consumer
  @agent-device/selectors/ast    the published parser surface; one consumer,
                                 src/sdk/selectors.ts

`packages/selectors/src/ast.ts` re-exports parseSelectorChain,
tryParseSelectorChain, isSelectorToken, the AST-taking findSelectorChainMatch
and resolveSelectorChain, isNodeVisible, isNodeEditable, and types
SelectorChain / SelectorDiagnostics. `formatSelectorFailure` keeps its
published `SelectorChain | string` first parameter as a shim here rather than
widening internal/resolve.ts back to a union — the compatibility obligation
sits at the boundary that owes it.

This is strictly narrower than main, where the AST was reachable from anywhere
in src/ via src/selectors/*. Two gates hold it there: facade-symbols.ts pins
./ast to exactly the v0.20.5 list, and package-boundaries.test.ts asserts
src/sdk/selectors.ts is the only file outside the package that imports it.

Restored alongside: the ./selectors export and tsdown entry/chunk group, the
.fallowrc.json entry, the package-exports supported-subpath list, and both
client-api.md sections. No CHANGELOG entry — nothing is removed any more.

pnpm check green: 602 unit files / 5278 tests, smoke 35 passed / 3 live
skipped, layering 71/71 (10 packages, 32 subpaths), depgraph 22/22, mutation
config 45/45, fallow clean over 129 changed files, package smoke imported all
12 published entry points with publint and attw passing. Verified functionally
against the built dist: the doc's parse -> findSelectorChainMatch example
returns the same shapes as before, resolveSelectorChain still returns an AST
`selector`, and formatSelectorFailure still accepts a chain.

* fix(selectors): correct the two expectations that still assume the removal

Review P1s on a792415a: restoring the public subpath left two gates asserting
it was gone.

- installed-package-metro.test.ts moved `agent-device/selectors` into the
  blocked-specifier list. It goes back to the subpath smoke set, running the
  same `isSelectorToken('||')` + `parseSelectorChain` check it ran before the
  removal, so the file's only remaining delta from main is a formatter reflow.
- owner-files-no-leak.test.ts asserted `dist/src/sdk-selectors.js` was absent.
  It requires the stable named chunk again, and still rejects an auto-numbered
  `selectors2.js` fallback — the pair is what proves the restored tsdown chunk
  group is doing its job, verified against a clean build.

PR body corrected: the removal is no longer described as intentional API
tightening.

* refactor(selectors): satisfy the widened fallow scope after rebase

main's #1591 (the follow-up filed from this review) removed `packages/**` from
.fallowrc.json's ignorePatterns, so the new package is audited for the first
time. Everything below is a finding fallow could not previously see.

Dead surface, all confirmed consumer-free:

- 12 type re-exports from the `.` façade whose shapes consumers only ever
  reach structurally.
- MAESTRO_TEXT_SELECTOR_KEYS / MAESTRO_STATE_SELECTOR_KEYS, orphaned by this
  branch's own MAESTRO_SELECTOR_PROJECTION change, and the test-util
  SELECTOR_VALUE_HAZARDS. All three are module-local now.
- IS_PREDICATE_USAGE_HINT fails --production because its only consumer is the
  is-argument-surface parity test. It gets a commented `ignoreExports` entry
  rather than deletion: the constant is what makes the daemon and CLI raise
  ONE hint instead of two copied strings (ADR 0010), so the test asserting
  that is the point, not an accident.

`fast-check` is now declared by the package that imports it.

Duplication, split by what could be proven:

- `isUsefulVisibilityAnchor` existed character-for-character in both
  packages/selectors and packages/maestro. Moved to
  @agent-device/contracts/snapshot, which both already depend on and which
  already owns this vocabulary. Safe because the `normalizeType` each copy
  called is itself character-identical to the contracts one — checked before
  moving, since a different normalizer would have silently changed which
  nodes anchor.
- maestro additionally reimplemented `normalizeType`, `buildSnapshotNodeMap`
  (as `buildSnapshotNodeByIndex`) and `findSnapshotAncestor`, all
  character-identical to contracts'. Deleted in favour of the shared ones.
- The three scroll-ancestor walks are NOT deduped. They are structurally the
  same walk but each uses a different scrollable predicate, and I have no
  evidence the three agree; collapsing them would be a Maestro-conformance
  change, not a cleanup. Both maestro sites now say so, and the work is filed
  separately.

`projectSelectorExpression` (15 cyclomatic / 22 cognitive, written by the
cutover) splits into a dispatcher plus `readAgreedTextValue` and
`projectSelectorTerms`; all three are under threshold.

Rebase note: the one conflict, in package-boundaries.test.ts, resolved to
NEITHER side — #1591 had already deleted `AdReplayVerifiedTargetGuard` as an
unused export, and this branch deletes the seven ReplaySelectorPort names, so
the conflicting block is empty.

* build: record fast-check for packages/selectors in the lockfile

Declaring the dependency in packages/selectors/package.json without
regenerating pnpm-lock.yaml made every CI job fail in its install step with
ERR_PNPM_OUTDATED_LOCKFILE. My local `pnpm install --frozen-lockfile` printed
"+ 1 dependencies were added: fast-check@^4.9.0" and exited 0, which read as
success but was the same mismatch CI refuses.

Regenerated with the pinned pnpm 11.17.0, not the 11.5.3 on this machine:
11.5.3 rewrites peer-dependency resolution keys repo-wide (dropping
`(supports-color@7.2.0)` suffixes) and produced a 222-line diff. With the
pinned version the diff is the 4 lines this change actually needs, plus
pnpm's alphabetical re-sort of the root selectors entry.
2026-08-04 19:06:39 +02:00
Michał Pierzchała 2e74b789fd feat: verify device cloud connections (#1564)
* feat: verify device cloud connections

* refactor: unify connect provider adapters

* refactor: separate connect verification facts

* fix: tighten connect provider verification

* fix: use neutral cloud connection wording

* perf: deduplicate local affected checks

* refactor: simplify affected check runner

* refactor: derive connect workflow from verification
2026-08-03 16:47:57 +02:00
Michał Pierzchała 99967c7f01 fix: restrict project config trust (#1565)
* fix: restrict project config trust

* fix: preserve daemon auth transport context

* refactor: simplify project config trust

* fix: restrict project config write sinks
2026-08-03 15:48:45 +02:00
Michał Pierzchała f8617a2db9 fix: read Android get text from target field (#1561) 2026-08-03 12:09:31 +02:00
Michał Pierzchała 480e3883b1 fix(daemon): reject unarmed close --save-script before teardown (#1558)
* fix(daemon): reject unarmed close --save-script before teardown

Live evidence (2026-08-02) showed a plain `open` followed by `close
--save-script` silently published a script: the close request armed
authoring at record time and published moments later in the same
request, folding the never-armed case into the ADR 0016 authoring
lifecycle. The resulting .ad carries selector fallback chains but no
recording-time target-v1 evidence, and nothing told the caller
evidence capture never ran — degraded replay verification with no
signal beats a loud refusal.

`assertTerminalRecordingCloseAllowed` (src/daemon/handlers/session-close.ts)
now rejects an unarmed `close --save-script` with INVALID_ARGS before
any teardown or filesystem work runs, the same seam that already
rejected ABORTED/PUBLISHED terminal recordings. The rejection does not
tear the session down, so a plain `close` retry still completes
cleanly; recovery names `open --save-script` since evidence can only
be captured from action zero. Repair transactions (ADR 0012) are a
disjoint lifecycle and are explicitly unaffected.

This is distinct from #1533 (an already-armed-then-aborted session
whose flag ingress re-enables recordSession and lets a *bare* close
publish); that case remains open.

* fix: review follow-ups for #1558 (help text, test strength, docs)

- Give replay --save-script its own help text instead of the shared
  open/close "arm on open, publish on close" description: replay's flag
  arms an ADR 0012 repair transaction, a disjoint lifecycle. Adds
  CommandSchema.flagDescriptionOverrides so a command can swap a shared
  flag's usageDescription without duplicating the FlagDefinition entry
  (which would have shown --save-script twice in `help replay`). Pinned
  in src/cli/parser/__tests__/cli-help-command-usage.test.ts (open/close
  keep the shared text unchanged; replay gets the new one).

- Strengthen the never-armed close --save-script regression test in
  session-close-shutdown.test.ts: the fixture now carries real
  cleanup-bearing state (an active iOS simulator recording, reusing
  makeIosSimulatorRecordingSession/recordingKillMock) with spies proving
  no teardown hook (recorder kill, runner stop) runs on the rejected
  request, then that a follow-up plain close does tear it down. The
  prior fixture had nothing for teardown to observably touch, so moving
  the guard after stopBestEffortSessionResources would have passed it
  silently. Also fixes a latent test-isolation leak this exposed: an
  earlier test set a persistent mockStopIosRunnerSession rejection
  (vi.clearAllMocks() clears call history, not implementations), which
  would have poisoned any later Apple-platform close test; scoped it to
  mockRejectedValueOnce.

- Point the migration guide (website/docs/docs/migrating-gestures.md) at
  `open --save-script` → interact → `close` instead of the now-rejected
  `open` → interact → `close --save-script`, matching the new guard and
  the corrected help text.

_Generated by [Claude Code](https://claude.ai/code)_
2026-08-03 09:10:40 +02:00
Michał Pierzchała 2c2df031ff feat: keep replay session active on request (#1554)
* feat: keep replay session active on request

* test: cover replay keep-session provider route

* fix: make replay session handoff reliable

* refactor(daemon): extract the replay terminal-lifecycle policy module (#1554 review)

session-replay-runtime.ts was already over the 500-line extract-before-adding-behavior
tripwire before this PR; the keep-session/repair terminal-close decision, its
live-session postcondition, and the dispatched-action count pushed it further past
budget. Move that policy into a focused session-replay-terminal-lifecycle.ts
(isExecutableReplayAction, resolveSuppressedTerminalCloseIndex,
countExecutedReplayActions, requireLiveSessionForKeepSession) so the runtime file
stays orchestration-only, and mirror its PR-added unit tests into
session-replay-terminal-lifecycle.test.ts. Pure extraction: no assertions changed.
2026-08-02 18:35:13 +02:00
Michał Pierzchała e5cebcd8e3 fix(test): tolerate dead session in iOS e2e full-tier cleanup (#1548)
* fix(test): tolerate dead session in iOS e2e cleanup

full:device-lifecycle reboots the simulator, and whether the daemon
session survives that is environment-sensitive: it does on CI but not
locally, so every all-green local full-tier run ended red in cleanup
with all three retries of each step failing ('permission setting
requires an active app in session' / 'No active session').

Two layers:
- finalizeLiveRun re-checks sessionExists instead of short-circuiting
  on sessionOpen, so a session that died mid-run skips cleanup entirely.
- cleanupSession treats SESSION_NOT_FOUND and the appless-session
  INVALID_ARGS failure as already-clean instead of burning retries.

Other cleanup failures still exhaust three attempts and fail loudly.

* fix(test): narrow appless-session cleanup guard to the mic-permission step

Scope sessionAlreadyClean's INVALID_ARGS tolerance to the microphone-
permission reset step and its exact known message instead of matching
any cleanup step whose message contains "requires an active app in
session" — that substring is also thrown by the unrelated location
setting, so the old check could have hidden a real failure there.

Extract the per-step retry policy into an exported retryCleanupStep so
it's unit-testable without spawning the CLI, and add a deterministic
regression (test/integration/ios-simulator-e2e-cleanup.test.ts)
covering: a dead session (SESSION_NOT_FOUND) stops retrying on any
step, the known mic-permission appless response stops retrying, and a
different INVALID_ARGS (wrong step or wrong message) still exhausts
all three retries and fails.

* test(ios-e2e): cover the finalization cleanup-gate decision

Extract the sessionOpen re-check + conditional cleanupSession call out
of finalizeLiveRun into an exported finalizeSessionCleanup(context,
runSessionExists, runCleanupSession) — same behavior, now driven by
injected callbacks instead of the module-level runStep-backed
bindings, so it's unit-testable without spawning the CLI.

Add two deterministic cases: sessionOpen starts true and the final
sessionExists() resolves false -> cleanupSession is never invoked;
sessionOpen true and sessionExists() resolves true -> cleanupSession
runs (the live path). Counterfactual (reverting the recheck to the old
`sessionOpen || sessionExists(...)` form) turns the first case red as
expected.
2026-08-02 11:36:25 +02:00
Karthik Varma 92b22229e6 feat(cloud-webdriver): BrowserStack device-feature capabilities, and fix cloud orientation (#1544)
* feat(cloud-webdriver): support BrowserStack device-feature capabilities

Adds the eight BrowserStack "device feature" session capabilities that had no
representation in agent-device: deviceOrientation, geoLocation, timezone,
language, locale, networkProfile, customNetwork, and resignApp.

These are vendor capabilities, so they are emitted inside `bstack:options`
rather than at the top level. BrowserStack's YAML config lists them unnested
and its SDK relocates them; agent-device talks to the hub directly, so it
nests them itself.

A single spec table drives both the flag reader and the capability builder, so
adding a capability is a table row rather than a branch in each. A structural
test asserts every field owns exactly one row, since a field the table forgets
would parse off the CLI, ride the profile, and then be silently dropped before
the hub ever saw it.

Rejects combinations the provider cannot act on unambiguously: an unknown
orientation is caught at the flag boundary instead of being forwarded to a hub
that accepts and then ignores it, --provider-no-resign-app is refused on
Android, and a named network profile cannot be combined with a custom network
shape.

Also fixes a latent shallow-merge bug in buildBrowserStackCapabilities: a
caller supplying its own `bstack:options` replaced the whole object and
silently dropped the project, build, and session labels. It is now merged
per key.

* fix(cloud-webdriver): rotate via WebDriver orientation endpoints

`setOrientation` on the cloud WebDriver path sent `mobile: rotate`, which is
not a driver command at all. UiAutomator2's own error enumerates its
extensions and `rotate` is absent from the list, so `agent-device orientation`
was a hard failure on every hosted provider.

It also forwarded agent-device's four-way rotation vocabulary verbatim
("landscape-left", "portrait-upside-down"), where the protocol accepts only
uppercase PORTRAIT/LANDSCAPE. Every other platform has a translation layer;
this path was the only one without one.

Now two transports, ordered by backend. `POST /rotation` takes exact four-way
degrees and leads on Android, since it is the only endpoint that can express
upside-down and left-versus-right. `POST /orientation` is two-way and leads on
XCUITest, which rejects `/rotation`. Each falls back to the other, because only
BrowserStack's UiAutomator2 is verified and a provider whose driver disagrees
should degrade rather than hard-fail.

Verified live against BrowserStack App Automate:
  POST /rotation {"x":0,"y":0,"z":0} -> 200 {"value":"ROTATION_0"}

The rotation-to-surface-index mapping moves to contracts/device-rotation.ts and
the existing adb path now reads from it, so the local and hosted mappings
cannot drift apart.

Note this rotates the current display, not persistent device rotation, so an
activity that does not pin its own orientation may still need rotating once it
is in the foreground.

The capability was declared "partial" without the transport existing, and no
test covered setOrientation on the cloud path; only adb and the Apple runner
were covered. Both gaps are now closed.

* fix(cloud-webdriver): narrow orientation fallback and gate provider-owned flags

Addresses review on #1544.

The orientation fallback caught every error, so a timeout, an auth rejection, a
dead session or a provider 5xx on the first transport was swallowed and retried
against the second. When that one also failed the caller got "rejected both
endpoints" with the real cause discarded. Fallback is now keyed on structured
unsupported-endpoint signals only — HTTP 404/405, or a W3C `unknown command` /
`unknown method` code — matching the repo rule of keying on typed details rather
than message text. Everything else rethrows unchanged.

Device-feature capabilities are BrowserStack-owned, but the flags were accepted
by any cloud provider, persisted into the generated profile, and then silently
dropped at session creation. `connect aws-device-farm` now rejects them with a
typed error naming each offending flag, raised before the provider's own
required-argument checks so the caller is told what is unsupported rather than
what else is missing. Ownership is modelled on the capability spec table, so a
new capability inherits the guard without a second list to maintain.

Adds provider-backed orientation scenarios driven through public daemon dispatch
against the fake WebDriver provider: the four-way endpoint on the happy path,
the documented collapse onto the two-way endpoint when the driver does not
implement `/rotation`, and a provider 5xx that must surface without consulting
the second transport. The fake server's route handling became a table in the
process — it had grown to ten branches in one function.

* fix(cloud-webdriver): read W3C error codes before status, enforce ownership at the runtime boundary

Addresses the second review pass on #1544.

The fallback classifier returned on any 404/405 before consulting the W3C error
code, so an HTTP 404 carrying `invalid session id` was masked as a missing route
and retried against the second transport. The structured code now takes
precedence whenever the driver sent one; bare status is consulted only when no
code exists. Two cases pin it: a 404 `invalid session id` and a 405 `timeout`
must both surface rather than fall through.

Provider ownership was enforced only in the CLI profile builder, which the typed
client and hand-authored remote-config profiles bypass entirely — both reach
session preparation without passing through `connect`, so the capabilities were
accepted and then dropped. The check now lives on the capability-ownership
module and runs inside AWS Device Farm's `prepareSession`, with the CLI builder
calling the same helper instead of its own copy. Covered by a scenario that
drives the runtime boundary directly and asserts the rejection happens before
any provider session is created.
2026-08-02 08:17:49 +02:00
Michał Pierzchała b9509fe006 refactor: extract the .ad script codec into packages/ad-script (#1478) (#1536)
* refactor: extract the .ad script codec into packages/ad-script

Moves the mutually-coupled .ad read/write codec (script.ts, script-utils.ts,
script-formatting.ts, open-script.ts) plus the target-v1 annotation SERDE
slice of target-identity.ts into a new private leaf package,
@agent-device/ad-script, exporting only `.`. This is option 1 from the P5
scoping dossier on #1478: the codec is shared by the daemon's session-script
publication writer, the future replay engine, the CLI's `replay export`, and
Maestro's failure-label formatting, so it can no longer live in root src/
once packages/ad-replay lands (R11 forbids a package reaching into root src),
and a second export subpath or writer-half duplication are both ruled out by
existing gates/tests.

target-identity.ts keeps only the record/replay-shared classification core
(classifyTargetBindingMatch, local-identity/ancestry-prefix matching),
importing its shared types from the new package. Every real consumer
(re-derived by grep, not the dossier's list alone) is rewired to
@agent-device/ad-script.

Refs #1478

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor: trim the ad-script façade to real consumers, lock the one-export boundary

- packages/ad-script/src/index.ts: drop parseReplaySeriesFlags,
  formatTargetAnnotationCommentLine, parseTargetAnnotationCommentLine,
  TargetAnnotationLineParseResult, and TargetRect from the public façade —
  none has a consumer outside the package (re-swept every remaining export
  by grep; everything else kept has at least one real external importer).
  The functions/types stay exported from their declaring internal modules
  for the package's own internal use (script.ts, script-formatting.ts).
- scripts/layering/package-boundaries.test.ts: add the parallel R11
  assertions "the real tree parses, declares, and passes R11" already makes
  for maestro/provider-webdriver/provider-limrun/xml — ad-script exports
  exactly `.`, depends on exactly contracts+kernel, and is declared in root
  package.json — plus ad-script entries in the deep-resolution rejection
  coverage. Verified the lock catches a regression: temporarily added a
  fake `./codec` export to packages/ad-script/package.json and confirmed
  both the export-key-list assertion and the deep-resolution-rejection
  assertion fail; removed the plant and reconfirmed green.

Refs #1478

Co-Authored-By: Claude <noreply@anthropic.com>

* fix: remove polynomial-redos ambiguity from the target-v1 annotation line regex

CodeQL js/polynomial-redos flagged TARGET_ANNOTATION_LINE_RE
(packages/ad-script/src/internal/target-annotation-serde.ts): the payload
group's `\s+(.*)` let `\s+` and the unconstrained `.*` both match whitespace,
so a run of separator whitespace that ultimately fails to complete the match
has many `\s+`/`.*` splits to backtrack through before concluding failure.

Anchor the payload group on `\S` (the exact complement of `\s`), so the
mandatory `\s+` separator and the payload's first character can never
overlap — the split point becomes unique and no backtracking is possible.

Behavior-preserving: the only caller (parseTargetAnnotationCommentLine)
always matches against an already-.trim()-ed line, whose last character
(whenever the tag matches at all) is never whitespace — so a payload section
`\S.*` would reject (content that is entirely whitespace) can never reach
this regex through the real call path. Verified against the frozen
replay-compat corpus and the full serde/parser test suites, unmodified.

Added a regression test with the exact adversarial shape CodeQL/the reviewer
cited (many tab pairs after the version digits), asserting sub-second parse.

Refs #1478

Co-Authored-By: Claude <noreply@anthropic.com>

* test(ad-script): pin the annotation-line pattern's linear rejection directly

The entry-point adversarial case matched greedily even with the retired
regex (trim strips edge whitespace and per-line input carries no newline),
so it proved nothing about the pattern. The regression surface is the
pattern itself: an interior tab run with an x-newline tail fails the match,
which the retired form re-split quadratically (3.7s at 100k tabs) and the
\S anchor rejects in one attempt.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-01 20:21:38 +02:00
Michał Pierzchała 8a6ddbc11d fix(test): repair Android replay fixtures against live device reality (#1538)
* fix(test): repair Android replay fixtures against live device reality

Three Android fixture defects from the #1482/#1484 full-tier suite, none of
which ever executed in CI (both nightlies since failed on adb infra before
the suite ran). All three verified live on a fresh API 36 emulator with a
pixel_7-geometry AVD and a Release fixture APK:

- 01-navigation-scroll.ad clicked label=Catalog, but the expo-router
  NativeTabs cart badge leaks '0 new notifications' into the tab's content
  description even while hidden, and unselected native tabs expose no child
  text node - exact match can never hit. Target the composed label the
  device actually exposes (deterministic at fixture start: cart is 0 after
  --relaunch). The badge does not leak on iOS, so the iOS twin keeps
  label="Catalog".

- checkout-form-android.ad opened by iOS display name 'Agent Device
  Tester'; Android open resolves packages (the APK label is
  'Agentdevicelab'), so APP_NOT_INSTALLED was guaranteed. Use the package
  id, matching gesture-lab-android.ad.

- gesture-lab-android.ad aimed every gesture at y=700, above the gesture
  card (its targets span y754-1329 on pixel_7 geometry; the home screen
  gained content above the card since authoring). Re-aim pans inside the
  exact-two-pointer zone, flings on the image clear of that zone, and
  pinch/rotate/transform at the card center. Verified: full suite passes
  2/2 via the public test command (20 + 32 steps replayed).

Refs #1478

* docs(test): pin the Android gesture fixture's validated emulator geometry

The re-aimed coordinates are validated on CI's profile (pixel_7 1080x2400
@420); any booted emulator can receive them via test-app:replay:android, so
the fixture and README now say which geometry the numbers mean and what a
mismatch failure looks like. The checkout twin is selector-driven and
unconstrained.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-01 13:50:48 +02:00
Michał Pierzchała 67f3d09d95 refactor(daemon): session script publication behind one capability (#1478 P4a) (#1532)
* refactor(daemon): add the tagged script-publication aggregate

First step of P4a. Nine co-resident optional SessionState fields encode two
lifecycles plus a shared output target, with nothing in the shape saying the
lifecycles are disjoint — so readers re-derived that from field combinations
and writers had to remember which siblings to clear.

The aggregate makes both invariants structural: a session publishes nothing,
authors ordinarily, or is under repair; and force lives inside the target, so
retargeting replaces the authorization along with the path.

Three corrections after an adversarial review of the first draft:

- The target is a default|explicit union, not a mandatory path. A bare
  'open --save-script' arms with no path and lets the writer resolve a
  daemon-owned destination at write time, and force can be granted before any
  path exists. Eagerly materializing a default path would have silently changed
  retarget semantics, because today's check requires a previously persisted
  path — so 'open --save-script --force' then 'close --save-script=out.ad' is
  not currently a retarget and the grant survives. That behavior is preserved
  here and flagged in the docblock as a probable #1258 gap; tightening it is a
  product change and belongs in its own commit.

- The repair status relation is not linear. A failed commit followed by
  'replay --from' demotes complete back to armed, so demoteRepairToArmed exists
  and deliberately RETAINS the close receipt: the platform close already
  succeeded for that operation identity, and dropping it would re-dispatch a
  close on retry — which is also how a migrator ends up reaching for the
  caller-computed platformCloseSucceeded boolean the brief forbids.

- The receipt doc no longer claims it is set only at close-succeeded and later,
  since the demotion path makes {armed, receipt set} reachable.

Still to come in this PR: both projections, and the writer migration. Note the
brief's seven-file writer inventory omits session-open.ts, which holds the only
two writers of the authoring armed/aborted states.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(daemon): migrate script publication onto the tagged aggregate (#1478 P4a)

The eight co-resident SessionState fields (scriptRecordingState, saveScriptPath,
saveScriptForce, saveScriptBoundary, saveScriptComplete, saveScriptCommitted,
repairPlatformCloseReceipt, repairSourcePath) are gone; SessionState.scriptPublication
holds the aggregate, and every writer migrated in this commit — no shadow state.

Two daemon-private projections own the writes, enforced by the R7 ownership gate:

- session-replay-transaction.ts (ReplaySessionTransaction): repair arm/demote/
  complete/abort, close receipts, and the uncommitted/boundary/sourcePath reads
  that idle-reap, tombstones, divergence-hold, and the recorder's exclusion key off.
- session-script-publication-capability.ts (SessionScriptPublication): authoring
  arm on open, the recorded --save-script flag ingress, active publication, the
  published transitions, and the effective per-target force decision (#1258).
  The writer keeps the commit transition so idempotence stays colocated with the
  atomic publish.

Failure/retry transitions pinned as the brief requires: platform-close failure
leaves state unchanged (no receipt, retry re-dispatches); publication failure
retains target+force+receipt (same-identity retry skips close dispatch); committed
and aborted are explicit terminal states that drop the receipt.

Design decisions resolved:

- Force retention across a default->explicit retarget is preserved as-is and
  still flagged in resolveScriptTarget's docblock as a probable #1258 gap;
  tightening it stays a separate product change.
- The never-armed 'close --save-script' whole-log publication folds into the
  authoring lifecycle (armed at the recorded close, published in the same
  request) instead of a fourth variant: every close path that reaches the write
  deletes the session, so the transient armed state cannot leak into
  'session save-script' eligibility, whose not-armed-before-this-journey
  rejection is untouched.

One real bug caught by the migrated tests and fixed in resolveScriptTarget: a
bare (pathless) re-arm collapsed an already-materialized explicit target back to
the daemon default, wiping the healed-sibling path on every per-step repair
re-arm and defeating the persisted-force preflight bypass. A bare re-arm now
keeps the previous target and only adds a live force grant.

R7 rows consolidated to one scriptPublication entry (three owners) and the
recordSession row narrowed; the R10 baseline drops to 22 writer-owned fields /
28 owner claims so the consolidation cannot regrow.

Gates: typecheck, lint, format, layering clean; 624 files / 5220 tests pass
(two known contention-flake timeouts reproduce only under full-suite load and
pass in isolation).

Refs #1478

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(daemon): satisfy the Fallow gate by extracting decisions, not suppressing

- scriptPublicationTarget is module-private; both public target reads
  (scriptTargetPath/scriptTargetForce) go through it and nothing else did.
- validatePublicationEligibility splits into a pure ineligibility classifier
  and an error table, so the four rejections read as one decision each.
- prepareSaveScriptSession hands its two arm-time rejections (authoring
  re-arm, EEXIST preflight) to rejectSaveScriptArming and keeps only the
  demote-and-arm flow.
- The repair-record-exclusion provider scenario extracts its three phases
  (arm-and-hold, exclusion contrast, healed-script contract) into named
  helpers; the test body is the journey again.

Refs #1478

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-01 11:10:09 +02:00
Michał Pierzchała cbe1a57094 refactor(replay-test): extract packages/replay-test (#1478 P3b) (#1525)
* refactor(replay-test): source the manifest device vocabulary from the kernel

`session-test-types.ts` reached `ReplayScriptMetadata['platform']` and
`['target']` through `replay/script.ts` — the native `.ad` engine. A
format-neutral scheduler must not name an engine module, and P5 relocates that
engine into `packages/ad-replay` regardless, so the import had to go before the
scheduler can move.

Both members already resolve to neutral kernel types
(`Exclude<PlatformSelector, 'web'>` and `DeviceTarget` from
`@agent-device/kernel/device`), so this re-sources them directly and the
manifest shape is unchanged. Only the import direction differs.

First increment of P3b; the scheduler still has request-global, engine and
daemon imports to port before the physical move.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(replay-test): inject the progress sink instead of reading a request global

The scheduler called `emitRequestProgress` in eight places, which reads a sink
out of a request-global `AsyncLocalStorage`. That is ambient authority a
format-neutral scheduler cannot hold once it lives in `packages/replay-test`,
and #1505 recorded it as a shrink-only R10 entry.

The host now injects the capability through the existing
`ReplayTestRuntimeDependencies` seam established in P3a, so no new seam is
invented. `session-replay.ts` supplies `emitProgress: emitRequestProgress`;
`src/request/progress.ts` keeps the sink and its AsyncLocalStorage binding for
every other caller.

The port is deliberately narrower than `RequestProgressSink`: it accepts only
`ReplayTestSuiteProgressEvent | ReplayTestProgressEvent`, so the scheduler is
not handed the ability to emit `CommandProgressEvent`.

Authority narrows again one hop down: `runReplayTestAttempt` spread the whole
dependency bag but uses three of its members and never publishes progress, so
it now takes `Pick<..., 'runReplay' | 'cleanupSession' | 'finalizeAttempt'>`.
That is why no runtime test fixture needed changing — the attempt runtime never
gained the capability in the first place.

Also drops the last two `replay/script.ts` type references from
`session-test-runtime.ts`, so the engine import is gone from that file too.

Reporter contract preserved: `session-test-reporter-values.test.ts` and
`session-test-reporter-values-maestro.test.ts` both pass unmodified (27 tests
green across the five scheduler suites). Typecheck clean.

Remaining scheduler boundary for P3b: `request/cancel.ts`, `replay/format.ts`,
`replay/script.ts` in discovery, `session-store.ts`, `daemon/types.ts`,
`replay-source-discovery.ts`, `core/dispatch*`, `utils/diagnostics.ts`.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(replay-test): ask the host whether the suite is canceled

The scheduler called `isRequestCanceled(requestId)` in five places. That both
reaches a request-global registry and forces the scheduler to name a daemon
request id as the cancellation key — neither survives the move into
`packages/replay-test`.

The host now binds the predicate to its own request and passes
`isCanceled: () => boolean`. The scheduler asks a question it is entitled to
ask and learns nothing about how cancellation is tracked. `shouldStopReplayTestExecution`
takes the capability rather than a request id, so no scheduler function threads
a daemon identifier for this purpose any more.

`session-test-attempt.ts` and `session-test.ts` no longer import
`request/cancel.ts` at all. It remains in `session-test-runtime.ts`, which does
something different — `registerRequestAbort`, `markRequestCanceled` and the
parent-abort relay are cancellation *binding*, which the brief assigns to the
daemon adapter, so that split is its own step.

Behavior preserved: both pinned reporter characterizations pass unmodified,
32/33 across the five scheduler suites. The one failure is the pre-existing
P2/#1506 discovery-ordering regression, unrelated and untouched here.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(replay-test): drop the dead request-tracking call from attempt ids

`buildReplayTestAttemptRequestId` wrapped its template in
`resolveRequestTrackingId`, pulling `request/cancel.ts` into the scheduler.

That wrapper substitutes a generated id only when its first argument is an
empty string. The template here always contains `:test:`, so it is never empty
and the wrapper always returned it unchanged — the call is unreachable in this
path. Probed all three shapes (explicit request id, suite-id fallback with a
shard, and degenerate empty inputs); every one returns the template verbatim.

Removing it takes `request/cancel.ts` out of discovery without altering a
single produced id. The scheduler mints attempt identity itself, which is what
the brief asks for.

Evidence the ids are byte-identical: the pinned reporter characterizations
assert exact session strings such as
`default:test:suite-reporter:1-02-retry:attempt-1` and pass unmodified —
30 tests green across the reporter, suite and discovery suites.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(replay-test): move cancellation binding and diagnostics to the host

`session-test-runtime.ts` held the last two request-globals in the scheduler:
`request/cancel.ts` (registerRequestAbort, markRequestCanceled,
clearRequestCanceled, plus the parent-abort relay) and `utils/diagnostics.ts`.

These are different in kind from the earlier ports. The brief gives the daemon
adapter the job of mapping an attempt id to daemon request identifiers and
*binding cancellation*, while timeout policy stays scheduler-owned. So the
scheduler now receives a per-attempt capability with exactly two verbs —
`cancel()` on timeout and `release()` when the attempt settles — and every
registry interaction, including `relayReplayTestAbortFromParent`, moved to
`session-replay.ts` next to the rest of the adapter.

Diagnostics became a narrow publish capability for the same reason:
`emitDiagnostic` reads a request-global scope. The level vocabulary is spelled
out at the seam rather than imported, so nothing engine- or daemon-shaped
crosses it.

The runtime fixtures drive the real exported host binding rather than a stub.
They assert cancellation through `isRequestCanceled`, and a stubbed binding
would have kept those assertions passing while proving nothing.

24 tests green across the runtime, suite and both reporter characterizations,
which pass unmodified. Typecheck, lint and oxfmt clean.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(replay-test): split discovery into host inspection and scheduler policy

discoverReplayTestEntries expanded paths, read every file, and called both
engines — readReplayScriptMetadata for .ad, inspectMaestroFlow for Maestro —
plus resolveReplayFormat to choose between them. Four imports a format-neutral
scheduler cannot hold.

Inspection is now the host's discoverSources capability. What stays in the
scheduler is the genuinely neutral half: which sources a --platform filter
runs, which it skips and with what message, and the empty-suite error.

The manifest carries exactly the four fields the scheduler consumes (platform,
target, retries, timeoutMs) plus the reporter's title, per the brief's
instruction not to add more without a demonstrated call site.

The platform tag is what removes the last format leak. The filter used to ask
resolveReplayFormat(...) === 'maestro' to decide whether a missing platform was
disqualifying. It now reads a tag: caller-bound means the invocation supplies
the platform, unspecified means the source declared none. Maestro is what
caller-bound looks like from the scheduler's side, and the format cannot be
recovered from it.

Discovery tests drive the real inspection capability, writing actual .ad and
Maestro sources — a stubbed host half would have kept them green while proving
nothing about the composition they exist to pin.

35 tests green across discovery, suite, runtime and both reporter
characterizations, which pass unmodified. The Maestro one is the direct check
that titles still flow, since they now arrive via the manifest.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(replay-test): build attempt ids from named segments; trim comments

Review feedback on the attempt-id builder: the comment explained a deletion
that git already records, and it sat above an opaque template literal.

The id is now a segment list joined on ':', so its shape is readable without
prose. Output is byte-identical — the reporter characterizations assert exact
session and attempt strings and pass unmodified.

Applied the same standard to four other docblocks in this PR that narrated
what the code used to do rather than what it does. The durable 'why' stays:
which side of the seam owns what, and why the vocabulary is neutral. The
migration history goes, since git carries it and these docblocks will outlive
the migration.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(replay-test): move shard device binding to the host

buildReplayTestShardPlan called listDeviceInventory to discover what to shard
across, and buildReplayTestShardFlags constructed daemon CommandFlags for the
nested request. Inventory enumeration, allowlists, simulator set paths,
explicit --device selectors and the too-few-devices error are host concerns;
what is scheduler-owned is deciding how many shards exist and which entries
each one runs.

The scheduler now receives resolved shard targets through a capability. The
target is neutral: id and name for session labels and progress metadata, plus
platform and target, which are already kernel vocabulary. DeviceInfo no longer
crosses into scheduling.

One behavior note: an explicit --device selector could in principle name a web
target, which is not a shardable device. That is now rejected with INVALID_ARGS
rather than widening the neutral platform vocabulary to carry something the
scheduler can never run. Implicit selection already filtered to mobile.

919 of 920 handler tests pass. The one failure, session-test-runner.test.ts
'binds each replay script to its declared platform metadata', fails identically
on clean origin/main in this container and is unrelated: directory discovery
walks with opendirSync/readSync and directory results are deduped but not
sorted, while glob results are sorted, so suite order is filesystem-dependent.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(replay-test): extract packages/replay-test behind a façade

Completes the P3b extraction. The scheduler, attempt runtime, discovery
policy, sharding distribution, artifacts and neutral types now live in
packages/replay-test/src/internal/, with one package-root export.

The façade takes a neutral ReplayTestSuiteRequest and returns a tagged
ReplayTestSuiteOutcome. DaemonRequest, DaemonResponse and CommandFlags no
longer reach the scheduler; the adapter translates flags and meta in, and the
outcome back to a daemon response. Eight flags were read by the scheduler and
each became a field it owns.

Host work moved to daemon adapters: source inspection (both engines and format
routing), shard device binding and shard-flag parsing, and artifacts-dir home
expansion, which is why the package can resolve paths without SessionStore.

The one remaining shared concern was the timing trace: the host writes video
lifecycle events into the same trace the scheduler owns. Rather than export a
writer from the façade, each attempt hands the host an appendTimingEvent
closure, so the trace format stays private and the authority is scoped to that
attempt.

Tests mirror the topology. Discovery tests split along the seam they now
cross: ordering, traversal and routing are pinned host-side against real files,
filtering policy is pinned in the package against fake sources. The runtime
tests assert the scheduler's cancellation obligation (cancel once on timeout,
always release) against a recording binding, and a new daemon test pins the
adapter's half — registry entries, the parent-abort relay, and detach on
release — so that coverage moved rather than disappeared.

R10 retargeted to packages/replay-test/src/ and the zone ranked alongside
maestro. R11 confirms zero root-src imports from the package.

914 of 915 handler and package tests pass. The one failure,
session-test-runner 'binds each replay script to its declared platform
metadata', fails identically on clean main here: directory discovery walks with
opendirSync and dedupes without sorting, while globs sort, so suite order is
filesystem-dependent.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(daemon): simplify replay-test request translation

Fallow flagged toReplayTestSuiteRequest at 14 cyclomatic in 18 lines. The
branches were self-inflicted: every req.flags?.x is one, and each optional
field was written as a conditional spread to avoid setting an undefined key.

exactOptionalPropertyTypes is not enabled, so assigning undefined to an
optional field is equivalent and the spreads bought nothing. Destructuring
flags once and extracting two flag readers removes most of the rest.

One correctness note on the simplification itself: the first version used
`artifactsDir && expandHome(...)`, which returns '' for an empty-string flag
where the previous code called expandHome(''). Replaced with an explicit
undefined check so the empty-string path is unchanged.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* test(live): share the replay test-suite harness across iOS and Android

Both live journeys invoked the public test command and then re-derived the
same value-contract assertions by hand — suite totals, per-script status,
replay counts, non-empty JUnit. Those are claims about the published suite
result and are identical on every platform, and they had already drifted: iOS
iterated with readReplayCommands inline, Android cast data.tests at the call
site.

The shared helper owns exactly that boundary. It takes the caller's runStep
rather than binding a context type, so it is not a platform-configured runner
and cannot template a platform's journey.

Everything a platform genuinely differs on stays with the caller: which
scripts run, the retry policy (iOS 2, Android none — itself a claim worth
keeping), which commands each script exercises, and the behavioral evidence.
Both callers keep every verify* call they had.

67 lines removed, 15 added.

Residual risk: this container has no iOS or Android devices, so the live suites
could not be executed here. Typecheck and lint pass; the harness needs a run on
real targets before the claim that behavior is unchanged is evidence rather
than inference.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* test: pin directory enumeration in the platform-binding suite test

The test wrote two scripts into a temp directory and assumed discovery would
return them in creation order. Directory expansion deliberately preserves
filesystem order to match Maestro — only glob expansion sorts, and 'preserves
Maestro directory filesystem order' pins that with a mocked opendirSync. So the
ordering contract is correct; this test's assumption about enumeration was not.

It passes on CI, where small directories usually enumerate in creation order,
and fails on filesystems that do not — identically on clean main, where the
platform-to-script binding appears reversed.

Pinning enumeration the way the discovery tests already do keeps the subject
intact (each script binds to ITS declared platform, and session numbering
follows discovery order) without depending on the host filesystem. The fs
import became a default import because vi.spyOn cannot redefine an ESM
namespace export.

915 of 915 handler and package tests now pass here.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* test: scope the enumeration spy and restore it in a finally

The spy I added restored only on the happy path and asserted on its argument
inside the mock implementation. Either would misfire for anything else sharing
the worker: an assertion thrown from inside fs, or a leaked global opendirSync,
surfaces as a worker crash with no failed test rather than a readable failure.

It now delegates to the real implementation for any directory but this suite's
own, and restores in a finally.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* fix(replay-test): put package tests where they are actually run

Review found the moved package tests were neither executed nor typechecked.
They sat under packages/replay-test/test/, but vitest's unit-core lane includes
packages/*/src/**/*.test.ts, and neither the root nor the package tsconfig
covers a top-level test directory. A plain unit-core run discovered zero files
under the package.

That is why they looked green: my earlier runs passed those paths explicitly on
the command line, which masked that the default run skipped them. The count is
the proof — 550 files/4741 tests before, 553/4753 now, and the delta is exactly
the three files and twelve tests that were being skipped.

The runtime test also imported runReplayTestAttempt from the package specifier,
which the facade does not export. It would have failed the moment it was
discovered. It now imports internally, like the rest of the internal tests.

Also removed replayTestAttemptFailure from the facade: zero consumers outside
the package, so exporting it widened the boundary for nothing. P3 asks for a
one-function facade.

553 test files and 4753 tests pass; lint and the layering guard are clean.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* style: format the facade after removing the unused export

A scripted edit removed the export line but left a stray blank line; oxfmt was
not re-run on that file afterward, so Lint & Format caught what pnpm lint
alone does not.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* fix(replay-test): typecheck the package and fix a type-only import

Review found the moved package tests were transpiled by vitest but never
typechecked: the root typecheck script builds six packages via tsc -b and
packages/replay-test was not among them, so its tsconfig was never used.

That hid a real TS2459. session-test-runtime.test.ts imported
ReplayTestAttemptOutcome from ../session-test-runtime.ts, which imports that
type but does not re-export it. It now imports from ../session-test-types.ts,
where the type is defined.

Adding the package to the tsc -b list closes the gap. Verified empirically
rather than assumed: planting a string-to-number error in a package test makes
typecheck fail, and removing it makes it pass. This is the second finding of
the same shape on this PR — first the tests were not discovered by vitest, now
they were not covered by typecheck — so the gate was confirmed to reach the
files rather than trusted to.

12 package tests pass, lint, format and the layering guard are clean, and
typecheck is clean with the package included.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-31 20:16:18 +02:00
Michał Pierzchała e669676bc8 test: fix Android observability scenario to Android contracts (#1531)
The full:observability-artifacts scenario (#1482/#1484) had never executed
end-to-end: both nightly Android Full Emulator Suite runs since merge died
on adb infra before the suite ran, and a live run fails deterministically
at its first perf assertion. Fixing that revealed four more latent
failures, each written against iOS or remote-daemon behavior the Android
live run does not have. Validated with two consecutive green full-tier
runs on a dedicated Pixel 9 Pro XL emulator.

- perf metrics: assert totalPssKb (Android's required meminfo field)
  instead of the Apple-only residentMemoryKb.
- presses: reveal the Quick-actions card with scroll steps before
  pressing home-open-catalog/home-open-settings — Android snapshots only
  contain on-screen nodes — and restore scroll top before waiting on the
  home title, since scroll position persists across tab switches.
- batch get: target id="dismiss-notice" (a node that owns its text);
  Android resolves the home-title container to a child's text (the
  subtitle), unlike iOS's container label.
- events: run an explicit snapshot so the timeline assertion holds when
  the scenario runs standalone under AGENT_DEVICE_ANDROID_E2E_SCENARIOS.
- artifacts: assert the local-client contract — trace-log tracked,
  downloadable, consumed; screen-recording inventory entries only exist
  for remote clients (artifacts without a client localPath are never
  tracked). Also stop consuming the download response body in the assert
  message before arrayBuffer() reads it.
2026-07-31 19:48:45 +02:00
Michał Pierzchała 14b71c96f1 test: narrow Limrun public type contract (#1528) 2026-07-31 18:02:34 +02:00
Michał Pierzchała 2e0ed0a41d fix: normalize physical iOS landscape taps (#1520)
* fix: normalize physical iOS landscape taps

* fix: preserve iOS selector tap fallback

* fix: normalize stale iOS screenshot dimensions
2026-07-31 15:36:21 +02:00
Michał Pierzchała da93191201 refactor: move Limrun provider behind package facade (#1518)
* refactor: move Limrun provider behind package facade

* fix: preserve Limrun public provider types

* fix: tighten Limrun provider facade boundaries

* test: harden Limrun compatibility coverage

* fix: narrow Limrun public type exports

* fix: narrow Limrun provider exports
2026-07-31 15:35:49 +02:00
Michał Pierzchała e3cdbcdb37 fix(test): give the iOS scroll search a capture-sized wait budget (#1514)
`assertElementTextAfterScrolling` waited 1000ms per attempt. The wait's whole
budget goes to its first capture, so that budget has to cover one snapshot of
the current surface; on a loaded simulator it does not, and every attempt fails
with `wait_capture_stalled` instead of reporting the element as off-screen.

This broke the iOS Smoke Tests lane on main. #1484 moved the smoke automation
scenario onto this helper for `automation-press` and `automation-longpress`;
before that it was only reached from the full lane, so the tight budget went
unnoticed. Every sibling wait in the same scenario already budgets 2500ms or
more (`label="Settings"` 10000, `alert wait` 5000, `wait text` 2500).

Two changes:

- Raise the per-attempt budget to 2500ms, matching the nearest sibling.
- Stop spending a scroll attempt on a stalled capture. A stall means the
  snapshot never came back, so the surface was never read — it is not evidence
  the element is off-screen, and scrolling on it moves the surface for an
  unrelated reason. Two stall retries absorb a slow runner; a genuine absence
  still consumes attempts and still fails.

The final assertion now carries the last wait's JSON, so a future failure says
whether it stalled or genuinely never found the element.


Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-31 09:10:44 +02:00
Michał Pierzchała b125435989 refactor: extract WebDriver provider package (#1504)
* refactor: extract webdriver provider package

* refactor: consolidate shared XML codec
2026-07-31 09:10:04 +02:00
Michał Pierzchała 32b9db2d7a test: add Android full emulator coverage (#1484)
* test: add Android full emulator coverage

* ci: package Android helpers before nightly coverage

* fix: harden Android nightly runtime evidence

* fix: expose trace artifacts in MCP schema

* style: format Android coverage manifest

* fix: address Android coverage review findings

* refactor: share live device coverage helpers

* fix: restore fixture landmarks in device smokes

* refactor: centralize live artifact assertions

* fix: normalize fixture canary visibility
2026-07-30 21:07:06 +02:00
Michał Pierzchała 0e51007b04 refactor: isolate maestro engine package (#1506)
* refactor: isolate maestro engine package

* perf: deepen maestro facade boundaries
2026-07-30 20:58:20 +02:00
Michał Pierzchała e70fac7b44 fix(ios): deny viewport during capability admission on Apple targets (#1507)
The capability matrix admitted `viewport` on every Apple simulator and
device while no Apple interactor implements `setViewport`, so the command
was routed to the device and only rejected inside dispatch with the
backend-shaped "viewport is not supported by this backend". Nothing can
serve it there: Apple screen geometry is fixed by the selected device type
and neither simctl nor XCTest exposes a resize primitive, so viewport is a
web-surface contract (its CLI summary and docs already say so).

Deny it in the descriptor's capability facet, matching the adjacent Android
denial, and give the Apple family an unsupported hint that names the surface
that does support it. Admission now fails before the request reaches the
device: "viewport is not supported on this device" with supportedOn: web.

The iOS simulator coverage manifest's known-gap row becomes a
capability-denial row owned by the static coverage test, which retires the
whole known-gap level and the `full:known-gaps` live scenario that existed
only to pin this failure.

Closes #1407
2026-07-30 20:22:17 +02:00
Michał Pierzchała 0ee2a86129 refactor: extract contracts workspace package (#1499)
* refactor: extract contracts workspace package

* fix: preserve screenshot diff result contract

* test: stabilize Android keyboard smoke
2026-07-30 17:07:46 +02:00
Michał Pierzchała cd9a7ce41b test(android): add comprehensive emulator E2E coverage (#1482)
* test(android): add catalog emulator smoke coverage

* test(android): use stable snapshot diff mutation

* test(android): assert actual back destination

* test(android): separate keyboard and fill IMEs

* fix(ci): keep Android timing report in one shell

* refactor(test): simplify simulator e2e coverage

* test(android): assert stable diff landmarks

* fix(android): release snapshot helper gracefully

* fix(android): fully release snapshot helper runtime

* test(android): report coverage classifications

* fix(android): stabilize accessibility root capture

* fix(android): bound UiAutomation connection

* ci: upload worktree daemon diagnostics

* fix(android): cancel stalled wait captures

* fix(android): bound helper fallback lifecycle

* fix(android): harden emulator e2e lifecycle

* fix: align e2e changes with kernel package

* test(android): prove alert helper reuse directly

* fix(android): cancel stalled settle captures

* fix(android): separate helper retirement budgets
2026-07-30 15:10:21 +02:00
Michał Pierzchała 2e6849ddb6 test: stabilize parallel provider scenarios (#1497) 2026-07-30 14:41:37 +02:00
Michał Pierzchała 117cf4ff68 test(ios): fix nightly fixture selector drift (#1492)
* test(ios): fix nightly fixture selector drift

* test(ios): harden fixture selector coverage
2026-07-30 12:14:57 +02:00
Michał Pierzchała 76453add71 refactor: pnpm workspace + @agent-device/kernel pilot (#1490 W0) (#1494)
* refactor: pnpm workspace + @agent-device/kernel pilot (#1490 W0)

Extend the workspace with packages/* and move the kernel behind an
enforced public API: packages/kernel with nine consumer-earned subpath
exports (errors, device, snapshot, contracts, collections, rect,
redaction, daemon-error, bounds — the last absorbed from utils as Rect
vocabulary). Every kernel import repo-wide becomes the
@agent-device/kernel/<sub> specifier; kernel tests move to
src/__tests__/kernel/ and exercise the package surface. The root
declares the package in devDependencies (workspace:*), tsdown bundles
it (noExternal) so the published artifact and its runtime dependency
manifest are unchanged.

Gate rewiring in the same change, per the W0 brief:
- R1 kernel-sink retires (physically subsumed); new R11
  package-boundaries guards no-root-back-imports, relative tunnelling
  past exports maps, undeclared workspace deps, and non-exported
  subpaths, with runtime resolution pins via import.meta.resolve.
- resolveImportEdges and mutation ownership follow workspace
  specifiers through exports maps, keeping R4 cycle checks, depgraph,
  and derived test ownership connected across the seam (kernel-errors
  still owns 495 tests). listSourceFiles includes packages/*/src.
- kernel becomes an unranked zone; mutation registry, stryker mutate
  globs, and the mutation-affected workflow path filter move to
  packages/kernel/src/errors.ts.
- check:affected gains packages/ ownership (manifests fail open);
  vitest and coverage include packages/*/src; fallow ignores
  packages/** (its resolver cannot follow workspace specifiers).
- The affected-selector CI job installs dependencies: its closure now
  crosses workspace specifiers, and the R8 relative exception is
  unsafe for production src files (Node ESM does not realpath, so dual
  specifier/relative loads would instantiate modules twice). The R8
  zero-dep set is pinned empty with that rationale.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUv7bvbWNryuXgSBuqTtep

* fix: address W0 review — mutation sandbox, exports-map resolution, tsc -b

Review findings on #1494, all five:

1. contracts-schema-public.test.ts reads the kernel source at its
   packages/ path (fs access invisible to the codemod and typecheck).
2. Mutation lane: Stryker sandboxes the tree but pnpm's node_modules
   symlink resolves @agent-device/* back to the real repo, so mutants
   in the sandbox never load and vitest.related finds no tests.
   vitest.mutation.config.ts now aliases each EXPORTED specifier to
   its source (derived from exports maps, never a wildcard), keeping
   resolution inside the mutated tree. Validated: kernel-errors module
   runs end to end (dry run 3,984 tests, mutants killed, exit 0).
3. Layering/depgraph resolve workspace specifiers through the
   exports-derived map (workspaceSpecifierTargets) instead of
   reconstructing paths, so '.'-facade packages resolve; the
   positional fallback remains only for map-less fixtures (P0 pin).
4. Per-package project references implemented: packages/kernel is
   composite (emitDeclarationOnly -> dist-types, gitignored), the root
   references it, and typecheck becomes tsc -b — probed to catch type
   errors on both sides under TypeScript 7 native.
5. R11's relative-route exception now requires membership in an actual
   R8 zero-dep job closure (zeroDepClosureFiles walks entries), not
   mere scripts/ placement — closing the dual-instantiation bypass.

Also from review discussion: daemon-error moves out of the kernel
package to src/client/ — its consumers (cli, client facade) rehydrate
wire DaemonErrors client-side; the daemon only produces them. Kernel
drops to 8 exported subpaths before any of them ship.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUv7bvbWNryuXgSBuqTtep

* refactor: one exports-map reader for mutation alias and ownership

Fallow flagged workspaceExportAliases (cognitive 15, CRAP 90). The
manifest-reading logic already exists as workspaceSpecifierTargets in
scripts/layering/package-boundaries.ts, so both the Stryker sandbox
alias table and the mutation ownership walker now consume it instead
of carrying near-clones. Behavior unchanged; mutation suite 45/45 and
changed-code fallow green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUv7bvbWNryuXgSBuqTtep

* fix: composite kernel without a root references edge

FreeRange runs plain `tsc -p tsconfig.json`, and a root `references`
entry makes non-build-mode TypeScript demand the referenced project's
built declarations (TS6305) — a standing "build first" tax on every
plain -p consumer (fr, editors). Keep the per-package composite
project and build it in typecheck (`tsc -b packages/kernel` before the
root and examples/sdk passes), but drop the root references edge: root
consumption resolves through exports to source, identical to runtime
and to the bundler. Probed: plain -p green with no prebuilt output;
kernel-side type errors still caught by its own build.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUv7bvbWNryuXgSBuqTtep

* fix: R11 uses the layering parser; mutation config is a fallow entry

Review blockers on #1494:

- R11's private single-quote regex could miss a double-quoted or
  re-export route into packages/*/src. specifierSites now delegates to
  the layering model's parseImports (both quote styles, side-effect
  imports, re-exports, dynamic imports), with direct regressions for
  each formerly-invisible form.
- vitest.mutation.config.ts becomes a declared fallow entry instead of
  a tolerated unused-file finding: the full-repo audit now reports it
  reachable (unused files 2 -> 1; the remainder predates this PR).

FreeRange clean-checkout evidence: with packages/kernel/dist-types and
every *.tsbuildinfo deleted, `pnpm check:freerange` reports 0 findings
on this head — the TS6305 topology died with the root references edge
in the previous commit; check:freerange has no build precondition.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FUv7bvbWNryuXgSBuqTtep

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-30 12:12:46 +02:00