1272 Commits

Author SHA1 Message Date
Michał Pierzchała 95b4623461 0.20.4 v0.20.4 2026-08-03 16:49:15 +02:00
Michał Pierzchała 2e74b789fd feat: verify device cloud connections (#1564)
* feat: verify device cloud connections

* refactor: unify connect provider adapters

* refactor: separate connect verification facts

* fix: tighten connect provider verification

* fix: use neutral cloud connection wording

* perf: deduplicate local affected checks

* refactor: simplify affected check runner

* refactor: derive connect workflow from verification
2026-08-03 16:47:57 +02:00
Michał Pierzchała 6ef7cc0d6d docs: add security policy (#1568) 2026-08-03 15:58:47 +02:00
Michał Pierzchała 99967c7f01 fix: restrict project config trust (#1565)
* fix: restrict project config trust

* fix: preserve daemon auth transport context

* refactor: simplify project config trust

* fix: restrict project config write sinks
2026-08-03 15:48:45 +02:00
Michał Pierzchała 2bdbef3f2b fix(daemon): distrust post-gesture stability that matches the pre-gesture baseline (#1563)
* fix(daemon): distrust post-gesture stability that matches the pre-gesture baseline

#1542 defect 2: post-gesture-stabilization.ts treated two consecutive
matching AX-signature polls as proof the screen settled. On iOS's AX-free
synthesized gesture lane, XCTest's tree can serve a stale-but-internally-
consistent read for a window after a scroll/swipe, so that "match" can be
false: the daemon then evaluates pre-gesture node positions on the very
next interaction.

Fix: capture the interaction-surface signature before the gesture
dispatches (reusing session.snapshot, no extra capture), and when a quiet
poll-to-poll match still equals that baseline, don't trust it — keep
polling past the normal 1.5s deadline up to a bounded 3.5s cap. On cap
expiry with the signature still identical, accept the result (a genuine
no-op gesture is the honest answer) but flag it via a new
post_gesture_snapshot_stale_accept diagnostic so a stale-accept is
observable in ndjson.

Baseline comparison is subset-tolerant (interactionSurfaceMatchesBaseline)
rather than whole-array equality: the pre-gesture baseline and the
post-gesture capture are routinely fetched with different snapshot scopes,
so naive equality reported "changed" from scope drift alone and never
caught the real staleness on first implementation — live-verified and
fixed before shipping.

Platform-scoped to Apple only (requiresPostGestureBaselineDistrust):
Android's persistent helper clears its accessibility-node cache before
every capture (AccessibilityTreeCapture.capture, #1254/#1259), so an
Android post-gesture read is fresh by construction and never computes a
baseline signature — latency and semantics unchanged, confirmed live
(checkout-form-android.ad + gesture-lab-android.ad 2/2 on Pixel_7_CI).

Does not close #1542: live validation on checkout-form.ad still fails at
step 11, but now for a distinct reason this fix correctly surfaces rather
than causes — a corrupted ScrollView-ancestor viewport frame
((18,381,366,109) vs the true (18,62,366,729)) that the off-screen guard's
findNearestScrollableAncestorRect trusts, independent of whether the
signature matches the baseline. gesture-lab.ad (iOS) remains 2/2 clean,
confirming no regression on the passing scenario.

_Generated by Claude Code_

* refactor(daemon): decompose the stabilization loop's diagnostics and capture pair

The distrust integration pushed capturePostGestureStabilizedResult over the
complexity gate (cyclomatic 15, cognitive 24); extracting the settle-diagnostic
branching and the capture+signature pair restores a clean fallow pass with no
behavior change.

* fix(daemon): require discriminating overlap for a post-gesture baseline match

PR review on #1563 (P1): interactionSurfaceMatchesBaseline returned true
whenever ANY shared entry was frozen, including the application/window
viewport root, whose rect is invariant under any gesture. In the exact
scope-drift case this PR supports, a broad pre-gesture baseline and a
narrow post-gesture selector capture can share only that root after a
real, successful scroll — the boolean predicate called that a baseline
match and extended the interaction to the 3.5s stale-read cap on zero
real evidence.

Fix: replace the boolean with classifyBaselineSurfaceEvidence, a
subset-tolerant classifier reusing this module's existing
InteractionSurfaceChange vocabulary ('changed' | 'unchanged' |
'ambiguous') instead of a bespoke boolean or an Application-only special
case. An entry only counts as evidence when it is `discriminating` —
excludes the viewport root (minimal local equivalent of
snapshot-occlusion.ts's isViewportRoot) and keyboard chrome (minimal
local equivalent of snapshot-chrome.ts's keyboard-container check), both
computed once at signature-build time since the flat signature-entry
representation has no ref/parentIndex to reuse those modules' full
ancestor-walk classifiers directly. Zero discriminating overlap is now
'ambiguous' (insufficient evidence) rather than a match, and
decidePostGestureStabilityVerdict falls through 'ambiguous' to 'trust' —
the safe default, same as 'changed'.

Tests: the reviewer's exact shape (signatures sharing only the
Application root, with the real content swapped) at three layers —
classifyBaselineSurfaceEvidence directly, decidePostGestureStabilityVerdict,
and the full capturePostGestureStabilizedResult async loop (proving no
cap-tax: settles in 2 capture attempts, not 3.5s). Also: root+one real
element both frozen still matches (guards against over-excluding), and
keyboard chrome excluded from discriminating overlap. All prior tests
kept green unchanged.

Counterfactual: reverted to the old boolean predicate and reran — 5
tests went red, including the async regression test, which didn't just
fail an assertion but timed out after 5s because the boolean predicate
extended the interaction to the 3.5s distrust cap the test's 1s timer
advance never covered — exactly the "extends to cap" failure mode the
review predicted. Restored, 37/37 green.

_Generated by Claude Code_

* fix(daemon): exclude keyboard descendants (not just the container) from baseline evidence

PR review on #1563 (two findings, blocking merge):

1. isKeyboardChromeKind excluded only the [Keyboard] container node itself.
   collectKeyboardChrome (src/core/snapshot-chrome.ts, the established
   source of truth) classifies the WHOLE keyboard window/subtree — keys,
   AND the "Next keyboard"/"Dictate" assistant buttons, which are documented
   siblings of the container, not descendants, so a container-descendant
   walk alone provably misses them. In the scope-drift case this PR
   supports, a successful scroll can leave only those keyboard descendants
   shared between a baseline and a later capture, and the narrower check
   called that a baseline match — extending a fresh result to the 3.5s
   stale-read cap.

   Fixed by exporting a narrow predicate, collectKeyboardChromeRefs(nodes),
   from snapshot-chrome.ts (returns collectKeyboardChrome(nodes).refs, no
   Android union — this caller has no appBundleId in scope and only needs
   the iOS half). buildInteractionSurfaceSignature computes it once per
   signature build and threads it into buildInteractionSurfaceEntry, so
   discriminating is now `!isViewportRootKind(node) && !keyboardChromeRefs
   .has(node.ref)` — reusing the real ancestor-walk classification instead
   of a per-node type check, no ancestry needed in the signature entries
   themselves.

2. post-gesture-stabilization.test.ts had grown to 550 LOC, past the
   repository's 500-line extraction tripwire (AGENTS.md: "past 500,
   extract before adding behavior... Tests are not exempt"). Split along
   subject lines: the pure decidePostGestureStabilityVerdict coverage
   moved to a new sibling post-gesture-stabilization-verdict.test.ts, and
   shared fixtures (pickupSnapshot, deliverySnapshot, applicationRootNode,
   keyboardWindowNodes, makeSession) moved to a new non-test
   post-gesture-stabilization-fixtures.ts. The async capturePostGesture-
   StabilizedResult loop tests stay in the original file. Assertions
   unchanged, only relocation, plus the new regression tests below.
   Resulting LOC: post-gesture-stabilization.test.ts 381, -verdict.test.ts
   208, -fixtures.ts 129 (interaction-outcome-policy.test.ts grew to 413,
   still under the tripwire).

Tests: the reviewer's exact regression — a shared overlap consisting only
of keyboard descendants (a key + the "Next keyboard" sibling button, NOT
the container) plus real content that changed (Pickup -> Delivery) — at
three layers: classifyBaselineSurfaceEvidence directly (ambiguous), the
verdict function (trust, elapsedMs: 0), and the full async capture loop
(settles in 2 attempts, no cap tax).

Counterfactual: reverted isNonDiscriminatingSurfaceNode to a container-only
check (normalizeType(node.type) === 'keyboard') and reran — 3 of the new
tests went red across all three layers, including the async test, which
timed out after 5s (not just a failed assertion) because the container-only
exclusion genuinely extended the interaction to the 3.5s distrust cap the
test's 1s timer advance never covers — the same "extends to cap" failure
shape as the review's finding 1. Restored, 40/40 green.

_Generated by [Claude Code](https://claude.ai/code)_
2026-08-03 15:46:04 +02:00
Michał Pierzchała 123521652c fix(ios): double-check off-screen click refusals against a direct element read (#1566)
* fix(ios): double-check off-screen click refusals against a direct element read

#1542: after an AX-free scroll on iOS, the off-screen interaction guard can
refuse a click even though the target is genuinely on-screen, because it
trusts a scroll-container ancestor's rect from the bulk accessibility tree,
which a keyboard-dismiss content-offset correction can leave stale/corrupted
while the target's own rect is already correct.

When the guard is about to refuse on iOS, it now takes a single fresh,
tree-independent XCUITest read of the target element (querySelector) and
trusts that read's live `hittable` + rect-vs-root-viewport signal instead,
if it positively confirms on-screen. Any failure to unambiguously re-resolve
the element (no id/label, not found, ambiguous, transport error) fails
closed exactly as before. Genuinely off-screen targets, and every other
platform, are unchanged: the backend method is gated to local (non-provider)
iOS sessions only, and only ever runs on the about-to-fail path.

The decision itself is a pure function (decideOffscreenRefusalDoubleCheck in
mobile-snapshot-semantics.ts) with counterfactual-proven tests: hardcoding it
to always trust the bulk verdict turns the rescue test red, and hardcoding
it to always trust the direct read (including on "unavailable") turns the
fail-closed/genuine-refusal test red.

Live-validated on a fresh-boot iOS simulator: checkout-form.ad 2/2 passes
(previously failing at step 11), gesture-lab.ad 2/2 (regression), and the
Android checkout-form/gesture-lab suite passes unchanged, proving no
cross-platform behavior change.

* fix(ios): tap the live rect after a rescued offscreen refusal; collapse the double-check to one backend hook

Review blockers 1+2 (interleaved by design — the soundness fix is expressed
through the collapsed hook's contract):

1. SOUNDNESS: a rescued refusal now returns the node PATCHED WITH THE LIVE
   RECT the backend confirmed, and every downstream use (tap point, response)
   reads from that returned node — never the original. In the frozen-tree
   manifestation (the whole bulk tree pinned at pre-gesture values), the
   original rect can be stale even when the rescue verdict is correct;
   tapping it would have silently landed at the wrong coordinate. New
   regression: offscreen-double-check.test.ts's frozen-tree case, with a
   counterfactual (revert to computing the point from the pre-guard node)
   proven red then reverted.

2. SURFACE: collapsed to ONE optional backend hook,
   `confirmOffscreenTargetVisible?(context, node, rootViewport): Promise<Rect
   | null>` — conceptually a boolean, but returns the live rect so item 1's
   fix has something to act on. Deleted decideOffscreenRefusalDoubleCheck,
   the OffscreenRefusalDoubleCheckSignal/Reading ADT, and resolution.ts's
   dual-signal reconciliation shell: the bulk side was hardcoded 'off-screen'
   at the only call site, so the two-signal model was dead weight. The shared
   guard is now: bulk-off-screen -> ask the hook -> a live rect proceeds
   (patched), anything else (including no hook) throws exactly as before.

The pure geometry boundary that decision reduces to (`isConfirmedOnScreenProbe`
in mobile-snapshot-semantics.ts, replacing the deleted ADT) is unit-tested
with two counterfactuals: ignoring `hittable` and ignoring the viewport
containment check each turn a test red (proved, then reverted).

`throwIfOffscreenInteractionTarget` is now exported (ADR 0011 registry
honesty, see the contracts commit) and directly unit-tested in
resolution.test.ts, mirroring the existing tryResolveRefNode pattern.

* refactor(ios): direct-ios-selector.ts back to pure gate/parse; reuse queryDirectIosSelector

Review blocker 3 (BOUNDARIES):

- direct-ios-selector.ts no longer does any runner I/O — it's back to pure
  gate/parse (readSimpleIosSelectorTarget, deriveDirectIosNodeSelector,
  isDirectIosSelectorFallbackError) plus the ONE shared eligibility
  predicate, isLocalIosRunnerSession(session, { skipPendingPostGestureStabilization
  }). Both the direct-selector tap fast path and the new offscreen
  double-check probe call this same function; the one behavioral difference
  between them (the tap fast path skips a session with a pending
  postGestureStabilization, the double-check does not) is now an explicit
  parameter instead of two separately-written gates.

- The probe I/O moved to a new sibling, src/daemon/offscreen-target-probe.ts,
  which reuses selector-runtime.ts's `queryDirectIosSelector` (now exported
  and decoupled from SelectorRuntimeParams — it takes a session + a bare
  {key, value} selector + AppleRunnerRequestOptions) rather than opening a
  second querySelector client. Node extraction (`readDirectIosSelectorNode`,
  the one `as SnapshotNode` cast) stays singular, inside selector-runtime.ts.

- interaction-runtime.ts wires confirmOffscreenTargetVisible only when
  isLocalIosRunnerSession(session, { skipPendingPostGestureStabilization:
  false }) — deliberately NOT skipping a pending post-gesture stabilization,
  since that is exactly the window the double-check exists to cover.

* docs(contracts): name the iOS offscreen rescue hook as part of the guarantee matrix

Review blocker 4 (GUARANTEE HONESTY): the shared offscreen cell
(RUNTIME_TREE_SHARED_GUARANTEES.offscreen, used by runtime-selector and
runtime-ref) and the native-ref path's offscreen cell still named
isNodeVisibleOnScreen as sole enforcement after #1542's double-check landed —
that understates what actually enforces the guarantee now.

Both cells' `via` now point at throwIfOffscreenInteractionTarget (exported
from resolution.ts in the prior commit for exactly this), the real
end-to-end enforcement point: isNodeVisibleOnScreen is the bulk-tree
decision it starts from, and on iOS a would-be refusal can still be
confirmed via the optional AgentDeviceBackend.confirmOffscreenTargetVisible
hook before erroring. The cell's comment states the rescue-only, fail-closed
shape explicitly per ADR 0011's matrix rules — this does not weaken the
cell, it extends its description to match reality.

iOS rescue policy stays OUT of resolution.ts's shared docstrings (the
"spine"): this registry file is where per-path enforcement detail belongs,
and the optional-method wiring in interaction-runtime.ts remains the only
cross-platform touch.

The registry's own gate test (interaction-guarantees.test.ts) still passes:
every `via` resolves to a real exported symbol.

* test(ios): move #1542 offscreen double-check tests out of interaction.test.ts

Review blocker 5 (TEST HOMES): AGENTS.md forbids adding to
daemon/handlers/__tests__/interaction.test.ts (it predates the
test-mirrors-source-topology rule and shrinks opportunistically). Reverts
the 172 lines added there in the original PR version; interaction.test.ts is
back to its pre-#1542 baseline (81 tests, unchanged).

The same assertions now live in their proper homes (see the prior three
commits for the sources they cover):
- pure decision pin: src/utils/__tests__/mobile-snapshot-semantics.test.ts
  (isConfirmedOnScreenProbe, with the two counterfactuals)
- direct-guard pin: src/commands/interaction/runtime/resolution.test.ts
  (throwIfOffscreenInteractionTarget, mirroring tryResolveRefNode)
- probe unit tests: src/daemon/__tests__/selector-runtime.test.ts
  (queryDirectIosSelector) and src/daemon/__tests__/direct-ios-selector.test.ts
  (isLocalIosRunnerSession, deriveDirectIosNodeSelector)
- probe integration: src/daemon/__tests__/offscreen-target-probe.test.ts
  (confirmIosOffscreenTargetVisible, mocked runner)
- end-to-end rescue/refuse, including the frozen-tree live-geometry
  regression + its counterfactual: new sibling
  src/commands/interaction/runtime/offscreen-double-check.test.ts (next to
  resolution.ts, using the same createInteractionDevice harness
  resolution.test.ts already uses)

* style: oxfmt formatting for resolution.test.ts
2026-08-03 14:13:09 +02:00
Michał Pierzchała f8617a2db9 fix: read Android get text from target field (#1561) 2026-08-03 12:09:31 +02:00
Michał Pierzchała c18636315a fix(ios): keyboard-dismiss content settle race (#1542) — partial, defect 2 needs a decision (#1559)
* fix(ios): wait for post-dismiss content settle before the next gesture (#1542)

Dismissing the keyboard can trigger the app's own ScrollView content-offset
correction (e.g. releasing the inset it grew to keep a focused field above
the keyboard). That correction is a separate, unsynchronized animation that
`keyboard.waitForNonExistence` knows nothing about — the keyboard AX element
can disappear well before the app visually settles. The very next command is
frequently a synthesized, AX-free drag (scroll/gesture, kept AX-free so it
still works under #1105-family AX degradation), which has no XCTest
quiescence wait of its own, so it can land mid-animation and net to zero —
the "scroll does nothing" symptom on the Form screen's checkout-form.ad leg.

Add a bounded, AX-free screenshot-stability wait to dismissKeyboard() so the
runner only returns once the screen has actually stopped changing (or a
generous cap elapses). The stopping decision is a pure function
(runnerScreenshotStabilitySettled) covered by unit tests under
AGENT_DEVICE_RUNNER_UNIT_TESTS; the surrounding capture/sleep loop is the
thin, untestable I/O shell around it.

Live-verified on iPhone 17 Pro / iOS 26.2: the scroll now visually lands at
the correct position (confirmed via screen-recording frame correlation)
instead of leaving content at its pre-scroll offset.

Not a full fix for #1542: the checkout-form.ad corpus leg still fails at the
same step, now because the daemon's shared post-gesture snapshot
stabilization (src/daemon/post-gesture-stabilization.ts) can read a
stale-but-internally-consistent AX tree after the AX-free scroll and
mistake "unchanged across polls" for "settled", so the following click's
off-screen guard sees pre-scroll node positions. That is a cross-platform,
cross-command stabilization semantics change and needs a design decision,
not a unilateral fix here — see the PR description.

* ci(ios): execute the screenshot-stability runner tests (#1559 review)
2026-08-03 09:11:05 +02:00
Michał Pierzchała 480e3883b1 fix(daemon): reject unarmed close --save-script before teardown (#1558)
* fix(daemon): reject unarmed close --save-script before teardown

Live evidence (2026-08-02) showed a plain `open` followed by `close
--save-script` silently published a script: the close request armed
authoring at record time and published moments later in the same
request, folding the never-armed case into the ADR 0016 authoring
lifecycle. The resulting .ad carries selector fallback chains but no
recording-time target-v1 evidence, and nothing told the caller
evidence capture never ran — degraded replay verification with no
signal beats a loud refusal.

`assertTerminalRecordingCloseAllowed` (src/daemon/handlers/session-close.ts)
now rejects an unarmed `close --save-script` with INVALID_ARGS before
any teardown or filesystem work runs, the same seam that already
rejected ABORTED/PUBLISHED terminal recordings. The rejection does not
tear the session down, so a plain `close` retry still completes
cleanly; recovery names `open --save-script` since evidence can only
be captured from action zero. Repair transactions (ADR 0012) are a
disjoint lifecycle and are explicitly unaffected.

This is distinct from #1533 (an already-armed-then-aborted session
whose flag ingress re-enables recordSession and lets a *bare* close
publish); that case remains open.

* fix: review follow-ups for #1558 (help text, test strength, docs)

- Give replay --save-script its own help text instead of the shared
  open/close "arm on open, publish on close" description: replay's flag
  arms an ADR 0012 repair transaction, a disjoint lifecycle. Adds
  CommandSchema.flagDescriptionOverrides so a command can swap a shared
  flag's usageDescription without duplicating the FlagDefinition entry
  (which would have shown --save-script twice in `help replay`). Pinned
  in src/cli/parser/__tests__/cli-help-command-usage.test.ts (open/close
  keep the shared text unchanged; replay gets the new one).

- Strengthen the never-armed close --save-script regression test in
  session-close-shutdown.test.ts: the fixture now carries real
  cleanup-bearing state (an active iOS simulator recording, reusing
  makeIosSimulatorRecordingSession/recordingKillMock) with spies proving
  no teardown hook (recorder kill, runner stop) runs on the rejected
  request, then that a follow-up plain close does tear it down. The
  prior fixture had nothing for teardown to observably touch, so moving
  the guard after stopBestEffortSessionResources would have passed it
  silently. Also fixes a latent test-isolation leak this exposed: an
  earlier test set a persistent mockStopIosRunnerSession rejection
  (vi.clearAllMocks() clears call history, not implementations), which
  would have poisoned any later Apple-platform close test; scoped it to
  mockRejectedValueOnce.

- Point the migration guide (website/docs/docs/migrating-gestures.md) at
  `open --save-script` → interact → `close` instead of the now-rejected
  `open` → interact → `close --save-script`, matching the new guard and
  the corrected help text.

_Generated by [Claude Code](https://claude.ai/code)_
2026-08-03 09:10:40 +02:00
Michał Pierzchała 2c2df031ff feat: keep replay session active on request (#1554)
* feat: keep replay session active on request

* test: cover replay keep-session provider route

* fix: make replay session handoff reliable

* refactor(daemon): extract the replay terminal-lifecycle policy module (#1554 review)

session-replay-runtime.ts was already over the 500-line extract-before-adding-behavior
tripwire before this PR; the keep-session/repair terminal-close decision, its
live-session postcondition, and the dispatched-action count pushed it further past
budget. Move that policy into a focused session-replay-terminal-lifecycle.ts
(isExecutableReplayAction, resolveSuppressedTerminalCloseIndex,
countExecutedReplayActions, requireLiveSessionForKeepSession) so the runtime file
stays orchestration-only, and mirror its PR-added unit tests into
session-replay-terminal-lifecycle.test.ts. Pure extraction: no assertions changed.
2026-08-02 18:35:13 +02:00
Michał Pierzchała 4551fb7aa3 test: make pid-liveness fixtures deterministic under load (#1556)
Three tests classify a fixture owner's liveness via classifyOwnerLiveness
(or the daemon-process equivalent), which re-reads the owner's process
start time via a real `ps -p <pid> -o lstart=` shell-out with a 1s timeout.
Under full-suite CPU contention that second read can miss its deadline and
return null, mismatching the value captured earlier and flipping a
genuinely-live owner to 'owner-process-dead'.

Pin the pid->start-time (and, for the daemon-client case, pid->command)
mapping to a deterministic value per test file instead of letting a second
real subprocess call race the first, without weakening the dead-owner path
(isProcessAlive stays real and un-mocked everywhere).
2026-08-02 17:59:28 +02:00
Michał Pierzchała 14d731c015 test: pin selector-port behavior ahead of the P5 extraction (#1478) (#1552)
* test: pin selector-port behavior ahead of the P5 extraction (#1478)

Pins, at existing root seams, the eight behavior cells the approved P5
amendment (issue #1478 comment 5156017698) requires the future
packages/ad-replay selector port (readSelectorExpression /
resolveRecordedTarget / buildSelectorCandidates) to preserve. Test-only —
no production code changes.

* test: consolidate duplicated cell-5/cell-7 coverage per review

Cell 7: relocate #1349's wait-landmark cases from selector-read.test.ts to
selector-wait.test.ts (the 1:1 topology location for selector-wait.ts),
replacing the weaker duplicate cell-7 cases added in the prior commit. The
relocated tests keep the stronger assertions (real computeTargetEvidence-
derived evidence, an initial no-match poll, observed-ancestry checks, and
the plain-timeout-vs-landmark-mismatch distinction).

Cell 5: the first case overlapped an existing later-alternative regression
in session-replay-target-classification.test.ts. Sharpened it (rather than
dropping it, since it is the only counterfactual-sensitive case for the
allowDisambiguation=false skip path) to isolate the branch the existing
regression's exact-tie fixture cannot reach, and paired it explicitly with
the second case as a same-fixture, flag-flipped contrast.
2026-08-02 12:26:18 +02:00
Michał Pierzchała 4fd04414e0 fix: clear the last polynomial-redos instance in swift-cache (#1549)
* fix: clear the last polynomial-redos instance in swift-cache

sanitizeCacheName used /^-+|-+$/g to trim edge dashes, the same
js/polynomial-redos pattern PR #1546 retired everywhere else. Replace
it with the linear-time trim used there, and add a counterfactual
regression test that fails against the old regex on a long interior
dash run.

* fix: keep sanitizeCacheName private, drive redos/fallback pins through compileSwiftSourceText

Addresses PR #1549 reviewer feedback: sanitizeCacheName was exported
solely so the regression test could import it, which docs/agents/testing.md's
test-interface rule forbids. Reverted the export and rewrote the test to
exercise the sanitizer through compileSwiftSourceText, an existing production
seam that already calls it.

- Timing pin: a cache name with a 100k-char interior dash run still resolves
  in sub-second time (the call may reject once it reaches disk I/O due to the
  OS path-component length limit, but that happens only after the now-fast
  sanitize step, so timing the settle either way still proves no catastrophic
  backtracking).
- Fallback pin: a cache name that sanitizes to nothing (e.g. '---') still
  produces the 'swift-helper' fallback, observed via the returned executable
  path.

Counterfactuals (see PR comment for full output):
- Restoring the retired `/^-+|-+$/g` regex trim made the timing pin fail:
  3428ms >= 1000ms.
- Removing the `|| 'swift-helper'` fallback made the fallback pin fail:
  basename did not start with 'swift-helper-'.
2026-08-02 12:26:06 +02:00
Michał Pierzchała 60400d04b7 feat(mutation): add target-annotation-serde + snapshot-occlusion kernels (#1553)
* feat(mutation): add target-annotation-serde + snapshot-occlusion kernels

Both are pure decision kernels the lane's own membership rule covers
(target-annotation-serde: parse/validate/normalize the .ad comment-line
codec, zero I/O; snapshot-occlusion: pure covered/not-covered decision
where a wrong answer silently blocks or mis-allows a tap) but were
excluded from KERNEL_MODULES.

Fixing the harness's packages/*/src blind spot was required, not
optional: test-scope.ts, ownership.ts, and vitest.mutation.config.ts
all hardcoded `src/` as the only place a kernel's tests could live.
target-annotation-serde's own tests live under
packages/ad-script/src/internal/__tests__/, so without this fix the
module would score 0% from day one — not from weak tests, but because
its test file was silently invisible to the lane. Widened the same
three places, plus mutation-affected.yml's path filter and
isTestFile/ownedTestFiles in ownership.ts, to also recognize
packages/*/src/**/*.test.ts (mirroring vitest.config.ts's own
unit-core project include list).

Triaged every surviving mutant from the initial run: real coverage
gaps got a new/adjusted test (kill-with-test), everything else is
documented equivalent with an inline comment at the mutation site
explaining the invariant that makes it unobservable (redundant
early-returns, JSON.stringify dropping undefined-valued keys,
Number.isFinite/isSafeInteger's total-function safety, caller-enforced
positiveRect/candidate invariants, etc). Baseline recorded from the
actual measured run, not inherited or guessed: 94.03% (315/335) and
89.74% (175/195).

* style: run the formatter over the four files the gate flagged
2026-08-02 11:36:43 +02:00
Michał Pierzchała e5cebcd8e3 fix(test): tolerate dead session in iOS e2e full-tier cleanup (#1548)
* fix(test): tolerate dead session in iOS e2e cleanup

full:device-lifecycle reboots the simulator, and whether the daemon
session survives that is environment-sensitive: it does on CI but not
locally, so every all-green local full-tier run ended red in cleanup
with all three retries of each step failing ('permission setting
requires an active app in session' / 'No active session').

Two layers:
- finalizeLiveRun re-checks sessionExists instead of short-circuiting
  on sessionOpen, so a session that died mid-run skips cleanup entirely.
- cleanupSession treats SESSION_NOT_FOUND and the appless-session
  INVALID_ARGS failure as already-clean instead of burning retries.

Other cleanup failures still exhaust three attempts and fail loudly.

* fix(test): narrow appless-session cleanup guard to the mic-permission step

Scope sessionAlreadyClean's INVALID_ARGS tolerance to the microphone-
permission reset step and its exact known message instead of matching
any cleanup step whose message contains "requires an active app in
session" — that substring is also thrown by the unrelated location
setting, so the old check could have hidden a real failure there.

Extract the per-step retry policy into an exported retryCleanupStep so
it's unit-testable without spawning the CLI, and add a deterministic
regression (test/integration/ios-simulator-e2e-cleanup.test.ts)
covering: a dead session (SESSION_NOT_FOUND) stops retrying on any
step, the known mic-permission appless response stops retrying, and a
different INVALID_ARGS (wrong step or wrong message) still exhausts
all three retries and fails.

* test(ios-e2e): cover the finalization cleanup-gate decision

Extract the sessionOpen re-check + conditional cleanupSession call out
of finalizeLiveRun into an exported finalizeSessionCleanup(context,
runSessionExists, runCleanupSession) — same behavior, now driven by
injected callbacks instead of the module-level runStep-backed
bindings, so it's unit-testable without spawning the CLI.

Add two deterministic cases: sessionOpen starts true and the final
sessionExists() resolves false -> cleanupSession is never invoked;
sessionOpen true and sessionExists() resolves true -> cleanupSession
runs (the live path). Counterfactual (reverting the recheck to the old
`sessionOpen || sessionExists(...)` form) turns the first case red as
expected.
2026-08-02 11:36:25 +02:00
Michał Pierzchała 2e4825ef64 refactor: tidy three post-extraction seams (#1551)
* refactor(replay): import REPLAY_VAR_KEY_RE from the codec package directly

vars.ts re-exported the constant for a single consumer, recorded-input.ts.
Point that consumer at @agent-device/ad-script and drop the shim, which also
makes script.ts's doc comment ("recorded-input.ts imports it from this
package") true.

* refactor(ad-script): import the target-annotation shape from contracts directly

The annotation shape types (TargetAncestryEntry, TargetAnnotationV1,
TargetScrollRegion, TargetVerification) live in @agent-device/contracts/replay;
the codec package re-exported them, and 21 files reached the shape through that
detour. Point every consumer — root src, root tests, and the package's own
tests — at contracts, then drop the re-export from the serde module and the
façade. Type-only, so nothing changes at runtime.

The package.json exports map is unchanged, so the R11 boundary assertion in
scripts/layering/package-boundaries.test.ts still holds as written.

* refactor(daemon): name the authoring-armed session read

`kind === 'authoring' && status === 'armed'` was spelled out at three handler
sites that all ask the same question. Give it a name next to
isSessionScriptPublished, mirroring how isRepairArmedSession is housed in the
repair projection, and route the three sites through it.

abortAuthoring's own guard keeps its inline check: that one is the transition's
legality test, not a session-level read.
2026-08-02 08:41:38 +02:00
Michał Pierzchała 634073a601 test(replay-test): cover failFast, retry exhaustion, plan-prep failure, and JUnit escaping (#1550)
Closes four untested failure paths flagged by the 2026-08-01 test-strength
audit of packages/replay-test:

- request.failFast=true now has a test proving the suite stops after the
  first failure and leaves the rest in the notRun bucket.
- Retry exhaustion (every attempt fails) is pinned to attempts === maxAttempts,
  guarding the attemptIndex <= retries loop bound.
- A malformed .ad source (bad env directive) is fed through discovery to
  prove the suite-level try/catch in session-test.ts converts the thrown
  AppError into a {status:'failed'} response instead of escaping uncaught.
- The JUnit reporter's escaping is exercised directly for the first time,
  round-tripping a title/message containing <, &, ", and a newline through
  parseXmlDocumentSync.

Each test was verified red: the production condition was temporarily broken,
the test observed failing, then the code was restored (see PR body for the
four before/after runs).
2026-08-02 08:41:14 +02:00
Karthik Varma 92b22229e6 feat(cloud-webdriver): BrowserStack device-feature capabilities, and fix cloud orientation (#1544)
* feat(cloud-webdriver): support BrowserStack device-feature capabilities

Adds the eight BrowserStack "device feature" session capabilities that had no
representation in agent-device: deviceOrientation, geoLocation, timezone,
language, locale, networkProfile, customNetwork, and resignApp.

These are vendor capabilities, so they are emitted inside `bstack:options`
rather than at the top level. BrowserStack's YAML config lists them unnested
and its SDK relocates them; agent-device talks to the hub directly, so it
nests them itself.

A single spec table drives both the flag reader and the capability builder, so
adding a capability is a table row rather than a branch in each. A structural
test asserts every field owns exactly one row, since a field the table forgets
would parse off the CLI, ride the profile, and then be silently dropped before
the hub ever saw it.

Rejects combinations the provider cannot act on unambiguously: an unknown
orientation is caught at the flag boundary instead of being forwarded to a hub
that accepts and then ignores it, --provider-no-resign-app is refused on
Android, and a named network profile cannot be combined with a custom network
shape.

Also fixes a latent shallow-merge bug in buildBrowserStackCapabilities: a
caller supplying its own `bstack:options` replaced the whole object and
silently dropped the project, build, and session labels. It is now merged
per key.

* fix(cloud-webdriver): rotate via WebDriver orientation endpoints

`setOrientation` on the cloud WebDriver path sent `mobile: rotate`, which is
not a driver command at all. UiAutomator2's own error enumerates its
extensions and `rotate` is absent from the list, so `agent-device orientation`
was a hard failure on every hosted provider.

It also forwarded agent-device's four-way rotation vocabulary verbatim
("landscape-left", "portrait-upside-down"), where the protocol accepts only
uppercase PORTRAIT/LANDSCAPE. Every other platform has a translation layer;
this path was the only one without one.

Now two transports, ordered by backend. `POST /rotation` takes exact four-way
degrees and leads on Android, since it is the only endpoint that can express
upside-down and left-versus-right. `POST /orientation` is two-way and leads on
XCUITest, which rejects `/rotation`. Each falls back to the other, because only
BrowserStack's UiAutomator2 is verified and a provider whose driver disagrees
should degrade rather than hard-fail.

Verified live against BrowserStack App Automate:
  POST /rotation {"x":0,"y":0,"z":0} -> 200 {"value":"ROTATION_0"}

The rotation-to-surface-index mapping moves to contracts/device-rotation.ts and
the existing adb path now reads from it, so the local and hosted mappings
cannot drift apart.

Note this rotates the current display, not persistent device rotation, so an
activity that does not pin its own orientation may still need rotating once it
is in the foreground.

The capability was declared "partial" without the transport existing, and no
test covered setOrientation on the cloud path; only adb and the Apple runner
were covered. Both gaps are now closed.

* fix(cloud-webdriver): narrow orientation fallback and gate provider-owned flags

Addresses review on #1544.

The orientation fallback caught every error, so a timeout, an auth rejection, a
dead session or a provider 5xx on the first transport was swallowed and retried
against the second. When that one also failed the caller got "rejected both
endpoints" with the real cause discarded. Fallback is now keyed on structured
unsupported-endpoint signals only — HTTP 404/405, or a W3C `unknown command` /
`unknown method` code — matching the repo rule of keying on typed details rather
than message text. Everything else rethrows unchanged.

Device-feature capabilities are BrowserStack-owned, but the flags were accepted
by any cloud provider, persisted into the generated profile, and then silently
dropped at session creation. `connect aws-device-farm` now rejects them with a
typed error naming each offending flag, raised before the provider's own
required-argument checks so the caller is told what is unsupported rather than
what else is missing. Ownership is modelled on the capability spec table, so a
new capability inherits the guard without a second list to maintain.

Adds provider-backed orientation scenarios driven through public daemon dispatch
against the fake WebDriver provider: the four-way endpoint on the happy path,
the documented collapse onto the two-way endpoint when the driver does not
implement `/rotation`, and a provider 5xx that must surface without consulting
the second transport. The fake server's route handling became a table in the
process — it had grown to ten branches in one function.

* fix(cloud-webdriver): read W3C error codes before status, enforce ownership at the runtime boundary

Addresses the second review pass on #1544.

The fallback classifier returned on any 404/405 before consulting the W3C error
code, so an HTTP 404 carrying `invalid session id` was masked as a missing route
and retried against the second transport. The structured code now takes
precedence whenever the driver sent one; bare status is consulted only when no
code exists. Two cases pin it: a 404 `invalid session id` and a 405 `timeout`
must both surface rather than fall through.

Provider ownership was enforced only in the CLI profile builder, which the typed
client and hand-authored remote-config profiles bypass entirely — both reach
session preparation without passing through `connect`, so the capabilities were
accepted and then dropped. The check now lives on the capability-ownership
module and runs inside AWS Device Farm's `prepareSession`, with the CLI builder
calling the same helper instead of its own copy. Covered by a scenario that
drives the runtime boundary directly and asserts the rejection happens before
any provider session is created.
2026-08-02 08:17:49 +02:00
Michał Pierzchała c7af6cd69d refactor(ios): drop the transport seam usbmux-first made dead (#1540)
The physical-device control exposed resolveRunnerTransport returning either a
network tunnel or usbmux, from before the route resolver decided transports.
Since #1517 the resolver returns a usbmux route for XCTest devices and for any
attached CoreDevice one, and only reaches the control after usbmux has reported
the device unattached — so the control is asked exclusively for a tunnel.

That left the usbmux arm with no live producer or consumer: the branch handling
it in the resolver was unreachable, and the XCTest implementation returning it
was called only by a test.

Collapse the union to the one shape that is resolved, rename the seam to say
what it does, delete the unreachable branch, and let XCTest reject a tunnel
lookup the way it already rejects app inventory and process lookup — it has no
CoreDevice tunnel, which is why the resolver never asks it for one.
2026-08-02 08:03:36 +02:00
Michał Pierzchała 5aba93f26b fix: restore scheduled workflow health (#1543)
* fix: avoid replay test slug ReDoS

* fix: restore scheduled replay and conformance health

* fix: trim replay slugs without regex backtracking

* chore: remove superseded replay slug changes
2026-08-01 21:16:30 +02:00
Michał Pierzchała b9509fe006 refactor: extract the .ad script codec into packages/ad-script (#1478) (#1536)
* refactor: extract the .ad script codec into packages/ad-script

Moves the mutually-coupled .ad read/write codec (script.ts, script-utils.ts,
script-formatting.ts, open-script.ts) plus the target-v1 annotation SERDE
slice of target-identity.ts into a new private leaf package,
@agent-device/ad-script, exporting only `.`. This is option 1 from the P5
scoping dossier on #1478: the codec is shared by the daemon's session-script
publication writer, the future replay engine, the CLI's `replay export`, and
Maestro's failure-label formatting, so it can no longer live in root src/
once packages/ad-replay lands (R11 forbids a package reaching into root src),
and a second export subpath or writer-half duplication are both ruled out by
existing gates/tests.

target-identity.ts keeps only the record/replay-shared classification core
(classifyTargetBindingMatch, local-identity/ancestry-prefix matching),
importing its shared types from the new package. Every real consumer
(re-derived by grep, not the dossier's list alone) is rewired to
@agent-device/ad-script.

Refs #1478

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor: trim the ad-script façade to real consumers, lock the one-export boundary

- packages/ad-script/src/index.ts: drop parseReplaySeriesFlags,
  formatTargetAnnotationCommentLine, parseTargetAnnotationCommentLine,
  TargetAnnotationLineParseResult, and TargetRect from the public façade —
  none has a consumer outside the package (re-swept every remaining export
  by grep; everything else kept has at least one real external importer).
  The functions/types stay exported from their declaring internal modules
  for the package's own internal use (script.ts, script-formatting.ts).
- scripts/layering/package-boundaries.test.ts: add the parallel R11
  assertions "the real tree parses, declares, and passes R11" already makes
  for maestro/provider-webdriver/provider-limrun/xml — ad-script exports
  exactly `.`, depends on exactly contracts+kernel, and is declared in root
  package.json — plus ad-script entries in the deep-resolution rejection
  coverage. Verified the lock catches a regression: temporarily added a
  fake `./codec` export to packages/ad-script/package.json and confirmed
  both the export-key-list assertion and the deep-resolution-rejection
  assertion fail; removed the plant and reconfirmed green.

Refs #1478

Co-Authored-By: Claude <noreply@anthropic.com>

* fix: remove polynomial-redos ambiguity from the target-v1 annotation line regex

CodeQL js/polynomial-redos flagged TARGET_ANNOTATION_LINE_RE
(packages/ad-script/src/internal/target-annotation-serde.ts): the payload
group's `\s+(.*)` let `\s+` and the unconstrained `.*` both match whitespace,
so a run of separator whitespace that ultimately fails to complete the match
has many `\s+`/`.*` splits to backtrack through before concluding failure.

Anchor the payload group on `\S` (the exact complement of `\s`), so the
mandatory `\s+` separator and the payload's first character can never
overlap — the split point becomes unique and no backtracking is possible.

Behavior-preserving: the only caller (parseTargetAnnotationCommentLine)
always matches against an already-.trim()-ed line, whose last character
(whenever the tag matches at all) is never whitespace — so a payload section
`\S.*` would reject (content that is entirely whitespace) can never reach
this regex through the real call path. Verified against the frozen
replay-compat corpus and the full serde/parser test suites, unmodified.

Added a regression test with the exact adversarial shape CodeQL/the reviewer
cited (many tab pairs after the version digits), asserting sub-second parse.

Refs #1478

Co-Authored-By: Claude <noreply@anthropic.com>

* test(ad-script): pin the annotation-line pattern's linear rejection directly

The entry-point adversarial case matched greedily even with the retired
regex (trim strips edge whitespace and per-line input carries no newline),
so it proved nothing about the pattern. The regression surface is the
pattern itself: an interior tab run with an x-newline tail fails the match,
which the retired form re-split quadratically (3.7s at 100k tabs) and the
\S anchor rejects in one attempt.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-01 20:21:38 +02:00
Michał Pierzchała 26246ea41e fix: memoize snapshot occlusion coverage checks to stop a daemon CPU-spin (#1541)
* fix: memoize snapshot occlusion coverage checks to stop a daemon CPU-spin

Root-caused a deterministic daemon hang reported against this branch's
`.ad` test/replay path (a two-fill Android form wedges the daemon at
~99% CPU indefinitely, blocking the checkout-form-android.ad live evidence).

Mechanism: `annotateCoveredSnapshotNodes` (src/snapshot/snapshot-occlusion.ts)
asks, for every overlay-classified candidate cover, whether THAT candidate is
itself covered by something later — via a recursive call back into
`findCoveringNode` (through `visibleCoverRect`). That recursive question was
never memoized: resolving position P's answer required resolving every later
position Q > P from scratch, and resolving Q required resolving every
position after IT from scratch again, giving O(2^overlayPositions.length)
work with no bound. A live CDP pause on the wedged daemon (`kill -USR1`,
`Debugger.pause` over the inspector) landed repeatedly inside exactly this
recursive triad (`findCoveringNode` -> `canCoverPoint` -> `visibleCoverRect`
-> `findCoveringNode`), matching the reporter's own `sample` profile
(role-normalization / `normalizeType` hot, called from inside this loop).

The pathological input is real, not synthetic: the second `fill` in a
two-field Android form runs while the on-screen IME keyboard is open, and
each individual key is classified `isAdditionalOverlayNode` — roughly 40
mutually-adjacent "overlay-like" nodes, confirmed by instrumenting the
function directly against the live repro (`nodes=62 overlayPositions=39`
right where the daemon stops responding). A/B against the p4a branch head
with matching instrumentation shows the equivalent snapshot there carries
zero overlay-classified nodes at the same point in the script and completes
in ~7s; extending the gap between fills with a genuine (non-instant) real
wait on p4a does not reproduce the 39-overlay state either, ruling out a
pure timing race as the sole explanation. `snapshot-occlusion.ts` itself is
untouched by the codec extraction, so the exponential blowup is a
pre-existing latent algorithmic defect — this PR's consumer rewiring is the
first thing to reliably land the fill-resolution snapshot in the
39-overlay-node regime; the exact mechanism connecting the codec/target-
identity import changes to that timing shift was not pinned to a single
line, and is called out as a residual question in the PR body.

Fix: cache `findCoveringNode`'s answer per position on the scan object,
scoped to one `annotateCoveredSnapshotNodes` call. The scan's own node list
is immutable input for the duration of one pass (byIndex is only ever
extended forward, never revised for a position already resolved), so a given
position's covered-by-something-later answer is provably stable across every
path that asks it — caching turns the unbounded double recursion into O(K)
resolutions of O(K) work each, i.e. O(K^2) instead of O(2^K).

Regression test (src/snapshot/__tests__/snapshot-occlusion.test.ts):
constructs a synthetic 40-node "keyboard" (mutually non-overlapping,
same-kind overlay-classified nodes, matching the live scale) and asserts
`annotateCoveredSnapshotNodes` returns well under a second. Verified the test
actually catches the regression: with the memoization reverted, the same
test times out (never returns) instead of failing an assertion — it hangs
exactly like the daemon did. Two existing-behavior sanity cases (a covered
touch target, an uncovered one) guard against a memoization bug silently
changing output.

Live verification: `node bin/agent-device.mjs test /tmp/m6.ad --platform
android` (the reported minimal repro) now passes in ~7.5s with no orphaned
daemon, down from a 180s timeout at ~99% CPU. The full two-script run
(`checkout-form-android.ad` + `gesture-lab-android.ad --platform android`)
passes in ~75s with no orphan.

Refs #1478

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(snapshot): make occlusion decisions read immutable input only

The memoized pass cached findCoveringNode by position while the annotation
loop mutated scan.nodes and scan.byIndex, so a cached answer could predate
annotations the caller predicate or ancestor classification would observe —
first-evaluation-wins order dependence (present, unmemoized, in the original
too). Decisions now evaluate exclusively against the caller's input; covered
positions are collected read-only and annotations applied in a separate
output pass that never feeds back.

Two invariants pinned: the input array and its nodes are never mutated, and
a chain (target under a covered sheet under a dialog) resolves identically
regardless of evaluation order.

Co-Authored-By: Claude <noreply@anthropic.com>

* test(snapshot): pin annotation-blindness through a mutation-sensitive predicate

The prior invariant tests pass against the mutable implementation too (it
copied the array up front, and the chain case never exercised the caller
predicate). This one fails against it: a predicate that also matches
annotated nodes would, through the ancestor walk over a mutable byIndex,
declassify a child overlay mid-pass and flip a later target's outcome.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-01 15:12:03 +02:00
Michał Pierzchała 4c2a30cc8f docs(agents): a green check is evidence only once you have seen it red (#1547)
* docs(agents): a green check is evidence only once you have seen it red

Three vacuous regression tests shipped in one day (an edge-run input the
retired regex handled in one pass, invariants the old implementation
already satisfied, an entry point whose trimming defused the flagged
pattern); review's counterfactual checks caught all three. The same proof
discipline already existed piecemeal for moved tests and structural gates —
name it once and point to the mechanical proof shapes.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(agents): trim the obvious, keep the earned

Dropped three bullets: open-before-close (CLI help and
device-verification.md own it), don't-remove-without-migration (subsumed by
the stronger no-fallback scope rule), and generic Node built-ins advice
(engines owns the version). Strengthened the oxfmt rule with the confirmed
mechanism: a path argument bypasses ignorePatterns, not just hides drift —
one path-scoped run re-quoted 44 excluded conformance corpus files.

Co-Authored-By: Claude <noreply@anthropic.com>

* Revert "docs(agents): trim the obvious, keep the earned"

This reverts commit 0229cbadba.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-01 15:11:26 +02:00
Michał Pierzchała e88b50f75a fix: clear the polynomial-redos class across main (#1546)
* fix: clear the polynomial-redos class across main

Three sites of the same CodeQL js/polynomial-redos family:

- packages/replay-test session-test-artifacts/-discovery slugs trimmed edge
  dashes with /^-+|-+$/g, which backtracks polynomially on long dash runs
  built from caller-supplied paths (alerts #27/#28). Replaced with a shared
  linear trimEdgeDashes.
- src/replay/target-identity.ts's target-v1 annotation line regex had the
  \s+(.*) ambiguity (the shape flagged as alert #29 on the #1536 copy).
  Anchored the payload group on \S so the split point is unique; the only
  caller matches against trimmed lines, so behavior is unchanged.

Adversarial regression test on the slug path (100k-char dash run,
sub-second); the annotation-regex adversarial case is covered on the #1536
package copy and the frozen replay-compat corpus passes here unchanged.

Co-Authored-By: Claude <noreply@anthropic.com>

* test: make the redos regression fail against the retired regex form

The edge-run input matched the old /^-+|-+$/g in one pass; the quadratic
case is an interior run (each dash restarts a -+$ attempt that fails at the
trailing byte). The slug pipeline collapses runs before trimming, so the
test targets trimEdgeDashes directly and asserts the input comes back
byte-identical.

Co-Authored-By: Claude <noreply@anthropic.com>

* fix: drop the import the test rewrite orphaned

Co-Authored-By: Claude <noreply@anthropic.com>

* test: pin the all-dash fallback identifiers

Artifact slug falls back to 'test', invocation id to 'suite', and a
session-name slug that trims to nothing is omitted without a dangling
separator.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-01 15:11:09 +02:00
Michał Pierzchała ef66dcdf25 refactor(daemon): serialize replay transactions behind a locked coordinator (#1478 P4b) (#1535)
* refactor(daemon): serialize replay transactions behind a locked coordinator

Adds session-replay-coordinator.ts, a ReplayCoordinator scoped to one
locked native .ad replay request, and routes every repair-transaction
write session-replay-runtime.ts and session-replay-resume.ts perform
through it: arm, demote-for-rerun, mark-complete, hold-on-divergence
stamping, the pendingRecordAndHeal corrective watermark (set + clear),
and reap-tombstone clearing. Neither file imports
session-replay-transaction.ts (P4a's ReplaySessionTransaction) or
writes session.pendingRecordAndHeal directly anymore.

Adds a minimal immutable ReplaySessionView (repairBoundary,
pendingRecordAndHeal) so the three readers this slice touches
(preflightReplayAgainstActiveRepair, isRepairArmedTerminalClose, the
entry-index resolution in prepareReplayPlan) stop taking mutable
SessionState.

Close-time sequencing (session-close.ts's platform-close receipt,
session-close-script.ts's commit/abort) stays a direct
ReplaySessionTransaction caller by design: commit/abort happen at
teardown, ordered against platform close and lease release, not
during a replay request.

Updates the R7 session-state ownership registry: pendingRecordAndHeal
moves from session-replay-resume.ts to session-replay-coordinator.ts.
The daemon-modularity baseline (writer-owned fields / owner claims)
is unchanged.

Refs #1478

Co-Authored-By: Claude <noreply@anthropic.com>

* docs: state the coordinator constraint, not the migration

Co-Authored-By: Claude <noreply@anthropic.com>

* fix(daemon): thread one bound resume-stamper instead of a second coordinator

buildAndPersistReplayDivergenceResume (session-replay-resume.ts)
constructed a SECOND ReplayCoordinator from a bare SessionStore +
session name, reachable from both divergence paths
(session-replay-target-verification.ts and the action-failure chain
through session-replay-runtime-failure.ts / session-replay-divergence.ts).
That let a lower handler manufacture repair authority by naming a
session instead of using the request's own locked coordinator.

Adds ReplayResumeStamper: a narrow capability bound to the coordinator
runReplayScriptFile already created, exposing only sessionExists() and
stampCorrectiveWatermark(). Threads it through ReplayStepContext and
the failure-wrapper params into both chains.
buildAndPersistReplayDivergenceResume now takes the stamper and holds
no SessionStore or coordinator-construction ability at all.

Adds src/daemon/__tests__/replay-coordinator-ownership.test.ts, an
oxc-parser AST structural test (same approach as
scripts/layering/session-state.ts) asserting: createReplayCoordinator
has exactly one production call site
(session-replay-runtime.ts); none of the five divergence-chain files
import the coordinator factory or session-replay-transaction.ts;
session-replay-resume.ts holds no session-store.ts import at all; and
the other four hold SessionStore only as a type. Verified the test
fails on a planted violation of each of the two structurally-distinct
invariants (coordinator-construction, SessionStore value-import) and
passes once removed.

Refs #1478

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-01 13:51:02 +02:00
Michał Pierzchała 8a6ddbc11d fix(test): repair Android replay fixtures against live device reality (#1538)
* fix(test): repair Android replay fixtures against live device reality

Three Android fixture defects from the #1482/#1484 full-tier suite, none of
which ever executed in CI (both nightlies since failed on adb infra before
the suite ran). All three verified live on a fresh API 36 emulator with a
pixel_7-geometry AVD and a Release fixture APK:

- 01-navigation-scroll.ad clicked label=Catalog, but the expo-router
  NativeTabs cart badge leaks '0 new notifications' into the tab's content
  description even while hidden, and unselected native tabs expose no child
  text node - exact match can never hit. Target the composed label the
  device actually exposes (deterministic at fixture start: cart is 0 after
  --relaunch). The badge does not leak on iOS, so the iOS twin keeps
  label="Catalog".

- checkout-form-android.ad opened by iOS display name 'Agent Device
  Tester'; Android open resolves packages (the APK label is
  'Agentdevicelab'), so APP_NOT_INSTALLED was guaranteed. Use the package
  id, matching gesture-lab-android.ad.

- gesture-lab-android.ad aimed every gesture at y=700, above the gesture
  card (its targets span y754-1329 on pixel_7 geometry; the home screen
  gained content above the card since authoring). Re-aim pans inside the
  exact-two-pointer zone, flings on the image clear of that zone, and
  pinch/rotate/transform at the card center. Verified: full suite passes
  2/2 via the public test command (20 + 32 steps replayed).

Refs #1478

* docs(test): pin the Android gesture fixture's validated emulator geometry

The re-aimed coordinates are validated on CI's profile (pixel_7 1080x2400
@420); any booted emulator can receive them via test-app:replay:android, so
the fixture and README now say which geometry the numbers mean and what a
mismatch failure looks like. The checkout twin is selector-driven and
unconstrained.

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-01 13:50:48 +02:00
Michał Pierzchała e9407284b5 test(daemon): pin resolveScriptTarget's retention table directly (#1539)
The repair suites pin these rules transitively (they caught the bare-re-arm
collapse during the P4a migration), but only through healed-sibling path
assertions downstream. A direct table on the transition makes a regression
name the retention rule it broke.

Refs #1478, #1258

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-01 11:27:02 +02:00
Michał Pierzchała 67f3d09d95 refactor(daemon): session script publication behind one capability (#1478 P4a) (#1532)
* refactor(daemon): add the tagged script-publication aggregate

First step of P4a. Nine co-resident optional SessionState fields encode two
lifecycles plus a shared output target, with nothing in the shape saying the
lifecycles are disjoint — so readers re-derived that from field combinations
and writers had to remember which siblings to clear.

The aggregate makes both invariants structural: a session publishes nothing,
authors ordinarily, or is under repair; and force lives inside the target, so
retargeting replaces the authorization along with the path.

Three corrections after an adversarial review of the first draft:

- The target is a default|explicit union, not a mandatory path. A bare
  'open --save-script' arms with no path and lets the writer resolve a
  daemon-owned destination at write time, and force can be granted before any
  path exists. Eagerly materializing a default path would have silently changed
  retarget semantics, because today's check requires a previously persisted
  path — so 'open --save-script --force' then 'close --save-script=out.ad' is
  not currently a retarget and the grant survives. That behavior is preserved
  here and flagged in the docblock as a probable #1258 gap; tightening it is a
  product change and belongs in its own commit.

- The repair status relation is not linear. A failed commit followed by
  'replay --from' demotes complete back to armed, so demoteRepairToArmed exists
  and deliberately RETAINS the close receipt: the platform close already
  succeeded for that operation identity, and dropping it would re-dispatch a
  close on retry — which is also how a migrator ends up reaching for the
  caller-computed platformCloseSucceeded boolean the brief forbids.

- The receipt doc no longer claims it is set only at close-succeeded and later,
  since the demotion path makes {armed, receipt set} reachable.

Still to come in this PR: both projections, and the writer migration. Note the
brief's seven-file writer inventory omits session-open.ts, which holds the only
two writers of the authoring armed/aborted states.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(daemon): migrate script publication onto the tagged aggregate (#1478 P4a)

The eight co-resident SessionState fields (scriptRecordingState, saveScriptPath,
saveScriptForce, saveScriptBoundary, saveScriptComplete, saveScriptCommitted,
repairPlatformCloseReceipt, repairSourcePath) are gone; SessionState.scriptPublication
holds the aggregate, and every writer migrated in this commit — no shadow state.

Two daemon-private projections own the writes, enforced by the R7 ownership gate:

- session-replay-transaction.ts (ReplaySessionTransaction): repair arm/demote/
  complete/abort, close receipts, and the uncommitted/boundary/sourcePath reads
  that idle-reap, tombstones, divergence-hold, and the recorder's exclusion key off.
- session-script-publication-capability.ts (SessionScriptPublication): authoring
  arm on open, the recorded --save-script flag ingress, active publication, the
  published transitions, and the effective per-target force decision (#1258).
  The writer keeps the commit transition so idempotence stays colocated with the
  atomic publish.

Failure/retry transitions pinned as the brief requires: platform-close failure
leaves state unchanged (no receipt, retry re-dispatches); publication failure
retains target+force+receipt (same-identity retry skips close dispatch); committed
and aborted are explicit terminal states that drop the receipt.

Design decisions resolved:

- Force retention across a default->explicit retarget is preserved as-is and
  still flagged in resolveScriptTarget's docblock as a probable #1258 gap;
  tightening it stays a separate product change.
- The never-armed 'close --save-script' whole-log publication folds into the
  authoring lifecycle (armed at the recorded close, published in the same
  request) instead of a fourth variant: every close path that reaches the write
  deletes the session, so the transient armed state cannot leak into
  'session save-script' eligibility, whose not-armed-before-this-journey
  rejection is untouched.

One real bug caught by the migrated tests and fixed in resolveScriptTarget: a
bare (pathless) re-arm collapsed an already-materialized explicit target back to
the daemon default, wiping the healed-sibling path on every per-step repair
re-arm and defeating the persisted-force preflight bypass. A bare re-arm now
keeps the previous target and only adds a live force grant.

R7 rows consolidated to one scriptPublication entry (three owners) and the
recordSession row narrowed; the R10 baseline drops to 22 writer-owned fields /
28 owner claims so the consolidation cannot regrow.

Gates: typecheck, lint, format, layering clean; 624 files / 5220 tests pass
(two known contention-flake timeouts reproduce only under full-suite load and
pass in isolation).

Refs #1478

Co-Authored-By: Claude <noreply@anthropic.com>

* refactor(daemon): satisfy the Fallow gate by extracting decisions, not suppressing

- scriptPublicationTarget is module-private; both public target reads
  (scriptTargetPath/scriptTargetForce) go through it and nothing else did.
- validatePublicationEligibility splits into a pure ineligibility classifier
  and an error table, so the four rejections read as one decision each.
- prepareSaveScriptSession hands its two arm-time rejections (authoring
  re-arm, EEXIST preflight) to rejectSaveScriptArming and keeps only the
  demote-and-arm flow.
- The repair-record-exclusion provider scenario extracts its three phases
  (arm-and-hold, exclusion contrast, healed-script contract) into named
  helpers; the test body is the journey again.

Refs #1478

Co-Authored-By: Claude <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-01 11:10:09 +02:00
Michał Pierzchała de1654127e fix(cli): drop the removed durationMs positional from swipe help (#1534)
swipe --help still advertised a trailing [durationMs] positional
that #1393 stopped accepting, so agents following the help text hit
the runtime's migration-hint rejection. Remove it from the swipe
command's help-schema positionals and defer arity enforcement to
swipePayloadFromPositionals so the migration-hint error keeps firing
unchanged.

Refs #1393

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-01 07:45:27 +02:00
Michał Pierzchała 9fea8ffde1 fix: treat Android permission prompts as pending alerts, not tap escapes (#1530)
* fix: treat Android permission prompts as pending alerts, not tap escapes

A click that raises a system permission dialog (e.g. the lifecycle
scenario's automation-request-microphone) previously failed the
assertAndroidPressStayedInApp escape guard with COMMAND_FAILED, so the
harness never reached its alert get / alert accept steps. But raising
the dialog is the intended press outcome, and alert-detection already
treats the permission-controller packages as a legitimate alert source.

The guard now returns a response warning for permission-prompt packages
(shared isAndroidPermissionPackage authority, which also covers the AOSP
permissioncontroller/packageinstaller packages the guard previously
missed) and keeps throwing for genuine escapes (settings, systemui,
launcher). The warning tells the agent to consume the dialog with
alert get / alert accept / alert dismiss, and tap/press/fill/longpress
plain-text CLI output now prints response warnings so plain-CLI agents
see it too (previously data.warning was JSON-only).

* style: format interaction.test.ts

* test: cover Android package installer prompts
2026-07-31 21:22:24 +02:00
Michał Pierzchała d0859460b9 fix(ios): diagnose a runner that cannot install, instead of blaming the screen (#1529)
* fix(ios): diagnose a runner that cannot install, instead of blaming the screen

A physical device that is not covered by the runner's provisioning profile
fails to install the XCTest runner. Every runner-backed command then failed
with 'the current screen is overwhelming the iOS accessibility capture' and
advice to run screenshot instead — which fails identically, because it needs
the same runner. The suggested remedy could never work and the stated cause
was never observed.

Classify the install failure and say what it is: the profile does not cover
this device, register it with the signing team. It is deliberately not an
infrastructure reason, because no retry registers a device.

Matching anchors on the CoreDevice error code and the English framework
strings; the installer prose around them is localized by macOS, so the real
log arrived partly in Polish.

The recycle-budget error also stops asserting a cause it never observed. It
now points at the runner log first and offers the heavy-screen reading second,
which is where it belongs — that case is real, it just is not the only one.

Observed on an iPhone that could not be added to the signing account.

* fix(ios): carry the classified reason into the early-exit hint

buildRunnerEarlyExitError classified the provisioning failure correctly, then
built its hint through resolveRunnerEarlyExitHint, which ignored the reason and
always fell back to connect-timeout and cache-recovery guidance. The shipped
error therefore still told people to retry a runner that can never install,
which is the misdiagnosis the previous commit set out to remove.

Thread the reason through, keeping the busy-connecting device special case, and
withhold the cache-recovery sentence for a provisioning failure: clearing
derived data cannot put a device into a profile.

The previous commit only tested classifyBootFailure and bootFailureHint, never
the function that assembles the error a user receives. The regression added
here exercises that production route on the captured xcodebuild output.

* fix(ios): narrow device provisioning diagnosis
2026-07-31 21:22:05 +02:00
Michał Pierzchała cbe1a57094 refactor(replay-test): extract packages/replay-test (#1478 P3b) (#1525)
* refactor(replay-test): source the manifest device vocabulary from the kernel

`session-test-types.ts` reached `ReplayScriptMetadata['platform']` and
`['target']` through `replay/script.ts` — the native `.ad` engine. A
format-neutral scheduler must not name an engine module, and P5 relocates that
engine into `packages/ad-replay` regardless, so the import had to go before the
scheduler can move.

Both members already resolve to neutral kernel types
(`Exclude<PlatformSelector, 'web'>` and `DeviceTarget` from
`@agent-device/kernel/device`), so this re-sources them directly and the
manifest shape is unchanged. Only the import direction differs.

First increment of P3b; the scheduler still has request-global, engine and
daemon imports to port before the physical move.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(replay-test): inject the progress sink instead of reading a request global

The scheduler called `emitRequestProgress` in eight places, which reads a sink
out of a request-global `AsyncLocalStorage`. That is ambient authority a
format-neutral scheduler cannot hold once it lives in `packages/replay-test`,
and #1505 recorded it as a shrink-only R10 entry.

The host now injects the capability through the existing
`ReplayTestRuntimeDependencies` seam established in P3a, so no new seam is
invented. `session-replay.ts` supplies `emitProgress: emitRequestProgress`;
`src/request/progress.ts` keeps the sink and its AsyncLocalStorage binding for
every other caller.

The port is deliberately narrower than `RequestProgressSink`: it accepts only
`ReplayTestSuiteProgressEvent | ReplayTestProgressEvent`, so the scheduler is
not handed the ability to emit `CommandProgressEvent`.

Authority narrows again one hop down: `runReplayTestAttempt` spread the whole
dependency bag but uses three of its members and never publishes progress, so
it now takes `Pick<..., 'runReplay' | 'cleanupSession' | 'finalizeAttempt'>`.
That is why no runtime test fixture needed changing — the attempt runtime never
gained the capability in the first place.

Also drops the last two `replay/script.ts` type references from
`session-test-runtime.ts`, so the engine import is gone from that file too.

Reporter contract preserved: `session-test-reporter-values.test.ts` and
`session-test-reporter-values-maestro.test.ts` both pass unmodified (27 tests
green across the five scheduler suites). Typecheck clean.

Remaining scheduler boundary for P3b: `request/cancel.ts`, `replay/format.ts`,
`replay/script.ts` in discovery, `session-store.ts`, `daemon/types.ts`,
`replay-source-discovery.ts`, `core/dispatch*`, `utils/diagnostics.ts`.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(replay-test): ask the host whether the suite is canceled

The scheduler called `isRequestCanceled(requestId)` in five places. That both
reaches a request-global registry and forces the scheduler to name a daemon
request id as the cancellation key — neither survives the move into
`packages/replay-test`.

The host now binds the predicate to its own request and passes
`isCanceled: () => boolean`. The scheduler asks a question it is entitled to
ask and learns nothing about how cancellation is tracked. `shouldStopReplayTestExecution`
takes the capability rather than a request id, so no scheduler function threads
a daemon identifier for this purpose any more.

`session-test-attempt.ts` and `session-test.ts` no longer import
`request/cancel.ts` at all. It remains in `session-test-runtime.ts`, which does
something different — `registerRequestAbort`, `markRequestCanceled` and the
parent-abort relay are cancellation *binding*, which the brief assigns to the
daemon adapter, so that split is its own step.

Behavior preserved: both pinned reporter characterizations pass unmodified,
32/33 across the five scheduler suites. The one failure is the pre-existing
P2/#1506 discovery-ordering regression, unrelated and untouched here.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(replay-test): drop the dead request-tracking call from attempt ids

`buildReplayTestAttemptRequestId` wrapped its template in
`resolveRequestTrackingId`, pulling `request/cancel.ts` into the scheduler.

That wrapper substitutes a generated id only when its first argument is an
empty string. The template here always contains `:test:`, so it is never empty
and the wrapper always returned it unchanged — the call is unreachable in this
path. Probed all three shapes (explicit request id, suite-id fallback with a
shard, and degenerate empty inputs); every one returns the template verbatim.

Removing it takes `request/cancel.ts` out of discovery without altering a
single produced id. The scheduler mints attempt identity itself, which is what
the brief asks for.

Evidence the ids are byte-identical: the pinned reporter characterizations
assert exact session strings such as
`default:test:suite-reporter:1-02-retry:attempt-1` and pass unmodified —
30 tests green across the reporter, suite and discovery suites.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(replay-test): move cancellation binding and diagnostics to the host

`session-test-runtime.ts` held the last two request-globals in the scheduler:
`request/cancel.ts` (registerRequestAbort, markRequestCanceled,
clearRequestCanceled, plus the parent-abort relay) and `utils/diagnostics.ts`.

These are different in kind from the earlier ports. The brief gives the daemon
adapter the job of mapping an attempt id to daemon request identifiers and
*binding cancellation*, while timeout policy stays scheduler-owned. So the
scheduler now receives a per-attempt capability with exactly two verbs —
`cancel()` on timeout and `release()` when the attempt settles — and every
registry interaction, including `relayReplayTestAbortFromParent`, moved to
`session-replay.ts` next to the rest of the adapter.

Diagnostics became a narrow publish capability for the same reason:
`emitDiagnostic` reads a request-global scope. The level vocabulary is spelled
out at the seam rather than imported, so nothing engine- or daemon-shaped
crosses it.

The runtime fixtures drive the real exported host binding rather than a stub.
They assert cancellation through `isRequestCanceled`, and a stubbed binding
would have kept those assertions passing while proving nothing.

24 tests green across the runtime, suite and both reporter characterizations,
which pass unmodified. Typecheck, lint and oxfmt clean.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(replay-test): split discovery into host inspection and scheduler policy

discoverReplayTestEntries expanded paths, read every file, and called both
engines — readReplayScriptMetadata for .ad, inspectMaestroFlow for Maestro —
plus resolveReplayFormat to choose between them. Four imports a format-neutral
scheduler cannot hold.

Inspection is now the host's discoverSources capability. What stays in the
scheduler is the genuinely neutral half: which sources a --platform filter
runs, which it skips and with what message, and the empty-suite error.

The manifest carries exactly the four fields the scheduler consumes (platform,
target, retries, timeoutMs) plus the reporter's title, per the brief's
instruction not to add more without a demonstrated call site.

The platform tag is what removes the last format leak. The filter used to ask
resolveReplayFormat(...) === 'maestro' to decide whether a missing platform was
disqualifying. It now reads a tag: caller-bound means the invocation supplies
the platform, unspecified means the source declared none. Maestro is what
caller-bound looks like from the scheduler's side, and the format cannot be
recovered from it.

Discovery tests drive the real inspection capability, writing actual .ad and
Maestro sources — a stubbed host half would have kept them green while proving
nothing about the composition they exist to pin.

35 tests green across discovery, suite, runtime and both reporter
characterizations, which pass unmodified. The Maestro one is the direct check
that titles still flow, since they now arrive via the manifest.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(replay-test): build attempt ids from named segments; trim comments

Review feedback on the attempt-id builder: the comment explained a deletion
that git already records, and it sat above an opaque template literal.

The id is now a segment list joined on ':', so its shape is readable without
prose. Output is byte-identical — the reporter characterizations assert exact
session and attempt strings and pass unmodified.

Applied the same standard to four other docblocks in this PR that narrated
what the code used to do rather than what it does. The durable 'why' stays:
which side of the seam owns what, and why the vocabulary is neutral. The
migration history goes, since git carries it and these docblocks will outlive
the migration.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(replay-test): move shard device binding to the host

buildReplayTestShardPlan called listDeviceInventory to discover what to shard
across, and buildReplayTestShardFlags constructed daemon CommandFlags for the
nested request. Inventory enumeration, allowlists, simulator set paths,
explicit --device selectors and the too-few-devices error are host concerns;
what is scheduler-owned is deciding how many shards exist and which entries
each one runs.

The scheduler now receives resolved shard targets through a capability. The
target is neutral: id and name for session labels and progress metadata, plus
platform and target, which are already kernel vocabulary. DeviceInfo no longer
crosses into scheduling.

One behavior note: an explicit --device selector could in principle name a web
target, which is not a shardable device. That is now rejected with INVALID_ARGS
rather than widening the neutral platform vocabulary to carry something the
scheduler can never run. Implicit selection already filtered to mobile.

919 of 920 handler tests pass. The one failure, session-test-runner.test.ts
'binds each replay script to its declared platform metadata', fails identically
on clean origin/main in this container and is unrelated: directory discovery
walks with opendirSync/readSync and directory results are deduped but not
sorted, while glob results are sorted, so suite order is filesystem-dependent.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(replay-test): extract packages/replay-test behind a façade

Completes the P3b extraction. The scheduler, attempt runtime, discovery
policy, sharding distribution, artifacts and neutral types now live in
packages/replay-test/src/internal/, with one package-root export.

The façade takes a neutral ReplayTestSuiteRequest and returns a tagged
ReplayTestSuiteOutcome. DaemonRequest, DaemonResponse and CommandFlags no
longer reach the scheduler; the adapter translates flags and meta in, and the
outcome back to a daemon response. Eight flags were read by the scheduler and
each became a field it owns.

Host work moved to daemon adapters: source inspection (both engines and format
routing), shard device binding and shard-flag parsing, and artifacts-dir home
expansion, which is why the package can resolve paths without SessionStore.

The one remaining shared concern was the timing trace: the host writes video
lifecycle events into the same trace the scheduler owns. Rather than export a
writer from the façade, each attempt hands the host an appendTimingEvent
closure, so the trace format stays private and the authority is scoped to that
attempt.

Tests mirror the topology. Discovery tests split along the seam they now
cross: ordering, traversal and routing are pinned host-side against real files,
filtering policy is pinned in the package against fake sources. The runtime
tests assert the scheduler's cancellation obligation (cancel once on timeout,
always release) against a recording binding, and a new daemon test pins the
adapter's half — registry entries, the parent-abort relay, and detach on
release — so that coverage moved rather than disappeared.

R10 retargeted to packages/replay-test/src/ and the zone ranked alongside
maestro. R11 confirms zero root-src imports from the package.

914 of 915 handler and package tests pass. The one failure,
session-test-runner 'binds each replay script to its declared platform
metadata', fails identically on clean main here: directory discovery walks with
opendirSync and dedupes without sorting, while globs sort, so suite order is
filesystem-dependent.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* refactor(daemon): simplify replay-test request translation

Fallow flagged toReplayTestSuiteRequest at 14 cyclomatic in 18 lines. The
branches were self-inflicted: every req.flags?.x is one, and each optional
field was written as a conditional spread to avoid setting an undefined key.

exactOptionalPropertyTypes is not enabled, so assigning undefined to an
optional field is equivalent and the spreads bought nothing. Destructuring
flags once and extracting two flag readers removes most of the rest.

One correctness note on the simplification itself: the first version used
`artifactsDir && expandHome(...)`, which returns '' for an empty-string flag
where the previous code called expandHome(''). Replaced with an explicit
undefined check so the empty-string path is unchanged.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* test(live): share the replay test-suite harness across iOS and Android

Both live journeys invoked the public test command and then re-derived the
same value-contract assertions by hand — suite totals, per-script status,
replay counts, non-empty JUnit. Those are claims about the published suite
result and are identical on every platform, and they had already drifted: iOS
iterated with readReplayCommands inline, Android cast data.tests at the call
site.

The shared helper owns exactly that boundary. It takes the caller's runStep
rather than binding a context type, so it is not a platform-configured runner
and cannot template a platform's journey.

Everything a platform genuinely differs on stays with the caller: which
scripts run, the retry policy (iOS 2, Android none — itself a claim worth
keeping), which commands each script exercises, and the behavioral evidence.
Both callers keep every verify* call they had.

67 lines removed, 15 added.

Residual risk: this container has no iOS or Android devices, so the live suites
could not be executed here. Typecheck and lint pass; the harness needs a run on
real targets before the claim that behavior is unchanged is evidence rather
than inference.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* test: pin directory enumeration in the platform-binding suite test

The test wrote two scripts into a temp directory and assumed discovery would
return them in creation order. Directory expansion deliberately preserves
filesystem order to match Maestro — only glob expansion sorts, and 'preserves
Maestro directory filesystem order' pins that with a mocked opendirSync. So the
ordering contract is correct; this test's assumption about enumeration was not.

It passes on CI, where small directories usually enumerate in creation order,
and fails on filesystems that do not — identically on clean main, where the
platform-to-script binding appears reversed.

Pinning enumeration the way the discovery tests already do keeps the subject
intact (each script binds to ITS declared platform, and session numbering
follows discovery order) without depending on the host filesystem. The fs
import became a default import because vi.spyOn cannot redefine an ESM
namespace export.

915 of 915 handler and package tests now pass here.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* test: scope the enumeration spy and restore it in a finally

The spy I added restored only on the happy path and asserted on its argument
inside the mock implementation. Either would misfire for anything else sharing
the worker: an assertion thrown from inside fs, or a leaked global opendirSync,
surfaces as a worker crash with no failed test rather than a readable failure.

It now delegates to the real implementation for any directory but this suite's
own, and restores in a finally.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* fix(replay-test): put package tests where they are actually run

Review found the moved package tests were neither executed nor typechecked.
They sat under packages/replay-test/test/, but vitest's unit-core lane includes
packages/*/src/**/*.test.ts, and neither the root nor the package tsconfig
covers a top-level test directory. A plain unit-core run discovered zero files
under the package.

That is why they looked green: my earlier runs passed those paths explicitly on
the command line, which masked that the default run skipped them. The count is
the proof — 550 files/4741 tests before, 553/4753 now, and the delta is exactly
the three files and twelve tests that were being skipped.

The runtime test also imported runReplayTestAttempt from the package specifier,
which the facade does not export. It would have failed the moment it was
discovered. It now imports internally, like the rest of the internal tests.

Also removed replayTestAttemptFailure from the facade: zero consumers outside
the package, so exporting it widened the boundary for nothing. P3 asks for a
one-function facade.

553 test files and 4753 tests pass; lint and the layering guard are clean.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* style: format the facade after removing the unused export

A scripted edit removed the export line but left a stray blank line; oxfmt was
not re-run on that file afterward, so Lint & Format caught what pnpm lint
alone does not.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* fix(replay-test): typecheck the package and fix a type-only import

Review found the moved package tests were transpiled by vitest but never
typechecked: the root typecheck script builds six packages via tsc -b and
packages/replay-test was not among them, so its tsconfig was never used.

That hid a real TS2459. session-test-runtime.test.ts imported
ReplayTestAttemptOutcome from ../session-test-runtime.ts, which imports that
type but does not re-export it. It now imports from ../session-test-types.ts,
where the type is defined.

Adding the package to the tsc -b list closes the gap. Verified empirically
rather than assumed: planting a string-to-number error in a package test makes
typecheck fail, and removing it makes it pass. This is the second finding of
the same shape on this PR — first the tests were not discovered by vitest, now
they were not covered by typecheck — so the gate was confirmed to reach the
files rather than trusted to.

12 package tests pass, lint, format and the layering guard are clean, and
typecheck is clean with the package included.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-31 20:16:18 +02:00
Michał Pierzchała e669676bc8 test: fix Android observability scenario to Android contracts (#1531)
The full:observability-artifacts scenario (#1482/#1484) had never executed
end-to-end: both nightly Android Full Emulator Suite runs since merge died
on adb infra before the suite ran, and a live run fails deterministically
at its first perf assertion. Fixing that revealed four more latent
failures, each written against iOS or remote-daemon behavior the Android
live run does not have. Validated with two consecutive green full-tier
runs on a dedicated Pixel 9 Pro XL emulator.

- perf metrics: assert totalPssKb (Android's required meminfo field)
  instead of the Apple-only residentMemoryKb.
- presses: reveal the Quick-actions card with scroll steps before
  pressing home-open-catalog/home-open-settings — Android snapshots only
  contain on-screen nodes — and restore scroll top before waiting on the
  home title, since scroll position persists across tab switches.
- batch get: target id="dismiss-notice" (a node that owns its text);
  Android resolves the home-title container to a child's text (the
  subtitle), unlike iOS's container label.
- events: run an explicit snapshot so the timeline assertion holds when
  the scenario runs standalone under AGENT_DEVICE_ANDROID_E2E_SCENARIOS.
- artifacts: assert the local-client contract — trace-log tracked,
  downloadable, consumed; screen-recording inventory entries only exist
  for remote clients (artifacts without a client localPath are never
  tracked). Also stop consuming the download response body in the assert
  message before arrayBuffer() reads it.
2026-07-31 19:48:45 +02:00
Michał Pierzchała f0fa81ad87 fix(ios): name Developer Mode and pairing as the real launch blockers (#1527)
A freshly paired iPhone that is unlocked, trusted and reported by Xcode as
available (paired) still cannot launch anything while Developer Mode is off.
The failure surfaced as a disk-image mount error carrying the default hint,
which tells the user to check that the device is unlocked, trusted and visible
in Xcode — all of which were already true. devicectl knows the actual cause and
says so: 'The operation failed because Developer Mode is disabled.'

Map that failure, and the unpaired one, onto hints that name what to do. The
pairing hint also mentions the device passcode, without which tapping Trust
leaves the device unpaired.

Observed on a real iPhone 13 that had never been used for development.
2026-07-31 18:04:54 +02:00
Michał Pierzchała 14b71c96f1 test: narrow Limrun public type contract (#1528) 2026-07-31 18:02:34 +02:00
Michał Pierzchała 4c1d62e2f7 test: stop two tests failing on timing and megapixels (#1526)
Both of these blocked unrelated PRs today and neither is auto-retryable:
#1419's contention retry declines find.test.ts as outside its enumerated
list, and slow-gate failures cannot be re-checked by a rerun. Each
occurrence costs a manual re-run.

find.test.ts 'wait captures fresh snapshots while polling' asserted exactly
2 dispatches. The wait loop polls at a 300ms interval against a 350ms
budget, and its last sleep consumes whatever remains, so it lands exactly on
remainingMs() === 0 — a sleep returning a millisecond early admits a third
poll. The count was never assertable. The test's actual subject is that each
poll re-captures instead of reusing the first tree, so it now asserts that:
at least two captures, every one a snapshot.

The sibling assertion at the top of the file is left alone: it takes its
count from mockResolvedValueOnce ordering, not from the clock.

apps.test.ts built 1206x2622 (iPhone 16 Pro, 3.2 megapixels) source PNGs to
prove two rescale ratios. Both sources are now 126x273 — still divisible by
3, so the /3 and 2/3 arithmetic stays exact — with expectations updated to
match. Measured on this machine: the density test 1021ms -> 34ms, the retry
test 1176ms -> 233ms, the file's test time 2.48s -> 1.00s.

Both suites pass: 18/18 find, 54/54 apps.


Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-31 16:16:09 +02:00
Michał Pierzchała 2e0ed0a41d fix: normalize physical iOS landscape taps (#1520)
* fix: normalize physical iOS landscape taps

* fix: preserve iOS selector tap fallback

* fix: normalize stale iOS screenshot dimensions
2026-07-31 15:36:21 +02:00
Michał Pierzchała da93191201 refactor: move Limrun provider behind package facade (#1518)
* refactor: move Limrun provider behind package facade

* fix: preserve Limrun public provider types

* fix: tighten Limrun provider facade boundaries

* test: harden Limrun compatibility coverage

* fix: narrow Limrun public type exports

* fix: narrow Limrun provider exports
2026-07-31 15:35:49 +02:00
Michał Pierzchała ac286d2a9e perf(devices): probe platform inventories concurrently (#1524)
An inventory lookup with no platform filter awaited each platform's toolchain
in turn, so it cost their sum. Opening a physical iPhone by name spent 6.7s in
resolve_target_device on a host with the Apple, Android and Vega toolchains
installed, most of it enumerating platforms the request could not target
(vega device list alone was 2.8s of it).

The probes are independent, so run them concurrently and concatenate in
selector order — the Linux local device still lands last, where it must be so
it does not displace connected Android/Apple devices in implicit selection.

Alternating A/B on one host against a cabled iPhone, three runs each:
sequential 4065/4015/3912ms, concurrent 2359/2365/2295ms.

A platform answering with a non-array still contributes nothing: spreading it
used to throw into the per-platform catch, and that is now explicit.
2026-07-31 13:49:06 +02:00
Michał Pierzchała 2ba6ecf599 fix(device): state when nothing is claimed and reap dead claims (#1519)
* fix(device): state when nothing is claimed and reap dead claims

Two defects found while asking which agent held a connected iPhone.

device status printed only "21 stale claims hidden" and no verdict: the
"No local advisory device claims found" line was gated on there being zero
stale claims too, so the one case where a user most needs to hear that nothing
holds the device is exactly the case that never said it. It now reports the
empty live set alongside the hidden-stale notice.

Nothing ever reaped claims whose owner died abruptly. Claims are released on
session close and daemon shutdown, but a killed process leaves its file
behind, and a real store had accumulated 21 of them spanning two weeks, every
owner dead. Daemon startup now prunes them, next to the existing web-browser
orphan cleanup.

Pruning is deliberately narrower than the CLI's stale filter: it removes only
owner-process-dead claims. owner-state-dir-gone describes a LIVE process whose
state dir vanished, and deleting that claim could hand its device to a second
session.

* fix(device): prune under the claim lock and record its diagnostics

Two review findings on the startup prune.

The scan classified a claim dead and then unlinked it, but claim paths are
derived from the device key: a concurrent daemon can prune the same dead claim
while a new session writes its live successor to that exact path, and the
unlink would take the successor with it. Liveness and owner token are now
re-checked while holding the per-device claim lock, matching what
clearAdvisoryDeviceClaim already does, and a file whose name is not the
canonical path for the key it contains is left alone.

Daemon startup runs outside any diagnostics scope, where emitDiagnostic returns
without recording, so neither a successful prune nor a failure produced the
promised event. The prune now opens its own scope and flushes, the way
emitFatalDiagnostic does.

* fix(device): prune after daemon info is published so its log survives

publishDaemonInfo truncates daemon.log, so the prune's diagnostic was written
and then wiped: every successful startup left no device_claim_prune event in
the log users are pointed at. The prune now runs after publication.

The regression starts the real runtime and reads the resulting daemon.log,
because the ordering is the bug — a test around the prune alone passes either
way.
2026-07-31 13:09:39 +02:00
Michał Pierzchała 6b972ae430 fix(ios): map usbmux connect result codes to the right verdict (#1523)
usbmuxd answers Connect with a result code, and every non-zero code collapsed
into one 'Failed to connect' error carrying a cable hint. Probing the daemon on
this host for the two codes that matter: an unknown DeviceID answers 2, and a
closed port on an attached device answers 3.

Result 2 means the device went away between ListDevices and Connect. It now
raises the same unattached verdict as a missing ListDevices entry, so a
CoreDevice device falls back to its network tunnel instead of failing with a
cable hint while Wi-Fi is available — the gap #1517 left open.

Result 3 means the device is reachable and only the runner port is not bound
yet, which is the normal state while the runner starts. Telling the user to
check the cable was wrong; it now says so.

Adds hermetic multi-device coverage for #1521: selection follows the UDID
rather than list position, and a UDID sharing a prefix with another device
never matches.
2026-07-31 12:30:25 +02:00
Michał Pierzchała b3cf29bc67 feat(ios): reach physical devices through usbmux first, tunnel as fallback (#1517)
* feat(ios): reach physical devices through usbmux first, tunnel as fallback

Physical iOS runner commands now resolve to usbmux whenever the device is
attached by cable, and fall back to the CoreDevice tunnel route only when
usbmuxd reports it unattached (#1403).

Measured on an iPhone 17 Pro: steady-state is a wash between the two routes,
but past the tunnel cache's 30s TTL the network route pays ~4.5s of devicectl
re-probe plus session re-establish on the next command, where the usbmux
session stays hot at ~440ms. Cabled devices now never pay that tax, because
the tunnel lookup, its cache, and the cache invalidation only run on the
fallback path.

Wi-Fi-only devices keep working: modern CoreDevice Wi-Fi runs over remoted and
never appears in usbmuxd, so the unattached verdict routes them to the tunnel.
That verdict is carried by usbmuxDeviceAttached:false and answered inside the
same connect attempt rather than by burning a retry, so an XCTest-backed
device — which has no tunnel — now surfaces its cable/trust/unlock hint
instead of retrying for the full budget.

Replaces the AGENT_DEVICE_IOS_RUNNER_ROUTE experiment override from #1510.

* fix(ios): make the unattached usbmux verdict terminal for xctest devices

waitForRunner recorded the unattached verdict as a generic connect failure and
retried it for the whole budget, so readiness preflight and read-only commands
on an XCTest device still burned 2x45s and lost the cable/trust/unlock hint —
the exact hang #1510 measured, which this PR claimed to fix but only fixed on
the sendRunnerCommandOnce path.

Retrying cannot attach a cable and an XCTest device has no tunnel to fall back
to, so the typed verdict is now thrown from the attempt, excluded from the
connect retry policy, and passed through waitForRunner unwrapped. The predicate
moved to runner-contract.ts because the retry policy needs it and importing the
transport there would close an import cycle.
2026-07-31 11:31:20 +02:00
Michał Pierzchała ab80fb696c test(replay-test): pin the retry-path reporter hint (#1516)
`emitReplayTestRetryProgress` publishes `hint` on every retried failure
(session-test-attempt.ts), and nothing asserted it. The existing hint coverage
coming out of #1505 pins only the final-attempt emit site, so deleting the
retry line left the entire suite green.

The retry fixture's first-attempt error carried no hint, so no assertion could
have caught it even in principle. It now carries one, and the retry result
value asserts it.

Counterfactual, run rather than assumed — with `hint: attempt.outcome.error.hint`
deleted from the retry emit site:

    ✓ a reporter sees the shipped suite-start, skip, and test-start values
    × reporter step and result sessions track the running attempt, not the start value
    ✓ a failing suite reaches the reporter with the failure message, hint fields, and exit code
    ✓ sharded runs give the reporter shard-scoped sessions and device identity
    -   "hint": "retry hint from the failed attempt",
    +   "hint": undefined,

Only the new pin fails; the final-path hint test keeps passing, which is the
gap. Production code is unchanged — `git diff` on session-test-attempt.ts is
empty.

This lands before P3b relocates session-test-attempt.ts into
packages/replay-test, so the move cannot drop the field silently.

Refs #1478


Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-31 10:13:01 +02:00
Michał Pierzchała 1f9f25ecc2 fix(android): re-capture past a transient helper content verdict (#1513)
A content verdict (system-window-only, content-poor-app-window,
empty-helper-output) means the capture mechanism worked but sampled the
screen mid-transition. That state resolves itself within a frame or two,
yet a single sample turned it into a hard command failure: the CI alert
dismiss in #1510's run failed because one capture landed while no
application window was attached.

Only polling waits rode these verdicts out (isUnreadableCaptureContentError);
every other command failed on the first mistimed sample. Re-capture a bounded
number of times before reporting the verdict, so the loud failure is preserved
for a screen that stays unreadable while a single mistimed sample no longer
fails a command.
2026-07-31 09:24:33 +02:00
Michał Pierzchała e3cdbcdb37 fix(test): give the iOS scroll search a capture-sized wait budget (#1514)
`assertElementTextAfterScrolling` waited 1000ms per attempt. The wait's whole
budget goes to its first capture, so that budget has to cover one snapshot of
the current surface; on a loaded simulator it does not, and every attempt fails
with `wait_capture_stalled` instead of reporting the element as off-screen.

This broke the iOS Smoke Tests lane on main. #1484 moved the smoke automation
scenario onto this helper for `automation-press` and `automation-longpress`;
before that it was only reached from the full lane, so the tight budget went
unnoticed. Every sibling wait in the same scenario already budgets 2500ms or
more (`label="Settings"` 10000, `alert wait` 5000, `wait text` 2500).

Two changes:

- Raise the per-attempt budget to 2500ms, matching the nearest sibling.
- Stop spending a scroll attempt on a stalled capture. A stall means the
  snapshot never came back, so the surface was never read — it is not evidence
  the element is off-screen, and scrolling on it moves the surface for an
  unrelated reason. Two stall retries absorb a slow runner; a genuine absence
  still consumes attempts and still fails.

The final assertion now carries the last wait's JSON, so a future failure says
whether it stalled or genuinely never found the element.


Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-31 09:10:44 +02:00
Michał Pierzchała b125435989 refactor: extract WebDriver provider package (#1504)
* refactor: extract webdriver provider package

* refactor: consolidate shared XML codec
2026-07-31 09:10:04 +02:00
Michał Pierzchała a3ab69a110 refactor(replay-test): neutralize the values crossing the scheduler seam (#1478 P3, part 1) (#1509)
* refactor(replay-test): neutralize the values crossing the scheduler seam

#1478 P3, part 1 of 2. Prepares the replay-test extraction by removing every
non-neutral value that crosses the scheduler seam, in place under `src/`, so the
physical move to `packages/replay-test` is a file move rather than a redesign.

`DaemonResponse` no longer crosses the seam. `session-test-types.ts` typed
`runReplay`/`finalizeAttempt` as returning a daemon response and the scheduler read
`.error.code`, `.error.details`, and `.data.replayed/.healed/.warnings/
.snapshotDiagnostics` off it throughout. That is invisible to R10 today only
because `checkDaemonTypesImporters` skips `src/daemon/`; once the files live in a
package they become external `daemon/types.ts` importers, which the ratchet only
lets shrink. Attempts now resolve as tagged `ReplayTestAttemptOutcome` values
carrying exactly what the scheduler consumes, including an `infrastructure` tag —
classifying an environmental failure needs platform boot-diagnostic vocabulary the
scheduler must not import, so the host decides and the scheduler reads the verdict.
`session-test-outcome.ts` is the one place a daemon response becomes an outcome.

Step events get a narrow per-attempt port. They were emitted from
`session-replay-runtime.ts` and `session-replay-maestro-observer.ts`, both reading
a request-global `AsyncLocalStorage` seeded per attempt. The scheduler now hands
each attempt an `onStep` sink, threaded the way `tracePath` already is; both
engines call it and `withReplayTestActionProgress`/`readReplayTestActionProgress`
are gone. A direct `replay` simply has no sink.

ADR 0012 divergence becomes a neutral leaf. `src/replay/divergence.ts` depended
only on kernel contracts and redaction, yet Maestro constructs divergences too and
CLI/MCP both render them, so P5 could not have moved it into `packages/ad-replay`.
It is now `@agent-device/contracts/divergence`; the renderer's output text is
unchanged.

The progress wire vocabulary moves to `@agent-device/contracts/progress`. It is
serialized by `request-progress-protocol.ts` and reconstructed by the CLI reporter
path, so it belongs below both; `src/request/progress.ts` keeps only the sink and
its AsyncLocalStorage binding.

Together these clear all four of replay-test's recorded R10 migration imports, so
the rule now enforces unconditionally for that module.

Behavior is unchanged. The shipped reporter contract — export spellings,
object/factory loading, hook names, timing/order, value fields, the synchronous
live-hook rule, awaited suite completion, error handling, exit codes — is
untouched, and `session-test-reporter-values.test.ts` passes unmodified. The
`--shard-all` `total`/`runnable` asymmetry is preserved as characterized.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

* test(replay-test): pin the Maestro reporter step path against the onStep port

Review finding on #1509: the native `.ad` reporter ratchet exercises only one of
the two `onStep` forwarding chains, so deleting a link in the Maestro chain would
silently stop `onTestStep` for every `test --maestro` run while every existing
reporter test stayed green. Same defect class as the dropped diagnosticId/logPath
(#1501) and the dropped reporter `hint` (#1505).

Maestro is one of P3's two required real adapters and its chain shares no links
with the native one below `runReplayScriptFile`:

  scheduler sink -> runReplayScriptFile -> runTypedMaestroReplayFile
                 -> createMaestroReplayObserver({ onStep }) -> actionStarted -> onStep

Adds a Maestro scenario driving `test --maestro` through the real session handler
and the real reporter registry. It asserts the step payload the engine produces
(`stepIndex`/`stepTotal`/`stepCommand`/`stepValue`, including that a value-less
command stays value-less) together with the attempt/session identity the scheduler
supplies, since that half of the event came from request-global AsyncLocalStorage
before P3. A second case drives a retry so step events must carry attempt-1's
session and then attempt-2's. The flow `name` also pins the reporter `title`, a
value only the Maestro path can produce.

New file rather than an addition to session-test-reporter-values.test.ts: that file
is the pinned characterization and must keep passing unmodified, and Maestro needs
its own vi.mock of core/dispatch for device resolution.

Counterfactual run, both links, each restored after:
  - dropping `onStep` from createMaestroReplayObserver in
    session-replay-maestro-runtime.ts
  - dropping the emitMaestroStep call from actionStarted in
    session-replay-maestro-observer.ts
Each dropped both onTestStep events ("expected [ 'onSuiteStart', 'onTestStart',
…(2) ] to deeply equal [ 'onSuiteStart', 'onTestStart', …(4) ]") and failed both
new cases, while session-test-reporter-values.test.ts passed all 4 — exactly the
hole the reviewer identified.

Test-only; no production change. Bundle output is byte-identical to b339c640f.

Refs #1478

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-31 08:48:20 +02:00
Michał Pierzchała 6b0b9cb0ef feat(ios): usbmux runner route override + #1403 transport experiment evidence (#1510)
* feat(ios): usbmux runner route override + #1403 transport experiment evidence

Adds AGENT_DEVICE_IOS_RUNNER_ROUTE=usbmux, an experimental override that
routes coredevice-backend physical devices' runner commands through usbmux,
and simplifies the xctest branch that awaited a no-op resolveRunnerTransport.

Documents the live #1403 experiment (iPhone 17 Pro, USB + Wi-Fi legs):
steady-state is a wash, but the >30s-idle tax drops from ~4.5s (tunnel
re-probe + session re-establish) to ~440ms because the usbmux session stays
hot; CoreDevice Wi-Fi devices never appear in usbmuxd, so the verdict is
usbmux-primary with network fallback rather than tunnel-code deletion.
Also records the cable-out failure gap (2x45s retry hang swallowing the
usbmux DEVICE_NOT_FOUND hint), which affects today's xctest backend too.

* fix(ios): document daemon scoping of the usbmux route override

The env is re-read per resolve but from the daemon's environment, which is
captured at daemon launch — a later CLI invocation cannot flip the route on
a running daemon. Correct the source comment and experiment doc, and lock
the read-point semantics with a regression test.
2026-07-31 08:36:19 +02:00
Michał Pierzchała 209ac83df5 fix(ci): drop pipefail from the nightly Android emulator script (#1512)
android-emulator-runner runs the script with /usr/bin/sh, which is dash on
ubuntu runners and rejects `-o pipefail`:

    /usr/bin/sh: 1: set: Illegal option -o pipefail

The Android full emulator suite has therefore failed on every nightly since
it landed in #1484, aborting before the first command. The sibling emulator
scripts in android.yml and perf-nightly.yml already use `set -eu`/`set -e`;
match them. The serial assignment stays guarded by `test -n".
2026-07-31 07:57:04 +02:00
Michał Pierzchała 1fc9169188 refactor(daemon): consolidate session-script test factories and drop saveScriptDefaultedHealedPath (#1508)
Preparatory slice for P4a (#1478). No aggregate, no transaction type, no
publication-writer migration — those land separately.

Two things:

1. Name the session-script session states in the shared test factories
   (`makeAuthoringSession`, `makeRepairArmedSession`,
   `makeRepairCompleteSession`) and route 40 inline session literals across
   11 test files through them. The `saveScript*` fields are not independent
   — recording without a boundary is ordinary authoring, a boundary without
   `saveScriptComplete` is an ARMED-but-uncommittable repair, and only the
   COMPLETE combination publishes — so re-deriving the combination per test
   buried the distinction each test was actually about. Pure refactor: the
   factories write today's fields and no assertion was weakened.

2. Delete `saveScriptDefaultedHealedPath`. It had zero production readers:
   the writer's refuse-on-exist guard has been uniform since #1235, so the
   flag was written in three places and never consulted. Removing it takes
   its R7 owner row with it and lowers the R10 baseline from 30/42 to 29/40
   (the ratchet fails on a drop too, so this cannot be deferred).


Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-31 07:24:22 +02:00