mirror of
https://github.com/callstack/agent-device.git
synced 2026-09-14 20:06:34 +08:00
refactor/adr19-boot-unit
48 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5bc3354113 | refactor(runtime): move readiness into platform owners | ||
|
|
123e2c607e | fix: separate boot admission from readiness | ||
|
|
cd3b4af782 | refactor: route boot through readiness runtime | ||
|
|
97c87eec4d |
refactor(registry): exhaustive platformExecution discriminator (ADR 0019 §6) (#1740)
* refactor(registry): make the platform-execution discriminator exhaustive
ADR 0019 §6 (amended): every command descriptor declares its platform-execution
mode explicitly. Adds the `none` mode to `CommandPlatformExecution`, removes the
silent `{ kind: 'legacy' }` default at registry entry, and annotates all 76
descriptors so the migration denominator is machine-readable.
Part of #1739 (wave 0)
* fix(registry): react-devtools executes delegated platform behavior
`react-devtools start` on a Limrun Android instance dispatches internal
`runtime port-reverse`, which reaches a provider device runtime, so ADR 0019 §6
`none` is false for it. Reclassify as `legacy` and add the derived coherence gate
that catches delegated platform execution: if a CLI route for command R
dispatches command D, R may declare `none` only when D is `none`.
Part of #1739 (wave 0)
* fix(registry): attribute CLI dispatches by occurrence, not command name
Subtracting attributed command NAMES let a stray dispatch hide behind a routed
one that names the same command, so the gate's totality claim did not hold.
Dispatch sites now carry their source offset and attribution subtracts
occurrences.
Part of #1739 (wave 0)
* fix(registry): unresolvable CLI daemon-send targets fail the gate
An unknown literal or computed command target resolved to undefined and never
entered the scan, so a dispatch could evade attribution by naming a target the
gate could not read. Daemon-send envelopes are now located by their send call and
an unresolvable target is reported instead of skipped.
Part of #1739 (wave 0)
* refactor(cli): own injected daemon dispatches at a typed construction seam
The syntactic scan recognized only a direct sendToDaemon call whose first
argument was an inline object literal, so a variable envelope or a computed
callee was omitted from every result. Rather than teach the scanner more shapes,
the CLI's injected dispatches now flow through one typed construction point whose
route/command pairs are declared, and the gate reads that declaration instead of
recovering it from syntax.
Part of #1739 (wave 0)
* fix: constrain injected daemon transport handoffs
|
||
|
|
74eab2a554 |
refactor: route selector-resolution structural stages into typed policy (#1744)
* refactor: route selector structural stages into typed policy #1649 landed the per-caller ambiguity matrix and deliberately left four structural columns out: occlusion, off-screen, hittable-ancestor promotion, and the poll budget were per-caller pipeline code, so declaring them would have been an unverifiable claim (nothing consumed them; flipping one left the suite green). This adds the missing half as a table with runners. `SELECTOR_PIPELINE_POLICIES` (src/core/selector-pipeline-policy.ts) gives each caller ONE row naming its ambiguity contract plus its four stages, and every stage is reached only through a runner that reads the row: - occlusion -> selectorPipelineCandidates (candidacy) and resolveSelectorPipelineTarget (refusal). Acting rows exclude covered nodes and refuse covered targets; `find` and the diagnosis probe keep them as candidates and refuse at the target; reads and `wait` ignore them. - promotion -> resolveSelectorPipelineTarget. The per-call-site `promoteToHittableAncestor: boolean` is gone: click/press/longpress name `promotedTarget`, fill/focus/scroll/drag endpoints and the native-ref preflight name `resolvedTarget`. `find`'s below-the-root variant is a declared value rather than a second local helper. - off-screen -> throwIfOffscreenInteractionTarget, which now takes the row and returns the node untouched (no iOS rescue round trip) for observation rows. - poll -> selectorPollBudget, which createWaitPolling derives its deadline and inter-poll delay from; the two wait loops carry a budget, every other row carries none and cannot be polled. Behavior is byte-identical. The acting refusal keeps its exact node, label and details in every branch (promotion declines to retarget away from a covered node, so the "both covered" case names the same node it always did), and `find` carries the occlusion verdict to the focus/type seam rather than raising it early, because find click/fill still delegate that refusal to the interaction leaf's own error shape. selector-pipeline-policy.test.ts drives EVERY row through EVERY runner, including the rows whose answer is "skip" — the half that used to be an absence of code, and an absence cannot fail. Each stage was proven red by flipping its cell (occlusion, promotion, off-screen, poll, plus the declare-only-what-is-enforced guard). The ADR 0011 occlusion/nonHittable `via` pointers for the runtime tree paths now name the runner that makes the decision, not the predicate it applies. Closes #1656; prework for #1739 (waves 4-5). * docs: state constraints instead of narrating the refactor Comment pass over #1656: drop the "used to be per-caller code" / "not module constants" / "rather than an omission" narration — a comment should say what a future edit must respect, not what the previous shape was — and compress the find occlusion-verdict and poll-budget notes to the constraint they actually carry. * refactor: make the selector pipeline the only door to the engine Review of #1744: the structural rows were declared but bypassable. Read and wait routes composed `selectorPipelineCandidates(row, nodes)` with the raw `resolveSelectorChainWithPolicy(..., row.resolution)` and never entered the promotion or off-screen stages, so flipping a read row's `promotion` or `offscreen` changed only the policy unit tests — production `get`/`is`/`wait` were unaffected, which is the unverifiable-column failure #1656 exists to remove. Callers could also pair one row's candidate set with another row's ambiguity contract, and `find list` reached the engine directly. The owning interface (src/core/selector-pipeline.ts) now runs every stage a row declares, skips included, and the stage functions are private to it: - `resolveSelectorPipeline` — single-target rows: candidacy, ambiguity, the replay-guard hook, promotion, occlusion, off-screen. - `listSelectorPipelineMatches` — `reject-candidates` rows, returning the candidate set AND the tree the row sees, so ranking and equivalence classification judge the same nodes candidacy produced. - `runNodePipelineStages` — the node stages for a target from a non-chain matcher (`@ref`, find's fuzzy locator) or a narrowed candidate set. A row whose off-screen stage refuses must supply a refusal shape, so flipping an observation row to `refuse` fails on its real route instead of silently observing. `find list` now names a `readList` row (the new `reject-candidates`/no-rect ambiguity row) instead of calling the engine. R17 selector-pipeline-ownership (scripts/layering/) makes the bypass structurally inexpressible: only the owner may import the engine entry points. Proven against a planted import in selector-read.ts, which the repo-wide scan rejects with the entry points that replace it. Flips now fail through REAL command routes, verified one at a time: readUnique.occlusion/offscreen/promotion and wait.occlusion via get attrs / is / wait; readAny.offscreen via is exists and find; readList.occlusion via find list; promotedTarget.promotion via runtime click. The wait route test needed an advancing clock first — with the frozen one a refused wait spun instead of failing, so the flip hung rather than asserting. * refactor: drop find's dead candidate binding The selector branch bound the row's candidate set and never read it: only the acting classification needs that tree, and find's locator branch brings its own matcher. Names what actually governs the locator target — the shared node stages below, not a candidate set it never had. * refactor: reserve the selector engine behind the pipeline owner Review of #1744 (three blockers). **Listing rows no longer claim stages they cannot run.** `find <q> list` resolves to a candidate SET, so promotion, the off-screen guard and a poll budget have nothing to apply to — a listing has no single element to retarget, keep on screen, or wait for. `readList` now declares only the two stages a listing executes (`SelectorListPolicy`: resolution + occlusion), and the narrower shape is load-bearing: `runNodePipelineStages` and `selectorPollBudget` take the full row, so handing them a listing row is a compile error rather than a silently skipped stage. Pinned with `@ts-expect-error` — widening `readList` makes the directives unused and fails the typecheck. **The engine door is a specifier, not a symbol.** R17's regex could not see a namespace import, a re-export, or a deferred `import()`, none of which mention the symbol it matched. The two engine entries moved to `@agent-device/selectors/engine`, and R19 enforces over the resolved import graph, where every one of those forms is the same edge. Proven on the repo-wide scan by planting each form into a shipped route: namespace import, dynamic import, and `export *` laundering all come back red. `resolveImportEdges` drops an edge whose specifier resolves to nothing, so a specifier rule goes quiet — not red — if the subpath is ever retired. The gate now says that out loud instead of scanning clean. **R19, not R17.** #1750 allocates R17/R18. Verified free against origin/main and that PR's diff, then validated by real merges in both directions: the uniqueness gate passes either way and the three ids stay distinct. The gate itself is new (`scripts/layering/rule-ids.ts`): two branches taking one free number do not conflict in git, so nothing caught R17 twice. Matching whole string literals is what separates a declaration from prose that names a rule, and it is what let the gate see #1750's `const RULE = '…'` shape — the first version missed it and would have been vacuous. `main`'s two pre-existing collisions (R11, R13) are listed as known, not pinned by equality, so #1750 lands in either order without breaking this. Also: the root façade now exposes no resolver at all, and its surface test pins both doors. * fix(layering): make each rule-id allowance expire with its collision Review of #1744: `KNOWN_RULE_ID_COLLISIONS` filtered the exact R11/R13 collision strings, so once #1750 renames those rules apart the entries would keep waving those very collisions through if anyone reintroduced them. "Inert" was wrong — a stale allowance fails open, permanently. `ruleIdCollisionFailures` now checks the transition from both sides: a collision nobody allowed fails, AND an allowance whose collision is absent from the scan fails as a stale allowance. The entry therefore has to be deleted in the same change that removes the collision, and the list burns down to empty, which admits nothing. #1750 is still open, so the transitional entries stay for now (option (b)). Verified against a scratch tree carrying that PR's rename: leaving the list untouched reports both entries as stale; deleting them is clean; and reintroducing `R11 names contracts-implementation-authority and package-boundaries` afterwards is rejected. The last of those is also a unit regression, so the post-transition guarantee is pinned rather than argued. |
||
|
|
f18f8b076f |
refactor(contracts): consolidate per-domain defineUse wrappers into one neutral defineUse (#1741)
* refactor(contracts): consolidate per-domain defineUse wrappers into one neutral defineUse
ADR 0019 §9: use declarations share one neutral defineUse; per-domain
currying wrappers around runtimeUse<PlatformRuntimeOperations>() add a
module per domain for no information. network-runtime-plan.ts,
logs-runtime-plan.ts, screen-recording-runtime-plan.ts, and
app-log-resource-recovery.ts each re-derived their own curried alias
(defineNetworkUse, appLogUse, defineScreenRecordingUse, and an inline
instantiation) from the same generic factory with the same type
parameter.
Export defineUse = runtimeUse<PlatformRuntimeOperations>() once from
platform-runtime.ts (where runtimeUse lives) and re-export it from the
platform facade. Every runtime-use declaration (networkDumpUse,
networkAdmissionUse, the app-log uses, the screen-recording uses, and
appLogRecoveryUse) now builds through that single export; the
per-domain wrappers are deleted.
Type-level only: no required/preferred keys changed, and no use
declaration was added, removed, or altered. Existing deepEqual
assertions in network-runtime-plan.test.ts, logs-runtime-plan.test.ts,
and screen-recording-runtime-plan.test.ts already pin every produced
use object's exact {required, preferred} shape, so they double as the
before/after regression proof that this refactor is behavior-neutral.
Part of #1739 (wave 0).
* fix(contracts): move defineUse into platform-runtime-operations.ts
Review: defining defineUse in platform-runtime.ts required importing the
concrete PlatformRuntimeOperations catalog into the generic runtimeUse
primitive module, while platform-runtime-operations.ts already imports
generic runtime types from platform-runtime.ts. That's an avoidable reverse
type dependency — the lower generic primitive depended on its concrete
aggregate catalog. The unchanged SCC file count didn't prove this harmless;
it counts cycle members, not newly introduced back-edges.
Move defineUse = runtimeUse<PlatformRuntimeOperations>() into
platform-runtime-operations.ts, alongside PlatformRuntimeOperations.
platform-runtime.ts no longer imports the concrete catalog. Re-export
defineUse through the platform facade from its new source module; every
call site keeps importing it from @agent-device/contracts/platform
unchanged, and the three contracts-internal call sites now import it
directly from platform-runtime-operations.ts.
Validation: tsc (full workspace + examples/sdk), check:layering (136/136,
type-cycle count unchanged at 46), and check:affected --run (473 files /
3939 tests) all clean at the new head.
|
||
|
|
602b7a2995 |
refactor: narrow perf API to actionable evidence (#1731)
* refactor: narrow perf API to actionable evidence * fix: address perf API review feedback * fix: preserve deprecated Android CPU metrics |
||
|
|
1b2e786128 | refactor: move screen recording onto platform runtime (#1724) | ||
|
|
cdc754e6ed |
perf: speed up iOS agent recovery and streamline CLI guidance (#1700)
* Avoid interactive children in parent taps * docs: streamline no-skill CLI help * perf: recover faster from sparse iOS trees * fix: preserve selector context for blocked parent taps * fix: preserve coordinate text-entry focus * fix: preserve thin parent touch targets * fix: fail closed for unscoped iOS typing * test: isolate replay lock fixture * test: share node integration process |
||
|
|
b15c502318 |
refactor: extract platform network runtime (#1702)
* refactor: extract platform network runtime * fix: preserve platform network recovery routes * test: guard network parser placement |
||
|
|
b1ed5353d1 |
refactor: extract platform log runtime (#1701)
* refactor: extract platform log runtime * fix: clear terminal app log recovery markers * fix: preserve scoped app log tooling * fix: preserve app log cancellation * fix: handle large changed coverage diffs * fix: harden Limrun runtime identity * refactor: tighten platform log runtime * fix: close app log trust gaps * fix: accept canonical session path aliases * refactor: extract durable capture kit * fix: refresh retained log marker admission * fix: rotate app logs after process relaunch |
||
|
|
13cc90ffc6 |
fix: harden Android snapshot and fill reliability (#1708)
* fix: harden Android automation reliability * test: isolate CLI flush integration * test: close Android review gaps * test: register CLI transport fixture * test: consolidate CLI subprocess fixture |
||
|
|
c06bed9f77 |
refactor: extract platform device inventory runtime (#1699)
* refactor: extract platform inventory runtime * fix: preserve scoped Apple inventory tooling * fix: preserve Apple tool cancellation * refactor: tighten platform inventory boundaries |
||
|
|
4b432fb59b |
feat: add HarmonyOS support (#1683)
* feat: add HarmonyOS device automation foundation Add HDC-backed discovery, snapshots, application lifecycle, and core mobile interactions. Route HarmonyOS through the platform registry and client contracts. Cover parsing and capability parity with focused tests. * feat: support HarmonyOS HAP deployment Install and reinstall signed HAP archives through HDC. Resolve bundle identities from module metadata and relaunch after package replacement. Extend deploy routing and capability coverage for HarmonyOS. * feat: add HarmonyOS single-pointer gestures Execute pan, fling, and swipe plans through HDC uiInput primitives. Derive scroll coordinates from the live ArkUI viewport. Keep unsupported multi-touch gestures explicitly rejected. * refactor: split session inventory command handling Separate session, device, capability, and app inventory response paths. Preserve the public inventory response contract while reducing handler complexity. * feat: support HarmonyOS keyboard actions Route HarmonyOS enter, return, and dismiss through HDC key events. Expose supported keyboard actions through the system command metadata. Keep keyboard visibility inspection explicitly unsupported. * fix: reject unsupported HarmonyOS drag gestures Keep drag unavailable until HDC can preserve source and destination hold semantics. * feat: add HarmonyOS app log streaming 1. Stream HarmonyOS app logs through PID-scoped hilog sessions.\n2. Record HarmonyOS app identity during bundle-id opens for app-scoped commands.\n3. Cover backend routing and bundle identity resolution. * feat: report HarmonyOS foreground app state 1. Read the foreground HarmonyOS mission through aa dump.\n2. Expose HarmonyOS appstate with package and ability metadata.\n3. Add parser coverage for foreground and missing-state cases. * fix: advertise appstate through capabilities 1. Classify appstate in the command descriptor capability matrix.\n2. Surface supported appstate commands in capability inventory.\n3. Cover the advertised Android capability contract. * feat: sample HarmonyOS process performance 1. Sample HarmonyOS process CPU and resident memory through HDC.\n2. Expose the verified metrics through the shared perf command.\n3. Keep frame and memory snapshot collection explicitly unavailable. * feat: clear HarmonyOS app state 1. Add HarmonyOS settings clear-app-state through bundle cleanup.\n2. Force stop the app before clearing data and cache.\n3. Reject all unverified HarmonyOS settings explicitly. * docs: document HarmonyOS support 1. Describe HarmonyOS HDC prerequisites and HAP installation.\n2. Add HarmonyOS to platform discovery and product documentation.\n3. Document verified performance limits for the public HDC surface. * fix: preserve HarmonyOS deploy session identity 1. Bind a resolved HarmonyOS bundle after install or reinstall.\n2. Keep app-scoped logs and observability available after deployment.\n3. Cover session identity preservation for HarmonyOS reinstall. * test: lock HarmonyOS capability boundary 1. Add an independent HarmonyOS capability-matrix oracle and exact advertised-command regression test. 2. Document current HDC-backed support and evidence-based unsupported command boundaries. * refactor: simplify HarmonyOS shared platform boundaries 1. Split device selection and settings dispatch into focused helpers without changing behavior. 2. Keep HarmonyOS serial selection and lock-policy classification covered by regression tests. 3. Remove Fallow complexity findings from the HarmonyOS diff against upstream main. * fix: bound default HarmonyOS HDC commands 1. Apply a 15 second timeout to ordinary HDC operations. 2. Preserve operation-specific timeout budgets for installation and capture paths. 3. Add regression coverage for default and overridden HDC timeouts. * feat: add HarmonyOS screen recording Implement physical-device whole-screen recording through the system recorder and HDC media transfer. Reject unsupported HarmonyOS recording scopes and export flags. Cover capability routing, media retrieval, cleanup, and simulator rejection. * feat: report HarmonyOS HDC readiness Add an HDC version check to the HarmonyOS doctor flow. Document HarmonyOS as a supported doctor platform and cover the result. * refactor: simplify HarmonyOS recording checks Reduce recording validation and test complexity without changing behavior. * test: cover HarmonyOS platform contracts Synchronize public platform expectations across CLI, MCP, replay, and inventory tests. Mock HarmonyOS inventory probes to preserve concurrent test behavior. * test: model HarmonyOS recording capability Require a physical HarmonyOS device in the independent capability parity oracle. * test: cover HarmonyOS input and lifecycle paths Exercise HDC input, lifecycle, installation, and relaunch command sequences. * test: cover HarmonyOS device observability paths Exercise discovery, screenshot validation, and process performance sampling. * docs: define HarmonyOS CI hardware policy Keep HDC hardware validation local and require mocked CI contract tests. * fix: honor HarmonyOS app inventory filters * fix: bound HarmonyOS app inventory classification 1. 限制应用元数据分类并发并为默认清单设置整体时限. 2. 将请求取消信号传递给 HarmonyOS 应用清单读取. 3. 补充失败时中止在飞读取且不继续排队的回归测试. * fix: preserve HarmonyOS inventory failure causes 1. 保留触发应用元数据分类失败的原始错误, 避免被取消同级任务覆盖. 2. 补充总时限中止在飞读取且不启动排队任务的回归测试. 3. 验证后序任务失败时保留默认筛选的恢复提示. |
||
|
|
18291ba8e2 |
perf: collapse app-driving startup turns (#1693)
* perf: collapse app-driving startup turns * fix: align foreground open guidance |
||
|
|
6c0fcb64a1 |
fix: reject distinct ambiguous mutation targets (#1667)
* fix: reject distinct ambiguous mutation targets * fix(ios): scope the raw-match rejection to mutating dispatches `RunnerTests+Interaction.findElement` applied the new fail-closed classification to `querySelector` as well as press/type, because the read call site takes the default `allowNonHittableFallback: false`. With one visible/hittable match and one non-hittable same-selector duplicate the query started returning AMBIGUOUS_MATCH where it previously selected the hittable element, and `queryDirectIosSelectorOrFallback` preserves that error for read callers — so `get`, `is`, and `wait` surfaced an error instead of their prior answer. `classifyDirectSelectorCandidates` now takes a `rawMatchPolicy`. Mutations keep `.rejectDistinctMatches` (the default, so no mutation call site changes); `queryElement` passes `.preferHittableMatch`, restoring the prior read rule: prefer the single hittable match, ambiguous only when hittable matches compete, and never adopt the Maestro coordinate fallback. The Maestro expected-point path is untouched. Covers the one-hittable + one-non-hittable read, competing hittable reads, and the non-hittable-only read. ADR 0011's amendment now states the scope. * test(ios): execute selector read ambiguity regression --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
4b89482a06 |
fix(daemon): open --foreground P1 hotfix — selector rejection, interactive snapshot, capture-failure masking (#1671)
* fix(daemon): open --foreground P1 hotfix — selector rejection, interactive snapshot, capture-failure masking Three P1s from the post-merge review of #1670, all at the session-open-foreground dispatch seam: 1. Explicit device selectors were silently overwritten: the resolved-device rewrite pinned --udid/--platform over whatever the caller passed, so `open --foreground --udid B` with sim A sole-booted silently opened A. Now fails fast with INVALID_ARGS (matching the existing app-positional rejection) on --udid/--device, and on --platform other than ios; an explicit --platform ios passes through. 2. The promised interactive snapshot was never requested: the composed snapshot dispatch forwarded the open request's flags untouched, without snapshotInteractiveOnly — so the capture was NOT the `snapshot -i` path the doc comment promised and returned no interactive presentation. The composed request now sets snapshotInteractiveOnly: true (the exact key the CLI maps -i to and the snapshot runtime reads as interactiveOnly). 3. A capture failure masked the successful open: returning the snapshot error discarded openResponse even though the session exists, so a retry of `open --foreground` failed with "close the current session first". Open success + snapshot failure now returns ok with an explicit initialSnapshotError {code, message} detail and a rendered warning that the session IS open and how to capture manually (snapshot -i). Regressions added for all three: explicit-selector rejection (udid/device/both/non-iOS platform + ios pass-through), the composed dispatch carrying snapshotInteractiveOnly, and the snapshot-failure path returning ok + warning + usable session. * fix(cli): render the composed open --foreground snapshot on default stdout and project initialSnapshotError through the public surfaces Post-merge review on #1671 found the daemon fixes never reached the public boundaries: openCliOutput ignored the nested snapshot (the one-call promise held only under --json), and initialSnapshotError was daemon-only — absent from AppOpenResult, Node normalization, and serializeOpenResult, with the normalized shape truncated to code+message. - default open output now renders the composed interactive tree through the same snapshotCliOutput path snapshot -i uses - AppOpenResult carries initialSnapshotError as the FULL daemon error (hint/details/diagnosticId/logPath preserved) through normalization and serialization Worker-authored; committed by the coordinating session after the worker stalled twice mid-push. Tests: output.test.ts + session-open-foreground (26 pass), typecheck, oxfmt. * refactor(client): one daemon-error normalizer + client-route regressions for initialSnapshotError Review follow-ups on #1671: normalizeInitialSnapshotError duplicated the target-shutdown error normalization and pushed the module over the fallow complexity threshold; consolidated into a single internal normalizeDaemonError (table-driven, full shape incl. retriable/supportedOn) used by both result paths, projected through normalizeOpenForegroundComposition. Client-route regressions: createAgentDeviceClient().apps.open now proves the full initialSnapshotError shape (hint/details/diagnosticId/logPath/retriable) survives normalization, and that a malformed one is dropped — deleting the boundary normalization fails these tests. Also rebased onto current main. * fix(daemon): a thrown initial-snapshot capture failure gets the same successful-open contract Review P1 on #1671: dispatchSnapshotViaRuntime rethrows ordinary capture/runner exceptions; the composition only handled a returned { ok: false }, so a thrown failure escaped to the router and failed the whole open after the session was created — retrying then wedged on the existing session. The catch normalizes the rejection (kernel normalizeError, same conversion the router applies) into the shared openWithInitialSnapshotFailure path: ok response, full-shape initialSnapshotError, session-usable warning. Rejecting-mock regression added alongside the returned-failure case. |
||
|
|
13bc70f24f |
refactor(find): resolve a mutating find's target once, not twice (#1654) (#1669)
* fix(find): resolve a mutating find's target once, not twice (#1654) A mutating `find click`/`find fill` captured the screen, matched by locator under the `findAct` policy, promoted to a hittable ancestor, minted `@eN` off the node it chose — and then re-entered the interaction leaf by bare `@ref`, which looked that ref up AGAIN via resolveSnapshotForRef. The second lookup reads the SESSION frame tree, not the fresh capture find matched against, so the two could disagree: it could hand the action a different node than find picked, or refuse as unresolvable a ref find had observed a moment earlier. find now passes the node itself. `internal.findPreresolvedTarget` carries the resolved node and its tree over the in-process invoke hop, and the ref branch adopts them instead of resolving the ref a second time. What is NOT skipped: the guarantees. Occlusion, hittable-ancestor promotion, and the off-screen guard all still run, on that node, at the same symbols the ADR 0011 `runtime-ref` cells name. Only the LOOKUP is replaced. Recording, ref-frame effects, settle/observation, and deferred-outcome marking are untouched — they live in the dispatch wrapper, not in resolution, which is why the invoke hop is kept rather than bypassed the way find focus/type bypass it. The channel is a second field rather than widening `findResolvedTarget` because they answer different questions: that flag governs ref-frame admission and staleness (policy), this governs resolution (lookup), and focus/type set neither. It is `internal`-only, so it never crosses the wire and carrying live node references is sound. ADR 0011 re-check: the `runtime-ref` disclosure cell gains a second producer, adoptPreresolvedRefTarget, reporting `exact`. Truthful rather than borrowed — the ref is minted off the node handed over. It cannot report label-fallback, which is right: label recovery is a property of looking a stale ref up, and this path performs no lookup. Tests pin the behavior by tap coordinates, so they name which tree the leaf resolved against, plus a control proving the ordinary @ref path is unchanged and a case where the session tree cannot resolve the ref at all — that one can only pass if no second lookup runs. Verified revert-sensitive: removing the short-circuit fails 3 of the 4, and the control stays green. Out of scope, unchanged: the ADR 0011 `native-ref` path (web provider clickRef/fillRef only), where the ref is the provider's own element handle. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0117wfrvC6MRDJUWdRBErNEF * test(find): prove the production route, and correct the divergence claim (#1654 review P2) The regression tests built `internal.findPreresolvedTarget` by hand and called handleInteractionCommands directly, so they proved the leaf consumes the channel but never that find ATTACHES it. Either producer could have stopped doing so and they would all have stayed green — the exact gap #1649's review named, reappearing one layer up. Four tests now drive the real handleFindCommands with invoke wired to the real handleInteractionCommands. Two assert click and fill each carry the selected node (separate producers, separate forwarding hops, so asserted separately); two advance the session tree between find's match and the leaf's resolution and assert the dispatched point is still find's node. That last pair also corrects the record. The claim that the leaf resolves against a different tree than find matched is too strong for the common path: `omitRefFrameSnapshot` (interaction-runtime.ts) makes find's internal dispatch skip the authorized frame tree and resolve against `session.snapshot`, which find's own capture just wrote — so the second lookup normally AGREES, and the two non-diverged tests above pass with or without the fix. The divergence is real but narrower: it needs something to advance `session.snapshot` between find's match and the leaf's resolution, which `refreshAndroidRefSnapshotIfFreshnessActive` does on the ref path. Measured at this head, the old code taps (60,720) "Delete" where find matched "Save" at (310,510) — a wrong-element mutation, now pinned by both click and fill. So the change's value is what #1654 asked for — one resolution end to end — plus closing that window, not a fix for a divergence occurring on every find. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0117wfrvC6MRDJUWdRBErNEF * fix(find): tighten resolved target provenance --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
04c33f9e1e |
feat(ios): expose AX custom actions on merged accessibility elements (#1665)
* feat(ios): expose AX custom actions on merged accessibility elements
Apps that merge a card into one accessibility element for VoiceOver (React
Native's `accessible` prop) publish the card's real affordances as
UIAccessibilityCustomActions rather than as child elements. Our snapshot showed
only the merged node, so an agent looking for a feed card's options control had
nothing to aim at and fell back to coordinate guessing.
`snapshot --actions` now names them:
@e8 [link] "feedItem-by-whiskers.test" actions: ["Reply", "Repost", "Open post options menu"]
Opt-in, because the AX server cannot serve custom actions in a bulk tree
request: adding the attribute makes testmanagerd's reply decoder reject the
nested arrays a custom action serializes into, drop the reply, and time the
request out (~65s vs ~110ms). Only per-element reads answer, at ~100ms each, so
the runner reads at most 12 labelled childless nodes and stops at the
capture-plan deadline. The request pins the private-AX backend, since no other
backend can read the attribute, and reports that as its own `requested-backend`
verdict so a deliberate pin never renders as a degradation warning.
Invocation is not shipped: the actions are readable but not invocable from the
runner. RunnerAXSnapshotBridge.h records the five APIs that were tried.
* fix(ios): disclose a capped custom-action pass, and read on-screen elements first
Two gaps in the first cut.
An element the bounded pass never reached rendered identically to one with no
custom actions, so a capped capture silently taught the reader that later feed
cards have no affordances — the exact mis-inference this feature exists to
prevent. The runner already counted reads against candidates; it now carries
both to the response, and the verdict renders one response-level line when the
pass was incomplete. A complete pass stays silent, and "never asked" stays
distinguishable from "read none" (the key is absent, not (0, 0)).
The obvious remedy for a capped pass would be a scoped re-run, but scope is
applied when the Swift walk builds nodes, long after the read pass, so it does
not redirect the budget at all — measured: reads=12 candidates=18 with and
without --scope. Rather than print a remedy that does nothing, the read pass now
orders candidates on-screen first. That makes the budget land on elements an
agent can act on, and makes the disclosed remedy true: scrolling changes the
on-screen set, so a re-run reads elements the previous pass could not.
Also states plainly, in the flag help and the tool/SDK field description, that
the names are for planning: nothing invokes them, so the affordance is reached
through the element's detail screen, the same control exposed elsewhere, or
coordinates.
* test(snapshot): pin the custom-action coverage pair in the verdict shape assertion
* chore(scripts): classify the --actions flag in the integration progress model
The completeness gate flagged snapshotCustomActions as unclassified, which is
what it is for. It gets its own bucket rather than joining the provider-scenario
table: the values come from the private AX client inside the runner process, and
the fake runner derives its behavior from fixture tables that cannot fabricate
custom actions, so there is no provider-backed scenario to claim. The owning
coverage is named instead — runner XCTest unit, snapshot-lines, snapshot-quality.
* fix(ios): fail closed on unserviceable --actions, bound each read, cap output, and count actions in identity
Four review findings.
1. `--actions` with `--raw`, or on any target that is not an iOS simulator, used
to succeed and return nodes with no actions — a requested capability silently
no-opped, indistinguishable from "this screen has none". Both now fail closed.
The raw pairing is rejected at the shared request seam (INVALID_ARGS) so CLI,
Node client and MCP answer alike before any device work; the platform case is
rejected once the session device is resolved (UNSUPPORTED_OPERATION), naming the
resolved target. `diff --actions` was already rejected as an unsupported flag.
2. The per-element AX read had no timeout, so one wedged element could consume
the whole capture budget. Each read now runs off-thread behind a 1s wait. A
timed-out element counts as unread, never as "read, and it has no actions", so
the existing partial-pass disclosure already covers it.
3. The element budget bounded element count only; one element could still return
an unbounded list of unbounded names. Capped at 8 names of 80 characters, and
clipped elements are counted into the coverage so a truncated list is disclosed
rather than silently presented as complete.
4. Action names were rendered unescaped, and no comparison key read them. Names
now get the same escaping as text previews plus control-character folding, so an
app-authored name cannot split or corrupt a line. `actions` joins the diff
comparable key, the unchanged-comparison projection, and — the sharper bug —
the snapshot presentation key, without which `snapshot` followed by
`snapshot --actions` on a still screen answered "unchanged" and never delivered
the actions that were explicitly requested.
* fix(ios): contain a hung custom-action read instead of accumulating orphans
The 1s read deadline frees the caller, but the underlying AX call is a
synchronous XPC round trip that cannot be cancelled — it keeps running. On a
global concurrent queue that meant repeated `snapshot --actions` against a
wedged element piled up orphaned reads, all using the shared XCAXClient
concurrently. The deadline was containment for the capture, not for the runner.
Since the call cannot be cancelled, contain it instead:
- every read runs on one dedicated serial queue, so a wedged call can never be
joined by a second concurrent user of the shared client;
- a single-flight guard refuses to dispatch at all while an abandoned read is
still outstanding, so a repeated capture adds no work — the dispatch counter
stands still;
- the read pass stops at that point rather than paying a deadline per element
on reads that would all be refused, and reports `blocked` so the capture stays
honest. That gets its own line, because the partial-pass remedy (scroll and
re-run) cannot clear a hang and would send the reader in circles.
Recovery needs no reset: when the hung call finally returns, in-flight drops to
zero and reads resume.
The regression drives a fake AX client that never returns, and asserts the three
things the fix exists for — exactly one in-flight read with no further
dispatches across repeated captures, immediate returns with the skip disclosed
instead of the scroll remedy, and reads working again once the wedge clears.
* ci(ios): execute the custom-action runner regressions instead of only compiling them
The iOS workflow runs a targeted -only-testing list, so a runner test that is
not named there is compiled by the build step and then never executed. All seven
custom-action tests were in that gap — including the containment regression,
which is the only executable proof that a hung AX read cannot accumulate
orphaned in-flight reads.
Red/green against the containment regression, with the fix reverted to its
pre-fix concurrency behavior (global concurrent queue, no single-flight guard,
no blocked exit):
RED in-flight 6 (want 1), dispatches 6 (want 1), each repeat paid the full
1.004s deadline (want <0.2s), blocked=false (want true), and the
in-flight drain never completed — "Exceeded timeout of 5 seconds".
GREEN 7/7 pass, containment regression in 1.02s.
|
||
|
|
e14c9d8d7b |
fix(cloud): unblock iOS snapshot and gate cloud fill on text-entry focus (#1658) (#1666)
* fix(cloud): unblock iOS snapshot and gate cloud fill on text-entry focus (#1658)
Two bugs isolated to the cloud-webdriver iOS path.
`snapshot`/`diff` refused every capture on a live BrowserStack session with
SESSION_NOT_FOUND, instantly and without a driver round trip. The app-session
guard they ran belongs to the local XCUITest runner, which must attach to a
target app; a cloud capture reads the provider's own driver session and needs
no app identity, so it now applies to local Apple targets only. The session
was empty in the first place because the provider open path skips local app
resolution wholesale — no simctl/devicectl reaches a hosted device — and
dropped an explicitly spelled bundle id along with it. A dotted, non-deep-link
target is the bundle id under the same convention resolveIosApp applies
locally, so a cloud `open com.example.app` now records it.
`fill` tapped and sent its keys in back-to-back requests. A WebView input —
an OAuth page in a Safari view controller — takes first responder
asynchronously, so the keys landed with nothing focused while the command
still answered "Filled N chars"; tapping and filling as two separate commands
worked only because the round trip between them gave the field time to focus.
The cloud interactor now waits on the same signal the Apple runner uses, the
software keyboard going from hidden to shown after its tap, and discloses what
it observed as `textEntryReadiness` so a fill with no witness cannot pass for
a filled field. Where keyboard visibility cannot witness the focus move —
back-to-back fills into one form, the shape that failed most often — it spends
the runner's full readiness budget rather than racing the app with a short
settle.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Zjbzf7HdziX9SzFpNP7WW
* fix(cloud): let a new bundle-id open replace the tracked cloud iOS app
Adopting an explicitly spelled bundle id on a provider-backed open (the
fix that makes snapshot/diff work at all) also made a previously dead
precedence rule live: the provider branch returned currentAppBundleId
first, so once a first open had populated it, `open com.a` followed by
`open com.b` left the session still reporting com.a to every
appBundleId-gated command.
The local path does the opposite, and is the convention this branch is
meant to mirror: resolveIosApp returns a dotted target unchanged and
never consults the session's current app. Only its deep-link branches
prefer the tracked id. Flip the provider branch to match — an explicit
bundle-id target wins, and everything the branch cannot name (deep
links, display names, bare open) still falls back to the tracked id.
* fix(cloud): fail a witness-less cloud fill instead of reporting it filled
Review follow-ups on #1658.
`not-observed` was still a success: it sent the keys and answered "Filled N
chars", and nothing renders `textEntryReadiness` in default CLI output — so the
exact silent success this branch exists to remove survived whenever focus never
happened. A tap that raises no keyboard now fails with
`text_entry_focus_not_observed` and sends no keys, leaving the field untouched
rather than half-written, and the readiness vocabulary keeps only outcomes that
describe a fill that did type.
The readiness budget was advertised but not enforced at the request boundary:
each keyboard probe inherited the client's 30s default, so one hung probe could
hold a 2s wait for far longer. Probes now carry their own bound, threaded
through the client as a per-request timeout override.
The probe also swallowed every error as "this driver cannot answer", which
degraded a dead session, an auth rejection, or a grid outage into a blind text
entry. Only a positively classified unimplemented route counts as unsupported
now — classified on the W3C error code rather than the status, since `unknown
command` and `invalid session id` share HTTP 404 — and everything else
propagates.
The provider scenario proved request ordering against a stub that always
accepted keys. Its fake now models the device: focus lands a beat after the
tap, and keys arriving while the keyboard is down are accepted and dropped,
exactly as an unfocused field does. The tests assert the field's own value, and
both go red against the pre-fix `fill`.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Zjbzf7HdziX9SzFpNP7WW
* fix(cloud): witness the tapped field's focus before a cloud fill types
Review of cc23f2b found two ways a fill could still report success
without evidence that OUR tap focused the field it was aimed at.
P1. `settled-keyboard-up` and `settled-unknown` both typed and returned
normal success. Keyboard visibility can only witness that *a* field took
focus, never *which*: filling a second field in an already-open form
reads the same before and after, so a missed tap left the first field
focused and `POST /keys` — which the driver routes to whatever holds
first responder — appended to it while every request returned 200.
Failing those closed outright would have broken ordinary multi-field
form fills, which do work: a live AWS Device Farm run types both fields
of a WebView login correctly. So witness focus properly instead. W3C
`GET /element/active` answers the question keyboard visibility cannot —
is the thing focused now the thing I tapped — and answers it whether or
not the keyboard was already up. That becomes the primary signal
(`focused-element`); the keyboard transition stays as the fallback for
drivers without the route, and a keyboard already up on such a driver
now refuses rather than typing.
The test is identity, not geometry. Containment of the tap point looks
like the obvious rule and is wrong: focusing a field can re-lay it out.
On a live iPhone 16, tapping Safari's collapsed address bar expands it
into a taller field that no longer covers the tapped point, and a
containment-only rule refused a fill that plainly worked. So a tap that
MOVES focus counts, with containment as the second half of the test —
re-filling the already-focused field moves nothing, and only geometry
tells that from a tap that missed. Both readings are taken before the
tap, since each is evidence only as a change.
P2. The 2s budget bounded the loop but not the calls inside it: every
probe got a fixed 1500ms, so one begun near the deadline finished well
past it. Both the probe timeout and the sleep are now capped by the
remaining budget.
Also fixes a related escape the review did not name: the poll loop had
no catch, so one transient grid error aborted a fill the next poll would
have satisfied. Probe failures are now tolerated within the budget, but
a budget that expires without a single answered probe rethrows, so a
dead session surfaces as itself rather than as "the tap missed".
The provider scenario gains the two-field case the review asked for: it
begins keyboard-up with the email field focused, misses the password
tap, and asserts no keys reach the email field.
Verified on AWS Device Farm iPhone 16 / iOS 18.0 at this exact tree:
address bar (the re-layout case) and both WebView login fields all
report `focused-element`, the second with the keyboard already up, and
the device reads back `tomsmith` and a 20-character password.
* fix(cloud): refuse a cloud fill no focus probe can witness, and bound the composite probe
Two blockers from the review of
|
||
|
|
cc943400a9 |
feat(daemon): [RFC] prototype foreground-attach convenience (#1670)
Prototype `open --foreground`: on a fresh session with no app argument, auto-resolves the target from the sole booted iOS simulator's sole foreground app (reusing the exact same ambiguity-detection probe that enriches the SESSION_NOT_FOUND hint), then attaches the initial interactive snapshot to the response by composing the existing snapshot-runtime dispatch. Collapses the documented 3-call snapshot-fails -> read-hint -> open -> snapshot-succeeds dance into a single call for the unambiguous case, while failing closed (AMBIGUOUS_MATCH) with no guessing otherwise. First-pass RFC, not reviewed — see PR body for the design tradeoff writeup, live before/after evidence, and scoped-out follow-ups. |
||
|
|
3937036e5e |
feat: support --settle on scroll and back (#1638) (#1650)
* feat: support --settle on scroll and back (#1638) Scroll-then-observe and back-then-observe are legitimate agent pairs, but the post-action observation registry never grew past the touch commands, so `--settle` on either was rejected with INVALID_ARGS — burning a tool call each in AppControlBench's bsky-16. Both commands now carry the `settle` descriptor trait, and every surface derives from it rather than a hand list: CLI allowed flags, MCP/SDK input fields, the flag-sourced timeout envelope, and MCP ref-pinning. The CLI flag/metadata helpers moved out of the interaction family into post-action-observation-grammar.ts (back is a system command), and SETTLE_REF_ISSUING_TOOLS became a derivation — a hand list would have silently stopped pinning the new commands' refs. settleAfterInteraction and the new settleObservationCommand are two entry points over one engine: same loop, storage, hints, and diff bounds, with the target-less path supplying its own baseline and no proximity point. The daemon reaches that command through the runtime surface, never by importing `commands/` (R2) — the same seam the touch handlers use for press/fill — and generic-settle.ts is loaded through a lazy `await import` returning a closure, so the interaction runtime subgraph stays out of this dispatcher's static graph (a static edge folded ~18 files into the daemon-server type cycle; R10 caught it). Both of generic-settle's orderings are load-bearing and tested: the baseline is frozen before dispatch (and before the Android dialog preflight), and the observation runs after markDeferredInteractionOutcome so settle's first capture folds in the #1542 stabilization rather than racing it. The ADR 0014 "a settled diff publishes refs" rule moved to settle-ref-issuance.ts, shared by both routes. One divergence is deliberate: scroll/back resolve no element, so the diff baseline is the session's STORED pre-action tree — "settled tree vs the last tree you observed" — not press's freshly resolved pre-action capture. Both commands also switch to preserve-daemon on timeout, which changes the non-settle path too: with --settle their dominant hang mode is now a wedged accessibility bridge, and a timed-out capture must not reset the daemon and lose every session (#1105). The reviewed-set gate records it. Live-validated on an iOS 26.2 simulator (Settings): scroll --settle settled in 1786ms with a +6/-6 diff carrying fresh refs; back --settle in 771ms with +15/-6. Alternating cost runs, one call vs the pair it replaces: scroll 2.9-3.0s vs 5.3-5.6s, back 3.1-3.2s vs 4.7-5.1s. Those include the #1627 deep-capture extension. * fix: render settled-diff refs paste-ready in CLI output A settled diff activates a PARTIAL ref frame (ADR 0014), which admits only the pinned `@eN~s<gen>` form of the refs it issued. The unchanged-interactive tail already rendered that way, but the diff's own added lines rendered the bare `@eN` embedded in the snapshot line — so a CLI caller who copied the ref the diff just handed them got `plain_ref_requires_complete_frame` and had to append the generation by hand. Added lines now render pinned when the response carries `refsGeneration`, exactly like the tail. Removed lines render verbatim: they name elements that just left the screen, and `SettleDiffLine` never gives them a ref. This is not new to scroll/back — press/click/fill/longpress had the same gap since #1101. MCP was never affected: its ref-pin store rewrites plain refs on the way in, which is why the model never sees a suffix. Live: `scroll down --settle` now emits `+ @e14~s218078 [cell] "Game Center"`, and `press @e14~s218078` copied straight out of that line taps successfully. * test: record the pinned-diff-ref bytes in the output-economy baseline Rendering added diff-line refs pinned costs 8 bytes in the two settle CLI text samples (two `~s<gen>` suffixes). The output-economy baseline is the tripwire for exactly this, so the increase takes an explicit reviewed waiver rather than a silent baseline bump — the same one the settled TAIL's pins already carry, for the same ADR 0014 reason. Only `bytes` moves: lines, refs, hints, and shape are unchanged, which is the evidence that this is a suffix on existing refs and not a new payload. Caught by CI, not locally: `pnpm test:unit` runs unit-core and subprocess-stub only, while the Coverage lane runs every vitest project. * test: prove the generic settle degrades when its runtime cannot be built `createGenericSettleRuntime` catches and returns undefined so an observation that cannot even start does not fail an action that already succeeded. That was a claim in a docstring with nothing behind it — the one changed line the coverage gate reported uncovered (95/96). The test puts the session in the state the catch exists for: the router handed us a session that is no longer in the store, so building the settle runtime throws SESSION_NOT_FOUND. The response keeps its scroll result and simply carries no settle payload. Removing the try/catch fails it. * build: teach fallow that vi.mock reaches pinOwnProcessStartTime dynamically Not from this PR: #1642 added `pinOwnProcessStartTime` on main, and its three consumers reach it the only way a Vitest module mock can — `vi.mock(path, async (importOriginal) => (await import('...')).pinOwnProcessStartTime(...))`. Dependency analysis cannot follow that dynamic import to a consumer, so the export reads as dead the moment any PR pulls that file into its audit scope. This PR is the one that did. The entry records the consumers by path and the reason, matching the daemon route-handler entry directly above it, which exists for the same dynamic-`import()` limitation. * refactor: adopt the best of the parallel #1653 implementation Two sessions independently built #1638 (PR #1650 and PR #1653) and converged on the same architecture — trait in the registry, one engine with two entry points, runtime-command seam, lazy import, preserve-daemon, stored-baseline honesty. #1650 continues; this folds in what #1653 did better: - The agent-facing help core loop (cli-help.ts) now names scroll and back as settle-capable. Without this, the benchmarked closed-grammar help line kept instructing agents that --settle is only for press/click/fill/longpress — actively steering the AppControlBench models away from what #1638 shipped. - issueSettleRefs moves into session-snapshot.ts, beside the partial-frame primitive it wraps, deleting the single-function settle-ref-issuance module. - Their seam tests: back reader→writer settle plumbing, back CLI settle rendering, and a trait-less generic command (home) ignoring a stray settle flag rather than observing or rejecting. What #1650 had that #1653 lacked, for the record: the SETTLE_REF_ISSUING_TOOLS registry derivation (without it, MCP never pins a scroll/back settle diff's refs and the partial frame rejects every follow-up), BackCommandResult.settle in contracts, back's MCP output schema, paste-ready pinned diff refs, and the docs/changelog/baseline surfaces. * bench: help-conformance case for settled scroll-to-find planning The #1638 extension of the closed --settle grammar to scroll/back is the feature's entire payoff — collapsing scroll-then-observe into one call — and the closed command list is an enumerated N whose enumerator is this bench. The regex over the help text proves the sentence exists; this case checks whether a model plans differently because of it. One focused case, deliberately not coached: a pinned visible-first snapshot (rendered by formatSnapshotText, pinned by the sample-producers gate) whose wanted row is summarized off-screen with no ref anywhere in the output. The tempting pre-#1638 plan is `scroll` plus a separate `snapshot -i`; acceptance is the single settled call. Scoring was verified against eight plan shapes in both directions before recording. Model-backed record (claude-haiku-4-5, 3 trials, current help): 0/3 — but the decomposition is the finding. Settle eligibility GENERALIZED (3/3 trials put --settle on scroll unprompted; the mutation-suffix framing concern did not materialize) and the two-call habit is residual (1/3). All three trials failed on `scroll @e3 down --settle` — the pre-existing #1366 scroll-takes-no-target confusion, which the live CLI recovers with a dedicated hint but a single-shot bench cannot. The recorded gap is therefore a first-30 doc gap (nothing teaches that scroll takes no target), not a settle-eligibility gap; tuning the case until it passes would just delete the evidence. |
||
|
|
870d12c406 |
fix: honest find contract — press/tap aliases, read-only list action, selector uniqueness (#1637)
* fix: honest find contract — press/tap aliases, read-only list, selector uniqueness (#1625) Three defects in find's contract, fixed together because they are one vocabulary: press/tap are the same action as click everywhere else in this CLI, yet find rejected them — agents using the vocabulary the tool itself established burned a tool call per attempt (four in one bench run). Both parsers now normalize press/tap to click; longpress/swipe stay real exclusions. The #1602 recovery hint told agents to run bare find to 'list matches', but bare find CLICKS a unique match — inspection guidance pointing at a mutation (the #1625 report: 'find Dictionary' navigated into Dictionary). find <q> list is the read-only surface that guidance needed: every match with its @ref, unique match included, never a tap. Captured UNSCOPED (the label-scope optimization narrows to the first match, exactly wrong for listing), published as an ADR 0014 partial frame authorizing every listed ref. Selector-shaped queries skipped the ambiguity check and took the first match silently — the mis-binding path the AMBIGUOUS_MATCH recovery advice itself pointed agents at, while --first/--last were documented as explicit opt-ins. Selector and text queries now share one contract: multiple matches reject with the #1597 candidates listing unless --first/--last narrows explicitly. The hint is rewritten around the new contract; docs and the MCP find output schema follow. Regressions at every layer: both parsers (alias, list token, unsupported-action hint shape), the daemon handler (selector ambiguity with candidates, --first opt-out, list returns all matches with zero action dispatches, unique-match list does not tap). * refactor: single-home the find read result and flatten parseFindArgs (fallow) The daemon's DaemonFindResult had drifted into an identical structural twin of the engine's FindReadCommandResult — the two grew the list variant in parallel and crossed the clone threshold. The shape now lives in contracts as FindReadResult (below both zones, per R2's own remedy) with the engine and daemon both aliasing it. parseFindArgs collapses the four bare single-token actions into one membership check and extracts the get sub-action parser, bringing it back under the complexity threshold instead of waiving it. * style: merge duplicate contracts import (lint) * fix: accept list on the MCP input surface and pin every listed ref (review) - FIND_ACTION_VALUES gains 'list' so field-metadata/MCP input no longer rejects the action the CLI parser accepts - FindCommandResponseData types 'matches' (public client response) - MCP mergeFindRefPins learns every matches[] ref, so a plain @eN press after find-list forwards pinned and the partial frame admits it - CLI/MCP text renders every listed match as its own pinned line via the snapshot-line role/label normalizers - regressions: daemon partial-frame scope, pin store, CLI output, MCP schema + find-list->press chain, typed client list response |
||
|
|
3e4828d68d |
feat: add scale-only screenshot sizing (#1617)
* feat: add scale-only screenshot sizing
* fix: refuse retired --max-size inputs on every released surface
Released sizing inputs must fail closed with migration guidance instead of
silently producing native-size artifacts:
- contracts: RETIRED_SCREENSHOT_MAX_SIZE declaration + SCREENSHOT_SCALE_LIMITS
as the single source for the scale bounds and migration messages
- .ad parser: released 'screenshot ... --max-size N' and 'record start ...
--max-size N' lines now refuse at parse time (frozen replay-compat witnesses)
- daemon: screenshot rejects old-client screenshotMaxSize like recording does;
the recording guard now shares the same contract data
- Node client: screenshot/record daemon writers refuse the removed { maxSize }
option before transport
- CLI: --max-size unknown-flag error carries the migration guidance
- config/env: stale screenshotMaxSize config keys and the retired
AGENT_DEVICE_SCREENSHOT_MAX_SIZE env var are refused for sizing commands
(other commands keep working)
Quality: numberField now reuses the canonical readOptionalNumber contract
helper (AppError bounds instead of plain Error); png-resize inlines one-use
wrappers and restores the worker-thread rationale; docs typo fixed.
* test: drop retired maxSize entries from the MCP undocumented-input allowlist
* fix: refuse retired maxSize at the MCP field-projection seam + release-provenance corpus witnesses
- readFieldInput silently dropped undeclared keys before the daemon writers
could refuse them, so an MCP call carrying { maxSize } reached transport and
returned native-size success. New retiredField() combinator declares the
removed key in the field map: the projection seam refuses it with the
canonical migration message and the JSON schema no longer advertises it.
Real-route MCP executor regressions cover screenshot and record.
- replay-compat corpus: derived v0.20.5 witnesses for the released screenshot
and record --max-size forms (SHA-256 pinned, new retired-capture-size
coverage surface) so check:replay-compat proves the shipped syntax refuses
with migration guidance instead of degrading silently.
---------
Co-authored-by: Michał Pierzchała <thymikee@gmail.com>
|
||
|
|
f93b259d15 |
fix: stop the slow-snapshot warning from firing on a single cold start (#1628)
* fix: stop the slow-snapshot warning from firing on a single cold start The session's first capture folds one-time startup (runner launch, helper install) into its duration, and nearest-rank p95 over a small sample set equals its largest one or two values — so one 12s cold start produced 'snapshots are slow in this run: p95 12417ms over 1 captures' with hints blaming device load or a stale daemon, inviting exactly the restart spirals the hints exist to prevent (observed on every AppControlBench run). The warning now judges only warm captures (first sample excluded) and only once at least three exist; the displayed stats still cover every sample, so the cold start remains visible as maxMs. * refactor: single-home the slow-snapshot warning policy (review) Push the warm-judging rule down into summarizeSnapshotTimingSamples so the interactive session path and all three replay handler paths share one policy, and summarizeSnapshotDiagnostics returns to a one-line delegate. Merge no longer judges slowness from lossy order-less reconstructed samples (a run's cold start comes back as both its p95 and max): it aggregates display stats and carries a warning only when a constituent run judged one itself. MIN_WARNING_SAMPLE_COUNT renamed to MIN_WARM_SAMPLE_COUNT — it gates warm samples, not total captures. Cold-start regression tests move to the shared layer; suite-aggregation tests now pin that individually-silent runs merge silent. * style: oxfmt * fix: quorum-gate the warm warning and make the merged message speak about warned runs (review) A single warm outlier could still fire the warning (nearest-rank p95 is the maximum through nineteen samples): chronic now additionally requires at least two slow warm captures. And the merged warning formatted its number from the reconstructed aggregate, so one slow run among many fast ones produced 'slow: p95 <fast number>' — the merged message now reports how many runs warned and the worst warned run's own p95, never the aggregate. Regressions for both: one-warm-outlier stays silent; a slow run merged with many fast ones warns with the slow run's number while the aggregate p95 sits below the threshold. |
||
|
|
a67c72c211 |
fix(ios): pin tap-outcome corroboration probes to the baseline's backend (#1634)
* fix(ios): pin tap-outcome corroboration probes to the baseline's backend The recorded-failure screens are exactly where the capture plan flips between XCTest and private-AX (the penalty boundary), so #1605's same-backend requirement failed closed right where XCTest tap false negatives actually happen: the baseline was captured via private-AX under penalty, the probe came back via tree, and a landed tap surfaced as XCTEST_RECORDED_FAILURE. In the AppControlBench bsky-16 run this fired four times, each sending the model into a re-observe/retry spiral. The comparison stays same-backend by design (backends are not comparable views of a screen); instead the probe is now CAPTURED the way its baseline was: a new internal preferredBackend option (never CLI-exposed) threads daemon -> runner, and a private-AX-preferred capture takes the exact penalized route — privateAX-first plan, 'deferred' verdict, no degradation warning, no settle budget reset. Live-verified on the deterministic repro (Bluesky drawer-menu press under penalty, seeded bench feed): errored with the backend-mismatch diagnostic before, corroborates as landed after, with no mismatch phase in the request diagnostics. Daemon tests cover pinned and unpinned baselines end to end through the dispatch context; the Swift plan gate is a pure function with an executed in-bundle test (added to the ios.yml regression list). * style: oxfmt * fix: exclude raw baselines from corroboration and prove the pin end to end (review) Raw baselines could not be pinned: the raw diagnostic plan keeps tree-first error propagation by contract and is never rerouted by the penalty or the preferred backend, so preserving 'raw: true' on the probe recreated exactly the backend-mismatch false failure this PR removes. Corroboration now declines raw baselines up front (they are diagnostics, not evidence baselines) with a regression pinning that no probe capture is dispatched at all. The wire is now regression-proven at every hop: a dispatch-level test drives dispatchCommand with the context flag and asserts the emitted RunnerCommand carries preferredBackend (red if handleSnapshotCommand or the interactor stops forwarding); the injected-transport test asserts the interactor's snapshot payload both ways; and a runner unit test decodes the wire JSON, projects it through the extracted snapshotOptions(from:), and composes it with the plan rule — pinned regular plan defers to privateAX-first, RAW plan stays untouched. Executed on-simulator; added to the ios.yml regression list. |
||
|
|
d5f99bab1c |
refactor: sink backend.ts's cycle-closing types below both zones (#1632) (#1636)
backend.ts imported RepeatedInput up from commands/command-input.ts and ScreenshotResultData up from utils/screenshot-result.ts — the interface hub typed in terms of the zones that depend on it, R6's textbook inversion shape. - RepeatedInput now lives in @agent-device/contracts/interaction; command-input.ts re-exports it for its existing importers. - ScreenshotResultData already had a byte-identical canonical declaration in contracts/snapshot-types.ts (exported via contracts/capture); the utils copy is now a re-export of it, deleting the duplicate outright. Measured member-by-member: the R9 type cycle collapses 76 -> 49 files. backend.ts, runtime-contract.ts, commands/runtime-types.ts, and commands/runtime-common.ts all leave the component (27 files stranded out at once); zone ceilings lowered to the measured values (commands 33 -> 14, platforms 7 -> 2, root 5 -> 3, daemon-server 20 -> 19) and CONTEXT.md's hub list recomputed (core/dispatch.ts 8, command-catalog.ts 7, resolution.ts 6, command-descriptor/registry.ts 6). No TYPE_INVERSION_BASELINE additions. |
||
|
|
8c800ae53f |
refactor(contracts): one viewport-root predicate for the whole repo (#1613)
* refactor(contracts): one viewport-root predicate for the whole repo "Is this the Application/Window root" was written nine times: three spellings normalizing `type|role|subrole`, five lowercasing `type` alone, and one comparing the normalized type for EQUALITY. Two of the nine sat in `contracts/snapshot-visibility.ts` itself, disagreeing with each other. Measured before collapsing, using #1592's method — ground the comparison in what each backend ACTUALLY emits, not in fixture strings. Over the 31 names iOS's `elementTypeName` can return, the 18 fully-qualified class names Android emits, and the 24 mapped/raw forms the macOS helper produces, the nine agreed on 71 of 73. The two exceptions are macOS window subroles, and the only spelling that disagreed is maestro's `===`, whose platform union is `android | ios` — so it can never see them. The duplication was textual, not behavioral, which is what made the collapse safe. `isViewportRootNode` reads role and subrole because the macOS helper is the only backend populating them and the only one able to emit a window whose `type` does not say so: `normalizedSnapshotType` returns the raw subrole for a non-standard window, so an `AXWindow` with subrole `AXSystemDialog` or `AXUnknown` reads as neither from `type` alone. Those two shapes are the whole behavioral delta of this change, at the six call sites that were type-only, and they are windows by role. `snapshot-viewport-root.test.ts` pins the predicate over those three emitted vocabularies. Red evidence: reverting the canonical definition to the type-only spelling fails 2 of 5 cells, to the equality spelling 4 of 5. Also drops two kernel re-declarations this made visible: maestro's local `containsPoint` and `rectsOverlap` were character-identical to `@agent-device/kernel/rect`'s `containsPoint` and `isRectVisibleInViewport`, in a file that already imports from that module. And `resolveViewportRect` loses three `as Rect` casts that only existed because `.filter()` cannot narrow `node.rect` — one `flatMap` states the same thing honestly. Deliberately NOT in this change: the three viewport RESOLVERS still diverge, and on Android that is a live defect rather than duplication. Filed separately with the measurement. * test(contracts): enumerate the macOS emitter's real vocabulary Review found the table claimed to pin "the vocabulary each backend actually emits" while omitting most of it. `normalizedSnapshotType` has three output classes and only two were represented: 1. thirteen roles mapped to fixed short names — six were missing (StaticText, TextField, TextArea, MenuBarItem, Menu, MenuItem); 2. AXWindow, whose output is the SUBROLE unless it is AXStandardWindow; 3. the `default:` arm, `subrole ?? role`, emitting the raw AX-prefixed value for every unmapped role. All three are now enumerated, and the table asserts its own completeness against the emitter's fixed-output set — a role added to that switch without being added here fails, which is the emitter-drift protection the docblock was promising but not delivering. Re-measuring over the complete tables also corrected the header's own numbers. The claim was "71 of 73 agree, 2 disagree"; over 75 names it is 71 agree and FOUR disagree, because AXSystemDialog and AXUnknown were absent from the old table. Those two are the behavioral delta of this PR — an AXWindow whose subrole is emitted as the type, invisible to the six type-only spellings and named exactly by `role` — so the incomplete table had been hiding the very rows that justify reading role/subrole. The other two (AXFloatingWindow, AXSystemFloatingWindow) remain inert: only the `===` spelling misses them and its platform union is `android | ios`. * test(contracts): derive the macOS fixed-output set from the emitter Two test-validity defects from review, both real. The raw-fallback row `{ type: 'AXSearchField', role: 'AXTextField', subrole: 'AXSearchField' }` was unreachable: the `AXTextField` arm returns `TextField` whatever the subrole, so no emitter run can produce it. Replaced with `{ type: 'AXSortButton', role: 'AXCell', subrole: 'AXSortButton' }` — a subrole on a genuinely unmapped role, which is what the `subrole ?? role` default arm actually emits. `MACOS_FIXED_OUTPUTS` was a hand-kept twin compared against a hand-kept table, which is circular: a new mapped Swift role is absent from BOTH, so they agree and the gate stays green. The "emitter-drift protection" the docblock promised did not exist. The set is now parsed out of `normalizedSnapshotType` in SnapshotTraversal.swift, so the comparison is against the emitter rather than against a copy of the table's own assumptions. `case "AXWindow"` returns a subrole expression rather than a literal and is deliberately outside the literal-return set. Red evidence: adding `case "AXDisclosureTriangle": return "DisclosureTriangle"` to the Swift switch fails with `expected [ 'DisclosureTriangle' ] to deeply equal []`; 6 pass once reverted. The parser throws rather than silently matching nothing if the function is renamed or moved. * chore: restore maestro conformance corpus to main 45 corpus YAMLs carried an unrelated quote-style churn ("Button" -> 'Button'). They were already modified in the worktree when this branch started and a `git add -A` swept them into the predicate commit. Nothing in this PR reads them. Restored verbatim to main. |
||
|
|
d8b309c6db |
refactor(contracts): name façade exports explicitly and retire the pin table (#1614)
* refactor(contracts): name façade exports explicitly and retire the pin table Thirteen of the fourteen `@agent-device/contracts` façades were bare `export *` barrels. `facades/snapshot.ts`, added by #1582, was the one exception — explicit named re-exports — and that is now the rule. Everything #1574 built to cope with `export *` goes with them: scripts/layering/facade-symbols.ts -980 (816 pinned names) scripts/layering/facade-exports.ts -192 (readFacadeExports) scripts/layering/facade-exports.test.ts -234 (star semantics) scripts/layering/package-boundaries.test.ts -55 `readFacadeExports` re-implemented ESM `GetExportedNames`/`ResolveExport` — star-chain resolution, ambiguity rejection, diamond binding identity, cycle guards, spec-accurate `default` filtering at the star rather than the source. All of it existed to enumerate what `export *` hides. 523 of the 816 pinned names belonged to contracts, i.e. to those thirteen files. Once a façade names its exports, the façade file IS the pin, and it is visible in the diff of the file that widened rather than in a separate table a reviewer has to cross-check. `readNamedExports` (20 lines) stays and is enough: it already throws on bare `export *` and on `export default`. The pin is replaced by one structural gate — no façade may contain a bare star — which reuses that rejection rather than adding a regex. Surface equivalence verified independently, not asserted: main's own `readFacadeExports` run over the new façades, compared against main's own `FACADE_SYMBOLS` table — 31 subpaths, 0 added, 0 removed. Red evidence for the new gate: planting `export * from '../request-progress.ts'` back into facades/progress.ts fails it with the file named and the reason quoted; 12 pass / 0 fail once reverted. Not included: the `lowerAndroidTouchPlan` tuple-assertion drive-by. It needs `sampleGestureOffsets` to carry a min-arity tuple through `.map()`, which TypeScript will not infer without a typed helper — a real change to the gesture-plan contract rather than a drive-by, so it stays out. * test(layering): assert façades stay exhaustive over their sources Review on #1614 caught this conversion silently narrowing the public surface. The explicit lists were generated against the surface at fork time; #1567 landed 13 exports meanwhile — `DragOptions`, the drag-gesture vocabulary (`COORDINATE_GESTURE_KINDS`, `CoordinateGesturePayload`, the three `DEFAULT_DRAG_*` constants, `DragGestureInput`, `DragGesturePayload`, `GestureCommandInput`, `buildDragGesturePlan`, `dragGesturePayloadFromPositionals`, `normalizeGestureCommandInput`) and `MultiTargetAnnotationV1`. The `export *` barrels had been forwarding all 13 automatically; the rebase dropped every one, and only a human diff caught it. The star-rejection gate could not: it only proves a façade does not WIDEN invisibly. Narrowing is the failure an explicit list newly makes possible, because `export *` could not narrow by construction. So the property the stars gave for free is now asserted directly — every name a re-exported source declares must appear in the façade. Scoped to `packages/*/src/facades/`, the barrels this PR converted. A hand-curated package `index.ts` is a different thing: `ad-replay` deliberately publishes two values out of a much larger `internal/`, and forcing exhaustiveness there would widen a surface its owner narrowed on purpose (#1555). A source that itself carries a bare `export *` is skipped — unknowable from that file alone, and reachable because the façade re-exports the starred module directly too, which IS checked. Red evidence: dropping `MultiTargetAnnotationV1` from facades/replay.ts — one of the 13 the old gate was blind to — fails with the file, the source and the symbol named. 13 pass / 0 fail once restored. * fix(layering): close the exhaustiveness gate's starred-source hole Two review findings, plus a third the gate caught on itself. P1 — the three `DEFAULT_DRAG_*` constants join the existing public-façade suppression, alongside `COORDINATE_GESTURE_KINDS` and `normalizePublicGesture` which the same conversion surfaced. All five are #1567's drag vocabulary, made individually visible to `--production` analysis for the first time because a bare star used to hide them from that exact check. Kept rather than narrowed, for the reason the existing entry already states: the façade's surface stays byte-identical to what the retired pin table asserted, and narrowing is a follow-up with its own review. P2 — the exhaustiveness gate skipped any source carrying a bare `export *`, which dropped that module's DIRECT exports from the check too. `gesture-plan.ts` stars `gesture-plan-types.ts`, so removing `buildDragGesturePlan` from the façade narrowed the public surface and still passed. `readDirectNamedExports` now reads exactly the names a module declares or re-exports BY NAME and ignores the star, so direct exports are checked while the starred set stays covered by the façade's own direct re-export of that module. Red evidence: removing `buildDragGesturePlan` from facades/interaction.ts now fails naming file, source and symbol; 13 pass / 0 fail restored. Third, and the reason the gate is worth having: rebasing onto main after #1612 merged silently dropped `TEXT_ENTRY_ROUTES`, `TextEntryRoute` and `TypeTextBackendResult` from the interaction façade — the same narrowing class as the #1567 one review caught by hand, one merge later. The gate failed on it before CI did. Restored. |
||
|
|
ee473b6adc |
refactor(daemon): give the Maestro fallback and ambiguous-match details real types (#1612)
Three places smuggled structured data through untyped bags and re-read it
with runtime guards. Each gets an explicit typed boundary.
A. The resolution-suppression rule was encoded twice in
interaction-touch-response.ts — a spread ternary in the runner-payload
branch and an unconditional destructure used conditionally in the runtime
branch, with the ADR 0012 rationale living on only one source variant.
Both branches now read one `suppressesResolutionDisclosure(source)`
predicate through one `applyResolutionDisclosurePolicy` helper, where the
reason is stated once. The union field is renamed
`maestroCoordinateFallbackDispatched` (the dispatch path that ran) and
hoisted into a shared base. handleFillCommand's two-arm interactor.fill
call collapses to one.
B. `Interactor.type` narrows from `Record<string, unknown> | void` to
`TypeTextBackendResult | void`; the Apple runner boundary is the single
place the wire payload becomes that type. `maestroFallbackDetails` returns
a typed `{ used, extra }` instead of a bag both call sites re-read.
C. `details.candidates` meant two incompatible things. The device-domain
resolvers now key their list `devices`, so the shared renderer drops its
shape-disambiguation guards and the device list actually renders.
|
||
|
|
a13a6832ee |
feat: add selector-targeted drag gestures (#1567)
* feat: add selector-targeted drag gestures * fix: address drag gesture review feedback * fix: satisfy drag review quality gates * fix(android): lower drag trajectories piecewise * test(replay): validate drag fixture selectors * fix(ios): ignore full-viewport chrome containers * test(drag): prove destination on live devices |
||
|
|
611858103e |
fix(ios): harden Bluesky-class interaction reliability (#1588)
* fix: type into focused iOS inputs without AX * fix: fill AX-hostile iOS text inputs * fix: keep scrolling containers from stealing taps * fix: stop agents after explicit task success * chore: format benchmark guidance * fix(ios): preserve fill semantics across fast paths * test: retire direct selector fill expectations * test: assert runtime selector fill evidence * fix(ios): preserve verified and Maestro fill paths * refactor(ios): isolate synthesized text entry * fix(client): preserve open diagnostic paths * fix(ios): expose structured text entry route * fix(packaging): strip text entry policy tests |
||
|
|
543e9f8c05 |
fix(ios): give keyboard dismiss a safe-area-tap fallback (#1598) (#1606)
* fix(ios): give keyboard dismiss a safe-area-tap fallback (#1598) The runner already tapped a keyboard's own Hide/Dismiss/Done key when the AX tree exposed one, but iPhone's default software keyboard has no such key, so `keyboard dismiss` returned UNSUPPORTED_OPERATION on the common case and agents proceeded with the keyboard (and any live QuickType predictive-text bar) still up. Live-validated on throwaway simulators before choosing a design: hardware escape key (no effect without a connected hardware keyboard), swipe-down starting on the keyboard (does not trigger UIKit's interactive dismissal on Settings/Safari/Contacts), and a private `performAction:onElement:value:error:` AX call (hung the runner for 90s on a guessed action name, force-killed by the daemon timeout) were all ruled out. The dismiss-key tap remains the primary mechanism (iPad, or any app with an inputAccessoryView Done/Cancel button); a new snapshot-derived safe-area tap is added as the disclosed last resort, computed to land outside both the keyboard and every currently-hittable element so it is a safe no-op even when it fails to dismiss. The response now discloses which mechanism actually fired (`mechanism: 'dismissKey' | 'safeAreaTap'`) across the CLI/daemon dispatch path, the SDK runtime.backend surface, and session-event summaries, so callers can tell a real dismiss-key press apart from a best-effort tap. UNSUPPORTED_OPERATION now says both mechanisms were tried. * fix: satisfy CI formatting and complexity gates oxfmt on three touched files; buildKeyboardActionSummary split so the dismiss wording (incl. the safeAreaTap mechanism disclosure) lives in its own helper below the complexity threshold. * fix: any-element obstacle rule for the safe-area dismiss tap (#1606 review P1) A role allowlist cannot prove a point is AX-empty: an unlabeled RN Pressable surfaces as a hittable Other, and a tappable parent can cover a point its static-text child does not. Every known element frame now counts as an obstacle regardless of role or hittability, with only ~window-sized structural frames exempt (isStructuralRootFrame, 95% coverage) — exempting those is what keeps the rule satisfiable, and a genuinely tappable full-screen backdrop staying exempt is the disclosed, accepted behavior of this fallback. One .any resolution replaces ten typed queries (single tree snapshot, no per-element isHittable round trips), so the stricter rule is also cheaper. * fix: drop the safe-area tap — background-tap dismissal is unsupported (#1606 review P1, round 2) No geometry or role query can prove a coordinate is side-effect-free: after the any-element rule, the structural-root exemption still deliberately removed full-screen actionable elements (RN Pressable backdrops) from the obstacle set, so the tap could navigate or submit — and report success because the mutation hid the keyboard. Per review, generic background-tap dismissal is now explicitly unsupported: the dismiss key is the only mechanism the runner vouches for, UNSUPPORTED_OPERATION says so and steers callers to press-the-next-target / keyboard enter, and the mechanism field narrows to 'dismissKey'. Unrecognized wire mechanisms degrade to the bare message and are dropped from event details. |
||
|
|
4f8dc3f31e |
refactor: move selector engine into workspace package (#1589)
* refactor: move selector engine into workspace package
* refactor(selectors): trim the package façade to its real consumers
Follow-up to the selector-package cutover, from a structural review of it.
- Drop 15 façade symbols with no consumer anywhere in the repo:
selectorUsesKey (added by the cutover, never called), isNodeVisible /
isNodeEditable (the real helpers are contracts/snapshot's), normalizeText,
splitIsSelectorArgs, IS_PREDICATE_REQUIRED_MESSAGE, four nested Replay
types, SelectorDisambiguationDisclosure, and the four kernel type
re-exports every consumer already imports from kernel directly.
- Delete SelectorCapturePolicyInput.selectorExpression, which
deriveSelectorCapturePolicy never read; the policy varies only by
predicate, so it takes one now. Two of the four tests asserted that the
unread parameter had no effect and could not fail; they go with it.
- Return the Maestro export vocabulary to the maestro package. The cutover
inlined MAESTRO_TEXT/STATE_SELECTOR_KEYS' values into the CLI call site,
leaving both constants dead in the package that owns the concept and no
gate over the two copies. MAESTRO_SELECTOR_PROJECTION is now the one
statement of it.
- Dedupe SelectorDiagnostics and SelectorDisambiguationDisclosure, declared
character-for-character twice across the AST/string seam, and name the two
shared option shapes once instead of five inline copies. The parser-side
resolution types take an Ast prefix so the twins read as twins.
- Delete three identity wrappers: parsePrivateSelector,
selectorExpressionToMaestro, and the formatSelectorFailure forwarder —
nothing passes it a chain any more, so the SelectorChain | string union
and its branch go too.
- Delete internal/index.ts, an AST barrel whose only consumer was one test
in the same directory (renamed to engine.test.ts), and the match.ts
pass-through that existed to feed it.
- ReplaySelectorGrammar had three variants for two behaviors; 'wait' and
'ordinary' were the same path. It is 'is' | 'positional' now.
- Drop the deleted src/sdk/selectors.ts from .fallowrc.json's entry list.
Behavior unchanged. pnpm check green: 598 unit files / 5278 tests, smoke
35 passed / 3 live skipped, layering 71/71, depgraph 22/22, mutation config
45/45, fallow clean, package smoke sound. Counterfactual: pointing
MAESTRO_SELECTOR_PROJECTION.textKeys at the state keys turns three
replay-maestro-export cells red; restored before commit.
* test(selectors): split the engine aggregation test by source concept
`internal/index.test.ts` (renamed `engine.test.ts` when its barrel went away)
was a 708-line aggregation over the whole engine — past the 500-line tripwire
and mirroring no source module, so it also ran as one serial unit.
It becomes five files that each mirror what they test, plus the parser cells
folded into the existing parse test:
resolve.test.ts alternative fallback, strict uniqueness,
first-match existence
resolve-disambiguation.test.ts ADR 0012 ranking: deepest, smallest-area,
winner-vs-challenger disclosure, tie fallback
resolve-viewport.test.ts the visibility half: on-screen beats
off-screen, including inside an off-screen
scroll container
match.test.ts per-key matching semantics (text, role,
focused, appname/windowtitle, decoded
newline labels)
arguments.test.ts where the selector ends and the command's
positionals begin, both grammars
parse.test.ts +6 grammar/escape cells beside the existing
property tests
The login-form tree shared by resolve.test.ts and match.test.ts moves to
`__tests__/login-form-nodes.ts` rather than being copied into both.
All 27 cells are carried over unchanged and still pass; no file now exceeds
224 lines. pnpm check green: 602 unit files / 5278 tests, layering 71/71,
depgraph 22/22, mutation config 45/45, fallow clean over 127 changed files.
* revert(selectors): keep agent-device/selectors public, behind one AST subpath
The cutover removed the `agent-device/selectors` public subpath as part of
tightening the API. It is in use, so the removal is reverted: the subpath ships
the same ten symbols v0.20.5 shipped, with the same signatures.
That has to coexist with the reason the package façade is string-only, so the
AST leaves through one named door instead of the main one:
@agent-device/selectors string-in/string-out; every in-repo consumer
@agent-device/selectors/ast the published parser surface; one consumer,
src/sdk/selectors.ts
`packages/selectors/src/ast.ts` re-exports parseSelectorChain,
tryParseSelectorChain, isSelectorToken, the AST-taking findSelectorChainMatch
and resolveSelectorChain, isNodeVisible, isNodeEditable, and types
SelectorChain / SelectorDiagnostics. `formatSelectorFailure` keeps its
published `SelectorChain | string` first parameter as a shim here rather than
widening internal/resolve.ts back to a union — the compatibility obligation
sits at the boundary that owes it.
This is strictly narrower than main, where the AST was reachable from anywhere
in src/ via src/selectors/*. Two gates hold it there: facade-symbols.ts pins
./ast to exactly the v0.20.5 list, and package-boundaries.test.ts asserts
src/sdk/selectors.ts is the only file outside the package that imports it.
Restored alongside: the ./selectors export and tsdown entry/chunk group, the
.fallowrc.json entry, the package-exports supported-subpath list, and both
client-api.md sections. No CHANGELOG entry — nothing is removed any more.
pnpm check green: 602 unit files / 5278 tests, smoke 35 passed / 3 live
skipped, layering 71/71 (10 packages, 32 subpaths), depgraph 22/22, mutation
config 45/45, fallow clean over 129 changed files, package smoke imported all
12 published entry points with publint and attw passing. Verified functionally
against the built dist: the doc's parse -> findSelectorChainMatch example
returns the same shapes as before, resolveSelectorChain still returns an AST
`selector`, and formatSelectorFailure still accepts a chain.
* fix(selectors): correct the two expectations that still assume the removal
Review P1s on a792415a: restoring the public subpath left two gates asserting
it was gone.
- installed-package-metro.test.ts moved `agent-device/selectors` into the
blocked-specifier list. It goes back to the subpath smoke set, running the
same `isSelectorToken('||')` + `parseSelectorChain` check it ran before the
removal, so the file's only remaining delta from main is a formatter reflow.
- owner-files-no-leak.test.ts asserted `dist/src/sdk-selectors.js` was absent.
It requires the stable named chunk again, and still rejects an auto-numbered
`selectors2.js` fallback — the pair is what proves the restored tsdown chunk
group is doing its job, verified against a clean build.
PR body corrected: the removal is no longer described as intentional API
tightening.
* refactor(selectors): satisfy the widened fallow scope after rebase
main's #1591 (the follow-up filed from this review) removed `packages/**` from
.fallowrc.json's ignorePatterns, so the new package is audited for the first
time. Everything below is a finding fallow could not previously see.
Dead surface, all confirmed consumer-free:
- 12 type re-exports from the `.` façade whose shapes consumers only ever
reach structurally.
- MAESTRO_TEXT_SELECTOR_KEYS / MAESTRO_STATE_SELECTOR_KEYS, orphaned by this
branch's own MAESTRO_SELECTOR_PROJECTION change, and the test-util
SELECTOR_VALUE_HAZARDS. All three are module-local now.
- IS_PREDICATE_USAGE_HINT fails --production because its only consumer is the
is-argument-surface parity test. It gets a commented `ignoreExports` entry
rather than deletion: the constant is what makes the daemon and CLI raise
ONE hint instead of two copied strings (ADR 0010), so the test asserting
that is the point, not an accident.
`fast-check` is now declared by the package that imports it.
Duplication, split by what could be proven:
- `isUsefulVisibilityAnchor` existed character-for-character in both
packages/selectors and packages/maestro. Moved to
@agent-device/contracts/snapshot, which both already depend on and which
already owns this vocabulary. Safe because the `normalizeType` each copy
called is itself character-identical to the contracts one — checked before
moving, since a different normalizer would have silently changed which
nodes anchor.
- maestro additionally reimplemented `normalizeType`, `buildSnapshotNodeMap`
(as `buildSnapshotNodeByIndex`) and `findSnapshotAncestor`, all
character-identical to contracts'. Deleted in favour of the shared ones.
- The three scroll-ancestor walks are NOT deduped. They are structurally the
same walk but each uses a different scrollable predicate, and I have no
evidence the three agree; collapsing them would be a Maestro-conformance
change, not a cleanup. Both maestro sites now say so, and the work is filed
separately.
`projectSelectorExpression` (15 cyclomatic / 22 cognitive, written by the
cutover) splits into a dispatcher plus `readAgreedTextValue` and
`projectSelectorTerms`; all three are under threshold.
Rebase note: the one conflict, in package-boundaries.test.ts, resolved to
NEITHER side — #1591 had already deleted `AdReplayVerifiedTargetGuard` as an
unused export, and this branch deletes the seven ReplaySelectorPort names, so
the conflicting block is empty.
* build: record fast-check for packages/selectors in the lockfile
Declaring the dependency in packages/selectors/package.json without
regenerating pnpm-lock.yaml made every CI job fail in its install step with
ERR_PNPM_OUTDATED_LOCKFILE. My local `pnpm install --frozen-lockfile` printed
"+ 1 dependencies were added: fast-check@^4.9.0" and exited 0, which read as
success but was the same mismatch CI refuses.
Regenerated with the pinned pnpm 11.17.0, not the 11.5.3 on this machine:
11.5.3 rewrites peer-dependency resolution keys repo-wide (dropping
`(supports-color@7.2.0)` suffixes) and produced a 222-line diff. With the
pinned version the diff is the 4 lines this change actually needs, plus
pnpm's alphabetical re-sort of the root selectors entry.
|
||
|
|
351ef7a14f |
refactor(android): enforce transport lowering in the type system (#1583)
`AndroidLoweredTouchPlan` widened the canonical two-sample trajectory to a plain sample array, so a plan that skipped `lowerAndroidTouchPlan` still satisfied the transport types. That is the mistake the lowering exists to prevent: an un-lowered plan injects a two-sample gesture, which is the sparse delivery #1572 removed from the shared plan in the first place. Transport samples are now a minimum-arity tuple. `sampleGestureOffsets` floors the frame count at three, so lowering always yields at least four samples, which makes "denser than the canonical endpoint pair" a true statement about the data rather than a comment. A canonical plan is no longer assignable, so skipping the lowering fails typecheck at every injection seam. Tightening the type caught three call sites that were passing un-lowered plans straight to the helper transport, which is the evidence the previous signature enforced nothing. `longPressPlan` now returns `AndroidLongPressTouchPlan` instead of the wide union it never produced, and the dual-pointer normalize test routes through the lowering like every other transport call. Also drops the unused `= 'default'` on `sampleGestureOffsets` so every caller states which platform sampling convention it wants, which is the point of having centralized the policy. Sample values are unchanged by construction, so Android injection stays bit-identical to #1572. Co-authored-by: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
bcaa106845 |
refactor: extract snapshot and replay identity semantics (#1582)
* refactor: extract snapshot and replay identity semantics * refactor: move pure rect primitives from contracts to kernel/rect containsPoint, pickLargestRect, and isRectVisibleInViewport are raw rectangle arithmetic with no snapshot awareness, so they belong beside rectContains/rectArea in @agent-device/kernel/rect rather than in the snapshot-semantics vocabulary. The node-aware resolveViewportRect folds into contracts/snapshot-visibility.ts, retiring snapshot-geometry.ts; after this split, everything behavioral in @agent-device/contracts/snapshot is policy that interprets the snapshot model. * refactor: restore ADR-0012 rationale docs and dedupe replay identity shapes The #1478/#1581 extraction moved the identity/structural helpers but compressed their invariant documentation to one-liners; the deliberate no-ancestry-exclusion rule on idMatchCountInTree, the fail-closed guard comparison, and the who-throws/who-detects contracts on the two divergence reason markers now travel with their definitions again. LocalIdentity and NodeStructuralDenotation move to contracts/target-annotation.ts (beside TargetAncestryEntry, which WaitLandmarkMismatchEvidence now references directly), so the guard shapes in contracts/replay.ts are nominal instead of hand-rolled structural twins; ad-script re-exports the vocabulary beside the readers that produce it. Also inlines the demoteNonUniqueId pass-through wrapper in session-target-evidence.ts. * style: fix oxfmt formatting in snapshot-visibility * refactor: move findSnapshotAncestor into contracts/snapshot-tree The last root value import from src/selectors: predicates.ts reached src/snapshot/snapshot-processing.ts for the ancestor walker. The walker is index-based tree traversal with no presentation policy, so it joins buildSnapshotNodeMap in contracts/snapshot-tree.ts; both consumers repoint to the façade and the non-contiguous-index/cycle coverage moves to the package test. src/selectors now has zero value imports from root src in production files. |
||
|
|
56d9ee605c |
fix(ios): preserve timed pan duration (#1572)
* fix(ios): preserve timed pan gesture execution * fix(gestures): encode linear pans as endpoint plans * fix(ci): pin wait contract exports * fix(android): lower endpoint gesture plans for touch transport * fix(gestures): preserve timed pan duration across adapters |
||
|
|
6baa5d97e1 | fix: make wait verdicts evidence-based (#1570) | ||
|
|
2e74b789fd |
feat: verify device cloud connections (#1564)
* feat: verify device cloud connections * refactor: unify connect provider adapters * refactor: separate connect verification facts * fix: tighten connect provider verification * fix: use neutral cloud connection wording * perf: deduplicate local affected checks * refactor: simplify affected check runner * refactor: derive connect workflow from verification |
||
|
|
99967c7f01 |
fix: restrict project config trust (#1565)
* fix: restrict project config trust * fix: preserve daemon auth transport context * refactor: simplify project config trust * fix: restrict project config write sinks |
||
|
|
123521652c |
fix(ios): double-check off-screen click refusals against a direct element read (#1566)
* fix(ios): double-check off-screen click refusals against a direct element read #1542: after an AX-free scroll on iOS, the off-screen interaction guard can refuse a click even though the target is genuinely on-screen, because it trusts a scroll-container ancestor's rect from the bulk accessibility tree, which a keyboard-dismiss content-offset correction can leave stale/corrupted while the target's own rect is already correct. When the guard is about to refuse on iOS, it now takes a single fresh, tree-independent XCUITest read of the target element (querySelector) and trusts that read's live `hittable` + rect-vs-root-viewport signal instead, if it positively confirms on-screen. Any failure to unambiguously re-resolve the element (no id/label, not found, ambiguous, transport error) fails closed exactly as before. Genuinely off-screen targets, and every other platform, are unchanged: the backend method is gated to local (non-provider) iOS sessions only, and only ever runs on the about-to-fail path. The decision itself is a pure function (decideOffscreenRefusalDoubleCheck in mobile-snapshot-semantics.ts) with counterfactual-proven tests: hardcoding it to always trust the bulk verdict turns the rescue test red, and hardcoding it to always trust the direct read (including on "unavailable") turns the fail-closed/genuine-refusal test red. Live-validated on a fresh-boot iOS simulator: checkout-form.ad 2/2 passes (previously failing at step 11), gesture-lab.ad 2/2 (regression), and the Android checkout-form/gesture-lab suite passes unchanged, proving no cross-platform behavior change. * fix(ios): tap the live rect after a rescued offscreen refusal; collapse the double-check to one backend hook Review blockers 1+2 (interleaved by design — the soundness fix is expressed through the collapsed hook's contract): 1. SOUNDNESS: a rescued refusal now returns the node PATCHED WITH THE LIVE RECT the backend confirmed, and every downstream use (tap point, response) reads from that returned node — never the original. In the frozen-tree manifestation (the whole bulk tree pinned at pre-gesture values), the original rect can be stale even when the rescue verdict is correct; tapping it would have silently landed at the wrong coordinate. New regression: offscreen-double-check.test.ts's frozen-tree case, with a counterfactual (revert to computing the point from the pre-guard node) proven red then reverted. 2. SURFACE: collapsed to ONE optional backend hook, `confirmOffscreenTargetVisible?(context, node, rootViewport): Promise<Rect | null>` — conceptually a boolean, but returns the live rect so item 1's fix has something to act on. Deleted decideOffscreenRefusalDoubleCheck, the OffscreenRefusalDoubleCheckSignal/Reading ADT, and resolution.ts's dual-signal reconciliation shell: the bulk side was hardcoded 'off-screen' at the only call site, so the two-signal model was dead weight. The shared guard is now: bulk-off-screen -> ask the hook -> a live rect proceeds (patched), anything else (including no hook) throws exactly as before. The pure geometry boundary that decision reduces to (`isConfirmedOnScreenProbe` in mobile-snapshot-semantics.ts, replacing the deleted ADT) is unit-tested with two counterfactuals: ignoring `hittable` and ignoring the viewport containment check each turn a test red (proved, then reverted). `throwIfOffscreenInteractionTarget` is now exported (ADR 0011 registry honesty, see the contracts commit) and directly unit-tested in resolution.test.ts, mirroring the existing tryResolveRefNode pattern. * refactor(ios): direct-ios-selector.ts back to pure gate/parse; reuse queryDirectIosSelector Review blocker 3 (BOUNDARIES): - direct-ios-selector.ts no longer does any runner I/O — it's back to pure gate/parse (readSimpleIosSelectorTarget, deriveDirectIosNodeSelector, isDirectIosSelectorFallbackError) plus the ONE shared eligibility predicate, isLocalIosRunnerSession(session, { skipPendingPostGestureStabilization }). Both the direct-selector tap fast path and the new offscreen double-check probe call this same function; the one behavioral difference between them (the tap fast path skips a session with a pending postGestureStabilization, the double-check does not) is now an explicit parameter instead of two separately-written gates. - The probe I/O moved to a new sibling, src/daemon/offscreen-target-probe.ts, which reuses selector-runtime.ts's `queryDirectIosSelector` (now exported and decoupled from SelectorRuntimeParams — it takes a session + a bare {key, value} selector + AppleRunnerRequestOptions) rather than opening a second querySelector client. Node extraction (`readDirectIosSelectorNode`, the one `as SnapshotNode` cast) stays singular, inside selector-runtime.ts. - interaction-runtime.ts wires confirmOffscreenTargetVisible only when isLocalIosRunnerSession(session, { skipPendingPostGestureStabilization: false }) — deliberately NOT skipping a pending post-gesture stabilization, since that is exactly the window the double-check exists to cover. * docs(contracts): name the iOS offscreen rescue hook as part of the guarantee matrix Review blocker 4 (GUARANTEE HONESTY): the shared offscreen cell (RUNTIME_TREE_SHARED_GUARANTEES.offscreen, used by runtime-selector and runtime-ref) and the native-ref path's offscreen cell still named isNodeVisibleOnScreen as sole enforcement after #1542's double-check landed — that understates what actually enforces the guarantee now. Both cells' `via` now point at throwIfOffscreenInteractionTarget (exported from resolution.ts in the prior commit for exactly this), the real end-to-end enforcement point: isNodeVisibleOnScreen is the bulk-tree decision it starts from, and on iOS a would-be refusal can still be confirmed via the optional AgentDeviceBackend.confirmOffscreenTargetVisible hook before erroring. The cell's comment states the rescue-only, fail-closed shape explicitly per ADR 0011's matrix rules — this does not weaken the cell, it extends its description to match reality. iOS rescue policy stays OUT of resolution.ts's shared docstrings (the "spine"): this registry file is where per-path enforcement detail belongs, and the optional-method wiring in interaction-runtime.ts remains the only cross-platform touch. The registry's own gate test (interaction-guarantees.test.ts) still passes: every `via` resolves to a real exported symbol. * test(ios): move #1542 offscreen double-check tests out of interaction.test.ts Review blocker 5 (TEST HOMES): AGENTS.md forbids adding to daemon/handlers/__tests__/interaction.test.ts (it predates the test-mirrors-source-topology rule and shrinks opportunistically). Reverts the 172 lines added there in the original PR version; interaction.test.ts is back to its pre-#1542 baseline (81 tests, unchanged). The same assertions now live in their proper homes (see the prior three commits for the sources they cover): - pure decision pin: src/utils/__tests__/mobile-snapshot-semantics.test.ts (isConfirmedOnScreenProbe, with the two counterfactuals) - direct-guard pin: src/commands/interaction/runtime/resolution.test.ts (throwIfOffscreenInteractionTarget, mirroring tryResolveRefNode) - probe unit tests: src/daemon/__tests__/selector-runtime.test.ts (queryDirectIosSelector) and src/daemon/__tests__/direct-ios-selector.test.ts (isLocalIosRunnerSession, deriveDirectIosNodeSelector) - probe integration: src/daemon/__tests__/offscreen-target-probe.test.ts (confirmIosOffscreenTargetVisible, mocked runner) - end-to-end rescue/refuse, including the frozen-tree live-geometry regression + its counterfactual: new sibling src/commands/interaction/runtime/offscreen-double-check.test.ts (next to resolution.ts, using the same createInteractionDevice harness resolution.test.ts already uses) * style: oxfmt formatting for resolution.test.ts |
||
|
|
2c2df031ff |
feat: keep replay session active on request (#1554)
* feat: keep replay session active on request * test: cover replay keep-session provider route * fix: make replay session handoff reliable * refactor(daemon): extract the replay terminal-lifecycle policy module (#1554 review) session-replay-runtime.ts was already over the 500-line extract-before-adding-behavior tripwire before this PR; the keep-session/repair terminal-close decision, its live-session postcondition, and the dispatched-action count pushed it further past budget. Move that policy into a focused session-replay-terminal-lifecycle.ts (isExecutableReplayAction, resolveSuppressedTerminalCloseIndex, countExecutedReplayActions, requireLiveSessionForKeepSession) so the runtime file stays orchestration-only, and mirror its PR-added unit tests into session-replay-terminal-lifecycle.test.ts. Pure extraction: no assertions changed. |
||
|
|
92b22229e6 |
feat(cloud-webdriver): BrowserStack device-feature capabilities, and fix cloud orientation (#1544)
* feat(cloud-webdriver): support BrowserStack device-feature capabilities
Adds the eight BrowserStack "device feature" session capabilities that had no
representation in agent-device: deviceOrientation, geoLocation, timezone,
language, locale, networkProfile, customNetwork, and resignApp.
These are vendor capabilities, so they are emitted inside `bstack:options`
rather than at the top level. BrowserStack's YAML config lists them unnested
and its SDK relocates them; agent-device talks to the hub directly, so it
nests them itself.
A single spec table drives both the flag reader and the capability builder, so
adding a capability is a table row rather than a branch in each. A structural
test asserts every field owns exactly one row, since a field the table forgets
would parse off the CLI, ride the profile, and then be silently dropped before
the hub ever saw it.
Rejects combinations the provider cannot act on unambiguously: an unknown
orientation is caught at the flag boundary instead of being forwarded to a hub
that accepts and then ignores it, --provider-no-resign-app is refused on
Android, and a named network profile cannot be combined with a custom network
shape.
Also fixes a latent shallow-merge bug in buildBrowserStackCapabilities: a
caller supplying its own `bstack:options` replaced the whole object and
silently dropped the project, build, and session labels. It is now merged
per key.
* fix(cloud-webdriver): rotate via WebDriver orientation endpoints
`setOrientation` on the cloud WebDriver path sent `mobile: rotate`, which is
not a driver command at all. UiAutomator2's own error enumerates its
extensions and `rotate` is absent from the list, so `agent-device orientation`
was a hard failure on every hosted provider.
It also forwarded agent-device's four-way rotation vocabulary verbatim
("landscape-left", "portrait-upside-down"), where the protocol accepts only
uppercase PORTRAIT/LANDSCAPE. Every other platform has a translation layer;
this path was the only one without one.
Now two transports, ordered by backend. `POST /rotation` takes exact four-way
degrees and leads on Android, since it is the only endpoint that can express
upside-down and left-versus-right. `POST /orientation` is two-way and leads on
XCUITest, which rejects `/rotation`. Each falls back to the other, because only
BrowserStack's UiAutomator2 is verified and a provider whose driver disagrees
should degrade rather than hard-fail.
Verified live against BrowserStack App Automate:
POST /rotation {"x":0,"y":0,"z":0} -> 200 {"value":"ROTATION_0"}
The rotation-to-surface-index mapping moves to contracts/device-rotation.ts and
the existing adb path now reads from it, so the local and hosted mappings
cannot drift apart.
Note this rotates the current display, not persistent device rotation, so an
activity that does not pin its own orientation may still need rotating once it
is in the foreground.
The capability was declared "partial" without the transport existing, and no
test covered setOrientation on the cloud path; only adb and the Apple runner
were covered. Both gaps are now closed.
* fix(cloud-webdriver): narrow orientation fallback and gate provider-owned flags
Addresses review on #1544.
The orientation fallback caught every error, so a timeout, an auth rejection, a
dead session or a provider 5xx on the first transport was swallowed and retried
against the second. When that one also failed the caller got "rejected both
endpoints" with the real cause discarded. Fallback is now keyed on structured
unsupported-endpoint signals only — HTTP 404/405, or a W3C `unknown command` /
`unknown method` code — matching the repo rule of keying on typed details rather
than message text. Everything else rethrows unchanged.
Device-feature capabilities are BrowserStack-owned, but the flags were accepted
by any cloud provider, persisted into the generated profile, and then silently
dropped at session creation. `connect aws-device-farm` now rejects them with a
typed error naming each offending flag, raised before the provider's own
required-argument checks so the caller is told what is unsupported rather than
what else is missing. Ownership is modelled on the capability spec table, so a
new capability inherits the guard without a second list to maintain.
Adds provider-backed orientation scenarios driven through public daemon dispatch
against the fake WebDriver provider: the four-way endpoint on the happy path,
the documented collapse onto the two-way endpoint when the driver does not
implement `/rotation`, and a provider 5xx that must surface without consulting
the second transport. The fake server's route handling became a table in the
process — it had grown to ten branches in one function.
* fix(cloud-webdriver): read W3C error codes before status, enforce ownership at the runtime boundary
Addresses the second review pass on #1544.
The fallback classifier returned on any 404/405 before consulting the W3C error
code, so an HTTP 404 carrying `invalid session id` was masked as a missing route
and retried against the second transport. The structured code now takes
precedence whenever the driver sent one; bare status is consulted only when no
code exists. Two cases pin it: a 404 `invalid session id` and a 405 `timeout`
must both surface rather than fall through.
Provider ownership was enforced only in the CLI profile builder, which the typed
client and hand-authored remote-config profiles bypass entirely — both reach
session preparation without passing through `connect`, so the capabilities were
accepted and then dropped. The check now lives on the capability-ownership
module and runs inside AWS Device Farm's `prepareSession`, with the CLI builder
calling the same helper instead of its own copy. Covered by a scenario that
drives the runtime boundary directly and asserts the rejection happens before
any provider session is created.
|
||
|
|
b125435989 |
refactor: extract WebDriver provider package (#1504)
* refactor: extract webdriver provider package * refactor: consolidate shared XML codec |
||
|
|
a3ab69a110 |
refactor(replay-test): neutralize the values crossing the scheduler seam (#1478 P3, part 1) (#1509)
* refactor(replay-test): neutralize the values crossing the scheduler seam
#1478 P3, part 1 of 2. Prepares the replay-test extraction by removing every
non-neutral value that crosses the scheduler seam, in place under `src/`, so the
physical move to `packages/replay-test` is a file move rather than a redesign.
`DaemonResponse` no longer crosses the seam. `session-test-types.ts` typed
`runReplay`/`finalizeAttempt` as returning a daemon response and the scheduler read
`.error.code`, `.error.details`, and `.data.replayed/.healed/.warnings/
.snapshotDiagnostics` off it throughout. That is invisible to R10 today only
because `checkDaemonTypesImporters` skips `src/daemon/`; once the files live in a
package they become external `daemon/types.ts` importers, which the ratchet only
lets shrink. Attempts now resolve as tagged `ReplayTestAttemptOutcome` values
carrying exactly what the scheduler consumes, including an `infrastructure` tag —
classifying an environmental failure needs platform boot-diagnostic vocabulary the
scheduler must not import, so the host decides and the scheduler reads the verdict.
`session-test-outcome.ts` is the one place a daemon response becomes an outcome.
Step events get a narrow per-attempt port. They were emitted from
`session-replay-runtime.ts` and `session-replay-maestro-observer.ts`, both reading
a request-global `AsyncLocalStorage` seeded per attempt. The scheduler now hands
each attempt an `onStep` sink, threaded the way `tracePath` already is; both
engines call it and `withReplayTestActionProgress`/`readReplayTestActionProgress`
are gone. A direct `replay` simply has no sink.
ADR 0012 divergence becomes a neutral leaf. `src/replay/divergence.ts` depended
only on kernel contracts and redaction, yet Maestro constructs divergences too and
CLI/MCP both render them, so P5 could not have moved it into `packages/ad-replay`.
It is now `@agent-device/contracts/divergence`; the renderer's output text is
unchanged.
The progress wire vocabulary moves to `@agent-device/contracts/progress`. It is
serialized by `request-progress-protocol.ts` and reconstructed by the CLI reporter
path, so it belongs below both; `src/request/progress.ts` keeps only the sink and
its AsyncLocalStorage binding.
Together these clear all four of replay-test's recorded R10 migration imports, so
the rule now enforces unconditionally for that module.
Behavior is unchanged. The shipped reporter contract — export spellings,
object/factory loading, hook names, timing/order, value fields, the synchronous
live-hook rule, awaited suite completion, error handling, exit codes — is
untouched, and `session-test-reporter-values.test.ts` passes unmodified. The
`--shard-all` `total`/`runnable` asymmetry is preserved as characterized.
Refs #1478
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8
* test(replay-test): pin the Maestro reporter step path against the onStep port
Review finding on #1509: the native `.ad` reporter ratchet exercises only one of
the two `onStep` forwarding chains, so deleting a link in the Maestro chain would
silently stop `onTestStep` for every `test --maestro` run while every existing
reporter test stayed green. Same defect class as the dropped diagnosticId/logPath
(#1501) and the dropped reporter `hint` (#1505).
Maestro is one of P3's two required real adapters and its chain shares no links
with the native one below `runReplayScriptFile`:
scheduler sink -> runReplayScriptFile -> runTypedMaestroReplayFile
-> createMaestroReplayObserver({ onStep }) -> actionStarted -> onStep
Adds a Maestro scenario driving `test --maestro` through the real session handler
and the real reporter registry. It asserts the step payload the engine produces
(`stepIndex`/`stepTotal`/`stepCommand`/`stepValue`, including that a value-less
command stays value-less) together with the attempt/session identity the scheduler
supplies, since that half of the event came from request-global AsyncLocalStorage
before P3. A second case drives a retry so step events must carry attempt-1's
session and then attempt-2's. The flow `name` also pins the reporter `title`, a
value only the Maestro path can produce.
New file rather than an addition to session-test-reporter-values.test.ts: that file
is the pinned characterization and must keep passing unmodified, and Maestro needs
its own vi.mock of core/dispatch for device resolution.
Counterfactual run, both links, each restored after:
- dropping `onStep` from createMaestroReplayObserver in
session-replay-maestro-runtime.ts
- dropping the emitMaestroStep call from actionStarted in
session-replay-maestro-observer.ts
Each dropped both onTestStep events ("expected [ 'onSuiteStart', 'onTestStart',
…(2) ] to deeply equal [ 'onSuiteStart', 'onTestStart', …(4) ]") and failed both
new cases, while session-test-reporter-values.test.ts passed all 4 — exactly the
hole the reviewer identified.
Test-only; no production change. Bundle output is byte-identical to
|
||
|
|
32b9db2d7a |
test: add Android full emulator coverage (#1484)
* test: add Android full emulator coverage * ci: package Android helpers before nightly coverage * fix: harden Android nightly runtime evidence * fix: expose trace artifacts in MCP schema * style: format Android coverage manifest * fix: address Android coverage review findings * refactor: share live device coverage helpers * fix: restore fixture landmarks in device smokes * refactor: centralize live artifact assertions * fix: normalize fixture canary visibility |
||
|
|
0e51007b04 |
refactor: isolate maestro engine package (#1506)
* refactor: isolate maestro engine package * perf: deepen maestro facade boundaries |
||
|
|
0ee2a86129 |
refactor: extract contracts workspace package (#1499)
* refactor: extract contracts workspace package * fix: preserve screenshot diff result contract * test: stabilize Android keyboard smoke |