mirror of
https://github.com/callstack/agent-device.git
synced 2026-09-14 20:06:34 +08:00
v0.20.10
1064 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
9fb307aa63 |
fix: derive transition backend from capture (#1851)
* fix: derive transition backend from capture * fix: retain iOS transition confirmation without provenance * test: guard backendless local settle latency * fix: confirm transitions from modal captures * fix: confirm transitions from tiny iOS modals * fix: recover settle transition baseline from session * fix: trust ref frames for settle transitions * chore: format transition settle tests * refactor: simplify settle transition decisions * test: isolate private ax recovery budget * test: keep settle test within size ratchet |
||
|
|
275a66ea23 |
test(ratchet): pin snapshot-handler.test.ts at main's 2654 lines (#1860)
#1847 grew the file 2652→2654 and merged minutes before #1843 pinned it at 2652 (measured against the merge-base #1843 had at the time). Both PRs were green alone and main is red together — the cross-PR growth the equality pin exists to catch, landing as a catch-up rather than a raise: the history rule agrees (2654 at the merge-base). |
||
|
|
b12a3e3cb3 |
test: pin test files over 1,000 lines at their exact length so they can only shrink (#1843)
* test: pin test files over 1,000 lines at their exact length so they can only shrink AGENTS.md has said for a while that past 1,000 lines is architecture debt and tests are not exempt; nothing enforced it, and the second-largest test file gained 55 lines in the PR before this one. This is the slow-test ratchet's shape for a reader's context instead of wall clock: the 26 test files over the tripwire are pinned at their exact length (R9-style equality pin, #1781 A6); growth fails, shrink fails until the pin is lowered in the same PR, a file that drops under the line leaves the list, and a new file may not cross it. One directory walk per unit run, ~250ms; the pin list emptying deletes it. * test(ratchet): hold giant test files to their merge-base length so pin edits cannot admit growth Review (P1): the equality pin compared measured lengths only against the pin map in the same checkout, so growing a file and raising its pin, or adding a new >1,000-line file with a pin, stayed green. The gate is now history-backed: every test file over the tripwire may be no longer than at the merge-base with origin/main (renames followed; new files may not cross the line), and no pin may exceed its file's base length — one git cat-file --batch spawn, parsed by bytes because the sizes are bytes. Both bypasses planted red against real git on a pinned file and on a fresh 1,001-line file with a pin added. * test(ratchet): a pin on a file at or under the tripwire is itself a finding Review: a new pin for an unchanged sub-tripwire file (900 pinned at 900) passed equality and history and grew the map. Pins now exist only for files over the tripwire — any other pin is red with 'remove it' — which also subsumes the old shrink-under-the-line message. Planted red in-file and against real git (a 186-line test pinned at 186). The android snapshot test pin bootstraps 1636→1660: main grew that file in #1846 before this gate exists, and history agrees (1660 at the merge-base). |
||
|
|
3f0f706f0b | refactor: migrate diff to request-bound runtime (#1847) | ||
|
|
ee13203a16 |
feat(ios): unify snapshot eligibility (#1850)
Make iOS regular snapshot eligibility one backend-neutral presentation rule. Acquire tree nodes conservatively, preserve interactive scroll containers, normalize surviving hierarchy, and keep raw membership plus daemon publication policy unchanged. Part of #1797. - iOS and macOS unit-enabled runner builds - 2 focused XCTest cases - 3 production-path publication tests - live Settings snapshots: 73 regular nodes and 167 raw nodes, both healthy tree captures |
||
|
|
294654a3a5 |
fix(android): resolve snapshot scope once and disclose the API 23 occlusion-scan gap (#1832 C1/C2) (#1846)
* fix(android): resolve snapshot scope once and disclose the API 23 occlusion-scan gap (#1832 C1/C2) - Android resolves --scope inside its projection only, under the shared scope specification (matchesSnapshotScope in @agent-device/contracts/snapshot: first document-order match over label/value/identifier, empty on no match). The daemon post-wire scopeSnapshotNodes pass skips the android backend, so scope has one owner and one no-match semantics instead of BFS+fallback followed by document-order+empty. - Golden table contracts/fixtures/snapshot-scope-policy.json is asserted against the predicate, the Android projection, and the daemon pass; the Swift runner twin (#1797) consumes the same table. - androidSnapshot.occlusionScanUnavailable discloses helper trees without drawing-order (API 23), where the covered-sibling pruner cannot run. Disclosure only; C1 stays open until occlusion moves to the daemon annotator. * fix(android): resolve scope over the presented tree and stop dropping it on interaction captures Adversarial review findings on the first commit: - BLOCKER: captureSnapshotData spread `snapshotScope: undefined` over flags, so an interaction capture (press/click/fill/longpress/hover --scope, --settle observation) reached the Android platform unscoped while buildSnapshotState still saw the scope. The post-wire pass used to rescue it; after skipping android it returned the unscoped tree. One effective scope now feeds both. - Scope resolves over the PRESENTED nodes of the requested projection, not the acquired tree, so an acquired match that membership drops no longer empties the snapshot, and Android matches the domain iOS's pass uses. - Slicing after the walk keeps ancestor context (hittable/collection/chrome) above the scope root, so scoped -i is a subset of unscoped -i; --depth stays scope-relative. - Shared findSnapshotScopeRange/reindexSnapshotNodes so the daemon pass and the Android projection run one implementation; scope slice extracted to ui-hierarchy-scope.ts (mirrors its test). - parseUiHierarchy moved to a test fixture module (it had no production caller left). - Golden rows sharpened (value row no longer matches via label on Android); the Android leg runs raw AND regular. CHANGELOG entry; docs wording corrected for iOS/@ref. * fix(layering): keep the contracts snapshot façade exhaustive over snapshot-scope * fix(android): scope to the first match whose subtree still has presented content Review P1 on #1846: with scope resolved strictly over presented nodes, `snapshot -i --scope panel` answered "no nodes" whenever the matched container was a structural view membership drops — even though the button inside it was exactly what was asked for — and `--depth 0` then hid a node the response prints at depth 0. The scope root is now the first document-order match whose subtree contributes at least one node to the requested projection, and the result is that subtree's presented nodes re-rooted at depth 0. Both failure modes die: a decorative match membership drops no longer empties the snapshot, and a dropped container still scopes to its content. `--depth` under scope filters the depths the response emits, so a node shown at depth 0 survives `--depth 0`. Tests: the golden legs stay raw+regular (bare TextViews cannot survive -i, so an -i leg would measure membership, not scope) with the projection interplay pinned by two dedicated tests on actionable shapes; the 'not re-scoped after the wire' case now runs a real parsed scoped tree instead of fabricated depth-0 siblings. |
||
|
|
f03c0309a1 |
fix: derive iOS transition snapshots from visible presentation (#1831)
* fix: project iOS transition semantics * fix: derive iOS transition semantics from visible state * fix: preserve iOS presentation context for scoped snapshots * fix: confirm broad iOS transition settlement * ci: run coordinate input regression on pull requests * test: mock migrated snapshot capture seam * fix: confirm transitions across snapshot backends * fix: arm transition confirmation after first capture * fix: settle against immutable action baseline |
||
|
|
6a8beb653e |
feat(mcp): compact server instructions in both eras + MCP-only help tool (#1839)
* feat(mcp): compact server instructions in both eras + MCP-only help tool (#1833) MCP-only clients got no workflow guidance: server/discover carried two sentences, legacy initialize carried nothing, and the CLI guides (agent-device --help, help <topic>) were unreachable over MCP. - MCP_SERVER_INSTRUCTIONS: one MCP-phrased workflow card (<2 KB, the Claude Code truncation limit) returned by server/discover and legacy initialize alike. - help tool, router-owned (not a command descriptor): no topic -> the CLI decision card; topic -> agent-device help <topic|command> text, prefixed with the one-line CLI->tool-property mapping; unknown topic -> isError listing the topics. listCommandTools() stays descriptor-only for the AI SDK; the router composes descriptors + help. - Move src/cli/parser/cli-help{,-overview}.ts to src/cli-schema/ so src/mcp (rank 3) can import the renderers without a layering back-edge into src/cli (rank 6). * fix(mcp): name terminal-only commands in help guides; colocate cli-help tests with their sources - The MCP guide preamble claimed every `agent-device <command>` line is a tool of that name; `help web` tells the reader to run `web setup` / `web doctor` and no `web` tool exists. The preamble now lists the exact CLI-only set (listCliCommandNames minus listMcpExposedCommandNames) — derived, not scanned out of prose where `device`/`web` are ordinary words. Regression: help web names `web` as terminal-only, and the listed set equals the registry difference. - cli-help-*.test.ts move from src/cli/parser/__tests__ to src/cli-schema/ to mirror the moved sources. * perf(mcp): tighten the guide card, tool description, and preamble Instructions card 1572 -> 1378 bytes (paid every session), tool description and preamble trimmed, HELP_TOOL built once as a const. Bundle delta vs main 3189 -> 2715 bytes; the remainder is the guide text itself, which the bundle carried in no MCP-phrased form before. |
||
|
|
4bba404249 |
fix(layering): keep the snapshot interactor seam out of the type cycle (#1838)
#1779 added src/daemon/handlers/snapshot-interactor-capture.ts as a
vi.mock seam between snapshot-capture and core/interactors. Both of its
edges are value imports, and it sits on the path
request-generic-dispatch -> snapshot-capture -> (seam) -> core/interactors
-> register-builtins -> command-catalog -> ... -> daemon-command-registry,
so it joined the largest type-level SCC (46 -> 47 files, daemon-server
16 -> 17) and R9/R10 have failed on main since
|
||
|
|
9a0d6dead2 |
refactor: split the test-suite command out of the replay handler (#1826)
* refactor: split the test-suite command out of the replay handler handleSessionReplayCommands becomes the routing decision alone; the test suite's harness-flag admission, request translation and scheduler run move to session-test-suite-command.ts beside the replay runtime they already sit next to. Pure move: no behavior change. * test: pin the replay handler's routing decisions session-replay.ts is a router now, so it gets a focused test of its own (AGENTS.md 1:1 source/test topology): replay reaches the script-source runtime, test reaches the suite command with the whole parameter set, and an unrelated command is declined. Both destinations are mocked so a wrong edge shows up as the wrong marker; each case was proven red by mutating the routing it pins. |
||
|
|
d76e0f94e9 |
refactor: migrate snapshot to device runtime (#1779)
* refactor: migrate snapshot to device runtime * refactor: complete snapshot runtime policy cutover * test: enforce snapshot owner-facts admission * refactor: consolidate desktop snapshot capture * fix: scroll to visible iOS smoke targets * fix: close snapshot cutover alias bypasses * fix: constrain snapshot admission identity flow * fix: enforce snapshot admission through owner facts * fix: adapt replay source tests to snapshot runtime |
||
|
|
a853734f0c |
fix(webdriver): give cloud session creation its own budget and stop leaking billed sessions (#1782)
* fix(webdriver): give cloud session creation its own budget and stop leaking billed sessions Cloud lease allocation ran under the generic 30s/1-retry request policy, so BrowserStack iOS real-device session creation (45-90s) aborted client-side at ~60s on most runs. Each timed-out POST /session still completed server-side and, being non-idempotent, was retried — leaving two billed provider sessions per failed open with no id to release them. - POST /session is its own phase: a 180s create budget (default), zero retries, and no request-bound abort, so the daemon always learns the session id. - lease_allocate carries a 300s allocation budget surfaced to providers as LeaseLifecycleContext.deadline, and a matching 330s client envelope that preserves the daemon on timeout (a reset would SIGKILL mid-create and orphan every billed session the daemon held). - The request's cancellation signal is ownership evidence: a session that completes after the requester left is released, not registered; a create that the transport gives up on surfaces typed evidence (provider + lease) so an operator can find and stop the maybe-orphaned session. Closes #1774 * refactor: one canceled-request error, and tighten the #1774 shapes Review pass over the session-create fix: - The canceled-request error had nine hand-rolled copies (src/request/cancel, maestro shared, exec, retry, install-source x2, and the new provider one). It now has one definition in @agent-device/kernel/errors: createRequestCanceledError(details?, cause?) + isRequestCanceledError + REQUEST_CANCELED_REASON. Callers add evidence or a sharper hint; the reason itself is not overridable, so nothing can build one the predicate misses. - lease_allocate's timeout bundle moves beside INSTALL_TIMEOUT_POLICY in the registry (same {...DEFAULT, envelopeMs, onTimeout} shape); the request timeout constant stays exported from timeout-policy like its siblings. - Transport: fetch helper returns Response's own ok/status; the timeout reason const is private behind isWebDriverRequestTimeout. - Client: one-use options type inlined; the two deadline helpers share one floor. - Session-manager tests: shared makeRuntime/jsonResponse/afterEach restore. Net -29 lines with the feature in. * chore: keep the canceled-request reason private to the kernel * fix: typed cancellation everywhere + own the AWS remote-access ARN through startup Second-order follow-ups the #1774 refactor made cheap: - markRequestCanceled aborts the request signal WITH the kernel's typed canceled error as its reason. Every signal.throwIfAborted(), aborted fetch, and 'throw signal.reason' in the daemon (20+ sites) now surfaces a canceled request as such instead of a bare DOMException that normalized to UNKNOWN — and no site has to know the factory exists. - AWS Device Farm prepareSession owns the remote-access ARN from the moment create-remote-access-session answers: a startup timeout, the allocation deadline, or a canceled request now stops it before the failure surfaces (previously a timed-out startup left a RUNNING billed session behind — the same leak class as the WebDriver session, one phase earlier). The startup wait is capped by LeaseLifecycleContext.deadline and wakes on cancellation. - BrowserStack's pre-session local app upload honors the request signal (an upload is not billed, so plain abort is right there). - lease_heartbeat/lease_release share lease_allocate's preserve-daemon policy: the rationale — the daemon owns billed sessions; a reset orphans them all — applies verbatim. Each AWS ownership test proven red without the guard (3/3). * refactor: dedupe billed-resource cleanup and lease-signal wiring Shrink pass — same behavior, less duplication: - releaseOnFailure(primaryError, release) in webdriver-utils replaces the two identical 'best-effort stop the billed resource, attach cleanupError to the primary AppError' helpers (WebDriver session + AWS remote-access ARN); shared errorMessage too. - The lease handler pulls the request signal from getRequestSignal(requestId) like every sibling handler, instead of threading a requestSignal arg through LeaseHandlerArgs and the request-handler chain. Drops the field, the wiring, and five mechanical test edits; the handler test now proves the request-bound signal (abort it, watch the provider's signal flip) rather than arg identity. - Inlined the one-use requestHeaders back into fetchWebDriver. Handler-signal test proven red without the wiring. * fix(lease): the daemon releases a lease allocated for a gone requester; honest release evidence Review follow-up. The provider was doing the daemon's job: it treated the request signal as 'ownership evidence, not an interrupt' and needed three paragraphs to say so. The daemon owns the request, so it now decides — generically, for every provider — what happens to a lease that finished allocating after its requester left: release it (provider + registry) and answer with the canceled error. - lease.ts: after allocate returns, isRequestCanceled(requestId) → releaseAllocationForGoneRequester(). Release evidence is claimed ONLY on a clean release (no warnings, no throw); a WEBDRIVER_SESSION_DELETE_FAILED release is reported released:false with providerSessionId + a stop-by-hand hint (thymikee's finding: the previous evidence was success-shaped even when DELETE failed). - WebDriverSessionManager: the createOwnedSession/releaseCanceledSession trio is gone; allocate is plain 'create with a budget; on failure clean up' again. - LeaseLifecycleContext.signal is just cancellation, like everywhere else; the ownership-semantics comments on the contract, client, registry, AWS prepare and utils shrink to what the code no longer says itself. - Tests: the two provider-level cancellation tests move to the daemon handler (where the logic now lives), plus the failing-DELETE regression; both proven red without the post-allocate check. * fix(aws): the allocation deadline bounds remote-access startup, not the 120s default Live iOS real-device run: startup needed ~128s and hit the standalone 120s default while the daemon's 300s allocation budget still had room — the new ownership guard correctly stopped the ARN, but the open failed for no reason. When the daemon supplies a deadline it is the bound; the default only applies standalone. Rerun: open in 112s, snapshot, clean close, session STOPPING. * test(aws): pin that the allocation deadline outlives the 120s startup default; drop empty import Review follow-ups on 7f9d1481a: a virtual-clock test (Date.now advanced 10s per poll, RUNNING at 150s, deadline 300s) that fails on the old min(default, deadline) logic and passes now; and the empty 'import {} from kernel/errors' left in maestro/shared.ts is removed. * refactor: finish the dedupe — one release path, kernel errorMessage, AWS on releaseOnFailure Code-quality review at 7f9d1481a: 1. aws-device-farm.ts still carried its own copy of releaseOnFailure (the dedupe commit's script aborted before reaching it and I mis-verified). Now uses the shared helper; private copy deleted. 2. Empty 'import {} from kernel/errors' in maestro/shared.ts removed (273870099). 3. errorMessage() lives in @agent-device/kernel/errors; the two copies this PR had added (lease.ts, webdriver-utils.ts) import it. Sweeping the pre-existing copies is a follow-up. 4. lease.ts has ONE release path: releaseLease(registry, provider, lease, request, ctx) → { released (registry), provider } used by both the lease_release case (wire shape unchanged) and the gone-requester branch, which folds a throwing provider release into releaseError. 'released' now means the same thing in both; the provider verdict is a separate 'providerReleased' (warnings-free, no throw) that drives the stop-by-hand hint. -~35 lines. 5. sessionCreateTimeoutMs is Omit-ed at the WebDriverTransportOptions boundary instead of Pick-ed back out internally. * fix(lease): 'released' on a canceled allocation means the billed session is confirmed gone Re-review at 3665ea06: unifying the release path had made the cancellation error report released:true from the daemon's registry record while the provider DELETE had failed — success-shaped again, with the operator verdict demoted to a second key. Fixed at the source of the ambiguity: - LeaseReleaseOutcome names its bookkeeping field registryReleased. - On the canceled error, 'released' is true only when registryReleased AND the provider released without warnings AND without throwing; the registry record is exposed as 'registryReleased'. The stop-by-hand hint keys on 'released'. - lease_release keeps its existing wire field ('released' = registry; provider cleanup rides in 'provider'), unchanged. - Regressions: failed DELETE and throwing release both pin released:false / registryReleased:true (+ providerSessionId, warnings|releaseError, hint); both proven red on registry-only semantics. * ci: retrigger default-setup CodeQL Run 32051017472 is wedged on GitHub's side: status=completed with Analyze (python) still queued and Analyze (java-kotlin) failed only at SARIF upload (503, 'No server is currently available'). It can be neither cancelled nor rerun, and default-setup CodeQL has no dispatchable workflow, so a new push is the only way to get a fresh run. No source change. * test(webdriver): assert the typed timeout contract on the shared-budget probe main's #1790 tightened this test to expect the raw TimeoutError DOMException, which this PR intentionally normalizes into AppError{reason: webdriver_request_timeout}. On the merge ref the two met and Coverage went red. The regression now asserts the structured contract and that the second request's budget is the shared remainder (~118ms of 200 after an 80ms first call). |
||
|
|
d0d5c8594c | fix: serve remote daemon request diagnostics to the caller (#1801) (#1814) | ||
|
|
4b44c1c53a |
chore(test): remove the contention retry and shrink the subprocess-stub project (#1781 A4) (#1827)
The enumerated single-retry policy (#1419) has fired zero times since it landed on 2026-07-29: 0 of 234 sampled Coverage-job lane envelopes (2026-08-11 to 2026-08-18) have retryCount > 0, and none of 17 recent failed runs was retried (5 refused "outside the enumerated retry list", 4 refused "unhandled error"). All three trackers its entries pointed at (#1098, #1414, #1419) are closed. It cost ~1,454 LOC, a per-run secret marker threaded through a setup file on every Vitest project, and a standing obligation for every future gate reporter to call the blocker bus. Delete the scripts, tests and fixtures, the check:contention-retry script and gate, the envelope artifact upload, and the runner-timeout setup file; test:coverage:ci is a plain `vitest run --coverage` again. lane-envelope.ts stays: the mutation, fuzz and concurrency-torture lanes build their envelopes from it. run-blocker-bus.ts goes: its only consumer was the retry's failure sink, and its only publisher already fails the run by setting process.exitCode. Keep the subprocess-stub project for the three files that really spawn (client-metro, fuzz harness, fuzz corpus-replay) and drop the three that run in 31/212/277ms in CI, which cannot contend for anything. The list is now a plain array in vitest.config.ts with the reason at each entry. Membership and the project's kill criterion live in #1823. Because test:coverage:ci is a bare vitest run, the gate manifest reads its projects directly, so OPAQUE_RUNNERS no longer needs it and an unrun Vitest project becomes unrepresentable rather than detected; the audit test now constructs that state by project-scoping the script. |
||
|
|
60f6356b04 |
fix: read replay scripts on the caller and ship them with the request (#1810)
* fix: read replay scripts on the caller and ship them with the request Closes #1802 * test: assert the caller-side replay path as a substring, not a hand-escaped regex * perf(cli): load the Maestro engine only when a replay entry is a flow The command registry evaluates every command family on CLI startup, so the replay script-source builder's static @agent-device/maestro import put the YAML parser on the --help path. It now loads on demand behind the format check, and the startup import-closure guard covers the engine the way it already covers node:http. * refactor: share the replay request field vocabulary across the CLI and client views The new replay script-source flags appear in both CliFlags and CommandExecutionOptions, which fallow flagged as a clone; ReplayRequestFields declares them once. The test-suite handler's missing-sources rejection now travels the typed-error path its sibling rejections already use, so the fix adds no branch to handleSessionReplayCommands. |
||
|
|
3908559fe2 |
fix: report real claim results from daemon stop (#1818)
* fix: report real claim results from daemon stop `daemon stop` typed `claimsReleased`/`claimsOrphaned` as the literal `[]` and every path hardcoded them, so a graceful stop that released a device claim still reported none (#1799 observation 3, #1320 acceptance). Graceful teardown now records each session's claim outcome — released after a clean teardown, orphaned when teardown left the claim in place — into the daemon shutdown report, and the CLI merges them alongside provider releases. Forced and not-running stops stay empty because they cannot know, and a report written before claim reporting still reads its provider releases. * fix: classify daemon stop claim results from the clear outcome `clearDeviceClaim` deliberately resolves without deleting when the on-disk claim is no longer the one it acquired, so the shutdown ledger's "the call resolved" test reported a successor's claim as released — a device the daemon never freed, counted as freed. `clearDeviceClaim` now returns a typed outcome (`deleted` | `absent` | `ownership-changed`) instead of nothing, and the ledger classifies from it: released only when absence is confirmed, and a new `superseded` bucket for a claim another owner had already taken over. Superseded is neither released (this daemon freed nothing) nor orphaned (no claim of ours remains to reconcile), so folding it into either would break that list's meaning; it also raises a warning so a device now owned elsewhere cannot pass silently. |
||
|
|
8d0de32ba4 |
fix(android): let a covering sibling hide only what its content covers (#1808)
* fix(android): let a sibling hide only what its content covers (#1806) pruneAndroidCoveredSubtrees credited a higher drawing-order sibling with painting its whole box as soon as it had any content anywhere inside it, or a label of its own. A full-screen DoraemonKit drag surface holding one 189px floating icon therefore condemned the entire app subtree, and an empty labelled match_parent placeholder did the same. Occlusion is now spatial. A subtree's footprint is the bounding box of what it presents (agent targets and labelled leaves); a sibling is covered when its footprint lies under a candidate's footprint. A node's own label is no longer paint evidence: a container's content-desc describes its children and an empty labelled View draws nothing. Only a touch target still hides its full box (scrims). Comparing footprint to footprint keeps stacked screens with matching margins registering as covered. Live on a Pixel 9 Pro XL API 37 emulator with a DoKit-shaped overlay added to the test app: snapshot -i went from 2 nodes + the sparse hint to the full app; helper-XML A/B across home/catalog/form/product-detail recovered every label with none lost, and non-overlay screens are byte-identical. * fix(android): measure occlusion by overlapped area, not bounding box Review on #1808: a bounding box of two corner controls spans the viewport, so a transparent overlay with a control in each corner still acquired a full-screen footprint and could prune the app beneath it. Footprints now keep their presented rects apart, and coverage is the overlapped area of the two unions (coordinate-compressed cell sweep). Scrollables count as presenting their box: they consume touches over it, which is what lets a real pushed screen (header, scrollable body, footer) still cover a drawer surface. Adds the disconnected-corner regression. * fix(android): count what a covered sibling shows, not only what it paints Fuzzing random sibling trees old-vs-new surfaced the one direction the footprint model could still regress: a container whose only painted content is small (one corner icon) but which also carries labelled containers or testID-only markers was condemned as soon as a touch surface covered that icon, since markers and container labels are not paint and never entered the footprint. Footprints now carry two rect sets. `paints` (touch targets, scrollables, labelled leaves) is what a candidate can cover with; it still excludes identifiers and container labels, or the DoKit fix would unwind. `shows` adds every labelled or identified node and is what a covered sibling must lose in full. Focusable-only nodes no longer paint their box either, matching #1733 for descendants as well as siblings. Adds the marker regression. Re-fuzzed 20k trees: new-prunes-more is down to 0.14 %, all of the class where everything the target shows lies under a higher touch/scroll surface. Live captures unchanged. * test(android): pin that focusability never paints a covering candidate's box A full-screen focusable wrapper holding one clickable icon is a covering candidate; the lower app content must survive. Fails when paintsOwnBox counts focus targets again. |
||
|
|
2b6d04a13e |
fix: enforce device claims for sessionless device mutations (#1809)
`boot` and `shutdown` never consulted the host-global device claim store, so a daemon in one state directory could terminate an emulator another daemon held a verified-live claim on and report success (#1799). Rather than adding a claim check to those two handlers, this makes the class unrepresentable: `CommandDescriptor` gains a REQUIRED `deviceClaimPolicy` trait (#1320's vocabulary), and the request-execution scope enforces it where the request runtime bindings create a device binding — the one seam through which any handler can obtain device operations, and already the place per-device deduplication lives. A `transient-exclusive` command acquires a command-scoped claim before operations reach the handler, refuses a foreign live claim with the existing DEVICE_IN_USE/DEVICE_CLAIM_LIVE_OWNER error, and releases in the scope's finally. Every other policy performs no claim-store I/O, so session-bound commands keep #1320's non-goal intact. |
||
|
|
4c5f693a03 | fix: declare frameworkTier on the hover descriptor (main parity gate red) (#1817) | ||
|
|
0f4f322878 |
fix: reject selector-shaped wait arguments instead of reading them as text (#1813)
`wait <condition> '<selector>' [timeoutMs]` (e.g. `wait open 'label="Open"' 25000`, `wait exists 'label="x"' 100`) and any unrecognized `key=value` token used to fall through parseWaitPositionals' text fallback and wait out the full timeout for literal text that could never appear on screen — reading as a false "element absent" instead of the caller's own argument mistake (#1035 is the sibling fix for click/press/fill/get). parseWaitPositionals now returns a typed `invalid` variant whenever a positional token is selector-shaped (a recognized key, or an unrecognized key=value) but the list doesn't form a valid selector expression, or a valid selector prefix is followed by unquoted trailing tokens. The message names the offending token, points condition words (exists/present/appears/gone/disappears) at the selector form, and always offers the explicit `wait text '<text>'` escape hatch. Bare text (single- and multi-word) and the explicit `text` keyword form are unaffected. Excluding `invalid` from the type consumed by selector-runtime's toWaitTarget makes the remaining kind-by-kind narrowing exhaustive without a runtime fallback branch. |
||
|
|
801734d433 |
feat(ai-sdk): add agent-device/ai-sdk tool set and document the MCP zero-code path (#1804)
* feat(ai-sdk): add agent-device/ai-sdk tool set and document the MCP zero-code path
Adds `createAgentDeviceTools()` under a new `agent-device/ai-sdk` subpath,
built from the same command registry the MCP server uses so both stay in
lockstep without a hand-maintained tool list. Introduces a `frameworkTier`
descriptor facet ('core' | 'extended') so the factory can default to a
curated perceive/act loop instead of handing a model dozens of tools.
`ai` is wired as an optional peer dependency, imported lazily inside the
factory rather than at module scope, so importing the subpath itself never
requires `ai` to be installed - only calling it does. The package's own
publishing gate (scripts/lib/shipped-imports.ts) is extended to recognize
peerDependencies as a valid resolution source, since this is the first
optional peer this package has shipped.
Also restructures the AI SDK doc around three tiers (zero-code via
@ai-sdk/mcp, the new typed tool set, hand-written tools) and fixes a stale
`needsApproval` reference in favor of the current `toolApproval` API.
* fix(layering): classify src/ai-sdk as a rank-4 zone
The layering guard requires every src/<folder>/ to be explicitly ranked or
unranked; the new src/ai-sdk/ subpath (added in the prior commit) was left
unclassified, failing CI's Layering Guard job. It sits at the same tier as
client/compat/daemon-server/metro/remote/sdk - a public integration surface
consuming mcp (3) and core (2), imported by nothing else in the tree.
* fix(ci): cover, exempt, and pack the new ai-sdk subpath
Fixes the remaining CI failures on the ai-sdk subpath commit:
- Coverage: src/ai-sdk/index.ts had no dedicated unit test (only manual/
integration verification), so changed-line coverage sat at 6.9% against
the 70% gate. Adds src/ai-sdk/__tests__/index.test.ts (core vs 'all' tool
filtering, session/platform pinning and schema hiding, error
normalization, toolApproval passthrough) with createCommandToolExecutor
and createAgentDeviceClient mocked the same way command-tools.test.ts
does, plus a dedicated missing-peer-dependency.test.ts that mocks `ai`
itself to throw, isolated to its own file so it doesn't affect the other
tests' use of the real, installed `ai` package. Changed-line coverage is
now 29/29 (100%).
- Fallow Code Quality: src/ai-sdk/index.ts and examples/sdk/ai-sdk-tools.ts
are entry points with no in-repo importer (reached only via package.json
exports / run directly), and the new subpath's exports are unused
internally by design - both need the same treatment src/sdk/*.ts and its
examples already have in .fallowrc.json.
- Integration Tests: test/integration/installed-package-metro.test.ts and
src/__tests__/package-exports.test.ts each hand-list every published
subpath and smoke-check it from a real packed install; added ./ai-sdk to
both so the new subpath is actually exercised, not just silently passing.
* fix(ai-sdk): hide MCP transport/config fields from the model too
createAgentDeviceTools() only removed session and mcpOutputFormat from tool
schemas. stateDir was still model-visible and reached the shared executor
as client configuration, letting a tool call redirect into a different
daemon state directory - defeating the "one pinned session" guarantee the
factory exists to provide. includeCost and responseLevel are MCP
tool-config knobs in the same category, irrelevant to this adapter.
Widens the hidden-field set to session/stateDir/mcpOutputFormat/
includeCost/responseLevel, and now strips them from the runtime input
inside execute() too (not just the schema), so the guarantee holds even if
a caller bypasses schema validation. The schema-properties filter and the
input filter now share one omitHidden() helper instead of two near-
duplicate implementations.
Addresses the P1 review comment on #1804.
|
||
|
|
20458cbcf6 |
fix(daemon): reject '.', '..' and empty session names at resolveSessionDir (#1815)
safeSessionName only rewrites characters outside [a-zA-Z0-9._-], so the names '.' and '..' survive unchanged and path.join resolves them to the sessions dir itself or its parent, the daemon state dir. A remote caller's --session .. would then land app.log / runner.log / requests/*.ndjson outside the sessions tree. SessionStore.resolveSessionDir is the one place a session name becomes a directory (AGENTS.md: session artifact paths come from session-store), so it now refuses such a name with INVALID_ARGS. Every request goes through it first thing in createRequestExecutionScope, before any artifact path is used, so this is also the admission-time rejection; every other caller passes an already admitted name. isSafeSessionSegment mirrors the predicate PR #1814 adds for its request-diagnostics route; whichever lands second takes the trivial merge. Regression tests were proven red against the pre-fix code: resolveSessionDir returned the sessions dir / state dir for '.', '..', '' and the request scope resolved runnerLogPath to <stateDir>/runner.log. |
||
|
|
8db36299e4 |
feat(web): add hover command for hover-gated UI (#1783) (#1786)
* feat(web): add hover command for hover-gated UI (#1783) Add a first-class `hover <x y|@ref|selector> [--settle]` verb, admitted on web only, that moves the pointer without pressing via the agent-browser backend (mouse move). It rides the existing targeted-touch pipeline (ref/selector/coordinate resolution, occlusion/off-screen guards, settle observation, response builder, recording) through a new optional Interactor/backend `hover` op that only the web provider implements. Touch platforms have no hover state: capabilities advertise it on web only and iOS/Android/Linux reject it at admission with a --platform web hint; longpress stays the mobile hold-gesture verb. Closes #1783 * fix(hover): native hoverRef route for web @ref, android coverage pin, revert skill edit Review follow-ups on #1786: - hover @ref on web now dispatches through the provider's own element handle (agent-browser `hover <ref>`) via a new backend hoverTarget, mirroring click/fill's ADR 0011 native-ref path — web ref frames carry no rects, so the coordinate route could never resolve them. The shared preflight + exact-ref dispatch is extracted into dispatchNativeRefInteraction and used by tap/fill/hover; the guarantee matrix native-ref row now lists hover. - Daemon regression test is production-faithful: rect-less web ref frame, scoped provider, asserts no coordinate dispatch. Selector→coordinate and provider hoverRef tests added. - Android emulator coverage summary pin 2/53 → 3/54. - skills/agent-device/SKILL.md reverted (out of scope, AGENTS.md rule). - Docs/help disclose that --settle with @ref on web shares click's existing limitation; use a selector or coordinates for the settled diff. * test: drive hover through the apple output guard; cover direct hover dispatch The provider-integration apple-leak guard partitions every public command into driven/skipped; hover was neither, which failed Integration Tests and took Coverage down with it. Drive it (it reaches the Apple capability refusal, which is scanned like any other error response). Also cover the direct-dispatch handleHoverCommand seam. * test(web): drive hover @ref in the provider-backed web scenario The integration-progress gate requires every public command to be referenced by a provider-backed scenario. Add hover @ref to the web desktop flow: it must reach the provider's hoverRef handle (never a coordinate) and be recorded on the session without fabricated x/y, like click @ref. |
||
|
|
681cad2222 |
test: assert the specific error code instead of any failure (#1781 B4) (#1790)
Converts the 20 test assertions across the repo that accepted ANY
failure (bare `expect(...).toThrow()`, bare `assert.throws(fn)`, bare
`assert.rejects(p)`) into assertions on the specific AppError `code`
each test is actually about, or — where the propagated error is
genuinely opaque (a mocked upstream failure whose identity, not its
shape, is the point) — identity assertions with a comment explaining
why.
Added a synchronous `assertThrowsAppError(fn, {code, message?})`
sibling to the existing `assertRejectsAppError` helper in
src/__tests__/test-utils/app-error.ts, exported via the test-utils
index, for the two src/ sites that needed it.
packages/provider-limrun and packages/provider-webdriver have no
test-utils dir and cannot import from src/, so those sites use
vitest's `expect(...).toThrow(expect.objectContaining({ code }))` or
an inline `assert.rejects(p, matcherFn)` instead.
Sites converted:
- packages/provider-limrun/src/app-log-runtime.test.ts:153-155
(bare `.toThrow()` x3 -> `UNSUPPORTED_OPERATION`)
- src/daemon/__tests__/app-log.test.ts:39 (bare `.toThrow()` ->
message match; plain Error, not AppError, from verified-file's
identity check)
- src/daemon/__tests__/resumable-upload-range.test.ts:13 (bare
`assert.throws(fn)` -> `INVALID_ARGS`); also fixed line 19's
`assert.throws(fn, value)`, a documented Node.js gotcha where a
string second argument is the failure message, not a matcher, so
it was equally bare in effect
- packages/provider-webdriver/src/webdriver-client.test.ts:229 (bare
`assert.rejects(p)` -> asserts the raw AbortSignal.timeout()
rejection's `name`, since the transport re-throws it unwrapped)
- src/daemon/handlers/__tests__/session-device-claims.test.ts:129,
151, 174 (bare `assert.rejects(p)` x3 -> identity assertions; each
test's point is device-claim rollback/retention around an opaque
mocked upstream failure)
- src/platforms/android/__tests__/settings.test.ts:109 (bare
`assert.rejects(p)` -> `UNSUPPORTED_OPERATION`)
- src/platforms/android/__tests__/snapshot.test.ts:1071, 1342 (bare
`assert.rejects(p)` x2 -> `COMMAND_FAILED` + message)
- src/platforms/android/__tests__/touch-helper-session.test.ts:526
(bare `assert.rejects(p)` -> `COMMAND_FAILED`, wrong-protocol
message)
- src/platforms/apple/core/__tests__/runner-command-retry.test.ts:472,
527, 550, 762, 881, 1016 (bare `assert.rejects(p)` x6 ->
`COMMAND_FAILED` with the recovery-path-specific details/message)
- src/platforms/apple/core/__tests__/runner-transport.test.ts:61
(bare `assert.rejects(p)` -> identity assertion; fetchWithTimeout
does not wrap fetch() failures into an AppError)
No repo-wide scanner/lint rule added (explicitly out of scope per
#1781); no test loosened.
|
||
|
|
0d3b7413c5 |
fix: prevent private AX subtree leaks at source (#1807)
* fix: prevent private AX subtree leaks at source * fix: preserve values in settle signals * fix: normalize settle signal semantics |
||
|
|
ed1d44fa62 |
fix: unify private AX scroll visibility (#1798)
* fix: unify iOS scroll snapshot visibility * fix: preserve composed scroll hints |
||
|
|
8b698e8efc |
fix: stabilize private AX settle snapshots (#1784)
* fix: stabilize private AX settle snapshots * fix: harden private AX settling |
||
|
|
856ff3886f |
fix(ios): preserve final-probe xcodebuild diagnostics (#1776)
* fix(ios): preserve final-probe xcodebuild diagnostics * refactor(ios): split runner startup transport |
||
|
|
f378050586 |
feat(snapshot): attach a fallback screenshot to sparse captures (#1764)
A sparse verdict already tells the caller to use a screenshot as visual truth, which made that screenshot the guaranteed next command on every unreadable screen — a second round trip to obey advice we authored. The user-facing `snapshot` dispatch now takes the shot itself and links the path in its warnings. The fallback is deliberately hung off `dispatchSnapshotViaRuntime` and skipped for internal observations: selector resolution, settle, and wait polling reach `captureSnapshot` directly, so a wait polling an unreadable screen cannot turn into a screenshot per poll. A failed shot is swallowed — the verdict's own warning still carries the manual remedy, so the fallback can never fail the snapshot that was asked for. Sparse captures also say when the screen is the app's problem. Only the `sparse-tree` reason code is evidence about the app: every backend reached the screen and it published no semantic content, which is the same emptiness assistive tech gets. `ax-rejected`, `budget`, `no-nodes` and `capture-failed` are limits of this tool and stay unattributed, so readers are not sent to file bugs against code that is not broken. |
||
|
|
13c482ce50 |
fix: preserve compact snapshot action labels (#1773)
* fix: preserve compact snapshot action labels * fix: narrow iOS implementation cell labels |
||
|
|
d8a7d03faf |
refactor: route application lifecycle through runtime facts (#1759)
* refactor: route application lifecycle through runtime facts Moves the canonical `open`, `prepare`, `close` and internal `runtime` descriptors behind package-owned lifecycle bindings admitted from device runtime facts, while daemon request/session policy and public response construction stay put. Based on main, which already carries the boot unit, the parametrized cutover gate and the apps unit. Readiness is package-owned there, so the Apple and Android bindings call ensureAppleReady/ensureAndroidReady rather than a root readiness bag; ensureAppleReady gained an onColdBootStart hook so open keeps warming the runner cache in parallel with a cold boot, and a narrow markBooted port publishes readiness' fresh observation so a flow still makes one simctl listing. Cutover rows take R24-R27, clear of the accepted catalog and the sibling install stack, and cutoverTableDefects rejects a duplicate rule id. Two defects this unit introduced are fixed here rather than shipped: `open <app> <url>` dropped the URL on a first open, and test-IME activation was first fatal on an unobtainable helper and then over-caught. Helper unavailability is a typed non-activation outcome now; fence, lock and post-record failures propagate. The duplication the unit had accumulated is gone: one runtime-admission module instead of five per-command copies, one direct-lifecycle binding factory instead of six hand-rolled packages, one transport-hint predicate, one session finalization path, and no identity-wrapper module. * fix: allocate lifecycle cutover rows after deployment * chore: preserve lifecycle union reconstruction * fix: reconcile lifecycle runtime stack * refactor: tighten lifecycle runtime topology * refactor: remove superseded runtime adapters * fix: preserve stacked runtime cutovers * test: preserve migrated runtime ownership * test: move Android deployment retry ownership * test: extract runtime hint fixtures * fix: preserve lifecycle stack invariants * fix: complete lifecycle runtime cutover * fix: remove lifecycle cutover residue |
||
|
|
66cca1a5b8 |
refactor: route install commands through platform runtime (#1758)
* refactor: route install commands through platform runtime * fix: preserve stacked runtime facts * fix: align deployment facts with shutdown runtime * refactor: simplify capability facts projection * style: format capability facts projection * fix: preserve migrated capability ownership * fix: propagate deployment artifact cancellation * refactor: move Harmony deployment mechanics into package * refactor: move Apple deployment tools into package * refactor: move Android deployment tools into package * refactor: inject deployment temporary storage * refactor: remove superseded deployment helpers * fix: preserve provider deployment transport |
||
|
|
5855dfc2e0 |
refactor: route shutdown through device runtime (#1757)
* refactor: route shutdown through device runtime * fix: cover shutdown cutover review gaps * fix: propagate shutdown cancellation * fix: move shutdown mechanics to platform owners * fix: pass device to shutdown fact fixture * test: cover shutdown facts in session state fixtures * test: simplify Android shutdown assertions * fix: preserve Apple shutdown cancellation |
||
|
|
5b1b1aad2e |
fix: fingerprint the built daemon's real import graph (#1545) (#1771)
* fix: fingerprint the built daemon's real import graph (#1545) The static-import regex in computeDaemonCodeSignature required whitespace around `from` and after `import`/`export`. The built daemon entry is a minified bundle (tsdown/rolldown `minify: true`), whose real import statements have zero whitespace (e.g. `import{o as e}from"../foo.js"`), so the regex never matched any dependency edge and the graph walk silently degraded to fingerprinting only the entry file (`graph:1`) instead of its true ~50-file runtime graph. That defeats the safety net the signature exists for: a daemon started from one build can look identical to a client running a materially different build as long as the entry file's own size/mtime happen to coincide, and the "code-signature mismatch" takeover this check drives was one of the mechanisms observed in #1545's worktree dev-daemon multiplication. * refactor: fingerprint import specifiers by shape, not by import grammar Replaces the two keyword-anchored regexes (one for static import/export/from clauses, one for dynamic import()) with a single pattern that matches any quoted, relative-path-shaped string literal. The prior approach had to track the exact whitespace/keyword shape bundlers emit around a specifier, which is exactly what broke for minified output in the immediately preceding commit — formatted source, minified builds, static imports, re-exports, and dynamic imports all put the same thing in the same place (a relative string literal), so matching on that directly is both simpler and immune to future formatting differences. A same-shaped string that isn't really an import is safe to over-match: it just fails to resolve to a file and gets dropped. Verified the file count is unchanged against this repo's real build (dist: graph:50, src: graph:741, both identical to the keyword-anchored fix). |
||
|
|
b59b4e5c96 |
fix(daemon): make request-timeout cleanup and hints platform-aware (#1751)
* fix(daemon): make request-timeout cleanup and hints platform-aware
Every local request timeout ran the Apple-runner xcodebuild pkill cleanup
and emitted Apple-specific hint wording ("Apple runner work was aborted",
"Timed-out Apple runner xcodebuild processes were terminated") regardless
of the session's actual platform, which was misleading on Android/web/
Harmony timeouts.
The client already carries the request's declared --platform on the exact
DaemonRequest object both timeout call sites hold (req.flags?.platform), so
no new plumbing or src/platforms import is needed. When the platform is
known non-Apple, the Apple pkill cleanup is skipped and the hint drops its
Apple-runner claim; when it's unknown, both stay on their historical
fail-safe behavior.
Part of #1739 (independent cleanups)
* fix(daemon): split timeout cleanup eligibility from hint evidence
Maintainer review on #1751 found a real correctness inversion in the
first pass: gating the Apple xcodebuild pkill cleanup on the request's
declared --platform flag is wrong in both directions.
- req.flags.platform is not authoritative for session-bound execution.
applyStripLockPolicy (request-lock-policy.ts) lets an existing
session's real device platform silently override a conflicting
declared selector under --session-lock strip, so a request declaring
a non-Apple platform can still legitimately execute Apple work.
Skipping cleanup on that declared flag would skip real cleanup —
the dangerous direction.
- The common session-bound request omits --platform entirely, so the
previous fail-safe (undeclared -> treat as Apple) left the hint
wrongly claiming Apple involvement on the motivating case
(Android/web/Harmony session timeouts) too.
Redesign: separate the two decisions.
- Cleanup eligibility is unconditional again for every local timeout,
matching the original pre-fix behavior: the pkill patterns are
Apple-process-name-specific, so sweeping them on a non-Apple host or
session matches nothing and costs a few no-op subprocess spawns,
never a wrong skip. There is no client-visible signal that proves a
session-bound request cannot touch an Apple runner, so eligibility
does not try to prove one.
- The hint may only name Apple-runner involvement on evidence this call
site actually has: an explicitly declared Apple platform selector, or
the cleanup itself having terminated a matching process
(appleCleanupEvidence). Anything else gets platform-neutral wording.
Adds production-seam route tests
(src/daemon/client/__tests__/daemon-client-timeout-route.test.ts) that
spy on the real exec seam and drive actual socket/HTTP timeouts through
sendRequest, covering the rebound-session and unknown-session cases a
pure formatter test cannot catch. Proved red against the prior commit
(
|
||
|
|
39cd4d346a |
refactor: route appstate through platform runtime (#1755)
* refactor: route appstate through platform runtime * test: keep appstate capability fixture below complexity limit * test: cover appstate required readiness fact * fix: align appstate facts with boot readiness * fix: keep appstate use declaration minimal * fix: close appstate parity and ownership gaps * docs: record final appstate size accounting * fix: merge neutral runtime imports * docs: align final appstate size totals * fix: remove stale runtime dependency edges * docs: correct appstate size accounting * refactor: keep runtime-use factory internal * docs: itemize runtime-use relocation * fix: move appstate queries into runtime packages * refactor: retire root foreground query paths * refactor: share Android foreground parser ownership * fix: preserve Android appstate parser precedence * docs: keep appstate evidence in review artifacts * fix: keep appstate runtime loading lazy * fix: fail closed for stale limrun appstate * fix: preserve limrun recovery and abort appstate * fix: narrow limrun exact-owner recovery * fix: allocate appstate cutover rule * fix: reconcile appstate with merged main * style: format harmony runtime test * fix: allocate appstate rule id * fix: allocate appstate layering rule * fix: remove stale app command admissions * fix: close appstate layering regressions * fix: align Harmony capability parity with runtime facts * test: cover limrun recovery-only readiness * fix: keep Limrun recovery binding app-log only * fix: parse Android app state in linear time |
||
|
|
c7565cb1f8 |
refactor(snapshot): clean snapshot ownership (#1754)
* refactor(snapshot): clean snapshot ownership * fix(snapshot): address ownership review feedback |
||
|
|
15133b587f |
fix: validate macOS recording finalization (#1732)
* fix: validate macOS recording finalization * fix: preserve local macOS recording ownership * fix: resolve recording teardown session keys |
||
|
|
f98814ff10 |
fix: preserve recording evidence across fault boundaries (#1734)
* fix(recording): preserve native evidence across fault boundaries * fix(android): retire evidence after active publish rollback |
||
|
|
c9e6df1d6f |
test(cli-help): gate every help example against the CLI schema (#1766)
Help is the agent-facing command contract, and agents copy its example lines verbatim, so an example the parser rejects is a broken command surface. `help react-native` advertised `open --metro-host/--metro-port` from 0.16.x until `open` gained the flags, and the same drift was still live in `help macos`, which showed `snapshot -i --platform macos --surface menubar` even though `--surface` is an `open`-only flag the session then keeps. Add a conformance test that reads every example back out of the rendered help (root usage, every topic, every command) and runs it through the production parser plus the command's CLI reader. Lines carrying placeholder syntax in an unquoted token are synopsis shapes and are skipped; everything else must parse. Fix the macOS menu-bar example and reword the two prose lines that opened with `agent-device ` but were not commands, so the rule stays total. |
||
|
|
b990271580 |
fix(ios): use current devicectl capture-screenshot syntax for physical devices (#1769)
The direct-capture fallback ran `devicectl device screenshot --device <id> <path>`, which current Xcode devicectl rejects with "Unknown option '--device'", forcing every physical-device screenshot into the slower XCTest runner path. Xcode's devicectl now requires `devicectl device capture screenshot --device <id> --destination <path>`. Fixes #1760 |
||
|
|
eabc936a0f |
refactor: route apps through request runtime (#1756)
* refactor: route apps through request runtime * test: remove stale apps adapter mock * fix: clean apps runtime replay artifacts * fix: remove stale runtime test exports * refactor: simplify apps runtime admission * test: exercise runtime use through facade * fix: close apps runtime admission gaps * fix: disambiguate doctor app inventory callback * fix: keep capability fixture below complexity limit * fix: keep HarmonyOS app inventory fail-closed * test: type HarmonyOS app admission fixture * fix: restore HarmonyOS app inventory parity * test: align HarmonyOS readiness fixture * fix: preserve HarmonyOS doctor app parity * test: type HarmonyOS doctor fixture * fix: isolate HarmonyOS doctor policy * fix: align apps cutover with shared rule catalog |
||
|
|
b13c06338e |
refactor: route boot through readiness runtime (#1747)
* refactor: route boot through readiness runtime * fix: separate boot admission from readiness * fix: register boot cutover policy * refactor(runtime): move readiness into platform owners * fix(test): tolerate provider temp cleanup races |
||
|
|
7f5dbd2e50 |
chore: drive unused production exports to zero (#1743)
* chore: drive unused production exports to zero `pnpm check:production-exports` has been failing on main with 21 findings. Each was investigated rather than blanket-suppressed; they split three ways. Genuinely dead, deleted: - `androidDeviceForSerial` (android/adb.ts) had zero references anywhere, tests included. - `streamAndroidLogcatWithAdb` (android/logcat.ts) had no production consumer and only a guard-clause test; its `captureAndroidLogcatWithAdb` sibling is the published SDK surface. Removed with its options type and test. Test-only aliases over live siblings, collapsed: - app-log-resource-store re-exposed four bound store methods; production used only `resolvePath`, tests used the other three. The sibling screen-recording-resource-store exports just the store, so this now matches: one export, all consumers call `appLogResourceStore.x`. - device-claims re-exported `canonicalLocalDeviceKey` for a single test, while production imports it from device-claim-paths directly. Dropped the re-export and pointed the test at the canonical module. Real consumers the analysis cannot see, exempted with the reason: - The nine remaining `src/cli/commands/*Command` handlers are reached only through `dedicatedCliCommandHandlerLoaders`, the dynamic import() table in router.ts. `deviceCommand` already carried an inline suppression for exactly this; replaced it with one config entry naming the table that enumerates all ten, matching the existing daemon route-handler entry. - `resolveVitestMaxWorkers` (vitest.config.ts), `DEVICE_CLAIM_IN_USE_SAMPLE` (bench sample producers) and the capture-kit `createAppLogLiveHandle` facade export joined the existing entries that already record their exact shape. - `**/*.fixtures.ts` is now an ignorePattern: all 15 build doubles for co-located tests, several import `vi`, and none is imported by production source. Pattern-matching them as test infrastructure also keeps this class of finding from recurring. Gate now reports zero. Unit suite, layering, fallow audit, MCP metadata, build, bundle-owner and package checks all pass. * chore: scope the fixture exemption to unused exports Review feedback on #1743: `ignorePatterns` removes a file from every Fallow mode and rule, but the false positive here is only production-unused-exports. Moved *.fixtures.ts to an ignoreExports entry so fixtures stay inside health, dead-code and cycle analysis. Kept `exports: ["*"]` rather than today's three symbols because the property is per-file — no fixture module has a production consumer — so a new fixture symbol should not reopen the finding. check:production-exports still reports zero, and a full `fallow --summary` returns identical totals (2 dead-code / 4 dupes / 123 health) with and without the change, so nothing was newly surfaced or newly hidden. * test(fallow): prove fixture policy scope |
||
|
|
97c87eec4d |
refactor(registry): exhaustive platformExecution discriminator (ADR 0019 §6) (#1740)
* refactor(registry): make the platform-execution discriminator exhaustive
ADR 0019 §6 (amended): every command descriptor declares its platform-execution
mode explicitly. Adds the `none` mode to `CommandPlatformExecution`, removes the
silent `{ kind: 'legacy' }` default at registry entry, and annotates all 76
descriptors so the migration denominator is machine-readable.
Part of #1739 (wave 0)
* fix(registry): react-devtools executes delegated platform behavior
`react-devtools start` on a Limrun Android instance dispatches internal
`runtime port-reverse`, which reaches a provider device runtime, so ADR 0019 §6
`none` is false for it. Reclassify as `legacy` and add the derived coherence gate
that catches delegated platform execution: if a CLI route for command R
dispatches command D, R may declare `none` only when D is `none`.
Part of #1739 (wave 0)
* fix(registry): attribute CLI dispatches by occurrence, not command name
Subtracting attributed command NAMES let a stray dispatch hide behind a routed
one that names the same command, so the gate's totality claim did not hold.
Dispatch sites now carry their source offset and attribution subtracts
occurrences.
Part of #1739 (wave 0)
* fix(registry): unresolvable CLI daemon-send targets fail the gate
An unknown literal or computed command target resolved to undefined and never
entered the scan, so a dispatch could evade attribution by naming a target the
gate could not read. Daemon-send envelopes are now located by their send call and
an unresolvable target is reported instead of skipped.
Part of #1739 (wave 0)
* refactor(cli): own injected daemon dispatches at a typed construction seam
The syntactic scan recognized only a direct sendToDaemon call whose first
argument was an inline object literal, so a variable envelope or a computed
callee was omitted from every result. Rather than teach the scanner more shapes,
the CLI's injected dispatches now flow through one typed construction point whose
route/command pairs are declared, and the gate reads that declaration instead of
recovering it from syntax.
Part of #1739 (wave 0)
* fix: constrain injected daemon transport handoffs
|
||
|
|
0fc0b6917f |
fix(install-source): preserve URL credential validation error (#1762)
* fix(install-source): preserve URL credential validation error * test(install-source): cover URL credential validation error |
||
|
|
52ef4ca1b1 |
fix(batch): preserve typed error recovery signals (#1761)
* fix(batch): preserve typed error recovery signals * test(batch): cover typed error recovery signals * fix(batch): omit unknown typed error signals * test(batch): keep unknown typed signals absent |
||
|
|
74eab2a554 |
refactor: route selector-resolution structural stages into typed policy (#1744)
* refactor: route selector structural stages into typed policy #1649 landed the per-caller ambiguity matrix and deliberately left four structural columns out: occlusion, off-screen, hittable-ancestor promotion, and the poll budget were per-caller pipeline code, so declaring them would have been an unverifiable claim (nothing consumed them; flipping one left the suite green). This adds the missing half as a table with runners. `SELECTOR_PIPELINE_POLICIES` (src/core/selector-pipeline-policy.ts) gives each caller ONE row naming its ambiguity contract plus its four stages, and every stage is reached only through a runner that reads the row: - occlusion -> selectorPipelineCandidates (candidacy) and resolveSelectorPipelineTarget (refusal). Acting rows exclude covered nodes and refuse covered targets; `find` and the diagnosis probe keep them as candidates and refuse at the target; reads and `wait` ignore them. - promotion -> resolveSelectorPipelineTarget. The per-call-site `promoteToHittableAncestor: boolean` is gone: click/press/longpress name `promotedTarget`, fill/focus/scroll/drag endpoints and the native-ref preflight name `resolvedTarget`. `find`'s below-the-root variant is a declared value rather than a second local helper. - off-screen -> throwIfOffscreenInteractionTarget, which now takes the row and returns the node untouched (no iOS rescue round trip) for observation rows. - poll -> selectorPollBudget, which createWaitPolling derives its deadline and inter-poll delay from; the two wait loops carry a budget, every other row carries none and cannot be polled. Behavior is byte-identical. The acting refusal keeps its exact node, label and details in every branch (promotion declines to retarget away from a covered node, so the "both covered" case names the same node it always did), and `find` carries the occlusion verdict to the focus/type seam rather than raising it early, because find click/fill still delegate that refusal to the interaction leaf's own error shape. selector-pipeline-policy.test.ts drives EVERY row through EVERY runner, including the rows whose answer is "skip" — the half that used to be an absence of code, and an absence cannot fail. Each stage was proven red by flipping its cell (occlusion, promotion, off-screen, poll, plus the declare-only-what-is-enforced guard). The ADR 0011 occlusion/nonHittable `via` pointers for the runtime tree paths now name the runner that makes the decision, not the predicate it applies. Closes #1656; prework for #1739 (waves 4-5). * docs: state constraints instead of narrating the refactor Comment pass over #1656: drop the "used to be per-caller code" / "not module constants" / "rather than an omission" narration — a comment should say what a future edit must respect, not what the previous shape was — and compress the find occlusion-verdict and poll-budget notes to the constraint they actually carry. * refactor: make the selector pipeline the only door to the engine Review of #1744: the structural rows were declared but bypassable. Read and wait routes composed `selectorPipelineCandidates(row, nodes)` with the raw `resolveSelectorChainWithPolicy(..., row.resolution)` and never entered the promotion or off-screen stages, so flipping a read row's `promotion` or `offscreen` changed only the policy unit tests — production `get`/`is`/`wait` were unaffected, which is the unverifiable-column failure #1656 exists to remove. Callers could also pair one row's candidate set with another row's ambiguity contract, and `find list` reached the engine directly. The owning interface (src/core/selector-pipeline.ts) now runs every stage a row declares, skips included, and the stage functions are private to it: - `resolveSelectorPipeline` — single-target rows: candidacy, ambiguity, the replay-guard hook, promotion, occlusion, off-screen. - `listSelectorPipelineMatches` — `reject-candidates` rows, returning the candidate set AND the tree the row sees, so ranking and equivalence classification judge the same nodes candidacy produced. - `runNodePipelineStages` — the node stages for a target from a non-chain matcher (`@ref`, find's fuzzy locator) or a narrowed candidate set. A row whose off-screen stage refuses must supply a refusal shape, so flipping an observation row to `refuse` fails on its real route instead of silently observing. `find list` now names a `readList` row (the new `reject-candidates`/no-rect ambiguity row) instead of calling the engine. R17 selector-pipeline-ownership (scripts/layering/) makes the bypass structurally inexpressible: only the owner may import the engine entry points. Proven against a planted import in selector-read.ts, which the repo-wide scan rejects with the entry points that replace it. Flips now fail through REAL command routes, verified one at a time: readUnique.occlusion/offscreen/promotion and wait.occlusion via get attrs / is / wait; readAny.offscreen via is exists and find; readList.occlusion via find list; promotedTarget.promotion via runtime click. The wait route test needed an advancing clock first — with the frozen one a refused wait spun instead of failing, so the flip hung rather than asserting. * refactor: drop find's dead candidate binding The selector branch bound the row's candidate set and never read it: only the acting classification needs that tree, and find's locator branch brings its own matcher. Names what actually governs the locator target — the shared node stages below, not a candidate set it never had. * refactor: reserve the selector engine behind the pipeline owner Review of #1744 (three blockers). **Listing rows no longer claim stages they cannot run.** `find <q> list` resolves to a candidate SET, so promotion, the off-screen guard and a poll budget have nothing to apply to — a listing has no single element to retarget, keep on screen, or wait for. `readList` now declares only the two stages a listing executes (`SelectorListPolicy`: resolution + occlusion), and the narrower shape is load-bearing: `runNodePipelineStages` and `selectorPollBudget` take the full row, so handing them a listing row is a compile error rather than a silently skipped stage. Pinned with `@ts-expect-error` — widening `readList` makes the directives unused and fails the typecheck. **The engine door is a specifier, not a symbol.** R17's regex could not see a namespace import, a re-export, or a deferred `import()`, none of which mention the symbol it matched. The two engine entries moved to `@agent-device/selectors/engine`, and R19 enforces over the resolved import graph, where every one of those forms is the same edge. Proven on the repo-wide scan by planting each form into a shipped route: namespace import, dynamic import, and `export *` laundering all come back red. `resolveImportEdges` drops an edge whose specifier resolves to nothing, so a specifier rule goes quiet — not red — if the subpath is ever retired. The gate now says that out loud instead of scanning clean. **R19, not R17.** #1750 allocates R17/R18. Verified free against origin/main and that PR's diff, then validated by real merges in both directions: the uniqueness gate passes either way and the three ids stay distinct. The gate itself is new (`scripts/layering/rule-ids.ts`): two branches taking one free number do not conflict in git, so nothing caught R17 twice. Matching whole string literals is what separates a declaration from prose that names a rule, and it is what let the gate see #1750's `const RULE = '…'` shape — the first version missed it and would have been vacuous. `main`'s two pre-existing collisions (R11, R13) are listed as known, not pinned by equality, so #1750 lands in either order without breaking this. Also: the root façade now exposes no resolver at all, and its surface test pins both doors. * fix(layering): make each rule-id allowance expire with its collision Review of #1744: `KNOWN_RULE_ID_COLLISIONS` filtered the exact R11/R13 collision strings, so once #1750 renames those rules apart the entries would keep waving those very collisions through if anyone reintroduced them. "Inert" was wrong — a stale allowance fails open, permanently. `ruleIdCollisionFailures` now checks the transition from both sides: a collision nobody allowed fails, AND an allowance whose collision is absent from the scan fails as a stale allowance. The entry therefore has to be deleted in the same change that removes the collision, and the list burns down to empty, which admits nothing. #1750 is still open, so the transitional entries stay for now (option (b)). Verified against a scratch tree carrying that PR's rename: leaving the list untouched reports both entries as stale; deleting them is clean; and reintroducing `R11 names contracts-implementation-authority and package-boundaries` afterwards is rejected. The last of those is also a unit regression, so the post-transition guarantee is pinned rather than argued. |
||
|
|
52402aec4d |
refactor(daemon): split touch interaction orchestration into semantic modules (#1748)
* refactor(daemon): split touch interaction orchestration Closes part of #1691: interaction-touch.ts becomes a router; press, fill, direct-iOS, shared runtime, Android readiness, and response projection each own one module. Behavior is unchanged. * test(daemon): split touch interaction coverage by module Redistributes all 87 discovered cases across the new module topology and re-keys the two touch-family fallow baseline entries to the paths that now hold the same (net one fewer) findings. * docs: point ADR 0014 at the merged Android readiness regression file * refactor(daemon): give targeted-touch admission its own module Keeps interaction-touch-press.ts inside the 300-line budget after the complexity decomposition: admission (surface/capability/button policy, target parsing, @ref staleness and mutation admission) answers its own question. * refactor(daemon): drop the redundant targeted-touch label alias * test(daemon): split touch suites along the new production seams Adds the press-admission suite the production module was missing and splits the four over-budget suites along new production seams (direct-iOS eligibility, Android ref freshness, touch payload). Every suite installs the full device mock set: three Android-session cases regressed to TOOL_MISSING on a runner without adb when the mock set was trimmed per file. |
||
|
|
943913a2b0 |
fix(android): stop empty focusable overlays from hiding app content (#1737)
* fix(android): stop empty focusable overlays from hiding app content (#1733) The covered-subtree pruner let any `hittable` sibling condemn a lower drawing-order sibling it geometrically covers. `hittable` is `clickable ?? focusable`, so a childless, textless, id-less full-screen focusable View qualified as "agent-visible content" and dropped the sibling holding the real UI. Telegram wraps every screen in exactly such a View, which is why `snapshot` returned 1 node on every Telegram screen while the helper and stock `uiautomator dump` both saw the full tree — the loss was in the TS parser, not the Android helper. Key covering candidacy on `clickable` instead: focusability is an accessibility-traversal property and says nothing about painting over a sibling. Also generalise the existing "never condemn a marker leaf" exemption to any childless, non-clickable sibling that carries its own text or identifier, since drawing order plus geometry cannot distinguish a transparent overlay from an opaque one. * fix(android): keep descendant coverage classification on hittable Review feedback on #1733: narrowing hasActionableDescendant from `hittable` to `clickable` was over-reach. The helper emits clickable/focusable only when true, so a focusable-only control parses as clickable=undefined, hittable=true — and D-pad/TV surfaces are built almost entirely from those. A foreground surface containing them would have stopped qualifying as covering, leaving background selector/ref targets retained and actionable behind it. Descendant classification returns to the established `hittable` contract, and the leaf exemption returns to `!hittable` so a genuinely covered affordance stays condemned. The Telegram fix needs only to exclude the empty focusable wrapper itself, which is self-classification in hasOwnAgentVisibleContent — that stays on `clickable`. Replaces the test that asserted the rejected behavior with a helper-shaped regression: a foreground descendant carrying focusable without any clickable attribute must still suppress a covered target. * refactor(android): split clickable and focusable instead of collapsing them `hittable: clickable ?? focusable` collapsed two independent Android facts and made behavior depend on how a producer encodes a false attribute. The snapshot helper omits `clickable`; stock UiAutomator writes `clickable="false"`. For one focusable control those encodings gave opposite answers — verified: identical trees pruned differently and projected opposite `hittable` values. No comment can hold an invariant the representation contradicts, which is what the previous two commits tried to do. The tree now stores `clickable` and `focusable` as separate non-optional booleans, and every decision reads a named predicate: isTouchTarget, isFocusTarget, isAgentTarget, hasSemanticContent, hasDirectOcclusionEvidence, hasDescendantOcclusionEvidence, isPresentationLeaf. `canCoverSibling` consumes one derived classification (hasOcclusionEvidence) rather than choosing between raw attributes, so the self-versus-descendant substitution that caused the last review round is no longer expressible. Removing `hittable` from the tree type made the typechecker find every construction site; the public field is derived once at projection from isAgentTarget. Behavior change: a focusable, non-clickable control encoded by stock UiAutomator now projects hittable=true and participates in occlusion, matching what the helper backend already did for the same control. The Android TV test asserted the old encoding-specific value; it now asserts inclusion plus the unified projection. One encoding-parity regression replaces the patch-specific test added last round. |