mirror of
https://github.com/callstack/agent-device.git
synced 2026-09-14 20:06:34 +08:00
v0.20.9
1423 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
0697915b10 | 0.20.9 v0.20.9 | ||
|
|
8b698e8efc |
fix: stabilize private AX settle snapshots (#1784)
* fix: stabilize private AX settle snapshots * fix: harden private AX settling |
||
|
|
75a4817789 | fix: install dependencies for iOS runner releases (#1785) | ||
|
|
4cfa34bd22 |
docs: reposition README around mobile app automation for AI agents (#1780)
* docs: reposition README around mobile app automation for AI agents Lead the README with the category (mobile app automation, testing, and verification for AI coding agents) and the three surfaces (CLI, MCP server, Node.js API) so humans and search engines can classify it, then keep verification and evidence as the differentiator. - Move per-platform transport/caveat sentences (HarmonyOS HDC/uitest, Vega VVD-only) out of the intro into "How it works" and add an inline guard comment plus an AGENTS.md rule so new platforms only add a name to the intro list. - Add MCP and Node.js quick starts next to the CLI walkthrough. - Add a works-with/proof line, a "What to ask your agent" prompt list, a product-ladder sentence, and two AEO-shaped FAQ entries. - Point the cloud/remote row at the remote proxy and device clouds docs. - Align npm, MCP registry, and docs-site descriptions with the same category phrasing. - Remove em dashes and evaluative filler per the humanizer skill. * docs: tighten README hero, link proof points, move install above capabilities * docs: add mobile-MCP FAQ distinction, sharpen hero support line, drop coming-soon * docs: make the build-on-top audience explicit in README * docs: trim README repetition and signposting * docs: close the session in finally in the README Node.js snippet |
||
|
|
856ff3886f |
fix(ios): preserve final-probe xcodebuild diagnostics (#1776)
* fix(ios): preserve final-probe xcodebuild diagnostics * refactor(ios): split runner startup transport |
||
|
|
f378050586 |
feat(snapshot): attach a fallback screenshot to sparse captures (#1764)
A sparse verdict already tells the caller to use a screenshot as visual truth, which made that screenshot the guaranteed next command on every unreadable screen — a second round trip to obey advice we authored. The user-facing `snapshot` dispatch now takes the shot itself and links the path in its warnings. The fallback is deliberately hung off `dispatchSnapshotViaRuntime` and skipped for internal observations: selector resolution, settle, and wait polling reach `captureSnapshot` directly, so a wait polling an unreadable screen cannot turn into a screenshot per poll. A failed shot is swallowed — the verdict's own warning still carries the manual remedy, so the fallback can never fail the snapshot that was asked for. Sparse captures also say when the screen is the app's problem. Only the `sparse-tree` reason code is evidence about the app: every backend reached the screen and it published no semantic content, which is the same emptiness assistive tech gets. `ax-rejected`, `budget`, `no-nodes` and `capture-failed` are limits of this tool and stay unattributed, so readers are not sent to file bugs against code that is not broken. |
||
|
|
13c482ce50 |
fix: preserve compact snapshot action labels (#1773)
* fix: preserve compact snapshot action labels * fix: narrow iOS implementation cell labels |
||
|
|
d8a7d03faf |
refactor: route application lifecycle through runtime facts (#1759)
* refactor: route application lifecycle through runtime facts Moves the canonical `open`, `prepare`, `close` and internal `runtime` descriptors behind package-owned lifecycle bindings admitted from device runtime facts, while daemon request/session policy and public response construction stay put. Based on main, which already carries the boot unit, the parametrized cutover gate and the apps unit. Readiness is package-owned there, so the Apple and Android bindings call ensureAppleReady/ensureAndroidReady rather than a root readiness bag; ensureAppleReady gained an onColdBootStart hook so open keeps warming the runner cache in parallel with a cold boot, and a narrow markBooted port publishes readiness' fresh observation so a flow still makes one simctl listing. Cutover rows take R24-R27, clear of the accepted catalog and the sibling install stack, and cutoverTableDefects rejects a duplicate rule id. Two defects this unit introduced are fixed here rather than shipped: `open <app> <url>` dropped the URL on a first open, and test-IME activation was first fatal on an unobtainable helper and then over-caught. Helper unavailability is a typed non-activation outcome now; fence, lock and post-record failures propagate. The duplication the unit had accumulated is gone: one runtime-admission module instead of five per-command copies, one direct-lifecycle binding factory instead of six hand-rolled packages, one transport-hint predicate, one session finalization path, and no identity-wrapper module. * fix: allocate lifecycle cutover rows after deployment * chore: preserve lifecycle union reconstruction * fix: reconcile lifecycle runtime stack * refactor: tighten lifecycle runtime topology * refactor: remove superseded runtime adapters * fix: preserve stacked runtime cutovers * test: preserve migrated runtime ownership * test: move Android deployment retry ownership * test: extract runtime hint fixtures * fix: preserve lifecycle stack invariants * fix: complete lifecycle runtime cutover * fix: remove lifecycle cutover residue |
||
|
|
66cca1a5b8 |
refactor: route install commands through platform runtime (#1758)
* refactor: route install commands through platform runtime * fix: preserve stacked runtime facts * fix: align deployment facts with shutdown runtime * refactor: simplify capability facts projection * style: format capability facts projection * fix: preserve migrated capability ownership * fix: propagate deployment artifact cancellation * refactor: move Harmony deployment mechanics into package * refactor: move Apple deployment tools into package * refactor: move Android deployment tools into package * refactor: inject deployment temporary storage * refactor: remove superseded deployment helpers * fix: preserve provider deployment transport |
||
|
|
5855dfc2e0 |
refactor: route shutdown through device runtime (#1757)
* refactor: route shutdown through device runtime * fix: cover shutdown cutover review gaps * fix: propagate shutdown cancellation * fix: move shutdown mechanics to platform owners * fix: pass device to shutdown fact fixture * test: cover shutdown facts in session state fixtures * test: simplify Android shutdown assertions * fix: preserve Apple shutdown cancellation |
||
|
|
5b1b1aad2e |
fix: fingerprint the built daemon's real import graph (#1545) (#1771)
* fix: fingerprint the built daemon's real import graph (#1545) The static-import regex in computeDaemonCodeSignature required whitespace around `from` and after `import`/`export`. The built daemon entry is a minified bundle (tsdown/rolldown `minify: true`), whose real import statements have zero whitespace (e.g. `import{o as e}from"../foo.js"`), so the regex never matched any dependency edge and the graph walk silently degraded to fingerprinting only the entry file (`graph:1`) instead of its true ~50-file runtime graph. That defeats the safety net the signature exists for: a daemon started from one build can look identical to a client running a materially different build as long as the entry file's own size/mtime happen to coincide, and the "code-signature mismatch" takeover this check drives was one of the mechanisms observed in #1545's worktree dev-daemon multiplication. * refactor: fingerprint import specifiers by shape, not by import grammar Replaces the two keyword-anchored regexes (one for static import/export/from clauses, one for dynamic import()) with a single pattern that matches any quoted, relative-path-shaped string literal. The prior approach had to track the exact whitespace/keyword shape bundlers emit around a specifier, which is exactly what broke for minified output in the immediately preceding commit — formatted source, minified builds, static imports, re-exports, and dynamic imports all put the same thing in the same place (a relative string literal), so matching on that directly is both simpler and immune to future formatting differences. A same-shaped string that isn't really an import is safe to over-match: it just fails to resolve to a file and gets dropped. Verified the file count is unchanged against this repo's real build (dist: graph:50, src: graph:741, both identical to the keyword-anchored fix). |
||
|
|
b59b4e5c96 |
fix(daemon): make request-timeout cleanup and hints platform-aware (#1751)
* fix(daemon): make request-timeout cleanup and hints platform-aware
Every local request timeout ran the Apple-runner xcodebuild pkill cleanup
and emitted Apple-specific hint wording ("Apple runner work was aborted",
"Timed-out Apple runner xcodebuild processes were terminated") regardless
of the session's actual platform, which was misleading on Android/web/
Harmony timeouts.
The client already carries the request's declared --platform on the exact
DaemonRequest object both timeout call sites hold (req.flags?.platform), so
no new plumbing or src/platforms import is needed. When the platform is
known non-Apple, the Apple pkill cleanup is skipped and the hint drops its
Apple-runner claim; when it's unknown, both stay on their historical
fail-safe behavior.
Part of #1739 (independent cleanups)
* fix(daemon): split timeout cleanup eligibility from hint evidence
Maintainer review on #1751 found a real correctness inversion in the
first pass: gating the Apple xcodebuild pkill cleanup on the request's
declared --platform flag is wrong in both directions.
- req.flags.platform is not authoritative for session-bound execution.
applyStripLockPolicy (request-lock-policy.ts) lets an existing
session's real device platform silently override a conflicting
declared selector under --session-lock strip, so a request declaring
a non-Apple platform can still legitimately execute Apple work.
Skipping cleanup on that declared flag would skip real cleanup —
the dangerous direction.
- The common session-bound request omits --platform entirely, so the
previous fail-safe (undeclared -> treat as Apple) left the hint
wrongly claiming Apple involvement on the motivating case
(Android/web/Harmony session timeouts) too.
Redesign: separate the two decisions.
- Cleanup eligibility is unconditional again for every local timeout,
matching the original pre-fix behavior: the pkill patterns are
Apple-process-name-specific, so sweeping them on a non-Apple host or
session matches nothing and costs a few no-op subprocess spawns,
never a wrong skip. There is no client-visible signal that proves a
session-bound request cannot touch an Apple runner, so eligibility
does not try to prove one.
- The hint may only name Apple-runner involvement on evidence this call
site actually has: an explicitly declared Apple platform selector, or
the cleanup itself having terminated a matching process
(appleCleanupEvidence). Anything else gets platform-neutral wording.
Adds production-seam route tests
(src/daemon/client/__tests__/daemon-client-timeout-route.test.ts) that
spy on the real exec seam and drive actual socket/HTTP timeouts through
sendRequest, covering the rebound-session and unknown-session cases a
pure formatter test cannot catch. Proved red against the prior commit
(
|
||
|
|
39cd4d346a |
refactor: route appstate through platform runtime (#1755)
* refactor: route appstate through platform runtime * test: keep appstate capability fixture below complexity limit * test: cover appstate required readiness fact * fix: align appstate facts with boot readiness * fix: keep appstate use declaration minimal * fix: close appstate parity and ownership gaps * docs: record final appstate size accounting * fix: merge neutral runtime imports * docs: align final appstate size totals * fix: remove stale runtime dependency edges * docs: correct appstate size accounting * refactor: keep runtime-use factory internal * docs: itemize runtime-use relocation * fix: move appstate queries into runtime packages * refactor: retire root foreground query paths * refactor: share Android foreground parser ownership * fix: preserve Android appstate parser precedence * docs: keep appstate evidence in review artifacts * fix: keep appstate runtime loading lazy * fix: fail closed for stale limrun appstate * fix: preserve limrun recovery and abort appstate * fix: narrow limrun exact-owner recovery * fix: allocate appstate cutover rule * fix: reconcile appstate with merged main * style: format harmony runtime test * fix: allocate appstate rule id * fix: allocate appstate layering rule * fix: remove stale app command admissions * fix: close appstate layering regressions * fix: align Harmony capability parity with runtime facts * test: cover limrun recovery-only readiness * fix: keep Limrun recovery binding app-log only * fix: parse Android app state in linear time |
||
|
|
9c22467832 |
refactor(ci): make gate ownership structural (#1429) (#1753)
* test(ci): prove every registered gate is owned and reachable (#1429) A check that silently stops running looks exactly like a green build. Two suites had already stopped: `check:tmpdir-leaks` (with its model tests) and `test:fixture-cache` are real package scripts that no workflow ran, reachable only through the `check:unit` aggregate CI never invokes. `CHECK_CATALOG` becomes the registry of every check and `pnpm gate <id>` the only way CI runs one, so finding what a lane runs is a scan for `pnpm gate` rather than an attempt to interpret shell. `pnpm check:gate-manifest` then asserts against the real workflows that every registered check is run by some qualifying lane (per unit, not per script name), that every check the real selector activates for a path is run by a lane that path would start (#1420's class), and that every Vitest project and suite script belongs to a check. The wiring that keeps those honest is asserted too: a gate id must name a registered check, an `if:` must be ruled on in GATE_CONDITIONS so `if: false` unowns what it guards, an action declared to run a gate is proven to, and a job whose steps the loader cannot open fails closed. It deliberately does not try to prove CI runs project code only through `pnpm gate`. Whether a shell block executes project code is not decidable from its text, so shell this model does not recognise earns no ownership credit — the failure direction is a check reported unowned, never one waved through. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ * test(ci): update the two suites that assert on rewired workflow text `scripts/mutation/workflow.test.ts` and `test/ci/trusted-fixture-artifact.test.mjs` read the workflow and action files and assert on their command text, so routing those steps through `pnpm gate <id>` moved what they were matching. They are the two suites the manifest cannot help with: it proves a gate is still run, not that a test asserting on how CI spells a command was updated with it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ * fix(ci): credit gates by execution shape, and keep every guard Three ways the manifest could report a gate as owned when it does not run. 1. Crediting was a substring scan over `run:`, which #1429 explicitly rules out — "do not infer reachability from a command name merely appearing in workflow text". `false && pnpm gate x`, a gate inside `if false; then … fi`, one named in a heredoc, and `echo pnpm gate x` all credited it. There is a live instance: conformance-regenerate.yml's "Fail if regeneration changed anything" step names `pnpm gate maestro-regenerate` inside an error message telling a human to run it, and that credited the gate. A gate now counts only as the first command segment of a line, and a body carrying shell structure earns nothing. Reachability inside a script is not decidable, so this does not try: unrecognised shape means no credit and the check reports unowned. `VAR=$(pnpm gate x …)` is read, since the assignment form is unambiguous and the gate runs. 2. Job-level `if:` was not modelled at all, though six live jobs carry one, so a job that cannot run still credited every gate inside it. Two conditions on the mutation lanes are now declared. 3. A caller's `if:` REPLACED the guard on a nested composite-action step (`guard[0] ?? step.condition`), so an outer `always()` erased an inner `if: false`. Steps carry every guard between the lane and the step. Also corrects two source comments that still claimed project code run outside the runner fails the manifest. It does not: such a step earns no credit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ * ci: add the run-gate action that names a gate structurally The seam the ownership proof will read instead of shell. A lane says which gate it runs in `with.gate`, a typed input the manifest reads straight out of the YAML and validates against CHECK_CATALOG. Nothing here is wired yet — the ~60 call sites and the model change follow. Added first so the target of that conversion is reviewable on its own. `args` cannot select which gate runs; it is appended after the id, so the worst a wrong value does is fail the gate it already named. There is no `|| true` and no output capture: the gate's exit code is the step's exit code, so a gate cannot run without being able to fail its lane. Part of #1429. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ * merge: main (#1770) and route its three new steps through the runner #1770 landed the orphan-check fix on main, wiring `check:tmpdir-leaks`, `check:tmpdir-leaks:test` and `test:fixture-cache` into Coverage, Layering Guard and Integration Tests. This branch had wired the same three through `pnpm gate`, so the merge produced two steps per check rather than a conflict — each check ran twice. Kept main's steps, with the placement and reasoning reviewed on #1770, and changed only their `run:` line to the canonical runner. Dropped this branch's duplicates. Net effect on CI is unchanged: the same three checks, in the same three lanes, once each. Gate manifest green after the merge: 47 checks wired across 33 lanes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ * fix(ci): address review — suite detection, freerange, glob, vacuous skip-list Six review findings plus the mutation blocker. [bug] `registered` was shape-only, so a `test:*` script running `node src/bin.ts test <dir>` resolved to a `script:` leaf and was invisible. Four `test:replay:*` scripts were owned only because someone hand-registered them; `test:replay:android` was neither registered nor reported while the nightly ran the same six .ad files by inlining them. A `test:*` script is now a suite by name. `replay-android` is registered, and the nightly runs the script instead of re-listing its files so the two cannot drift. The nightly invokes it inside `reactivecircus/android-emulator-runner`'s `script:` input — shell handed to a third-party action this loader does not read — so the suite executes but cannot be credited. Recorded in UNPROVABLE_OWNERS with that exact reason rather than assumed. The fixed detector also found a second orphan the review did not name: `test:integration:progress`. That one is a reporter whose `--check` sibling is the registered gate, so it is declared in REPORTING_SCRIPTS — a declaration that itself fails when inert. [bug] `freerange` defaulted to localRunnable, so fail-open ran `fr` (a Bun binary) on the pre-push path. Now false. [suggestion] The `--run` skip-list asserted `build:android-snapshot-helper`, a name `android-helpers` no longer uses, so it could not fail. Derived from the catalog instead. [suggestion] `matchesGlob` joined `**` splits with `.*`, making the adjacent slash mandatory — GitHub's `**` matches zero directories, so `src/**/*.test.ts` did not match `src/a.test.ts`. Pinned against `packages/*/src/**/*.test.ts`. [suggestion] Deleted the unwired `run-gate` action. It had no callers, was absent from GATE_ACTIONS, and its comment described a system that had not shipped. It returns with the rewiring, not before. [suggestion] Collapsed the module headers that narrated discarded designs. Mutation: `daemon entrypoint publishes HTTP metadata and cleans up on shutdown` is the only test here that spawns a real daemon process. It takes ~1.1s alone but exceeds Vitest's 5s default inside Stryker's dry run, which aborts the sweep before a single mutant runs. Given 30s. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ * fix(mutation): order sandbox aliases longest-first so subpaths resolve Every shard of the mutation sweep aborted in Stryker's dry run with: Cannot find package '@agent-device/selectors/engine' imported from .tmp/stryker/sandbox-*/src/core/selector-pipeline.ts The alias was generated correctly; it just never won. Vite matches a STRING alias by prefix and takes the first hit, and `workspaceSpecifierTargets` emitted the bare `@agent-device/selectors` ahead of the subpath entries. The bare entry therefore captured `@agent-device/selectors/engine` and rewrote it to `…/src/index.ts/engine`, which does not exist; Node fell back to real package resolution, could not find the subpath inside the sandbox, and the dry run failed before a single mutant ran — so the shard uploaded an empty envelope instead of a report and the ratchet failed for want of one. Sorting longest specifier first makes the most specific alias win: @agent-device/selectors/engine -> packages/selectors/src/engine.ts @agent-device/selectors/ast -> packages/selectors/src/ast.ts @agent-device/selectors -> packages/selectors/src/index.ts `/ast` never tripped this because nothing in a related test set imported it; `selector-pipeline.ts` introduced the first subpath import that mattered (#1744), so the mutation lane has been unable to run since that landed. Any PR touching `scripts/mutation/**` — which fails open into the full sweep — would have hit it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ * refactor: derive gate ownership from workflow structure * fix: run gates without optional arguments * fix: resolve mutation workspace subpaths exactly --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
c7565cb1f8 |
refactor(snapshot): clean snapshot ownership (#1754)
* refactor(snapshot): clean snapshot ownership * fix(snapshot): address ownership review feedback |
||
|
|
07d528086b |
fix: harden iOS alert smoke scenario (#1767)
* fix: harden iOS alert smoke scenario * fix: address iOS alert smoke review feedback * fix: present fixture alert after React commit |
||
|
|
85aef7af75 |
fix(ios): make custom-action reads atomically single-flight (#1749)
* fix(ios): relax custom-action refusal timing assertion * fix(ios): make custom-action admission atomic |
||
|
|
15133b587f |
fix: validate macOS recording finalization (#1732)
* fix: validate macOS recording finalization * fix: preserve local macOS recording ownership * fix: resolve recording teardown session keys |
||
|
|
3eda37b0e0 |
fix(ci): run the three checks no workflow was running (#1770)
`check:tmpdir-leaks`, `check:tmpdir-leaks:test` and `test:fixture-cache` are real package scripts with real assertions that no CI lane executed. All three are reachable only through `check:unit`, an aggregate no workflow invokes, so a regression in any of them could not fail a PR — and a check that silently stops running looks exactly like a green build. Verified by transitive script reachability over every workflow and composite action on main: zero references to any of the three anywhere in .github. Placement: - `check:tmpdir-leaks` joins Coverage, the lane whose instrumented suite is what would leak a run directory. - `check:tmpdir-leaks:test` joins Layering Guard rather than sitting next to the check it covers. vitest-tmpdir-global-setup.test.ts proves the lifecycle by spawning a real nested `vitest run`, and Coverage already loses runs to worker-fork teardown errors; starting a nested Vitest beside the full instrumented suite is a contention risk with nothing to gain. - `test:fixture-cache` joins Integration Tests, beside the fixture-app fallback smoke that covers the same artifact contract. All three pass at this commit, so this wires green gates rather than red ones. `check:tmpdir-leaks` was also shown non-vacuous: planting an abandoned agent-device-test-run-* directory makes it fail and name the directory. Part of #1429, which stays open for the ownership proof that would have caught this class automatically. Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
f98814ff10 |
fix: preserve recording evidence across fault boundaries (#1734)
* fix(recording): preserve native evidence across fault boundaries * fix(android): retire evidence after active publish rollback |
||
|
|
c9e6df1d6f |
test(cli-help): gate every help example against the CLI schema (#1766)
Help is the agent-facing command contract, and agents copy its example lines verbatim, so an example the parser rejects is a broken command surface. `help react-native` advertised `open --metro-host/--metro-port` from 0.16.x until `open` gained the flags, and the same drift was still live in `help macos`, which showed `snapshot -i --platform macos --surface menubar` even though `--surface` is an `open`-only flag the session then keeps. Add a conformance test that reads every example back out of the rendered help (root usage, every topic, every command) and runs it through the production parser plus the command's CLI reader. Lines carrying placeholder syntax in an unquoted token are synopsis shapes and are skipped; everything else must parse. Fix the macOS menu-bar example and reword the two prose lines that opened with `agent-device ` but were not commands, so the rule stays total. |
||
|
|
b990271580 |
fix(ios): use current devicectl capture-screenshot syntax for physical devices (#1769)
The direct-capture fallback ran `devicectl device screenshot --device <id> <path>`, which current Xcode devicectl rejects with "Unknown option '--device'", forcing every physical-device screenshot into the slower XCTest runner path. Xcode's devicectl now requires `devicectl device capture screenshot --device <id> --destination <path>`. Fixes #1760 |
||
|
|
eabc936a0f |
refactor: route apps through request runtime (#1756)
* refactor: route apps through request runtime * test: remove stale apps adapter mock * fix: clean apps runtime replay artifacts * fix: remove stale runtime test exports * refactor: simplify apps runtime admission * test: exercise runtime use through facade * fix: close apps runtime admission gaps * fix: disambiguate doctor app inventory callback * fix: keep capability fixture below complexity limit * fix: keep HarmonyOS app inventory fail-closed * test: type HarmonyOS app admission fixture * fix: restore HarmonyOS app inventory parity * test: align HarmonyOS readiness fixture * fix: preserve HarmonyOS doctor app parity * test: type HarmonyOS doctor fixture * fix: isolate HarmonyOS doctor policy * fix: align apps cutover with shared rule catalog |
||
|
|
b13c06338e |
refactor: route boot through readiness runtime (#1747)
* refactor: route boot through readiness runtime * fix: separate boot admission from readiness * fix: register boot cutover policy * refactor(runtime): move readiness into platform owners * fix(test): tolerate provider temp cleanup races |
||
|
|
ba48146520 |
refactor(layering): one parametrized runtime-command-cutover gate (ADR 0019 §8) (#1745)
* refactor(layering): fold the four cutover policies into one parametrized gate ADR 0019 §8: the per-command cutover gates consolidate into one parametrized runtime-command-cutover gate driven by a table of migrated commands. Adding a migrated command adds a row; the mechanism carries one planted-red proof instead of one per command. Part of #1739 (wave 0) * fix(layering): scope cutover calls to lexical owners * chore: format cutover ownership model |
||
|
|
8f98d23f14 |
refactor(layering): give each colliding rule id its own number (#1750)
* refactor(layering): give each colliding rule id its own number
R11 and R13 each named two unrelated rules. report() groups violations by the
rule string and titles every annotation `Layering drift (${rule})`, so a shared
number made the guard's output ambiguous about which rule fired.
Reference counts decided which rule keeps its number. R11 package-boundaries is
named in ~30 places (CONTEXT.md, ADR 0019, testing.md, the mutation and
affected-check configs, four package source comments, its own tests) against one
for the contracts rule; R13 platform-package-substrate is the RULE in three
policy files plus CONTEXT.md, ADR 0019 and model.ts against two for the devices
cutover. Both keepers stay put and the two newest rules move up:
R11 contracts-implementation-authority -> R18
R13 device-inventory-cutover -> R17
R17/R18 follow the namespace's order-of-addition convention (R14 #1701 < R15
#1702 < R16 #1724): device-inventory-cutover landed in #1699 and
contracts-implementation-authority in #1701. #1656 took R19 for
selector-pipeline-ownership on the same reading.
The rule-map header in check.ts is renumbered and reordered back into numeric
order, and gains the R18 entry the contracts rule never had -- without it a
reader looking up an R18 violation finds nothing where they used to find the
wrong rule. deviceInventoryCutoverSummary() was also the only OK-line summary
not leading with its rule number, which is what made the number unreadable from
the success line in the first place.
Also corrects a normative ADR reference. ADR 0019's platform-package import
rules -- contracts-to-platform, sibling-platform, root/daemon, raw-process --
are R13's, as CONTEXT.md:420 already says. The R11 attribution predates
platform-package-policy (#1697, a day before #1699), when R11 was the only
package rule.
* chore(layering): retire the expired R11/R13 collision allowances
KNOWN_RULE_ID_COLLISIONS was opened for exactly the two collisions the previous
commit renames apart, and ruleIdCollisionFailures expires an allowance on
contact: once the collision is gone the entry fails as stale, because a list
still naming it would wave it back through if anyone reintroduced it.
Both entries are therefore deleted in the change that removes the collisions,
leaving the empty list that admits nothing. The namespace is now one-to-one
across R2-R19.
|
||
|
|
7f5dbd2e50 |
chore: drive unused production exports to zero (#1743)
* chore: drive unused production exports to zero `pnpm check:production-exports` has been failing on main with 21 findings. Each was investigated rather than blanket-suppressed; they split three ways. Genuinely dead, deleted: - `androidDeviceForSerial` (android/adb.ts) had zero references anywhere, tests included. - `streamAndroidLogcatWithAdb` (android/logcat.ts) had no production consumer and only a guard-clause test; its `captureAndroidLogcatWithAdb` sibling is the published SDK surface. Removed with its options type and test. Test-only aliases over live siblings, collapsed: - app-log-resource-store re-exposed four bound store methods; production used only `resolvePath`, tests used the other three. The sibling screen-recording-resource-store exports just the store, so this now matches: one export, all consumers call `appLogResourceStore.x`. - device-claims re-exported `canonicalLocalDeviceKey` for a single test, while production imports it from device-claim-paths directly. Dropped the re-export and pointed the test at the canonical module. Real consumers the analysis cannot see, exempted with the reason: - The nine remaining `src/cli/commands/*Command` handlers are reached only through `dedicatedCliCommandHandlerLoaders`, the dynamic import() table in router.ts. `deviceCommand` already carried an inline suppression for exactly this; replaced it with one config entry naming the table that enumerates all ten, matching the existing daemon route-handler entry. - `resolveVitestMaxWorkers` (vitest.config.ts), `DEVICE_CLAIM_IN_USE_SAMPLE` (bench sample producers) and the capture-kit `createAppLogLiveHandle` facade export joined the existing entries that already record their exact shape. - `**/*.fixtures.ts` is now an ignorePattern: all 15 build doubles for co-located tests, several import `vi`, and none is imported by production source. Pattern-matching them as test infrastructure also keeps this class of finding from recurring. Gate now reports zero. Unit suite, layering, fallow audit, MCP metadata, build, bundle-owner and package checks all pass. * chore: scope the fixture exemption to unused exports Review feedback on #1743: `ignorePatterns` removes a file from every Fallow mode and rule, but the false positive here is only production-unused-exports. Moved *.fixtures.ts to an ignoreExports entry so fixtures stay inside health, dead-code and cycle analysis. Kept `exports: ["*"]` rather than today's three symbols because the property is per-file — no fixture module has a production consumer — so a new fixture symbol should not reopen the finding. check:production-exports still reports zero, and a full `fallow --summary` returns identical totals (2 dead-code / 4 dupes / 123 health) with and without the change, so nothing was newly surfaced or newly hidden. * test(fallow): prove fixture policy scope |
||
|
|
97c87eec4d |
refactor(registry): exhaustive platformExecution discriminator (ADR 0019 §6) (#1740)
* refactor(registry): make the platform-execution discriminator exhaustive
ADR 0019 §6 (amended): every command descriptor declares its platform-execution
mode explicitly. Adds the `none` mode to `CommandPlatformExecution`, removes the
silent `{ kind: 'legacy' }` default at registry entry, and annotates all 76
descriptors so the migration denominator is machine-readable.
Part of #1739 (wave 0)
* fix(registry): react-devtools executes delegated platform behavior
`react-devtools start` on a Limrun Android instance dispatches internal
`runtime port-reverse`, which reaches a provider device runtime, so ADR 0019 §6
`none` is false for it. Reclassify as `legacy` and add the derived coherence gate
that catches delegated platform execution: if a CLI route for command R
dispatches command D, R may declare `none` only when D is `none`.
Part of #1739 (wave 0)
* fix(registry): attribute CLI dispatches by occurrence, not command name
Subtracting attributed command NAMES let a stray dispatch hide behind a routed
one that names the same command, so the gate's totality claim did not hold.
Dispatch sites now carry their source offset and attribution subtracts
occurrences.
Part of #1739 (wave 0)
* fix(registry): unresolvable CLI daemon-send targets fail the gate
An unknown literal or computed command target resolved to undefined and never
entered the scan, so a dispatch could evade attribution by naming a target the
gate could not read. Daemon-send envelopes are now located by their send call and
an unresolvable target is reported instead of skipped.
Part of #1739 (wave 0)
* refactor(cli): own injected daemon dispatches at a typed construction seam
The syntactic scan recognized only a direct sendToDaemon call whose first
argument was an inline object literal, so a variable envelope or a computed
callee was omitted from every result. Rather than teach the scanner more shapes,
the CLI's injected dispatches now flow through one typed construction point whose
route/command pairs are declared, and the gate reads that declaration instead of
recovering it from syntax.
Part of #1739 (wave 0)
* fix: constrain injected daemon transport handoffs
|
||
|
|
0fc0b6917f |
fix(install-source): preserve URL credential validation error (#1762)
* fix(install-source): preserve URL credential validation error * test(install-source): cover URL credential validation error |
||
|
|
52ef4ca1b1 |
fix(batch): preserve typed error recovery signals (#1761)
* fix(batch): preserve typed error recovery signals * test(batch): cover typed error recovery signals * fix(batch): omit unknown typed error signals * test(batch): keep unknown typed signals absent |
||
|
|
74eab2a554 |
refactor: route selector-resolution structural stages into typed policy (#1744)
* refactor: route selector structural stages into typed policy #1649 landed the per-caller ambiguity matrix and deliberately left four structural columns out: occlusion, off-screen, hittable-ancestor promotion, and the poll budget were per-caller pipeline code, so declaring them would have been an unverifiable claim (nothing consumed them; flipping one left the suite green). This adds the missing half as a table with runners. `SELECTOR_PIPELINE_POLICIES` (src/core/selector-pipeline-policy.ts) gives each caller ONE row naming its ambiguity contract plus its four stages, and every stage is reached only through a runner that reads the row: - occlusion -> selectorPipelineCandidates (candidacy) and resolveSelectorPipelineTarget (refusal). Acting rows exclude covered nodes and refuse covered targets; `find` and the diagnosis probe keep them as candidates and refuse at the target; reads and `wait` ignore them. - promotion -> resolveSelectorPipelineTarget. The per-call-site `promoteToHittableAncestor: boolean` is gone: click/press/longpress name `promotedTarget`, fill/focus/scroll/drag endpoints and the native-ref preflight name `resolvedTarget`. `find`'s below-the-root variant is a declared value rather than a second local helper. - off-screen -> throwIfOffscreenInteractionTarget, which now takes the row and returns the node untouched (no iOS rescue round trip) for observation rows. - poll -> selectorPollBudget, which createWaitPolling derives its deadline and inter-poll delay from; the two wait loops carry a budget, every other row carries none and cannot be polled. Behavior is byte-identical. The acting refusal keeps its exact node, label and details in every branch (promotion declines to retarget away from a covered node, so the "both covered" case names the same node it always did), and `find` carries the occlusion verdict to the focus/type seam rather than raising it early, because find click/fill still delegate that refusal to the interaction leaf's own error shape. selector-pipeline-policy.test.ts drives EVERY row through EVERY runner, including the rows whose answer is "skip" — the half that used to be an absence of code, and an absence cannot fail. Each stage was proven red by flipping its cell (occlusion, promotion, off-screen, poll, plus the declare-only-what-is-enforced guard). The ADR 0011 occlusion/nonHittable `via` pointers for the runtime tree paths now name the runner that makes the decision, not the predicate it applies. Closes #1656; prework for #1739 (waves 4-5). * docs: state constraints instead of narrating the refactor Comment pass over #1656: drop the "used to be per-caller code" / "not module constants" / "rather than an omission" narration — a comment should say what a future edit must respect, not what the previous shape was — and compress the find occlusion-verdict and poll-budget notes to the constraint they actually carry. * refactor: make the selector pipeline the only door to the engine Review of #1744: the structural rows were declared but bypassable. Read and wait routes composed `selectorPipelineCandidates(row, nodes)` with the raw `resolveSelectorChainWithPolicy(..., row.resolution)` and never entered the promotion or off-screen stages, so flipping a read row's `promotion` or `offscreen` changed only the policy unit tests — production `get`/`is`/`wait` were unaffected, which is the unverifiable-column failure #1656 exists to remove. Callers could also pair one row's candidate set with another row's ambiguity contract, and `find list` reached the engine directly. The owning interface (src/core/selector-pipeline.ts) now runs every stage a row declares, skips included, and the stage functions are private to it: - `resolveSelectorPipeline` — single-target rows: candidacy, ambiguity, the replay-guard hook, promotion, occlusion, off-screen. - `listSelectorPipelineMatches` — `reject-candidates` rows, returning the candidate set AND the tree the row sees, so ranking and equivalence classification judge the same nodes candidacy produced. - `runNodePipelineStages` — the node stages for a target from a non-chain matcher (`@ref`, find's fuzzy locator) or a narrowed candidate set. A row whose off-screen stage refuses must supply a refusal shape, so flipping an observation row to `refuse` fails on its real route instead of silently observing. `find list` now names a `readList` row (the new `reject-candidates`/no-rect ambiguity row) instead of calling the engine. R17 selector-pipeline-ownership (scripts/layering/) makes the bypass structurally inexpressible: only the owner may import the engine entry points. Proven against a planted import in selector-read.ts, which the repo-wide scan rejects with the entry points that replace it. Flips now fail through REAL command routes, verified one at a time: readUnique.occlusion/offscreen/promotion and wait.occlusion via get attrs / is / wait; readAny.offscreen via is exists and find; readList.occlusion via find list; promotedTarget.promotion via runtime click. The wait route test needed an advancing clock first — with the frozen one a refused wait spun instead of failing, so the flip hung rather than asserting. * refactor: drop find's dead candidate binding The selector branch bound the row's candidate set and never read it: only the acting classification needs that tree, and find's locator branch brings its own matcher. Names what actually governs the locator target — the shared node stages below, not a candidate set it never had. * refactor: reserve the selector engine behind the pipeline owner Review of #1744 (three blockers). **Listing rows no longer claim stages they cannot run.** `find <q> list` resolves to a candidate SET, so promotion, the off-screen guard and a poll budget have nothing to apply to — a listing has no single element to retarget, keep on screen, or wait for. `readList` now declares only the two stages a listing executes (`SelectorListPolicy`: resolution + occlusion), and the narrower shape is load-bearing: `runNodePipelineStages` and `selectorPollBudget` take the full row, so handing them a listing row is a compile error rather than a silently skipped stage. Pinned with `@ts-expect-error` — widening `readList` makes the directives unused and fails the typecheck. **The engine door is a specifier, not a symbol.** R17's regex could not see a namespace import, a re-export, or a deferred `import()`, none of which mention the symbol it matched. The two engine entries moved to `@agent-device/selectors/engine`, and R19 enforces over the resolved import graph, where every one of those forms is the same edge. Proven on the repo-wide scan by planting each form into a shipped route: namespace import, dynamic import, and `export *` laundering all come back red. `resolveImportEdges` drops an edge whose specifier resolves to nothing, so a specifier rule goes quiet — not red — if the subpath is ever retired. The gate now says that out loud instead of scanning clean. **R19, not R17.** #1750 allocates R17/R18. Verified free against origin/main and that PR's diff, then validated by real merges in both directions: the uniqueness gate passes either way and the three ids stay distinct. The gate itself is new (`scripts/layering/rule-ids.ts`): two branches taking one free number do not conflict in git, so nothing caught R17 twice. Matching whole string literals is what separates a declaration from prose that names a rule, and it is what let the gate see #1750's `const RULE = '…'` shape — the first version missed it and would have been vacuous. `main`'s two pre-existing collisions (R11, R13) are listed as known, not pinned by equality, so #1750 lands in either order without breaking this. Also: the root façade now exposes no resolver at all, and its surface test pins both doors. * fix(layering): make each rule-id allowance expire with its collision Review of #1744: `KNOWN_RULE_ID_COLLISIONS` filtered the exact R11/R13 collision strings, so once #1750 renames those rules apart the entries would keep waving those very collisions through if anyone reintroduced them. "Inert" was wrong — a stale allowance fails open, permanently. `ruleIdCollisionFailures` now checks the transition from both sides: a collision nobody allowed fails, AND an allowance whose collision is absent from the scan fails as a stale allowance. The entry therefore has to be deleted in the same change that removes the collision, and the list burns down to empty, which admits nothing. #1750 is still open, so the transitional entries stay for now (option (b)). Verified against a scratch tree carrying that PR's rename: leaving the list untouched reports both entries as stale; deleting them is clean; and reintroducing `R11 names contracts-implementation-authority and package-boundaries` afterwards is rejected. The last of those is also a unit regression, so the post-transition guarantee is pinned rather than argued. |
||
|
|
057ab1c82d |
fix(layering): stop double-reporting contracts-authority violations (#1746)
* fix(layering): stop double-reporting contracts-authority violations main()'s violation list spread checkContractsImplementationAuthority(sources) twice, so every R11 contracts-implementation-authority finding was printed and ::error-annotated twice on a red run — inflating the headline violation count and producing duplicate GitHub annotations on the same file:line. Verified by planting a `setTimeout` call in a contracts production source: the rule reported 2 identical violations before and 1 after, with the extra annotation gone. `pnpm check:layering` stays green (136/136 policy tests). Nothing in the suite covers main()'s assembly of the violation list — the policy tests all call their rule functions directly — so neither a duplicated nor a dropped entry there is currently detectable. * test(layering): hold main() to wiring every rule exactly once The duplicate this branch removed survived because nothing enumerates the guard's rules: main()'s violation list is hand-written, and the per-policy tests call their rule functions directly, never seeing the wiring. A lost spread is the dangerous version of the same gap — the rule stops being enforced and the run still prints OK. Make the file's own bindings the oracle: every in-scope `check*` value, local or imported, must be spread into main()'s violation list exactly once. That covers both directions plus a third case — a policy written and never wired in. Fails closed if main() or the array is renamed, so the instrument cannot pass by finding nothing. Test-only rather than an R17 inside the guard: a self-referential rule is defeated by dropping its own spread, which is exactly the failure it exists to catch. Verified by mutating the real check.ts in both directions (re-planting the duplicate, then dropping checkZeroDepJobs) — each turns the run red, and the restored file is green at 143/143. * test(layering): discover layering suites by glob instead of by hand check:layering named its 14 test files one by one, so adding a policy test meant remembering to register it — and twice nobody did. Both halves of the R16 record cutover shipped with tests that have never run: scripts/layering/record-runtime-mechanics-policy.test.ts (2 tests) scripts/layering/record-runtime-registry-policy.test.ts (1 test) Their policies are live in the guard; only the tests were dormant. All three pass, so nothing had rotted — the coverage was simply never being collected. Glob the directory the way mutation:test already globs its own, which makes the filesystem the enumeration and retires the registration step. 143 -> 146 tests, still green. This is the same defect as the duplicate spread this branch opened with, one level up: a hand-maintained list with nothing checking it against reality. * refactor(layering): register guard rules in a keyed table Replaces the AST wiring guard with a construction that cannot express the defect, per review on #1746. The parser was the wrong instrument: it reconstructed one array's shape from TypeScript syntax, so it only recognised top-level function declarations and imports whose local name matched /^check[A-Z]/. A const-defined or aliased rule was invisible to it, a helper named checkX was a false positive, and naming and syntax became part of the interface — all to detect a mistake rather than prevent it. Rules now live in a keyed table over a shared context, executed once via Object.values. An object cannot hold a key twice, so double registration is unrepresentable rather than merely detected, and oxlint's no-dupe-keys rejects the attempt at the source. LayeringRuleId makes a missing key a type error, and LAYERING_RULE_IDS gives the catalog to check exhaustiveness against. Call sites and order are unchanged, so grouped output and the success line are byte-identical. One regression test remains, through the production interface: scripts/ is outside tsconfig.json's `include`, so the Record's exhaustiveness is an editor signal rather than a CI gate, and the catalog assertion is what fails the build when wiring goes missing. Verified by mutation: dropping an entry and registering an uncatalogued one both fail the test, a duplicated key fails oxlint, and re-planting the original contracts violation reports it exactly once. Net -133 LOC. |
||
|
|
52402aec4d |
refactor(daemon): split touch interaction orchestration into semantic modules (#1748)
* refactor(daemon): split touch interaction orchestration Closes part of #1691: interaction-touch.ts becomes a router; press, fill, direct-iOS, shared runtime, Android readiness, and response projection each own one module. Behavior is unchanged. * test(daemon): split touch interaction coverage by module Redistributes all 87 discovered cases across the new module topology and re-keys the two touch-family fallow baseline entries to the paths that now hold the same (net one fewer) findings. * docs: point ADR 0014 at the merged Android readiness regression file * refactor(daemon): give targeted-touch admission its own module Keeps interaction-touch-press.ts inside the 300-line budget after the complexity decomposition: admission (surface/capability/button policy, target parsing, @ref staleness and mutation admission) answers its own question. * refactor(daemon): drop the redundant targeted-touch label alias * test(daemon): split touch suites along the new production seams Adds the press-admission suite the production module was missing and splits the four over-budget suites along new production seams (direct-iOS eligibility, Android ref freshness, touch payload). Every suite installs the full device mock set: three Android-session cases regressed to TOOL_MISSING on a runner without adb when the mock set was trimmed per file. |
||
|
|
943913a2b0 |
fix(android): stop empty focusable overlays from hiding app content (#1737)
* fix(android): stop empty focusable overlays from hiding app content (#1733) The covered-subtree pruner let any `hittable` sibling condemn a lower drawing-order sibling it geometrically covers. `hittable` is `clickable ?? focusable`, so a childless, textless, id-less full-screen focusable View qualified as "agent-visible content" and dropped the sibling holding the real UI. Telegram wraps every screen in exactly such a View, which is why `snapshot` returned 1 node on every Telegram screen while the helper and stock `uiautomator dump` both saw the full tree — the loss was in the TS parser, not the Android helper. Key covering candidacy on `clickable` instead: focusability is an accessibility-traversal property and says nothing about painting over a sibling. Also generalise the existing "never condemn a marker leaf" exemption to any childless, non-clickable sibling that carries its own text or identifier, since drawing order plus geometry cannot distinguish a transparent overlay from an opaque one. * fix(android): keep descendant coverage classification on hittable Review feedback on #1733: narrowing hasActionableDescendant from `hittable` to `clickable` was over-reach. The helper emits clickable/focusable only when true, so a focusable-only control parses as clickable=undefined, hittable=true — and D-pad/TV surfaces are built almost entirely from those. A foreground surface containing them would have stopped qualifying as covering, leaving background selector/ref targets retained and actionable behind it. Descendant classification returns to the established `hittable` contract, and the leaf exemption returns to `!hittable` so a genuinely covered affordance stays condemned. The Telegram fix needs only to exclude the empty focusable wrapper itself, which is self-classification in hasOwnAgentVisibleContent — that stays on `clickable`. Replaces the test that asserted the rejected behavior with a helper-shaped regression: a foreground descendant carrying focusable without any clickable attribute must still suppress a covered target. * refactor(android): split clickable and focusable instead of collapsing them `hittable: clickable ?? focusable` collapsed two independent Android facts and made behavior depend on how a producer encodes a false attribute. The snapshot helper omits `clickable`; stock UiAutomator writes `clickable="false"`. For one focusable control those encodings gave opposite answers — verified: identical trees pruned differently and projected opposite `hittable` values. No comment can hold an invariant the representation contradicts, which is what the previous two commits tried to do. The tree now stores `clickable` and `focusable` as separate non-optional booleans, and every decision reads a named predicate: isTouchTarget, isFocusTarget, isAgentTarget, hasSemanticContent, hasDirectOcclusionEvidence, hasDescendantOcclusionEvidence, isPresentationLeaf. `canCoverSibling` consumes one derived classification (hasOcclusionEvidence) rather than choosing between raw attributes, so the self-versus-descendant substitution that caused the last review round is no longer expressible. Removing `hittable` from the tree type made the typechecker find every construction site; the public field is derived once at projection from isAgentTarget. Behavior change: a focusable, non-clickable control encoded by stock UiAutomator now projects hittable=true and participates in occlusion, matching what the helper backend already did for the same control. The Android TV test asserted the old encoding-specific value; it now asserts inclusion plus the unified projection. One encoding-parity regression replaces the patch-specific test added last round. |
||
|
|
f18f8b076f |
refactor(contracts): consolidate per-domain defineUse wrappers into one neutral defineUse (#1741)
* refactor(contracts): consolidate per-domain defineUse wrappers into one neutral defineUse
ADR 0019 §9: use declarations share one neutral defineUse; per-domain
currying wrappers around runtimeUse<PlatformRuntimeOperations>() add a
module per domain for no information. network-runtime-plan.ts,
logs-runtime-plan.ts, screen-recording-runtime-plan.ts, and
app-log-resource-recovery.ts each re-derived their own curried alias
(defineNetworkUse, appLogUse, defineScreenRecordingUse, and an inline
instantiation) from the same generic factory with the same type
parameter.
Export defineUse = runtimeUse<PlatformRuntimeOperations>() once from
platform-runtime.ts (where runtimeUse lives) and re-export it from the
platform facade. Every runtime-use declaration (networkDumpUse,
networkAdmissionUse, the app-log uses, the screen-recording uses, and
appLogRecoveryUse) now builds through that single export; the
per-domain wrappers are deleted.
Type-level only: no required/preferred keys changed, and no use
declaration was added, removed, or altered. Existing deepEqual
assertions in network-runtime-plan.test.ts, logs-runtime-plan.test.ts,
and screen-recording-runtime-plan.test.ts already pin every produced
use object's exact {required, preferred} shape, so they double as the
before/after regression proof that this refactor is behavior-neutral.
Part of #1739 (wave 0).
* fix(contracts): move defineUse into platform-runtime-operations.ts
Review: defining defineUse in platform-runtime.ts required importing the
concrete PlatformRuntimeOperations catalog into the generic runtimeUse
primitive module, while platform-runtime-operations.ts already imports
generic runtime types from platform-runtime.ts. That's an avoidable reverse
type dependency — the lower generic primitive depended on its concrete
aggregate catalog. The unchanged SCC file count didn't prove this harmless;
it counts cycle members, not newly introduced back-edges.
Move defineUse = runtimeUse<PlatformRuntimeOperations>() into
platform-runtime-operations.ts, alongside PlatformRuntimeOperations.
platform-runtime.ts no longer imports the concrete catalog. Re-export
defineUse through the platform facade from its new source module; every
call site keeps importing it from @agent-device/contracts/platform
unchanged, and the three contracts-internal call sites now import it
directly from platform-runtime-operations.ts.
Validation: tsc (full workspace + examples/sdk), check:layering (136/136,
type-cycle count unchanged at 46), and check:affected --run (473 files /
3939 tests) all clean at the new head.
|
||
|
|
bd233d54dc |
refactor(daemon): absorb androidAdbExecutor DI slot into the generic provider scope (#1742)
* refactor(daemon): absorb androidAdbExecutor DI slot into the generic provider scope request-handler-chain.ts threaded a platform-concrete androidAdbExecutor parameter through the generic RequestHandlerChainParams, flattened out of request-router.ts's already-existing generic per-request provider-injection mechanism (RequestPlatformProviderScope, built by withRequestPlatformProviderScope in request-platform-providers.ts) and passed down as a separately named chain-level slot. RequestHandlerChainParams now carries the neutral providerScope: RequestPlatformProviderScope instead of a platform-named field. The session route (the sole in-chain consumer) reads params.providerScope.androidAdbExecutor when building handleSessionCommands' own params, which is unchanged. request-router.ts passes the whole resolved providerScope straight through instead of extracting androidAdbExecutor out of it first. The generic chain signature no longer names any platform. ## Why this is behavior-neutral - The exact same RequestPlatformProviderScope object (and its androidAdbExecutor field, resolved the same way by withRequestPlatformProviderScope) now reaches the session handler; only the path it travels through RequestHandlerChainParams changed, not its value or resolution. - session.ts, session-doctor.ts, session-doctor-android.ts, session-native-perf.ts, and session-observability.ts — everything downstream of handleSessionCommands — are untouched. - Full provider-integration coverage (android-lifecycle.test.ts's "Android Settings flow uses scripted ADB provider", doctor.test.ts, android-recording.test.ts, etc.) exercises the real end-to-end wiring through the daemon route and passes unchanged. ## Validation - New wiring-equivalence test (request-handler-chain-provider-scope.test.ts) mocks handleSessionCommands and proves the exact androidAdbExecutor function reference passed into providerScope reaches it unchanged for a session-routed request, and that an empty provider scope forwards undefined (as before). - pnpm exec tsc -p tsconfig.json — clean - pnpm check:layering — 136/136 structural tests pass - pnpm check:affected --run — 80 test files / 337 tests passed (including the full provider-integration suite), plus format, lint, layering, build/declarations, and vitest-related checks Part of #1739 (wave 0). * test: trim history-narration header comment in provider-scope wiring test Per review: comments should state constraints code can't show, not narrate migration history that goes stale at merge. Reduce the header to what the tests prove. |
||
|
|
73b54c15c4 |
docs: amend ADR 0019 with broader-migration governance (#1738)
Amends ADR 0019 for migration beyond the adoption checkpoint: - status: recordings command unit recorded as completed; no further command unit is authorized by Status — subsequent units are planned, budgeted, and authorized through the successor tracking issue - section 6: explicit none platform-execution mode with invariants; the cutover gate applies to platform-executing descriptors only and rejects a false none declaration; silent registry-entry mode defaults prohibited - section 8: move-dominated size accounting, deprecated surfaces die on legacy at the next major, two evidence tiers, one parametrized cutover gate, capability buckets deleted per unit - section 9: one bind per handler, side-effect-free facts inspection, single defineUse, preferred operations require a recorded measurement - section 10: cross-cutting facets land with their first consuming unit, evidence-gated startup recovery, two-phase gateway shutdown, teardown steps ride their owning domain, all-edge-kinds end-state layering rule over production daemon modules |
||
|
|
f5d9789764 |
feat: enforce local device claims and reconcile stale owners (#1735)
* feat: enforce local device claims * fix: address device claim review feedback * fix: persist canonical daemon claim state directory |
||
|
|
602b7a2995 |
refactor: narrow perf API to actionable evidence (#1731)
* refactor: narrow perf API to actionable evidence * fix: address perf API review feedback * fix: preserve deprecated Android CPU metrics |
||
|
|
b8dd6a5854 | refactor: tighten capture ownership boundaries (#1736) | ||
|
|
167f93ba8c | 0.20.8 v0.20.8 | ||
|
|
b0d4b40467 |
chore: stop publishing skills to npm (#1730)
* chore: stop publishing skills to npm * fix: align simulator skill startup * docs: align agent setup with open-first workflow |
||
|
|
57fc0f99fa | test: fail closed on unknown recording provider commands (#1728) | ||
|
|
62001cf210 |
refactor(record): derive session recording from the publication lifecycle (#1719)
* refactor(record): derive session recording from the publication lifecycle `SessionState.recordSession` stored an answer the script-publication aggregate already contained. Every writer set both, but nothing made them agree, and #1533 was the consequence: a `--save-script` ingress re-armed the flag behind an ABORTED status, and a bare `close` published a recording the caller had been told was aborted. That fix routed every write through one rule, which made the two agree without making disagreement unrepresentable. The field remained a second source of truth, and its doc comments had to carry the invariant that a type could enforce. Remove the field and derive the answer. `isRecordingPublication` reads recording off the lifecycle: ordinary authoring records only while ARMED; a repair transaction records for its whole lifetime, terminal statuses included. That last clause is deliberately exact rather than merely safe — `armRepairStep` armed the old flag and neither `abortRepair` nor `commitRepair` ever cleared it, so narrowing it would silently stop evidence capture for a committed repair. Whether it should is a real question, and a behavior change, so it is left alone here. What this buys, beyond one less field: - `buildNextOpenSession` and `finalizeOrdinaryCloseScript` make no recording decision at all now, so no surface can arm recording without moving the lifecycle that authorizes it. - The writer's publication gate is answered entirely by the aggregate. Its separate ABORTED check is gone: a terminal authoring lifecycle is already not recording, so one question replaces two that could disagree. - The R7 ownership ratchet drops from 23 writer-owned fields / 29 owner claims to 22 / 26, and the layering manifest loses the entry whose comment documented the smell ("deliberately set on its own by paths that record without arming a publication"). Behavior-preserving: the derivation reproduces what the flag held at every transition. The test fixtures that armed `recordSession` with no publication state described a shape production stopped producing at #1478; they now carry the lifecycle that causes recording. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PFW9gJqz1wEHoowkdd1nFW * test(close-script): flush queued event-log writes before removing the tmp root CI failed the Coverage lane with ENOTEMPTY removing the test's tmp root, in `afterEach` rather than in an assertion. `SessionStore.recordAction` QUEUES its event-log append (`queueEventLogWrite`) instead of writing it, and every close path in this file records an action. Nothing awaited that write, so `fs.rmSync(root, {recursive: true})` could race it: the pending append recreates `<root>/sessions/<name>/` while rmSync is walking, and the final rmdir fails ENOTEMPTY. It needs CI's parallel load to lose the race — the file passes 12/12 in isolation locally. Await `flushSessionEventLogWrites()` before removing. The hazard is latent in any test that records actions and then removes its tmp root; this fixes the file that failed rather than sweeping the pattern, which deserves its own change. Not added to the #1419 contention-retry list: that list requires a concrete spawn/wait mechanism named per entry, and this file has none. The race was a real teardown bug, not lane contention. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PFW9gJqz1wEHoowkdd1nFW * docs: correct ADR 0016 on recording vs publication for repair Review caught a real overstatement. The amendment claimed evidence capture and publication authorization are "the same question asked of the same state". That holds for ordinary authoring — ARMED both records and publishes, ABORTED and PUBLISHED do neither — but not for repair: `isRecordingPublication` is true for every repair status including `committed` and `aborted`, while the writer additionally applies `isRepairArmedWriteBlocked`, refusing a committed transaction and one that is not yet committable. State it as it is: both decisions derive from the same aggregate, but they remain distinct predicates, and collapsing them would republish a committed repair or commit an incomplete prefix. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PFW9gJqz1wEHoowkdd1nFW --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
1b2e786128 | refactor: move screen recording onto platform runtime (#1724) | ||
|
|
e1684091e4 |
fix: complete request cancellation propagation (#1709)
* fix: route duration waits through cancellation-aware sleep AI-assisted implementation. The code and validation evidence were reviewed before submission. * refactor: expose shared cancellation-aware wait sleep AI-assisted implementation. The code and validation evidence were reviewed before submission. * fix: forward request cancellation through snapshot runtime AI-assisted implementation. The code and validation evidence were reviewed before submission. * test: cover duration wait cancellation authorities AI-assisted implementation. The code and validation evidence were reviewed before submission. * test: cover snapshot request cancellation propagation AI-assisted implementation. The code and validation evidence were reviewed before submission. * style: normalize wait cancellation test ending AI-assisted cleanup. The remote content was compared byte-for-byte with the reviewed local test file. * style: normalize snapshot cancellation test ending AI-assisted cleanup. The remote content was compared byte-for-byte with the reviewed local test file. |
||
|
|
c7242f877f | refactor: extract durable capture resource lifecycle (#1720) | ||
|
|
191d49c7db |
fix: restore gzip artifact uploads (#1727)
* fix: restore gzip artifact uploads * test: cover resumable gzip uploads |
||
|
|
1b75e102d7 | perf: speed up device inventory and status (#1723) | ||
|
|
338aa2a0d5 |
refactor: route every native selector resolution through the policy interface (#1715)
* refactor: route every native selector resolution through the policy interface #1649 declared the per-caller ambiguity matrix; four native call sites still bypassed it, spreading `selectorResolutionKnobs(row)` into a raw `resolveSelectorChain` instead of naming the row. That left the "one interface" claim aspirational: a caller could restate its contract as engine knobs and nothing would notice. - `is` non-exists, `get text`/`get attrs`, find's read actions, and the covered-selector diagnosis probe now call `resolveSelectorChainWithPolicy` with their existing row. Semantics are byte-identical: the knob-backed branch of that interface forwards to the same engine call the call sites built by hand. - The façade drops `resolveSelectorChain` and `selectorResolutionKnobs`, so no knob-taking resolver is reachable from outside the package and a call site cannot re-acquire the knobs even by accident. `requireUnique`/`disambiguateAmbiguous` are now named in exactly one function, which `resolve-with-policy.ts` and the replay resolver both derive through. - `get` names the two rows it may consume as a type, so pointing it at any other ambiguity contract is a compile error. Tests: selector-read-policy.test.ts pins which row each read command consumes, end to end, on one ambiguous fixture — the only tree the rows disagree on. Each assertion was proven red by re-pointing its caller at a neighbouring row. The knob-consistency check moves into the package beside the now-private helper. Test call sites that used the raw resolver move to `resolveRecordedTarget`, the same knobs and the path that actually replays a recorded chain. Extracting the failure branch drops `resolveSelectorInteractionTarget` below the complexity threshold; its `fallow-ignore` waiver is removed (verified load-bearing before the extraction, unnecessary after). Closes #1630. Structural stages (occlusion, off-screen, promotion, poll budget) stay per-caller pipeline code, tracked in #1656. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD * test: observe which node find's row selected, not just that one existed #1715 review, P2: the find row assertion was only half a pin. `find exists` returns `found: true` for any resolved node, and the `list` call it leaned on goes through listFindMatches — a path that consumes no policy row at all. So repointing findFirstLocatorMatch at `readText` left both assertions green while selection silently moved from the document-order head to the tiebreak winner. Assert through `find get_attrs`, which returns the ref of the node the row actually selected. Both neighbouring rows are now red: `readText` fails '@e3' !== '@e2' (the move the old test missed), `readUnique` fails by refusing the ambiguous screen. `exists` stays as a second, weaker assertion on the same resolution. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD * refactor: route is exists through the matrix, collapse the double match pass Follow-up tightening on the same seam. `is exists` reached findSelectorChainMatch directly while the `readAny` row's own doc claimed to serve "`exists` and find's read-only actions" — true of the docs, not of the code, which is the unverifiable-claim shape #1656's review called out. It now names `readAny`, the row it always described. Equivalent by construction: both take the first alternative with any match under requireRect: false, and disclose that alternative's count. That leaves the root façade with no consumer for findSelectorChainMatch, so it goes the way of resolveSelectorChain — dropped from the string-only façade, kept on the published ./ast surface. Its façade-twin type SelectorChainMatch dies with it (fallow caught it). resolveSelectorChainWithPolicy matched twice on the uniqueness path: once via resolveSelectorChain, then again to fill matchedNodes. Hoisting the single list call above the row switch removes that second pass, collapses two duplicated ambiguous literals into one helper, and drops a `?? [resolution.node]` fallback that was unreachable — a resolution implies its alternative matched, so the list is never null there. While hoisting: the resolved arm's matchedNodes can describe a different alternative than resolution.selector, because uniqueness skips an ambiguous alternative to try the next one. Unreachable today (only first-match callers read it, where both come from one list), and left as-is rather than silently changed — but the doc claimed "the alternative it came from", so it now says what is actually true. Tests: is exists gets a caller-level pin on the shared ambiguous fixture — passes with matches: 2 where its fail-closed siblings refuse — proven red by pointing it at readUnique. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD * test: discriminate is exists's row by alternative, guard the façade structurally #1715 review, second regression-validity gap. The `is exists` pin observed only `pass: true` and `matches: 2` on a fixture whose first alternative was merely TIEBREAKABLE — so disambiguation succeeded there and reported the same count first-match would. `readAny`, `readText`, and the pre-migration raw lookup all produced that, and only the readUnique swap I had checked went red. One mutation proven is not the same as the row being pinned. `exists` exposes no node ref, so the row has to be read off WHICH alternative answered. New fixture: alternative one matches two nodes that are genuinely indistinguishable (same depth, same area, both on screen) so the tiebreak declines; alternative two matches exactly one. First-match answers from alternative one; every uniqueness row skips the undecidable alternative and answers from alternative two. Asserting the selector now separates them — readText and readUnique both fail with `id="save-unique"` where `label="Save"` is expected. Restoring the raw lookup stays behaviourally invisible, though: findSelectorChainMatch is equivalent to the readAny row it migrated to, which is precisely why that migration preserved semantics. No fixture assertion can catch that revert, so the guard is structural — the façade's export list must not carry resolveSelectorChain, findSelectorChainMatch, or selectorResolutionKnobs. Follows the packages/maestro index.test.ts absence-assertion precedent. Verified red by re-exporting the lookup. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HuKzQWn6WQcMYaAZVvJzdD * fix: cover selector routes in device replays * test: simplify selector replay regression --------- Co-authored-by: Claude <noreply@anthropic.com> |