mirror of
https://github.com/callstack/agent-device.git
synced 2026-09-14 20:06:34 +08:00
agent/ci-only-affected-coverage
1486 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
2836006ac6 | perf: keep affected coverage on CI | ||
|
|
d29dc22861 |
fix(android): warn when a permission revoke kills the session app (#1856)
* fix(android): warn when a permission revoke kills the session app settings permission deny|reset maps to pm revoke, and Android kills the app's process whenever a runtime permission it currently holds is revoked, so a grant -> deny/reset sequence silently left the session on the launcher and the next selector failed with no hint. The revoke path now reads the prior grant state from dumpsys package first and, when it was granted, returns wasGranted: true plus a warning naming open <app> --relaunch; the settings CLI output renders response warnings, and commands.md documents the behavior next to the pm revoke mapping. Closes #1796 * fix(android): state the revoke-kill consequence conditionally Review finding: the warning asserted the app had been killed, inferred only from the prior grant state. dumpsys reports granted=true for any user profile while pm revoke acts on the current one, and the app need not have been running, so the claim could be false. State the platform rule and make the consequence conditional; the relaunch guidance is unchanged. * fix(android): model the prior grant state as granted/not_granted/unknown A failed or unparseable dumpsys read as "nothing granted", so the response asserted the app was untouched when the state was simply unknown, and the grant scan matched every granted=true line — install permissions and other users' blocks included — so another profile's grant could claim a kill that never happened. Both directions of the same defect. The read now resolves the acting user (am get-current-user) and walks the dump's nesting (Packages: > User <id>: > runtime permissions:), and reports priorGrantState: granted | not_granted | unknown. unknown carries the same relaunch guidance without claiming what the state was; only not_granted is silent. * fix(android): address the foreground user in every permission mutation The tri-state read scoped state to am get-current-user, but the mutations ran bare pm grant/revoke and clear-permission-flags. PackageManagerShellCommand defaults those to UserHandle.USER_SYSTEM, so on a device whose foreground user is nonzero the command read one user's state and edited user 0 — leaving the running app's permission untouched while reporting on a user it did not change. Proven on a Pixel 7 / API 36 emulator with the foreground user switched to 10: a bare pm revoke flipped User 0 to granted=false and left User 10 granted=true. The foreground user is now resolved once and passed as --user to pm grant/revoke, pm clear-permission-flags, and appops set, and the state read takes that same id. When it cannot be resolved the mutation keeps the platform default and the state is reported unknown rather than guessed. * test(android): pin the user-scoped permission argv in the provider scenario The scripted ADB provider answered only the unscoped pm grant/revoke form, and the Settings contract asserted the unscoped transcript entry, so the provider lane could not see which user a permission mutation addressed. * test(android): extract the settings contract out of the lifecycle monolith The user-scoped argv assertions pushed android-lifecycle.test.ts past its size ratchet, whose instruction is to extract rather than grow a file over the tripwire. assertAndroidSettingsContract moves to a sibling module and the pin drops 1597 -> 1559. * refactor(android): shrink the permission path to one concept per file Size/design pass on the #1796 change: - settings.ts was 505 lines (past the 500 extract-before-adding tripwire); the permission family moves to settings-permission.ts and the dispatcher drops to 265. - permission-grant-state.ts loses topLevelSection (a nestedBlock with an indent-0 header), its single-use line reader, and androidPriorGrantState (one map lookup at its only production call site). - the grants map narrows to 'granted' | 'not_granted': unknown was never a value, absence is what carries it, so the tests read the map directly. - the permission tests move to settings-permission.test.ts and consolidate into argv/tri-state/photos/rejection tables; the parser tests fold seven cases into two. Every red-proof re-run after the consolidation: dropping --user reds 7 argv/photos cases, and the pre-fix state model reds 13 across both files. * fix(android): refuse permission mutations that cannot name their user The fallback issued bare pm/appops commands when am get-current-user did not answer, which is the #1796 defect itself: those default to UserHandle.USER_SYSTEM, so a session running as user 10 had user 0 edited while the response reported only priorGrantState: unknown. It was also a fallback added without approval, and the docs' claim that every mutation names its user was false on that path. Resolving the acting user is now a prerequisite: setAndroidSetting permission fails with COMMAND_FAILED and a recovery hint, issuing no pm, appops or clear-permission-flags call at all. The test that locked the fallback in is replaced by one asserting the empty mutation call list for grant, deny and reset. |
||
|
|
393eb30a28 |
ci: give check:affected real Apple ownership rules and route ios.yml on them (#1781 A9-2) (#1857)
* ci: give check:affected real Apple ownership rules and route ios.yml on them (#1781 A9-2) Device-lane ownership by platform family in the affected selector (scripts/check-affected/device-lanes.ts): a TypeScript-only Apple change now carries replay-ios/replay-ios-device/replay-macos in a narrow plan, other families own only their own lanes, shared runtime surface owns every lane, unit tests own none. Golden tables (contracts/fixtures) own the parity unit test and both runner builds instead of failing open. ios.yml pull_request paths-ignore is routed on that ownership; the gate manifest asserts the list against the selector over every tracked path both ways (scripts/gate/routing.ts, ROUTED_LANES). push to main is unfiltered. Path coverage exempts declared manual-only checks the way owned does. * ci: tighten routing assertion shape (fallow: unused exports, complexity) * ci: name parked checks in check:affected --run skips * ci: bound the routed-lane exemption to sibling workflows (review of #1857) The exact-name .github exemption was unbounded: naming the lane's own setup-apple-runner-build or boot-ios-test-simulator action skipped the lane that runs them and the manifest stayed green. Lane now carries the transitive composite-action closure plus its own workflow file (Lane.uses, same walk declaredGates does), and the exemption refuses anything in it. Also: an unowned path under an ignored root (a non-TS fixture under a family root) asked for the ignore entry to be removed, which would un-route every sibling in that tree; it now asks for a selector owner. Both cases pinned, both proven red against the pre-fix code. Documents GitHub's 300-changed-file path-filter limit in docs/agents/testing.md. * ci: close the routed-lane exemption over composite-action support files Lane.uses recorded only each composite action's action.yml, so a support file the descriptor executes was exemptible as if it were an unrelated sibling workflow: ios.yml uses setup-fixture-app, whose action.yml runs "$GITHUB_ACTION_PATH/fetch-artifact.sh", and that script runs its siblings resolve-artifact-name.sh and trusted-artifact.mjs — references that exist only inside shell, one level past anything YAML parsing sees. The closure unit is the action's directory now. It needs no shell model and cannot miss a file however deep the reference chain runs; the coarseness is harmless because a file in an action's own directory belongs to that action. All three files pinned, red against the descriptor-only closure. |
||
|
|
735ab7672a |
refactor(daemon): one capture-input builder and one admit-then-bind step (#1876)
Behaviour-neutral. No descriptor changes platform execution, the cutover table is untouched, and no contract surface is added. - buildRuntimeCaptureInput moves to its own module so every request-bound capture consumer builds CaptureSnapshotInput one way. - The admit-then-bind sequence in the snapshot/diff resolver becomes one named step, ready for the selector units' second caller. - handlers/find.ts splits into focused target-capture and match-resolution concepts (600 -> 346 lines); behaviour unchanged. Co-authored-by: agent <agent@local> |
||
|
|
d8e03aea9b |
refactor: migrate screenshot to request-bound runtime (#1878)
Retires the last dispatchCommand edges for screen capture: the generic-route command, the sparse-snapshot fallback, and the Android snapshot-timeout evidence capture all admit exact owner facts and bind once (ADR 0019, cutover rule R39). --overlay-refs becomes part of the declared use, so a target that can capture pixels but not a tree is refused before anything is written to disk. |
||
|
|
67ce19b50c |
fix(daemon,kernel): device-selection safety — identity conflicts and ambiguity fail instead of retargeting (#1880)
* fix(daemon): session-lock identity conflicts fail instead of changing device identity
`--session-lock strip` resolved every conflict by deleting the offending selector, including
--udid/--serial/--device. The command then ran against the BOUND session's device rather than the
one the caller named, and the error hint that produced this state actively recommended strip. A
wrong-device tap that reports success is worse than any loud failure, so:
- a conflict on a device IDENTITY selector now fails under both reject and strip; strip keeps
resolving platform/scope selectors (--platform, --target, --ios-simulator-device-set,
--android-device-allowlist), which is what it exists for;
- the structured error carries both sides (requestedDevice, boundDevice) so a caller can choose a
recovery without parsing prose;
- the fresh-session hint offers the two real recoveries — close the bound session if the requested
device is intended, drop the selector if the bound device is — and never mentions strip for an
identity conflict. The existing bound-session hint already stated both and is unchanged.
Admission-layer only: this is device/session resolution policy, evaluated before any interaction
dispatch path is selected, so no ADR 0011 guarantee cells move.
One table pins the crossing: fresh vs existing session, matching/conflicting/absent identity,
reject vs strip, binding vs inventory command, Android serial / Android udid / Apple udid, with the
exact error details and hint asserted. 7 of its 13 rows are red against the previous policy.
* fix(kernel): validate platform-specific device flags before resolution
`--udid` enters Apple resolution unconditionally, so `--platform android --udid emulator-5580`
answered "No Apple device with UDID emulator-5580" — an answer about a platform the request had
explicitly excluded, which reads as a missing device rather than a mistyped flag. Both directions
now fail as INVALID_ARGS naming the correct flag (--serial for Android/HarmonyOS, --udid for Apple).
Requests that name no platform keep the existing DEVICE_NOT_FOUND behavior, since nothing
contradicts the selector there.
* fix(kernel): refuse ambiguous singular device resolution instead of picking one
resolveDevice answers with exactly one device, so every caller of it needs one concrete device. When
the request carried no identity and several candidates were equally preferred, it returned the first
by discovery order (or alphabetically for Apple) — a successful response describing a device the
caller never selected. That is worse than any loud failure, and reads are not safer than writes: a
snapshot of the wrong emulator is a wrong answer that looks right.
Ambiguity is now a refusal at the resolver boundary, not a command-kind allowlist:
- established preference tiers are preserved (virtual over physical, the Apple kind/target rank,
then booted over offline); only what survives them equally is ambiguous, since the comparator's
remaining tie-breaks are name and discovery order, which encode nothing about intent;
- one booted emulator beside offline candidates still resolves, as do explicit --device/--udid/
--serial and any command inside an existing session, whose identity is already fixed;
- multi-device commands ("devices") never enter singular resolution;
- the error reuses the declared device-candidate details domain (AMBIGUOUS_MATCH + "devices"), so
CLI and MCP already render the bounded list, with a hint naming the platform-appropriate selector.
* fix(kernel): drop the mismatched article from the selector-flag hint
"Use --serial X to select a android device by serial" — the platform name is interpolated, so no
article fits every value. Names the flag's platform family instead.
|
||
|
|
debda139a3 | fix: stabilize global flag declaration type (#1879) | ||
|
|
ef2094c9d7 |
docs: keep size review in CI and local feedback fast (#1842)
* feat(size): measure a base ref in one command (pnpm size --base <ref>) The Size workflow already compares base and PR builds; locally that needed a manual checkout, install, build, --json, and --compare dance, so budgets were negotiated late. --base <ref> does the workflow's recipe in a detached worktree under .tmp/size-base/<sha> (kept for reuse, other bases pruned) and compares against it: first run ~1-2 min, later runs against the same base ~3s. Documents the local caveat: npm tarball/unpacked rows compare a fresh base against a working tree that may carry locally built helper artifacts. * feat(tooling): pnpm pr:evidence — one paste-ready, SHA-stamped evidence block for PR bodies Composes what the repo already measures instead of hand-transcribing it after every rebase: exact merge-base and head, changed-file areas, the affected selector's plan (local vs GitHub-authoritative, fail-open summarized), the layering guard verdict, depgraph counts with a real delta against the base (a throwaway git worktree, no install — the script analyzes its cwd while its imports resolve from this checkout), and, behind flags, the changed-line coverage table and pnpm size --base. It claims nothing about CI: the last line links the head's checks. ~20s default tier. The pure model (grouping, report parsing, rendering) has node:test coverage registered as the pr-evidence-model gate, run in the Affected-check Selector job next to the selector it reads. * fix(tooling): pr:evidence measures pristine head/base worktrees from an os.tmpdir scratch; size --base gets a per-SHA lock, completion stamp, and non-destructive eviction Review (three P1s): - pr:evidence created its scratch under an untracked .tmp/ that a fresh checkout lacks (ENOENT). Scratch now lives under os.tmpdir(), which exists by construction; a real entrypoint regression runs the whole pipeline with --base HEAD (no origin/main needed) and asserts JSON shape plus cleanup of both worktrees and the scratch. - Untracked or uncommitted production files could move the layering/depgraph numbers the block labels as HEAD's. Head is now measured from a pristine worktree of the head commit exactly like base, and the affected plan takes the head SHA (the literal HEAD folds the working tree in). The dirty flag now counts untracked files and says they are not in the block. - size --base force-pruned other cached bases without locking and trusted a dist/src that could be half-built. Per-SHA .lock (pid, O_EXCL) held from before the worktree exists until the base report is read; a live lock on the same base fails fast, a stale one is replaced; eviction skips worktrees whose lock owner is alive; dist/.size-base-complete marks a finished build. Orchestration tests run the real script against a throwaway git repo with pnpm/npm shimmed on PATH (build once, reuse, live lock, stale lock, interrupted build, guarded vs idle eviction). Also fixes the /tmp → /private/tmp realpath mismatch those tests surfaced (git lists worktrees by real path, so the registration check removed a live worktree). * fix(tooling): symlink-identity locks with compare-then-unlink; evict under the victim's lock; pr:evidence registers worktrees on add and cleans up exhaustively Review (three P1s): - Lock creation/takeover races: the lock is now a symlink whose target is the owner identity (pid:nonce), created with its identity in one syscall (no empty-file window), taken over only by compare-then-unlink on the exact identity judged stale, and verified after creation; release unlinks only a link that still names this run. Real overlapping-process tests: two runs on one base (exactly one builds, the other fails fast), and a takeover race against a simulated other taker across delays straddling the acquire window (a live lock is never unlinked, both never proceed). - Cross-base eviction: a victim is removed only while holding its own lock, acquired through the same path, so a run wanting it after the check finds it locked rather than half-removed; a live-locked victim is skipped. - pr:evidence worktrees: withWorktrees registers each worktree the moment its add succeeds and sweeps every resource on the way out, collecting failures instead of stopping at the first; planted reds for both (second add fails → first removed; removal of the middle one throws → the others still go). * test(size): serialize the size --base orchestration file with the other real spawners Caught running the full unit suite on the rebased branch: the file passed in isolation but intermittently failed under broad file parallelism, where it took 14s versus ~5.5s alone. It spawns node scripts/size-report.mjs per case, which spawns git and the shimmed package managers under it — the SUBPROCESS_STUB_TESTS class exactly (starved spawns surface as a vitest test timeout instead of the orchestration assertion the case is about), so it joins that serialized project with its spawn named at the entry, per docs/agents/testing.md. No rerun layer is involved: the flake is removed, not retried. Two full-suite runs green after. * refactor(size): extract the base-cache claim protocol and make stale takeover atomic Review (P1 + architecture): Stale-claim removal was compare-then-unlink (readlink then unlink; lstat then rm for a stray file), so another taker could replace the observed entry with its live claim between the two syscalls and this run would delete the replacement. Removal now happens only while holding the entry's takeover mutex — an atomically created directory — and re-verifies the claim inside it. A replacement can appear only by creating one on a free path (the abandoned claim occupies it until the unlink) or by another takeover (needs the mutex), so removal cannot delete a replacement. A mutex leaked by a process killed inside its sub-millisecond critical section is reclaimed by age, and even a wrong reclamation is contained: both takers re-verify inside, and the winner is still decided by the atomic symlink() that follows. The protocol moves out of size-report.mjs into scripts/size-base-cache.mjs (AGENTS.md: extract past 500 LOC) — 719 → 536, with the entry lifecycle (claim → evict others → ensure worktree → build if unstamped → measure → release) owned by the module behind withPreparedBaseWorktree. Mirrored tests in scripts/__tests__/size-base-cache.test.ts plant every dangerous interleaving directly on the filesystem: replacement-after-observation, a takeover held by another run, age reclamation, release-after-retarget, and a stray non-symlink. They need no subprocess and run in 9ms, so the raced single-process case was dropped from the orchestration file, which keeps only what real processes can show. Planted red: removing the mutex makes the contended case delete the claim it must not touch. * ci(size): preserve the reporter's whole module graph, and gate that it stays whole The Size workflow measures the base commit with the PR's reporter, so it copies the reporter out of the tree before checking the base out. Extracting size-base-cache.mjs made the reporter a two-file graph while the step still copied one file, and the base measurement died with ERR_MODULE_NOT_FOUND — after every deterministic gate had passed, because nothing local reproduces that copy. The step now copies the scripts directory, so a further split cannot leave an import behind, and size-report-preserved-closure.test.ts holds it to the reporter's real relative-import closure and to running the preserved copy rather than the checked-out tree. Planted red: restoring the single-file copy fails both cases, naming scripts/size-base-cache.mjs. Verified by running the reporter from a copied directory exactly as the workflow does. * fix(size): the takeover mutex has one holder for life; split report publishing out of the reporter Review (P1 + architecture): Age-based reclamation of the takeover mutex reintroduced the split ownership the mutex exists to prevent: a holder that is merely slow — paused or SIGSTOPed past any threshold — could have its mutex force-removed and replaced, putting two takers inside the supposedly exclusive section, where either could unlink the claim the other had just created; the unconditional pathname-based release could also delete the replacement mutex. The mutex is now a symlink naming its holder, created in one syscall, never reclaimed at any age, and released only by the run that owns it. A mutex leaked by a process killed inside a three-syscall critical section wedges one cache entry with the path to clear in the message, rather than silently deleting another run's live claim. Planted red: restoring age reclamation displaces a day-old delayed holder, which the new case pins. Publishing the report to a PR is a separate question from measuring and formatting it, so it moves to scripts/size-report-comment.mjs with the marker and retry policy it owns; its existing regression drives it through the real script unchanged. scripts/size-report.mjs is 386 LOC — under the 500 tripwire and below the 512 it had on base. * test: prove delayed size cache holder is preserved * docs: keep size review in CI |
||
|
|
d07b837621 |
test: classify the runner XCTests — pure decisions to a macOS host lane, simulator semantics gated os(iOS) (#1781 A7) (#1861)
Every declared AgentDeviceRunnerUITests method now belongs to a lane, and the #if guard is the classification: AGENT_DEVICE_RUNNER_UNIT_TESTS alone means a pure runner decision (runs on the macOS host on every PR — ci.yml's existing compile job now executes the bundle it builds), '&& os(iOS)' means runner/XCTest semantics (simulator lanes only). check:xctest-selection evaluates the guards per platform, derives each lane's reach, and fails on a flagged identifier that is undeclared or uncompiled on that lane, on a declared test no lane reaches (found the two tvOS-only tests, dark since birth — widened to os(tvOS) || os(macOS)), and on testCommand reaching any lane. The host and nightly lanes assert executed == derived reach, so a missing -D flag or a guard that compiles a file out reads red, not as a smaller green. One duplicate test deleted (sparse-verdict assertions folded into its twin). |
||
|
|
e4c3b420a4 |
test: refuse foreign-pid signals from unit-test workers (Coverage fork death, #1824) (#1854)
* ci: w3-1824 experiment — trace fork signals and plant pid sentinels in the Coverage job
Temporary instrumentation for #1824. Every vitest fork logs each real
process.kill it sends to a foreign pid (and every kill/pkill it spawns);
the Coverage job parks sentinel processes on the pids the Apple runner
tests fabricate (4141/4242/4343/4444) and reports which of them survive
the run. Reverted before this PR leaves draft.
* test: refuse foreign-pid signals from unit-test workers
A vitest worker may signal only itself and the processes it spawned.
src/__tests__/hermetic-signal-setup.ts records any other process.kill,
answers it with ESRCH (so best-effort kill paths proceed as if the pid
were dead), and fails the sending test by name in afterEach.
The senders this catches today are the Apple runner tests, which
fabricate runner child pids (4242, 4141, 4343, 4444) and mocked the
liveness reads in host-process.ts but not the signal writes:
killRunnerProcessTree delivered real SIGINT/SIGTERM/SIGKILL to those
pids and their process groups — 146 signals per run of
runner-session.test.ts. On the CI runner the sibling vitest forks live
in that pid band, so the Coverage job periodically lost one fork
mid-file with no test attributed (issue #1824, 6 of the last 40 red CI
runs).
The group-signal write moves behind signalProcessGroupBestEffort in
host-process.ts, next to signalPidsBestEffort, so the runner tests mock
the signal seam in the same place they already mock the liveness reads.
Refs #1824
* Revert "ci: w3-1824 experiment — trace fork signals and plant pid sentinels in the Coverage job"
This reverts commit
|
||
|
|
db5d26bffc |
fix(utils): throw AppError from app-log-files and verified-file guards (#1853)
* fix(utils): throw AppError from app-log-files and verified-file guards The symlink, not-a-regular-file, identity-changed and identity-race guards threw plain Error, so they surfaced as UNKNOWN with a misleading hint and tests could only assert them by message. They are now COMMAND_FAILED with a recovery hint (ADR 0010); the affected tests assert through assertThrowsAppError, and the two race guards gain planted-interleaving coverage. Closes #1792 * test(utils): pin the recovery hints the AppError conversion added The tightened assertions supplied only code and message, and the helper did not look at the hint, so deleting either hint constant left every targeted test green — vacuous for the half of #1792 that ADR 0010 actually cares about. assertThrowsAppError/assertRejectsAppError now accept a hint, checked against normalizeError's view so a dropped hint surfaces as the misleading per-code default rather than passing, and each guard pins its exact text as a literal (importing the constant would compare it to itself). |
||
|
|
99754dc69c |
fix(layering): resolve relative imports inside workspace packages referencing #1781 (#1872)
* fix(layering): resolve relative imports inside workspace packages resolveTargetFile() dropped any relative import whose resolved path didn't start with src/. Since #1490 W0 added packages/*/src/** to the source set, every intra-package relative import (e.g. a facade re-exporting a sibling file) was silently invisible to the layering graph — R4 value-cycle rejection and depgraph reverse reachability both stopped at the package facade. Resolve relative specifiers that land under packages/<name>/src/ too, while still refusing anything outside src/ and packages/*/src/. Verified against the real tree: 448 previously-invisible value edges and 436 type-only edges now resolve, but none of them close a new R4 value cycle or grow the R9 type-cycle SCC, so the R6/R9 baselines are unchanged. Refs #1781 * style: apply oxfmt to regression test |
||
|
|
8c06965d28 |
fix(daemon): cap events.ndjson with cursor-safe rotation (#1867)
* fix(daemon): cap events.ndjson with cursor-safe rotation Rotate events.ndjson to events.ndjson.1 once it reaches AGENT_DEVICE_EVENT_LOG_MAX_BYTES (default 5 MB), keeping one rotated generation. Cursors stay absolute across rotation through a sidecar window offset, so a persisted nextCursor still names the same event; a cursor older than the retained window fails with COMMAND_FAILED and details.reason EVENT_LOG_CURSOR_EXPIRED instead of returning a wrong page. Closes #1788 * fix(daemon): verify the events.ndjson window against the files on disk Rotation recorded only a dropped-line offset, written after the rename, so a reader landing in that window mapped every absolute cursor a whole generation too far (reproduced: 5 of 58 reads returned event 17 for cursor 9), and a missing rotated file or stale sidecar shifted cursors permanently and silently. The sidecar now records each retained generation's first absolute line index, line count, and first-line digest, and is written before the rename it describes. The reader identifies each file on disk by digest, derives its start from the matching record, and checks the recorded line count and generation contiguity; anything unverifiable raises a typed EVENT_LOG_WINDOW_UNVERIFIED instead of a guessed offset. A torn snapshot (rotation landing mid-read from the threadpool) is retried, not interpreted. A corrupt sidecar fails reads typed and never blocks appends, and rotation no longer does synchronous whole-file I/O. * refactor(daemon): split event-log window placement and share one line splitter |
||
|
|
6984a1e095 |
fix(layering): list the whole zone when R10's type-cycle ceiling is exceeded (#1852)
* fix(layering): list the whole zone when R10's type-cycle ceiling is exceeded The per-zone R10 violation named members.find(<zone match>) — the alphabetically-first zone member, a file that had been in the cycle all along — so the +1 in #1825 x #1779 was found only by diffing largestTypeCycleMembers between commits. The ceiling records a count, not a membership, so the gate cannot name the joining file; it now lists every member of the over-budget zone and annotates the ceiling table instead. Closes #1837 * fix(layering): state the zone overflow in net terms Review nit: the overflow is net growth over the ceiling, not a join count (two joins and one departure print "1"), so the message no longer claims N members joined. |
||
|
|
79dddaf781 | refactor: tighten viewport runtime facts (#1873) | ||
|
|
f3d5b3d92c |
refactor(daemon): admit-before-bind as an admitted-plan token; retire the R32 syntax policy (#1841)
* refactor(daemon): admit-before-bind as an identity-keyed admitted-plan token; retire the R32 syntax policy admitRuntimePlan (was inspectRequiredRuntimeUse) takes the plan and, on success, mints an AdmittedRuntimePlan: a nominal class instance with nothing readable on it. Its payload — a frozen copy of the device the facts were read for, and the plan — lives in a module-private WeakMap keyed by the token's exact identity, and the only way to read it is unwrapAdmittedRuntimePlan, which refuses anything not minted here. The snapshot owning interface (resolveBoundSnapshotCaptureRuntime, #1847) admits through it and its private binder takes only the token: no bare plan, no separate device, and no look-alike — a spread lacks the #private member (not assignable), a Proxy around a real token types as the token but is a different identity (refused at unwrap), Object.assign/defineProperty throw on the frozen instance, and the class value is not exported so its constructor is not nameable. That retires scripts/layering/runtime-command-cutover-snapshot.ts — R32's per-command AST policy (call-shape recognition of the admission and a text sniff for a local admission) — and the source-regex test beside the descriptor tests. The generic row keeps retirement, narrowing, and singular execution; the manufactured-proof column now also rejects casts to AdmittedRuntimePlan. Planted reds: token degraded to a plain public shape → 2 unused @ts-expect-error directives; unwrap reading the token surface via getters → the Proxy regression fails; getter-based branded literal → the runtime retarget test fails. * docs(agents): the ADR 0019 unit checklist teaches the shipped admission API #1836 documented inspectRequiredRuntimeUse with a forward note pointing here; this PR makes admitRuntimePlan real, so the row now teaches it plus the identity-keyed unwrap the binder uses, and points at the shared snapshot/diff owning interface as the model. |
||
|
|
37b1bc8cbd |
refactor: migrate viewport to request runtime (#1864)
* refactor: migrate viewport to request runtime * fix: preserve viewport cutover evidence |
||
|
|
fda81c5121 | 0.20.10 v0.20.10 | ||
|
|
9fb307aa63 |
fix: derive transition backend from capture (#1851)
* fix: derive transition backend from capture * fix: retain iOS transition confirmation without provenance * test: guard backendless local settle latency * fix: confirm transitions from modal captures * fix: confirm transitions from tiny iOS modals * fix: recover settle transition baseline from session * fix: trust ref frames for settle transitions * chore: format transition settle tests * refactor: simplify settle transition decisions * test: isolate private ax recovery budget * test: keep settle test within size ratchet |
||
|
|
275a66ea23 |
test(ratchet): pin snapshot-handler.test.ts at main's 2654 lines (#1860)
#1847 grew the file 2652→2654 and merged minutes before #1843 pinned it at 2652 (measured against the merge-base #1843 had at the time). Both PRs were green alone and main is red together — the cross-PR growth the equality pin exists to catch, landing as a catch-up rather than a raise: the history rule agrees (2654 at the merge-base). |
||
|
|
b12a3e3cb3 |
test: pin test files over 1,000 lines at their exact length so they can only shrink (#1843)
* test: pin test files over 1,000 lines at their exact length so they can only shrink AGENTS.md has said for a while that past 1,000 lines is architecture debt and tests are not exempt; nothing enforced it, and the second-largest test file gained 55 lines in the PR before this one. This is the slow-test ratchet's shape for a reader's context instead of wall clock: the 26 test files over the tripwire are pinned at their exact length (R9-style equality pin, #1781 A6); growth fails, shrink fails until the pin is lowered in the same PR, a file that drops under the line leaves the list, and a new file may not cross it. One directory walk per unit run, ~250ms; the pin list emptying deletes it. * test(ratchet): hold giant test files to their merge-base length so pin edits cannot admit growth Review (P1): the equality pin compared measured lengths only against the pin map in the same checkout, so growing a file and raising its pin, or adding a new >1,000-line file with a pin, stayed green. The gate is now history-backed: every test file over the tripwire may be no longer than at the merge-base with origin/main (renames followed; new files may not cross the line), and no pin may exceed its file's base length — one git cat-file --batch spawn, parsed by bytes because the sizes are bytes. Both bypasses planted red against real git on a pinned file and on a fresh 1,001-line file with a pin added. * test(ratchet): a pin on a file at or under the tripwire is itself a finding Review: a new pin for an unchanged sub-tripwire file (900 pinned at 900) passed equality and history and grew the map. Pins now exist only for files over the tripwire — any other pin is red with 'remove it' — which also subsumes the old shrink-under-the-line message. Planted red in-file and against real git (a 186-line test pinned at 186). The android snapshot test pin bootstraps 1636→1660: main grew that file in #1846 before this gate exists, and history agrees (1660 at the merge-base). |
||
|
|
3f0f706f0b | refactor: migrate diff to request-bound runtime (#1847) | ||
|
|
ee13203a16 |
feat(ios): unify snapshot eligibility (#1850)
Make iOS regular snapshot eligibility one backend-neutral presentation rule. Acquire tree nodes conservatively, preserve interactive scroll containers, normalize surviving hierarchy, and keep raw membership plus daemon publication policy unchanged. Part of #1797. - iOS and macOS unit-enabled runner builds - 2 focused XCTest cases - 3 production-path publication tests - live Settings snapshots: 73 regular nodes and 167 raw nodes, both healthy tree captures |
||
|
|
294654a3a5 |
fix(android): resolve snapshot scope once and disclose the API 23 occlusion-scan gap (#1832 C1/C2) (#1846)
* fix(android): resolve snapshot scope once and disclose the API 23 occlusion-scan gap (#1832 C1/C2) - Android resolves --scope inside its projection only, under the shared scope specification (matchesSnapshotScope in @agent-device/contracts/snapshot: first document-order match over label/value/identifier, empty on no match). The daemon post-wire scopeSnapshotNodes pass skips the android backend, so scope has one owner and one no-match semantics instead of BFS+fallback followed by document-order+empty. - Golden table contracts/fixtures/snapshot-scope-policy.json is asserted against the predicate, the Android projection, and the daemon pass; the Swift runner twin (#1797) consumes the same table. - androidSnapshot.occlusionScanUnavailable discloses helper trees without drawing-order (API 23), where the covered-sibling pruner cannot run. Disclosure only; C1 stays open until occlusion moves to the daemon annotator. * fix(android): resolve scope over the presented tree and stop dropping it on interaction captures Adversarial review findings on the first commit: - BLOCKER: captureSnapshotData spread `snapshotScope: undefined` over flags, so an interaction capture (press/click/fill/longpress/hover --scope, --settle observation) reached the Android platform unscoped while buildSnapshotState still saw the scope. The post-wire pass used to rescue it; after skipping android it returned the unscoped tree. One effective scope now feeds both. - Scope resolves over the PRESENTED nodes of the requested projection, not the acquired tree, so an acquired match that membership drops no longer empties the snapshot, and Android matches the domain iOS's pass uses. - Slicing after the walk keeps ancestor context (hittable/collection/chrome) above the scope root, so scoped -i is a subset of unscoped -i; --depth stays scope-relative. - Shared findSnapshotScopeRange/reindexSnapshotNodes so the daemon pass and the Android projection run one implementation; scope slice extracted to ui-hierarchy-scope.ts (mirrors its test). - parseUiHierarchy moved to a test fixture module (it had no production caller left). - Golden rows sharpened (value row no longer matches via label on Android); the Android leg runs raw AND regular. CHANGELOG entry; docs wording corrected for iOS/@ref. * fix(layering): keep the contracts snapshot façade exhaustive over snapshot-scope * fix(android): scope to the first match whose subtree still has presented content Review P1 on #1846: with scope resolved strictly over presented nodes, `snapshot -i --scope panel` answered "no nodes" whenever the matched container was a structural view membership drops — even though the button inside it was exactly what was asked for — and `--depth 0` then hid a node the response prints at depth 0. The scope root is now the first document-order match whose subtree contributes at least one node to the requested projection, and the result is that subtree's presented nodes re-rooted at depth 0. Both failure modes die: a decorative match membership drops no longer empties the snapshot, and a dropped container still scopes to its content. `--depth` under scope filters the depths the response emits, so a node shown at depth 0 survives `--depth 0`. Tests: the golden legs stay raw+regular (bare TextViews cannot survive -i, so an -i leg would measure membership, not scope) with the projection interplay pinned by two dedicated tests on actionable shapes; the 'not re-scoped after the wire' case now runs a real parsed scoped tree instead of fabricated depth-0 siblings. |
||
|
|
72d421fe36 |
docs(agents): ADR 0019 unit checklist, owning-seam mock rule, worktree and rebase guidance (#1836)
* docs(agents): ADR 0019 unit checklist, owning-seam mock rule, worktree and rebase guidance Retro follow-up (item 2). Adds docs/agents/adr-0019-unit.md — the order of operations for one command unit with the declaration site for each step, the evidence a unit review must carry, and what 'done' is not — so the pattern rediscovered during the snapshot unit (#1779) is written down once. testing.md: mock the seam the code under test consumes (fake inspectFacts / bindDevice), not the generic dispatchCommand mock; a migrating command moves its tests off the dispatch mock in the same PR. AGENTS.md: fresh-worktree preflight (pnpm install + build in the worktree; layering scan reads tracked files only) and concurrent-agent hygiene (one full gate per host, verify subagent edits with git -C, one PR per worktree). pull-requests.md: two readiness claims (published-and-reported vs merge-ready) and the rebase rule — main has no up-to-date protection; rebase on conflict or when `check:affected --base <merge-base> --head origin/main` names your surface. * docs(agents): name the admitted-plan token in the ADR 0019 unit checklist (#1841) * docs(agents): merge-ready owes live evidence only for changed device-facing paths * docs(agents): the unit checklist documents the admission API on main; #1841 updates the row when it lands |
||
|
|
a70cdee360 |
refactor(ios): route snapshot backends through presentation (#1848)
## Summary Route every iOS capture-plan backend through one SnapshotPresentation boundary while preserving each backend's current output semantics. SnapshotAcquisition now carries nodes and attempt-level facts, PresentationOptions is the stable policy input, and only the presentation module assembles wire-facing nodes. Part of #1797. Touches 10 files within the existing iOS snapshot module and its architecture vocabulary; scope did not expand beyond the planned command family. ## Validation - Unit-enabled iOS runner build and focused presentation XCTest passed. - Removing the custom-action handoff made the focused test fail with exactly two assertions, proving the routing check is non-vacuous; restoring it returned to 1/1 green. - macOS runner build passed for the shared Swift path. - XCTest selection and repository formatting checks passed. |
||
|
|
f03c0309a1 |
fix: derive iOS transition snapshots from visible presentation (#1831)
* fix: project iOS transition semantics * fix: derive iOS transition semantics from visible state * fix: preserve iOS presentation context for scoped snapshots * fix: confirm broad iOS transition settlement * ci: run coordinate input regression on pull requests * test: mock migrated snapshot capture seam * fix: confirm transitions across snapshot backends * fix: arm transition confirmation after first capture * fix: settle against immutable action baseline |
||
|
|
6a8beb653e |
feat(mcp): compact server instructions in both eras + MCP-only help tool (#1839)
* feat(mcp): compact server instructions in both eras + MCP-only help tool (#1833) MCP-only clients got no workflow guidance: server/discover carried two sentences, legacy initialize carried nothing, and the CLI guides (agent-device --help, help <topic>) were unreachable over MCP. - MCP_SERVER_INSTRUCTIONS: one MCP-phrased workflow card (<2 KB, the Claude Code truncation limit) returned by server/discover and legacy initialize alike. - help tool, router-owned (not a command descriptor): no topic -> the CLI decision card; topic -> agent-device help <topic|command> text, prefixed with the one-line CLI->tool-property mapping; unknown topic -> isError listing the topics. listCommandTools() stays descriptor-only for the AI SDK; the router composes descriptors + help. - Move src/cli/parser/cli-help{,-overview}.ts to src/cli-schema/ so src/mcp (rank 3) can import the renderers without a layering back-edge into src/cli (rank 6). * fix(mcp): name terminal-only commands in help guides; colocate cli-help tests with their sources - The MCP guide preamble claimed every `agent-device <command>` line is a tool of that name; `help web` tells the reader to run `web setup` / `web doctor` and no `web` tool exists. The preamble now lists the exact CLI-only set (listCliCommandNames minus listMcpExposedCommandNames) — derived, not scanned out of prose where `device`/`web` are ordinary words. Regression: help web names `web` as terminal-only, and the listed set equals the registry difference. - cli-help-*.test.ts move from src/cli/parser/__tests__ to src/cli-schema/ to mirror the moved sources. * perf(mcp): tighten the guide card, tool description, and preamble Instructions card 1572 -> 1378 bytes (paid every session), tool description and preamble trimmed, HELP_TOOL built once as a const. Bundle delta vs main 3189 -> 2715 bytes; the remainder is the guide text itself, which the bundle carried in no MCP-phrased form before. |
||
|
|
0fb38f1da2 |
test: prune abandoned test-run tmp directories at run setup (#1834)
* test: prune abandoned test-run tmp directories at run setup A run killed before its teardown (tool-timeout SIGKILL, OOM, cancelled job) left /tmp/agent-device-test-run-<pid>-* behind, and check:tmpdir-leaks — which runs after test:unit in check:unit — flagged every dead-pid directory it found. It could not tell this run's leak from a historical one, so one killed run made every later, otherwise-green gate on the host fail. Both TMPDIR redirection entry points (the Vitest global setup and the node --test wrapper) now prune dead-pid run directories before creating their own, printing one [tmpdir] line when they did; the post-run check keeps its semantics and can now only ever name the run that just finished. Live owners (a concurrent run in another worktree) are never touched. The root/prefix constants move into check-tmpdir-leaks-model.ts, next to the liveness classification, so the setup can import the prune without a cycle. * test(tmpdir): a run directory is live while any process still holds it as TMPDIR, not only while its owner runs Review (P1): owner-pid liveness alone would prune a directory out from under the orphaned children of a SIGKILLed run — the node --test chain, Vitest forks, or a daemon a test spawned all keep running with that TMPDIR. The liveness model now reads every process's TMPDIR (ps -E on macOS, /proc/<pid>/environ on Linux) and treats a run directory as live while its owner pid is alive OR any process's TMPDIR points into it; both the prune and the post-run leak check use it. Regression: a wrapped probe spawns a detached long-lived child, only the wrapper is SIGKILLed, the next prune preserves the directory; after every consumer exits, the next prune removes it. Planted red with owner-only liveness: the orphaned directory is pruned. |
||
|
|
423927fdd8 |
chore(mutation): shrink to report-only — drop the ratchet, baseline and graduation (#1457, #1781) (#1828)
* chore(mutation): shrink the lane to report-only (#1457, #1781 wave 2) The mutation harness's two real catches (#1474, #1475) both came from humans reading the weekly score report. The ratchet half never operated: the baseline was committed exactly twice ( |
||
|
|
9d6154eecb |
ci: park perf-nightly to dispatch and stop the coverage-gate cascade double-red (#1781 A3, A5) (#1822)
A3: perf-nightly writes a report and compares nothing, so it structurally cannot catch a regression. iOS wall-clock medians swing up to +122% night-to-night at n=5 (a comparator would print noise), no doc/issue reads the report, and the iOS job holds a macOS runner ~22min nightly. Parked to workflow_dispatch following the #1781 A1 pattern (replays-manual.yml); it declares no gate-manifest check, so no declarations.ts change is needed. `pnpm perf` / scripts/perf are untouched. A5: the "Enforce changed-line coverage gate" step ran `if: always()`, so when the preceding "Run coverage" step failed, lcov.info was never written and this step failed too with "no lcov report" -- a cascade double-red, not a coverage verdict. 16 of the last 17 red instances (60d) were this cascade; the step now runs only when Run coverage succeeded. |
||
|
|
8300fa131e |
refactor(ios): establish snapshot presentation seam (#1845)
Introduce RawAXNode and PresentedNode so acquisition backends can no longer construct the wire-facing snapshot shape directly. Preserve current output while #1797 moves semantics behind the seam. Non-vacuity: setting PresentedNode.label to nil made testSnapshotPresentationPreservesCurrentWireShape execute once and fail on the missing label field; restoring the production mapping made the same focused XCTest pass. |
||
|
|
4bba404249 |
fix(layering): keep the snapshot interactor seam out of the type cycle (#1838)
#1779 added src/daemon/handlers/snapshot-interactor-capture.ts as a
vi.mock seam between snapshot-capture and core/interactors. Both of its
edges are value imports, and it sits on the path
request-generic-dispatch -> snapshot-capture -> (seam) -> core/interactors
-> register-builtins -> command-catalog -> ... -> daemon-command-registry,
so it joined the largest type-level SCC (46 -> 47 files, daemon-server
16 -> 17) and R9/R10 have failed on main since
|
||
|
|
9a0d6dead2 |
refactor: split the test-suite command out of the replay handler (#1826)
* refactor: split the test-suite command out of the replay handler handleSessionReplayCommands becomes the routing decision alone; the test suite's harness-flag admission, request translation and scheduler run move to session-test-suite-command.ts beside the replay runtime they already sit next to. Pure move: no behavior change. * test: pin the replay handler's routing decisions session-replay.ts is a router now, so it gets a focused test of its own (AGENTS.md 1:1 source/test topology): replay reaches the script-source runtime, test reaches the suite command with the whole parameter set, and an unrelated command is declined. Both destinations are mocked so a wrong edge shows up as the wrong marker; each case was proven red by mutating the routing it pins. |
||
|
|
d76e0f94e9 |
refactor: migrate snapshot to device runtime (#1779)
* refactor: migrate snapshot to device runtime * refactor: complete snapshot runtime policy cutover * test: enforce snapshot owner-facts admission * refactor: consolidate desktop snapshot capture * fix: scroll to visible iOS smoke targets * fix: close snapshot cutover alias bypasses * fix: constrain snapshot admission identity flow * fix: enforce snapshot admission through owner facts * fix: adapt replay source tests to snapshot runtime |
||
|
|
a853734f0c |
fix(webdriver): give cloud session creation its own budget and stop leaking billed sessions (#1782)
* fix(webdriver): give cloud session creation its own budget and stop leaking billed sessions Cloud lease allocation ran under the generic 30s/1-retry request policy, so BrowserStack iOS real-device session creation (45-90s) aborted client-side at ~60s on most runs. Each timed-out POST /session still completed server-side and, being non-idempotent, was retried — leaving two billed provider sessions per failed open with no id to release them. - POST /session is its own phase: a 180s create budget (default), zero retries, and no request-bound abort, so the daemon always learns the session id. - lease_allocate carries a 300s allocation budget surfaced to providers as LeaseLifecycleContext.deadline, and a matching 330s client envelope that preserves the daemon on timeout (a reset would SIGKILL mid-create and orphan every billed session the daemon held). - The request's cancellation signal is ownership evidence: a session that completes after the requester left is released, not registered; a create that the transport gives up on surfaces typed evidence (provider + lease) so an operator can find and stop the maybe-orphaned session. Closes #1774 * refactor: one canceled-request error, and tighten the #1774 shapes Review pass over the session-create fix: - The canceled-request error had nine hand-rolled copies (src/request/cancel, maestro shared, exec, retry, install-source x2, and the new provider one). It now has one definition in @agent-device/kernel/errors: createRequestCanceledError(details?, cause?) + isRequestCanceledError + REQUEST_CANCELED_REASON. Callers add evidence or a sharper hint; the reason itself is not overridable, so nothing can build one the predicate misses. - lease_allocate's timeout bundle moves beside INSTALL_TIMEOUT_POLICY in the registry (same {...DEFAULT, envelopeMs, onTimeout} shape); the request timeout constant stays exported from timeout-policy like its siblings. - Transport: fetch helper returns Response's own ok/status; the timeout reason const is private behind isWebDriverRequestTimeout. - Client: one-use options type inlined; the two deadline helpers share one floor. - Session-manager tests: shared makeRuntime/jsonResponse/afterEach restore. Net -29 lines with the feature in. * chore: keep the canceled-request reason private to the kernel * fix: typed cancellation everywhere + own the AWS remote-access ARN through startup Second-order follow-ups the #1774 refactor made cheap: - markRequestCanceled aborts the request signal WITH the kernel's typed canceled error as its reason. Every signal.throwIfAborted(), aborted fetch, and 'throw signal.reason' in the daemon (20+ sites) now surfaces a canceled request as such instead of a bare DOMException that normalized to UNKNOWN — and no site has to know the factory exists. - AWS Device Farm prepareSession owns the remote-access ARN from the moment create-remote-access-session answers: a startup timeout, the allocation deadline, or a canceled request now stops it before the failure surfaces (previously a timed-out startup left a RUNNING billed session behind — the same leak class as the WebDriver session, one phase earlier). The startup wait is capped by LeaseLifecycleContext.deadline and wakes on cancellation. - BrowserStack's pre-session local app upload honors the request signal (an upload is not billed, so plain abort is right there). - lease_heartbeat/lease_release share lease_allocate's preserve-daemon policy: the rationale — the daemon owns billed sessions; a reset orphans them all — applies verbatim. Each AWS ownership test proven red without the guard (3/3). * refactor: dedupe billed-resource cleanup and lease-signal wiring Shrink pass — same behavior, less duplication: - releaseOnFailure(primaryError, release) in webdriver-utils replaces the two identical 'best-effort stop the billed resource, attach cleanupError to the primary AppError' helpers (WebDriver session + AWS remote-access ARN); shared errorMessage too. - The lease handler pulls the request signal from getRequestSignal(requestId) like every sibling handler, instead of threading a requestSignal arg through LeaseHandlerArgs and the request-handler chain. Drops the field, the wiring, and five mechanical test edits; the handler test now proves the request-bound signal (abort it, watch the provider's signal flip) rather than arg identity. - Inlined the one-use requestHeaders back into fetchWebDriver. Handler-signal test proven red without the wiring. * fix(lease): the daemon releases a lease allocated for a gone requester; honest release evidence Review follow-up. The provider was doing the daemon's job: it treated the request signal as 'ownership evidence, not an interrupt' and needed three paragraphs to say so. The daemon owns the request, so it now decides — generically, for every provider — what happens to a lease that finished allocating after its requester left: release it (provider + registry) and answer with the canceled error. - lease.ts: after allocate returns, isRequestCanceled(requestId) → releaseAllocationForGoneRequester(). Release evidence is claimed ONLY on a clean release (no warnings, no throw); a WEBDRIVER_SESSION_DELETE_FAILED release is reported released:false with providerSessionId + a stop-by-hand hint (thymikee's finding: the previous evidence was success-shaped even when DELETE failed). - WebDriverSessionManager: the createOwnedSession/releaseCanceledSession trio is gone; allocate is plain 'create with a budget; on failure clean up' again. - LeaseLifecycleContext.signal is just cancellation, like everywhere else; the ownership-semantics comments on the contract, client, registry, AWS prepare and utils shrink to what the code no longer says itself. - Tests: the two provider-level cancellation tests move to the daemon handler (where the logic now lives), plus the failing-DELETE regression; both proven red without the post-allocate check. * fix(aws): the allocation deadline bounds remote-access startup, not the 120s default Live iOS real-device run: startup needed ~128s and hit the standalone 120s default while the daemon's 300s allocation budget still had room — the new ownership guard correctly stopped the ARN, but the open failed for no reason. When the daemon supplies a deadline it is the bound; the default only applies standalone. Rerun: open in 112s, snapshot, clean close, session STOPPING. * test(aws): pin that the allocation deadline outlives the 120s startup default; drop empty import Review follow-ups on 7f9d1481a: a virtual-clock test (Date.now advanced 10s per poll, RUNNING at 150s, deadline 300s) that fails on the old min(default, deadline) logic and passes now; and the empty 'import {} from kernel/errors' left in maestro/shared.ts is removed. * refactor: finish the dedupe — one release path, kernel errorMessage, AWS on releaseOnFailure Code-quality review at 7f9d1481a: 1. aws-device-farm.ts still carried its own copy of releaseOnFailure (the dedupe commit's script aborted before reaching it and I mis-verified). Now uses the shared helper; private copy deleted. 2. Empty 'import {} from kernel/errors' in maestro/shared.ts removed (273870099). 3. errorMessage() lives in @agent-device/kernel/errors; the two copies this PR had added (lease.ts, webdriver-utils.ts) import it. Sweeping the pre-existing copies is a follow-up. 4. lease.ts has ONE release path: releaseLease(registry, provider, lease, request, ctx) → { released (registry), provider } used by both the lease_release case (wire shape unchanged) and the gone-requester branch, which folds a throwing provider release into releaseError. 'released' now means the same thing in both; the provider verdict is a separate 'providerReleased' (warnings-free, no throw) that drives the stop-by-hand hint. -~35 lines. 5. sessionCreateTimeoutMs is Omit-ed at the WebDriverTransportOptions boundary instead of Pick-ed back out internally. * fix(lease): 'released' on a canceled allocation means the billed session is confirmed gone Re-review at 3665ea06: unifying the release path had made the cancellation error report released:true from the daemon's registry record while the provider DELETE had failed — success-shaped again, with the operator verdict demoted to a second key. Fixed at the source of the ambiguity: - LeaseReleaseOutcome names its bookkeeping field registryReleased. - On the canceled error, 'released' is true only when registryReleased AND the provider released without warnings AND without throwing; the registry record is exposed as 'registryReleased'. The stop-by-hand hint keys on 'released'. - lease_release keeps its existing wire field ('released' = registry; provider cleanup rides in 'provider'), unchanged. - Regressions: failed DELETE and throwing release both pin released:false / registryReleased:true (+ providerSessionId, warnings|releaseError, hint); both proven red on registry-only semantics. * ci: retrigger default-setup CodeQL Run 32051017472 is wedged on GitHub's side: status=completed with Analyze (python) still queued and Analyze (java-kotlin) failed only at SARIF upload (503, 'No server is currently available'). It can be neither cancelled nor rerun, and default-setup CodeQL has no dispatchable workflow, so a new push is the only way to get a fresh run. No source change. * test(webdriver): assert the typed timeout contract on the shared-budget probe main's #1790 tightened this test to expect the raw TimeoutError DOMException, which this PR intentionally normalizes into AppError{reason: webdriver_request_timeout}. On the merge ref the two met and Coverage went red. The regression now asserts the structured contract and that the second request's budget is the shared remainder (~118ms of 200 after an 80ms first call). |
||
|
|
d0d5c8594c | fix: serve remote daemon request diagnostics to the caller (#1801) (#1814) | ||
|
|
ef6ec2995b |
chore(layering): document R12/R18/R19, retire R8, make R9 shrink mandatory (#1781 A6) (#1825)
* chore(layering): document R12/R18/R19, retire R8, make R9 shrink mandatory (#1781 A6) The A6 review kept `check:layering` in full (15/15 planted violations fired, no other enforcer exists) and left four follow-throughs. R12 bin-alias-fast-path, R18 contracts-implementation-authority and R19 selector-pipeline-ownership were live rules with no ADR or CONTEXT anchor — they now carry one each, in the same list as R7/R9/R10/R13. R8 zero-dep-job-closure is retired: no CI job sets `install-deps: false` and ci.yml records why each keeps it enabled, so the invariant has no subjects. R11's relative-into-packages exception existed only because a zero-dep closure cannot coexist with specifier loads, so it retires with R8; the route is now closed to every caller. R1 was retired the same way at #1490. R9 was growth-only and merely suggested lowering the ceiling, which is headroom the next change spends without a number moving. It is now an equality pin like R6 and the R10 R7 counts, and the committed baseline drops 47 -> 46 (daemon-server ceiling 17 -> 16) to match the measurement. ADR 0019 §6 now says each runtime-command-cutover row is deleted when that command's migration is declared closed. * chore(layering): rename R9 to type-cycle-size now that it fails both ways (#1781 A6) |
||
|
|
4b44c1c53a |
chore(test): remove the contention retry and shrink the subprocess-stub project (#1781 A4) (#1827)
The enumerated single-retry policy (#1419) has fired zero times since it landed on 2026-07-29: 0 of 234 sampled Coverage-job lane envelopes (2026-08-11 to 2026-08-18) have retryCount > 0, and none of 17 recent failed runs was retried (5 refused "outside the enumerated retry list", 4 refused "unhandled error"). All three trackers its entries pointed at (#1098, #1414, #1419) are closed. It cost ~1,454 LOC, a per-run secret marker threaded through a setup file on every Vitest project, and a standing obligation for every future gate reporter to call the blocker bus. Delete the scripts, tests and fixtures, the check:contention-retry script and gate, the envelope artifact upload, and the runner-timeout setup file; test:coverage:ci is a plain `vitest run --coverage` again. lane-envelope.ts stays: the mutation, fuzz and concurrency-torture lanes build their envelopes from it. run-blocker-bus.ts goes: its only consumer was the retry's failure sink, and its only publisher already fails the run by setting process.exitCode. Keep the subprocess-stub project for the three files that really spawn (client-metro, fuzz harness, fuzz corpus-replay) and drop the three that run in 31/212/277ms in CI, which cannot contend for anything. The list is now a plain array in vitest.config.ts with the reason at each entry. Membership and the project's kill criterion live in #1823. Because test:coverage:ci is a bare vitest run, the gate manifest reads its projects directly, so OPAQUE_RUNNERS no longer needs it and an unrun Vitest project becomes unrepresentable rather than detected; the audit test now constructs that state by project-scoping the script. |
||
|
|
60f6356b04 |
fix: read replay scripts on the caller and ship them with the request (#1810)
* fix: read replay scripts on the caller and ship them with the request Closes #1802 * test: assert the caller-side replay path as a substring, not a hand-escaped regex * perf(cli): load the Maestro engine only when a replay entry is a flow The command registry evaluates every command family on CLI startup, so the replay script-source builder's static @agent-device/maestro import put the YAML parser on the --help path. It now loads on demand behind the format check, and the startup import-closure guard covers the engine the way it already covers node:http. * refactor: share the replay request field vocabulary across the CLI and client views The new replay script-source flags appear in both CliFlags and CommandExecutionOptions, which fallow flagged as a clone; ReplayRequestFields declares them once. The test-suite handler's missing-sources rejection now travels the typed-error path its sibling rejections already use, so the fix adds no branch to handleSessionReplayCommands. |
||
|
|
3908559fe2 |
fix: report real claim results from daemon stop (#1818)
* fix: report real claim results from daemon stop `daemon stop` typed `claimsReleased`/`claimsOrphaned` as the literal `[]` and every path hardcoded them, so a graceful stop that released a device claim still reported none (#1799 observation 3, #1320 acceptance). Graceful teardown now records each session's claim outcome — released after a clean teardown, orphaned when teardown left the claim in place — into the daemon shutdown report, and the CLI merges them alongside provider releases. Forced and not-running stops stay empty because they cannot know, and a report written before claim reporting still reads its provider releases. * fix: classify daemon stop claim results from the clear outcome `clearDeviceClaim` deliberately resolves without deleting when the on-disk claim is no longer the one it acquired, so the shutdown ledger's "the call resolved" test reported a successor's claim as released — a device the daemon never freed, counted as freed. `clearDeviceClaim` now returns a typed outcome (`deleted` | `absent` | `ownership-changed`) instead of nothing, and the ledger classifies from it: released only when absence is confirmed, and a new `superseded` bucket for a claim another owner had already taken over. Superseded is neither released (this daemon freed nothing) nor orphaned (no claim of ours remains to reconcile), so folding it into either would break that list's meaning; it also raises a warning so a device now owned elsewhere cannot pass silently. |
||
|
|
8d0de32ba4 |
fix(android): let a covering sibling hide only what its content covers (#1808)
* fix(android): let a sibling hide only what its content covers (#1806) pruneAndroidCoveredSubtrees credited a higher drawing-order sibling with painting its whole box as soon as it had any content anywhere inside it, or a label of its own. A full-screen DoraemonKit drag surface holding one 189px floating icon therefore condemned the entire app subtree, and an empty labelled match_parent placeholder did the same. Occlusion is now spatial. A subtree's footprint is the bounding box of what it presents (agent targets and labelled leaves); a sibling is covered when its footprint lies under a candidate's footprint. A node's own label is no longer paint evidence: a container's content-desc describes its children and an empty labelled View draws nothing. Only a touch target still hides its full box (scrims). Comparing footprint to footprint keeps stacked screens with matching margins registering as covered. Live on a Pixel 9 Pro XL API 37 emulator with a DoKit-shaped overlay added to the test app: snapshot -i went from 2 nodes + the sparse hint to the full app; helper-XML A/B across home/catalog/form/product-detail recovered every label with none lost, and non-overlay screens are byte-identical. * fix(android): measure occlusion by overlapped area, not bounding box Review on #1808: a bounding box of two corner controls spans the viewport, so a transparent overlay with a control in each corner still acquired a full-screen footprint and could prune the app beneath it. Footprints now keep their presented rects apart, and coverage is the overlapped area of the two unions (coordinate-compressed cell sweep). Scrollables count as presenting their box: they consume touches over it, which is what lets a real pushed screen (header, scrollable body, footer) still cover a drawer surface. Adds the disconnected-corner regression. * fix(android): count what a covered sibling shows, not only what it paints Fuzzing random sibling trees old-vs-new surfaced the one direction the footprint model could still regress: a container whose only painted content is small (one corner icon) but which also carries labelled containers or testID-only markers was condemned as soon as a touch surface covered that icon, since markers and container labels are not paint and never entered the footprint. Footprints now carry two rect sets. `paints` (touch targets, scrollables, labelled leaves) is what a candidate can cover with; it still excludes identifiers and container labels, or the DoKit fix would unwind. `shows` adds every labelled or identified node and is what a covered sibling must lose in full. Focusable-only nodes no longer paint their box either, matching #1733 for descendants as well as siblings. Adds the marker regression. Re-fuzzed 20k trees: new-prunes-more is down to 0.14 %, all of the class where everything the target shows lies under a higher touch/scroll surface. Live captures unchanged. * test(android): pin that focusability never paints a covering candidate's box A full-screen focusable wrapper holding one clickable icon is a covering candidate; the lower app content must survive. Fails when paintsOwnBox counts focus targets again. |
||
|
|
f843dc2df1 |
fix(scroll): keep saturated scroll gestures out of the status bar; gate Android replays from android/emulator (#1781 A1) (#1820)
* fix(scroll): keep saturated scroll gestures out of the status bar; gate Android replays from android/emulator (#1781 A1) `pnpm gate replay-android` failed 4/8 whenever it ran after the full-tier Android E2E (replays-nightly run 32107665052, job 95620294899): 05-app-lifecycle, 06-swipe-gestures and both fixture replays diverged under "A system surface covers the app". The E2E was not the cause. Reproduced on a pixel_7 / API 36 AVD with the same cutout geometry CI's `avdmanager --device pixel_7` produces (status bar 136px, not the 63px of a plain 1080x2400 skin): - `03-scroll-discovery.ad` runs `scroll up 3`. The scroll planner clamps travel to the viewport minus a 5% band, so the touch-down landed at y=120 — inside the 136px status bar — and pulled the notification shade instead of scrolling. On API 36 the app window is edge-to-edge, so the reported viewport starts at y=0 and includes that bar. - The shade then covered every replay until `04`'s `back` closed it. Native readdir order on the runner (03, 05, 06, fixture/02, fixture/01, 04, 01, 02) put four files in that window; the last green run (2026-07-30) had 04 right after 03, so the pull was masked. Fix in the product, not the lane: DEFAULT_EDGE_PADDING_FRACTION 0.05 -> 0.1 in the TS scroll planner and its Swift port. Every real Pixel has a cutout (5.7% of a Pixel 7's height) and an iPhone's Dynamic Island status bar is 6.9%, so any saturated `scroll up` opened the shade / Notification Center for real agents too. Parity vectors updated in both suites plus a Pixel 7 regression vector (1080x2400, amount 3 -> touch-down y=240 > 136). Second contamination the same order exposed once the shade was gone: `fixture/02-selector-routes-covered-diagnosis.ad` is a #1715 reproduction recipe that FAILS BY DESIGN at step 9 (covered-target refusal) and leaves the device in landscape, yet the gate enumerated `test/integration/replays/android` recursively. iOS keeps gate replays in `replays/ios/simulator` and fixture recipes in `replays/ios/fixture`; Android now mirrors that: the six Settings replays move to `replays/android/emulator`, `test:replay:android` points there, and `fixture/` stays E2E-owned (`full:fixture-replays` already runs 01 by path). android.yml and the workflow-evidence fixture follow the path; the replay-compat manifest keeps the historical paths it pins at released tags. Verified live (Pixel 7 geometry, API 36, --retries 0): control run at main head in CI order reproduces exactly CI's 4/8; with the fix, `pnpm gate replay-android` 6/6 in both native and CI order, and `03` leaves Settings on screen (scroll up 3 now touches down at y=240). * test(scroll): drive the TS and Swift scroll-plan parity vectors from one table (#1820 review) The two suites hand-mirrored the same vectors and #1820 had to edit both by hand — the drift class the repo already closes for the tap-point rule via contracts/fixtures/tap-point-policy.json. The scroll vectors (plus both planner constants, pinned behaviourally on a 1000px axis) now live in contracts/fixtures/scroll-gesture.json; scroll-gesture.test.ts and RunnerTests+ScrollGesture.swift iterate it. Verified: vitest 10/10; the four XCTests run on an iOS 26.2 simulator with the unit flag on (Executed 4 tests, 0 failures). Also: test/ci/android-workflow-evidence.json says what it guards. Follow-up for content-safe viewport bounds + discovery order: #1821. |
||
|
|
2b6d04a13e |
fix: enforce device claims for sessionless device mutations (#1809)
`boot` and `shutdown` never consulted the host-global device claim store, so a daemon in one state directory could terminate an emulator another daemon held a verified-live claim on and report success (#1799). Rather than adding a claim check to those two handlers, this makes the class unrepresentable: `CommandDescriptor` gains a REQUIRED `deviceClaimPolicy` trait (#1320's vocabulary), and the request-execution scope enforces it where the request runtime bindings create a device binding — the one seam through which any handler can obtain device operations, and already the place per-device deduplication lives. A `transient-exclusive` command acquires a command-scoped claim before operations reach the handler, refuses a foreign live claim with the existing DEVICE_IN_USE/DEVICE_CLAIM_LIVE_OWNER error, and releases in the scope's finally. Every other policy performs no claim-store I/O, so session-bound commands keep #1320's non-goal intact. |
||
|
|
4c5f693a03 | fix: declare frameworkTier on the hover descriptor (main parity gate red) (#1817) | ||
|
|
1234590a69 |
fix(ios-runner): reject CGRect.infinite in navigation frame guards (#1816)
topLeadingNavigationFallbackPoint and isTopNavigationControlFrame only validated width/height, never the origin. CGRect.infinite has a finite (if absurd, ~-9e307) origin, so it slipped past the isFinite/>0 guard: the fallback point collapsed to roughly (-9e307, -9e307), and an unresolved element frame was classified as sitting in the nav header band. Extract isUsableNavigationFrame(_:), which adds the isInfinite check TapPointPolicy.isAllowed already uses for the same rect, and share it between both helpers. Fixes #1812 |
||
|
|
142d156338 |
ci(ios): run the full XCTest suite nightly and check the PR test list (#1781 A7) (#1789)
* ci(ios): run the full XCTest suite nightly and check the PR test list (#1781 A7) * fix(ci): skip the runner server entry point in the nightly and validate both test flags * docs(ci): restate the nightly lane cost and timeout honestly * docs(ci): stop quoting XCTest counts that drift between commits * ci(ios): tighten the nightly timeout to the measured suite duration |
||
|
|
0f4f322878 |
fix: reject selector-shaped wait arguments instead of reading them as text (#1813)
`wait <condition> '<selector>' [timeoutMs]` (e.g. `wait open 'label="Open"' 25000`, `wait exists 'label="x"' 100`) and any unrecognized `key=value` token used to fall through parseWaitPositionals' text fallback and wait out the full timeout for literal text that could never appear on screen — reading as a false "element absent" instead of the caller's own argument mistake (#1035 is the sibling fix for click/press/fill/get). parseWaitPositionals now returns a typed `invalid` variant whenever a positional token is selector-shaped (a recognized key, or an unrecognized key=value) but the list doesn't form a valid selector expression, or a valid selector prefix is followed by unquoted trailing tokens. The message names the offending token, points condition words (exists/present/appears/gone/disappears) at the selector form, and always offers the explicit `wait text '<text>'` escape hatch. Bare text (single- and multi-word) and the explicit `text` keyword form are unaffected. Excluding `invalid` from the type consumed by selector-runtime's toWaitTarget makes the remaining kind-by-kind narrowing exhaustive without a runtime fallback branch. |
||
|
|
801734d433 |
feat(ai-sdk): add agent-device/ai-sdk tool set and document the MCP zero-code path (#1804)
* feat(ai-sdk): add agent-device/ai-sdk tool set and document the MCP zero-code path
Adds `createAgentDeviceTools()` under a new `agent-device/ai-sdk` subpath,
built from the same command registry the MCP server uses so both stay in
lockstep without a hand-maintained tool list. Introduces a `frameworkTier`
descriptor facet ('core' | 'extended') so the factory can default to a
curated perceive/act loop instead of handing a model dozens of tools.
`ai` is wired as an optional peer dependency, imported lazily inside the
factory rather than at module scope, so importing the subpath itself never
requires `ai` to be installed - only calling it does. The package's own
publishing gate (scripts/lib/shipped-imports.ts) is extended to recognize
peerDependencies as a valid resolution source, since this is the first
optional peer this package has shipped.
Also restructures the AI SDK doc around three tiers (zero-code via
@ai-sdk/mcp, the new typed tool set, hand-written tools) and fixes a stale
`needsApproval` reference in favor of the current `toolApproval` API.
* fix(layering): classify src/ai-sdk as a rank-4 zone
The layering guard requires every src/<folder>/ to be explicitly ranked or
unranked; the new src/ai-sdk/ subpath (added in the prior commit) was left
unclassified, failing CI's Layering Guard job. It sits at the same tier as
client/compat/daemon-server/metro/remote/sdk - a public integration surface
consuming mcp (3) and core (2), imported by nothing else in the tree.
* fix(ci): cover, exempt, and pack the new ai-sdk subpath
Fixes the remaining CI failures on the ai-sdk subpath commit:
- Coverage: src/ai-sdk/index.ts had no dedicated unit test (only manual/
integration verification), so changed-line coverage sat at 6.9% against
the 70% gate. Adds src/ai-sdk/__tests__/index.test.ts (core vs 'all' tool
filtering, session/platform pinning and schema hiding, error
normalization, toolApproval passthrough) with createCommandToolExecutor
and createAgentDeviceClient mocked the same way command-tools.test.ts
does, plus a dedicated missing-peer-dependency.test.ts that mocks `ai`
itself to throw, isolated to its own file so it doesn't affect the other
tests' use of the real, installed `ai` package. Changed-line coverage is
now 29/29 (100%).
- Fallow Code Quality: src/ai-sdk/index.ts and examples/sdk/ai-sdk-tools.ts
are entry points with no in-repo importer (reached only via package.json
exports / run directly), and the new subpath's exports are unused
internally by design - both need the same treatment src/sdk/*.ts and its
examples already have in .fallowrc.json.
- Integration Tests: test/integration/installed-package-metro.test.ts and
src/__tests__/package-exports.test.ts each hand-list every published
subpath and smoke-check it from a real packed install; added ./ai-sdk to
both so the new subpath is actually exercised, not just silently passing.
* fix(ai-sdk): hide MCP transport/config fields from the model too
createAgentDeviceTools() only removed session and mcpOutputFormat from tool
schemas. stateDir was still model-visible and reached the shared executor
as client configuration, letting a tool call redirect into a different
daemon state directory - defeating the "one pinned session" guarantee the
factory exists to provide. includeCost and responseLevel are MCP
tool-config knobs in the same category, irrelevant to this adapter.
Widens the hidden-field set to session/stateDir/mcpOutputFormat/
includeCost/responseLevel, and now strips them from the runtime input
inside execute() too (not just the schema), so the guarantee holds even if
a caller bypasses schema validation. The schema-properties filter and the
input filter now share one omitHidden() helper instead of two near-
duplicate implementations.
Addresses the P1 review comment on #1804.
|
||
|
|
20458cbcf6 |
fix(daemon): reject '.', '..' and empty session names at resolveSessionDir (#1815)
safeSessionName only rewrites characters outside [a-zA-Z0-9._-], so the names '.' and '..' survive unchanged and path.join resolves them to the sessions dir itself or its parent, the daemon state dir. A remote caller's --session .. would then land app.log / runner.log / requests/*.ndjson outside the sessions tree. SessionStore.resolveSessionDir is the one place a session name becomes a directory (AGENTS.md: session artifact paths come from session-store), so it now refuses such a name with INVALID_ARGS. Every request goes through it first thing in createRequestExecutionScope, before any artifact path is used, so this is also the admission-time rejection; every other caller passes an already admitted name. isSafeSessionSegment mirrors the predicate PR #1814 adds for its request-diagnostics route; whichever lands second takes the trivial merge. Regression tests were proven red against the pre-fix code: resolveSessionDir returned the sessions dir / state dir for '.', '..', '' and the request scope resolved runnerLogPath to <stateDir>/runner.log. |