Commit Graph

161 Commits

Author SHA1 Message Date
Michał Pierzchała 8c06965d28 fix(daemon): cap events.ndjson with cursor-safe rotation (#1867)
* fix(daemon): cap events.ndjson with cursor-safe rotation

Rotate events.ndjson to events.ndjson.1 once it reaches
AGENT_DEVICE_EVENT_LOG_MAX_BYTES (default 5 MB), keeping one rotated
generation. Cursors stay absolute across rotation through a sidecar
window offset, so a persisted nextCursor still names the same event; a
cursor older than the retained window fails with COMMAND_FAILED and
details.reason EVENT_LOG_CURSOR_EXPIRED instead of returning a wrong
page.

Closes #1788

* fix(daemon): verify the events.ndjson window against the files on disk

Rotation recorded only a dropped-line offset, written after the rename,
so a reader landing in that window mapped every absolute cursor a whole
generation too far (reproduced: 5 of 58 reads returned event 17 for
cursor 9), and a missing rotated file or stale sidecar shifted cursors
permanently and silently.

The sidecar now records each retained generation's first absolute line
index, line count, and first-line digest, and is written before the
rename it describes. The reader identifies each file on disk by digest,
derives its start from the matching record, and checks the recorded line
count and generation contiguity; anything unverifiable raises a typed
EVENT_LOG_WINDOW_UNVERIFIED instead of a guessed offset. A torn snapshot
(rotation landing mid-read from the threadpool) is retried, not
interpreted. A corrupt sidecar fails reads typed and never blocks
appends, and rotation no longer does synchronous whole-file I/O.

* refactor(daemon): split event-log window placement and share one line splitter
2026-08-19 12:53:51 +02:00
Michał Pierzchała f3d5b3d92c refactor(daemon): admit-before-bind as an admitted-plan token; retire the R32 syntax policy (#1841)
* refactor(daemon): admit-before-bind as an identity-keyed admitted-plan token; retire the R32 syntax policy

admitRuntimePlan (was inspectRequiredRuntimeUse) takes the plan and, on
success, mints an AdmittedRuntimePlan: a nominal class instance with nothing
readable on it. Its payload — a frozen copy of the device the facts were read
for, and the plan — lives in a module-private WeakMap keyed by the token's
exact identity, and the only way to read it is unwrapAdmittedRuntimePlan,
which refuses anything not minted here. The snapshot owning interface
(resolveBoundSnapshotCaptureRuntime, #1847) admits through it and its private
binder takes only the token: no bare plan, no separate device, and no
look-alike — a spread lacks the #private member (not assignable), a Proxy
around a real token types as the token but is a different identity (refused
at unwrap), Object.assign/defineProperty throw on the frozen instance, and the
class value is not exported so its constructor is not nameable.

That retires scripts/layering/runtime-command-cutover-snapshot.ts — R32's
per-command AST policy (call-shape recognition of the admission and a text
sniff for a local admission) — and the source-regex test beside the descriptor
tests. The generic row keeps retirement, narrowing, and singular execution;
the manufactured-proof column now also rejects casts to AdmittedRuntimePlan.

Planted reds: token degraded to a plain public shape → 2 unused
@ts-expect-error directives; unwrap reading the token surface via getters →
the Proxy regression fails; getter-based branded literal → the runtime
retarget test fails.

* docs(agents): the ADR 0019 unit checklist teaches the shipped admission API

#1836 documented inspectRequiredRuntimeUse with a forward note pointing here;
this PR makes admitRuntimePlan real, so the row now teaches it plus the
identity-keyed unwrap the binder uses, and points at the shared snapshot/diff
owning interface as the model.
2026-08-19 10:46:25 +02:00
Michał Pierzchała b12a3e3cb3 test: pin test files over 1,000 lines at their exact length so they can only shrink (#1843)
* test: pin test files over 1,000 lines at their exact length so they can only shrink

AGENTS.md has said for a while that past 1,000 lines is architecture debt and
tests are not exempt; nothing enforced it, and the second-largest test file
gained 55 lines in the PR before this one. This is the slow-test ratchet's
shape for a reader's context instead of wall clock: the 26 test files over the
tripwire are pinned at their exact length (R9-style equality pin, #1781 A6);
growth fails, shrink fails until the pin is lowered in the same PR, a file
that drops under the line leaves the list, and a new file may not cross it.
One directory walk per unit run, ~250ms; the pin list emptying deletes it.

* test(ratchet): hold giant test files to their merge-base length so pin edits cannot admit growth

Review (P1): the equality pin compared measured lengths only against the pin
map in the same checkout, so growing a file and raising its pin, or adding a
new >1,000-line file with a pin, stayed green. The gate is now history-backed:
every test file over the tripwire may be no longer than at the merge-base with
origin/main (renames followed; new files may not cross the line), and no pin
may exceed its file's base length — one git cat-file --batch spawn, parsed by
bytes because the sizes are bytes. Both bypasses planted red against real git
on a pinned file and on a fresh 1,001-line file with a pin added.

* test(ratchet): a pin on a file at or under the tripwire is itself a finding

Review: a new pin for an unchanged sub-tripwire file (900 pinned at 900)
passed equality and history and grew the map. Pins now exist only for files
over the tripwire — any other pin is red with 'remove it' — which also
subsumes the old shrink-under-the-line message. Planted red in-file and
against real git (a 186-line test pinned at 186). The android snapshot test
pin bootstraps 1636→1660: main grew that file in #1846 before this gate
exists, and history agrees (1660 at the merge-base).
2026-08-18 19:40:34 +02:00
Michał Pierzchała ee13203a16 feat(ios): unify snapshot eligibility (#1850)
Make iOS regular snapshot eligibility one backend-neutral presentation rule.

Acquire tree nodes conservatively, preserve interactive scroll containers, normalize surviving hierarchy, and keep raw membership plus daemon publication policy unchanged. Part of #1797.

- iOS and macOS unit-enabled runner builds
- 2 focused XCTest cases
- 3 production-path publication tests
- live Settings snapshots: 73 regular nodes and 167 raw nodes, both healthy tree captures
2026-08-18 18:57:29 +02:00
Michał Pierzchała 72d421fe36 docs(agents): ADR 0019 unit checklist, owning-seam mock rule, worktree and rebase guidance (#1836)
* docs(agents): ADR 0019 unit checklist, owning-seam mock rule, worktree and rebase guidance

Retro follow-up (item 2). Adds docs/agents/adr-0019-unit.md — the order of
operations for one command unit with the declaration site for each step, the
evidence a unit review must carry, and what 'done' is not — so the pattern
rediscovered during the snapshot unit (#1779) is written down once.

testing.md: mock the seam the code under test consumes (fake inspectFacts /
bindDevice), not the generic dispatchCommand mock; a migrating command moves
its tests off the dispatch mock in the same PR.

AGENTS.md: fresh-worktree preflight (pnpm install + build in the worktree;
layering scan reads tracked files only) and concurrent-agent hygiene (one full
gate per host, verify subagent edits with git -C, one PR per worktree).

pull-requests.md: two readiness claims (published-and-reported vs merge-ready)
and the rebase rule — main has no up-to-date protection; rebase on conflict or
when `check:affected --base <merge-base> --head origin/main` names your surface.

* docs(agents): name the admitted-plan token in the ADR 0019 unit checklist (#1841)

* docs(agents): merge-ready owes live evidence only for changed device-facing paths

* docs(agents): the unit checklist documents the admission API on main; #1841 updates the row when it lands
2026-08-18 18:43:27 +02:00
Michał Pierzchała a70cdee360 refactor(ios): route snapshot backends through presentation (#1848)
## Summary

Route every iOS capture-plan backend through one SnapshotPresentation boundary while preserving each backend's current output semantics.

SnapshotAcquisition now carries nodes and attempt-level facts, PresentationOptions is the stable policy input, and only the presentation module assembles wire-facing nodes. Part of #1797.

Touches 10 files within the existing iOS snapshot module and its architecture vocabulary; scope did not expand beyond the planned command family.

## Validation

- Unit-enabled iOS runner build and focused presentation XCTest passed.
- Removing the custom-action handoff made the focused test fail with exactly two assertions, proving the routing check is non-vacuous; restoring it returned to 1/1 green.
- macOS runner build passed for the shared Swift path.
- XCTest selection and repository formatting checks passed.
2026-08-18 18:24:44 +02:00
Michał Pierzchała 6a8beb653e feat(mcp): compact server instructions in both eras + MCP-only help tool (#1839)
* feat(mcp): compact server instructions in both eras + MCP-only help tool (#1833)

MCP-only clients got no workflow guidance: server/discover carried two
sentences, legacy initialize carried nothing, and the CLI guides
(agent-device --help, help <topic>) were unreachable over MCP.

- MCP_SERVER_INSTRUCTIONS: one MCP-phrased workflow card (<2 KB, the
  Claude Code truncation limit) returned by server/discover and legacy
  initialize alike.
- help tool, router-owned (not a command descriptor): no topic -> the CLI
  decision card; topic -> agent-device help <topic|command> text, prefixed
  with the one-line CLI->tool-property mapping; unknown topic -> isError
  listing the topics. listCommandTools() stays descriptor-only for the AI
  SDK; the router composes descriptors + help.
- Move src/cli/parser/cli-help{,-overview}.ts to src/cli-schema/ so
  src/mcp (rank 3) can import the renderers without a layering back-edge
  into src/cli (rank 6).

* fix(mcp): name terminal-only commands in help guides; colocate cli-help tests with their sources

- The MCP guide preamble claimed every `agent-device <command>` line is a
  tool of that name; `help web` tells the reader to run `web setup` /
  `web doctor` and no `web` tool exists. The preamble now lists the exact
  CLI-only set (listCliCommandNames minus listMcpExposedCommandNames) —
  derived, not scanned out of prose where `device`/`web` are ordinary
  words. Regression: help web names `web` as terminal-only, and the listed
  set equals the registry difference.
- cli-help-*.test.ts move from src/cli/parser/__tests__ to src/cli-schema/
  to mirror the moved sources.

* perf(mcp): tighten the guide card, tool description, and preamble

Instructions card 1572 -> 1378 bytes (paid every session), tool
description and preamble trimmed, HELP_TOOL built once as a const.
Bundle delta vs main 3189 -> 2715 bytes; the remainder is the guide text
itself, which the bundle carried in no MCP-phrased form before.
2026-08-18 17:48:36 +02:00
Michał Pierzchała 0fb38f1da2 test: prune abandoned test-run tmp directories at run setup (#1834)
* test: prune abandoned test-run tmp directories at run setup

A run killed before its teardown (tool-timeout SIGKILL, OOM, cancelled job)
left /tmp/agent-device-test-run-<pid>-* behind, and check:tmpdir-leaks — which
runs after test:unit in check:unit — flagged every dead-pid directory it
found. It could not tell this run's leak from a historical one, so one killed
run made every later, otherwise-green gate on the host fail.

Both TMPDIR redirection entry points (the Vitest global setup and the
node --test wrapper) now prune dead-pid run directories before creating their
own, printing one [tmpdir] line when they did; the post-run check keeps its
semantics and can now only ever name the run that just finished. Live owners
(a concurrent run in another worktree) are never touched.

The root/prefix constants move into check-tmpdir-leaks-model.ts, next to the
liveness classification, so the setup can import the prune without a cycle.

* test(tmpdir): a run directory is live while any process still holds it as TMPDIR, not only while its owner runs

Review (P1): owner-pid liveness alone would prune a directory out from under
the orphaned children of a SIGKILLed run — the node --test chain, Vitest forks,
or a daemon a test spawned all keep running with that TMPDIR. The liveness
model now reads every process's TMPDIR (ps -E on macOS, /proc/<pid>/environ
on Linux) and treats a run directory as live while its owner pid is alive OR
any process's TMPDIR points into it; both the prune and the post-run leak
check use it. Regression: a wrapped probe spawns a detached long-lived child,
only the wrapper is SIGKILLed, the next prune preserves the directory; after
every consumer exits, the next prune removes it. Planted red with owner-only
liveness: the orphaned directory is pruned.
2026-08-18 17:48:12 +02:00
Michał Pierzchała 423927fdd8 chore(mutation): shrink to report-only — drop the ratchet, baseline and graduation (#1457, #1781) (#1828)
* chore(mutation): shrink the lane to report-only (#1457, #1781 wave 2)

The mutation harness's two real catches (#1474, #1475) both came from humans
reading the weekly score report. The ratchet half never operated: the baseline
was committed exactly twice (8cce0ef6b, 60400d04b), both times with
`stableRuns: 0, gating: false`, and was never updated after the very fixes it
triggered — the weekly job computed a new baseline and then `git checkout --`d
it, uploading a proposal nobody applied in 3+ weeks. A gate nobody arms is
harness weight; the report is the part that paid.

Deletes ratchet.ts + ratchet.test.ts, mutation-baselines/, and every
baseline/graduation/gating path in run.ts (`--update`, `mutation:baseline`).
run.ts now exits non-zero only on a harness failure, never on a score. The
report renders the per-kernel table (kernel, score, killed, survived, total,
timeouts) plus the surviving mutants a strengthening PR works from.

Kernel scoping stays: stryker.config.json and KERNEL_MODULES are untouched.

* fix(mutation): restore denominator coverage and publish the table before judging the shard set

Review of #1828:
- `report.test.ts` re-asserts that Ignored/CompileError/RuntimeError leave the
  denominator — the one behaviour `ratchet.test.ts` covered and nothing replaced.
  A `tally()` edit that counted tool noise would have deflated every published
  score with a green `mutation:test`.
- `assertShardsCoverModules` now runs after `emit()`, so an incomplete shard set
  still publishes the kernels that completed instead of only an error string.
  This makes the workflow comments' claim about the job summary true rather than
  re-wording them down.

* chore(mutation): trigger the affected lane on exactly the paths that can select mutants

The PR lane returns an empty matrix unless the diff touches the harness, so the
kernel-source and `**/*.test.ts` triggers only bought a 1-4 min no-op job on
~96% of PRs. `on.pull_request.paths` is now exactly `LANE_TOOLING` plus the
workflow file, asserted in both directions by workflow.test.ts against the
exported constant — a missing path would let a harness change merge unproven,
an extra one starts a job that can only answer `[]`.

Also drops the workflow header's contradictory scope paragraph: it claimed the
lane selects on kernel sources and any test reaching one, which has not been
true since the ratchet went.

* fix(mutation): score and publish a short shard set before failing on the count

The expected-count check ran inside readShardedReports, before anything was
summarized, so on the weekly's real `--expect-shards 10` one dead shard threw
away the nine that had reported — the earlier reorder only moved the
zero-mutants check. The merge now returns the shard count, and both verdicts
run after emit() with the same exit code and `score` stage.

Regression uses the weekly argument shape (`--expect-shards 10`, one shard
present) and asserts the reporting kernel's row reaches stdout while the run
still fails.
2026-08-18 17:47:29 +02:00
Michał Pierzchała 8300fa131e refactor(ios): establish snapshot presentation seam (#1845)
Introduce RawAXNode and PresentedNode so acquisition backends can no longer construct the wire-facing snapshot shape directly. Preserve current output while #1797 moves semantics behind the seam.

Non-vacuity: setting PresentedNode.label to nil made testSnapshotPresentationPreservesCurrentWireShape execute once and fail on the missing label field; restoring the production mapping made the same focused XCTest pass.
2026-08-18 17:39:26 +02:00
Michał Pierzchała ef6ec2995b chore(layering): document R12/R18/R19, retire R8, make R9 shrink mandatory (#1781 A6) (#1825)
* chore(layering): document R12/R18/R19, retire R8, make R9 shrink mandatory (#1781 A6)

The A6 review kept `check:layering` in full (15/15 planted violations fired,
no other enforcer exists) and left four follow-throughs.

R12 bin-alias-fast-path, R18 contracts-implementation-authority and R19
selector-pipeline-ownership were live rules with no ADR or CONTEXT anchor —
they now carry one each, in the same list as R7/R9/R10/R13.

R8 zero-dep-job-closure is retired: no CI job sets `install-deps: false` and
ci.yml records why each keeps it enabled, so the invariant has no subjects.
R11's relative-into-packages exception existed only because a zero-dep closure
cannot coexist with specifier loads, so it retires with R8; the route is now
closed to every caller. R1 was retired the same way at #1490.

R9 was growth-only and merely suggested lowering the ceiling, which is
headroom the next change spends without a number moving. It is now an equality
pin like R6 and the R10 R7 counts, and the committed baseline drops 47 -> 46
(daemon-server ceiling 17 -> 16) to match the measurement.

ADR 0019 §6 now says each runtime-command-cutover row is deleted when that
command's migration is declared closed.

* chore(layering): rename R9 to type-cycle-size now that it fails both ways (#1781 A6)
2026-08-18 15:35:46 +02:00
Michał Pierzchała 4b44c1c53a chore(test): remove the contention retry and shrink the subprocess-stub project (#1781 A4) (#1827)
The enumerated single-retry policy (#1419) has fired zero times since it
landed on 2026-07-29: 0 of 234 sampled Coverage-job lane envelopes
(2026-08-11 to 2026-08-18) have retryCount > 0, and none of 17 recent
failed runs was retried (5 refused "outside the enumerated retry list",
4 refused "unhandled error"). All three trackers its entries pointed at
(#1098, #1414, #1419) are closed. It cost ~1,454 LOC, a per-run secret
marker threaded through a setup file on every Vitest project, and a
standing obligation for every future gate reporter to call the blocker
bus.

Delete the scripts, tests and fixtures, the check:contention-retry
script and gate, the envelope artifact upload, and the runner-timeout
setup file; test:coverage:ci is a plain `vitest run --coverage` again.
lane-envelope.ts stays: the mutation, fuzz and concurrency-torture lanes
build their envelopes from it. run-blocker-bus.ts goes: its only
consumer was the retry's failure sink, and its only publisher already
fails the run by setting process.exitCode.

Keep the subprocess-stub project for the three files that really spawn
(client-metro, fuzz harness, fuzz corpus-replay) and drop the three that
run in 31/212/277ms in CI, which cannot contend for anything. The list
is now a plain array in vitest.config.ts with the reason at each entry.
Membership and the project's kill criterion live in #1823.

Because test:coverage:ci is a bare vitest run, the gate manifest reads
its projects directly, so OPAQUE_RUNNERS no longer needs it and an
unrun Vitest project becomes unrepresentable rather than detected; the
audit test now constructs that state by project-scoping the script.
2026-08-18 15:35:25 +02:00
Michał Pierzchała 142d156338 ci(ios): run the full XCTest suite nightly and check the PR test list (#1781 A7) (#1789)
* ci(ios): run the full XCTest suite nightly and check the PR test list (#1781 A7)

* fix(ci): skip the runner server entry point in the nightly and validate both test flags

* docs(ci): restate the nightly lane cost and timeout honestly

* docs(ci): stop quoting XCTest counts that drift between commits

* ci(ios): tighten the nightly timeout to the measured suite duration
2026-08-18 12:00:43 +02:00
Michał Pierzchała ccf64f6797 ci: move parked device replay suites to a dispatch-only workflow (#1781 A1) (#1794)
* ci: move parked device replay suites to a dispatch-only workflow (#1781 A1)

Both full-tier device jobs have failed every scheduled run since 2026-07-24: the
Android suite inside full-tier scenarios that had never executed end to end, the
iOS suite on varying steps. They move to .github/workflows/replays-manual.yml,
which has no `schedule:`, so the schedule stops emitting a guaranteed failure while
the suites stay runnable on demand.

A job-level `if: github.event_name == 'workflow_dispatch'` would have looked the
same and lied: `workflowLanes()` decides `qualifying` per workflow FILE and never
reads job-level `if:`, so the manifest kept reporting replay-android, replay-ios,
and replay-ios-device as scheduled-lane owners — the silent-owner-loss failure the
manifest exists to catch. A separate file is what the file-level model already
reads correctly.

Those three checks now have no pull_request/schedule owner, so they are declared as
MANUAL_ONLY_OWNERS rather than folded into UNPROVABLE_OWNERS, whose claim ("it runs,
this loader cannot see it") is no longer true for replay-android. check:gate-manifest
drops from 48 to 46 wired checks and names the three on every run. Two tests pin it:
a dispatch-only lane is non-qualifying however many gates it declares, and every
manual-only declaration must name a registered check that no qualifying lane owns, so
a re-scheduled lane cannot keep a stale exemption.

* ci: attest manual-only checks against their dispatch lane (#1781 A1)

Review P1: MANUAL_ONLY_OWNERS was a negative allowlist — it proved each entry named a
registered check no qualifying lane owned, but nothing tied the entry to a lane that can
still run it. Deleting a parked job, or its run-gate step, would have left the manifest
green and still printing the check as manual-only: parked coverage silently becoming
deleted coverage.

Each entry now names its dispatch lane, and a new 'manual-only' audit assertion resolves
that name against the derived model: the lane must exist, must still be dispatch-only, and
must still declare the gate. replay-android carries an explicit `opaque` flag because its
gate sits inside the third-party emulator action's `script:` (#1429), so the job's
existence is the whole attestation the model can make — and the flag says so rather than
letting an unreadable lane look like a declaring one.

Four regressions pin both directions: deleting a declaration reports the check as unowned;
deleting the parked job fails with 'no workflow defines'; re-scheduling the lane fails
until the entry is dropped; and a parked lane that loses its run-gate step fails unless the
entry is opaque.

* ci: make manual-only mean dispatch-only, not merely non-qualifying (#1781 A1)

Review follow-up: the attestation checked `qualifying === false`, which is true of any lane
that is not pull_request/schedule. Swapping `workflow_dispatch` for `push` in
replays-manual.yml would have kept the audit green and the checks printed as manual-only,
while the runs nobody starts by hand quietly started themselves on every push.

The lane model now keeps the trigger names instead of collapsing them into that one bit, and
the manual-only assertion requires `workflow_dispatch` and nothing else. Three planted
regressions cover the gap the review named: a parked lane re-triggered by `push` fails, a
parked lane with no trigger at all fails, and the loader test pins that trigger kinds survive
into the model (a push lane reads `[push]`, the nightly reads `[schedule, workflow_dispatch]`).
2026-08-18 09:59:34 +02:00
Michał Pierzchała 39cd4d346a refactor: route appstate through platform runtime (#1755)
* refactor: route appstate through platform runtime

* test: keep appstate capability fixture below complexity limit

* test: cover appstate required readiness fact

* fix: align appstate facts with boot readiness

* fix: keep appstate use declaration minimal

* fix: close appstate parity and ownership gaps

* docs: record final appstate size accounting

* fix: merge neutral runtime imports

* docs: align final appstate size totals

* fix: remove stale runtime dependency edges

* docs: correct appstate size accounting

* refactor: keep runtime-use factory internal

* docs: itemize runtime-use relocation

* fix: move appstate queries into runtime packages

* refactor: retire root foreground query paths

* refactor: share Android foreground parser ownership

* fix: preserve Android appstate parser precedence

* docs: keep appstate evidence in review artifacts

* fix: keep appstate runtime loading lazy

* fix: fail closed for stale limrun appstate

* fix: preserve limrun recovery and abort appstate

* fix: narrow limrun exact-owner recovery

* fix: allocate appstate cutover rule

* fix: reconcile appstate with merged main

* style: format harmony runtime test

* fix: allocate appstate rule id

* fix: allocate appstate layering rule

* fix: remove stale app command admissions

* fix: close appstate layering regressions

* fix: align Harmony capability parity with runtime facts

* test: cover limrun recovery-only readiness

* fix: keep Limrun recovery binding app-log only

* fix: parse Android app state in linear time
2026-08-12 16:55:50 +02:00
Michał Pierzchała 9c22467832 refactor(ci): make gate ownership structural (#1429) (#1753)
* test(ci): prove every registered gate is owned and reachable (#1429)

A check that silently stops running looks exactly like a green build. Two
suites had already stopped: `check:tmpdir-leaks` (with its model tests) and
`test:fixture-cache` are real package scripts that no workflow ran, reachable
only through the `check:unit` aggregate CI never invokes.

`CHECK_CATALOG` becomes the registry of every check and `pnpm gate <id>` the
only way CI runs one, so finding what a lane runs is a scan for `pnpm gate`
rather than an attempt to interpret shell. `pnpm check:gate-manifest` then
asserts against the real workflows that every registered check is run by some
qualifying lane (per unit, not per script name), that every check the real
selector activates for a path is run by a lane that path would start (#1420's
class), and that every Vitest project and suite script belongs to a check.

The wiring that keeps those honest is asserted too: a gate id must name a
registered check, an `if:` must be ruled on in GATE_CONDITIONS so `if: false`
unowns what it guards, an action declared to run a gate is proven to, and a
job whose steps the loader cannot open fails closed.

It deliberately does not try to prove CI runs project code only through
`pnpm gate`. Whether a shell block executes project code is not decidable from
its text, so shell this model does not recognise earns no ownership credit —
the failure direction is a check reported unowned, never one waved through.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* test(ci): update the two suites that assert on rewired workflow text

`scripts/mutation/workflow.test.ts` and `test/ci/trusted-fixture-artifact.test.mjs`
read the workflow and action files and assert on their command text, so routing
those steps through `pnpm gate <id>` moved what they were matching.

They are the two suites the manifest cannot help with: it proves a gate is still
run, not that a test asserting on how CI spells a command was updated with it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* fix(ci): credit gates by execution shape, and keep every guard

Three ways the manifest could report a gate as owned when it does not run.

1. Crediting was a substring scan over `run:`, which #1429 explicitly rules
   out — "do not infer reachability from a command name merely appearing in
   workflow text". `false && pnpm gate x`, a gate inside `if false; then … fi`,
   one named in a heredoc, and `echo pnpm gate x` all credited it. There is a
   live instance: conformance-regenerate.yml's "Fail if regeneration changed
   anything" step names `pnpm gate maestro-regenerate` inside an error message
   telling a human to run it, and that credited the gate.

   A gate now counts only as the first command segment of a line, and a body
   carrying shell structure earns nothing. Reachability inside a script is not
   decidable, so this does not try: unrecognised shape means no credit and the
   check reports unowned. `VAR=$(pnpm gate x …)` is read, since the assignment
   form is unambiguous and the gate runs.

2. Job-level `if:` was not modelled at all, though six live jobs carry one, so
   a job that cannot run still credited every gate inside it. Two conditions on
   the mutation lanes are now declared.

3. A caller's `if:` REPLACED the guard on a nested composite-action step
   (`guard[0] ?? step.condition`), so an outer `always()` erased an inner
   `if: false`. Steps carry every guard between the lane and the step.

Also corrects two source comments that still claimed project code run outside
the runner fails the manifest. It does not: such a step earns no credit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* ci: add the run-gate action that names a gate structurally

The seam the ownership proof will read instead of shell. A lane says which
gate it runs in `with.gate`, a typed input the manifest reads straight out of
the YAML and validates against CHECK_CATALOG.

Nothing here is wired yet — the ~60 call sites and the model change follow.
Added first so the target of that conversion is reviewable on its own.

`args` cannot select which gate runs; it is appended after the id, so the
worst a wrong value does is fail the gate it already named. There is no
`|| true` and no output capture: the gate's exit code is the step's exit code,
so a gate cannot run without being able to fail its lane.

Part of #1429.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* merge: main (#1770) and route its three new steps through the runner

#1770 landed the orphan-check fix on main, wiring `check:tmpdir-leaks`,
`check:tmpdir-leaks:test` and `test:fixture-cache` into Coverage, Layering
Guard and Integration Tests. This branch had wired the same three through
`pnpm gate`, so the merge produced two steps per check rather than a conflict
— each check ran twice.

Kept main's steps, with the placement and reasoning reviewed on #1770, and
changed only their `run:` line to the canonical runner. Dropped this branch's
duplicates. Net effect on CI is unchanged: the same three checks, in the same
three lanes, once each.

Gate manifest green after the merge: 47 checks wired across 33 lanes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* fix(ci): address review — suite detection, freerange, glob, vacuous skip-list

Six review findings plus the mutation blocker.

[bug] `registered` was shape-only, so a `test:*` script running
`node src/bin.ts test <dir>` resolved to a `script:` leaf and was invisible.
Four `test:replay:*` scripts were owned only because someone hand-registered
them; `test:replay:android` was neither registered nor reported while the
nightly ran the same six .ad files by inlining them. A `test:*` script is now
a suite by name. `replay-android` is registered, and the nightly runs the
script instead of re-listing its files so the two cannot drift.

  The nightly invokes it inside `reactivecircus/android-emulator-runner`'s
  `script:` input — shell handed to a third-party action this loader does not
  read — so the suite executes but cannot be credited. Recorded in
  UNPROVABLE_OWNERS with that exact reason rather than assumed.

  The fixed detector also found a second orphan the review did not name:
  `test:integration:progress`. That one is a reporter whose `--check` sibling
  is the registered gate, so it is declared in REPORTING_SCRIPTS — a
  declaration that itself fails when inert.

[bug] `freerange` defaulted to localRunnable, so fail-open ran `fr` (a Bun
binary) on the pre-push path. Now false.

[suggestion] The `--run` skip-list asserted `build:android-snapshot-helper`,
a name `android-helpers` no longer uses, so it could not fail. Derived from
the catalog instead.

[suggestion] `matchesGlob` joined `**` splits with `.*`, making the adjacent
slash mandatory — GitHub's `**` matches zero directories, so
`src/**/*.test.ts` did not match `src/a.test.ts`. Pinned against
`packages/*/src/**/*.test.ts`.

[suggestion] Deleted the unwired `run-gate` action. It had no callers, was
absent from GATE_ACTIONS, and its comment described a system that had not
shipped. It returns with the rewiring, not before.

[suggestion] Collapsed the module headers that narrated discarded designs.

Mutation: `daemon entrypoint publishes HTTP metadata and cleans up on
shutdown` is the only test here that spawns a real daemon process. It takes
~1.1s alone but exceeds Vitest's 5s default inside Stryker's dry run, which
aborts the sweep before a single mutant runs. Given 30s.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* fix(mutation): order sandbox aliases longest-first so subpaths resolve

Every shard of the mutation sweep aborted in Stryker's dry run with:

  Cannot find package '@agent-device/selectors/engine' imported from
    .tmp/stryker/sandbox-*/src/core/selector-pipeline.ts

The alias was generated correctly; it just never won. Vite matches a STRING
alias by prefix and takes the first hit, and `workspaceSpecifierTargets`
emitted the bare `@agent-device/selectors` ahead of the subpath entries. The
bare entry therefore captured `@agent-device/selectors/engine` and rewrote it
to `…/src/index.ts/engine`, which does not exist; Node fell back to real
package resolution, could not find the subpath inside the sandbox, and the dry
run failed before a single mutant ran — so the shard uploaded an empty
envelope instead of a report and the ratchet failed for want of one.

Sorting longest specifier first makes the most specific alias win:

  @agent-device/selectors/engine -> packages/selectors/src/engine.ts
  @agent-device/selectors/ast    -> packages/selectors/src/ast.ts
  @agent-device/selectors        -> packages/selectors/src/index.ts

`/ast` never tripped this because nothing in a related test set imported it;
`selector-pipeline.ts` introduced the first subpath import that mattered
(#1744), so the mutation lane has been unable to run since that landed. Any
PR touching `scripts/mutation/**` — which fails open into the full sweep —
would have hit it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* refactor: derive gate ownership from workflow structure

* fix: run gates without optional arguments

* fix: resolve mutation workspace subpaths exactly

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-12 16:02:18 +02:00
Michał Pierzchała 8f98d23f14 refactor(layering): give each colliding rule id its own number (#1750)
* refactor(layering): give each colliding rule id its own number

R11 and R13 each named two unrelated rules. report() groups violations by the
rule string and titles every annotation `Layering drift (${rule})`, so a shared
number made the guard's output ambiguous about which rule fired.

Reference counts decided which rule keeps its number. R11 package-boundaries is
named in ~30 places (CONTEXT.md, ADR 0019, testing.md, the mutation and
affected-check configs, four package source comments, its own tests) against one
for the contracts rule; R13 platform-package-substrate is the RULE in three
policy files plus CONTEXT.md, ADR 0019 and model.ts against two for the devices
cutover. Both keepers stay put and the two newest rules move up:

  R11 contracts-implementation-authority -> R18
  R13 device-inventory-cutover           -> R17

R17/R18 follow the namespace's order-of-addition convention (R14 #1701 < R15
#1702 < R16 #1724): device-inventory-cutover landed in #1699 and
contracts-implementation-authority in #1701. #1656 took R19 for
selector-pipeline-ownership on the same reading.

The rule-map header in check.ts is renumbered and reordered back into numeric
order, and gains the R18 entry the contracts rule never had -- without it a
reader looking up an R18 violation finds nothing where they used to find the
wrong rule. deviceInventoryCutoverSummary() was also the only OK-line summary
not leading with its rule number, which is what made the number unreadable from
the success line in the first place.

Also corrects a normative ADR reference. ADR 0019's platform-package import
rules -- contracts-to-platform, sibling-platform, root/daemon, raw-process --
are R13's, as CONTEXT.md:420 already says. The R11 attribution predates
platform-package-policy (#1697, a day before #1699), when R11 was the only
package rule.

* chore(layering): retire the expired R11/R13 collision allowances

KNOWN_RULE_ID_COLLISIONS was opened for exactly the two collisions the previous
commit renames apart, and ruleIdCollisionFailures expires an allowance on
contact: once the collision is gone the entry fails as stale, because a list
still naming it would wave it back through if anyone reintroduced it.

Both entries are therefore deleted in the change that removes the collisions,
leaving the empty list that admits nothing. The namespace is now one-to-one
across R2-R19.
2026-08-12 11:42:15 +02:00
Michał Pierzchała 52402aec4d refactor(daemon): split touch interaction orchestration into semantic modules (#1748)
* refactor(daemon): split touch interaction orchestration

Closes part of #1691: interaction-touch.ts becomes a router; press, fill,
direct-iOS, shared runtime, Android readiness, and response projection each
own one module. Behavior is unchanged.

* test(daemon): split touch interaction coverage by module

Redistributes all 87 discovered cases across the new module topology and
re-keys the two touch-family fallow baseline entries to the paths that now
hold the same (net one fewer) findings.

* docs: point ADR 0014 at the merged Android readiness regression file

* refactor(daemon): give targeted-touch admission its own module

Keeps interaction-touch-press.ts inside the 300-line budget after the
complexity decomposition: admission (surface/capability/button policy, target
parsing, @ref staleness and mutation admission) answers its own question.

* refactor(daemon): drop the redundant targeted-touch label alias

* test(daemon): split touch suites along the new production seams

Adds the press-admission suite the production module was missing and splits the
four over-budget suites along new production seams (direct-iOS eligibility,
Android ref freshness, touch payload). Every suite installs the full device mock
set: three Android-session cases regressed to TOOL_MISSING on a runner without
adb when the mock set was trimmed per file.
2026-08-11 20:06:59 +02:00
Michał Pierzchała 73b54c15c4 docs: amend ADR 0019 with broader-migration governance (#1738)
Amends ADR 0019 for migration beyond the adoption checkpoint:

- status: recordings command unit recorded as completed; no further
  command unit is authorized by Status — subsequent units are planned,
  budgeted, and authorized through the successor tracking issue
- section 6: explicit none platform-execution mode with invariants; the
  cutover gate applies to platform-executing descriptors only and rejects
  a false none declaration; silent registry-entry mode defaults prohibited
- section 8: move-dominated size accounting, deprecated surfaces die on
  legacy at the next major, two evidence tiers, one parametrized cutover
  gate, capability buckets deleted per unit
- section 9: one bind per handler, side-effect-free facts inspection,
  single defineUse, preferred operations require a recorded measurement
- section 10: cross-cutting facets land with their first consuming unit,
  evidence-gated startup recovery, two-phase gateway shutdown, teardown
  steps ride their owning domain, all-edge-kinds end-state layering rule
  over production daemon modules
2026-08-11 16:24:09 +02:00
Michał Pierzchała f5d9789764 feat: enforce local device claims and reconcile stale owners (#1735)
* feat: enforce local device claims

* fix: address device claim review feedback

* fix: persist canonical daemon claim state directory
2026-08-11 16:18:45 +02:00
Michał Pierzchała 62001cf210 refactor(record): derive session recording from the publication lifecycle (#1719)
* refactor(record): derive session recording from the publication lifecycle

`SessionState.recordSession` stored an answer the script-publication
aggregate already contained. Every writer set both, but nothing made them
agree, and #1533 was the consequence: a `--save-script` ingress re-armed
the flag behind an ABORTED status, and a bare `close` published a
recording the caller had been told was aborted.

That fix routed every write through one rule, which made the two agree
without making disagreement unrepresentable. The field remained a second
source of truth, and its doc comments had to carry the invariant that a
type could enforce.

Remove the field and derive the answer. `isRecordingPublication` reads
recording off the lifecycle: ordinary authoring records only while ARMED;
a repair transaction records for its whole lifetime, terminal statuses
included. That last clause is deliberately exact rather than merely safe —
`armRepairStep` armed the old flag and neither `abortRepair` nor
`commitRepair` ever cleared it, so narrowing it would silently stop
evidence capture for a committed repair. Whether it should is a real
question, and a behavior change, so it is left alone here.

What this buys, beyond one less field:

- `buildNextOpenSession` and `finalizeOrdinaryCloseScript` make no
  recording decision at all now, so no surface can arm recording without
  moving the lifecycle that authorizes it.
- The writer's publication gate is answered entirely by the aggregate. Its
  separate ABORTED check is gone: a terminal authoring lifecycle is
  already not recording, so one question replaces two that could disagree.
- The R7 ownership ratchet drops from 23 writer-owned fields / 29 owner
  claims to 22 / 26, and the layering manifest loses the entry whose
  comment documented the smell ("deliberately set on its own by paths that
  record without arming a publication").

Behavior-preserving: the derivation reproduces what the flag held at every
transition. The test fixtures that armed `recordSession` with no
publication state described a shape production stopped producing at #1478;
they now carry the lifecycle that causes recording.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFW9gJqz1wEHoowkdd1nFW

* test(close-script): flush queued event-log writes before removing the tmp root

CI failed the Coverage lane with ENOTEMPTY removing the test's tmp root,
in `afterEach` rather than in an assertion.

`SessionStore.recordAction` QUEUES its event-log append
(`queueEventLogWrite`) instead of writing it, and every close path in this
file records an action. Nothing awaited that write, so `fs.rmSync(root,
{recursive: true})` could race it: the pending append recreates
`<root>/sessions/<name>/` while rmSync is walking, and the final rmdir
fails ENOTEMPTY. It needs CI's parallel load to lose the race — the file
passes 12/12 in isolation locally.

Await `flushSessionEventLogWrites()` before removing. The hazard is latent
in any test that records actions and then removes its tmp root; this fixes
the file that failed rather than sweeping the pattern, which deserves its
own change.

Not added to the #1419 contention-retry list: that list requires a
concrete spawn/wait mechanism named per entry, and this file has none. The
race was a real teardown bug, not lane contention.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFW9gJqz1wEHoowkdd1nFW

* docs: correct ADR 0016 on recording vs publication for repair

Review caught a real overstatement. The amendment claimed evidence capture
and publication authorization are "the same question asked of the same
state". That holds for ordinary authoring — ARMED both records and
publishes, ABORTED and PUBLISHED do neither — but not for repair:
`isRecordingPublication` is true for every repair status including
`committed` and `aborted`, while the writer additionally applies
`isRepairArmedWriteBlocked`, refusing a committed transaction and one that
is not yet committable.

State it as it is: both decisions derive from the same aggregate, but they
remain distinct predicates, and collapsing them would republish a committed
repair or commit an incomplete prefix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFW9gJqz1wEHoowkdd1nFW

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-11 10:58:03 +02:00
Michał Pierzchała 1b2e786128 refactor: move screen recording onto platform runtime (#1724) 2026-08-11 10:24:57 +02:00
Michał Pierzchała c7242f877f refactor: extract durable capture resource lifecycle (#1720) 2026-08-11 10:08:27 +02:00
Michał Pierzchała fa9a350361 docs: prefer design fixes over regression-only guards (#1722)
* ci: require simplicity review for large tooling changes

* docs: prefer design constraints over regression-only fixes

* docs: simplify design-first guidance
2026-08-10 21:15:41 +02:00
Michał Pierzchała 05a1d76f2e test: add daemon RPC wire-surface compatibility gate (#1717)
* test: gate daemon RPC wire compatibility against the last released tag (#1432)

ADR 0006 fixes exactly when DAEMON_RPC_PROTOCOL_VERSION must be bumped, and
nothing checked that it was. The runtime guard (readRemoteDaemonHealth) refuses
a mismatched peer, but only fires when someone remembered the bump — a wire
change that skipped it left both sides advertising protocol 2 while parsing
different payloads, which is the failure ADR 0006 exists to prevent.

Local daemons cannot skew (isReusableDaemonInfo takes over on any package
version mismatch). Cross-machine is skewed by design — proxy, cloud/limrun, a
remote macOS host — and ADR 0006 explicitly rules package version out as the
compatibility gate there, so the one boundary where skew is intended was the
one boundary with no gate.

test/wire-compat/surface.ts declares the wire surface grouped by the ADR bullet
each group serves, quoting it, with an `uncovered` note where a bullet is only
partly digestible (the /health and /rpc literals inside http-server.ts stay
reviewer-owned: a moved route 404s at connect time rather than misparsing).
ledger.json records what each declaration hashes to, at which protocol version.

Two gates, split for the same reason the replay-compat corpus splits:
- unit-core holds the ledger to its source and prints the digest to paste;
- Released-Surface Compatibility reads the ledger at the last RELEASED tag and
  requires the drift since then to carry a bump or a compatibleChanges ack.

From one commit a bumped ledger and an unbumped one are both just an edited
file, so only a released baseline can tell them apart. Acks are keyed by the
digest they cover, so one "added an optional field" cannot launder later
changes. Digests ignore comments and formatting; the manifest's closure is
derived from the AST, so a field typed by an unlisted sibling fails rather than
sitting outside the gate.

CI cost: one added job (checkout + toolchain + two node scripts, ~1 min),
mirroring the existing full-history replay-compat job.

* test: close wire-surface overclaim and make the closure fail closed (#1432)

Addresses both review P1s on #1717.

P1 — the manifest materially overclaimed ADR 0006 coverage. It quoted all four
bullets while digesting only the payload TYPES, so the producer and consumer
seams could break a skewed peer without moving a listed digest. Now listed on
both sides of every boundary: JSON-RPC method sets and the projections that
turn each method's params into a DaemonRequest, createRpcError/sendJson/
writeRpcResponseEnvelope, resolveToken and the auth-hook types, upload
preflight/finalize/308 handlers and the resumable ticket shape, artifact route
and download/inventory framing, REST error mapping, and the client's own
payload builder, lease-method mapping, response parser and error projection.
57 -> 117 declarations.

What stays out is now named rather than implied: createDaemonHttpServer's
dispatch wiring and the /health and /rpc literals inside it. Everything it
dispatches WITH is digested individually, and a moved route 404s at connect
time rather than misparsing — the loud failure, not the silent one.

P1 — imported and re-exported payload shapes escaped the closure.
declarationHomes() scanned only the manifest's own files and the walk
continued silently when a name could not be placed, so a listed type could
gain foo?: ImportedShape from a new module and stay green. Resolution is now
explicit and fails closed: relative imports, workspace specifiers (through the
owning package's own exports map, so a re-pointed export cannot drop a type),
and facade re-export chains. Every referenced name must land on a listed
declaration, a waiver with a written reason, a declared external module, or the
TS/Node global set. Fixed two extractor blind spots the walk exposed: a
declaration's own generic parameters and `as const` were being reported as
references.

Planted-red proofs (wire-mutations.test.ts): 13 cases independently mutate
method naming, response serialization, response parsing, auth projection,
upload ticket shape, 308 framing, artifact framing, REST error mapping, and
progress framing, each asserting the digest moves; 3 probes prove the closure
really reaches across a package boundary, a facade re-export, and a plain
relative import. Mutations apply inside the declaration's own span — a
whole-file replace silently hit a sibling sharing the substring, which is how
the first draft of one case passed vacuously.

The largest waiver pair (InternalRequestOptions, CommandFlags) rests on ADR
0006's own additive rule: they reach the peer inside DaemonRequest's untyped
flags/input bags, and the decision says a new flag needs no bump. Digesting
them would fire the gate on every new CLI flag and train reviewers to
rubber-stamp acks.

* test: list the consumer half of the auxiliary HTTP boundaries (#1432)

Addresses the remaining review P1 on #1717. The manifest claimed both sides of
response/upload/artifact framing while listing nothing from upload-client.ts,
daemon-artifacts.ts, or the health consumer in daemon-client-transport.ts, so
those parsers could narrow without moving a listed digest or protocol 2.

Now listed (117 -> 141 declarations):

- /health consumer: RemoteDaemonHealth, readHealthPayload, readDaemonHttpHealth,
  readRemoteDaemonHealth. This is the sharpest of the three — narrowing the
  reader or the comparison disables the very refusal ADR 0006 exists to
  guarantee, and nothing else in the repo would notice.
- /upload consumer: UploadResponse, UploadPreflightResponse, UploadPreflightResult,
  parseUploadPreflightResult, requestUploadPreflight, uploadDirectArtifact,
  tryDirectUploadWithResume, shouldRetryDirectUpload, finalizeDirectUpload,
  uploadLegacyArtifact, ARTIFACT_HASH_ALGORITHM, isStringRecord, and
  PreparedUploadArtifact — whose sha256/sizeBytes/fileName/artifactType/
  contentType fields ARE the preflight body the daemon parses.
- /artifacts/* consumer: DaemonArtifactEndpoint, buildDaemonArtifactUrl,
  isRemoteDaemon, DownloadRemoteArtifactParams, downloadRemoteArtifact,
  materializeRemoteArtifacts, resolveMaterializedArtifactPath.

Running the closure fail-closed over the new files surfaced three more stops,
each decided rather than skipped: PreparedUploadArtifact listed (it is payload),
UploadProgressSink waived (client-local rendering, never leaves the process),
and src/daemon/types.ts#DaemonArtifact waived as a re-export alias of the listed
kernel type, matching its DaemonRequest/DaemonResponse siblings.

10 more planted-red mutations cover the new seams: health version-read and
mismatch-refusal defeated, RemoteDaemonHealth field dropped, preflight parser
narrowed, preflight/legacy response shapes narrowed, finalize body key renamed,
ticket field renamed, artifact tenant header dropped, artifact URL moved. A
fourth closure probe proves the upload-consumer files are genuinely reached by
the walk rather than merely listed. 22 -> 33 tests.

The README now states the coverage as a producer/consumer table per boundary,
so the claim is checkable at a glance instead of asserted in prose.

* test: list the client half of the resumable 308 contract (#1432)

Addresses the third review P1 on #1717. Listing the daemon's
handleResumableUpload proved it still PRODUCES 308; nothing proved the client
still CONSUMES the released one. src/remote/upload-stream.ts owns that half and
was entirely outside the manifest, so a newer client could stop accepting
`upload-offset`, change how it reads `Range: bytes=0-N`, or emit a different
resumed `Content-Range` without moving one of the 141 listed digests.

Now listed (141 -> 151): UploadStreamResponse, streamFileToHttpRequest,
streamFileToHttpRequestAttempt, buildUploadRequestHeaders, isUploadResumeStatus,
isUploadRedirectStatus, parseUploadResumeOffset, parseNonNegativeIntegerHeader,
firstHeaderValue, MAX_UPLOAD_REDIRECTS.

streamFileToHttpRequestAttempt is listed despite its size, unlike
createDaemonHttpServer which stays in `uncovered`. The distinction is stated at
the declaration: the HTTP server only dispatches to handlers that are each
digested, while the attempt loop IS the resume state machine — it decides
whether a 308 continues the upload and what the next request carries, so its
sequencing alone can break a released daemon while every helper keeps its digest.

6 new planted-red mutations prove the client half moves the ledger: a dropped
`upload-offset` fallback, narrowed Range parsing, a changed resumed
Content-Range, 308 no longer treated as continue, a narrowed UploadStreamResponse,
and dropped header-value coercion. 33 -> 39 tests.

Closure fail-closed surfaced two more stops: UploadStreamProgressOptions waived
(local byte-progress rendering) and URL/URLSearchParams added to the global set.

README now carries a `/upload` resume row in the producer/consumer table, and
names the pattern behind three rounds of review: the coverage sentence kept
getting written ahead of the coverage, so the table and the `uncovered` notes
are the claims to trust — they are checkable against surface.ts, prose is not.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-10 20:52:29 +02:00
Michał Pierzchała 4f9aded0b8 fix(record): make an aborted authoring recording terminal by construction (#1712)
* fix(record): make an aborted authoring recording terminal by construction

A second successful `open` on an `open --save-script` session aborts the
recording: the aggregate goes to `authoring{aborted}`, `recordSession` is
cleared, and the caller is warned. `close --save-script` then refuses it with
"Retry with plain close; it will tear down the session without writing."

That promise was not kept. When the second `open` itself carried
`--save-script`, the recorder's shared flag ingress re-armed `recordSession`
while leaving the status terminal, and a bare `close` published the full
session log — the writer gated only on `recordSession` and the repair variant,
so nothing on the ordinary authoring path refused an aborted lifecycle.

The abort is now terminal by construction rather than inert by ordering:

- `isAuthoringAborted` gives the pure aggregate one home for the question.
- `applyRecordedSaveScriptFlags` takes no branch for an aborted lifecycle:
  it neither re-arms recording nor retargets the output path.
- `SessionScriptWriter` asks one `isPublicationWriteBlocked` question covering
  all three reasons to publish nothing, so every path reaching the writer
  (bare `close`, teardown, idle-reap, active publication) refuses it.

Armed recordings, published recordings, and every repair transaction are
unaffected; the control tests for those stay green against the pre-fix code
while the five new regressions go red.

Closes #1533

* docs: record the #1533 resolution in ADR 0016 and fix a stale symbol reference

The ADR 0016 close-time amendment still described #1533 as unresolved, and
the session-close.ts note named `isAuthoringAbortedWriteBlocked` — a private
helper that was folded into `isPublicationWriteBlocked` when the writer's
three sequential guards collapsed into one predicate, so the symbol names
nothing in the tree.

* fix(record): arm recording through the publication lifecycle, not around it

`recordSession` is an evidence-capture flag, but three surfaces set it
directly without consulting the publication aggregate, so it could
contradict a terminal ABORTED authoring status. The #1533 fix closed the
recorded-action ingress and made the writer refuse an ABORTED lifecycle,
then documented the remaining contradiction as acceptable — the writer's
own comment noted that "something can re-arm that boolean behind the
terminal status".

That something was live: `buildNextOpenSession` re-armed recording for any
`open --save-script`, and `applyOrdinaryScriptRecordingOpenOutcome` only
aborts a lifecycle that is still ARMED. A third `open --save-script` on an
already-ABORTED session therefore left `recordSession` true behind the
terminal status. The writer gate hid the publication symptom, but the
session kept paying recording-time costs for a recording that can never
publish: `recordSession` disables the direct iOS selector fast paths for
click and get, forcing every interaction onto the snapshot route.

Route the flag through one rule owned by the publication projection
(`recordSessionAfterSaveScriptFlag`), which answers "not recording" for an
ABORTED lifecycle on every surface that handles it — the re-open builder,
the close finalizer, and the recorded-action ingress. The writer's gate is
unchanged and still correct; it now stands on the aggregate alone rather
than as a net under a known drift, so the comments defending the drift are
replaced by statements of the rule.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFW9gJqz1wEHoowkdd1nFW

* docs: describe the full #1533 surface in the changelog entry

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFW9gJqz1wEHoowkdd1nFW

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-10 18:01:21 +02:00
Michał Pierzchała f569b91d65 docs: record platform runtime adoption checkpoint (#1703)
* docs: record platform runtime checkpoint revision

* docs: correct platform runtime checkpoint evidence

* docs: finalize platform runtime checkpoint
2026-08-10 17:58:43 +02:00
Michał Pierzchała b1ed5353d1 refactor: extract platform log runtime (#1701)
* refactor: extract platform log runtime

* fix: clear terminal app log recovery markers

* fix: preserve scoped app log tooling

* fix: preserve app log cancellation

* fix: handle large changed coverage diffs

* fix: harden Limrun runtime identity

* refactor: tighten platform log runtime

* fix: close app log trust gaps

* fix: accept canonical session path aliases

* refactor: extract durable capture kit

* fix: refresh retained log marker admission

* fix: rotate app logs after process relaunch
2026-08-10 17:58:42 +02:00
Michał Pierzchała c06bed9f77 refactor: extract platform device inventory runtime (#1699)
* refactor: extract platform inventory runtime

* fix: preserve scoped Apple inventory tooling

* fix: preserve Apple tool cancellation

* refactor: tighten platform inventory boundaries
2026-08-10 12:51:59 +02:00
Michał Pierzchała 44c298d7f3 docs: adopt request-bound platform runtime (#1697)
* docs: adopt request-bound platform runtime

* docs(adr): record the rejected process-local live-handle ledger alternative
2026-08-09 17:50:23 +02:00
vw2x 4b432fb59b feat: add HarmonyOS support (#1683)
* feat: add HarmonyOS device automation foundation

Add HDC-backed discovery, snapshots, application lifecycle, and core mobile interactions.

Route HarmonyOS through the platform registry and client contracts.

Cover parsing and capability parity with focused tests.

* feat: support HarmonyOS HAP deployment

Install and reinstall signed HAP archives through HDC.

Resolve bundle identities from module metadata and relaunch after package replacement.

Extend deploy routing and capability coverage for HarmonyOS.

* feat: add HarmonyOS single-pointer gestures

Execute pan, fling, and swipe plans through HDC uiInput primitives.

Derive scroll coordinates from the live ArkUI viewport.

Keep unsupported multi-touch gestures explicitly rejected.

* refactor: split session inventory command handling

Separate session, device, capability, and app inventory response paths.

Preserve the public inventory response contract while reducing handler complexity.

* feat: support HarmonyOS keyboard actions

Route HarmonyOS enter, return, and dismiss through HDC key events.

Expose supported keyboard actions through the system command metadata.

Keep keyboard visibility inspection explicitly unsupported.

* fix: reject unsupported HarmonyOS drag gestures

Keep drag unavailable until HDC can preserve source and destination hold semantics.

* feat: add HarmonyOS app log streaming

1. Stream HarmonyOS app logs through PID-scoped hilog sessions.\n2. Record HarmonyOS app identity during bundle-id opens for app-scoped commands.\n3. Cover backend routing and bundle identity resolution.

* feat: report HarmonyOS foreground app state

1. Read the foreground HarmonyOS mission through aa dump.\n2. Expose HarmonyOS appstate with package and ability metadata.\n3. Add parser coverage for foreground and missing-state cases.

* fix: advertise appstate through capabilities

1. Classify appstate in the command descriptor capability matrix.\n2. Surface supported appstate commands in capability inventory.\n3. Cover the advertised Android capability contract.

* feat: sample HarmonyOS process performance

1. Sample HarmonyOS process CPU and resident memory through HDC.\n2. Expose the verified metrics through the shared perf command.\n3. Keep frame and memory snapshot collection explicitly unavailable.

* feat: clear HarmonyOS app state

1. Add HarmonyOS settings clear-app-state through bundle cleanup.\n2. Force stop the app before clearing data and cache.\n3. Reject all unverified HarmonyOS settings explicitly.

* docs: document HarmonyOS support

1. Describe HarmonyOS HDC prerequisites and HAP installation.\n2. Add HarmonyOS to platform discovery and product documentation.\n3. Document verified performance limits for the public HDC surface.

* fix: preserve HarmonyOS deploy session identity

1. Bind a resolved HarmonyOS bundle after install or reinstall.\n2. Keep app-scoped logs and observability available after deployment.\n3. Cover session identity preservation for HarmonyOS reinstall.

* test: lock HarmonyOS capability boundary

1. Add an independent HarmonyOS capability-matrix oracle and exact advertised-command regression test.
2. Document current HDC-backed support and evidence-based unsupported command boundaries.

* refactor: simplify HarmonyOS shared platform boundaries

1. Split device selection and settings dispatch into focused helpers without changing behavior.
2. Keep HarmonyOS serial selection and lock-policy classification covered by regression tests.
3. Remove Fallow complexity findings from the HarmonyOS diff against upstream main.

* fix: bound default HarmonyOS HDC commands

1. Apply a 15 second timeout to ordinary HDC operations.
2. Preserve operation-specific timeout budgets for installation and capture paths.
3. Add regression coverage for default and overridden HDC timeouts.

* feat: add HarmonyOS screen recording

Implement physical-device whole-screen recording through the system recorder and HDC media transfer.

Reject unsupported HarmonyOS recording scopes and export flags.

Cover capability routing, media retrieval, cleanup, and simulator rejection.

* feat: report HarmonyOS HDC readiness

Add an HDC version check to the HarmonyOS doctor flow.

Document HarmonyOS as a supported doctor platform and cover the result.

* refactor: simplify HarmonyOS recording checks

Reduce recording validation and test complexity without changing behavior.

* test: cover HarmonyOS platform contracts

Synchronize public platform expectations across CLI, MCP, replay, and inventory tests.

Mock HarmonyOS inventory probes to preserve concurrent test behavior.

* test: model HarmonyOS recording capability

Require a physical HarmonyOS device in the independent capability parity oracle.

* test: cover HarmonyOS input and lifecycle paths

Exercise HDC input, lifecycle, installation, and relaunch command sequences.

* test: cover HarmonyOS device observability paths

Exercise discovery, screenshot validation, and process performance sampling.

* docs: define HarmonyOS CI hardware policy

Keep HDC hardware validation local and require mocked CI contract tests.

* fix: honor HarmonyOS app inventory filters

* fix: bound HarmonyOS app inventory classification

1. 限制应用元数据分类并发并为默认清单设置整体时限.
2. 将请求取消信号传递给 HarmonyOS 应用清单读取.
3. 补充失败时中止在飞读取且不继续排队的回归测试.

* fix: preserve HarmonyOS inventory failure causes

1. 保留触发应用元数据分类失败的原始错误, 避免被取消同级任务覆盖.
2. 补充总时限中止在飞读取且不启动排队任务的回归测试.
3. 验证后序任务失败时保留默认筛选的恢复提示.
2026-08-09 10:29:20 +02:00
Michał Pierzchała ac9e4d0f04 test: measure oracle liveness suite-wide; pin the one dead-path oracle (#1679)
* test: measure oracle liveness suite-wide; pin the one dead-path oracle

- docs/agents/oracle-negation-spike.md: assertion-negation sweep over 677
  test files (5,687 verdicts). Zero vacuous tests: all 150 negation
  survivors decompose into assert.rejects-validator artifacts (112),
  helper-oracles (31), in-file-fake breakage (6), and one conditional
  oracle. Records the companion mock-coupled coverage-uniqueness numbers
  and the follow-ups they motivate (diff-scoped mutation gate, provider
  seam closures, transcript provenance).
- watchos-sentinel: the non-watchOS test's only assertion sat in a catch
  block that never fires (tvOS interactor creation succeeds), so no
  assertion executed on the observed path. Pin creation success instead.
  Red-run proof: the old shape survived the negation sweep; the new shape
  fails under it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(android): inject fake adb through the provider scope, not PATH stubs

Adds withFakeAdb to test-utils: a scripted in-process AndroidAdbProvider
installed through the production withAndroidAdbProvider seam — the same
scope the daemon installs per request — replacing PATH-stub shell scripts
that spawn a real subprocess per adb call. No PATH mutation, no spawns,
no real subprocess waits.

Converts settings.test.ts (15 tests, 23ms; waiver said "waits real
settings-apply poll time") and notifications.test.ts (2 tests, 9ms).
Assertions move from args-log regex greps to structural checks on the
recorded call list; the fake receives device-scoped args with the
-s serial pair stripped, so serial routing is enforced by the scoped
provider matching device.id instead of asserted per call.

Remaining PATH-stub files convert next; their contention-retry waiver
entries lift together with the conversions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(android): convert device-input-state to fake adb provider injection

10 PATH-stub cases move to withFakeAdb through the production provider
scope; the 2 tests that already inject an executor directly are
unchanged. Cross-invocation shell STATE_FILE state becomes a closure
boolean; args-log regex asserts become structural checks on recorded
calls. 12/12 green at 386ms — the residue is dismissAndroidKeyboard's
two fixed 120ms retry sleeps, not stub subprocess waits.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(android): convert app-lifecycle-install adb stubbing to fake provider

The adb half of every case moves to withFakeAdb through the production
provider scope; installs take the documented exec-shaped fallback
(exec(['install','-r',...])), matching what the PATH stub saw minus the
serial pair. bundletool/zip/unzip stay real or PATH-stubbed — they run
via runCmd outside the adb seam, so this file remains in the serialized
subprocess-stub lane with its waiver reason to be corrected from adb to
bundletool. 13/13 green at ~130ms; no case enters a retry/poll loop.

Conversion note: manifest identity's `unzip -p` failure is silently
swallowed (readZipEntry catch -> undefined, aapt fallback) — an
invisible degradation path worth a future explicit diagnostic.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(android): convert input-actions adb stubbing to fake provider

9 PATH-stub cases move to withFakeAdb; the 3 tests already injecting
providers directly are unchanged. Chunked shell-input assertions become
ordered deepEqual on the recorded calls; never-called negatives and
call-count checks preserved 1:1. 12/12 green.

File time drops to 2.2s, all of it production sleeps:
verifyAndroidFilledText unconditionally waits its [0,150,350]ms
verification cadence even when the first inspection matches, so each
fill verification pass costs ~500ms with an instant fake. A
budget-derived cadence there (testing.md pattern 1) would put this file
near 25ms; flagged as follow-up rather than changed here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(android): extract shared oracles; make fake adb failure-faithful

Three test-utils extractions applied across the six converted files:

- assertRejectsAppError collapses the hand-rolled AppError code+message
  rejection validator (10 sites here; ~30 more repo-wide can adopt it
  incrementally). Validators asserting details or multiple differently-
  flagged regexes stay explicit on purpose.
- withFakeAdb gains a `provider` option for extra capabilities
  (snapshotHelperArtifact, reverse, ...), replacing input-actions'
  nested re-scoping bridge.
- withFakeAdb now mirrors the local executor's contract: a scripted
  nonzero exit throws androidAdbResultError unless the call site passed
  allowFailure. Provider-scoped exec bypasses exec.ts's throw-on-close-
  failure, so returning {exitCode:1} took a different production path
  than the PATH-stub `exit 1` these fakes replaced. All 75 tests hold
  under the corrected semantics.

Also swaps settings' inline emulator DeviceInfo literals for the shared
ANDROID_EMULATOR fixture.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test: lift five converted Android files from the contention-retry waiver

settings, notifications, device-input-state, input-actions, and
app-lifecycle-open no longer stub binaries on PATH or spawn
subprocesses, so their contention mechanism is gone: they leave
CONTENTION_RETRY_FILES and, through the derived SUBPROCESS_STUB_TESTS
constant, the serialized subprocess-stub project (17 -> 12 files).
app-lifecycle-install stays with its reason corrected: adb is now
in-process, but bundletool stays PATH-stubbed and zip/unzip spawn for
.aab packaging paths.

Full unit suite green at the new membership: 638 files, 5,724 tests,
with the five files running at unit-core's default parallelism.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test: apply review findings to the fake-adb conversion batch

- app-lifecycle-open: the missing-package launch failure returns
  {stderr, exitCode: 1} and lets withFakeAdb's throw path produce the
  production-shaped androidAdbResultError instead of hand-modeling the
  thrown AppError — the drift the helper exists to eliminate.
- withScriptedAdb deleted: the six converted files were its only
  callers, and a live PATH-stub export invites new tests back into the
  serialized lane this batch shrank. withMockedAdb stays (dispatch and
  runtime-hints tests still stub other binaries).
- android-snapshot-helper gains androidSnapshotHelperScriptResponse so
  the version-probe detection and versionCode reply have one source of
  truth; input-actions' local copy delegates to it.
- withFakeAdb's provider option becomes a distributed Omit over the
  AndroidAdbProvider union, so touch without gestureViewport is a
  compile error at the fake's boundary (planted and verified) instead
  of a TypeError inside production gesture planning.
- spike-doc re-run checklist restores wider than the codemod globs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test: drop the consumer-less FakeAdbScript barrel re-export

Fallow's dead-code gate flagged it: scripts are always passed as inline
lambdas, so only FakeAdbResponse needs a name at the barrel. The type
stays exported from fake-adb.ts where the withFakeAdb signature uses it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(apple): inject fake xcrun through the tool-provider scope, not PATH stubs

withFakeAppleTool mirrors withFakeAdb for the Apple seam: a scripted
provider installed via the production withAppleToolProvider scope, flat
simctl/devicectl invocations recorded exactly as the PATH-stub shell
scripts saw them, throw-on-nonzero fidelity matching exec.ts unless the
call site passed allowFailure, and the canned `simctl privacy help`
listing served by default (the block withMockedXcrun injected into
every script). screenshot-status-bar.test.ts converts as the exemplar:
3/3 green at 9ms with deepEqual call-sequence assertions replacing the
args-log regexes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(apple): convert apps.test.ts xcrun stubbing to fake tool provider

All withMockedXcrun scripts and hand-rolled PATH stubs move to
withFakeAppleTool; args-log regexes become structural call assertions
(exact deepEqual where order is deterministic, presence checks where
the 5s simulatorBootedMemo TTL makes boot-probe order test-dependent).
12 hand-rolled AppError validators collapse into assertRejectsAppError.
54/54 green; file test time 1172ms -> ~400ms with no test over 201ms.

Five .ipa install tests keep a minimal PATH stub for unzip only:
install-artifact.ts:112 and install-source.ts:438 call runCmd('unzip')
directly, outside the Apple tool provider seam — the file therefore
stays in the serialized subprocess-stub lane with its waiver reason
corrected from xcrun to unzip.

Also observed: getSimctlPrivacyServices caches per PATH+simulatorSetPath
and simulatorBootedMemo keys on deviceId|setPath, so neither cache
accounts for the provider scope — worked around per test, follow-up
worthy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test: lift six Apple waivers; fix format and fallow findings from CI

- interactions, simulator, screenshot, physical-device-screenshot,
  devicectl, and screenshot-status-bar leave CONTENTION_RETRY_FILES:
  the first five stopped stubbing PATH binaries in earlier refactors
  (measured 3-64ms per file, no subprocess activity), and
  screenshot-status-bar now injects through the fake tool provider.
  apps.test.ts stays with its reason corrected to the unzip PATH stub
  (xcrun is in-process; install-artifact.ts:112 / install-source.ts:438
  call runCmd('unzip') outside the Apple seam). Serialized lane 12 -> 6.
- oxfmt: fake-apple-tool.ts and contention-retry.ts were pushed
  unformatted (local check piped through tail masked the failure).
- fallow complexity: the three fake-script arrows in apps.test.ts drop
  under threshold via shared predicates (isSimctlMainScreenScale,
  isSimctlScreenshot, isDevicectlDevice), which also deduplicate the
  screenshot pair.

Full unit suite green at the new membership: 638 files, 5,724 tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-08 09:01:05 +02:00
Michał Pierzchała 6c0fcb64a1 fix: reject distinct ambiguous mutation targets (#1667)
* fix: reject distinct ambiguous mutation targets

* fix(ios): scope the raw-match rejection to mutating dispatches

`RunnerTests+Interaction.findElement` applied the new fail-closed
classification to `querySelector` as well as press/type, because the read
call site takes the default `allowNonHittableFallback: false`. With one
visible/hittable match and one non-hittable same-selector duplicate the
query started returning AMBIGUOUS_MATCH where it previously selected the
hittable element, and `queryDirectIosSelectorOrFallback` preserves that
error for read callers — so `get`, `is`, and `wait` surfaced an error
instead of their prior answer.

`classifyDirectSelectorCandidates` now takes a `rawMatchPolicy`. Mutations
keep `.rejectDistinctMatches` (the default, so no mutation call site
changes); `queryElement` passes `.preferHittableMatch`, restoring the
prior read rule: prefer the single hittable match, ambiguous only when
hittable matches compete, and never adopt the Maestro coordinate fallback.
The Maestro expected-point path is untouched.

Covers the one-hittable + one-non-hittable read, competing hittable reads,
and the non-hittable-only read. ADR 0011's amendment now states the scope.

* test(ios): execute selector read ambiguity regression

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-08 08:57:32 +02:00
Michał Pierzchała ac52281448 fix(test): deterministic temp-dir cleanup across node --test lanes (#1661)
* fix(test): deterministic temp-dir cleanup across node --test lanes

node --test has no global setup/teardown hook, so unlike Vitest (#1593) every
node --test package.json script (maestro:conformance, mutation:test,
check:affected:test, check:coverage-changed:test, check:layering,
depgraph:test, check:tmpdir-leaks:test, check:contention-retry,
test:fixture-cache, test:smoke(:web), test:integration:node,
test:concurrency-torture) still created scratch directories against the
real, unredirected os.tmpdir(), with cleanup only as reliable as each call
site's own try/finally — which a crash, OOM, or timeout kill bypasses
entirely.

Add scripts/node-test-tmpdir.ts: it wraps the whole `node --test`
invocation as a child process, redirecting TMPDIR to one disposable,
pid-tagged directory (shared root/prefix with the Vitest lane) and removing
it from the process 'exit' event, which fires on normal completion, a
thrown error, or a forwarded SIGINT/SIGTERM alike. Every node --test script
now runs through it. check-tmpdir-leaks.ts already scans by root/prefix, so
it covers both mechanisms with no changes to its detection logic.

Verified: a node --test process that mkdtemp's then gets SIGKILL'd leaves a
directory behind unwrapped; wrapped and SIGTERM'd, TMPDIR is redirected and
the directory is gone with no orphaned processes. All 13 wrapped lanes and
the full Vitest suite (5,591 tests) pass with zero residual
agent-device-test-run-* directories after the run.

Fixes #1595

* test(tmpdir): ratchet every node --test script through the wrapper

The 13 lanes wrapped in package.json were a one-time hand sweep with
nothing enforcing the pattern going forward — a 14th node --test script
added later without scripts/node-test-tmpdir.ts would silently reopen
#1595 for that one lane.

Add a structural check to scripts/node-test-tmpdir.test.ts (now part of
check:tmpdir-leaks:test) that reads package.json and fails if any script
invokes `node ... --test` without routing through the wrapper. Dumb
string matching over the scripts map, no shell parsing, with an explicit
(currently empty) NODE_TEST_WRAPPER_BYPASS_ALLOWLIST for any lane that
must legitimately bypass it. Verified it both passes on the current
package.json and fails when a synthetic unwrapped `node --test` script is
added.

* fix(test): preserve the Swift cache and close the raw node --test bypasses

Review on #1661 found two gaps:

1. The wrapper only overrode TMPDIR, so it discarded and forced a
   recompile of the durable Swift compiler cache every run instead of
   mirroring vitest-tmpdir-global-setup.ts's carve-out for it. Read
   os.tmpdir() before the child's TMPDIR redirect takes effect and set
   AGENT_DEVICE_SWIFT_CACHE_DIR from that (only when unset), same as the
   Vitest lane — the two now share one durable cache instead of each
   discarding and recompiling their own. Added a probe assertion
   (scripts/node-test-tmpdir.test.ts) that fails without the fix and
   passes with it (verified both ways).

2. docs/agents/testing.md documented raw `node --test` commands for the
   iOS smoke files, and the android/ios/conformance-regenerate/nightly
   workflows invoked `node --test` directly outside package.json. Routed
   all of them through scripts/node-test-tmpdir.ts so the documented
   local commands and CI lanes get the same crash/timeout-safe cleanup
   the package.json scripts already have.
2026-08-07 13:26:20 +02:00
Michał Pierzchała d81ac0a092 fix(ios): corroborate recorded tap outcomes (#1605)
* fix(ios): corroborate recorded tap outcomes

* fix(ios): preserve corroborated tap target identity

* fix(ios): suppress corroborated tap retries

* fix: require comparable iOS tap evidence

* fix: bound iOS tap corroboration baseline

* test(ios): deterministic injection seam for recorded-tap-failure corroboration (#1605 merge gate)

The field failure cannot be reproduced on this head: the tap false-failures
were a downstream symptom of XCTest-channel saturation, which the #1587
capture fixes removed. The seam records a real XCTIssue AFTER the real
gesture inside the per-command failure-count window, so
xctestRecordedFailureResponse and target invalidation fire byte-for-byte
like the field failure. Armed via a decrementing /tmp flag file (the daemon
regenerates tampered xctestrun templates, so env plumbing cannot reach a
daemon-spawned runner); compiled only under AGENT_DEVICE_RUNNER_UNIT_TESTS.

Live evidence on a daemon-spawned runner (Bluesky, ad-bsky-repro sim):
- landed case: injected failure on a real Search-tab tap -> success with
  the corroboration warning, screen verifiably on Search, no redispatch,
  runner serving next commands; flag consumed exactly once.
- unchanged case: injected failure on a dead-coordinate tap -> capture
  unchanged -> XCTEST_RECORDED_FAILURE preserved with the new honest hint;
  runner still usable.
- field-shape race (relaunch -> full snapshot -> immediate press, 5
  attempts): no natural recorded failure occurs on this head — the hostile
  tree needed for channel saturation is gone, corroborating the causal
  story.

* fix: reconcile tap corroboration with current interaction semantics
2026-08-05 18:38:07 +02:00
Michał Pierzchała 36f44ca2cc docs: clarify iOS drag synthesis profiles (#1616)
* docs: clarify iOS drag synthesis profiles

* refactor: consolidate Apple gesture event lifecycle
2026-08-05 14:21:58 +02:00
Thiago Brezinski a13a6832ee feat: add selector-targeted drag gestures (#1567)
* feat: add selector-targeted drag gestures

* fix: address drag gesture review feedback

* fix: satisfy drag review quality gates

* fix(android): lower drag trajectories piecewise

* test(replay): validate drag fixture selectors

* fix(ios): ignore full-viewport chrome containers

* test(drag): prove destination on live devices
2026-08-05 12:37:02 +02:00
Michał Pierzchała 23a3016e9b fix: stop the unit suite from leaking temp directories (#1593)
* fix: stop the unit suite from leaking temp directories

~650 test call sites across the unit suite create scratch directories via
fs.mkdtemp(path.join(os.tmpdir(), ...)) or shared factories (makeSessionStore)
with no cleanup, ever. Over time this accumulated 1.16M+ orphaned directories
in the real system tmpdir, slow enough to make tools that enumerate $TMPDIR at
startup (e.g. opencode) take 1-2 minutes to launch.

Rather than migrate every call site, redirect os.tmpdir() itself for the
lifetime of the whole `vitest run` invocation: scripts/vitest-tmpdir-global-setup.ts
wires in as vitest's globalSetup/globalTeardown, points TMPDIR at one
/tmp-rooted directory (verified: env mutations here propagate to every forked
worker, confirmed empirically), and removes it in one recursive rm after every
worker across every project finishes. Since os.tmpdir() reads TMPDIR on every
call, this covers all ~650 call sites without touching any of them.

Rooted at /tmp rather than nested inside the current (already deep, on macOS)
os.tmpdir(): that broke real AF_UNIX socket tests (runner-usbmux.test.ts) by
pushing socket paths past the 104-byte sun_path limit. A per-file afterAll
hook was tried first but proved unreliable — 5 of 7 workers in one run never
ran it before their process was torn down; the global setup/teardown pair
(one process, confirmed single execution) is the mechanism that's actually
guaranteed to run once.

Also adds:
- scripts/check-tmpdir-leaks.ts: CI/local guard asserting no
  agent-device-test-run-* directory survives a run (a leftover one means a
  worker was killed before cleanup could run).
- src/__tests__/test-utils/tmp-dir.ts: documented mkdtempForTest /
  mkdtempForTestSync helpers, the discoverable way to get a scratch dir going
  forward (mirrors the src/utils/exec.ts pattern for node:child_process).
- scripts/check-test-tmpdir-helper.ts: ratchet guard capping raw
  fs.mkdtemp/mkdtempSync call sites in test files at today's count (632); it
  can only shrink as call sites migrate to the helper.

* fix: make check-tmpdir-leaks scan the same root the fix actually uses

check-tmpdir-leaks.ts was scanning os.tmpdir() for leftover run
directories, but vitest-tmpdir-global-setup.ts creates them under a
hard-coded /tmp. On macOS those are different paths (TMPDIR is a deep
per-user /var/folders/.../T/ directory) — the guard could never find a
leak on the exact platform the original leak happened on, only on
Linux CI where os.tmpdir() already is /tmp.

Export TEST_RUN_TMP_ROOT and TEST_RUN_TMP_PREFIX from the global-setup
module and import them in the leak check instead of recomputing a path
that can drift. Switched from a fixed pid-based directory name to
fs.mkdtempSync so a same-named leftover from a prior killed run (or,
on a shared machine, another user) can't collide with a live run.

Also fixes two stale comments (in this file and ci.yml) that still
described the per-file afterAll hook design that was abandoned in
favor of the global setup/teardown pair, and notes the check only
covers vitest runs, not the node --test lanes (test:smoke,
test:integration:node).

Verified live: with the old code the guard reported no leaks even with
a real orphaned /tmp/agent-device-test-run-* directory present (left
by a command that got killed mid-run); with this fix it correctly
found and reported it.

* simplify: drop the tmpdir ratchet guard, keep the leak check local-only

Two guards were more than this needed:

- check:test-tmpdir-helper (ratchet on raw fs.mkdtemp call counts)
  protects nothing a bug could actually trigger — the leak is already
  fixed architecturally regardless of call-site count, so this was
  pure style/discoverability nudging. Dropped the script and its
  check:tooling/CI wiring; kept mkdtempForTest/mkdtempForTestSync in
  tmp-dir.ts as the documented option without enforcing it.

- check:tmpdir-leaks in CI added little: GitHub-hosted runners are
  destroyed after each job, so a leftover directory there is harmless
  by construction, and a worker getting killed mid-run would already
  surface as a job failure some other way. Its real value is local,
  on the long-lived dev machines where the original leak actually
  accumulated — kept it wired into check:unit, dropped the CI step.

* fix: don't flag a concurrent vitest run's tmpdir as a leak

check-tmpdir-leaks.ts reported every agent-device-test-run-* directory
as a leak, but a concurrent vitest run in another worktree legitimately
keeps its own directory present until its own teardown finishes. On a
machine that regularly runs several worktrees at once, that made
check:unit fail on unrelated in-progress work.

Embed the owning process's pid in the directory name (still random-
suffixed via mkdtempSync, so same-pid reuse across separate runs can't
collide) and have the leak check skip any directory whose pid is still
alive (process.kill(pid, 0)) — only directories whose owning process
already exited without running its globalTeardown are real leaks.

Split the pure logic into check-tmpdir-leaks-model.ts (findLeakedRunDirectories,
with an injectable liveness check for testing) so it has a real
regression suite, including the concurrent-run case, instead of only
being exercised by hand.

* refactor: migrate raw fs.mkdtemp call sites to mkdtempForTest(Sync)

Migrates 629 raw fs.mkdtemp(Sync)(path.join(os.tmpdir(), PREFIX)) call
sites across 168 test files to the mkdtempForTest / mkdtempForTestSync
helpers (src/__tests__/test-utils/tmp-dir.ts), so there's one
documented, discoverable way to get a scratch dir in a test — cleanup
already didn't depend on the call-site shape (the global TMPDIR
redirect covers any of them), this is purely for consistency and
discoverability, same reasoning as src/utils/exec.ts for
node:child_process. Existing manual per-test cleanup (fs.rm in
finally/afterEach/onTestFinished blocks) is untouched — the global
teardown is a fallback for killed workers, not a replacement for tests
cleaning up after themselves.

Migrated with a one-off AST-based codemod (oxc-parser, since regex
mismatched multi-line calls and complex prefix expressions like
`options?.tempPrefix ?? 'default-'`) rather than by hand across 168
files. The codemod isn't included — it doesn't need to survive this
commit. Caught and fixed one real bug in it during review: a small
number of files declare a second import statement later in the file,
after some of the matched call sites, which broke a naive "insert
after the textually-last ImportDeclaration" placement; fixed to insert
after the top contiguous import block instead, plus a self-check that
re-parses every generated file before writing it.

Also fixes 5 fallow dead-code findings the branch introduced: the
vitest globalSetup functions (setup/teardown) are only referenced by
the config-string path vitest.config.ts hands to globalSetup, invisible
to static analysis — suppressed with the documented convention.
isProcessAlive didn't need to be exported (nothing outside the module
uses it). And dropped a barrel re-export of the new helpers from
test-utils/index.ts: nothing actually imports through the barrel
(matching the existing makeSessionStore convention, which is imported
directly from store-factory.ts everywhere despite also being
barrel-exported), so the re-export was genuinely dead.

Documents the convention in docs/agents/testing.md.

Verified: full unit suite (5308 tests) passes except the one
pre-existing, unrelated package-exports.test.ts failure; typecheck,
lint, and format all clean; fallow audit clean against the PR base.

* fix: correct fallow suppression token and drop unused barrel re-export

These were meant to be part of 397cdd6d7 (verified locally before that
commit) but didn't actually get staged — caught by CI's Fallow Code
Quality check re-running against the pushed commit, which still had
the plural 'unused-exports' token (fallow expects singular
'unused-export') and the dead barrel re-export.

* fix: migrate the two mkdtemp call sites the rebase silently reintroduced

Rebasing onto main pulled in #1594's two new test cases in this file,
added independently of this branch's migration, still using raw
fs.mkdtempSync(path.join(os.tmpdir(), ...)). Git's line-based merge
found no textual conflict with this branch's removal of the os import
(the changes touch non-overlapping regions), so it silently produced
a file that doesn't typecheck. Migrated both to mkdtempForTestSync for
consistency with the rest of the file, caught by CI's Typecheck,
Fallow Code Quality, and FreeRange checks re-running against the
pushed commit.

* test: pin Vitest tmpdir lifecycle

* fix: preserve the Swift cache across test runs
2026-08-05 08:56:39 +02:00
Michał Pierzchała 4f8dc3f31e refactor: move selector engine into workspace package (#1589)
* refactor: move selector engine into workspace package

* refactor(selectors): trim the package façade to its real consumers

Follow-up to the selector-package cutover, from a structural review of it.

- Drop 15 façade symbols with no consumer anywhere in the repo:
  selectorUsesKey (added by the cutover, never called), isNodeVisible /
  isNodeEditable (the real helpers are contracts/snapshot's), normalizeText,
  splitIsSelectorArgs, IS_PREDICATE_REQUIRED_MESSAGE, four nested Replay
  types, SelectorDisambiguationDisclosure, and the four kernel type
  re-exports every consumer already imports from kernel directly.
- Delete SelectorCapturePolicyInput.selectorExpression, which
  deriveSelectorCapturePolicy never read; the policy varies only by
  predicate, so it takes one now. Two of the four tests asserted that the
  unread parameter had no effect and could not fail; they go with it.
- Return the Maestro export vocabulary to the maestro package. The cutover
  inlined MAESTRO_TEXT/STATE_SELECTOR_KEYS' values into the CLI call site,
  leaving both constants dead in the package that owns the concept and no
  gate over the two copies. MAESTRO_SELECTOR_PROJECTION is now the one
  statement of it.
- Dedupe SelectorDiagnostics and SelectorDisambiguationDisclosure, declared
  character-for-character twice across the AST/string seam, and name the two
  shared option shapes once instead of five inline copies. The parser-side
  resolution types take an Ast prefix so the twins read as twins.
- Delete three identity wrappers: parsePrivateSelector,
  selectorExpressionToMaestro, and the formatSelectorFailure forwarder —
  nothing passes it a chain any more, so the SelectorChain | string union
  and its branch go too.
- Delete internal/index.ts, an AST barrel whose only consumer was one test
  in the same directory (renamed to engine.test.ts), and the match.ts
  pass-through that existed to feed it.
- ReplaySelectorGrammar had three variants for two behaviors; 'wait' and
  'ordinary' were the same path. It is 'is' | 'positional' now.
- Drop the deleted src/sdk/selectors.ts from .fallowrc.json's entry list.

Behavior unchanged. pnpm check green: 598 unit files / 5278 tests, smoke
35 passed / 3 live skipped, layering 71/71, depgraph 22/22, mutation config
45/45, fallow clean, package smoke sound. Counterfactual: pointing
MAESTRO_SELECTOR_PROJECTION.textKeys at the state keys turns three
replay-maestro-export cells red; restored before commit.

* test(selectors): split the engine aggregation test by source concept

`internal/index.test.ts` (renamed `engine.test.ts` when its barrel went away)
was a 708-line aggregation over the whole engine — past the 500-line tripwire
and mirroring no source module, so it also ran as one serial unit.

It becomes five files that each mirror what they test, plus the parser cells
folded into the existing parse test:

  resolve.test.ts                 alternative fallback, strict uniqueness,
                                  first-match existence
  resolve-disambiguation.test.ts  ADR 0012 ranking: deepest, smallest-area,
                                  winner-vs-challenger disclosure, tie fallback
  resolve-viewport.test.ts        the visibility half: on-screen beats
                                  off-screen, including inside an off-screen
                                  scroll container
  match.test.ts                   per-key matching semantics (text, role,
                                  focused, appname/windowtitle, decoded
                                  newline labels)
  arguments.test.ts               where the selector ends and the command's
                                  positionals begin, both grammars
  parse.test.ts                   +6 grammar/escape cells beside the existing
                                  property tests

The login-form tree shared by resolve.test.ts and match.test.ts moves to
`__tests__/login-form-nodes.ts` rather than being copied into both.

All 27 cells are carried over unchanged and still pass; no file now exceeds
224 lines. pnpm check green: 602 unit files / 5278 tests, layering 71/71,
depgraph 22/22, mutation config 45/45, fallow clean over 127 changed files.

* revert(selectors): keep agent-device/selectors public, behind one AST subpath

The cutover removed the `agent-device/selectors` public subpath as part of
tightening the API. It is in use, so the removal is reverted: the subpath ships
the same ten symbols v0.20.5 shipped, with the same signatures.

That has to coexist with the reason the package façade is string-only, so the
AST leaves through one named door instead of the main one:

  @agent-device/selectors        string-in/string-out; every in-repo consumer
  @agent-device/selectors/ast    the published parser surface; one consumer,
                                 src/sdk/selectors.ts

`packages/selectors/src/ast.ts` re-exports parseSelectorChain,
tryParseSelectorChain, isSelectorToken, the AST-taking findSelectorChainMatch
and resolveSelectorChain, isNodeVisible, isNodeEditable, and types
SelectorChain / SelectorDiagnostics. `formatSelectorFailure` keeps its
published `SelectorChain | string` first parameter as a shim here rather than
widening internal/resolve.ts back to a union — the compatibility obligation
sits at the boundary that owes it.

This is strictly narrower than main, where the AST was reachable from anywhere
in src/ via src/selectors/*. Two gates hold it there: facade-symbols.ts pins
./ast to exactly the v0.20.5 list, and package-boundaries.test.ts asserts
src/sdk/selectors.ts is the only file outside the package that imports it.

Restored alongside: the ./selectors export and tsdown entry/chunk group, the
.fallowrc.json entry, the package-exports supported-subpath list, and both
client-api.md sections. No CHANGELOG entry — nothing is removed any more.

pnpm check green: 602 unit files / 5278 tests, smoke 35 passed / 3 live
skipped, layering 71/71 (10 packages, 32 subpaths), depgraph 22/22, mutation
config 45/45, fallow clean over 129 changed files, package smoke imported all
12 published entry points with publint and attw passing. Verified functionally
against the built dist: the doc's parse -> findSelectorChainMatch example
returns the same shapes as before, resolveSelectorChain still returns an AST
`selector`, and formatSelectorFailure still accepts a chain.

* fix(selectors): correct the two expectations that still assume the removal

Review P1s on a792415a: restoring the public subpath left two gates asserting
it was gone.

- installed-package-metro.test.ts moved `agent-device/selectors` into the
  blocked-specifier list. It goes back to the subpath smoke set, running the
  same `isSelectorToken('||')` + `parseSelectorChain` check it ran before the
  removal, so the file's only remaining delta from main is a formatter reflow.
- owner-files-no-leak.test.ts asserted `dist/src/sdk-selectors.js` was absent.
  It requires the stable named chunk again, and still rejects an auto-numbered
  `selectors2.js` fallback — the pair is what proves the restored tsdown chunk
  group is doing its job, verified against a clean build.

PR body corrected: the removal is no longer described as intentional API
tightening.

* refactor(selectors): satisfy the widened fallow scope after rebase

main's #1591 (the follow-up filed from this review) removed `packages/**` from
.fallowrc.json's ignorePatterns, so the new package is audited for the first
time. Everything below is a finding fallow could not previously see.

Dead surface, all confirmed consumer-free:

- 12 type re-exports from the `.` façade whose shapes consumers only ever
  reach structurally.
- MAESTRO_TEXT_SELECTOR_KEYS / MAESTRO_STATE_SELECTOR_KEYS, orphaned by this
  branch's own MAESTRO_SELECTOR_PROJECTION change, and the test-util
  SELECTOR_VALUE_HAZARDS. All three are module-local now.
- IS_PREDICATE_USAGE_HINT fails --production because its only consumer is the
  is-argument-surface parity test. It gets a commented `ignoreExports` entry
  rather than deletion: the constant is what makes the daemon and CLI raise
  ONE hint instead of two copied strings (ADR 0010), so the test asserting
  that is the point, not an accident.

`fast-check` is now declared by the package that imports it.

Duplication, split by what could be proven:

- `isUsefulVisibilityAnchor` existed character-for-character in both
  packages/selectors and packages/maestro. Moved to
  @agent-device/contracts/snapshot, which both already depend on and which
  already owns this vocabulary. Safe because the `normalizeType` each copy
  called is itself character-identical to the contracts one — checked before
  moving, since a different normalizer would have silently changed which
  nodes anchor.
- maestro additionally reimplemented `normalizeType`, `buildSnapshotNodeMap`
  (as `buildSnapshotNodeByIndex`) and `findSnapshotAncestor`, all
  character-identical to contracts'. Deleted in favour of the shared ones.
- The three scroll-ancestor walks are NOT deduped. They are structurally the
  same walk but each uses a different scrollable predicate, and I have no
  evidence the three agree; collapsing them would be a Maestro-conformance
  change, not a cleanup. Both maestro sites now say so, and the work is filed
  separately.

`projectSelectorExpression` (15 cyclomatic / 22 cognitive, written by the
cutover) splits into a dispatcher plus `readAgreedTextValue` and
`projectSelectorTerms`; all three are under threshold.

Rebase note: the one conflict, in package-boundaries.test.ts, resolved to
NEITHER side — #1591 had already deleted `AdReplayVerifiedTargetGuard` as an
unused export, and this branch deletes the seven ReplaySelectorPort names, so
the conflicting block is empty.

* build: record fast-check for packages/selectors in the lockfile

Declaring the dependency in packages/selectors/package.json without
regenerating pnpm-lock.yaml made every CI job fail in its install step with
ERR_PNPM_OUTDATED_LOCKFILE. My local `pnpm install --frozen-lockfile` printed
"+ 1 dependencies were added: fast-check@^4.9.0" and exited 0, which read as
success but was the same mismatch CI refuses.

Regenerated with the pinned pnpm 11.17.0, not the 11.5.3 on this machine:
11.5.3 rewrites peer-dependency resolution keys repo-wide (dropping
`(supports-color@7.2.0)` suffixes) and produced a 222-line diff. With the
pinned version the diff is the 4 lines this change actually needs, plus
pnpm's alphabetical re-sort of the root selectors entry.
2026-08-04 19:06:39 +02:00
Michał Pierzchała 351ef7a14f refactor(android): enforce transport lowering in the type system (#1583)
`AndroidLoweredTouchPlan` widened the canonical two-sample trajectory to a
plain sample array, so a plan that skipped `lowerAndroidTouchPlan` still
satisfied the transport types. That is the mistake the lowering exists to
prevent: an un-lowered plan injects a two-sample gesture, which is the sparse
delivery #1572 removed from the shared plan in the first place.

Transport samples are now a minimum-arity tuple. `sampleGestureOffsets` floors
the frame count at three, so lowering always yields at least four samples,
which makes "denser than the canonical endpoint pair" a true statement about
the data rather than a comment. A canonical plan is no longer assignable, so
skipping the lowering fails typecheck at every injection seam.

Tightening the type caught three call sites that were passing un-lowered plans
straight to the helper transport, which is the evidence the previous signature
enforced nothing. `longPressPlan` now returns `AndroidLongPressTouchPlan`
instead of the wide union it never produced, and the dual-pointer normalize
test routes through the lowering like every other transport call.

Also drops the unused `= 'default'` on `sampleGestureOffsets` so every caller
states which platform sampling convention it wants, which is the point of
having centralized the policy.

Sample values are unchanged by construction, so Android injection stays
bit-identical to #1572.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-04 15:09:27 +02:00
Michał Pierzchała 56d9ee605c fix(ios): preserve timed pan duration (#1572)
* fix(ios): preserve timed pan gesture execution

* fix(gestures): encode linear pans as endpoint plans

* fix(ci): pin wait contract exports

* fix(android): lower endpoint gesture plans for touch transport

* fix(gestures): preserve timed pan duration across adapters
2026-08-04 14:07:53 +02:00
Michał Pierzchała 016577e6bc docs: remove superseded architecture proposals and design prototypes (#1478 P7) (#1580)
P7 cleanup per #1478's ratified defer decision: delete the daemon-modularity
proposal, module-interface-principles (durable kernel folded into CONTEXT.md),
the pre-package Maestro debt map, and the daemon-boundary prototypes with their
package scripts; refresh CONTEXT.md's R10 bullet to post-arc state. Also
reconciles the interaction façade symbol pin with #1570's wait symbols, which
had left check:layering red on main via #1570 × #1574 branch-race skew.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-04 11:14:51 +02:00
Michał Pierzchała 2e74b789fd feat: verify device cloud connections (#1564)
* feat: verify device cloud connections

* refactor: unify connect provider adapters

* refactor: separate connect verification facts

* fix: tighten connect provider verification

* fix: use neutral cloud connection wording

* perf: deduplicate local affected checks

* refactor: simplify affected check runner

* refactor: derive connect workflow from verification
2026-08-03 16:47:57 +02:00
Michał Pierzchała 99967c7f01 fix: restrict project config trust (#1565)
* fix: restrict project config trust

* fix: preserve daemon auth transport context

* refactor: simplify project config trust

* fix: restrict project config write sinks
2026-08-03 15:48:45 +02:00
Michał Pierzchała 480e3883b1 fix(daemon): reject unarmed close --save-script before teardown (#1558)
* fix(daemon): reject unarmed close --save-script before teardown

Live evidence (2026-08-02) showed a plain `open` followed by `close
--save-script` silently published a script: the close request armed
authoring at record time and published moments later in the same
request, folding the never-armed case into the ADR 0016 authoring
lifecycle. The resulting .ad carries selector fallback chains but no
recording-time target-v1 evidence, and nothing told the caller
evidence capture never ran — degraded replay verification with no
signal beats a loud refusal.

`assertTerminalRecordingCloseAllowed` (src/daemon/handlers/session-close.ts)
now rejects an unarmed `close --save-script` with INVALID_ARGS before
any teardown or filesystem work runs, the same seam that already
rejected ABORTED/PUBLISHED terminal recordings. The rejection does not
tear the session down, so a plain `close` retry still completes
cleanly; recovery names `open --save-script` since evidence can only
be captured from action zero. Repair transactions (ADR 0012) are a
disjoint lifecycle and are explicitly unaffected.

This is distinct from #1533 (an already-armed-then-aborted session
whose flag ingress re-enables recordSession and lets a *bare* close
publish); that case remains open.

* fix: review follow-ups for #1558 (help text, test strength, docs)

- Give replay --save-script its own help text instead of the shared
  open/close "arm on open, publish on close" description: replay's flag
  arms an ADR 0012 repair transaction, a disjoint lifecycle. Adds
  CommandSchema.flagDescriptionOverrides so a command can swap a shared
  flag's usageDescription without duplicating the FlagDefinition entry
  (which would have shown --save-script twice in `help replay`). Pinned
  in src/cli/parser/__tests__/cli-help-command-usage.test.ts (open/close
  keep the shared text unchanged; replay gets the new one).

- Strengthen the never-armed close --save-script regression test in
  session-close-shutdown.test.ts: the fixture now carries real
  cleanup-bearing state (an active iOS simulator recording, reusing
  makeIosSimulatorRecordingSession/recordingKillMock) with spies proving
  no teardown hook (recorder kill, runner stop) runs on the rejected
  request, then that a follow-up plain close does tear it down. The
  prior fixture had nothing for teardown to observably touch, so moving
  the guard after stopBestEffortSessionResources would have passed it
  silently. Also fixes a latent test-isolation leak this exposed: an
  earlier test set a persistent mockStopIosRunnerSession rejection
  (vi.clearAllMocks() clears call history, not implementations), which
  would have poisoned any later Apple-platform close test; scoped it to
  mockRejectedValueOnce.

- Point the migration guide (website/docs/docs/migrating-gestures.md) at
  `open --save-script` → interact → `close` instead of the now-rejected
  `open` → interact → `close --save-script`, matching the new guard and
  the corrected help text.

_Generated by [Claude Code](https://claude.ai/code)_
2026-08-03 09:10:40 +02:00
Michał Pierzchała 4c2a30cc8f docs(agents): a green check is evidence only once you have seen it red (#1547)
* docs(agents): a green check is evidence only once you have seen it red

Three vacuous regression tests shipped in one day (an edge-run input the
retired regex handled in one pass, invariants the old implementation
already satisfied, an entry point whose trimming defused the flagged
pattern); review's counterfactual checks caught all three. The same proof
discipline already existed piecemeal for moved tests and structural gates —
name it once and point to the mechanical proof shapes.

Co-Authored-By: Claude <noreply@anthropic.com>

* docs(agents): trim the obvious, keep the earned

Dropped three bullets: open-before-close (CLI help and
device-verification.md own it), don't-remove-without-migration (subsumed by
the stronger no-fallback scope rule), and generic Node built-ins advice
(engines owns the version). Strengthened the oxfmt rule with the confirmed
mechanism: a path argument bypasses ignorePatterns, not just hides drift —
one path-scoped run re-quoted 44 excluded conformance corpus files.

Co-Authored-By: Claude <noreply@anthropic.com>

* Revert "docs(agents): trim the obvious, keep the earned"

This reverts commit 0229cbadba.

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-01 15:11:26 +02:00
Michał Pierzchała b3cf29bc67 feat(ios): reach physical devices through usbmux first, tunnel as fallback (#1517)
* feat(ios): reach physical devices through usbmux first, tunnel as fallback

Physical iOS runner commands now resolve to usbmux whenever the device is
attached by cable, and fall back to the CoreDevice tunnel route only when
usbmuxd reports it unattached (#1403).

Measured on an iPhone 17 Pro: steady-state is a wash between the two routes,
but past the tunnel cache's 30s TTL the network route pays ~4.5s of devicectl
re-probe plus session re-establish on the next command, where the usbmux
session stays hot at ~440ms. Cabled devices now never pay that tax, because
the tunnel lookup, its cache, and the cache invalidation only run on the
fallback path.

Wi-Fi-only devices keep working: modern CoreDevice Wi-Fi runs over remoted and
never appears in usbmuxd, so the unattached verdict routes them to the tunnel.
That verdict is carried by usbmuxDeviceAttached:false and answered inside the
same connect attempt rather than by burning a retry, so an XCTest-backed
device — which has no tunnel — now surfaces its cable/trust/unlock hint
instead of retrying for the full budget.

Replaces the AGENT_DEVICE_IOS_RUNNER_ROUTE experiment override from #1510.

* fix(ios): make the unattached usbmux verdict terminal for xctest devices

waitForRunner recorded the unattached verdict as a generic connect failure and
retried it for the whole budget, so readiness preflight and read-only commands
on an XCTest device still burned 2x45s and lost the cable/trust/unlock hint —
the exact hang #1510 measured, which this PR claimed to fix but only fixed on
the sendRunnerCommandOnce path.

Retrying cannot attach a cable and an XCTest device has no tunnel to fall back
to, so the typed verdict is now thrown from the attempt, excluded from the
connect retry policy, and passed through waitForRunner unwrapped. The predicate
moved to runner-contract.ts because the retry policy needs it and importing the
transport there would close an import cycle.
2026-07-31 11:31:20 +02:00
Michał Pierzchała 6b0b9cb0ef feat(ios): usbmux runner route override + #1403 transport experiment evidence (#1510)
* feat(ios): usbmux runner route override + #1403 transport experiment evidence

Adds AGENT_DEVICE_IOS_RUNNER_ROUTE=usbmux, an experimental override that
routes coredevice-backend physical devices' runner commands through usbmux,
and simplifies the xctest branch that awaited a no-op resolveRunnerTransport.

Documents the live #1403 experiment (iPhone 17 Pro, USB + Wi-Fi legs):
steady-state is a wash, but the >30s-idle tax drops from ~4.5s (tunnel
re-probe + session re-establish) to ~440ms because the usbmux session stays
hot; CoreDevice Wi-Fi devices never appear in usbmuxd, so the verdict is
usbmux-primary with network fallback rather than tunnel-code deletion.
Also records the cable-out failure gap (2x45s retry hang swallowing the
usbmux DEVICE_NOT_FOUND hint), which affects today's xctest backend too.

* fix(ios): document daemon scoping of the usbmux route override

The env is re-read per resolve but from the daemon's environment, which is
captured at daemon launch — a later CLI invocation cannot flip the route on
a running daemon. Correct the source comment and experiment doc, and lock
the read-point semantics with a regression test.
2026-07-31 08:36:19 +02:00
Michał Pierzchała 0e51007b04 refactor: isolate maestro engine package (#1506)
* refactor: isolate maestro engine package

* perf: deepen maestro facade boundaries
2026-07-30 20:58:20 +02:00
Michał Pierzchała 0ee2a86129 refactor: extract contracts workspace package (#1499)
* refactor: extract contracts workspace package

* fix: preserve screenshot diff result contract

* test: stabilize Android keyboard smoke
2026-07-30 17:07:46 +02:00