1019 Commits

Author SHA1 Message Date
Michał Pierzchała 5c08f1664f 0.19.2 v0.19.2 2026-07-09 19:32:16 +02:00
Michał Pierzchała a3885351c2 fix: stabilize android maestro gestures (#1171)
* fix: stabilize android maestro gestures

* fix: address maestro android gesture review
2026-07-09 19:31:11 +02:00
Michał Pierzchała 8878399272 feat: include unchanged interactive refs in settle output (#1167)
* feat: include unchanged interactive refs in settle output

Benchmarks (gpt-5.4-mini + claude-haiku, July 2026) showed 27% of --settle
actions were followed by a fallback snapshot -i because a change-only diff
omits refs for elements that did not change: after a modal dismiss the diff
shows only removals, so the next button to press is invisible.

Add an unchanged-interactive tail to SettleObservation, attached only when
the diff's added lines carry zero refs (the modal-dismiss/toast-only
signature). It lists the settled tree's remaining hittable, uncovered
elements so the response stays actionable without an extra round trip.
Rides the CLI text, MCP digest view, ref pinning, and output schema the same
way the diff's added-line refs already do.

* refactor: address fallow audit findings on the settle tail

- drop the unused export on buildSettleTail (tests exercise the trigger
  through the public interaction path and the filter via
  buildSettleTailEntries)
- extract the digest tail capping from interactionSettleView into a
  module-private helper to stay under the complexity gate
2026-07-09 19:18:47 +02:00
Michał Pierzchała 9a2277c045 fix: alias launch/relaunch to open and suggest canonical commands for unknown names (#1166)
* fix: suggest canonical commands for unknown command names

Agents commonly guess command names that don't exist, e.g. relaunch/launch
instead of `open <app> --relaunch`, burning turns on Unknown command errors
that only say "run --help". Add a curated alias-to-canonical-shape map for
the most common guesses (launch/relaunch/start/restart, touch, input/
settext/entertext, screencap/capture, dismiss), backed by a nearest-name
edit-distance fallback derived from the live command registry so
suggestions can't drift. Also hint that `open` takes the app/bundle id as
a positional when an unknown flag looks like a bundle-id guess (e.g.
--bundle-id), and apply the same suggestion to `help <unknown>`.

Suggestions are display-only; nothing auto-executes and the error code
stays INVALID_ARGS.

* fix: address review — dead export, case-insensitive suggestions, tighter nearest-name matching

- Drop the export on getNearestCommandNames (module-private; only
  suggestCommandFor uses it) to satisfy the Fallow unused-export gate.
- Lowercase the input token before both the curated-map lookup and the
  nearest-name pass, so RELAUNCH/Relaunch/TAP/Touch get the same hint as
  their lowercase forms. Added a curated `tap` entry: lowercase `tap` is
  normalized to press before the unknown-command check, so the entry only
  catches case variants like TAP.
- Tighten the nearest-name fallback: exact prefix matches win outright
  (`clos` now suggests only `close`, not "one of: close, logs"),
  otherwise only ties at the minimum edit distance are kept, and 1-2
  character tokens never get a suggestion (`ls` no longer suggests `is`).
- Share the "open <app> --relaunch" example string between the curated
  map and the unknown-flag hint, and extend the registry-drift tests to
  parse each curated example end-to-end (validates open --relaunch as a
  registered flag) plus assert keyboard dismiss is a real keyboard action.

* feat: promote launch and relaunch to true open aliases

Follow the tap -> press precedent: `relaunch <app>` now runs
`open <app>` with --relaunch injected, and `launch <app>` runs a plain
`open <app>` (no forced restart — that would silently destroy app
state). Both are normalized in normalizeCommandAlias before parsing, so
command identity stays `open` for daemon requests and telemetry, all
other args/flags pass through to open's normal validation (URL targets
still get the daemon's existing --relaunch guidance), and an explicit
--relaunch stays idempotent. Alias matching is now case-insensitive
(TAP, RELAUNCH, Launch), so the curated tap suggestion entry is dead
and removed along with launch/relaunch; start/restart and the rest of
the map stay suggestion-only since start is genuinely ambiguous.
2026-07-09 19:00:41 +02:00
Michał Pierzchała 9dbecd02a6 fix: make manual-qa help self-sufficient and slim routine help paths (#1168)
* fix: make manual-qa help self-sufficient and slim routine help paths

July 2026 benchmarks (gpt-5.4-mini/haiku/sonnet driving the CLI) showed
agents burning 14-41KB of tool output on help before a routine QA flow:
help react-native routed "generic navigation, selectors, refs,
verification" to help workflow (31.7KB), one run read the bare
`agent-device help` mid-task (16.3KB), and several fell back to
per-command help because manual-qa lacked concrete command shapes.

- manual-qa gets a "Command shapes" block covering the full routine QA
  loop (open --relaunch, deep-link open, snapshot -i, press/fill --settle,
  wait text, close) plus a quoting note for apostrophe/quote labels, so
  agents don't need to read help workflow or per-command help for a
  normal pass.
- react-native's routing block no longer sends routine QA to help
  workflow; it now points routine flows at manual-qa and reserves
  workflow for deep exploration/debugging.
- physical-device and the shared topic-footer routing also point at
  manual-qa alongside workflow.
- Trimmed the bare `agent-device help` output: cut Agent Quickstart
  lines that duplicated topic-specific detail already owned by
  react-native/workflow/remote/web/tv, dropped 3 redundant Examples, and
  fixed an alignment bug where one long usage string (tv-remote) forced
  padding whitespace onto every other Commands: row. The "Default app
  loop" line (shown to flip Haiku from 0/3 to 3/3 in earlier benchmarks)
  stays intact and early.

* fix: satisfy oxfmt and restore pinned-ref staleness nuance in help workflow

- pnpm format:check failed on the new renderAlignedSection columnWidth
  line; reformatted with pnpm format.
- Reviewer noted the only content genuinely lost from the bare-help
  Quickstart trim was the "pinned refs get exact staleness warnings"
  nuance; restored it as one line in help workflow's Snapshots and refs
  section (its natural owner), keeping the bare-help line short.
2026-07-09 18:46:33 +02:00
Michał Pierzchała c325842f62 fix: reap idle daemons and take over stale runner leases (#1169)
* fix: reap idle daemons and take over stale runner leases

Each AGENT_DEVICE_STATE_DIR spawns its own daemon that never exits on its
own; deleted codex/claude sandboxes leave orphaned daemons accumulating
(10+ observed). A stale-but-still-running orphan also keeps holding its
iOS runner lease, so a fresh daemon for the same device fails with
COMMAND_FAILED "already owned by another agent-device daemon".

- Daemon self-reaps after an idle window (default 5 minutes, matching the
  iOS runner idle-stop default) once it has no open sessions, no
  in-flight requests, and no active recording. AGENT_DEVICE_DAEMON_IDLE_TIMEOUT_MS
  overrides the window; 0 disables it.
- A runner lease whose owner PID is dead, or whose owner
  AGENT_DEVICE_STATE_DIR no longer exists, is now reclaimed automatically
  instead of erroring; a genuinely live owner with an existing state dir
  still gets the existing rejection + hint.

* fix: gate lease takeover on proven-dead owners and fail closed on stat errors

Review follow-ups on #1169:

- Replace fs.existsSync (which never throws and swallows EACCES/IO errors
  into "gone") with fs.statSync + error-code inspection: only ENOENT/
  ENOTDIR count as proof the owner state dir is gone; any other stat
  error fails closed and classifies the owner as alive.
- Split stale classification by reason (owner-process-dead vs
  owner-state-dir-gone). Adoption (readStaleRunnerLease ->
  tryAdoptRunnerSessionFromLease) is now strictly PID-dead-gated:
  a dir-gone-but-alive owner may still hold a live runner connection,
  so its lease routes through the force-stop path (kill leased runner
  processes, rebuild) instead of being silently adopted - no two
  masters.
- Tests: EACCES stat error keeps the busy rejection; dir-gone+PID-alive
  refuses adoption before probing; dir-gone force-stop asserts a fresh
  runner launch instead of adopting the old runner pid.

* test: make idle reap tests deterministic
2026-07-09 18:42:03 +02:00
Michał Pierzchała 3667f6ece5 fix: point unknown selector keys at role=/label= forms (#1165)
* fix: point unknown selector keys at role=/label= forms

press 'button="Push Article"' errored with a nonsense suggestion
(text="button=\"Push Article\"") because isSelectorToken rejects
`button` as a key, splitSelectorFromArgs returns null, and the point
fallback wraps the whole raw token in text=.

Add detectUnknownSelectorKeyToken to spot a key=value token whose key
isn't a recognized selector key, and use it in readPointTarget to
throw a targeted error before the numeric parse: role=<key>
label=<quoted value> when the key looks like an accessibility role
word (isRoleHintWord, mirroring ROLE_LABELS), otherwise label=<quoted
value>. wait-positionals.ts has no analogous point-fallback path, so
it needs no change.

* fix: fold unquoted multi-word values into the unknown-key suggestion

Review follow-up: readPointTarget only inspected positionals[0], so an
unquoted multi-word value split across positionals (press 'button=Push'
'Article') dropped the trailing tokens and confidently suggested the
wrong completion (label="Push"). Fold trailing positionals into the
suggested value like mergeRestIntoSelectorValue does, unless the value
was fully quoted (button="Push Article") and therefore complete — that
distinction keeps fill's trailing text argument out of the suggestion.

Also: drop the dead `text` entry from ROLE_HINT_WORDS (valid selector
key, short-circuits in ALL_KEYS first) and note the set is a superset
of ROLE_LABELS rather than a mirror; reject whitespace-only values in
detectUnknownSelectorKeyToken.
2026-07-09 18:28:29 +02:00
Michał Pierzchała 477c684cf6 fix: read runner cache package version from project root (#1170) 2026-07-09 14:19:05 +02:00
Michał Pierzchała ff7fc5ba61 0.19.1 v0.19.1 2026-07-08 21:49:10 +02:00
Michał Pierzchała 888984169b fix: make record app-scoped by default (#1163)
* fix: reject recording for failed iOS simulator session

* fix: make record app-scoped by default
2026-07-08 21:48:35 +02:00
Michał Pierzchała 0d2b0353ff docs: clarify agent setup and text entry guidance (#1164) 2026-07-08 21:36:33 +02:00
Szymon Dziedzic cfef0a4bca feat: add session event timeline (#1032)
* feat: add session event timeline

* fix: support cursor-only event reads

* refactor: simplify event log formatting

* refactor: trim event log helpers

* docs: document session event timeline

* refactor: tighten session event log internals

* fix: redact event log action positionals by default

* fix: align event log after rebase

* test: cover events in provider output guard

* fix: harden session event privacy

* fix: harden event message redaction

* fix: harden session event logging

---------

Co-authored-by: Michał Pierzchała <thymikee@gmail.com>
2026-07-08 21:35:45 +02:00
Michał Pierzchała cf31fb3f7b fix: harden iOS XCTest recovery paths (#1158) 2026-07-08 21:11:28 +02:00
Michał Pierzchała 2717047b86 fix: normalize iOS simulator screenshot density (#1160)
* fix: normalize iOS simulator screenshot density

* fix: avoid density metadata after screenshot downscale

* fix: satisfy screenshot density CI gates

* fix: harden screenshot metadata collection

* refactor: centralize screenshot density policy

* refactor: reuse screenshot density support check
2026-07-08 21:05:24 +02:00
Michał Pierzchała e4115ec5ab chore: migrate to TypeScript 7 (#1161) 2026-07-08 21:04:54 +02:00
Michał Pierzchała 106c238697 fix: narrow client result contracts (#1155) 2026-07-08 18:35:27 +02:00
Michał Pierzchała f21727d065 fix: handle Android IME overlays in snapshots (#1157)
* fix: handle Android IME overlays in snapshots

* fix: satisfy Android IME CI guards

* fix: detect localized Gboard tutorial overlays

* fix: keep Android IME overlay handling passive
2026-07-08 18:29:47 +02:00
Michał Pierzchała 9dabe5b1c1 refactor: derive command identity from descriptors (#1151)
* refactor: derive client-backed cli routing

* refactor: derive command identity from descriptors
2026-07-08 17:55:00 +02:00
Michał Pierzchała b91eaad885 refactor: make iOS synthesized gesture policy explicit (#1152)
* refactor: make iOS synthesized gesture policy explicit

* test: harden settle observation under coverage

* fix: preserve first-command synthesized drag behavior

* refactor: simplify synthesized frame policy

* refactor: inline synthesized command policies

* refactor: simplify sequence synthesized context

* refactor: clarify synthesized drag fallback policy

* refactor: keep synthesized gesture policy runner-local
2026-07-08 17:15:42 +02:00
Michał Pierzchała f18d0b2e92 fix: improve settle observation guidance (#1154) 2026-07-08 17:12:19 +02:00
Michał Pierzchała bbc577c11a fix: derive interaction response data transforms (#1149)
* fix: derive interaction wire projection

* fix: derive wire projection from command descriptors

* refactor: clarify response data transform naming

* test: guard response transform field ownership
2026-07-08 14:13:15 +02:00
Michał Pierzchała b0c70ad4e4 feat: support repack dev server prepare (#1145) 2026-07-08 11:14:36 +02:00
Michał Pierzchała 694266d802 fix: keep iOS synthesized drags off AX (#1148)
* fix: keep iOS synthesized drags off AX

* fix: address iOS synthesized drag review
2026-07-08 11:06:57 +02:00
Michał Pierzchała 7f61df30ae feat: add TV remote command (#1147)
* feat: add TV remote command

* feat: improve TV remote ergonomics

* test: cover tv-remote provider scenario

* fix: preserve focused Android TV nodes

* docs: tighten PR description guidance

* fix: remove d-pad command alias

* docs: clarify tv-remote hold syntax

* feat: add tv-remote longpress CLI sugar
2026-07-08 10:59:48 +02:00
Michał Pierzchała bc7dcc8345 fix: keep XCTest tree snapshots on main (#1144)
* fix: keep XCTest tree snapshots on main

* fix: address iOS runner snapshot review
2026-07-07 16:59:00 +02:00
Michał Pierzchała 9dfebbe3be fix: omit interaction uptime from wire responses (#1142)
* fix: omit interaction uptime from wire responses

* fix: simplify interaction wire sanitizer
2026-07-07 16:53:44 +02:00
Michał Pierzchała 69a8f6f3fe fix: soften recovered snapshot warning (#1146) 2026-07-07 16:52:25 +02:00
Michał Pierzchała ef4b66d4dc test: remove slow-test ratchet pins (#1143) 2026-07-07 14:46:02 +02:00
Michał Pierzchała 9009c5aff7 0.19.0 v0.19.0 2026-07-07 13:17:01 +02:00
Michał Pierzchała 8f28c31c86 fix: recover completed Android recording from pending-only manifest (#1141)
* fix: recover completed Android recording from a pending-only manifest

When the daemon crashes in the brief window between writing the pending recovery
manifest (before screenrecord starts) and upgrading it to a `current` manifest, the
screenrecord process can still finish and leave a complete MP4 on the device. record
stop previously discarded it as stale because the pending-only recovery path never
checked for an on-device file, unlike the `current` path which already recovers a
finished recording. Extend the pending-only path to recover the completed file with the
same finished-recording warning, and skip the stop signal when the recovered recording
has no tracked pid (a pending chunk never records one, and probing an empty pid is
unsafe).

* fix: treat JSON arrays as invalid Android recovery manifests

isRecord accepted arrays (typeof [] === 'object'), so a stray `[]` recovery manifest
was classified as blocked rather than deleted, wedging every subsequent record stop.
Reject arrays and null so a non-object manifest is cleaned up like other malformed
metadata.
2026-07-07 12:04:58 +02:00
Michał Pierzchała 0dcc1aa553 fix: normalize interaction response wire shapes (#1114)
* fix: normalize interaction response wire shapes

* fix: reduce interaction response complexity
2026-07-07 11:29:01 +02:00
Michał Pierzchała 7e583c4136 fix: simplify Android recording recovery (#1135)
* fix: harden android recording recovery

* fix: reduce android recording recovery fallow complexity

* test: fix android recording recovery rebase

* fix: block uncertain android recording fallback

* fix: address android recording recovery review

* fix: address android recording recovery review followup

* fix: simplify Android recording recovery

* fix: address android recovery ownership review

* refactor: reuse android recovery manifest helpers

* refactor: split android pending recovery resolution

* fix: clarify scoped android recovery hint
2026-07-07 10:50:57 +02:00
Michał Pierzchała d5a7af0f4d fix: hint on iOS runner main-thread timeouts (#1140) 2026-07-07 10:26:49 +02:00
Michał Pierzchała b052eb9a37 fix: speed up iOS text entry (#1139)
* fix: speed up iOS text entry

* fix: reverify settled iOS text entry
2026-07-07 09:50:40 +02:00
Michał Pierzchała 8ef4e73408 refactor: derive command exposure lists from descriptors (#1137) 2026-07-07 08:00:33 +02:00
Michał Pierzchała 5c5fa012f7 feat: --settle returns the settled diff in the interaction response (#1101) (#1106)
* feat: --settle returns the settled diff in the interaction response (#1101)

press/click/fill/longpress --settle executes the action, waits for the UI
to go quiet (wait stable's loop, shared via stable-capture.ts), and returns
the settled diff vs the pre-action tree in the same response — one round
trip instead of the interact -> observe pair.

- payload: changed lines only (bounded), summary counts, added-line refs,
  refsGeneration; best-effort (settled:false + hint on never-quiet content,
  never an action failure); --verify shares the settle captures
- ref issuance: the settled tree becomes the session snapshot; a
  diff-carrying settle response clears snapshotRefsStale and the MCP layer
  merge-only re-pins added-line refs at the settle generation
- grammar: --settle + --settle-quiet <ms> + --timeout <ms> (flag-sourced
  descriptor budget with new envelope:'widen' semantics mirroring wait)
- ADR 0011: new settleObservation guarantee classified on every path with
  contract scenarios per enforced/delegated cell

* test: give the two contention-flaky doctor scenarios explicit budgets

The doctor provider scenarios sit at ~5s of real daemon-harness work on a
loaded host and flake at vitest's 5s default during full-suite runs (the
known contention flake AGENTS.md documents). Same in-file precedent as the
Metro-probe scenario's 10s budget.

* fix: move SettleParams to contracts to satisfy the layering DAG

daemon/handlers/interaction-flags.ts imported the type across the
daemon -> commands boundary (R2 commands-floor). The tuning params are
part of the interaction contract like SettleObservation, so they live
in contracts/interaction.ts and both layers import from there.

* feat: keep settle diffs content-first — drop Key nodes, added lines win the cap

Bluesky dogfood: a fill that summons the iOS keyboard spent 49 of the 80
capped diff lines spelling out QWERTY keys, and a screen transition with
269 removals could starve out the added lines entirely. Key-type nodes
are now filtered from both diff sides (the [keyboard] container line
still signals presence), and under truncation added lines — the ones
carrying fresh refs — win slots over removals.

* docs: state the core loop in the top-level help starting point

Benchmarked with headless haiku/sonnet agents given only --help: both
models skipped the help-workflow pointer and started with plain
snapshot (38KB payloads they then had to re-read from files). One
core-loop line at the starting point is what teaches snapshot -i and
--settle to models that never read a second help page.

* fix: preserve settle digest refs for mcp

* fix: reduce settle fallow complexity

* fix: surface settle output in CLI text

* fix: complete settle handling for longpress

* refactor: localize daemon timeout envelopes

* refactor: deepen post-action observation

* refactor: centralize post-action observation planning

* refactor: derive settle capability from descriptors

* refactor: trim settle descriptor helpers
2026-07-06 20:18:44 +02:00
Michał Pierzchała db69124c00 feat: add capabilities command (#1133) 2026-07-06 19:35:46 +02:00
Michał Pierzchała 54f6d45b32 refactor: extract host process primitives (#1134) 2026-07-06 19:01:32 +02:00
Michał Pierzchała 126d85ac8a fix: reap managed web browser orphans (#1112)
* fix: reap managed web browser orphans

* fix: narrow managed web browser reaping

* fix: harden managed web cleanup
2026-07-06 15:51:36 +02:00
Michał Pierzchała be4bd092b6 fix: recover Android recordings after daemon restart (#1129)
* fix: recover android recordings after daemon restart

* refactor: reduce recording recovery complexity

* fix: address android recording recovery feedback
2026-07-06 15:28:27 +02:00
Michał Pierzchała 7475415ac5 build: strip runner unit-test blocks from package (#1128)
* build: strip runner unit-test blocks from package

* ci: fix package size and fallow checks

* fix: resolve packaged recording scripts

* fix: keep recording script resolver internal
2026-07-06 12:52:05 +02:00
Michał Pierzchała ea69dc3767 fix: harden apple runner recovery (#1126)
* fix: harden apple runner recovery

* fix: address runner recovery ci failures

* fix: clean up runner recycle review findings
2026-07-06 11:40:54 +02:00
Michał Pierzchała f9721e7e8b test: split the Android platform test aggregation and share the scripted adb stub (#1103)
* test: split the Android platform test aggregation and share the scripted adb stub

AGENTS.md names the platform index.test.ts aggregations as offenders to
shrink opportunistically; this splits the 2,735-line Android one along
its (already well-factored) source modules, every test moved verbatim
(92 tests before and after):

- ui-hierarchy.test.ts (22): parseUiHierarchy/androidUiNodes
- app-lifecycle-install.test.ts (13): install/resolve/infer/launch
  component parsing
- app-lifecycle-open.test.ts (19): open/close, deep links, launch args,
  TV category, fallback resolve-activity
- input-actions.test.ts (11): type/fill/swipe/scroll/rotate
- settings.test.ts (14): appearance/clear-app-state/fingerprint/
  permissions
- notifications.test.ts (2), app-parsers.test.ts (1)
- keyboard state/dismiss tests (10) appended to the existing
  device-input-state.test.ts

Consistency fix folded in: the file carried a local withMockedAdb fork
because it needs scripted per-subcommand adb responses, which the shared
arg-recorder helper cannot express. The fork now lives in
src/__tests__/test-utils/mocked-binaries.ts as withScriptedAdb next to
withMockedAdb, and hands each call a fresh copy of the shared
ANDROID_EMULATOR fixture.

The copy matters: the Android TV test mutated the callback's device
(device.target = 'tv'), which the old per-call object literal absorbed
silently. With a shared fixture that mutation leaked into the next test
and flipped its launch to LEANBACK. The helper now clones per call and
the TV test builds { ...device, target: 'tv' } instead of mutating.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqeW8sA2ZnvnftdvpCqFMS

* test: serialize the scripted-adb group and repoint its slow-test pins

Review follow-up for the android index.test.ts split: the monolith
implicitly serialized the env-mutating adb-stub tests (PATH,
AGENT_DEVICE_TEST_ARGS_FILE) in one worker, and the split let vitest
run them across parallel files. Make the contract explicit:

- new android-adb vitest project runs the six scripted-adb test files
  in a single fork (singleFork), keeping the pre-split execution
  semantics; ui-hierarchy and app-parsers stay in the parallel unit
  project (pure parsing, no env mutation)
- test/test:unit scripts run both projects
- the five slow-test ratchet pins that referenced index.test.ts keys
  now point at the split file names, so the pinned real-time offenders
  keep their exemption instead of failing at 2x budget under load; the
  reporter's own pinned-key fixture updated to match

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqeW8sA2ZnvnftdvpCqFMS

* test: use vitest 4 android adb serialization

* docs: update unit project readiness guidance

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-06 11:23:03 +02:00
Michał Pierzchała 8a8ebcafda fix: name split declaration chunks (#1127) 2026-07-06 11:17:49 +02:00
Michał Pierzchała e4139b6802 ci: deepen node 22 packaged smoke (#1125) 2026-07-06 10:24:12 +02:00
Michał Pierzchała 318e813cc8 test: split interaction runtime tests (#1113) 2026-07-06 08:22:45 +02:00
Michał Pierzchała 915f5ba374 build: support packaged CLI on Node 22.12 (#1116) 2026-07-06 07:59:37 +02:00
Michał Pierzchała 83d54614d8 fix: bound iOS capture stalls and make runner recovery session-preserving (#1105) (#1107)
* fix: bound iOS capture stalls and make runner recovery session-preserving (#1105)

Runner (Swift):
- Coalesce duplicate transport sends of one commandId onto the in-flight
  execution instead of enqueueing them again behind it (capture pileup).
- Fail fast with RUNNER_BUSY while watchdog-abandoned main-thread work is
  draining; escalate to RUNNER_WEDGED past 120s so the daemon recycles.
- Carry the capture-plan deadline into the query-sweep and private-AX
  ladder tiers so chained recovery cannot stack past the watchdog.
- Penalize the tree backend after a slow (>5s) or abandoned capture and
  lead subsequent regular plans with private-AX for that bundle (sticky,
  120s), stamped recovered/budget so the deferral stays observable.

Daemon (TS):
- Per-request runner recycle budget: at most one invalidate+reboot per
  request, then fail fast with an actionable, session-preserving hint.
- RUNNER_WEDGED joins the runner-fatal invalidation reasons.
- Interaction commands (click/fill/longpress/press/type/get/is) preserve
  the daemon on request timeout like snapshot/wait/find: resetting it
  destroyed every healthy app session the daemon owned.

* fix: suppress AX-broken-screen snapshot issues so the runner survives capture

XCTest records 'Failed to get matching snapshot: kAXErrorIllegalArgument'
issues for every XCUIApplication query on AX-broken screens; after a few
of them the test case tears down the moment the in-flight command
completes, killing the long-lived runner after every capture of the
screen (the restart loop behind #1105). The capture plan already
classifies and recovers from AX failures, so this issue class is noise:
swallow exactly it in record(_:); everything else still records and
still drives XCTEST_RECORDED_FAILURE.

* feat: time-slice the XCTest tree capture on a worker thread

The tree snapshot XPC is a single blocking call whose duration moves
with live content (4s to minutes on Bluesky profile screens); no
in-process budget could bound it on the main thread. Run it on a worker
bounded to an 8s slice: on timeout the plan penalizes the tree backend,
skips the XCTest-backed tiers while the abandoned XPC drains (they
would block behind it inside testmanagerd), and recovers through the
private AX backend, which does not use testmanagerd.

* tune: lower the tree-backend penalty threshold to 3s

The Bluesky profile tree grind measures ~4.5s before kAXErrorIllegalArgument,
just under the old 5s threshold, so every capture re-paid the doomed grind
(9s each). At 3s the second capture onward defers to private AX (2.4s
snapshot, 4.9s press on the live repro).

* fix: harden the AX-issue suppression per review

- Require the kAXError token: 'Failed to get matching snapshot: Timed out
  while evaluating UI query.' is a genuinely-hung-query signal and must
  keep recording (and keep driving XCTEST_RECORDED_FAILURE). Sibling AX
  server codes (kAXErrorCannotComplete, ...) are deliberately included:
  any AX-server rejection inside a matching-snapshot fetch is the same
  capture-plan noise.
- State honestly that the override is suite-global and why (tap-triggered
  queries record the same noise; command outcomes stay honest via their
  own error paths).
- Lock-guarded suppressed-issue counter following the file's existing
  abandoned-work counter pattern, logged with each suppression.
- Unit-test the pure classifier (record(_:) itself is not invoked: the
  must-record variants would record real failures in the test run).
2026-07-05 10:08:15 +02:00
Michał Pierzchała a0556583c3 docs: tighten agent operating guide (#1104)
* docs: tighten agent operating guide

* docs: update agent guide cli paths
2026-07-05 08:17:34 +02:00
Michał Pierzchała 0159975f2e test: split the args.test.ts aggregation along source topology (#1102)
* test: split the args.test.ts aggregation along source topology

AGENTS.md file-size tripwires now apply to tests with no exemption, and
test files are expected to mirror source topology 1:1. args.test.ts was
a 2,503-line aggregation in src/utils/__tests__ while the code it
exercises lives in src/cli/parser. Split it into six focused files with
every test moved verbatim (142 tests before and after):

- src/cli/parser/__tests__/args-parse-interaction.test.ts (29 tests):
  parseArgs shapes for press/click/swipe/gesture/type/record/screenshot
  and friends
- src/cli/parser/__tests__/args-parse-session.test.ts (41 tests):
  parseArgs shapes for session/daemon/device flags, passthrough,
  install/metro/connect/proxy/auth and friends
- src/cli/parser/__tests__/args-validation.test.ts (17 tests): strict/
  compat modes, rejections, deterministic errors
- src/cli/parser/__tests__/cli-help-topics.test.ts (15 tests): global
  usage and help topics
- src/cli/parser/__tests__/cli-help-command-usage.test.ts (35 tests):
  per-command usage copy
- src/utils/__tests__/command-schema-guards.test.ts (5 tests): schema/
  catalog/capability guards and the cli.ts dispatch-literal walk (the
  oxc-parser helpers live here)

AGENTS.md testing-matrix and help-source pointers updated to the new
paths, including the stale src/utils/cli-help.ts and cli-flags.ts
locations (both live under src/cli/parser/).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FqeW8sA2ZnvnftdvpCqFMS

* docs: fix cli parser paths in agent guide

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-07-04 21:46:10 +02:00