1032 Commits

Author SHA1 Message Date
Michał Pierzchała fb7dbfe6fe 0.19.3 v0.19.3 2026-07-10 18:01:57 +02:00
Michał Pierzchała 6e21fedc08 refactor: remove production-unused exports (#1203) 2026-07-10 18:00:43 +02:00
devin-ai-integration[bot] 47134bf764 feat: add derived fail-open check:affected selector (#1195)
* feat: add derived fail-open check:affected selector

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor: simplify selector for complexity gate; add docs

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: fail open on ambiguous non-source fixtures; guard catalog against real package.json/vitest.config

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* refactor: use src/utils/exec.ts process helpers in check:affected runner

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(check:affected): SkillGym ownership, honest catalog, working-tree discovery

- Add SkillGym ownership for skills/ and test/skillgym/; stop short-circuiting
  their Markdown as docs-only (findings 2 & 4).
- Drop the fabricated GitHub 'SkillGym' job: it is a local-only gate, now
  localRunnable with no CI job, guarded by a workflow-existence self-test (3).
- Fold working-tree (staged/unstaged/untracked) state into local discovery and
  disable rename detection so both rename paths classify (1).
- Add run.test.ts entrypoint regressions (real diff/status/rename discovery,
  --run order/skip/stop-on-failure).

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(check:affected): union staged + unstaged diffs so they cannot cancel

A single `git diff HEAD` nets index against working tree, so a staged add
and an unstaged delete of the same file cancel and hide it. Collect
`--cached` (staged) and unstaged diffs separately and union them; add a
cancellation regression test.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(check:affected): cover required suite gates

* refactor(check:affected): delegate tests to vitest

---------

Co-authored-by: Michał Pierzchała <thymikee@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-10 17:53:52 +02:00
Michał Pierzchała fb1117f229 ci: ratchet against production-unused exports (#1202)
* ci: ratchet against test-only exports

Three exported-and-unit-tested-but-unreferenced-in-production incidents
this week (#1166 getNearestCommandNames, #1167 buildSettleTail, #1199
clearMetroSessionHints) — the first two were caught by fallow's dead-code
check because they had zero importers anywhere; #1199 was missed because a
test file imports the export, and fallow's default reachability graph
counts a test import as "used".

Adds a second, stricter pass reusing fallow's own --production mode
(entry.exclude test/story/dev files) via scripts/test-only-exports/check.ts:
an export alive in fallow's default graph but dead in its production graph,
with no other reference anywhere in its own file, has no production call
site — exactly the #1199 shape. Ratchets against a checked-in baseline
(scripts/test-only-exports-baseline.json, 77 entries); new findings fail
`pnpm check:test-only-exports` (wired into CI's Fallow job and
check:tooling). A `// test-seam: <reason>` comment above an export is the
escape hatch for intentional test seams.

Also extends .fallowrc.json's ignoreExports for seven daemon route handlers
(src/daemon/handlers/*.ts) that are genuinely production-reachable through
request-handler-chain.ts's `typeof import()` lazy-load pattern, which
fallow's static import graph can't trace as a named-export consumer —
without this they were false positives in the production-mode pass.

* fix: harden test-only-exports ratchet per review

Addresses the two should-fixes and all five minors from the independent
review of #1202:

- Replace the regex own-file occurrence count with an oxc-parser AST walk
  (typescript@7 ships no JS scanner API, so the review's fallback tool
  suggestion is the primary): identifiers are counted as AST nodes deduped
  by source span, so mentions in JSDoc/block comments, strings, and
  template-literal text no longer masquerade as call sites (review finding
  1, both constructed cases re-verified fixed), and a `//` inside a string
  no longer hides real usages (finding 6). Span dedupe keeps barrel
  re-exports (`export { x } from`) counting once. The sharper count
  surfaced one organic false negative on main: `selector` in
  src/commands/index.ts was previously exempted because the regex matched
  "selector" inside the './...selector-read.ts' import path string; it is
  now baselined alongside its sibling `ref` (same re-export line).
- Make the baseline shrink-only (finding 2): --update-baseline refuses new
  findings with the same wire/delete/annotate message, so the `// test-seam:`
  annotation in the reviewed source diff is the only acceptance path;
  CONTRIBUTING no longer documents baseline regeneration as an acceptance
  option and now describes baseline growth as a deliberate manual edit.
- Stale baseline entries now emit a `::warning` CI annotation (finding 3).
- Commit a re-runnable fixture test (finding 4): check.test.ts mirrors
  scripts/layering/model.test.ts, builds a synthetic package with a
  clearMetroSessionHints-shaped export (JSDoc self-mention included),
  asserts it is flagged, and asserts the annotated twin passes; wired
  before the check in pnpm check:test-only-exports.
- Mark the unreadable/unparseable-file fallbacks CONSERVATIVE: per
  CONTRIBUTING's convention (finding 5).
- Document the dynamic property access (obj[name]) blind spot in the
  script header and CONTRIBUTING (finding 7).

* fix: harden test-only export ratchet

* refactor: use native Fallow export gate

* chore: refresh production export baseline
2026-07-10 17:14:54 +02:00
Michał Pierzchała f53d572f87 fix: align Maestro swipe semantics across platforms (#1179)
* fix: preserve explicit Android Maestro swipe lanes

* fix: align Maestro swipe semantics across platforms

* fix: avoid replaying iOS Maestro gestures

* refactor: make swipe coordinate policies explicit
2026-07-10 16:41:54 +02:00
devin-ai-integration[bot] db492cbaed test(output-economy): add routine-workflow output-behavior oracle (#1190)
* test(output-economy): add routine-workflow output-behavior oracle

Add a deterministic routine-workflow measurement (#1180) that pairs
response bytes with follow-up behavior: fallback-observation count,
retry count, and whether an actionable failure preserves the session.
Refs chain across one recorded checkout session and counts derive from
the real formatters, so dropping settled-diff refs, the unchanged-
interactive tail, or a recovery handle fails the suite. Adds a matching
non-gating help-conformance next-command case. Response defaults
unchanged.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(output-economy): make routine-workflow ref-surfacing depend on rendered output

Address review on #1190:
- Drop the raw mutation-confirm/failure MCP data samples that leaked e4/e5
  and mislabeled the surface; e4/e5 now surface only from the rendered CLI
  settled-diff, so dropping added refs genuinely raises the fallback count.
- Track the failure once as its projection-invariant normalized payload
  (workflow.failure.shared.json) instead of duplicate cli/mcp raw copies.
- Reuse the shared REF_TOKEN_PATTERN from economy-metrics.ts.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(output-economy): reuse shared fixtures in routine-workflow oracle

Address review finding #2 on #1190: chain the workflow onto the shared
per-surface fixtures instead of copy-pasting them.
- orient/recheck reuse SNAPSHOT_RESULT + SNAPSHOT_DAEMON_RESULT; the
  first mutation reuses SETTLE_ADDED_REF_RESULT, so session identity and
  ref generations come from ./fixtures.ts and the two suites cannot drift.
- Only the genuinely workflow-specific pieces remain local: the unchanged
  recheck, a tail retargeted onto a settled-diff ref (SETTLE_TAIL_RESULT
  taps an unsurfaced @e6 and can't chain), the in-session timeout failure,
  and its recovered retry. routine-workflow.ts drops ~110 LOC.
- Rendered-output ref guard preserved: @e5 (settled diff) and @e7 (tail)
  surface only from formatter output; recovery semantics unchanged.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Michał Pierzchała <thymikee@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-10 15:27:24 +02:00
devin-ai-integration[bot] 1f14e224d3 refactor(daemon): route perf metrics sampling body through PlatformPlugin facet (#1191)
* refactor(daemon): route perf metrics sampling body through PlatformPlugin facet

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(daemon): cover the shipped perf sampler dispatch path via the facet

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci: re-trigger checks (flaky settle-observation integration test)

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(daemon): pin perf sampler selection to the facet tag, not the platform

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Michał Pierzchała <thymikee@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-10 15:26:30 +02:00
devin-ai-integration[bot] ae74c51abd chore: add agent-efficiency regression guards (#1174)
* chore: ratchet architecture dependency graph

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test: ratchet agent-facing output economy

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* feat: derive command navigation explanations

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: keep efficiency checks fallow-clean

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(layering): enforce back-edge ceiling monotonicity and cover root src files

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* ci(layering): pin back-edge-ceiling ratchet to PR merge-base

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* test(output-economy): baseline-independent actionability floors, policy-derived error, like-for-like screenshot surfaces

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix(explain): resolve true CLI aliases, canonical usage, and derived owners

Surface true CLI aliases from parser normalization (long-press, metrics,
tap, launch, relaunch) distinct from catalog keys, preserving implied-flag
semantics (relaunch => open --relaunch). Extract the canonical single-line
usage builder to src/utils/cli-usage.ts so schemas without usageOverride
include positionals and flags. Replace guessed handler paths with a
completeness-checked daemon-route owner map keyed by the closed
DaemonCommandRoute union, fixing silently-dropped non-kebab routes
(reactNative, recordTrace) and generic dispatch. Add table-driven coverage
for aliases, synthesized usage, split-family/route-variant/dispatch owners,
structured output, and explain:command CLI exit/stdout/stderr.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: enforce exact ratchets and compact command explain

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: colocate command ownership metadata

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: bind daemon owners to production routes

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: preserve generic dispatch bundling

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

* fix: enforce monotonic output budgets

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>

---------

Co-authored-by: Michał Pierzchała <thymikee@gmail.com>
Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
2026-07-10 11:54:09 +02:00
Michał Pierzchała 66fe801377 docs: ADR for interactive replay and resolution disclosure (#1177)
* docs: add ADR 0012 for interactive replay, resolution disclosure, and retiring --update healing

Records the decision to retire --update healing as a silent actor (repurposing
its candidate machinery as ranked suggestions), disclose selector
disambiguation in every interaction response, verify replay steps against
record-time identity evidence, and add an interactive replay --from loop with
a structured divergence report for all callers.

* docs(adr-0012): ground in live replay evidence; require step provenance for --from

Adds hands-on evidence from driving replay on the RN playground (silent
text-mode success, app-state divergence heal cannot fix, Maestro step-index
shift from runFlow flattening, per-format hint/code inconsistency, recordings
carrying zero observation steps), makes step provenance (source file + line,
including through Maestro runFlow inlining) a requirement of the divergence
report plus an optional replay --list-steps dry-run, and adds a one-line
text-mode success summary as decision 4d.

* docs: make interactive replay ADR implementable

* docs: tighten interactive replay contracts

* docs(adr-0012): demote geometry to disambiguation signal, fix matchCount, define matching algorithm

Reworks the target-v1 contract per review: identity is recorded id, else
role + normalized label, plus a leaf-anchored ancestry prefix (K=8, nearest
ancestors kept, root-side truncation only); absolute rects are demoted to
never-compared diagnostics with the ±8 tolerance removed rather than tuned;
duplicates disambiguate by recorded sibling order among the matching set,
then viewport-relative order within the recorded scroll region — never
absolute pixels; ties are identity-unverifiable divergences with candidates
listed. matchCount is redefined as the replay-time recorded-selector match
count (0..N, always present), with selector-miss (0) and identity-mismatch
(>=1, no identity candidate) as distinct classes in an explicit six-path
verification classification. Also inlines the quantitative benchmark numbers
(3.67->1.00 snapshots, 14.3 vs 23.3/26.7 commands, 38/38 in 539s) so the
evidence is durable without the external harness directory.

* docs(adr-0012): unify positional-signal candidate domains between record and replay

Fixes the P1 domain mismatch: sibling becomes a genuine same-parent child
index (parent already captured as ancestry[0], no new field; identical by
definition on both sides, non-isolating when the same index recurs under
different parents); viewportOrder gets one region-scoped domain — the
identity set partitioned by scroll region, ordinal within the recorded
partition on both sides, unavailable (never compared cross-region) when the
recorded region no longer exists; document order (pre-order index) is the
canonical total order making every ordering deterministic, including equal
rect centers. Residual ties stay identity-unverifiable with candidates
listed. Record-time write, replay verification, and mandatory validation
updated in lockstep; the six-path classification is unchanged.

* docs(adr-0012): conditional matchCount, dependency-ordered migration, writer invariant, suggestion ranking contract
2026-07-10 11:39:45 +02:00
Michał Pierzchała b3b0f590b3 fix: add curated suggestions for open-url and close-session (#1175)
Last night's help-conformance benchmark showed agents guessing
open-url and close-session with no "did you mean" hint (the
Levenshtein fallback doesn't cover either token — both exceed the
edit-distance threshold for their length). Add curated entries
pointing them at `open <url>` and `close`.
2026-07-10 10:06:54 +02:00
Michał Pierzchała 983625fc5d feat: fix codex runner, add --override-doc grading, port skillgym quiz cases to help-conformance bench (#1176)
* feat: fix codex runner, add override-doc grading, port skillgym quiz cases

The 2026-07-09 evaluation of scripts/help-conformance-bench.mjs found it
structurally right but broken for the codex runner (two bugs), thin on
coverage (4 cases), sequential, and unable to grade a draft help rewrite
without a rebuild.

- Fix runCodex: (1) codex exec reads stdin until EOF when not attached to
  a TTY, and execFile never closes the child's stdin, so every codex call
  hung until RUN_TIMEOUT_MS with empty output — close stdin right after
  spawn. (2) `-o outFile` writes the same final JSON that codex also
  prints to stdout, so concatenating both produced two back-to-back JSON
  objects that broke every JSON.parse candidate and silently zeroed
  extractCommands() — prefer the clean -o payload, fall back to stdout
  only when it's empty.
- Add `--override-doc <topicId>=<path>` (repeatable): loads a topic's
  text from a file instead of shelling out to
  `node bin/agent-device.mjs help <topic>`, so a draft help rewrite can
  be A/B graded with zero rebuild.
- Port three cases from
  test/skillgym/suites/agent-device-smoke-suite.ts
  (settle-diff-is-observation, sample-output-settled-diff-next-target,
  sample-output-not-settled-needs-observe) as self-contained
  "next-command quiz" cases, generalizing the scorer to support regex
  matchers/forbidden patterns alongside the existing named expectations.
  Fixture output text matches the CURRENT settle rendering in
  src/commands/interaction/output.ts, including the "unchanged
  interactive (N):" tail added by #1167/#1172.
- Parallelize the runner x case matrix with a concurrency cap
  (HELP_BENCH_CONCURRENCY, default 4); results still print in the
  original matrix order.
- Extend test/skillgym/README.md's existing pointer to this bench with
  the new flags.

Validated with real LLM calls (both runners, all 7 cases, 14 calls,
~$0.25 total): 13/14 pass; the one fail (claude-haiku-4-5 on
dogfood-mode) is a genuine model miss (returned an empty command plan
asking for the app name instead of committing to a generic plan), not a
bench bug. `--override-doc` demonstrated live: stripping the dogfood
doc's evidence-command examples regresses codex:gpt-5.4-mini from 3/3 to
2/3 on the same case, showing the flag both loads and changes grading.

* fix: apply live-doc post-processing to --override-doc, fail fast on bad overrides

Review findings on the initial version (all reproduced):

- HIGH: an override for the --help:first30 doc id skipped the live path's
  firstLines(text, 30) cap, so a 49-line draft leaked lines 31-49 into the
  prompt — grading content a live run never shows, on the doc id every case
  uses. loadDoc now splits source (live shell-out vs override file) from
  post-processing, and the post-processing applies to both, so an override
  differs ONLY in where the text comes from.
- MEDIUM: an --override-doc topic id no selected case uses was silently
  ignored (exit 0, real doc graded). Now fails fast listing the valid doc
  ids for the selection.
- LOW: a missing override file threw a raw ENOENT stack trace; expected
  failures now print one clean Error line. Added --help usage text that
  documents last-wins semantics for repeated same-topic overrides and the
  post-processing parity.

Guard tests (scripts/__tests__/help-conformance-bench.test.ts, wired into
the unit-core vitest project by explicit path): a 49-line fixture whose
prompt must keep line 30 and drop line 31, unknown-topic fail-fast with
valid ids listed, clean no-stack error for a missing file, and last-wins
for repeated overrides. All spawn the script in --dry-run with every
required doc overridden, so they need no LLM calls and no built CLI.

Live re-validation: a 33-line override of --help:first30 whose lines
31-33 instruct the model to emit a sentinel command; neither
claude-haiku-4-5 nor codex:gpt-5.4-mini emitted it (both scored 4/4,
matching the live-doc baseline), proving the cap applies end-to-end.
2026-07-10 09:06:58 +02:00
Michał Pierzchała b8ac75893a fix: make settle tail list real actionable targets (#1172)
* fix: make settle tail list real actionable targets

Post-merge benchmark of #1167 (React Navigation prevent-remove flow, iOS
sim, claude-haiku/sonnet) found the unchanged-interactive tail regressing
to chrome-only noise in exactly the cases it was built for:

- buildSettleTailEntries required `hittable === true`, stricter than what
  `snapshot -i` itself shows for the same interactive-only capture. Right
  after a dismiss animation, real buttons commonly report `hittable:
  false`/undefined while application/window containers pass, so the tail
  surfaced two useless chrome lines and dropped the actionable button.
  Fixed by dropping the hittable requirement and excluding structural
  application/window roles instead.
- withoutKeyboardKeys only stripped `Key` nodes; real keyboard chrome
  (shift/Emoji/return/Dictate/Next keyboard) are XCUIElementTypeButton
  nodes and leaked through as fresh added-line refs, which suppressed the
  tail trigger for exactly the post-fill case it exists for. Fixed by
  detecting the whole keyboard subtree structurally (via parentIndex, not
  a locale-fragile label list): descendants collapse out of the diff, and
  the keyboard container's own ref no longer counts as a "meaningful"
  added ref for the trigger decision.

Nobody presses shift via settle diff refs, so collapsing keyboard chrome
does not block a user from explicitly targeting the keyboard itself.

* fix: classify the whole iOS keyboard window as settle chrome

Live-device review of the first Bug B fix found "Next keyboard" and
"Dictate" still leaking into the settled diff as added refs: on a real
iPhone 17 Pro simulator they live in a SIBLING subtree of the [Keyboard]
container (the candidate bar), not under it, so the container-descendant
walk missed them. A raw hierarchy capture shows the software keyboard in
its own dedicated window hosting both the container and the candidate
bar, so the structural rule is now: every node inside a window that has
a [Keyboard] descendant is keyboard chrome. Conservative guard: a window
also hosting an editable text node outside the container (iOS puts
inputAccessoryView composers in the keyboard window) is never
window-classified — those fall back to the container-descendant walk.

The live run also showed the filled field re-labeling itself with its
new value (ancestor wrappers inherit it), which produced added refs that
suppressed the tail even with chrome fixed. The trigger now also ignores
self-echo refs — added lines whose settled node rect contains the action
point — since they re-describe the acted-on element, not a new target.

The fill-keyboard provider fixture is now a trimmed REAL capture from
the benchmark flow (sibling candidate bar, main app window absent from
the interactive settled capture, self-echo relabels) instead of a
hand-built tree that hid the sibling-branch shape.
2026-07-09 21:07:15 +02:00
Michał Pierzchała c78bfd1e62 fix: guide agents past iOS keyboard dismissal and get-text guesses (#1173)
* fix: guide agents past iOS keyboard dismissal and get-text guesses

Tonight's benchmark leaderboard showed two recurring agent-UX misses:
keyboard dismiss failing after fill (5x, now the #1 failed command) and
get-text guessed as a command name (1x).

- iOS keyboardDismiss now returns a hint explaining the on-screen keyboard
  does not block agent-device interactions, so agents should press the next
  target directly instead of retrying dismiss, and use keyboard enter only
  when submission is actually wanted.
- help manual-qa and help workflow recovery text no longer teach the false
  "dismiss to unblock the target" pattern.
- The command-suggestion curated map now maps get-text/gettext/get_text to
  the real get text command shape; the generic edit-distance fallback did
  not produce any suggestion for these guesses.

* fix: soften keyboard guidance for genuinely covered targets

Review on #1173 flagged the unconditional "does not block" claim as false
for targets visually covered by the keyboard: direct-selector press
resolution is isHittable-gated and falls through to ELEMENT_NOT_FOUND, the
tree path allows non-hittable taps with only a no-visible-effect hint, and
bottom-pinned-button-under-keyboard (#291/#469/#957) is a real case where
dismissal was the remedy.

All three strings (runner hint, help manual-qa, help workflow) now say the
keyboard USUALLY does not block presses and name concrete fallbacks: scroll
the target into view, or keyboard enter when submission is wanted.
2026-07-09 21:03:55 +02:00
Michał Pierzchała 5c08f1664f 0.19.2 v0.19.2 2026-07-09 19:32:16 +02:00
Michał Pierzchała a3885351c2 fix: stabilize android maestro gestures (#1171)
* fix: stabilize android maestro gestures

* fix: address maestro android gesture review
2026-07-09 19:31:11 +02:00
Michał Pierzchała 8878399272 feat: include unchanged interactive refs in settle output (#1167)
* feat: include unchanged interactive refs in settle output

Benchmarks (gpt-5.4-mini + claude-haiku, July 2026) showed 27% of --settle
actions were followed by a fallback snapshot -i because a change-only diff
omits refs for elements that did not change: after a modal dismiss the diff
shows only removals, so the next button to press is invisible.

Add an unchanged-interactive tail to SettleObservation, attached only when
the diff's added lines carry zero refs (the modal-dismiss/toast-only
signature). It lists the settled tree's remaining hittable, uncovered
elements so the response stays actionable without an extra round trip.
Rides the CLI text, MCP digest view, ref pinning, and output schema the same
way the diff's added-line refs already do.

* refactor: address fallow audit findings on the settle tail

- drop the unused export on buildSettleTail (tests exercise the trigger
  through the public interaction path and the filter via
  buildSettleTailEntries)
- extract the digest tail capping from interactionSettleView into a
  module-private helper to stay under the complexity gate
2026-07-09 19:18:47 +02:00
Michał Pierzchała 9a2277c045 fix: alias launch/relaunch to open and suggest canonical commands for unknown names (#1166)
* fix: suggest canonical commands for unknown command names

Agents commonly guess command names that don't exist, e.g. relaunch/launch
instead of `open <app> --relaunch`, burning turns on Unknown command errors
that only say "run --help". Add a curated alias-to-canonical-shape map for
the most common guesses (launch/relaunch/start/restart, touch, input/
settext/entertext, screencap/capture, dismiss), backed by a nearest-name
edit-distance fallback derived from the live command registry so
suggestions can't drift. Also hint that `open` takes the app/bundle id as
a positional when an unknown flag looks like a bundle-id guess (e.g.
--bundle-id), and apply the same suggestion to `help <unknown>`.

Suggestions are display-only; nothing auto-executes and the error code
stays INVALID_ARGS.

* fix: address review — dead export, case-insensitive suggestions, tighter nearest-name matching

- Drop the export on getNearestCommandNames (module-private; only
  suggestCommandFor uses it) to satisfy the Fallow unused-export gate.
- Lowercase the input token before both the curated-map lookup and the
  nearest-name pass, so RELAUNCH/Relaunch/TAP/Touch get the same hint as
  their lowercase forms. Added a curated `tap` entry: lowercase `tap` is
  normalized to press before the unknown-command check, so the entry only
  catches case variants like TAP.
- Tighten the nearest-name fallback: exact prefix matches win outright
  (`clos` now suggests only `close`, not "one of: close, logs"),
  otherwise only ties at the minimum edit distance are kept, and 1-2
  character tokens never get a suggestion (`ls` no longer suggests `is`).
- Share the "open <app> --relaunch" example string between the curated
  map and the unknown-flag hint, and extend the registry-drift tests to
  parse each curated example end-to-end (validates open --relaunch as a
  registered flag) plus assert keyboard dismiss is a real keyboard action.

* feat: promote launch and relaunch to true open aliases

Follow the tap -> press precedent: `relaunch <app>` now runs
`open <app>` with --relaunch injected, and `launch <app>` runs a plain
`open <app>` (no forced restart — that would silently destroy app
state). Both are normalized in normalizeCommandAlias before parsing, so
command identity stays `open` for daemon requests and telemetry, all
other args/flags pass through to open's normal validation (URL targets
still get the daemon's existing --relaunch guidance), and an explicit
--relaunch stays idempotent. Alias matching is now case-insensitive
(TAP, RELAUNCH, Launch), so the curated tap suggestion entry is dead
and removed along with launch/relaunch; start/restart and the rest of
the map stay suggestion-only since start is genuinely ambiguous.
2026-07-09 19:00:41 +02:00
Michał Pierzchała 9dbecd02a6 fix: make manual-qa help self-sufficient and slim routine help paths (#1168)
* fix: make manual-qa help self-sufficient and slim routine help paths

July 2026 benchmarks (gpt-5.4-mini/haiku/sonnet driving the CLI) showed
agents burning 14-41KB of tool output on help before a routine QA flow:
help react-native routed "generic navigation, selectors, refs,
verification" to help workflow (31.7KB), one run read the bare
`agent-device help` mid-task (16.3KB), and several fell back to
per-command help because manual-qa lacked concrete command shapes.

- manual-qa gets a "Command shapes" block covering the full routine QA
  loop (open --relaunch, deep-link open, snapshot -i, press/fill --settle,
  wait text, close) plus a quoting note for apostrophe/quote labels, so
  agents don't need to read help workflow or per-command help for a
  normal pass.
- react-native's routing block no longer sends routine QA to help
  workflow; it now points routine flows at manual-qa and reserves
  workflow for deep exploration/debugging.
- physical-device and the shared topic-footer routing also point at
  manual-qa alongside workflow.
- Trimmed the bare `agent-device help` output: cut Agent Quickstart
  lines that duplicated topic-specific detail already owned by
  react-native/workflow/remote/web/tv, dropped 3 redundant Examples, and
  fixed an alignment bug where one long usage string (tv-remote) forced
  padding whitespace onto every other Commands: row. The "Default app
  loop" line (shown to flip Haiku from 0/3 to 3/3 in earlier benchmarks)
  stays intact and early.

* fix: satisfy oxfmt and restore pinned-ref staleness nuance in help workflow

- pnpm format:check failed on the new renderAlignedSection columnWidth
  line; reformatted with pnpm format.
- Reviewer noted the only content genuinely lost from the bare-help
  Quickstart trim was the "pinned refs get exact staleness warnings"
  nuance; restored it as one line in help workflow's Snapshots and refs
  section (its natural owner), keeping the bare-help line short.
2026-07-09 18:46:33 +02:00
Michał Pierzchała c325842f62 fix: reap idle daemons and take over stale runner leases (#1169)
* fix: reap idle daemons and take over stale runner leases

Each AGENT_DEVICE_STATE_DIR spawns its own daemon that never exits on its
own; deleted codex/claude sandboxes leave orphaned daemons accumulating
(10+ observed). A stale-but-still-running orphan also keeps holding its
iOS runner lease, so a fresh daemon for the same device fails with
COMMAND_FAILED "already owned by another agent-device daemon".

- Daemon self-reaps after an idle window (default 5 minutes, matching the
  iOS runner idle-stop default) once it has no open sessions, no
  in-flight requests, and no active recording. AGENT_DEVICE_DAEMON_IDLE_TIMEOUT_MS
  overrides the window; 0 disables it.
- A runner lease whose owner PID is dead, or whose owner
  AGENT_DEVICE_STATE_DIR no longer exists, is now reclaimed automatically
  instead of erroring; a genuinely live owner with an existing state dir
  still gets the existing rejection + hint.

* fix: gate lease takeover on proven-dead owners and fail closed on stat errors

Review follow-ups on #1169:

- Replace fs.existsSync (which never throws and swallows EACCES/IO errors
  into "gone") with fs.statSync + error-code inspection: only ENOENT/
  ENOTDIR count as proof the owner state dir is gone; any other stat
  error fails closed and classifies the owner as alive.
- Split stale classification by reason (owner-process-dead vs
  owner-state-dir-gone). Adoption (readStaleRunnerLease ->
  tryAdoptRunnerSessionFromLease) is now strictly PID-dead-gated:
  a dir-gone-but-alive owner may still hold a live runner connection,
  so its lease routes through the force-stop path (kill leased runner
  processes, rebuild) instead of being silently adopted - no two
  masters.
- Tests: EACCES stat error keeps the busy rejection; dir-gone+PID-alive
  refuses adoption before probing; dir-gone force-stop asserts a fresh
  runner launch instead of adopting the old runner pid.

* test: make idle reap tests deterministic
2026-07-09 18:42:03 +02:00
Michał Pierzchała 3667f6ece5 fix: point unknown selector keys at role=/label= forms (#1165)
* fix: point unknown selector keys at role=/label= forms

press 'button="Push Article"' errored with a nonsense suggestion
(text="button=\"Push Article\"") because isSelectorToken rejects
`button` as a key, splitSelectorFromArgs returns null, and the point
fallback wraps the whole raw token in text=.

Add detectUnknownSelectorKeyToken to spot a key=value token whose key
isn't a recognized selector key, and use it in readPointTarget to
throw a targeted error before the numeric parse: role=<key>
label=<quoted value> when the key looks like an accessibility role
word (isRoleHintWord, mirroring ROLE_LABELS), otherwise label=<quoted
value>. wait-positionals.ts has no analogous point-fallback path, so
it needs no change.

* fix: fold unquoted multi-word values into the unknown-key suggestion

Review follow-up: readPointTarget only inspected positionals[0], so an
unquoted multi-word value split across positionals (press 'button=Push'
'Article') dropped the trailing tokens and confidently suggested the
wrong completion (label="Push"). Fold trailing positionals into the
suggested value like mergeRestIntoSelectorValue does, unless the value
was fully quoted (button="Push Article") and therefore complete — that
distinction keeps fill's trailing text argument out of the suggestion.

Also: drop the dead `text` entry from ROLE_HINT_WORDS (valid selector
key, short-circuits in ALL_KEYS first) and note the set is a superset
of ROLE_LABELS rather than a mirror; reject whitespace-only values in
detectUnknownSelectorKeyToken.
2026-07-09 18:28:29 +02:00
Michał Pierzchała 477c684cf6 fix: read runner cache package version from project root (#1170) 2026-07-09 14:19:05 +02:00
Michał Pierzchała ff7fc5ba61 0.19.1 v0.19.1 2026-07-08 21:49:10 +02:00
Michał Pierzchała 888984169b fix: make record app-scoped by default (#1163)
* fix: reject recording for failed iOS simulator session

* fix: make record app-scoped by default
2026-07-08 21:48:35 +02:00
Michał Pierzchała 0d2b0353ff docs: clarify agent setup and text entry guidance (#1164) 2026-07-08 21:36:33 +02:00
Szymon Dziedzic cfef0a4bca feat: add session event timeline (#1032)
* feat: add session event timeline

* fix: support cursor-only event reads

* refactor: simplify event log formatting

* refactor: trim event log helpers

* docs: document session event timeline

* refactor: tighten session event log internals

* fix: redact event log action positionals by default

* fix: align event log after rebase

* test: cover events in provider output guard

* fix: harden session event privacy

* fix: harden event message redaction

* fix: harden session event logging

---------

Co-authored-by: Michał Pierzchała <thymikee@gmail.com>
2026-07-08 21:35:45 +02:00
Michał Pierzchała cf31fb3f7b fix: harden iOS XCTest recovery paths (#1158) 2026-07-08 21:11:28 +02:00
Michał Pierzchała 2717047b86 fix: normalize iOS simulator screenshot density (#1160)
* fix: normalize iOS simulator screenshot density

* fix: avoid density metadata after screenshot downscale

* fix: satisfy screenshot density CI gates

* fix: harden screenshot metadata collection

* refactor: centralize screenshot density policy

* refactor: reuse screenshot density support check
2026-07-08 21:05:24 +02:00
Michał Pierzchała e4115ec5ab chore: migrate to TypeScript 7 (#1161) 2026-07-08 21:04:54 +02:00
Michał Pierzchała 106c238697 fix: narrow client result contracts (#1155) 2026-07-08 18:35:27 +02:00
Michał Pierzchała f21727d065 fix: handle Android IME overlays in snapshots (#1157)
* fix: handle Android IME overlays in snapshots

* fix: satisfy Android IME CI guards

* fix: detect localized Gboard tutorial overlays

* fix: keep Android IME overlay handling passive
2026-07-08 18:29:47 +02:00
Michał Pierzchała 9dabe5b1c1 refactor: derive command identity from descriptors (#1151)
* refactor: derive client-backed cli routing

* refactor: derive command identity from descriptors
2026-07-08 17:55:00 +02:00
Michał Pierzchała b91eaad885 refactor: make iOS synthesized gesture policy explicit (#1152)
* refactor: make iOS synthesized gesture policy explicit

* test: harden settle observation under coverage

* fix: preserve first-command synthesized drag behavior

* refactor: simplify synthesized frame policy

* refactor: inline synthesized command policies

* refactor: simplify sequence synthesized context

* refactor: clarify synthesized drag fallback policy

* refactor: keep synthesized gesture policy runner-local
2026-07-08 17:15:42 +02:00
Michał Pierzchała f18d0b2e92 fix: improve settle observation guidance (#1154) 2026-07-08 17:12:19 +02:00
Michał Pierzchała bbc577c11a fix: derive interaction response data transforms (#1149)
* fix: derive interaction wire projection

* fix: derive wire projection from command descriptors

* refactor: clarify response data transform naming

* test: guard response transform field ownership
2026-07-08 14:13:15 +02:00
Michał Pierzchała b0c70ad4e4 feat: support repack dev server prepare (#1145) 2026-07-08 11:14:36 +02:00
Michał Pierzchała 694266d802 fix: keep iOS synthesized drags off AX (#1148)
* fix: keep iOS synthesized drags off AX

* fix: address iOS synthesized drag review
2026-07-08 11:06:57 +02:00
Michał Pierzchała 7f61df30ae feat: add TV remote command (#1147)
* feat: add TV remote command

* feat: improve TV remote ergonomics

* test: cover tv-remote provider scenario

* fix: preserve focused Android TV nodes

* docs: tighten PR description guidance

* fix: remove d-pad command alias

* docs: clarify tv-remote hold syntax

* feat: add tv-remote longpress CLI sugar
2026-07-08 10:59:48 +02:00
Michał Pierzchała bc7dcc8345 fix: keep XCTest tree snapshots on main (#1144)
* fix: keep XCTest tree snapshots on main

* fix: address iOS runner snapshot review
2026-07-07 16:59:00 +02:00
Michał Pierzchała 9dfebbe3be fix: omit interaction uptime from wire responses (#1142)
* fix: omit interaction uptime from wire responses

* fix: simplify interaction wire sanitizer
2026-07-07 16:53:44 +02:00
Michał Pierzchała 69a8f6f3fe fix: soften recovered snapshot warning (#1146) 2026-07-07 16:52:25 +02:00
Michał Pierzchała ef4b66d4dc test: remove slow-test ratchet pins (#1143) 2026-07-07 14:46:02 +02:00
Michał Pierzchała 9009c5aff7 0.19.0 v0.19.0 2026-07-07 13:17:01 +02:00
Michał Pierzchała 8f28c31c86 fix: recover completed Android recording from pending-only manifest (#1141)
* fix: recover completed Android recording from a pending-only manifest

When the daemon crashes in the brief window between writing the pending recovery
manifest (before screenrecord starts) and upgrading it to a `current` manifest, the
screenrecord process can still finish and leave a complete MP4 on the device. record
stop previously discarded it as stale because the pending-only recovery path never
checked for an on-device file, unlike the `current` path which already recovers a
finished recording. Extend the pending-only path to recover the completed file with the
same finished-recording warning, and skip the stop signal when the recovered recording
has no tracked pid (a pending chunk never records one, and probing an empty pid is
unsafe).

* fix: treat JSON arrays as invalid Android recovery manifests

isRecord accepted arrays (typeof [] === 'object'), so a stray `[]` recovery manifest
was classified as blocked rather than deleted, wedging every subsequent record stop.
Reject arrays and null so a non-object manifest is cleaned up like other malformed
metadata.
2026-07-07 12:04:58 +02:00
Michał Pierzchała 0dcc1aa553 fix: normalize interaction response wire shapes (#1114)
* fix: normalize interaction response wire shapes

* fix: reduce interaction response complexity
2026-07-07 11:29:01 +02:00
Michał Pierzchała 7e583c4136 fix: simplify Android recording recovery (#1135)
* fix: harden android recording recovery

* fix: reduce android recording recovery fallow complexity

* test: fix android recording recovery rebase

* fix: block uncertain android recording fallback

* fix: address android recording recovery review

* fix: address android recording recovery review followup

* fix: simplify Android recording recovery

* fix: address android recovery ownership review

* refactor: reuse android recovery manifest helpers

* refactor: split android pending recovery resolution

* fix: clarify scoped android recovery hint
2026-07-07 10:50:57 +02:00
Michał Pierzchała d5a7af0f4d fix: hint on iOS runner main-thread timeouts (#1140) 2026-07-07 10:26:49 +02:00
Michał Pierzchała b052eb9a37 fix: speed up iOS text entry (#1139)
* fix: speed up iOS text entry

* fix: reverify settled iOS text entry
2026-07-07 09:50:40 +02:00
Michał Pierzchała 8ef4e73408 refactor: derive command exposure lists from descriptors (#1137) 2026-07-07 08:00:33 +02:00
Michał Pierzchała 5c5fa012f7 feat: --settle returns the settled diff in the interaction response (#1101) (#1106)
* feat: --settle returns the settled diff in the interaction response (#1101)

press/click/fill/longpress --settle executes the action, waits for the UI
to go quiet (wait stable's loop, shared via stable-capture.ts), and returns
the settled diff vs the pre-action tree in the same response — one round
trip instead of the interact -> observe pair.

- payload: changed lines only (bounded), summary counts, added-line refs,
  refsGeneration; best-effort (settled:false + hint on never-quiet content,
  never an action failure); --verify shares the settle captures
- ref issuance: the settled tree becomes the session snapshot; a
  diff-carrying settle response clears snapshotRefsStale and the MCP layer
  merge-only re-pins added-line refs at the settle generation
- grammar: --settle + --settle-quiet <ms> + --timeout <ms> (flag-sourced
  descriptor budget with new envelope:'widen' semantics mirroring wait)
- ADR 0011: new settleObservation guarantee classified on every path with
  contract scenarios per enforced/delegated cell

* test: give the two contention-flaky doctor scenarios explicit budgets

The doctor provider scenarios sit at ~5s of real daemon-harness work on a
loaded host and flake at vitest's 5s default during full-suite runs (the
known contention flake AGENTS.md documents). Same in-file precedent as the
Metro-probe scenario's 10s budget.

* fix: move SettleParams to contracts to satisfy the layering DAG

daemon/handlers/interaction-flags.ts imported the type across the
daemon -> commands boundary (R2 commands-floor). The tuning params are
part of the interaction contract like SettleObservation, so they live
in contracts/interaction.ts and both layers import from there.

* feat: keep settle diffs content-first — drop Key nodes, added lines win the cap

Bluesky dogfood: a fill that summons the iOS keyboard spent 49 of the 80
capped diff lines spelling out QWERTY keys, and a screen transition with
269 removals could starve out the added lines entirely. Key-type nodes
are now filtered from both diff sides (the [keyboard] container line
still signals presence), and under truncation added lines — the ones
carrying fresh refs — win slots over removals.

* docs: state the core loop in the top-level help starting point

Benchmarked with headless haiku/sonnet agents given only --help: both
models skipped the help-workflow pointer and started with plain
snapshot (38KB payloads they then had to re-read from files). One
core-loop line at the starting point is what teaches snapshot -i and
--settle to models that never read a second help page.

* fix: preserve settle digest refs for mcp

* fix: reduce settle fallow complexity

* fix: surface settle output in CLI text

* fix: complete settle handling for longpress

* refactor: localize daemon timeout envelopes

* refactor: deepen post-action observation

* refactor: centralize post-action observation planning

* refactor: derive settle capability from descriptors

* refactor: trim settle descriptor helpers
2026-07-06 20:18:44 +02:00
Michał Pierzchała db69124c00 feat: add capabilities command (#1133) 2026-07-06 19:35:46 +02:00