Commit Graph

841 Commits

Author SHA1 Message Date
Hubert Gancarczyk 0285e91ee1 test(cli): expect the scripts.bash message the schema now prints 2026-09-11 13:19:18 +02:00
Hubert Gancarczyk 627e70aebc Merge remote-tracking branch 'origin/main' into feat/flow-bash-scripts
# Conflicts:
#	packages/tool-server/src/tools/flows/flow-run.ts
#	packages/tool-server/src/tools/flows/flow-utils.ts
2026-09-11 13:06:27 +02:00
Hubert Gancarczyk a3c8e858d4 chore: remove output reference from docs 2026-09-11 11:51:50 +02:00
Hubert Gancarczyk aa029bd6ec chore: remove comments 2026-09-11 11:40:54 +02:00
Kacper Kapuściak e1d7654af5 docs: add a Limits section to the physical iOS devices page (#1050)
## Summary

- Adds a `## Limits` section to
`packages/docs/docs/features/physical-ios-devices.mdx`, matching the
shape of the other feature pages.
- Moves the unsupported-capability bullets out of "What works" into the
new section.
- Adds limits that were undocumented: iPhone only, USB cable only,
home/volume/action buttons only, Enter and Backspace only.

## Docs

This PR is a docs-only change. `npx docusaurus build` and `npm run
format` pass.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
  * Clarified limitations when using physical iOS devices.
* Documented supported hardware: iPhones are available, while iPads are
not currently shown.
* Added details about USB-cable-only control, supported home, volume,
and action buttons, and available text, Enter, and Backspace keys.
* Moved simulator-specific limitations into a dedicated “Limits”
section, including unavailable rotation, shake, permission changes,
debugging, profiling, inspection, recording, and paste capabilities.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-09-11 11:33:33 +02:00
Hubert Gancarczyk ed02a5151d chore: refine docs 2026-09-11 10:17:56 +02:00
Hubert Gancarczyk ca66a79cd0 fix(flows): give a cancelled run the settle every step had before
Fifth review found that a cancel landing while a stderr consumer was
still delivering the script's last lines - after "step 1" but before the
error - made the reason "step 1: seeding": 76b619e23 let a cancel end the
post-exit wait at once, and ede786396 then took the line as it stood.
4edac0170 gave the error, because its 500 ms settle ran whatever the
caller did. A cancel now ends the wait only once that first settle has
passed, so the consumer gets as long as it always had and a cancel waits
no longer than it did before the wait could stretch. The abort leaves
the race once it has fired, so the loop does not spin while it waits.
2026-09-11 01:25:22 +02:00
Hubert Gancarczyk ede7863963 fix(flows): keep what stderr had said when a cancel ends the wait
Fourth review found that a cancel landing within 500 ms of a stderr
consumer's late line dropped that line from the reason: the settle ended
on the abort before stderr had been quiet for 500 ms, and the reason fell
back to the line as bash exited - empty, the consumer not having written
yet. A cancel says nothing about whether a job is still writing, so a
settle it ended now takes the stderr line as it stands.
2026-09-11 01:01:07 +02:00
Hubert Gancarczyk 0a0026419c fix(flows): read pending stderr before judging it quiet after a stall
Fourth review found that a block of the tool server's event loop across
the 500 ms mark - a synchronous call in an I/O callback - made the reason
freeze on the line before the script's error. After such a block the
due timer runs before the poll that reads what arrived meanwhile, so both
the quiet watch and the settle judged stderr quiet on stale data: with a
stderr consumer and a surviving job, the reason ended with "step 1:
seeding" instead of "FATAL: the real error". Each now lets one turn go by
after a timer, so pending pipe data is read before the judgement, and the
watch stops once the settle ends so a late check cannot take a line the
stop provoked.
2026-09-11 00:58:33 +02:00
Hubert Gancarczyk 83d24967d5 docs(flows): say when the reason takes the line bash exited on
Fourth review found that f22f4b46b's docs sentence said the reason takes
the line bash exited on whenever stderr is never quiet for 500 ms. When
the streams close on their own while stderr is still running, every line
counts, which is what the code does and what the sentence before it
implies. The pre-exit line applies only when the wait for the log
reaches its limit with stderr still running; the page now says that.
2026-09-11 00:53:57 +02:00
Hubert Gancarczyk e19801e6ec fix(flows): remove an empty directory the step may not list
Third review found that removeTree (851daeec2) failed on an empty
directory without read permission - `mkdir -m 000`, or a umask of 0777 -
because opendir needs read permission where rmdir asks only the parent.
The step noted EACCES, and every later sweep failed the same way, where
the recursive rm it replaced removed the directory. An EACCES from
opendir, before anything was listed, now falls through to that final rm;
a directory with entries still fails into the note as before.
2026-09-11 00:26:46 +02:00
Hubert Gancarczyk f22f4b46bd fix(flows): count stderr after bash exits only until it first goes quiet
Third review found that fd41a28b5 let a job's stray stderr line become a
bash step's reason: a mock server that logs requests on stdout keeps the
settle open, and its one stderr line a second after bash exited
("listening on :8080") was taken over the script's error, which
4edac0170 reported. After bash exits, stderr carries the script's own
lines late - a consumer working through its backlog - and a job's lines;
the first come as one run from the exit, the second whenever the job
writes. The reason now takes the line stderr stood on when it first went
quiet for 500 ms after the exit. Streams that close while stderr is
still running on count in full; stderr that never goes quiet is a job
still writing, and then the line is the one bash exited on. The docs say
the same.
2026-09-11 00:25:16 +02:00
Hubert Gancarczyk b66344d233 docs(flows): place the settle's cancel boundary at the script's exit
Third review found that the comments added by 24ba6d713 put the line
between a cancel that ends the post-exit wait and one that does not at
the script's answer. The code reads the abort after the script's
process has exited, a few milliseconds later; the comments now say so.
2026-09-11 00:19:44 +02:00
Hubert Gancarczyk 87b69c2663 test(flows): pin where the batched remove moves a directory up
Third review found that the only case guarding 8e3d9b15b's threshold
runs on macOS alone, while CI runs unit tests on Linux, where no file
system takes a 765-byte name: putting REMOVE_TREE_HOIST_AT_BYTES back to
512 left every CI case green. The fs mock now records each rename, and
a case pins that a directory whose path has passed 257 bytes is moved
before it is walked - it fails at 512 on any host.

The cleanup helper moves below the imports and drops a chmod no case
needs, and the long-name case's comment no longer says one name alone
passes 1 024 bytes.
2026-09-11 00:19:44 +02:00
Hubert Gancarczyk 4c1c627630 fix(flows): remove a read-only directory that cannot be moved up
Second review found that removeTree's move-up (096afddd1) left an empty
read-only directory behind for good. Moving a directory to another
parent needs write permission on the directory itself, which a plain
rmdir never asks for, so the rename was refused and the whole removal
failed where the recursive rm it replaced succeeded. A refused move now
leaves the directory where it is and walks it there: its own path still
fits, and an empty directory needs nothing below it named.
2026-09-10 23:59:55 +02:00
Hubert Gancarczyk 8e3d9b15bc fix(flows): move a directory up before a long name can overrun its path
Second review found that removeTree (096afddd1) moved a directory up only
once its path passed 512 bytes, on the premise that a name adds at most
255. APFS takes 255 characters, which in UTF-8 is up to 765 bytes, so a
file or directory with a long multibyte name under a 490-byte path got
ENAMETOOLONG and was left in $TMPDIR for good. The directory is now moved
once its path passes 257 bytes, so any name below it fits in the 1 023
bytes macOS takes.

The docstring and the deep-tree case no longer say the rmSync it replaced
walked such a tree: on Node 20, which CI runs, it fails the same way. The
cleanup cases now clear their roots with `rm -rf`, since a throw from
fs.rmSync in a finally would hide the failure that left the tree there.
2026-09-10 23:58:21 +02:00
Hubert Gancarczyk 33249a81d4 fix(flows): never follow a link at the top of a tree being removed
Second review found that removeTree (851daeec2) opened the directory it
was handed with opendir, which follows a symbolic link. The stale sweep
hands it every entry of $TMPDIR whose name carries the executor's prefix
and a past stamp, so a link planted under such a name had the directory
it pointed to emptied - where the recursive rm it replaced removed the
link alone. A script that replaced its own exchange directory with a
link got the same. removeTree now lstats the top and removes anything
that is not a directory as itself; below the top a directory entry
already says whether it is a link.
2026-09-10 23:56:41 +02:00
Hubert Gancarczyk 24ba6d7134 fix(flows): wait as before for what a run cancelled mid-script leaves
Second review found that 76b619e23 ended the post-exit wait at once for
every cancelled run, including one cancelled before the script answered.
That run was already stopped, and a job in a group of its own outlives
the stop: before, the wait still saw it writing and marked the log cut;
after, the log stopped short with no flag and no note. The settle now
takes the abort only when it had not happened by the script's exit, so
only a cancel during the wait itself ends it. The cancel case also pins
that a job still writing then marks the log cut.
2026-09-10 23:54:56 +02:00
Hubert Gancarczyk ccf4772f72 docs(flows): stop counting the waits outside a step's timeout
Second review found four problems in the sentences this series added to
the script-step reference. "Three waits sit outside it" undercounted, as
"two" had: the stop's grace and the removal of the private directory
also sit outside `timeout`. The page no longer counts them. "Argent
stops reading the log" was not what happens at the 3 s mark - Argent
stops waiting, and still reads what a stopped job writes - so it now
says that. A 29-word sentence is split under the 25-word limit and
names stderr, which is what now decides the line. The post-exit wait
uses the same tense as the paragraph it repeats.
2026-09-10 23:53:21 +02:00
Hubert Gancarczyk fd41a28b55 fix(flows): let only stderr decide a bash reason's line after the settle
Second review found that a1913b052 read the settle's "cut" as a sign that
stderr was still being written, when a job chattering on stdout alone
also ran the settle to its limit. With a stderr consumer in front of the
script's error and such a job, the reason took the line as bash exited -
empty, the consumer not having written yet - and dropped the error that
4edac0170 reported. The line as bash exited is now used only when stderr
itself was still being written as the settle ended.
2026-09-10 23:52:09 +02:00
Hubert Gancarczyk 55eef49a4a fix(flows): keep a refused bash's late stderr line in its refusal
Second review found that 4a5f9b5bd, by answering the probe once the
candidate had exited and its stdout had ended, cut off the stderr line a
refusal quotes when that line arrives after the candidate is gone: a
wrapper that sends stderr through `tee`, or hands it to a job. The early
answer is only needed for a candidate that answered with a version, so a
refused one waits for `close` or the settle again, as it did before.
The probe-timing case's grandchild now lives 1 s instead of 10.
2026-09-10 23:52:09 +02:00
Hubert Gancarczyk 096afddd11 fix(flows): remove a script's tree that runs past the longest path
Review found a regression in 851daeec2: removeTree built full paths, so a
tree the script left deeper than the system's longest path - 1 024 bytes
on macOS - failed with ENAMETOOLONG and stayed in $TMPDIR for good, since
every later sweep failed at the same place. The rmSync it replaced walked
such a tree by descriptor and removed it, after freezing the tool server.

removeTree now moves a directory whose path has grown past 512 bytes up
under the directory being removed before it walks it, which keeps every
path it hands the system short. The main thread still does not stall.
2026-09-10 23:20:50 +02:00
Hubert Gancarczyk dee4705a0a docs(flows): name the third wait outside a script step's timeout
Review found that 06b7891da added a wait outside `timeout` - after the
script exits, a step waits up to 3 s for a command that still writes to
the log - while the page still said two waits sit outside it. The page
now says three and describes the third. The claim that the step reports
each wait is dropped: this one is reported only when it cuts the log.
2026-09-10 23:17:10 +02:00
Hubert Gancarczyk ce129259b8 test(flows): give the log-cut case room before its job writes
Review found the case that pins the log-cut note tight under load: its
job wrote its first line 200 ms after bash exited, and a stall of that
size between the runner's answer and the read of stderr would put the
job's line in the reason. The job now waits 300 ms, still inside the
500 ms quiet window the case needs. The case's name and comment no
longer say Argent stops the job, which it cannot do everywhere.
2026-09-10 23:17:10 +02:00
Hubert Gancarczyk 345db5605a fix(flows): stop the log-cut note claiming Argent stopped the process
Review found that 06b7891da's note and docs said Argent stopped a process
still writing when the settle gave up. On Windows the tree stop reaches
nothing once the runner has exited, and on macOS and Linux it cannot
reach a job in a process group of its own - the page's own background-job
rules say both. What holds everywhere is that Argent stops reading, which
is what the note and the docs paragraph now say. The note no longer names
the 3 s limit, which a cancelled run does not reach.
2026-09-10 23:13:48 +02:00
Hubert Gancarczyk 4a5f9b5bd6 fix(flows): answer the bash probe at the candidate's exit, not a settle later
Review found that 979c9c473, by piping a candidate's stderr to quote it,
made a candidate that leaves a job holding stderr cost the probe's 250 ms
settle on every .sh step: `close` now waited for stderr as well. The
answer is on stdout, so the probe now answers once the candidate has
exited and its stdout has ended, as it did before stderr was piped.
2026-09-10 23:12:31 +02:00
Hubert Gancarczyk 1aba3405e0 test(flows): slow only the abandoned directory in the sweep-wait case
Review found that 851daeec2 weakened "has finished its sweep by the time
the step resolves": its spy delayed every fs.promises.rm call, and since
the step now removes its own exchange directory through fs.promises.rm
too, that removal outlasted the sweep and the case passed with the
executor's `await pendingSweep` deleted. The spy now delays only the
abandoned directory's removal, so the case fails again without the wait.
2026-09-10 23:12:31 +02:00
Hubert Gancarczyk 76b619e231 fix(flows): let a cancelled run end the wait for a job still writing
Review found that 06b7891da's longer settle ignored the request's abort:
the timer and the abort listener are removed when the script's process
exits, so a run cancelled while a job the script left running was still
writing waited up to 3 s more (before: at most 0.5 s). The settle now
takes the request's signal and ends at once; if output was still
arriving, the log is marked cut as it is at the limit.
2026-09-10 23:11:13 +02:00
Hubert Gancarczyk d52571da08 fix(flows): quote the reason a refused bash gave, not the last line it wrote
Review found that 979c9c473 quoted the last non-blank line a refused bash
candidate wrote to stderr. asdf, the version manager the fix was for,
writes its reason first ("No version is set for command bash") and then
lists the versions it has, one per line, so the refusal quoted
"bash 5.2.37" - a line that reads as if the candidate were a bash.

The refusal now quotes the first non-blank line. The docs say which line
is quoted, and that only a candidate that exits without a version gets
one.
2026-09-10 23:08:16 +02:00
Hubert Gancarczyk a1913b0529 fix(flows): keep a stderr consumer's late error in a bash step's reason
Review found a regression in edbdf4dfb: when stderr went through a consumer
(`exec 2> >(…)`) and a quiet background job held the streams open, the
reason took the stderr line as it stood when bash exited - often nothing,
since the consumer had not written yet - and dropped the script's error
that 4edac0170 reported.

The line now follows how the settle ended: the last line overall when the
streams closed on their own, the line stderr ended on before Argent's stop
when what held them went quiet, and the line as bash exited only when a
process was still writing at the settle's limit. A job's answer to the
stop still never becomes the reason.
2026-09-10 23:06:28 +02:00
Hubert Gancarczyk 851daeec2d fix(flows): remove a step's private directory without stalling the tool server
After every .sh step the executor removed the step's exchange directory with
a synchronous recursive rmSync on the tool server's main thread. A script
that left many files there - a fixture it unpacked, a clone - froze every
request, device socket and flow on the host: 100 000 files held the event
loop for about 3 s.

An awaited fs.promises.rm of the whole tree is no cure: it starts one
operation per entry at once and their completions come back in bursts the
loop runs in one turn, stalling it for up to 1.2 s on the same tree. The
directory, and each one the stale sweep collects, is now walked a batch of
64 names at a time with at most 64 removals in flight; the worst stall for
100 000 files is about 20 ms. Each entry still goes through fs.promises.rm
for its force and Windows read-only handling, and a removal that fails
still becomes the step's "could not be removed" note.
2026-09-10 22:33:39 +02:00
Hubert Gancarczyk 06b7891da1 fix(flows): let a process still writing the log finish before the stop
After the script's process exited, the executor waited at most 500 ms for
the output pipes and then stopped the process group. A stderr consumer
still working through its backlog - `exec 2> >(…)`, the plain-bash
timestamp idiom - was killed mid-drain: the log lost its tail with
scriptLogTruncated still false, and a failing step's reason ended with a
progress line instead of the script's error.

The settle now stretches while the tree keeps writing: it ends when the
streams close, after 500 ms of quiet, or 3 s after the exit. A tree still
writing at that limit is stopped, the log is marked cut, and the step
carries a note saying so.
2026-09-10 22:20:38 +02:00
Hubert Gancarczyk 979c9c4732 fix(flows): probe a bash candidate in the step's directory, and say why it failed
The bash version probe ran in the tool server's own working directory and
dropped the candidate's stderr. A version-manager shim picks its bash from
the directory it starts in (asdf reads .tool-versions), so a shim on PATH
was refused and the step fell through to /bin/bash - Apple's 3.2 on a Mac -
with nothing said, and a shim pinned in scripts.bash errored with "is not a
bash" or passed depending on where the tool server was started.

The probe now runs in the step's working directory, the refusal quotes the
last line the candidate wrote to stderr, and a step that runs under a later
candidate after a refused PATH bash carries a note naming the refusal.
2026-09-10 22:14:48 +02:00
Hubert Gancarczyk edbdf4dfb6 fix(flows): keep a job bash left running out of the step's reason
A failing script that left a background job running took its reason from
whatever that job wrote to stderr after bash exited, often the line that
Argent's own SIGTERM provoked ("mock-server: SIGTERM received, closing").
The reason now takes the stderr line as it stood when the runner reported
bash's exit whenever the stop had to end a process still holding stderr.
A step whose streams close on their own keeps the last line overall.
2026-09-10 22:06:37 +02:00
Ignacy Łątka b27f222386 fix(flows): refuse to record a step whose native-devtools precheck blocked (#1116)
Fixes #935.

**Cause:** a blocked native-devtools precheck is a *resolved* result, so
`flow-add-step` read "it returned" as "it ran" and wrote the step - a
`restart-app` that terminated and launched nothing became a passing
`launch:`, while the runner scores that same result an error at replay.

**Fix:** the recorder reads the runner's own
`isNativeDevtoolsBlockResult` right after the sub-invoke and fails the
call, recording nothing. Keyed on `params.command`, so it covers all
nine precheck tools, not just the `restart-app` in the report.

**Left alone:** `isDebuggerNotConnectedResult` is the same shape on the
recorder path and is filed separately as #964;
`nestedOrchestratorOutcome` is #719's scope.

**Docs:** none needed - no tool, CLI, config key or flow directive
changed, and `flow-add-step`'s description already states that an error
means nothing was recorded.

<details>
<summary>Verification</summary>

New
`packages/tool-server/test/flows/flow-record-native-devtools-gate.test.ts`
- all four block statuses on `restart-app`, the raw `tool:` spelling on
`native-full-hierarchy`, plus two negative controls (a green
`restart-app` still records `launch:`; a `gesture-tap` answering
`{status: "connect_pending"}` still records).

Discriminates - with the source hunk stashed:

```
Tests  5 failed | 2 passed (7)
AssertionError: promise resolved "{ ...(5) }" instead of rejecting
+   "message": "Step added to \"nd-rec\" flow",
+   "recorded": "1. tool: native-full-hierarchy {\"bundleId\":\"com.example.app\"}",
```

Restored:

```
Test Files  1 passed (1)
     Tests  7 passed (7)
```

- `npx tsc --build` - clean
- `npm run typecheck:tests -w @argent/tool-server` - clean
- `npm test -w @argent/tool-server` - 369 files, 4839 passed, 1 skipped
- `npx prettier --check` on both files - unchanged
- eslint on the two files could not run locally
(`packages/docs/tsconfig.json` needs its own `npm ci`); CI covers it.

</details>

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 16:56:23 +02:00
Ignacy Łątka 3dbf7bd3cb fix(screenshot-diff): classify an undecodable PNG and name the failing side (#1147)
Fixes #818

**Cause** - `decodePngFile` let the pngjs/fs error through as-is, so an
undecodable input answered with a bare library string, no failure code
(bucketed `ARGENT_UNCLASSIFIED_FAILURE`), and no hint which of the two
sides was bad - both decode in one `Promise.all`.

**Fix** - wrap the decode in a `FailureError` carrying
`SCREENSHOT_DIFF_INPUT_INVALID` / `screenshot_diff_decode_failed` and
the offending path, matching `validateInputSources` in the same tool.
`flow`'s `snapshot` goes through the same `diffPngFiles`, so it gets it
too.

**Class** - the other `PNG.sync.read` sites decode bytes the tool itself
just captured, not a caller path: `vega-screen.ts` (emulator's own
output), `flow-visual.ts` `cropPngFile` (the capture it just wrote),
`flow-pixels.ts` (deliberately soft - absence of visual evidence). Left
as-is.

**Docs** - none needed: no tool schema, CLI, config key or flow
directive changed.

<details>
<summary>Verification</summary>

Three-case table test (`test/screenshot-diff.test.ts`), each failing on
the un-fixed code with exactly the strings the issue reports:

```
non-PNG bytes  expected Error: Invalid file signature to be an instance of FailureError
truncated PNG  expected Error: There are some read requests waitn… to be an instance of FailureError
a directory    expected Error: EISDIR: illegal operation on a dir… to be an instance of FailureError
```

Proven by stashing the source fix, running the file (3 failed | 21
passed), restoring it (24 passed).

- `npx tsc --build` clean
- `npm run typecheck:tests -w @argent/tool-server` clean
- `npm test -w @argent/tool-server` -> 368 files, 4835 passed | 1
skipped
- `npx prettier --write` on both changed files - unchanged

</details>
2026-09-10 16:29:32 +02:00
Ignacy Łątka a3312cab97 fix(tool-server): admit a digit-leading bundleId in the four app-lifecycle schemas (#1101)
Fixes #1024

**Cause.** `launch-app`, `restart-app`, `reinstall-app`,
`settings-permissions` and `open-url` each carried their own copy of
`/^[A-Za-z_][A-Za-z0-9._-]*$/`, which rejects an iOS bundle id such as
`9gag.app` - with a message that says digits are allowed, so the agent
reading it has nothing to act on.

**Fix.** One shared `BUNDLE_ID_PATTERN` / `BUNDLE_ID_MESSAGE` in
`src/utils/bundle-id.ts`. The head class becomes `[A-Za-z0-9_]` (the
same class `native-devtools`' handshake validator already uses); `-` and
`.` stay out, which is the whole of the flag-injection guard the old
comments cited. The message names the head rule it enforces.

**Class check.** No copy of the pattern remains (`grep
BUNDLE_ID_PATTERN`). `open-url` was a fifth copy that landed on main
while this branch was open - the catalogue-wide test caught it.
`ACTIVITY_PATTERN` in `launch-app` / `restart-app` keeps its letter-only
head deliberately: an activity is a Java class name and cannot start
with a digit.

**Docs.** No docs change needed - `reference/tools.mdx` lists these
tools by one-line summary and documents no parameter constraints.

<details>
<summary>Verification</summary>

New `test/bundle-id-pattern.test.ts` walks the whole catalogue rather
than one tool, so a sixth `bundleId` constraint has to satisfy the same
rules. Discrimination proved by restoring the old pattern and message in
the shared module and re-running:

```
 × derives a sane MCP JSON schema with bundleId required for every action
 × launch-app: admits a digit-leading bundle id and still refuses a flag
 × launch-app: names the head rule when it rejects one
 × restart-app / reinstall-app / settings-permissions / open-url: both, each
 Tests  11 failed | 61 passed (72)
```

With the fix in place, from the worktree root:

```
npx tsc --build                                  OK
npm run typecheck:tests -w @argent/tool-server   OK
npm test -w @argent/tool-server                  434 files, 6174 passed | 1 skipped
npx prettier --write <changed files>             unchanged
npx knip                                         bundle-id.ts absent from the report
```

`npx eslint` could not run in this checkout
(`packages/docs/tsconfig.json` cannot resolve `@docusaurus/tsconfig`
without a separate `npm ci` in that workspace, a pre-existing
local-environment issue); the ESLint CI job passes.

</details>

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 16:12:38 +02:00
Ignacy Łątka 9e196a97be fix(installer): quote the project root in the manual-install advice (#1099)
Fixes #1053

**Cause.** `installLocally`'s failure advice built `cd ${projectRoot} &&
${cmdStr}` with the root interpolated raw. A project path with a space
is cut in two when the user pastes it; one holding `$`, a backtick or
`\` still runs, aimed elsewhere.

**Fix.** `shellQuotePath` in `package-manager.ts` single-quotes the path
(escaping embedded `'`), and the advice uses it.

**Class.** This is the only discovered path interpolated into printed
shell advice - the two other "Install manually with" lines print
`cmdStr` alone. `cmdStr`'s one variable argument is `--from <tarball>`,
a path the reader typed themselves and gets back in the form they gave;
left as is.

**Docs.** No tool, CLI flag, config key or flow file changes - nothing
in `packages/docs/` needs an update.

**Verified.** New test in `install-runner.test.ts` drives a failing
local install with the root `.../My 'Project' $x`, extracts the printed
`cd`, and runs it under `/bin/sh` - `pwd` must come back as the project
root. It reproduces the issue's error on the unfixed line.

<details>
<summary>Test discriminates - fix reverted</summary>

```
 FAIL  test/install-runner.test.ts > installLocally failure handling > quotes the project root in the manual-install advice
/bin/sh: line 0: cd: /var/folders/.../argent-install-advice-UW3dHe/My: No such file or directory
 Test Files  1 failed (1)
      Tests  1 failed | 6 passed (7)
```

with the fix in place:

```
 Test Files  1 passed (1)
      Tests  7 passed (7)
```
</details>

<details>
<summary>Checks</summary>

- `npx tsc --build` - clean
- `npm run typecheck:tests -w @argent/installer` - clean
- `npm test -w @argent/installer` - 17 files, 556 tests passed
- `npx knip` - 30 lines, identical to the same run on the unchanged tree
- `npx eslint` on the three changed `src/` files - clean; `npx prettier
--write` applied
</details>

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 16:11:36 +02:00
Ignacy Łątka b14d1fb478 fix(settings-permissions): 400, not 500, for a permission the platform cannot map (#1100)
Fixes #1032

**Cause** - both `settings-permissions` platform arms rejected an
unmappable permission with a bare `FailureError`. The HTTP dispatcher
keys the status on the error *class*, never on `error_kind`, so it fell
through to **500** and the agent read caller input as the retryable
class.

**Fix** - throw `InvalidToolInputError` instead (same message, same
`SETTINGS_PERMISSION_UNSUPPORTED` signal), which the dispatcher maps to
**400** while keeping the granular telemetry bucket. Same shape as the
keyboard backends' "no keycode for this character" arm.

Both sites of this class inside `settings-permissions` are fixed (iOS +
Android; there is no third platform arm). The two sites the issue also
names in `devices/boot-device.ts` are left alone:
`BOOT_IOS_UNSUPPORTED_HOST` (iOS asked for on a non-macOS host) and
`VEGA_ALREADY_RUNNING` (a device-state conflict a `force:true` retry
does resolve) each need a different class - a design call in its own
right, not this fix.

No docs change: `packages/docs` documents what the tool does, not
per-tool HTTP status codes.

<details>
<summary>Verification</summary>

Discriminating test - reverted just the two source files to `main` and
re-ran the suite:

```
FAIL test/settings-permissions.test.ts > settings-permissions iOS branch > notifications is rejected as unsupported without calling simctl
  expected FailureError: Permission 'notifications' … to be an instance of InvalidToolInputError
FAIL test/settings-permissions.test.ts > settings-permissions Android branch > reminders is rejected as unsupported (no Android equivalent)
  expected FailureError: Permission 'reminders' has … to be an instance of InvalidToolInputError
Tests  2 failed | 59 passed (61)
```

With the fix restored: `61 passed (61)`.

Also run from the worktree root:

- `npx tsc --build` - clean
- `npm run typecheck:tests -w @argent/tool-server` - clean
- `npm test -w @argent/tool-server` - full package suite green
- `npx prettier --check` on the three changed files - clean
</details>

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 15:49:15 +02:00
Ignacy Łątka 15e369eff5 fix(tool-server): stop asking about --remote-debugging-port on a port that answered (#1138)
Fixes #880

**Cause.** `discoverPrimaryPage`'s no-page-target branch is only reached
after `GET /json/list` succeeded, so the app provably has
`--remote-debugging-port` (a missing flag fails earlier, as
`CHROMIUM_CDP_UNREACHABLE`) - yet the message asked about the flag and
withheld the window hint its `devtools://` sibling gives.

**Fix.** One message: the endpoint answered but exposes no page target,
i.e. the app is running with no window. Same shape as the sibling branch
two lines up.

| Before | After |
| --- | --- |
| `Chromium CDP on port N reported no page targets. Is the app started
with --remote-debugging-port=N?` | `Chromium CDP on port N answered but
exposes no page target (the app is running with no window). Open an app
window and retry.` |

Why it matters: `CHROMIUM_CDP_NO_PAGE_TARGET` maps to `cdp_unreachable`,
whose Chromium guidance says "see detail" - this string *is* the detail,
and its relaunch implication leaves a second copy contending for the
single-instance lock while the windowless one still holds the port.

**Class.** The other `--remote-debugging-port` question in the file
(`fetchJson`'s connect catch) is correct - it fires only when the port
did not answer - and is left. No docs/skills change: no prose surface
quotes this string, and no tool, CLI flag, config key or flow behaviour
changed.

<details>
<summary>Verification</summary>

New test
`packages/tool-server/test/chromium-no-page-target-message.test.ts`
drives `discoverPrimaryPage` against a local HTTP server serving a
`/json/list` with no `page` entry (and, for the sibling, only a
`devtools://` page).

Discriminates - source fix stashed:

```
FAIL test/chromium-no-page-target-message.test.ts > reports a windowless app when the list holds no page target at all
AssertionError: expected 'Chromium CDP on port 55564 reported n…' not to match /--remote-debugging-port/
+ Received: "Chromium CDP on port 55564 reported no page targets. Is the app started with --remote-debugging-port=55564?"
```

Fix restored: 2 passed.

Ran from the worktree root:

- `npx tsc --build` - clean
- `npm run typecheck:tests -w packages/tool-server` - clean
- `npm test -w packages/tool-server` - 4833 passed, 1 skipped, 1 failed:
`lens-tools-platform-gate.test.ts` timing out at 5000 ms, the known
full-suite/cold-cache flake. Reproduced on a clean tree and passes on
3/3 warm reruns with the fix applied; it imports nothing this change
touches.
- `npx prettier --write` on both changed files - unchanged

</details>

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-09-10 15:40:33 +02:00
Hubert Gancarczyk 4edac01700 fix(flows): read a Windows runner that ended past its limit as the timeout
The parent classifies a runner that ended with no verdict by how it ended,
and reads the clock rather than its own timer for one case: a SIGKILL past
the step's limit is the deadline watchdog's stop, not an unexplained kill,
because a stall in the tool server's event loop can hold the timer behind
the exit it is racing.

Windows has no signal for that rule to see. The watchdog's `taskkill` ends
the runner with exit code 1, so a stalled tool server's deadline stop there
read as "The script stopped its own process with exit code 1" - the case
the Windows runner kept reporting as `exit` after the runner itself stopped
sending bash's forced exit. Past the limit, a runner that ended with no
verdict on Windows is now that same stop. POSIX is unchanged: the rule
still needs a signal there.

The deadline case asserts the whole failure now, so a wrong verdict says
which side produced it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X2EFhLrnjAkfgPkN1fva1k
2026-09-10 15:24:16 +02:00
Hubert Gancarczyk 8bcc26ec0c test(flows): let every Windows-run bash case watch a pid Windows knows
Two more cases read their descendant's pid from `sleep & echo $!`: the time
limit and the cancellation, both of which then ask `process.kill(pid, 0)`
whether it is gone. Under Git Bash `$!` is an MSYS number, not a Windows pid,
and a small one can belong to an unrelated Windows process - the time-limit
case passed on one Windows run and failed on the next at "no process left
behind".

One helper now starts the descendant as a Node process that writes its own
`process.pid`, as the lifeline case already did, and the deadline case uses
it too. The POSIX-only cases keep `$!`, which is a real pid there.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X2EFhLrnjAkfgPkN1fva1k
2026-09-10 15:18:20 +02:00
Hubert Gancarczyk 4d946c7a6c fix(flows): report a Windows deadline kill as the timeout it is
The deadline watchdog is the child's own bound for a tool server whose event
loop has stalled. On POSIX it stops the step with one group SIGKILL, which
takes the runner and bash in the same instant. Windows has no group, so it
runs `taskkill /t`, and that takes the tree one process at a time: bash can
die first, and the runner's main thread then reported bash's forced exit as
the script's own - "The script exited with code 1" - which the parent, its
loop just woken, accepted as a verdict already sent.

The watchdog now sets a flag shared with the runner's main thread before it
stops anything, and `finish` sends nothing once it is set. The parent's own
timer, which is always due before the watchdog's deadline, then reports the
step as the timeout it was. Setting the flag first means bash cannot die
before the flag is visible. The POSIX path is unchanged: the group SIGKILL
left no thread to send from.

Found once the deadline case finally ran on Windows: the descendant it watches
now dies inside the margin there, and the verdict alone was wrong.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X2EFhLrnjAkfgPkN1fva1k
2026-09-10 15:18:20 +02:00
Hubert Gancarczyk 409ce3386f test(flows): run on Windows only the bash cases whose semantics it has
With the IPC fix in, ten bash cases still failed on the Windows runner.
None is a product failure; each asserts something Windows does not do, or
measures it in a way that cannot work there.

- A named pipe as the document: Git Bash's `mkfifo` makes a Cygwin FIFO
  that Node sees as no file at all. The directory row still runs there.
- Three CRLF cases: Git for Windows' bash drops a carriage return from the
  line it reads, so a CRLF script runs there as its LF twin - all three
  ran clean on the runner. POSIX only.
- The stray-sibling case wrote a literal carriage return inside `$'...'`,
  which that same rule turned into an empty word, so the redirection
  emptied the real document. It now writes U+F00D on Windows, the name
  msys2 gives such a file, as the case beside it already does.
- Three signal-verdict cases: the step dies by a signal, and Windows has
  none to send. POSIX only, as the file already says of signal verdicts.
- The deadline-watchdog case found its descendant through `sleep & echo
  $!`, and under Git Bash that pid is an MSYS number `process.kill` cannot
  ask Windows about, so it read as dead from the start. A Node descendant
  writes its own pid now, as the lifeline case does - which makes this the
  case that runs the watchdog's `taskkill` arm on Windows.
- The sweep's 12 ms stall bound sits under Windows' ~15.6 ms timer tick,
  so a 5 ms heartbeat cannot resolve it there. Skipped on Windows.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X2EFhLrnjAkfgPkN1fva1k
2026-09-10 15:09:49 +02:00
Hubert Gancarczyk 38d36ce8d1 fix(flows): stop writing raw frames onto a Windows IPC channel
Every bash step on Windows ended as "The script runner exited before it
started the script (exit code 1)", bash included - its log carried Git
Bash's `child_copy ... Win32 error 299` and `dofork ... died unexpectedly`,
the marks of a parent taken away mid-fork. The same failure on the head
that predates dropping #970 shows it is this branch's own, never seen in
CI before because only the description check ran on it.

A breadcrumb build put `disconnect` first in every step's log. In bash
mode `announceStarted` writes `started` synchronously onto the channel's
descriptor, so that a script killing the runner on its first line cannot
race it. That descriptor carries newline-delimited JSON on POSIX, but on
Windows libuv frames each IPC message with a header of its own, so the
raw line is a frame the parent cannot read: it closes its end, the runner
reads `disconnect`, and in bash mode `exitOnParentDisconnect` stops the
whole tree with `taskkill /t /f` - exit code 1, before any verdict.

`sendSynchronously` now has no descriptor on Windows, so `started` and
every verdict go through the queued `process.send` there, which frames
them. The POSIX path is unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X2EFhLrnjAkfgPkN1fva1k
2026-09-10 14:58:25 +02:00
Hubert Gancarczyk 0a129b204b test(flows): hold the mode and group-kill cases on Linux as well
Both were written against macOS and had never run on Linux: until this
branch targeted `main`, only the description check ran on its pull request.

The mode case read a file's mode with `stat -f '%Lp' "$1" || stat -c '%a'
"$1"`. GNU `stat -f` is `--file-system`, so on Linux it treats `%Lp` as a
file it cannot find and still prints the file system `$1` sits on to
stdout before the fallback runs - and that report became the document the
script wrote. GNU's `-c` comes first now; BSD `stat` refuses it with
nothing on stdout.

The body `kill 0` case pinned macOS's answer, where the signal reaches bash
alone and arrives as a plain signal. On Linux it reaches the whole process
group, the runner that leads it included, which is the case the runner
already reports as a group signal. Both reports name `kill 0` and the
remedy, so the case now accepts either.

Reproduced in a `node:20-bookworm` container (GNU bash 5.2.15) before the
change - `expected 'exit' to be 'signal'`, and a document of `undefined`
- and green there after it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X2EFhLrnjAkfgPkN1fva1k
2026-09-10 14:54:03 +02:00
Hubert Gancarczyk e09e200eb5 feat(flows): let a flow run select a booted remote iOS simulator (#1060)
## What

A booted remote iOS simulator is now a device a flow run can be pointed
at - both by auto-detection and by `--platform ios-remote`.

## Why

It was invisible to flow device resolution. With a remote simulator
booted alongside local ones:

```
$ argent flow run checkout
7 booted devices matched - pass --device or --platform to disambiguate.
Available devices: 79EE0744 (ios, Booted), A1E0DF35 (ios, Booted), emulator-5554 (android, device), ...
```

Three local iPhones and four Android emulators, and the remote simulator
absent entirely. Narrowing to it did not work either - `--platform
ios-remote` was rejected by the parameter's enum. The only way onto a
remote simulator was typing the full `remote:<uuid>` id.

The cause is one type. `RawDevice.platform` was typed from the tuple of
platforms an author can *write* in a flow file, which has no
`ios-remote`, so nothing forced the case to be handled and two functions
quietly fell through:

- `isBooted` hit its `default: return false`, so a remote row never
counted as booted.
- `deviceEntryId` fell through to `serial`, a field a remote row does
not carry (it uses `udid`, prefixed `remote:`). Even once counted, it
had no id - and a resolution error printed it as `? (ios-remote,
Booted)`, hiding the very id the caller needed.

## How

The tuple served two different jobs at once - what an author may write,
and what a run may be pointed at - so it is now two:

- `LAUNCH_PLATFORMS` - launch-map keys and `when:` guards. Unchanged.
- `SELECTABLE_PLATFORMS` - device selection. Adds `ios-remote`.

`ios-remote` is selectable but deliberately not writable. A flow says
what it drives, not which machine hosts the simulator, so `when: {
platform: ios-remote }` stays a parse error.

For the same reason `ios` and `ios-remote` do not overlap: `ios` selects
local simulators only. The flow file is identical either way, so which
host runs it is an explicit choice rather than something that changes
with whatever happens to be booted.

## Scope

Selection only. A remote flow run is still coordinate-only - selector
steps, `await: { idle: true }` and `snapshot: { cropOn }` need a UI tree
source that `ios-remote` does not have yet, and `launch:` maps still
have no `ios-remote` key. Those are separate changes.

## Verification

End-to-end against a booted cloud simulator (`remote:73A22194-…`, iPhone
17 Pro / iOS 26.5) driven through the branch tool-server, with three
local simulators and four Android emulators also up:

| Run | Result |
| --- | --- |
| `platform: "ios-remote"` | resolves `remote:73A22194-…`, flow passes |
| no platform | 8 booted devices listed, including `remote:73A22194-…
(ios-remote, Booted)` - named by its id, not `?` |
| `platform: "ios"` | 3 local simulators only, the remote one absent |

`npm run build` and the tool-server suite pass.

## Docs

`flow-yaml.mdx` and the `argent flow run --help` text both list
`ios-remote` and note that `ios` never picks a remote simulator.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added support for selecting remote iOS simulators with `--platform
ios-remote`.
* Remote simulators are automatically detected when booted and
identified by their device ID.
  * Local `ios` selection continues to target only local simulators.

* **Documentation**
* Updated command-line help and flow documentation to describe remote
iOS simulator support and platform selection behavior.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Hubert Gancarczyk <claude-hubert.gancarczyk@swmansion.com>
2026-09-10 14:31:02 +02:00
Hubert Gancarczyk 4c71bc911d fix(flow): raise the flow tree's depth cap and say why targeting failed (#721)
Rebased onto `main` (was stacked on #720), with a single hunk dropped —
see
**Restack notes** at the bottom. Two commits: `a71c0f0` is exactly this
PR's own
delta from the stacked version, `54fe761` is a review pass over it on
the
post-rebase base. Squash on merge.

## Depth cap: 40 → 100

The flow selector tree capped raw `UIView` nesting at **40**. That does
not clear the invisible wrapper layers React Native stacks on every
screen — nested navigators (`RNSScreenStackView` → `RNSScreenView`,
often two deep) plus a drawer/root wrapper routinely bury on-screen
content **40–60 levels down** before the first tappable view.

In a deeply nested production app, plainly visible interactive elements
sat at depths **41–62**. So `id:` and `text:` selectors silently failed
to resolve, and only coordinate taps worked — which is the origin of
coordinate-heavy flows on those apps.

The overflow is **silent**: nothing reports "truncated", selectors just
stop matching. So the new cap carries generous headroom rather than
hugging the deepest observed measurement.

**What it costs.** The tree itself is internal to selector resolution —
consumed by `selectorToFrame`/`evaluateCondition`, never returned — so
no tool result carries it. Text derived from it does reach the agent,
though: `assertReason`'s `text` arm quotes the matched node's hoisted
`subtreeText` verbatim into a failing `assert`/`await` reason, and
nothing truncates it downstream. Descendants a depth-40 read dropped now
hoist, so that quoted string grows with the cap. That is the trade —
under the old cap those selectors did not resolve at all. Otherwise the
cap only grows the `getFullHierarchy` payload over the native-devtools
socket, which is already field-limited: ~11KB at depth 40, ~15KB at 48,
and nothing at all past the tree's real depth.

## Targeting errors that could not be acted on

**`Provide bundleId explicitly`** — `resolveNativeTargetApp`'s own
errors close with that advice. A flow selector step cannot follow it:
the call hardcodes auto-targeting. Each failure now carries the remedy
that actually exists:

- **ambiguous connected set** — foreground the intended app with
`launch-app` (which does not terminate, so the instrumentation it
already has survives), and clear the others with `xcrun simctl
terminate` (argent exposes no terminate tool, and `restart-app` would
just bring that app back to the front). The command is resolved through
`simctlTargetForUdid`, so it carries the device's own CoreSimulator
`--set` — a simulator from a configured `ios.additionalDeviceSets`
resolves — and passes the same entitlement gate an actual spawn does. An
external provider (#735) that grants `native-devtools` but withholds
`simctl` gets the clause dropped whole, rather than a command argent
itself would refuse and its device-set path quoted back.
- **a lone connected app that is not foreground** — same foreground
remedy; it is already instrumented, so a relaunch would only discard
state.
- **the state probe failed while connections are live** — which app is
missing decides the remedy. While the app the flow launched is still in
the live connections map it is merely suspended: foreground it, do not
relaunch, since a relaunch discards the state the flow built up and
cannot help when another app is the stale one. When it is *absent* from
that map it is the one thing a relaunch fixes, so the reason names it
and sends it to `restart-app` — `launch-app` cannot instrument a process
that is already gone.
- **no connected app** — this is the one state a relaunch fixes, through
`restart-app` or a flow `launch` step.

**`launch-app, or a flow launch step`** — this was wrong in a way that
wastes a cycle. `launch-app` does **not** terminate first, so against an
app already running from Metro/Expo, Xcode, or its home-screen icon it
only foregrounds that same uninstrumented process, and the next read
fails identically. Only `restart-app` (terminate + relaunch) guarantees
an instrumented launch whatever was already running.

These reasons are repeated per step — the recorder embeds one in the
warning for every captured tap, and a failing `await:` repeats it per
poll — so the two that enumerate connected apps are capped at two
entries plus a count of what was left out, and the whole set is held
under a documented ceiling by a test that iterates every branch. That
ceiling is 800 characters. It is set by the `unregistered` reason, which
measures 738 plus the bundle id, so it holds for the ids seen in
practice and not for every one: a 63-character id clears it. Every
figure quoted at `MAX_LISTED_APPS` and `MAX_TARGETING_REASON_CHARS` was
re-measured against the current code.

`FailureError` data is preserved when these messages wrap an underlying
error, so the failure code still reaches callers that switch on it.

## Recorder: role-only selectors

A tapped node with no identifier and no visible text replays on role
alone, which holds only while that element keeps winning
`selectorToFrame`'s ranking. The raised depth cap makes this more common
— an unlabeled icon the old cap truncated away, which left `nodeAtPoint`
to pick its `testID` container, is now present and is the smaller frame
under the tap. The recorder now warns instead of recording that
silently.

## Launch gate: why the wait is per bundle

`treeSourceGate` already waits for every bundle, `com.apple.*` included,
and withholds only the *verdict* for a non-injectable one. This PR
documents why that wait is load-bearing rather than a readiness
formality: it is what ties the launched bundle to the app a later
selector step auto-targets, since `resolveNativeTargetApp(api,
undefined)` never compares the two. Skipping it per bundle could hand
the next step a different app's hierarchy and still report the run
green.

The matching fallback advice lands in the no-connected-app reason: for a
`com.apple.*` bundle that never connects, drive the app with raw point
taps and `tool: await-ui-element` steps, which read the AX tree instead
of selectors.

## Restack notes

- `a71c0f0` keeps the stacked version's per-file insert/delete counts
unchanged. `54fe761` is the review pass on top: the entitlement gate
above, the launched-app-gone remedy, four comments this branch left
stale or overstated, and the re-measured figures.
- No docs update needed — nothing in `packages/docs` documents the depth
cap, the targeting reasons or the recorder warning.
- Conflict resolved in `flow-ios-tree.ts`: #944 (physical iOS support)
landed new device-tree functions in the same region. Both blocks kept —
the targeting helpers sit next to `queryFullHierarchyTree`, which they
serve; the device block follows.
- **One hunk dropped:** a 35-line test in
`flow-hidden-blank-reads.test.ts` asserting that an `exists` miss does
not quote hoisted `subtreeText`. It is written against
`compatibilityMissNote` / "typographic variant", which #720 introduces
and `main` does not have. Worth re-adding to that PR, or here once #720
lands.
- Verified on this base: `tsc --build` and `typecheck:tests` clean; 254
tests across the 4 touched test files pass; full tool-server suite 433
files / 6190 passed / 1 skipped (one run showed a single unrelated load
flake, green on a clean re-run); eslint, prettier and both knip passes
clean. All 8 CI checks green on `54fe761`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

https://claude.ai/code/session_014dA4nMpWrzECNhV6if4YXZ


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- iOS hierarchy inspection now supports substantially deeper views,
improving selector and condition handling for deeply nested content.
- iOS targeting failures provide clearer, actionable recovery guidance,
including relevant app and device-state details.
- System-app flows can continue successfully when apps do not establish
a connection, including coordinate-based actions.

- **Bug Fixes**
- Captured tap steps now warn when selectors rely only on an element’s
role.
- Related selector warnings are combined into a single concise warning.
  - iOS failure messages no longer include misleading recovery advice.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Hubert Gancarczyk <claude-hubert.gancarczyk@swmansion.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-10 14:30:34 +02:00
Hubert Gancarczyk 0f5a96c87b docs(flows): describe a bash step's log and its stderr reason
The reference gains the log paragraph in the shared script section - it
never described the `.mjs` log either - and its bash section teaches the
two-part reason, the `echo … >&2; exit 1` idiom and what
`set -euo pipefail` already gives. The skill's `flow-yaml.md` says the
same in its own words, and the CRLF rule points at the line bash writes
to the log. `flow-skill-docs.test.ts` fails if `$ARGENT_REASON` comes back
to either, or if either stops teaching the stderr reason.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X2EFhLrnjAkfgPkN1fva1k
2026-09-10 14:29:15 +02:00
Hubert Gancarczyk 58822a6abd feat(flows): report a bash step's output, and take its reason from stderr
A `.sh` step now reports what it prints the way a `.mjs` step does on
`main`: bash inherits the runner's stdout and stderr, which are the pipes
`ScriptLogCapture` reads, so the log, its limits, the drain and the secret
scrub apply with no bash-specific code.

The failure reason comes from stderr rather than from a file. bash has no
`throw`; what it has is an exit code and stderr, and every command a
script runs already writes its error there. So a non-zero exit reports the
runner's exit line followed by the last non-blank line the script wrote to
stderr - under `set -euo pipefail`, the error of the command that stopped
it, with no Argent-specific code in the script.

The runner never reads stderr: bash writes straight into the parent's
pipe. The executor's capture keeps the line (`LastLineTracker`), from the
raw text ahead of the scrub and the limits, so a script that floods stderr
and then says why still says it. Only the head of a line is kept - 1,000
characters and the usual omission marker - so a cut always lands at the
end of the message, where `redactTruncated` repairs one, and the line
joins the failure text before it is clamped and redacted. A line of
nothing but whitespace is blank at any length. stdout is never the
reason, and only the `exit` kind gets the line.

`$ARGENT_REASON` goes, and with it the runner's bounded read of the
reason file, its kept-count marker and `REASON_KEPT_RE`, its UTF-8
refusal, the CRLF hint for a reason that strayed, and the second exchange
file. `ARGENT_REASON` is an ordinary environment name again; the CRLF hint
for `$ARGENT_OUTPUT` stays.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01X2EFhLrnjAkfgPkN1fva1k
2026-09-10 14:29:15 +02:00