mirror of
https://github.com/callstack/agent-device.git
synced 2026-09-14 20:06:34 +08:00
v0.20.7
1372 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
700d85ec54 | 0.20.7 v0.20.7 | ||
|
|
c2c81549d9 |
feat: add simulator verification skills (#1716)
* feat: add simulator verification skills * chore: simplify simulator skills * docs: refine simulator skill guidance * test: guard simulator skill workflows * style: format simulator skill contract test |
||
|
|
fa9a350361 |
docs: prefer design fixes over regression-only guards (#1722)
* ci: require simplicity review for large tooling changes * docs: prefer design constraints over regression-only fixes * docs: simplify design-first guidance |
||
|
|
05a1d76f2e |
test: add daemon RPC wire-surface compatibility gate (#1717)
* test: gate daemon RPC wire compatibility against the last released tag (#1432) ADR 0006 fixes exactly when DAEMON_RPC_PROTOCOL_VERSION must be bumped, and nothing checked that it was. The runtime guard (readRemoteDaemonHealth) refuses a mismatched peer, but only fires when someone remembered the bump — a wire change that skipped it left both sides advertising protocol 2 while parsing different payloads, which is the failure ADR 0006 exists to prevent. Local daemons cannot skew (isReusableDaemonInfo takes over on any package version mismatch). Cross-machine is skewed by design — proxy, cloud/limrun, a remote macOS host — and ADR 0006 explicitly rules package version out as the compatibility gate there, so the one boundary where skew is intended was the one boundary with no gate. test/wire-compat/surface.ts declares the wire surface grouped by the ADR bullet each group serves, quoting it, with an `uncovered` note where a bullet is only partly digestible (the /health and /rpc literals inside http-server.ts stay reviewer-owned: a moved route 404s at connect time rather than misparsing). ledger.json records what each declaration hashes to, at which protocol version. Two gates, split for the same reason the replay-compat corpus splits: - unit-core holds the ledger to its source and prints the digest to paste; - Released-Surface Compatibility reads the ledger at the last RELEASED tag and requires the drift since then to carry a bump or a compatibleChanges ack. From one commit a bumped ledger and an unbumped one are both just an edited file, so only a released baseline can tell them apart. Acks are keyed by the digest they cover, so one "added an optional field" cannot launder later changes. Digests ignore comments and formatting; the manifest's closure is derived from the AST, so a field typed by an unlisted sibling fails rather than sitting outside the gate. CI cost: one added job (checkout + toolchain + two node scripts, ~1 min), mirroring the existing full-history replay-compat job. * test: close wire-surface overclaim and make the closure fail closed (#1432) Addresses both review P1s on #1717. P1 — the manifest materially overclaimed ADR 0006 coverage. It quoted all four bullets while digesting only the payload TYPES, so the producer and consumer seams could break a skewed peer without moving a listed digest. Now listed on both sides of every boundary: JSON-RPC method sets and the projections that turn each method's params into a DaemonRequest, createRpcError/sendJson/ writeRpcResponseEnvelope, resolveToken and the auth-hook types, upload preflight/finalize/308 handlers and the resumable ticket shape, artifact route and download/inventory framing, REST error mapping, and the client's own payload builder, lease-method mapping, response parser and error projection. 57 -> 117 declarations. What stays out is now named rather than implied: createDaemonHttpServer's dispatch wiring and the /health and /rpc literals inside it. Everything it dispatches WITH is digested individually, and a moved route 404s at connect time rather than misparsing — the loud failure, not the silent one. P1 — imported and re-exported payload shapes escaped the closure. declarationHomes() scanned only the manifest's own files and the walk continued silently when a name could not be placed, so a listed type could gain foo?: ImportedShape from a new module and stay green. Resolution is now explicit and fails closed: relative imports, workspace specifiers (through the owning package's own exports map, so a re-pointed export cannot drop a type), and facade re-export chains. Every referenced name must land on a listed declaration, a waiver with a written reason, a declared external module, or the TS/Node global set. Fixed two extractor blind spots the walk exposed: a declaration's own generic parameters and `as const` were being reported as references. Planted-red proofs (wire-mutations.test.ts): 13 cases independently mutate method naming, response serialization, response parsing, auth projection, upload ticket shape, 308 framing, artifact framing, REST error mapping, and progress framing, each asserting the digest moves; 3 probes prove the closure really reaches across a package boundary, a facade re-export, and a plain relative import. Mutations apply inside the declaration's own span — a whole-file replace silently hit a sibling sharing the substring, which is how the first draft of one case passed vacuously. The largest waiver pair (InternalRequestOptions, CommandFlags) rests on ADR 0006's own additive rule: they reach the peer inside DaemonRequest's untyped flags/input bags, and the decision says a new flag needs no bump. Digesting them would fire the gate on every new CLI flag and train reviewers to rubber-stamp acks. * test: list the consumer half of the auxiliary HTTP boundaries (#1432) Addresses the remaining review P1 on #1717. The manifest claimed both sides of response/upload/artifact framing while listing nothing from upload-client.ts, daemon-artifacts.ts, or the health consumer in daemon-client-transport.ts, so those parsers could narrow without moving a listed digest or protocol 2. Now listed (117 -> 141 declarations): - /health consumer: RemoteDaemonHealth, readHealthPayload, readDaemonHttpHealth, readRemoteDaemonHealth. This is the sharpest of the three — narrowing the reader or the comparison disables the very refusal ADR 0006 exists to guarantee, and nothing else in the repo would notice. - /upload consumer: UploadResponse, UploadPreflightResponse, UploadPreflightResult, parseUploadPreflightResult, requestUploadPreflight, uploadDirectArtifact, tryDirectUploadWithResume, shouldRetryDirectUpload, finalizeDirectUpload, uploadLegacyArtifact, ARTIFACT_HASH_ALGORITHM, isStringRecord, and PreparedUploadArtifact — whose sha256/sizeBytes/fileName/artifactType/ contentType fields ARE the preflight body the daemon parses. - /artifacts/* consumer: DaemonArtifactEndpoint, buildDaemonArtifactUrl, isRemoteDaemon, DownloadRemoteArtifactParams, downloadRemoteArtifact, materializeRemoteArtifacts, resolveMaterializedArtifactPath. Running the closure fail-closed over the new files surfaced three more stops, each decided rather than skipped: PreparedUploadArtifact listed (it is payload), UploadProgressSink waived (client-local rendering, never leaves the process), and src/daemon/types.ts#DaemonArtifact waived as a re-export alias of the listed kernel type, matching its DaemonRequest/DaemonResponse siblings. 10 more planted-red mutations cover the new seams: health version-read and mismatch-refusal defeated, RemoteDaemonHealth field dropped, preflight parser narrowed, preflight/legacy response shapes narrowed, finalize body key renamed, ticket field renamed, artifact tenant header dropped, artifact URL moved. A fourth closure probe proves the upload-consumer files are genuinely reached by the walk rather than merely listed. 22 -> 33 tests. The README now states the coverage as a producer/consumer table per boundary, so the claim is checkable at a glance instead of asserted in prose. * test: list the client half of the resumable 308 contract (#1432) Addresses the third review P1 on #1717. Listing the daemon's handleResumableUpload proved it still PRODUCES 308; nothing proved the client still CONSUMES the released one. src/remote/upload-stream.ts owns that half and was entirely outside the manifest, so a newer client could stop accepting `upload-offset`, change how it reads `Range: bytes=0-N`, or emit a different resumed `Content-Range` without moving one of the 141 listed digests. Now listed (141 -> 151): UploadStreamResponse, streamFileToHttpRequest, streamFileToHttpRequestAttempt, buildUploadRequestHeaders, isUploadResumeStatus, isUploadRedirectStatus, parseUploadResumeOffset, parseNonNegativeIntegerHeader, firstHeaderValue, MAX_UPLOAD_REDIRECTS. streamFileToHttpRequestAttempt is listed despite its size, unlike createDaemonHttpServer which stays in `uncovered`. The distinction is stated at the declaration: the HTTP server only dispatches to handlers that are each digested, while the attempt loop IS the resume state machine — it decides whether a 308 continues the upload and what the next request carries, so its sequencing alone can break a released daemon while every helper keeps its digest. 6 new planted-red mutations prove the client half moves the ledger: a dropped `upload-offset` fallback, narrowed Range parsing, a changed resumed Content-Range, 308 no longer treated as continue, a narrowed UploadStreamResponse, and dropped header-value coercion. 33 -> 39 tests. Closure fail-closed surfaced two more stops: UploadStreamProgressOptions waived (local byte-progress rendering) and URL/URLSearchParams added to the global set. README now carries a `/upload` resume row in the producer/consumer table, and names the pattern behind three rounds of review: the coverage sentence kept getting written ahead of the coverage, so the table and the `uncovered` notes are the claims to trust — they are checkable against surface.ts, prose is not. --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
cdc754e6ed |
perf: speed up iOS agent recovery and streamline CLI guidance (#1700)
* Avoid interactive children in parent taps * docs: streamline no-skill CLI help * perf: recover faster from sparse iOS trees * fix: preserve selector context for blocked parent taps * fix: preserve coordinate text-entry focus * fix: preserve thin parent touch targets * fix: fail closed for unscoped iOS typing * test: isolate replay lock fixture * test: share node integration process |
||
|
|
4f9aded0b8 |
fix(record): make an aborted authoring recording terminal by construction (#1712)
* fix(record): make an aborted authoring recording terminal by construction
A second successful `open` on an `open --save-script` session aborts the
recording: the aggregate goes to `authoring{aborted}`, `recordSession` is
cleared, and the caller is warned. `close --save-script` then refuses it with
"Retry with plain close; it will tear down the session without writing."
That promise was not kept. When the second `open` itself carried
`--save-script`, the recorder's shared flag ingress re-armed `recordSession`
while leaving the status terminal, and a bare `close` published the full
session log — the writer gated only on `recordSession` and the repair variant,
so nothing on the ordinary authoring path refused an aborted lifecycle.
The abort is now terminal by construction rather than inert by ordering:
- `isAuthoringAborted` gives the pure aggregate one home for the question.
- `applyRecordedSaveScriptFlags` takes no branch for an aborted lifecycle:
it neither re-arms recording nor retargets the output path.
- `SessionScriptWriter` asks one `isPublicationWriteBlocked` question covering
all three reasons to publish nothing, so every path reaching the writer
(bare `close`, teardown, idle-reap, active publication) refuses it.
Armed recordings, published recordings, and every repair transaction are
unaffected; the control tests for those stay green against the pre-fix code
while the five new regressions go red.
Closes #1533
* docs: record the #1533 resolution in ADR 0016 and fix a stale symbol reference
The ADR 0016 close-time amendment still described #1533 as unresolved, and
the session-close.ts note named `isAuthoringAbortedWriteBlocked` — a private
helper that was folded into `isPublicationWriteBlocked` when the writer's
three sequential guards collapsed into one predicate, so the symbol names
nothing in the tree.
* fix(record): arm recording through the publication lifecycle, not around it
`recordSession` is an evidence-capture flag, but three surfaces set it
directly without consulting the publication aggregate, so it could
contradict a terminal ABORTED authoring status. The #1533 fix closed the
recorded-action ingress and made the writer refuse an ABORTED lifecycle,
then documented the remaining contradiction as acceptable — the writer's
own comment noted that "something can re-arm that boolean behind the
terminal status".
That something was live: `buildNextOpenSession` re-armed recording for any
`open --save-script`, and `applyOrdinaryScriptRecordingOpenOutcome` only
aborts a lifecycle that is still ARMED. A third `open --save-script` on an
already-ABORTED session therefore left `recordSession` true behind the
terminal status. The writer gate hid the publication symptom, but the
session kept paying recording-time costs for a recording that can never
publish: `recordSession` disables the direct iOS selector fast paths for
click and get, forcing every interaction onto the snapshot route.
Route the flag through one rule owned by the publication projection
(`recordSessionAfterSaveScriptFlag`), which answers "not recording" for an
ABORTED lifecycle on every surface that handles it — the re-open builder,
the close finalizer, and the recorded-action ingress. The writer's gate is
unchanged and still correct; it now stands on the aggregate alone rather
than as a net under a known drift, so the comments defending the drift are
replaced by statements of the rule.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFW9gJqz1wEHoowkdd1nFW
* docs: describe the full #1533 surface in the changelog entry
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFW9gJqz1wEHoowkdd1nFW
---------
Co-authored-by: Claude <noreply@anthropic.com>
|
||
|
|
f569b91d65 |
docs: record platform runtime adoption checkpoint (#1703)
* docs: record platform runtime checkpoint revision * docs: correct platform runtime checkpoint evidence * docs: finalize platform runtime checkpoint |
||
|
|
b15c502318 |
refactor: extract platform network runtime (#1702)
* refactor: extract platform network runtime * fix: preserve platform network recovery routes * test: guard network parser placement |
||
|
|
b1ed5353d1 |
refactor: extract platform log runtime (#1701)
* refactor: extract platform log runtime * fix: clear terminal app log recovery markers * fix: preserve scoped app log tooling * fix: preserve app log cancellation * fix: handle large changed coverage diffs * fix: harden Limrun runtime identity * refactor: tighten platform log runtime * fix: close app log trust gaps * fix: accept canonical session path aliases * refactor: extract durable capture kit * fix: refresh retained log marker admission * fix: rotate app logs after process relaunch |
||
|
|
4279d4c580 |
docs: align snapshot fallback and actions guidance (#1713)
* docs: align snapshot fallback and actions guidance
The snapshot guide claimed a zero-node XCTest result fails without ever
switching to AX, but regular iOS capture has an explicit recursive-tree →
query-sweep → private-AX recovery plan (ADR 0004,
RunnerTests+SnapshotCapturePlan.swift). The public CLI also exposes
`--actions`, `--force-full`, and `--timeout`, while both website reference
pages published a three-flag snapshot usage line.
- Add a schema-derived gate: `commands.md` must publish the exact usage
`buildCommandUsage('snapshot', getCliCommandSchema('snapshot'))` produces,
so the canonical invocation cannot drift from the command schema again. The
flag list is never restated in the test. Proven red against the pre-fix
`commands.md`.
- Extract the fence walker both doc checks now share, and prove the new gate
fails on a planted usage drift.
- Publish the canonical snapshot usage in the command reference and describe
`--actions` as iOS-simulator-only and planning-only.
- Replace the "Backends (iOS)" list with an iOS capture behavior section
written from ADR 0004 and the live capture plan: regular visible strategy
with a bounded recovery ladder, raw diagnostic strategy preserving strict
capture failures, and recovered/sparse/degraded output staying observable
through quality warnings. Capture tiers are documented as internal, not as
user-selectable backends.
Custom-action discovery stays separate from invocation: the runner can read
names but cannot trigger them (RunnerAXSnapshotBridge.h), so neither page
implies otherwise.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014VzyMim5q4jbp1xDMm3jVo
* docs: state that --actions and --raw are rejected as a pair
Both pages described the combination as a silent no-op ("returns a raw
tree without them"), which cannot happen: `customActionFlagsResponse` in
src/daemon/request-router.ts rejects `snapshotCustomActions` + `snapshotRaw`
with INVALID_ARGS at the shared request seam, before any session or device
work, so CLI, Node client, and MCP all get the same answer. Pinned by
src/daemon/__tests__/request-router-custom-action-flags.test.ts.
The underlying reason was right and is kept — custom actions are only
readable through the private-AX capture path, which the raw diagnostic
strategy does not take — but the user-visible outcome is a rejection, not a
degraded capture, so both pages now say to choose one flag or the other.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014VzyMim5q4jbp1xDMm3jVo
---------
Co-authored-by: Claude <noreply@anthropic.com>
|
||
|
|
b7470698a0 |
fix(test): include workspace packages in changed-line coverage (#1711)
The full Vitest coverage run instruments both root `src/` and `packages/*/src/**`, but the changed-line gate's prefilter rejected every package path before consulting LCOV. A PR could add uncovered lines to contracts, selectors, kernel, or provider packages while the 70% changed-line gate reported no package denominator. Make the source-root predicate package-generic so it accepts both coverage roots, while package tests (`.test.ts`, `__tests__/`), package-level `test/`, and `.tsx` stay excluded as before. Scoring, waivers, the excluded-line tally, and branch reporting are unchanged. Pinned at both levels: the classification matrix and a pure-model scoring regression in model.test.ts, plus a temporary-repository regression in run.test.ts proving the executable gate fails and names the uncovered package path and line. Claude-Session: https://claude.ai/code/session_01HgngjVMdk2eSKQ9d1gYLGo Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
4af1307024 |
test: cap Vitest workers for parallel worktrees (#1710)
* test: cap vitest workers for parallel worktrees * perf: leave Vitest workers uncapped in CI |
||
|
|
057f0e6eb1 |
fix(android): shell-quote free-text arguments reaching the device shell (#1645)
Text entry (input text) and clipboard write (cmd clipboard set text) now quote their free-form text argument with the same shellQuoteIfNeeded helper app-lifecycle.ts already uses for deep-link URLs and launch arguments, and app-lifecycle.ts's local duplicate of that helper is retired in favor of the shared one. Multi-word clipboard writes also now arrive at the device as a single argument instead of being re-tokenized into separate ones. Updates the provider-scenario test harness's scripted clipboard-state simulator to unwrap shell quoting the same way a device shell does, so it keeps modelling what the device actually receives. |
||
|
|
13cc90ffc6 |
fix: harden Android snapshot and fill reliability (#1708)
* fix: harden Android automation reliability * test: isolate CLI flush integration * test: close Android review gaps * test: register CLI transport fixture * test: consolidate CLI subprocess fixture |
||
|
|
c06bed9f77 |
refactor: extract platform device inventory runtime (#1699)
* refactor: extract platform inventory runtime * fix: preserve scoped Apple inventory tooling * fix: preserve Apple tool cancellation * refactor: tighten platform inventory boundaries |
||
|
|
44c298d7f3 |
docs: adopt request-bound platform runtime (#1697)
* docs: adopt request-bound platform runtime * docs(adr): record the rejected process-local live-handle ledger alternative |
||
|
|
4b432fb59b |
feat: add HarmonyOS support (#1683)
* feat: add HarmonyOS device automation foundation Add HDC-backed discovery, snapshots, application lifecycle, and core mobile interactions. Route HarmonyOS through the platform registry and client contracts. Cover parsing and capability parity with focused tests. * feat: support HarmonyOS HAP deployment Install and reinstall signed HAP archives through HDC. Resolve bundle identities from module metadata and relaunch after package replacement. Extend deploy routing and capability coverage for HarmonyOS. * feat: add HarmonyOS single-pointer gestures Execute pan, fling, and swipe plans through HDC uiInput primitives. Derive scroll coordinates from the live ArkUI viewport. Keep unsupported multi-touch gestures explicitly rejected. * refactor: split session inventory command handling Separate session, device, capability, and app inventory response paths. Preserve the public inventory response contract while reducing handler complexity. * feat: support HarmonyOS keyboard actions Route HarmonyOS enter, return, and dismiss through HDC key events. Expose supported keyboard actions through the system command metadata. Keep keyboard visibility inspection explicitly unsupported. * fix: reject unsupported HarmonyOS drag gestures Keep drag unavailable until HDC can preserve source and destination hold semantics. * feat: add HarmonyOS app log streaming 1. Stream HarmonyOS app logs through PID-scoped hilog sessions.\n2. Record HarmonyOS app identity during bundle-id opens for app-scoped commands.\n3. Cover backend routing and bundle identity resolution. * feat: report HarmonyOS foreground app state 1. Read the foreground HarmonyOS mission through aa dump.\n2. Expose HarmonyOS appstate with package and ability metadata.\n3. Add parser coverage for foreground and missing-state cases. * fix: advertise appstate through capabilities 1. Classify appstate in the command descriptor capability matrix.\n2. Surface supported appstate commands in capability inventory.\n3. Cover the advertised Android capability contract. * feat: sample HarmonyOS process performance 1. Sample HarmonyOS process CPU and resident memory through HDC.\n2. Expose the verified metrics through the shared perf command.\n3. Keep frame and memory snapshot collection explicitly unavailable. * feat: clear HarmonyOS app state 1. Add HarmonyOS settings clear-app-state through bundle cleanup.\n2. Force stop the app before clearing data and cache.\n3. Reject all unverified HarmonyOS settings explicitly. * docs: document HarmonyOS support 1. Describe HarmonyOS HDC prerequisites and HAP installation.\n2. Add HarmonyOS to platform discovery and product documentation.\n3. Document verified performance limits for the public HDC surface. * fix: preserve HarmonyOS deploy session identity 1. Bind a resolved HarmonyOS bundle after install or reinstall.\n2. Keep app-scoped logs and observability available after deployment.\n3. Cover session identity preservation for HarmonyOS reinstall. * test: lock HarmonyOS capability boundary 1. Add an independent HarmonyOS capability-matrix oracle and exact advertised-command regression test. 2. Document current HDC-backed support and evidence-based unsupported command boundaries. * refactor: simplify HarmonyOS shared platform boundaries 1. Split device selection and settings dispatch into focused helpers without changing behavior. 2. Keep HarmonyOS serial selection and lock-policy classification covered by regression tests. 3. Remove Fallow complexity findings from the HarmonyOS diff against upstream main. * fix: bound default HarmonyOS HDC commands 1. Apply a 15 second timeout to ordinary HDC operations. 2. Preserve operation-specific timeout budgets for installation and capture paths. 3. Add regression coverage for default and overridden HDC timeouts. * feat: add HarmonyOS screen recording Implement physical-device whole-screen recording through the system recorder and HDC media transfer. Reject unsupported HarmonyOS recording scopes and export flags. Cover capability routing, media retrieval, cleanup, and simulator rejection. * feat: report HarmonyOS HDC readiness Add an HDC version check to the HarmonyOS doctor flow. Document HarmonyOS as a supported doctor platform and cover the result. * refactor: simplify HarmonyOS recording checks Reduce recording validation and test complexity without changing behavior. * test: cover HarmonyOS platform contracts Synchronize public platform expectations across CLI, MCP, replay, and inventory tests. Mock HarmonyOS inventory probes to preserve concurrent test behavior. * test: model HarmonyOS recording capability Require a physical HarmonyOS device in the independent capability parity oracle. * test: cover HarmonyOS input and lifecycle paths Exercise HDC input, lifecycle, installation, and relaunch command sequences. * test: cover HarmonyOS device observability paths Exercise discovery, screenshot validation, and process performance sampling. * docs: define HarmonyOS CI hardware policy Keep HDC hardware validation local and require mocked CI contract tests. * fix: honor HarmonyOS app inventory filters * fix: bound HarmonyOS app inventory classification 1. 限制应用元数据分类并发并为默认清单设置整体时限. 2. 将请求取消信号传递给 HarmonyOS 应用清单读取. 3. 补充失败时中止在飞读取且不继续排队的回归测试. * fix: preserve HarmonyOS inventory failure causes 1. 保留触发应用元数据分类失败的原始错误, 避免被取消同级任务覆盖. 2. 补充总时限中止在飞读取且不启动排队任务的回归测试. 3. 验证后序任务失败时保留默认筛选的恢复提示. |
||
|
|
fcf429e0bd |
docs: split device cloud integration guides (#1695)
* docs: split device cloud integration guides * docs: refine device cloud integration copy |
||
|
|
18291ba8e2 |
perf: collapse app-driving startup turns (#1693)
* perf: collapse app-driving startup turns * fix: align foreground open guidance |
||
|
|
e18a183ac9 |
fix: harden artifact ingestion boundaries (#1692)
* fix: harden artifact ingestion boundaries * fix: bound archive inspection and upload expiry * fix: preserve upload preflight expiry |
||
|
|
588a419524 |
feat(mcp): serve the stateless 2026-07-28 revision alongside the legacy handshake (#1678)
* feat(mcp): serve the stateless 2026-07-28 revision alongside the legacy handshake MCP 2026-07-28 drops the initialize handshake: each request carries its protocol version and client capabilities in `_meta`, and clients probe `server/discover` to tell a modern server from a handshake-only one. agent-device answered that probe with -32601, so a dual-era client fell back to `initialize` and a modern-only client had no way to connect at all. Serve both eras from the one stdio process, which is what the spec calls a dual-era server: - `server/discover` advertises the supported revisions, the tools capability, and server identity. - A request declaring a protocol version in `_meta` is served modern: its result carries `resultType: "complete"` and `_meta["io.modelcontextprotocol/serverInfo"]`. - `tools/list` and `server/discover` return `ttlMs`/`cacheScope`, so clients can cache the 55-tool ~223KB list instead of re-fetching it every start. The list was already emitted sorted, which is the other half of what makes it cacheable. - A declared revision we do not implement is rejected with `UnsupportedProtocolVersionError` (-32022) naming the ones we do. Also fixes legacy version negotiation, which the era split surfaced: `initialize` returned 2025-11-25 whatever the client asked for, so a client pinned to 2025-06-18 was answered with a revision it had not requested — the lifecycle contract's cue to disconnect. It now echoes the requested revision when we implement it, and otherwise names the newest legacy one we do. Legacy responses are otherwise byte-identical: `initialize` and `ping` are still served, and no cache, `resultType`, or `_meta` field is added to them. The stdio transport, the tool set, and every tool's schema are untouched, so the CLI, Node, and daemon surfaces are unaffected. Era handling lives in its own module so the router stays a dispatcher. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TKWswqMZ6Z9UAQDAda83Xz * fix(mcp): keep 2025 revisions on the legacy wire contract and require modern metadata Review found the era model was too loose in two ways. Membership was one flat set, so a request declaring 2025-11-25 or 2025-06-18 through modern `_meta` was served the 2026-only envelope (`resultType`, `serverInfo`, cache hints) — fields absent from those revisions' schemas — and `initialize` would echo 2026-07-28, agreeing to a revision whose handshake the modern era removed. Split modern and legacy membership so the declared revision picks the wire contract: 2025 declared through modern framing is answered legacy-shaped, `server/discover` requires a modern revision because it exists in no legacy one, and `initialize` negotiates only within the legacy set. Modern request metadata is now required rather than guessed. `_meta` carries `protocolVersion` and `clientCapabilities` as required fields, so `server/discover` without them is malformed instead of being promoted to modern, and a half-declared `_meta` is rejected as invalid params (-32602) rather than having its lenient handling locked in by tests. Adds black-box router cases for declared-2025 requests, initialize(2026), `server/discover` with missing and with legacy metadata, and a declared revision without client capabilities. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TKWswqMZ6Z9UAQDAda83Xz * fix(mcp): validate supplied client identity and gate the methods 2026-07-28 removed Two protocol-boundary gaps from review. `clientInfo` was read but never checked. The field is optional in 2026-07-28, so its absence is fine, but a supplied one must be an `Implementation` — `clientInfo: 42` was accepted and the request served. A present value now has to carry string `name` and `version`, matching how `clientCapabilities` is already validated; omitting it stays legal. `initialize` and `ping` were served regardless of era, so a request carrying valid modern `_meta` could call methods its own revision deleted and get a `resultType: "complete"` envelope back — with `initialize` reporting a legacy `protocolVersion` inside a modern result. Both are now gated by the resolved era and answer -32601 to modern-framed callers, while metadata-free legacy calls keep working unchanged. Adds black-box router cases for modern-framed `initialize`/`ping` (each paired with its still-working legacy call) and for malformed `clientInfo`, including the omitted-is-legal case. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TKWswqMZ6Z9UAQDAda83Xz * fix(mcp): validate every recognized clientInfo field, not just the required ones The previous round checked `name` and `version` but let malformed recognized optional fields through, so `{name:'c', version:'1', websiteUrl:42}` — and the same for `title`, `description`, and `icons` — was accepted and served. `Implementation` validation now type-checks each recognized field when present: `title`, `description`, and `websiteUrl` as strings, and `icons` as an array of `Icon`, where `src` is required and `mimeType`, `sizes`, and `theme` are typed when supplied (`theme` against its `light`/`dark` union). Unrecognized keys still pass — `_meta` payloads carry extension fields, and rejecting those would reject the future. Adds regressions across both layers: a wrong scalar per optional field, a wrong icons container, an icon entry missing `src`, and each malformed typed icon member. The positive cases pin the other direction — a fully populated clientInfo carrying an extension key must still be served, so the validator cannot harden into rejecting what the spec allows. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TKWswqMZ6Z9UAQDAda83Xz * fix(mcp): accept schema-valid empty strings in clientInfo identity fields `isImplementation` reused `stringField`, which requires a non-empty string, for the required `name` and `version`. The 2026-07-28 schema declares both as plain `string` with no minimum length, so `{name: "", version: ""}` is a conforming `Implementation` and was being answered -32602. Required now means present and a string. `Icon.src` gets the same treatment. Its `format: uri` annotation is not something this server enforces — any other non-URI string is accepted — so rejecting the empty one alone was arbitrary rather than stricter. Adds positive regressions at both layers for empty `name`/`version` and an empty `Icon.src`, alongside the existing malformed cases, so the validator is pinned against over-rejection as well as under-rejection. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01TKWswqMZ6Z9UAQDAda83Xz --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
9c25bc66f4 |
docs(cli): advertise open --foreground and snapshot --actions in the workflow card (#1682)
* docs(cli): advertise open --foreground and snapshot --actions in the workflow card open --foreground (#1670/#1671) and snapshot -i --actions (#1665) shipped with no mention in the compact `help workflow` card, so a planning model never discovers either. Add one terse line each: the foreground fast-path in Bootstrap, and the merged-element custom-action guidance in Validation and evidence. Stays under the 9,000-byte compact-card budget (8493 -> 8908 bytes). Adds two help-conformance bench cases per the repo's changed-guidance rule: foreground-attach-single-sim (correct plan starts with `open --foreground` in an unambiguous single-sim scenario, fail-closed alternative forbidden) and merged-card-actions-not-directly-invokable (a merged Bluesky-style feed card's actions list is evidence, not a selector). Both use a real pinned sample rebuilt through the production snapshot renderer. * fix(scripts): accept flag order in the foreground-attach conformance matcher Flag order after `open` isn't semantically meaningful (`open --platform ios --foreground` is exactly as correct as `open --foreground --platform ios`), but startsWithForegroundOpen required --foreground to be the literal next token after `open`. Rescoring the completed repeat=3 bench report shows this docked codex:gpt-5.4-mini on all 3 trials even though its plan was config-order noise, not a real deviation -- the no-positional/no-device guarantee already comes from the forbidden checks. Loosened to require --foreground anywhere on the open line; foreground-attach-single-sim now scores 54/54 across both runners. * fix: close workflow help conformance gaps |
||
|
|
ac9e4d0f04 |
test: measure oracle liveness suite-wide; pin the one dead-path oracle (#1679)
* test: measure oracle liveness suite-wide; pin the one dead-path oracle - docs/agents/oracle-negation-spike.md: assertion-negation sweep over 677 test files (5,687 verdicts). Zero vacuous tests: all 150 negation survivors decompose into assert.rejects-validator artifacts (112), helper-oracles (31), in-file-fake breakage (6), and one conditional oracle. Records the companion mock-coupled coverage-uniqueness numbers and the follow-ups they motivate (diff-scoped mutation gate, provider seam closures, transcript provenance). - watchos-sentinel: the non-watchOS test's only assertion sat in a catch block that never fires (tvOS interactor creation succeeds), so no assertion executed on the observed path. Pin creation success instead. Red-run proof: the old shape survived the negation sweep; the new shape fails under it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf * test(android): inject fake adb through the provider scope, not PATH stubs Adds withFakeAdb to test-utils: a scripted in-process AndroidAdbProvider installed through the production withAndroidAdbProvider seam — the same scope the daemon installs per request — replacing PATH-stub shell scripts that spawn a real subprocess per adb call. No PATH mutation, no spawns, no real subprocess waits. Converts settings.test.ts (15 tests, 23ms; waiver said "waits real settings-apply poll time") and notifications.test.ts (2 tests, 9ms). Assertions move from args-log regex greps to structural checks on the recorded call list; the fake receives device-scoped args with the -s serial pair stripped, so serial routing is enforced by the scoped provider matching device.id instead of asserted per call. Remaining PATH-stub files convert next; their contention-retry waiver entries lift together with the conversions. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf * test(android): convert device-input-state to fake adb provider injection 10 PATH-stub cases move to withFakeAdb through the production provider scope; the 2 tests that already inject an executor directly are unchanged. Cross-invocation shell STATE_FILE state becomes a closure boolean; args-log regex asserts become structural checks on recorded calls. 12/12 green at 386ms — the residue is dismissAndroidKeyboard's two fixed 120ms retry sleeps, not stub subprocess waits. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf * test(android): convert app-lifecycle-install adb stubbing to fake provider The adb half of every case moves to withFakeAdb through the production provider scope; installs take the documented exec-shaped fallback (exec(['install','-r',...])), matching what the PATH stub saw minus the serial pair. bundletool/zip/unzip stay real or PATH-stubbed — they run via runCmd outside the adb seam, so this file remains in the serialized subprocess-stub lane with its waiver reason to be corrected from adb to bundletool. 13/13 green at ~130ms; no case enters a retry/poll loop. Conversion note: manifest identity's `unzip -p` failure is silently swallowed (readZipEntry catch -> undefined, aapt fallback) — an invisible degradation path worth a future explicit diagnostic. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf * test(android): convert input-actions adb stubbing to fake provider 9 PATH-stub cases move to withFakeAdb; the 3 tests already injecting providers directly are unchanged. Chunked shell-input assertions become ordered deepEqual on the recorded calls; never-called negatives and call-count checks preserved 1:1. 12/12 green. File time drops to 2.2s, all of it production sleeps: verifyAndroidFilledText unconditionally waits its [0,150,350]ms verification cadence even when the first inspection matches, so each fill verification pass costs ~500ms with an instant fake. A budget-derived cadence there (testing.md pattern 1) would put this file near 25ms; flagged as follow-up rather than changed here. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf * test(android): extract shared oracles; make fake adb failure-faithful Three test-utils extractions applied across the six converted files: - assertRejectsAppError collapses the hand-rolled AppError code+message rejection validator (10 sites here; ~30 more repo-wide can adopt it incrementally). Validators asserting details or multiple differently- flagged regexes stay explicit on purpose. - withFakeAdb gains a `provider` option for extra capabilities (snapshotHelperArtifact, reverse, ...), replacing input-actions' nested re-scoping bridge. - withFakeAdb now mirrors the local executor's contract: a scripted nonzero exit throws androidAdbResultError unless the call site passed allowFailure. Provider-scoped exec bypasses exec.ts's throw-on-close- failure, so returning {exitCode:1} took a different production path than the PATH-stub `exit 1` these fakes replaced. All 75 tests hold under the corrected semantics. Also swaps settings' inline emulator DeviceInfo literals for the shared ANDROID_EMULATOR fixture. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf * test: lift five converted Android files from the contention-retry waiver settings, notifications, device-input-state, input-actions, and app-lifecycle-open no longer stub binaries on PATH or spawn subprocesses, so their contention mechanism is gone: they leave CONTENTION_RETRY_FILES and, through the derived SUBPROCESS_STUB_TESTS constant, the serialized subprocess-stub project (17 -> 12 files). app-lifecycle-install stays with its reason corrected: adb is now in-process, but bundletool stays PATH-stubbed and zip/unzip spawn for .aab packaging paths. Full unit suite green at the new membership: 638 files, 5,724 tests, with the five files running at unit-core's default parallelism. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf * test: apply review findings to the fake-adb conversion batch - app-lifecycle-open: the missing-package launch failure returns {stderr, exitCode: 1} and lets withFakeAdb's throw path produce the production-shaped androidAdbResultError instead of hand-modeling the thrown AppError — the drift the helper exists to eliminate. - withScriptedAdb deleted: the six converted files were its only callers, and a live PATH-stub export invites new tests back into the serialized lane this batch shrank. withMockedAdb stays (dispatch and runtime-hints tests still stub other binaries). - android-snapshot-helper gains androidSnapshotHelperScriptResponse so the version-probe detection and versionCode reply have one source of truth; input-actions' local copy delegates to it. - withFakeAdb's provider option becomes a distributed Omit over the AndroidAdbProvider union, so touch without gestureViewport is a compile error at the fake's boundary (planted and verified) instead of a TypeError inside production gesture planning. - spike-doc re-run checklist restores wider than the codemod globs. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf * test: drop the consumer-less FakeAdbScript barrel re-export Fallow's dead-code gate flagged it: scripts are always passed as inline lambdas, so only FakeAdbResponse needs a name at the barrel. The type stays exported from fake-adb.ts where the withFakeAdb signature uses it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf * test(apple): inject fake xcrun through the tool-provider scope, not PATH stubs withFakeAppleTool mirrors withFakeAdb for the Apple seam: a scripted provider installed via the production withAppleToolProvider scope, flat simctl/devicectl invocations recorded exactly as the PATH-stub shell scripts saw them, throw-on-nonzero fidelity matching exec.ts unless the call site passed allowFailure, and the canned `simctl privacy help` listing served by default (the block withMockedXcrun injected into every script). screenshot-status-bar.test.ts converts as the exemplar: 3/3 green at 9ms with deepEqual call-sequence assertions replacing the args-log regexes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf * test(apple): convert apps.test.ts xcrun stubbing to fake tool provider All withMockedXcrun scripts and hand-rolled PATH stubs move to withFakeAppleTool; args-log regexes become structural call assertions (exact deepEqual where order is deterministic, presence checks where the 5s simulatorBootedMemo TTL makes boot-probe order test-dependent). 12 hand-rolled AppError validators collapse into assertRejectsAppError. 54/54 green; file test time 1172ms -> ~400ms with no test over 201ms. Five .ipa install tests keep a minimal PATH stub for unzip only: install-artifact.ts:112 and install-source.ts:438 call runCmd('unzip') directly, outside the Apple tool provider seam — the file therefore stays in the serialized subprocess-stub lane with its waiver reason corrected from xcrun to unzip. Also observed: getSimctlPrivacyServices caches per PATH+simulatorSetPath and simulatorBootedMemo keys on deviceId|setPath, so neither cache accounts for the provider scope — worked around per test, follow-up worthy. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf * test: lift six Apple waivers; fix format and fallow findings from CI - interactions, simulator, screenshot, physical-device-screenshot, devicectl, and screenshot-status-bar leave CONTENTION_RETRY_FILES: the first five stopped stubbing PATH binaries in earlier refactors (measured 3-64ms per file, no subprocess activity), and screenshot-status-bar now injects through the fake tool provider. apps.test.ts stays with its reason corrected to the unzip PATH stub (xcrun is in-process; install-artifact.ts:112 / install-source.ts:438 call runCmd('unzip') outside the Apple seam). Serialized lane 12 -> 6. - oxfmt: fake-apple-tool.ts and contention-retry.ts were pushed unformatted (local check piped through tail masked the failure). - fallow complexity: the three fake-script arrows in apps.test.ts drop under threshold via shared predicates (isSimctlMainScreenScale, isSimctlScreenshot, isDevicectlDevice), which also deduplicate the screenshot pair. Full unit suite green at the new membership: 638 files, 5,724 tests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
6c0fcb64a1 |
fix: reject distinct ambiguous mutation targets (#1667)
* fix: reject distinct ambiguous mutation targets * fix(ios): scope the raw-match rejection to mutating dispatches `RunnerTests+Interaction.findElement` applied the new fail-closed classification to `querySelector` as well as press/type, because the read call site takes the default `allowNonHittableFallback: false`. With one visible/hittable match and one non-hittable same-selector duplicate the query started returning AMBIGUOUS_MATCH where it previously selected the hittable element, and `queryDirectIosSelectorOrFallback` preserves that error for read callers — so `get`, `is`, and `wait` surfaced an error instead of their prior answer. `classifyDirectSelectorCandidates` now takes a `rawMatchPolicy`. Mutations keep `.rejectDistinctMatches` (the default, so no mutation call site changes); `queryElement` passes `.preferHittableMatch`, restoring the prior read rule: prefer the single hittable match, ambiguous only when hittable matches compete, and never adopt the Maestro coordinate fallback. The Maestro expected-point path is untouched. Covers the one-hittable + one-non-hittable read, competing hittable reads, and the non-hittable-only read. ADR 0011's amendment now states the scope. * test(ios): execute selector read ambiguity regression --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
919700fd42 |
perf(cli): keep node:http and ICU off the CLI startup path (-54% warm call) (#1681)
* perf(cli): keep node:http and ICU off the CLI startup path (-54% warm call) A warm `appstate` cost 146ms; 79ms of that was startup work no local command needs. `node:http` was value-imported by four modules inside the eager import closure of `src/cli.ts`, so every invocation initialized undici -- and, under `NODE_USE_SYSTEM_CA=1`, the macOS trust store behind it. That is ~70ms of mostly off-CPU time (a blocking trustd lookup, invisible to --cpu-prof), paid by commands that never speak HTTP: the default daemon transport is a plain `node:net` socket. Cheap daemon-only commands paid it in full; only commands slow enough to absorb it, like `snapshot`, did not. Separately, `option-schema.ts` sorted its specs with `localeCompare` at module scope. The first `localeCompare` in a process costs ~5.4ms of ICU collator initialization (every later one is ~0.002ms), so the CLI paid ICU init purely to alphabetize an internal list. A case-insensitive comparator reproduces the exact same order for the ASCII flag-key vocabulary -- verified against all 159 keys. Both HTTP modules are now loaded on demand, via `.default` so they stay the same mutable module object a static import yields (the transport tests stub `http.request` on it). Behavior is unchanged: the HTTP transport, remote daemons and artifact upload/download all still work. Interleaved A/B, alternating prebuilt dists in one loop, 5-min load avg < 10, n=30 per arm: appstate 146.4ms -> 67.6ms -53.8% session list 137.9ms -> 57.7ms -58.2% snapshot -i 195.5ms -> 188.8ms -3.4% (device-work bound) Without NODE_USE_SYSTEM_CA the appstate win is 73.9ms -> 57.9ms (-22%), so this is not purely an artifact of that setting. Measured and rejected: splitting command metadata from runtime handlers (the facet restructure deferred by #1660). registry.js is 142kB of the 643kB eager closure but only ~1.6ms of module evaluation, and ALL module compilation is ~9.5ms -- the whole refactor could not clear the 20ms bar set for attempting it. A unit test walks the eager closure of src/cli.ts and fails if any file in it value-imports node:http/node:https again; type-only imports stay allowed. * test(cli): classify startup-closure imports with the ES module record The regex-based guard only recognized `import ... from` / `export ... from` forms, so a bare side-effect `import 'node:http'` anywhere in the eager closure would reintroduce the full undici + system-CA initialization cost while the guard stayed green -- the one form that binds nothing yet still evaluates the module. Classification now comes from oxc-parser's ES module record, the parser this repo already uses for import analysis in scripts/lib/shipped-imports. A side-effect import is exactly an entry-less static import, so it falls out of the record rather than needing another regex branch; type-only imports and re-exports are identified by their own `isType` flag instead of being pattern-matched. A fixture table now pins all eleven import forms -- side-effect, default, namespace, named, value re-export, star re-export, mixed value+type, type-only default/named/re-export, and dynamic -- so the classifier's contract is asserted directly rather than only through the closure walk. Verified red end to end: a side-effect import inserted into a real closure module fails the guard and names the file. * test(cli): follow workspace subpath imports in the startup-closure guard The walk stopped at the package boundary, so a `node:http` import inside any @agent-device/* module the CLI evaluates would have reintroduced the cost invisibly -- the same blind spot the module-record rewrite closed for side-effect imports, just relocated. Workspace specifiers now resolve through each package's exports map. The anti-vacuity assertion checks the closure actually contains package files, so a resolver that silently returned null for every workspace specifier would fail rather than leave the guard green over a src-only closure. Verified red both ways: a side-effect import inserted into packages/kernel/src/errors.ts, and one in a src/ module, each fail the guard and name the offending file. * test(cli): drop a redundant cast in the workspace exports lookup * test(cli): classify a top-level dynamic import as eager, and share one HTTP loader "On demand" is a claim about SCOPE, not syntax, and the guard only knew syntax. `import('node:http')` at MODULE TOP LEVEL runs during module evaluation -- so a top-level `await import(...)`, an `import(...).then(...)`, or an immediately-invoked top-level function reintroduced the whole undici plus trust-store cost while the guard stayed green. A dynamic import is lazy only when its nearest enclosing function scope is not the module itself. The guard now walks each closure file's OXC AST carrying an `eager` flag that drops on entry to a FunctionDeclaration/FunctionExpression/ArrowFunction body, and reports `import()` calls still reached with it set. Immediately-invoked top-level functions are detected by unwrapping the callee's ParenthesizedExpression, so `(async () => { ... })()` is eager too. The one shape that still reads as lazy -- a top-level function invoked indirectly at load time -- is stated as a known limitation in the test rather than left implied. Verified red by inserting each shape into a REACHABLE closure module (daemon-client-transport.ts): top-level await, `.then(...)`, and the top-level IIFE each fail the guard and name the file. The lazy case is proven by the suite itself rather than a fixture alone -- the shared loader below lives in the closure, so its function-local `await import('node:http')` would turn the guard red if scope tracking were wrong. The three lazy-load sites had drifted into three copies of the same ternary. They now share `loadNodeHttpRequester` in src/utils/node-http.ts, next to the other node:http helpers, which is where the reason for the indirection is documented once. This also clears the Fallow complexity finding the previous push introduced. Fixture table extended to 19 cases covering both axes: static form (side-effect, default, namespace, named, value/star re-export, mixed, type-only x3) and dynamic scope (top-level x3, function/arrow/method-local, and the exact ternary the shipped loader uses). |
||
|
|
e6b4fa2810 |
fix: isolate concurrent remote connections (#1675)
* fix: isolate concurrent remote connections * refactor: harden remote connection state * fix(cli): scope every emitted connect command to its own session `scopeNextSteps` only reached `ConnectReadiness.nextSteps`, so two command-bearing outputs still shipped unscoped: - `providerArtifactNotes()` emitted `agent-device artifacts --json` as a prose note, which never passes through that helper. - `buildDeferredRuntimeNotice()` emitted `agent-device metro prepare --remote-config <path>` independently in connection.ts. On the shared-host concurrency path this branch fixes, following either one resolves against the host-global active connection, so the artifacts instruction can return another job's provider video and log URLs (#1659). Both producers now take the connection state and format through one exported `scopeCommand` helper, which is the single place a suggested command is bound to its originating session. The metro config path is shell-quoted alongside the session name. Coverage: the BrowserStack route asserts human and JSON shapes carry one `--session` per suggested command and that each emitted session resolves back through `readRemoteConnectionState` to the connection that printed it; `connection status` pins the scoped deferred-metro `nextStep`. --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
4b89482a06 |
fix(daemon): open --foreground P1 hotfix — selector rejection, interactive snapshot, capture-failure masking (#1671)
* fix(daemon): open --foreground P1 hotfix — selector rejection, interactive snapshot, capture-failure masking Three P1s from the post-merge review of #1670, all at the session-open-foreground dispatch seam: 1. Explicit device selectors were silently overwritten: the resolved-device rewrite pinned --udid/--platform over whatever the caller passed, so `open --foreground --udid B` with sim A sole-booted silently opened A. Now fails fast with INVALID_ARGS (matching the existing app-positional rejection) on --udid/--device, and on --platform other than ios; an explicit --platform ios passes through. 2. The promised interactive snapshot was never requested: the composed snapshot dispatch forwarded the open request's flags untouched, without snapshotInteractiveOnly — so the capture was NOT the `snapshot -i` path the doc comment promised and returned no interactive presentation. The composed request now sets snapshotInteractiveOnly: true (the exact key the CLI maps -i to and the snapshot runtime reads as interactiveOnly). 3. A capture failure masked the successful open: returning the snapshot error discarded openResponse even though the session exists, so a retry of `open --foreground` failed with "close the current session first". Open success + snapshot failure now returns ok with an explicit initialSnapshotError {code, message} detail and a rendered warning that the session IS open and how to capture manually (snapshot -i). Regressions added for all three: explicit-selector rejection (udid/device/both/non-iOS platform + ios pass-through), the composed dispatch carrying snapshotInteractiveOnly, and the snapshot-failure path returning ok + warning + usable session. * fix(cli): render the composed open --foreground snapshot on default stdout and project initialSnapshotError through the public surfaces Post-merge review on #1671 found the daemon fixes never reached the public boundaries: openCliOutput ignored the nested snapshot (the one-call promise held only under --json), and initialSnapshotError was daemon-only — absent from AppOpenResult, Node normalization, and serializeOpenResult, with the normalized shape truncated to code+message. - default open output now renders the composed interactive tree through the same snapshotCliOutput path snapshot -i uses - AppOpenResult carries initialSnapshotError as the FULL daemon error (hint/details/diagnosticId/logPath preserved) through normalization and serialization Worker-authored; committed by the coordinating session after the worker stalled twice mid-push. Tests: output.test.ts + session-open-foreground (26 pass), typecheck, oxfmt. * refactor(client): one daemon-error normalizer + client-route regressions for initialSnapshotError Review follow-ups on #1671: normalizeInitialSnapshotError duplicated the target-shutdown error normalization and pushed the module over the fallow complexity threshold; consolidated into a single internal normalizeDaemonError (table-driven, full shape incl. retriable/supportedOn) used by both result paths, projected through normalizeOpenForegroundComposition. Client-route regressions: createAgentDeviceClient().apps.open now proves the full initialSnapshotError shape (hint/details/diagnosticId/logPath/retriable) survives normalization, and that a malformed one is dropped — deleting the boundary normalization fails these tests. Also rebased onto current main. * fix(daemon): a thrown initial-snapshot capture failure gets the same successful-open contract Review P1 on #1671: dispatchSnapshotViaRuntime rethrows ordinary capture/runner exceptions; the composition only handled a returned { ok: false }, so a thrown failure escaped to the router and failed the whole open after the session was created — retrying then wedged on the existing session. The catch normalizes the rejection (kernel normalizeError, same conversion the router applies) into the shared openWithInitialSnapshotFailure path: ok response, full-shape initialSnapshotError, session-usable warning. Rejecting-mock regression added alongside the returned-failure case. |
||
|
|
13bc70f24f |
refactor(find): resolve a mutating find's target once, not twice (#1654) (#1669)
* fix(find): resolve a mutating find's target once, not twice (#1654) A mutating `find click`/`find fill` captured the screen, matched by locator under the `findAct` policy, promoted to a hittable ancestor, minted `@eN` off the node it chose — and then re-entered the interaction leaf by bare `@ref`, which looked that ref up AGAIN via resolveSnapshotForRef. The second lookup reads the SESSION frame tree, not the fresh capture find matched against, so the two could disagree: it could hand the action a different node than find picked, or refuse as unresolvable a ref find had observed a moment earlier. find now passes the node itself. `internal.findPreresolvedTarget` carries the resolved node and its tree over the in-process invoke hop, and the ref branch adopts them instead of resolving the ref a second time. What is NOT skipped: the guarantees. Occlusion, hittable-ancestor promotion, and the off-screen guard all still run, on that node, at the same symbols the ADR 0011 `runtime-ref` cells name. Only the LOOKUP is replaced. Recording, ref-frame effects, settle/observation, and deferred-outcome marking are untouched — they live in the dispatch wrapper, not in resolution, which is why the invoke hop is kept rather than bypassed the way find focus/type bypass it. The channel is a second field rather than widening `findResolvedTarget` because they answer different questions: that flag governs ref-frame admission and staleness (policy), this governs resolution (lookup), and focus/type set neither. It is `internal`-only, so it never crosses the wire and carrying live node references is sound. ADR 0011 re-check: the `runtime-ref` disclosure cell gains a second producer, adoptPreresolvedRefTarget, reporting `exact`. Truthful rather than borrowed — the ref is minted off the node handed over. It cannot report label-fallback, which is right: label recovery is a property of looking a stale ref up, and this path performs no lookup. Tests pin the behavior by tap coordinates, so they name which tree the leaf resolved against, plus a control proving the ordinary @ref path is unchanged and a case where the session tree cannot resolve the ref at all — that one can only pass if no second lookup runs. Verified revert-sensitive: removing the short-circuit fails 3 of the 4, and the control stays green. Out of scope, unchanged: the ADR 0011 `native-ref` path (web provider clickRef/fillRef only), where the ref is the provider's own element handle. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0117wfrvC6MRDJUWdRBErNEF * test(find): prove the production route, and correct the divergence claim (#1654 review P2) The regression tests built `internal.findPreresolvedTarget` by hand and called handleInteractionCommands directly, so they proved the leaf consumes the channel but never that find ATTACHES it. Either producer could have stopped doing so and they would all have stayed green — the exact gap #1649's review named, reappearing one layer up. Four tests now drive the real handleFindCommands with invoke wired to the real handleInteractionCommands. Two assert click and fill each carry the selected node (separate producers, separate forwarding hops, so asserted separately); two advance the session tree between find's match and the leaf's resolution and assert the dispatched point is still find's node. That last pair also corrects the record. The claim that the leaf resolves against a different tree than find matched is too strong for the common path: `omitRefFrameSnapshot` (interaction-runtime.ts) makes find's internal dispatch skip the authorized frame tree and resolve against `session.snapshot`, which find's own capture just wrote — so the second lookup normally AGREES, and the two non-diverged tests above pass with or without the fix. The divergence is real but narrower: it needs something to advance `session.snapshot` between find's match and the leaf's resolution, which `refreshAndroidRefSnapshotIfFreshnessActive` does on the ref path. Measured at this head, the old code taps (60,720) "Delete" where find matched "Save" at (310,510) — a wrong-element mutation, now pinned by both click and fill. So the change's value is what #1654 asked for — one resolution end to end — plus closing that window, not a fix for a divergence occurring on every find. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0117wfrvC6MRDJUWdRBErNEF * fix(find): tighten resolved target provenance --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
04c33f9e1e |
feat(ios): expose AX custom actions on merged accessibility elements (#1665)
* feat(ios): expose AX custom actions on merged accessibility elements
Apps that merge a card into one accessibility element for VoiceOver (React
Native's `accessible` prop) publish the card's real affordances as
UIAccessibilityCustomActions rather than as child elements. Our snapshot showed
only the merged node, so an agent looking for a feed card's options control had
nothing to aim at and fell back to coordinate guessing.
`snapshot --actions` now names them:
@e8 [link] "feedItem-by-whiskers.test" actions: ["Reply", "Repost", "Open post options menu"]
Opt-in, because the AX server cannot serve custom actions in a bulk tree
request: adding the attribute makes testmanagerd's reply decoder reject the
nested arrays a custom action serializes into, drop the reply, and time the
request out (~65s vs ~110ms). Only per-element reads answer, at ~100ms each, so
the runner reads at most 12 labelled childless nodes and stops at the
capture-plan deadline. The request pins the private-AX backend, since no other
backend can read the attribute, and reports that as its own `requested-backend`
verdict so a deliberate pin never renders as a degradation warning.
Invocation is not shipped: the actions are readable but not invocable from the
runner. RunnerAXSnapshotBridge.h records the five APIs that were tried.
* fix(ios): disclose a capped custom-action pass, and read on-screen elements first
Two gaps in the first cut.
An element the bounded pass never reached rendered identically to one with no
custom actions, so a capped capture silently taught the reader that later feed
cards have no affordances — the exact mis-inference this feature exists to
prevent. The runner already counted reads against candidates; it now carries
both to the response, and the verdict renders one response-level line when the
pass was incomplete. A complete pass stays silent, and "never asked" stays
distinguishable from "read none" (the key is absent, not (0, 0)).
The obvious remedy for a capped pass would be a scoped re-run, but scope is
applied when the Swift walk builds nodes, long after the read pass, so it does
not redirect the budget at all — measured: reads=12 candidates=18 with and
without --scope. Rather than print a remedy that does nothing, the read pass now
orders candidates on-screen first. That makes the budget land on elements an
agent can act on, and makes the disclosed remedy true: scrolling changes the
on-screen set, so a re-run reads elements the previous pass could not.
Also states plainly, in the flag help and the tool/SDK field description, that
the names are for planning: nothing invokes them, so the affordance is reached
through the element's detail screen, the same control exposed elsewhere, or
coordinates.
* test(snapshot): pin the custom-action coverage pair in the verdict shape assertion
* chore(scripts): classify the --actions flag in the integration progress model
The completeness gate flagged snapshotCustomActions as unclassified, which is
what it is for. It gets its own bucket rather than joining the provider-scenario
table: the values come from the private AX client inside the runner process, and
the fake runner derives its behavior from fixture tables that cannot fabricate
custom actions, so there is no provider-backed scenario to claim. The owning
coverage is named instead — runner XCTest unit, snapshot-lines, snapshot-quality.
* fix(ios): fail closed on unserviceable --actions, bound each read, cap output, and count actions in identity
Four review findings.
1. `--actions` with `--raw`, or on any target that is not an iOS simulator, used
to succeed and return nodes with no actions — a requested capability silently
no-opped, indistinguishable from "this screen has none". Both now fail closed.
The raw pairing is rejected at the shared request seam (INVALID_ARGS) so CLI,
Node client and MCP answer alike before any device work; the platform case is
rejected once the session device is resolved (UNSUPPORTED_OPERATION), naming the
resolved target. `diff --actions` was already rejected as an unsupported flag.
2. The per-element AX read had no timeout, so one wedged element could consume
the whole capture budget. Each read now runs off-thread behind a 1s wait. A
timed-out element counts as unread, never as "read, and it has no actions", so
the existing partial-pass disclosure already covers it.
3. The element budget bounded element count only; one element could still return
an unbounded list of unbounded names. Capped at 8 names of 80 characters, and
clipped elements are counted into the coverage so a truncated list is disclosed
rather than silently presented as complete.
4. Action names were rendered unescaped, and no comparison key read them. Names
now get the same escaping as text previews plus control-character folding, so an
app-authored name cannot split or corrupt a line. `actions` joins the diff
comparable key, the unchanged-comparison projection, and — the sharper bug —
the snapshot presentation key, without which `snapshot` followed by
`snapshot --actions` on a still screen answered "unchanged" and never delivered
the actions that were explicitly requested.
* fix(ios): contain a hung custom-action read instead of accumulating orphans
The 1s read deadline frees the caller, but the underlying AX call is a
synchronous XPC round trip that cannot be cancelled — it keeps running. On a
global concurrent queue that meant repeated `snapshot --actions` against a
wedged element piled up orphaned reads, all using the shared XCAXClient
concurrently. The deadline was containment for the capture, not for the runner.
Since the call cannot be cancelled, contain it instead:
- every read runs on one dedicated serial queue, so a wedged call can never be
joined by a second concurrent user of the shared client;
- a single-flight guard refuses to dispatch at all while an abandoned read is
still outstanding, so a repeated capture adds no work — the dispatch counter
stands still;
- the read pass stops at that point rather than paying a deadline per element
on reads that would all be refused, and reports `blocked` so the capture stays
honest. That gets its own line, because the partial-pass remedy (scroll and
re-run) cannot clear a hang and would send the reader in circles.
Recovery needs no reset: when the hung call finally returns, in-flight drops to
zero and reads resume.
The regression drives a fake AX client that never returns, and asserts the three
things the fix exists for — exactly one in-flight read with no further
dispatches across repeated captures, immediate returns with the skip disclosed
instead of the scroll remedy, and reads working again once the wedge clears.
* ci(ios): execute the custom-action runner regressions instead of only compiling them
The iOS workflow runs a targeted -only-testing list, so a runner test that is
not named there is compiled by the build step and then never executed. All seven
custom-action tests were in that gap — including the containment regression,
which is the only executable proof that a hung AX read cannot accumulate
orphaned in-flight reads.
Red/green against the containment regression, with the fix reverted to its
pre-fix concurrency behavior (global concurrent queue, no single-flight guard,
no blocked exit):
RED in-flight 6 (want 1), dispatches 6 (want 1), each repeat paid the full
1.004s deadline (want <0.2s), blocked=false (want true), and the
in-flight drain never completed — "Exceeded timeout of 5 seconds".
GREEN 7/7 pass, containment regression in 1.02s.
|
||
|
|
5f90d908e2 |
fix(ios): wait for hidden-keyboard synthesized text to commit before responding (#1676)
* fix(test): wait for typed text to settle in the hidden-keyboard runner test
testBareTypeUsesTappedInputWhenSoftwareKeyboardIsHidden read textField.value
in one shot right after executeTypeCommand returned. The simulator commits
synthesized keystrokes after the command responds, so on a loaded CI machine
the read landed mid-word — observed failures reported ("h") and
("hardware-ke"). It failed on 3 of 5 runs of a branch carrying zero Swift
changes and passed on re-run.
Poll the value until it holds the expected text (10s ceiling) and assert on
the last value read, so a real regression still fails with what the field
actually held. Both reads in the test use the same helper; the assertions are
unchanged.
* Revert "fix(test): wait for typed text to settle in the hidden-keyboard runner test"
This reverts commit
|
||
|
|
f5b6e8f28c |
fix: enforce replay timeout envelope (#1674)
* fix: enforce replay timeout envelope * fix: preserve replay timeout metadata |
||
|
|
7b4e461220 |
fix(test): clean Swift toolchain temporary directories (#1664)
* fix(test): clean Swift toolchain temporary directories * fix(test): wait for Swift toolchain shutdown * fix(test): terminate Swift toolchain process groups |
||
|
|
e14c9d8d7b |
fix(cloud): unblock iOS snapshot and gate cloud fill on text-entry focus (#1658) (#1666)
* fix(cloud): unblock iOS snapshot and gate cloud fill on text-entry focus (#1658)
Two bugs isolated to the cloud-webdriver iOS path.
`snapshot`/`diff` refused every capture on a live BrowserStack session with
SESSION_NOT_FOUND, instantly and without a driver round trip. The app-session
guard they ran belongs to the local XCUITest runner, which must attach to a
target app; a cloud capture reads the provider's own driver session and needs
no app identity, so it now applies to local Apple targets only. The session
was empty in the first place because the provider open path skips local app
resolution wholesale — no simctl/devicectl reaches a hosted device — and
dropped an explicitly spelled bundle id along with it. A dotted, non-deep-link
target is the bundle id under the same convention resolveIosApp applies
locally, so a cloud `open com.example.app` now records it.
`fill` tapped and sent its keys in back-to-back requests. A WebView input —
an OAuth page in a Safari view controller — takes first responder
asynchronously, so the keys landed with nothing focused while the command
still answered "Filled N chars"; tapping and filling as two separate commands
worked only because the round trip between them gave the field time to focus.
The cloud interactor now waits on the same signal the Apple runner uses, the
software keyboard going from hidden to shown after its tap, and discloses what
it observed as `textEntryReadiness` so a fill with no witness cannot pass for
a filled field. Where keyboard visibility cannot witness the focus move —
back-to-back fills into one form, the shape that failed most often — it spends
the runner's full readiness budget rather than racing the app with a short
settle.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Zjbzf7HdziX9SzFpNP7WW
* fix(cloud): let a new bundle-id open replace the tracked cloud iOS app
Adopting an explicitly spelled bundle id on a provider-backed open (the
fix that makes snapshot/diff work at all) also made a previously dead
precedence rule live: the provider branch returned currentAppBundleId
first, so once a first open had populated it, `open com.a` followed by
`open com.b` left the session still reporting com.a to every
appBundleId-gated command.
The local path does the opposite, and is the convention this branch is
meant to mirror: resolveIosApp returns a dotted target unchanged and
never consults the session's current app. Only its deep-link branches
prefer the tracked id. Flip the provider branch to match — an explicit
bundle-id target wins, and everything the branch cannot name (deep
links, display names, bare open) still falls back to the tracked id.
* fix(cloud): fail a witness-less cloud fill instead of reporting it filled
Review follow-ups on #1658.
`not-observed` was still a success: it sent the keys and answered "Filled N
chars", and nothing renders `textEntryReadiness` in default CLI output — so the
exact silent success this branch exists to remove survived whenever focus never
happened. A tap that raises no keyboard now fails with
`text_entry_focus_not_observed` and sends no keys, leaving the field untouched
rather than half-written, and the readiness vocabulary keeps only outcomes that
describe a fill that did type.
The readiness budget was advertised but not enforced at the request boundary:
each keyboard probe inherited the client's 30s default, so one hung probe could
hold a 2s wait for far longer. Probes now carry their own bound, threaded
through the client as a per-request timeout override.
The probe also swallowed every error as "this driver cannot answer", which
degraded a dead session, an auth rejection, or a grid outage into a blind text
entry. Only a positively classified unimplemented route counts as unsupported
now — classified on the W3C error code rather than the status, since `unknown
command` and `invalid session id` share HTTP 404 — and everything else
propagates.
The provider scenario proved request ordering against a stub that always
accepted keys. Its fake now models the device: focus lands a beat after the
tap, and keys arriving while the keyboard is down are accepted and dropped,
exactly as an unfocused field does. The tests assert the field's own value, and
both go red against the pre-fix `fill`.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015Zjbzf7HdziX9SzFpNP7WW
* fix(cloud): witness the tapped field's focus before a cloud fill types
Review of cc23f2b found two ways a fill could still report success
without evidence that OUR tap focused the field it was aimed at.
P1. `settled-keyboard-up` and `settled-unknown` both typed and returned
normal success. Keyboard visibility can only witness that *a* field took
focus, never *which*: filling a second field in an already-open form
reads the same before and after, so a missed tap left the first field
focused and `POST /keys` — which the driver routes to whatever holds
first responder — appended to it while every request returned 200.
Failing those closed outright would have broken ordinary multi-field
form fills, which do work: a live AWS Device Farm run types both fields
of a WebView login correctly. So witness focus properly instead. W3C
`GET /element/active` answers the question keyboard visibility cannot —
is the thing focused now the thing I tapped — and answers it whether or
not the keyboard was already up. That becomes the primary signal
(`focused-element`); the keyboard transition stays as the fallback for
drivers without the route, and a keyboard already up on such a driver
now refuses rather than typing.
The test is identity, not geometry. Containment of the tap point looks
like the obvious rule and is wrong: focusing a field can re-lay it out.
On a live iPhone 16, tapping Safari's collapsed address bar expands it
into a taller field that no longer covers the tapped point, and a
containment-only rule refused a fill that plainly worked. So a tap that
MOVES focus counts, with containment as the second half of the test —
re-filling the already-focused field moves nothing, and only geometry
tells that from a tap that missed. Both readings are taken before the
tap, since each is evidence only as a change.
P2. The 2s budget bounded the loop but not the calls inside it: every
probe got a fixed 1500ms, so one begun near the deadline finished well
past it. Both the probe timeout and the sleep are now capped by the
remaining budget.
Also fixes a related escape the review did not name: the poll loop had
no catch, so one transient grid error aborted a fill the next poll would
have satisfied. Probe failures are now tolerated within the budget, but
a budget that expires without a single answered probe rethrows, so a
dead session surfaces as itself rather than as "the tap missed".
The provider scenario gains the two-field case the review asked for: it
begins keyboard-up with the email field focused, misses the password
tap, and asserts no keys reach the email field.
Verified on AWS Device Farm iPhone 16 / iOS 18.0 at this exact tree:
address bar (the re-layout case) and both WebView login fields all
report `focused-element`, the second with the keyboard already up, and
the device reads back `tomsmith` and a 20-character password.
* fix(cloud): refuse a cloud fill no focus probe can witness, and bound the composite probe
Two blockers from the review of
|
||
|
|
9b4239b2f3 |
fix: heal dangling XCTestDevices redirect and prove lock-owner death before reclaim (#1672)
* fix(ios): heal dangling XCTestDevices symlink during device-set reconcile fs.existsSync follows symlinks, so a redirect symlink whose target set was deleted read as absent: reconcile skipped the unlink and the backup rename failed with ENOTDIR, wedging every XCTest command on the machine until manual cleanup. Detect symlinks with lstat (throwIfNoEntry) in both the reconcile check and unlinkIfSymlink so the stale link is removed and the backup restored. * fix: prove owner death before reclaiming locks, and detect zombies A zombified lock owner passes kill(pid, 0) and still reports its original lstart, so a killed-but-unreaped daemon held the XCTest device-set lock forever. Conversely a ps read lost to CPU contention condemned a live owner and let waiters steal a held lock (the flake pinOwnProcessStartTime papers over in tests). Unify both liveness surfaces on classifyOwnerLiveness: condemn zombies via ps state, treat failed ps reads as no-proof rather than death, and reclaim null-start-time owners whose pid provably started after the lock was acquired (ps etime bound). Lock timeout errors now report ownerLiveness. * fix: keep null-start-time lock owners fail-closed Review P1 on #1672: the started-after-acquisition bound compared a wall-clock birth estimate (Date.now() - ps etime) against the persisted wall-clock acquiredAtMs, so a clock step larger than the fixed 30s slack could condemn a live null-start owner and let a waiter steal the held lock — the failure class this PR eliminates. There is no same-clock-domain proof of birth order to be had here (macOS ps etime is itself wall-derived), so drop the bound: an alive pid with no recorded start time is never reclaimed. Regression covers the steal at both the classifier and the lock level. |
||
|
|
a158434a9c |
feat(cli): compact workflow help card + version header (#1663)
* feat(cli): compact workflow help card + version header Shrinks the per-task agent protocol tax of the help/skill surface. `agent-device help workflow` drops from 41025 to 8466 bytes (-79%) by moving depth into new `help scripting` (save-script, secret-safe fills, batch JSON, replay divergence/repair) and `help gestures` (multi-touch shapes/quirks) topics, and folding a few paragraphs into topics that already owned the subject (help debugging, help physical-device, help validate). Content is moved, not deleted. Every `help <topic>` first line is now `agent-device <version> — <topic>`, so the skill router reads the CLI version off the mandatory first help read instead of a separate `agent-device --version` call. SKILL.md is updated to do that and stays a thin router otherwise. The compact card also gains two terse behavioral rules: chain confident consecutive steps with `&&` (falling back to one command at a time when uncertain), and confirm the requested end state is actually visible on screen before declaring a task done. help-conformance-bench (22 cases x 2 runners) improves after the change: 29/44 -> 32/44 passing checks. * fix(cli): review follow-ups on the compact workflow card (#1663) Three fixes from PR review: - Extend the help-conformance plan validator to split a command line on unquoted && and validate each chained segment independently, so a plan that follows the workflow card's "chain confident consecutive steps with &&" guidance is accepted instead of rejected as one shell-projection violation. && inside a quoted selector value (e.g. label="A && B") is not a chain boundary and does not split. Adds unit tests for the splitter and a chains-confident-consecutive- settle-steps conformance case. batch stays out of this: it is deliberately stop-only. - Replace the literal @ref placeholder the compact card used in its own "snapshot -s @ref" example with a concrete ref (snapshot -s @e12 (the current concrete ref)), matching the same card's rule against placeholder targets. Reverts the test to demand the concrete shape. - Give help scripting and help gestures real conformance cases instead of waivers: a secret-safe recorded-fill + publish case, and an Android transform-then-verify case whose exact verification text only appears in the gestures topic. Removes both waivers. help-conformance-bench (25 cases x 2 runners, repeat=1) after these fixes: two full runs landed at 32/50 and 33/50. That is on par with the pre-change baseline (29/44) once the topic-untouched cases' run-to-run swings are accounted for (confirmed noise: one case with zero exposure to any change here flipped 10/10 -> 1/10 on a runner API error, and another swung across all three post-fix runs). The new scripting case now passes 8/8 for both runners; the new chaining case correctly reports the model's choice not to chain as a soft signal, not a validator failure. * fix(cli): update session.test.ts help pointer for moved script-authoring content * fix(cli): reject empty && chain operands in the plan validator (#1663) splitOnUnquotedAnd() previously trimmed and filtered out empty segments, so a plan with a leading (`&& press ...`), trailing (`press ... &&`), or doubled (`a && && b`) operator passed validPlanCommands even though a real shell rejects all three as a syntax error. The validator would bless a plan that fails at execution. Empty segments are now surfaced as an `empty-chain-operand` issue instead of being silently dropped. The quoted-&& non-split behavior (label="A && B") is unchanged, and a normal single command with no chain still parses identically to before. Adds regression tests for all three empty-operand shapes plus the quoted-&& case. |
||
|
|
ac52281448 |
fix(test): deterministic temp-dir cleanup across node --test lanes (#1661)
* fix(test): deterministic temp-dir cleanup across node --test lanes node --test has no global setup/teardown hook, so unlike Vitest (#1593) every node --test package.json script (maestro:conformance, mutation:test, check:affected:test, check:coverage-changed:test, check:layering, depgraph:test, check:tmpdir-leaks:test, check:contention-retry, test:fixture-cache, test:smoke(:web), test:integration:node, test:concurrency-torture) still created scratch directories against the real, unredirected os.tmpdir(), with cleanup only as reliable as each call site's own try/finally — which a crash, OOM, or timeout kill bypasses entirely. Add scripts/node-test-tmpdir.ts: it wraps the whole `node --test` invocation as a child process, redirecting TMPDIR to one disposable, pid-tagged directory (shared root/prefix with the Vitest lane) and removing it from the process 'exit' event, which fires on normal completion, a thrown error, or a forwarded SIGINT/SIGTERM alike. Every node --test script now runs through it. check-tmpdir-leaks.ts already scans by root/prefix, so it covers both mechanisms with no changes to its detection logic. Verified: a node --test process that mkdtemp's then gets SIGKILL'd leaves a directory behind unwrapped; wrapped and SIGTERM'd, TMPDIR is redirected and the directory is gone with no orphaned processes. All 13 wrapped lanes and the full Vitest suite (5,591 tests) pass with zero residual agent-device-test-run-* directories after the run. Fixes #1595 * test(tmpdir): ratchet every node --test script through the wrapper The 13 lanes wrapped in package.json were a one-time hand sweep with nothing enforcing the pattern going forward — a 14th node --test script added later without scripts/node-test-tmpdir.ts would silently reopen #1595 for that one lane. Add a structural check to scripts/node-test-tmpdir.test.ts (now part of check:tmpdir-leaks:test) that reads package.json and fails if any script invokes `node ... --test` without routing through the wrapper. Dumb string matching over the scripts map, no shell parsing, with an explicit (currently empty) NODE_TEST_WRAPPER_BYPASS_ALLOWLIST for any lane that must legitimately bypass it. Verified it both passes on the current package.json and fails when a synthetic unwrapped `node --test` script is added. * fix(test): preserve the Swift cache and close the raw node --test bypasses Review on #1661 found two gaps: 1. The wrapper only overrode TMPDIR, so it discarded and forced a recompile of the durable Swift compiler cache every run instead of mirroring vitest-tmpdir-global-setup.ts's carve-out for it. Read os.tmpdir() before the child's TMPDIR redirect takes effect and set AGENT_DEVICE_SWIFT_CACHE_DIR from that (only when unset), same as the Vitest lane — the two now share one durable cache instead of each discarding and recompiling their own. Added a probe assertion (scripts/node-test-tmpdir.test.ts) that fails without the fix and passes with it (verified both ways). 2. docs/agents/testing.md documented raw `node --test` commands for the iOS smoke files, and the android/ios/conformance-regenerate/nightly workflows invoked `node --test` directly outside package.json. Routed all of them through scripts/node-test-tmpdir.ts so the documented local commands and CI lanes get the same crash/timeout-safe cleanup the package.json scripts already have. |
||
|
|
cc943400a9 |
feat(daemon): [RFC] prototype foreground-attach convenience (#1670)
Prototype `open --foreground`: on a fresh session with no app argument, auto-resolves the target from the sole booted iOS simulator's sole foreground app (reusing the exact same ambiguity-detection probe that enriches the SESSION_NOT_FOUND hint), then attaches the initial interactive snapshot to the response by composing the existing snapshot-runtime dispatch. Collapses the documented 3-call snapshot-fails -> read-hint -> open -> snapshot-succeeds dance into a single call for the unambiguous case, while failing closed (AMBIGUOUS_MATCH) with no guessing otherwise. First-pass RFC, not reviewed — see PR body for the design tradeoff writeup, live before/after evidence, and scoped-out follow-ups. |
||
|
|
4f510ebd18 |
fix(daemon): SESSION_NOT_FOUND hint names the detected foreground app (#1662)
* fix(daemon): SESSION_NOT_FOUND hint names the detected foreground app iOS snapshot/diff without an app session always emitted the generic "Run open first" hint, even when the environment was unambiguous — costing an agent 2-3 extra turns to discover the bundle id itself (27/30 benchmark tasks hit this). When exactly one iOS simulator is booted and exactly one installed app is currently running on it, the hint now names the exact runnable `open` command instead. Any ambiguity (0/2+ booted simulators, 0/2+ running apps, or a probe failure) keeps the existing generic hint — the detection never guesses. - listBootedIosSimulators (devices.ts): cheap, bounded simctl probe for every currently booted simulator, used to detect true device-level ambiguity independent of what device resolution silently picked. - detectSoleRunningIosSimulatorApp (app-resolution.ts): cross-references `launchctl list` running UIKitApplication jobs against `simctl listapps` to find the sole real installed app running, filtering out simulator- internal services without a version-fragile bundle-id denylist. - ios-app-session-hint.ts composes both into the hint text, wired into snapshot-runtime.ts's existing SESSION_NOT_FOUND guard. * fix(daemon): say the detected app is running, not in the foreground launchctl only confirms the app's process is alive, not that it's the frontmost scene — a resident-but-backgrounded app (launched, then sent home) would match the same way. The actionable open command is correct either way, but "in the foreground" overclaimed what the probe actually verified. Say "running" instead. * fix(daemon): bound and catch the hint probe, pin the printed command Two precision fixes from review: - The probe was best-effort in spirit but not in fact: allowFailure only suppresses a non-zero exit code, not a timeout or spawn-failure rejection (exec.ts always rejects on those). listSimulatorApps' listapps call had no timeout at all, so it could hang the error path indefinitely; the launchctl probe had no catch, so a rejection would propagate and replace the deterministic SESSION_NOT_FOUND with a probe failure. Threaded an optional timeoutMs through listSimulatorApps/listSimulatorAppMetadata (including its plutil fallback) and wrapped both detectSoleRunningIosSimulatorApp and buildIosOpenCommandHint in try/catch so any throw falls back to the generic hint, never propagates. - The printed command now always pins --udid (the sole-booted-device check only proves uniqueness within the probed device's own simulator set, not that a fresh CLI invocation without --ios-simulator-device-set would resolve to the same device) and echoes --ios-simulator-device-set when the detected device carries one, so the hint is the exact command that was live-validated, not a shorter one that happens to work by coincidence. Added a length guard: wire-level error details are redacted before send and silently truncate any string over 400 chars, so a hint that would exceed a safe budget (a long device-set path pushes it there) now falls back instead of printing a truncated, non-copy-pasteable command. Also: "in the foreground" overclaimed what the probe verifies (a resident but backgrounded app would match the same running-process check) — reworded to "running". |
||
|
|
35f4cd63d7 |
perf(cli): lazy-load dedicated client command handlers (#1660)
router.ts statically imported all 10 "dedicated" command handlers
(connect, disconnect, connection, auth, daemon, device, proxy, replay,
screenshot, diff) at module top level, so every real CLI invocation paid
for loading all of them -- including replay.ts's @agent-device/maestro
subtree -- even when dispatching to an unrelated command like snapshot
or press. The generic client-backed commands (snapshot, press, tap,
fill, ...) already lazy-load their handler via a dynamic import('./generic.ts');
this brings the 10 dedicated handlers in line with that existing pattern.
ClientCommandHandlerMap (router-types.ts) is repointed to describe the
lazy-loader shape (Record<CliCommandName, () => Promise<ClientCommandHandler>>)
and the loader map is checked against it via `satisfies`, so a typo'd or
stale command key still fails to compile, matching the safety the old
eager map had.
Measured via git-stash-built before/after dist trees, interleaved 1:1 to
cancel drift, run against a live iOS session (own throwaway simulator):
the CLI-side "pre-daemon" phase (node boot + import graph + arg parse,
before the process contacts the daemon) drops from a ~121-125ms median
to a ~76-80ms median -- about a 37% cut -- consistently across snapshot -i
and press --settle. Total command wall time drops 10-21% depending on
background system load (daemon/device RPC time is unaffected by this
change and dominates more under contention).
|
||
|
|
f4fc089099 | 0.20.6 v0.20.6 | ||
|
|
628406390b |
fix(daemon): report the device claim retained by a failed close (#1647)
* fix(daemon): report the device claim retained by a failed close When close cannot confirm the device was released, it deliberately keeps the advisory claim (handing an unconfirmed device to the next session would be worse) but deletes the session record on the next line regardless, leaving a claim naming a session the daemon no longer tracks with no trace. Emit a warn diagnostic naming the device key and session, mirroring the open path's existing rollbackNewSessionClaim handling. Retention policy is unchanged. * test(daemon): pin claim retention on the cleanup-failure branch too The retention decision reads `platformCloseError ?? cleanupAggregate`, and the existing pair only drove the first input. A best-effort cleanup failure — a wedged perfetto stop, a dead helper — is the branch operators hit more often and reaches the same retention through a different value, so narrowing the diagnostic to the platform-close branch left every existing test green. Verified red against exactly that: gating the emit on `platformCloseError` fails this test alone, 28 others unaffected. Live evidence for the device-facing path is still outstanding; this closes the untested residual the review named, not that requirement. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Rva4YGtSCAKJqH5PbpcCU --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
f4ebd6f14f |
fix(ios): synthesize hidden-keyboard text through responder (#1657)
* fix(ios): synthesize hidden-keyboard text through responder * docs: document iOS text synthesis failure |
||
|
|
fd1ed578d5 |
fix(ios): keep the baseline presentation when a tap corroboration has no request flags (#1646)
* fix(ios): keep the baseline presentation when a tap corroboration has no request flags matchingCaptureFlags dropped the baseline snapshot's scope/depth/raw whenever the incoming request carried no flags, so the post-action corroboration capture ran at the default presentation and could never match a non-default baseline's presentationKey. The corroboration then silently declined to engage, leaking the raw XCTEST_RECORDED_FAILURE it exists to eliminate. * docs(test): correct tap corroboration reachability |
||
|
|
b19ee116b5 |
fix(remote): stop persisting the daemon bearer token, and authenticate forced-reconnect release correctly (#1648)
* fix(remote): stop persisting the daemon bearer token in connection state ADR 0007 requires generated connection profiles to strip daemon and Metro bearer tokens; only the Metro half was honored. `connect` was writing the daemon bearer token into the 0600 connection-state file, and every later command read it back out. Stop writing `authToken` into `RemoteConnectionState['daemon']` and resolve it at each reader from the existing flag -> environment (AGENT_DEVICE_DAEMON_AUTH_TOKEN) -> remote-config-profile chain instead, matching src/cli/auth-session.ts's precedence. Behavior change: a user who ran `connect --daemon-auth-token <value>` and relied on later commands picking the token back up from the state file will now get an auth failure. They must export AGENT_DEVICE_DAEMON_AUTH_TOKEN, set daemonAuthToken in their remote config, or pass --daemon-auth-token on each command. website/docs/docs/remote-proxy.md is updated to show the supported env-var workflow. * fix(remote): authenticate forced-reconnect lease release with the previous endpoint's own credential connect --force released the previous connection's lease using the new connection's ambient daemonAuthToken instead of the previous endpoint's own credential, and swallowed the resulting auth failure — silently orphaning the old lease when replacing a connection with a differently-authenticated one. Resolve the release token from the previous connection's own remote-config profile first, fall back to the ambient token only when the two connections share the same daemon endpoint, and otherwise skip the release and surface an actionable notice (tenant, run id, lease id, endpoint) through the existing connect notice channel instead of hiding the failure. * fix(remote): stop merging ambient env defaults into the previous lease's own token resolvePreviousOwnDaemonAuthToken read the previous connection's profile through resolveRemoteConfigProfile, which folds AGENT_DEVICE_DAEMON_AUTH_TOKEN (and other env defaults) into the result. When the previous config file declared no token and the new connection's credential came from that same global env var, it was misclassified as belonging to the previous endpoint and sent there on forced-reconnect release — recreating the credential leak the prior fix was meant to close, just via env instead of --daemon-auth-token. Read the previous profile with the new readRemoteConfigFile (a provenance- preserving, file-only load with no ambient env/CLI merging), so only a token the previous config file itself declares can satisfy rule 1. Rules 2 and 3 are unchanged. * fix(remote): verify the previous config file still speaks for its endpoint Rule 1 reads the previous connection's own config file to recover a credential that provably belongs to the previous endpoint. It re-read `previous.remoteConfigPath` and trusted whatever token that file holds *now* — but a config path is routinely reused, so "connect to A from ./remote.json, re-point ./remote.json at B, connect --force" classified B's token as A's own and sent it to A during lease release. Same cross-endpoint leak the env-merge fix closed, arriving through the file instead of the environment. The file must now still vouch for the previous endpoint, by either of two independent facts: its bytes still hash to the `remoteConfigHash` recorded at connect time (so it is literally the declaration that stood up the previous connection), or — if it changed — it still declares the same daemon base URL. The second is what keeps an ordinary credential rotation releasing its lease instead of orphaning one; endpoint equality, not the fact of an edit, is what separates rotation from re-pointing. Endpoint comparison runs both sides through `buildRemoteConnectionDaemonState`, the same normalizer that produced the stored `daemon.baseUrl`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Rva4YGtSCAKJqH5PbpcCU * fix(remote): bind previous config token to its endpoint --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
09a0881e6c |
refactor(ios): runner error classification as data; recovery tested at the transport seam (#1644)
* refactor(ios): runner error classification as data; recovery tested at the transport seam (#1631) Error-retry classification lived in four places consulted from one catch dispatch: two overlapping substring chains in runner-contract.ts, a code-keyed fatality chain in runner-session.ts, and an inline composite in runner-lifecycle.ts. RUNNER_ERROR_RULES is now the one declaration (mirroring RUNNER_COMMAND_TRAIT_MANIFEST): each rule names an error shape and its verdicts across the four axes (retryable, connect-retry, session-fatal reason, restart-before-send); the predicates keep their signatures and derive from the table. One deliberate widening, noted at the declaration: restart-before-send matching is now case-insensitive like the other axes (the message is our own transport literal). runner-command-recovery.ts gets its first direct coverage — through the real stack: a scripted fake iOS runner (local HTTP server) stands in for the XCTest runner process, executeRunnerCommandWithSession and the transport fetch run for real, and invalidation is observed through the module's existing invalidateSession parameter. Seven scenarios (retained response, runner-reported failure, in-flight, unknown lifecycle, probe failure, missing command id, read-only completed-without-response) plus an 11-case classification suite. Zero vi.mocks of runner internals. * test(ios): prove the recovery wiring and the recycled-pid guarantee (#1644 review) P1 was right, and I verified it before fixing: disabling the shipped recovery callsite in runner-lifecycle.ts left all seven recovery tests green. They call handleRunnerTransportErrorAfterCommandSend directly, so they prove the module, not the wiring. runner-recovery-wiring.test.ts enters at runAppleRunnerCommand — the production facade and provider seam — and fakes only session creation (the xcodebuild spawn). Real: command-id assignment, provider resolution, executeRunnerCommand's catch/classification, the recovery callsite, the recovery module, executeRunnerCommandWithSession, the transport fetch, and response parsing. Disabling the callsite now turns all three red. Entering one layer lower silently defeats it: the id is assigned by the facade and recovery declines to probe a command without one. P2: runner-lease-recycled-pid.test.ts covers #1621 with a REAL spawned process and real readProcessStartTime — no mock choreography. It asserts through the injected cleanup adapter rather than by observing a kill, so a regression reports a pid instead of signalling a live process; removing the identity guard turns it red. Only the refusal direction, which is contention-proof: the recorded start time never matches, so a timed-out verification read yields the same verdict. The opposite direction needs a second live `ps` read inside production code and flakes under load, so it stays with the mocked cases in runner-session.test.ts where pinning is the point. The fake runner now scripts per command rather than as one queue, because production sends a readiness `uptime` probe before a mutating command and `status` during recovery; a rigid queue coupled every test to that order. * test(ios): route the recycled-pid test through exec.ts and off the real lease tree (#1644 review) Two fixes to the same test: - Process execution goes through utils/exec.ts's runCmdBackground, per AGENTS.md's hard rule, instead of node:child_process spawn. The kill-induced `wait` rejection is owned deliberately. - The lease root defaults to the developer's REAL ~/.agent-device tree, so the test was writing leases there. It now redirects AGENT_DEVICE_IOS_RUNNER_LEASE_DIR to a temp dir per test and restores it, writing nothing outside the sandbox. (Verified by inspecting and cleaning the leases the earlier revision left behind.) |
||
|
|
10ff339d14 |
refactor: declare selector resolution policy as data (#1649)
* refactor: declare selector resolution policy as data (#1630) Five native consumers of "resolve a selector against the screen" each hand-declared their ambiguity contract as inline requireUnique/ disambiguateAmbiguous literals, so the repo's real policy matrix was only discoverable by reading four files. SELECTOR_RESOLUTION_POLICIES (packages/selectors) now declares one row per caller — ambiguity kind plus the structural columns (rect, occlusion, off-screen guard, promotion, poll) — and selectorResolutionKnobs turns a row into the engine knobs it stands for. Callers consume rows; zero ambiguity literals remain in src. Semantics are unchanged by construction: each row was read off its call site. The matrix names what was previously implicit — act and get text disambiguate, is/get attrs fail closed, exists/find-reads and wait take the first match, mutating find rejects candidates unless narrowed (#1625). `reject-candidates` is declaration-only and rejected by selectorResolutionKnobs at the type level, because find enforces it through its own narrowing rather than engine knobs. resolution-policy-parity.test.ts gate-tests the matrix against the callers (ADR 0011's declared-plus-gate-tested pattern): knobs must match the named ambiguity contract, every claimed structural column must appear in the caller's source, the read/wait pipelines must genuinely lack the machinery they disclaim, and no caller may reintroduce an inline literal. Verified revert-sensitive: flipping readUnique to disambiguate and faking wait's occlusion column each fail it. Out of scope, unchanged, per the issue: the Maestro engine (ADR 0015) and the open click-implicit-wait product decision. * refactor: route wait and mutating find through the policy interface (#1649 review) P1 was right: the first head declared seven rows but genuinely routed five. selector-wait.ts never imported its row (it called listSelectorChainMatches directly), findAct consumed only requireRect while its ambiguity contract stayed bespoke, and the parity test sniffed marker strings in source files — so it stayed green across exactly that gap. Asserting about the layer I had edited instead of the behavior it produces. resolveSelectorChainWithPolicy is now the one policy-driven entry: it returns a discriminated outcome (none / resolved / ambiguous) because the rows genuinely disagree about what several matches mean, which is what previously forced each caller to re-derive its contract inline. wait and find's selector branch both route through it; find additionally asserts its row still says reject-candidates rather than assuming. The parity test is rebuilt on fixture trees driven through that interface — no source sniffing. Wiring verified revert-sensitive: flipping the wait row fails the policy tests, and flipping findAct fails REAL find handler tests (ambiguous-candidate listing), which is the proof the previous version could not produce. One behavior nuance the fixture work surfaced and now pins: disambiguation declines on genuinely indistinguishable candidates (the tiebreak is evidence, not a coin flip), so an acting row surfaces ambiguity there rather than binding one silently. * fix(test): let fallow see the host-process mock helper's real consumers Rebase onto main brought #1642's host-process-mock.ts into this PR's fallow scope, where its export reports as unused. It is not: three suites consume it, but only through `(await import(...)).pinOwnProcessStartTime` inside vi.mock factories — vitest hoists those above static imports, so the dynamic form is required and fallow cannot trace it statically. Documented suppression rather than a restructure that would break the hoisting contract. Latent on main rather than introduced here: the audit gate is changed-files-only, so main sees the file in scope only from a PR whose diff contains it. * fix: keep every candidate when a policy resolves one winner (#1649 review P1) A real regression I introduced, not a test gap: routing wait through the policy interface collapsed the candidate set to the winner, and the #1349 landmark check is satisfied when SOME match carries the recorded identity. A first same-selector impostor therefore hid a later genuine landmark and timed the wait out. The resolved outcome now carries `matchedNodes` — the full candidate set of the alternative the winner came from — so a policy that picks one node no longer throws the rest away. wait passes that straight to the landmark check, restoring the original semantics. Regression test added at the within-one-poll shape the existing suite did not cover (both candidates in the SAME capture, impostor first); verified it goes red against the singleton reconstruction it replaces. * refactor: declare only the policy fields the matrix enforces (#1649 review) The occlusion / offscreenGuard / promotion / poll columns were never consumed by resolveSelectorChainWithPolicy or selectorResolutionKnobs: changing any of them left behavior and the suite green, so they were unverifiable claims that read as truth. (My earlier source-sniffing test "verified" them by grepping caller files for marker strings — which is why it also stayed green when a row was disconnected entirely.) The matrix now declares exactly what it enforces: the ambiguity contract and the rect requirement, both consumed by the resolution interface and pinned behaviorally. A new test asserts every row's field set, so an unenforceable column cannot reappear without coverage — verified by re-adding one and watching it fail. Routing the structural stages into typed behavior is tracked in #1656 with the constraint that each field must be consumed, not merely declared. * fix(selectors): flatten the policy outcome at the package boundary `PolicyResolutionOutcome.resolution` was typed as `AstSelectorResolution` and the root façade returned it unchanged, so the parser AST #1589 confined to `@agent-device/selectors/ast` came back through a nested field. `selector-wait.ts` reading `outcome.resolution.selector.raw` was the runtime proof. The existing boundary gate reads exported *names*, so it could not see this. The public outcome now lives beside `SelectorResolution` in public-resolution-types.ts with its selector as text; the parser-side shape is renamed `AstPolicyResolutionOutcome` and stays package-private, and the façade wrapper flattens on the way out — the same treatment `resolveSelectorChain` already gave `AstSelectorResolution`. Two new pins, both verified red against the shape they replace: a behavioral one asserting the façade returns selector text under every policy row, and a structural one asserting resolution shapes are re-exported from public-resolution-types.ts rather than from a parser-side module — which is what distinguishes the leak from a correct re-export in a name list. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Rva4YGtSCAKJqH5PbpcCU --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
0033002ed5 |
fix(daemon): say why a no-effect claim was withheld (#1655)
* fix(daemon): say why a no-effect claim was withheld A vetoed gestureNoEffect claim is invisible from outside: the response looks exactly like a gesture that worked. #1620 spent weeks unable to tell a cross-backend pair from capture drift from real movement for precisely this reason, and its own suggested next step was to log the corroboration inputs and re-run. Every veto on an accept-stale verdict now emits post_gesture_no_effect_vetoed with a reason: baseline_rebased (the pair is cross-backend, #1569) or surface_divergence carrying onlyInBaseline/onlyInCurrent/rectMismatched/shared from a new summarizeDiscriminatingSurfaceDivergence — one-sided keys are membership drift, rectMismatched is movement. Measured live against the seeded Bluesky fixture (167 nodes, depth-56, private-ax): inert scrolls corroborate and the warning fires 2/2, successful gestures never claim 2/2, zero rebases, and the one veto observed was rectMismatched:4 shared:158 onlyIn*:0 from ambient feed motion over an artificial 105s gap. #1620's truncation-drift half does not reproduce; membership is stable under the remembered depth cap. Also: marking no longer stores an empty baseline signature. '[]' is truthy, so an invented baseline was being rebased onto a post-gesture capture; 'no usable baseline' now has one representation, matching markPendingInteractionOutcome. * test(daemon): pin the veto diagnostic's payload, not just its phase The claim tests counted `post_gesture_no_effect_vetoed` events by phase. The operator-facing behavior this PR adds is the *reason* and the divergence counts, and both survive an event tally: emitting the wrong reason, or dropping the counts entirely, keeps every count at 1. The four veto tests now assert the whole emitted payload, read back out of a diagnostics trace file so the assertion pins the serialized NDJSON line rather than an in-memory event. Backend rebase pins `reason: baseline_rebased` with no counts riding along (the pair is cross-backend; there is nothing comparable to count over). Both divergence cases pin `reason: surface_divergence` with exact onlyInBaseline/onlyInCurrent/rectMismatched/shared — the numbers that separate scope drift from movement, and the two membership-drift directions from each other. Verified red both ways the reviewer named: dropping the counts fails 2 tests, forcing the reason to surface_divergence fails 2. The old count assertions stayed green under both. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Rva4YGtSCAKJqH5PbpcCU --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
d85072d935 |
perf(cli): route command aliases through the help fast path (#1641)
* perf(cli): route command aliases through the help fast path bin.ts's `--help` fast path resolved aliases through a hand-written two-entry table that had drifted out of sync with the real CLI_COMMAND_ALIASES registry (five entries). `tap`, `launch`, and `relaunch` missed the table and silently fell through to a full runCli() bootstrap just to print static help text (~150-165ms vs ~45-50ms for aliases already in the table). Delegate to the shared normalizeCliCommandAlias registry instead of the stale local table, so every alias the registry knows about gets the fast path automatically. * test(cli): add R12 layering guard for bin.ts's alias delegation The unit test added for the alias fast-path fix (cli-help-alias-fast-path.test.ts) calls normalizeCliCommandAlias directly, so it stays green even if bin.ts itself reverts to a hand-rolled table — it pins the registry composition, not bin.ts's own wiring, and bin.ts cannot be safely unit-imported (it runs unguarded top-level dispatch on import and is deliberately excluded from coverage). Add an AST-based structural guard instead, in the style already established by scripts/layering/session-state.ts, facade-exports.ts, and zero-dep-jobs.ts (oxc-parser's module/program records, not a line scan, so a fixture's string literal can't produce a false hit). R12 asserts two facts about src/bin.ts: it holds a value import of normalizeCliCommandAlias from commands/cli-command-aliases.ts, and it contains none of the registry's own alias tokens as string literals. The token list is read out of the registry's own source (CLI_COMMAND_ALIASES's `alias:` property values), not hard-coded, so a future sixth alias is covered automatically. Both facts were false on the pre-fix bin.ts, verified by reverting locally and capturing the failure before restoring the fix. Wired into the existing check:layering chain (already part of check:tooling), next to R7's session-state ownership rule, which pins the same "delegate to your single owner" shape. * test(cli): pin the alias-resolver call into buildCommandUsageText (R12 P2) Maintainer review of R12 (PR #1641): import-presence and literal-absence alone let bin.ts regress to buildCommandUsageText(helpTarget) while the normalizeCliCommandAlias import stays in place, used harmlessly elsewhere (or not at all) — the real-tree gate stayed green through that exact regression. Add a third fact: bin.ts's call to buildCommandUsageText must receive, as its argument, a call to the LOCAL binding the resolver was imported as (aliasResolverLocalName + usageTextCallsResolver, both AST-based). Binding by local name rather than the literal export name means a renamed import (`as resolveAlias`) still verifies, and an unrelated same-named local cannot be mistaken for it. Verified by reverting locally to exactly the missed regression — import left in place, call reverted to buildCommandUsageText(helpTarget) — and confirming R12 now fails where the two-fact version passed; restored after. Two negative fixtures pin the scenario going forward: import present but unused, and import present but used only unrelated to the call. * test(cli): make R12's delegation fact universal and value-bound The previous fact 3 asked whether *any* `buildCommandUsageText(resolver(...))` existed in bin.ts. That quantifier is satisfied by a decoy call while the line that actually ships resolves nothing: void buildCommandUsageText(normalizeCliCommandAlias('open')); const commandHelp = buildCommandUsageText(helpTarget); Fact 3 now requires EVERY `buildCommandUsageText` call to receive the imported resolver applied to the fast path's own help-target binding, which rejects both lines above independently. The help-target name is read from bin.ts (the variable initialized by `resolveSimpleHelpTarget`), so renaming it re-points the guard instead of disarming it. Because fact 3 claims binding identity by name, it also now rejects a local shadow of the resolver and an ambiguous second help-target declaration — a same-named local would otherwise let the composition read as delegation while calling something that resolves nothing. The predicate returns the reason rather than a boolean, so the gate names which of the several distinct failures happened. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017Rva4YGtSCAKJqH5PbpcCU --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
3937036e5e |
feat: support --settle on scroll and back (#1638) (#1650)
* feat: support --settle on scroll and back (#1638) Scroll-then-observe and back-then-observe are legitimate agent pairs, but the post-action observation registry never grew past the touch commands, so `--settle` on either was rejected with INVALID_ARGS — burning a tool call each in AppControlBench's bsky-16. Both commands now carry the `settle` descriptor trait, and every surface derives from it rather than a hand list: CLI allowed flags, MCP/SDK input fields, the flag-sourced timeout envelope, and MCP ref-pinning. The CLI flag/metadata helpers moved out of the interaction family into post-action-observation-grammar.ts (back is a system command), and SETTLE_REF_ISSUING_TOOLS became a derivation — a hand list would have silently stopped pinning the new commands' refs. settleAfterInteraction and the new settleObservationCommand are two entry points over one engine: same loop, storage, hints, and diff bounds, with the target-less path supplying its own baseline and no proximity point. The daemon reaches that command through the runtime surface, never by importing `commands/` (R2) — the same seam the touch handlers use for press/fill — and generic-settle.ts is loaded through a lazy `await import` returning a closure, so the interaction runtime subgraph stays out of this dispatcher's static graph (a static edge folded ~18 files into the daemon-server type cycle; R10 caught it). Both of generic-settle's orderings are load-bearing and tested: the baseline is frozen before dispatch (and before the Android dialog preflight), and the observation runs after markDeferredInteractionOutcome so settle's first capture folds in the #1542 stabilization rather than racing it. The ADR 0014 "a settled diff publishes refs" rule moved to settle-ref-issuance.ts, shared by both routes. One divergence is deliberate: scroll/back resolve no element, so the diff baseline is the session's STORED pre-action tree — "settled tree vs the last tree you observed" — not press's freshly resolved pre-action capture. Both commands also switch to preserve-daemon on timeout, which changes the non-settle path too: with --settle their dominant hang mode is now a wedged accessibility bridge, and a timed-out capture must not reset the daemon and lose every session (#1105). The reviewed-set gate records it. Live-validated on an iOS 26.2 simulator (Settings): scroll --settle settled in 1786ms with a +6/-6 diff carrying fresh refs; back --settle in 771ms with +15/-6. Alternating cost runs, one call vs the pair it replaces: scroll 2.9-3.0s vs 5.3-5.6s, back 3.1-3.2s vs 4.7-5.1s. Those include the #1627 deep-capture extension. * fix: render settled-diff refs paste-ready in CLI output A settled diff activates a PARTIAL ref frame (ADR 0014), which admits only the pinned `@eN~s<gen>` form of the refs it issued. The unchanged-interactive tail already rendered that way, but the diff's own added lines rendered the bare `@eN` embedded in the snapshot line — so a CLI caller who copied the ref the diff just handed them got `plain_ref_requires_complete_frame` and had to append the generation by hand. Added lines now render pinned when the response carries `refsGeneration`, exactly like the tail. Removed lines render verbatim: they name elements that just left the screen, and `SettleDiffLine` never gives them a ref. This is not new to scroll/back — press/click/fill/longpress had the same gap since #1101. MCP was never affected: its ref-pin store rewrites plain refs on the way in, which is why the model never sees a suffix. Live: `scroll down --settle` now emits `+ @e14~s218078 [cell] "Game Center"`, and `press @e14~s218078` copied straight out of that line taps successfully. * test: record the pinned-diff-ref bytes in the output-economy baseline Rendering added diff-line refs pinned costs 8 bytes in the two settle CLI text samples (two `~s<gen>` suffixes). The output-economy baseline is the tripwire for exactly this, so the increase takes an explicit reviewed waiver rather than a silent baseline bump — the same one the settled TAIL's pins already carry, for the same ADR 0014 reason. Only `bytes` moves: lines, refs, hints, and shape are unchanged, which is the evidence that this is a suffix on existing refs and not a new payload. Caught by CI, not locally: `pnpm test:unit` runs unit-core and subprocess-stub only, while the Coverage lane runs every vitest project. * test: prove the generic settle degrades when its runtime cannot be built `createGenericSettleRuntime` catches and returns undefined so an observation that cannot even start does not fail an action that already succeeded. That was a claim in a docstring with nothing behind it — the one changed line the coverage gate reported uncovered (95/96). The test puts the session in the state the catch exists for: the router handed us a session that is no longer in the store, so building the settle runtime throws SESSION_NOT_FOUND. The response keeps its scroll result and simply carries no settle payload. Removing the try/catch fails it. * build: teach fallow that vi.mock reaches pinOwnProcessStartTime dynamically Not from this PR: #1642 added `pinOwnProcessStartTime` on main, and its three consumers reach it the only way a Vitest module mock can — `vi.mock(path, async (importOriginal) => (await import('...')).pinOwnProcessStartTime(...))`. Dependency analysis cannot follow that dynamic import to a consumer, so the export reads as dead the moment any PR pulls that file into its audit scope. This PR is the one that did. The entry records the consumers by path and the reason, matching the daemon route-handler entry directly above it, which exists for the same dynamic-`import()` limitation. * refactor: adopt the best of the parallel #1653 implementation Two sessions independently built #1638 (PR #1650 and PR #1653) and converged on the same architecture — trait in the registry, one engine with two entry points, runtime-command seam, lazy import, preserve-daemon, stored-baseline honesty. #1650 continues; this folds in what #1653 did better: - The agent-facing help core loop (cli-help.ts) now names scroll and back as settle-capable. Without this, the benchmarked closed-grammar help line kept instructing agents that --settle is only for press/click/fill/longpress — actively steering the AppControlBench models away from what #1638 shipped. - issueSettleRefs moves into session-snapshot.ts, beside the partial-frame primitive it wraps, deleting the single-function settle-ref-issuance module. - Their seam tests: back reader→writer settle plumbing, back CLI settle rendering, and a trait-less generic command (home) ignoring a stray settle flag rather than observing or rejecting. What #1650 had that #1653 lacked, for the record: the SETTLE_REF_ISSUING_TOOLS registry derivation (without it, MCP never pins a scroll/back settle diff's refs and the partial frame rejects every follow-up), BackCommandResult.settle in contracts, back's MCP output schema, paste-ready pinned diff refs, and the docs/changelog/baseline surfaces. * bench: help-conformance case for settled scroll-to-find planning The #1638 extension of the closed --settle grammar to scroll/back is the feature's entire payoff — collapsing scroll-then-observe into one call — and the closed command list is an enumerated N whose enumerator is this bench. The regex over the help text proves the sentence exists; this case checks whether a model plans differently because of it. One focused case, deliberately not coached: a pinned visible-first snapshot (rendered by formatSnapshotText, pinned by the sample-producers gate) whose wanted row is summarized off-screen with no ref anywhere in the output. The tempting pre-#1638 plan is `scroll` plus a separate `snapshot -i`; acceptance is the single settled call. Scoring was verified against eight plan shapes in both directions before recording. Model-backed record (claude-haiku-4-5, 3 trials, current help): 0/3 — but the decomposition is the finding. Settle eligibility GENERALIZED (3/3 trials put --settle on scroll unprompted; the mutation-suffix framing concern did not materialize) and the two-call habit is residual (1/3). All three trials failed on `scroll @e3 down --settle` — the pre-existing #1366 scroll-takes-no-target confusion, which the live CLI recovers with a dedicated hint but a single-shot bench cannot. The recorded gap is therefore a first-30 doc gap (nothing teaches that scroll takes no target), not a settle-eligibility gap; tuning the case until it passes would just delete the evidence. |
||
|
|
8030fc10d6 |
fix: make Android record stop survive static screens, slow moov finalization, and dead-pid recovery (#1651)
* fix: make Android record stop survive static screens, slow moov finalization, and dead-pid recovery Three live-reproduced defects on a loaded Pixel_7_CI emulator shared the "pulled file is not a playable MP4" / "manifest could not be verified" symptom family: - The Swift video validator required duration > 0, permanently rejecting valid single-frame recordings of fully static screens (AVFoundation reports their duration as 0). Whether a run passed depended on whether anything — even the status-bar clock — changed during the window. - screenrecord finalizes by patching a front-reserved moov in place, so the remote file size never changes; the copy path now re-pulls with escalating delays (750/1500/3000ms) and detects finalization from the pulled bytes via a mandatory ftyp+moov container sniff, which also preserves the truncation detection the duration check provided by accident. The local waitForStableFile call is gone: a pull is complete when adb returns. - toybox `ps -p <missing-pid>` exits 1 with empty output — the normal pid-gone signature — but the recovery liveness probe read every non-zero exit as an uncertain adb failure, making a finished recording behind a live-status manifest unrecoverable forever. Empty-output failures now corroborate via the full process list: healthy listing with the pid absent recovers the finished recording; listing failure stays conservatively uncertain (a stale verdict deletes the manifest, so transport health is proven first). Transport failures keep stderr and exec-layer timeouts throw, which is what makes the empty-output signature safe to trust. Provider-scenario coverage: in-place same-size finalization landing past the first retry, and dead-pid recovery on a responsive device — both fail on the previous implementation with the live-observed errors. * refactor: extract Android liveness probe and split recording scenarios per file-shape rule record-trace-android-recovery.ts (763 lines) and android-recording.test.ts (1,764 lines) were both past the 500-LOC extraction tripwire. The screenrecord liveness/probe concept this PR modified now lives in record-trace-android-liveness.ts, and the new scenarios moved to test files mirroring the source modules they cover (record-trace-android-copy.test.ts, record-trace-android-liveness.test.ts) with shared scenario plumbing in the android-recording-fixtures.ts sibling. android-recording.test.ts shrinks to 1,476 lines — below its pre-PR size. |