mirror of
https://github.com/callstack/agent-device.git
synced 2026-09-14 20:06:34 +08:00
v0.20.10
443 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a853734f0c |
fix(webdriver): give cloud session creation its own budget and stop leaking billed sessions (#1782)
* fix(webdriver): give cloud session creation its own budget and stop leaking billed sessions Cloud lease allocation ran under the generic 30s/1-retry request policy, so BrowserStack iOS real-device session creation (45-90s) aborted client-side at ~60s on most runs. Each timed-out POST /session still completed server-side and, being non-idempotent, was retried — leaving two billed provider sessions per failed open with no id to release them. - POST /session is its own phase: a 180s create budget (default), zero retries, and no request-bound abort, so the daemon always learns the session id. - lease_allocate carries a 300s allocation budget surfaced to providers as LeaseLifecycleContext.deadline, and a matching 330s client envelope that preserves the daemon on timeout (a reset would SIGKILL mid-create and orphan every billed session the daemon held). - The request's cancellation signal is ownership evidence: a session that completes after the requester left is released, not registered; a create that the transport gives up on surfaces typed evidence (provider + lease) so an operator can find and stop the maybe-orphaned session. Closes #1774 * refactor: one canceled-request error, and tighten the #1774 shapes Review pass over the session-create fix: - The canceled-request error had nine hand-rolled copies (src/request/cancel, maestro shared, exec, retry, install-source x2, and the new provider one). It now has one definition in @agent-device/kernel/errors: createRequestCanceledError(details?, cause?) + isRequestCanceledError + REQUEST_CANCELED_REASON. Callers add evidence or a sharper hint; the reason itself is not overridable, so nothing can build one the predicate misses. - lease_allocate's timeout bundle moves beside INSTALL_TIMEOUT_POLICY in the registry (same {...DEFAULT, envelopeMs, onTimeout} shape); the request timeout constant stays exported from timeout-policy like its siblings. - Transport: fetch helper returns Response's own ok/status; the timeout reason const is private behind isWebDriverRequestTimeout. - Client: one-use options type inlined; the two deadline helpers share one floor. - Session-manager tests: shared makeRuntime/jsonResponse/afterEach restore. Net -29 lines with the feature in. * chore: keep the canceled-request reason private to the kernel * fix: typed cancellation everywhere + own the AWS remote-access ARN through startup Second-order follow-ups the #1774 refactor made cheap: - markRequestCanceled aborts the request signal WITH the kernel's typed canceled error as its reason. Every signal.throwIfAborted(), aborted fetch, and 'throw signal.reason' in the daemon (20+ sites) now surfaces a canceled request as such instead of a bare DOMException that normalized to UNKNOWN — and no site has to know the factory exists. - AWS Device Farm prepareSession owns the remote-access ARN from the moment create-remote-access-session answers: a startup timeout, the allocation deadline, or a canceled request now stops it before the failure surfaces (previously a timed-out startup left a RUNNING billed session behind — the same leak class as the WebDriver session, one phase earlier). The startup wait is capped by LeaseLifecycleContext.deadline and wakes on cancellation. - BrowserStack's pre-session local app upload honors the request signal (an upload is not billed, so plain abort is right there). - lease_heartbeat/lease_release share lease_allocate's preserve-daemon policy: the rationale — the daemon owns billed sessions; a reset orphans them all — applies verbatim. Each AWS ownership test proven red without the guard (3/3). * refactor: dedupe billed-resource cleanup and lease-signal wiring Shrink pass — same behavior, less duplication: - releaseOnFailure(primaryError, release) in webdriver-utils replaces the two identical 'best-effort stop the billed resource, attach cleanupError to the primary AppError' helpers (WebDriver session + AWS remote-access ARN); shared errorMessage too. - The lease handler pulls the request signal from getRequestSignal(requestId) like every sibling handler, instead of threading a requestSignal arg through LeaseHandlerArgs and the request-handler chain. Drops the field, the wiring, and five mechanical test edits; the handler test now proves the request-bound signal (abort it, watch the provider's signal flip) rather than arg identity. - Inlined the one-use requestHeaders back into fetchWebDriver. Handler-signal test proven red without the wiring. * fix(lease): the daemon releases a lease allocated for a gone requester; honest release evidence Review follow-up. The provider was doing the daemon's job: it treated the request signal as 'ownership evidence, not an interrupt' and needed three paragraphs to say so. The daemon owns the request, so it now decides — generically, for every provider — what happens to a lease that finished allocating after its requester left: release it (provider + registry) and answer with the canceled error. - lease.ts: after allocate returns, isRequestCanceled(requestId) → releaseAllocationForGoneRequester(). Release evidence is claimed ONLY on a clean release (no warnings, no throw); a WEBDRIVER_SESSION_DELETE_FAILED release is reported released:false with providerSessionId + a stop-by-hand hint (thymikee's finding: the previous evidence was success-shaped even when DELETE failed). - WebDriverSessionManager: the createOwnedSession/releaseCanceledSession trio is gone; allocate is plain 'create with a budget; on failure clean up' again. - LeaseLifecycleContext.signal is just cancellation, like everywhere else; the ownership-semantics comments on the contract, client, registry, AWS prepare and utils shrink to what the code no longer says itself. - Tests: the two provider-level cancellation tests move to the daemon handler (where the logic now lives), plus the failing-DELETE regression; both proven red without the post-allocate check. * fix(aws): the allocation deadline bounds remote-access startup, not the 120s default Live iOS real-device run: startup needed ~128s and hit the standalone 120s default while the daemon's 300s allocation budget still had room — the new ownership guard correctly stopped the ARN, but the open failed for no reason. When the daemon supplies a deadline it is the bound; the default only applies standalone. Rerun: open in 112s, snapshot, clean close, session STOPPING. * test(aws): pin that the allocation deadline outlives the 120s startup default; drop empty import Review follow-ups on 7f9d1481a: a virtual-clock test (Date.now advanced 10s per poll, RUNNING at 150s, deadline 300s) that fails on the old min(default, deadline) logic and passes now; and the empty 'import {} from kernel/errors' left in maestro/shared.ts is removed. * refactor: finish the dedupe — one release path, kernel errorMessage, AWS on releaseOnFailure Code-quality review at 7f9d1481a: 1. aws-device-farm.ts still carried its own copy of releaseOnFailure (the dedupe commit's script aborted before reaching it and I mis-verified). Now uses the shared helper; private copy deleted. 2. Empty 'import {} from kernel/errors' in maestro/shared.ts removed (273870099). 3. errorMessage() lives in @agent-device/kernel/errors; the two copies this PR had added (lease.ts, webdriver-utils.ts) import it. Sweeping the pre-existing copies is a follow-up. 4. lease.ts has ONE release path: releaseLease(registry, provider, lease, request, ctx) → { released (registry), provider } used by both the lease_release case (wire shape unchanged) and the gone-requester branch, which folds a throwing provider release into releaseError. 'released' now means the same thing in both; the provider verdict is a separate 'providerReleased' (warnings-free, no throw) that drives the stop-by-hand hint. -~35 lines. 5. sessionCreateTimeoutMs is Omit-ed at the WebDriverTransportOptions boundary instead of Pick-ed back out internally. * fix(lease): 'released' on a canceled allocation means the billed session is confirmed gone Re-review at 3665ea06: unifying the release path had made the cancellation error report released:true from the daemon's registry record while the provider DELETE had failed — success-shaped again, with the operator verdict demoted to a second key. Fixed at the source of the ambiguity: - LeaseReleaseOutcome names its bookkeeping field registryReleased. - On the canceled error, 'released' is true only when registryReleased AND the provider released without warnings AND without throwing; the registry record is exposed as 'registryReleased'. The stop-by-hand hint keys on 'released'. - lease_release keeps its existing wire field ('released' = registry; provider cleanup rides in 'provider'), unchanged. - Regressions: failed DELETE and throwing release both pin released:false / registryReleased:true (+ providerSessionId, warnings|releaseError, hint); both proven red on registry-only semantics. * ci: retrigger default-setup CodeQL Run 32051017472 is wedged on GitHub's side: status=completed with Analyze (python) still queued and Analyze (java-kotlin) failed only at SARIF upload (503, 'No server is currently available'). It can be neither cancelled nor rerun, and default-setup CodeQL has no dispatchable workflow, so a new push is the only way to get a fresh run. No source change. * test(webdriver): assert the typed timeout contract on the shared-budget probe main's #1790 tightened this test to expect the raw TimeoutError DOMException, which this PR intentionally normalizes into AppError{reason: webdriver_request_timeout}. On the merge ref the two met and Coverage went red. The regression now asserts the structured contract and that the second request's budget is the shared remainder (~118ms of 200 after an 80ms first call). |
||
|
|
d0d5c8594c | fix: serve remote daemon request diagnostics to the caller (#1801) (#1814) | ||
|
|
f378050586 |
feat(snapshot): attach a fallback screenshot to sparse captures (#1764)
A sparse verdict already tells the caller to use a screenshot as visual truth, which made that screenshot the guaranteed next command on every unreadable screen — a second round trip to obey advice we authored. The user-facing `snapshot` dispatch now takes the shot itself and links the path in its warnings. The fallback is deliberately hung off `dispatchSnapshotViaRuntime` and skipped for internal observations: selector resolution, settle, and wait polling reach `captureSnapshot` directly, so a wait polling an unreadable screen cannot turn into a screenshot per poll. A failed shot is swallowed — the verdict's own warning still carries the manual remedy, so the fallback can never fail the snapshot that was asked for. Sparse captures also say when the screen is the app's problem. Only the `sparse-tree` reason code is evidence about the app: every backend reached the screen and it published no semantic content, which is the same emptiness assistive tech gets. `ax-rejected`, `budget`, `no-nodes` and `capture-failed` are limits of this tool and stay unattributed, so readers are not sent to file bugs against code that is not broken. |
||
|
|
13c482ce50 |
fix: preserve compact snapshot action labels (#1773)
* fix: preserve compact snapshot action labels * fix: narrow iOS implementation cell labels |
||
|
|
d8a7d03faf |
refactor: route application lifecycle through runtime facts (#1759)
* refactor: route application lifecycle through runtime facts Moves the canonical `open`, `prepare`, `close` and internal `runtime` descriptors behind package-owned lifecycle bindings admitted from device runtime facts, while daemon request/session policy and public response construction stay put. Based on main, which already carries the boot unit, the parametrized cutover gate and the apps unit. Readiness is package-owned there, so the Apple and Android bindings call ensureAppleReady/ensureAndroidReady rather than a root readiness bag; ensureAppleReady gained an onColdBootStart hook so open keeps warming the runner cache in parallel with a cold boot, and a narrow markBooted port publishes readiness' fresh observation so a flow still makes one simctl listing. Cutover rows take R24-R27, clear of the accepted catalog and the sibling install stack, and cutoverTableDefects rejects a duplicate rule id. Two defects this unit introduced are fixed here rather than shipped: `open <app> <url>` dropped the URL on a first open, and test-IME activation was first fatal on an unobtainable helper and then over-caught. Helper unavailability is a typed non-activation outcome now; fence, lock and post-record failures propagate. The duplication the unit had accumulated is gone: one runtime-admission module instead of five per-command copies, one direct-lifecycle binding factory instead of six hand-rolled packages, one transport-hint predicate, one session finalization path, and no identity-wrapper module. * fix: allocate lifecycle cutover rows after deployment * chore: preserve lifecycle union reconstruction * fix: reconcile lifecycle runtime stack * refactor: tighten lifecycle runtime topology * refactor: remove superseded runtime adapters * fix: preserve stacked runtime cutovers * test: preserve migrated runtime ownership * test: move Android deployment retry ownership * test: extract runtime hint fixtures * fix: preserve lifecycle stack invariants * fix: complete lifecycle runtime cutover * fix: remove lifecycle cutover residue |
||
|
|
66cca1a5b8 |
refactor: route install commands through platform runtime (#1758)
* refactor: route install commands through platform runtime * fix: preserve stacked runtime facts * fix: align deployment facts with shutdown runtime * refactor: simplify capability facts projection * style: format capability facts projection * fix: preserve migrated capability ownership * fix: propagate deployment artifact cancellation * refactor: move Harmony deployment mechanics into package * refactor: move Apple deployment tools into package * refactor: move Android deployment tools into package * refactor: inject deployment temporary storage * refactor: remove superseded deployment helpers * fix: preserve provider deployment transport |
||
|
|
5b1b1aad2e |
fix: fingerprint the built daemon's real import graph (#1545) (#1771)
* fix: fingerprint the built daemon's real import graph (#1545) The static-import regex in computeDaemonCodeSignature required whitespace around `from` and after `import`/`export`. The built daemon entry is a minified bundle (tsdown/rolldown `minify: true`), whose real import statements have zero whitespace (e.g. `import{o as e}from"../foo.js"`), so the regex never matched any dependency edge and the graph walk silently degraded to fingerprinting only the entry file (`graph:1`) instead of its true ~50-file runtime graph. That defeats the safety net the signature exists for: a daemon started from one build can look identical to a client running a materially different build as long as the entry file's own size/mtime happen to coincide, and the "code-signature mismatch" takeover this check drives was one of the mechanisms observed in #1545's worktree dev-daemon multiplication. * refactor: fingerprint import specifiers by shape, not by import grammar Replaces the two keyword-anchored regexes (one for static import/export/from clauses, one for dynamic import()) with a single pattern that matches any quoted, relative-path-shaped string literal. The prior approach had to track the exact whitespace/keyword shape bundlers emit around a specifier, which is exactly what broke for minified output in the immediately preceding commit — formatted source, minified builds, static imports, re-exports, and dynamic imports all put the same thing in the same place (a relative string literal), so matching on that directly is both simpler and immune to future formatting differences. A same-shaped string that isn't really an import is safe to over-match: it just fails to resolve to a file and gets dropped. Verified the file count is unchanged against this repo's real build (dist: graph:50, src: graph:741, both identical to the keyword-anchored fix). |
||
|
|
b59b4e5c96 |
fix(daemon): make request-timeout cleanup and hints platform-aware (#1751)
* fix(daemon): make request-timeout cleanup and hints platform-aware
Every local request timeout ran the Apple-runner xcodebuild pkill cleanup
and emitted Apple-specific hint wording ("Apple runner work was aborted",
"Timed-out Apple runner xcodebuild processes were terminated") regardless
of the session's actual platform, which was misleading on Android/web/
Harmony timeouts.
The client already carries the request's declared --platform on the exact
DaemonRequest object both timeout call sites hold (req.flags?.platform), so
no new plumbing or src/platforms import is needed. When the platform is
known non-Apple, the Apple pkill cleanup is skipped and the hint drops its
Apple-runner claim; when it's unknown, both stay on their historical
fail-safe behavior.
Part of #1739 (independent cleanups)
* fix(daemon): split timeout cleanup eligibility from hint evidence
Maintainer review on #1751 found a real correctness inversion in the
first pass: gating the Apple xcodebuild pkill cleanup on the request's
declared --platform flag is wrong in both directions.
- req.flags.platform is not authoritative for session-bound execution.
applyStripLockPolicy (request-lock-policy.ts) lets an existing
session's real device platform silently override a conflicting
declared selector under --session-lock strip, so a request declaring
a non-Apple platform can still legitimately execute Apple work.
Skipping cleanup on that declared flag would skip real cleanup —
the dangerous direction.
- The common session-bound request omits --platform entirely, so the
previous fail-safe (undeclared -> treat as Apple) left the hint
wrongly claiming Apple involvement on the motivating case
(Android/web/Harmony session timeouts) too.
Redesign: separate the two decisions.
- Cleanup eligibility is unconditional again for every local timeout,
matching the original pre-fix behavior: the pkill patterns are
Apple-process-name-specific, so sweeping them on a non-Apple host or
session matches nothing and costs a few no-op subprocess spawns,
never a wrong skip. There is no client-visible signal that proves a
session-bound request cannot touch an Apple runner, so eligibility
does not try to prove one.
- The hint may only name Apple-runner involvement on evidence this call
site actually has: an explicitly declared Apple platform selector, or
the cleanup itself having terminated a matching process
(appleCleanupEvidence). Anything else gets platform-neutral wording.
Adds production-seam route tests
(src/daemon/client/__tests__/daemon-client-timeout-route.test.ts) that
spy on the real exec seam and drive actual socket/HTTP timeouts through
sendRequest, covering the rebound-session and unknown-session cases a
pure formatter test cannot catch. Proved red against the prior commit
(
|
||
|
|
c7565cb1f8 |
refactor(snapshot): clean snapshot ownership (#1754)
* refactor(snapshot): clean snapshot ownership * fix(snapshot): address ownership review feedback |
||
|
|
b13c06338e |
refactor: route boot through readiness runtime (#1747)
* refactor: route boot through readiness runtime * fix: separate boot admission from readiness * fix: register boot cutover policy * refactor(runtime): move readiness into platform owners * fix(test): tolerate provider temp cleanup races |
||
|
|
f5d9789764 |
feat: enforce local device claims and reconcile stale owners (#1735)
* feat: enforce local device claims * fix: address device claim review feedback * fix: persist canonical daemon claim state directory |
||
|
|
602b7a2995 |
refactor: narrow perf API to actionable evidence (#1731)
* refactor: narrow perf API to actionable evidence * fix: address perf API review feedback * fix: preserve deprecated Android CPU metrics |
||
|
|
b0d4b40467 |
chore: stop publishing skills to npm (#1730)
* chore: stop publishing skills to npm * fix: align simulator skill startup * docs: align agent setup with open-first workflow |
||
|
|
1b2e786128 | refactor: move screen recording onto platform runtime (#1724) | ||
|
|
1b75e102d7 | perf: speed up device inventory and status (#1723) | ||
|
|
cdc754e6ed |
perf: speed up iOS agent recovery and streamline CLI guidance (#1700)
* Avoid interactive children in parent taps * docs: streamline no-skill CLI help * perf: recover faster from sparse iOS trees * fix: preserve selector context for blocked parent taps * fix: preserve coordinate text-entry focus * fix: preserve thin parent touch targets * fix: fail closed for unscoped iOS typing * test: isolate replay lock fixture * test: share node integration process |
||
|
|
b1ed5353d1 |
refactor: extract platform log runtime (#1701)
* refactor: extract platform log runtime * fix: clear terminal app log recovery markers * fix: preserve scoped app log tooling * fix: preserve app log cancellation * fix: handle large changed coverage diffs * fix: harden Limrun runtime identity * refactor: tighten platform log runtime * fix: close app log trust gaps * fix: accept canonical session path aliases * refactor: extract durable capture kit * fix: refresh retained log marker admission * fix: rotate app logs after process relaunch |
||
|
|
4b432fb59b |
feat: add HarmonyOS support (#1683)
* feat: add HarmonyOS device automation foundation Add HDC-backed discovery, snapshots, application lifecycle, and core mobile interactions. Route HarmonyOS through the platform registry and client contracts. Cover parsing and capability parity with focused tests. * feat: support HarmonyOS HAP deployment Install and reinstall signed HAP archives through HDC. Resolve bundle identities from module metadata and relaunch after package replacement. Extend deploy routing and capability coverage for HarmonyOS. * feat: add HarmonyOS single-pointer gestures Execute pan, fling, and swipe plans through HDC uiInput primitives. Derive scroll coordinates from the live ArkUI viewport. Keep unsupported multi-touch gestures explicitly rejected. * refactor: split session inventory command handling Separate session, device, capability, and app inventory response paths. Preserve the public inventory response contract while reducing handler complexity. * feat: support HarmonyOS keyboard actions Route HarmonyOS enter, return, and dismiss through HDC key events. Expose supported keyboard actions through the system command metadata. Keep keyboard visibility inspection explicitly unsupported. * fix: reject unsupported HarmonyOS drag gestures Keep drag unavailable until HDC can preserve source and destination hold semantics. * feat: add HarmonyOS app log streaming 1. Stream HarmonyOS app logs through PID-scoped hilog sessions.\n2. Record HarmonyOS app identity during bundle-id opens for app-scoped commands.\n3. Cover backend routing and bundle identity resolution. * feat: report HarmonyOS foreground app state 1. Read the foreground HarmonyOS mission through aa dump.\n2. Expose HarmonyOS appstate with package and ability metadata.\n3. Add parser coverage for foreground and missing-state cases. * fix: advertise appstate through capabilities 1. Classify appstate in the command descriptor capability matrix.\n2. Surface supported appstate commands in capability inventory.\n3. Cover the advertised Android capability contract. * feat: sample HarmonyOS process performance 1. Sample HarmonyOS process CPU and resident memory through HDC.\n2. Expose the verified metrics through the shared perf command.\n3. Keep frame and memory snapshot collection explicitly unavailable. * feat: clear HarmonyOS app state 1. Add HarmonyOS settings clear-app-state through bundle cleanup.\n2. Force stop the app before clearing data and cache.\n3. Reject all unverified HarmonyOS settings explicitly. * docs: document HarmonyOS support 1. Describe HarmonyOS HDC prerequisites and HAP installation.\n2. Add HarmonyOS to platform discovery and product documentation.\n3. Document verified performance limits for the public HDC surface. * fix: preserve HarmonyOS deploy session identity 1. Bind a resolved HarmonyOS bundle after install or reinstall.\n2. Keep app-scoped logs and observability available after deployment.\n3. Cover session identity preservation for HarmonyOS reinstall. * test: lock HarmonyOS capability boundary 1. Add an independent HarmonyOS capability-matrix oracle and exact advertised-command regression test. 2. Document current HDC-backed support and evidence-based unsupported command boundaries. * refactor: simplify HarmonyOS shared platform boundaries 1. Split device selection and settings dispatch into focused helpers without changing behavior. 2. Keep HarmonyOS serial selection and lock-policy classification covered by regression tests. 3. Remove Fallow complexity findings from the HarmonyOS diff against upstream main. * fix: bound default HarmonyOS HDC commands 1. Apply a 15 second timeout to ordinary HDC operations. 2. Preserve operation-specific timeout budgets for installation and capture paths. 3. Add regression coverage for default and overridden HDC timeouts. * feat: add HarmonyOS screen recording Implement physical-device whole-screen recording through the system recorder and HDC media transfer. Reject unsupported HarmonyOS recording scopes and export flags. Cover capability routing, media retrieval, cleanup, and simulator rejection. * feat: report HarmonyOS HDC readiness Add an HDC version check to the HarmonyOS doctor flow. Document HarmonyOS as a supported doctor platform and cover the result. * refactor: simplify HarmonyOS recording checks Reduce recording validation and test complexity without changing behavior. * test: cover HarmonyOS platform contracts Synchronize public platform expectations across CLI, MCP, replay, and inventory tests. Mock HarmonyOS inventory probes to preserve concurrent test behavior. * test: model HarmonyOS recording capability Require a physical HarmonyOS device in the independent capability parity oracle. * test: cover HarmonyOS input and lifecycle paths Exercise HDC input, lifecycle, installation, and relaunch command sequences. * test: cover HarmonyOS device observability paths Exercise discovery, screenshot validation, and process performance sampling. * docs: define HarmonyOS CI hardware policy Keep HDC hardware validation local and require mocked CI contract tests. * fix: honor HarmonyOS app inventory filters * fix: bound HarmonyOS app inventory classification 1. 限制应用元数据分类并发并为默认清单设置整体时限. 2. 将请求取消信号传递给 HarmonyOS 应用清单读取. 3. 补充失败时中止在飞读取且不继续排队的回归测试. * fix: preserve HarmonyOS inventory failure causes 1. 保留触发应用元数据分类失败的原始错误, 避免被取消同级任务覆盖. 2. 补充总时限中止在飞读取且不启动排队任务的回归测试. 3. 验证后序任务失败时保留默认筛选的恢复提示. |
||
|
|
e18a183ac9 |
fix: harden artifact ingestion boundaries (#1692)
* fix: harden artifact ingestion boundaries * fix: bound archive inspection and upload expiry * fix: preserve upload preflight expiry |
||
|
|
6c0fcb64a1 |
fix: reject distinct ambiguous mutation targets (#1667)
* fix: reject distinct ambiguous mutation targets * fix(ios): scope the raw-match rejection to mutating dispatches `RunnerTests+Interaction.findElement` applied the new fail-closed classification to `querySelector` as well as press/type, because the read call site takes the default `allowNonHittableFallback: false`. With one visible/hittable match and one non-hittable same-selector duplicate the query started returning AMBIGUOUS_MATCH where it previously selected the hittable element, and `queryDirectIosSelectorOrFallback` preserves that error for read callers — so `get`, `is`, and `wait` surfaced an error instead of their prior answer. `classifyDirectSelectorCandidates` now takes a `rawMatchPolicy`. Mutations keep `.rejectDistinctMatches` (the default, so no mutation call site changes); `queryElement` passes `.preferHittableMatch`, restoring the prior read rule: prefer the single hittable match, ambiguous only when hittable matches compete, and never adopt the Maestro coordinate fallback. The Maestro expected-point path is untouched. Covers the one-hittable + one-non-hittable read, competing hittable reads, and the non-hittable-only read. ADR 0011's amendment now states the scope. * test(ios): execute selector read ambiguity regression --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
919700fd42 |
perf(cli): keep node:http and ICU off the CLI startup path (-54% warm call) (#1681)
* perf(cli): keep node:http and ICU off the CLI startup path (-54% warm call) A warm `appstate` cost 146ms; 79ms of that was startup work no local command needs. `node:http` was value-imported by four modules inside the eager import closure of `src/cli.ts`, so every invocation initialized undici -- and, under `NODE_USE_SYSTEM_CA=1`, the macOS trust store behind it. That is ~70ms of mostly off-CPU time (a blocking trustd lookup, invisible to --cpu-prof), paid by commands that never speak HTTP: the default daemon transport is a plain `node:net` socket. Cheap daemon-only commands paid it in full; only commands slow enough to absorb it, like `snapshot`, did not. Separately, `option-schema.ts` sorted its specs with `localeCompare` at module scope. The first `localeCompare` in a process costs ~5.4ms of ICU collator initialization (every later one is ~0.002ms), so the CLI paid ICU init purely to alphabetize an internal list. A case-insensitive comparator reproduces the exact same order for the ASCII flag-key vocabulary -- verified against all 159 keys. Both HTTP modules are now loaded on demand, via `.default` so they stay the same mutable module object a static import yields (the transport tests stub `http.request` on it). Behavior is unchanged: the HTTP transport, remote daemons and artifact upload/download all still work. Interleaved A/B, alternating prebuilt dists in one loop, 5-min load avg < 10, n=30 per arm: appstate 146.4ms -> 67.6ms -53.8% session list 137.9ms -> 57.7ms -58.2% snapshot -i 195.5ms -> 188.8ms -3.4% (device-work bound) Without NODE_USE_SYSTEM_CA the appstate win is 73.9ms -> 57.9ms (-22%), so this is not purely an artifact of that setting. Measured and rejected: splitting command metadata from runtime handlers (the facet restructure deferred by #1660). registry.js is 142kB of the 643kB eager closure but only ~1.6ms of module evaluation, and ALL module compilation is ~9.5ms -- the whole refactor could not clear the 20ms bar set for attempting it. A unit test walks the eager closure of src/cli.ts and fails if any file in it value-imports node:http/node:https again; type-only imports stay allowed. * test(cli): classify startup-closure imports with the ES module record The regex-based guard only recognized `import ... from` / `export ... from` forms, so a bare side-effect `import 'node:http'` anywhere in the eager closure would reintroduce the full undici + system-CA initialization cost while the guard stayed green -- the one form that binds nothing yet still evaluates the module. Classification now comes from oxc-parser's ES module record, the parser this repo already uses for import analysis in scripts/lib/shipped-imports. A side-effect import is exactly an entry-less static import, so it falls out of the record rather than needing another regex branch; type-only imports and re-exports are identified by their own `isType` flag instead of being pattern-matched. A fixture table now pins all eleven import forms -- side-effect, default, namespace, named, value re-export, star re-export, mixed value+type, type-only default/named/re-export, and dynamic -- so the classifier's contract is asserted directly rather than only through the closure walk. Verified red end to end: a side-effect import inserted into a real closure module fails the guard and names the file. * test(cli): follow workspace subpath imports in the startup-closure guard The walk stopped at the package boundary, so a `node:http` import inside any @agent-device/* module the CLI evaluates would have reintroduced the cost invisibly -- the same blind spot the module-record rewrite closed for side-effect imports, just relocated. Workspace specifiers now resolve through each package's exports map. The anti-vacuity assertion checks the closure actually contains package files, so a resolver that silently returned null for every workspace specifier would fail rather than leave the guard green over a src-only closure. Verified red both ways: a side-effect import inserted into packages/kernel/src/errors.ts, and one in a src/ module, each fail the guard and name the offending file. * test(cli): drop a redundant cast in the workspace exports lookup * test(cli): classify a top-level dynamic import as eager, and share one HTTP loader "On demand" is a claim about SCOPE, not syntax, and the guard only knew syntax. `import('node:http')` at MODULE TOP LEVEL runs during module evaluation -- so a top-level `await import(...)`, an `import(...).then(...)`, or an immediately-invoked top-level function reintroduced the whole undici plus trust-store cost while the guard stayed green. A dynamic import is lazy only when its nearest enclosing function scope is not the module itself. The guard now walks each closure file's OXC AST carrying an `eager` flag that drops on entry to a FunctionDeclaration/FunctionExpression/ArrowFunction body, and reports `import()` calls still reached with it set. Immediately-invoked top-level functions are detected by unwrapping the callee's ParenthesizedExpression, so `(async () => { ... })()` is eager too. The one shape that still reads as lazy -- a top-level function invoked indirectly at load time -- is stated as a known limitation in the test rather than left implied. Verified red by inserting each shape into a REACHABLE closure module (daemon-client-transport.ts): top-level await, `.then(...)`, and the top-level IIFE each fail the guard and name the file. The lazy case is proven by the suite itself rather than a fixture alone -- the shared loader below lives in the closure, so its function-local `await import('node:http')` would turn the guard red if scope tracking were wrong. The three lazy-load sites had drifted into three copies of the same ternary. They now share `loadNodeHttpRequester` in src/utils/node-http.ts, next to the other node:http helpers, which is where the reason for the indirection is documented once. This also clears the Fallow complexity finding the previous push introduced. Fixture table extended to 19 cases covering both axes: static form (side-effect, default, namespace, named, value/star re-export, mixed, type-only x3) and dynamic scope (top-level x3, function/arrow/method-local, and the exact ternary the shipped loader uses). |
||
|
|
4b89482a06 |
fix(daemon): open --foreground P1 hotfix — selector rejection, interactive snapshot, capture-failure masking (#1671)
* fix(daemon): open --foreground P1 hotfix — selector rejection, interactive snapshot, capture-failure masking Three P1s from the post-merge review of #1670, all at the session-open-foreground dispatch seam: 1. Explicit device selectors were silently overwritten: the resolved-device rewrite pinned --udid/--platform over whatever the caller passed, so `open --foreground --udid B` with sim A sole-booted silently opened A. Now fails fast with INVALID_ARGS (matching the existing app-positional rejection) on --udid/--device, and on --platform other than ios; an explicit --platform ios passes through. 2. The promised interactive snapshot was never requested: the composed snapshot dispatch forwarded the open request's flags untouched, without snapshotInteractiveOnly — so the capture was NOT the `snapshot -i` path the doc comment promised and returned no interactive presentation. The composed request now sets snapshotInteractiveOnly: true (the exact key the CLI maps -i to and the snapshot runtime reads as interactiveOnly). 3. A capture failure masked the successful open: returning the snapshot error discarded openResponse even though the session exists, so a retry of `open --foreground` failed with "close the current session first". Open success + snapshot failure now returns ok with an explicit initialSnapshotError {code, message} detail and a rendered warning that the session IS open and how to capture manually (snapshot -i). Regressions added for all three: explicit-selector rejection (udid/device/both/non-iOS platform + ios pass-through), the composed dispatch carrying snapshotInteractiveOnly, and the snapshot-failure path returning ok + warning + usable session. * fix(cli): render the composed open --foreground snapshot on default stdout and project initialSnapshotError through the public surfaces Post-merge review on #1671 found the daemon fixes never reached the public boundaries: openCliOutput ignored the nested snapshot (the one-call promise held only under --json), and initialSnapshotError was daemon-only — absent from AppOpenResult, Node normalization, and serializeOpenResult, with the normalized shape truncated to code+message. - default open output now renders the composed interactive tree through the same snapshotCliOutput path snapshot -i uses - AppOpenResult carries initialSnapshotError as the FULL daemon error (hint/details/diagnosticId/logPath preserved) through normalization and serialization Worker-authored; committed by the coordinating session after the worker stalled twice mid-push. Tests: output.test.ts + session-open-foreground (26 pass), typecheck, oxfmt. * refactor(client): one daemon-error normalizer + client-route regressions for initialSnapshotError Review follow-ups on #1671: normalizeInitialSnapshotError duplicated the target-shutdown error normalization and pushed the module over the fallow complexity threshold; consolidated into a single internal normalizeDaemonError (table-driven, full shape incl. retriable/supportedOn) used by both result paths, projected through normalizeOpenForegroundComposition. Client-route regressions: createAgentDeviceClient().apps.open now proves the full initialSnapshotError shape (hint/details/diagnosticId/logPath/retriable) survives normalization, and that a malformed one is dropped — deleting the boundary normalization fails these tests. Also rebased onto current main. * fix(daemon): a thrown initial-snapshot capture failure gets the same successful-open contract Review P1 on #1671: dispatchSnapshotViaRuntime rethrows ordinary capture/runner exceptions; the composition only handled a returned { ok: false }, so a thrown failure escaped to the router and failed the whole open after the session was created — retrying then wedged on the existing session. The catch normalizes the rejection (kernel normalizeError, same conversion the router applies) into the shared openWithInitialSnapshotFailure path: ok response, full-shape initialSnapshotError, session-usable warning. Rejecting-mock regression added alongside the returned-failure case. |
||
|
|
04c33f9e1e |
feat(ios): expose AX custom actions on merged accessibility elements (#1665)
* feat(ios): expose AX custom actions on merged accessibility elements
Apps that merge a card into one accessibility element for VoiceOver (React
Native's `accessible` prop) publish the card's real affordances as
UIAccessibilityCustomActions rather than as child elements. Our snapshot showed
only the merged node, so an agent looking for a feed card's options control had
nothing to aim at and fell back to coordinate guessing.
`snapshot --actions` now names them:
@e8 [link] "feedItem-by-whiskers.test" actions: ["Reply", "Repost", "Open post options menu"]
Opt-in, because the AX server cannot serve custom actions in a bulk tree
request: adding the attribute makes testmanagerd's reply decoder reject the
nested arrays a custom action serializes into, drop the reply, and time the
request out (~65s vs ~110ms). Only per-element reads answer, at ~100ms each, so
the runner reads at most 12 labelled childless nodes and stops at the
capture-plan deadline. The request pins the private-AX backend, since no other
backend can read the attribute, and reports that as its own `requested-backend`
verdict so a deliberate pin never renders as a degradation warning.
Invocation is not shipped: the actions are readable but not invocable from the
runner. RunnerAXSnapshotBridge.h records the five APIs that were tried.
* fix(ios): disclose a capped custom-action pass, and read on-screen elements first
Two gaps in the first cut.
An element the bounded pass never reached rendered identically to one with no
custom actions, so a capped capture silently taught the reader that later feed
cards have no affordances — the exact mis-inference this feature exists to
prevent. The runner already counted reads against candidates; it now carries
both to the response, and the verdict renders one response-level line when the
pass was incomplete. A complete pass stays silent, and "never asked" stays
distinguishable from "read none" (the key is absent, not (0, 0)).
The obvious remedy for a capped pass would be a scoped re-run, but scope is
applied when the Swift walk builds nodes, long after the read pass, so it does
not redirect the budget at all — measured: reads=12 candidates=18 with and
without --scope. Rather than print a remedy that does nothing, the read pass now
orders candidates on-screen first. That makes the budget land on elements an
agent can act on, and makes the disclosed remedy true: scrolling changes the
on-screen set, so a re-run reads elements the previous pass could not.
Also states plainly, in the flag help and the tool/SDK field description, that
the names are for planning: nothing invokes them, so the affordance is reached
through the element's detail screen, the same control exposed elsewhere, or
coordinates.
* test(snapshot): pin the custom-action coverage pair in the verdict shape assertion
* chore(scripts): classify the --actions flag in the integration progress model
The completeness gate flagged snapshotCustomActions as unclassified, which is
what it is for. It gets its own bucket rather than joining the provider-scenario
table: the values come from the private AX client inside the runner process, and
the fake runner derives its behavior from fixture tables that cannot fabricate
custom actions, so there is no provider-backed scenario to claim. The owning
coverage is named instead — runner XCTest unit, snapshot-lines, snapshot-quality.
* fix(ios): fail closed on unserviceable --actions, bound each read, cap output, and count actions in identity
Four review findings.
1. `--actions` with `--raw`, or on any target that is not an iOS simulator, used
to succeed and return nodes with no actions — a requested capability silently
no-opped, indistinguishable from "this screen has none". Both now fail closed.
The raw pairing is rejected at the shared request seam (INVALID_ARGS) so CLI,
Node client and MCP answer alike before any device work; the platform case is
rejected once the session device is resolved (UNSUPPORTED_OPERATION), naming the
resolved target. `diff --actions` was already rejected as an unsupported flag.
2. The per-element AX read had no timeout, so one wedged element could consume
the whole capture budget. Each read now runs off-thread behind a 1s wait. A
timed-out element counts as unread, never as "read, and it has no actions", so
the existing partial-pass disclosure already covers it.
3. The element budget bounded element count only; one element could still return
an unbounded list of unbounded names. Capped at 8 names of 80 characters, and
clipped elements are counted into the coverage so a truncated list is disclosed
rather than silently presented as complete.
4. Action names were rendered unescaped, and no comparison key read them. Names
now get the same escaping as text previews plus control-character folding, so an
app-authored name cannot split or corrupt a line. `actions` joins the diff
comparable key, the unchanged-comparison projection, and — the sharper bug —
the snapshot presentation key, without which `snapshot` followed by
`snapshot --actions` on a still screen answered "unchanged" and never delivered
the actions that were explicitly requested.
* fix(ios): contain a hung custom-action read instead of accumulating orphans
The 1s read deadline frees the caller, but the underlying AX call is a
synchronous XPC round trip that cannot be cancelled — it keeps running. On a
global concurrent queue that meant repeated `snapshot --actions` against a
wedged element piled up orphaned reads, all using the shared XCAXClient
concurrently. The deadline was containment for the capture, not for the runner.
Since the call cannot be cancelled, contain it instead:
- every read runs on one dedicated serial queue, so a wedged call can never be
joined by a second concurrent user of the shared client;
- a single-flight guard refuses to dispatch at all while an abandoned read is
still outstanding, so a repeated capture adds no work — the dispatch counter
stands still;
- the read pass stops at that point rather than paying a deadline per element
on reads that would all be refused, and reports `blocked` so the capture stays
honest. That gets its own line, because the partial-pass remedy (scroll and
re-run) cannot clear a hang and would send the reader in circles.
Recovery needs no reset: when the hung call finally returns, in-flight drops to
zero and reads resume.
The regression drives a fake AX client that never returns, and asserts the three
things the fix exists for — exactly one in-flight read with no further
dispatches across repeated captures, immediate returns with the skip disclosed
instead of the scroll remedy, and reads working again once the wedge clears.
* ci(ios): execute the custom-action runner regressions instead of only compiling them
The iOS workflow runs a targeted -only-testing list, so a runner test that is
not named there is compiled by the build step and then never executed. All seven
custom-action tests were in that gap — including the containment regression,
which is the only executable proof that a hung AX read cannot accumulate
orphaned in-flight reads.
Red/green against the containment regression, with the fix reverted to its
pre-fix concurrency behavior (global concurrent queue, no single-flight guard,
no blocked exit):
RED in-flight 6 (want 1), dispatches 6 (want 1), each repeat paid the full
1.004s deadline (want <0.2s), blocked=false (want true), and the
in-flight drain never completed — "Exceeded timeout of 5 seconds".
GREEN 7/7 pass, containment regression in 1.02s.
|
||
|
|
9b4239b2f3 |
fix: heal dangling XCTestDevices redirect and prove lock-owner death before reclaim (#1672)
* fix(ios): heal dangling XCTestDevices symlink during device-set reconcile fs.existsSync follows symlinks, so a redirect symlink whose target set was deleted read as absent: reconcile skipped the unlink and the backup rename failed with ENOTDIR, wedging every XCTest command on the machine until manual cleanup. Detect symlinks with lstat (throwIfNoEntry) in both the reconcile check and unlinkIfSymlink so the stale link is removed and the backup restored. * fix: prove owner death before reclaiming locks, and detect zombies A zombified lock owner passes kill(pid, 0) and still reports its original lstart, so a killed-but-unreaped daemon held the XCTest device-set lock forever. Conversely a ps read lost to CPU contention condemned a live owner and let waiters steal a held lock (the flake pinOwnProcessStartTime papers over in tests). Unify both liveness surfaces on classifyOwnerLiveness: condemn zombies via ps state, treat failed ps reads as no-proof rather than death, and reclaim null-start-time owners whose pid provably started after the lock was acquired (ps etime bound). Lock timeout errors now report ownerLiveness. * fix: keep null-start-time lock owners fail-closed Review P1 on #1672: the started-after-acquisition bound compared a wall-clock birth estimate (Date.now() - ps etime) against the persisted wall-clock acquiredAtMs, so a clock step larger than the fixed 30s slack could condemn a live null-start owner and let a waiter steal the held lock — the failure class this PR eliminates. There is no same-clock-domain proof of birth order to be had here (macOS ps etime is itself wall-derived), so drop the bound: an alive pid with no recorded start time is never reclaimed. Regression covers the steal at both the classifier and the lock level. |
||
|
|
cc943400a9 |
feat(daemon): [RFC] prototype foreground-attach convenience (#1670)
Prototype `open --foreground`: on a fresh session with no app argument, auto-resolves the target from the sole booted iOS simulator's sole foreground app (reusing the exact same ambiguity-detection probe that enriches the SESSION_NOT_FOUND hint), then attaches the initial interactive snapshot to the response by composing the existing snapshot-runtime dispatch. Collapses the documented 3-call snapshot-fails -> read-hint -> open -> snapshot-succeeds dance into a single call for the unambiguous case, while failing closed (AMBIGUOUS_MATCH) with no guessing otherwise. First-pass RFC, not reviewed — see PR body for the design tradeoff writeup, live before/after evidence, and scoped-out follow-ups. |
||
|
|
8030fc10d6 |
fix: make Android record stop survive static screens, slow moov finalization, and dead-pid recovery (#1651)
* fix: make Android record stop survive static screens, slow moov finalization, and dead-pid recovery Three live-reproduced defects on a loaded Pixel_7_CI emulator shared the "pulled file is not a playable MP4" / "manifest could not be verified" symptom family: - The Swift video validator required duration > 0, permanently rejecting valid single-frame recordings of fully static screens (AVFoundation reports their duration as 0). Whether a run passed depended on whether anything — even the status-bar clock — changed during the window. - screenrecord finalizes by patching a front-reserved moov in place, so the remote file size never changes; the copy path now re-pulls with escalating delays (750/1500/3000ms) and detects finalization from the pulled bytes via a mandatory ftyp+moov container sniff, which also preserves the truncation detection the duration check provided by accident. The local waitForStableFile call is gone: a pull is complete when adb returns. - toybox `ps -p <missing-pid>` exits 1 with empty output — the normal pid-gone signature — but the recovery liveness probe read every non-zero exit as an uncertain adb failure, making a finished recording behind a live-status manifest unrecoverable forever. Empty-output failures now corroborate via the full process list: healthy listing with the pid absent recovers the finished recording; listing failure stays conservatively uncertain (a stale verdict deletes the manifest, so transport health is proven first). Transport failures keep stderr and exec-layer timeouts throw, which is what makes the empty-output signature safe to trust. Provider-scenario coverage: in-place same-size finalization landing past the first retry, and dead-pid recovery on a responsive device — both fail on the previous implementation with the live-observed errors. * refactor: extract Android liveness probe and split recording scenarios per file-shape rule record-trace-android-recovery.ts (763 lines) and android-recording.test.ts (1,764 lines) were both past the 500-LOC extraction tripwire. The screenrecord liveness/probe concept this PR modified now lives in record-trace-android-liveness.ts, and the new scenarios moved to test files mirroring the source modules they cover (record-trace-android-copy.test.ts, record-trace-android-liveness.test.ts) with shared scenario plumbing in the android-recording-fixtures.ts sibling. android-recording.test.ts shrinks to 1,476 lines — below its pre-PR size. |
||
|
|
d5f11f6e2f |
refactor: import package types directly — no internal re-export laundering (#1640)
* refactor: import package types directly instead of re-exporting from internal modules
Post-#1636 review feedback: internal src modules were re-exporting package
types (export type { X } from '@agent-device/...'), giving one declaration
several import paths and hiding its provenance. New rule applied repo-wide:
internal modules import directly from the owning package; only published
entry surfaces (src/sdk/* entries, client-types, finders, metro composition,
remote-config-schema) may re-export.
Eleven internal re-exports removed and ~110 import sites redirected to the
packages, the big two being CommandFlags (core/dispatch chain, 39 sites) and
SessionAction (daemon/types.ts, 19 sites). Two were already dead
(RefFrameEffect via daemon-command-registry, DiffSnapshotCommandResult via
capture/runtime/snapshot). Entry-surface chains now re-export from the
package rather than laundering through a second internal module
(client-types/client-metro MetroBridgeScope).
Side effect: the R9 type cycle shrinks again, 49 -> 47 (daemon-server
19 -> 17); ceilings lowered to match.
* refactor: drop command-schema's CliFlags re-export (#1640 review P2)
The one consumer (cli/parser/args.ts, a multi-line import the sweep's
single-line scan missed) now imports CliFlags from contracts/command;
FlagDefinition/FlagKey stay — they are src-declared types, not package
laundering.
|
||
|
|
3e4828d68d |
feat: add scale-only screenshot sizing (#1617)
* feat: add scale-only screenshot sizing
* fix: refuse retired --max-size inputs on every released surface
Released sizing inputs must fail closed with migration guidance instead of
silently producing native-size artifacts:
- contracts: RETIRED_SCREENSHOT_MAX_SIZE declaration + SCREENSHOT_SCALE_LIMITS
as the single source for the scale bounds and migration messages
- .ad parser: released 'screenshot ... --max-size N' and 'record start ...
--max-size N' lines now refuse at parse time (frozen replay-compat witnesses)
- daemon: screenshot rejects old-client screenshotMaxSize like recording does;
the recording guard now shares the same contract data
- Node client: screenshot/record daemon writers refuse the removed { maxSize }
option before transport
- CLI: --max-size unknown-flag error carries the migration guidance
- config/env: stale screenshotMaxSize config keys and the retired
AGENT_DEVICE_SCREENSHOT_MAX_SIZE env var are refused for sizing commands
(other commands keep working)
Quality: numberField now reuses the canonical readOptionalNumber contract
helper (AppError bounds instead of plain Error); png-resize inlines one-use
wrappers and restores the worker-thread rationale; docs typo fixed.
* test: drop retired maxSize entries from the MCP undocumented-input allowlist
* fix: refuse retired maxSize at the MCP field-projection seam + release-provenance corpus witnesses
- readFieldInput silently dropped undeclared keys before the daemon writers
could refuse them, so an MCP call carrying { maxSize } reached transport and
returned native-size success. New retiredField() combinator declares the
removed key in the field map: the projection seam refuses it with the
canonical migration message and the JSON schema no longer advertises it.
Real-route MCP executor regressions cover screenshot and record.
- replay-compat corpus: derived v0.20.5 witnesses for the released screenshot
and record --max-size forms (SHA-256 pinned, new retired-capture-size
coverage surface) so check:replay-compat proves the shipped syntax refuses
with migration guidance instead of degrading silently.
---------
Co-authored-by: Michał Pierzchała <thymikee@gmail.com>
|
||
|
|
d5f99bab1c |
refactor: sink backend.ts's cycle-closing types below both zones (#1632) (#1636)
backend.ts imported RepeatedInput up from commands/command-input.ts and ScreenshotResultData up from utils/screenshot-result.ts — the interface hub typed in terms of the zones that depend on it, R6's textbook inversion shape. - RepeatedInput now lives in @agent-device/contracts/interaction; command-input.ts re-exports it for its existing importers. - ScreenshotResultData already had a byte-identical canonical declaration in contracts/snapshot-types.ts (exported via contracts/capture); the utils copy is now a re-export of it, deleting the duplicate outright. Measured member-by-member: the R9 type cycle collapses 76 -> 49 files. backend.ts, runtime-contract.ts, commands/runtime-types.ts, and commands/runtime-common.ts all leave the component (27 files stranded out at once); zone ceilings lowered to the measured values (commands 33 -> 14, platforms 7 -> 2, root 5 -> 3, daemon-server 20 -> 19) and CONTEXT.md's hub list recomputed (core/dispatch.ts 8, command-catalog.ts 7, resolution.ts 6, command-descriptor/registry.ts 6). No TYPE_INVERSION_BASELINE additions. |
||
|
|
ee473b6adc |
refactor(daemon): give the Maestro fallback and ambiguous-match details real types (#1612)
Three places smuggled structured data through untyped bags and re-read it
with runtime guards. Each gets an explicit typed boundary.
A. The resolution-suppression rule was encoded twice in
interaction-touch-response.ts — a spread ternary in the runner-payload
branch and an unconditional destructure used conditionally in the runtime
branch, with the ADR 0012 rationale living on only one source variant.
Both branches now read one `suppressesResolutionDisclosure(source)`
predicate through one `applyResolutionDisclosurePolicy` helper, where the
reason is stated once. The union field is renamed
`maestroCoordinateFallbackDispatched` (the dispatch path that ran) and
hoisted into a shared base. handleFillCommand's two-arm interactor.fill
call collapses to one.
B. `Interactor.type` narrows from `Record<string, unknown> | void` to
`TypeTextBackendResult | void`; the Apple runner boundary is the single
place the wire payload becomes that type. `maestroFallbackDetails` returns
a typed `{ used, extra }` instead of a bag both call sites re-read.
C. `details.candidates` meant two incompatible things. The device-domain
resolvers now key their list `devices`, so the shared renderer drops its
shape-disambiguation guards and the device list actually renders.
|
||
|
|
23a3016e9b |
fix: stop the unit suite from leaking temp directories (#1593)
* fix: stop the unit suite from leaking temp directories ~650 test call sites across the unit suite create scratch directories via fs.mkdtemp(path.join(os.tmpdir(), ...)) or shared factories (makeSessionStore) with no cleanup, ever. Over time this accumulated 1.16M+ orphaned directories in the real system tmpdir, slow enough to make tools that enumerate $TMPDIR at startup (e.g. opencode) take 1-2 minutes to launch. Rather than migrate every call site, redirect os.tmpdir() itself for the lifetime of the whole `vitest run` invocation: scripts/vitest-tmpdir-global-setup.ts wires in as vitest's globalSetup/globalTeardown, points TMPDIR at one /tmp-rooted directory (verified: env mutations here propagate to every forked worker, confirmed empirically), and removes it in one recursive rm after every worker across every project finishes. Since os.tmpdir() reads TMPDIR on every call, this covers all ~650 call sites without touching any of them. Rooted at /tmp rather than nested inside the current (already deep, on macOS) os.tmpdir(): that broke real AF_UNIX socket tests (runner-usbmux.test.ts) by pushing socket paths past the 104-byte sun_path limit. A per-file afterAll hook was tried first but proved unreliable — 5 of 7 workers in one run never ran it before their process was torn down; the global setup/teardown pair (one process, confirmed single execution) is the mechanism that's actually guaranteed to run once. Also adds: - scripts/check-tmpdir-leaks.ts: CI/local guard asserting no agent-device-test-run-* directory survives a run (a leftover one means a worker was killed before cleanup could run). - src/__tests__/test-utils/tmp-dir.ts: documented mkdtempForTest / mkdtempForTestSync helpers, the discoverable way to get a scratch dir going forward (mirrors the src/utils/exec.ts pattern for node:child_process). - scripts/check-test-tmpdir-helper.ts: ratchet guard capping raw fs.mkdtemp/mkdtempSync call sites in test files at today's count (632); it can only shrink as call sites migrate to the helper. * fix: make check-tmpdir-leaks scan the same root the fix actually uses check-tmpdir-leaks.ts was scanning os.tmpdir() for leftover run directories, but vitest-tmpdir-global-setup.ts creates them under a hard-coded /tmp. On macOS those are different paths (TMPDIR is a deep per-user /var/folders/.../T/ directory) — the guard could never find a leak on the exact platform the original leak happened on, only on Linux CI where os.tmpdir() already is /tmp. Export TEST_RUN_TMP_ROOT and TEST_RUN_TMP_PREFIX from the global-setup module and import them in the leak check instead of recomputing a path that can drift. Switched from a fixed pid-based directory name to fs.mkdtempSync so a same-named leftover from a prior killed run (or, on a shared machine, another user) can't collide with a live run. Also fixes two stale comments (in this file and ci.yml) that still described the per-file afterAll hook design that was abandoned in favor of the global setup/teardown pair, and notes the check only covers vitest runs, not the node --test lanes (test:smoke, test:integration:node). Verified live: with the old code the guard reported no leaks even with a real orphaned /tmp/agent-device-test-run-* directory present (left by a command that got killed mid-run); with this fix it correctly found and reported it. * simplify: drop the tmpdir ratchet guard, keep the leak check local-only Two guards were more than this needed: - check:test-tmpdir-helper (ratchet on raw fs.mkdtemp call counts) protects nothing a bug could actually trigger — the leak is already fixed architecturally regardless of call-site count, so this was pure style/discoverability nudging. Dropped the script and its check:tooling/CI wiring; kept mkdtempForTest/mkdtempForTestSync in tmp-dir.ts as the documented option without enforcing it. - check:tmpdir-leaks in CI added little: GitHub-hosted runners are destroyed after each job, so a leftover directory there is harmless by construction, and a worker getting killed mid-run would already surface as a job failure some other way. Its real value is local, on the long-lived dev machines where the original leak actually accumulated — kept it wired into check:unit, dropped the CI step. * fix: don't flag a concurrent vitest run's tmpdir as a leak check-tmpdir-leaks.ts reported every agent-device-test-run-* directory as a leak, but a concurrent vitest run in another worktree legitimately keeps its own directory present until its own teardown finishes. On a machine that regularly runs several worktrees at once, that made check:unit fail on unrelated in-progress work. Embed the owning process's pid in the directory name (still random- suffixed via mkdtempSync, so same-pid reuse across separate runs can't collide) and have the leak check skip any directory whose pid is still alive (process.kill(pid, 0)) — only directories whose owning process already exited without running its globalTeardown are real leaks. Split the pure logic into check-tmpdir-leaks-model.ts (findLeakedRunDirectories, with an injectable liveness check for testing) so it has a real regression suite, including the concurrent-run case, instead of only being exercised by hand. * refactor: migrate raw fs.mkdtemp call sites to mkdtempForTest(Sync) Migrates 629 raw fs.mkdtemp(Sync)(path.join(os.tmpdir(), PREFIX)) call sites across 168 test files to the mkdtempForTest / mkdtempForTestSync helpers (src/__tests__/test-utils/tmp-dir.ts), so there's one documented, discoverable way to get a scratch dir in a test — cleanup already didn't depend on the call-site shape (the global TMPDIR redirect covers any of them), this is purely for consistency and discoverability, same reasoning as src/utils/exec.ts for node:child_process. Existing manual per-test cleanup (fs.rm in finally/afterEach/onTestFinished blocks) is untouched — the global teardown is a fallback for killed workers, not a replacement for tests cleaning up after themselves. Migrated with a one-off AST-based codemod (oxc-parser, since regex mismatched multi-line calls and complex prefix expressions like `options?.tempPrefix ?? 'default-'`) rather than by hand across 168 files. The codemod isn't included — it doesn't need to survive this commit. Caught and fixed one real bug in it during review: a small number of files declare a second import statement later in the file, after some of the matched call sites, which broke a naive "insert after the textually-last ImportDeclaration" placement; fixed to insert after the top contiguous import block instead, plus a self-check that re-parses every generated file before writing it. Also fixes 5 fallow dead-code findings the branch introduced: the vitest globalSetup functions (setup/teardown) are only referenced by the config-string path vitest.config.ts hands to globalSetup, invisible to static analysis — suppressed with the documented convention. isProcessAlive didn't need to be exported (nothing outside the module uses it). And dropped a barrel re-export of the new helpers from test-utils/index.ts: nothing actually imports through the barrel (matching the existing makeSessionStore convention, which is imported directly from store-factory.ts everywhere despite also being barrel-exported), so the re-export was genuinely dead. Documents the convention in docs/agents/testing.md. Verified: full unit suite (5308 tests) passes except the one pre-existing, unrelated package-exports.test.ts failure; typecheck, lint, and format all clean; fallow audit clean against the PR base. * fix: correct fallow suppression token and drop unused barrel re-export These were meant to be part of 397cdd6d7 (verified locally before that commit) but didn't actually get staged — caught by CI's Fallow Code Quality check re-running against the pushed commit, which still had the plural 'unused-exports' token (fallow expects singular 'unused-export') and the dead barrel re-export. * fix: migrate the two mkdtemp call sites the rebase silently reintroduced Rebasing onto main pulled in #1594's two new test cases in this file, added independently of this branch's migration, still using raw fs.mkdtempSync(path.join(os.tmpdir(), ...)). Git's line-based merge found no textual conflict with this branch's removal of the os import (the changes touch non-overlapping regions), so it silently produced a file that doesn't typecheck. Migrated both to mkdtempForTestSync for consistency with the rest of the file, caught by CI's Typecheck, Fallow Code Quality, and FreeRange checks re-running against the pushed commit. * test: pin Vitest tmpdir lifecycle * fix: preserve the Swift cache across test runs |
||
|
|
611858103e |
fix(ios): harden Bluesky-class interaction reliability (#1588)
* fix: type into focused iOS inputs without AX * fix: fill AX-hostile iOS text inputs * fix: keep scrolling containers from stealing taps * fix: stop agents after explicit task success * chore: format benchmark guidance * fix(ios): preserve fill semantics across fast paths * test: retire direct selector fill expectations * test: assert runtime selector fill evidence * fix(ios): preserve verified and Maestro fill paths * refactor(ios): isolate synthesized text entry * fix(client): preserve open diagnostic paths * fix(ios): expose structured text entry route * fix(packaging): strip text entry policy tests |
||
|
|
65450dd3a7 |
fix: flush stdout/stderr before every CLI process.exit() (#1596) (#1603)
* fix: flush stdout/stderr before every CLI process.exit() (#1596) Node only flushes process.stdout/stderr synchronously to a file or TTY; on a pipe (the normal condition for this CLI when driven as a subprocess) a write queued right before process.exit() can be silently dropped. handleRunCliFailure's --debug daemon-log-tail dump made this reachable from the exact path that renders a SESSION_NOT_FOUND error right after a daemon replace, matching field reports of the driving process going silent immediately after "Replacing daemon ... unreachable" plus the SESSION_NOT_FOUND error. Add exitAfterFlush() and route every process.exit() in src/cli.ts and src/bin.ts through it, so a piped caller always receives the full structured error (with its "run open first" hint) before the process terminates. Also bound the --debug log-tail dump to a byte cap instead of an unbounded 200 lines. Verified directly against the real CLI (piped subprocess, pre-fix vs post-fix): a live daemon with a seeded >64KB log truncates its --debug error output before the fix and delivers it in full after. * fix: satisfy CI gates on #1596 (format, fallow, coverage) - oxfmt formatting on the new integration test file. - Register the two exit-flush regression fixtures (support/exit-naive.ts, support/exit-after-flush.ts) as fallow entry points: they're run as real subprocesses via a string path (runCmdSync), which fallow's static dependency analysis can't follow, same as the existing test/contention-retry-fixtures/* entries. exit-payload.ts becomes reachable transitively through their static imports. Also switched the integration test's local PAYLOAD_MARKER duplicate to import the one fallow flagged as unused from exit-payload.ts. - Added real unit coverage for the new exitAfterFlush code paths, since node --test integration files aren't measured by the vitest coverage gate: src/utils/__tests__/process-exit.test.ts exercises the already-drained, backlogged-then-drains, and never-drains/timeout branches directly against a fake stream; src/__tests__/cli-exit-paths.test.ts drives runCli() for --version, bare help, no-command, and web to cover their exitAfterFlush call sites, plus a --debug case with a >64KB seeded daemon.log proving printDaemonLogTailOnError's new byte cap actually trims the oldest lines. Changed-line coverage gate now passes at 92.59% (was 59.26%); the two remaining uncovered lines are the bottom-of-file `isDirectRun` catch handler, which only runs when cli.ts is executed as the literal entry script and is not reachable by importing it as a module in a test (the same shape as bin.ts's already-excluded top-level fast paths). |
||
|
|
8ba5f9b8de |
fix: surface AMBIGUOUS_MATCH candidates and name find's supported actions (#1602)
* fix: surface AMBIGUOUS_MATCH candidates and name find's supported actions (#1597) AMBIGUOUS_MATCH errors now list the matching candidates (ref, role, label/identifier) rendered the same way as snapshot -i lines, capped at 5 with a "+N more" marker. buildAmbiguousMatchError (the single producer, src/daemon/handlers/find.ts) reuses formatSnapshotLine to build the list; formatAmbiguousMatchCandidateLines (src/utils/output.ts) renders it unconditionally on both text surfaces an agent actually reads (CLI printHumanError and MCP formatToolErrorText) — previously the candidates lived only in details, which neither surface printed. find's "Unsupported find action: X" (e.g. from `find <text> press`) now attaches a hint naming every action find actually supports and the two-step recovery shape: run find "<text>" to resolve the ref, then dispatch the gesture as its own command (press @eNN). The hint is a single exported constant (UNSUPPORTED_FIND_ACTION_HINT) shared by both throw sites — packages/selectors' raw-token parser and the CLI's typed reader (src/commands/interaction/selectors.ts) — so they can't drift. Matching semantics are unchanged; ambiguous rejection stays by-design. The help-conformance corpus's AMBIGUOUS_MATCH quiz is updated: its premise ("candidate refs were not shown") no longer holds, but with 3 identically-labeled candidates the lesson (don't guess a specific ref) still holds. * fix: guard the AMBIGUOUS_MATCH candidate renderer against device-domain shapes Review on #1602 (P2): formatAmbiguousMatchCandidateLines ran for every normalized error and stringified details.candidates unconditionally, but device-domain AMBIGUOUS_MATCH/APP_NOT_INSTALLED errors (findBootedAppleSimulatorWithApp, src/core/dispatch-resolve.ts) reuse that key for { id, name } device objects with no `matches` field — CLI and MCP would have printed "Candidates: [object Object]" for those. The renderer now requires numeric details.matches AND every candidate to be a string before rendering anything, restricting it to buildAmbiguousMatchError's element-match shape; unrecognized shapes render nothing, same as before this feature existed. Added regression tests against the exact device-error shape on both text surfaces. Also unexports AMBIGUOUS_MATCH_CANDIDATE_LIMIT (fallow flagged it as an unused production export) — it has no consumer outside find.ts. |
||
|
|
8d526f2402 |
perf(ios): halve hostile-screen capture cost under the XCTest-channel penalty (#1587)
* perf(ios): stop re-paying known-failing work on penalized private-AX captures Live-measured on the Bluesky bench feed (139-171 nodes), each private-AX capture wasted ~1.35s of its ~1.65s runner-side cost re-doing work a prior capture already proved futile: - ~1000ms: the viewport read is XCTest main-thread work; under the channel penalty it reliably burns its full timeout and falls back to the root frame anyway. Honor the penalty in privateAXSnapshotViewport the same way capture plans do. - ~310ms: the depth ladder re-paid the kAXErrorIllegalArgument rejection of the default depth on every capture. Remember the accepted rung per bundle (penalty-duration TTL, cleared on target process change); expiry re-probes the full depth so screens that recover are not capped forever. Explicit --depth requests bypass the memory in both directions. Steady-state hostile-screen captures drop 2.65s -> ~0.55s CLI wall, and press --settle round-trips drop ~5.4s -> ~2s (settle needs two captures). * fix(snapshot): distinguish penalty-deferred captures from genuine recoveries A capture whose backend was PRE-selected by the XCTest-channel penalty was stamped with the same recovered verdict as one that ground through a live failure. Two costs followed on hostile screens (Bluesky bench: 130 repeats per 30-task run): - the daemon repeated the full fell-back warning on every capture, long after the arming capture had already said it once, and - settle's one-shot private-AX budget reset fired on every loop even though the capture paid no grind to give the budget back for. The deferred plan now stamps reasonCode 'deferred' (a new code; older daemons drop unknown codes and keep today's behavior). The daemon keeps the verdict recovered but suppresses the repeated warning, keeps the depth-cap line, and skips the settle budget reset for deferred captures. * style: wrap reasonCode union to satisfy oxfmt * fix(ios): bind accepted-depth memory to the process, pin deferred through the wire parser Review follow-up (#1587): - The depth memory was cleared only in refreshCachedTargetIfProcessChanged; resetTargetAfterExternalRelaunch -> invalidateCachedTarget drops the cached PID without clearing it, so the next activation had no old PID to compare and could reuse a stale shallow rung for up to 120s (also A->B->A when A restarted while inactive). The memory now stores the PID it was learned under and only matches the same live process; recording without a PID is refused. Every invalidation path is covered automatically because they all drop currentAppProcessIdentifier. - The deferred settle/warning tests constructed typed verdicts directly, so removing 'deferred' from the accepted reason-code set would silently restore the repeated warning and budget reset while tests stayed green. They now parse a raw runner-wire object through readSnapshotQualityVerdict (red on base: parser strips the code -> warning re-appears, reset fires). - Extracted shouldReadPrivateAXViewportViaXCTest() and pinned the penalized viewport skip with an in-bundle regression test. |
||
|
|
bcaa106845 |
refactor: extract snapshot and replay identity semantics (#1582)
* refactor: extract snapshot and replay identity semantics * refactor: move pure rect primitives from contracts to kernel/rect containsPoint, pickLargestRect, and isRectVisibleInViewport are raw rectangle arithmetic with no snapshot awareness, so they belong beside rectContains/rectArea in @agent-device/kernel/rect rather than in the snapshot-semantics vocabulary. The node-aware resolveViewportRect folds into contracts/snapshot-visibility.ts, retiring snapshot-geometry.ts; after this split, everything behavioral in @agent-device/contracts/snapshot is policy that interprets the snapshot model. * refactor: restore ADR-0012 rationale docs and dedupe replay identity shapes The #1478/#1581 extraction moved the identity/structural helpers but compressed their invariant documentation to one-liners; the deliberate no-ancestry-exclusion rule on idMatchCountInTree, the fail-closed guard comparison, and the who-throws/who-detects contracts on the two divergence reason markers now travel with their definitions again. LocalIdentity and NodeStructuralDenotation move to contracts/target-annotation.ts (beside TargetAncestryEntry, which WaitLandmarkMismatchEvidence now references directly), so the guard shapes in contracts/replay.ts are nominal instead of hand-rolled structural twins; ad-script re-exports the vocabulary beside the readers that produce it. Also inlines the demoteNonUniqueId pass-through wrapper in session-target-evidence.ts. * style: fix oxfmt formatting in snapshot-visibility * refactor: move findSnapshotAncestor into contracts/snapshot-tree The last root value import from src/selectors: predicates.ts reached src/snapshot/snapshot-processing.ts for the ancestor walker. The walker is index-based tree traversal with no presentation policy, so it joins buildSnapshotNodeMap in contracts/snapshot-tree.ts; both consumers repoint to the façade and the non-contiguous-index/cycle coverage moves to the package test. src/selectors now has zero value imports from root src in production files. |
||
|
|
761317deb7 |
refactor(daemon): extract native .ad replay to packages/ad-replay (#1478 P5) (#1555)
* refactor(replay): move the dependency-free engine leaves into packages/ad-replay Stage A of the #1478 P5 extraction: vars, plan-digest (+canonical-json, sole consumer), the target-identity classification core, report-action, and suggestion-ranking move verbatim; imports updated. The package facade temporarily re-exports the moved symbols so root consumers keep compiling; a later stage narrows it to inspectAdReplay/runAdReplay only. * chore(layering): register packages/ad-replay in the workspace and DAG * refactor(replay): define the three-operation replay selector port with dual adapters (#1478 P5) * refactor(daemon): route replay handlers through the selector port (#1478 P5) * refactor(replay): split target verification into engine policy and daemon authority (#1478 P5) * refactor(replay): move the .ad step loop behind inspectAdReplay/runAdReplay (#1478 P5) * refactor(replay): lock the ad-replay façade to its real consumers (#1478 P5) * test(replay): prove shared-id demotion on both selector-port adapters (#1555 review) * fix(replay): restore invalid replayBackend rejection on the native path (#1555 review) * refactor(replay): move shared .ad vocabulary to its owner, packages/ad-script (#1555 review) * refactor(replay): neutral step/run outcomes and digest/resume behind inspectAdReplay (#1555 review) P1 "do not smuggle daemon wire failures through a generic": drop the TResponse generic from AdReplayStepRuntime/runAdReplay. executeStep and handleActionFailure now return neutral tagged AdReplayStepOutcome/ AdReplayStepFailure values (kind/message/artifactPaths only); runAdReplay returns a neutral completed/failed AdReplayRunOutcome. The engine never holds or returns a DaemonResponse. The daemon adapter (createAdReplayStepRuntime, session-replay-runtime.ts) keeps its real wire response in a local side-map as it builds each neutral outcome, and runReplayScriptFile reads it back once runAdReplay reports which step failed, so the final response is byte-identical to before this split. P1 "parsing/planning/digest/resume must also occur behind runAdReplay": relocate computeReplayPlanDigest's call site and the --from/--plan-digest resume-point math (resolveReplayEntryIndex) behind inspectAdReplay's manifest as planDigest and a resolveEntryIndex closure. Neither is a new top-level export -- inspectAdReplay/runAdReplay stay the only two. Timing is preserved exactly (still called eagerly in prepareReplayPlan, before prepareReplaySession's coordinator-mutating side effects) since moving resume validation to run inside runAdReplay itself would let a rejected --from request mutate coordinator/session state first -- a real ordering hazard, not just a cosmetic one. computeReplayPlanDigest/ReplayPlanDigestMetadata/resolveReplayEntryIndex leave the ad-replay façade; request-router-repair-expired.test.ts and prepareReplayPlan read the digest/resume result off the manifest instead. * refactor(replay): relocate classifyTargetBindingMatch and pin the ad-replay façade (#1555 review) P1 "complete the binding façade instead of documenting deviations": classifyTargetBindingMatch never had a real consumer reachable through inspectAdReplay/runAdReplay -- both its callers (the daemon's record-time self-check in session-target-evidence.ts and its replay-time classification wrapper in session-replay-target-classification.ts) are daemon files that imported it directly. It interprets TargetAnnotationV1 evidence semantics shared beyond the engine, so it moves to packages/ad-script alongside target-annotation-identity.ts (new target-annotation-classification.ts + its test), and both daemon call sites now import it from there instead of @agent-device/ad-replay. One deviation remains and is reported rather than papered over per the review's own instruction: the four target-verification policy functions (planPreDispatchTargetVerification, planPostResolutionTargetVerification, deriveReplayTargetGuardMismatchEvidence, deriveWaitLandmarkMismatchEvidence) and the ReplaySelectorPort type family stay exported. Their sole caller, session-replay-target-verification.ts, interleaves these pure decisions with daemon-only async work (capture, SessionStore, coordinator/resume stamping, wire shaping) that must stay outside the engine by design; moving their call sites to live only behind runAdReplay would require restructuring that whole orchestration into new fine-grained AdReplayStepRuntime capabilities, which is out of scope for this pass. See packages/ad-replay/src/index.ts's header comment for the full reasoning. P1 "add the reviewer-required exact exported-symbol gate": adds readNamedExports (scripts/layering/package-boundaries.ts), a small parser over a façade's `export { .. } from`, `export type { .. } from`, and direct-declaration forms, and pins @agent-device/ad-replay's exact 21-symbol export list in package-boundaries.test.ts. Plant-verified: a stray `export const` addition failed the assertion; removed it and the gate went green again. * refactor(replay): drive target verification from the engine step loop (#1555 review) Moves the verify-then-dispatch decision flow into packages/ad-replay's step loop so the four target-verification policy functions (plan{PostResolution,PreDispatch}TargetVerification, derive{ReplayTargetGuardMismatch,WaitLandmark}MismatchEvidence) become engine-private and leave the ad-replay façade. The daemon (session-replay-target-verification.ts) shrinks to the narrow AdReplayStepRuntime capabilities the engine drives: routing (beginTargetVerification), capture (captureObservation), classification (classifyTarget), dispatch (dispatchStep), and wire-building (buildRecordedUnverifiableFailure, buildTargetBindingFailure, buildPostDispatchTargetBindingFailure). Wire output and replay-compat stay byte-identical; the exact-symbol façade gate is updated to the shrunken export list. * refactor(daemon): decompose the replay adapter's two over-threshold functions (#1555) * refactor(replay): fold #1554's keep-session terminal-lifecycle policy into the ad-replay engine Rebasing p5/extract-ad-replay onto main pulled in #1554's --keep-session feature, which had grown its own daemon-side terminal-close-suppression predicate (session-replay-terminal-lifecycle.ts's resolveSuppressedTerminalCloseIndex/countExecutedReplayActions) independently of this branch's own engine-side one (step-loop.ts's isRepairArmedTerminalCloseAction). Both are the same decision family — replay --keep-session and an active --save-script repair now share ONE structural resolution (resolveSuppressedTerminalCloseIndex, generalized to "terminal among EXECUTABLE actions" rather than the old physical-last-index check) and one suppression check inside runAdReplay, gated on keepSession OR runtime.isRepairArmed(). AdReplayRunRequest grew a keepSession field; the neutral 'replayed' count in AdReplayRunOutcome is now computed inline in the loop instead of the daemon's old actions.length - entryIndex approximation. requireLiveSessionForKeepSession (the --keep-session live-session postcondition) stays daemon-side, inlined into session-replay-runtime.ts, since it inspects SessionStore state the engine never sees. The daemon-only session-replay-terminal-lifecycle.ts this arrived with is deleted entirely — its isExecutableReplayAction was a duplicate of the engine's own. runReplayScriptFile's Maestro-format routing (including the new --keep-session Maestro rejection) was extracted into routeMaestroReplay to keep the function under fallow's complexity threshold after re-threading keepSession through it. Added packages/ad-replay/src/internal/__tests__/step-loop.test.ts covering the unified suppression decision (both keepSession and repair-armed) directly against runAdReplay, including the terminal-among-executable-actions case with a trailing nested replay marker. The daemon-level integration tests (6 tests in session-replay-terminal-lifecycle.test.ts, exercising the same behavior through runReplayScriptFile) and the SDK provider-scenario test (active-session-script-publication.test.ts) needed no changes and pass unmodified. * refactor(daemon): decompose session-replay-runtime.ts into three modules (#1555) Splits the ~1096-line replay runtime into cohesive pieces, keeping session-replay-runtime.ts as thin orchestration (~240 LOC): - session-replay-runtime-engine-adapter.ts: the AdReplayStepRuntime adapter (createAdReplayStepRuntime, the build*Failure capability implementations, and the lastResponse/lastObservation side-map mechanics), extracted verbatim. - session-replay-runtime-plan.ts: extended with the plan-side helpers (validateReplayBackendFlag, inspectReplayPlanManifest, resolveReplayPlanEntryIndex, prepareReplayPlan, routeMaestroReplay) alongside the buildReplayMetadataFlags helper already there — buildReplayMetadataFlags is now module-private since its one caller moved into the same file. Also introduces ReplayScriptFileParams, named here (instead of derived via Parameters<typeof runReplayScriptFile>) so routeMaestroReplay can reference the shape without importing back from session-replay-runtime.ts. - session-replay-runtime-session.ts (new): session preparation (prepareReplaySession and its coordinator arming/repair-preflight helpers), extracted verbatim. Coordinator ownership is unchanged: createReplayCoordinator is still constructed only in session-replay-runtime.ts, matching replay-coordinator-ownership.test.ts's allowlist as-is — every extracted module receives the already-constructed ReplayCoordinator as a parameter. Pure move; no behavior change. * test(replay): cover pre-step artifact ordering and resume-before-mutation (#1555) Two invariants found during the P5 decomposition pass now have direct counterfactual-verified coverage: - packages/ad-replay/src/internal/__tests__/step-loop.test.ts: a post-dispatch target-binding mismatch (dispatchWithGuard) must report the accumulated PRE-step artifact snapshot it was called with, never the artifacts the failed dispatch itself produced. Verified red by swapping the buildPostDispatchTargetBindingFailure call to outcome.artifactPaths. - src/daemon/handlers/__tests__/session-replay-runtime-plan.test.ts: a rejected --from/--plan-digest resume must never reach prepareReplaySession's coordinator-mutating writes (the R2 ordering invariant) — a pre-armed repair transaction and corrective-resume watermark are asserted byte-for-byte unchanged after rejection. Verified red by calling prepareReplaySession before honoring the plan-validation rejection. * fix(ad-replay): enforce the exact two-entrypoint facade (#1555 review P1) packages/ad-replay/src/index.ts now exports exactly two value symbols, inspectAdReplay and runAdReplay, and zero types — formatReplaySuccessMessage (presentation) moves beside its one caller in session-replay-runtime.ts, and every type a root daemon file needs is derived structurally off the two entrypoints in the one new src/daemon/ad-replay-facade-types.ts module instead of being named off the façade. scripts/layering/package-boundaries.ts's readNamedExports is rewritten on oxc-parser's own static-export table instead of a regex, so it can no longer silently miss a widening export form: a bare `export *` re-export or an `export default` now throws (an un-enumerable, and therefore un-pinnable, export), while `export * as ns` and every other enumerable form is still counted. The pinned exact-symbol assertion in package-boundaries.test.ts is narrowed to ['inspectAdReplay', 'runAdReplay']. * fix(ad-replay): translate wire failures before the engine boundary (#1555 review P1) AdReplayDispatchOutcome's guard-mismatch/landmark-mismatch variants carried a generic `details: Record<string, unknown> | undefined` bag straight off the wire response — a daemon wire projection crossing into the engine even though the outcome itself was already a neutral type. The daemon adapter (session-replay-runtime-engine-adapter.ts) now narrows that bag into the typed AdReplayGuardMismatchEvidence/AdReplayLandmarkMismatchEvidence shapes (observed identity, expected/observed structural denotation, ancestry entries, match count) before returning the outcome; the unknown-parsing readers move there with the wire-reading responsibility they always were. target-verification.ts's deriveReplayTargetGuardMismatchEvidence/ deriveWaitLandmarkMismatchEvidence now consume only the typed values — no `unknown`-valued record type remains on any engine-crossing signature. * fix(ad-replay): move variable semantics/planning behind runAdReplay (#1555 review P1) The daemon assembled the `${VAR}` scope (buildPreparedReplayScope) and interpolated actions at two independent call sites: dispatch's own (invokeReplayAction) and target verification's separate one (resolveTargetVerificationEntry) — duplicated orchestration the P5 design assigns to the engine. runAdReplay's request now carries the raw scope INPUTS (varSources: plain builtins/file/shell/cli-env data, plus actionLines/actionSourcePaths/ resolvedPath for interpolation-error location) instead of a built scope; the engine builds the scope and resolves each action exactly once per step, handing the RESOLVED action to dispatchStep/beginTargetVerification while every other capability still receives the ORIGINAL recorded action (a target-binding divergence reports the recorded selector, never an expanded ${VAR}). This is the one resolution site now — session-replay-action-runtime.ts's invokeReplayAction and session-replay-target-verification.ts's resolveTargetVerificationEntry no longer hold a scope or call resolveReplayAction themselves. Scrub-value collection (collectReplayScrubbableVarValues, for divergence-report redaction) is kept single-sourced in the engine too: it's computed from the engine's own live scope and threaded to each build-failure/handleActionFailure capability as an explicit scrubVars argument, rather than the daemon recomputing it from a second scope object (which would have gone stale, since expandedBuiltinNames tracking now only happens engine-side). The Maestro replay path's own daemon-side vars usage is unrelated (a different engine) and is out of scope here. * fix(ad-script): make ${VAR} interpolation a linear scanner CodeQL flagged the interpolation regex's fallback group as js/polynomial-redos once vars.ts moved into packages/ (library-input classification): every ${NAME:- prefix of an unclosed input rescanned to end-of-string, quadratic overall — 1,857 ms measured on 20k repetitions of '${A:-['. Replaced with a single-pass scanner; failed fallback scans emit their span verbatim and resume after it (escape-pair alignment is identical from every candidate start inside the span, so no later candidate can terminate where the failed scan could not). Equivalence: 200k-trial differential fuzz against the retired regex over the adversarial alphabet, zero mismatches; both adversarial shapes now resolve in 1-2 ms. * refactor(ad-replay): typed façade replaces the zero-type rule (#1555 structural-quality review) Reverses the exact-two-value zero-type export shape #1555's second review pass established: it forced every root type derivation through one shim (src/daemon/ad-replay-facade-types.ts) and left four daemon-side twin types (TargetVerificationEntry, TargetClassificationOutcome, TargetBindingFailureEvidence, ReplayVerifiedTargetGuard) plus a toDaemonEvidence copy translator shadowing the engine's own shapes. packages/ad-replay/src/index.ts now exports inspectAdReplay/runAdReplay (unchanged, still the only two values) plus the neutral vocabulary their signatures are built from, by name — following packages/maestro's façade precedent. The exact-symbol gate in scripts/layering/package-boundaries.test.ts is widened to pin the full sorted list (values + types). The four daemon twins are deleted; session-replay-target-verification.ts and session-replay-runtime-engine-adapter.ts now use the engine's own AdReplayVerificationEntry/AdReplayTargetClassification/ AdReplayTargetBindingEvidence/AdReplayVerifiedTargetGuard directly. TargetBindingDivergenceBuilt's array fields are now readonly-compatible, so toDaemonEvidence's copy is gone — evidence flows through unchanged. * fix(ad-replay): honor the selector port's own contract in the parse gate target-verification.ts's planPreDispatchTargetVerification used resolveRecordedTarget (operation 2, resolve) over an empty node tree purely to read its parse-invalid reason — a resolve call standing in for a parse call, even though readSelectorExpression (operation 1, parse) exists to answer exactly that question and was already unused inside the engine. Replaced with port.readSelectorExpression('ordinary', [token]). The mapping is not 'invalid' -> skip: production's 'ordinary'/'wait' grammars only ever record a boundary once it has already parsed, so a single malformed token can only come back 'not-applicable' there ('invalid' is unreachable from this call site on the production adapter). Both non-'expression' outcomes map to skip, matching the historical behavior (a single parse-invalid reason covered both cases). platform dropped from the function's params — it was only ever threaded to the resolve call this replaces. Added a contract-suite cell pinning the exact (diverging) discriminant each adapter reports for a selector-shaped-but-malformed bare token, and why the divergence is harmless for the one real consumer. * refactor(ad-replay): split step-loop.ts and shrink the daemon adapter (#1555 structural-quality review) step-loop.ts (810 LOC) splits three ways, following packages/maestro's own precedent: - internal/runtime-port-types.ts: the AdReplayStepRuntime boundary vocabulary (all the neutral types the engine/daemon exchange). - internal/verify-dispatch.ts: verifyAndDispatchStep + its dispatchNoGuard/ dispatchWithGuard helpers. - internal/step-loop.ts: runAdReplay itself plus the terminal-close/ executable-action structural logic (isExecutableReplayAction, resolveSuppressedTerminalCloseIndex). packages/ad-replay/src/index.ts's type exports now source from runtime-port-types.ts. step-loop.test.ts's AdReplayStepRuntime import moves to the new path (no assertion changes). src/daemon/handlers/session-replay-runtime-engine-adapter.ts (553 LOC after item 1's twin removal) shrinks to 294 via two further extractions: - session-replay-dispatch-narrowing.ts: the wire `details` bag -> typed evidence narrowing and dispatch-failure classification. - session-replay-runtime-step-support.ts: ReplayStepContext (moved here to avoid a cycle with the adapter, which re-exports it by name) plus the failure-wrapping/diagnostics-support helpers. Final LOC: adapter 294, dispatch-narrowing 148, step-support 153, step-loop 225, verify-dispatch 246, runtime-port-types 374. * test(ad-replay): package-local tests for resume.ts/target-verification.ts + terminal-lifecycle test rename resume.test.ts covers resolveReplayEntryIndex directly (previously only exercised transitively through the daemon's session-replay-runtime-plan tests): no --from/--plan-digest, the paired-flags requirement, in-range --from, out-of-range rejection, stale-digest rejection, the authorized empty-tail boundary (actionCount + 1) gated on a matching watermark, and the unperformed-record-and-heal growth check. Counterfactual run and restored: widening describeOutOfRangeResumeFrom's bound turns the out-of-range/ empty-tail-without-watermark assertions red (2 failures observed). target-verification.test.ts covers all four engine policy functions directly: the two plan* pre-capture gates and the two derive* post-dispatch evidence builders, including item 2's own new decision surface (a fake ReplaySelectorPort proving both non-'expression' readSelectorExpression outcomes map to skip). Counterfactual run and restored: narrowing the check to the literal `'invalid' -> skip` reading turns the 'not-applicable' case red (reports recorded-unverifiable instead of skip). session-replay-terminal-lifecycle.test.ts renamed to session-replay-runtime-keep-session.test.ts: its production module (session-replay-terminal-lifecycle.ts) was already deleted by the #1554 fold-in, and its six cases drive the full runReplayScriptFile round trip against a real SessionStore (including daemon-only postconditions the engine's step loop never reaches) rather than testing engine policy through the façade in isolation — the engine's own terminal-close-suppression decision already has direct, cheaper coverage in step-loop.test.ts. No assertion changes; both files' header comments cross-reference the split. * refactor(ad-replay): compute scrub values once per step, one name end to end collectReplayScrubbableVarValues(scope) was called fresh at 5 separate return points inside one verifyAndDispatchStep invocation plus once more in handleActionFailure — always the same result, since nothing between them mutates scope. step-loop.ts's runAdReplay now computes scrubVars ONCE per step, right after resolveReplayAction (the one call that can grow the scope's expanded-builtins set), and threads it as a plain readonly AdReplayScrubValue[] value; verify-dispatch.ts no longer imports ReplayVarScope or collectReplayScrubbableVarValues at all. "One name" end to end: the daemon's TargetBindingDivergenceContext.scrubVars and withReplayFailureDiagnostics's scrubVars param used a separately-derived ReturnType<typeof collectReplayScrubbableVarValues> (mutable array) instead of the engine's own AdReplayScrubValue, requiring a [...scrubVars] copy at every daemon call site to satisfy the mutable-array type. Both now use readonly AdReplayScrubValue[]/readonly ReplayVarScrubEntry[] (structurally identical, already readonly-safe downstream — scrubReplayVarValues and createReplayDivergenceSanitizer already accepted readonly arrays), so the four [...scrubVars] copies in session-replay-runtime-engine-adapter.ts are gone. * fix(daemon): make lastObservation genuinely per-step, not per-run createAdReplayStepRuntime's lastObservation closure lives for the whole replay run (one factory call covers every step), but was never reset between steps. Every current buildTargetBindingFailure call site happens to be preceded by this same step's own captureObservation, so the `lastObservation ?? { reason: 'observation-missing' }` fallback could never actually fire — but if it ever did (a future call path reaching buildTargetBindingFailure without capturing first), it would silently attach the PREVIOUS step's screen instead of reporting the missing-capture condition the fallback message claims. armStep runs exactly once per step, before any of that step's other capabilities (verified against step-loop.ts's runAdReplay loop order) — the natural per-step boundary. It now clears lastObservation first. No behavior change on any reachable path today (full daemon + ad-replay suite: 1766/1766 green); an unrelated device-claim-prune contention flake was observed once and did not reproduce on isolated or full-suite reruns. * docs(ad-replay): fix decayed review-changelog comments naming defunct symbols Four comments named symbols/paths that no longer exist, left behind by earlier review passes describing PR history rather than the current constraint: - session-replay-runtime-step-support.ts / session-replay-runtime.ts (2 sites): referenced a function called executeStep, which was never reintroduced under that name after the P5 split — the actual mechanism is the runtime's dispatch/build-failure capabilities recording into the lastResponse side-map. - session-replay-runtime.ts: referenced an engine collectArtifactPaths capability that does not exist — artifactPaths is a daemon-side Set the adapter mutates via collectReplayActionArtifactPaths. - packages/ad-replay/src/internal/selector-port.ts: pointed at ./testing/in-memory-selector-port.ts, the in-memory adapter's pre-stage-D location — it has lived at src/__tests__/test-utils/in-memory-replay-selector-port.ts since. - session-replay-repair-hint.ts / session-replay-runtime-step-support.ts (2 sites): named target-identity.ts, which does not exist (the real file is target-identity-node.ts); the second site additionally mislabeled classifyReplayTarget as engine-side when it is daemon-side (session-replay-target-classification.ts). Comment-only; no behavior change. * refactor(ad-script): move declaredScriptPlatform to its natural shared owner packages/ad-replay/src/internal/inspect.ts's declaredScriptPlatform and src/daemon/replay-device-selection.ts's readScriptReplaySelection each kept their own copy of the same "platform declared before the first open" scan over runtime/open actions — .ad script semantics, not engine or daemon policy, needed independently by ad-replay's plan-digest precedence and the daemon's device-selection platform resolution. Verified this was a genuine duplicate (not the single-sourced state I initially reported): readScriptReplaySelection's platform-tracking loop computes the identical result via a differently-shaped traversal fused with its own app-target scan. resolveDeclaredScriptPlatform now lives in packages/ad-script (its natural owner: the one package both ad-replay and the daemon already depend on, avoiding the R11 issue that justified the original duplication). The daemon's app-target scan stays its own separate pass; fusing it back into the shared function would smuggle a daemon-only concern into ad-script for no measurable cost (the actions array is small, and the shared function already stops at the same point the app-target scan needs to look). * docs(ad-replay): fix package.json description to match the current façade Described "target-identity, variable substitution, plan-digest, and report primitives" — the wide pre-#1555-review façade shape. Vars/identity/report vocabulary moved to ad-script/daemon across the P5 and #1555 review passes; the package now exports exactly inspectAdReplay/runAdReplay plus the neutral AdReplayStepRuntime vocabulary. Description updated to match. * refactor(daemon): fold the step-support fragment back into the engine adapter A simplicity audit judged session-replay-runtime-step-support.ts a size-target fragment, not a concern boundary: four unrelated concerns, one consumer, and a header comment admitting it existed to satisfy the <300 LOC metric. Folded back; the previously-exported helpers are module-private again; the adapter's honest size is renegotiated from the plan metric (dispatch-narrowing stays extracted — it has one nameable job). |
||
|
|
99967c7f01 |
fix: restrict project config trust (#1565)
* fix: restrict project config trust * fix: preserve daemon auth transport context * refactor: simplify project config trust * fix: restrict project config write sinks |
||
|
|
123521652c |
fix(ios): double-check off-screen click refusals against a direct element read (#1566)
* fix(ios): double-check off-screen click refusals against a direct element read #1542: after an AX-free scroll on iOS, the off-screen interaction guard can refuse a click even though the target is genuinely on-screen, because it trusts a scroll-container ancestor's rect from the bulk accessibility tree, which a keyboard-dismiss content-offset correction can leave stale/corrupted while the target's own rect is already correct. When the guard is about to refuse on iOS, it now takes a single fresh, tree-independent XCUITest read of the target element (querySelector) and trusts that read's live `hittable` + rect-vs-root-viewport signal instead, if it positively confirms on-screen. Any failure to unambiguously re-resolve the element (no id/label, not found, ambiguous, transport error) fails closed exactly as before. Genuinely off-screen targets, and every other platform, are unchanged: the backend method is gated to local (non-provider) iOS sessions only, and only ever runs on the about-to-fail path. The decision itself is a pure function (decideOffscreenRefusalDoubleCheck in mobile-snapshot-semantics.ts) with counterfactual-proven tests: hardcoding it to always trust the bulk verdict turns the rescue test red, and hardcoding it to always trust the direct read (including on "unavailable") turns the fail-closed/genuine-refusal test red. Live-validated on a fresh-boot iOS simulator: checkout-form.ad 2/2 passes (previously failing at step 11), gesture-lab.ad 2/2 (regression), and the Android checkout-form/gesture-lab suite passes unchanged, proving no cross-platform behavior change. * fix(ios): tap the live rect after a rescued offscreen refusal; collapse the double-check to one backend hook Review blockers 1+2 (interleaved by design — the soundness fix is expressed through the collapsed hook's contract): 1. SOUNDNESS: a rescued refusal now returns the node PATCHED WITH THE LIVE RECT the backend confirmed, and every downstream use (tap point, response) reads from that returned node — never the original. In the frozen-tree manifestation (the whole bulk tree pinned at pre-gesture values), the original rect can be stale even when the rescue verdict is correct; tapping it would have silently landed at the wrong coordinate. New regression: offscreen-double-check.test.ts's frozen-tree case, with a counterfactual (revert to computing the point from the pre-guard node) proven red then reverted. 2. SURFACE: collapsed to ONE optional backend hook, `confirmOffscreenTargetVisible?(context, node, rootViewport): Promise<Rect | null>` — conceptually a boolean, but returns the live rect so item 1's fix has something to act on. Deleted decideOffscreenRefusalDoubleCheck, the OffscreenRefusalDoubleCheckSignal/Reading ADT, and resolution.ts's dual-signal reconciliation shell: the bulk side was hardcoded 'off-screen' at the only call site, so the two-signal model was dead weight. The shared guard is now: bulk-off-screen -> ask the hook -> a live rect proceeds (patched), anything else (including no hook) throws exactly as before. The pure geometry boundary that decision reduces to (`isConfirmedOnScreenProbe` in mobile-snapshot-semantics.ts, replacing the deleted ADT) is unit-tested with two counterfactuals: ignoring `hittable` and ignoring the viewport containment check each turn a test red (proved, then reverted). `throwIfOffscreenInteractionTarget` is now exported (ADR 0011 registry honesty, see the contracts commit) and directly unit-tested in resolution.test.ts, mirroring the existing tryResolveRefNode pattern. * refactor(ios): direct-ios-selector.ts back to pure gate/parse; reuse queryDirectIosSelector Review blocker 3 (BOUNDARIES): - direct-ios-selector.ts no longer does any runner I/O — it's back to pure gate/parse (readSimpleIosSelectorTarget, deriveDirectIosNodeSelector, isDirectIosSelectorFallbackError) plus the ONE shared eligibility predicate, isLocalIosRunnerSession(session, { skipPendingPostGestureStabilization }). Both the direct-selector tap fast path and the new offscreen double-check probe call this same function; the one behavioral difference between them (the tap fast path skips a session with a pending postGestureStabilization, the double-check does not) is now an explicit parameter instead of two separately-written gates. - The probe I/O moved to a new sibling, src/daemon/offscreen-target-probe.ts, which reuses selector-runtime.ts's `queryDirectIosSelector` (now exported and decoupled from SelectorRuntimeParams — it takes a session + a bare {key, value} selector + AppleRunnerRequestOptions) rather than opening a second querySelector client. Node extraction (`readDirectIosSelectorNode`, the one `as SnapshotNode` cast) stays singular, inside selector-runtime.ts. - interaction-runtime.ts wires confirmOffscreenTargetVisible only when isLocalIosRunnerSession(session, { skipPendingPostGestureStabilization: false }) — deliberately NOT skipping a pending post-gesture stabilization, since that is exactly the window the double-check exists to cover. * docs(contracts): name the iOS offscreen rescue hook as part of the guarantee matrix Review blocker 4 (GUARANTEE HONESTY): the shared offscreen cell (RUNTIME_TREE_SHARED_GUARANTEES.offscreen, used by runtime-selector and runtime-ref) and the native-ref path's offscreen cell still named isNodeVisibleOnScreen as sole enforcement after #1542's double-check landed — that understates what actually enforces the guarantee now. Both cells' `via` now point at throwIfOffscreenInteractionTarget (exported from resolution.ts in the prior commit for exactly this), the real end-to-end enforcement point: isNodeVisibleOnScreen is the bulk-tree decision it starts from, and on iOS a would-be refusal can still be confirmed via the optional AgentDeviceBackend.confirmOffscreenTargetVisible hook before erroring. The cell's comment states the rescue-only, fail-closed shape explicitly per ADR 0011's matrix rules — this does not weaken the cell, it extends its description to match reality. iOS rescue policy stays OUT of resolution.ts's shared docstrings (the "spine"): this registry file is where per-path enforcement detail belongs, and the optional-method wiring in interaction-runtime.ts remains the only cross-platform touch. The registry's own gate test (interaction-guarantees.test.ts) still passes: every `via` resolves to a real exported symbol. * test(ios): move #1542 offscreen double-check tests out of interaction.test.ts Review blocker 5 (TEST HOMES): AGENTS.md forbids adding to daemon/handlers/__tests__/interaction.test.ts (it predates the test-mirrors-source-topology rule and shrinks opportunistically). Reverts the 172 lines added there in the original PR version; interaction.test.ts is back to its pre-#1542 baseline (81 tests, unchanged). The same assertions now live in their proper homes (see the prior three commits for the sources they cover): - pure decision pin: src/utils/__tests__/mobile-snapshot-semantics.test.ts (isConfirmedOnScreenProbe, with the two counterfactuals) - direct-guard pin: src/commands/interaction/runtime/resolution.test.ts (throwIfOffscreenInteractionTarget, mirroring tryResolveRefNode) - probe unit tests: src/daemon/__tests__/selector-runtime.test.ts (queryDirectIosSelector) and src/daemon/__tests__/direct-ios-selector.test.ts (isLocalIosRunnerSession, deriveDirectIosNodeSelector) - probe integration: src/daemon/__tests__/offscreen-target-probe.test.ts (confirmIosOffscreenTargetVisible, mocked runner) - end-to-end rescue/refuse, including the frozen-tree live-geometry regression + its counterfactual: new sibling src/commands/interaction/runtime/offscreen-double-check.test.ts (next to resolution.ts, using the same createInteractionDevice harness resolution.test.ts already uses) * style: oxfmt formatting for resolution.test.ts |
||
|
|
4551fb7aa3 |
test: make pid-liveness fixtures deterministic under load (#1556)
Three tests classify a fixture owner's liveness via classifyOwnerLiveness (or the daemon-process equivalent), which re-reads the owner's process start time via a real `ps -p <pid> -o lstart=` shell-out with a 1s timeout. Under full-suite CPU contention that second read can miss its deadline and return null, mismatching the value captured earlier and flipping a genuinely-live owner to 'owner-process-dead'. Pin the pid->start-time (and, for the daemon-client case, pid->command) mapping to a deterministic value per test file instead of letting a second real subprocess call race the first, without weakening the dead-owner path (isProcessAlive stays real and un-mocked everywhere). |
||
|
|
4fd04414e0 |
fix: clear the last polynomial-redos instance in swift-cache (#1549)
* fix: clear the last polynomial-redos instance in swift-cache sanitizeCacheName used /^-+|-+$/g to trim edge dashes, the same js/polynomial-redos pattern PR #1546 retired everywhere else. Replace it with the linear-time trim used there, and add a counterfactual regression test that fails against the old regex on a long interior dash run. * fix: keep sanitizeCacheName private, drive redos/fallback pins through compileSwiftSourceText Addresses PR #1549 reviewer feedback: sanitizeCacheName was exported solely so the regression test could import it, which docs/agents/testing.md's test-interface rule forbids. Reverted the export and rewrote the test to exercise the sanitizer through compileSwiftSourceText, an existing production seam that already calls it. - Timing pin: a cache name with a 100k-char interior dash run still resolves in sub-second time (the call may reject once it reaches disk I/O due to the OS path-component length limit, but that happens only after the now-fast sanitize step, so timing the settle either way still proves no catastrophic backtracking). - Fallback pin: a cache name that sanitizes to nothing (e.g. '---') still produces the 'swift-helper' fallback, observed via the returned executable path. Counterfactuals (see PR comment for full output): - Restoring the retired `/^-+|-+$/g` regex trim made the timing pin fail: 3428ms >= 1000ms. - Removing the `|| 'swift-helper'` fallback made the fallback pin fail: basename did not start with 'swift-helper-'. |
||
|
|
b125435989 |
refactor: extract WebDriver provider package (#1504)
* refactor: extract webdriver provider package * refactor: consolidate shared XML codec |
||
|
|
a3ab69a110 |
refactor(replay-test): neutralize the values crossing the scheduler seam (#1478 P3, part 1) (#1509)
* refactor(replay-test): neutralize the values crossing the scheduler seam
#1478 P3, part 1 of 2. Prepares the replay-test extraction by removing every
non-neutral value that crosses the scheduler seam, in place under `src/`, so the
physical move to `packages/replay-test` is a file move rather than a redesign.
`DaemonResponse` no longer crosses the seam. `session-test-types.ts` typed
`runReplay`/`finalizeAttempt` as returning a daemon response and the scheduler read
`.error.code`, `.error.details`, and `.data.replayed/.healed/.warnings/
.snapshotDiagnostics` off it throughout. That is invisible to R10 today only
because `checkDaemonTypesImporters` skips `src/daemon/`; once the files live in a
package they become external `daemon/types.ts` importers, which the ratchet only
lets shrink. Attempts now resolve as tagged `ReplayTestAttemptOutcome` values
carrying exactly what the scheduler consumes, including an `infrastructure` tag —
classifying an environmental failure needs platform boot-diagnostic vocabulary the
scheduler must not import, so the host decides and the scheduler reads the verdict.
`session-test-outcome.ts` is the one place a daemon response becomes an outcome.
Step events get a narrow per-attempt port. They were emitted from
`session-replay-runtime.ts` and `session-replay-maestro-observer.ts`, both reading
a request-global `AsyncLocalStorage` seeded per attempt. The scheduler now hands
each attempt an `onStep` sink, threaded the way `tracePath` already is; both
engines call it and `withReplayTestActionProgress`/`readReplayTestActionProgress`
are gone. A direct `replay` simply has no sink.
ADR 0012 divergence becomes a neutral leaf. `src/replay/divergence.ts` depended
only on kernel contracts and redaction, yet Maestro constructs divergences too and
CLI/MCP both render them, so P5 could not have moved it into `packages/ad-replay`.
It is now `@agent-device/contracts/divergence`; the renderer's output text is
unchanged.
The progress wire vocabulary moves to `@agent-device/contracts/progress`. It is
serialized by `request-progress-protocol.ts` and reconstructed by the CLI reporter
path, so it belongs below both; `src/request/progress.ts` keeps only the sink and
its AsyncLocalStorage binding.
Together these clear all four of replay-test's recorded R10 migration imports, so
the rule now enforces unconditionally for that module.
Behavior is unchanged. The shipped reporter contract — export spellings,
object/factory loading, hook names, timing/order, value fields, the synchronous
live-hook rule, awaited suite completion, error handling, exit codes — is
untouched, and `session-test-reporter-values.test.ts` passes unmodified. The
`--shard-all` `total`/`runnable` asymmetry is preserved as characterized.
Refs #1478
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RXQLYV7etZx3gcXsUsrQJ8
* test(replay-test): pin the Maestro reporter step path against the onStep port
Review finding on #1509: the native `.ad` reporter ratchet exercises only one of
the two `onStep` forwarding chains, so deleting a link in the Maestro chain would
silently stop `onTestStep` for every `test --maestro` run while every existing
reporter test stayed green. Same defect class as the dropped diagnosticId/logPath
(#1501) and the dropped reporter `hint` (#1505).
Maestro is one of P3's two required real adapters and its chain shares no links
with the native one below `runReplayScriptFile`:
scheduler sink -> runReplayScriptFile -> runTypedMaestroReplayFile
-> createMaestroReplayObserver({ onStep }) -> actionStarted -> onStep
Adds a Maestro scenario driving `test --maestro` through the real session handler
and the real reporter registry. It asserts the step payload the engine produces
(`stepIndex`/`stepTotal`/`stepCommand`/`stepValue`, including that a value-less
command stays value-less) together with the attempt/session identity the scheduler
supplies, since that half of the event came from request-global AsyncLocalStorage
before P3. A second case drives a retry so step events must carry attempt-1's
session and then attempt-2's. The flow `name` also pins the reporter `title`, a
value only the Maestro path can produce.
New file rather than an addition to session-test-reporter-values.test.ts: that file
is the pinned characterization and must keep passing unmodified, and Maestro needs
its own vi.mock of core/dispatch for device resolution.
Counterfactual run, both links, each restored after:
- dropping `onStep` from createMaestroReplayObserver in
session-replay-maestro-runtime.ts
- dropping the emitMaestroStep call from actionStarted in
session-replay-maestro-observer.ts
Each dropped both onTestStep events ("expected [ 'onSuiteStart', 'onTestStart',
…(2) ] to deeply equal [ 'onSuiteStart', 'onTestStart', …(4) ]") and failed both
new cases, while session-test-reporter-values.test.ts passed all 4 — exactly the
hole the reviewer identified.
Test-only; no production change. Bundle output is byte-identical to
|
||
|
|
0ee2a86129 |
refactor: extract contracts workspace package (#1499)
* refactor: extract contracts workspace package * fix: preserve screenshot diff result contract * test: stabilize Android keyboard smoke |
||
|
|
809968d940 |
Simplify screenshot diff output (#1495)
* refactor: simplify screenshot diff output * test: update screenshot diff cli output * chore: remove unused screenshot geometry helpers * fix: version screenshot diff result contract * fix: preserve screenshot diff result compatibility |
||
|
|
7ef83bc8bc | refactor: restore pngjs codec (#1498) | ||
|
|
76453add71 |
refactor: pnpm workspace + @agent-device/kernel pilot (#1490 W0) (#1494)
* refactor: pnpm workspace + @agent-device/kernel pilot (#1490 W0) Extend the workspace with packages/* and move the kernel behind an enforced public API: packages/kernel with nine consumer-earned subpath exports (errors, device, snapshot, contracts, collections, rect, redaction, daemon-error, bounds — the last absorbed from utils as Rect vocabulary). Every kernel import repo-wide becomes the @agent-device/kernel/<sub> specifier; kernel tests move to src/__tests__/kernel/ and exercise the package surface. The root declares the package in devDependencies (workspace:*), tsdown bundles it (noExternal) so the published artifact and its runtime dependency manifest are unchanged. Gate rewiring in the same change, per the W0 brief: - R1 kernel-sink retires (physically subsumed); new R11 package-boundaries guards no-root-back-imports, relative tunnelling past exports maps, undeclared workspace deps, and non-exported subpaths, with runtime resolution pins via import.meta.resolve. - resolveImportEdges and mutation ownership follow workspace specifiers through exports maps, keeping R4 cycle checks, depgraph, and derived test ownership connected across the seam (kernel-errors still owns 495 tests). listSourceFiles includes packages/*/src. - kernel becomes an unranked zone; mutation registry, stryker mutate globs, and the mutation-affected workflow path filter move to packages/kernel/src/errors.ts. - check:affected gains packages/ ownership (manifests fail open); vitest and coverage include packages/*/src; fallow ignores packages/** (its resolver cannot follow workspace specifiers). - The affected-selector CI job installs dependencies: its closure now crosses workspace specifiers, and the R8 relative exception is unsafe for production src files (Node ESM does not realpath, so dual specifier/relative loads would instantiate modules twice). The R8 zero-dep set is pinned empty with that rationale. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FUv7bvbWNryuXgSBuqTtep * fix: address W0 review — mutation sandbox, exports-map resolution, tsc -b Review findings on #1494, all five: 1. contracts-schema-public.test.ts reads the kernel source at its packages/ path (fs access invisible to the codemod and typecheck). 2. Mutation lane: Stryker sandboxes the tree but pnpm's node_modules symlink resolves @agent-device/* back to the real repo, so mutants in the sandbox never load and vitest.related finds no tests. vitest.mutation.config.ts now aliases each EXPORTED specifier to its source (derived from exports maps, never a wildcard), keeping resolution inside the mutated tree. Validated: kernel-errors module runs end to end (dry run 3,984 tests, mutants killed, exit 0). 3. Layering/depgraph resolve workspace specifiers through the exports-derived map (workspaceSpecifierTargets) instead of reconstructing paths, so '.'-facade packages resolve; the positional fallback remains only for map-less fixtures (P0 pin). 4. Per-package project references implemented: packages/kernel is composite (emitDeclarationOnly -> dist-types, gitignored), the root references it, and typecheck becomes tsc -b — probed to catch type errors on both sides under TypeScript 7 native. 5. R11's relative-route exception now requires membership in an actual R8 zero-dep job closure (zeroDepClosureFiles walks entries), not mere scripts/ placement — closing the dual-instantiation bypass. Also from review discussion: daemon-error moves out of the kernel package to src/client/ — its consumers (cli, client facade) rehydrate wire DaemonErrors client-side; the daemon only produces them. Kernel drops to 8 exported subpaths before any of them ship. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FUv7bvbWNryuXgSBuqTtep * refactor: one exports-map reader for mutation alias and ownership Fallow flagged workspaceExportAliases (cognitive 15, CRAP 90). The manifest-reading logic already exists as workspaceSpecifierTargets in scripts/layering/package-boundaries.ts, so both the Stryker sandbox alias table and the mutation ownership walker now consume it instead of carrying near-clones. Behavior unchanged; mutation suite 45/45 and changed-code fallow green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FUv7bvbWNryuXgSBuqTtep * fix: composite kernel without a root references edge FreeRange runs plain `tsc -p tsconfig.json`, and a root `references` entry makes non-build-mode TypeScript demand the referenced project's built declarations (TS6305) — a standing "build first" tax on every plain -p consumer (fr, editors). Keep the per-package composite project and build it in typecheck (`tsc -b packages/kernel` before the root and examples/sdk passes), but drop the root references edge: root consumption resolves through exports to source, identical to runtime and to the bundler. Probed: plain -p green with no prebuilt output; kernel-side type errors still caught by its own build. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FUv7bvbWNryuXgSBuqTtep * fix: R11 uses the layering parser; mutation config is a fallow entry Review blockers on #1494: - R11's private single-quote regex could miss a double-quoted or re-export route into packages/*/src. specifierSites now delegates to the layering model's parseImports (both quote styles, side-effect imports, re-exports, dynamic imports), with direct regressions for each formerly-invisible form. - vitest.mutation.config.ts becomes a declared fallow entry instead of a tolerated unused-file finding: the full-repo audit now reports it reachable (unused files 2 -> 1; the remainder predates this PR). FreeRange clean-checkout evidence: with packages/kernel/dist-types and every *.tsbuildinfo deleted, `pnpm check:freerange` reports 0 findings on this head — the TS6305 topology died with the root references edge in the previous commit; check:freerange has no build precondition. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FUv7bvbWNryuXgSBuqTtep --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
12fd36e786 |
test: add scroll-edge-state suite, close mutation gap (#1455) (#1474)
* test: add scroll-edge-state suite, close mutation gap (#1455) scroll-edge-state.ts had no dedicated test file and scored 28.75% mutation (92/320) — the worst of the five tracked decision kernels. Adds src/utils/scroll-edge-state.test.ts covering every exported function's decision branches (captureScrollEdgeState, runScrollEdgePasses, formatScrollEdgeMessage) plus the private container-selection/scope/signature logic they exercise. Raises the scoped mutation score to 95.31% (305/320); the 15 remaining survivors are named, justified-equivalent mutants (documented in the PR). * refactor(test): split scroll-edge-state suite, drop signature test-gaming Splits the 1,323-line scroll-edge-state.test.ts into three behavior-focused sibling suites (container selection, scope naming, pass orchestration) plus a shared fixtures module, per AGENTS.md's per-file LOC tripwire and sibling-test convention. Also removes the dead `signature`/buildScrollStateSignature/collectSubtreeNodes/ hasAncestor production code: it had zero consumers anywhere in the repo, and the tests pinning its private string format were the mutation-score-gaming flagged in review — with it gone, there's nothing left to game. Addresses review feedback on PR #1474 (#1455). * refactor(test): split scroll-edge-state suite further, tighten unreachable branch Splits container-selection.test.ts (562 LOC, over the 500-LOC extract tripwire) into four narrower suites by durable question: container existence, hidden-edge decisions, target resolution, and broad selection. Adds a scrollNode() fixture builder to cut inline node-literal repetition across the suite. Removes the "captureNodes resolves to undefined" test: the real contract (both production call sites, and captureScrollEdgeState's own captureNodes type) always resolves to a real array, so this exercised an unreachable branch behind a type-unsound cast. Tightens analyzeScrollEdgeState's parameter type to drop the now-provably-unreachable `| undefined`. Simplifies buildScrollContainerScope to filter out non-string identifier/label fields instead of mapping them to an empty-string placeholder — behaviorally identical (the placeholder always failed isUsefulScope's length check anyway) but removes a StringLiteral mutation surface that a synthetic fallback could never legitimately kill without coupling a test to Stryker's specific replacement text. Addresses further review feedback on PR #1474 (#1455). |
||
|
|
4e4ecdea0d |
test(ios): expand simulator e2e coverage (#1408)
* test(ios): expand simulator e2e coverage * test(ios): make coverage checks host portable * test(ios): handle deep link confirmation * test(ios): fix deep link prompt selector * ci: stabilize full simulator nightly * test(ios): stabilize permission prompt lifecycle * test(e2e): wait for route-specific landmarks * test(e2e): reset permissions from inactive app * test(e2e): redeliver trusted cold deep links * test(ios): verify orientation native readback * test(ios): stabilize simulator permission coverage * test(ios): wait for tab target after deep link * test(ios): paginate full event timeline * test(ios): simplify event pagination coverage * test(ios): stabilize simulator e2e coverage * test(ios): model simulator recorder lifetime * ci(test-app): cache fixture dependencies * fix(ci): isolate test app cache by node * fix(ios): settle fixture route navigation * test(ci): waive unbenchmarked ios system UI help * fix(ios): tolerate delayed simulator scale lookup * 0.20.1 * test(ci): remove superseded system UI waiver * fix(ios): harden simulator e2e reliability * chore: clarify Apple runner CI steps * fix(ios): wait before fixture home snapshot * fix(ios): require exact catalog navigation * chore(ci): format rebased workflows * fix(ios): retry unobserved fixture navigation * refactor(test): remove iOS e2e workarounds * fix(ci): verify fixture artifact provenance * fix(ci): align fixture artifact fingerprints * fix(ci): use unified Android helper packager * perf(ci): scope fixture build concurrency * chore: format fixture artifact tests * test(ios): update split Apple coverage owner * fix(ios): accept deep-link confirmation alerts |
||
|
|
8cce0ef6b8 |
test: ratchet mutation score over enumerated decision kernels (#1441)
* test: ratchet mutation score over enumerated decision kernels Adds a Stryker (vitest runner) mutation lane scoped to the decision kernels, a per-module baseline with tool/config provenance, and a ratchet that only lets scores rise. Non-gating until two consecutive stable weekly sweeps. Refs #1415 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * chore: declare the mutation test-scope seam for production-export analysis Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test: own kernel tests in the mutation registry and ship the #1430 lane envelope - restore bench:help-conformance, broken by a formatting-path edit - kernel test files select their module on PRs (registry `tests` + workflow paths), asserted to reach the kernel through the import graph - every mutation run writes the standard scheduled-lane artifact envelope - move src/utils/__tests__/errors.test.ts beside its source per the mirror rule Refs #1415, #1430 Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test: derive kernel test ownership and land the scheduled-lane health monitor Ownership of a kernel's tests is now computed from the static import graph (scripts/mutation/ownership.ts) instead of a hand-listed set, so a test that reaches a kernel indirectly -- src/__tests__/daemon-error.test.ts through src/daemon.ts -- selects that kernel on a PR. The PR lane triggers on every src test and shards the derived modules, keeping wall clock at one module. The lane envelope (#1430) is now written on every exit path with the stage it reached, so a crash before any mutant runs is distinguishable from a lane that never ran. Adds the derived cadence monitor (scripts/lane-health, daily workflow): scheduled lanes are enumerated from .github/workflows/ and reported dark, failing, or never-run against their own cron cadence. * fix: merge only Stryker reports from a shard directory The shard artifacts now carry the lane envelope beside mutation.json, and the merge globbed every .json under the download path, so the ratchet job fed the envelope to the report parser and died after the mutants had already run. * fix(mutation): fail on incomplete shard sets and envelope pre-run failures Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mutation): downgrade a passing envelope when a later lane step fails Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(mutation): shard by registry, defer the PR lane, drop the bundled watcher Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * docs: describe registry sharding and the deferred PR mutation lane Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(mutation): make the pre-graduation tooling exception select real mutants Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * test(mutation): give the worktree fixture commits their own identity Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: Michał Pierzchała <thymikee@gmail.com> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> |