mirror of
https://github.com/callstack/agent-device.git
synced 2026-09-14 20:06:34 +08:00
main
5 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3bf3ff130a |
fix: Linux input/a11y defects from #1935 (click miss, typed '=', GTK4 text) (#1949)
* diag: instrument Linux CI to gather evidence for #1935 input/a11y defects Temporary — adds a diagnostic step that dumps raw AT-SPI interfaces/actions for gnome-calculator's digit buttons, tests a raw xdotool click at a button's own rect (bypassing our promotion logic), and isolates the typed '=' character in several configurations. Will be removed once the real fixes land. * diag: harden diagnostic step against bash -e and AT-SPI registration races The prior version crashed 7s in: GH Actions runs steps under bash -e, and an unguarded python3 heredoc threw (iterating a dict instead of a list when the app wasn't found yet), aborting the rest of the script silently under continue-on-error. Guards every fallible command, and replaces the fixed 2s sleep with inspect.py's own poll-until-found loop. * diag: test WINDOW coordtype and static Text.get_text call (round 3) Round 2 proved Component.get_extents(SCREEN) returns (0,0) for every non-toplevel widget (real click miss confirmed on-screen), and Text.get_text() throws — a documented PyGObject binding collision with the deprecated 1-arg Accessible.get_text(). This narrows to the two candidate fixes before writing them: does CoordType.WINDOW give usable relative offsets, and does Atspi.Text.get_text(accessible, ...) (static call) return the real typed text. * fix(linux): resolve click-miss, dropped '=', and GTK4 text exposure defects Three defects surfaced by CI on #1935 (Linux Smoke lane), all confirmed live via instrumented CI runs before being fixed here: 1. Click misses its target: Component.get_extents(Atspi.CoordType.SCREEN) returns (0, 0) as the origin for every non-toplevel widget under this GTK4 build — confirmed by a raw click at the computed rect center landing on the window's own header-bar button instead of the intended digit button. CoordType.WINDOW gives correct, distinct per-widget offsets, so get_rect() now computes screen-absolute rects as that offset plus the enclosing top-level frame's own (correct) screen origin, threaded through traverse_node() alongside the existing window-title tracking. Complementary hardening: role "label" is now excluded from `hittable`, since GTK4 wraps every button's caption in a same-rect "label" child, and the shared cross-platform promotion logic in interaction-targeting.ts would otherwise retarget a click from the button onto that non-interactive label. 2. Typed '=' never arrives: a single isolated synthetic keystroke sent right after a focus change is unreliably delivered — confirmed live, both `xdotool type -- "="` and `xdotool key equal` sent alone produced no character at all, while multi-character bursts always landed in full. typeLinux and sendKey now wait a short settle margin before dispatching to xdotool/ydotool, absorbing the race regardless of which action last changed focus. 3. GTK4 apps expose no editable text: accessible.get_text_iface().get_text() throws "Atspi.Accessible.get_text() takes exactly 1 argument (3 given)" — a documented PyGObject binding collision between Text.get_text and the deprecated 1-argument Accessible.get_text, silently swallowed as "no text" by the broad exception handler. get_text_value() now calls the unbound Atspi.Text.get_text(accessible, ...) form, which correctly returns the real content. The Linux smoke replay is restored to exercise all three fixes together (click a resolved digit button, type a full calculation including the '=' keystroke, wait on the computed result through the tree) instead of staying at the weakened, contract-tier assertions the defects had forced. The coverage manifest promotes click and type from command-contract to live accordingly. * fix(linux): drop unproven keyboard-settle and hittable changes per review Addresses thymikee's review on #1949 (both points correct): P1: the keyboard settle (typeLinux/sendKey) was unjustified. The cited diagnostic evidence for a dropped '=' actually shows the opposite — "100+55=" and "5=5" both computed correctly with zero settle, proving '=' was delivered in every multi-character burst tested. Sending '=' alone to an empty entry showing a blank display is normal calculator semantics (nothing to evaluate), not a lost keystroke. The likelier explanation for the original "100+55" screenshot (run 32487868346) is that its attempt-3 hit the already-fixed mousemove --sync hang, not an independent keyboard-dispatch defect. Reverted; no keyboard-dispatch change was needed. P2: the `role_name != "label"` hittable narrowing was extra surface beyond what the click-miss fix required. The corrected AT-SPI coordinates alone fix the observed miss — the button and its same-rect label child resolve to nearly identical centers, so descendant promotion still lands inside the button either way, and the replay can't distinguish which node it actually targeted. Reverted; only the coordinate fix remains. |
||
|
|
81409f1a7c |
refactor: migrate type to the request-bound device runtime (#1935)
* refactor: migrate type to the request-bound device runtime
Wave 5 unit 2 for #1739 (ADR 0019), find's last blocker. `type "text"` and
`find <q> type "text"` now reach the device through one admitted, request-bound
`typeText` operation instead of the `handleTypeCommand` interactor leaf and its
dispatch-table arm.
- New `TypeTextRuntimeOperations` contract riding the same `Interactor` seam as
focus/screenshot/element-text; the operation returns the interactor's own
closed `TypeTextBackendResult`, so Apple route evidence passes through and
every other owner types blind, exactly as before. The iOS synthesized-type
commit wait (#1676) is Apple-interactor-internal and moves nowhere.
- The interaction backend's `typeText` member exists only when the `type`
handler admitted and bound a runtime — no caller can fall back to legacy
dispatch, so the command keeps exactly one execution path (R41).
- `executeBoundTypeText` reproduces the retired leaf byte-for-byte: leading-ref
rejection with the same hint, space-joined positionals, the 0-10000 delay
bound, and only textEntryRoute surviving from the owner's result. Its parse
pins moved from the dispatch-level tests into the daemon runtime test.
- Exact-owner facts replace the capability bucket (apple sim+device, android
all-but-simulator-row, harmonyos emulator+device, linux device, web device,
vega unavailable, providers wherever their interactor is reachable) and
`type` leaves HARMONYOS_SUPPORTED_COMMANDS / WEB_INTERACTION_COMMANDS.
- The Linux desktop replay types a digit on real hardware; the coverage
manifest promotes `type` contract -> live with the two-sided count pins.
- Android/webdriver facts helpers extracted (androidTouchFact, interactorCell)
to keep inspectFacts under the complexity gate.
`find` stays legacy: both of its direct execution legs now share bound
runtimes, so the atomic R35 cutover is next.
* refactor(type): single-pass daemon routing, shared binder source, owner cell tests
Review follow-ups on #1935.
- The type handler now calls the bound executor directly: the interaction-
runtime hop validated and formatted what executeBoundTypeText validates and
formats again, so it is gone — no boundTypeText backend member, no second
result rebuild. The ADR 0014 frame expiry moves to the handler.
- New contracts/interactor-operation-binding.ts: one local resolver and one
fail-closed provider resolver shared by the screenshot, focus, and type
binders — three private copies retired, provider error text preserved.
- provider-limrun/interaction-operations.ts: the interactor-backed interaction
cells move out of the app-log owner (586 -> 563 lines, below its pre-unit
size); text interaction is composed by that owner, not defined in it.
- Every owner runtime test now pins the focusPoint/typeText fact cells and
bound-operation presence for its exact kinds: apple, android (incl. the
synthetic-simulator refusal), harmonyos, linux, web, vega (refusal + hint),
webdriver (reachability-gated, incl. inactive session), limrun (live +
recovery). The webdriver unsupported-capability row documents that
interaction gates on interactor reachability, not capture declarations.
- The Linux replay assertion is now change-sensitive: type "555" then wait for
a 555 node — no calculator button carries that label, so the wait passes only
if the keystrokes landed in the display; deleting the type step turns it red.
* fix(replay): give the calculator focus before the Linux type assertion
The change-sensitive wait exposed what the review predicted: the typed digits
never landed, because `focus 100 100` clicks the DESKTOP and takes keyboard
focus away from the calculator. The old broad assertion masked exactly this.
The retries then wedged on a latent quirk: attempt-1 leaves the pointer at
(100,100), and the next attempt's `xdotool mousemove --sync` to the same point
waits for a motion event that never comes, so every retry dies at the focus
step with a 10s timeout — which is why the lane reported step 7, not the
failing wait.
New tail: `focus 100 100` (R40 evidence + survival assert), then
`click "label=1"` — a resolved press inside the window that restores keyboard
focus, proves pointer input lands in the app, and moves the pointer off
(100,100) so retries cannot trip the mousemove no-op hang — then `type "55"`
and `wait "label=155 || text=155 || value=155"`. No button is labelled 155, so
the wait passes only if the typed keystrokes reached the display.
* refactor(type): delete the fallow SDK typeText surface, drop dead surface fields
Thermo-nuclear review follow-ups (reviewed at 82fb8c2dc; the three-hop relay
it names was already deleted in
|
||
|
|
46eff36f85 |
refactor: migrate focus to the request-bound device runtime (#1925)
* refactor: migrate focus to the request-bound device runtime Wave 5's first unit (#1739, ADR 0019). `focus x y` and `find <q> focus` now reach the device through one admitted, request-bound `focusPoint` operation instead of the `handleFocusCommand` interactor leaf and its dispatch-table arm. - New `FocusRuntimeOperations` contract with local and provider interactor binders, mirroring the screenshot/element-text seam rather than inventing a second way for one operation class to reach its mechanics. - Exact-owner facts replace the capability bucket: apple simulator/device, android emulator/device/unknown, harmonyos emulator/device, linux device, web device, vega none, providers wherever their interactor is reachable. That is the retired bucket's cell table, restated as facts. - `focus` leaves BASE_COMMAND_CAPABILITY_MATRIX and both hand-maintained overlays (HARMONYOS_SUPPORTED_COMMANDS, WEB_INTERACTION_COMMANDS). - R40 is the new parametrized cutover row; `focusPoint` has exactly one owner. - The `x y` positional parse moves to utils and is shared with the still-legacy touch siblings, so a migrated command cannot drift from them. `find` stays legacy: this unit owns its focus leg only, its `type` leg still dispatches, and R35 waits on the Wave 5 `type` unit. * test(focus): cover the owning interactor binders, lower the find ratchet Review follow-ups on #1925. P1: focus-runtime.test.ts bound a fake focusPoint, so deleting the interactor call inside bindLocalFocusInteractor left focus a successful no-op with every test green. Adds packages/contracts/src/focus-runtime.test.ts, which executes both binders and asserts resolver context, positional (x, y) forwarding, the structured missing-provider failure, and that an already-cancelled request never resolves an interactor at all. Two planted mutants confirm it bites: removing `await interactor.focus(input.point.x, input.point.y)` and transposing its two arguments each fail exactly the two forwarding tests, while the daemon-level focus and find suites stay green — which is the gap the reviewer named. Coverage: find.test.ts shrank to 1204 lines when its focus assertion moved off the dispatch mock; the ratchet pin follows it down. * test(focus): add live Linux focus coverage to the desktop replay The Linux `focus` claim rested on the provider scenario at command-contract level. The desktop replay runs on real Linux hardware in the Smoke lane, so it now runs a coordinate focus and re-asserts the session survived it. Coordinate, not selector: the step exists to prove the migrated `focusPoint` path executes on real hardware, so it must not be able to fail on match ambiguity or CI layout drift. Reclassifies focus contract -> live in the Linux coverage manifest and updates the two pinned counts. The manifest gate is two-sided — a live claim must name a command the replay actually invokes — so the claim cannot drift from the file. |
||
|
|
59d28e8446 |
refactor: add provider-first device lab tests (#542)
* refactor: add provider-first device lab tests * refactor: tighten device lab provider seams * test: cover provider lab contracts * docs: record device lab harness direction * ci: run device lab integration tests * test: move device lab under integration * test: extract device lab helpers * refactor: centralize apps filter defaults * test: drop lab-covered unit tests * test: fold platform happy paths into device lab * test: reuse device lab helpers * test: move device lab to in-process harness * test: replace session handler cases with device lab * test: harden device lab scenario contracts * docs: define unit test retention policy * test: expand provider device lab coverage * test: harden provider device lab coverage * test: cover manifest install and runner session contracts * chore: remove unused provider cleanup code * test: split android find device lab scenario * test: track provider lab architecture progress * test: clarify provider lab roadmap progress * test: advance provider lab session coverage * test: move menubar click routing to device lab * test: move menubar snapshots to device lab * refactor: centralize screenshot flag plumbing * refactor: colocate screenshot flag metadata * test: cover all public commands in device lab * test: move macos wait success to device lab * test: drop redundant perf and diff units * test: move push payload paths to device lab * test: move network parsing to device lab * test: move log cleanup to device lab * test: move log restart and boot to device lab * test: move ios physical boot to device lab * test: cover perf startup in device lab * test: extract android and ios device lab worlds * test: trim device lab world surface * test: split snapshot capture unit coverage * test: deepen device lab coverage and trim handler units * test: clean up device lab migration scaffolding * test: report device lab public command coverage * refactor: make Apple provider seams semantic * refactor: tighten device inventory and Linux provider seams * refactor: tighten request provider scoping * refactor: add semantic macos host provider * test: broaden device lab find coverage * test: cover workflow flags in device lab * refactor: promote linux input provider seam * test: clarify device lab flag coverage * test: classify snapshot force-full progress * test: enforce device lab progress in ci * test: stabilize device lab ci * test: move packaged metro smoke to integration * test: drop stale provider seam coverage * test: harden provider scope regression coverage * refactor: remove stale platform barrels * refactor: keep linux clipboard and screenshots semantic * refactor: move macos host tools behind provider * fix: honor remote artifact output paths * test: deepen runtime coverage for daemon and runner paths * test: share loopback test helpers * refactor: make daemon runtime importable * fix: honor replay target metadata * chore: tighten final device lab quality gates * test: share device lab setup helpers * test: remove generic apple lab fallback * test: deduplicate device lab helpers * chore: tighten fallow duplication signal * refactor: share apple diagnostic helpers * fix: detect active android ime during fill verification * test: consolidate provider-backed integration suite * ci: fix fallow and iOS smoke setup * chore: consolidate cleanup after ci fixes * test: split vitest unit and integration projects * docs: mention MCP discovery metadata * docs: add agent skills context pointers * fix: close provider recording coverage gaps * fix: restore mcp compatibility smoke * test: cover provider edge regressions * test: consolidate loopback helpers * docs: remove stale provider routing reference * fix: harden final provider review issues * chore: defer mcp cleanup from provider refactor |
||
|
|
caf0e834b8 |
feat: add Linux desktop automation support via AT-SPI2 (#356)
* feat: add Linux desktop automation support via AT-SPI2 (Phase 1+2) Add Linux as a first-class platform using AT-SPI2 accessibility framework via node-gtk for accessibility tree snapshots. This mirrors the macOS desktop automation approach using accessibility snapshots. New files: - src/platforms/linux/atspi-bridge.ts: Core AT-SPI2 bridge using node-gtk with lazy loading, recursive tree traversal (max 1500 nodes, depth 12) - src/platforms/linux/role-map.ts: AT-SPI2 role normalization (~100 roles mapped to existing snapshot type conventions) - src/platforms/linux/snapshot.ts: Snapshot entry point with surface, scope, depth, and interactive-only filtering support - src/platforms/linux/devices.ts: Local device discovery for Linux - src/platforms/linux/node-gtk.d.ts: Type declarations for node-gtk Integration: - Extended Platform type with 'linux', backend union with 'linux-atspi' - Wired snapshot into dispatch.ts and snapshot-capture.ts - Added Linux device discovery to dispatch-resolve.ts - Added stub interactor (input actions deferred to Phase 3) - Added 'linux' to CLI --platform flag - node-gtk added as optional dependency (only installs on Linux) https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * refactor: code review cleanup for Linux platform support - Extract SnapshotBackend type alias to replace repeated string union across 5 files (snapshot.ts, snapshot-capture.ts, session-replay-heal.ts, interaction.test.ts) - Remove duplicate scope/interactive/depth filtering from linux/snapshot.ts — let the existing buildSnapshotState pipeline handle it, same as Android - Extract isDesktopBackend() helper in snapshot-capture.ts to consolidate the "skip mobile semantics" pattern for macos-helper and linux-atspi - Collapse 17 repetitive throw statements in Linux interactor stubs into a linuxStub() factory function https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * feat: Linux input synthesis, screenshots, and app lifecycle (Phase 3+4) Add xdotool/ydotool input actions (tap, swipe, scroll, type, fill, right/middle click, long press, double click), screenshot capture via grim/scrot, and app lifecycle management (open, close, back, home). Wire Linux interactors with real implementations and fix device discovery order so Linux doesn't displace Android in auto-selection. https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * refactor: consolidate Linux env detection, simplify input actions - Extract linux-env.ts with cached display server + input tool detection so every action avoids repeated `which` lookups - Add moveTo/clickButton/sendKey helpers to eliminate repeated mousemove boilerplate across 5 mouse actions - Make scrollLinux respect amount/pixels options instead of hardcoded scroll count - Have backLinux/homeLinux reuse sendKey instead of duplicating tool detection https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * feat: add Linux CI smoke test with Xvfb and AT-SPI2 Add GitHub Actions workflow that boots a virtual X11 display (Xvfb), installs AT-SPI2 accessibility tooling and xdotool, opens gnome-calculator, takes screenshots, and captures an accessibility snapshot. Screenshots are uploaded as artifacts for visual verification. Also adds 'linux' to replay script metadata platforms and a test:replay:linux script to package.json. https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * fix: remove pre-session screenshot from Linux replay test The replay runner requires an active session before any commands can run. Move the screenshot after the open command that creates the session. https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * fix: Linux CI — add missing node-gtk build deps and AT-SPI2 env - Add gobject-introspection, libcairo2-dev, build-essential for node-gtk native compilation - Split AT-SPI2 registry start into its own step so it picks up DBUS_SESSION_BUS_ADDRESS from GITHUB_ENV - Set GTK_A11Y=atspi, GTK_MODULES=gail:atk-bridge, NO_AT_BRIDGE=0 to ensure GTK apps expose their accessibility tree on headless CI - Set GSETTINGS_BACKEND=memory to avoid dconf failures - Add node-gtk verification step to catch build failures early https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * fix: explicitly rebuild node-gtk native module in Linux CI pnpm install silently skips failed optional dependency builds and the pnpm cache may not include the native binary. Force a rebuild after install to ensure the node-gtk .node binding is compiled against the system GI/cairo headers. https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * fix: use node-pre-gyp directly to build node-gtk from source pnpm rebuild doesn't trigger node-pre-gyp properly for optional deps. Run node-pre-gyp install --fallback-to-build --update-binary directly inside the node-gtk package directory to force compilation when no prebuilt binary exists for the current Node ABI (v127 / Node 22). https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * refactor: replace node-gtk with Python subprocess for AT-SPI2 node-gtk is a native C++ addon that requires compilation against specific Node ABI versions and GObject Introspection headers. This proved unreliable on CI (no prebuilt binaries for Node 22 ABI v127, silent optional dep build failures, pnpm cache staleness). Replace it with a Python helper script (atspi-dump.py) that uses PyGObject — the reference GObject Introspection consumer. python3-gi is trivially installable on any Linux distro with no compilation step. The Node bridge spawns `python3 atspi-dump.py` and parses JSON output. - Remove node-gtk from optionalDependencies - Remove node-gtk.d.ts type stub - Add atspi-dump.py (~200 lines) doing the same tree traversal - Rewrite atspi-bridge.ts to use subprocess instead of in-process GI - Simplify CI workflow: no more native build deps or rebuild steps https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * chore: drop pre-installed packages from Linux CI apt-get python3-gi, gir1.2-atspi-2.0, at-spi2-core, and dbus-x11 are already present on Ubuntu GitHub Actions runners. https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * feat: surface support for Linux, unit tests, stronger CI assertions - Allow --surface desktop and --surface frontmost-app on Linux (previously only macOS could use --surface) - Add unit tests for atspi-bridge (9 tests: JSON parsing, role normalization, null coercion, error handling, arg forwarding) - Add unit tests for role-map (3 tests: common roles, case normalization, PascalCase fallback) - Improve .py script path resolution (walk upward instead of hardcoded relative paths) - CI replay test now asserts snapshot contains calculator UI nodes via is-exists https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * docs: add cross-platform snapshot traversal contract Document the shared schema, traversal rules, surface semantics, and normalized role types that all snapshot backends (Swift, Python, Android) must conform to. This serves as the single source of truth when adding or modifying platform backends. https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * fix: update lockfile after removing node-gtk optional dependency pnpm-lock.yaml still referenced node-gtk after it was removed from package.json, causing pnpm install --frozen-lockfile to fail in CI. https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * fix: address review findings in Linux platform code - atspi-dump.py: use ctx dict for traversal limits instead of globals, fix rect filter (width/height <= 0 should use `or`), add surface validation - input-actions.ts: make sendKey scancodes required to prevent silent no-op on ydotool, fix ydotool longPress/swipe to use click --down/--up - app-lifecycle.ts: use pkill -x (exact match) instead of pkill -f - linux-env.ts: emit diagnostic warning when falling back to xdotool on Wayland https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * fix(ci): explicitly install all Linux a11y dependencies Ubuntu runners may not have at-spi2-core, python3-gi, gir1.2-atspi-2.0, or dbus-x11 pre-installed. Install them explicitly instead of assuming they exist. Also make the verify step's tree dump non-fatal since no apps are running at that point. https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * fix: quote multi-word role value in Linux smoke test selector The selector parser tokenizes on whitespace, so `role=push button` was split into two tokens causing a parse failure. Use single quotes inside the selector: `role='push button'`. https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * fix: use valid selector keys in Linux smoke test appName is not a valid selector key. The supported keys are: id, role, text, label, value, visible, hidden, editable, selected, enabled, hittable. Simplified to use label and role only. https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * chore: cleanup pass — menubar warning, fix contract doc example - snapshot.ts: emit diagnostic warning when menubar surface is requested on Linux (falls back to desktop silently otherwise) - SNAPSHOT_CONTRACT.md: fix unmapped role example to use a role that isn't actually mapped (was "color chooser" which maps to Dialog) https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * chore: add Python bytecache to gitignore https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * feat: harden Linux platform — capability matrix, CI, error handling P0: Add explicit Linux capability matrix with 3-way platform routing (Apple/Linux/Android) in isCommandSupportedOnDevice. Linux now correctly blocks unsupported commands (clipboard, rotate, scrollIntoView, etc.) at capability level rather than throwing at runtime. Includes tests. P0: Expand Linux CI to run typecheck + unit tests before smoke tests. Add AT-SPI2 registry health probe with fail-fast on missing registry. P1: Harden atspi-dump.py — arg parsing now produces JSON errors on bad int values, and a top-level catch wraps unexpected exceptions in JSON. P1: Add 10s per-action timeout to xdotool/ydotool input commands to prevent indefinite hangs. P1: Tighten smoke test selectors to calculator-specific signals (digit labels) instead of generic role='push button'. P2: Document Linux surface mapping, supported commands, and known limitations in SNAPSHOT_CONTRACT.md. https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * fix: apply depth/interactive filtering to Linux snapshots Linux snapshots were bypassing snapshotInteractiveOnly and snapshotDepth filtering that macOS-helper gets via shapeDesktopSurfaceSnapshot. Route Linux through the same function so snapshot -i and --depth flags work. Renamed shapeMacOsSurfaceSnapshot → shapeDesktopSurfaceSnapshot since it's now shared between macOS and Linux desktop backends. https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * fix: address review findings — error reporting, Wayland, timeout - app-lifecycle.ts: emit diagnostic on fire-and-forget app launch failure instead of silently swallowing errors - linux-env.ts: make xdotool on Wayland a hard error instead of a broken fallback (xdotool doesn't work on Wayland) - atspi-bridge.ts: increase Python subprocess timeout from 15s to 30s for safety on slow/loaded systems with large a11y trees https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * feat: appName/windowTitle selectors, clipboard, input-action tests Selectors: - Add appname and windowtitle as selector keys for desktop platforms. Both macOS and Linux snapshots already populate these fields — now they're usable in selector expressions (e.g., "label=OK appname=Calc"). Keys are case-insensitive. Clipboard: - Implement readLinuxClipboard/writeLinuxClipboard using xclip/xsel (X11) or wl-copy/wl-paste (Wayland) with descriptive TOOL_MISSING errors. Enable clipboard in Linux capability matrix. 7 unit tests. Input action tests: - Add 18 unit tests covering xdotool and ydotool code paths: press, right/middle click, double click, sendKey, type, scroll, swipe, focus, fill. Tests mock runCmd and verify correct tool + args. https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT * chore: cache tool detection for screenshot/clipboard, extract get_app_info helper Avoid repeated `which` calls on every screenshot/clipboard operation by caching the resolved tool on first use, matching the input-action pattern. Extract duplicated app_name/pid retrieval in atspi-dump.py into get_app_info. https://claude.ai/code/session_01H9hrmueNF5pcBM8JeX81mT --------- Co-authored-by: Claude <noreply@anthropic.com> |