mirror of
https://github.com/callstack/agent-device.git
synced 2026-09-14 20:06:34 +08:00
codex/runner-diagnostics-proxy-docs
4 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
3a02e514e1 |
refactor(ios): consolidate series batching onto the sequence runner command (#768)
* refactor(ios): consolidate series batching onto the sequence runner command Closes #767 Routes every Apple multi-press variant (plain, double-tap, hold, jitter) and swipe series through budget-chunked sequence requests, retiring the daemon-side tapSeries and dragSeries senders: - Add a doubleTap step kind to the sequence allowlist on both ends, mirroring the retired tapSeries doubleTapAt branch. - The single doubleTap interactor sends a one-step sequence and parses the result, surfacing step failures as errors. - Swipe series unroll ping-pong daemon-side into per-step endpoints; the runner's coordinate-drag path ignores durationMs exactly as the daemon-sent (non-synthesized) dragSeries did. - Extract runIosSequenceChunks so press and swipe share the chunking, aggregation, and global step-index rebasing. - Keep tapSeries/dragSeries runner handlers for wire compatibility with older daemons, annotated like interactionFrame; remove both from the preflight-skip allowlist (daemon never sends them) and update ADR 0005 / protocol-optimizations docs. This also closes the latent watchdog exposure where press --count N --interval-ms M routed to tapSeries and executed all pauses inside one 30s-watchdog main-thread block with no chunking. Behavior note: plain tap series now use the synthesized HID tap path on iOS non-tv (with runner-side tapAt fallback), matching the individual tap command instead of the retired tapSeries' XCUICoordinate taps. https://claude.ai/code/session_01VokBZWESTDgcnbYwS4DkJo * refactor(ios): drop dead series wire surface from the daemon - Remove chunkRunnerSequenceSteps: superseded by the budget-aware chunker; no production callers remained. - Remove tapSeries/dragSeries from the RunnerCommand union along with their orphaned fields (count, intervalMs, doubleTap, pauseMs, pattern) and protocol fixtures: this type is the send surface of the current daemon, which no longer sends either command. The Swift runner keeps serving both for wire compatibility with older daemons. - Retarget the ready-mutation preflight test from tapSeries to sequence. https://claude.ai/code/session_01VokBZWESTDgcnbYwS4DkJo * refactor(ios): remove retired series and frame wire commands entirely Drops the runner-side wire compatibility for tapSeries, dragSeries, and interactionFrame now that no daemon path sends them (series fuse into sequence since this branch; interactionFrame was fused into scroll in #760): - Swift: delete the three handler cases, performDragSeries, runSeries (no remaining callers), the CommandType enum cases, journal-retention and traits entries, and the Command fields (count, intervalMs, doubleTap, pauseMs, pattern) that existed only for them. The never-sent synthesized dragSeries branch goes with it. - TS: drop interactionFrame from the RunnerCommand union and isReadOnlyRunnerCommand, and its protocol fixture. - Update stale perf scenario labels referencing the retired commands. Verified dead before removal: no dynamic command construction anywhere (runner-command-recovery only echoes in-flight command ids), no raw-string references in Swift, no docs references. Helpers shared with live paths (synthesizedDragAt, doubleTapAt, keyboardAvoidingDragPoints, sleepFor) all retain callers. Compat: an old daemon paired with a runner built from these sources gets a CommandType decode rejection; the source-fingerprint check rebuilds a matching runner on the next session. https://claude.ai/code/session_01VokBZWESTDgcnbYwS4DkJo --------- Co-authored-by: Claude <noreply@anthropic.com> |
||
|
|
5492cf4642 |
refactor(ios): single CommandTraits table for runner command classification (#642)
* refactor(ios): single CommandTraits table for runner command classification Replace the three hand-maintained switches in RunnerTests+Lifecycle.swift (isInteractionCommand / isReadOnlyCommand / isRunnerLifecycleCommand) with one source of truth: CommandType.traits, an exhaustive switch returning a CommandTraits struct (interaction / readOnly / lifecycle axes), collocated with CommandType in RunnerTests+Models.swift. Pure refactor: every command's classification is reproduced verbatim, and the three predicates become one-line lookups with unchanged signatures, so call sites are untouched. The exhaustive switch makes it a compile error to add a CommandType without classifying it, closing the drift that historically let tapSeries/dragSeries/keyboardReturn fall out of isInteractionCommand. readOnly is a 3-state enum (.always/.never/.conditional); .conditional preserves alert's action-dependent read-only behavior, resolved in isReadOnlyCommand. Classification feeds ADR-0002 session invalidation (the read-only retry that nulls currentApp/currentBundleId), so behavior is intentionally unchanged. Adds the "Runner command traits" term to CONTEXT.md. * docs(ios): note CommandTraits.readOnly .conditional is alert-only (review follow-up) * fix(ios): classify tapSeries/dragSeries/keyboardReturn as interaction commands (#643) * fix(ios): classify tapSeries/dragSeries/keyboardReturn as interaction commands tapSeries and dragSeries are the series forms of tap/drag (already interaction commands); keyboardReturn is the sibling of keyboardDismiss (already an interaction command). All three were missing from the historical isInteractionCommand switch — a drift the new CommandTraits table (#642) makes visible. Classifying them as interaction commands gives them the foreground-guard + stabilization preflight that their single-shot/sibling forms already get. Behavior change: these three commands now re-activate a backgrounded target to foreground and pay the stabilization delays before running. Ships separately from the CommandTraits refactor (#642) and should land after that bakes. mouseClick left unchanged: macOS-only and the foreground guard interacts with bespoke macOS activation, so it needs a macOS smoke check first. * test: cover iOS runner series commands in perf harness |
||
|
|
2068f604bb | fix: improve ios selector reads and maestro reliability (#636) | ||
|
|
45cfad5cc5 |
feat: e2e command perf benchmark harness + nightly CI (#630)
* feat: add e2e command perf benchmark harness + nightly CI Adds scripts/perf, a cheap end-to-end perf benchmark that drives the built CLI through an ordered Settings tour of ~24 commands for N rounds, on a fully isolated daemon/state-dir and self-cleaning device, and emits JSON + Markdown reports. Per-command timing comes from wrapping each batchable command in its own single-step batch (daemon durationMs) plus wall-clock around the process. Wires a scheduled + workflow_dispatch CI job (perf-nightly.yml) that reuses the cached iOS XCUITest runner (setup-apple-replay) and the Android replay host, and runs the CLI from source via --experimental-strip-types (no dist build). * refactor(perf): drive the harness CLI via runCmdSync, not spawnSync Review (P2): repo rule is to spawn processes through src/utils/exec.ts, not node:child_process directly. Switch the perf harness's invokeCli to runCmdSync (allowFailure so non-zero exits are recorded as samples) and add a maxBuffer option to ExecOptions/runCmdSync (snapshot payloads exceed Node's ~1MB default). * perf(harness): warm the runner after open so the first measured command is clean The first interaction after open/relaunch pays the one-time iOS XCUITest runner startup (~10s+ cold) and a per-relaunch first-AX-query settle cost (~4s). That was landing on the first measured command each round (snapshot -i), inflating it ~10x vs the next snapshot. Run an untimed warmup snapshot -i after establishSession, after each round's reset-open, and after every freshRoot relaunch, so no measured command absorbs runner startup. Noted in the report header. * refactor(perf): address review + fix Fallow CI - exec.ts: extract spawnRejectionError + commandCloseFailure helpers, deduping the error/close handler clones (Fallow duplication ✗ that surfaced once the maxBuffer change pulled exec.ts into the audit scope). - .fallowrc: exclude scripts/perf/** (non-shipped benchmark tooling, like examples/ test-app) so its naturally-moderate functions don't trip the complexity gate. - config.ts: drop unused exports CLI_BIN/DEFAULT_OUT_DIR; add readIntValue so --n/--rounds/--warmup report the actual flag + reject non-integers clearly. - harness.ts: extract toSample(); type sampleError param as CliResult. - scenario.ts: ScenarioStep is now a discriminated union on execMode (removes step.step!/ step.args ?? []). - comment/legend rewords (platform defaults are local-convenience/CI-overridden; elements = node count). check:fallow now green; typecheck/lint/unit pass. * perf(harness): downgrade sample ok when a batch step reports ok:false Defensive belt-and-suspenders for the Codex review note: stop-only batch already surfaces a failed step as a top-level failure (caught by invokeCli), but if an on-error=continue mode ever keeps the batch ok while a step fails, don't silently count that step as a successful sample — derive ok from the step's own result.ok. |