Files
callstack__agent-device/scripts/perf/platform-profiles.ts
Michał Pierzchała 45cfad5cc5 feat: e2e command perf benchmark harness + nightly CI (#630)
* feat: add e2e command perf benchmark harness + nightly CI

Adds scripts/perf, a cheap end-to-end perf benchmark that drives the built
CLI through an ordered Settings tour of ~24 commands for N rounds, on a fully
isolated daemon/state-dir and self-cleaning device, and emits JSON + Markdown
reports. Per-command timing comes from wrapping each batchable command in its
own single-step batch (daemon durationMs) plus wall-clock around the process.

Wires a scheduled + workflow_dispatch CI job (perf-nightly.yml) that reuses the
cached iOS XCUITest runner (setup-apple-replay) and the Android replay host, and
runs the CLI from source via --experimental-strip-types (no dist build).

* refactor(perf): drive the harness CLI via runCmdSync, not spawnSync

Review (P2): repo rule is to spawn processes through src/utils/exec.ts, not
node:child_process directly. Switch the perf harness's invokeCli to runCmdSync
(allowFailure so non-zero exits are recorded as samples) and add a maxBuffer
option to ExecOptions/runCmdSync (snapshot payloads exceed Node's ~1MB default).

* perf(harness): warm the runner after open so the first measured command is clean

The first interaction after open/relaunch pays the one-time iOS XCUITest runner
startup (~10s+ cold) and a per-relaunch first-AX-query settle cost (~4s). That was
landing on the first measured command each round (snapshot -i), inflating it ~10x
vs the next snapshot. Run an untimed warmup snapshot -i after establishSession, after
each round's reset-open, and after every freshRoot relaunch, so no measured command
absorbs runner startup. Noted in the report header.

* refactor(perf): address review + fix Fallow CI

- exec.ts: extract spawnRejectionError + commandCloseFailure helpers, deduping the
  error/close handler clones (Fallow duplication ✗ that surfaced once the maxBuffer
  change pulled exec.ts into the audit scope).
- .fallowrc: exclude scripts/perf/** (non-shipped benchmark tooling, like examples/
  test-app) so its naturally-moderate functions don't trip the complexity gate.
- config.ts: drop unused exports CLI_BIN/DEFAULT_OUT_DIR; add readIntValue so
  --n/--rounds/--warmup report the actual flag + reject non-integers clearly.
- harness.ts: extract toSample(); type sampleError param as CliResult.
- scenario.ts: ScenarioStep is now a discriminated union on execMode (removes step.step!/
  step.args ?? []).
- comment/legend rewords (platform defaults are local-convenience/CI-overridden;
  elements = node count). check:fallow now green; typecheck/lint/unit pass.

* perf(harness): downgrade sample ok when a batch step reports ok:false

Defensive belt-and-suspenders for the Codex review note: stop-only batch already
surfaces a failed step as a top-level failure (caught by invokeCli), but if an
on-error=continue mode ever keeps the batch ok while a step fails, don't silently
count that step as a successful sample — derive ok from the step's own result.ok.
2026-05-31 14:37:59 +02:00

78 lines
3.1 KiB
TypeScript

import type { PerfConfig } from './config.ts';
import type { Platform } from './types.ts';
// Local-convenience defaults for ad-hoc runs; CI always overrides them (--device / --serial).
// The iOS UDID is a specific local "iPhone 17" sim; the Android serial is a dedicated emulator
// port. Pass --udid/--device/--serial to target your own device.
const DEFAULT_IOS_UDID = 'D74E0B66-57EB-4EC1-92DC-DA0A30581FE7';
const DEFAULT_ANDROID_SERIAL = 'emulator-5556';
export type ProfileSelectors = {
// A row on the Settings root that pushes a large sub-screen (big a11y tree).
deepScreen: string;
// The Settings search field (for press/focus; auto-picks a match).
searchField: string;
// A selector that uniquely targets the EDITABLE search field (for fill).
searchFieldEditable: string;
// iOS exposes an editable search field at the Settings root (fill works without focusing
// first; focusing then filling can hang). Android only reveals the editable after tapping
// the search card, so it must press the search entry before fill/type.
searchEditableAtRoot: boolean;
// A label reliably visible on the Settings root, for get/is (selector form).
anchorLabel: string;
// Plain text of the anchor, for wait text / find (not a selector).
anchorText: string;
};
export type ResolvedProfile = {
platform: Platform;
deviceName: string;
udid?: string;
serial?: string;
platformFlags: string[]; // --platform; applied to every call (only conflicts if it mismatches a locked session)
selectorFlags: string[]; // device selectors — ONLY on the session-establishing open / selectorless boot
appTarget: string; // `open` target for Settings
selectors: ProfileSelectors;
};
export function resolveProfile(cfg: PerfConfig): ResolvedProfile {
if (cfg.platform === 'ios') {
// Prefer targeting by device name (CI boots a named simulator); fall back to a UDID.
const useName = cfg.device !== undefined;
const udid = useName ? undefined : (cfg.udid ?? DEFAULT_IOS_UDID);
return {
platform: 'ios',
deviceName: cfg.device ?? 'iPhone 17',
udid,
platformFlags: ['--platform', 'ios'],
selectorFlags: useName ? ['--device', cfg.device!] : ['--udid', udid!],
appTarget: 'settings',
selectors: {
deepScreen: 'label="General"',
searchField: 'label="Search"',
searchFieldEditable: 'label="Search" editable',
searchEditableAtRoot: true,
anchorLabel: 'label="General"',
anchorText: 'General',
},
};
}
const serial = cfg.serial ?? DEFAULT_ANDROID_SERIAL;
return {
platform: 'android',
deviceName: cfg.serial ? `android (${serial})` : 'Pixel_9_Pro_XL_API_37',
serial,
platformFlags: ['--platform', 'android'],
selectorFlags: ['--serial', serial, '--android-device-allowlist', serial],
appTarget: 'com.android.settings',
selectors: {
deepScreen: 'text="Network & internet"',
searchField: 'text="Search Settings"',
searchFieldEditable: 'editable',
searchEditableAtRoot: false,
anchorLabel: 'label="Network & internet"',
anchorText: 'Network & internet',
},
};
}