Files
callstack__agent-device/scripts/lib/contention-retry.ts
Michał Pierzchała ac9e4d0f04 test: measure oracle liveness suite-wide; pin the one dead-path oracle (#1679)
* test: measure oracle liveness suite-wide; pin the one dead-path oracle

- docs/agents/oracle-negation-spike.md: assertion-negation sweep over 677
  test files (5,687 verdicts). Zero vacuous tests: all 150 negation
  survivors decompose into assert.rejects-validator artifacts (112),
  helper-oracles (31), in-file-fake breakage (6), and one conditional
  oracle. Records the companion mock-coupled coverage-uniqueness numbers
  and the follow-ups they motivate (diff-scoped mutation gate, provider
  seam closures, transcript provenance).
- watchos-sentinel: the non-watchOS test's only assertion sat in a catch
  block that never fires (tvOS interactor creation succeeds), so no
  assertion executed on the observed path. Pin creation success instead.
  Red-run proof: the old shape survived the negation sweep; the new shape
  fails under it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(android): inject fake adb through the provider scope, not PATH stubs

Adds withFakeAdb to test-utils: a scripted in-process AndroidAdbProvider
installed through the production withAndroidAdbProvider seam — the same
scope the daemon installs per request — replacing PATH-stub shell scripts
that spawn a real subprocess per adb call. No PATH mutation, no spawns,
no real subprocess waits.

Converts settings.test.ts (15 tests, 23ms; waiver said "waits real
settings-apply poll time") and notifications.test.ts (2 tests, 9ms).
Assertions move from args-log regex greps to structural checks on the
recorded call list; the fake receives device-scoped args with the
-s serial pair stripped, so serial routing is enforced by the scoped
provider matching device.id instead of asserted per call.

Remaining PATH-stub files convert next; their contention-retry waiver
entries lift together with the conversions.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(android): convert device-input-state to fake adb provider injection

10 PATH-stub cases move to withFakeAdb through the production provider
scope; the 2 tests that already inject an executor directly are
unchanged. Cross-invocation shell STATE_FILE state becomes a closure
boolean; args-log regex asserts become structural checks on recorded
calls. 12/12 green at 386ms — the residue is dismissAndroidKeyboard's
two fixed 120ms retry sleeps, not stub subprocess waits.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(android): convert app-lifecycle-install adb stubbing to fake provider

The adb half of every case moves to withFakeAdb through the production
provider scope; installs take the documented exec-shaped fallback
(exec(['install','-r',...])), matching what the PATH stub saw minus the
serial pair. bundletool/zip/unzip stay real or PATH-stubbed — they run
via runCmd outside the adb seam, so this file remains in the serialized
subprocess-stub lane with its waiver reason to be corrected from adb to
bundletool. 13/13 green at ~130ms; no case enters a retry/poll loop.

Conversion note: manifest identity's `unzip -p` failure is silently
swallowed (readZipEntry catch -> undefined, aapt fallback) — an
invisible degradation path worth a future explicit diagnostic.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(android): convert input-actions adb stubbing to fake provider

9 PATH-stub cases move to withFakeAdb; the 3 tests already injecting
providers directly are unchanged. Chunked shell-input assertions become
ordered deepEqual on the recorded calls; never-called negatives and
call-count checks preserved 1:1. 12/12 green.

File time drops to 2.2s, all of it production sleeps:
verifyAndroidFilledText unconditionally waits its [0,150,350]ms
verification cadence even when the first inspection matches, so each
fill verification pass costs ~500ms with an instant fake. A
budget-derived cadence there (testing.md pattern 1) would put this file
near 25ms; flagged as follow-up rather than changed here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(android): extract shared oracles; make fake adb failure-faithful

Three test-utils extractions applied across the six converted files:

- assertRejectsAppError collapses the hand-rolled AppError code+message
  rejection validator (10 sites here; ~30 more repo-wide can adopt it
  incrementally). Validators asserting details or multiple differently-
  flagged regexes stay explicit on purpose.
- withFakeAdb gains a `provider` option for extra capabilities
  (snapshotHelperArtifact, reverse, ...), replacing input-actions'
  nested re-scoping bridge.
- withFakeAdb now mirrors the local executor's contract: a scripted
  nonzero exit throws androidAdbResultError unless the call site passed
  allowFailure. Provider-scoped exec bypasses exec.ts's throw-on-close-
  failure, so returning {exitCode:1} took a different production path
  than the PATH-stub `exit 1` these fakes replaced. All 75 tests hold
  under the corrected semantics.

Also swaps settings' inline emulator DeviceInfo literals for the shared
ANDROID_EMULATOR fixture.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test: lift five converted Android files from the contention-retry waiver

settings, notifications, device-input-state, input-actions, and
app-lifecycle-open no longer stub binaries on PATH or spawn
subprocesses, so their contention mechanism is gone: they leave
CONTENTION_RETRY_FILES and, through the derived SUBPROCESS_STUB_TESTS
constant, the serialized subprocess-stub project (17 -> 12 files).
app-lifecycle-install stays with its reason corrected: adb is now
in-process, but bundletool stays PATH-stubbed and zip/unzip spawn for
.aab packaging paths.

Full unit suite green at the new membership: 638 files, 5,724 tests,
with the five files running at unit-core's default parallelism.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test: apply review findings to the fake-adb conversion batch

- app-lifecycle-open: the missing-package launch failure returns
  {stderr, exitCode: 1} and lets withFakeAdb's throw path produce the
  production-shaped androidAdbResultError instead of hand-modeling the
  thrown AppError — the drift the helper exists to eliminate.
- withScriptedAdb deleted: the six converted files were its only
  callers, and a live PATH-stub export invites new tests back into the
  serialized lane this batch shrank. withMockedAdb stays (dispatch and
  runtime-hints tests still stub other binaries).
- android-snapshot-helper gains androidSnapshotHelperScriptResponse so
  the version-probe detection and versionCode reply have one source of
  truth; input-actions' local copy delegates to it.
- withFakeAdb's provider option becomes a distributed Omit over the
  AndroidAdbProvider union, so touch without gestureViewport is a
  compile error at the fake's boundary (planted and verified) instead
  of a TypeError inside production gesture planning.
- spike-doc re-run checklist restores wider than the codemod globs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test: drop the consumer-less FakeAdbScript barrel re-export

Fallow's dead-code gate flagged it: scripts are always passed as inline
lambdas, so only FakeAdbResponse needs a name at the barrel. The type
stays exported from fake-adb.ts where the withFakeAdb signature uses it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(apple): inject fake xcrun through the tool-provider scope, not PATH stubs

withFakeAppleTool mirrors withFakeAdb for the Apple seam: a scripted
provider installed via the production withAppleToolProvider scope, flat
simctl/devicectl invocations recorded exactly as the PATH-stub shell
scripts saw them, throw-on-nonzero fidelity matching exec.ts unless the
call site passed allowFailure, and the canned `simctl privacy help`
listing served by default (the block withMockedXcrun injected into
every script). screenshot-status-bar.test.ts converts as the exemplar:
3/3 green at 9ms with deepEqual call-sequence assertions replacing the
args-log regexes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test(apple): convert apps.test.ts xcrun stubbing to fake tool provider

All withMockedXcrun scripts and hand-rolled PATH stubs move to
withFakeAppleTool; args-log regexes become structural call assertions
(exact deepEqual where order is deterministic, presence checks where
the 5s simulatorBootedMemo TTL makes boot-probe order test-dependent).
12 hand-rolled AppError validators collapse into assertRejectsAppError.
54/54 green; file test time 1172ms -> ~400ms with no test over 201ms.

Five .ipa install tests keep a minimal PATH stub for unzip only:
install-artifact.ts:112 and install-source.ts:438 call runCmd('unzip')
directly, outside the Apple tool provider seam — the file therefore
stays in the serialized subprocess-stub lane with its waiver reason
corrected from xcrun to unzip.

Also observed: getSimctlPrivacyServices caches per PATH+simulatorSetPath
and simulatorBootedMemo keys on deviceId|setPath, so neither cache
accounts for the provider scope — worked around per test, follow-up
worthy.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

* test: lift six Apple waivers; fix format and fallow findings from CI

- interactions, simulator, screenshot, physical-device-screenshot,
  devicectl, and screenshot-status-bar leave CONTENTION_RETRY_FILES:
  the first five stopped stubbing PATH binaries in earlier refactors
  (measured 3-64ms per file, no subprocess activity), and
  screenshot-status-bar now injects through the fake tool provider.
  apps.test.ts stays with its reason corrected to the unzip PATH stub
  (xcrun is in-process; install-artifact.ts:112 / install-source.ts:438
  call runCmd('unzip') outside the Apple seam). Serialized lane 12 -> 6.
- oxfmt: fake-apple-tool.ts and contention-retry.ts were pushed
  unformatted (local check piped through tail masked the failure).
- fallow complexity: the three fake-script arrows in apps.test.ts drop
  under threshold via shared predicates (isSimctlMainScreenScale,
  isSimctlScreenshot, isDevicectlDevice), which also deduplicate the
  screenshot pair.

Full unit suite green at the new membership: 638 files, 5,724 tests.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019S5sZnmPn4A9Ct7sTJdfAf

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-08 09:01:05 +02:00

321 lines
12 KiB
TypeScript

// Single-retry policy for contention-flaky test files (issue #1419, umbrella #1412 Track B).
//
// Some unit-test files stub a real binary and then spawn or wait on it, so under
// host load their production code takes a generic timeout path and the file
// fails for a reason that has nothing to do with the diff (AGENTS.md, "Known
// environment traps"). Rerunning the whole suite hides real regressions and
// burns runner minutes, so the policy is deliberately narrow:
//
// - the retry set is ENUMERATED here, never a glob: a glob would silently
// enroll every future file under a directory,
// - only tests the runner itself aborted at their timeout retry; an assertion
// failure in a listed file fails the job on the first run,
// - one retry, of the failed files only, and every retry is reported in the
// job summary so a permanently flaky file cannot hide behind a green check,
// - each entry is an owned waiver in the ADR 0011 sense: a tracking issue plus
// a review date, and an expired entry fails the gate until it is renewed or
// removed.
import { runnerTimedOut } from './runner-timeout-meta.ts';
/** One enumerated retry-eligible file, owned like an ADR 0011 waiver. */
export type ContentionRetryEntry = {
/** Repo-relative test file path. Exact path — never a glob. */
file: string;
/**
* Why this file spawns or waits. Required: the retry list may only grow with a
* concrete contention mechanism named at the entry, not by habit.
*/
reason: string;
/** Issue tracking the removal of the underlying real-time wait. */
trackingIssue: string;
/** ISO date (YYYY-MM-DD). Past this, the gate fails until renewed or removed. */
reviewBy: string;
/**
* True for files that also run serialized in the `subprocess-stub` Vitest
* project: they stub a binary on PATH, so two of them running at once race
* over the same stub budget. Serialization bounds that contention; it does not
* remove the real waits, which is why they are retry-eligible too.
*/
serializedStub?: true;
};
const SUBPROCESS_STUB_ISSUE = 'https://github.com/callstack/agent-device/issues/1098';
const CONTENTION_ISSUE = 'https://github.com/callstack/agent-device/issues/1419';
const FUZZ_ISSUE = 'https://github.com/callstack/agent-device/issues/1414';
const REVIEW_BY = '2026-10-31';
/**
* The retry-eligible set: every `SUBPROCESS_STUB_TESTS` file (vitest.config.ts —
* serializing them bounds contention but does not remove the real waits) plus
* the repeat offenders observed failing timeout-shaped on unrelated diffs.
* `contention-retry-policy.test.ts` asserts the subprocess-stub half stays in
* sync with the vitest projects.
*/
export const CONTENTION_RETRY_FILES: readonly ContentionRetryEntry[] = [
{
file: 'src/platforms/android/__tests__/app-lifecycle-install.test.ts',
reason:
'Stubs bundletool on PATH (adb is in-process) and spawns zip/unzip for .aab packaging paths.',
trackingIssue: SUBPROCESS_STUB_ISSUE,
reviewBy: REVIEW_BY,
serializedStub: true,
},
{
file: 'src/daemon/__tests__/runtime-hints.test.ts',
reason: 'Stubs platform binaries on PATH and spawns them to derive runtime hints.',
trackingIssue: SUBPROCESS_STUB_ISSUE,
reviewBy: REVIEW_BY,
serializedStub: true,
},
{
file: 'src/platforms/apple/core/__tests__/apps.test.ts',
reason:
'Stubs unzip on PATH for .ipa extraction (xcrun is in-process) and spawns it per install case.',
trackingIssue: SUBPROCESS_STUB_ISSUE,
reviewBy: REVIEW_BY,
serializedStub: true,
},
{
file: 'src/__tests__/client-metro.test.ts',
reason: 'Stubs npx plus the package managers and spawns a real Metro dev server per case.',
trackingIssue: SUBPROCESS_STUB_ISSUE,
reviewBy: REVIEW_BY,
serializedStub: true,
},
{
file: 'scripts/fuzz/harness.test.ts',
reason: 'Spawns a node subprocess or worker per case; one target hangs on purpose (#1414).',
trackingIssue: FUZZ_ISSUE,
reviewBy: REVIEW_BY,
serializedStub: true,
},
{
file: 'scripts/fuzz/corpus-replay.test.ts',
reason: 'Replays the fuzz corpus through the worker watchdog, waiting its per-case budget.',
trackingIssue: FUZZ_ISSUE,
reviewBy: REVIEW_BY,
serializedStub: true,
},
{
file: 'src/daemon/__tests__/request-router-open.test.ts',
reason: 'Drives real open routing over keyed locks, waiting on lock/session settle budgets.',
trackingIssue: CONTENTION_ISSUE,
reviewBy: REVIEW_BY,
},
{
file: 'src/platforms/apple/core/__tests__/runner-client.test.ts',
reason: 'Drives runner transport connect/retry against a real socket, waiting retry backoff.',
trackingIssue: CONTENTION_ISSUE,
reviewBy: REVIEW_BY,
},
{
file: 'src/platforms/apple/core/__tests__/runner-xctestrun.test.ts',
reason: 'Stubs xcodebuild on PATH and spawns it for xctestrun preparation and cache checks.',
trackingIssue: CONTENTION_ISSUE,
reviewBy: REVIEW_BY,
},
{
file: 'scripts/__tests__/help-conformance-bench.test.ts',
reason: 'Spawns the real bench script per case in --dry-run mode and waits for it to exit.',
trackingIssue: CONTENTION_ISSUE,
reviewBy: REVIEW_BY,
},
];
/**
* The `subprocess-stub` Vitest project's file set, derived from the one list so
* the retry policy and the execution contract cannot drift apart.
* `vitest.config.ts` consumes this for both the unit-core exclude and the
* subprocess-stub include.
*/
export const SUBPROCESS_STUB_TESTS: readonly string[] = CONTENTION_RETRY_FILES.filter(
(entry) => entry.serializedStub,
).map((entry) => entry.file);
function isRetryEligibleFile(file: string): boolean {
return CONTENTION_RETRY_FILES.some((entry) => entry.file === normalizeTestFile(file));
}
/** Absolute or `./`-prefixed paths arrive from reporters; the list is repo-relative. */
export function normalizeTestFile(file: string, repoRoot?: string): string {
let normalized = file.replaceAll('\\', '/');
if (repoRoot) {
const prefix = `${repoRoot.replaceAll('\\', '/').replace(/\/$/, '')}/`;
if (normalized.startsWith(prefix)) normalized = normalized.slice(prefix.length);
}
return normalized.replace(/^\.\//, '');
}
/** Entries whose review date has passed, in list order. */
export function expiredRetryEntries(now: Date): readonly ContentionRetryEntry[] {
const today = now.toISOString().slice(0, 10);
return CONTENTION_RETRY_FILES.filter((entry) => entry.reviewBy < today);
}
/** An error as the lane reporter receives it from Vitest. */
export type ReportedError = {
name?: unknown;
message?: unknown;
expected?: unknown;
actual?: unknown;
};
/** A test the runner aborted at its timeout — the only failure this lane retries. */
export function isRunnerTimeout(
errors: readonly ReportedError[],
meta: unknown,
token: string | undefined,
): boolean {
return runnerTimedOut(meta, token) && errors.length > 0 && errors.every(hasNoAssertionDiff);
}
/** A timeout that also failed an assertion is a regression, not contention. */
function hasNoAssertionDiff(error: ReportedError): boolean {
return error.expected === undefined && error.actual === undefined;
}
/** One failed test case as read from a runner report. */
export type TestFailure = {
file: string;
testName: string;
message: string;
/** Classified by the reporter from the live error object, never from text. */
timeout: boolean;
};
/**
* A failure that is not a failed test case — an unhandled error, a module that
* failed to load, a coverage-threshold miss, a nonzero exit nothing else
* explains. These can never be retried away: they block the plan outright, so a
* timeout retry can never erase a second, unrelated failure from the same run.
*/
export type RunBlocker = { kind: string; detail: string };
/**
* Failures as written by the lane's reporter (`contention-retry-reporter.ts`),
* narrowed at the trust boundary. A run that dies before writing the file parses
* to zero failures, which the policy reads as "nothing retry-eligible": the job
* fails, the safe direction.
*/
export function parseFailureReport(
report: unknown,
repoRoot?: string,
): { failures: readonly TestFailure[]; blockers: readonly RunBlocker[] } {
const { failures, blockers } = report as { failures?: unknown; blockers?: unknown };
return {
failures: Array.isArray(failures)
? failures.map((entry) => readFailure(entry, repoRoot)).filter((entry) => entry !== undefined)
: [],
blockers: Array.isArray(blockers)
? blockers.map((entry) => readBlocker(entry)).filter((entry) => entry !== undefined)
: [],
};
}
function readBlocker(entry: unknown): RunBlocker | undefined {
const { kind, detail } = entry as { kind?: unknown; detail?: unknown };
if (typeof kind !== 'string') return undefined;
return { kind, detail: typeof detail === 'string' ? detail : '' };
}
function readFailure(entry: unknown, repoRoot: string | undefined): TestFailure | undefined {
const { file, testName, message, timeout } = entry as Partial<Record<keyof TestFailure, unknown>>;
if (typeof file !== 'string') return undefined;
return {
file: normalizeTestFile(file, repoRoot),
testName: typeof testName === 'string' ? testName : 'unknown test',
message: typeof message === 'string' ? message : '',
timeout: timeout === true,
};
}
export type RetryPlan =
| {
retry: false;
reason: string;
blocked: readonly TestFailure[];
blockers: readonly RunBlocker[];
}
| { retry: true; files: readonly string[]; failures: readonly TestFailure[] };
/**
* Decide whether a failed run may retry. Retry requires *every* failure to be a
* timeout in a listed file: one assertion failure anywhere fails the job, so a
* real regression can never be papered over by a rerun.
*/
export function planContentionRetry(
failures: readonly TestFailure[],
blockers: readonly RunBlocker[] = [],
): RetryPlan {
if (blockers.length > 0) {
return {
retry: false,
reason: `the run failed for a reason a rerun cannot re-check (${blockers
.map((blocker) => blocker.kind)
.join(', ')})`,
blocked: failures,
blockers,
};
}
if (failures.length === 0) {
return { retry: false, reason: 'no test failures to retry', blocked: [], blockers };
}
const blocked = failures.filter(
(failure) => !isRetryEligibleFile(failure.file) || !failure.timeout,
);
if (blocked.length > 0) {
const assertionShaped = blocked.filter((failure) => isRetryEligibleFile(failure.file));
return {
retry: false,
reason:
assertionShaped.length > 0
? 'a listed file failed for a non-timeout reason (assertion failures never retry)'
: 'a failure landed outside the enumerated retry list',
blocked,
blockers,
};
}
const files = [...new Set(failures.map((failure) => normalizeTestFile(failure.file)))].sort();
return { retry: true, files, failures };
}
export type RetryOutcome = 'passed' | 'failed';
export type RetrySummaryInput = {
plan: RetryPlan;
outcome?: RetryOutcome;
};
/** Markdown for the job summary: every retried file is named, always. */
export function formatRetrySummary({ plan, outcome }: RetrySummaryInput): string {
const lines = ['## Contention retry (#1419)', ''];
if (!plan.retry) {
lines.push(`No retry: ${plan.reason}.`, '');
for (const blocker of plan.blockers) {
lines.push(`- **${blocker.kind}** — ${blocker.detail}`);
}
for (const failure of plan.blocked) {
lines.push(`- \`${normalizeTestFile(failure.file)}\` — ${failure.testName}`);
}
return `${lines.join('\n').trimEnd()}\n`;
}
lines.push(
`Retried ${plan.files.length} timeout-shaped file(s) once — outcome: **${outcome ?? 'pending'}**.`,
'',
'| File | Failing test(s) | Tracking issue | Review by |',
'| --- | --- | --- | --- |',
);
for (const file of plan.files) {
const entry = CONTENTION_RETRY_FILES.find((candidate) => candidate.file === file);
const tests = plan.failures
.filter((failure) => normalizeTestFile(failure.file) === file)
.map((failure) => failure.testName)
.join(', ');
lines.push(
`| \`${file}\` | ${tests} | ${entry?.trackingIssue ?? '—'} | ${entry?.reviewBy ?? '—'} |`,
);
}
return `${lines.join('\n')}\n`;
}