mirror of
https://github.com/callstack/agent-device.git
synced 2026-09-14 20:06:34 +08:00
9c22467832
* test(ci): prove every registered gate is owned and reachable (#1429) A check that silently stops running looks exactly like a green build. Two suites had already stopped: `check:tmpdir-leaks` (with its model tests) and `test:fixture-cache` are real package scripts that no workflow ran, reachable only through the `check:unit` aggregate CI never invokes. `CHECK_CATALOG` becomes the registry of every check and `pnpm gate <id>` the only way CI runs one, so finding what a lane runs is a scan for `pnpm gate` rather than an attempt to interpret shell. `pnpm check:gate-manifest` then asserts against the real workflows that every registered check is run by some qualifying lane (per unit, not per script name), that every check the real selector activates for a path is run by a lane that path would start (#1420's class), and that every Vitest project and suite script belongs to a check. The wiring that keeps those honest is asserted too: a gate id must name a registered check, an `if:` must be ruled on in GATE_CONDITIONS so `if: false` unowns what it guards, an action declared to run a gate is proven to, and a job whose steps the loader cannot open fails closed. It deliberately does not try to prove CI runs project code only through `pnpm gate`. Whether a shell block executes project code is not decidable from its text, so shell this model does not recognise earns no ownership credit — the failure direction is a check reported unowned, never one waved through. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ * test(ci): update the two suites that assert on rewired workflow text `scripts/mutation/workflow.test.ts` and `test/ci/trusted-fixture-artifact.test.mjs` read the workflow and action files and assert on their command text, so routing those steps through `pnpm gate <id>` moved what they were matching. They are the two suites the manifest cannot help with: it proves a gate is still run, not that a test asserting on how CI spells a command was updated with it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ * fix(ci): credit gates by execution shape, and keep every guard Three ways the manifest could report a gate as owned when it does not run. 1. Crediting was a substring scan over `run:`, which #1429 explicitly rules out — "do not infer reachability from a command name merely appearing in workflow text". `false && pnpm gate x`, a gate inside `if false; then … fi`, one named in a heredoc, and `echo pnpm gate x` all credited it. There is a live instance: conformance-regenerate.yml's "Fail if regeneration changed anything" step names `pnpm gate maestro-regenerate` inside an error message telling a human to run it, and that credited the gate. A gate now counts only as the first command segment of a line, and a body carrying shell structure earns nothing. Reachability inside a script is not decidable, so this does not try: unrecognised shape means no credit and the check reports unowned. `VAR=$(pnpm gate x …)` is read, since the assignment form is unambiguous and the gate runs. 2. Job-level `if:` was not modelled at all, though six live jobs carry one, so a job that cannot run still credited every gate inside it. Two conditions on the mutation lanes are now declared. 3. A caller's `if:` REPLACED the guard on a nested composite-action step (`guard[0] ?? step.condition`), so an outer `always()` erased an inner `if: false`. Steps carry every guard between the lane and the step. Also corrects two source comments that still claimed project code run outside the runner fails the manifest. It does not: such a step earns no credit. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ * ci: add the run-gate action that names a gate structurally The seam the ownership proof will read instead of shell. A lane says which gate it runs in `with.gate`, a typed input the manifest reads straight out of the YAML and validates against CHECK_CATALOG. Nothing here is wired yet — the ~60 call sites and the model change follow. Added first so the target of that conversion is reviewable on its own. `args` cannot select which gate runs; it is appended after the id, so the worst a wrong value does is fail the gate it already named. There is no `|| true` and no output capture: the gate's exit code is the step's exit code, so a gate cannot run without being able to fail its lane. Part of #1429. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ * merge: main (#1770) and route its three new steps through the runner #1770 landed the orphan-check fix on main, wiring `check:tmpdir-leaks`, `check:tmpdir-leaks:test` and `test:fixture-cache` into Coverage, Layering Guard and Integration Tests. This branch had wired the same three through `pnpm gate`, so the merge produced two steps per check rather than a conflict — each check ran twice. Kept main's steps, with the placement and reasoning reviewed on #1770, and changed only their `run:` line to the canonical runner. Dropped this branch's duplicates. Net effect on CI is unchanged: the same three checks, in the same three lanes, once each. Gate manifest green after the merge: 47 checks wired across 33 lanes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ * fix(ci): address review — suite detection, freerange, glob, vacuous skip-list Six review findings plus the mutation blocker. [bug] `registered` was shape-only, so a `test:*` script running `node src/bin.ts test <dir>` resolved to a `script:` leaf and was invisible. Four `test:replay:*` scripts were owned only because someone hand-registered them; `test:replay:android` was neither registered nor reported while the nightly ran the same six .ad files by inlining them. A `test:*` script is now a suite by name. `replay-android` is registered, and the nightly runs the script instead of re-listing its files so the two cannot drift. The nightly invokes it inside `reactivecircus/android-emulator-runner`'s `script:` input — shell handed to a third-party action this loader does not read — so the suite executes but cannot be credited. Recorded in UNPROVABLE_OWNERS with that exact reason rather than assumed. The fixed detector also found a second orphan the review did not name: `test:integration:progress`. That one is a reporter whose `--check` sibling is the registered gate, so it is declared in REPORTING_SCRIPTS — a declaration that itself fails when inert. [bug] `freerange` defaulted to localRunnable, so fail-open ran `fr` (a Bun binary) on the pre-push path. Now false. [suggestion] The `--run` skip-list asserted `build:android-snapshot-helper`, a name `android-helpers` no longer uses, so it could not fail. Derived from the catalog instead. [suggestion] `matchesGlob` joined `**` splits with `.*`, making the adjacent slash mandatory — GitHub's `**` matches zero directories, so `src/**/*.test.ts` did not match `src/a.test.ts`. Pinned against `packages/*/src/**/*.test.ts`. [suggestion] Deleted the unwired `run-gate` action. It had no callers, was absent from GATE_ACTIONS, and its comment described a system that had not shipped. It returns with the rewiring, not before. [suggestion] Collapsed the module headers that narrated discarded designs. Mutation: `daemon entrypoint publishes HTTP metadata and cleans up on shutdown` is the only test here that spawns a real daemon process. It takes ~1.1s alone but exceeds Vitest's 5s default inside Stryker's dry run, which aborts the sweep before a single mutant runs. Given 30s. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ * fix(mutation): order sandbox aliases longest-first so subpaths resolve Every shard of the mutation sweep aborted in Stryker's dry run with: Cannot find package '@agent-device/selectors/engine' imported from .tmp/stryker/sandbox-*/src/core/selector-pipeline.ts The alias was generated correctly; it just never won. Vite matches a STRING alias by prefix and takes the first hit, and `workspaceSpecifierTargets` emitted the bare `@agent-device/selectors` ahead of the subpath entries. The bare entry therefore captured `@agent-device/selectors/engine` and rewrote it to `…/src/index.ts/engine`, which does not exist; Node fell back to real package resolution, could not find the subpath inside the sandbox, and the dry run failed before a single mutant ran — so the shard uploaded an empty envelope instead of a report and the ratchet failed for want of one. Sorting longest specifier first makes the most specific alias win: @agent-device/selectors/engine -> packages/selectors/src/engine.ts @agent-device/selectors/ast -> packages/selectors/src/ast.ts @agent-device/selectors -> packages/selectors/src/index.ts `/ast` never tripped this because nothing in a related test set imported it; `selector-pipeline.ts` introduced the first subpath import that mattered (#1744), so the mutation lane has been unable to run since that landed. Any PR touching `scripts/mutation/**` — which fails open into the full sweep — would have hit it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ * refactor: derive gate ownership from workflow structure * fix: run gates without optional arguments * fix: resolve mutation workspace subpaths exactly --------- Co-authored-by: Claude <noreply@anthropic.com>
169 lines
7.1 KiB
YAML
169 lines
7.1 KiB
YAML
name: Conformance Differential
|
|
|
|
# Layer 3 of the Maestro conformance oracle: run a small set of flows through
|
|
# BOTH real Maestro and agent-device on a live device and compare app-observable
|
|
# outcomes, plus assert engine-side timing invariants over agent-device's own
|
|
# replay trace. This is the only layer that can catch settle-loop ordering, which
|
|
# has no reflectable upstream constant.
|
|
#
|
|
# Scheduled: this path is proven end-to-end (run 29504440599 executed both
|
|
# engines against the real fixture app). Scenarios that currently diverge are
|
|
# DECLARED via `knownDivergence` in differential/scenarios.ts with a tracking
|
|
# issue, so the schedule stays green on known gaps and fails only on undeclared
|
|
# ones — the same contract layer 1 has. A declaration that stops reproducing also
|
|
# fails, so a fix cannot land while leaving the oracle blind.
|
|
|
|
on:
|
|
schedule:
|
|
- cron: '0 5 * * *'
|
|
workflow_dispatch:
|
|
inputs:
|
|
only:
|
|
description: 'Single scenario id to run (default: all)'
|
|
required: false
|
|
|
|
permissions:
|
|
contents: read
|
|
# setup-fixture-app looks up the cached test-app binary via the artifacts API.
|
|
actions: read
|
|
|
|
concurrency:
|
|
group: ci-${{ github.workflow }}-${{ github.ref }}
|
|
cancel-in-progress: true
|
|
|
|
env:
|
|
AGENT_DEVICE_CLI: '--experimental-strip-types src/bin.ts'
|
|
DIFFERENTIAL_ONLY: ${{ github.event.inputs.only || '' }}
|
|
# CI should not phone home, and it keeps `maestro --version` to just the
|
|
# version instead of prefixing an analytics notice.
|
|
MAESTRO_CLI_NO_ANALYTICS: '1'
|
|
|
|
jobs:
|
|
differential-ios:
|
|
name: iOS Conformance Differential
|
|
runs-on: macos-26
|
|
timeout-minutes: 90
|
|
env:
|
|
IOS_RUNTIME_VERSION: '26.2'
|
|
AGENT_DEVICE_STATE_DIR: ${{ github.workspace }}/.tmp/agent-device-state
|
|
AGENT_DEVICE_IOS_RUNNER_DERIVED_PATH: ${{ github.workspace }}/.tmp/ios-runner-derived
|
|
steps:
|
|
- name: Checkout
|
|
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
|
|
|
|
- name: Setup toolchain
|
|
uses: ./.github/actions/setup-node-pnpm
|
|
|
|
- name: Restore and build iOS XCTest runner
|
|
uses: ./.github/actions/setup-apple-runner-build
|
|
with:
|
|
derived-path: ${{ env.AGENT_DEVICE_IOS_RUNNER_DERIVED_PATH }}
|
|
cache-key-prefix: ios-runner-prebuilt
|
|
cache-key-suffix: -ios-${{ env.IOS_RUNTIME_VERSION }}
|
|
gate: swift-runner-ios
|
|
xcuitest-platform: ios
|
|
xcuitest-destination: generic/platform=iOS Simulator
|
|
|
|
- name: Boot simulator
|
|
uses: ./.github/actions/boot-ios-test-simulator
|
|
with:
|
|
runtime-version: ${{ env.IOS_RUNTIME_VERSION }}
|
|
preferred-device-name: iPhone 17 Pro
|
|
|
|
# Installs the cached Release test-app and refreshes its JS. Shared with any
|
|
# other job that needs a controlled app to drive (see #320) — the artifact
|
|
# is repo-wide, so once the producer builds a fingerprint every workflow
|
|
# gets it as a download.
|
|
- name: Setup fixture app
|
|
id: fixture-app
|
|
uses: ./.github/actions/setup-fixture-app
|
|
with:
|
|
device-name: iPhone 17 Pro
|
|
|
|
# The action guarantees *an* app is installed; this asserts it is the one
|
|
# the scenarios actually target, so a fixture rename cannot leave the
|
|
# differential driving the wrong app.
|
|
- name: Verify the installed app is the one the scenarios target
|
|
run: |
|
|
set -euo pipefail
|
|
EXPECTED=$(node --experimental-strip-types -e \
|
|
"import('./packages/maestro/test/conformance/differential/scenarios.ts').then(m => console.log(m.DIFFERENTIAL_APP_ID))")
|
|
ACTUAL="${{ steps.fixture-app.outputs.app-id }}"
|
|
echo "expected=$EXPECTED actual=$ACTUAL (source=${{ steps.fixture-app.outputs.source }})"
|
|
if [ "$ACTUAL" != "$EXPECTED" ]; then
|
|
echo "::error::Installed fixture app is $ACTUAL but the scenarios target $EXPECTED; they would fail vacuously."
|
|
exit 1
|
|
fi
|
|
|
|
# Pin the Maestro CLI to the SAME version the oracle pins for layers 1-2,
|
|
# read from pinned-upstream.json so the two can never drift apart. The
|
|
# installer honours MAESTRO_VERSION; assert what actually landed.
|
|
- name: Install pinned Maestro CLI
|
|
run: |
|
|
set -euo pipefail
|
|
MAESTRO_VERSION=$(node -p "require('./scripts/maestro-conformance/pinned-upstream.json').version")
|
|
export MAESTRO_VERSION
|
|
echo "Pinning Maestro CLI to $MAESTRO_VERSION (from pinned-upstream.json)"
|
|
curl -fsSL "https://get.maestro.mobile.dev" | bash
|
|
echo "$HOME/.maestro/bin" >> "$GITHUB_PATH"
|
|
|
|
- name: Verify the Maestro CLI matches the oracle pin
|
|
run: |
|
|
set -euo pipefail
|
|
EXPECTED=$(node -p "require('./scripts/maestro-conformance/pinned-upstream.json').version")
|
|
# `maestro --version` can print notices (e.g. the analytics banner)
|
|
# around the version, so match the semver line rather than the whole
|
|
# output. Empty ACTUAL fails the comparison below, as it should.
|
|
ACTUAL=$(maestro --version 2>/dev/null | grep -oE '[0-9]+\.[0-9]+\.[0-9]+' | tail -1 || true)
|
|
echo "expected=$EXPECTED actual=$ACTUAL"
|
|
if [ "$ACTUAL" != "$EXPECTED" ]; then
|
|
echo "::error::Layer-3 Maestro CLI is $ACTUAL but the oracle pins $EXPECTED. A differential against a different Maestro proves nothing about the pinned behavior."
|
|
exit 1
|
|
fi
|
|
|
|
# --trace-root lets the runner find agent-device's replay-timing.ndjson and
|
|
# assert engine-side invariants (e.g. the settle loop latches instead of
|
|
# burning its budget) — outcome parity alone cannot see that.
|
|
- name: Run differential
|
|
if: env.DIFFERENTIAL_ONLY == ''
|
|
uses: ./.github/actions/run-gate
|
|
with:
|
|
gate: maestro-differential
|
|
args: |
|
|
--platform
|
|
ios
|
|
--out-dir
|
|
${{ github.workspace }}/.tmp/conformance-differential
|
|
--trace-root
|
|
${{ github.workspace }}/.agent-device
|
|
|
|
- name: Run filtered differential
|
|
if: env.DIFFERENTIAL_ONLY != ''
|
|
uses: ./.github/actions/run-gate
|
|
with:
|
|
gate: maestro-differential
|
|
args: |
|
|
--platform
|
|
ios
|
|
--out-dir
|
|
${{ github.workspace }}/.tmp/conformance-differential
|
|
--trace-root
|
|
${{ github.workspace }}/.agent-device
|
|
--only
|
|
${{ env.DIFFERENTIAL_ONLY }}
|
|
|
|
# Keep the trace, not just the verdict: the engine-side invariants are
|
|
# computed FROM replay-timing.ndjson, so without it a report saying
|
|
# "invariant violated: tapRetries was 0" cannot be audited or debugged
|
|
# after the runner is gone.
|
|
- name: Upload report and trace
|
|
if: always()
|
|
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
|
|
with:
|
|
name: conformance-differential-ios
|
|
include-hidden-files: true
|
|
path: |
|
|
.tmp/conformance-differential/differential-report.json
|
|
.agent-device/**/replay-timing.ndjson
|
|
if-no-files-found: ignore
|