Files
callstack__agent-device/.github/workflows/test-app-build-cache.yml
Michał Pierzchała 9c22467832 refactor(ci): make gate ownership structural (#1429) (#1753)
* test(ci): prove every registered gate is owned and reachable (#1429)

A check that silently stops running looks exactly like a green build. Two
suites had already stopped: `check:tmpdir-leaks` (with its model tests) and
`test:fixture-cache` are real package scripts that no workflow ran, reachable
only through the `check:unit` aggregate CI never invokes.

`CHECK_CATALOG` becomes the registry of every check and `pnpm gate <id>` the
only way CI runs one, so finding what a lane runs is a scan for `pnpm gate`
rather than an attempt to interpret shell. `pnpm check:gate-manifest` then
asserts against the real workflows that every registered check is run by some
qualifying lane (per unit, not per script name), that every check the real
selector activates for a path is run by a lane that path would start (#1420's
class), and that every Vitest project and suite script belongs to a check.

The wiring that keeps those honest is asserted too: a gate id must name a
registered check, an `if:` must be ruled on in GATE_CONDITIONS so `if: false`
unowns what it guards, an action declared to run a gate is proven to, and a
job whose steps the loader cannot open fails closed.

It deliberately does not try to prove CI runs project code only through
`pnpm gate`. Whether a shell block executes project code is not decidable from
its text, so shell this model does not recognise earns no ownership credit —
the failure direction is a check reported unowned, never one waved through.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* test(ci): update the two suites that assert on rewired workflow text

`scripts/mutation/workflow.test.ts` and `test/ci/trusted-fixture-artifact.test.mjs`
read the workflow and action files and assert on their command text, so routing
those steps through `pnpm gate <id>` moved what they were matching.

They are the two suites the manifest cannot help with: it proves a gate is still
run, not that a test asserting on how CI spells a command was updated with it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* fix(ci): credit gates by execution shape, and keep every guard

Three ways the manifest could report a gate as owned when it does not run.

1. Crediting was a substring scan over `run:`, which #1429 explicitly rules
   out — "do not infer reachability from a command name merely appearing in
   workflow text". `false && pnpm gate x`, a gate inside `if false; then … fi`,
   one named in a heredoc, and `echo pnpm gate x` all credited it. There is a
   live instance: conformance-regenerate.yml's "Fail if regeneration changed
   anything" step names `pnpm gate maestro-regenerate` inside an error message
   telling a human to run it, and that credited the gate.

   A gate now counts only as the first command segment of a line, and a body
   carrying shell structure earns nothing. Reachability inside a script is not
   decidable, so this does not try: unrecognised shape means no credit and the
   check reports unowned. `VAR=$(pnpm gate x …)` is read, since the assignment
   form is unambiguous and the gate runs.

2. Job-level `if:` was not modelled at all, though six live jobs carry one, so
   a job that cannot run still credited every gate inside it. Two conditions on
   the mutation lanes are now declared.

3. A caller's `if:` REPLACED the guard on a nested composite-action step
   (`guard[0] ?? step.condition`), so an outer `always()` erased an inner
   `if: false`. Steps carry every guard between the lane and the step.

Also corrects two source comments that still claimed project code run outside
the runner fails the manifest. It does not: such a step earns no credit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* ci: add the run-gate action that names a gate structurally

The seam the ownership proof will read instead of shell. A lane says which
gate it runs in `with.gate`, a typed input the manifest reads straight out of
the YAML and validates against CHECK_CATALOG.

Nothing here is wired yet — the ~60 call sites and the model change follow.
Added first so the target of that conversion is reviewable on its own.

`args` cannot select which gate runs; it is appended after the id, so the
worst a wrong value does is fail the gate it already named. There is no
`|| true` and no output capture: the gate's exit code is the step's exit code,
so a gate cannot run without being able to fail its lane.

Part of #1429.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* merge: main (#1770) and route its three new steps through the runner

#1770 landed the orphan-check fix on main, wiring `check:tmpdir-leaks`,
`check:tmpdir-leaks:test` and `test:fixture-cache` into Coverage, Layering
Guard and Integration Tests. This branch had wired the same three through
`pnpm gate`, so the merge produced two steps per check rather than a conflict
— each check ran twice.

Kept main's steps, with the placement and reasoning reviewed on #1770, and
changed only their `run:` line to the canonical runner. Dropped this branch's
duplicates. Net effect on CI is unchanged: the same three checks, in the same
three lanes, once each.

Gate manifest green after the merge: 47 checks wired across 33 lanes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* fix(ci): address review — suite detection, freerange, glob, vacuous skip-list

Six review findings plus the mutation blocker.

[bug] `registered` was shape-only, so a `test:*` script running
`node src/bin.ts test <dir>` resolved to a `script:` leaf and was invisible.
Four `test:replay:*` scripts were owned only because someone hand-registered
them; `test:replay:android` was neither registered nor reported while the
nightly ran the same six .ad files by inlining them. A `test:*` script is now
a suite by name. `replay-android` is registered, and the nightly runs the
script instead of re-listing its files so the two cannot drift.

  The nightly invokes it inside `reactivecircus/android-emulator-runner`'s
  `script:` input — shell handed to a third-party action this loader does not
  read — so the suite executes but cannot be credited. Recorded in
  UNPROVABLE_OWNERS with that exact reason rather than assumed.

  The fixed detector also found a second orphan the review did not name:
  `test:integration:progress`. That one is a reporter whose `--check` sibling
  is the registered gate, so it is declared in REPORTING_SCRIPTS — a
  declaration that itself fails when inert.

[bug] `freerange` defaulted to localRunnable, so fail-open ran `fr` (a Bun
binary) on the pre-push path. Now false.

[suggestion] The `--run` skip-list asserted `build:android-snapshot-helper`,
a name `android-helpers` no longer uses, so it could not fail. Derived from
the catalog instead.

[suggestion] `matchesGlob` joined `**` splits with `.*`, making the adjacent
slash mandatory — GitHub's `**` matches zero directories, so
`src/**/*.test.ts` did not match `src/a.test.ts`. Pinned against
`packages/*/src/**/*.test.ts`.

[suggestion] Deleted the unwired `run-gate` action. It had no callers, was
absent from GATE_ACTIONS, and its comment described a system that had not
shipped. It returns with the rewiring, not before.

[suggestion] Collapsed the module headers that narrated discarded designs.

Mutation: `daemon entrypoint publishes HTTP metadata and cleans up on
shutdown` is the only test here that spawns a real daemon process. It takes
~1.1s alone but exceeds Vitest's 5s default inside Stryker's dry run, which
aborts the sweep before a single mutant runs. Given 30s.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* fix(mutation): order sandbox aliases longest-first so subpaths resolve

Every shard of the mutation sweep aborted in Stryker's dry run with:

  Cannot find package '@agent-device/selectors/engine' imported from
    .tmp/stryker/sandbox-*/src/core/selector-pipeline.ts

The alias was generated correctly; it just never won. Vite matches a STRING
alias by prefix and takes the first hit, and `workspaceSpecifierTargets`
emitted the bare `@agent-device/selectors` ahead of the subpath entries. The
bare entry therefore captured `@agent-device/selectors/engine` and rewrote it
to `…/src/index.ts/engine`, which does not exist; Node fell back to real
package resolution, could not find the subpath inside the sandbox, and the dry
run failed before a single mutant ran — so the shard uploaded an empty
envelope instead of a report and the ratchet failed for want of one.

Sorting longest specifier first makes the most specific alias win:

  @agent-device/selectors/engine -> packages/selectors/src/engine.ts
  @agent-device/selectors/ast    -> packages/selectors/src/ast.ts
  @agent-device/selectors        -> packages/selectors/src/index.ts

`/ast` never tripped this because nothing in a related test set imported it;
`selector-pipeline.ts` introduced the first subpath import that mattered
(#1744), so the mutation lane has been unable to run since that landed. Any
PR touching `scripts/mutation/**` — which fails open into the full sweep —
would have hit it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SkS4S8XXrfkJ8TD1VBKkvJ

* refactor: derive gate ownership from workflow structure

* fix: run gates without optional arguments

* fix: resolve mutation workspace subpaths exactly

---------

Co-authored-by: Claude <noreply@anthropic.com>
2026-08-12 16:02:18 +02:00

235 lines
9.9 KiB
YAML

name: Test App Build Cache
# Builds a Release binary of examples/test-app and publishes it as a workflow
# artifact keyed by the Expo fingerprint, so CI e2e jobs install a ready binary
# (via .github/actions/setup-fixture-app) instead of compiling from scratch.
#
# Release, not dev-client: the JS bundle is embedded, so a consuming job needs no
# Metro. Its JS is refreshed on retrieval with @expo/repack-app, so keying on the
# native-only fingerprint is correct -- a JS-only change reuses the same native
# binary. Local development does not use this; it caches on disk via the
# expo-build-disk-cache provider in app.config.js.
#
# Only builds when the fingerprint has no trusted artifact yet. Runs on every
# push to main and every PR so a same-head consumer can wait for a scheduled
# producer (workflow_dispatch cannot reach a workflow that is not yet on main).
on:
push:
branches:
- main
pull_request:
workflow_dispatch:
permissions:
contents: read
# Look up whether this fingerprint's artifact already exists across runs.
actions: read
jobs:
fingerprint:
name: Resolve native fingerprint
runs-on: ubuntu-latest
timeout-minutes: 10
outputs:
matrix: ${{ steps.fingerprint.outputs.matrix }}
steps:
- name: Checkout
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
- name: Setup toolchain
uses: ./.github/actions/setup-node-pnpm
with:
cache-dependency-path: examples/test-app/pnpm-lock.yaml
- name: Setup test app dependencies
uses: ./.github/actions/setup-test-app-dependencies
- name: Typecheck test app
uses: ./.github/actions/run-gate
with: { gate: test-app-typecheck }
- name: Resolve native fingerprint
id: fingerprint
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
run: |
set -euo pipefail
ARTIFACT_NAME_RESOLVER=".github/actions/setup-fixture-app/resolve-artifact-name.sh"
IOS_NAME="$(sh "$ARTIFACT_NAME_RESOLVER" ios)"
ANDROID_NAME="$(sh "$ARTIFACT_NAME_RESOLVER" android)"
EXPECTED_HEAD_SHA="${{ github.event.pull_request.head.sha || github.sha }}"
TRUSTED_ARTIFACT=".github/actions/setup-fixture-app/trusted-artifact.mjs"
IOS_ARTIFACT="$(node "$TRUSTED_ARTIFACT" find "${{ github.repository }}" "$IOS_NAME" "$EXPECTED_HEAD_SHA")"
ANDROID_ARTIFACT="$(node "$TRUSTED_ARTIFACT" find "${{ github.repository }}" "$ANDROID_NAME" "$EXPECTED_HEAD_SHA")"
IOS_BUILD=true
if [ -n "$IOS_ARTIFACT" ] ||
{ [ "${{ github.event_name }}" = "pull_request" ] &&
[ "${{ github.event.pull_request.head.repo.full_name }}" != "${{ github.repository }}" ]; }; then
IOS_BUILD=false
fi
ANDROID_BUILD=true
if [ -n "$ANDROID_ARTIFACT" ]; then
ANDROID_BUILD=false
fi
MATRIX="$(jq -cn \
--arg iosName "$IOS_NAME" \
--argjson iosBuild "$IOS_BUILD" \
--arg androidName "$ANDROID_NAME" \
--argjson androidBuild "$ANDROID_BUILD" \
'{
include: [
{
name: "iOS Release",
platform: "ios",
runsOn: (if $iosBuild then "macos-26" else "ubuntu-latest" end),
artifactName: $iosName,
build: $iosBuild
},
{
name: "Android Release",
platform: "android",
runsOn: "ubuntu-latest",
artifactName: $androidName,
build: $androidBuild
}
]
}')"
echo "matrix=$MATRIX" >> "$GITHUB_OUTPUT"
release:
name: ${{ matrix.name }}
needs: fingerprint
runs-on: ${{ matrix.runsOn }}
timeout-minutes: 60
concurrency:
group: test-app-${{ matrix.artifactName }}-${{ github.event.pull_request.number || github.ref_name }}
cancel-in-progress: false
strategy:
# Independent caches; one platform failing must not withhold the other's.
fail-fast: false
matrix: ${{ fromJSON(needs.fingerprint.outputs.matrix) }}
steps:
- name: Checkout
if: matrix.build
uses: actions/checkout@de0fac2e4500dabe0009e67214ff5f5447ce83dd # v6.0.2
- name: Setup toolchain
if: matrix.build
uses: ./.github/actions/setup-node-pnpm
with:
cache-dependency-path: |
pnpm-lock.yaml
examples/test-app/pnpm-lock.yaml
- name: Setup test app dependencies
if: matrix.build
uses: ./.github/actions/setup-test-app-dependencies
# The generated Android project is safe to assemble without an emulator.
# Keep Gradle's heavy dependency/transforms cache, but never let an
# untrusted fork PR write it.
- name: Setup Gradle cache
if: matrix.platform == 'android' && matrix.build
uses: gradle/actions/setup-gradle@ed408507eac070d1f99cc633dbcf757c94c7933a # v4.4.3
with:
cache-read-only: ${{ github.event_name == 'pull_request' && github.event.pull_request.head.repo.full_name != github.repository }}
# `--device generic` builds for the simulator without booting one or
# installing. Release embeds the bundle (no Metro) and links a universal
# x86_64+arm64 slice, so the artifact runs on any developer's simulator.
# Fork PRs cannot safely publish this binary; their iOS workflow builds the
# same Release app inline on a cache miss, so compiling it here would double
# the macOS cost without adding signal.
- name: Build the iOS Release app
if: |
matrix.platform == 'ios' &&
matrix.build &&
(github.event_name != 'pull_request' ||
github.event.pull_request.head.repo.full_name == github.repository)
run: |
set -euo pipefail
pnpm --dir examples/test-app exec expo run:ios \
--configuration Release \
--device generic \
--no-bundler \
--output "${{ github.workspace }}/.tmp/test-app-build"
# Deliberately no `set -o pipefail`: `yes` is killed by SIGPIPE once
# sdkmanager stops reading, and pipefail would report that 141 as the
# step's own failure.
- name: Install Android SDK packages
if: matrix.platform == 'android' && matrix.build
run: |
SDK_ROOT="${ANDROID_HOME:-${ANDROID_SDK_ROOT:-/usr/local/lib/android/sdk}}"
SDKMANAGER="$SDK_ROOT/cmdline-tools/latest/bin/sdkmanager"
if [ ! -x "$SDKMANAGER" ]; then
SDKMANAGER="$SDK_ROOT/cmdline-tools/bin/sdkmanager"
fi
if [ ! -x "$SDKMANAGER" ]; then
echo "sdkmanager not found under $SDK_ROOT" >&2
exit 1
fi
yes | "$SDKMANAGER" --licenses >/dev/null
"$SDKMANAGER" "platforms;android-36" "build-tools;36.0.0"
# CNG's generated Android project exposes the ordinary Gradle Release task;
# it builds and packages every ABI without device resolution. Installation
# and launch are deliberately proven by the Android consumer instead.
- name: Build the Android Release apk
if: matrix.platform == 'android' && matrix.build
run: |
set -euo pipefail
pnpm --dir examples/test-app exec expo prebuild --platform android --no-install
(
cd examples/test-app/android
./gradlew :app:assembleRelease
)
mkdir -p "${{ github.workspace }}/.tmp/test-app-build"
cp examples/test-app/android/app/build/outputs/apk/release/app-release.apk \
"${{ github.workspace }}/.tmp/test-app-build/app-release.apk"
- name: Stage the binary for upload
id: stage
if: |
matrix.build &&
(matrix.platform == 'android' ||
github.event_name != 'pull_request' ||
github.event.pull_request.head.repo.full_name == github.repository)
run: |
set -euo pipefail
SRC="${{ github.workspace }}/.tmp/test-app-build"
BINARY="$(find "$SRC" -maxdepth 1 \( -name '*.app' -o -name '*.apk' \) | head -1)"
if [ -z "$BINARY" ]; then
echo "::error::Build produced no .app or .apk under $SRC."
exit 1
fi
# An APK must span both the maintainer's arm64 and CI's x86_64; a
# regression here (e.g. an unexpected abiFilter) is otherwise silent.
if [ "${{ matrix.platform }}" = "android" ] && ! unzip -l "$BINARY" | grep -q "lib/arm64-v8a/"; then
echo "::error::Release APK has no arm64-v8a libraries."
exit 1
fi
# Tar the binary: artifact zips drop the executable bit, which would
# leave an installed .app unable to launch.
mkdir -p .tmp/test-app-artifact
tar -czf .tmp/test-app-artifact/binary.tar.gz -C "$SRC" "$(basename "$BINARY")"
echo "publishing ${{ matrix.artifactName }} ($(du -h .tmp/test-app-artifact/binary.tar.gz | cut -f1))"
# Fork PRs run with a read-only token but can still upload artifacts, and a
# consumer serves whatever it finds by name -- so a fork could publish a
# tampered binary that CI then executes. Only branches in this repo publish.
# Android still builds on forks; iOS gets equivalent build/install signal
# from the smoke workflow's inline cache-miss fallback.
- name: Publish to the build cache
if: |
matrix.build &&
(github.event_name != 'pull_request' ||
github.event.pull_request.head.repo.full_name == github.repository)
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
with:
name: ${{ matrix.artifactName }}
path: .tmp/test-app-artifact/binary.tar.gz
if-no-files-found: error
compression-level: 0 # already gzipped