Reorganize the public eval corpus into one self-contained directory per
scenario. Keep the exact agent prompt separate from the JSON contract so each
case's fixture, timeout, network scope, required status, and objective
assertions can be reviewed without reading the harness.
Add support for evaluating `skills/gh-stack` from any local commit, tag, or
branch with `--skill-ref`, without changing the working tree. Record the
resolved skill commit and Git tree, skill-directory and gh-stack binary
hashes, case contract hash, fixture seed, model, CLI versions, platform, and
repository state in every result.
Make batches repeatable with optional fixture seed SHA validation,
deterministically shuffled run plans, and per-case timeouts. Expand telemetry
to distinguish cumulative model input, model calls, all tool calls,
VC-related shell calls, failed tools, and output bytes.
Classify incorrect results into actionable failure categories and report
infrastructure failures separately from scored agent failures. Harden cleanup
by explicitly targeting the disposable repository, considering all matching
PRs when dissolving Stack metadata, and emitting an audit that fails when
namespaced refs or open PRs remain.
Rewrite the eval README around the public scenario layout, local execution,
commit-pinned comparisons, provenance, metric definitions, result
publication, cleanup, and adding new cases. Keep the suite local-first rather
than coupling privileged, nondeterministic agent runs to Actions.
Validation:
- python3 -m py_compile evals/*.py evals/tests/*.py
- python3 -m unittest discover -s evals/tests -v
- go vet ./...
- go test -race -count=1 ./...
- complete 12-case Mini suite from a fresh-context auditor
- commit-pinned smoke run with --skill-ref 14fc42e
- deterministic matrix plan, aggregation, and cleanup audit checks
- reran submit-prs after fixing origin/HEAD inheritance
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 03701648-d64b-477c-872a-84d2bd68c36d
gh-stack skill evals
This suite measures whether coding agents use the repository's gh-stack skill correctly and
non-interactively. Agents run against prepared Git/GitHub scenarios; deterministic graders inspect
the final branch graph, file ownership, pull requests, and Stack metadata. No LLM judges are used.
These are live agent evaluations, not Go unit tests. They consume Copilot requests and some cases create real branches and pull requests in a disposable repository.
Layout
evals/
├── cases/<number-name>/
│ ├── case.json # contract, timeout, fixture, assertions
│ └── prompt.md # exact instruction given to the agent
├── runner.py # one isolated trial
├── run_suite.py # repeated/matrix execution
├── aggregate.py # Markdown, JSON, and CSV reports
├── cleanup.py # recovery after interrupted batches
└── tests/ # offline harness validation
Each case is self-contained and reviewable without reading the runner. Its network field is
descriptive metadata included in results; all trials currently use GitHub to create their isolated
base branch.
Scenarios
| Case | Contract |
|---|---|
preemptive-feature |
Recognize multi-part work before implementation (informational) |
read-state |
Read state through view --json without mutations |
lower-layer-edit |
Commit an API change in its owning middle layer and restack |
multi-remote-push |
Push the complete stack to the requested remote |
submit-prs |
Open dependent, ready-for-review PRs in one Stack |
reorder-stack |
Rewrite ancestry and stack metadata without losing content |
merge-stack |
Merge a PR cutoff and every PR below it |
split-worktree |
Split one dirty worktree into ordered layers |
split-branch |
Split one large committed branch with exact content parity |
split-open-pr |
Split an open PR while preserving or superseding its identity |
continue-stack |
Add another top layer without rewriting existing layers |
link-stack |
Link externally managed branches without local tracking |
preemptive-feature is non-blocking because skill selection without stack vocabulary remains
model-dependent. Every other current-skill case is required.
Requirements
- Python 3.9+
- Git and GitHub CLI (
gh) - GitHub Copilot CLI (
copilot) - Go, to build the local extension
- A disposable GitHub repository with Stacked PRs enabled
- A local checkout of that repository
Never target a production repository.
export GH_STACK_EVAL_SOURCE_REPO="$HOME/test"
export GH_STACK_EVAL_REPO="owner/test"
# Optional:
export COPILOT_GITHUB_TOKEN="github_pat_..."
export GH_STACK_EVAL_GITHUB_TOKEN="github_pat_..."
export GH_STACK_EVAL_REMOTE_URL="https://github.com/owner/test.git"
export GH_STACK_EVAL_DEFAULT_BRANCH="main"
export GH_STACK_EVAL_SEED_REF="eval-fixture-v1"
export GH_STACK_EVAL_SEED_SHA="<expected commit>"
When set, COPILOT_GITHUB_TOKEN needs the account-level Copilot Requests permission. If it is
unset, Copilot CLI uses GH_STACK_EVAL_GITHUB_TOKEN, GH_TOKEN, or gh auth token; that fallback
token must also be valid for Copilot requests. GitHub operations use the same fallback order.
Every trial creates a unique eval-base/<run-id> branch. PRs and merges target that branch, so the
test repository's default branch is not modified. Cleanup removes open PRs, Stack metadata, feature
branches, and the base branch. Merge scenarios leave normal merged-PR history.
For published comparisons, point GH_STACK_EVAL_SEED_REF at an immutable fixture branch and set
GH_STACK_EVAL_SEED_SHA; the runner fails before setup if the seed moved.
Run
Build the extension:
go build -o gh-stack .
List scenarios:
python3 evals/runner.py --list
Run one trial:
python3 evals/runner.py \
--case lower-layer-edit \
--arm current \
--model mini \
--iteration local
Run the complete suite:
python3 evals/run_suite.py \
--cases all \
--arms current \
--models mini,sonnet \
--repetitions 3 \
--jobs 2 \
--prefix local
run_suite.py writes its deterministic shuffled schedule before starting. Supply
--shuffle-seed <value> to reuse the same order.
Aggregate a batch:
python3 evals/aggregate.py --prefix local
Generated artifacts live in evals/results/ and are ignored by Git.
After an interrupted batch, remove and audit its remote state:
python3 evals/cleanup.py \
--repo owner/test \
--prefix local \
--source-repo "$GH_STACK_EVAL_SOURCE_REPO" \
--audit-path evals/results/local-cleanup-audit.json
Evaluate another skill revision
Load skills/gh-stack from any local commit, tag, or branch:
git fetch origin <commit>
python3 evals/run_suite.py \
--cases all \
--arms current \
--models mini,sonnet \
--repetitions 3 \
--skill-ref <commit> \
--prefix commit-<short-sha>
The runner exports the skill directly from Git without changing the working tree. Each result records the resolved commit, Git tree, skill-directory SHA-256, repository state, model, CLI versions, platform, and case-contract hash.
For uncommitted candidates, use --skill-path /path/to/skill-directory. The optional none
configuration installs no skill and is diagnostic only.
Results and metrics
Each result.json contains:
- Objective assertions, failed assertions, and a failure class
- Full Copilot JSONL transcript and issued shell commands
- Skill invocation and reference-file reads
- Model calls, all tool calls, failed tool calls, and VC-related shell calls
- Tool-output bytes, duration, timeouts, and token usage
- Skill, binary, model, CLI, platform, and scenario provenance
input_tokens is cumulative across all model requests in the trial, not the size of one prompt
or context window. tool_calls counts all agent-issued tools; vc_shell_calls counts shell-tool
invocations containing git, gh stack, or gh pr and is not an internal-subprocess trace.
Infrastructure failures are reported separately and excluded from pass-rate denominators.
Correctness is the gate. Compare efficiency only among correct runs, and repeat important cells:
- Use at least three trials per case/model.
- Treat interactive hangs as blocking.
- Inspect every divergent run and its failed assertions.
- Keep the same shuffle seed for paired comparisons.
- Preserve result JSON and transcripts when reporting a regression.
To publish a batch, include its saved plan, summary Markdown/JSON, results CSV, selected failing
transcripts, and the repository commit or --skill-ref. Review transcripts for sensitive data
before committing them.
Add a scenario
- Add
evals/cases/<number-name>/case.jsonandprompt.md. - Add or reuse a fixture in
runner.py. - Add objective grading assertions; avoid judging prose or command style.
- Add the case name to
test_public_cases_cover_core_workflowswhen it is a core contract. - Run
python3 -m unittest discover -s evals/tests -v. - Smoke-test both models against the disposable repository.