Files
dotnet__skills/.github/workflows/eval-quality.yml
Abhitej John 3e9d74ade4 Use distinct stimuli for eval inference
Collapse repeated runs to one majority-direction vote per stimulus, preserve run-level reliability evidence, and align authoring checks, reports, and guidance with that inference unit.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>

Copilot-Session: 1c6716a2-172b-461a-adb5-127b1ba7e96b
2026-08-18 18:08:00 -07:00

84 lines
3.0 KiB
YAML

name: eval-quality
# Structural quality gate for evaluation specs and their fixtures.
#
# Each failing check corresponds to a defect that has already cost a real
# evaluation result on this repo — see eng/eval-quality/README.md. All of them
# are structural (file existence, git state, declared counts, YAML keys), so the
# gate cannot fire spuriously on well-written prose. Judgement calls such as
# orphaned fixtures and grandfathered underpowered evals are reported but never
# fail the build.
on:
pull_request:
paths:
- "tests/**"
- "plugins/**"
- "eng/eval-quality/**"
# check_floor_agreement() reads MIN_CREDIBLE_STIMULI out of the adapter, so
# the one diff that can break the floor's two-language agreement is a diff
# to the adapter. Without this the guard would never run on it.
- "eng/vally-adapter/**"
- ".github/workflows/eval-quality.yml"
- ".gitignore"
push:
branches: [main]
# Kept in sync with the pull_request paths above. The gate reads plugins/*
# (skills with no eval), .gitignore (fixtures excluded from the index), and
# eng/vally-adapter (the stimulus floor it must agree with), so omitting any of
# them here would let a direct push or a squash-merge that touches only
# those land on main without the gate ever running.
paths:
- "tests/**"
- "plugins/**"
- "eng/eval-quality/**"
- "eng/vally-adapter/**"
- ".github/workflows/eval-quality.yml"
- ".gitignore"
workflow_dispatch:
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
permissions:
contents: read
jobs:
check:
runs-on: ubuntu-latest
steps:
- name: Checkout repository
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v6
with:
persist-credentials: false
# The underpowered allowlist is shrink-only, which needs the base
# revision to compare against. Shallow history has no merge base.
fetch-depth: 0
- name: Set up Python
uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
with:
python-version: "3.12"
- name: Install PyYAML
run: python -m pip install --quiet pyyaml
# The gate is only trustworthy if it has been shown to fire. This injects
# each defect into a scratch tree and asserts the gate rejects it.
- name: Self-test the gate
run: python eng/eval-quality/selftest_eval_quality.py
- name: Check eval quality
env:
BASE_REF: ${{ github.base_ref }}
run: |
# On a pull request, additionally reject new entries in the
# underpowered allowlist relative to the base branch — otherwise a PR
# could add a below-floor eval and exempt it in the same change.
if [ -n "$BASE_REF" ]; then
python eng/eval-quality/check_eval_quality.py --base-ref "origin/$BASE_REF"
else
python eng/eval-quality/check_eval_quality.py
fi