36 "# noqa: E402 / BLE001" markers in scripts/ and evals/ previously
annotated a lint gate that didn't exist (no config, no CI step) — a
false signal to anyone reading the code. Add ruff.toml selecting E, F,
BLE (the codes the noqa markers assume) and wire one CI step after
py_compile.
Style-only codes (E401/E501/E701/E702/E731) are disabled outright:
this codebase's small check_*.py helpers deliberately use compact
one-line style (semicolon joins, single-line if/lambda bodies), and
those codes fire on that choice rather than a defect.
Everything else that fired is a real, pre-existing finding, scoped to
its file via per-file-ignores rather than fixed here (out of scope for
this plan — baseline must pass today via rule selection, not source
edits). Worth a maintainer's attention later:
evals/check_contrib.py F821 `subprocess` referenced only in
a `from __future__ import
annotations` return type, never
imported; F401 unused `Path`.
evals/check_pattern_coverage.py F401 unused `Path`.
evals/check_seeded_docs.py BLE001 two bare `except Exception`.
evals/run_model_parity.py F401 unused `os`.
scripts/banned_phrase_scan.py F401 two unused `_lang` imports;
F601 `"needless to say"` key literal
repeated in BANNED_PHRASES (same
value both times, harmless but dead).
scripts/calibrate_score.py F401 unused `POLES`.
scripts/harvest_classify.py F401 unused `DATE_FLOOR`.
scripts/structure_scan.py F401 two unused `_lang` imports.
scripts/suggest.py F401 unused `re`.
scripts/voice_profile.py F401 unused `math`.
scripts/wiki_sync.py F541 three f-strings with no
placeholders.
Baseline is deliberately minimal; tightening is maintainer taste later.
adversarial-evals.json is hand-merged and hand-numbered with no
validation: a duplicate id runs twice silently and collapses in the
EXPECTED_XFAIL comparison, and a malformed row crashes at runtime
instead of failing a gate cleanly. evals/check_evals_schema.py checks
row uniqueness, id shape, required keys, and per-target shape
(script command/assertions against check_assertion's real type set,
skill rows against build_shared_benchmark.py's loader), registered as
the new "evals-schema" gate (DOC-24).
CI previously named only 4 gate steps by hand while ~20 other blocking
gates ran only transitively through DOC-* rows, so deleting one DOC
row could silently drop a gate from CI. The workflow now parses
--list-gates and runs every blocking gate directly, skipping the
handful whose commands template a live rewrite session's own output
check (transformed.txt/original.txt aren't repo artifacts; their
scanner logic is already exercised by the suite's FP-*/STRUCT-*/SIL-*/
PRES-*/ROB-*/AS-* rows).
CONTRIBUTING.md, an issue form, and a PR template link into the existing
contribute pipeline (references/contribute.md, references/commands/contribute.md,
CLAUDE.md) rather than restating it, plus one README pointer. Keeps the docs
summary-free so they can't drift from the gate-checked flow they link to.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K6CYksdLbXbTAxcAQjvHz5
Insert `from __future__ import annotations` in the 17 files that use
PEP 604/585 annotations in module-level positions evaluated at import
time, so scripts/banned_phrase_scan.py and friends no longer raise
TypeError on Python 3.8/3.9. Add a 3.8 leg to the CI matrix so the
floor claim in README.md is actually gated, and correct the two
imprecise "439 deterministic cases" references to "440 deterministic
script cases (439 pass, 1 documented xfail)".
New scripts must carry the future-import until the floor is raised;
if the maintainer later chooses 3.10+, delete the CI 3.8 leg and
README claim together.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K6CYksdLbXbTAxcAQjvHz5
- evals/run_behavioral.sh: one-command behavioral run (prepare -> run_local ->
judge -> null-score coercion -> benchmark); holdback sealed behind
UNSLOP_CONFIRM_HOLDBACK=1
- evals/CHECKS.md: gate matrix mirroring --list-gates plus parallel check
protocol for small-context sub-agents; DOC-03 + check_gates_doc.py pin the
doc to the live matrix
- DOC-02 + check_skill_examples.py: SKILL.md example outputs must pass the
skill's own blocking gates (the old crisp example tripped both)
- CI runs all deterministic gates including strict-leakage validate and
taboo parity
- CLAUDE.md/AGENTS.md: consistent decision table for smallest useful eval
coverage, live-gate guidance, explicit behavioral-run rule; maintenance
commands are now an eval-first procedure; CRITIQUE.md counts point at the
runner instead of going stale
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Generate evals/shared-benchmark.json from the 26 behavioral skill cases via
evals/build_shared_benchmark.py, adding what the deterministic suite can't
measure: with_skill/without_skill variants for lift, tune/holdout/holdback
splits against overfitting, LLM judge assertions, deterministic script
backstops that reuse banned_phrase_scan.py / validate_preservation.py over
each run's output.md, and three ablations. CI checks the manifest stays in
sync. Running the harness itself (install + judge) is documented in
evals/BEHAVIORAL-EVALS.md and left as a local step (needs model credentials).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Robustness:
- every script now exits cleanly (code 2) on a missing/unreadable input file
instead of raising a traceback (was only fixed for validate_preservation)
- runner adds a per-case timeout and survives a missing command/crash instead
of aborting the whole run
- gated two more context-sensitive single words behind their jargon
collocations: "harness" (a horse harness vs "harness the power of") and
"foster" (a foster family vs "foster a culture of")
CI:
- .github/workflows/evals.yml compiles the scripts and runs the regression
harness on every push and PR; only an undocumented FAIL breaks the build
New eval cases (script target, all passing):
- FP-07/08 literal harness/foster; FP-09 smart-quote exemption
- REC-01/02/03 recall guards proving the gating still flags real jargon and
stacked slop (so de-noising can't silently gut detection)
- PRES-05 comma-format equality, PRES-06 $2.4M == $2.4 million, PRES-07 dropped
quarter is a lost fact
- ROB-04..07 missing-file robustness for each script, ROB-08 empty stdin
Suite: 53 cases (35 script: 34 PASS / 1 XFAIL / 0 FAIL, 18 behavioral skill).