Files
Nuno Sabino 4822dc3876 variant-analysis: convert the skill to a dynamic workflow (#232)
* Converted skill into a dynamic workflow. Still working on the tests

* Added gradio test with injected vulns

* Fix grader

* Fix trailing whitespaces

* Bump version number

* Remove trailing whitespaces from a git patch...

* Run pre-commit

* Add claude evals

* Address PR claude review

* variant-analysis: fix problems found by testing #232 before release (#237)

* variant-analysis: fix three workflow defects found in a cold run

Prose args killed the run on the first line. The model wrote
`bug: ...; root: /path; lang: python` instead of an object, and the invocation
died with `args.bug is required` before a single agent started. Parse that
shape, and say in whenToUse that args is a JSON object.

The baseline command was not shell-safe. The pattern went through
JSON.stringify, which looks like quoting but yields a double-quoted string where
$(...) and backticks still expand -- and the pattern is model-generated from
codebase content. The root was not quoted at all, so any path with a space broke
the command. Single-quote both.

The sweep had no size floor. It spawned 25 agents against a 5-file fixture,
re-reading in parallel what one agent holds at once. The eval's own negative
result already said so: five small synthetic codebases showed no difference
between the workflow and the skill alone because the fan-out had nothing to buy.
Below 40 source files, sweep two axes in one round -- 7 agents on the same
fixture. The baseline gate reports the file count, and a single-round sweep is
now reported as the deliberate bound it is rather than as a truncated one.

The report stage now has to emit `**Location:**` fields. Without them the
grader falls through to a permissive path its own docstring calls over-counting,
which is what happened on the cold run: a real report scored through the
fallback and nothing said so.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* variant-analysis: score construct spans, not line proximity

A cold run scored a correct report as wrong. The report flagged a helper at
lines 4 and 7 of a file whose safe site began at line 10; LINE_WINDOW=30
credited it as the safe site being reported as real, and the run failed. Two
different functions three lines apart, conflated.

Ground truth now records a `span` per site -- the function's real line range --
and a reported location has to fall inside it. verify_fixtures.py fails if a
span stops containing its own anchor line, so a stale hand-edit cannot
reintroduce the failure silently. LINE_WINDOW drops 30 -> 12 as the fallback for
entries carrying no span.

Line-less mentions now lean opposite ways for recall and precision, and both
directions favour not failing a run that did the work. A report naming the right
file without a line is still credited for recall. It is no longer treated as
claiming the decoy: the decoy's file in the real fixture also holds a genuine
upstream finding, so any run reporting the real one without a line number was
marked as having flagged the decoy.

Three self-tests added, all reduced from the cold run. Both fixes were
mutation-checked: reverting the span logic and reverting require_line each fail
the suite.

Also removeprefix("./") for lstrip("./"), which took a character set and ate the
leading dot of paths like .github/scripts/x.py.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* variant-analysis: surface loose scoring, plumb --strict-decoy, parse the workflow

summarize.py prints a `loose` column counting runs scored through score.py's
permissive fallback. A score built on it is worth less than one built on
location fields, and that was invisible.

--strict-decoy was documented in the README and implemented in score.py but
unreachable from eval.sh, which exited 2 on the unknown option. Plumbed through.
The usage header also advertised `--codebase go`, left over from the five
synthetic codebases; gradio is the only one, and passing both modes needs
quoting.

run_fixtures.sh now runs `node --check` on the workflow. It is the only
JavaScript in the repo and nothing in CI parses it, so a syntax error would
surface only inside a paid eval.sh run. Skipped, not failed, where node is
absent.

setup-gradio.sh reported "the checkout is not at $SHA" for any failed
apply --check, including a checkout at the right SHA whose patch is already
partly applied -- reachable, since the unpatched probe only looks at one of the
three files. Name both causes and the recovery.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* variant-analysis: drop eval graders no arm can fail, correct the firing claim

The skill-not-fired graders on cases 06-07 set `arm: both`, which makes them
scored, and neither arm can fail them: the baseline arm has no plugin so Skill
never fires, and the with-plugin arm does not fire on these shapes either. The
suite's own guidance says a grader no arm can fail is worth deleting rather than
reweighting. The type: llm grader on each case carries the real check.

The "skill does not fire" limitation was overstated as a property of the skill.
A 9-run cold run across three prompt shapes locates the actual cause: it fires
2/3 on a conversational prompt and 3/3 on the description's trigger language
when there is a codebase on disk, and 0/3 on an inline candidate panel -- which
is the shape of every case in this directory. With nothing to sweep, declining
the skill is arguably correct. Giving these cases files on disk would fix the
saturated delta and the trigger rate at once; that is the highest-value change
left here and it is not a small one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* variant-analysis: describe the trigger that actually fires

The skill description was generic where the measured trigger is specific: a bug
just found in a named file, and the question of where else it occurs. It now
leads with that situation and names the bare conversational form, which is what
fired 2/3 in a cold run. The old description was diagnosed as the reason the
skill never fired; it was not, but it was still vague.

The README's entry-point table claimed the skill is "best for a narrow search
where you want a say in each generalization" and triggers on its own. Measured
on a real codebase, Claude reaches for the workflow in 4 of 5 firing runs and
the skill in 1 of 9 -- so ask for the skill by name if you want to weigh in.
Also records the size floor, and that args is a JSON object.

tests/README.md documents spans, the recall/precision asymmetry on line-less
mentions, and the loose column.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Document dynamic workflow layout; variant-analysis 2.0.1

AGENTS.md described only skills/<skill>/workflows/, the prose step-by-step kind,
so the plugin-root workflows/*.js layout that ships as /<plugin>:<workflow> was
undocumented -- and variant-analysis is the first plugin in the repo to use it.
Names both, says which one a "Phase 1 / for each / repeat until" SKILL.md
belongs in, and records that ${CLAUDE_PLUGIN_ROOT} is unavailable inside a
workflow script.

Version bumped 2.0.0 -> 2.0.1 since these are behavioural changes on top of an
unmerged 2.0.0. Squash it back to 2.0.0 if you would rather ship one version.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* variant-analysis: close gaps found by review of the fix PR

A line-less claim on the safe site's file fell into the gap between the strict
accusation check and the permissive examined check: it stayed out of
decoy_reported_as_real (correct -- it names no line), matched `known`
permissively so it dropped out of unreviewed_findings, and then satisfied
decoy_examined_and_ruled_out. A run was credited with correctly ruling out the
site it had just listed under Findings, and passed even under --strict-decoy.
Now surfaced as decoy_claimed_without_line, kept visible in
unreviewed_findings, and it blocks the ruled-out credit without counting as a
false positive.

The small-tree bound could drop expansion axes with no record in the artifact.
With axesPerRound=2 and one round, a 6-axis root cause left four
generalizations unattempted and only the live progress log said so; REPORT.md
was indistinguishable from an exhausted sweep. The report prompt and the return
value now carry swept/total axes and name the unswept ones.

Spans are exact def..return, which left no room for a decorator directly above
a def. RECALL_PAD=3 covers that on the recall side only; the safe site gets no
slack, since padding it walks back into the conflation the spans fixed.

verify_fixtures.py now requires a span on every entry and validates the range.
Without that, a dropped span silently reverted the grader to a proximity window
with a green suite, while ground-truth's own comment documented a guarantee that
no longer held.

source_file_count was `rg --files | wc -l`, which counts assets and fixtures.
A 25-source-file project behind 300 fixtures reported 325 and missed the floor
it was built for. The prompt and the schema now ask for source files only.

`node --check` runs against an .mjs copy. On a .js file whose first statement is
`export`, it only passes on Node ~22.7+, and lint.yml pins no Node version.

Also: pinned the extraction-mode labels as constants with a self-test, so
renaming one cannot leave summarize.py's loose column reading zero forever;
fixed a self-test fixture whose span did not contain its own anchor line, a
shape verify_fixtures.py now rejects; dropped a dead condition in parseArgs;
third-person skill description per AGENTS.md.

score.py self-test 16 -> 18 checks, summarize.py 6 -> 7. The label-rename and
span-removal mutations were both confirmed to fail the suite.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Fix node check. Claude workflows have syntax like top-level returns that will trip the linter

* Add trailing newline

---------

Co-authored-by: kz-tob <kara.zaffarano@trailofbits.com>
Co-authored-by: Clinton Thomas <1033162+KernelClint@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 09:13:41 -04:00

810 lines
30 KiB
Python
Executable File

#!/usr/bin/env python3
"""Grade a variant-analysis report against ground truth.
Grades the ARTIFACT, not the transcript. A run that talks convincingly about
finding variants but writes no report scores as a failure, not as zero findings.
That distinction is the whole point: an eval that reads the response text will
pass a run that skipped the work.
Usage:
score.py --report REPORT.md --codebase cpp [--ground-truth ground-truth.json]
score.py --self-test
"""
import argparse
import json
import pathlib
import re
import sys
# A finding must name a file with one of these extensions to be counted. Adding a
# codebase in a language that is missing here does not score zero silently:
# reported_locations() raises GradingError when Location fields parse to nothing.
PATH_RE = re.compile(
r"[\w./\\-]+\.(?:c|h|cpp|hpp|cc|go|js|mjs|ts|tsx|java|kt|py|rb|rs|php|cs|swift|scala)\b",
re.IGNORECASE,
)
# How real reports declare where a finding lives. Observed across actual runs:
# "**Location:** `src/a.cpp:22`" and "- **Location:** `/abs/path/handlers/a.go:23`".
# Both markdown spellings occur: "**Location:**" (colon inside the bold markers,
# which is what the template produces) and "**Location**:".
LOCATION_RE = re.compile(
r"^\s*[-*]?\s*\*\*\s*(?:location|file)\s*:?\s*\*\*\s*:?",
re.IGNORECASE,
)
# A block or row carrying one of these is an entry the report itself rejected.
# Without this, a triage table row like
# | 3 | `handlers/status.go:37` | REFUTED | allowlist severs the flow |
# scores as the decoy being reported as real — the opposite of what happened.
# Matched against block HEADERS and individual table rows only, never against
# block prose. A real finding routinely explains the safe fix ("use argv
# separation instead"), and matching that text inside the body silently voided
# the whole finding. "safe" is dropped for the same reason: too ambiguous to be
# a verdict token.
REFUTED_RE = re.compile(
r"\b(refuted|false[ -]positive|not a variant|not vulnerable|"
r"not exploitable|ruled out|no finding)\b",
re.IGNORECASE,
)
class GradingError(Exception):
"""The report could not be graded at all — distinct from scoring zero."""
def split_sections(text):
"""Map each '## Heading' to its body."""
sections = {}
current = None
buf = []
for line in text.splitlines():
if line.startswith("## "):
if current is not None:
sections[current] = "\n".join(buf)
current = line[3:].strip().lower()
buf = []
else:
buf.append(line)
if current is not None:
sections[current] = "\n".join(buf)
return sections
def find_section(sections, *keywords):
for name, body in sections.items():
if any(k in name for k in keywords):
return body
return None
def paths_in(text):
"""Normalized (path, line) pairs mentioned in a chunk of report text.
Line is None when the report named a file without one. Ranges like
"foo.py:279-285" keep the first number.
"""
if not text:
return set()
out = set()
for m in PATH_RE.finditer(text):
# removeprefix, not lstrip("./"): lstrip takes a character SET, so it also ate
# the leading dot of `.github/scripts/x.py`.
p = m.group(0).replace("\\", "/").removeprefix("./")
tail = text[m.end() : m.end() + 12]
lm = re.match(r"[:# ]L?(\d+)", tail)
out.add((p, int(lm.group(1)) if lm else None))
return out
def split_blocks(body):
"""Split a section body into '### ' blocks, with the preamble first."""
blocks = []
current = []
for line in body.splitlines():
if line.startswith("### "):
blocks.append("\n".join(current))
current = [line]
else:
current.append(line)
blocks.append("\n".join(current))
return blocks
def reported_locations(findings_body):
"""Paths the report asserts are real findings.
Scraping every path in the Findings section over-counts badly: it picks up
entry-point files named while tracing data flow ("flows unmodified from
`main.cpp:28`") and rows in triage tables the report itself refuted.
So prefer explicit '**Location:**' declarations inside non-refuted blocks,
which is how every real report observed so far marks a finding. Fall back to
permissive line scanning only when a report uses no Location fields at all,
and report which mode was used so a surprising score can be traced.
"""
strict = set()
location_lines = 0
location_paths = 0
for block in split_blocks(findings_body):
lines = block.splitlines()
if not any(line.strip() for line in lines):
continue
# Counted before the refutation checks: this is about whether PATH_RE can
# read the report at all, which has nothing to do with the verdicts in it.
for ln in lines:
if LOCATION_RE.match(ln):
location_lines += 1
location_paths += len(paths_in(ln))
# Header-only refutation check: "### Ruled out: foo.go" voids the block,
# but a body sentence about the safe alternative does not.
header = next((line for line in lines if line.startswith("### ")), "")
if header and REFUTED_RE.search(header):
continue
# Status-row refutation. variant-report-template.md puts the verdict in its
# own `| Severity | Confidence | Status |` row, separated from both the
# header and the **Location:** line — so a decoy written up exactly per the
# template scored as reported-real, failing a run that triaged correctly.
# Only rows carrying no path of their own count as a verdict on the block:
# a triage row that names a file is a verdict on *that* file and is handled
# by the per-line check below, and treating it as block-wide would void the
# real findings listed beside it.
rows = [ln for ln in lines if ln.lstrip().startswith("|") and not paths_in(ln)]
if any(REFUTED_RE.search(ln) for ln in rows):
continue
for line in lines:
if LOCATION_RE.match(line) and not REFUTED_RE.search(line):
strict |= paths_in(line)
if strict:
return strict, STRICT_MODE
# Location fields were declared but PATH_RE recognized nothing in *any* of
# them: the report names files in a language the extension allowlist does not
# cover. Falling through to permissive scanning would score every such run as
# "found nothing", indistinguishable from a run that genuinely found nothing —
# so add the codebase's extensions to PATH_RE instead. Note this fires only
# when zero locations parsed; a report whose locations all parsed and were all
# refuted legitimately scores zero rather than raising.
if location_lines and not location_paths:
raise GradingError(
f"the report declares {location_lines} **Location:** field(s) but no "
"recognizable file path was extracted from any of them — the codebase's "
"language is probably missing from PATH_RE's extension allowlist"
)
loose = set()
for line in findings_body.splitlines():
if REFUTED_RE.search(line):
continue
loose |= paths_in(line)
return loose, PERMISSIVE_MODE
# How far a reported line may sit from the ground-truth line and still be the same
# construct, when ground truth gives no explicit span. A report may cite the def, the
# sink inside it, or a range.
#
# Tightened from 30 after a cold run scored wrong: a report flagged `_concat_file` at
# lines 4 and 7 of a fixture whose safe site began at line 10, and a 30-line window
# credited it as "the safe site reported as real" — two different functions three lines
# apart, conflated. Prefer an explicit `span` in ground truth; this is the fallback for
# entries that do not carry one.
LINE_WINDOW = 12
# Spans are exact (`def` line through the closing `return`), which leaves no room for a
# report that cites a decorator or an overload stub sitting immediately above the def.
# Pad the *recall* side only: crediting a variant a line or two early costs nothing,
# whereas padding the safe site would walk straight back into the conflation this span
# work exists to prevent.
RECALL_PAD = 3
# The extraction mode summarize.py counts in its `loose` column. Defined here, where it is
# produced, so a rename cannot leave that column silently reading zero forever.
PERMISSIVE_MODE = "permissive-lines"
STRICT_MODE = "location-fields"
def same_file(reported_path, truth_file):
"""Directory-aware path comparison.
Suffix matching in both directions already covers every legitimate spelling:
an absolute path from the run's cwd, the repo-relative path, and a bare
basename (`flagging.py` is a suffix of `gradio/flagging.py`). A bare-basename
fallback on top of that would only ever fire for the case suffix matching
deliberately rejects — a same-named file in a *different* directory, such as
gradio's `client/python/gradio_client/flagging.py` against a ground truth of
`gradio/flagging.py` — handing out free true positives in any repo with
duplicated filenames.
"""
truth = truth_file.replace("\\", "/")
r = reported_path.replace("\\", "/")
return r == truth or r.endswith("/" + truth) or truth.endswith("/" + r)
def truth_span(truth, pad=0):
"""The line range a reported location must fall in to be this construct.
An explicit `span: [start, end]` in ground truth is the construct's real
boundaries — the function it lives in. Without one, fall back to a window
around the recorded line. Returns None when ground truth records no line at
all, meaning any line in the right file counts.
"""
span = truth.get("span")
if span:
return int(span[0]) - pad, int(span[1]) + pad
line = truth.get("line")
if line is None:
return None
return line - LINE_WINDOW - pad, line + LINE_WINDOW + pad
def matches(reported, truth, require_line=False, pad=0):
"""True if a reported location refers to the ground-truth construct.
File match alone is NOT enough. A real codebase puts several unrelated
constructs in one file: gradio's screen_recording_utils.py holds this eval's
decoy at line 14 and a genuine upstream finding at line 279. Scoring on
filename alone counted that upstream finding as "the decoy reported as real"
— inverting the result on a run that had done nothing wrong.
`require_line` sets which way a line-less mention leans, and the two callers
lean opposite ways on purpose. Both directions favour not failing a correct
run:
- Recall (did it find the planted variant?) stays permissive. A report that
names the right file without a line is credited; refusing to would be
harsher than the evidence supports.
- The decoy-reported-as-real check is strict. A line-less mention of a file
that happens to hold the decoy is not evidence the decoy was claimed, and
treating it as such fails a run for a sentence about a different function.
This was a live false-failure path: the decoy's file also holds a genuine
upstream finding, so any run that reported the real one without a line
number was marked as having flagged the decoy.
"""
lo_hi = truth_span(truth, pad)
for r, line in reported:
if not same_file(r, truth["file"]):
continue
if lo_hi is None:
return True
if line is None:
if require_line:
continue
return True
if lo_hi[0] <= line <= lo_hi[1]:
return True
return False
def grade(report_text, entry):
sections = split_sections(report_text)
findings_body = find_section(sections, "finding", "variant", "confirmed")
fp_body = find_section(sections, "false positive", "ruled out", "not a variant")
if findings_body is None:
raise GradingError(
"no findings section in the report — expected a '## Findings' heading. "
"The run did not produce a gradeable artifact."
)
reported, extraction_mode = reported_locations(findings_body)
# "Examined" is deliberately permissive: any mention anywhere counts as
# having looked at it, including a refuted row inside Findings.
examined = paths_in(findings_body) | paths_in(fp_body)
vulns = entry["vulnerabilities"]
decoy = entry["decoy"]
found = [v for v in vulns if matches(reported, v, pad=RECALL_PAD)]
missed = [v for v in vulns if not matches(reported, v, pad=RECALL_PAD)]
# Strict on the accusation, permissive on "did it look at it". See matches().
decoy_reported = matches(reported, decoy, require_line=True)
decoy_examined = matches(examined, decoy)
# A claim on the decoy's file with NO line is the gap between those two. Strictness
# keeps it out of decoy_reported, which is right — it is not evidence the decoy was
# named. But it must not then be laundered into "examined and correctly ruled out":
# that credited a run for triaging the very site it had just listed under Findings.
# Surfaced instead, and it blocks the ruled-out credit without counting as a
# false positive.
decoy_line_less_claim = sorted(
p for p, ln in reported if ln is None and same_file(p, decoy["file"])
)
# Findings that are none of the three injected sites.
#
# These are NOT false positives. The fixture is a real 772-file project that
# contains its own issues: a run found `create_subprocess_shell` with an
# interpolated path at screen_recording_utils.py:279, which is upstream
# gradio code and a genuine instance of the same root cause. Ground truth
# only knows what was injected, so it cannot judge these — calling them false
# positives punished the workflow for sweeping wider than the baseline, which
# is the exact behaviour the eval exists to reward.
#
# They are surfaced for a human to read and deliberately excluded from the
# verdict. Only the injected decoy is a definite false positive.
known = list(vulns) + [decoy]
unreviewed = sorted(
f"{p}:{ln}" if ln else p
for p, ln in reported
if not any(matches({(p, ln)}, k) for k in known)
)
# A line-less claim on the decoy's file matches `known` permissively, so it would
# drop out of `unreviewed` too and leave no trace anywhere in the artifact. Put it
# back: a human reading the score needs to see the claim that was made.
for p in decoy_line_less_claim:
if p not in unreviewed:
unreviewed.append(p)
unreviewed.sort()
non_seed = [v for v in vulns if not v.get("is_seed")]
non_seed_found = [v for v in found if not v.get("is_seed")]
return {
"codebase": entry["name"],
"extraction_mode": extraction_mode,
"true_positives": len(found),
"found_ids": [v["id"] for v in found],
"missed_ids": [v["id"] for v in missed],
"new_variants_found": len(non_seed_found),
"new_variants_total": len(non_seed),
"non_seed_recall": f"{len(non_seed_found)}/{len(non_seed)}",
"decoy_reported_as_real": decoy_reported,
"decoy_examined_and_ruled_out": (
decoy_examined and not decoy_reported and not decoy_line_less_claim
),
"decoy_claimed_without_line": decoy_line_less_claim,
"unreviewed_findings": unreviewed,
"false_positives": 1 if decoy_reported else 0,
}
def verdict(score, require_decoy_examined=False):
"""Pass criteria. Kept separate from grading so thresholds are visible.
Keyed on NEW variants, not total true positives. The seed bug is handed to
the run, so whether it reappears under '## Findings' or under '## Original
Vulnerability' is a report-formatting convention — both were observed in
real runs, and scoring on the total penalized the one that followed the
template correctly. What the eval is actually measuring is whether the
second, unseeded vulnerability was found.
"""
reasons = []
if score["new_variants_found"] < score["new_variants_total"]:
reasons.append(
f"found {score['new_variants_found']}/{score['new_variants_total']} "
f"new variants; missed {', '.join(score['missed_ids'])}"
)
if score["decoy_reported_as_real"]:
reasons.append("decoy reported as a real finding")
# unreviewed_findings deliberately does NOT fail the run: they are findings in
# real upstream code that ground truth cannot adjudicate. Read them by hand.
if require_decoy_examined and not score["decoy_examined_and_ruled_out"]:
reasons.append("decoy was never examined (not in the ruled-out section)")
return (not reasons), reasons
# --------------------------------------------------------------------------
# Self-test: proves the grader still discriminates. A grader that cannot fail
# is worth nothing, so this asserts on both directions and on a fixed count.
# --------------------------------------------------------------------------
SELF_TEST_ENTRY = {
"name": "selftest",
"vulnerabilities": [
{"id": "v1", "file": "src/a.py", "line": 1, "is_seed": True},
{"id": "v2", "file": "src/b.py", "line": 2, "is_seed": False},
],
"decoy": {"id": "d", "file": "src/decoy.py", "line": 3},
}
PERFECT = """
## Findings
### Variant #1
**Location:** `src/a.py:1`
### Variant #2
**Location:** `src/b.py:2`
## False Positive Patterns
| src/decoy.py | 1 | guarded before comparison |
"""
MISSED_ONE = """
## Findings
### Variant #1
**Location:** `src/a.py:1`
## False Positive Patterns
none
"""
DECOY_AS_REAL = """
## Findings
### Variant #1
**Location:** `src/a.py:1`
### Variant #2
**Location:** `src/b.py:2`
### Variant #3
**Location:** `src/decoy.py:3`
"""
SPURIOUS = """
## Findings
### Variant #1
**Location:** `src/a.py:1`
### Variant #2
**Location:** `src/b.py:2`
### Variant #3
**Location:** `src/unrelated.py:9`
"""
NO_FINDINGS_SECTION = """
## Summary
I looked at everything and found two variants. Trust me.
"""
DECOY_NOT_EXAMINED = """
## Findings
### Variant #1
**Location:** `src/a.py:1`
### Variant #2
**Location:** `src/b.py:2`
## False Positive Patterns
none
"""
# The next three are reduced from reports real runs actually produced. Each one
# scored wrong before the extraction rewrite.
# go/workflow: decoy listed in a triage table inside Findings, marked REFUTED.
# Previously scored as "decoy reported as real".
REFUTED_IN_TABLE = """
## Findings
| # | Location | Verdict | Note |
|---|---|---|---|
| 1 | `src/a.py:1` | CONFIRMED | seed |
| 2 | `src/b.py:2` | CONFIRMED | variant |
| 3 | `src/decoy.py:3` | REFUTED | allowlist severs the flow |
### 1. SEED -- the original
- **Location:** `src/a.py:1`
### 2. VARIANT -- the new one
- **Location:** `src/b.py:2`
"""
# cpp/baseline: entry point named while tracing data flow inside an
# exploitability checklist. Previously scored as a spurious finding.
FLOW_MENTION = """
## Findings
### Variant #1
**Location:** `src/b.py:2`
**Exploitability:**
- [x] User-controlled data — flows unmodified from `src/main.py:28`
## False Positive Patterns
| `src/decoy.py` | 1 | guarded |
"""
# cpp/baseline: seed in its own section per the template, only the new variant
# under Findings. Previously scored 1/2 true positives and failed.
SEED_IN_OWN_SECTION = """
## Original Vulnerability
**Location:** `src/a.py:1`
## Findings
### Variant #1
**Location:** `src/b.py:2`
## False Positive Patterns
| `src/decoy.py` | 1 | guarded |
"""
# The decoy written up exactly as variant-report-template.md prescribes: verdict in
# its own Status row, **Location:** on a separate line. Distinct from
# REFUTED_IN_TABLE, where the refuted row carries the path itself.
REFUTED_STATUS_ROW = """
## Findings
### Variant #1
**Location:** `src/a.py:1`
### Variant #2
**Location:** `src/b.py:2`
### Variant #3: Decoy -- argv form
| Severity | Confidence | Status |
|----------|------------|--------|
| N/A | High | Refuted |
**Location:** `src/decoy.py:3`
**Analysis:** uses argv form, not shell. Prefer argv separation everywhere.
"""
# A triage table inside a block must not void the block's own finding just because
# one of its rows refutes a different file.
REFUTED_ROW_BESIDE_REAL = """
## Findings
### Variant #1
**Location:** `src/a.py:1`
### Variant #2
| # | Location | Verdict |
|---|---|---|
| a | `src/decoy.py:3` | REFUTED |
**Location:** `src/b.py:2`
"""
# A language PATH_RE does not know. Must be ungradeable, not a silent zero.
UNKNOWN_LANGUAGE = """
## Findings
### Variant #1
**Location:** `src/a.erl:1`
### Variant #2
**Location:** `src/b.erl:2`
"""
# Same basename, different directory. Must NOT credit the ground-truth variant.
SAME_BASENAME_ELSEWHERE = """
## Findings
### Variant #1
**Location:** `src/a.py:1`
### Variant #2
**Location:** `vendor/pkg/b.py:2`
"""
# Both from a cold run of the workflow, and both scored wrong before this pass.
# A real finding in a *neighbouring construct* of the file that holds the safe site.
# The safe site spans lines 20-30; this finding is at line 7, in a different function.
# A 30-line proximity window credited it as the safe site being reported as real, failing
# a run whose only mistake was existing in the same file.
NEIGHBOUR_CONSTRUCT = """
## Findings
### Variant #1
**Location:** `src/a.py:1`
### Variant #2
**Location:** `src/b.py:2`
### Variant #3 — helper that builds the argument list
**Location:** `src/decoy.py:7`
"""
# The safe site's file named with NO line number, for a genuine issue elsewhere in it.
# Must not count as claiming the safe site: that inverted the verdict on correct runs,
# because the decoy's file in the real fixture also holds an upstream finding.
SPANNED_FILE_NO_LINE = """
## Findings
### Variant #1
**Location:** `src/a.py:1`
### Variant #2
**Location:** `src/b.py:2`
### Variant #3 — unrelated issue, line not pinned
**Location:** `src/decoy.py`
"""
def self_test():
checks = 0
s = grade(PERFECT, SELF_TEST_ENTRY)
ok, why = verdict(s, require_decoy_examined=True)
assert ok, f"perfect report should pass: {why}"
assert s["true_positives"] == 2, s
assert s["decoy_examined_and_ruled_out"], s
assert s["non_seed_recall"] == "1/1", s
checks += 1
s = grade(MISSED_ONE, SELF_TEST_ENTRY)
ok, why = verdict(s)
assert not ok, "missing a variant must fail"
assert s["true_positives"] == 1, s
assert s["missed_ids"] == ["v2"], s
assert s["non_seed_recall"] == "0/1", s
checks += 1
s = grade(DECOY_AS_REAL, SELF_TEST_ENTRY)
ok, why = verdict(s)
assert not ok, "reporting the decoy as real must fail"
assert s["decoy_reported_as_real"], s
assert s["false_positives"] == 1, s
checks += 1
s = grade(SPURIOUS, SELF_TEST_ENTRY)
ok, why = verdict(s)
assert ok, "an unreviewed finding must NOT fail the run"
assert s["unreviewed_findings"] == ["src/unrelated.py:9"], s
assert s["false_positives"] == 0, "an unreviewed finding is not a false positive"
checks += 1
try:
grade(NO_FINDINGS_SECTION, SELF_TEST_ENTRY)
except GradingError:
checks += 1
else: # pragma: no cover
raise AssertionError("a report with no findings section must not grade as 0")
s = grade(DECOY_NOT_EXAMINED, SELF_TEST_ENTRY)
ok, _ = verdict(s, require_decoy_examined=False)
assert ok, "not examining the decoy is only a failure under the strict flag"
ok, _ = verdict(s, require_decoy_examined=True)
assert not ok, "strict mode must require the decoy to be examined"
checks += 1
# Regressions from real runs.
s = grade(REFUTED_IN_TABLE, SELF_TEST_ENTRY)
ok, why = verdict(s, require_decoy_examined=True)
assert not s["decoy_reported_as_real"], f"a REFUTED table row is not a finding: {s}"
assert s["decoy_examined_and_ruled_out"], s
assert s["extraction_mode"] == "location-fields", s
assert ok, f"a run that finds both and refutes the decoy must pass: {why}"
checks += 1
s = grade(FLOW_MENTION, SELF_TEST_ENTRY)
assert s["unreviewed_findings"] == [], f"a data-flow mention is not a finding: {s}"
ok, _ = verdict(s)
assert ok, f"should pass: {s}"
checks += 1
s = grade(SEED_IN_OWN_SECTION, SELF_TEST_ENTRY)
assert s["new_variants_found"] == 1, s
ok, why = verdict(s)
assert ok, f"seed outside Findings is a convention, not a miss: {why}"
checks += 1
s = grade(REFUTED_STATUS_ROW, SELF_TEST_ENTRY)
assert not s["decoy_reported_as_real"], f"a Refuted status row voids the block: {s}"
assert s["decoy_examined_and_ruled_out"], s
ok, why = verdict(s, require_decoy_examined=True)
assert ok, f"the shipped template's refutation shape must pass: {why}"
checks += 1
s = grade(REFUTED_ROW_BESIDE_REAL, SELF_TEST_ENTRY)
assert s["true_positives"] == 2, f"a refuted row about another file is not a block verdict: {s}"
assert not s["decoy_reported_as_real"], s
checks += 1
try:
grade(UNKNOWN_LANGUAGE, SELF_TEST_ENTRY)
except GradingError:
checks += 1
else: # pragma: no cover
raise AssertionError("an unparseable language must be ungradeable, not zero")
s = grade(SAME_BASENAME_ELSEWHERE, SELF_TEST_ENTRY)
assert s["new_variants_found"] == 0, f"a same-named file elsewhere is not the variant: {s}"
assert s["unreviewed_findings"] == ["vendor/pkg/b.py:2"], s
checks += 1
# Cold-run regressions. The decoy's span must contain its own anchor line, which is
# what verify_fixtures.py enforces on real ground truth — so move the line with it
# rather than writing a fixture the repo declares invalid.
spanned = json.loads(json.dumps(SELF_TEST_ENTRY))
spanned["decoy"]["line"] = 22
spanned["decoy"]["span"] = [20, 30]
s = grade(NEIGHBOUR_CONSTRUCT, spanned)
assert not s["decoy_reported_as_real"], (
f"a finding outside the safe site's span is not that safe site: {s}"
)
assert s["unreviewed_findings"] == ["src/decoy.py:7"], s
ok, why = verdict(s)
assert ok, f"a run that found both and flagged a neighbouring construct must pass: {why}"
checks += 1
# A line-less claim on the safe site's file: not an accusation, but not a clean
# triage either. It must stay out of decoy_reported_as_real, stay visible in the
# artifact, and block the ruled-out credit.
s = grade(SPANNED_FILE_NO_LINE, spanned)
assert not s["decoy_reported_as_real"], (
f"a line-less mention of the safe site's file is not a claim about it: {s}"
)
assert s["decoy_claimed_without_line"] == ["src/decoy.py"], s
assert "src/decoy.py" in s["unreviewed_findings"], (
f"the claim must remain visible somewhere in the artifact: {s}"
)
assert not s["decoy_examined_and_ruled_out"], (
f"a site claimed under Findings was not 'correctly ruled out': {s}"
)
ok, why = verdict(s, require_decoy_examined=True)
assert not ok, "strict mode must not pass a run that claimed the safe site line-lessly"
ok, _ = verdict(s)
assert ok, "without --strict-decoy it is not a hard failure, since nothing was accused"
checks += 1
# Recall stays permissive in the same situation: a line-less mention of a real
# variant's file is still credited. The asymmetry is the point.
s = grade(SPANNED_FILE_NO_LINE.replace("src/b.py:2", "src/b.py"), spanned)
assert s["new_variants_found"] == 1, f"line-less recall must still be credited: {s}"
checks += 1
# Spans are exact def..return, so a decorator directly above the def is outside them.
# RECALL_PAD covers that on the recall side only; the safe site gets no such slack.
padded = json.loads(json.dumps(SELF_TEST_ENTRY))
padded["vulnerabilities"][1]["line"] = 20
padded["vulnerabilities"][1]["span"] = [20, 28]
s = grade(PERFECT.replace("src/b.py:2", "src/b.py:18"), padded)
assert s["new_variants_found"] == 1, (
f"a decorator line just above the def must still credit the variant: {s}"
)
s = grade(PERFECT.replace("src/b.py:2", "src/b.py:14"), padded)
assert s["new_variants_found"] == 0, f"but not an unrelated line 6 above it: {s}"
checks += 1
# The label summarize.py's `loose` column counts. Pinned here so a rename cannot
# leave that column silently reading zero.
assert PERMISSIVE_MODE == "permissive-lines", PERMISSIVE_MODE
assert STRICT_MODE == "location-fields", STRICT_MODE
s = grade("## Findings\nA bug in `src/a.py:1` and `src/b.py:2`.\n", SELF_TEST_ENTRY)
assert s["extraction_mode"] == PERMISSIVE_MODE, (
f"a report with no Location fields must report the permissive mode: {s}"
)
checks += 1
expected = 18
if checks != expected:
raise AssertionError(f"self-test ran {checks} assertions, expected {expected}")
print(f"score.py self-test: {checks}/{expected} checks passed")
def main():
ap = argparse.ArgumentParser(description=__doc__)
ap.add_argument("--report")
ap.add_argument("--codebase")
ap.add_argument(
"--ground-truth",
default=str(pathlib.Path(__file__).parent / "ground-truth.json"),
)
ap.add_argument("--strict-decoy", action="store_true")
ap.add_argument("--self-test", action="store_true")
args = ap.parse_args()
if args.self_test:
self_test()
return 0
if not args.report or not args.codebase:
ap.error("--report and --codebase are required unless --self-test")
truth = json.loads(pathlib.Path(args.ground_truth).read_text())
entry = next(
(c for c in truth["codebases"] if c["name"] == args.codebase),
None,
)
if entry is None:
print(f"unknown codebase: {args.codebase}", file=sys.stderr)
return 2
path = pathlib.Path(args.report)
if not path.exists():
print(
json.dumps(
{
"codebase": args.codebase,
"error": f"no report at {path} — the run produced no artifact",
"gradeable": False,
}
)
)
return 3
try:
score = grade(path.read_text(), entry)
except GradingError as exc:
print(json.dumps({"codebase": args.codebase, "error": str(exc), "gradeable": False}))
return 3
ok, reasons = verdict(score, require_decoy_examined=args.strict_decoy)
score["gradeable"] = True
score["pass"] = ok
score["fail_reasons"] = reasons
print(json.dumps(score, indent=2))
return 0 if ok else 1
if __name__ == "__main__":
sys.exit(main())