Files
callstack__agent-device/scripts/mutation/report.ts
Michał Pierzchała 423927fdd8 chore(mutation): shrink to report-only — drop the ratchet, baseline and graduation (#1457, #1781) (#1828)
* chore(mutation): shrink the lane to report-only (#1457, #1781 wave 2)

The mutation harness's two real catches (#1474, #1475) both came from humans
reading the weekly score report. The ratchet half never operated: the baseline
was committed exactly twice (8cce0ef6b, 60400d04b), both times with
`stableRuns: 0, gating: false`, and was never updated after the very fixes it
triggered — the weekly job computed a new baseline and then `git checkout --`d
it, uploading a proposal nobody applied in 3+ weeks. A gate nobody arms is
harness weight; the report is the part that paid.

Deletes ratchet.ts + ratchet.test.ts, mutation-baselines/, and every
baseline/graduation/gating path in run.ts (`--update`, `mutation:baseline`).
run.ts now exits non-zero only on a harness failure, never on a score. The
report renders the per-kernel table (kernel, score, killed, survived, total,
timeouts) plus the surviving mutants a strengthening PR works from.

Kernel scoping stays: stryker.config.json and KERNEL_MODULES are untouched.

* fix(mutation): restore denominator coverage and publish the table before judging the shard set

Review of #1828:
- `report.test.ts` re-asserts that Ignored/CompileError/RuntimeError leave the
  denominator — the one behaviour `ratchet.test.ts` covered and nothing replaced.
  A `tally()` edit that counted tool noise would have deflated every published
  score with a green `mutation:test`.
- `assertShardsCoverModules` now runs after `emit()`, so an incomplete shard set
  still publishes the kernels that completed instead of only an error string.
  This makes the workflow comments' claim about the job summary true rather than
  re-wording them down.

* chore(mutation): trigger the affected lane on exactly the paths that can select mutants

The PR lane returns an empty matrix unless the diff touches the harness, so the
kernel-source and `**/*.test.ts` triggers only bought a 1-4 min no-op job on
~96% of PRs. `on.pull_request.paths` is now exactly `LANE_TOOLING` plus the
workflow file, asserted in both directions by workflow.test.ts against the
exported constant — a missing path would let a harness change merge unproven,
an extra one starts a job that can only answer `[]`.

Also drops the workflow header's contradictory scope paragraph: it claimed the
lane selects on kernel sources and any test reaching one, which has not been
true since the ratchet went.

* fix(mutation): score and publish a short shard set before failing on the count

The expected-count check ran inside readShardedReports, before anything was
summarized, so on the weekly's real `--expect-shards 10` one dead shard threw
away the nine that had reported — the earlier reorder only moved the
zero-mutants check. The merge now returns the shard count, and both verdicts
run after emit() with the same exit code and `score` stage.

Regression uses the weekly argument shape (`--expect-shards 10`, one shard
present) and asserts the reporting kernel's row reaches stdout while the run
still fails.
2026-08-18 17:47:29 +02:00

71 lines
2.3 KiB
TypeScript

// Markdown rendering for the mutation lane: GitHub job summary and terminal
// output share one renderer, so the artifact and the console never disagree.
//
// The lane reports and never gates (#1457), so the report is the whole product:
// a per-kernel score table plus the surviving mutants a test-strengthening PR
// would have to kill.
import { moduleById } from './modules.ts';
import type { ModuleScore } from './score.ts';
const DEFAULT_MAX_SURVIVING_LISTED = 20;
export type Provenance = { readonly strykerVersion: string; readonly configHash: string };
function renderRow(score: ModuleScore): string {
const module = moduleById(score.module);
return (
`| \`${score.module}\`${module.label} | ${score.score}% | ${score.killed} | ` +
`${score.survived} | ${score.total} | ${score.timeout} |`
);
}
function renderDetail(score: ModuleScore, maxListed: number): string[] {
const lines = [
'',
`### \`${score.module}\`${score.survived} surviving mutant(s) at ${score.score}%`,
'',
];
for (const mutant of score.surviving.slice(0, maxListed)) {
lines.push(`- \`${mutant.file}:${mutant.line}\` ${mutant.mutator}`);
}
if (score.surviving.length > maxListed) {
lines.push(`- …and ${score.surviving.length - maxListed} more`);
}
return lines;
}
export type RenderOptions = {
readonly title?: string;
readonly maxSurvivingListed?: number;
};
export function renderReport(
scores: readonly ModuleScore[],
provenance: Provenance,
options: RenderOptions = {},
): string {
const maxListed = options.maxSurvivingListed ?? DEFAULT_MAX_SURVIVING_LISTED;
const lines: string[] = [
`## ${options.title ?? 'Mutation score — decision kernels'}`,
'',
`Stryker \`${provenance.strykerVersion}\` · config \`${provenance.configHash}\` · ` +
'report only — this lane never fails a build.',
'',
'| Kernel | Score | Killed | Survived | Total | Timeout |',
'| --- | --- | --- | --- | --- | --- |',
...scores.map(renderRow),
];
for (const score of scores) {
if (score.surviving.length > 0) lines.push(...renderDetail(score, maxListed));
}
lines.push(
'',
'A low score names tests worth strengthening (#1474, #1475 were written from this table); ' +
'it is an input for a human-authored PR, not a verdict.',
);
return `${lines.join('\n')}\n`;
}