Transcript diff (pnpm eval:diff) shows with/without outputs with
ANSI-highlighted signals — green for matched methods, red for
hallucinations, yellow for anti-patterns.
Calibration (pnpm eval:calibrate) compares scorer ship/no-ship
decisions against human labels with 80% agreement gate integrated
into --fail-on-regression (requires 10+ labels to activate).