Files
Clayton Kim 7a16699f7c Spike plan 015: measure per-author silhouette reference stability
Read-only spike harness computes the five silhouette_scan.py metrics
over three synthetic voice authors plus the human silhouette fixtures
(as a 4th pseudo-author), then checks extraction (median/IQR
degeneracy), stability (jackknife + subsample swings vs. the
pre-registered 50%-of-fence threshold), and discrimination
(nearest-reference accuracy over held-out docs).

Recommendation: no-go for now. Discrimination is the decisive test and
finds no working signal once leakage (human_fixtures is itself a
human_reference.json build source) and degenerate-tie artifacts are
accounted for -- 2 of 3 independent synthetic authors never matched
their own held-out doc to their own reference. Nothing under scripts/
or evals/ is touched; this is measurement only.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K6CYksdLbXbTAxcAQjvHz5
2026-07-07 06:18:43 -07:00
..