workflow_dispatch action that runs the LongCoT benchmark using
pi-mono (@mariozechner/pi-coding-agent) as the inference harness.
Defaults (difficulty=longcot, thinking=high, no tools/scaffolding,
2 retries) mirror the paper for 1:1 result comparison; README
covers iteration with longcot-mini + slice inputs.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>