Dependency installs and CI jobs both executed code this repo never reviewed. This closes those paths without changing what anything resolves to. pnpm 9 ran every dependency's install lifecycle script and had no allow-list model, so CI executed install hooks from the whole tree on every PR. pnpm 10.34.5 blocks them by default and carries the fix for the fail-open integrity check in CVE-2026-50021. minimumReleaseAge holds freshly published versions out of resolution for 48 hours, the window in which a registry compromise is typically caught and yanked. The two `*` peer ranges in reactor-cli were the only place the workspace opted out of range discipline: any future major satisfied them, including a hijacked one. Both now carry carets on the versions already resolved, and the README install lines are pinned to match, since a 0.x caret does not cross the minor line. Third-party actions take the commit SHA their movable tags resolved to, with a monthly grouped Dependabot entry so those pins do not go stale. GitHub-maintained actions stay on tags as a stated trust decision rather than a claim that pinning them would buy nothing. The benchmark jobs needed the most work, since they run an external repository's code with five LLM provider keys in scope. Dispatch inputs now travel through the step environment as quoted variables instead of being interpolated into shell. The pi version is hardcoded rather than dispatchable: npm accepts package sources after the `@` — an alias, a repository, a tarball URL — so an input in that position chose a package rather than a version. The remaining ref override is documented as an operator escape hatch whose clone target is fixed. The publish job holds a credential that can publish under our name and installed npm at latest before using it. It now pins an exact version above the floor OIDC trusted publishing requires, and fails at that step if the pin does not take. CI also fails on a tampered or unsigned tarball now, checked per publishable package rather than once at the root where npm would only see dev tooling. The advisory audit runs alongside it as a signal, not a gate. Nothing re-resolves: no version or integrity line in the lockfile moves.
LongCoT benchmark runner
This directory hosts the harness-under-test bridge between the LongCoT benchmark (paper · repo) and our first candidate RLM harness, pi-mono (@mariozechner/pi-coding-agent — https://github.com/badlogic/pi-mono).
The GitHub Action lives at .github/workflows/longcot-bench.yml. This README is how you invoke it and what the knobs mean.
What it does
Per run, the workflow:
- Clones LongCoT at the ref you pick,
uv syncs its deps (includesrdkit,chess,sympy, the Python SDKs). npm i -g @mariozechner/pi-coding-agent@0.73.1to get thepibinary. The version is hardcoded in the workflow; change it there.- Runs
run_pi.py(this directory) — our shim. For each selected question it shells out to:with the question prompt on stdin, parses thepi --mode json \ --model <model> --thinking <thinking> \ --no-tools --no-skills --no-extensions --no-context-files --no-session \ --system-prompt "<system_prompt>"message_endJSONL event, and writes one line per question toresponses/*.jsonlin the schema LongCoT'srun_eval.pyconsumes. - Runs
uv run python run_eval.py responses/*.jsonl(unlessrun_eval=false) to score against each domain's deterministic verifier. - Writes a summary table to the job summary and uploads
responses/+results/as artifacts (30-day retention).
We use the library API (longcot.load_questions) rather than run_inference.py, and we use the paper's retry discipline (2 retries on API error = up to 3 attempts per question).
How to trigger
GitHub UI → Actions tab → LongCoT Benchmark → Run workflow, pick inputs, go.
Or via gh:
gh workflow run longcot-bench.yml \
-f model=anthropic/claude-opus-4-5 \
-f difficulty=longcot-mini \
-f max_questions=20
Inputs
| input | default | notes |
|---|---|---|
model |
anthropic/claude-opus-4-5 |
pi model string. Examples: openai/gpt-5.2, google/gemini-3-pro, openrouter/deepseek/deepseek-v3.2. |
thinking |
high |
off / minimal / low / medium / high / xhigh. Paper specifies "highest setting if available"; xhigh is literally pi's highest but high is the provider-native max for most APIs. |
domain |
all |
or one of logic, cs, chemistry, chess, math. |
difficulty |
longcot |
Paper-match: longcot = medium+hard, ~1995 q. Iteration-friendly: longcot-mini = easy, ~507 q. Also accepts easy/medium/hard/all. |
max_questions |
0 |
0 means no cap. Slices after deterministic shuffle. |
offset |
0 |
Skip the first N questions after shuffle — handy for resuming. |
seed |
0 |
Shuffle seed. Same seed → same slice across runs. |
concurrency |
8 |
Parallel pi subprocesses. Matches LongCoT's default num_workers. |
system_prompt |
You are a helpful assistant. |
Pi's default is a coding-agent prompt; we override so the model sees only the problem. |
longcot_ref |
fb96494 (main @ 2026-04-20) |
Git ref of LongCoT repo (branch/tag/SHA). Defaults to a pinned commit, not main: this job runs the checked-out repo's Python with five LLM provider keys in scope. The clone target repository is fixed, so this only selects a ref within it. Pass main explicitly to take upstream HEAD, understanding that you are running unreviewed code with the keys attached. |
run_eval |
true |
If false, produce responses JSONL but skip scoring. |
fallback_judge |
true |
If false, pass --no-fallback to run_eval.py (disables Gemini judge for math/chem borderline cases). |
During early iteration — use longcot-mini
The default (longcot) is set to match the paper's headline numbers for 1:1 comparison. A full longcot run is ~1995 questions × potentially tens of thousands of reasoning tokens each, and can push several hours of wall time.
When iterating on the harness, set difficulty=longcot-mini and max_questions=20 or so. You'll get feedback in minutes, not hours.
# Fast smoke test while iterating:
gh workflow run longcot-bench.yml \
-f model=anthropic/claude-opus-4-5 \
-f difficulty=longcot-mini \
-f max_questions=20 \
-f concurrency=4
# Paper-match headline run — expect 3-4+ hours:
gh workflow run longcot-bench.yml \
-f model=openai/gpt-5.2 \
-f difficulty=longcot
Paper fidelity
Defaults are calibrated to the LongCoT paper so our results are directly comparable to Tables 2-6 and Figure 4.
Matching:
- Single-shot, no pass@k / self-consistency.
- Provider-default temperature / top-p (pi-mono's CLI doesn't expose these; the paper uses "default" — consistent by omission).
- Highest reasoning effort available (
--thinking high). - No tools, no scaffolding, no context files, no skills.
- 2 independent retries on API errors.
- Deterministic per-domain verification, with optional LLM fallback for ambiguous math/chemistry cases.
Known divergences (accepted for now):
- Max output tokens: the paper lets models generate up to the provider's max (e.g. 128K for GPT 5.2). Pi-mono's CLI doesn't expose a
max_tokensflag, so we inherit pi-ai's internal default. This may cap generation below the provider max for long-horizon traces. Revisit by either calling@mariozechner/pi-aidirectly or upstreaming a pi CLI flag. - Fallback judge model: the paper uses GPT-5-mini for ambiguous-case extraction; the LongCoT repo defaults to Gemini. We use the repo default (Gemini) for simplicity.
- System prompt: the paper doesn't specify one. We send
"You are a helpful assistant."because pi's default is a coding-agent prompt and some form of override is required.
Secrets the action reads
Set any of these as repo secrets (unset secrets are fine — only the provider you actually use needs to be present):
ANTHROPIC_API_KEYOPENAI_API_KEYOPENROUTER_API_KEYXAI_API_KEYGOOGLE_API_KEY(also exposed asGEMINI_API_KEYfor the LongCoT fallback judge)
Artifacts
Every run uploads two artifacts (30-day retention):
longcot-responses-<run_id>— the rawresponses/*.jsonl(one line per question withresponse_text,usage,reasoning, errors).longcot-results-<run_id>— the scoredresults/*.json(totals, accuracy,overall_accuracy, per-question verdicts).
Download with gh run download <run-id> -n longcot-results-<run_id>.
Running locally
You can run the shim outside GitHub Actions — useful for one-off debugging or running against your local pi install.
# 1. Clone LongCoT, set up its env
git clone https://github.com/LongHorizonReasoning/longcot
cd longcot
uv sync
# 2. Install pi
npm install -g @mariozechner/pi-coding-agent
# 3. Set a provider key
export ANTHROPIC_API_KEY=sk-ant-...
# 4. Run the shim (adjust the path to run_pi.py for your checkout)
uv run python /path/to/prose/.github/scripts/longcot/run_pi.py \
--model anthropic/claude-opus-4-5 \
--difficulty longcot-mini \
--max-questions 5 \
--concurrency 2
# 5. Score it
uv run python run_eval.py responses/<your-jsonl>
--dry-run on run_pi.py resolves the question list, prints the plan, and exits without calling pi — handy for verifying the slice before spending API budget.
Swapping the harness under test
Pi-mono is the current subject. Replacing it means:
- Swap the
npm installstep in the workflow for however the new harness installs. - Rewrite the per-question invocation inside
run_pi.py— specificallypi_command(...)(the CLI) andrun_pi_once(...)(the subprocess call + output parsing). The JSONL output schema (whatrun_eval.pyconsumes) stays the same.
Keeping run_pi.py's input/output contract stable across harness changes is the whole point of this layout — harness-specific code is isolated to one function.