Commit Graph

130 Commits

Author SHA1 Message Date
github-actions[bot] 1ee6ee852e Version Packages 2026-06-01 17:11:11 +00:00
molebox cedf84b4bc Use native default when model is omitted 2026-06-01 18:48:13 +02:00
molebox aa66c4d35b Add native default model policy 2026-06-01 17:01:36 +02:00
github-actions[bot] 3db1716a00 Version Packages 2026-05-30 00:34:54 +00:00
Allen Zhou 2b7eddf90f Revert "Merge pull request #141 from vercel-labs/feat/enable-agent-web-sources"
This reverts commit daaeccb42c, reversing
changes made to cb2280db61.
2026-05-29 17:30:14 -07:00
github-actions[bot] e0db9a3813 Version Packages 2026-05-29 15:03:07 +00:00
molebox 74be0aa25b Merge main into feat/enable-agent-web-sources 2026-05-29 16:51:52 +02:00
github-actions[bot] a53721752b Version Packages 2026-05-28 18:45:17 +00:00
molebox 2d27942cc3 Enable source-capable web tools for agent runs 2026-05-28 11:15:10 +02:00
Allen Zhou a9efa3ad1f Fix Codex profile reasoning_effort and verbosity defaults
The Codex CLI defaults both `model_reasoning_effort` and `model_verbosity`
to "low", but `gpt-5.2-codex` (the default Codex model) only accepts
"medium" for both. Out-of-the-box `codex exec` against the AI Gateway
fails with:

  Unsupported value: 'low' is not supported with the 'gpt-5.2-codex'
  model. Supported values are: 'medium'.

The error covers both the `reasoning.effort` and `text.verbosity`
request parameters, depending on which the model rejects first.

Set both fields to "medium" in two places:
- the generated profile config in ~/.codex/default.config.toml
- explicit -c flags on `codex exec`, since CLI flags have the highest
  precedence and we observed the profile-only setting being silently
  overridden by the CLI's "low" default in some Codex versions.

`generateCodexConfig` now accepts an optional `reasoningEffort`
parameter so callers can override per-run via
`model: "gpt-5.2-codex?reasoningEffort=high"`.

Verified end-to-end against the Vercel AI Gateway: a previously-failing
`codex exec` smoke run now completes in ~31s and returns a real
response instead of erroring at `turn.failed`.

Also added `vercel-agent-eval-*.tgz` to .gitignore so local `npm pack`
artifacts don't leak into commits.
2026-05-27 20:19:58 -07:00
github-actions[bot] 8bb80230c4 Version Packages 2026-05-28 02:55:06 +00:00
molebox 5950d74405 Fix Codex profile config 2026-05-27 21:36:53 +02:00
github-actions[bot] 10cfa27c51 Version Packages 2026-05-13 08:02:51 +00:00
Allen Zhou ea8d7abba6 Neutralize sandbox workspace path 2026-05-06 16:52:56 -07:00
github-actions[bot] 5294461898 Version Packages 2026-05-05 23:02:49 +00:00
Allen Zhou c52126f198 Keep agent config validation strict 2026-05-05 15:58:10 -07:00
Allen Zhou 07614ec3b7 Add response-only harness support 2026-05-05 15:37:24 -07:00
github-actions[bot] f5cba1ea67 Version Packages (#122)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-04-26 08:51:13 -04:00
Jude Gao 384133b982 [CLI] Surface AI Gateway errors during failure classification (#121) 2026-04-26 02:22:46 -04:00
github-actions[bot] f7e79f7a1f Version Packages (#120)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-04-24 23:33:25 -04:00
Jude Gao 660ea3ea20 [CLI] Remove auto-retry of non-model failures (#118) 2026-04-24 21:52:45 -04:00
github-actions[bot] 4b81d0f676 Version Packages (#116)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-04-17 03:09:12 +08:00
Jude Gao a3c2136f03 [Sandbox] Reconnect on terminated streams to avoid spurious failures on long commands (#115) 2026-04-16 15:06:15 -04:00
github-actions[bot] 69db6fca7b Version Packages (#114)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-04-16 11:39:53 -04:00
Jude Gao a209ae099f [Claude Code] Forward cliPackage and effort agentOptions to the CLI (#113) 2026-04-16 11:38:02 -04:00
github-actions[bot] 4385c53c84 Version Packages (#111)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-04-15 10:53:20 -04:00
Jude Gao f838bd7363 [OpenCode] Pass timeout to provider config (#112) 2026-04-15 10:41:25 -04:00
Jude Gao 481637dd6e Auto-retry non-model failures with configurable retry rounds (#110) 2026-04-14 23:26:27 -04:00
github-actions[bot] 01c6e88322 Version Packages (#109)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-04-14 22:01:08 -04:00
Jude Gao 5185640dde [OpenCode] Deep-merge vercel provider config and use user-space binary path (#108) 2026-04-14 21:57:07 -04:00
github-actions[bot] 38faa84631 Version Packages (#107)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-04-14 15:31:11 -04:00
Rune Botten df0dfb6657 fix(codex): append to config.toml instead of overwriting (#104) 2026-04-14 15:28:28 -04:00
Jerilyn Zheng 8d138a28e7 Add agentOptions support to ExperimentConfig (#106)
Allows experiment config files to pass agent-specific options (like
binaryUrl and extraProviders) at runtime via a new agentOptions field.
Previously these could only be set at agent registration time, making
configs for unreleased models non-replicable.

The options flow: ExperimentConfig → runner → AgentRunOptions → agent.run().

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 15:27:00 -04:00
Jude Gao d4c0a01a08 gpt 5.4 integration test 2026-03-21 13:43:14 -04:00
github-actions[bot] 90f33d8e6c Version Packages (#100)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-03-21 11:20:08 -04:00
Jude Gao ec11c4a6b5 [CLI] Add override flag to dotenv config to allow shell env vars to take precedence (#99) 2026-03-21 11:07:10 -04:00
github-actions[bot] 7c6fee55eb Version Packages (#98)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-03-19 10:51:29 -04:00
Jude Gao 4815babe17 Bump minimatch to 10.2.4 to resolve ReDoS CVE (#97)
* Bump minimatch to 10.2.4 to resolve ReDoS CVE

* Fix lockfile: use npm instead of pnpm
2026-03-19 10:49:47 -04:00
github-actions[bot] c9c0db1d7d Version Packages (#93)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-03-19 01:21:31 -04:00
Jude Gao 6ced2ea189 Strip model prefix from Codex CLI flag for direct API (#95)
* Use built-in OpenAI provider for Codex

* Strip openai/ prefix from --model CLI flag for direct API path

The config.toml correctly strips the provider prefix for direct OpenAI API
usage, but the --model CLI flag still passed the prefixed name (e.g.
"openai/gpt-5.2-codex"), causing a "model not found" error.
2026-03-19 01:03:52 -04:00
Yunfei He 0f9ba7ad7e feat: support CLAUDE_CODE_OAUTH_TOKEN for Claude Code agent (#55)
* feat: support CLAUDE_CODE_OAUTH_TOKEN for Claude Code agent

Allow Claude Pro/Max subscribers to authenticate using their OAuth token
instead of requiring a separate ANTHROPIC_API_KEY. When CLAUDE_CODE_OAUTH_TOKEN
is set in the environment, it takes precedence over ANTHROPIC_API_KEY for
non-AI-Gateway configurations.

Closes #54

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor: address PR review comments for OAuth token support

Refactor nested ternary for sandbox env to if/else for readability,
add clarifying comment about credential consistency, and add unit
tests for getApiKeyEnvVar() precedence (gateway > oauth > direct).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-17 19:33:55 -04:00
github-actions[bot] fb9ac72ea9 Version Packages 2026-02-25 21:39:31 +00:00
Allen Zhou 5aa83e4efd best effort transcript caputre 2026-02-25 15:38:48 -06:00
github-actions[bot] 49548ba3a0 Version Packages 2026-02-25 19:50:13 +00:00
Allen Zhou eb0eea919a Vercel Sandbox config 2026-02-25 13:49:38 -06:00
github-actions[bot] 512f29972e Version Packages 2026-02-20 22:47:26 +00:00
Allen Zhou 097490384c asserting agent behavior 2026-02-20 13:59:49 -08:00
github-actions[bot] e514739637 Version Packages (#78)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-02-17 22:57:42 -08:00
Jude Gao 065afc909f Revert "[Classifier] Add "eval" failure type for flawed eval tests (#79)" (#82)
This reverts commit 683e681b7d.
2026-02-17 17:10:37 -08:00
Jude Gao 330ec5e8b7 [Classifier] Switch to Haiku 4.5 and parallelize classification (#81)
* [Classifier] Upgrade classification model to Claude Sonnet 4.6

* [Classifier] Parallelize classification with p-limit and add dashboard progress

Run up to 4 classifications concurrently using p-limit instead of a
sequential for-loop. The dashboard now shows "classifying N/M…" so
users can track per-eval progress during the classification phase.

* [Classifier] Switch to Haiku 4.5 for classification

Haiku is faster and cheaper while still capable enough for the
classification task, especially with 4 concurrent classifications.
2026-02-17 17:06:55 -08:00