Commit Graph

138 Commits

Author SHA1 Message Date
github-actions[bot] bb7cbed5dd Version Packages 2026-06-11 16:46:15 +00:00
molebox c016ea0319 Fix Claude prompt consumed by variadic --allowedTools
The comma-separated token from #150 prevents the tools list itself from
splitting, but --allowedTools is variadic: it keeps capturing positional
tokens until the next flag. With the prompt directly after the value,
claude 2.1.112 consumed it as another tool name and failed with 'Input
must be provided either through stdin or as a prompt argument when using
--print' — caught by a live a0-local smoke run, invisible to the unit
tests because they asserted the broken order.

Emit --allowedTools first so the always-present
--dangerously-skip-permissions terminates the variadic capture before
the trailing prompt. Verified live: the reordered invocation accepts the
prompt. Default-off argument construction is byte-identical. New test
asserts the token after the allowedTools value is always a flag.
2026-06-11 17:00:37 +02:00
github-actions[bot] 7a23a1377a Version Packages 2026-06-11 14:13:27 +00:00
molebox 17d246965c Include webResearch in result-reuse fingerprints when enabled
Conditional include (same pattern as native-default modelPolicy) so
default-off fingerprints are byte-identical to existing releases, while
research and non-research configs never share a fingerprint — result
reuse must not serve a cached parametric-only result for a research
run, or vice versa.
2026-06-11 15:43:16 +02:00
molebox 1d8321f0c6 Thread webResearch through ExperimentConfig and the runner
The adapters read options.webResearch, but runExperiment built
AgentRunOptions from an explicit field list that never included it, so
the option was unreachable for experiment-config consumers (a0-local
calls runExperiment, not executeAgent). Adds the field to
ExperimentConfig/ResolvedExperimentConfig/RunnableExperimentConfig, the
zod schema (z.object strips unknown keys, so schema membership is
required for validateConfig not to drop it), resolveConfig, and both
agent.run call sites (runExperiment and runSingleEval).

Still default-off: absent config yields webResearch: undefined, which
leaves every adapter branch untaken.
2026-06-11 15:37:29 +02:00
molebox 084d895f66 Add opt-in webResearch option for agent web tools
Safe redo of #141 (reverted in #144). webResearch defaults to false, so
command construction is byte-identical for existing consumers; coding
evals are unaffected unless they opt in.

The #141 breakage is fixed and regression-tested: Claude Code's
--allowedTools is variadic, so WebSearch/WebFetch are passed as a single
comma-separated value instead of separate tokens that consumed the
trailing positional prompt.

Verified against AI Gateway with live spikes: Claude Code WebSearch
executes (tool_use/tool_result events), OpenCode Exa websearch executes,
and Codex researches via shell even though no web_search items appear
through the responses wire (setting kept for direct-OpenAI runs and
future gateway support).
2026-06-11 15:03:59 +02:00
github-actions[bot] c4961d7cbc Version Packages 2026-06-11 06:26:45 +00:00
molebox b4841d6791 Fix OpenCode observed model extraction for OpenCode >= 1.17.0
OpenCode 1.17.0 rewrote its logging pipeline and removed the
service=llm log lines the adapter scraped for providerID/modelID,
so native-default runs silently lost model observation.

Fall back to 'opencode export <sessionID>' when log scraping yields
nothing: the session id comes from the --format json event stream and
the exported assistant message carries providerID/modelID. Observation
never fails the run. The log scrape stays as the first, cheaper source
for OpenCode <= 1.16.x.
2026-06-10 16:54:43 +02:00
github-actions[bot] 1ee6ee852e Version Packages 2026-06-01 17:11:11 +00:00
molebox cedf84b4bc Use native default when model is omitted 2026-06-01 18:48:13 +02:00
molebox aa66c4d35b Add native default model policy 2026-06-01 17:01:36 +02:00
github-actions[bot] 3db1716a00 Version Packages 2026-05-30 00:34:54 +00:00
Allen Zhou 2b7eddf90f Revert "Merge pull request #141 from vercel-labs/feat/enable-agent-web-sources"
This reverts commit daaeccb42c, reversing
changes made to cb2280db61.
2026-05-29 17:30:14 -07:00
github-actions[bot] e0db9a3813 Version Packages 2026-05-29 15:03:07 +00:00
molebox 74be0aa25b Merge main into feat/enable-agent-web-sources 2026-05-29 16:51:52 +02:00
github-actions[bot] a53721752b Version Packages 2026-05-28 18:45:17 +00:00
molebox 2d27942cc3 Enable source-capable web tools for agent runs 2026-05-28 11:15:10 +02:00
Allen Zhou a9efa3ad1f Fix Codex profile reasoning_effort and verbosity defaults
The Codex CLI defaults both `model_reasoning_effort` and `model_verbosity`
to "low", but `gpt-5.2-codex` (the default Codex model) only accepts
"medium" for both. Out-of-the-box `codex exec` against the AI Gateway
fails with:

  Unsupported value: 'low' is not supported with the 'gpt-5.2-codex'
  model. Supported values are: 'medium'.

The error covers both the `reasoning.effort` and `text.verbosity`
request parameters, depending on which the model rejects first.

Set both fields to "medium" in two places:
- the generated profile config in ~/.codex/default.config.toml
- explicit -c flags on `codex exec`, since CLI flags have the highest
  precedence and we observed the profile-only setting being silently
  overridden by the CLI's "low" default in some Codex versions.

`generateCodexConfig` now accepts an optional `reasoningEffort`
parameter so callers can override per-run via
`model: "gpt-5.2-codex?reasoningEffort=high"`.

Verified end-to-end against the Vercel AI Gateway: a previously-failing
`codex exec` smoke run now completes in ~31s and returns a real
response instead of erroring at `turn.failed`.

Also added `vercel-agent-eval-*.tgz` to .gitignore so local `npm pack`
artifacts don't leak into commits.
2026-05-27 20:19:58 -07:00
github-actions[bot] 8bb80230c4 Version Packages 2026-05-28 02:55:06 +00:00
molebox 5950d74405 Fix Codex profile config 2026-05-27 21:36:53 +02:00
github-actions[bot] 10cfa27c51 Version Packages 2026-05-13 08:02:51 +00:00
Allen Zhou ea8d7abba6 Neutralize sandbox workspace path 2026-05-06 16:52:56 -07:00
github-actions[bot] 5294461898 Version Packages 2026-05-05 23:02:49 +00:00
Allen Zhou c52126f198 Keep agent config validation strict 2026-05-05 15:58:10 -07:00
Allen Zhou 07614ec3b7 Add response-only harness support 2026-05-05 15:37:24 -07:00
github-actions[bot] f5cba1ea67 Version Packages (#122)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-04-26 08:51:13 -04:00
Jude Gao 384133b982 [CLI] Surface AI Gateway errors during failure classification (#121) 2026-04-26 02:22:46 -04:00
github-actions[bot] f7e79f7a1f Version Packages (#120)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-04-24 23:33:25 -04:00
Jude Gao 660ea3ea20 [CLI] Remove auto-retry of non-model failures (#118) 2026-04-24 21:52:45 -04:00
github-actions[bot] 4b81d0f676 Version Packages (#116)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-04-17 03:09:12 +08:00
Jude Gao a3c2136f03 [Sandbox] Reconnect on terminated streams to avoid spurious failures on long commands (#115) 2026-04-16 15:06:15 -04:00
github-actions[bot] 69db6fca7b Version Packages (#114)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-04-16 11:39:53 -04:00
Jude Gao a209ae099f [Claude Code] Forward cliPackage and effort agentOptions to the CLI (#113) 2026-04-16 11:38:02 -04:00
github-actions[bot] 4385c53c84 Version Packages (#111)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-04-15 10:53:20 -04:00
Jude Gao f838bd7363 [OpenCode] Pass timeout to provider config (#112) 2026-04-15 10:41:25 -04:00
Jude Gao 481637dd6e Auto-retry non-model failures with configurable retry rounds (#110) 2026-04-14 23:26:27 -04:00
github-actions[bot] 01c6e88322 Version Packages (#109)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-04-14 22:01:08 -04:00
Jude Gao 5185640dde [OpenCode] Deep-merge vercel provider config and use user-space binary path (#108) 2026-04-14 21:57:07 -04:00
github-actions[bot] 38faa84631 Version Packages (#107)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-04-14 15:31:11 -04:00
Rune Botten df0dfb6657 fix(codex): append to config.toml instead of overwriting (#104) 2026-04-14 15:28:28 -04:00
Jerilyn Zheng 8d138a28e7 Add agentOptions support to ExperimentConfig (#106)
Allows experiment config files to pass agent-specific options (like
binaryUrl and extraProviders) at runtime via a new agentOptions field.
Previously these could only be set at agent registration time, making
configs for unreleased models non-replicable.

The options flow: ExperimentConfig → runner → AgentRunOptions → agent.run().

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 15:27:00 -04:00
Jude Gao d4c0a01a08 gpt 5.4 integration test 2026-03-21 13:43:14 -04:00
github-actions[bot] 90f33d8e6c Version Packages (#100)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-03-21 11:20:08 -04:00
Jude Gao ec11c4a6b5 [CLI] Add override flag to dotenv config to allow shell env vars to take precedence (#99) 2026-03-21 11:07:10 -04:00
github-actions[bot] 7c6fee55eb Version Packages (#98)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-03-19 10:51:29 -04:00
Jude Gao 4815babe17 Bump minimatch to 10.2.4 to resolve ReDoS CVE (#97)
* Bump minimatch to 10.2.4 to resolve ReDoS CVE

* Fix lockfile: use npm instead of pnpm
2026-03-19 10:49:47 -04:00
github-actions[bot] c9c0db1d7d Version Packages (#93)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-03-19 01:21:31 -04:00
Jude Gao 6ced2ea189 Strip model prefix from Codex CLI flag for direct API (#95)
* Use built-in OpenAI provider for Codex

* Strip openai/ prefix from --model CLI flag for direct API path

The config.toml correctly strips the provider prefix for direct OpenAI API
usage, but the --model CLI flag still passed the prefixed name (e.g.
"openai/gpt-5.2-codex"), causing a "model not found" error.
2026-03-19 01:03:52 -04:00
Yunfei He 0f9ba7ad7e feat: support CLAUDE_CODE_OAUTH_TOKEN for Claude Code agent (#55)
* feat: support CLAUDE_CODE_OAUTH_TOKEN for Claude Code agent

Allow Claude Pro/Max subscribers to authenticate using their OAuth token
instead of requiring a separate ANTHROPIC_API_KEY. When CLAUDE_CODE_OAUTH_TOKEN
is set in the environment, it takes precedence over ANTHROPIC_API_KEY for
non-AI-Gateway configurations.

Closes #54

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor: address PR review comments for OAuth token support

Refactor nested ternary for sandbox env to if/else for readability,
add clarifying comment about credential consistency, and add unit
tests for getApiKeyEnvVar() precedence (gateway > oauth > direct).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-17 19:33:55 -04:00
github-actions[bot] fb9ac72ea9 Version Packages 2026-02-25 21:39:31 +00:00