Commit Graph

235 Commits

Author SHA1 Message Date
Jude Gao cb441f58be fix(ci): remove stale pnpm-lock.yaml to fix changeset publish
Changesets was detecting pnpm-lock.yaml and trying to publish with
pnpm, which isn't installed in CI. The project uses npm.
@vercel/agent-eval@0.10.0
2026-04-14 15:42:58 -04:00
github-actions[bot] 38faa84631 Version Packages (#107)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
2026-04-14 15:31:11 -04:00
Rune Botten df0dfb6657 fix(codex): append to config.toml instead of overwriting (#104) 2026-04-14 15:28:28 -04:00
Jerilyn Zheng 8d138a28e7 Add agentOptions support to ExperimentConfig (#106)
Allows experiment config files to pass agent-specific options (like
binaryUrl and extraProviders) at runtime via a new agentOptions field.
Previously these could only be set at agent registration time, making
configs for unreleased models non-replicable.

The options flow: ExperimentConfig → runner → AgentRunOptions → agent.run().

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-14 15:27:00 -04:00
Jude Gao d4c0a01a08 gpt 5.4 integration test 2026-03-21 13:43:14 -04:00
github-actions[bot] 90f33d8e6c Version Packages (#100)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
@vercel/agent-eval@0.9.5
2026-03-21 11:20:08 -04:00
Jude Gao ec11c4a6b5 [CLI] Add override flag to dotenv config to allow shell env vars to take precedence (#99) 2026-03-21 11:07:10 -04:00
github-actions[bot] 7c6fee55eb Version Packages (#98)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
@vercel/agent-eval@0.9.4
2026-03-19 10:51:29 -04:00
Jude Gao 4815babe17 Bump minimatch to 10.2.4 to resolve ReDoS CVE (#97)
* Bump minimatch to 10.2.4 to resolve ReDoS CVE

* Fix lockfile: use npm instead of pnpm
2026-03-19 10:49:47 -04:00
github-actions[bot] c9c0db1d7d Version Packages (#93)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
@vercel/agent-eval@0.9.3
2026-03-19 01:21:31 -04:00
Jude Gao bb69d68f47 Fix changeset package name to match workspace (#96) 2026-03-19 01:18:42 -04:00
Jude Gao 9dcc63abc6 Use built-in OpenAI provider for Codex (#94) 2026-03-19 01:11:09 -04:00
Jude Gao 6ced2ea189 Strip model prefix from Codex CLI flag for direct API (#95)
* Use built-in OpenAI provider for Codex

* Strip openai/ prefix from --model CLI flag for direct API path

The config.toml correctly strips the provider prefix for direct OpenAI API
usage, but the --model CLI flag still passed the prefixed name (e.g.
"openai/gpt-5.2-codex"), causing a "model not found" error.
2026-03-19 01:03:52 -04:00
Yunfei He 0f9ba7ad7e feat: support CLAUDE_CODE_OAUTH_TOKEN for Claude Code agent (#55)
* feat: support CLAUDE_CODE_OAUTH_TOKEN for Claude Code agent

Allow Claude Pro/Max subscribers to authenticate using their OAuth token
instead of requiring a separate ANTHROPIC_API_KEY. When CLAUDE_CODE_OAUTH_TOKEN
is set in the environment, it takes precedence over ANTHROPIC_API_KEY for
non-AI-Gateway configurations.

Closes #54

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor: address PR review comments for OAuth token support

Refactor nested ternary for sandbox env to if/else for readability,
add clarifying comment about credential consistency, and add unit
tests for getApiKeyEnvVar() precedence (gateway > oauth > direct).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
2026-03-17 19:33:55 -04:00
Allen Zhou 91207502a1 Merge pull request #89 from vercel-labs/changeset-release/main
Version Packages
@vercel/agent-eval@0.9.2
2026-02-25 13:40:35 -08:00
github-actions[bot] fb9ac72ea9 Version Packages 2026-02-25 21:39:31 +00:00
Allen Zhou 5aa83e4efd best effort transcript caputre 2026-02-25 15:38:48 -06:00
Allen Zhou 45ed0f2e6d Merge pull request #87 from vercel-labs/changeset-release/main
Version Packages
@vercel/agent-eval@0.9.1
2026-02-25 11:51:19 -08:00
github-actions[bot] 49548ba3a0 Version Packages 2026-02-25 19:50:13 +00:00
Allen Zhou eb0eea919a Vercel Sandbox config 2026-02-25 13:49:38 -06:00
Allen Zhou 8507f17559 Merge pull request #86 from vercel-labs/changeset-release/main
Version Packages
@vercel/agent-eval@0.9.0
2026-02-20 14:47:44 -08:00
github-actions[bot] 512f29972e Version Packages 2026-02-20 22:47:26 +00:00
Allen Zhou 3a4544ff03 Merge pull request #85 from vercel-labs/allen/inject-o11y
[o11y] asserting agent behavior
2026-02-20 14:46:06 -08:00
Allen Zhou 097490384c asserting agent behavior 2026-02-20 13:59:49 -08:00
github-actions[bot] e514739637 Version Packages (#78)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
@vercel/agent-eval@0.8.0
2026-02-17 22:57:42 -08:00
Jude Gao 065afc909f Revert "[Classifier] Add "eval" failure type for flawed eval tests (#79)" (#82)
This reverts commit 683e681b7d.
2026-02-17 17:10:37 -08:00
Jude Gao 330ec5e8b7 [Classifier] Switch to Haiku 4.5 and parallelize classification (#81)
* [Classifier] Upgrade classification model to Claude Sonnet 4.6

* [Classifier] Parallelize classification with p-limit and add dashboard progress

Run up to 4 classifications concurrently using p-limit instead of a
sequential for-loop. The dashboard now shows "classifying N/M…" so
users can track per-eval progress during the classification phase.

* [Classifier] Switch to Haiku 4.5 for classification

Haiku is faster and cheaper while still capable enough for the
classification task, especially with 4 concurrent classifications.
2026-02-17 17:06:55 -08:00
Jude Gao 620fb473ad [CLI] run-all subcommand options like --dry are intercepted by the parent program (#80)
* [CLI] run-all subcommand options like --dry are intercepted by the parent program

* Add changeset
2026-02-17 16:46:04 -08:00
Jude Gao 683e681b7d [Classifier] Add "eval" failure type for flawed eval tests (#79) 2026-02-17 16:28:03 -08:00
Jude Gao c8bcde36d1 [Runner] Retry eval attempts on 429 rate limiting with exponential backoff (#77)
* [Runner] Retry eval attempts on 429 rate limiting with exponential backoff

* Add changeset

* Add StartRateLimiter and switch to anomaly-based retry detection

- Add StartRateLimiter class to throttle sandbox starts across experiments
  (20 starts per 2s window) to prevent 429s at the source
- Replace 429 string matching with anomaly detection: retry any failure
  that completes in <5s, since real evals take minutes
- Move progress events (eval:start, eval:complete, earlyExit) outside
  runAttempt so they fire correctly after retries

* Fix runner tests: exclude timeouts from retry, bump mock durations above anomaly threshold
2026-02-17 12:06:10 -08:00
Paolo Ricciuti aca7f85003 Merge pull request #76 from vercel-labs/changeset-release/main
Version Packages
@vercel/agent-eval@0.7.1
2026-02-16 20:19:53 +01:00
github-actions[bot] 28b64c8bee Version Packages 2026-02-13 16:08:45 +00:00
Jude Gao 9558ee90b9 Remove debug console.log from saveResults (#75)
Removes a debug console.log statement that was left in the saveResults function. The statement was logging internal state (copyFiles, hasGeneratedFiles, hasDeletedFiles, options, and runData) during result persistence, which is unnecessary noise in production output.

Co-authored-by: Claude Haiku 4.5 <noreply@anthropic.com>
2026-02-13 11:08:00 -05:00
github-actions[bot] a784dcf962 Version Packages (#72)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
@vercel/agent-eval@0.7.0
2026-02-12 16:56:35 -05:00
Jude Gao 087415c73d [O11y] Add transcript parsers for Gemini and Cursor agents (#74)
* [O11y] Add transcript parsers for Gemini and Cursor agents

* Handle Gemini direct-API transcript format

The Gemini agent has two output formats: the CLI format (via OpenCode
framework with part.tool/part.state) and the direct-API format (with
tool_name/parameters and separate tool_result events). The initial
parser only handled the CLI format, leaving direct-API transcripts
with all-zero metrics. This adds handling for both.

* Aggregate Gemini direct-API delta messages into single turns

Contiguous assistant delta messages are now merged into one message
event so totalTurns reflects actual conversation turns instead of 0.

* Add changeset
2026-02-12 16:50:18 -05:00
Jude Gao be7ca1560e [Agents] Add Cursor CLI agent with direct API support (#73)
* [Agents] Add Cursor CLI agent with direct API support

Add Cursor CLI as a new agent to the eval framework, enabling testing against Cursor models using direct API access via `CURSOR_API_KEY`. The implementation follows the established pattern from existing agents (Claude Code, Codex, Gemini) and integrates seamlessly with the sandbox infrastructure.

Cursor CLI uses a different interaction model than other agents. Unlike the `agent chat` subcommand syntax, the correct invocation is to pass the prompt as the first positional argument, then use the `--print` flag to enable non-interactive mode (required for scripts and headless execution). The `--force` flag is added to automatically approve all CLI operations without waiting for user confirmation. This combination ensures the agent completes successfully in sandboxed environments without hanging on interactive prompts.

The agent installation uses Cursor's official shell-based installation method (`curl https://cursor.com/install -fsSL | bash`), which handles proper setup of the CLI environment. Once installed, prompts are executed as `agent "<prompt>" --print --force`, capturing both stdout and stderr as output. The implementation includes transcript extraction that filters for JSON-formatted events when available, falling back to raw text output if no structured format is detected. This matches the approach used by other agents but adapts to Cursor's plain-text output characteristics.

The agent uses `composer-1.5` as the default model and integrates with all eval framework features including file capture, script validation, and test execution. Integration tests verify both basic eval execution and result output structure matching. The agent is fully typed in the framework's type system and added to the configuration schema alongside other direct-API agents.

Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>

* Fix missing --model and --output-format flags, update env var docs

- Pass --model to Cursor CLI so user-configured model is actually used
- Add --output-format stream-json for structured JSONL transcripts
- Update changeset to reflect correct default model (composer-1.5)
- Add GEMINI_API_KEY and CURSOR_API_KEY to README env var table

---------

Co-authored-by: Claude Haiku 4.5 <noreply@anthropic.com>
2026-02-12 15:08:22 -05:00
Jude Gao 8f198d4d18 [Agents] Add Gemini CLI agent with direct API and stream-json transcript support (#71) 2026-02-12 14:15:34 -05:00
Paolo Ricciuti 0b7bb881ca Merge pull request #70 from vercel-labs/changeset-release/main
Version Packages
@vercel/agent-eval@0.6.2
2026-02-12 16:50:44 +01:00
github-actions[bot] b58bc3d8d7 Version Packages 2026-02-12 14:56:51 +00:00
Paolo Ricciuti 1f83ac4097 Merge pull request #69 from vercel-labs/track-newly-created-files
fix: add all the files to track newly created files
2026-02-12 15:56:04 +01:00
paoloricciuti 93c1a6390a fix: add all the files to track newly created files 2026-02-12 15:24:14 +01:00
Paolo Ricciuti d37b219fb4 Merge pull request #68 from vercel-labs/changeset-release/main
Version Packages
@vercel/agent-eval@0.6.1
2026-02-12 09:43:51 +01:00
github-actions[bot] 37dbeaf29c Version Packages 2026-02-11 20:56:57 +00:00
Paolo Ricciuti fc91cfb7e2 Merge pull request #64 from vercel-labs/save-changed-files
fix: use git to get changes files
2026-02-11 21:56:10 +01:00
Paolo Ricciuti 326c87126c Merge pull request #67 from vercel-labs/copilot/sub-pr-64
docs: document copyFiles configuration option
2026-02-11 21:55:45 +01:00
copilot-swe-agent[bot] 1b07100d3c docs: remove "When to use" section from copyFiles docs
Co-authored-by: paoloricciuti <26281609+paoloricciuti@users.noreply.github.com>
2026-02-11 20:54:53 +00:00
copilot-swe-agent[bot] 08505b5cc1 docs: document copyFiles configuration option
Co-authored-by: paoloricciuti <26281609+paoloricciuti@users.noreply.github.com>
2026-02-11 20:49:57 +00:00
copilot-swe-agent[bot] aeee9aaa32 Initial plan 2026-02-11 20:48:21 +00:00
github-actions[bot] 68adf6525d Version Packages (#66)
Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
@vercel/agent-eval@0.6.0
2026-02-11 15:28:24 -05:00
Jude Gao cf50218fe2 Make failure classifier optional (#65) 2026-02-11 15:22:13 -05:00