* [Classifier] Upgrade classification model to Claude Sonnet 4.6
* [Classifier] Parallelize classification with p-limit and add dashboard progress
Run up to 4 classifications concurrently using p-limit instead of a
sequential for-loop. The dashboard now shows "classifying N/M…" so
users can track per-eval progress during the classification phase.
* [Classifier] Switch to Haiku 4.5 for classification
Haiku is faster and cheaper while still capable enough for the
classification task, especially with 4 concurrent classifications.
- Change wire_api from "chat" to "responses" (required by newer Codex CLI)
- Fix direct OpenAI API auth by using codex login --with-api-key
- Fix model name for direct API (remove openai/ prefix)
- Update integration tests to assert agent success (not just completion)
- Fix test isolation with beforeAll hooks for project initialization
Codex via Vercel AI Gateway still has a separate tools.7.type API error
that needs to be fixed on the AI Gateway side.
Adds Docker as a sandbox backend for running evals locally without
requiring Vercel credentials.
Features:
- Auto-detection: uses Vercel if VERCEL_TOKEN present, else Docker
- Explicit override via SANDBOX_BACKEND=docker|vercel env var
- Non-root execution (uid 1000) for security and Claude Code compatibility
- npm global installs configured for non-root user
New files:
- src/lib/docker-sandbox.ts - Docker sandbox implementation
- src/lib/docker-sandbox.test.ts - Integration tests
- scripts/test-docker-sandbox.ts - Manual test script
- scripts/test-docker-eval-flow.ts - Full eval flow test
- DOCKER.md - Beginner-friendly setup guide
API additions:
- createSandbox() - Factory function with auto-detection
- resolveBackend() - Determines which backend to use
- getSandboxBackendInfo() - Returns backend info for display
New npm scripts:
- test:sandbox - Quick Docker sandbox test
- test:docker - Docker unit tests
- test:docker:integration - Full integration tests with Docker
- Rename package from @judegao/eval to agent-eval
- Reset version to 0.0.1
- Update all imports, templates, and docs to use new package name
- Point repository and git remote to vercel-labs/agent-eval
- Fix duplicate evals keys in README config example
- Rewrite environment variables docs with summary table
- Add Vercel employees section for vc env pull workflow
Co-Authored-By: Claude (claude-fudge-eap-cc) <noreply@anthropic.com>
- Add support for direct ANTHROPIC_API_KEY and OPENAI_API_KEY authentication
- New agent naming: 'vercel-ai-gateway/claude-code' vs 'claude-code' for direct
- Change default model to opus
- Add ESLint configuration
- Add runner tests
- Clean up legacy agent.ts code
- Update README with dual authentication documentation
- Add Codex CLI agent alongside Claude Code
- Route all agents through Vercel AI Gateway (single API key)
- Extract shared agent utilities to reduce duplication
- Simplify model types (accept any string, let CLI validate)
- Remove list command (use --dry instead)
- Update README for framework authors and A/B testing
- Rename package from @vercel/eval-framework to @judegao/eval
- Update all references in init templates and docs
- Fix unused SandboxFile import in agent.ts
- Fix Sandbox interface to include options parameter
- Use EXCLUDED_FILES constant in fixture.ts
- Use dirname() instead of join(path, '..') in init.ts
- Make experimentConfigSchema private in config.ts
- Remove unused existsSync import in results.ts
- Consolidate duplicate imports in cli.ts
- Use TEST_FILE_PATTERNS constant in sandbox.ts
- Fix JSDoc comment ordering in sandbox.ts
CLI commands:
- `eval init <name>` - placeholder for project scaffolding
- `eval run <config>` - run experiments with dry run support
- `eval list` - list available eval fixtures
Fixture discovery:
- Discover fixtures in evals/ directory
- Validate required files (PROMPT.md, EVAL.ts, package.json)
- Validate package.json has type: module
- Load fixture contents and prompt text
- Get/read fixture files with exclusion patterns
Includes comprehensive tests for all functionality.