6.2 KiB
User Simulation for Dynamic Evaluation
ADK projects.
agents-cli eval dataset synthesizeloads and runs the agent through ADK, so it is unavailable on other frameworks. The rest of the eval loop (eval generate,eval grade, metrics, dataset schema) is framework-agnostic.
File paths below reference the scaffolded layout. Adjust for your project structure if not using
/google-agents-cli-scaffold.
When to Use
Use user simulation when fixed prompts are impractical — the agent may ask for information in different orders or respond in unexpected ways. Instead of hand-recording every user/agent turn, let agents-cli eval dataset synthesize ask the Vertex AI evaluation service to generate user scenarios for your agent and then play each scenario against an LLM-backed user simulator. The resulting traces (with full agent_data.turns populated) drop straight into agents-cli eval grade.
A user scenario is a starting_prompt (the user's opening message) plus a free-text conversation_plan (how the simulated user should behave for the rest of the conversation). You don't author these yourself in the agents-cli flow — eval dataset synthesize generates them from your agent's tools and instructions.
For deterministic, hand-authored eval cases (e.g., regression coverage), use the recorded-turns format instead: write
agent_data.turnsdirectly in your dataset and runagents-cli eval generateto play it back. Seereferences/dataset_schema.md.agents-cli eval generaterequires either a top-levelpromptoragent_dataon every case; it does not play hand-authoreduser_scenariocases.
Running eval dataset synthesize
# Synthesize 3 scenarios (default), simulate them, write traces to artifacts/traces/traces_<ts>.json
agents-cli eval dataset synthesize
# Steer scenario generation with an instruction and environment context
agents-cli eval dataset synthesize \
-n 5 \
--max-turns 8 \
--instruction "Customer asking about refunds" \
--environment-context "E-commerce support; orders are visible by order_id"
# Use a custom model for scenario generation (default: service default)
agents-cli eval dataset synthesize --model gemini-2.5-pro
CLI flags exposed by agents-cli eval dataset synthesize:
| Flag | What it controls |
|---|---|
-n / --count |
Number of scenarios to generate (default 3) |
--instruction |
Natural-language steering for scenario generation |
--environment-context |
World context the simulator can rely on (e.g., available data) |
--model |
Model used for scenario generation (server-side; not the simulated user model) |
--max-turns |
Cap on user↔agent turns per scenario (default 5) |
-o / --output |
Output path; defaults to artifacts/traces/traces_<ts>.json |
synthesize runs your agent locally and reads its config from the agent's .env (the whole file — GOOGLE_GENAI_USE_VERTEXAI, GEMINI_API_KEY, GOOGLE_CLOUD_*, app vars); there are no --project / --region flags. On Vertex AI, GOOGLE_CLOUD_LOCATION also picks the endpoint for the server-side scenario-generation call, which only supports a subset of eval regions — keep it global (the scaffold default) unless you know your region is supported.
Simulator internals are NOT user-configurable from agents-cli. The LLM-backed user simulator that plays the user side runs inside _synthesize_runner.py with hardcoded ADK defaults (gemini-2.5-flash for the user voice, default thinking config, no custom_instructions). Only --max-turns reaches it (as LlmBackedUserSimulatorConfig.max_allowed_invocations). There is no eval_config.yaml key, no --simulator-model flag, and no way to override custom_instructions or model_configuration short of editing _synthesize_runner.py directly.
What synthesize writes
A single JSON EvaluationDataset file at the output path. Each case has:
eval_case_id— server-generated UUIDuser_scenario— the generated{starting_prompt, conversation_plan}(preserved for traceability)agent_data.turns— the full simulated conversation: user events, agent responses, tool calls, tool responses
Because agent_data.turns is fully populated, the file is already a graded-ready trace. Skip eval generate and go straight to eval grade:
agents-cli eval dataset synthesize
agents-cli eval grade # reads artifacts/traces/ by default
If synthesize fails for some scenarios, the failing cases land in the output with empty agent_data.turns and a stderr warning; the rest still pass through to eval grade.
Compatible Metrics
Synthesized traces are multi-turn and have no ground-truth response, so only the three multi-turn metrics apply (every other built-in 400s on a multi-turn trace, and reference-based metrics have nothing to match against):
| Metric | Why it works |
|---|---|
multi_turn_task_success |
Adaptive rubric judges whether the simulated user's goal was met |
multi_turn_trajectory_quality |
Adaptive rubric on agent reasoning across turns |
multi_turn_tool_use_quality |
Adaptive rubric on tool calls across turns |
Example tests/eval/eval_config.yaml for grading synthesized traces:
metrics_to_run:
- multi_turn_task_success
- multi_turn_trajectory_quality
- multi_turn_tool_use_quality
Run with:
agents-cli eval grade --config tests/eval/eval_config.yaml
The eval_config.yaml file is read by eval run, eval grade, and eval submit. eval dataset synthesize ignores it.
Notes
- Scenario quality depends entirely on agent metadata.
generate_conversation_scenariosreads your agent's instructions and tool descriptions to generate plausible user behaviors. Vague tool descriptions produce vague scenarios. Tighten tool docstrings before running synthesize on a new agent. --max-turnsis a hard cap. The simulated user can stop earlier (when its goal is met or it gives up);--max-turnsonly prevents runaway loops.- Re-running synthesize generates new scenarios. There is no seed flag — each invocation produces fresh scenarios. For repeatable regression coverage, write
agent_data.turnsdirectly (seereferences/dataset_schema.md) instead of relying onsynthesize.