Files

6.2 KiB

User Simulation for Dynamic Evaluation

ADK projects. agents-cli eval dataset synthesize loads and runs the agent through ADK, so it is unavailable on other frameworks. The rest of the eval loop (eval generate, eval grade, metrics, dataset schema) is framework-agnostic.

File paths below reference the scaffolded layout. Adjust for your project structure if not using /google-agents-cli-scaffold.

When to Use

Use user simulation when fixed prompts are impractical — the agent may ask for information in different orders or respond in unexpected ways. Instead of hand-recording every user/agent turn, let agents-cli eval dataset synthesize ask the Vertex AI evaluation service to generate user scenarios for your agent and then play each scenario against an LLM-backed user simulator. The resulting traces (with full agent_data.turns populated) drop straight into agents-cli eval grade.

A user scenario is a starting_prompt (the user's opening message) plus a free-text conversation_plan (how the simulated user should behave for the rest of the conversation). You don't author these yourself in the agents-cli flow — eval dataset synthesize generates them from your agent's tools and instructions.

For deterministic, hand-authored eval cases (e.g., regression coverage), use the recorded-turns format instead: write agent_data.turns directly in your dataset and run agents-cli eval generate to play it back. See references/dataset_schema.md. agents-cli eval generate requires either a top-level prompt or agent_data on every case; it does not play hand-authored user_scenario cases.


Running eval dataset synthesize

# Synthesize 3 scenarios (default), simulate them, write traces to artifacts/traces/traces_<ts>.json
agents-cli eval dataset synthesize

# Steer scenario generation with an instruction and environment context
agents-cli eval dataset synthesize \
  -n 5 \
  --max-turns 8 \
  --instruction "Customer asking about refunds" \
  --environment-context "E-commerce support; orders are visible by order_id"

# Use a custom model for scenario generation (default: service default)
agents-cli eval dataset synthesize --model gemini-2.5-pro

CLI flags exposed by agents-cli eval dataset synthesize:

Flag What it controls
-n / --count Number of scenarios to generate (default 3)
--instruction Natural-language steering for scenario generation
--environment-context World context the simulator can rely on (e.g., available data)
--model Model used for scenario generation (server-side; not the simulated user model)
--max-turns Cap on user↔agent turns per scenario (default 5)
-o / --output Output path; defaults to artifacts/traces/traces_<ts>.json

synthesize runs your agent locally and reads its config from the agent's .env (the whole file — GOOGLE_GENAI_USE_VERTEXAI, GEMINI_API_KEY, GOOGLE_CLOUD_*, app vars); there are no --project / --region flags. On Vertex AI, GOOGLE_CLOUD_LOCATION also picks the endpoint for the server-side scenario-generation call, which only supports a subset of eval regions — keep it global (the scaffold default) unless you know your region is supported.

Simulator internals are NOT user-configurable from agents-cli. The LLM-backed user simulator that plays the user side runs inside _synthesize_runner.py with hardcoded ADK defaults (gemini-2.5-flash for the user voice, default thinking config, no custom_instructions). Only --max-turns reaches it (as LlmBackedUserSimulatorConfig.max_allowed_invocations). There is no eval_config.yaml key, no --simulator-model flag, and no way to override custom_instructions or model_configuration short of editing _synthesize_runner.py directly.


What synthesize writes

A single JSON EvaluationDataset file at the output path. Each case has:

  • eval_case_id — server-generated UUID
  • user_scenario — the generated {starting_prompt, conversation_plan} (preserved for traceability)
  • agent_data.turns — the full simulated conversation: user events, agent responses, tool calls, tool responses

Because agent_data.turns is fully populated, the file is already a graded-ready trace. Skip eval generate and go straight to eval grade:

agents-cli eval dataset synthesize
agents-cli eval grade   # reads artifacts/traces/ by default

If synthesize fails for some scenarios, the failing cases land in the output with empty agent_data.turns and a stderr warning; the rest still pass through to eval grade.


Compatible Metrics

Synthesized traces are multi-turn and have no ground-truth response, so only the three multi-turn metrics apply (every other built-in 400s on a multi-turn trace, and reference-based metrics have nothing to match against):

Metric Why it works
multi_turn_task_success Adaptive rubric judges whether the simulated user's goal was met
multi_turn_trajectory_quality Adaptive rubric on agent reasoning across turns
multi_turn_tool_use_quality Adaptive rubric on tool calls across turns

Example tests/eval/eval_config.yaml for grading synthesized traces:

metrics_to_run:
  - multi_turn_task_success
  - multi_turn_trajectory_quality
  - multi_turn_tool_use_quality

Run with:

agents-cli eval grade --config tests/eval/eval_config.yaml

The eval_config.yaml file is read by eval run, eval grade, and eval submit. eval dataset synthesize ignores it.


Notes

  • Scenario quality depends entirely on agent metadata. generate_conversation_scenarios reads your agent's instructions and tool descriptions to generate plausible user behaviors. Vague tool descriptions produce vague scenarios. Tighten tool docstrings before running synthesize on a new agent.
  • --max-turns is a hard cap. The simulated user can stop earlier (when its goal is met or it gives up); --max-turns only prevents runaway loops.
  • Re-running synthesize generates new scenarios. There is no seed flag — each invocation produces fresh scenarios. For repeatable regression coverage, write agent_data.turns directly (see references/dataset_schema.md) instead of relying on synthesize.