Files
google__adk-docs/docs/live/models.md
Kaz Sato 77ac9368c6 docs(live): restore Gemini 3.1 Flash Live and fix broken anchors (#2208)
* docs(live): restore Gemini 3.1 Flash Live and fix broken anchors

Verify the models and limits content merged in #2086 against upstream
documentation.

The platform limits table is correct as merged and stays as is: audio-only
sessions cap at 15 minutes and audio+video at 2 minutes on both backends, a
connection lasts ~10 minutes, and Agent Platform additionally defaults a
conversation session to 10 minutes. The earlier claim that Agent Platform
capped every session at 10 minutes conflated the connection limit with the
session limit.

Restore the coverage that was dropped:

- `gemini-3.1-flash-live-preview` is a real model, released 2026-03-26 and
  documented on AI Studio. It is not available on Agent Platform, which
  supports no Gemini Live 3.x model.
- Mark launch stages. `gemini-live-2.5-flash-native-audio` is the only GA Live
  model; the AI Studio IDs are all preview.
- Give the per-model feature table a second column, since with one column it
  said nothing. 3.1 supports neither proactivity/affective dialog nor
  non-blocking tools, and configures thinking with `thinking_level` rather
  than `thinking_budget`.
- State the `global` location restriction as fact rather than as something to
  go check: Live 2.5 models are not served there.

Also fix two anchors that do not resolve, and a deprecated model ID:

- `live/configuration.md` linked to `#response-modalities`; the heading is
  `## Response modes`.
- `live/evaluation.md` linked to `#audio-user-simulation-live-agents`; the
  heading is `## Audio user simulation for live agents`.
- `tutorials/multi-tool-agent.md` suggested `gemini-2.0-flash-live-001`, which
  was shut down on 2025-12-09.

* Apply suggestion from @joefernandez

---------

Co-authored-by: Joe Fernandez <931947+joefernandez@users.noreply.github.com>
2026-09-08 11:13:48 -07:00

9.5 KiB

Supported models for live agents

Supported in ADKPython v0.1.0

Live agents require a model that can hold a bidirectional connection; a standard Gemini model will not. For the models ADK supports outside live agents, and for non-Gemini providers, see Models for agents.

Live models

Live agents run on models that take audio in and produce audio out, end to end, with no intermediate text-to-speech stage. That is what gives them human-like speech with natural prosody, and it is what a standard Gemini model cannot do over a bidirectional connection.

Model AI Studio Agent Platform
Gemini 2.5 Flash Live gemini-2.5-flash-native-audio-preview-12-2025 (Preview) gemini-live-2.5-flash-native-audio (GA)
Gemini 3.1 Flash Live gemini-3.1-flash-live-preview (Preview) Not available

Gemini 2.5 Flash Live is one model with a different ID on each backend; the features are the same either way. gemini-live-2.5-flash-native-audio is ADK's LlmAgent.DEFAULT_LIVE_MODEL, the only Live model that is publicly available, and the model used in this section's examples.

Gemini 3.1 Flash Live is the newer model and is lower latency, but it is AI Studio only and it drops features that 2.5 has — see Per-model feature support before you switch.

Choosing a backend

Live models are reached through one of two backends. ADK talks to both with the same code; you switch with environment variables, so you can develop on one and deploy on the other.

AI Studio Agent Platform
Full name Google AI Studio Gemini Enterprise Agent Platform
Best for Prototyping, development Production, enterprise
Auth API key (GOOGLE_API_KEY) Cloud credentials (GOOGLE_CLOUD_PROJECT, GOOGLE_CLOUD_LOCATION)
Setup API key only Cloud project setup
Limits Session duration and concurrency Session duration and concurrency

Switch with the GOOGLE_GENAI_USE_ENTERPRISE environment variable (FALSE for AI Studio, TRUE for Agent Platform); no code changes. See the quickstarts for setup.

!!! note "Agent Platform: the global location is not supported"

Live models are not available at `GOOGLE_CLOUD_LOCATION=global`. Use a regional
endpoint such as `us-central1`, `us-east1`, or `asia-northeast1`, and check it against
the endpoint-locations table in
[Agent Platform locations](https://docs.cloud.google.com/gemini-enterprise-agent-platform/resources/locations)
before deploying.

These models produce audio directly, with natural prosody, and detect the conversation language on their own. What you configure on top — voices, transcription, turn detection — is described in Configuration.

One property is fixed at the model level: Live models produce audio only. They do not support the TEXT response modality, so to get text alongside speech you use audio transcription.

Per-model feature support

A few RunConfig and tool settings depend on which model you are running:

Feature Gemini 2.5 Flash Live Gemini 3.1 Flash Live
Proactivity and affective dialog Opt-in via RunConfig Not supported
response_scheduling on tools Supported Not supported; function calling is synchronous, so the model stays silent until you return the tool response
Thinking control thinking_budget thinking_level (minimal, low, medium, high)

!!! warning "Moving from 2.5 to 3.1"

Leaving `RunConfig.proactivity` or `RunConfig.enable_affective_dialog` set is the most
common upgrade failure — remove them. Two more differences bite client code: a single
server event can now carry several content parts at once, so iterate over
`event.content.parts` instead of reading `parts[0]`; and turn coverage now defaults to
including all detected audio activity and video frames, which changes token costs if you
stream video continuously. See the upstream
[migration notes](https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-live-preview#migrating-from-gemini-25-flash-live).

Platform limits and quotas

Both backends cap how long a connection and a session can run and how many sessions run at once. These numbers change, so treat the upstream documentation as authoritative and verify before you rely on a limit in production.

Limit AI Studio Agent Platform
Session duration, audio-only 15 min 15 min
Session duration, audio + video 2 min 2 min
Connection lifetime ~10 min ~10 min
Concurrent sessions See rate limits Up to 1,000 per project on pay-as-you-go; no limit with Provisioned Throughput

Agent Platform additionally caps a conversation session at 10 minutes by default, separately from the audio-only limit above.

Enabling context window compression lets a session be extended past the duration limits. On Agent Platform, request concurrent-session increases from the Cloud Console Quotas page under "Bidi generate content concurrent requests". Verify the current numbers against the AI Studio, Gemini API rate limits, and Agent Platform documentation.

How to handle model names

Read the model name from an environment variable rather than hard-coding it. The same model has a different ID on AI Studio and Agent Platform, so an .env var is what lets one codebase target both backends, and it insulates you from model deprecations.

Recommended Pattern:

import os
from google.adk.agents import Agent

# Use environment variable with fallback to a sensible default
agent = Agent(
    name="my_agent",
    model=os.getenv("DEMO_AGENT_MODEL", "gemini-live-2.5-flash-native-audio"),
    tools=[...],
    instruction="..."
)

Why use environment variables:

  • Backend-specific IDs: The same model is named differently on AI Studio and Agent Platform, so moving between them means changing the model ID. An env var keeps that out of your code
  • Model availability changes: Models are released and deprecated regularly. A live agent written a year ago should not be pinned in code to a model that no longer exists
  • Environment-specific configuration: Use different models for development, staging, and production

Configuration in .env file:

# AI Studio
DEMO_AGENT_MODEL=gemini-2.5-flash-native-audio-preview-12-2025

# AI Studio, if you do not need proactivity, affective dialog, or non-blocking tools
# DEMO_AGENT_MODEL=gemini-3.1-flash-live-preview

# Agent Platform
# DEMO_AGENT_MODEL=gemini-live-2.5-flash-native-audio

!!! note "Environment Variable Loading Order"

When using `.env` files with `python-dotenv`, you must call `load_dotenv()` **before** importing any modules that read environment variables. Otherwise, `os.getenv()` will return `None` and fall back to the default value, ignoring your `.env` configuration.

**Correct order in `main.py`:**

```python
from dotenv import load_dotenv
from pathlib import Path

# Load .env file BEFORE importing agent
load_dotenv(Path(__file__).parent / ".env")

# Now safe to import modules that use environment variables
from google_search_agent.agent import agent
```

**Incorrect order (will not work):**

```python
from dotenv import load_dotenv
from google_search_agent.agent import agent  # Agent reads env var here

# Too late! Agent already initialized with default model
load_dotenv(Path(__file__).parent / ".env")
```

This is a Python import behavior: when you import a module, its top-level code executes immediately. If your agent module calls `os.getenv("DEMO_AGENT_MODEL")` at import time, the `.env` file must already be loaded.

Selecting the right model:

  1. Choose a backend: AI Studio for prototyping, Agent Platform for production. This picks the ID column in the table above, and on Agent Platform it settles the model too — Gemini 2.5 Flash Live is the only Live model there
  2. Check current availability: Refer to the model table above and the official documentation
  3. Configure environment variable: Set the model name in your .env file and read it from there when constructing the agent

Model compatibility and availability

For the latest information on model compatibility and availability:

Always verify model availability and feature support in the official documentation before deploying to production.