Files
google__adk-docs/docs/live/configuration.md
Kaz Sato 03eccf55e0 docs(live): decompose the dev guide and fix staleness vs adk-python main (#2086)
* docs(live): decompose the development guide into capability pages

Split dev-guide/part1-5 into Sessions, Events, Tools, Workflows, Audio and
video, Configuration, Voice, Supported models, and Build a custom server.
Rewrite index.md as the section Overview with a streaming-type decision table.

Implements Phase 2 of the Live Interactions<>ADK documentation revamp.

* docs(live): drop half-cascade model coverage

Half-cascade models are no longer supported for live agents. Remove the
Native Audio vs Half-Cascade architecture framing from Supported models and
the half-cascade caveats from Voice configuration. The eight prebuilt Live
API voices are kept, relabeled as native-audio voices alongside the extended
Text-to-Speech list.

* docs(live): retire the five-part dev guide and rewire navigation

Delete live/dev-guide/ and live/streaming-tools.md now that their content
lives in the capability pages. Regroup the Live nav into Get started / Build /
Ship / Reference, repoint every partN.md cross-link at its new page and
anchor, and add direct redirects for the removed paths (mkdocs-redirects does
not chain, so streaming/* keys point at final destinations).

* docs(live): point at the API reference instead of pinned source

Swap the RunConfig, Event, SequentialAgent, LiveRequestQueue and
Runner.run_live source-reference notes for Python API reference links.
Implementation pointers with line ranges are left as source links, since they
document internals with no public reference equivalent.

* docs(live): fix docs against adk-python main and drop the bidi-demo links

The bidi-demo sample was removed from adk-samples, so all the source links in
docs/live/ were dead. The sample is not shipped here either, so remove every
reference to it instead of repointing the links.

The code snippets themselves are unchanged. What goes away is only the
scaffolding that pointed at the sample:

- 32 code fences lose their linked 'Demo implementation: file.py:NN-MM' title
  and become plain language-tagged fences.
- The 'Complete Demo Implementation' note in custom-server.md and the 'Demo
  Implementation' note in events.md are dropped; both existed only to link out.
- The 'Learn More' note in tools.md and the model setup step in models.md keep
  their guidance but no longer cite the sample's files.
- Prose that named the demo ('The bidi-demo demonstrates how to...') is
  rewritten to describe the pattern directly.
- The Bidi Demo card and its screenshot are removed from the Live demos section
  of index.md; LensMosaic remains.

Staleness fixes verified against adk-python main:

- StreamingMode.BIDI is inert. Only run_async() reads RunConfig.streaming_mode;
  run_live() never does. Remove it from every run_live()-facing sample and
  rewrite the 'StreamingMode: BIDI or SSE' section around the Runner method you
  call. Keeps the old anchor via attr_list.
- configuration.md: run_live(session=...) is gone; use user_id/session_id.
- tools.md: streaming tools are registered lazily on first model call, not
  scanned up front; the input_stream queue is created only for tools annotated
  with LiveRequestQueue, and stop_streaming resets it to None. The old
  runners.py / function_tool.py line references pointed at unrelated code.
- sessions.md: document DEFAULT_MAX_RECONNECT_ATTEMPTS = 5 and the go_away
  reconnect trigger; correct 'automatic closure in SSE mode', which really only
  happens for the internal queue under support_cfc.
- events.md: audio artifacts require RunConfig.save_live_blob=True;
  get_author_for_event() also keys off llm_response.input_transcription.
- configuration.md: document history_config and the
  initial_history_in_client_content=True that ADK sets when seeding history.

Not changed: get-started/streaming-java.md still sets StreamingMode.BIDI, which
could not be verified without an adk-java checkout.

* Refresh the Live API supported-model list

Checked against the Gemini Live API and Agent Platform model docs:

- models.md: replace the model list with a platform/model/stage table covering
  gemini-3.1-flash-live-preview (Preview, Gemini Live API only),
  gemini-2.5-flash-native-audio-preview-12-2025 (Preview), and
  gemini-live-2.5-flash-native-audio (now GA, not "public preview").
- Document what Gemini 3.1 Live does not support: proactivity, affective
  dialog, async function calling, thinking_budget (it uses thinking_level),
  plus multi-part server events and the turn-coverage default change.
- Note that no Gemini 3.x Live model exists on Agent Platform, and that Live
  API models are unavailable in the `global` location.
- voice.md: replace the Platform Compatibility text, which wrongly said
  proactivity and affective dialog are unavailable on Agent Platform, with a
  per-model support table.
- configuration.md: CFC's model check is a literal `gemini-2` prefix match, so
  it rejects Gemini 3.x; refresh the runners.py line anchor.
- bidi-demo: same model table in the README, the 3.1 option and the regional
  location requirement in .env.example, and an expanded model comment in
  agent.py. The default stays on 2.5 native audio because the demo exposes
  proactivity and affective dialog toggles. Re-anchored the agent.py line
  links in models.md, tools.md, and sessions.md.

* docs(live): align docs with current Live API model capabilities

Verified docs/live/ and docs/runtime/runconfig.md against the Gemini Live
API capabilities guide, the Agent Platform Live API docs, and ADK 2.6.3.

Model consistency:

- response_modalities=["TEXT"] was presented as a valid live configuration
  in configuration.md, events.md and sessions.md. Every Live API model ADK
  supports is a native audio model, and those accept AUDIO only. Reframed
  around AUDIO plus output audio transcription, and kept TEXT where it is
  actually correct: the run_async() / SSE path.
- docs/runtime/runconfig.md configured response_modalities=["AUDIO","TEXT"]
  in all three language samples. A session accepts exactly one modality.
- events.md snippets read event.content.parts[0], which drops content on
  gemini-3.1-flash-live-preview because it sends multiple parts per server
  event -- the failure models.md already warns about. All four snippets now
  iterate over parts.
- tools.md gave the streaming-tools root agent model="gemini-flash-latest",
  which has no Live API support, so the example could not run under
  run_live() on either platform. That alias is still used for the one-shot
  generate_content call inside the tool, where it is correct.
- configuration.md "Standard Gemini Models (1.5 Series) Accessed via SSE"
  described a retired model family and labelled gemini-pro-latest /
  gemini-flash-latest as 1.5 with 2M context.
- sessions.md: document that send_client_content is seeding-only on Gemini
  3.x Live, and that ADK reroutes single-part text to send_realtime_input.
- models.md: gemini-live-2.5-flash-native-audio is the only GA Live API
  model on Agent Platform, not the only one.

Coverage and links:

- configuration.md: document explicit_vad_signal, translation_config,
  avatar_config and model_input_context.
- voice.md: note that ADK picks the live API version (v1alpha / v1beta1),
  so proactivity and affective dialog need no http_options.
- Replace redirecting upstream URLs with their current targets:
  live-guide -> live-api/capabilities, live-session ->
  live-api/session-management, live -> live-api, and
  cloud.google.com/vertex-ai -> the Agent Platform equivalents.

Verified correct, left alone: session and context limits, audio and video
specs, the proactivity / affective dialog model matrix, thinking_level vs
thinking_budget, the support_cfc gemini-2 prefix check, and ADK's AUDIO
default in run_live().

* docs(live): trim the response-modality and SSE material

Every Live API model ADK supports is a native audio model, so a live
session's response modality is always AUDIO and there is nothing to
choose. Shrink the section to the one thing that still matters --
reading text off event.output_transcription.

StreamingMode is only read by run_async(); the SSE tutorial that grew
around it here (protocol diagrams, progressive-streaming walkthrough,
mode-selection table, 1.5-series model list) duplicates
runtime/runconfig.md and describes models that no longer exist. Keep
the inert-BIDI warning and the run_live()/run_async() split, drop the
rest.

Document explicit_vad_signal, translation_config, avatar_config and
model_input_context, which had no coverage at all.

* docs(live): cut duplicated and non-ADK material

Six sections carried weight that did not belong to them:

- sessions.md 'Best Practices for Live API Connection and Session
  Management' restated the Session Resumption and Context Window
  Compression sections verbatim, down to the RunConfig snippets.
  Deleted.
- sessions.md 'Concurrency and Thread Safety' + 'Message Ordering
  Guarantees' explained asyncio.Queue at length and reproduced the
  upstream task already in custom-server.md. Condensed to the three
  properties that actually affect calling code, with a pointer to
  the private _queue attribute dropped.
- sessions.md 'Architectural Patterns for Managing Quotas' was an
  ASCII decision tree and a comparison table for two patterns that
  reduce to one sentence each.
- index.md 'Real-world applications' spent five industry vignettes
  making one point.
- events.md 'Deserializing on the Client' pasted 80 lines of the
  bidi-demo's UI code, calling helpers that no longer exist anywhere
  in these docs. Reduced to the event-shape handling it was meant to
  show.
- audio-video.md 'Handling Image Input at the Client' was 130 lines
  of getUserMedia/canvas/FileReader boilerplate plus a seven-point
  recap of it.

Also fix two dead absolute links: /agents/multi-agents/#workflow-agents-as-orchestrators
(the page now redirects to workflows/index.md and the anchor is gone)
and /live/streaming-tools/ (no such page; the content is in tools.md).

* docs(live): restructure the live docs around ADK ownership

The live section had accumulated content it did not own: backend limits
restated on capability pages, Web Audio API implementation presented as
ADK guidance, and shared concepts re-explained rather than linked.

Applies one rule throughout: if a fact would still be true with the ADK
source deleted, it belongs on models.md or behind an upstream link, not
on a capability page.

- audio-video.md is now the format contract only (505 -> 121). The
  browser mic-capture, ring-buffer playback, and camera-frame code was
  Web Audio API with no ADK in it, had no counterpart in adk-python, and
  no test anywhere. Deleted rather than relocated. The twelve numbered
  'Key Implementation Details' lists restated the code comments directly
  above them; deleted. The streaming-tool lifecycle section duplicated
  tools.md; replaced with a link.
- custom-server.md gains 'Connect a client': what adk web handles
  (16 kHz capture, 24 kHz playback, 1 fps JPEG, transcripts, barge-in),
  where it stops, and the /run_live wire protocol, which was previously
  undocumented. Keeps the one JS snippet that shows ADK's event shape.
  Drops 'Client-side patterns'.
- sessions.md hands its platform-limits table and quota numbers to
  models.md, keeping the session-pool design guidance. The same figures
  had been stated in three places across two pages.
- models.md gains 'Platform limits and quotas' as the single source, and
  loses the 'Key characteristics' list that restated configuration.md.
- configuration.md drops the 'Platform Support' column, which read
  'Both' on 13 of 15 rows and labelled the two exceptions as platform
  constraints when they are model constraints.
- tools.md compresses 'Tool execution context' to the one fact that is
  live-specific: an InvocationContext spans the whole run_live() loop,
  not a single turn.
- workflows.md points at graphs/index.md, the ADK 2.0 graph workflow
  page, rather than the v0.1.0 multi-agent umbrella.
- Six internal links used absolute paths, which mkdocs does not
  validate, so --strict had been silently ignoring them. Now relative.
- Fixes class.="grid cards" in get-started/index.md, which was breaking
  the card grid.

* docs(live): standardize page leads and cut duplicated RunConfig prose

Every live page opened by narrating its own table of contents ("This page
covers X, Y, and Z"), which duplicates the rendered TOC, ages badly when
a heading changes, and spends a paragraph before the reader gets a fact.
evaluation.md already did the better thing: state the shared baseline,
link the canonical page, then cover only the delta. That is now the
convention across the section.

- sessions.md, events.md, configuration.md, audio-video.md,
  workflows.md, tools.md, models.md and get-started/index.md now name
  their non-live counterpart in the lead instead of listing their own
  headings. Three pages had no outbound link to the shared concept at
  all: tools.md to Custom Tools, models.md to Models for agents, and
  workflows.md pointed at the v0.1.0 umbrella rather than graph
  workflows.
- configuration.md drops the custom_metadata section (85 lines) for a
  pointer plus the one live-specific consequence: a run_live() call is a
  single invocation, so metadata is stamped on the whole session rather
  than one turn. runtime/runconfig.md already owns the field.
- configuration.md trims max_llm_calls and save_live_blob to the facts
  that are live-specific — max_llm_calls does not apply to run_live() at
  all, and save_live_blob writes ~1.92 MB per minute per session to two
  services — and drops the generic use-case and best-practice lists.
- custom-server.md replaces 'Key concepts', which re-pasted all three
  code blocks from the complete example directly above it, with prose
  explaining why the two tasks must run concurrently.

Live section: 2820 -> 2211 lines.

* docs(live): reframe pages around capabilities, fix eval config key

* Apply batched suggestions from code review

Co-authored-by: Joe Fernandez <931947+joefernandez@users.noreply.github.com>

* Apply suggestion from @joefernandez

* Apply batched suggestions from code review

Co-authored-by: Joe Fernandez <931947+joefernandez@users.noreply.github.com>

---------

Co-authored-by: Stephen Allen <stephenaallen@google.com>
Co-authored-by: Joe Fernandez <931947+joefernandez@users.noreply.github.com>
2026-09-01 17:15:50 -07:00

20 KiB

Configuration for live agents

Supported in ADKPython v0.1.0Java v0.2.0

RunConfig is where you shape a live session: how the agent sounds, how it transcribes speech, when it decides a turn is over, how much history it keeps, and what limits it runs under. You pass it to Runner.run_live(), and it applies to that session only. Two users of the same agent can run with completely different configurations.

RunConfig is not live-specific; Runtime configuration documents the full class and the fields that apply to run_async(). What follows is the subset that matters under run_live(), plus the voice-facing settings that only exist in a live session.

RunConfig Parameter Quick Reference

This table provides a quick reference for the RunConfig parameters that matter most to live agents:

Parameter Type Purpose Reference
response_modalities list[str] Output format. Live agents must use AUDIO — Live models do not accept TEXT Details
streaming_mode StreamingMode Chunked or single-shot delivery on the run_async() path; not read by run_live() Details
session_resumption SessionResumptionConfig Enable automatic reconnection Details
context_window_compression ContextWindowCompressionConfig Unlimited session duration Details
history_config HistoryConfig Control how prior conversation history is replayed to the Live server Details
max_llm_calls int Limit total LLM calls per session Details
save_live_blob bool Persist audio/video streams Details
custom_metadata dict[str, Any] Attach metadata to invocation events Details
speech_config SpeechConfig Voice and language configuration Voice and language
input_audio_transcription AudioTranscriptionConfig Transcribe user speech Audio transcription
output_audio_transcription AudioTranscriptionConfig Transcribe model speech Audio transcription
realtime_input_config RealtimeInputConfig VAD configuration Voice activity detection
explicit_vad_signal bool Emit voice activity events from the model Details
proactivity ProactivityConfig Enable proactive audio (model-specific) Proactivity and affective dialog
enable_affective_dialog bool Emotional adaptation (model-specific) Proactivity and affective dialog
translation_config TranslationConfig Real-time speech-to-speech translation (translation models only) Details
avatar_config AvatarConfig Render the agent as an animated avatar Details

For more details on configuration options, see RunConfig in the Python API reference.

Import Paths:

All configuration type classes referenced in the table above are imported from google.genai.types:

from google.genai import types
from google.adk.agents.run_config import RunConfig, StreamingMode

# Configuration types are accessed via types module
run_config = RunConfig(
    session_resumption=types.SessionResumptionConfig(),
    context_window_compression=types.ContextWindowCompressionConfig(...),
    speech_config=types.SpeechConfig(...),
    # etc.
)

The RunConfig class itself and StreamingMode enum are imported from google.adk.agents.run_config.

Response modes

The response_modalities setting controls the output format, and a session gets exactly one. For live agents the value is always ["AUDIO"], because every Live model ADK supports accepts no other modality. ADK fills this in for you when you leave it unset, so most live applications never touch the field.

!!! warning "Migrating from response_modalities=["TEXT"]"

Older ADK samples and half-cascade models allowed a text-only live session. That no
longer works: `run_live()` with `["TEXT"]` fails against current Live models, which
only produce audio.

**To get text out of a live agent, read
[`event.output_transcription`](#audio-transcription)**: transcription is enabled
by default in ADK, so deleting the `response_modalities` line is usually the whole fix.

`["TEXT"]` is still correct on the `run_async()` path, which runs on standard Gemini
models. See [Bidi-streaming or SSE](#streamingmode-bidi-or-sse).

Response modality only affects model output — you can always send text, voice, or video input (if the model supports that input modality) regardless of it.

Bidi-streaming or SSE

ADK can reach Gemini over two different endpoints, and the Runner method you call is what picks one:

  • runner.run_live(): ADK opens a WebSocket to the Live API (the bidirectional streaming endpoint via live.connect()). This is what the rest of this guide covers, and it is required for real-time audio and video
  • runner.run_async(): ADK uses HTTP to the standard Gemini API (the unary/streaming endpoint via generate_content_async()). Set RunConfig.streaming_mode = StreamingMode.SSE to stream that response back chunk by chunk

The two model sets barely overlap. Standard Gemini models such as gemini-flash-latest do not hold a bidirectional connection, and the models in Supported models are meant to be driven with run_live(), so choosing a model is part of choosing a Runner method.

!!! warning "Python: StreamingMode.BIDI does not switch ADK to the Live API"

In **Python**, `RunConfig.streaming_mode` is read only on the `run_async()` code path,
where it chooses between a single complete response (`StreamingMode.NONE`, the default)
and chunked delivery (`StreamingMode.SSE`). The `run_live()` path never reads it, so
setting `streaming_mode=StreamingMode.BIDI` has no effect and fails silently. **Calling
`run_live()` is what gets you bidirectional streaming.** ADK's own Python `StreamingMode`
docstring says as much: BIDI "is not used in the standard execution path", and the real
bidirectional behavior "uses a completely different code path that doesn't rely on
`streaming_mode`".

**Java differs.** ADK Java's flow does read `StreamingMode.BIDI`, and the Java quickstart
sets it explicitly on the `RunConfig` it passes to `runLive()`. Follow each language's
quickstart rather than porting the setting across.
# Live API: no streaming_mode needed, calling run_live() is what selects it
run_config = RunConfig(response_modalities=["AUDIO"])
async for event in runner.run_live(..., run_config=run_config):
    ...

This choice affects only how ADK talks to Gemini. Your client-facing architecture is independent: you can build WebSocket servers, REST APIs, or SSE endpoints on either path.

Runtime configuration covers the run_async() and SSE path: streaming_mode values, progressive SSE streaming, and the language-specific configuration.

Miscellaneous Controls

ADK provides additional RunConfig options to control session behavior, manage costs, and persist audio data for debugging and compliance purposes.

run_config = RunConfig(
    # Limit total LLM calls per invocation
    max_llm_calls=500,  # Default: 500 (prevents runaway loops)
                        # 0 or negative = unlimited (use with caution)

    # Save audio/video artifacts for debugging/compliance
    save_live_blob=True,  # Default: False

    # Attach custom metadata to events
    custom_metadata={"user_tier": "premium", "session_type": "support"},  # Default: None
)

max_llm_calls

max_llm_calls caps LLM invocations per invocation context, and Runtime configuration documents it in full.

It does not apply to run_live(). The parameter only guards the run_async() path, so a live session gets no automatic cost ceiling from it. Budget your own: cap session duration, count turns, watch usage_metadata on model events (Metadata), and put a circuit breaker in front of the loop.

save_live_blob

save_live_blob=True persists the session's audio to the session service as references and to the artifact service as files. Despite the name, only audio is persisted today, not video.

Enable it for debugging voice behavior, or for audit trails in regulated environments. Leave it off otherwise: 16 kHz PCM input runs about 1.92 MB per minute per session, written to two services, and that accumulates fast on a voice workload. If you need it in production, sample a fraction of sessions rather than all of them, and set a retention policy on the artifact service — ADK does not expire these for you.

!!! warning "save_live_audio is deprecated"

ADK migrates `save_live_audio=True` to `save_live_blob=True` automatically and warns,
but the shim will be removed in a future release. Update to `save_live_blob`.

history_config

When ADK opens a new Live API connection for a session that already has conversation history, it replays that history to the server. That history includes the model's own past turns, so the server has to be told not to answer them again. ADK handles this for you: before connecting, it sets live_connect_config.history_config.initial_history_in_client_content = True whenever there is history to send and no session resumption handle is in play.

from google.genai import types

# ADK sets this automatically; override only if you need the opposite behavior.
run_config = RunConfig(
    history_config=types.HistoryConfig(
        initial_history_in_client_content=True,
    ),
)

What this means in practice:

  • You normally do nothing. ADK only fills in the value when you have not set one, so an explicit history_config on RunConfig always wins.
  • Reconnections skip history entirely. When ADK reconnects with a session resumption handle, the server already holds the state for that session, so ADK sends no history and does not touch history_config.
  • Symptom if it goes wrong: setting initial_history_in_client_content=False while seeding history makes the model respond to the replayed turns, producing a burst of duplicate answers at the start of the connection.

custom_metadata

custom_metadata attaches an arbitrary JSON-serializable dict to every Event in the invocation, and it behaves the same in a live session as anywhere else — see Runtime configuration.

run_config = RunConfig(
    response_modalities=["AUDIO"],
    custom_metadata={"user_tier": "premium", "session_type": "support"},
)

The live-specific consequence is scope: one run_live() call is one invocation, so the metadata is stamped on every event for the entire streaming session rather than a single turn. Read it back with event.custom_metadata.

!!! warning "Do not put sensitive data in custom_metadata"

Every event carrying this metadata is persisted to the session service. Keep PII,
credentials, and other sensitive values out of it, and encrypt them if you have no
alternative.

RunConfig carries a few more fields that only take effect on the run_live() path. ADK passes them straight through to the live connection, so their exact behavior is defined by the Live API rather than by ADK:

Field Type What it does
explicit_vad_signal bool Asks the model to emit explicit voice activity signals. ADK surfaces them on event.voice_activity instead of inferring turn boundaries from content
translation_config types.TranslationConfig Enables real-time speech-to-speech translation. Takes target_language_code (BCP-47) and echo_target_language. Only supported by translation models such as gemini-3.5-live-translate-preview — not by the models in Supported models
avatar_config types.AvatarConfig Renders the agent as an animated avatar. Takes avatar_name (a prebuilt avatar) or customized_avatar, plus audio_bitrate_bps / video_bitrate_bps
from google.genai import types

run_config = RunConfig(
    response_modalities=["AUDIO"],
    explicit_vad_signal=True,
)

One more field is not live-specific but is often useful in a live session:

  • model_input_context (list[types.Content]): transient context injected into the LLM request for the current invocation only. The Runner does not persist it to the session, which makes it a clean way to supply per-turn grounding (a document the user just opened, a page they are viewing) without polluting conversation history.

Compositional function calling (support_cfc)

Compositional Function Calling (CFC) is a run_async() / SSE feature, not a live one: it applies to the current Live models only in theory, since none of them satisfy its model requirement. Leave support_cfc for the SSE path and use standard function calling in live sessions (see Tools). For the parameter itself, see Runtime configuration.

Audio transcription

The Live API transcribes both sides of the conversation for you, so you can show captions, log conversations, and support accessibility without a separate speech-to-text service. Transcription is on by default in ADK for both input (user speech) and output (model speech). Set a field to None to turn that direction off.

from google.genai import types
from google.adk.agents.run_config import RunConfig

# On by default. This is equivalent to setting both to AudioTranscriptionConfig().
run_config = RunConfig(response_modalities=["AUDIO"])

# Turn off user-input transcription, keep model-output transcription.
run_config = RunConfig(
    response_modalities=["AUDIO"],
    input_audio_transcription=None,
)

Transcriptions arrive as types.Transcription objects on event.input_transcription and event.output_transcription, separate from event.content. They stream in fragments: .text holds the latest fragment and .finished marks the last one for the turn. Concatenate the fragments to build the full transcript.

async for event in runner.run_live(...):
    if event.input_transcription and event.input_transcription.text:
        update_caption(
            event.input_transcription.text,
            is_user=True,
            is_final=event.input_transcription.finished,
        )
    if event.output_transcription and event.output_transcription.text:
        update_caption(
            event.output_transcription.text,
            is_user=False,
            is_final=event.output_transcription.finished,
        )

For the event structure, see Transcription events.

!!! note "Multi-agent sessions always transcribe"

When the root agent has `sub_agents`, `run_live()` enables both input and output
transcription even if you set them to `None`. Agent transfer needs the text transcript
to pass conversation context to the next agent, so it cannot be disabled
([`runners.py`](https://github.com/google/adk-python/blob/main/src/google/adk/runners.py)).

Voice and language

Set speech_config to choose the model's voice and language. You can set it in two places:

  • On the agent, by passing a Gemini instance with a speech_config. Use this to give each agent in a multi-agent workflow its own voice.
  • On the session, by setting RunConfig.speech_config. Use this for one voice across the whole session.

When both are set, the agent-level voice wins. With neither set, the Live API picks a default voice.

from google.genai import types
from google.adk.agents import Agent
from google.adk.models.google_llm import Gemini
from google.adk.agents.run_config import RunConfig

# Agent-level voice (wins over RunConfig).
agent = Agent(
    model=Gemini(
        model="gemini-live-2.5-flash-native-audio",
        speech_config=types.SpeechConfig(
            voice_config=types.VoiceConfig(
                prebuilt_voice_config=types.PrebuiltVoiceConfig(voice_name="Puck")
            ),
            language_code="en-US",
        ),
    ),
    instruction="You are a helpful assistant.",
)

# Session-level default voice, used by any agent without its own.
run_config = RunConfig(
    response_modalities=["AUDIO"],
    speech_config=types.SpeechConfig(
        voice_config=types.VoiceConfig(
            prebuilt_voice_config=types.PrebuiltVoiceConfig(voice_name="Kore")
        ),
    ),
)

voice_name selects a prebuilt voice. Live models support eight (Puck, Charon, Kore, Fenrir, Aoede, Leda, Orus, Zephyr) plus the extended Text-to-Speech voice list. For the current list and per-backend availability, see the Gemini Live API voice documentation. An unsupported voice returns an error at connection time.

language_code (for example en-US, ja-JP) sets the language and accent. Live models often infer the language from the conversation and may ignore it.

Voice activity detection (VAD)

VAD detects when the user starts and stops speaking so the model can take turns naturally, including handling interruptions. It is on by default on all Live models, and most applications need no configuration.

Disable automatic VAD when your application decides turn boundaries itself: push-to-talk, client-side VAD, or any UX where the user signals when they are done. When you disable it, you must send manual ActivityStart/ActivityEnd signals with send_activity_start() / send_activity_end(), and your client must translate its own turn signals into those calls on the server.

from google.genai import types
from google.adk.agents.run_config import RunConfig

run_config = RunConfig(
    response_modalities=["AUDIO"],
    realtime_input_config=types.RealtimeInputConfig(
        automatic_activity_detection=types.AutomaticActivityDetection(disabled=True)
    ),
)

A client that runs its own VAD sends those signals to your server, which forwards them with send_activity_start() / send_activity_end(). See Connect a client.

Proactivity and affective dialog

Some Live models offer two conversational features, both off by default:

  • Proactive audio (proactivity) lets the model decide when to respond, offer suggestions unprompted, or ignore irrelevant input.
  • Affective dialog (enable_affective_dialog) lets the model detect emotion in the user's tone and adapt its response.
from google.genai import types
from google.adk.agents.run_config import RunConfig

run_config = RunConfig(
    response_modalities=["AUDIO"],
    proactivity=types.ProactivityConfig(proactive_audio=True),
    enable_affective_dialog=True,
)

Both behaviors are probabilistic and make responses less predictable, so leave them off for formal or high-precision contexts and while debugging.

These settings apply to gemini-live-2.5-flash-native-audio. Some Live models build the behavior in and ignore both settings, so you do not need to set them. See Supported models.