Files
google__adk-docs/docs/live/sessions.md
Kaz Sato 03eccf55e0 docs(live): decompose the dev guide and fix staleness vs adk-python main (#2086)
* docs(live): decompose the development guide into capability pages

Split dev-guide/part1-5 into Sessions, Events, Tools, Workflows, Audio and
video, Configuration, Voice, Supported models, and Build a custom server.
Rewrite index.md as the section Overview with a streaming-type decision table.

Implements Phase 2 of the Live Interactions<>ADK documentation revamp.

* docs(live): drop half-cascade model coverage

Half-cascade models are no longer supported for live agents. Remove the
Native Audio vs Half-Cascade architecture framing from Supported models and
the half-cascade caveats from Voice configuration. The eight prebuilt Live
API voices are kept, relabeled as native-audio voices alongside the extended
Text-to-Speech list.

* docs(live): retire the five-part dev guide and rewire navigation

Delete live/dev-guide/ and live/streaming-tools.md now that their content
lives in the capability pages. Regroup the Live nav into Get started / Build /
Ship / Reference, repoint every partN.md cross-link at its new page and
anchor, and add direct redirects for the removed paths (mkdocs-redirects does
not chain, so streaming/* keys point at final destinations).

* docs(live): point at the API reference instead of pinned source

Swap the RunConfig, Event, SequentialAgent, LiveRequestQueue and
Runner.run_live source-reference notes for Python API reference links.
Implementation pointers with line ranges are left as source links, since they
document internals with no public reference equivalent.

* docs(live): fix docs against adk-python main and drop the bidi-demo links

The bidi-demo sample was removed from adk-samples, so all the source links in
docs/live/ were dead. The sample is not shipped here either, so remove every
reference to it instead of repointing the links.

The code snippets themselves are unchanged. What goes away is only the
scaffolding that pointed at the sample:

- 32 code fences lose their linked 'Demo implementation: file.py:NN-MM' title
  and become plain language-tagged fences.
- The 'Complete Demo Implementation' note in custom-server.md and the 'Demo
  Implementation' note in events.md are dropped; both existed only to link out.
- The 'Learn More' note in tools.md and the model setup step in models.md keep
  their guidance but no longer cite the sample's files.
- Prose that named the demo ('The bidi-demo demonstrates how to...') is
  rewritten to describe the pattern directly.
- The Bidi Demo card and its screenshot are removed from the Live demos section
  of index.md; LensMosaic remains.

Staleness fixes verified against adk-python main:

- StreamingMode.BIDI is inert. Only run_async() reads RunConfig.streaming_mode;
  run_live() never does. Remove it from every run_live()-facing sample and
  rewrite the 'StreamingMode: BIDI or SSE' section around the Runner method you
  call. Keeps the old anchor via attr_list.
- configuration.md: run_live(session=...) is gone; use user_id/session_id.
- tools.md: streaming tools are registered lazily on first model call, not
  scanned up front; the input_stream queue is created only for tools annotated
  with LiveRequestQueue, and stop_streaming resets it to None. The old
  runners.py / function_tool.py line references pointed at unrelated code.
- sessions.md: document DEFAULT_MAX_RECONNECT_ATTEMPTS = 5 and the go_away
  reconnect trigger; correct 'automatic closure in SSE mode', which really only
  happens for the internal queue under support_cfc.
- events.md: audio artifacts require RunConfig.save_live_blob=True;
  get_author_for_event() also keys off llm_response.input_transcription.
- configuration.md: document history_config and the
  initial_history_in_client_content=True that ADK sets when seeding history.

Not changed: get-started/streaming-java.md still sets StreamingMode.BIDI, which
could not be verified without an adk-java checkout.

* Refresh the Live API supported-model list

Checked against the Gemini Live API and Agent Platform model docs:

- models.md: replace the model list with a platform/model/stage table covering
  gemini-3.1-flash-live-preview (Preview, Gemini Live API only),
  gemini-2.5-flash-native-audio-preview-12-2025 (Preview), and
  gemini-live-2.5-flash-native-audio (now GA, not "public preview").
- Document what Gemini 3.1 Live does not support: proactivity, affective
  dialog, async function calling, thinking_budget (it uses thinking_level),
  plus multi-part server events and the turn-coverage default change.
- Note that no Gemini 3.x Live model exists on Agent Platform, and that Live
  API models are unavailable in the `global` location.
- voice.md: replace the Platform Compatibility text, which wrongly said
  proactivity and affective dialog are unavailable on Agent Platform, with a
  per-model support table.
- configuration.md: CFC's model check is a literal `gemini-2` prefix match, so
  it rejects Gemini 3.x; refresh the runners.py line anchor.
- bidi-demo: same model table in the README, the 3.1 option and the regional
  location requirement in .env.example, and an expanded model comment in
  agent.py. The default stays on 2.5 native audio because the demo exposes
  proactivity and affective dialog toggles. Re-anchored the agent.py line
  links in models.md, tools.md, and sessions.md.

* docs(live): align docs with current Live API model capabilities

Verified docs/live/ and docs/runtime/runconfig.md against the Gemini Live
API capabilities guide, the Agent Platform Live API docs, and ADK 2.6.3.

Model consistency:

- response_modalities=["TEXT"] was presented as a valid live configuration
  in configuration.md, events.md and sessions.md. Every Live API model ADK
  supports is a native audio model, and those accept AUDIO only. Reframed
  around AUDIO plus output audio transcription, and kept TEXT where it is
  actually correct: the run_async() / SSE path.
- docs/runtime/runconfig.md configured response_modalities=["AUDIO","TEXT"]
  in all three language samples. A session accepts exactly one modality.
- events.md snippets read event.content.parts[0], which drops content on
  gemini-3.1-flash-live-preview because it sends multiple parts per server
  event -- the failure models.md already warns about. All four snippets now
  iterate over parts.
- tools.md gave the streaming-tools root agent model="gemini-flash-latest",
  which has no Live API support, so the example could not run under
  run_live() on either platform. That alias is still used for the one-shot
  generate_content call inside the tool, where it is correct.
- configuration.md "Standard Gemini Models (1.5 Series) Accessed via SSE"
  described a retired model family and labelled gemini-pro-latest /
  gemini-flash-latest as 1.5 with 2M context.
- sessions.md: document that send_client_content is seeding-only on Gemini
  3.x Live, and that ADK reroutes single-part text to send_realtime_input.
- models.md: gemini-live-2.5-flash-native-audio is the only GA Live API
  model on Agent Platform, not the only one.

Coverage and links:

- configuration.md: document explicit_vad_signal, translation_config,
  avatar_config and model_input_context.
- voice.md: note that ADK picks the live API version (v1alpha / v1beta1),
  so proactivity and affective dialog need no http_options.
- Replace redirecting upstream URLs with their current targets:
  live-guide -> live-api/capabilities, live-session ->
  live-api/session-management, live -> live-api, and
  cloud.google.com/vertex-ai -> the Agent Platform equivalents.

Verified correct, left alone: session and context limits, audio and video
specs, the proactivity / affective dialog model matrix, thinking_level vs
thinking_budget, the support_cfc gemini-2 prefix check, and ADK's AUDIO
default in run_live().

* docs(live): trim the response-modality and SSE material

Every Live API model ADK supports is a native audio model, so a live
session's response modality is always AUDIO and there is nothing to
choose. Shrink the section to the one thing that still matters --
reading text off event.output_transcription.

StreamingMode is only read by run_async(); the SSE tutorial that grew
around it here (protocol diagrams, progressive-streaming walkthrough,
mode-selection table, 1.5-series model list) duplicates
runtime/runconfig.md and describes models that no longer exist. Keep
the inert-BIDI warning and the run_live()/run_async() split, drop the
rest.

Document explicit_vad_signal, translation_config, avatar_config and
model_input_context, which had no coverage at all.

* docs(live): cut duplicated and non-ADK material

Six sections carried weight that did not belong to them:

- sessions.md 'Best Practices for Live API Connection and Session
  Management' restated the Session Resumption and Context Window
  Compression sections verbatim, down to the RunConfig snippets.
  Deleted.
- sessions.md 'Concurrency and Thread Safety' + 'Message Ordering
  Guarantees' explained asyncio.Queue at length and reproduced the
  upstream task already in custom-server.md. Condensed to the three
  properties that actually affect calling code, with a pointer to
  the private _queue attribute dropped.
- sessions.md 'Architectural Patterns for Managing Quotas' was an
  ASCII decision tree and a comparison table for two patterns that
  reduce to one sentence each.
- index.md 'Real-world applications' spent five industry vignettes
  making one point.
- events.md 'Deserializing on the Client' pasted 80 lines of the
  bidi-demo's UI code, calling helpers that no longer exist anywhere
  in these docs. Reduced to the event-shape handling it was meant to
  show.
- audio-video.md 'Handling Image Input at the Client' was 130 lines
  of getUserMedia/canvas/FileReader boilerplate plus a seven-point
  recap of it.

Also fix two dead absolute links: /agents/multi-agents/#workflow-agents-as-orchestrators
(the page now redirects to workflows/index.md and the anchor is gone)
and /live/streaming-tools/ (no such page; the content is in tools.md).

* docs(live): restructure the live docs around ADK ownership

The live section had accumulated content it did not own: backend limits
restated on capability pages, Web Audio API implementation presented as
ADK guidance, and shared concepts re-explained rather than linked.

Applies one rule throughout: if a fact would still be true with the ADK
source deleted, it belongs on models.md or behind an upstream link, not
on a capability page.

- audio-video.md is now the format contract only (505 -> 121). The
  browser mic-capture, ring-buffer playback, and camera-frame code was
  Web Audio API with no ADK in it, had no counterpart in adk-python, and
  no test anywhere. Deleted rather than relocated. The twelve numbered
  'Key Implementation Details' lists restated the code comments directly
  above them; deleted. The streaming-tool lifecycle section duplicated
  tools.md; replaced with a link.
- custom-server.md gains 'Connect a client': what adk web handles
  (16 kHz capture, 24 kHz playback, 1 fps JPEG, transcripts, barge-in),
  where it stops, and the /run_live wire protocol, which was previously
  undocumented. Keeps the one JS snippet that shows ADK's event shape.
  Drops 'Client-side patterns'.
- sessions.md hands its platform-limits table and quota numbers to
  models.md, keeping the session-pool design guidance. The same figures
  had been stated in three places across two pages.
- models.md gains 'Platform limits and quotas' as the single source, and
  loses the 'Key characteristics' list that restated configuration.md.
- configuration.md drops the 'Platform Support' column, which read
  'Both' on 13 of 15 rows and labelled the two exceptions as platform
  constraints when they are model constraints.
- tools.md compresses 'Tool execution context' to the one fact that is
  live-specific: an InvocationContext spans the whole run_live() loop,
  not a single turn.
- workflows.md points at graphs/index.md, the ADK 2.0 graph workflow
  page, rather than the v0.1.0 multi-agent umbrella.
- Six internal links used absolute paths, which mkdocs does not
  validate, so --strict had been silently ignoring them. Now relative.
- Fixes class.="grid cards" in get-started/index.md, which was breaking
  the card grid.

* docs(live): standardize page leads and cut duplicated RunConfig prose

Every live page opened by narrating its own table of contents ("This page
covers X, Y, and Z"), which duplicates the rendered TOC, ages badly when
a heading changes, and spends a paragraph before the reader gets a fact.
evaluation.md already did the better thing: state the shared baseline,
link the canonical page, then cover only the delta. That is now the
convention across the section.

- sessions.md, events.md, configuration.md, audio-video.md,
  workflows.md, tools.md, models.md and get-started/index.md now name
  their non-live counterpart in the lead instead of listing their own
  headings. Three pages had no outbound link to the shared concept at
  all: tools.md to Custom Tools, models.md to Models for agents, and
  workflows.md pointed at the v0.1.0 umbrella rather than graph
  workflows.
- configuration.md drops the custom_metadata section (85 lines) for a
  pointer plus the one live-specific consequence: a run_live() call is a
  single invocation, so metadata is stamped on the whole session rather
  than one turn. runtime/runconfig.md already owns the field.
- configuration.md trims max_llm_calls and save_live_blob to the facts
  that are live-specific — max_llm_calls does not apply to run_live() at
  all, and save_live_blob writes ~1.92 MB per minute per session to two
  services — and drops the generic use-case and best-practice lists.
- custom-server.md replaces 'Key concepts', which re-pasted all three
  code blocks from the complete example directly above it, with prose
  explaining why the two tasks must run concurrently.

Live section: 2820 -> 2211 lines.

* docs(live): reframe pages around capabilities, fix eval config key

* Apply batched suggestions from code review

Co-authored-by: Joe Fernandez <931947+joefernandez@users.noreply.github.com>

* Apply suggestion from @joefernandez

* Apply batched suggestions from code review

Co-authored-by: Joe Fernandez <931947+joefernandez@users.noreply.github.com>

---------

Co-authored-by: Stephen Allen <stephenaallen@google.com>
Co-authored-by: Joe Fernandez <931947+joefernandez@users.noreply.github.com>
2026-09-01 17:15:50 -07:00

15 KiB

Sessions for live agents

Supported in ADKPython v0.1.0

A live agent is a connection that stays open while the user talks, listens, interrupts, and falls silent.

Live agents use the same Session, SessionService, and state model as any ADK agent, all covered in Conversational context. What a live session adds is a connection: one that can drop, time out, or outlive the model's context window. For what comes back out of that connection, see Events; for the settings that shape it, see Configuration.

Set up a live application

A live application has two kinds of objects: ones you create once at startup and reuse for every session, and ones you create fresh per session.

Create once, reuse everywhere:

  • Agent: your model, tools, and instructions. Stateless and reusable.
  • SessionService: stores conversation history so sessions survive reconnects and restarts.
  • Runner: the runtime that drives the agent and yields events.
import os
from google.adk.agents import Agent
from google.adk.runners import Runner
from google.adk.sessions import InMemorySessionService
from google.adk.tools import google_search

APP_NAME = "live-agent"

agent = Agent(
    name="google_search_agent",
    model=os.getenv("DEMO_AGENT_MODEL", "gemini-live-2.5-flash-native-audio"),
    tools=[google_search],
    instruction="You are a helpful assistant that can search the web.",
)

runner = Runner(
    app_name=APP_NAME,
    agent=agent,
    session_service=InMemorySessionService(),
)

InMemorySessionService loses state when the process stops. For production, use DatabaseSessionService (SQLite, PostgreSQL, or MySQL) or VertexAiSessionService (managed on Google Cloud). See Session services.

Create per session:

  • A Session, fetched or created before the loop runs.
  • A RunConfig, which can differ per user (voice, transcription, limits).
  • A LiveRequestQueue, the channel you send user input through.
from google.adk.agents.live_request_queue import LiveRequestQueue
from google.adk.agents.run_config import RunConfig
from google.genai import types

# Get-or-create handles both new conversations and reconnections.
session = await session_service.get_session(
    app_name=APP_NAME, user_id=user_id, session_id=session_id
)
if not session:
    await session_service.create_session(
        app_name=APP_NAME, user_id=user_id, session_id=session_id
    )

run_config = RunConfig(
    response_modalities=["AUDIO"],
    session_resumption=types.SessionResumptionConfig(),
)

live_request_queue = LiveRequestQueue()

user_id and session_id are arbitrary strings you define; ADK generates a UUID if you pass session_id=None. The session must exist before you call run_live() with the same identifiers, or run_live() raises ValueError: Session not found.

!!! warning "One queue per session"

Never reuse a `LiveRequestQueue` across sessions. The close signal persists in the queue
and would carry over, corrupting the next session. Create a fresh queue for every
`run_live()` call.

LiveRequestQueue

LiveRequestQueue is your channel for sending messages to the agent. Every message is a LiveRequest, a single container for the different kinds of input:

class LiveRequest(BaseModel):
    content: Optional[Content] = None            # Text and structured data
    blob: Optional[Blob] = None                  # Audio/video bytes
    activity_start: Optional[ActivityStart] = None  # Manual turn start
    activity_end: Optional[ActivityEnd] = None      # Manual turn end
    close: bool = False                          # Graceful termination

content and blob are mutually exclusive. Use the convenience methods rather than building LiveRequest objects yourself; they set the right field and keep you within that constraint.

Method Sends Mode
send_content(content) Text, as a discrete turn Turn-by-turn; triggers a response
send_realtime(blob) Audio, image, or video bytes Continuous streaming
send_activity_start() / send_activity_end() Manual turn boundaries Only when automatic VAD is disabled
close() Termination signal Ends the session
from google.genai import types

# Text turn.
live_request_queue.send_content(types.Content(parts=[types.Part(text=user_text)]))

# Audio chunk (streamed continuously).
live_request_queue.send_realtime(
    types.Blob(mime_type="audio/pcm;rate=16000", data=audio_data)
)

For audio, image, and video formats, see Audio and video. For manual turn control with activity signals, see Voice activity detection.

!!! note "Send one text Part per call"

Send a single text `Part` per `send_content()` call. Some Live models treat a multi-part
`Content` as conversation seeding (priming history) rather than a turn to respond to, so
one Part per call keeps behavior consistent across models.

Concurrency and ordering

LiveRequestQueue wraps an asyncio.Queue, which has three consequences:

  • Send methods are synchronous. They call put_nowait() underneath, so they never block and never need await.
  • Delivery is FIFO and uncoalesced. Requests reach the model in send order, one per call.
  • The queue is unbounded. Sending faster than the model consumes grows memory rather than applying backpressure, so cap your own send rate for high-rate audio or video.

Create the queue inside an async context so it binds to the event loop that runs run_live(). asyncio.Queue is safe for concurrent access within a single event loop thread; to feed it from another thread, use loop.call_soon_threadsafe().

The run_live() loop

run_live() is an async generator. It yields Event objects the moment they are generated, with no buffering or polling, while you send new input concurrently through the queue. That concurrency is what makes interruption work: the agent can be speaking while the user starts talking over it.

async for event in runner.run_live(
    user_id=user_id,
    session_id=session_id,
    live_request_queue=live_request_queue,
    run_config=run_config,
):
    await websocket.send_text(event.model_dump_json(exclude_none=True, by_alias=True))

run_live() opens the Live API connection when you call it, streams both directions while the loop runs, and closes the connection when you call live_request_queue.close(). For the event types it yields and how to handle them, see Events.

When run_live() exits

Exit condition Trigger Graceful
Manual close live_request_queue.close() Yes
Workflow complete Last agent in a live workflow calls task_completed() Yes
Session timeout Live API duration limit reached (without compression) Connection closed
Early exit end_invocation set by a tool or callback Yes
Error Connection failure or unhandled exception No

Always call close() when the session ends, even on error. Skipping it leaves the Live API without a graceful termination signal, which can strand "zombie" sessions that count against your concurrent-session quota until they time out.

try:
    await asyncio.gather(upstream_task(), downstream_task())
except WebSocketDisconnect:
    pass  # Client disconnected normally.
finally:
    live_request_queue.close()  # Always close the queue.

For error handling inside the loop, see Error events. For the full upstream/downstream server pattern, see Custom server.

What gets saved to the session

When run_live() exits, only some events persist to the ADK Session:

  • Saved: final (non-partial) transcriptions, usage metadata, function calls and responses, and most control events. Audio files are saved only when save_live_blob is True.
  • Ephemeral: raw audio bytes (inline_data) and partial transcriptions, yielded for real-time playback and display but not stored.

ADK Session vs Live API session

Two different things share the word "session":

  • ADK Session (managed by SessionService) is persistent conversation storage. It survives across many run_live() calls and application restarts.
  • Live API session (managed by the Live API backend) is a transient streaming context that exists only while the loop runs.

When run_live() starts, ADK loads history from the ADK Session, initializes a new Live API session with it, and updates the ADK Session as events occur. When the loop ends, the Live API session is destroyed and the ADK Session persists. The next call rebuilds a Live API session from the stored history. This separation is what lets conversations continue across network drops and restarts.

At the transport layer, one more distinction matters for reliability:

  • A connection is the WebSocket link between ADK and the Live API. It can time out.
  • A session is the conversation context, which can span multiple connections through session resumption.

Platform limits

Both backends cap connection duration, session duration, and concurrent sessions. The exact numbers differ by backend and change over time, so Supported models tracks them in one place.

Two of those caps change how you write the code. Context window compression lifts the session-duration limit, and the concurrent-session ceiling is what you design against in Concurrent sessions.

Session resumption

The Live API closes each WebSocket connection after about 10 minutes. Session resumption migrates the conversation across connections so it continues past that limit. Enable it and ADK handles all reconnection for you, caching resumption handles, detecting closures, and reconnecting in the background. Your run_live() loop keeps yielding events without interruption.

from google.genai import types

run_config = RunConfig(session_resumption=types.SessionResumptionConfig())

ADK manages the ADK-to-Live-API connection only. Your application still owns its own client connections (for example, the user's WebSocket to your server) and any client-side reconnect logic.

How ADK reconnects:

  1. The Live API sends session_resumption_update messages; ADK caches the latest handle.
  2. Before the limit, the Live API may send a go_away warning; ADK reconnects before the drop, so the handover is invisible.
  3. When a connection closes gracefully, ADK's loop reconnects with the cached handle and the session continues with full context.
sequenceDiagram
    participant App as Your Application
    participant ADK as ADK (run_live)
    participant API as Live API

    App->>ADK: run_live(run_config with session_resumption)
    ADK->>API: WebSocket connect()
    Note over ADK,API: Streaming (0-10 min)
    API-->>ADK: session_resumption_update { handle }
    ADK->>ADK: Cache handle
    Note over API: ~10 min: connection closes gracefully
    ADK->>API: reconnect(handle)
    API-->>ADK: Session resumed with full context
    Note over App,API: Loop continues, uninterrupted

!!! warning "Reconnection attempts are capped"

ADK retries a maximum of **5 consecutive** reconnections
([`DEFAULT_MAX_RECONNECT_ATTEMPTS`](https://github.com/google/adk-python/blob/main/src/google/adk/flows/llm_flows/base_llm_flow.py)).
The counter resets on each successful reconnect, so a long conversation is limited only to
five *failures in a row*, not five reconnects total. ADK retries only when a resumption
handle exists; without `session_resumption` enabled, the first drop propagates straight
out of `run_live()`, and your application must handle it.

Skip resumption only for short sessions (under 10 minutes), stateless request-response interactions, or development where a fresh session per run aids debugging.

Context window compression

Long conversations hit two limits: the session duration caps, and the model's context window (varies by model). Context window compression addresses both. It compresses older conversation history with a sliding window when the token count crosses a threshold, keeping recent turns in full. Enabling it removes the session duration limits. The trade-off: older context becomes a summary, not verbatim history.

from google.genai import types
from google.adk.agents.run_config import RunConfig

# For a 128k-context model.
run_config = RunConfig(
    context_window_compression=types.ContextWindowCompressionConfig(
        trigger_tokens=100000,  # Start compressing near ~78% of the window.
        sliding_window=types.SlidingWindow(
            target_tokens=80000,  # Compress down to ~62%, keeping recent turns.
        ),
    )
)

Set trigger_tokens to roughly 70-80% of the model's context window for headroom, and target_tokens to 60-70% so each compression frees enough room for several turns. Test with your own conversation patterns. Enable compression when sessions must run longer than the platform limits or may exceed the token limit; leave it off for short sessions or when precise recall of early turns is critical.

Concurrent sessions

Each user needs their own Live API session, and both backends cap concurrent sessions. Your concurrent-session ceiling is a hard cap on simultaneous users. For the current ceilings and how to request increases, see Supported models.

Design for the ceiling:

  • One session per user is the default and correct choice while peak concurrency fits inside the quota.
  • A session pool (a fixed set of sessions handed out through a queue) keeps you inside the quota when peak concurrency exceeds it, at the cost of wait time. Reset per-session state on release so conversations do not leak between users.

Either way, count active sessions yourself and queue or reject new connections before the platform does. A quota rejection surfaces as a connection failure, a worse experience than a visible queue position.