* docs(live): decompose the development guide into capability pages
Split dev-guide/part1-5 into Sessions, Events, Tools, Workflows, Audio and
video, Configuration, Voice, Supported models, and Build a custom server.
Rewrite index.md as the section Overview with a streaming-type decision table.
Implements Phase 2 of the Live Interactions<>ADK documentation revamp.
* docs(live): drop half-cascade model coverage
Half-cascade models are no longer supported for live agents. Remove the
Native Audio vs Half-Cascade architecture framing from Supported models and
the half-cascade caveats from Voice configuration. The eight prebuilt Live
API voices are kept, relabeled as native-audio voices alongside the extended
Text-to-Speech list.
* docs(live): retire the five-part dev guide and rewire navigation
Delete live/dev-guide/ and live/streaming-tools.md now that their content
lives in the capability pages. Regroup the Live nav into Get started / Build /
Ship / Reference, repoint every partN.md cross-link at its new page and
anchor, and add direct redirects for the removed paths (mkdocs-redirects does
not chain, so streaming/* keys point at final destinations).
* docs(live): point at the API reference instead of pinned source
Swap the RunConfig, Event, SequentialAgent, LiveRequestQueue and
Runner.run_live source-reference notes for Python API reference links.
Implementation pointers with line ranges are left as source links, since they
document internals with no public reference equivalent.
* docs(live): fix docs against adk-python main and drop the bidi-demo links
The bidi-demo sample was removed from adk-samples, so all the source links in
docs/live/ were dead. The sample is not shipped here either, so remove every
reference to it instead of repointing the links.
The code snippets themselves are unchanged. What goes away is only the
scaffolding that pointed at the sample:
- 32 code fences lose their linked 'Demo implementation: file.py:NN-MM' title
and become plain language-tagged fences.
- The 'Complete Demo Implementation' note in custom-server.md and the 'Demo
Implementation' note in events.md are dropped; both existed only to link out.
- The 'Learn More' note in tools.md and the model setup step in models.md keep
their guidance but no longer cite the sample's files.
- Prose that named the demo ('The bidi-demo demonstrates how to...') is
rewritten to describe the pattern directly.
- The Bidi Demo card and its screenshot are removed from the Live demos section
of index.md; LensMosaic remains.
Staleness fixes verified against adk-python main:
- StreamingMode.BIDI is inert. Only run_async() reads RunConfig.streaming_mode;
run_live() never does. Remove it from every run_live()-facing sample and
rewrite the 'StreamingMode: BIDI or SSE' section around the Runner method you
call. Keeps the old anchor via attr_list.
- configuration.md: run_live(session=...) is gone; use user_id/session_id.
- tools.md: streaming tools are registered lazily on first model call, not
scanned up front; the input_stream queue is created only for tools annotated
with LiveRequestQueue, and stop_streaming resets it to None. The old
runners.py / function_tool.py line references pointed at unrelated code.
- sessions.md: document DEFAULT_MAX_RECONNECT_ATTEMPTS = 5 and the go_away
reconnect trigger; correct 'automatic closure in SSE mode', which really only
happens for the internal queue under support_cfc.
- events.md: audio artifacts require RunConfig.save_live_blob=True;
get_author_for_event() also keys off llm_response.input_transcription.
- configuration.md: document history_config and the
initial_history_in_client_content=True that ADK sets when seeding history.
Not changed: get-started/streaming-java.md still sets StreamingMode.BIDI, which
could not be verified without an adk-java checkout.
* Refresh the Live API supported-model list
Checked against the Gemini Live API and Agent Platform model docs:
- models.md: replace the model list with a platform/model/stage table covering
gemini-3.1-flash-live-preview (Preview, Gemini Live API only),
gemini-2.5-flash-native-audio-preview-12-2025 (Preview), and
gemini-live-2.5-flash-native-audio (now GA, not "public preview").
- Document what Gemini 3.1 Live does not support: proactivity, affective
dialog, async function calling, thinking_budget (it uses thinking_level),
plus multi-part server events and the turn-coverage default change.
- Note that no Gemini 3.x Live model exists on Agent Platform, and that Live
API models are unavailable in the `global` location.
- voice.md: replace the Platform Compatibility text, which wrongly said
proactivity and affective dialog are unavailable on Agent Platform, with a
per-model support table.
- configuration.md: CFC's model check is a literal `gemini-2` prefix match, so
it rejects Gemini 3.x; refresh the runners.py line anchor.
- bidi-demo: same model table in the README, the 3.1 option and the regional
location requirement in .env.example, and an expanded model comment in
agent.py. The default stays on 2.5 native audio because the demo exposes
proactivity and affective dialog toggles. Re-anchored the agent.py line
links in models.md, tools.md, and sessions.md.
* docs(live): align docs with current Live API model capabilities
Verified docs/live/ and docs/runtime/runconfig.md against the Gemini Live
API capabilities guide, the Agent Platform Live API docs, and ADK 2.6.3.
Model consistency:
- response_modalities=["TEXT"] was presented as a valid live configuration
in configuration.md, events.md and sessions.md. Every Live API model ADK
supports is a native audio model, and those accept AUDIO only. Reframed
around AUDIO plus output audio transcription, and kept TEXT where it is
actually correct: the run_async() / SSE path.
- docs/runtime/runconfig.md configured response_modalities=["AUDIO","TEXT"]
in all three language samples. A session accepts exactly one modality.
- events.md snippets read event.content.parts[0], which drops content on
gemini-3.1-flash-live-preview because it sends multiple parts per server
event -- the failure models.md already warns about. All four snippets now
iterate over parts.
- tools.md gave the streaming-tools root agent model="gemini-flash-latest",
which has no Live API support, so the example could not run under
run_live() on either platform. That alias is still used for the one-shot
generate_content call inside the tool, where it is correct.
- configuration.md "Standard Gemini Models (1.5 Series) Accessed via SSE"
described a retired model family and labelled gemini-pro-latest /
gemini-flash-latest as 1.5 with 2M context.
- sessions.md: document that send_client_content is seeding-only on Gemini
3.x Live, and that ADK reroutes single-part text to send_realtime_input.
- models.md: gemini-live-2.5-flash-native-audio is the only GA Live API
model on Agent Platform, not the only one.
Coverage and links:
- configuration.md: document explicit_vad_signal, translation_config,
avatar_config and model_input_context.
- voice.md: note that ADK picks the live API version (v1alpha / v1beta1),
so proactivity and affective dialog need no http_options.
- Replace redirecting upstream URLs with their current targets:
live-guide -> live-api/capabilities, live-session ->
live-api/session-management, live -> live-api, and
cloud.google.com/vertex-ai -> the Agent Platform equivalents.
Verified correct, left alone: session and context limits, audio and video
specs, the proactivity / affective dialog model matrix, thinking_level vs
thinking_budget, the support_cfc gemini-2 prefix check, and ADK's AUDIO
default in run_live().
* docs(live): trim the response-modality and SSE material
Every Live API model ADK supports is a native audio model, so a live
session's response modality is always AUDIO and there is nothing to
choose. Shrink the section to the one thing that still matters --
reading text off event.output_transcription.
StreamingMode is only read by run_async(); the SSE tutorial that grew
around it here (protocol diagrams, progressive-streaming walkthrough,
mode-selection table, 1.5-series model list) duplicates
runtime/runconfig.md and describes models that no longer exist. Keep
the inert-BIDI warning and the run_live()/run_async() split, drop the
rest.
Document explicit_vad_signal, translation_config, avatar_config and
model_input_context, which had no coverage at all.
* docs(live): cut duplicated and non-ADK material
Six sections carried weight that did not belong to them:
- sessions.md 'Best Practices for Live API Connection and Session
Management' restated the Session Resumption and Context Window
Compression sections verbatim, down to the RunConfig snippets.
Deleted.
- sessions.md 'Concurrency and Thread Safety' + 'Message Ordering
Guarantees' explained asyncio.Queue at length and reproduced the
upstream task already in custom-server.md. Condensed to the three
properties that actually affect calling code, with a pointer to
the private _queue attribute dropped.
- sessions.md 'Architectural Patterns for Managing Quotas' was an
ASCII decision tree and a comparison table for two patterns that
reduce to one sentence each.
- index.md 'Real-world applications' spent five industry vignettes
making one point.
- events.md 'Deserializing on the Client' pasted 80 lines of the
bidi-demo's UI code, calling helpers that no longer exist anywhere
in these docs. Reduced to the event-shape handling it was meant to
show.
- audio-video.md 'Handling Image Input at the Client' was 130 lines
of getUserMedia/canvas/FileReader boilerplate plus a seven-point
recap of it.
Also fix two dead absolute links: /agents/multi-agents/#workflow-agents-as-orchestrators
(the page now redirects to workflows/index.md and the anchor is gone)
and /live/streaming-tools/ (no such page; the content is in tools.md).
* docs(live): restructure the live docs around ADK ownership
The live section had accumulated content it did not own: backend limits
restated on capability pages, Web Audio API implementation presented as
ADK guidance, and shared concepts re-explained rather than linked.
Applies one rule throughout: if a fact would still be true with the ADK
source deleted, it belongs on models.md or behind an upstream link, not
on a capability page.
- audio-video.md is now the format contract only (505 -> 121). The
browser mic-capture, ring-buffer playback, and camera-frame code was
Web Audio API with no ADK in it, had no counterpart in adk-python, and
no test anywhere. Deleted rather than relocated. The twelve numbered
'Key Implementation Details' lists restated the code comments directly
above them; deleted. The streaming-tool lifecycle section duplicated
tools.md; replaced with a link.
- custom-server.md gains 'Connect a client': what adk web handles
(16 kHz capture, 24 kHz playback, 1 fps JPEG, transcripts, barge-in),
where it stops, and the /run_live wire protocol, which was previously
undocumented. Keeps the one JS snippet that shows ADK's event shape.
Drops 'Client-side patterns'.
- sessions.md hands its platform-limits table and quota numbers to
models.md, keeping the session-pool design guidance. The same figures
had been stated in three places across two pages.
- models.md gains 'Platform limits and quotas' as the single source, and
loses the 'Key characteristics' list that restated configuration.md.
- configuration.md drops the 'Platform Support' column, which read
'Both' on 13 of 15 rows and labelled the two exceptions as platform
constraints when they are model constraints.
- tools.md compresses 'Tool execution context' to the one fact that is
live-specific: an InvocationContext spans the whole run_live() loop,
not a single turn.
- workflows.md points at graphs/index.md, the ADK 2.0 graph workflow
page, rather than the v0.1.0 multi-agent umbrella.
- Six internal links used absolute paths, which mkdocs does not
validate, so --strict had been silently ignoring them. Now relative.
- Fixes class.="grid cards" in get-started/index.md, which was breaking
the card grid.
* docs(live): standardize page leads and cut duplicated RunConfig prose
Every live page opened by narrating its own table of contents ("This page
covers X, Y, and Z"), which duplicates the rendered TOC, ages badly when
a heading changes, and spends a paragraph before the reader gets a fact.
evaluation.md already did the better thing: state the shared baseline,
link the canonical page, then cover only the delta. That is now the
convention across the section.
- sessions.md, events.md, configuration.md, audio-video.md,
workflows.md, tools.md, models.md and get-started/index.md now name
their non-live counterpart in the lead instead of listing their own
headings. Three pages had no outbound link to the shared concept at
all: tools.md to Custom Tools, models.md to Models for agents, and
workflows.md pointed at the v0.1.0 umbrella rather than graph
workflows.
- configuration.md drops the custom_metadata section (85 lines) for a
pointer plus the one live-specific consequence: a run_live() call is a
single invocation, so metadata is stamped on the whole session rather
than one turn. runtime/runconfig.md already owns the field.
- configuration.md trims max_llm_calls and save_live_blob to the facts
that are live-specific — max_llm_calls does not apply to run_live() at
all, and save_live_blob writes ~1.92 MB per minute per session to two
services — and drops the generic use-case and best-practice lists.
- custom-server.md replaces 'Key concepts', which re-pasted all three
code blocks from the complete example directly above it, with prose
explaining why the two tasks must run concurrently.
Live section: 2820 -> 2211 lines.
* docs(live): reframe pages around capabilities, fix eval config key
* Apply batched suggestions from code review
Co-authored-by: Joe Fernandez <931947+joefernandez@users.noreply.github.com>
* Apply suggestion from @joefernandez
* Apply batched suggestions from code review
Co-authored-by: Joe Fernandez <931947+joefernandez@users.noreply.github.com>
---------
Co-authored-by: Stephen Allen <stephenaallen@google.com>
Co-authored-by: Joe Fernandez <931947+joefernandez@users.noreply.github.com>
20 KiB
Build live streaming agent with Java
Build a Java agent that holds a low-latency, two-way voice conversation with ADK Streaming.
You will set up Java and Maven, define the project dependencies, and build a
ScienceTeacherAgent. You will test it as text streaming in the Dev UI first, then turn on
live audio and talk to it.
Create your first agent
Prerequisites
-
In this getting started guide, you will be programming in Java. Check if Java is installed on your machine. Ideally, you should be using Java 17 or more (you can check that by typing java -version)
-
You need the Maven build tool for Java, so install Maven before going further (Cloud Top and Cloud Shell already have it; your local development environment may not).
Prepare the project structure
To get started with ADK Java, let’s create a Maven project with the following directory structure:
adk-agents/
├── pom.xml
└── src/
└── main/
└── java/
└── agents/
└── ScienceTeacherAgent.java
Follow the instructions in Installation page to add pom.xml for using the ADK package.
!!! Note Feel free to use whichever name you like for the root directory of your project (instead of adk-agents)
Running a compilation
Let’s see if Maven is happy with this build, by running a compilation (mvn compile command):
$ mvn compile
[INFO] Scanning for projects...
[INFO]
[INFO] --------------------< adk-agents:adk-agents >--------------------
[INFO] Building adk-agents 1.0-SNAPSHOT
[INFO] from pom.xml
[INFO] --------------------------------[ jar ]---------------------------------
[INFO]
[INFO] --- resources:3.3.1:resources (default-resources) @ adk-demo ---
[INFO] skip non existing resourceDirectory /home/user/adk-demo/src/main/resources
[INFO]
[INFO] --- compiler:3.13.0:compile (default-compile) @ adk-demo ---
[INFO] Nothing to compile - all classes are up to date.
[INFO] ------------------------------------------------------------------------
[INFO] BUILD SUCCESS
[INFO] ------------------------------------------------------------------------
[INFO] Total time: 1.347 s
[INFO] Finished at: 2025-05-06T15:38:08Z
[INFO] ------------------------------------------------------------------------
Looks like the project is set up properly for compilation!
Creating an agent
Create the ScienceTeacherAgent.java file under the src/main/java/agents/ directory with the following content:
package samples.liveaudio;
import com.google.adk.agents.BaseAgent;
import com.google.adk.agents.LlmAgent;
/** Science teacher agent. */
public class ScienceTeacherAgent {
// Field expected by the Dev UI to load the agent dynamically
// (the agent must be initialized at declaration time)
public static final BaseAgent ROOT_AGENT = initAgent();
// Please fill in the latest model id that supports live API from
// https://adk.dev/live/get-started/streaming-python/#supported-models
public static BaseAgent initAgent() {
return LlmAgent.builder()
.name("science-app")
.description("Science teacher agent")
.model("...") // Pleaase fill in the latest model id for live API
.instruction("""
You are a helpful science teacher that explains
science concepts to kids and teenagers.
""")
.build();
}
}
We will use Dev UI to run this agent later. For the tool to automatically recognize the agent, its Java class has to comply with the following two rules:
- The agent should be stored in a global public static variable named ROOT_AGENT of type BaseAgent and initialized at declaration time.
- The agent definition has to be a static method so it can be loaded during the class initialization by the dynamic compiling classloader.
Run agent with Dev UI
Dev UI is a web server where you can quickly run and test your agents for development purpose, without building your own UI application for the agents.
Define environment variables
To run the server, you’ll need to export two environment variables:
- a Gemini key that you can get from AI Studio,
- a variable to specify we’re not using Agent Platform this time.
export GOOGLE_GENAI_USE_ENTERPRISE=FALSE
export GOOGLE_API_KEY=YOUR_API_KEY
Run Dev UI
Run the following command from the terminal to launch the Dev UI.
mvn exec:java \
-Dexec.mainClass="com.google.adk.web.AdkWebServer" \
-Dexec.args="--adk.agents.source-dir=." \
-Dexec.classpathScope="compile"
Step 1: Open the URL provided (usually http://localhost:8080 or
http://127.0.0.1:8080) directly in your browser.
Step 2. In the top-left corner of the UI, you can select your agent in the dropdown. Select "science-app".
!!!note "Troubleshooting"
If you do not see "science-app" in the dropdown menu, make sure you
are running the `mvn` command from the root of your maven project.
!!! warning "Caution: ADK Web for development only"
ADK Web is ***not meant for use in production deployments***. You should
use ADK Web for development and debugging purposes only.
Try Dev UI with voice and video
With your favorite browser, navigate to: http://127.0.0.1:8080/
You should see the following interface:
Click the microphone button to enable the voice input, and ask a question What's the electron? in voice. You will hear the answer in voice in real-time.
To try with video, reload the web browser, click the camera button to enable the video input, and ask questions like "What do you see?". The agent will answer what they see in the video input.
Caveat
- You can not use text chat with the native-audio models. You will see errors when entering text messages on
adk web.
Stop the tool
Stop the tool by pressing Ctrl-C on the console.
Run agent with a custom live audio app
Now, let's try audio streaming with the agent and a custom live audio application.
A Maven pom.xml build file for Live Audio
Replace your existing pom.xml with the following.
<?xml version="1.0" encoding="UTF-8"?>
<project xmlns="http://maven.apache.org/POM/4.0.0"
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance"
xsi:schemaLocation="http://maven.apache.org/POM/4.0.0 http://maven.apache.org/xsd/maven-4.0.0.xsd">
<modelVersion>4.0.0</modelVersion>
<groupId>com.google.adk.samples</groupId>
<artifactId>google-adk-sample-live-audio</artifactId>
<version>0.1.0</version>
<name>Google ADK - Sample - Live Audio</name>
<description>
A sample application demonstrating a live audio conversation using ADK,
runnable via samples.liveaudio.LiveAudioRun.
</description>
<packaging>jar</packaging>
<properties>
<project.build.sourceEncoding>UTF-8</project.build.sourceEncoding>
<java.version>17</java.version>
<auto-value.version>1.11.0</auto-value.version>
<!-- Main class for exec-maven-plugin -->
<exec.mainClass>samples.liveaudio.LiveAudioRun</exec.mainClass>
<google-adk.version>1.6.0</google-adk.version>
</properties>
<dependencyManagement>
<dependencies>
<dependency>
<groupId>com.google.cloud</groupId>
<artifactId>libraries-bom</artifactId>
<version>26.53.0</version>
<type>pom</type>
<scope>import</scope>
</dependency>
</dependencies>
</dependencyManagement>
<dependencies>
<dependency>
<groupId>com.google.adk</groupId>
<artifactId>google-adk</artifactId>
<version>${google-adk.version}</version>
</dependency>
<dependency>
<groupId>commons-logging</groupId>
<artifactId>commons-logging</artifactId>
<version>1.2</version> <!-- Or use a property if defined in a parent POM -->
</dependency>
</dependencies>
<build>
<plugins>
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-compiler-plugin</artifactId>
<version>3.13.0</version>
<configuration>
<source>${java.version}</source>
<target>${java.version}</target>
<parameters>true</parameters>
<annotationProcessorPaths>
<path>
<groupId>com.google.auto.value</groupId>
<artifactId>auto-value</artifactId>
<version>${auto-value.version}</version>
</path>
</annotationProcessorPaths>
</configuration>
</plugin>
<plugin>
<groupId>org.codehaus.mojo</groupId>
<artifactId>build-helper-maven-plugin</artifactId>
<version>3.6.0</version>
<executions>
<execution>
<id>add-source</id>
<phase>generate-sources</phase>
<goals>
<goal>add-source</goal>
</goals>
<configuration>
<sources>
<source>.</source>
</sources>
</configuration>
</execution>
</executions>
</plugin>
<plugin>
<groupId>org.codehaus.mojo</groupId>
<artifactId>exec-maven-plugin</artifactId>
<version>3.2.0</version>
<configuration>
<mainClass>${exec.mainClass}</mainClass>
<classpathScope>runtime</classpathScope>
</configuration>
</plugin>
</plugins>
</build>
</project>
Creating Live Audio Run tool
Create the LiveAudioRun.java file under the src/main/java/ directory with the following content. This tool runs the agent on it with live audio input and output.
package samples.liveaudio;
import com.google.adk.agents.LiveRequestQueue;
import com.google.adk.agents.RunConfig;
import com.google.adk.events.Event;
import com.google.adk.runner.Runner;
import com.google.adk.sessions.InMemorySessionService;
import com.google.common.collect.ImmutableList;
import com.google.genai.types.Blob;
import com.google.genai.types.Modality;
import com.google.genai.types.PrebuiltVoiceConfig;
import com.google.genai.types.Content;
import com.google.genai.types.Part;
import com.google.genai.types.SpeechConfig;
import com.google.genai.types.VoiceConfig;
import io.reactivex.rxjava3.core.Flowable;
import java.io.ByteArrayOutputStream;
import java.io.InputStream;
import java.net.URL;
import javax.sound.sampled.AudioFormat;
import javax.sound.sampled.AudioInputStream;
import javax.sound.sampled.AudioSystem;
import javax.sound.sampled.DataLine;
import javax.sound.sampled.LineUnavailableException;
import javax.sound.sampled.Mixer;
import javax.sound.sampled.SourceDataLine;
import javax.sound.sampled.TargetDataLine;
import java.util.UUID;
import java.util.concurrent.ExecutorService;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ConcurrentMap;
import java.util.concurrent.Executors;
import java.util.concurrent.Future;
import java.util.concurrent.TimeUnit;
import java.util.concurrent.atomic.AtomicBoolean;
import agents.ScienceTeacherAgent;
/** Main class to demonstrate running the {@link LiveAudioAgent} for a voice conversation. */
public final class LiveAudioRun {
private final String userId;
private final String sessionId;
private final Runner runner;
private static final javax.sound.sampled.AudioFormat MIC_AUDIO_FORMAT =
new javax.sound.sampled.AudioFormat(16000.0f, 16, 1, true, false);
private static final javax.sound.sampled.AudioFormat SPEAKER_AUDIO_FORMAT =
new javax.sound.sampled.AudioFormat(24000.0f, 16, 1, true, false);
private static final int BUFFER_SIZE = 4096;
public LiveAudioRun() {
this.userId = "test_user";
String appName = "LiveAudioApp";
this.sessionId = UUID.randomUUID().toString();
InMemorySessionService sessionService = new InMemorySessionService();
this.runner = new Runner(ScienceTeacherAgent.ROOT_AGENT, appName, null, sessionService);
ConcurrentMap<String, Object> initialState = new ConcurrentHashMap<>();
var unused =
sessionService.createSession(appName, userId, initialState, sessionId).blockingGet();
}
private void runConversation() throws Exception {
System.out.println("Initializing microphone input and speaker output...");
RunConfig runConfig =
RunConfig.builder()
.setStreamingMode(RunConfig.StreamingMode.BIDI)
.setResponseModalities(ImmutableList.of(new Modality("AUDIO")))
.setSpeechConfig(
SpeechConfig.builder()
.voiceConfig(
VoiceConfig.builder()
.prebuiltVoiceConfig(
PrebuiltVoiceConfig.builder().voiceName("Aoede").build())
.build())
.languageCode("en-US")
.build())
.build();
LiveRequestQueue liveRequestQueue = new LiveRequestQueue();
Flowable<Event> eventStream =
this.runner.runLive(
runner.sessionService().createSession(userId, sessionId).blockingGet(),
liveRequestQueue,
runConfig);
AtomicBoolean isRunning = new AtomicBoolean(true);
AtomicBoolean conversationEnded = new AtomicBoolean(false);
ExecutorService executorService = Executors.newFixedThreadPool(2);
// Task for capturing microphone input
Future<?> microphoneTask =
executorService.submit(() -> captureAndSendMicrophoneAudio(liveRequestQueue, isRunning));
// Task for processing agent responses and playing audio
Future<?> outputTask =
executorService.submit(
() -> {
try {
processAudioOutput(eventStream, isRunning, conversationEnded);
} catch (Exception e) {
System.err.println("Error processing audio output: " + e.getMessage());
e.printStackTrace();
isRunning.set(false);
}
});
// Wait for user to press Enter to stop the conversation
System.out.println("Conversation started. Press Enter to stop...");
System.in.read();
System.out.println("Ending conversation...");
isRunning.set(false);
try {
// Give some time for ongoing processing to complete
microphoneTask.get(2, TimeUnit.SECONDS);
outputTask.get(2, TimeUnit.SECONDS);
} catch (Exception e) {
System.out.println("Stopping tasks...");
}
liveRequestQueue.close();
executorService.shutdownNow();
System.out.println("Conversation ended.");
}
private void captureAndSendMicrophoneAudio(
LiveRequestQueue liveRequestQueue, AtomicBoolean isRunning) {
TargetDataLine micLine = null;
try {
DataLine.Info info = new DataLine.Info(TargetDataLine.class, MIC_AUDIO_FORMAT);
if (!AudioSystem.isLineSupported(info)) {
System.err.println("Microphone line not supported!");
return;
}
micLine = (TargetDataLine) AudioSystem.getLine(info);
micLine.open(MIC_AUDIO_FORMAT);
micLine.start();
System.out.println("Microphone initialized. Start speaking...");
byte[] buffer = new byte[BUFFER_SIZE];
int bytesRead;
while (isRunning.get()) {
bytesRead = micLine.read(buffer, 0, buffer.length);
if (bytesRead > 0) {
byte[] audioChunk = new byte[bytesRead];
System.arraycopy(buffer, 0, audioChunk, 0, bytesRead);
Blob audioBlob = Blob.builder().data(audioChunk).mimeType("audio/pcm").build();
liveRequestQueue.realtime(audioBlob);
}
}
} catch (LineUnavailableException e) {
System.err.println("Error accessing microphone: " + e.getMessage());
e.printStackTrace();
} finally {
if (micLine != null) {
micLine.stop();
micLine.close();
}
}
}
private void processAudioOutput(
Flowable<Event> eventStream, AtomicBoolean isRunning, AtomicBoolean conversationEnded) {
SourceDataLine speakerLine = null;
try {
DataLine.Info info = new DataLine.Info(SourceDataLine.class, SPEAKER_AUDIO_FORMAT);
if (!AudioSystem.isLineSupported(info)) {
System.err.println("Speaker line not supported!");
return;
}
final SourceDataLine finalSpeakerLine = (SourceDataLine) AudioSystem.getLine(info);
finalSpeakerLine.open(SPEAKER_AUDIO_FORMAT);
finalSpeakerLine.start();
System.out.println("Speaker initialized.");
for (Event event : eventStream.blockingIterable()) {
if (!isRunning.get()) {
break;
}
AtomicBoolean audioReceived = new AtomicBoolean(false);
processEvent(event, audioReceived);
event.content().ifPresent(content -> content.parts().ifPresent(parts -> parts.forEach(part -> playAudioData(part, finalSpeakerLine))));
}
speakerLine = finalSpeakerLine; // Assign to outer variable for cleanup in finally block
} catch (LineUnavailableException e) {
System.err.println("Error accessing speaker: " + e.getMessage());
e.printStackTrace();
} finally {
if (speakerLine != null) {
speakerLine.drain();
speakerLine.stop();
speakerLine.close();
}
conversationEnded.set(true);
}
}
private void playAudioData(Part part, SourceDataLine speakerLine) {
part.inlineData()
.ifPresent(
inlineBlob ->
inlineBlob
.data()
.ifPresent(
audioBytes -> {
if (audioBytes.length > 0) {
System.out.printf(
"Playing audio (%s): %d bytes%n",
inlineBlob.mimeType(),
audioBytes.length);
speakerLine.write(audioBytes, 0, audioBytes.length);
}
}));
}
private void processEvent(Event event, java.util.concurrent.atomic.AtomicBoolean audioReceived) {
event
.content()
.ifPresent(
content ->
content
.parts()
.ifPresent(parts -> parts.forEach(part -> logReceivedAudioData(part, audioReceived))));
}
private void logReceivedAudioData(Part part, AtomicBoolean audioReceived) {
part.inlineData()
.ifPresent(
inlineBlob ->
inlineBlob
.data()
.ifPresent(
audioBytes -> {
if (audioBytes.length > 0) {
System.out.printf(
" Audio (%s): received %d bytes.%n",
inlineBlob.mimeType(),
audioBytes.length);
audioReceived.set(true);
} else {
System.out.printf(
" Audio (%s): received empty audio data.%n",
inlineBlob.mimeType());
}
}));
}
public static void main(String[] args) throws Exception {
LiveAudioRun liveAudioRun = new LiveAudioRun();
liveAudioRun.runConversation();
System.out.println("Exiting Live Audio Run.");
}
}
Run the Live Audio Run tool
To run Live Audio Run tool, use the following command on the adk-agents directory:
mvn compile exec:java
Then you should see:
$ mvn compile exec:java
...
Initializing microphone input and speaker output...
Conversation started. Press Enter to stop...
Speaker initialized.
Microphone initialized. Start speaking...
With this message, the tool is ready to take voice input. Talk to the agent with a question like What's the electron?.
!!! Caution When you observe the agent keep speaking by itself and doesn't stop, try using earphones to suppress the echoing.
Next, see Configuration to set the voice and turn detection, and Tools to give your live agent tools.
