* Add MLflow scorers integration page Adds a new integration page at docs/integrations/mlflow-scorers.md covering MLflow's five Google ADK scorers (ToolTrajectory, ResponseMatch, ResponseEvaluation, Safety, Hallucination) for agent evaluation. The integration wraps ADK's TrajectoryEvaluator, RougeEvaluator, FinalResponseMatchV2Evaluator, SafetyEvaluatorV1, and HallucinationsV1Evaluator behind MLflow's scorer interface, so ADK users can evaluate agents inside mlflow.genai.evaluate() runs without leaving the ADK ecosystem. Complements the existing MLflow Tracing and MLflow AI Gateway integration pages by covering evaluation, the third leg of the MLflow stack for ADK. Signed-off-by: debu-sinha <debusinha2009@gmail.com> * Trigger CLA re-check Signed-off-by: debu-sinha <debusinha2009@gmail.com> * Use version-agnostic Gemini aliases and refine copy * Update category tags for other pages --------- Signed-off-by: debu-sinha <debusinha2009@gmail.com> Co-authored-by: Kristopher Overholt <koverholt@google.com>
5.0 KiB
catalog_title, catalog_description, catalog_icon, catalog_tags
| catalog_title | catalog_description | catalog_icon | catalog_tags | ||
|---|---|---|---|---|---|
| LangWatch | Observability, tracing, evaluation, and prompt optimization for ADK agents | /integrations/assets/langwatch.png |
|
LangWatch observability for ADK
LangWatch is an open-source LLMOps platform for observability, evaluation, and prompt optimization. It provides comprehensive tracing for ADK agents using OpenInference instrumentation, allowing you to monitor, debug, and improve your agents in development and production.
Overview
LangWatch captures traces from ADK using its built-in OpenTelemetry support, giving you:
- Automatic tracing - Capture every agent run, tool call, and model request with full context
- Online evaluation - Continuously score production traffic for quality and safety
- Guardrails - Block or modify harmful responses in real-time
- Prompt management - Version, test, and optimize prompts with built-in A/B testing
- Datasets and experiments - Build evaluation sets from real traces and run batch experiments
Installation
Install the required packages:
pip install langwatch openinference-instrumentation-google-adk google-adk
Setup
Sign up at langwatch.ai or self-host the platform, then set your API key:
export LANGWATCH_API_KEY="your-langwatch-api-key"
export GOOGLE_API_KEY="your-gemini-api-key"
Initialize tracing:
import langwatch
from openinference.instrumentation.google_adk import GoogleADKInstrumentor
langwatch.setup(
instrumentors=[GoogleADKInstrumentor()]
)
That's it. All ADK agent activity will now be traced and sent to your LangWatch dashboard automatically.
Observe
With tracing initialized, run your ADK agent as usual and all interactions will appear in LangWatch:
import langwatch
from google.adk.agents import Agent
from google.adk.runners import InMemoryRunner
from google.genai import types
from openinference.instrumentation.google_adk import GoogleADKInstrumentor
langwatch.setup(
instrumentors=[GoogleADKInstrumentor()]
)
# Define a tool
def get_weather(city: str) -> dict:
"""Retrieves the current weather report for a specified city.
Args:
city (str): The name of the city.
Returns:
dict: status and result or error msg.
"""
if city.lower() == "new york":
return {
"status": "success",
"report": (
"The weather in New York is sunny with a temperature of 25 degrees"
" Celsius (77 degrees Fahrenheit)."
),
}
else:
return {
"status": "error",
"error_message": f"Weather information for '{city}' is not available.",
}
# Create an agent with tools
agent = Agent(
name="weather_agent",
model="gemini-flash-latest",
description="Agent to answer questions about the weather.",
instruction="You must use the available tools to find an answer.",
tools=[get_weather],
)
app_name = "weather_app"
user_id = "test_user"
session_id = "test_session"
runner = InMemoryRunner(agent=agent, app_name=app_name)
session_service = runner.session_service
await session_service.create_session(
app_name=app_name,
user_id=user_id,
session_id=session_id,
)
# Run the agent — all interactions will be traced
async for event in runner.run_async(
user_id=user_id,
session_id=session_id,
new_message=types.Content(
role="user",
parts=[types.Part(text="What is the weather in New York?")],
),
):
if event.is_final_response():
print(event.content.parts[0].text.strip())
Adding Custom Metadata
Use the @langwatch.trace() decorator to attach additional context to your
traces:
@langwatch.trace(name="ADK Weather Agent")
def run_agent(user_message: str):
current_trace = langwatch.get_current_trace()
if current_trace:
current_trace.update(
metadata={
"user_id": "user_123",
"agent_name": "weather_agent",
"environment": "production",
}
)
user_msg = types.Content(
role="user", parts=[types.Part(text=user_message)]
)
for event in runner.run(
user_id="demo-user",
session_id="demo-session",
new_message=user_msg,
):
if event.is_final_response():
return event.content.parts[0].text
return "No response generated"