Adds a cached_content field on ChatGoogle (and a per-call kwarg on
ainvoke) that gets threaded into the GenerateContentConfigDict before
each generate_content call. This lets callers point Gemini at an
explicit CachedContent resource (e.g. "cachedContents/abc123") instead
of relying on implicit caching, which is constrained by the ~5-minute
TTL window.
Token accounting already pulls cached_content_token_count from
response.usage_metadata into ChatInvokeUsage.prompt_cached_tokens, so
the savings show up in usage stats without further work.
Cache creation itself (client.caches.create) is left to the caller —
this PR only adds the forwarding hook so explicit caching becomes
opt-in usable. A follow-up can wire agent-side cache lifecycle if
useful.
- add gemini-3-flash-preview-lite to VerifiedGeminiModels and token mappings
- list both gemini-3-flash-preview[-lite] alongside existing -latest aliases in CLOUD.md and api-v2.md SupportedLLMs
- swap example code and recommendations (examples/, AGENTS.md, quickstart.md, bug report placeholder) to gemini-3-flash-preview / -lite
- update Google CI test + evaluate_tasks judge LLM to gemini-3-flash-preview / -lite
The 'type' key is already handled by an earlier elif branch
(elif key == 'type'), making its presence in the later elif key in [...]
block dead code. Remove it to eliminate confusion.
Fixes#4703
- Add UTM params to all cloud-bound links across README, CLI, and error messages
- Rewrite README Open Source vs Cloud section: position cloud browsers as
recommended pairing for OSS users, remove separate Use Both section
- Rewrite error messages for use_cloud=True and ChatBrowserUse() to clearly
state what is wrong and what to do next
- Add missing URLs: invalid API key now links to key page, insufficient
credits now links to billing page
- Add cloud browser nudge on captcha detection (logger.warning)
- Add cloud browser nudge on local browser launch failure
litellm versions 1.82.7 and 1.82.8 were backdoored on March 24, 2026
by TeamPCP via a compromised Trivy CI/CD pipeline. browser-use 0.12.3
shipped litellm>=1.82.2 (unpinned) as a core dependency, exposing
~6,900 users to the backdoored versions during the 4-hour window.
This commit:
- Removes litellm entirely from pyproject.toml (core and optional)
- Keeps ChatLiteLLM wrapper intact with a docstring noting
`pip install litellm` is required separately
- litellm is already lazy-imported inside methods, so users who
don't use ChatLiteLLM are never affected
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
OpenAI's completion_tokens already includes reasoning_tokens as a subset,
so adding them again was incorrectly inflating token counts by ~2x for
all reasoning models (o1, o3, o4-mini, etc.).
This aligns with other OpenAI-compatible providers (OpenRouter, Vercel,
Cerebras, Groq) which use completion_tokens directly.
Note: This is different from Google Gemini where thinking_tokens ARE
reported separately and need to be added.
Fixes#4065
Fix indentation so structured-output request executes even when add_schema_to_system_prompt is false; add regression tests for choices=null in both structured and non-structured paths.
Avoid indexing response.choices[0] when proxies return choices=null/empty; raise a clear ModelProviderError (with base_url hint) and use the validated first choice for content/finish_reason
- agent/views.py: Move RateLimitError import inside format_error() function
to avoid loading openai SDK (~800ms) at module level
- llm/messages.py: Replace openai.BaseModel with pydantic.BaseModel directly
to remove unnecessary openai dependency
- filesystem/file_system.py: Move reportlab imports inside sync_to_disk_sync()
to avoid ~40ms startup cost when PDF generation is not used
- utils.py: Convert OpenAIBadRequestError and GroqBadRequestError to lazy
loaders to avoid loading SDKs at module level
This improves import time for users who don't use OpenAI provider,
especially when using Anthropic, Google, or other providers.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Co-Authored-By: Claude Opus 4.5 <noreply@anthropic.com>