9281 Commits

Author SHA1 Message Date
Magnus Müller d21b3da4ac Trim HistoryItem freeze follow-up (#4902)
Keeps HistoryItem frozen while removing the expanded byte-prefix
explanation and byte-stability regression test added in #4890.

<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Keep `HistoryItem` frozen but trim docs and tests. Removes the
byte-prefix caching explanation and drops the byte-stability regression
test from #4890.

- **Refactors**
  - Shortened the `HistoryItem` docstring to a single line.
- Removed `tests/ci/test_history_item_byte_stable.py` to stop enforcing
byte-prefix rendering in unit tests.

<sup>Written for commit 51598efd56.
Summary will update on new commits. <a
href="https://cubic.dev/pr/browser-use/browser-use/pull/4902?utm_source=github">Review
in cubic</a></sup>

<!-- End of auto-generated description by cubic. -->
0.12.8
2026-05-23 12:25:54 -07:00
Magnus Müller 3f4207a466 Merge branch 'main' into trim-pr-4890 2026-05-23 12:25:45 -07:00
Magnus Müller d14a1c6d8c Revert evaluate restriction on restricted profiles (#4901)
Reverts #4871.

<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Reverts the `evaluate()` guard from #4871 to allow JavaScript evaluation
on profiles with `allowed_domains`, `prohibited_domains`, or
`block_ip_addresses` configured. Removes the restriction logic in
`browser_use/tools/service.py` and deletes
`tests/ci/security/test_evaluate_restricted.py`.

<sup>Written for commit af9d406419.
Summary will update on new commits. <a
href="https://cubic.dev/pr/browser-use/browser-use/pull/4901?utm_source=github">Review
in cubic</a></sup>

<!-- End of auto-generated description by cubic. -->
2026-05-23 12:25:38 -07:00
Magnus Müller 682c67dd92 Merge branch 'main' into revert-pr-4871 2026-05-23 12:25:25 -07:00
MagMueller 51598efd56 Trim HistoryItem freeze change tests 2026-05-23 12:20:55 -07:00
MagMueller af9d406419 Revert "fix(tools): refuse evaluate() on restricted browser profiles (#4871)"
This reverts commit d6a87c0961, reversing
changes made to eeff4d1984.
2026-05-23 12:13:38 -07:00
Magnus Müller a402567971 Revert ChatGoogle cached_content forwarding (#4900)
## Summary
- Revert #4889 (`ChatGoogle.cached_content` field and per-call
`cached_content` forwarding).
- Remove the now-unused cached-content unit test.

## Why
Cloud explicit caching does not need the OSS convenience API: the
gateway injects `cached_content` through a copied `ChatGoogle.config`
dict on its request-local model. Keeping the public field/kwarg exposes
a sharp edge because Gemini rejects cachedContent calls that also send
live `system_instruction`, which normal `ChatGoogle` messages can still
do.

## Tests
- `uv run pytest tests/ci/models/test_llm_google.py -q`
- `uv run pre-commit run --all-files`


<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Reverts `ChatGoogle` `cached_content` forwarding to Gemini and removes
the field and per-call passthrough. This avoids Gemini errors when
`cachedContent` is sent with live `system_instruction`; the gateway
still injects caching, so the obsolete unit test is removed.

<sup>Written for commit 22939a2580.
Summary will update on new commits. <a
href="https://cubic.dev/pr/browser-use/browser-use/pull/4900?utm_source=github">Review
in cubic</a></sup>

<!-- End of auto-generated description by cubic. -->
2026-05-23 12:05:58 -07:00
MagMueller 22939a2580 Revert "feat(llm/google): forward cached_content into generate_content (#4889)"
This reverts commit 5a745a8502, reversing
changes made to 640360e9b7.
2026-05-23 12:03:38 -07:00
Magnus Müller dddc13030c Bump version to 0.12.8 (#4899)
## Summary
- bump browser-use package version to 0.12.8 for the next release

## Tests
- uv run ruff check pyproject.toml
-
PATH="/Users/magnus/.local/share/uv/python/cpython-3.11-macos-aarch64-none/bin:/usr/local/bin:/System/Cryptexes/App/usr/bin:/usr/bin:/bin:/usr/sbin:/sbin:/var/run/com.apple.security.cryptexd/codex.system/bootstrap/usr/local/bin:/var/run/com.apple.security.cryptexd/codex.system/bootstrap/usr/bin:/var/run/com.apple.security.cryptexd/codex.system/bootstrap/usr/appleinternal/bin:/opt/pmk/env/global/bin:/opt/homebrew/bin:/Users/magnus/.codex/tmp/arg0/codex-arg0LWaL9w:/opt/homebrew/lib/node_modules/@openai/codex/node_modules/@openai/codex-darwin-arm64/vendor/aarch64-apple-darwin/path:/Users/magnus/.superset/bin:/Users/magnus/.whatdidido:/Users/magnus/.browser-use-env/bin:/Users/magnus/.local/bin:/Users/magnus/.bun/bin:/Users/magnus/.orbstack/bin:/Users/magnus/.orbstack/bin"
uv run pre-commit run --files pyproject.toml

<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Bumps `browser-use` to 0.12.8 to publish the next patch release. Updates
the `version` in pyproject.toml from 0.12.7 to 0.12.8.

<sup>Written for commit e973caeae1.
Summary will update on new commits. <a
href="https://cubic.dev/pr/browser-use/browser-use/pull/4899?utm_source=github">Review
in cubic</a></sup>

<!-- End of auto-generated description by cubic. -->
2026-05-23 11:24:58 -07:00
MagMueller e973caeae1 Bump version to 0.12.8 2026-05-23 11:24:04 -07:00
Magnus Müller 3818b1437c Move user request before agent history (#4897)
## Summary
- move the rendered <user_request> block before <agent_history> in the
per-step agent user message
- keep <agent_state> for file system, todo, plan, sensitive data, and
file path context
- update system prompt docs and prompt-shape tests so the cacheable
prefix order is locked
- add a simple guard that the rendered agent user message contains
exactly one </agent_history> marker, which the LLM gateway cache
splitter depends on

## Why
Putting the stable user request before the append-only agent history
gives implicit and explicit provider caches a better stable prefix. The
end of <agent_history> is still the expanding part; compaction naturally
creates a new stable prefix once the compacted memory block changes.

## Tests
- uv run ruff check browser_use/agent/prompts.py
tests/ci/test_prompt_step_meta_suffix.py
- uv run pytest tests/ci/test_prompt_step_meta_suffix.py -q
-
PATH="/Users/magnus/.local/share/uv/python/cpython-3.11-macos-aarch64-none/bin:$PATH"
uv run pre-commit run --files browser_use/agent/prompts.py
browser_use/agent/system_prompts/system_prompt.md
browser_use/agent/system_prompts/system_prompt_anthropic_flash.md
browser_use/agent/system_prompts/system_prompt_no_thinking.md
tests/ci/test_prompt_step_meta_suffix.py

<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Moves the `<user_request>` block before `<agent_history>` in each step’s
user message to create a more stable prefix for provider/LLM caches.
Updates system prompts and tests to lock this order and adds a guard
that enforces a single `</agent_history>` marker.

- **Refactors**
- Render `<user_request>` before `<agent_history>` in the user message;
remove `<user_request>` from `<agent_state>`.
- Keep `<agent_state>` focused on file system, todo, plan, sensitive
data, and available file paths; keep per-step metadata at the tail.
- Update `system_prompt.md`, `system_prompt_anthropic_flash.md`, and
`system_prompt_no_thinking.md` to reflect the new input order and
include the `<user_request>` block before history.
- Add a guard to ensure exactly one `</agent_history>` marker and extend
tests to lock ordering and suffix stability.

<sup>Written for commit 90a052371d.
Summary will update on new commits. <a
href="https://cubic.dev/pr/browser-use/browser-use/pull/4897?utm_source=github">Review
in cubic</a></sup>

<!-- End of auto-generated description by cubic. -->
2026-05-23 11:16:53 -07:00
MagMueller 90a052371d Move user request before agent history 2026-05-23 09:48:15 -07:00
Saurav Panda 5a745a8502 feat(llm/google): forward cached_content into generate_content (#4889)
## Context

Closes part of #4887 (item #4 — explicit CachedContent support).

Gemini's implicit context cache is great for the steady-state of an
agent loop but has a ~5-minute TTL and is per-call best-effort. For long
sessions (or sessions that go quiet for >5 minutes and resume), explicit
`CachedContent` is the deterministic alternative: you create a named
cache once with the prefix, then reference it on subsequent calls and
Google bills the cached portion at the discounted rate, period.

The wrapper in `browser_use/llm/google/chat.py` already parses
`cached_content_token_count` out of `response.usage_metadata` into
`ChatInvokeUsage.prompt_cached_tokens`, so the savings are visible — it
just had no way to *send* a cache reference.

## Change

- New `cached_content: str | None = None` field on `ChatGoogle`.
- New `cached_content` kwarg on `ainvoke()` — overrides the instance
default for one call.
- When set, threads into `config['cached_content']` before all three
`generate_content` calls (string output, native JSON, fallback JSON).

Per-call kwarg takes precedence over the instance default, so you can
pin a cache for the whole session and still override per-call when
needed.

## What's NOT in this PR

- Cache creation (`client.caches.create(...)`) is left to the caller.
That gives users full control over TTL, contents, and lifecycle without
coupling the LLM wrapper to agent-state decisions.
- Agent-side automatic cache lifecycle (create on first step, reuse on
subsequent steps, refresh near TTL) is an obvious follow-up but a
separate design discussion — it touches `agent/service.py` and depends
on what the agent considers its "stable prefix."

## Usage

```python
from google.genai import types
from google import genai

client = genai.Client(api_key=...)
cache = await client.aio.caches.create(
    model='gemini-2.5-flash',
    config=types.CreateCachedContentConfig(
        contents=[...prefix turns...],
        system_instruction='...',
        ttl='3600s',
    ),
)

llm = ChatGoogle(model='gemini-2.5-flash', cached_content=cache.name)
# Every ainvoke() now references that cache.

# Per-call override:
await llm.ainvoke(messages, cached_content='cachedContents/other')
```

## Test plan

- [x] New unit test `test_cached_content_threaded_into_config` —
verifies the kwarg flows into `generate_content`'s `config` dict, the
per-call kwarg overrides the instance default, and when unset the key is
absent (`pytest tests/ci/models/test_llm_google.py`).
- [x] Existing google tests still pass.
- [ ] Real-API smoke test against a Gemini model with an actual
`cachedContents/...` resource (requires API key + paid quota — not in
CI).

<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Adds explicit Gemini CachedContent support to `ChatGoogle` by forwarding
a `cached_content` name into all `generate_content` calls. Addresses
#4887 (item 4) to enable deterministic cache reuse with discounted
billing.

- **New Features**
- Add `cached_content` on `ChatGoogle` with per-call override via
`ainvoke(..., cached_content=...)`.
  - Thread to `config['cached_content']` for all paths; omit when unset.
- Cache creation is caller-managed; usage already counts
`prompt_cached_tokens`.

- **Migration**
- No breaking changes; behavior is unchanged when `cached_content` is
unset.
- To adopt: create a Gemini cache and pass its name via
`ChatGoogle(cached_content=...)` or `ainvoke(..., cached_content=...)`.

<sup>Written for commit ac5ff6dd15.
Summary will update on new commits. <a
href="https://cubic.dev/pr/browser-use/browser-use/pull/4889?utm_source=github">Review
in cubic</a></sup>

<!-- End of auto-generated description by cubic. -->
2026-05-22 18:40:39 -07:00
Saurav Panda ac5ff6dd15 Merge branch 'main' into prompt-cache/gemini-cached-content 2026-05-22 18:32:45 -07:00
Saurav Panda 640360e9b7 agent(prompts): move per-step metadata out of <agent_state> into a tail block (#4891)
## Context

Closes part of #4887 (item #3 — strip per-step metadata from anything
prefix-stable).

`AgentMessagePrompt._get_agent_state_description()` was rendering two
per-step-varying values inside `<agent_state>`:

- `Step{N+1} maximum:{M}` — changes every step.
- `datetime.now().strftime('%Y-%m-%d')` — changes daily.

The user message currently looks like:

```
<agent_history>...</agent_history>          ← grows append-only (prefix-stable if HistoryItem is stable)
<agent_state>...<step_info>...</...> </agent_state>   ← cache miss starts here today
<browser_state>...</browser_state>
<read_state>...</read_state>
```

So the cache boundary already lands at `<agent_state>` and the step
counter inside it isn't actively bursting a live cache. **But**: the
layout meant that any future move of more-stable `<agent_state>` fields
(user_request, file_system, todo_contents) into the system prompt — or
anywhere we'd want to cache them — would still leave per-step varying
bytes sitting inside the would-be prefix. That silently caps how far the
cache can ever extend.

## Change

Pull the step counter + date into a new helper
`_get_step_meta_description()` and append it at the very tail of
`get_user_message()`, after `<agent_state>`, `<browser_state>`,
`<read_state>`, `<page_specific_actions>`, and the unavailable-skills
info block. The new layout:

```
<agent_history>...</agent_history>
<agent_state>...</agent_state>              ← no more <step_info> inside
<browser_state>...</browser_state>
<read_state>...</read_state>
<page_specific_actions>...</page_specific_actions>
[unavailable_skills_info]
<step_info>Step{N} maximum:{M}\nToday:{YYYY-MM-DD}</step_info>   ← suffix, explicitly per-step
```

Everything above `<step_info>` is now eligible to be treated as the
cacheable region — when/if we want to push that boundary further out, no
per-step varying bytes are in the way.

## Tests

New regression tests at `tests/ci/test_prompt_step_meta_suffix.py`:
- `<step_info>` appears after both `<agent_state>` and
`<browser_state>`.
- `<step_info>` does not leak back into `<agent_state>`.
- Bytes before `<step_info>` are byte-identical across two different
step numbers (proves the step counter isn't in the prefix).
- `<agent_state>` block is byte-identical across step numbers.

## Test plan

- [x] New tests pass.
- [x] Existing prompt / message_manager tests still pass (`pytest
tests/ci -k 'prompt or message_manager or agent_message'`).
- [x] pyright + ruff clean via pre-commit.
- [ ] Eyeball one real agent loop to confirm the model still parses
`<step_info>` correctly at the tail (no expected change in behavior —
the LLM doesn't care about position).

<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Moved per-step metadata (step counter and date) out of `<agent_state>`
into a trailing `<step_info>` block so the user-message prefix is stable
for caching. Preps the prompt layout for deeper caching and covers part
of #4887.

- **Refactors**
- Added `_get_step_meta_description()` and append it at the end of
`get_user_message()` after agent, browser, read, page actions, and
unavailable-skills blocks.
- Removed per-step `<step_info>` from `<agent_state>` so all bytes
before `<step_info>` are stable across steps.
- Added tests to lock ordering, prevent leakage into `<agent_state>`,
and verify a byte-identical prefix and `<agent_state>` across step
numbers.

<sup>Written for commit b06b47a23a.
Summary will update on new commits. <a
href="https://cubic.dev/pr/browser-use/browser-use/pull/4891?utm_source=github">Review
in cubic</a></sup>

<!-- End of auto-generated description by cubic. -->
2026-05-22 18:32:16 -07:00
Saurav Panda b06b47a23a Merge branch 'main' into prompt-cache/relocate-per-step-metadata 2026-05-22 18:11:00 -07:00
Saurav Panda 67f7b9e084 agent(history): freeze HistoryItem + lock byte-prefix property (#4890)
## Context

Closes part of #4887 (item #1 — make the per-step transcript
append-only).

For Gemini's implicit cache (and similar provider caches) to actually
hit step over step, the rendered transcript for steps 1..N-1 must be
byte-identical at step N and at step N+1. The agent today appends
`HistoryItem`s immutably in practice — every site I checked constructs a
new item and `.append()`s it — but nothing in the type itself prevents a
future caller from mutating one in place. That kind of mutation would
silently kill the cache from that byte onward, and we'd only notice it
as a slow erosion of cached-token-ratio.

## Change

1. Mark `HistoryItem` `frozen=True` via pydantic `ConfigDict`, so any
mutation fails loud at runtime instead of silently burning input tokens
later.
2. Add a regression test set
(`tests/ci/test_history_item_byte_stable.py`) that asserts the cache
property directly:
- `render(items[:N])` is a strict byte-prefix of `render(items[:N+1])`
   - `to_string()` is deterministic for identical inputs
- the prefix property holds across mixed entry shapes (normal steps,
errors, system messages, follow-up tasks)
- conditional field inclusion in `to_string()` doesn't collapse to
ambiguous output (a populated field shouldn't render identically to a
missing one)

The tests pass against current behavior — they're locking it down, not
changing it.

## Known limitation (intentionally not in this PR)

`MessageManager.agent_history_description` does have one real
cache-buster I noticed during the audit: when `max_history_items` is set
and exceeded, the compaction logic at
`message_manager/service.py:173-186` rewrites earlier bytes — the `[...
N previous steps omitted...]` count changes every step past the cap, so
the prefix is not stable past the boundary.

Fixing that properly means either (a) capping the omitted-count at a
stable value, or (b) restructuring the compaction to omit at a fixed
cutoff. Either way it's a bigger design call than this hygiene PR, so I
left it for a follow-up. Calling it out in the commit message.

## Test plan

- [x] New byte-stability tests pass (`pytest
tests/ci/test_history_item_byte_stable.py`).
- [x] Existing message_manager / history tests still pass (`pytest
tests/ci -k 'message_manager or history'`).
- [x] `pyright` + `ruff` clean via pre-commit.

<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Freeze `HistoryItem` to make agent step transcripts append-only and
preserve byte-prefix stability for provider caches (e.g., Gemini). Adds
regression tests to lock deterministic rendering and strict prefix
growth. Part of #4887.

- **Refactors**
- Set `frozen=True` in `pydantic` `ConfigDict` for `HistoryItem`;
mutation now raises instead of silently changing rendered bytes.
- Added regression tests to enforce: strict byte-prefix growth across
steps, deterministic `to_string()`, stability with mixed entry shapes,
and disambiguation of optional fields.

<sup>Written for commit 8236cc7506.
Summary will update on new commits. <a
href="https://cubic.dev/pr/browser-use/browser-use/pull/4890?utm_source=github">Review
in cubic</a></sup>

<!-- End of auto-generated description by cubic. -->
2026-05-22 18:09:49 -07:00
Saurav Panda a3bae60ed2 agent(prompts): move per-step metadata out of <agent_state> into a tail block
The step counter (Step N maximum:M) and datetime.now() were rendered
inside <agent_state>, ahead of <browser_state> in the user message.
The cache miss already happens at the <agent_state> boundary today, so
this isn't a live cache regression — but the layout meant that any
future move of more-stable agent_state fields into the system prompt
would still leave per-step varying bytes in the middle of the prefix,
silently capping how far the cache could extend.

Pull both fields into a new _get_step_meta_description() and append it
at the very end of get_user_message(), after <agent_state>,
<browser_state>, <read_state>, <page_specific_actions>, and unavailable-
skills info. Everything above this tail block is now eligible to be
treated as the cacheable region.

Adds regression tests that lock the layout:
- <step_info> must appear after <agent_state> and <browser_state>
- <step_info> must not leak back into <agent_state>
- bytes before <step_info> must be identical across two different step
  numbers (the step counter must not be in the prefix)
2026-05-22 17:58:13 -07:00
Saurav Panda 8236cc7506 agent(history): freeze HistoryItem + lock byte-prefix property in tests
For Gemini's implicit cache (and similar provider caches) to actually
hit step over step, the rendered transcript for steps 1..N-1 must be
byte-identical at step N and at step N+1. The agent already appends
HistoryItems immutably in practice, but nothing in the type prevented
a future caller from mutating one in place — which would silently kill
the cache from that byte onward.

Mark HistoryItem frozen so any mutation now fails loud at runtime
rather than slowly burning input tokens. Add a regression test set
that asserts the cache property directly:

- render(items[:N]) must be a strict byte-prefix of render(items[:N+1])
- to_string() is deterministic for identical inputs
- the prefix property holds across mixed entry shapes (normal steps,
  errors, system messages, follow-up tasks)
- conditional field inclusion doesn't collapse to ambiguous output

Known limitation, not addressed here: max_history_items compaction in
MessageManager.agent_history_description rewrites earlier bytes once
the cap is exceeded (the omitted-count message changes). That's a
larger redesign and deserves its own PR.
2026-05-22 17:55:43 -07:00
Saurav Panda df063f9068 feat(llm/google): forward cached_content into generate_content
Adds a cached_content field on ChatGoogle (and a per-call kwarg on
ainvoke) that gets threaded into the GenerateContentConfigDict before
each generate_content call. This lets callers point Gemini at an
explicit CachedContent resource (e.g. "cachedContents/abc123") instead
of relying on implicit caching, which is constrained by the ~5-minute
TTL window.

Token accounting already pulls cached_content_token_count from
response.usage_metadata into ChatInvokeUsage.prompt_cached_tokens, so
the savings show up in usage stats without further work.

Cache creation itself (client.caches.create) is left to the caller —
this PR only adds the forwarding hook so explicit caching becomes
opt-in usable. A follow-up can wire agent-side cache lifecycle if
useful.
2026-05-22 17:54:04 -07:00
Saurav Panda fb9357fa1c fix(tokens): add OpenRouter pricing fallback 2026-05-22 13:58:42 -07:00
Saurav Panda db39960431 fix(tokens): add OpenRouter pricing fallback 2026-05-22 13:53:29 -07:00
Saurav Panda 2b30c4852f chore(llm): recommend gemini-3-flash-preview in examples and tests (#4885)
## Summary
- Adds `gemini-3-flash-preview-lite` to `VerifiedGeminiModels` and token
mappings (the non-lite preview was already verified).
- Lists both `gemini-3-flash-preview[-lite]` alongside the existing
`gemini-flash-latest` / `-lite-latest` aliases in `CLOUD.md` and
`skills/cloud/references/api-v2.md` `SupportedLLMs` — the old aliases
stay valid.
- Switches all example code / recommendations (`examples/`, `AGENTS.md`,
`skills/open-source/references/quickstart.md`, bug-report placeholder)
to `gemini-3-flash-preview`.
- Updates the Google CI button-click test and the `evaluate_tasks` judge
LLM to `gemini-3-flash-preview` / `gemini-3-flash-preview-lite`.

The thinking-config branch at `browser_use/llm/google/chat.py:268`
(`is_gemini_3_flash`) already handles both new names — defaults
`thinking_budget=-1` for parity with the old `gemini-flash` substring
path.

## Test plan
- [ ] `uv run pyright` clean (verified locally)
- [ ] pre-commit hooks pass (verified locally)
- [ ] `uv run pytest -vxs tests/ci/models/test_llm_google.py` against
`GOOGLE_API_KEY` — confirm the new `gemini-3-flash-preview` test passes
- [ ] `uv run pytest -vxs tests/ci/evaluate_tasks.py` (or whatever CI
invokes it) — confirm judge LLM swap still works
- [ ] Spot-check `examples/models/gemini.py` runs end-to-end against the
new model

<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Switch our recommended Google model to `gemini-3-flash-preview`, add
`gemini-3.1-flash-lite`, and support `gemini-3.1-pro-preview`. Examples,
docs, and CI now default to the new Flash models; existing
`gemini-flash-*-latest` aliases still work.

- **New Features**
- Added `gemini-3-flash-preview`, `gemini-3.1-flash-lite`, and
`gemini-3.1-pro-preview` to verified models; mapped the first two in
`MODEL_TO_LITELLM`.
- Listed `gemini-3-flash-preview` and `gemini-3.1-flash-lite` in
`CLOUD.md` and API `SupportedLLMs`.

- **Bug Fixes**
- Route `gemini-3.1-flash-*` and `gemini-3.1-pro-*` through the correct
thinking paths.
- Map `gemini-3.1-flash-lite` to LiteLLM slug
`gemini/gemini-3.1-flash-lite-preview` to avoid lookup failures while
keeping the user-facing name unchanged.

<sup>Written for commit 5ef0d9dbac.
Summary will update on new commits. <a
href="https://cubic.dev/pr/browser-use/browser-use/pull/4885?utm_source=github">Review
in cubic</a></sup>

<!-- End of auto-generated description by cubic. -->
2026-05-22 11:54:50 -07:00
Saurav Panda 5ef0d9dbac feat(llm/google): add gemini-3.1-pro-preview to verified models
Routes through the existing Gemini 3 Pro thinking branch by extending
is_gemini_3_pro to match 'gemini-3.1-pro' alongside 'gemini-3-pro'.
2026-05-22 11:48:10 -07:00
Saurav Panda 81f5c49854 fix(tokens): map gemini-3.1-flash-lite to litellm preview slug
LiteLLM only recognizes gemini/gemini-3.1-flash-lite-preview today; the GA
slug fails the model lookup. Keeps the user-facing name unchanged.
2026-05-22 11:03:14 -07:00
Saurav Panda 9e8dfdd18f fix(llm/google): route gemini-3.1-flash through gemini-3 flash thinking branch 2026-05-22 10:50:35 -07:00
Saurav Panda 5c54ef4197 chore(llm): rename flash-lite recommendation to gemini-3.1-flash-lite 2026-05-22 10:49:12 -07:00
Saurav Panda b00c41b66b chore(llm): recommend gemini-3-flash-preview in examples and tests
- add gemini-3-flash-preview-lite to VerifiedGeminiModels and token mappings
- list both gemini-3-flash-preview[-lite] alongside existing -latest aliases in CLOUD.md and api-v2.md SupportedLLMs
- swap example code and recommendations (examples/, AGENTS.md, quickstart.md, bug report placeholder) to gemini-3-flash-preview / -lite
- update Google CI test + evaluate_tasks judge LLM to gemini-3-flash-preview / -lite
2026-05-22 10:47:58 -07:00
Saurav Panda 2ce9de7734 feat: add client header to GoogleChat (#4884)
Added header per integration
[guidelines](https://ai.google.dev/gemini-api/docs/partner-integration).

<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Set the `x-goog-api-client` header for Google Chat requests to meet
Google partner integration guidelines. It identifies the client as
`browser-use/{version}` with a fallback of `unknown`.

- **New Features**
  - Add `x-goog-api-client` header with value `browser-use/{version}`.
- Normalize and merge `http_options`; always set/override the header in
client params.

- **Bug Fixes**
- Handle both `types.HttpOptions` and `types.HttpOptionsDict`,
preserving `timeout` and existing headers.
- Add tests covering `None`, Pydantic, and dict `http_options` to ensure
the header is set and prefixed with `browser-use/`.

<sup>Written for commit 115199d2ba.
Summary will update on new commits. <a
href="https://cubic.dev/pr/browser-use/browser-use/pull/4884?utm_source=github">Review
in cubic</a></sup>

<!-- End of auto-generated description by cubic. -->
2026-05-22 09:59:21 -07:00
Mark McDonald 115199d2ba fix: handle types.HttpOptionsDict better, add more tests 2026-05-22 14:30:40 +08:00
Mark McDonald 2eb3e3cded feat: add client header to GoogleChat
Added header per integration
[guidelines](https://ai.google.dev/gemini-api/docs/partner-integration).
2026-05-22 14:20:24 +08:00
Saurav Panda 5f99e737c5 chore(llm): default ChatBrowserUse to bu-2-0 (#4876)
## Summary
- Switch the `ChatBrowserUse` default `model` from `bu-latest` to
`bu-2-0`.
- Remap the `bu-latest` alias (both in `chat.py` normalization and the
`custom_pricing.py` pricing table) to `bu-2-0`, so anyone pinned to
`bu-latest` is upgraded along with the default.
- Also remap the `smart` pricing alias to `bu-2-0` so it tracks
`bu-latest`. Note: any historical usage records still emitting
`model="smart"` will be re-priced at `bu-2-0` rates (input $0.60/M,
output $3.50/M vs. bu-1-0's $0.20/M and $2.00/M).
- Refresh the docstring/usage example accordingly. `bu-1-0` is still
selectable as an explicit opt-in.

## Test plan
- [ ] `uv run pytest -vxs tests/ci/models/test_llm_browseruse.py`
(exercises the `bu-latest` alias path)
- [ ] `uv run pyright`
- [ ] Spot-check: `ChatBrowserUse()` (no args) resolves `.model ==
'bu-2-0'`
- [ ] Spot-check: `ChatBrowserUse(model='bu-latest').model == 'bu-2-0'`
2026-05-20 15:19:11 -07:00
Saurav Panda 83c8690f0b chore(llm): point smart pricing alias at bu-2-0
Keep the 'smart' alias in lockstep with bu-latest now that both
default to bu-2-0.
2026-05-20 12:05:39 -07:00
Saurav Panda 02daa7d80d chore(llm): default ChatBrowserUse to bu-2-0
Make bu-2-0 the default model and point bu-latest at it (was bu-1-0).
Pricing alias for bu-latest now tracks bu-2-0 too.
2026-05-20 12:05:04 -07:00
Saurav Panda d6a87c0961 fix(tools): refuse evaluate() on restricted browser profiles (#4871)
## Summary

Closes **GHSA-p6cf-6hf7-xh7q** (critical) and **GHSA-hvch-rcrr-7g39**
(critical).

The agent's `evaluate()` action calls `Runtime.evaluate` directly
through CDP. `SecurityWatchdog` only subscribes to navigation events
(`NavigateToUrlEvent`, `NavigationCompleteEvent`, `TabCreatedEvent`), so
JS running inside an already-allowed page could:

- `fetch()` arbitrary internal URLs (SSRF inside the agent's network)
- Read `document.cookie` / `localStorage` / `sessionStorage` from any
allowed origin's context
- Read `window.location` of cross-frame contexts
- Otherwise act as if `allowed_domains` / `block_ip_addresses` weren't
configured

When a profile is configured with `allowed_domains` or
`block_ip_addresses`, the operator has signalled "this agent is
constrained". An unmediated JS evaluation primitive contradicts that
signal.

## Changes

- `browser_use/tools/service.py:evaluate` — refuse the action with an
explanatory `ActionResult(error=...)` when `profile.allowed_domains` is
truthy or `profile.block_ip_addresses` is True.
- Empty `allowed_domains=[]` is treated as "no restriction" elsewhere
(e.g. `SecurityWatchdog._is_url_allowed:195`). Behaves consistently —
does not refuse in that case.

## Test plan

- [x] `uv run pytest -vxs tests/ci/security/test_evaluate_restricted.py`
— 5 tests:
  - refused when `allowed_domains=['example.com']`
  - refused when `block_ip_addresses=True`
  - refused when both
- proceeds to CDP on unrestricted profile (stub raises so the test
observes the guard did NOT short-circuit)
- proceeds when `allowed_domains=[]` (consistent with SecurityWatchdog
treating empty as no-restriction)
- [x] pyright / ruff check / ruff format — clean.

## Notes

This is the chosen approach from the design options I sketched (refuse
when restricted, vs route-through-event, vs lexical scan). Trade-off:
agents running on locked-down profiles lose `evaluate()` entirely.
Acceptable since restricted profiles are typically server-side /
production agents that don't need arbitrary JS.

If a future use case needs evaluate() inside restricted profiles, the
cleaner fix is to route through a watchdog-vetoable
`ExecuteJavaScriptEvent` — that's a larger refactor.

**Do not auto-merge** — please review.

<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Refuses `evaluate()` on restricted browser profiles to prevent bypassing
domain/IP controls. Closes GHSA-p6cf-6hf7-xh7q and GHSA-hvch-rcrr-7g39
by blocking SSRF and cross-origin data access via `Runtime.evaluate`.

- **Bug Fixes**
- In `browser_use/tools/service.py`, `evaluate()` now returns
`ActionResult(error=...)` when `profile.allowed_domains`,
`profile.prohibited_domains`, or `profile.block_ip_addresses` is set;
`allowed_domains=[]` stays unrestricted.
- Added `tests/ci/security/test_evaluate_restricted.py` covering refusal
for `allowed_domains`, `prohibited_domains`, `block_ip_addresses`, and
both; and pass-through on unrestricted and empty-`allowed_domains`.

- **Refactors**
  - Tightened the guard’s inline comment; no behavior change.

<sup>Written for commit f4d7ab4979.
Summary will update on new commits. <a
href="https://cubic.dev/pr/browser-use/browser-use/pull/4871?utm_source=github">Review
in cubic</a></sup>

<!-- End of auto-generated description by cubic. -->
2026-05-19 20:53:30 -07:00
Saurav Panda f4d7ab4979 Merge branch 'main' into fix/sec-evaluate-restricted 2026-05-19 12:04:46 -07:00
Saurav Panda eeff4d1984 docs: clarify integration example placement (#4856)
## Summary
- add `examples/integrations/README.md` with placement rules for
integration examples, package integrations, custom-function examples,
and external community projects
- add a link from `.github/CONTRIBUTING.md` so contributors can find the
guidance before opening integration PRs
- avoid adding any third-party example code directly; this only defines
the contribution boundary

Resolves #4744.

## Tests
- `uv run pre-commit run --all-files`

<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Adds `examples/integrations/README.md` with placement rules, an example
checklist, and a community listing format, and links it from
`.github/CONTRIBUTING.md`. No third-party code added; this only sets the
contribution boundary. Resolves #4744.

<sup>Written for commit 7a33d8d251.
Summary will update on new commits. <a
href="https://cubic.dev/pr/browser-use/browser-use/pull/4856?utm_source=github">Review
in cubic</a></sup>

<!-- End of auto-generated description by cubic. -->
2026-05-19 11:45:13 -07:00
Saurav Panda 7a33d8d251 Merge branch 'main' into felix/integration-examples-guidance 2026-05-19 11:41:09 -07:00
Saurav Panda 0fa8768d56 Merge branch 'main' into fix/sec-evaluate-restricted 2026-05-19 11:40:00 -07:00
Saurav Panda e5f9b46f95 refactor(tools): tighten evaluate() guard comment
Drop the running narrative; the function name + the one-line why is
enough. Error message stays actionable but no longer reiterates the
threat model.

No behavior change; 6/6 tests pass.
2026-05-18 18:42:29 -07:00
Saurav Panda 157779338a fix(daemon): restrict unix socket file to owner-only access (#4870)
## Summary

The per-session HMAC auth token (commit `de14b9aa`, pre-0.12.6) gates
command dispatch on the daemon, but `asyncio.start_unix_server` creates
the socket file with the process umask — which is `0o022` by default,
leaving the socket at `0o755` (`srwxr-xr-x`). On multi-user hosts a
co-tenant can `connect()` and probe behavior even though the handshake
will ultimately fail.

This is the remaining sub-issue tracked by **GHSA-x6mv-rq4m-58g8** (the
RCE primitive itself was already closed by the auth-token fix).

## Changes

- `browser_use/skill_cli/daemon.py:399-415` — wrap `start_unix_server`
with `os.umask(0o077)` and explicitly `os.chmod(sock_path, 0o600)` after
bind. Matches the posture of the auth-token file at line 353.

## Test plan

- [x] `uv run pytest -vxs tests/ci/security/test_daemon_socket_perms.py`
— new test starts a real Daemon, waits for the bind, asserts os.stat
returns 0o600, signals shutdown. Pre-fix this failed with 0o755. Uses
/tmp for BROWSER_USE_HOME so the AF_UNIX path stays under the 104-byte
macOS cap.
- [x] pyright / ruff check / ruff format — clean.

## Notes

Part of the post-0.12.7 security cleanup. **Do not auto-merge** — please
review.

<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Locks down the daemon’s Unix socket to owner-only access to block
co-tenant probing on multi-user hosts. Addresses the remaining sub-issue
from GHSA-x6mv-rq4m-58g8.

- **Bug Fixes**
- Enforce owner-only permissions on the Unix socket: wrap
`asyncio.start_unix_server` with `os.umask(0o077)` and explicitly
`chmod(sock_path, 0o600)` after bind; logs a warning if chmod fails.
Aligns with the auth-token file posture.
- Add `tests/ci/security/test_daemon_socket_perms.py` to assert the
socket mode is `0o600` after startup (skipped on Windows).

<sup>Written for commit 8c41cf795a.
Summary will update on new commits. <a
href="https://cubic.dev/pr/browser-use/browser-use/pull/4870?utm_source=github">Review
in cubic</a></sup>

<!-- End of auto-generated description by cubic. -->
2026-05-18 18:38:23 -07:00
Saurav Panda 8f2967a5c0 fix(tools): include prohibited_domains in evaluate guard
SecurityWatchdog enforces prohibited_domains as a navigation restriction
alongside allowed_domains. Without prohibited_domains in the evaluate()
guard, a profile using only the deny-list could still load an allowed page
and have the agent fetch() into a blocked domain via JS.

Add prohibited_domains to the restriction check.
2026-05-18 18:37:12 -07:00
Saurav Panda a7ce680948 fix(tools): refuse evaluate() on restricted browser profiles
The agent's evaluate() action called Runtime.evaluate directly through
CDP. SecurityWatchdog only subscribes to navigation events, so JS
running inside an already-allowed page could fetch() arbitrary internal
URLs, read cookies and localStorage from any allowed origin's context,
and otherwise act as if allowed_domains / block_ip_addresses weren't
configured.

When a profile has allowed_domains or block_ip_addresses set, the
operator has signalled the agent is constrained — exposing an unmediated
JS evaluation primitive contradicts that signal. Refuse evaluate()
outright on such profiles; agents that need JS can run on an unrestricted
profile.

Empty allowed_domains=[] is treated as 'no restriction' elsewhere in the
codebase (e.g. SecurityWatchdog); evaluate() behaves consistently and
does not refuse in that case.
2026-05-18 18:30:04 -07:00
Saurav Panda 8c41cf795a fix(daemon): restrict unix socket file to owner-only access
The per-session HMAC auth token (de14b9aa) gates command dispatch, but
asyncio.start_unix_server creates the socket file with the process
umask, leaving it 0o755 by default. On multi-user hosts a co-tenant
can connect() and probe behavior even though the handshake will
ultimately fail.

Set umask to 0o077 around start_unix_server and chmod 0o600 after,
matching the auth-token file's posture.
2026-05-18 18:28:10 -07:00
Saurav Panda 18aae0b752 Bump version from 0.12.6 to 0.12.7 (#4869)
<!-- This is an auto-generated description by cubic. -->
## Summary by cubic
Bumps `browser-use` version from 0.12.6 to 0.12.7 to publish a patch
release. Only updates pyproject.toml.

<sup>Written for commit c7085726d2.
Summary will update on new commits. <a
href="https://cubic.dev/pr/browser-use/browser-use/pull/4869?utm_source=github">Review
in cubic</a></sup>

<!-- End of auto-generated description by cubic. -->
0.12.7
2026-05-18 18:01:12 -07:00
Saurav Panda c7085726d2 Bump version from 0.12.6 to 0.12.7 2026-05-18 17:41:49 -07:00
Saurav Panda 6f3d3ff610 fix(downloads): sanitize attacker-controlled filenames and verify containment (#4867)
## Summary

Closes **GHSA-rv9j-wqjp-2fv4** (critical, `triage`) and the duplicate
medium reports **GHSA-66xh-g88g-2h8j** / **GHSA-hpr4-fqgr-xhj9**.

`DownloadsWatchdog` joined attacker-controlled filenames from CDP
(`Page.downloadWillBegin.suggestedFilename`) and `Content-Disposition`
headers directly into the configured `downloads_path`. Strings like
`../../escape.bin` or `/etc/shadow.bak` would `os.path.join` outside the
downloads directory, writing the fetched bytes (also attacker-controlled
— the response body IS the exploit content) to an arbitrary location
with the agent's process privileges.

`download_file_from_url` triggers passively for any
`Content-Disposition: attachment` response, so this is reachable from
**any visited site** — `allowed_domains` does **not** mitigate it.

## Changes

Two new private helpers on `DownloadsWatchdog`:

- `_sanitize_download_filename(name)` — keep only basename, normalize
Windows separators, strip null bytes, fall back to `'download'` for
empty / pure-traversal inputs.
- `_is_path_contained(path, dir)` — `os.path.realpath` containment check
for the on-disk sinks.

**Sanitizer wired at every attacker-controlled filename ingress:**

| Site | What it reads |
|------|---------------|
| `download_will_begin_handler` | CDP `suggestedFilename` (cache +
events) |
| `_handle_cdp_download` | same field, separate code path |
| Network-monitor `Content-Disposition` parser |
`re.search(...).group(1)` |
| `download_file_from_url(suggested_filename=...)` | upstream-passed
filename |
| `_handle_download` | Playwright `download.suggested_filename` |

**Containment check wired at every on-disk write site:**

- `download_file_from_url` write (line ~755)
- `_handle_download` (Playwright `save_as` path)
- `trigger_pdf_download` write (defense in depth — already basename'd,
but pinned)

## Test plan

- [x] `uv run pytest -vxs
tests/ci/security/test_download_filename_sanitization.py` — 16 new tests
covering: relative traversal, absolute Unix paths, Windows backslash
paths, mixed separators, pure-traversal fallback, null-byte stripping,
empty/None fallback, normal filename preservation, Unicode preservation;
containment helper (inside / nested / escape / dir-itself /
sibling-dir); `_get_unique_filename` collision handling on sanitized
input.
- [x] `uv run pytest -vx tests/ci/security/` — 80/80 pass (existing
security suite unchanged).
- [x] `uv run pyright` / `ruff check` / `ruff format` — clean.
- [ ] Full `uv run pytest -vxs tests/ci` on CI.

## Notes

This is the most reachable of the post-0.12.6 critical advisories — no
prompt injection or domain bypass needed, any visited site can trigger
it. Recommend prioritizing this PR's review.

**Do not auto-merge** — please review.

<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Fixes a critical path traversal in download handling by sanitizing
attacker-controlled filenames and enforcing realpath containment before
writing files. Prevents arbitrary file writes via CDP
`suggestedFilename` and `Content-Disposition` from any visited site
(fixes GHSA-rv9j-wqjp-2fv4 and duplicates); trims comments/docstrings
with no behavior change.

- **Bug Fixes**
- Added `_sanitize_download_filename` and `_is_path_contained` helpers.
- Sanitized filenames from CDP events, Playwright
`download.suggested_filename`, `Content-Disposition`, and
`download_file_from_url`.
- Enforced containment at all write sites (`download_file_from_url`,
Playwright save path, PDF export); refuse writes outside `downloads_dir`
(covers symlink escapes).
- Added tests for traversal, absolute paths, Windows/mixed separators,
null bytes, empty/None, unicode, containment behavior, and
`_get_unique_filename` collision handling.

<sup>Written for commit f0413fbd8c.
Summary will update on new commits. <a
href="https://cubic.dev/pr/browser-use/browser-use/pull/4867?utm_source=github">Review
in cubic</a></sup>

<!-- End of auto-generated description by cubic. -->
2026-05-18 17:38:01 -07:00
Saurav Panda f0413fbd8c refactor(downloads): tighten download-sanitization comments
Drop the per-call-site rationale text; the helper names already convey
what's happening, and the WHY lives in the commit history. Shrink both
helper docstrings to one line.

No behavior change; 16/16 tests still pass.
2026-05-18 17:31:09 -07:00
Saurav Panda 29b03e43b4 Merge branch 'main' into fix/sec-download-path-traversal 2026-05-18 17:28:52 -07:00
Saurav Panda 6a275d3962 fix(security): canonicalize non-standard IPv4 forms in block_ip_addresses (#4866)
## Summary

Closes **GHSA-xrfv-gg9f-wwjp** and **GHSA-g27c-8gp4-28cv** (both high,
`triage`, duplicate reports).

`SecurityWatchdog._is_ip_address` only recognized IP strings that
`ipaddress.ip_address()` accepts — i.e. the canonical dotted-quad form
and full IPv6. Chromium and the kernel resolver, however, also accept
several non-standard IPv4 representations:

| Form | Resolves to |
|------|-------------|
| `http://2130706433/` | `127.0.0.1` (decimal int) |
| `http://0x7f000001/` | `127.0.0.1` (hex) |
| `http://0177.0.0.1/` | `127.0.0.1` (octal) |
| `http://127.1/`      | `127.0.0.1` (short-form) |
| `http://127.0.1/`    | `127.0.0.1` (short-form) |

`block_ip_addresses=True` was therefore trivially bypassed by
re-encoding the IP in any of these forms.

## Changes

- `browser_use/browser/watchdogs/security_watchdog.py:_is_ip_address` —
after `ipaddress.ip_address()` fails, fall back to `socket.inet_aton`.
`inet_aton` accepts the same liberal IPv4 forms the kernel resolver
does, so the classifier matches the browser's behavior.
- IPv6 brackets stripped defensively before parsing.

## Test plan

- [x] `uv run pytest -vxs tests/ci/security/test_ip_blocking.py` — 34/34
pass.
- [x] New `TestNonStandardIPv4Representations` class covers: decimal,
hex, octal, short-form blocking; lookalike-domain non-interference
(`127.0.0.1.evil.com`, `2130706433.evil.com`); interaction with
`block_ip_addresses=False`; interaction with `allowed_domains=['*']`.
- [x] Existing `test_ipv4_lookalike_domains_allowed` was codifying the
buggy behavior for `1.2.3` (which IS a short-form IPv4 == 1.2.0.3 that
Chromium resolves). Removed that assertion and documented the
cross-reference.
- [x] `uv run pyright` / `ruff check` / `ruff format` — clean.
- [ ] Full `uv run pytest -vxs tests/ci` on CI.

## Notes

Part of the post-0.12.6 security advisory triage. **Do not auto-merge**
— please review.

<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Closes GHSA-xrfv-gg9f-wwjp and GHSA-g27c-8gp4-28cv by canonicalizing
non-standard, percent-encoded, Unicode-variant, and IDNA-dot IPv4
hostnames so `block_ip_addresses=True` can’t be bypassed. Aligns
detection with browser/kernel behavior by stripping IPv6 brackets,
percent-decoding, NFKC-normalizing, folding IDNA label separators to
`.`, and falling back to `socket.inet_aton`; also simplifies the
`_is_ip_address` docstring with no behavior change.

- **Bug Fixes**
- In `_is_ip_address`, strip `[]`, percent-decode, NFKC-normalize, fold
`。`/`。`/`.` to `.`, try `ipaddress.ip_address()`, then fall back to
`socket.inet_aton` to catch decimal/hex/octal/short IPv4 (including
`%XX`-encoded and Unicode-digit forms); catch all exceptions; lookalike
domains (e.g., `127.0.0.1.evil.com`) remain allowed.

- **Tests**
- Added coverage for decimal/hex/octal/short IPv4 forms, percent-encoded
hosts, Unicode digit variants, IDNA dot separators (`。` `。` `.`), IDN
domains not misclassified, and malformed Unicode or `%` escapes; removed
the outdated `1.2.3` assertion.

<sup>Written for commit 4866656bba.
Summary will update on new commits. <a
href="https://cubic.dev/pr/browser-use/browser-use/pull/4866?utm_source=github">Review
in cubic</a></sup>

<!-- End of auto-generated description by cubic. -->
2026-05-18 17:28:40 -07:00