Files
MagMueller 6f8b0626b6 docs: point bu-30b-a3b-preview at self-hosting instead of Cloud
Browser Use Cloud does not serve `browser-use/bu-30b-a3b-preview`. The gateway
still has the route (cloud `backend/llm_use/gateway/pricing.py` MODAL_MODELS,
`service.py` _call_modal), but the Modal app behind it,
`browser-use-llm-prod` in browser-use/deploy-llm, was last deployed on
2025-12-16 and every production call since returns upstream HTTP 503, which the
gateway reports to the caller as a generic 500.

The weights are public and in use: https://huggingface.co/browser-use/bu-30b-a3b-preview
is a public repo with 2.34k downloads in the last month. So instead of dropping
the model, say what it actually is - open weights you host yourself.

- `examples/models/bu_oss.py` now starts from a vLLM server and talks to it
  through `ChatOpenAI`, so it needs no BROWSER_USE_API_KEY and does not depend
  on the dead Cloud route.
- The `ChatBrowserUse` docstring says Cloud does not serve it.
- The skills model table gains a self-hosting section with the vLLM command
  from the model card, and loses the priceless OSS row.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018oComYHNbeV4e21v22Bbdn
2026-09-12 00:22:16 -07:00

44 lines
1.1 KiB
Python

"""Run the open-weights Browser Use model on your own GPU.
`browser-use/bu-30b-a3b-preview` is published under
https://huggingface.co/browser-use/bu-30b-a3b-preview. It is open weights you host
yourself, not a model Browser Use Cloud serves for you, so it needs no
BROWSER_USE_API_KEY and has no per-token price.
Setup:
1. pip install vllm
2. vllm serve browser-use/bu-30b-a3b-preview --max-model-len 65536 --host 0.0.0.0 --port 8000
3. python examples/models/bu_oss.py
Point BU_OSS_BASE_URL at the server if it is not on localhost.
"""
import os
from dotenv import load_dotenv
from browser_use import Agent, ChatOpenAI
load_dotenv()
try:
from lmnr import Laminar
Laminar.initialize()
except ImportError:
pass
# Any OpenAI-compatible server works; vLLM is what the model card recommends.
llm = ChatOpenAI(
model='browser-use/bu-30b-a3b-preview',
base_url=os.getenv('BU_OSS_BASE_URL', 'http://localhost:8000/v1'),
api_key=os.getenv('BU_OSS_API_KEY', 'not-needed'),
)
agent = Agent(
task='Find the number of stars of browser-use and stagehand. Tell me which one has more stars :)',
llm=llm,
flash_mode=True,
)
agent.run_sync()