Picks up browser-use/benchmark#24. The luna bar is now the library's default
configuration (31.0%) labelled plain 'Luna', rather than an explicitly
effort-matched xhigh arm (35.0%). Default is 'low' -- ChatOpenAI hardcodes it at
browser_use/llm/openai/chat.py:39.
Propagates the regenerated plot from browser-use/benchmark#24. The committed
copies predate 2026-05-01 and showed retired models (ChatBrowserUse-2, gpt-5,
gpt-5-mini, gemini-2.5-flash, claude-opus-4-6) against Cloud v3 bu-ultra.
Now: Cloud v4 Opus 4.8 85.0% and Cloud v4 default (gpt-5.6-luna @ xhigh) 78.0%
against the open-source library on browser-use 0.13.7, topping out at
claude-opus-4-7 74.0%. Every bar measured 2026-08-04 at a uniform
task_timeout=3600.