Skip to main content
Browser Use Cloud has V2, V3, and V4 agent interfaces. The measurements below compare their behavior on the stated evaluation. See the developer overview when choosing an interface for a new integration. For current model, browser, and network rates, use the pricing page.

At a glance

V4 Agent — most accurate

V4 writes and runs code to complete long, multi-step browser tasks. It can research across many pages, compare options, follow detailed instructions, collect records, and save the results as a spreadsheet or in any format. Best for:
  • Bulk data collection
  • Complex tasks that span many pages and sites
  • Long, difficult instructions
Example tasks:
  • “Collect every document with its title, reference, and deadline.”
  • “Create a spreadsheet of all the pricing plans across these five competitors.”
In this evaluation, V4 was the most accurate and slower than V3.

V3 Agent — fastest

V3 looks at the page and acts step-by-step like a human would. In this evaluation, it was the fastest and less accurate than V4. Best for:
  • Simple tasks, such as finding publicly available information
  • Workloads where speed is essential
  • Open-ended research
Example tasks:
  • “Get a shipping quote for a 3 lb package between two zip codes.”
  • “Check whether this product is in stock and what it costs right now.”
In this evaluation, V3 was faster and less accurate than V4.

V2 Agent — first generation

V2 is based on the open-source agent architecture. In this evaluation it was less accurate than V3 and V4.

V4 completes the hardest tasks

V4 completed 76% of difficult, real-world browser tasks — nine points ahead of V3 and twenty-two ahead of V2. Browser Use Cloud success rate on hard web tasks: V2 Agent 54.25%, V3 Agent 67.25%, V4 Agent 76.47%

How we measured this

In July 2026, we ran each agent four times on Internal Bench Hard: 106 difficult tasks on real websites. Every run used Claude Opus 4.8, and the same independent judge scored each result. We left out tasks blocked by websites’ anti-bot measures. Accuracy is the share of included tasks the judge marked successful. Speed compares the average duration of those matched evaluation runs; simpler tasks usually run faster.