# Browser Use Benchmarks

## The most accurate web agents, with the best-in-class stealth browsers.

Compare Browser Use against other web automation frameworks and cloud browser providers on real-world accuracy and stealth.

### OnlineMind2Web Benchmarking

| Web Agent                         | Accuracy |
|-----------------------------------|----------|
| Browser Use Cloud (v3)           | 97%      |
| ABP + Opus 4.6                   | 86%      |
| TinyFish                          | 81%      |
| Navigator                         | 78%      |
| Gemini CUA                       | 69%      |
| Stagehand (Gemini 2.5 CU)       | 65%      |
| OpenAI Operator                  | 61%      |
| Sonnet 4.0 CU                    | 61%      |
| Stagehand (Sonnet 4.5)           | 55%      |

#### About this benchmark

[Online-Mind2Web](https://github.com/OSU-NLP-Group/Online-Mind2Web) is the standard browser agent benchmark. 300 tasks across 136 live websites — shopping, finance, travel, government, and more. We run all 300 tasks. No tasks removed.

#### Methodology

- Evaluation: All tasks run on live websites.
- Scoring: Agentic judge built on Claude Agent SDK, aligned with human judges.
- Date: March 2026.

### Provider Accuracy

| Provider                          | Accuracy |
|-----------------------------------|----------|
| Browser Use Cloud (v3)           | 97%      |
| ABP + Opus 4.6                   | 86%      |
| TinyFish                          | 81%      |
| Navigator                         | 78%      |
| Gemini CUA                       | 69%      |
| Stagehand (Gemini 2.5 CU)       | 65%      |
| OpenAI Operator                  | 61%      |
| Sonnet 4.0 CU                    | 61%      |
| Stagehand (Sonnet 4.5)           | 55%      |
