Agentic benchmark
BrowseComp leaderboard
Every model the catalog carries a published BrowseComp value for, ranked by that value.
A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.
BrowseComp ranking
43 models with a published BrowseComp value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Web search and evidence synthesis |
|---|---|---|---|
| 1 | Shanghai Artificial Intelligence Laboratory | 92.5 | |
| 2 | OpenAI | 92.2 | |
| 3 | OpenAI | 91.5 | |
| 4 | Moonshot AI | 91.2 | |
| 5 | Anthropic | 90.8 | |
| 6 | OpenAI | 90.1 | |
| 7 | OpenAI | 89.3 | |
| 8 | Anthropic | 88 | |
| 9 | OpenAI | 87.5 | |
| 10 | OAOrnith-1.5-397B | Ornith AI | 86.6 |
| 11 | Anthropic | 84.7 | |
| 12 | OpenAI | 84.4 | |
| 13 | Anthropic | 84.3 | |
| 14 | Anthropic | 83.7 | |
| 15 | MiniMax | 83.52 | |
| 16 | DeepSeek | 83.4 | |
| 17 | DSdots3-note Preview | Dots Studio | 83.3 |
| 17 | OpenAI | 83.3 | |
| 19 | Moonshot AI | 83.2 | |
| 20 | OpenAI | 82.7 | |
| 21 | Anthropic | 79.3 | |
| 22 | TMInkling-Small | Thinking Machines Lab | 77.4 |
| 23 | TMInkling | Thinking Machines Lab | 77.1 |
| 24 | StepFun | 75.82 | |
| 25 | InternScience | 75.51 | |
| 26 | DeepSeek | 73.2 | |
| 27 | InclusionAI | 72.2 | |
| 28 | Z.AI | 68 | |
| 29 | OAOrnith-1.5-35B-A3B | Ornith AI | 67.6 |
| 30 | InternScience | 66.8 | |
| 31 | OpenAI | 65.8 | |
| 32 | Alibaba | 63.8 | |
| 33 | Alibaba | 62 | |
| 34 | Alibaba | 61 | |
| 34 | Alibaba | 61 | |
| 36 | Moonshot AI | 60.6 | |
| 36 | Moonshot AI | 60.6 | |
| 38 | OAOrnith-1.5-9B | Ornith AI | 56.4 |
| 39 | Z.AI | 52 | |
| 40 | Upstage | 49.2 | |
| 41 | Meituan | 48.62 | |
| 42 | NVIDIA | 44.4 | |
| 43 | NVIDIA | 36.81 |
Evidence key: Observed
Rows are ordered by the value BrowseComp published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.