Skip to main content
ModelScale

Agentic benchmark

BrowseComp leaderboard

Every model the catalog carries a published BrowseComp value for, ranked by that value.

CategoryAgentic
MeasureWeb search and evidence synthesis
TasksResearch questions requiring browsing
DifficultyHard web research
Published byBrowseComp

A benchmark for web-browsing agents that must search, inspect sources, gather evidence, and return the correct answer to research-oriented questions.

BrowseComp ranking

43 models with a published BrowseComp value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published BrowseComp value
RankModelProviderWeb search and evidence synthesis
1Atria Dawn PreviewShanghai Artificial Intelligence Laboratory92.5
2GPT-5.6 SolOpenAI92.2
3GPT-6 AstraOpenAI91.5
4Kimi K3Moonshot AI91.2
5Claude Opus 5Anthropic90.8
6GPT-5.5 ProOpenAI90.1
7GPT-5.4 ProOpenAI89.3
8Claude Mythos 5Anthropic88
9GPT-5.6 TerraOpenAI87.5
10OAOrnith-1.5-397BOrnith AI86.6
11Claude Sonnet 5Anthropic84.7
12GPT-5.5OpenAI84.4
13Claude Opus 4.8Anthropic84.3
14Claude Opus 4.6Anthropic83.7
15MiniMax M3MiniMax83.52
16DeepSeek V4 Pro 0813DeepSeek83.4
17DSdots3-note PreviewDots Studio83.3
17GPT-5.6 LunaOpenAI83.3
19Kimi K2.6Moonshot AI83.2
20GPT-5.4OpenAI82.7
21Claude Opus 4.7 (Adaptive)Anthropic79.3
22TMInkling-SmallThinking Machines Lab77.4
23TMInklingThinking Machines Lab77.1
24Step 3.7 FlashStepFun75.82
25Agents-A1InternScience75.51
26DeepSeek V4 Flash 0731DeepSeek73.2
27Ling 3.0 FlashInclusionAI72.2
28GLM-5.1Z.AI68
29OAOrnith-1.5-35B-A3BOrnith AI67.6
30Agents-A1-4BInternScience66.8
31GPT-5.2OpenAI65.8
32Qwen3.5-122B-A10BAlibaba63.8
33Qwen3.5 397BAlibaba62
34Qwen3.5-27BAlibaba61
34Qwen3.5-35B-A3BAlibaba61
36Kimi K2.5Moonshot AI60.6
36Kimi K2.5 (Reasoning)Moonshot AI60.6
38OAOrnith-1.5-9BOrnith AI56.4
39GLM-4.7Z.AI52
40Solar Pro 4Upstage49.2
41LongCat-Flash-Lite-SparseMeituan48.62
42Nemotron 3 UltraNVIDIA44.4
43Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA36.81

Evidence key: Observed

Rows are ordered by the value BrowseComp published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard