Skip to main content
ModelScale

Agentic benchmark

DeepSearchQA leaderboard

Every model the catalog carries a published DeepSearchQA value for, ranked by that value.

CategoryAgentic
MeasureSearch / open / find browser-agent evaluation
TasksAgentic browsing and list-answer questions
DifficultyAgentic web research

An agentic browsing benchmark where models search the web, gather evidence, and answer list-style questions using browser tools.

DeepSearchQA ranking

18 models with a published DeepSearchQA value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published DeepSearchQA value
RankModelProviderSearch / open / find browser-agent evaluation
1Atria Dawn PreviewShanghai Artificial Intelligence Laboratory96.0
2Claude Opus 5Anthropic95.0
2Kimi K3Moonshot AI95.0
4Claude Opus 4.8Anthropic93.1
5Step 3.7 FlashStepFun92.8
6Kimi K2.6Moonshot AI92.5
7APApodex 1.1Apodex92.4
8DSdots3-note PreviewDots Studio92.1
9Muse Spark 1.3Meta89.4
10Muse Spark 1.1Meta84.9
11Kimi K2.5Moonshot AI77.1
12Muse SparkMeta74.8
13Muse Glimmer 30BMeta74.6
14Claude Opus 4.6Anthropic73.7
15GPT-5.4OpenAI73.6
16Gemini 3.1 ProGoogle69.7
17Grok 4.20xAI62.8
18Mercury 2.5Inception34.0

Evidence key: Observed

Rows are ordered by the value Muse Spark Eval Methodology published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard