Skip to main content
ModelScale

Agentic benchmark

WideResearch leaderboard

Every model the catalog carries a published WideResearch value for, ranked by that value.

CategoryAgentic
MeasureMulti-source research evaluation
TasksOpen-ended research tasks
DifficultyBroad research-agent workflows

A broad research-agent benchmark for open-ended information gathering, synthesis, and answer construction across wide search spaces.

WideResearch ranking

15 models with a published WideResearch value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published WideResearch value
RankModelProviderMulti-source research evaluation
1Hy4 previewTencent83.9
2Atria Dawn PreviewShanghai Artificial Intelligence Laboratory81.9
2Qwen3.8 MaxAlibaba81.9
4Kimi K2.6Moonshot AI80.8
4OAOrnith-1.5-397BOrnith AI80.8
6DSdots3-note PreviewDots Studio78.9
7Claude Opus 4.5Anthropic76.4
8Qwen3.6 PlusAlibaba74.3
9Qwen3.5 397BAlibaba74.0
10Ling 3.0 FlashInclusionAI73.6
11Kimi K2.5Moonshot AI72.7
12GLM-5Z.AI69.8
13OAOrnith-1.5-35B-A3BOrnith AI67.8
14Qwen3.6-35B-A3BAlibaba60.1
15OAOrnith-1.5-9BOrnith AI59.5

Evidence key: Observed

Rows are ordered by the value Qwen3.6 launch benchmarks published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard