Skip to main content
ModelScale

Agentic benchmark

ResearchClawBench leaderboard

Every model the catalog carries a published ResearchClawBench value for, ranked by that value.

CategoryAgentic
MeasureEnd-to-end autonomous research evaluation with RADS scoring
Tasks40 tasks across 10 scientific domains
DifficultyScientific research re-discovery

An end-to-end autonomous scientific research benchmark with 40 tasks across 10 scientific domains, where agents receive related literature and raw data, then attempt to rediscover the hidden target paper.

ResearchClawBench ranking

19 models with a published ResearchClawBench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published ResearchClawBench value
RankModelProviderEnd-to-end autonomous research evaluation with RADS scoring
1Claude Opus 4.8Anthropic21.1
2Claude Opus 4.7Anthropic20.7
2GLM-5.2Z.AI20.7
4Claude Opus 4.6Anthropic19.9
5MiniMax M3MiniMax19.8
6Qwen3.7 MaxAlibaba18.7
7GLM-5.1Z.AI18.2
8Gemini 3.5 FlashGoogle18.0
8Kimi K2.6Moonshot AI18.0
8Qwen3.6 PlusAlibaba18.0
11GPT-5.5OpenAI17.0
12MiMo-V2.5Xiaomi16.9
13GPT-5.4OpenAI15.3
13MiMo-V2-ProXiaomi15.3
15Qwen3.5 397BAlibaba14.2
16Kimi K2.5Moonshot AI14.0
17Grok 4.1xAI13.5
18Gemini 3.1 ProGoogle13.3
19Grok 4.3xAI12.4

Evidence key: Observed

Rows are ordered by the value ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard