Agentic benchmark
ResearchClawBench leaderboard
Every model the catalog carries a published ResearchClawBench value for, ranked by that value.
An end-to-end autonomous scientific research benchmark with 40 tasks across 10 scientific domains, where agents receive related literature and raw data, then attempt to rediscover the hidden target paper.
ResearchClawBench ranking
19 models with a published ResearchClawBench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | End-to-end autonomous research evaluation with RADS scoring |
|---|---|---|---|
| 1 | Anthropic | 21.1 | |
| 2 | Anthropic | 20.7 | |
| 2 | Z.AI | 20.7 | |
| 4 | Anthropic | 19.9 | |
| 5 | MiniMax | 19.8 | |
| 6 | Alibaba | 18.7 | |
| 7 | Z.AI | 18.2 | |
| 8 | 18.0 | ||
| 8 | Moonshot AI | 18.0 | |
| 8 | Alibaba | 18.0 | |
| 11 | OpenAI | 17.0 | |
| 12 | Xiaomi | 16.9 | |
| 13 | OpenAI | 15.3 | |
| 13 | Xiaomi | 15.3 | |
| 15 | Alibaba | 14.2 | |
| 16 | Moonshot AI | 14.0 | |
| 17 | xAI | 13.5 | |
| 18 | 13.3 | ||
| 19 | xAI | 12.4 |
Evidence key: Observed
Rows are ordered by the value ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.