Skip to main content
ModelScale

Reasoning benchmark

LongBench v2 leaderboard

Every model the catalog carries a published LongBench v2 value for, ranked by that value.

CategoryReasoning
MeasureExtended-context retrieval and reasoning
TasksLong-context tasks
DifficultyHard long-context
Published byLongBench v2

A long-context benchmark that measures whether models can actually use extended context windows for reasoning and retrieval.

LongBench v2 ranking

14 models with a published LongBench v2 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published LongBench v2 value
RankModelProviderExtended-context retrieval and reasoning
1Qwen3.8 MaxAlibaba66.3
2Claude Opus 4.5Anthropic64.4
3Qwen3.5 397BAlibaba63.2
4Qwen3.6 PlusAlibaba62
5Nemotron 3 UltraNVIDIA61.9
6Kimi K2.5Moonshot AI61
7GLM-5Z.AI60.8
8Qwen3.5-27BAlibaba60.6
9Agents-A1InternScience60.2
9Qwen3.5-122B-A10BAlibaba60.2
11Qwen3.5-35B-A3BAlibaba59
12Agents-A1-4BInternScience52.1
13OPMiniCPM5-2BOpenBMB43.7
14LLaDA2.2-miniInclusionAI34.99

Evidence key: Observed

Rows are ordered by the value LongBench v2 published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Reasoning capability leaderboard