Skip to main content
ModelScale

Reasoning benchmark

AA-LCR leaderboard

Artificial Analysis Long Context Reasoning. Every model the catalog carries a published AA-LCR value for, ranked by that value.

CategoryReasoning
MeasureAccuracy
TasksLong-context reasoning tasks
DifficultyLong-context reasoning

A display-only Artificial Analysis long-context reasoning evaluation.

AA-LCR ranking

172 models with a published AA-LCR value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Artificial Analysis Long Context Reasoning value
RankModelProviderAccuracy
1Kimi K3Moonshot AI88.7
2Claude Fable 5.1Anthropic85.3
3GPT-5.5OpenAI84.3
4DeepSeek V4.1 FlashDeepSeek84.0
4GPT-5.6 SolOpenAI84.0
6GPT-5.6 LunaOpenAI83.7
7GPT-5.3 CodexOpenAI83.3
7Muse Glimmer 30BMeta83.3
9GPT-5.6 TerraOpenAI83.0
9MiniMax M3MiniMax83.0
9Muse Spark 1.3Meta83.0
12GPT-5.2OpenAI82.7
13Claude Fable 5Anthropic82.3
13GPT-5.2-CodexOpenAI82.3
15Claude Sonnet 5Anthropic82.0
15Gemini 3.1 ProGoogle82.0
15GPT-5.4OpenAI82.0
15Qwen3.8-27BAlibaba82.0
19Gemini 3.7 FlashGoogle81.7
20Gemini 3.8 FlashGoogle81.3
21Kimi K2.6Moonshot AI81.0
22GPT-6 AstraOpenAI80.7
22Qwen 3.6 Max (preview)Alibaba80.7
24DeepSeek V4 Pro 0813DeepSeek80.3
24Grok 4.6xAI80.3
24Qwen3.8 Max PreviewAlibaba80.3
27Gemini 3.6 FlashGoogle80.0
27GPT-5.1OpenAI80.0
29DeepSeek V4 Flash 0731DeepSeek79.7
29GLM-5.3Z.AI79.7
29MiMo-V2.5-ProXiaomi79.7
29Qwen3.8-Flash-NextAlibaba79.7
33APApodex 1.1Apodex79.3
33APApodex 1.1 MiniApodex79.3
33Claude Opus 5Anthropic79.3
33Grok 4.5xAI79.3
33Kimi K2.7 CodeMoonshot AI79.3
38Hy3Tencent79.0
38Muse Spark 1.2Meta79.0
38Qwen3.7 MaxAlibaba79.0
41Claude Opus 4.7 (Adaptive)Anthropic78.7
42GLM-5.2Z.AI78.3
42Ling 3.0 Flash VLInclusionAI78.3
42MiniMax M2.7MiniMax78.3
42Qwen3.6 PlusAlibaba78.3
46GPT-5 (high)OpenAI78.2
47Claude Opus 4.6 (Adaptive)Anthropic78.0
47Kimi K2.5Moonshot AI78.0
47Kimi K2.5 (Reasoning)Moonshot AI78.0
47Muse SparkMeta78.0
51Claude Opus 4.8Anthropic77.7
51Muse Spark 1.1Meta77.7
51Qwen3.5-27BAlibaba77.7
54Claude Opus 4.5 ThinkingAnthropic77.3
54TMInklingThinking Machines Lab77.3
54Qwen3.6-27BAlibaba77.3
57GPT-5.4 miniOpenAI77.0
57PMTernary Bonsai 2 27BPrism ML77.0
59GPT-5.4 nanoOpenAI76.7
60MCQuasar 438BMultiverse Computing76.3
60Qwen3.5-122B-A10BAlibaba76.3
62Claude 4.1 Opus ThinkingAnthropic76.0
62Gemini 3 ProGoogle76.0
62Gemini 3.5 Flash-LiteGoogle76.0
62GPT-5 (medium)OpenAI76.0
66Claude Opus 4.7Anthropic75.7
66GLM-5Z.AI75.7
66TMInkling-SmallThinking Machines Lab75.7
69MiMo-V2-OmniXiaomi75.0
70o3OpenAI74.7
71Grok 4.1 Fast (Reasoning)xAI74.0
72GLM-5.1Z.AI73.7
72Grok 4 Fast (Reasoning)xAI73.7
72Step 3.7 FlashStepFun73.7
75Ling 3.0 FlashInclusionAI73.0
75Ling 3.0 Flash FP8InclusionAI73.0
75Qwen3.7 PlusAlibaba73.0
78Qwen3.5-35B-A3BAlibaba72.0
79GLM-5-TurboZ.AI71.7
79Qwen3.6-35B-A3BAlibaba71.7
81GLM-4.7Z.AI71.0
81Solar Pro 4Upstage71.0
83Claude Opus 4.5Anthropic70.7
84GLM-5V-TurboZ.AI70.3
85Gemma 4 31BGoogle69.7
86Gemini 3.5 FlashGoogle69.3
86GPT-5.1-CodexOpenAI69.3
86GPT-5.1-Codex-MaxOpenAI69.3
86Mistral Medium 3.5 128BMistral69.3
90Gemini 2.5 ProGoogle69.0
91Claude Sonnet 4.6Anthropic68.3
91GPT-4.1OpenAI68.3
91MiMo-V2-ProXiaomi68.3
94Grok 4xAI68.0
94Mercury 2.5Inception68.0
96Claude Opus 4.6Anthropic67.0
96Nemotron 3 UltraNVIDIA67.0
98Hy3 PreviewTencent66.7
99STA.X K2SK Telecom66.0
100Gemma 4 26B A4BGoogle65.7
100Nemotron 3 Super 100BNVIDIA65.7
102o1OpenAI65.0
103Grok 4.3xAI64.3
103Qwen3.5 397BAlibaba64.3
103Qwen3.5 397B (Reasoning)Alibaba64.3
106Gemma 4 12BGoogle63.7
107Solar Open 2Upstage62.3
108K-ExaoneLG AI Research61.3
109Ling 3.0 TinyInclusionAI60.3
110OPMiniCPM5-2BOpenBMB59.0
111DeepSeek V3.1 (Reasoning)DeepSeek56.7
112K-EXAONE 2.0LG AI Research56.2
113DeepSeek-R1DeepSeek55.7
114Gemini 3 FlashGoogle55.3
115Grok Code Fast 1xAI53.0
115Kimi K2Moonshot AI53.0
117Command A+Cohere52.7
118GPT-OSS 120BOpenAI52.0
119Llama 4 MaverickMeta50.0
119Qwen3 MaxAlibaba50.0
121Gemini 2.5 FlashGoogle49.9
122Mistral Small 4Mistral49.7
122Mistral Small 4 (Reasoning)Mistral49.7
124GPT-4oOpenAI49.3
125Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA49.2
126Granite 4.2 30BIBM49.0
127DeepSeek V3.1DeepSeek47.0
128GLM-4.5-AirZ.AI46.7
129DeepSeek V3.2DeepSeek45.7
130Granite 4.2 8BIBM45.0
131Claude 4 SonnetAnthropic44.0
131GPT-4.1 miniOpenAI44.0
133Nemotron 3 Nano Omni 30B A3BNVIDIA39.7
134CECeleris-1Celeris38.0
134Nemotron 3 Nano 30BNVIDIA38.0
134Trinity-Large-PreviewArcee AI38.0
134Trinity-Large-ThinkingArcee AI38.0
138North Mini CodeCohere37.3
139MiMo-V2-FlashXiaomi36.3
140Mistral Large 3Mistral36.0
141GPT-OSS 20BOpenAI34.7
142Solar Pro 3Upstage32.3
143Gemma 4 E4BGoogle32.0
144Grok 4.1 FastxAI31.3
144Ling 2.6 FlashInclusionAI31.3
144Mistral Medium 3Mistral31.3
147DeepSeek V3DeepSeek29.3
148Claude 3 HaikuAnthropic27.7
148Llama 4 ScoutMeta27.7
150GLM-4.6Z.AI26.3
151Llama 3.1 405BMeta25.3
152Granite 4.2 3BIBM24.3
153Nova ProAmazon21.0
154GPT-4.1 nanoOpenAI20.3
155Gemma 4 E2BGoogle16.3
156FAUltravox v0.6 Llama 3.3 70BFixie AI15.7
157DeepSeek R1 Distill Qwen 32BDeepSeek8.7
158Gemma 3 27BGoogle7.3
159Granite-4.0-H-1BIBM7.0
160Granite-4.0-1BIBM6.0
161LFM2.5-2.6BLiquidAI5.7
162Exaone 4.0 1.2BLG AI Research0.0
162Granite-4.0-350MIBM0.0
162Granite-4.0-H-350MIBM0.0
162LFM2.5-8B-A1BLiquidAI0.0
162LFM2.5-VL-1.6B-ExtractLiquidAI0.0
162Phi-4Microsoft0.0
162Qwen3-Omni-30B-A3B-InstructAlibaba0.0
162Qwen3-Omni-30B-A3B-ThinkingAlibaba0.0
162SASarvam 105BSarvam0.0
162SASarvam 30BSarvam0.0
162Solar Pro 2Upstage0.0

Evidence key: ObservedLast good

Rows are ordered by the value Artificial Analysis model benchmarks published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Reasoning capability leaderboard