Skip to main content
ModelScale

Knowledge benchmark

AA-Omniscience Accuracy leaderboard

Artificial Analysis Omniscience Accuracy. Every model the catalog carries a published AA-Omniscience Accuracy value for, ranked by that value.

CategoryKnowledge
MeasureAccuracy
TasksKnowledge questions
DifficultyBroad knowledge

A display-only Artificial Analysis knowledge metric for the proportion of correctly answered questions.

AA-Omniscience Accuracy ranking

171 models with a published AA-Omniscience Accuracy value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Artificial Analysis Omniscience Accuracy value
RankModelProviderAccuracy
1Claude Fable 5.1Anthropic67.2
2Claude Fable 5Anthropic65.4
3GPT-6 AstraOpenAI62.6
4Claude Opus 5Anthropic60.9
5GPT-5.6 SolOpenAI59.4
6GPT-5.5OpenAI58.0
7Gemini 3 ProGoogle55.8
8Gemini 3.7 FlashGoogle55.3
9Gemini 3.1 ProGoogle54.9
10Gemini 3.8 FlashGoogle54.6
11GPT-5.3 CodexOpenAI52.9
12Muse Spark 1.1Meta52.1
13Gemini 3.5 FlashGoogle51.9
14Grok 4.5xAI51.6
15GPT-5.4OpenAI50.8
16Gemini 3.6 FlashGoogle50.0
17Muse SparkMeta49.6
18DeepSeek V4 Pro 0813DeepSeek49.1
19Claude Opus 4.7 (Adaptive)Anthropic48.9
20Claude Opus 4.8Anthropic48.8
21Grok 4.6xAI48.2
22Kimi K3Moonshot AI47.6
23Claude Opus 4.6 (Adaptive)Anthropic47.0
24GPT-5.6 TerraOpenAI46.8
25Claude Opus 4.5 ThinkingAnthropic46.6
26DeepSeek V4.1 FlashDeepSeek46.4
27Claude Opus 4.6Anthropic45.8
27Gemini 3 FlashGoogle45.8
29Muse Spark 1.2Meta45.4
30Claude Opus 4.7Anthropic44.7
31GPT-5.2OpenAI44.3
32Muse Spark 1.3Meta43.6
33GPT-5.6 LunaOpenAI42.7
34TMInklingThinking Machines Lab41.6
35GPT-5.2-CodexOpenAI41.1
36Claude Opus 4.5Anthropic40.9
37Grok 4xAI40.5
38DeepSeek V4 Flash 0731DeepSeek40.4
39GPT-5 (high)OpenAI40.3
40Claude Sonnet 5Anthropic40.1
41GPT-5.1-CodexOpenAI39.9
41GPT-5.1-Codex-MaxOpenAI39.9
43Kimi K2.7 CodeMoonshot AI39.6
44GPT-5 (medium)OpenAI39.5
45Gemini 2.5 ProGoogle39.1
46Claude Sonnet 4.6Anthropic38.6
46o3OpenAI38.6
48Qwen 3.6 Max (preview)Alibaba37.9
49GPT-5.1OpenAI37.7
50GPT-5.4 miniOpenAI37.5
51Kimi K2.5Moonshot AI35.2
51Kimi K2.5 (Reasoning)Moonshot AI35.2
53Grok 4.3xAI34.6
54o1OpenAI34.5
55GLM-5.3Z.AI33.9
56TMInkling-SmallThinking Machines Lab33.2
57Kimi K2.6Moonshot AI32.6
58Hy3Tencent32.0
59APApodex 1.1Apodex31.7
59APApodex 1.1 MiniApodex31.7
59Qwen3.8 Max PreviewAlibaba31.7
62Hy3 PreviewTencent31.5
63Qwen3.7 MaxAlibaba31.1
64DeepSeek-R1DeepSeek30.5
65Gemini 3.5 Flash-LiteGoogle29.5
66GLM-4.7Z.AI29.3
66GLM-5V-TurboZ.AI29.3
68DeepSeek V3.1 (Reasoning)DeepSeek29.0
69GLM-5-TurboZ.AI28.4
70GPT-4.1OpenAI27.8
71Kimi K2Moonshot AI27.4
72Muse Glimmer 30BMeta27.0
73MiniMax M2.7MiniMax26.8
74MiMo-V2-ProXiaomi26.6
75Qwen3.6 PlusAlibaba26.4
76GLM-5Z.AI26.3
77Gemini 2.5 FlashGoogle26.1
78Step 3.7 FlashStepFun25.8
79GPT-5.4 nanoOpenAI25.7
80DeepSeek V3DeepSeek25.5
81Grok 4.1 Fast (Reasoning)xAI25.1
82Mistral Large 3Mistral25.0
83Llama 4 MaverickMeta24.9
84Mistral Medium 3.5 128BMistral24.7
85Qwen3.5 397BAlibaba24.5
85Qwen3.5 397B (Reasoning)Alibaba24.5
85Qwen3.8-Flash-NextAlibaba24.5
88Qwen3 MaxAlibaba24.4
88Qwen3.5-122B-A10BAlibaba24.4
90GLM-5.2Z.AI24.3
90Nemotron 3 Super 100BNVIDIA24.3
92DeepSeek V3.2DeepSeek24.0
93GLM-5.1Z.AI23.7
94Grok Code Fast 1xAI23.5
95Llama 3.1 405BMeta23.2
96DeepSeek V3.1DeepSeek23.1
97Grok 4 Fast (Reasoning)xAI22.8
98Claude 4 SonnetAnthropic22.7
99Qwen3.7 PlusAlibaba22.5
99Trinity-Large-PreviewArcee AI22.5
99Trinity-Large-ThinkingArcee AI22.5
102MiMo-V2.5-ProXiaomi22.4
103Mercury 2.5Inception22.0
104GPT-OSS 120BOpenAI21.8
105Mistral Small 4Mistral21.7
105Mistral Small 4 (Reasoning)Mistral21.7
107Nemotron 3 UltraNVIDIA21.6
108GLM-4.6Z.AI21.4
109Qwen3.5-27BAlibaba20.7
110GPT-4.1 miniOpenAI20.3
111Nemotron Ultra 253BNVIDIA20.1
111Qwen3.5-35B-A3BAlibaba20.1
113Gemma 4 31BGoogle20.0
114GPT-4oOpenAI19.9
114Mistral Large 2Mistral19.9
116Qwen3.6-27BAlibaba19.6
117MiMo-V2-OmniXiaomi19.3
118Gemma 4 26B A4BGoogle19.1
119FAUltravox v0.6 Llama 3.3 70BFixie AI19.0
120North Mini CodeCohere18.9
120Solar Pro 4Upstage18.9
122Qwen3.6-35B-A3BAlibaba18.8
123STA.X K2SK Telecom18.6
124Solar Pro 3Upstage18.5
125Mistral Medium 3Mistral18.3
126Ling 3.0 FlashInclusionAI18.2
126Ling 3.0 Flash FP8InclusionAI18.2
128Claude 3 HaikuAnthropic17.6
128SASarvam 105BSarvam17.6
130Nemotron 3 Nano 30BNVIDIA17.3
131Grok 4.1 FastxAI17.2
132Nova ProAmazon16.9
133MiniMax M3MiniMax16.7
134K-ExaoneLG AI Research16.4
135GLM-4.5-AirZ.AI16.3
136Solar Pro 2Upstage16.1
137GPT-OSS 20BOpenAI16.0
138Gemma 4 12BGoogle15.6
138Ling 2.6 FlashInclusionAI15.6
138MiMo-V2-FlashXiaomi15.6
138Qwen3.8-27BAlibaba15.6
142MCQuasar 438BMultiverse Computing15.5
143Llama 4 ScoutMeta15.2
143Nemotron 3 Nano Omni 30B A3BNVIDIA15.2
145Qwen3-Omni-30B-A3B-ThinkingAlibaba14.6
146Ling 3.0 Flash VLInclusionAI14.4
146Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA14.4
148Qwen3-Omni-30B-A3B-InstructAlibaba14.3
149Phi-4Microsoft14.1
150GPT-4.1 nanoOpenAI13.7
151K-EXAONE 2.0LG AI Research13.1
152Gemma 3 27BGoogle13.0
153SASarvam 30BSarvam12.6
154Granite 4.2 8BIBM11.2
155CECeleris-1Celeris11.0
156Exaone 4.0 32BLG AI Research10.6
157Granite 4.2 30BIBM10.1
158LFM2.5-8B-A1BLiquidAI9.4
159Granite 4.2 3BIBM9.2
160Command A+Cohere8.9
161Gemma 4 E4BGoogle8.6
162Ling 3.0 TinyInclusionAI8.5
163OPMiniCPM5-2BOpenBMB8.4
164Gemma 4 E2BGoogle6.6
165Granite-4.0-1BIBM6.2
166LFM2.5-VL-1.6B-ExtractLiquidAI5.8
167Granite-4.0-H-1BIBM5.2
168Exaone 4.0 1.2BLG AI Research5.0
169LFM2.5-2.6BLiquidAI4.4
170Granite-4.0-350MIBM3.9
171Granite-4.0-H-350MIBM3.8

Evidence key: ObservedLast good

Rows are ordered by the value Artificial Analysis model benchmarks published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Knowledge capability leaderboard