Skip to main content
ModelScale

Knowledge benchmark

AA-Omniscience Hallucination Rate leaderboard

Artificial Analysis Omniscience Hallucination Rate. Every model the catalog carries a published AA-Omniscience Hallucination Rate value for, ranked by that value.

CategoryKnowledge
MeasureHallucination rate
TasksKnowledge questions
DifficultyFactuality

A display-only Artificial Analysis factuality metric for the rate of incorrect answers among non-correct responses.

AA-Omniscience Hallucination Rate ranking

171 models with a published AA-Omniscience Hallucination Rate value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Artificial Analysis Omniscience Hallucination Rate value
RankModelProviderHallucination rate
1Qwen3-Omni-30B-A3B-InstructAlibaba97.6
2Ling 2.6 FlashInclusionAI96.7
3DeepSeek V4.1 FlashDeepSeek96.5
4SASarvam 30BSarvam96.3
5LFM2.5-VL-1.6B-ExtractLiquidAI95.7
6DeepSeek V4 Pro 0813DeepSeek94.1
6GPT-OSS 20BOpenAI94.1
8Granite-4.0-1BIBM93.5
9SASarvam 105BSarvam93.4
10DeepSeek V3.2DeepSeek93.3
10GPT-4.1OpenAI93.3
12Gemini 2.5 FlashGoogle93.0
12GLM-4.7Z.AI93.0
12Solar Pro 2Upstage93.0
15GLM-4.5-AirZ.AI92.9
16CECeleris-1Celeris92.8
17GPT-4.1 miniOpenAI92.7
18GPT-5.6 LunaOpenAI92.6
19Gemini 3 FlashGoogle92.4
20GPT-5.6 SolOpenAI92.2
21Gemma 3 27BGoogle92.1
22DeepSeek V4 Flash 0731DeepSeek91.7
22Exaone 4.0 1.2BLG AI Research91.7
22GPT-5.4OpenAI91.7
25Gemini 3 ProGoogle91.5
26Gemini 2.5 ProGoogle90.9
27GPT-OSS 120BOpenAI90.8
28GPT-5.4 miniOpenAI90.2
28FAUltravox v0.6 Llama 3.3 70BFixie AI90.2
30DeepSeek V3DeepSeek90.0
31Qwen3 MaxAlibaba89.9
32GPT-5.3 CodexOpenAI89.2
33GPT-5.5OpenAI89.0
33Qwen3-Omni-30B-A3B-ThinkingAlibaba89.0
35Llama 4 MaverickMeta88.9
36K-ExaoneLG AI Research88.8
37Solar Pro 3Upstage88.2
38Granite-4.0-H-350MIBM88.1
38o3OpenAI88.1
40GPT-5.6 TerraOpenAI87.9
41Qwen3.5-122B-A10BAlibaba87.1
42Nemotron 3 Super 100BNVIDIA87.0
43Gemma 4 26B A4BGoogle86.4
44Mistral Large 3Mistral86.0
45Trinity-Large-PreviewArcee AI85.9
45Trinity-Large-ThinkingArcee AI85.9
47DeepSeek V3.1DeepSeek85.7
47Nemotron 3 Nano Omni 30B A3BNVIDIA85.7
49Qwen3.5-35B-A3BAlibaba85.4
50Gemma 4 31BGoogle85.0
50Step 3.7 FlashStepFun85.0
52Muse SparkMeta84.2
53DeepSeek-R1DeepSeek83.4
54Nemotron 3 Nano 30BNVIDIA83.3
55GPT-5 (medium)OpenAI83.2
55North Mini CodeCohere83.2
57Qwen3.5 397BAlibaba82.7
57Qwen3.5 397B (Reasoning)Alibaba82.7
59GPT-4.1 nanoOpenAI82.6
60DeepSeek V3.1 (Reasoning)DeepSeek82.5
61Kimi K2.7 CodeMoonshot AI82.4
62Grok 4.1 FastxAI82.3
63GPT-5 (high)OpenAI82.2
64Exaone 4.0 32BLG AI Research82.1
65Muse Glimmer 30BMeta81.9
66Granite-4.0-H-1BIBM81.7
67Mistral Medium 3.5 128BMistral81.6
68Qwen3.5-27BAlibaba81.5
69GPT-5.2OpenAI81.2
69Nemotron Ultra 253BNVIDIA81.2
69Phi-4Microsoft81.2
72Gemma 4 12BGoogle81.0
73Claude 3 HaikuAnthropic80.5
74Claude Opus 4.6Anthropic80.1
75Llama 4 ScoutMeta79.4
76Grok Code Fast 1xAI79.3
77APApodex 1.1Apodex78.4
77APApodex 1.1 MiniApodex78.4
79Nova ProAmazon77.7
80GPT-5.1-CodexOpenAI77.2
80GPT-5.1-Codex-MaxOpenAI77.2
82Kimi K2Moonshot AI76.6
83Claude Opus 4.5Anthropic76.2
84Granite-4.0-350MIBM76.1
85MiMo-V2-FlashXiaomi76.0
86GPT-5.4 nanoOpenAI74.2
87Hy3Tencent74.1
88GPT-5.2-CodexOpenAI73.4
88Grok 4.1 Fast (Reasoning)xAI73.4
90Hy3 PreviewTencent73.0
91Claude Fable 5.1Anthropic72.6
92o1OpenAI69.6
93GLM-5V-TurboZ.AI68.8
94Claude Sonnet 4.6Anthropic68.5
95Grok 4 Fast (Reasoning)xAI68.3
96TMInklingThinking Machines Lab67.7
96Mistral Large 2Mistral67.7
98GLM-4.6Z.AI67.6
99Mercury 2.5Inception67.0
100Mistral Small 4Mistral66.5
100Mistral Small 4 (Reasoning)Mistral66.5
102Kimi K2.5Moonshot AI65.7
102Kimi K2.5 (Reasoning)Moonshot AI65.7
104Gemini 3.7 FlashGoogle64.5
104Grok 4xAI64.5
106Claude Fable 5Anthropic63.6
107TMInkling-SmallThinking Machines Lab63.0
108Claude Opus 4.6 (Adaptive)Anthropic62.8
109GLM-5-TurboZ.AI62.6
110Claude Opus 4.5 ThinkingAnthropic61.0
111Mistral Medium 3Mistral60.9
112Claude Opus 5Anthropic60.8
113Gemini 3.5 FlashGoogle60.7
114Gemini 3.6 FlashGoogle55.6
115Gemini 3.8 FlashGoogle55.2
116Claude Opus 4.7Anthropic54.1
116Grok 4.5xAI54.1
118Kimi K3Moonshot AI53.2
119Llama 3.1 405BMeta52.4
120GPT-5.1OpenAI51.9
121GPT-6 AstraOpenAI51.3
122Gemini 3.1 ProGoogle50.9
123Qwen3.6-35B-A3BAlibaba50.5
124Muse Spark 1.1Meta50.0
125Qwen3.6-27BAlibaba49.3
126MiMo-V2-OmniXiaomi48.9
127LFM2.5-8B-A1BLiquidAI46.9
128Qwen 3.6 Max (preview)Alibaba46.2
129Qwen3.8-Flash-NextAlibaba45.3
130Ling 3.0 FlashInclusionAI44.1
130Ling 3.0 Flash FP8InclusionAI44.1
132Claude Opus 4.7 (Adaptive)Anthropic42.3
133Claude 4 SonnetAnthropic41.0
134Kimi K2.6Moonshot AI40.5
135Claude Sonnet 5Anthropic39.4
136Claude Opus 4.8Anthropic39.3
137GPT-4oOpenAI37.9
138Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA37.6
139MiniMax M2.7MiniMax35.6
140GLM-5Z.AI35.3
141Qwen3.6 PlusAlibaba34.6
142Gemini 3.5 Flash-LiteGoogle34.4
143Grok 4.6xAI34.3
144Muse Spark 1.2Meta33.3
145STA.X K2SK Telecom33.0
146Muse Spark 1.3Meta32.9
147Gemma 4 E2BGoogle32.4
148Granite 4.2 8BIBM32.0
149Gemma 4 E4BGoogle30.9
150Ling 3.0 TinyInclusionAI30.5
151Qwen3.8-27BAlibaba30.3
152MiMo-V2-ProXiaomi30.0
153GLM-5.1Z.AI29.9
154Nemotron 3 UltraNVIDIA29.7
155GLM-5.3Z.AI29.6
156Qwen3.8 Max PreviewAlibaba28.8
157Qwen3.7 PlusAlibaba27.7
158GLM-5.2Z.AI26.3
158Granite 4.2 3BIBM26.3
160Granite 4.2 30BIBM25.6
160Qwen3.7 MaxAlibaba25.6
162Grok 4.3xAI25.0
163MiMo-V2.5-ProXiaomi24.7
164Solar Pro 4Upstage24.4
165K-EXAONE 2.0LG AI Research22.6
166Ling 3.0 Flash VLInclusionAI22.0
167OPMiniCPM5-2BOpenBMB21.9
168MCQuasar 438BMultiverse Computing21.4
169MiniMax M3MiniMax18.4
170LFM2.5-2.6BLiquidAI16.0
171Command A+Cohere14.2

Evidence key: ObservedLast good

Rows are ordered by the value Artificial Analysis model benchmarks published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Knowledge capability leaderboard