Skip to main content
ModelScale

Instruction Following benchmark

AA-IFBench leaderboard

Artificial Analysis IFBench. Every model the catalog carries a published AA-IFBench value for, ranked by that value.

CategoryInstruction Following
MeasureConstraint satisfaction accuracy
TasksVerifiable instruction constraints
DifficultyInstruction precision

A display-only Artificial Analysis IFBench score.

AA-IFBench ranking

130 models with a published AA-IFBench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Artificial Analysis IFBench value
RankModelProviderConstraint satisfaction accuracy
1MiniMax M3MiniMax82.9
2Nemotron 3 UltraNVIDIA81.4
3Grok 4.3xAI81.3
4Qwen3.7 MaxAlibaba80.5
5MiMo-V2.5-ProXiaomi79.9
6Qwen3.7 PlusAlibaba78.0
7GPT-5.2-CodexOpenAI77.6
8Gemini 3.1 ProGoogle77.1
9Qwen 3.6 Max (preview)Alibaba76.6
10DeepSeek V4 Pro 0813DeepSeek76.5
11Gemini 3.5 FlashGoogle76.3
11GLM-5.1Z.AI76.3
13Kimi K2.6Moonshot AI76.0
14GPT-5.4 nanoOpenAI75.9
14GPT-5.5OpenAI75.9
14Muse SparkMeta75.9
17MiniMax M2.7MiniMax75.7
17Qwen3.5-122B-A10BAlibaba75.7
19Gemma 4 31BGoogle75.6
19Qwen3.5-27BAlibaba75.6
21GPT-5.2OpenAI75.4
21GPT-5.3 CodexOpenAI75.4
23Qwen3.6 PlusAlibaba75.2
24Command A+Cohere73.9
24GPT-5.4OpenAI73.9
26Gemma 4 12BGoogle73.5
27GLM-5.2Z.AI73.3
27GPT-5.4 miniOpenAI73.3
29GLM-5-TurboZ.AI73.2
30GPT-5 (high)OpenAI73.1
31GPT-5.1OpenAI72.9
32GPT-5.6 SolOpenAI72.7
33Qwen3.5-35B-A3BAlibaba72.5
34Gemma 4 26B A4BGoogle72.4
35GLM-5Z.AI72.3
36Nemotron 3 Super 100BNVIDIA71.5
37o3OpenAI71.4
38GPT-5.6 TerraOpenAI71.2
38Solar Pro 3Upstage71.2
40Nemotron 3 Nano 30BNVIDIA71.1
41GPT-5 (medium)OpenAI70.6
42Gemini 3 ProGoogle70.4
43o1OpenAI70.3
44Kimi K2.5Moonshot AI70.2
44Kimi K2.5 (Reasoning)Moonshot AI70.2
46GPT-5.1-CodexOpenAI70.0
46GPT-5.1-Codex-MaxOpenAI70.0
48GPT-OSS 120BOpenAI69.0
49MiMo-V2-ProXiaomi68.8
49Mistral Medium 3.5 128BMistral68.8
51GLM-4.7Z.AI67.9
52Qwen3.6-27BAlibaba67.6
53Step 3.7 FlashStepFun67.3
54GPT-OSS 20BOpenAI65.1
55K-ExaoneLG AI Research64.7
56Qwen3.6-35B-A3BAlibaba64.4
57Claude Fable 5Anthropic63.5
58Nemotron 3 Nano Omni 30B A3BNVIDIA63.2
59Kimi K2.7 CodeMoonshot AI63.1
60Claude Opus 4.8Anthropic62.2
61GLM-5V-TurboZ.AI61.1
62Claude Opus 4.7 (Adaptive)Anthropic58.6
63Claude Opus 4.5 ThinkingAnthropic58.0
64North Mini CodeCohere57.6
65Ling 2.6 FlashInclusionAI57.4
66Trinity-Large-PreviewArcee AI56.3
66Trinity-Large-ThinkingArcee AI56.3
68LFM2.5-8B-A1BLiquidAI55.6
69Claude 4.1 Opus ThinkingAnthropic55.4
70Gemini 3 FlashGoogle55.1
71Grok 4xAI53.7
72MiMo-V2-OmniXiaomi53.5
73Claude Opus 4.6 (Adaptive)Anthropic53.1
74Grok 4.1 Fast (Reasoning)xAI52.7
75Qwen3.5 397BAlibaba51.6
75Qwen3.5 397B (Reasoning)Alibaba51.6
77Grok 4 Fast (Reasoning)xAI50.5
78DeepSeek V3.2DeepSeek49.0
79Gemini 2.5 ProGoogle48.7
80Mistral Small 4Mistral48.2
80Mistral Small 4 (Reasoning)Mistral48.2
82FAUltravox v0.6 Llama 3.3 70BFixie AI47.1
83Claude 4 SonnetAnthropic45.4
84Claude Opus 4.6Anthropic44.6
85Gemma 4 E4BGoogle44.2
86Qwen3 MaxAlibaba44.1
87Claude Opus 4.7Anthropic43.6
88Qwen3-Omni-30B-A3B-ThinkingAlibaba43.4
89Claude Opus 4.5Anthropic43.0
89GPT-4.1OpenAI43.0
89Llama 4 MaverickMeta43.0
92DeepSeek V3.1 (Reasoning)DeepSeek41.5
92Kimi K2Moonshot AI41.5
94Grok Code Fast 1xAI41.4
95Claude Sonnet 4.6Anthropic41.2
96MiMo-V2-FlashXiaomi39.9
97DeepSeek-R1DeepSeek39.6
98Llama 4 ScoutMeta39.5
99Mistral Medium 3Mistral39.3
100Gemini 2.5 FlashGoogle39.0
100Llama 3.1 405BMeta39.0
102GPT-4.1 miniOpenAI38.3
103Nemotron Ultra 253BNVIDIA38.2
104Nova ProAmazon38.1
105Gemma 4 E2BGoogle38.0
106DeepSeek V3.1DeepSeek37.8
107GLM-4.5-AirZ.AI37.6
108GLM-4.6Z.AI36.7
109Grok 4.1 FastxAI36.5
110Mistral Large 3Mistral36.2
111Claude 3 HaikuAnthropic36.1
112DeepSeek V3DeepSeek34.8
113SASarvam 105BSarvam34.4
114GPT-4oOpenAI34.3
115Solar Pro 2Upstage33.7
116Exaone 4.0 32BLG AI Research33.5
117LFM2.5-VL-1.6B-ExtractLiquidAI33.1
118GPT-4.1 nanoOpenAI32.0
119Gemma 3 27BGoogle31.8
120Mistral Large 2Mistral31.2
120Qwen3-Omni-30B-A3B-InstructAlibaba31.2
122GPT-4o miniOpenAI31.0
123SASarvam 30BSarvam26.5
124Granite-4.0-H-1BIBM26.2
125Exaone 4.0 1.2BLG AI Research25.3
126Phi-4Microsoft23.5
127DeepSeek R1 Distill Qwen 32BDeepSeek22.9
128Granite-4.0-1BIBM20.5
129Granite-4.0-H-350MIBM17.6
130Granite-4.0-350MIBM15.9

Evidence key: ObservedLast good

Rows are ordered by the value Artificial Analysis IFBench Benchmark Leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards