Skip to main content
ModelScale

Instruction Following benchmark

IFBench leaderboard

Instruction Following Benchmark. Every model the catalog carries a published IFBench value for, ranked by that value.

CategoryInstruction Following
Published byBenchLM

IFBench evaluates precise instruction-following generalization on 58 challenging, verifiable out-of-domain constraints. Unlike IFEval which tests familiar constraint types, IFBench specifically measures how well models follow novel instructions they haven't been optimized for, exposing overfitting to common instruction patterns.

IFBench ranking

42 models with a published IFBench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Instruction Following Benchmark value
RankModelProviderIFBench
1MAI-Thinking-1Microsoft85
2Qwen3.8 MaxAlibaba82.8
3TMInkling-SmallThinking Machines Lab82.2
4Nemotron 3 UltraNVIDIA81.7
5Qwen3.8-Omni-FlashAlibaba81.5
6Grok 4.3xAI81.3
6Qwen3.8-Flash-NextAlibaba81.3
8DSdots3-note PreviewDots Studio80.4
9Solar Open 2Upstage80
10TMInklingThinking Machines Lab79.8
11Qwen3.8-27BAlibaba79.5
12Granite 4.2 8BIBM79.33
13Qwen3.7 MaxAlibaba79.1
13Qwen3.7 PlusAlibaba79.1
15Granite 4.2 30BIBM77.17
16Mercury 2.5Inception77
16Muse Glimmer 30BMeta77
18Gemini 3.5 FlashGoogle76.3
19STA.X K2SK Telecom75.9
20Qwen3.6 PlusAlibaba75.8
21Ling 3.0 FlashInclusionAI74.5
22Granite 4.2 3BIBM74.33
23Nemotron 3 Nano Omni 30B A3BNVIDIA74.2
24PMTernary Bonsai 2 27BPrism ML74
25Ling 3.0 Flash FP8InclusionAI73.4
26Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA72.88
27K-EXAONE 2.0LG AI Research72.6
28Agents-A1-4BInternScience69.1
29OPMiniCPM5-2BOpenBMB66.3
30Hy3 PreviewTencent63.1
31LFM2.5-2.6BLiquidAI59.17
32Claude Opus 4.5Anthropic58
33Ling 2.6 FlashInclusionAI57
34LFM2.5-8B-A1BLiquidAI56.47
35Solar Pro 3Upstage55.78
36ZYZAYA1-8BZyphra52.56
37OPMiniCPM5-1BOpenBMB46.67
38LFM2.5-230MLiquidAI38.4
39KAKanana-2 1.3B InstructKakao34.69
40KAKanana-2 3B InstructKakao33.33
41LFM2.5-VL-3BLiquidAI25.8
42LLaDA2.2-miniInclusionAI24.93

Evidence key: Observed

Rows are ordered by the value BenchLM published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards