Instruction Following benchmark
IFBench leaderboard
Instruction Following Benchmark. Every model the catalog carries a published IFBench value for, ranked by that value.
IFBench evaluates precise instruction-following generalization on 58 challenging, verifiable out-of-domain constraints. Unlike IFEval which tests familiar constraint types, IFBench specifically measures how well models follow novel instructions they haven't been optimized for, exposing overfitting to common instruction patterns.
IFBench ranking
42 models with a published IFBench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | IFBench |
|---|---|---|---|
| 1 | Microsoft | 85 | |
| 2 | Alibaba | 82.8 | |
| 3 | TMInkling-Small | Thinking Machines Lab | 82.2 |
| 4 | NVIDIA | 81.7 | |
| 5 | Alibaba | 81.5 | |
| 6 | xAI | 81.3 | |
| 6 | Alibaba | 81.3 | |
| 8 | DSdots3-note Preview | Dots Studio | 80.4 |
| 9 | Upstage | 80 | |
| 10 | TMInkling | Thinking Machines Lab | 79.8 |
| 11 | Alibaba | 79.5 | |
| 12 | IBM | 79.33 | |
| 13 | Alibaba | 79.1 | |
| 13 | Alibaba | 79.1 | |
| 15 | IBM | 77.17 | |
| 16 | Inception | 77 | |
| 16 | Meta | 77 | |
| 18 | 76.3 | ||
| 19 | STA.X K2 | SK Telecom | 75.9 |
| 20 | Alibaba | 75.8 | |
| 21 | InclusionAI | 74.5 | |
| 22 | IBM | 74.33 | |
| 23 | NVIDIA | 74.2 | |
| 24 | PMTernary Bonsai 2 27B | Prism ML | 74 |
| 25 | InclusionAI | 73.4 | |
| 26 | NVIDIA | 72.88 | |
| 27 | LG AI Research | 72.6 | |
| 28 | InternScience | 69.1 | |
| 29 | OPMiniCPM5-2B | OpenBMB | 66.3 |
| 30 | Tencent | 63.1 | |
| 31 | LiquidAI | 59.17 | |
| 32 | Anthropic | 58 | |
| 33 | InclusionAI | 57 | |
| 34 | LiquidAI | 56.47 | |
| 35 | Upstage | 55.78 | |
| 36 | ZYZAYA1-8B | Zyphra | 52.56 |
| 37 | OPMiniCPM5-1B | OpenBMB | 46.67 |
| 38 | LiquidAI | 38.4 | |
| 39 | KAKanana-2 1.3B Instruct | Kakao | 34.69 |
| 40 | KAKanana-2 3B Instruct | Kakao | 33.33 |
| 41 | LiquidAI | 25.8 | |
| 42 | InclusionAI | 24.93 |
Evidence key: Observed
Rows are ordered by the value BenchLM published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.