Instruction Following benchmark
IFEval leaderboard
Instruction-Following Eval. Every model the catalog carries a published IFEval value for, ranked by that value.
A benchmark of 541 prompts built from 25 verifiable instruction types. It tests whether a model follows checkable constraints such as keyword, length, casing, and response-format requirements.
IFEval ranking
33 models with a published IFEval value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Constrained generation |
|---|---|---|---|
| 1 | Alibaba | 95 | |
| 2 | InternScience | 94.82 | |
| 3 | InternScience | 94.8 | |
| 4 | Alibaba | 94.6 | |
| 5 | Alibaba | 94.3 | |
| 5 | Alibaba | 94.3 | |
| 7 | DSdots3-note Preview | Dots Studio | 93.9 |
| 7 | Moonshot AI | 93.9 | |
| 7 | OpenAI | 93.9 | |
| 10 | Alibaba | 93.4 | |
| 11 | Z.AI | 92.6 | |
| 11 | Alibaba | 92.6 | |
| 13 | LG AI Research | 92.4 | |
| 14 | OpenAI | 92.2 | |
| 15 | Alibaba | 91.9 | |
| 16 | LiquidAI | 91.84 | |
| 17 | PMTernary Bonsai 2 27B | Prism ML | 91.31 |
| 18 | Anthropic | 90.9 | |
| 19 | OpenAI | 88.5 | |
| 20 | OpenAI | 87.4 | |
| 21 | OPMiniCPM5-2B | OpenBMB | 86.7 |
| 22 | DeepSeek | 86.1 | |
| 23 | ZYZAYA1-8B | Zyphra | 85.58 |
| 24 | OpenAI | 83.2 | |
| 25 | LiquidAI | 82.3 | |
| 26 | KAKanana-2 3B Instruct | Kakao | 80.96 |
| 27 | CECeleris-1 | Celeris | 80.8 |
| 28 | OPMiniCPM5-1B | OpenBMB | 80.41 |
| 29 | KAKanana-2 1.3B Instruct | Kakao | 77.63 |
| 30 | JEMellum2-12B-A2.5B-Thinking | JetBrains | 76.5 |
| 31 | JEMellum2-12B-A2.5B-Instruct | JetBrains | 75.8 |
| 32 | LiquidAI | 71.71 | |
| 33 | LiquidAI | 61.16 |
Evidence key: Observed
Rows are ordered by the value Instruction-Following Evaluation for Large Language Models published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.