Agentic benchmark
BFCL v4 leaderboard
Berkeley Function Calling Leaderboard v4. Every model the catalog carries a published BFCL v4 value for, ranked by that value.
A function-calling benchmark for tool selection, schema adherence, and argument correctness.
BFCL v4 ranking
22 models with a published BFCL v4 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Tool invocation and schema evaluation |
|---|---|---|---|
| 1 | BTBTL-3 | Bad Theory Labs | 88.5 |
| 2 | Shanghai Artificial Intelligence Laboratory | 77.0 | |
| 3 | Alibaba | 75.0 | |
| 4 | BTBTL-4 | Bad Theory Labs | 73.5 |
| 5 | InclusionAI | 73.0 | |
| 6 | Alibaba | 72.9 | |
| 7 | PAPokee-Isaac 28B | Pokee AI | 70.9 |
| 8 | OPMiniCPM5-2B | OpenBMB | 66.6 |
| 9 | IBM | 61.4 | |
| 10 | InclusionAI | 60.8 | |
| 11 | LiquidAI | 56.9 | |
| 12 | IBM | 52.4 | |
| 13 | IBM | 52.4 | |
| 14 | LiquidAI | 49.7 | |
| 15 | InclusionAI | 47.7 | |
| 16 | JEMellum2-12B-A2.5B-Thinking | JetBrains | 45.6 |
| 17 | JEMellum2-12B-A2.5B-Instruct | JetBrains | 44.2 |
| 18 | ZYZAYA1-8B | Zyphra | 39.2 |
| 19 | LiquidAI | 32.5 | |
| 20 | OPMiniCPM5-1B | OpenBMB | 25.1 |
| 21 | LiquidAI | 21.1 | |
| 22 | LiquidAI | 21.0 |
Evidence key: Observed
Rows are ordered by the value Trinity-Large-Thinking: Scaling an Open Source Frontier Agent published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.