Skip to main content
ModelScale

Agentic benchmark

BFCL v4 leaderboard

Berkeley Function Calling Leaderboard v4. Every model the catalog carries a published BFCL v4 value for, ranked by that value.

CategoryAgentic
MeasureTool invocation and schema evaluation
TasksFunction-calling tasks
DifficultyAdvanced tool use

A function-calling benchmark for tool selection, schema adherence, and argument correctness.

BFCL v4 ranking

22 models with a published BFCL v4 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Berkeley Function Calling Leaderboard v4 value
RankModelProviderTool invocation and schema evaluation
1BTBTL-3Bad Theory Labs88.5
2Atria Dawn PreviewShanghai Artificial Intelligence Laboratory77.0
3Qwen3.7 MaxAlibaba75.0
4BTBTL-4Bad Theory Labs73.5
5Ling 3.0 FlashInclusionAI73.0
6Qwen3.7 PlusAlibaba72.9
7PAPokee-Isaac 28BPokee AI70.9
8OPMiniCPM5-2BOpenBMB66.6
9Granite 4.2 30BIBM61.4
10LLaDA2.2-flashInclusionAI60.8
11LFM2.5-2.6BLiquidAI56.9
12Granite 4.2 3BIBM52.4
13Granite 4.2 8BIBM52.4
14LFM2.5-8B-A1BLiquidAI49.7
15LLaDA2.2-miniInclusionAI47.7
16JEMellum2-12B-A2.5B-ThinkingJetBrains45.6
17JEMellum2-12B-A2.5B-InstructJetBrains44.2
18ZYZAYA1-8BZyphra39.2
19LFM2.5-VL-3BLiquidAI32.5
20OPMiniCPM5-1BOpenBMB25.1
21LFM2.5-VL-450MLiquidAI21.1
22LFM2.5-230MLiquidAI21.0

Evidence key: Observed

Rows are ordered by the value Trinity-Large-Thinking: Scaling an Open Source Frontier Agent published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard