Skip to main content
ModelScale

Agentic benchmark

BFCL v3 leaderboard

Berkeley Function Calling Leaderboard v3. Every model the catalog carries a published BFCL v3 value for, ranked by that value.

CategoryAgentic
MeasureTool invocation and schema evaluation
TasksFunction-calling tasks
DifficultyAdvanced tool use

A function-calling benchmark for tool selection, schema adherence, and argument correctness, covering single-turn, parallel, irrelevance and multi-turn subsets.

BFCL v3 ranking

1 model with a published BFCL v3 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Berkeley Function Calling Leaderboard v3 value
RankModelProviderTool invocation and schema evaluation
1PMTernary Bonsai 2 27BPrism ML74.9

Evidence key: Observed

Rows are ordered by the value The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard