Agentic benchmark
MCP Atlas leaderboard
Every model the catalog carries a published MCP Atlas value for, ranked by that value.
A benchmark for tool-calling over Model Context Protocol integrations and external tools.
MCP Atlas ranking
37 models with a published MCP Atlas value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Interactive tool-calling evaluation |
|---|---|---|---|
| 1 | Meta | 88.1 | |
| 2 | Anthropic | 85.8 | |
| 3 | Moonshot AI | 84.2 | |
| 4 | Tencent | 83.7 | |
| 5 | 83.6 | ||
| 6 | Anthropic | 82.2 | |
| 7 | OAOrnith-1.5-397B | Ornith AI | 80 |
| 8 | TMInkling-Small | Thinking Machines Lab | 79.6 |
| 9 | Anthropic | 77.3 | |
| 10 | Z.AI | 76.8 | |
| 11 | Alibaba | 76.4 | |
| 12 | Moonshot AI | 76 | |
| 13 | Meta | 75.5 | |
| 14 | OpenAI | 75.3 | |
| 15 | MiniMax | 74.2 | |
| 16 | TMInkling | Thinking Machines Lab | 74.1 |
| 17 | DeepSeek | 73.6 | |
| 18 | Alibaba | 73.2 | |
| 19 | Z.AI | 71.8 | |
| 20 | OpenAI | 70.6 | |
| 21 | OAOrnith-1.5-35B-A3B | Ornith AI | 70.2 |
| 22 | DeepSeek | 69 | |
| 23 | InclusionAI | 65.5 | |
| 24 | Alibaba | 62.8 | |
| 25 | Upstage | 61.4 | |
| 26 | Upstage | 58.2 | |
| 27 | OpenAI | 57.7 | |
| 28 | OpenAI | 56.1 | |
| 29 | Moonshot AI | 55.9 | |
| 30 | OAOrnith-1.5-9B | Ornith AI | 54.2 |
| 31 | Alibaba | 48.2 | |
| 32 | InclusionAI | 46.21 | |
| 33 | Alibaba | 46.1 | |
| 34 | Meituan | 45.6 | |
| 35 | Anthropic | 42.3 | |
| 36 | Z.AI | 31.1 | |
| 37 | Moonshot AI | 29.5 |
Evidence key: Observed
Rows are ordered by the value Introducing GPT-5.4 mini and nano published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.