Skip to main content
ModelScale

Agentic benchmark

MCP-Tasks leaderboard

Every model the catalog carries a published MCP-Tasks value for, ranked by that value.

CategoryAgentic
MeasureInteractive tool-use evaluation
TasksMCP-integrated tool tasks
DifficultyAdvanced MCP workflows

A Model Context Protocol task benchmark used in Qwen's launch tables to measure practical execution over MCP-style tools and integrations.

MCP-Tasks ranking

5 models with a published MCP-Tasks value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published MCP-Tasks value
RankModelProviderInteractive tool-use evaluation
1Qwen3.5 397BAlibaba74.2
2Qwen3.6 PlusAlibaba74.1
3Claude Opus 4.5Anthropic71.8
4GLM-5Z.AI60.8
5Kimi K2.5Moonshot AI59.1

Evidence key: Observed

Rows are ordered by the value Qwen3.6 launch benchmarks published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard