Skip to main content
ModelScale

Agentic benchmark

QwenClawBench leaderboard

Every model the catalog carries a published QwenClawBench value for, ranked by that value.

CategoryAgentic
MeasureEnd-to-end agent evaluation
TasksReal-world agent workflows
DifficultyBroad real-world agentic execution

Qwen's internal OpenClaw-style benchmark for measuring broad real-world agent performance across practical productivity and research tasks.

QwenClawBench ranking

10 models with a published QwenClawBench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published QwenClawBench value
RankModelProviderEnd-to-end agent evaluation
1Qwen3.7 MaxAlibaba64.3
2Qwen3.7 PlusAlibaba61.8
3Qwen 3.6 Max (preview)Alibaba59.0
4Qwen3.6 PlusAlibaba57.2
5Kimi K2.5Moonshot AI54.3
6GLM-5Z.AI54.1
7Qwen3.6-27BAlibaba53.4
8Qwen3.6-35B-A3BAlibaba52.6
9Claude Opus 4.5Anthropic52.3
10Qwen3.5 397BAlibaba51.8

Evidence key: Observed

Rows are ordered by the value Qwen3.6 launch benchmarks published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard