Skip to main content
ModelScale

Agentic benchmark

VITA-Bench leaderboard

Every model the catalog carries a published VITA-Bench value for, ranked by that value.

CategoryAgentic
MeasureEnd-to-end interactive agent evaluation
TasksInteractive consumer-service agent tasks
DifficultyLong-horizon real-world workflows

An interactive real-world agent benchmark grounded in practical consumer-service tasks such as delivery, in-store consumption, and online travel workflows.

VITA-Bench ranking

12 models with a published VITA-Bench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published VITA-Bench value
RankModelProviderEnd-to-end interactive agent evaluation
1Qwen3.7 MaxAlibaba47.9
2Qwen3.7 PlusAlibaba45.6
3Qwen3.6 PlusAlibaba44.3
4Qwen3.5 397BAlibaba43.7
5Agents-A1-4BInternScience40.3
6Agents-A1InternScience38.8
7Qwen3.6-35B-A3BAlibaba35.6
8Claude Opus 4.5Anthropic23.3
9LongCat-Flash-Lite-SparseMeituan21.7
10DeepSeek V3.2DeepSeek18.5
11Claude Sonnet 4.5Anthropic17.0
12GLM-4.7Z.AI15.5

Evidence key: Observed

Rows are ordered by the value VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard