Skip to main content
ModelScale

Agentic benchmark

Agents' Last Exam leaderboard

Every model the catalog carries a published Agents' Last Exam value for, ranked by that value.

CategoryAgentic
MeasureProvider-reported task score
TasksAgent tasks
DifficultyAdvanced agentic work

An agent benchmark reported in DeepSeek's V4 Flash 0731 launch comparison.

Agents' Last Exam ranking

11 models with a published Agents' Last Exam value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Agents' Last Exam value
RankModelProviderProvider-reported task score
1GPT-6 AstraOpenAI59.3
2Qwen3.8 MaxAlibaba52.4
3Qwen3.8-Flash-NextAlibaba51.2
4Qwen3.8-27BAlibaba42.9
5DeepSeek V4.1 FlashDeepSeek31.8
6GLM-5.3Z.AI28.5
7Gemini 3.7 FlashGoogle26.3
7GLM-5.3-FlashZ.AI26.3
9DeepSeek V4 Pro 0813DeepSeek25.7
10DeepSeek V4 Flash 0731DeepSeek25.2
11Hy4 previewTencent22.8

Evidence key: Observed

Rows are ordered by the value DeepSeek V4 Flash 0731 update published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard