Skip to main content
ModelScale

Agentic benchmark

Gert Labs leaderboard

Gert Labs Composite Game Benchmark. Every model the catalog carries a published Gert Labs value for, ranked by that value.

CategoryAgentic
MeasureComposite game leaderboard
TasksNovel game environments
DifficultyAgentic coding and decision-making
Published byGert Labs rankings

A game-environment benchmark that evaluates AI models in novel games covering strategic planning, resource management, spatial reasoning, cooperation, and theory of mind.

Gert Labs ranking

52 models with a published Gert Labs value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Gert Labs Composite Game Benchmark value
RankModelProviderComposite game leaderboard
1Claude Opus 4.8Anthropic72.97
2GPT-5.5OpenAI72.93
3Claude Opus 4.7Anthropic65.59
4GPT-5.4OpenAI64.89
5Qwen3.7 MaxAlibaba64.27
6Claude Opus 4.5Anthropic64.23
7Gemini 3 ProGoogle63.23
8Claude Sonnet 4.6Anthropic62.92
9MiMo-V2.5-ProXiaomi62.70
10Claude Opus 4.6Anthropic61.85
10Gemini 3.5 FlashGoogle61.85
12GLM-5.1Z.AI60.11
13GPT-5.3 CodexOpenAI57.47
14Gemini 3.1 ProGoogle56.87
15Kimi K2.6Moonshot AI56.82
16Gemini 3 FlashGoogle56.63
17Qwen3.6-27BAlibaba54.84
18GPT-5.2-CodexOpenAI51.79
19Step 3.7 FlashStepFun51.57
20GLM-5Z.AI50.99
21Qwen3.6 PlusAlibaba50.60
22GPT-5.1-CodexOpenAI49.68
23Grok Build 0.1xAI49.15
24Claude Sonnet 4.5Anthropic48.51
25Grok 4.1 FastxAI47.32
26MiMo-V2.5Xiaomi46.89
27Qwen3.5 397BAlibaba46.76
28GPT-5.2OpenAI46.54
29Kimi K2.5Moonshot AI45.88
30Grok 4.3xAI43.86
31Qwen3 MaxAlibaba43.74
32Qwen3.6-35B-A3BAlibaba42.65
33Grok 4xAI42.34
34Gemini 2.5 ProGoogle42.01
35GPT-5.1OpenAI41.24
36MiniMax M2.7MiniMax40.40
37GLM-4.7Z.AI39.95
38Claude 4 SonnetAnthropic39.66
39Qwen3.5-27BAlibaba39.41
40Mistral Medium 3.5 128BMistral39.10
41Gemini 3.1 Flash-LiteGoogle38.46
42Grok 4.20xAI38.36
43Hy3 PreviewTencent36.91
44MiMo-V2-ProXiaomi36.68
45Gemma 4 31BGoogle35.26
46Kimi K2.5 (Reasoning)Moonshot AI32.58
47Trinity-Large-ThinkingArcee AI32.55
48GLM-5V-TurboZ.AI30.76
49GPT-OSS 120BOpenAI29.61
50DeepSeek V3.2DeepSeek29.57
51Qwen3.5-35B-A3BAlibaba28.96
52GPT-4.1OpenAI25.65

Evidence key: Observed

Rows are ordered by the value Gert Labs rankings published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard