Agentic benchmark
Gert Labs leaderboard
Gert Labs Composite Game Benchmark. Every model the catalog carries a published Gert Labs value for, ranked by that value.
A game-environment benchmark that evaluates AI models in novel games covering strategic planning, resource management, spatial reasoning, cooperation, and theory of mind.
Gert Labs ranking
52 models with a published Gert Labs value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Composite game leaderboard |
|---|---|---|---|
| 1 | Anthropic | 72.97 | |
| 2 | OpenAI | 72.93 | |
| 3 | Anthropic | 65.59 | |
| 4 | OpenAI | 64.89 | |
| 5 | Alibaba | 64.27 | |
| 6 | Anthropic | 64.23 | |
| 7 | 63.23 | ||
| 8 | Anthropic | 62.92 | |
| 9 | Xiaomi | 62.70 | |
| 10 | Anthropic | 61.85 | |
| 10 | 61.85 | ||
| 12 | Z.AI | 60.11 | |
| 13 | OpenAI | 57.47 | |
| 14 | 56.87 | ||
| 15 | Moonshot AI | 56.82 | |
| 16 | 56.63 | ||
| 17 | Alibaba | 54.84 | |
| 18 | OpenAI | 51.79 | |
| 19 | StepFun | 51.57 | |
| 20 | Z.AI | 50.99 | |
| 21 | Alibaba | 50.60 | |
| 22 | OpenAI | 49.68 | |
| 23 | xAI | 49.15 | |
| 24 | Anthropic | 48.51 | |
| 25 | xAI | 47.32 | |
| 26 | Xiaomi | 46.89 | |
| 27 | Alibaba | 46.76 | |
| 28 | OpenAI | 46.54 | |
| 29 | Moonshot AI | 45.88 | |
| 30 | xAI | 43.86 | |
| 31 | Alibaba | 43.74 | |
| 32 | Alibaba | 42.65 | |
| 33 | xAI | 42.34 | |
| 34 | 42.01 | ||
| 35 | OpenAI | 41.24 | |
| 36 | MiniMax | 40.40 | |
| 37 | Z.AI | 39.95 | |
| 38 | Anthropic | 39.66 | |
| 39 | Alibaba | 39.41 | |
| 40 | Mistral | 39.10 | |
| 41 | 38.46 | ||
| 42 | xAI | 38.36 | |
| 43 | Tencent | 36.91 | |
| 44 | Xiaomi | 36.68 | |
| 45 | 35.26 | ||
| 46 | Moonshot AI | 32.58 | |
| 47 | Arcee AI | 32.55 | |
| 48 | Z.AI | 30.76 | |
| 49 | OpenAI | 29.61 | |
| 50 | DeepSeek | 29.57 | |
| 51 | Alibaba | 28.96 | |
| 52 | OpenAI | 25.65 |
Evidence key: Observed
Rows are ordered by the value Gert Labs rankings published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.