Agentic benchmark
CyberGym leaderboard
Every model the catalog carries a published CyberGym value for, ranked by that value.
A cybersecurity task benchmark for evaluating defensive cyber workflows and vulnerability-oriented agent performance.
CyberGym ranking
24 models with a published CyberGym value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Vulnerability reproduction and PoC generation |
|---|---|---|---|
| 1 | DeepSeek | 88.1 | |
| 2 | SAFugu Cyber | Sakana AI | 86.9 |
| 3 | Shanghai Artificial Intelligence Laboratory | 86.5 | |
| 4 | 86.2 | ||
| 5 | Z.AI | 84.5 | |
| 5 | OpenAI | 84.5 | |
| 7 | Anthropic | 83.8 | |
| 8 | DeepSeek | 83.3 | |
| 9 | 83.2 | ||
| 10 | Anthropic | 83.1 | |
| 11 | OpenAI | 81.8 | |
| 11 | OpenAI | 81.8 | |
| 13 | OpenAI | 79.0 | |
| 14 | Tencent | 78.4 | |
| 15 | OpenAI | 77.9 | |
| 16 | DeepSeek | 76.7 | |
| 17 | Anthropic | 73.1 | |
| 18 | Z.AI | 68.7 | |
| 19 | Anthropic | 66.6 | |
| 20 | Anthropic | 65.2 | |
| 21 | Meta | 59.0 | |
| 22 | Anthropic | 50.6 | |
| 23 | Meta | 43.5 | |
| 24 | Z.AI | 43.2 |
Evidence key: Observed
Rows are ordered by the value CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.