Skip to main content
ModelScale

Agentic benchmark

CyberGym leaderboard

Every model the catalog carries a published CyberGym value for, ranked by that value.

CategoryAgentic
MeasureVulnerability reproduction and PoC generation
Tasks1,507 vulnerability analysis instances
DifficultyReal-world cybersecurity

A cybersecurity task benchmark for evaluating defensive cyber workflows and vulnerability-oriented agent performance.

CyberGym ranking

24 models with a published CyberGym value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published CyberGym value
RankModelProviderVulnerability reproduction and PoC generation
1DeepSeek V4.1 FlashDeepSeek88.1
2SAFugu CyberSakana AI86.9
3Atria Dawn PreviewShanghai Artificial Intelligence Laboratory86.5
4Gemini 3.8 Flash CyberGoogle86.2
5GLM-5.3Z.AI84.5
5GPT-5.6 SolOpenAI84.5
7Claude Mythos 5Anthropic83.8
8DeepSeek V4 Pro 0813DeepSeek83.3
9Gemini 3.5 Flash CyberGoogle83.2
10Claude Mythos PreviewAnthropic83.1
11GPT-5.5OpenAI81.8
11GPT-5.6 TerraOpenAI81.8
13GPT-5.4OpenAI79.0
14Hy4 previewTencent78.4
15GPT-5.6 LunaOpenAI77.9
16DeepSeek V4 Flash 0731DeepSeek76.7
17Claude Opus 4.7 (Adaptive)Anthropic73.1
18GLM-5.1Z.AI68.7
19Claude Opus 4.6Anthropic66.6
20Claude Sonnet 4.6Anthropic65.2
21Muse Spark 1.1Meta59.0
22Claude Opus 4.5Anthropic50.6
23Muse SparkMeta43.5
24GLM-5Z.AI43.2

Evidence key: Observed

Rows are ordered by the value CyberGym: Evaluating AI Agents' Real-World Cybersecurity Capabilities at Scale published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Agentic capability leaderboard