Skip to main content
ModelScale

external benchmark

ExploitBench leaderboard

ExploitBench v8-bench. Every model the catalog carries a published ExploitBench value for, ranked by that value.

Categoryexternal
MeasureCapability coverage percentage over 16 flags
TasksV8 exploit synthesis runs
DifficultyBrowser exploitation and cybersecurity
Published byExploitBench

A cybersecurity benchmark for evaluating LLM agents on full-control V8 exploit synthesis using 16 measured exploit capability flags.

ExploitBench ranking

8 models with a published ExploitBench value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published ExploitBench v8-bench value
RankModelProviderCapability coverage percentage over 16 flags
1GPT-6 AstraOpenAI100
2Claude Mythos 5Anthropic78
3GPT-5.6 SolOpenAI74
4Claude Mythos PreviewAnthropic69
5GLM-5.3Z.AI54
6GPT-5.6 TerraOpenAI53
7GPT-5.6 LunaOpenAI33
8Kimi K3Moonshot AI32

Evidence key: Observed

Rows are ordered by the value ExploitBench published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards