Skip to main content
ModelScale

external benchmark

CyberBench leaderboard

Vals CyberBench. Every model the catalog carries a published CyberBench value for, ranked by that value.

Categoryexternal
MeasureAccuracy score
TasksOSS-Fuzz PoC and patch-verification tasks
DifficultyAutonomous cybersecurity exploit reproduction
Published byCyberBench

Can autonomous agents craft PoC inputs that trigger OSS-Fuzz vulnerabilities—and stop crashing after the fix?

No model in the catalog has a published CyberBench score.

The benchmark is defined by CyberBench, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards