Skip to main content
ModelScale

Coding benchmark

FrontierCode 1.1 Main leaderboard

Every model the catalog carries a published FrontierCode 1.1 Main value for, ranked by that value.

CategoryCoding
MeasureRepository task completion with maintainer rubrics
Tasks100 private Main tasks (150 in Extended)
DifficultyFrontier coding-agent quality

Cognition's 100-task software-engineering benchmark for whether coding agents produce mergeable, production-quality pull requests, scored for correctness, tests, scope, style, and maintainability through maintainer-authored rubrics.

FrontierCode 1.1 Main ranking

13 models with a published FrontierCode 1.1 Main value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published FrontierCode 1.1 Main value
RankModelProviderRepository task completion with maintainer rubrics
1Claude Fable 5Anthropic53.5
2Claude Opus 5Anthropic53.4
3GPT-6 AstraOpenAI53.3
4SWE-2Cognition50.0
5Claude Opus 4.8Anthropic46.5
6Gemini 3.7 FlashGoogle43.6
7GPT-5.5OpenAI43.0
8Claude Sonnet 5Anthropic42.7
9SWE-1.7Cognition42.3
10Claude Opus 4.7Anthropic38.5
11GPT-5.4 miniOpenAI27.0
12Claude Opus 4.6Anthropic26.9
13Claude Sonnet 4.6Anthropic24.3

Evidence key: Observed

Rows are ordered by the value FrontierCode leaderboard published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard