Skip to main content
ModelScale

Coding benchmark

FLTEval leaderboard

Every model the catalog carries a published FLTEval value for, ranked by that value.

CategoryCoding
MeasureLean 4 repository task completion
TasksFLT project pull requests
DifficultyFormal verification / proof engineering

A repository-level Lean 4 proof engineering benchmark that measures whether a model can complete formal proofs and correctly define new mathematical concepts inside realistic FLT project pull requests.

No model in the catalog has a published FLTEval score.

The benchmark is defined by Leanstral: Open-Source foundation for trustworthy vibe-coding, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards · Coding capability leaderboard