Coding benchmark
FLTEval leaderboard
Every model the catalog carries a published FLTEval value for, ranked by that value.
CategoryCoding
MeasureLean 4 repository task completion
TasksFLT project pull requests
DifficultyFormal verification / proof engineering
A repository-level Lean 4 proof engineering benchmark that measures whether a model can complete formal proofs and correctly define new mathematical concepts inside realistic FLT project pull requests.
No model in the catalog has a published FLTEval score.
The benchmark is defined by Leanstral: Open-Source foundation for trustworthy vibe-coding, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.