Coding benchmark
SciCode leaderboard
Scientific Code Benchmark. Every model the catalog carries a published SciCode value for, ranked by that value.
SciCode evaluates language models on generating code for realistic scientific research problems across 16 subfields of physics, math, chemistry, biology, and material science. Problems decompose into 338 subproblems requiring domain knowledge recall, scientific reasoning, and precise code synthesis. Based on real scripts from published research.
SciCode ranking
26 models with a published SciCode value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | SciCode |
|---|---|---|---|
| 1 | SASakana Fugu | Sakana AI | 60.1 |
| 2 | SASakana Fugu-Ultra | Sakana AI | 58.7 |
| 3 | Alibaba | 53.5 | |
| 4 | 53.1 | ||
| 5 | Moonshot AI | 52.2 | |
| 6 | Alibaba | 51.3 | |
| 7 | TMInkling-Small | Thinking Machines Lab | 48.7 |
| 7 | Moonshot AI | 48.7 | |
| 9 | xAI | 47.3 | |
| 10 | Alibaba | 47 | |
| 11 | NVIDIA | 44.6 | |
| 12 | Meta | 43.6 | |
| 13 | InclusionAI | 41.24 | |
| 14 | Tencent | 41.2 | |
| 15 | STA.X K2 | SK Telecom | 41 |
| 16 | InclusionAI | 40.37 | |
| 17 | IBM | 38.76 | |
| 18 | Inception | 38 | |
| 19 | LG AI Research | 37.4 | |
| 20 | IBM | 36.09 | |
| 21 | NVIDIA | 32 | |
| 22 | NVIDIA | 31.38 | |
| 23 | InternScience | 29.6 | |
| 24 | InclusionAI | 27 | |
| 25 | OPMiniCPM5-2B | OpenBMB | 26.3 |
| 26 | IBM | 24.11 |
Evidence key: Observed
Rows are ordered by the value BenchLM published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.