Skip to main content
ModelScale

Coding benchmark

SciCode leaderboard

Scientific Code Benchmark. Every model the catalog carries a published SciCode value for, ranked by that value.

CategoryCoding
Published byBenchLM

SciCode evaluates language models on generating code for realistic scientific research problems across 16 subfields of physics, math, chemistry, biology, and material science. Problems decompose into 338 subproblems requiring domain knowledge recall, scientific reasoning, and precise code synthesis. Based on real scripts from published research.

SciCode ranking

26 models with a published SciCode value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published Scientific Code Benchmark value
RankModelProviderSciCode
1SASakana FuguSakana AI60.1
2SASakana Fugu-UltraSakana AI58.7
3Qwen3.7 MaxAlibaba53.5
4Gemini 3.5 FlashGoogle53.1
5Kimi K2.6Moonshot AI52.2
6Qwen3.7 PlusAlibaba51.3
7TMInkling-SmallThinking Machines Lab48.7
7Kimi K2.5Moonshot AI48.7
9Grok 4.3xAI47.3
10Qwen 3.6 Max (preview)Alibaba47
11Nemotron 3 UltraNVIDIA44.6
12Muse Glimmer 30BMeta43.6
13Ling 3.0 FlashInclusionAI41.24
14Hy3 PreviewTencent41.2
15STA.X K2SK Telecom41
16Ling 3.0 Flash FP8InclusionAI40.37
17Granite 4.2 30BIBM38.76
18Mercury 2.5Inception38
19K-EXAONE 2.0LG AI Research37.4
20Granite 4.2 8BIBM36.09
21Nemotron 3 Nano Omni 30B A3BNVIDIA32
22Nemotron 3.5 Lightning 30B A3B NVFP4NVIDIA31.38
23Agents-A1-4BInternScience29.6
24Ling 2.6 FlashInclusionAI27
25OPMiniCPM5-2BOpenBMB26.3
26Granite 4.2 3BIBM24.11

Evidence key: Observed

Rows are ordered by the value BenchLM published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Coding capability leaderboard