Skip to main content
ModelScale

Knowledge benchmark

FrontierScience Research leaderboard

Every model the catalog carries a published FrontierScience Research value for, ranked by that value.

CategoryKnowledge
MeasureResearch evaluation
TasksScientific research problems
DifficultyFrontier scientific research

A research-focused FrontierScience evaluation variant for scientific investigation and problem solving.

FrontierScience Research ranking

4 models with a published FrontierScience Research value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published FrontierScience Research value
RankModelProviderResearch evaluation
1APApodex 1.1Apodex63.3
2APApodex 1.1 MiniApodex51.7
3GPT-5.4 ProOpenAI36.7
4Agents-A1-4BInternScience33.3

Evidence key: Observed

Rows are ordered by the value Muse Spark Eval Methodology published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Knowledge capability leaderboard