Skip to main content
ModelScale

Knowledge benchmark

LABBench2 leaderboard

LABBench2: An Improved Benchmark for AI Systems Performing Biology Research. Every model the catalog carries a published LABBench2 value for, ranked by that value.

CategoryKnowledge
MeasureAggregate accuracy
TasksNearly 1,900 biology-research tasks
DifficultyReal-world biology research

A benchmark of realistic biology-research tasks involving literature, figures, tables, databases, and bioinformatics files.

LABBench2 ranking

6 models with a published LABBench2 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published LABBench2: An Improved Benchmark for AI Systems Performing Biology Research value
RankModelProviderAggregate accuracy
1Gemini 3.8 FlashGoogle86.2
2Claude Opus 5Anthropic84.2
3Gemini 3.7 FlashGoogle82.1
3GPT-5.6 SolOpenAI82.1
5GPT-5.6 TerraOpenAI81.2
6Claude Sonnet 5Anthropic80.1

Evidence key: Observed

Rows are ordered by the value LABBench2: An Improved Benchmark for AI Systems Performing Biology Research published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards · Knowledge capability leaderboard