Skip to main content
ModelScale

Multilingual benchmark

NOVA-63 leaderboard

Every model the catalog carries a published NOVA-63 value for, ranked by that value.

CategoryMultilingual
MeasureCross-lingual benchmark
TasksBroad multilingual evaluation
DifficultyBroad multilingual capability

A broad multilingual benchmark row from Qwen's launch comparisons intended to measure cross-lingual capability beyond a single language family.

NOVA-63 ranking

7 models with a published NOVA-63 value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published NOVA-63 value
RankModelProviderCross-lingual benchmark
1Qwen3.5 397BAlibaba59.1
2Qwen3.7 MaxAlibaba59.0
3Qwen3.7 PlusAlibaba58.8
4Qwen3.6 PlusAlibaba57.9
5Claude Opus 4.5Anthropic56.7
6Kimi K2.5Moonshot AI56.0
7GLM-5Z.AI55.1

Evidence key: Observed

Rows are ordered by the value Qwen3.6 launch benchmarks published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards