Knowledge benchmark
MMMLU leaderboard
Every model the catalog carries a published MMMLU value for, ranked by that value.
A multilingual MMLU-style benchmark reported in provider evaluation tables.
MMMLU ranking
5 models with a published MMMLU value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.
| Rank | Model | Provider | Exact match |
|---|---|---|---|
| 1 | INInterfaze Beta | Interfaze | 90.9 |
| 2 | Alibaba | 90.3 | |
| 3 | Alibaba | 89.0 | |
| 4 | LG AI Research | 86.6 | |
| 5 | 83.4 |
Evidence key: Observed
Rows are ordered by the value MMMLU published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.