Skip to main content
ModelScale

korean benchmark

KMMLU-Hard leaderboard

Every model the catalog carries a published KMMLU-Hard value for, ranked by that value.

Categorykorean
MeasureMultiple choice questions
Tasks~5,000 questions
DifficultyAdvanced Korean reasoning

A filtered hard subset of KMMLU containing ~5,000 questions that most models get wrong.

No model in the catalog has a published KMMLU-Hard score.

The benchmark is defined by Evaluating LLMs on Hard Korean Queries, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.

All leaderboards