Knowledge benchmark
OpenBookQA leaderboard
Every model the catalog carries a published OpenBookQA value for, ranked by that value.
CategoryKnowledge
Measure4-way multiple choice
TasksElementary science questions
DifficultyElementary science reasoning
A science question-answering benchmark that tests whether models can apply a small open-book set of elementary science facts to multi-step reasoning questions.
No model in the catalog has a published OpenBookQA score.
The benchmark is defined by Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.