Knowledge benchmark
TruthfulQA leaderboard
Every model the catalog carries a published TruthfulQA value for, ranked by that value.
CategoryKnowledge
MeasureQuestion answering
TasksTruthfulness and misconception resistance
DifficultyHallucination and factuality stress test
A benchmark designed to measure whether language models produce truthful answers instead of repeating common misconceptions or misleading falsehoods.
No model in the catalog has a published TruthfulQA score.
The benchmark is defined by TruthfulQA: Measuring How Models Mimic Human Falsehoods, but the catalog carries no value for it yet. An absent value is shown as absent here rather than as a zero.