Skip to main content
ModelScale

Multimodal & Grounded benchmark

A-OKVQA leaderboard

A Benchmark for Visual Question Answering using World Knowledge. Every model the catalog carries a published A-OKVQA value for, ranked by that value.

CategoryMultimodal & Grounded
MeasureMultiple choice and direct answer
TasksKnowledge-grounded visual question answering
DifficultyCommonsense and world knowledge about images

A visual question answering benchmark whose questions cannot be answered from the image alone and require commonsense or world knowledge.

A-OKVQA ranking

1 model with a published A-OKVQA value, ordered by that value, highest first. Models the source has not scored on this benchmark are not listed — they are unmeasured, not last.

Models ranked by their published A Benchmark for Visual Question Answering using World Knowledge value
RankModelProviderMultiple choice and direct answer
1PMTernary Bonsai 2 27BPrism ML86.8

Evidence key: Observed

Rows are ordered by the value A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge published, highest first. The source does not state whether a higher value is the better result, so this page does not either: for a benchmark that measures a rate of failure, read the table from the bottom.

All leaderboards