Search papers, labs, and topics across Lattice.
This paper introduces PerceptionBench, a benchmark designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs) by isolating perceptual errors from reasoning and domain knowledge failures. By analyzing responses from frontier MLLMs across 42 existing benchmarks, the authors develop an error taxonomy that identifies ten atomic perceptual capabilities, leading to the creation of 3,000 verified questions aimed at assessing these capabilities. The findings reveal that atomic perception remains a significant challenge, with no model achieving over 60% accuracy and notable variability in capability profiles, highlighting the need for targeted improvements in MLLMs' visual perception abilities.
Atomic visual perception in MLLMs is largely unsolved, with no model surpassing 60% accuracy on a new benchmark designed to isolate perceptual capabilities.
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60\% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs.