Search papers, labs, and topics across Lattice.
This study evaluates the reasoning capabilities of Large Vision Language Models (LVLMs) using visual illusions as a diagnostic tool, addressing a gap in existing assessments that typically focus on perception or specific domains. By creating the IllusionReasoning benchmark, which features real-world illusion images and annotated question-answer pairs, the authors reveal that the reasoning abilities of various LVLMs are overestimated. The findings suggest that LVLMs struggle with integrating perceptual and reasoning tasks, highlighting the need for further optimization in their design.
LVLMs may not be as adept at reasoning as previously thought, with new benchmarks revealing significant gaps in their performance.
Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either focus solely on perception or rely on specific domains such as maths or coding. Evaluation for reasoning capabilities that align with an open-world environment is still required, especially one that considers perception and reasoning jointly. To bridge this gap, we propose to evaluate LVLMs by exploiting visual illusions as a diagnostic tool. Visual illusions are phenomena in which the human visual system misinterprets objective signals, resulting in an understanding that deviates from reality. We constructed IllusionReasoning, a benchmark of illusion images collected from the real world, incorporating diverse annotated question-answer pairs. Based on IllusionReasoning, we show that the reasoning capabilities of a wide range of LVLMs are not as advanced as claimed. Our work provides new insights into LVLMs and offers future direction for optimisation.