Search papers, labs, and topics across Lattice.
This paper introduces Geo3R, a novel framework designed to address spatial reasoning hallucinations in Multimodal Large Language Models (MLLMs) by incorporating geometric evidence and structured 3D reasoning. The authors identify a critical gap in existing methods that fail to effectively bridge 2D visual representations with 3D spatial realities, leading to frequent hallucinations in specific scenarios such as perspective effects and object orientation. Experimental results demonstrate that Geo3R significantly reduces these hallucinations across 18 tasks in three benchmark scenarios, outperforming current state-of-the-art models without requiring additional training.
Spatial reasoning hallucinations in MLLMs can be drastically reduced by integrating geometric evidence, challenging the effectiveness of traditional mitigation methods.
Despite remarkable progress in visual understanding, Multimodal Large Language Models (MLLMs) remain prone to hallucinations when reasoning about spatial relationships, often producing judgments that contradict the true 3D structure of the scene. Though several existing works have proposed to mitigate hallucinations, our analysis indicates that they show limited effectiveness in spatial reasoning, as they fail to bridge the fundamental gap between 2D visual representations and 3D spatial reality. Based on this finding, we define hallucinations arising from insufficient spatial structure modeling as spatial reasoning hallucination, a subcategory of relation hallucination that existing mitigation methods fail to address. We further identify three typical scenarios where such hallucinations frequently occur: perspective effects, object orientation, and viewpoint changes. To this end, we propose Geo3R, a training-free, plug-and-play framework that incorporates geometric evidence and structured 3D reasoning to mitigate spatial reasoning hallucination. Experiments on three benchmarks, covering 18 tasks across all three scenarios, show that Geo3R substantially reduces spatial reasoning hallucination across diverse MLLMs without additional training, outperforming existing models and methods.