Search papers, labs, and topics across Lattice.
This paper introduces ConsiSpace, a geometry-consistency-aware framework designed to enhance video spatial reasoning by integrating spatial consistency as both an organizational principle and a learning signal. By developing a geometry-consistent memory (GCM) that incorporates implicit evidence tokens and explicit geometric cues, the framework effectively organizes spatial evidence for improved reasoning. The approach, which employs unified consistency self-supervised reinforcement learning (UC-SSRL) post-supervised fine-tuning, achieves significant performance improvements across three benchmarks, with an average score increase of 12.6 points over existing models.
Geometry-consistency awareness in video spatial reasoning can lead to a 12.6-point performance boost over current state-of-the-art models.
Video spatial reasoning is essential for navigation-oriented perception and long-video question answering, where models must infer spatial relations across long horizons under changing viewpoints. However, existing multimodal large language models (MLLMs) remain largely semantic-centric, and often fail to reliably aggregate consistent spatial evidence from redundant video observations, leading to inefficient or unstable reasoning. To address these issues, we propose ConsiSpace, a geometry-consistency-aware framework for geometry-sensitive video spatial reasoning that turns spatial consistency into both an evidence organization principle and an explicit post-SFT learning signal. We build a geometry-consistent memory (GCM) including implicit evidence tokens and explicit geometric cues, and leverage efficient organization strategies to compactly preserve task-related spatial evidence. Furthermore, we utilize unified consistency self-supervised reinforcement learning (UC-SSRL) after supervised fine-tuning to improve cross-view stability, with answer-, metric-, and topology-consistency rewards. Extensive experiments on three spatial-reasoning benchmarks, VSI-Bench, OSI-Bench, and MMSI-Video-Bench, show consistent gains, improving the average score by 12.6 points over the strongest baselines.