Search papers, labs, and topics across Lattice.
This paper addresses the issue of ego-motion ambiguity in Vision-Language Models (VLMs) that leads to Kinematic Collapse, where models fail to accurately interpret physical motion under large displacements. The authors introduce Dyn-3D, a benchmark that employs counterfactual 3D rendering to separate visual changes from actual kinematic properties, and propose the TempoVista framework with the Kinematic-GSPO algorithm to integrate physical ground truth into policy optimization. Experimental results show that this approach enhances motion estimation and spatial reasoning by leveraging camera dynamics for geometric calibration, marking a significant advancement in VLM capabilities.
Ego-motion ambiguity in VLMs can lead to severe spatial reasoning failures, but a new benchmark and framework show how to ground visual representations in 3D space effectively.
As Vision-Language Models (VLMs) tackle dynamic 3D spatial reasoning, ego-motion perception becomes essential to resolve monocular scale ambiguity. However, current models often overfit to smooth trajectory priors rather than genuinely understanding physical motion. Consequently, their spatial reasoning degrades severely under large displacements, a phenomenon we term Kinematic Collapse. This failure stems from spurious visual-motion correlations in natural videos and a lack of explicit physical supervision. To evaluate this, we introduce Dyn-3D, a benchmark using counterfactual 3D rendering to rigorously decouple visual changes from true kinematic properties. Furthermore, we propose the TempoVista framework, featuring the Kinematic-GSPO algorithm. By embedding metric physical ground truth into policy optimization, TempoVista explicitly grounds visual representations in 3D space. Experiments demonstrate that our approach significantly improves both motion estimation and robust spatial reasoning by utilizing camera dynamics as an effective geometric calibration signal.