Search papers, labs, and topics across Lattice.
This paper introduces the VSI-Super-Wild benchmark, designed to evaluate spatial supersensing capabilities of multimodal models in real-world settings, moving beyond the limitations of synthetic video datasets. By analyzing 6,980 human-verified question-answer pairs from 442 long-form videos, the authors reveal that current models struggle with coherent world-state tracking, particularly as complexity and temporal horizons increase. The study identifies four key failure modes鈥攕patial collapse, semantic shortcuts, insufficient update, and instance confusion鈥攖hat highlight the need for improved mechanisms to integrate objects, agents, and environments into a cohesive spatial model.
Despite advancements in static image understanding, models falter in tracking dynamic world states over time, exposing critical gaps in spatial reasoning.
Humans can efficiently parse continuous sensory streams, from hours to years, scaffolding an internal world model that grounds spatial reasoning and prediction. To mimic this capacity, spatial supersensing challenges multimodal models to move beyond linguistic understanding toward true world modeling. However, their benchmark relies on synthetic long videos, formed by concatenating random short clips, and is mostly limited to household scenes, leaving real-world continuity and diversity underexplored. To address the gap, we introduce $\textbf{VSI-Super-Wild}$, a large-scale benchmark for evaluating spatial supersensing over long temporal horizons in diverse in-the-wild scenes. Notably, inspired by cognitive studies on how humans structure experience, we systematically probe the full triad of world state: the agent (observer), objects (scene items), and the environment (places and global layout). In total, VSI-Super-Wild contains $\textbf{6,980}$ human-verified question-answer pairs derived from $\textbf{442}$ real-world videos spanning 8 scene categories, including long-form recordings exceeding 4 hours. Results on VSI-Super-Wild expose a fundamental disconnect: despite advances in static image understanding, models consistently fail at tasks that require coherent world-state tracking over time. We characterize how performance degrades with world-state complexity and temporal horizon, and diagnose four failure modes: spatial collapse, semantic shortcuts, insufficient update, and instance confusion. This taxonomy reveals that models lack mechanisms to bind objects, agents, and environments into a unified spatial world model, a fundamental gap that defines the path forward for spatial supersensing.