Search papers, labs, and topics across Lattice.
This paper introduces a novel evaluation paradigm called Retrospective Physical Process Reasoning, aimed at assessing Vision Language Models' ability to infer hidden physical processes from sparse observations. The authors present RetroHolmes, a benchmark featuring object-centric image pairs with reachability labels and causal sequences, which reveals significant failure modes in current models, such as judgment bias and reliance on beliefs over physical evidence. By employing visual simulation as an analysis-by-synthesis method, the study underscores the necessity of physically grounded representations for improving physical reasoning in AI systems.
Vision Language Models exhibit judgment bias and belief dominance over physical evidence, revealing critical flaws in their reasoning capabilities.
Humans can infer hidden physical processes from sparse observations, yet current evaluation protocols for Vision Language Models fail to assess whether such physical reasoning is genuinely captured. To address this gap, we introduce Retrospective Physical Process Reasoning, a new evaluation paradigm to reason backward from outcomes under explicit physical constraints. Building on the paradigm, we present RetroHolmes, the first real-world benchmark for Retrospective Physical Process Reasoning, comprising object-centric image pairs annotated with reachability labels and causal step sequences across diverse physical transitions. Using RetroHolmes, we analyze state of the art Vision Language Models and uncover systematic failure modes, including judgment bias in reachability assessment and belief dominance over physical evidence, mirroring sycophancy behavior observed in large language models. We further demonstrate a simple analysis-by-synthesis instantiation with visual simulation as an intermediate step, validating the diagnostic value of RetroHolmes and highlighting the importance of physically grounded intermediate representations for physical reasoning.