Search papers, labs, and topics across Lattice.
To overcome the calibration and metric-scale ambiguities inherent in evaluating video physics, the authors construct Principia, a benchmark that tests Newtonian reasoning via calibration-independent relational consistency between paired objects in the same scene. Across eight physical phenomena spanning translational, rotational, collisional, and oscillatory dynamics, the authors define an image-space consistency metric to quantify physical violations without 3D camera parameters. Evaluating six frontier video generators reveals that despite high visual quality scores (~0.8 on VBench), no model exceeds 0.42 on Principia, while leading vision-language models struggle to detect these physical violations (peaking at 67% accuracy).
Today's top video generators score near 0.8 on standard benchmarks like VBench, but none surpass 0.42 when tested on whether two co-occurring objects obey the same basic laws of physics.
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate, object scale, and camera calibration, all of which are often ambiguous or unavailable in generated video. We propose a different approach. When two objects in the same scene obey the same physical law, their motions must satisfy predictable relationships, and these relationships hold independent of calibration. We introduce Principia, a benchmark that evaluates Newtonian physics through relational consistency between paired objects. Principia spans eight phenomena - gravity, restitution, friction, rotational inertia, projectile motion, momentum, pendulum, and mass-spring oscillation - across translational, rotational, collisional, and oscillatory dynamics, using real-world scenes recorded under controlled protocols. We also introduce a calibration-independent consistency score that quantifies physical violation directly in image space. Across thousands of generations from six state-of-the-art video generators, no model exceeds 0.42 on Principia despite all scoring around 0.8 on VBench. Vision-language models are evaluated on their ability to detect relational physics violations, with the best model achieving only 67% accuracy and most performing near chance level.