Search papers, labs, and topics across Lattice.
This paper introduces AcoustiTrace, a diagnostic benchmark designed to assess the acoustic physical realism of audio-video generation systems by evaluating sound generation, propagation, and reception across eight measurable dimensions. The authors constructed a large-scale dataset of real-world audio-video recordings and acoustically annotated RGB-D observations to facilitate targeted evaluations and prompt suites. Experiments demonstrate that even state-of-the-art generators frequently violate fundamental acoustic principles, highlighting the need for improved model refinement guided by AcoustiTrace's diagnostics.
Even the best audio-video generators struggle with basic acoustic principles, revealing a critical gap in their performance.
Recent audio-video generators can produce semantically plausible and apparently synchronized sound, yet may still violate the acoustic processes implied by visible events and environments. Existing benchmarks provide limited support for attributing such violations to particular acoustic processes and quantifying their severity. We introduce AcoustiTrace, a diagnostic benchmark that formalizes acoustic physical realism in audio-video generation. AcoustiTrace organizes text-to-audio-video (T2AV) and image-to-audio-video (I2AV) evaluation around the acoustic process, covering sound generation, propagation environment, and acoustic reception through eight dimensions grounded in measurable acoustic quantities. Based on these evaluation dimensions, we construct a large-scale dataset organized around acoustic mechanisms, comprising real-world audio-video recordings and acoustically annotated RGB-D observations, and use it to develop targeted prompt suites and validated evaluators. Experiments reveal that even leading generators still struggle with fundamental acoustic processes despite producing plausible sound events. Finally, we show that the diagnostics AcoustiTrace provides for specific acoustic relations can guide model refinement toward more physically faithful audio and open new directions for incorporating acoustic principles into training objectives, reward modeling, and candidate selection.