Search papers, labs, and topics across Lattice.
This paper introduces the Audio-Visual Modal Sound Field (AV-MSF), a novel approach for reconstructing object-level acoustic representations using only a few impact sound recordings alongside multi-view images. By integrating 3D Gaussian Splatting with dense visual features, AV-MSF captures the acoustic properties of objects in a geometry-aware manner, significantly enhancing the efficiency of sound modeling without the need for extensive datasets or costly simulations. Experimental results demonstrate that AV-MSF achieves state-of-the-art performance in impact sound rendering and enables practical applications such as contact localization and object sound editing.
AV-MSF achieves state-of-the-art impact sound rendering with minimal data, revolutionizing how we model acoustic properties of objects.
While modern 3D reconstruction excels at modeling object geometry and appearance, it largely ignores the rich acoustic cues revealed through physical interaction. Object impact sounds convey material, stiffness, and structural properties that complement vision, yet existing impact sound modeling approaches either rely on expensive physics-based simulation or require large datasets to generalize in a purely data-driven manner. We introduce Audio-Visual Modal Sound Field (AV-MSF), a novel object-level acoustic representation reconstructed from multi-view images and only a few impact sound recordings. AV-MSF builds on 3D Gaussian Splatting integrated with dense 3D visual feature to provide a strong geometry-aware prior, and represents the impact sound field using compact, physically meaningful modal parameters, enabling robust few-shot reconstruction. Experiments on two real-world datasets show that AV-MSF achieves state-of-the-art impact sound rendering, outperforming both physics-based and data-driven baselines. Furthermore, we demonstrate downstream applications enabled by our representation, including contact localization and object sound editing.