Search papers, labs, and topics across Lattice.
This paper introduces FillGauss, a novel framework for generating impact sounds by integrating 3D Gaussian Splatting with internal state conditioning, addressing the overlooked influence of internal filling states on acoustic properties. The authors present the FillImpact dataset, comprising over 5,000 acoustic recordings from diverse objects with varying internal contents, which is crucial for training models that respect physical laws of sound generation. Experimental results show that FillGauss achieves high-fidelity sound synthesis that aligns with physical principles, setting a new benchmark in cross-modal audio generation.
Internal filling states significantly influence impact sound, and FillGauss captures this complexity to generate high-fidelity audio that adheres to physical laws.
Synthesizing physically plausible impact sounds from visual observations remains a great challenge in multi-modal AI. Existing 3D-aware audio generation methods primarily model the surface geometry of hollow rigid bodies. However, they fundamentally overlook internal filling states, a critical physical factor that drastically modulates acoustic resonance and damping. To address this issue, we have defined a new task called Fine-Grained Filling-Aware Impact Sound Generation. As a foundational step, we first introduce the fine-grained fill-aware dataset (FillImpact), a pioneering multi-modal collection comprising over 5,000 rigorous acoustic recordings from 88 diverse real-world objects. It captures impact interactions with varying internal contents (i.e., water, rice), a continuous range of fill levels, and distinct striker materials. Furthermore, comprehensive acoustic analysis confirms that the collected data closely aligns with established physical laws governing acoustic resonance and damping, indicating its suitability for physically grounded modeling. Building on this dataset, we propose a novel generative framework (FillGauss) that integrates 3D Gaussian Splatting (3DGS) with internal state conditioning for sound generation. By fusing 3DGS geometric features, precise 3D spatial strike coordinates, and fine-grained textual physical conditions within a latent diffusion architecture, FillGauss enables position-aware, striker-aware, and filling-aware audio generation. Extensive experiments demonstrate that our approach could generate high-fidelity impact sounds that adhere to underlying physical principles, establishing a new state-of-the-art for physically grounded cross-modal audio generation.