Search papers, labs, and topics across Lattice.
This paper introduces PhysWave, a physics-guided latent diffusion model designed for controllable text-to-First-Order Ambisonics (FOA) audio generation. By integrating a shared waypoint-caption representation and incorporating differentiable acoustic priors for direction and distance consistency, PhysWave enhances spatial audio generation while maintaining high audio quality. The model's performance is validated through a newly constructed 300K-clip FOA dataset, demonstrating significant improvements in spatial consistency and usability for users in the gaming and film industries.
PhysWave achieves spatially consistent audio generation by combining physics-based priors with a unified control framework, revolutionizing text-to-FOA applications.
Text-to-spatial audio generation, such as text-to-First-Order Ambisonics (FOA), provides a convenient way to create spatial audio for billion-dollar gaming and film industries. However, existing text-to-FOA methods are largely data-driven and may produce audio that violates acoustic relations between source direction and distance. They also separate descriptive and parametric control, forcing users to trade usability for precision. In this paper, we present PhysWave, a physics-guided latent diffusion model for controllable text-to-FOA generation. PhysWave unifies natural-language and trajectory control through a shared waypoint-caption representation, and augments diffusion training with two differentiable acoustic priors: spherical-harmonic direction consistency and inverse-square distance consistency. To support dynamic spatial generation, we further construct a 300K-clip FOA dataset with diverse sound categories and source trajectories. Extensive results show that the proposed priors help PhysWave generate spatially consistent FOA audio while maintaining competitive audio quality. Further analyses show that these physics priors improve spatial consistency during training and can also be used as inference-time guidance for training-free spatial refinement.