Search papers, labs, and topics across Lattice.
This paper presents a novel approach to spatial audio generation by leveraging visually guided First-Order Ambisonics (FOA) for real-world speech scenes, addressing the limitations of high-quality spatial capture. The authors introduce the YT-SPEECH dataset, which consists of aligned $360^\circ$ video and omnidirectional audio, and develop a two-stage Localizer-Renderer framework that enhances the reconstruction of directional FOA components. Experimental results demonstrate significant improvements in reconstruction fidelity, spatial accuracy, and perceptual speech quality compared to existing methods and ablated models.
By integrating visual cues with audio processing, this framework achieves unprecedented spatial audio fidelity in complex speech environments.
Spatial audio is a key component of immersive $360^\circ$ media, yet high-quality spatial capture remains limited in real-world speech-dominant scenes. We study visually guided First-Order Ambisonics (FOA) speech spatialization in the wild: given aligned $360^\circ$ video and an omnidirectional audio track, we recover the missing directional FOA components. To support this task, we introduce YT-SPEECH, a speech-oriented $360^\circ$ video-FOA dataset curated from YouTube. We propose a two-stage Localizer-Renderer framework, where an audio-visual segmentation backbone provides frame-wise spatial heatmaps and a conditional complex-domain U-Net reconstructs directional FOA signals from the omnidirectional channel. A confidence-based gating strategy stabilizes conditioning under ambiguous acoustic conditions. Experiments show improved reconstruction fidelity, spatial accuracy, and perceptual speech quality relative to ablated variants and prior approaches.