Search papers, labs, and topics across Lattice.
This paper introduces CSAVocoder, a causal GAN-based spatial audio vocoder that enhances the conversion of mel-spectrograms into high-fidelity spatial audio waveforms. By integrating a Spatial Adaptor that combines multi-channel mel-spectrograms with dynamic source-listener pose information and employing a spatial consistency discriminator, the framework effectively preserves inter-channel cues while maintaining audio quality. Experimental results demonstrate that CSAVocoder achieves superior spatial fidelity and real-time performance compared to existing methods, making it suitable for streaming applications.
CSAVocoder achieves real-time spatial audio generation with improved fidelity by effectively leveraging dynamic pose information and inter-channel cues.
Spatial audio vocoders are able to convert mel-spectrograms produced by generative models into spatial audio waveforms. Most neural vocoders are designed for monaural audio, and direct extensions to spatial audio can degrade spatial quality by ignoring inter-channel cues. We present CSAVocoder, a causal GAN-based spatial audio vocoder that jointly optimizes waveform fidelity and spatial rendering. Our framework introduces a Spatial Adaptor that fuses multi-channel mel-spectrograms with dynamic source-listener pose information, together with a spatial consistency discriminator that supervises inter-channel cues. To meet real-time requirements, we design a strictly causal, stateful generator that supports efficient streaming inference with constant memory overhead. Experiments on large-scale spatial audio datasets show that CSAVocoder improves spatial fidelity at competitive audio quality and real-time performance.