Search papers, labs, and topics across Lattice.
This paper introduces Diff-VS, a diffusion model for vocal separation based on the Elucidated Diffusion Model (EDM) framework. The model operates on complex STFT spectrograms and incorporates a U-Net architecture tailored for music processing. Diff-VS achieves objective metrics comparable to discriminative baselines and perceptual quality on par with state-of-the-art systems, suggesting generative models can be competitive in music source separation.
Generative models can now compete with discriminative methods in music source separation, thanks to a diffusion-based architecture with music-informed design.
While diffusion models are best known for their performance in generative tasks, they have also been successfully applied to many other tasks, including audio source separation. However, current generative approaches to music source separation often underperform on standard objective metrics. In this paper, we address this issue by introducing a novel generative vocal separation model based on the Elucidated Diffusion Model (EDM) framework. Our model processes complex short-time Fourier transform spectrograms and employs an improved U-Net architecture based on music-informed design choices. Our approach matches discriminative baselines on objective metrics and achieves perceptual quality comparable to state-of-the-art systems, as assessed by proxy subjective metrics. We hope these results encourage broader exploration of generative methods for music source separation