Search papers, labs, and topics across Lattice.
This paper adapts a diffusion-based generative model, initially designed for multi-instrument music synthesis, to perform voice conversion for both speech and singing. By integrating phonetic posteriorgrams (PPGs) and pitch contours into the conditioning process, the model achieves comparable or superior performance to dedicated voice conversion systems in terms of naturalness and performer similarity. However, the study also identifies challenges related to phonetic fidelity and vocal quality when using instrumental training data, while showcasing the effectiveness of off-the-shelf feature extractors for self-supervised training.
Voice conversion can now leverage a unified diffusion model, achieving high naturalness and performer similarity while navigating the complexities of speech and singing.
Recent diffusion-based generative models have achieved strong results in domain-specific audio generation tasks such as speech, singing, and instrumental music synthesis. However, these models are typically specialized and do not generalize well to mixed or intermediate audio types. In this work, we adapt a diffusion-based model originally designed for multi-instrument music synthesis to voice conversion, covering both speech and singing within a unified framework. Specifically, we extend musical note-based conditioning to include phonetic posteriorgrams (PPGs) and pitch contours, and reinterpret timbre conditioning as speaker or singer identity via feature-wise linear modulation. Experiments show that the adapted model matches or surpasses a dedicated voice conversion system in terms of naturalness and performer similarity, while maintaining accurate pitch control across speech and singing. At the same time, we observe limitations in phonetic fidelity and a degradation in vocal quality when incorporating instrumental training data. Furthermore, we demonstrate that off-the-shelf feature extractors provide effective conditioning signals, enabling large-scale self-supervised training without manual annotations. These results highlight the potential of cross-domain model transfer towards unified audio generation systems capable of handling speech, singing, and music. Qualitative samples can be found on our project page: https://benadar293.github.io/voice-conversion