Search papers, labs, and topics across Lattice.
This paper introduces VocalRender, a score-native singing voice synthesis system that synthesizes singing directly from lyrics, pitches, symbolic note values, and tempo, thereby streamlining the composition process. By employing an interleaved lyric-note representation and an autoregressive diffusion model, VocalRender eliminates the need for explicit duration prediction while generating continuous acoustic latents. Trained on a comprehensive 2,300-hour singing dataset, it significantly outperforms existing systems, achieving a notable improvement in naturalness and speaker similarity across various benchmarks.
VocalRender achieves a remarkable $0.42$ improvement in naturalness over the strongest baseline, revolutionizing singing voice synthesis for real-world composition.
Existing singing voice synthesis systems often require predefined durations, explicit duration prediction, or time-aligned acoustic guidance, which limits their compatibility with practical composition workflows. We propose VocalRender, a score-native system that directly synthesizes singing from lyrics, pitches, symbolic note values, and tempo. It uses an interleaved lyric--note representation and an autoregressive diffusion model to generate continuous acoustic latents while predicting the output length, eliminating the need for explicit duration prediction. Trained on a 2,300-hour singing dataset, VocalRender achieves strong intelligibility, strong melody control, and high speaker similarity across both in-domain and out-of-domain benchmarks. Notably, it outperforms the strongest baseline by $0.42$ points in naturalness CMOS, demonstrating the effectiveness of our proposed score-native architecture.