Search papers, labs, and topics across Lattice.
The WanSong model introduces a novel diffusion-based approach for generating high-fidelity, long-form songs, overcoming the limitations of traditional autoregressive and cascaded methods. This model can produce multilingual songs of up to 5 minutes in a single run, outputting both vocals and background music simultaneously. Key findings demonstrate that WanSong not only enhances generation efficiency but also facilitates faster inference and customization for various editing tasks through step-distillation techniques.
WanSong achieves high-fidelity song generation in a single diffusion pass, revolutionizing how we think about music creation models.
Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address these needs, we present \textbf{WanSong}, a simple yet powerful approach for long-form, commercial-grade song generation. Unlike autoregressive (AR) and cascaded multi-stage pipelines (\eg, AR followed by diffusion), \textbf{WanSong} is a pure diffusion-based model that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems (vocals and background music) in a single run. In addition, our diffusion framework enables faster inference through step-distillation, and offers an efficient pathway for fine-tuning and customization to support downstream editing tasks.