Search papers, labs, and topics across Lattice.
This paper introduces SingDance, a novel video diffusion framework that enables the generation of personalized singing-and-dancing videos by conditioning on audio, a reference image, and a text prompt. By treating vocal articulation as a semantic role鈥攅ither as a source or listener鈥攖he method effectively combines music-conditioned body motion with speech-driven articulation, allowing for compositional zero-shot generation. Experiments reveal that SingDance achieves impressive motion-beat alignment, visual fidelity, and lip synchronization while requiring significantly fewer parameters than existing speech-driven models.
Compositional zero-shot singing-and-dancing generation is now possible, allowing for unprecedented flexibility in video creation from audio and visual prompts.
Generating personalized dance videos from a reference image, text prompt, and audio track requires music-conditioned body motion. Singing-and-dancing adds a second requirement: the visible subject must also articulate the vocals. Existing music-conditioned methods focus primarily on choreography, while speech-driven models generally assume that the visible subject produces the input voice, leaving this combined setting largely underexplored. We introduce SingDance, a unified video diffusion framework that formulates controllable vocal articulation as a semantic role: the visible subject is either the source, who produces the vocal signal, or the listener, who receives it from an off-screen performer. Hard-compact routing selects task-relevant speech, music, and role conditions, which are composed through frame-wise joint audio injection; source and listener retain the same speech pathway. Training uses asymmetric supervision: on-screen speaking and curated off-screen conversational-response videos establish role control, while instrumental and song-based dancing-only videos establish music-conditioned body motion. The target Song/Source configuration is never observed during training. At inference, assigning the source role to a song composes separately learned articulation and song-conditioned dance capabilities, enabling compositional zero-shot singing-and-dancing. Experiments demonstrate strong motion--beat alignment and visual fidelity, reliable paired switching of vocal articulation while preserving music-aligned body motion, and highly competitive lip synchronization with substantially fewer generation-time parameters than the strongest speech-driven baseline evaluated.