Search papers, labs, and topics across Lattice.
This paper introduces CineDub, a unified diffusion-based model that enables precise multi-speaker dialogue dubbing directly from uncropped videos, addressing the limitations of existing hierarchical and holistic methods. By employing the Implicitly-Coupled Holistic Conditioning (ICHC) paradigm, CineDub effectively resolves speaker ambiguity and enhances temporal alignment through cross-modal training. Experimental results demonstrate that CineDub achieves state-of-the-art performance in both single-speaker dubbing and multi-speaker dialogue scenarios, while also facilitating coherent joint speech and audio generation.
CineDub achieves unprecedented accuracy in multi-speaker dialogue dubbing from uncropped videos, setting a new standard for video dubbing technology.
Automatic video dubbing in the wild remains fundamentally limited by two competing constraints: hierarchical methods depend on brittle, multi-stage preprocessing pipelines that severely restrict data scalability and practical deployment, while holistic approaches operating on uncropped video suffer from weak temporal alignment and speaker-utterance ambiguity in multi-speaker settings. To overcome these limitations, we propose CineDub, a unified diffusion-based model that achieves precise multi-speaker dialogue dubbing directly from uncropped videos, without face cropping or speaker diarization. Central to our approach is the Implicitly-Coupled Holistic Conditioning (ICHC) paradigm, where holistic visual representations and a semantic-bundled transcription format are encoded independently, yet implicitly coupled through cross-modal training to resolve speaker ambiguity and enable precise multi-speaker multi-turn dialogue dubbing. Building on the unified temporal cues captured by holistic visual features, we further extend CineDub to joint speech and audio generation. We introduce an Ambient-to-Linguistic Curriculum Learning (ALC) to mitigate sub-task degradation, and a decoupled textual branch control mechanism to resolve cross-prompt interference during simultaneous generation. We also release two in-the-wild benchmarks, CineDub-Multi for multi-speaker dialogue dubbing and CineDub-SA for video-to-speech-and-audio (V2SA) generation, to enable evaluation under realistic conditions. Experiments show that CineDub achieves state-of-the-art results on established single-speaker dubbing and video-to-audio benchmarks while excelling in multi-speaker dialogue dubbing and acoustically coherent joint generation.