Search papers, labs, and topics across Lattice.
DuplexGen introduces a novel framework for synthetic dialogue speech that separates the generation of content, timing, and acoustics, allowing for more natural conversational dynamics. By utilizing a large language model (LLM) to create dialogue scripts and employing two full-duplex conversational models that interact in real time, the framework enables timing to emerge organically rather than being artificially imposed. The resulting high-fidelity text-to-speech synthesis demonstrates significant improvements in mimicking real dialogue dynamics compared to traditional methods, as evidenced by a newly constructed patient-clinician conversational speech corpus with detailed annotations.
DuplexGen achieves more authentic conversational dynamics by allowing timing to emerge organically rather than relying on rigid, handcrafted rules.
Synthetic conversational speech has become an important resource for developing and evaluating conversational speech systems. However, existing dialogue synthesis pipelines typically generate dialogue content first and then insert interruptions, overlap, and backchannels using handcrafted markers or timing rules, making conversational timing prescribed rather than interaction-driven. We present DuplexGen, a dialogue synthesis framework that explicitly decouples content, timing, and acoustics. An LLM first generates the dialogue script, and then two full-duplex conversational models perform the script while listening to each other in real time. This allows conversational timing to emerge naturally while preserving the scripted content. Finally, a high-fidelity text-to-speech model re-renders the interaction without altering its timing. As a demonstration of the proposed framework, we construct a patient--clinician conversational speech corpus with construction-time annotations, including word timestamps, speaker activity, overlap regions, and interaction events. Experimental results show that the proposed framework produces conversational dynamics closer to real dialogue than conventional stitching-based synthesis.