Search papers, labs, and topics across Lattice.
This paper introduces an extended DAIEN-TTS framework for zero-shot text-to-speech synthesis that effectively disentangles and models speech, background noise, and reverberation, allowing for independent control over timbre and acoustic environment. By employing a flow-matching-based approach and a speech-environment separation module, the system enhances the naturalness and speaker similarity of synthesized speech while maintaining high fidelity to environmental characteristics. Experiments demonstrate that DAIEN-TTS outperforms previous models in generating personalized speech with realistic noise and reverberation, showcasing its potential for real-world applications.
DAIEN-TTS achieves unprecedented control over speech synthesis by disentangling speaker characteristics from environmental factors, resulting in highly natural and contextually relevant audio outputs.
Recent zero-shot text-to-speech (TTS) systems achieve remarkable naturalness and speaker similarity but typically require high-quality speaker prompts and either strip away or entangle the acoustic environment with speaker characteristics, limiting their real-world applicability. We present an extended DAIEN-TTS, an environment-aware zero-shot TTS framework that disentangles and jointly models speech, background noise, and reverberation, enabling independent control over timbre and acoustic environment through separate speaker and environment prompts. Built upon the flow-matching-based F5-TTS, it uses a speech-environment separation module to decompose environmental speech into speech, noise, and reverberation components, which are injected into the Diffusion Transformer for environment-aware generation. Training uses simulated data constructed by mixing clean speech with noise and room impulse responses, together with a cross-speaker conditioning strategy that suppresses speaker information leakage from the environment branch. When real-world data are available, the system can be further fine-tuned to bridge the simulated-to-real domain gap.At inference, a triple classifier-free guidance mechanism enables fine-grained control over speech, noise, and reverberation, and a signal-to-noise-ratio adaptation strategy aligns the synthesized speech with the environment prompt. Experiments on simulated and real-world test sets show that DAIEN-TTS generates environmental personalized speech with high naturalness, strong speaker similarity, and faithful noise and reverberation reproduction, while offering controllability beyond prior environment-aware TTS systems.