Search papers, labs, and topics across Lattice.
This paper introduces FineCombo-TTS, a novel framework for controllable text-to-speech synthesis that integrates both reference speech and text descriptions to achieve enhanced flexibility and precision in acoustic attribute control. By employing a Conditional Flow Matching-based Speech Variance Predictor, the model effectively learns a unified acoustic representation, allowing for fine-grained transformations between reference and target speech guided by textual input. Experimental results indicate that FineCombo-TTS significantly outperforms existing methods in terms of control and expressiveness, marking a substantial advancement in TTS technology.
FineCombo-TTS achieves unprecedented precision in speech synthesis by seamlessly integrating text descriptions with reference speech, allowing for fine-grained control over acoustic attributes.
Controllable text-to-speech (TTS) has become a key research focus. However, methods based on either reference speech or text descriptions lack flexibility and precise control, and recent joint approaches remain loosely coupled, with speech modeling timbre and text controlling global style. We propose FineCombo-TTS, a unified framework for speech synthesis grounded in reference speech and guided by text descriptions, enabling flexible and precise control over acoustic attributes. Instead of explicit attribute disentanglement, we learn a unified acoustic representation and introduce a Conditional Flow Matching (CFM)-based Speech Variance Predictor to model fine-grained reference-to-target transformations guided by text descriptions. To support relative attribute control, we construct FineEdit, a structured paired dataset that explicitly encodes source-to-target attribute variations. Experiments demonstrate that our approach achieves flexible, precise, and expressive controllable TTS.