Search papers, labs, and topics across Lattice.
This paper introduces EmoTra-TTS, a novel text-to-speech system that enables smooth intra-utterance emotion transitions by addressing the limitations of existing models that rely on static emotional embeddings. By employing a multi-pass flow blending pipeline and dual-stage Valence-Arousal-Dominance conditioning, the system achieves significant improvements in the quality of emotional transitions without adding latency or substantial parameter overhead. The results demonstrate a 30%-87% relative enhancement in emotion transition quality, validated through pairwise preference tests against state-of-the-art systems.
EmoTra-TTS achieves up to 87% improvement in emotional transition quality, redefining how TTS systems can convey nuanced human emotions in speech.
Psychological research on emotion dynamics has established that human affect is a continuous, evolving process: emotions rise, decay, and transition within seconds. Current emotional text-to-speech (TTS) systems, however, condition on a single discrete label or static embedding per utterance, fundamentally misaligning with the temporal nature of affect. While recent LLM-based TTS systems may implicitly vary prosody through text understanding, such variation is neither explicitly controllable nor precise enough for targeted intra-utterance transitions. We address three challenges: (1) a multi-pass flow blending pipeline synthesizes frame-aligned transition audio, circumventing the scarcity of natural intra-utterance transitions; (2) dual-stage Valence-Arousal-Dominance (VAD) conditioning guides prosodic planning in the LLM and acoustic realization in the flow decoder via frame-level VAD embeddings; (3) direction-magnitude decoupled injection structurally separates emotion direction from injection magnitude, preventing content degradation. EmoTra-TTS adds only +0.43% parameters with no latency overhead, achieves 30%-87% relative improvement on emotion transition quality, corroborated by 64.4%-79.5% overall win rates in pairwise preference tests against four SOTA baselines and two commercial systems.