Search papers, labs, and topics across Lattice.
This paper introduces a personalized and culturally adaptive emotional Text-to-Speech (TTS) framework that utilizes an Interactive Genetic Algorithm to optimize individual emotion perception spaces. By moving beyond generic emotion labels and employing a tailored approach to the arousal-valence model, the system achieves a more accurate emotional expression in speech that resonates with diverse listeners. Evaluations with participants from Japan, China, and Indonesia demonstrate significant improvements in emotional alignment compared to traditional methods that rely on averaged representations.
Emotional TTS can be dramatically enhanced by personalizing expression based on individual and cultural perception, leading to more authentic interactions.
The rise of conversational AI has increased interest in emotional Text-to-Speech (TTS). Most systems rely on discrete emotion labels, which fail to capture the nuanced nature of human affect. Recent models employ dimensional representations such as Russell's arousal-valence (A-V) model, offering finer control. However, emotional perception varies across individuals and cultures, which may cause mismatches between modeled and perceived emotions. We propose a personalized and culturally adaptive emotional TTS framework that performs interactive optimization of individualized A-V perception spaces using an Interactive Genetic Algorithm. By adapting emotion representations to each listener, the system produces speech with more perceptually aligned emotional expression than models using averaged A-V values. Evaluations with Japanese, Chinese, and Indonesian participants highlight the importance of personalization and cultural adaptation for moving beyond one-size-fits-all emotional TTS.