Search papers, labs, and topics across Lattice.
This study investigates the geometric properties of emotion control in text-to-speech (TTS) models by comparing speech language model (SLM) and conditional flow-matching (CFM) modules as activation steering sites for mixed-emotion synthesis. Through linear probing and local intrinsic dimensionality analysis, the authors reveal that SLM provides a clean, low-dimensional subspace for emotion-specific representations, while CFM suffers from poor generalization due to entangled speaker-emotion representations. The findings indicate that while joint steering can enhance emotion intensity, it compromises proportional control and speech quality, underscoring the significance of representation geometry in TTS systems.
SLM's low-dimensional emotion-specific subspace outperforms CFM's entangled representations, revealing critical trade-offs in emotion steering for TTS models.
While prior work has explored emotion control in hybrid text-to-speech systems, the geometric properties of these modules, and their implications for steerability, remain poorly understood. We present the first comparative study of speech language model (SLM) and conditional flow-matching (CFM) modules as activation steering sites for mixed emotion speech synthesis. We first characterize emotion representations using linear probing and local intrinsic dimensionality (LID), and then evaluate single-site and joint steering for mixed-emotion synthesis. Our results show that SLM offers a clean, low-dimensional emotion-specific subspace with strong speaker--emotion disentanglement, while CFM exhibitspoor cross-speaker generalization due to speaker--emotion entanglement. Joint steering increases emotion intensity but degrades proportional control and speech quality on in-distribution data. These findings provide practical guidance for multi-site activation steering in hybrid TTS systems and highlight the importance of representation geometry in controllable speech generation.