Search papers, labs, and topics across Lattice.
This paper introduces AutoSIFT, a framework for controllable speech generation that allows for fine-grained editing of speaking styles while maintaining the integrity of non-verbal prosody and speaker-specific nuances. By decomposing speaking style into known text-describable categories and unknown residual styles, AutoSIFT effectively enables users to modify specific attributes such as emotion or age without losing the subtleties of the original speech. The key result demonstrates that AutoSIFT can achieve natural and expressive speech generation, significantly enhancing customization in applications like film dubbing and voice acting.
AutoSIFT allows for precise control over speech styles, enabling users to modify attributes like emotion while preserving the nuanced prosody of the original voice.
State-of-the-art text-to-speech (TTS) models achieve impressive naturalness and expressiveness, yet fine-grained, disentangled control over speaking styles remains challenging. In professional scenarios such as film dubbing, game voice acting, and video content generation, users often need to modify a specific style category, such as emotion, age, or gender, while preserving all others. Existing style-controllable TTS methods typically rely on either text-described styles or speech-reference style transfer, making it difficult to jointly control explicit semantic attributes and preserve subtle, text-undescribed prosodic details. We propose AutoSIFT, a controllable speech generation framework for category-level style editing. AutoSIFT decomposes speaking style into known text-describable categories and unknown residual styles that capture non-verbal prosody and speaker-specific nuances. It consists of a generalized Style Disentangler, which extracts category-aware style prototypes from reference speech, and an Arbitrary Style Infiller, which selectively infills unspecified style categories from the reference. By replacing only text-specified style categories while preserving residual speech-derived styles, AutoSIFT enables natural, expressive, and highly customizable speech generation.