Search papers, labs, and topics across Lattice.
This paper addresses the challenge of generating direction-following text-to-speech (TTS) by introducing a scalable pseudo-triplet construction pipeline that creates training data from existing utterances. The method allows the TTS system to produce modified utterances that reflect specific performance directions while preserving the speaker's identity and linguistic content. Experimental results show that using pseudo-triplets alone yields stable modifications, and combining them with recorded data enhances direction alignment without sacrificing speaker similarity.
Pseudo-triplet construction enables TTS systems to generate nuanced voice modifications that respond to performance directions while maintaining speaker identity.
Voice actors often re-read the same script while modifying their delivery in response to performance directions. We study this setting as direction-following TTS, where a system generates a new utterance that reflects a given direction relative to a reference utterance while preserving speaker identity and linguistic content. A key challenge is the lack of training data capturing such relative modifications. To address this, we propose a scalable pseudo-triplet construction pipeline that generates~(reference utterance, direction text, modified utterance) triplets. It generates controlled style variations using an impression-controllable TTS model and uses an LLM to produce natural language directions from estimated impression differences. Experimental results demonstrate that pseudo-triplets alone enable stable speaker-preserving modification, and that combining pseudo and recorded data further improves direction alignment while maintaining speaker similarity. Audio examples are available on our demo page https://ntt-hilab-gensp.github.io/IS2026pseudo/