Search papers, labs, and topics across Lattice.
This paper introduces Iterative Semantic Token Purification (ISTP), a novel training procedure that alternates between speech-to-unit (S2U) and text-to-unit (T2U) models to refine semantic speech tokens. By leveraging text predictability, the method enhances the alignment between S2U and T2U outputs, effectively minimizing speaker and duration variations while preserving linguistic content. Experiments demonstrate that the refined tokens significantly improve intelligibility in voice conversion and text-to-speech synthesis, achieving better cross-speaker consistency and reduced speaker information leakage.
Semantic speech tokens can be refined to enhance intelligibility and consistency across speakers, with significant implications for voice conversion and TTS applications.
Semantic speech tokens should preserve linguistic content while suppressing speaker- and duration-dependent variation inherited from acoustic inputs. We propose Iterative Semantic Token Purification (ISTP), an alternating speech-to-unit (S2U) and text-to-unit (T2U) training procedure guided by text predictability. Starting from an initial S2U tokenizer, each iteration trains a T2U model on its deduplicated token sequences. The decoded T2U predictions then serve as connectionist temporal classification targets for a newly initialized S2U model, whose outputs supervise the next T2U model. This cycle progressively aligns the two token generators and biases the token space toward information recoverable from text. Experiments on Mandarin and English show substantially improved S2U--T2U agreement. Independently trained de-tokenizers further show that the refined S2U and T2U tokens retain sufficient content for high-intelligibility voice conversion and text-to-speech synthesis. In voice conversion, the generated speaking rate follows the reference more closely. The refined tokens also exhibit substantially improved cross-speaker consistency and reduced probe-recoverable speaker information.