Search papers, labs, and topics across Lattice.
UR-BERT is a novel text-to-speech (TTS) encoder that utilizes a Romanization-based approach to scale TTS systems to 495 languages, overcoming the limitations of traditional grapheme-to-phoneme methods. By integrating a speech token prediction objective during training, UR-BERT enhances phonetic accuracy and text-speech alignment, leading to improved performance across diverse languages. Experimental results indicate that TTS systems leveraging UR-BERT significantly outperform existing text encoder baselines and exhibit robust generalization capabilities for previously unseen languages.
Scaling TTS to 495 languages is now possible with a single Romanization-based encoder that outperforms traditional methods in phonetic fidelity.
We propose UR-BERT, a Romanized transcription-based text-to-speech (TTS) encoder for massively multilingual TTS systems. Conventional grapheme-to-phoneme (G2P)-based approaches are limited to around 100 languages due to the availability of reliable G2P resources. In contrast, UR-BERT scales to 495 languages by unifying diverse writing systems into a shared Romanization representation. To further enhance phonetic fidelity and text-speech alignment, we introduce a speech token prediction objective during training, which encourages the encoder to learn speech-aware phonetic representations in a data-efficient manner. Experiments show that TTS systems built on UR-BERT consistently outperform recent text encoder baselines across a wide range of languages and resource conditions, and demonstrate strong generalization to unseen languages.