Search papers, labs, and topics across Lattice.
This study systematically analyzes the language sensitivity of codec-based self-supervised learning (SSL) models using neural audio codec tokens (NACs). The findings reveal that while the downstream performance of these models is largely insensitive to the language used for NAC training, it is significantly affected by the language of SSL pre-training. This indicates that a single NAC can be effectively reused across different languages, but emphasizes the necessity of aligning the SSL pre-training language with the target language for optimal performance.
Downstream performance hinges on SSL pre-training language, not the NAC training language, enabling efficient cross-language model reuse.
Neural audio codecs (NACs) have become popular for obtaining speech representations as discrete tokens. Beyond compression, discrete tokens can be used to train self-supervised learning (SSL) models. Such models, referred to as codec-based SSL models, reduce data storage and computational cost, enabling scalable SSL pre-training. However, their language sensitivity remains unclear. When the language changes, codec-based SSL models may require retraining, which undermines their efficiency. In this paper, we present a systematic analysis of language sensitivity by varying either the NAC training language or the SSL pre-training language while keeping the other fixed. Experimental results show that downstream performance is insensitive to the NAC training language but strongly dependent on the SSL pre-training language. These findings suggest that a single NAC can be reused across languages, while aligning the SSL pre-training language with the target language is crucial.