Search papers, labs, and topics across Lattice.
This study investigates the semantic token codebook of the Mimi codec, revealing that the initial ABX experiment does not accurately reflect the mapping of semantic tokens to phonetic realizations. By realigning the Mimi representations with TIMIT corpus transcriptions, the authors demonstrate a clear mapping of the 2048 token IDs to various phonetic levels, including quadphone and subphone realizations. This finding enhances the understanding of how semantic representations can be effectively linked to phonetic structures in neural language models.
Realigning the Mimi codec reveals a precise mapping of semantic tokens to phonetic realizations, challenging previous assumptions about their representation.
In this paper, we focus on the dictionary of 2048 tokens used in Mimi semantic token codebook, the neural codec of the Moshi language model. We show that the ABX experiment carried out with Mimi fails to capture the mapping of the semantic tokens to phone realisations. By realigning Mimi representations to the TIMIT corpus transcriptions, we show that the 2048 tokens IDs of the semantic codebook map to quadphone, triphone, biphone, phone and subphone realisations.