Search papers, labs, and topics across Lattice.
This study conducts a comprehensive statistical analysis of token sequences generated by 13 neural audio codecs (NACs) across various architectures and noise conditions, revealing that acoustic conditions and codec types significantly influence token statistics. The research identifies unigram entropy as the most indicative metric of codec meta-category, while also demonstrating that clean-to-noise JSD correlates with mel-cepstral distortion, particularly under challenging DEMAND noise conditions. Notably, the findings highlight distinct degradation patterns in RVQ codecs compared to non-VQ codecs, offering valuable insights into the behavior of NACs under different acoustic environments.
Unigram entropy emerges as a key predictor of codec performance, revealing stark differences in how various architectures handle noise conditions.
Neural audio codecs (NACs) convert speech into discrete token sequences, and prior work has reported that these sequences follow language-like statistical laws. This paper analyzes the token statistics of 13 NACs spanning multi-codebook residual vector quantization (RVQ), single-codebook VQ, and non-VQ designs, evaluated on three corpora under clean, white-noise, and real-world DEMAND-noise conditions. Zipf and Heaps parameters, unigram entropy, codebook occupancy, and Jensen-Shannon divergence (JSD) are estimated from matched token samples with explicit fit-validity safeguards and family-conditional $n$-gram orders. Corpus identity explains little variance in any metric, whereas acoustic condition and quantizer meta-category dominate in a metric-dependent way, and unigram entropy is the metric most strongly associated with meta-category. Clean-to-noise JSD computed at a common unigram order is associated with mel-cepstral distortion most clearly under DEMAND noise. The collapse and explosion degradation signatures previously reported for RVQ codecs concentrate in RVQ cells under white and DEMAND noise, respectively; explosion also occurs in non-VQ codecs, and single-codebook VQ codecs shift in occupancy and distribution shape without either signature. These results provide architecture-conditioned conventions for applying language-statistical analysis to NAC tokens.