Search papers, labs, and topics across Lattice.
This study evaluates the effectiveness of different representation paradigms for predicting Mean Opinion Scores (MOS) in synthesized speech, emphasizing the limitations of self-supervised learning (SSL) models that focus primarily on semantics. By benchmarking against acoustic-only neural audio codecs (NACs) and unified NACs that combine semantic and acoustic features, the authors reveal that integrating both aspects significantly enhances performance. The key finding underscores that relying solely on semantic understanding is insufficient; a balanced approach that incorporates both semantic and acoustic fidelity is crucial for accurate MOS prediction.
Integrating semantic understanding with acoustic fidelity can dramatically elevate the accuracy of speech quality assessments, challenging the dominance of self-supervised learning models.
Mean Opinion Score (MOS) is the gold standard for evaluating synthesized speech naturalness. However, current automatic MOS predictors are dominated by self-supervised learning (SSL) models that prioritize high-level semantics, potentially compromising their ability to capture critical acoustic details. In this paper, we systematically investigate representations from three paradigms: SSLs, acoustic-only neural audio codecs (NACs), and unified NACs that integrate semantics into reconstruction-based architectures. Extensive benchmarking on the standard BVCC and multiple out-of-domain (OOD) datasets demonstrates that features synergizing semantic understanding with fine-grained acoustic modeling achieve a higher performance upper bound in speech quality assessment. Ultimately, our findings highlight that semantics alone are not enough; a dual focus on semantic content and acoustic fidelity is essential for robust MOS prediction.