Search papers, labs, and topics across Lattice.
This study evaluates the effectiveness of existing Automated Text-to-Speech (TTS) evaluators by deconstructing the concept of "naturalness" into ten linguistically grounded perceptual dimensions. By creating a meta-evaluation benchmark with 860 utterances annotated by trained linguists, the authors found that traditional Mean Opinion Score (MOS) predictors primarily assess acoustic signal quality, while Audio Large Language Models (Audio-LLMs) exhibit prompt-dependent detection that fails to generalize across all dimensions. The findings highlight significant limitations in current TTS evaluation methods, underscoring the need for more nuanced approaches that capture the complexity of human speech perception.
Current TTS evaluators miss the mark, with MOS predictors focusing solely on sound quality and Audio-LLMs struggling to generalize across speech dimensions.
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct"naturalness"into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLM judges reveal that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dimensions. Neither class reliably captures a breadth of linguistically structured speech errors. Our dataset, annotation schema, and evaluation code are publicly released to support more targeted and interpretable TTS evaluation.