Search papers, labs, and topics across Lattice.
This study critically evaluates the effectiveness of ASR-roundtrip evaluation as a proxy for TTS intelligibility in Chinese news contexts, revealing that it can obscure significant reading errors tied to contextual or conventional nuances. Through a detailed audit of 110 high-risk TTS cases, the authors identified 46 instances where ASR masked false negatives, alongside 9 exposed errors, highlighting the limitations of relying solely on ASR for accurate TTS assessment. The findings underscore the necessity for more robust evaluation methods, as demonstrated by the superior performance of Qwen3-ASR in recovering masked cases compared to Paraformer.
ASR-roundtrip evaluation can miss nearly half of the critical reading errors in Chinese news TTS, revealing a significant gap in current assessment methods.
ASR-roundtrip evaluation is widely used as a scalable proxy for text-to-speech (TTS) intelligibility, but it can produce false negatives for reading errors perceived by listeners. We study Chinese news TTS spans whose correct reading depends on context or domain conventions, such as sports scores, aircraft models, technical units, and membership names. In these cases, Raw TTS can choose a plausible but wrong reading while ASR transcribes the audio as the intended or surface-correct text. A targeted audit over 110 high-risk MiMo TTS cases, reported with a complete denominator, confirms 46 masked false negatives, 9 exposed TTS errors, and 55 cases with no Raw TTS error. A span-isolation diagnostic re-exposes 18/46 previously masked errors. A Raw-only CosyVoice audit on the same targeted pool confirms 51 masked cases. Across the 97 TTS-specific audio files labeled confirmed masked across the two audits, Qwen3-ASR surface-recovers 40 cases, whereas Paraformer does so in only 2. The results suggest that ASR-roundtrip is useful for screening but insufficient as standalone ground truth for Chinese news reading-risk evaluation.