Search papers, labs, and topics across Lattice.
This study introduces a multi-metric benchmarking framework for evaluating text-to-speech (TTS) systems, focusing on low-resource and underrepresented languages across various speech domains. By integrating subjective and objective evaluation methods, the framework assesses four state-of-the-art TTS systems on their performance in formal, conversational, literary, and emotional speech contexts. Key findings indicate that emotional speech synthesis poses significant challenges, while conversational speech demonstrates superior acoustic fidelity, highlighting the need for tailored evaluation approaches in diverse linguistic contexts.
Emotional speech synthesis proves to be the toughest challenge for TTS systems, with performance varying dramatically across different speech domains.
Recent advances in neural text-to-speech (TTS) systems have substantially improved speech naturalness and intelligibility across many languages. However, comprehensive evaluation methodologies that jointly assess perceptual quality, speaker similarity, and acoustic fidelity across diverse speech domains remain limited, particularly for low-resource and underrepresented languages. This paper presents a reproducible, multi-metric benchmarking framework for systematic evaluation of modern TTS systems through domain-specific analysis. The proposed framework integrates complementary subjective and objective evaluation protocols and is demonstrated through a comprehensive case study on a representative low-resource language spanning four speech domains: Formal, Conversational, Literary/Storytelling, and Emotional. Four state-of-the-art TTS systems -- Indic-Parler-TTS, MMS-TTS, Microsoft Edge TTS, and Google Gemini TTS -- are evaluated using MUSHRA listening tests, ABX discrimination tests, speaker similarity scoring with Resemblyzer, and acoustic analyses based on mel-cepstral distortion (MCD) and F0 RMSE over 960 audio pairs. Results reveal substantial variation in TTS performance across speech domains, with emotional speech consistently presenting the greatest synthesis challenge (mean MCD 12.03 dB; mean F0 RMSE 889 cents), while conversational speech achieves the highest overall acoustic fidelity. Beyond the empirical findings, this work provides a reproducible evaluation framework, publicly releasing evaluation scripts, result tables, and executable Colab notebooks to support standardized benchmarking and future research on TTS evaluation for low-resource languages.