ETHKITMay 27, 2026arXiv:2605.28227

Why We Need Speech to Evaluate Speech Translation

Maike Züfle, Danni Liu, Vilém Zouhar, Jan Niehues

AI Summary

The authors demonstrate that current speech translation evaluation metrics, including text-based and speech-based methods, fail to adequately assess speech-specific information like gender agreement and prosody. They train SpeechCOMET, a quality estimation model with speech encoders, and evaluate a SpeechLLM, finding that while they perform well on standard quality estimation, they still struggle with speech-specific nuances. The analysis identifies issues with encoder preservation of speech features, model reliance on speech source signals, and insufficient training data, advocating for dedicated speech-focused resources.

Key Contribution

Current speech translation evaluation metrics are blind to critical speech-specific information, even when given the audio signal.

Abstract

Speech translation models are increasingly capable of preserving speech-specific information (e.g., speaker gender, prosody, and emphasis), yet evaluation metrics remain blind to such phenomena. We meta-evaluate both text- and speech-based quality estimation metrics on two contrastive datasets targeting gender agreement and prosody, and find that both fall short, even when given direct access to the speech signal. We then train SpeechCOMET, a family of quality estimation models with speech encoders, and evaluate a state-of-the-art SpeechLLM as a judge. Both match or exceed text-based COMET on standard quality estimation, but neither consistently assesses speech-specific phenomena. We identify three causes: (1) speech-specific features are not reliably preserved in current encoders, (2) models tend to ignore the speech source signal, and (3) quality estimation training data contains too few relevant examples. We release all models and code, and argue that progress requires dedicated speech-specific training data and models that genuinely condition on speech.

Eval Frameworks & Benchmarks Natural Language Processing Speech & Audio

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Why We Need Speech to Evaluate Speech Translation

Related Papers