Search papers, labs, and topics across Lattice.
The RW-Voice-EQ Bench introduces a comprehensive evaluation framework for voice AI systems that assesses performance across multiple dimensions, including text-to-speech, speech-to-speech, speech understanding, and automatic speech recognition. This benchmark reveals significant dimension-specific performance discrepancies, highlighting that traditional metrics often overlook critical aspects such as vocal affect and robustness in real-world conditions. The findings underscore the necessity of a multidimensional approach to accurately evaluate and improve voice AI capabilities beyond simplistic aggregate scores.
Voice AI systems may excel in one area while failing in others, revealing a critical need for multidimensional evaluation beyond traditional benchmarks.
Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation. To this end, we introduce the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech (TTS), speech-to-speech (STS), speech understanding (SU), and automatic speech recognition (ASR). Our evaluations indicate that performance is highly dimension-specific. For TTS, naturalness, expressiveness, identity stability, and reliability are largely independent evaluation dimensions. For STS, access to audio does not guarantee use of vocal affect, and some agents remain largely transcript-driven. For SU, models perform unevenly across paralinguistic tasks. For ASR, real world accent, emotion, noise, and conversational conditions expose failures that are not captured by established clean-speech benchmarks. Together, these results show that voice AI should be evaluated as a profile of acoustic, expressive, interactional, and robustness capabilities rather than by a single aggregate score.