Search papers, labs, and topics across Lattice.
The authors develop EviSI, an LLM evaluation agent that adapts Multidimensional Quality Metrics (MQM) error-tagging principles to assess semantic fidelity and oral expression in simultaneous interpreting. This addresses a core failure of metrics like BLEU and COMET, which inappropriately penalize the valid summarization and reformulation strategies necessary to manage streaming latency. At the system level, EviSI significantly outperforms conventional baselines, achieving a mean Kendall agreement with human rankings of 0.707 for English-to-Chinese and 0.467 for Chinese-to-English, though segment-level agreement remains mixed.
Standard metrics routinely misclassify valid real-time summarization as translation failure, but structuring LLM evaluation around deterministic MQM error penalties recovers human system rankings with up to 0.707 Kendall concordance.
Simultaneous speech-to-speech translation requires understanding, translation and spoken delivery while the source stream continues. To support timely delivery and limit accumulated delay, systems adopt reformulation and summarization, which can preserve meaning while departing from written references. BLEU and COMET may not reliably distinguish such variation from semantic loss. We introduce EviSI, a large language model evaluation agent adapting the error analysis and penalty principles of Multidimensional Quality Metrics (MQM). It constructs shared source evidence, assesses semantic fidelity and oral expression, reconciles overlapping errors and scores deterministically. EviSI recovers the aggregate human system ranking for English to Chinese. Mean Kendall agreement with human system rankings within corpora reaches 0.707 for English to Chinese and 0.467 for Chinese to English, exceeding evaluated baselines. An extension across five directions shows positive concordance with COMET without human ratings. Individual output agreement with humans remains mixed.