Search papers, labs, and topics across Lattice.
This paper audits the reliance of large audio-language models (LALMs) on protocol-level shortcuts in speech evaluation, revealing that high agreement with human ratings can be misleading. The authors analyze three common evaluation protocols and find that several LALMs, including Qwen3-Omni-Thinking, exhibit significant inaccuracies when relying on specialist labels or structured descriptions instead of directly analyzing audio. These findings underscore the need for a joint assessment of models and evaluation protocols to ensure the validity of LALM judges, particularly highlighting that aggregate agreement may not reflect true performance.
LALMs can achieve high agreement with human evaluators while still relying on misleading shortcuts, risking the integrity of speech evaluations.
Large audio-language models (LALMs) are increasingly used as automatic judges for speech evaluation. However, high agreement with human ratings does not guarantee that their verdicts are grounded in the audio. A judge may instead rely on specialist labels or reference data supplied by the evaluation protocol itself, taking a shortcut in place of listening to the audio. In this paper, we audit such protocol-level ``shortcuts''in LALM judges across three common deployment protocols: feature-blueprint judging, where the audio is replaced by a structured text description of acoustic features, reference-conditioned judging, and pairwise A/B comparison. Across six judges and four attributes, we find that several LALMs rely on protocol-level shortcuts. For example, in feature-blueprint judging, incorrect specialist labels reduce five judges'emotion accuracy to 0.10 or below, and in concatenated A/B comparisons, Qwen3-Omni-Thinking often picks the same slot regardless of order swaps. These results indicate that aggregate agreement can overstate the validity of LALM judges unless the model and the evaluation protocol are assessed jointly, and that each model-protocol pair should be evaluated with a matched shortcut probe.