Search papers, labs, and topics across Lattice.
This paper introduces MRMAD, a novel benchmark designed to evaluate the ability of large audio-language models (LALMs) to perceive and reason about audio degradation across multiple rounds of dialogue. By framing the evaluation as multi-turn interactions with various audio inputs, MRMAD assesses models on their capacity to identify degradation types, compare severity, and articulate changes in audio quality. The findings indicate that while LALMs can recognize basic content, they struggle significantly with diagnosing and reasoning about audio degradations, highlighting a critical gap in their understanding of acoustic phenomena.
Current LALMs can identify basic audio content but often fail to diagnose and reason about degradation, revealing a significant blind spot in their capabilities.
Large audio-language models (LALMs) have shown promising progress in understanding speech, music, and general sound events, yet their ability to reason about how audio signals are degraded remains underexplored. Existing benchmarks primarily evaluate semantic understanding, event recognition, or high-level audio reasoning, leaving a basic question unanswered: Do LALMs understand the differences in audio quality? We introduce MRMAD, a Multi-Round Multi-Audio Degradation benchmark for evaluating audio degradation perception and understanding in LALMs. MRMAD spans speech, music, and sound, and frames evaluation as multi-turn dialogues over multiple audio inputs, requiring models to identify degradation types, compare severity, and perceive corruption changes across turns. Unlike current single-turn audio-language benchmarks, MRMAD evaluates whether LALMs can maintain consistent degradation hypotheses with new evidence and explain low-level acoustic phenomena in natural language. Through a systematic evaluation of 18 representative LALMs from non-thinking to reasoning and Omni models, we find that current models often recognize coarse content while failing to diagnose, compare, or reason about degradations reliably. MRMAD reveals an important yet overlooked aspect of audio-language understanding and provides a diagnostic foundation for building future LALMs that are robust to real-world acoustic conditions.