Search papers, labs, and topics across Lattice.
This paper introduces MMAC, a comprehensive benchmark designed to enhance the evaluation of audio captioning by providing multi-dimensional assessments across various capabilities. By analyzing 5,638 audio clips from diverse sources, the benchmark evaluates model-generated captions based on information coverage and consistency with reference labels. The findings reveal significant variations in performance across different evaluation dimensions, highlighting the need for more nuanced assessments in audio captioning tasks.
MMAC reveals stark differences in audio captioning performance across multiple dimensions, challenging the adequacy of current evaluation methods.
With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended and fine-grained free-form descriptions. Existing evaluations often focus on generation quality or task performance, making it difficult to diagnose information coverage and description reliability. We propose MMAC, a \textbf{M}assive \textbf{M}ulti-dimensional benchmark for \textbf{A}udio \textbf{C}aptioning. MMAC contains 5,638 audio clips from more than 20 data sources, covering 6 capability categories and 15 evaluation dimensions. Given a model-generated caption, MMAC checks whether it mentions relevant information in the target dimension and whether the mentioned content is consistent with the reference label. We evaluate representative open-source and proprietary AudioLLMs. Results show clear differences across evaluation dimensions, information coverage, and description reliability. We will release the MMAC benchmark and evaluation code.