Search papers, labs, and topics across Lattice.
This paper introduces temporal music grounding, a novel task that evaluates how well audio-language models can link musical notes and events to specific time spans in audio input. The authors develop MusicGroundingBench, a benchmark suite that allows for precise symbolic-to-audio alignment using algorithmically generated piano MIDI, and they demonstrate that current models struggle with this task, although targeted training can significantly improve performance. The findings highlight the importance of grounding in enhancing music understanding within audio-language models and establish a new framework for future research in this area.
Current audio-language models struggle with temporal music grounding, but targeted training can lead to significant performance improvements.
Large audio-language models can produce fluent and musically plausible responses, yet it often remains unclear whether those responses are grounded in the audio input. We introduce temporal music grounding, a task in which a model returns one or more time spans corresponding to a queried musical note, event, or pattern. To evaluate this capability, we present MusicGroundingBench, a controlled benchmark suite built by rendering algorithmically generated piano MIDI to audio, yielding exact symbolic-to-audio alignment. The suite comprises two subsets: MGBench-3N, which evaluates note-level grounding in clips containing up to three notes, and MGBench-2B, which evaluates structured grounding and short-form music understanding in two-bar excerpts. Experiments show that temporal music grounding remains challenging for current audio-language models, whereas task-specific training yields substantial gains. We further report exploratory evidence on the relationship between grounding supervision and music understanding. These results establish MusicGroundingBench as a controlled testbed for assessing whether audio-language models ground their responses in temporally localized musical evidence.