Search papers, labs, and topics across Lattice.
The paper introduces MusTBENCH, a new benchmark to evaluate temporal grounding in music-based Large Audio-Language Models (LALMs) using five question-answering tasks validated by music experts. They find that existing LALMs struggle with precise temporal grounding and propose MusT, a four-stage temporal optimization recipe (music encoder adaptation, LLM adaptation, LLM supervised fine-tuning, and RL-based optimization) to improve performance. Experiments show MusT significantly improves temporal grounding over strong baselines on MusTBENCH.
Current music-understanding LLMs can't tell you *when* something happens in a song, but a new benchmark and training recipe, MusTBENCH and MusT, can help them learn.
Recent Large Audio-Language Models (LALMs) have demonstrated promising abilities in understanding musical content. However, whether their responses are grounded in the correct temporal regions of the audio remains underexplored. This limitation is particularly critical for music understanding, where key information often occurs as temporally localized events, such as instrument entries and rhythmic transitions. To address this gap, we introduce MusTBENCH, a music-expert-validated benchmark designed to evaluate temporal grounding in LALMs through five temporally grounded question-answering tasks. To further improve temporal grounding in existing models, we propose MusT, a novel four-stage temporal optimization recipe spanning music encoder adaptation, LLM adaptation, LLM supervised fine-tuning, and RL-based optimization. Experiments on MusTBENCH show that existing LALMs struggle with precise temporal grounding, while MusT brings significant improvements over strong baselines. These results establish temporal grounding as a key missing capability in current LALMs and position MusTBENCH as a challenging benchmark for future research in temporally grounded music understanding.