Search papers, labs, and topics across Lattice.
This paper introduces the Rule-based Instructions for Music Editing (RIME) framework, which addresses the challenge of agentic post-production in music by generating realistic paired edit-instruction data from existing music datasets. The authors highlight the inadequacy of current multimodal LLMs in handling post-production tasks, demonstrating persistent performance issues even when evaluated with the newly generated data. RIME not only provides a structured approach to refine music production workflows but also shows potential for enhancing agent performance through supervised fine-tuning, paving the way for more effective collaborative music production systems.
RIME reveals that existing multimodal LLMs struggle with music post-production, highlighting a critical gap in their capabilities that could redefine how music is produced collaboratively.
Almost every piece of recorded music you have ever heard was modified before it reached you; commercial releases rarely spring fully-formed from the mind of a musician. Despite the promise of music generation models for one-shot output, such fine-grained iterative refinement workflows are a complementary problem, and largely out of their reach. There is also a gap for musicians: while they can express what they want to hear, not all have the facility with studio production tools to implement the complex set of actions needed to realize these intuitions. We formalize this task as agentic post-production, wherein individual aspects of a song are targeted, refined, and combined into a final track. We argue the bottleneck is data: existing corpora do not reflect how realistic post-production chains map onto the vocabulary musicians and engineers actually use. We argue there is a language for modifying recorded music that is dense, consistent, and learnable. We introduce the Rule-based Instructions for Music Editing (RIME) framework, which generates realistic paired edit-instruction data from any baseline music dataset grounded in canonical methods, design patterns, and constraints derived from real production workflows. RIME leverages POEMS, a new toolkit that combines stem separation, mixing, and common studio effects for use by multimodal agents. We use RIME and POEMS to generate 3,000 pairs of edit instructions and ground truth audio, and use this data to evaluate existing multimodal LLMs as agents on this task, showing persistent challenges in current models'post-production capabilities. We also demonstrate RIME's ability to improve post-production agent performance via supervised fine-tuning. We see RIME as an early step toward iterative musical agents, collaborative systems that could transform music production much as interactive coding agents have reshaped software engineering.