Search papers, labs, and topics across Lattice.
This paper presents SPHERE, an innovative approach to automatic music upmixing that utilizes audio language model (ALM) post-training to derive spatial mixing parameters from multi-stem recordings. By employing a combination of rejection sampling supervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR), the authors introduce a reward suite based on music mixing conventions that enhances the quality of the output mix. The key finding reveals that expert domain knowledge can effectively be encoded as verifiable rewards in language models, circumventing the need for task-specific architectures and yielding superior mixing results.
Encoding expert music mixing conventions into language models can significantly elevate the quality of automatic music upmixing.
In this paper, we study the task of automatic music upmixing, wherein a system predicts spatial mixing parameters from a multi-stem recording. Different from existing methods that rely on task-specific music encoders, we approach this task via audio language model (ALM) post-training, leveraging rich representations from existing ALMs, which encode both music semantics and mixing knowledge. Specifically, we propose a post-training recipe that first employs rejection sampling SFT, followed by reinforcement learning (RL) with verifiable rewards (RLVR) via GRPO. We propose Sphere (Spatial Heuristic Rewards), a deterministic reward suite inspired by music mixing conventions, to guide our post-training. It consists of 6 perceptually-motivated sub-rewards and encourages the output mix to be centered, balanced and spacious. More broadly, our results suggest that expert domain knowledge can be encoded as verifiable rewards and distilled into language models, without task-specific architectures.