Search papers, labs, and topics across Lattice.
This paper redefines automatic music mixing by introducing a sequential stem blending approach, inspired by the methodologies of human mix engineers who blend tracks one at a time. By employing a latent flow matching model conditioned on submix context, the authors enable the effective processing of an arbitrary number of input tracks in a sequential manner. Experimental results indicate that this method outperforms traditional parallelized architectures, showcasing its potential for more nuanced and context-aware audio mixing.
Sequential blending of audio stems could revolutionize automatic music mixing, outperforming traditional methods by leveraging human-like processing strategies.
Automatic music mixing, the task of automatically combining individual audio tracks into a cohesive mixture, is typically addressed by parallelized architectures that process all input tracks in a single pass. In this work, inspired by how human mix engineers process stems one at a time, we propose a paradigm shift and ask whether automatic music mixing can be reformulated as a sequential stem blending task, where each stem is blended into a growing submix. Specifically, we train a latent flow matching model conditioned on the submix context, enabling sequential processing of an arbitrary number of input tracks. To train the model, we introduce a degradation-based data synthesis strategy that simulates realistic stem blending scenarios from existing multitrack and source separation datasets. Experimental results on both stem blending and automatic music mixing benchmarks demonstrate the effectiveness of the proposed approach. We provide audio examples on the accompanying demo page\footnote{https://sequential-mixing-demo.vercel.app/}.