Search papers, labs, and topics across Lattice.
This paper introduces a novel approach for joint vocal-accompaniment separation using latent flow matching, leveraging a pretrained variational autoencoder (VAE) to map audio mixtures and sources into a compact latent space. By decoupling semantic separation from acoustic velocity prediction through a Separation Encoder and a Velocity Decoder, the authors effectively streamline the separation process. Experiments demonstrate that their method, enhanced by latent adversarial post-training, significantly improves perceptual quality and separation metrics while reducing sampling costs.
Latent adversarial refinement can enhance vocal-accompaniment separation quality while slashing sampling costs by leveraging a compact latent space.
Generative modeling provides a flexible way to model mixture-conditioned source distributions, but iterative diffusion and flow matching models are costly for long music signals. This paper studies joint vocal-accompaniment separation through latent flow matching, where a pretrained variational autoencoder (VAE) maps mixtures and sources into a compact latent space and a flow matching model generates vocal and accompaniment latents jointly. The proposed framework decouples semantic separation from acoustic velocity prediction through a Separation Encoder and a Velocity Decoder. To reduce sampling cost, we further apply latent adversarial post-training inspired by Flow2GAN for few-step generation. Experiments show that latent adversarial refinement can improve perceptual and separation metrics under a reduced sampling budget.