Search papers, labs, and topics across Lattice.
This paper introduces InvFlowFD, a novel perceptual music quality metric that operates without the need for background sets or reference samples, relying solely on a pre-trained Flow Matching backbone. By employing unconditional Flow Matching inversion through Euler integration, the method effectively detects artificial distortions and ranks music generation models based on human perceptual judgments. The results indicate that InvFlowFD demonstrates a strong correlation with human perception of sound quality, outperforming existing metrics in flexibility and applicability.
InvFlowFD achieves reference-free music quality assessment, aligning closely with human perception while eliminating the need for background data.
Existing reference-free methods for evaluating music perceptual quality alleviate the need for paired noisy-clean data, but they still rely on a background set, which is used to compute aggregated statistics of clean audio samples. In this work, we propose a novel approach that eliminates this requirement, achieving background-set-free and reference-free quality estimation using only a pre-trained Flow Matching backbone. We demonstrate that unconditional Flow Matching inversion via simple Euler integration is sufficient to detect various artificial distortions and accurately rank music generation models against human perceptual judgments. We introduce InvFlowFD, which performs flow inversion and compares a group of inverted samples to the prior distribution. We evaluate our method against prior work, quantitatively and with a thorough human study. Results suggest that InvFlowFD is highly correlated with human perception of sound distortions, as well as generative models'quality, while being more flexible and less restrictive than existing metrics.