Search papers, labs, and topics across Lattice.
This study investigates genre-induced shortcut learning in music evaluation models, revealing that deep neural networks often rely on spurious correlations between genre features and aesthetic scores rather than true perceptual qualities. By analyzing the SongEval dataset, the authors find that these biases lead to significant overestimation of pop music and undervaluation of high-quality samples from other genres, misaligning model predictions with human preferences. To mitigate this issue, they introduce a novel training objective that reweights challenging samples and regularizes performance across genres, resulting in improved genre-invariant representations and better alignment with human aesthetic judgments.
Genre biases in music evaluation models can lead to significant misalignments with human preferences, but a new training approach can correct this skew.
Music aesthetics scoring plays a critical role in applications such as dataset curation, generative model evaluation, and reward modeling for music generation. Recent approaches rely on deep neural networks trained on human-annotated ratings, but these models may exploit spurious correlations rather than capturing perceptually meaningful aesthetics. In this work, we identify a previously underexplored failure mode in music evaluation models: genre-induced shortcut learning. Through a systematic analysis of SongEval, we show that biases in training data lead to strong correlations between genre-related features and predicted scores, causing the model to use them as a proxy for aesthetics. This results in systematic overestimation of pop music and undervaluation of high-quality samples from other genres, leading to predictions that are inconsistent with human preferences. To address this issue, we propose a training objective that jointly reweights hard samples and regularizes group-level performance, encouraging the model to learn genre-invariant representations of musicality. Experimental results demonstrate that our method reduces genre-dependent bias and improves alignment with human preferences, as reflected by gains in both cross-genre and within-genre preference alignment.