Search papers, labs, and topics across Lattice.
This paper conducts an audit of two large-scale generative music systems, Suno and Lyria 3, to assess their musical homogenization compared to human-produced music across four genres. By generating and analyzing 100 tracks per system and genre using 72 music information retrieval features, the study uncovers distinct homogenizing tendencies: Lyria reduces acoustic diversity within genres, while Suno blurs genre distinctions without compressing intra-genre variation. The findings highlight that both systems exhibit learned priors that do not align with user prompts, raising concerns about the implications for musical diversity and economic equity in AI-generated music.
AI music systems are not just homogenizing sounds; they are reshaping the very landscape of musical value and recognition.
This paper audits whether large-scale generative music systems exhibit measurable musical homogenization relative to human-produced music, and develops a justice-centered account of why this matters. We audit two commercially deployed systems (Suno and Lyria 3) across four genres (Afrobeats, K-pop, Dance Pop, and Heavy Metal). For each system and genre, we generate 100 tracks and compare them against human corpora of equal size, using 72 music information retrieval (MIR) features and multiple diagnostics of dispersion, redundancy, and separability. We define homogenization as reduced acoustic variation in standard computational audio features including rhythm and timing, timbre/spectral shape, and dynamics, both within genres and across genre boundaries. We also generate tracks using only a genre name as the prompt, with no additional instructions, to reveal each system's default musical tendencies. The results show two structurally distinct homogenizing tendencies. Lyria reduces within-genre acoustic diversity, while Suno collapses the acoustic distinctions between genres without compressing within-genre spread. Neither system follows user prompts faithfully, indicating that the observed patterns reflect learned priors rather than prompt constraints. The two systems do not converge on a common acoustic profile and are more acoustically distant from each other than two random human subsamples would typically be. Nevertheless, a standard classifier distinguishes AI from human tracks near-perfectly on MIR features alone. We argue that these patterns matter not as an aesthetic curiosity but as a justice-relevant condition, shaping which musical styles become legible, valued, and economically rewarded as generated outputs increasingly circulate at scale.