Search papers, labs, and topics across Lattice.
This study evaluates how well generative music models adhere to emotional cues through a unified evaluation pipeline, using 1000 tracks from the GTZAN dataset. By extracting semantic audio descriptions and estimating valence and arousal with specialized models, the authors generate music using three systems and assess their emotional fidelity. Results indicate that text-conditioned generation consistently outperforms audio conditioning, with significant variations in emotion-following across different genres, highlighting the need for direct evaluation of affective controllability in generative music models.
Text-conditioned generative music models outperform their audio-conditioned counterparts in faithfully following emotional cues, revealing critical insights into genre-specific emotion-following dynamics.
Recent generative music models offer increasingly fine-grained control through text and audio conditioning, yet how faithfully they follow intended emotional cues remains an open question. We address this gap with a unified evaluation pipeline for emotion-following in generated music. Using all 1000 tracks in GTZAN, we extract semantic audio descriptions with DashengLM, an audio captioning model, and estimate source-track valence and arousal with Music2Emotion, a music emotion recognition model. We construct affect-aware text prompts by combining descriptions with top-ranked emotion tags and generate 30-second outputs with three systems, Stable Audio Open, MusicGen, and InspireMusic, evaluating both text- and audio-conditioned generation. To measure emotion-following, we compute valence and arousal on generated audio and compare them with the source tracks using absolute error and Euclidean distance in valence-arousal space. Text-conditioned generation consistently outperforms audio conditioning, with MusicGen (text) and InspireMusic (text) achieving the best performance, while audio-conditioned variants prove less stable. We further find that valence is preserved more reliably than arousal and that emotion-following varies substantially across genres. These findings underscore the importance of evaluating affective controllability directly rather than relying solely on general quality or prompt-relevance metrics.