Search papers, labs, and topics across Lattice.
This paper introduces Diff-Symbo, a novel approach for text-controlled symbolic music generation utilizing a latent diffusion model (LDM) to enhance quality, diversity, and duration of the compositions. By creating a dataset of 19,345 text templates and implementing a music information encoder, the authors significantly improve the controllability and compositional consistency of the generated music. Experimental results demonstrate that Diff-Symbo outperforms existing models like GPT-4 and MuseCoco in key metrics, establishing a new standard for high-quality symbolic music generation.
Diff-Symbo achieves unprecedented quality and diversity in text-controlled music generation, outperforming leading models by leveraging a novel latent diffusion framework.
Text-controlled symbolic music generation has recently gained research attention due to its versatile, flexible and straightforward approach to music composition. However, previous approaches tend to generate symbolic music with compromising quality, diversity, controllability and limited duration. In this paper, we present Diff-Symbo, an innovative method that uses latent diffusion model (LDM) to generate high-quality, diverse and long-duration symbolic music. To address the lack of text-symbolic music dataset, we develop a comprehensive dataset with 19,345 text templates by employing large language model. Furthermore, we design a music information encoder to reduce the training overhead while extracting more effective control representations. Given textual descriptions, our proposed method leverages LDM to improve the quality and diversity of music generation. Our method also improves the duration and the compositional consistency of music generation through an autoregressive approach. Experimental results show significant improvements of Diff-Symbo in text controllability, duration, and the quality of generated music compared to the baseline models such as GPT-4, MuseCoco and Multitrack Music Transformer (MMT). As one of the pioneer models in this field, Diff-Symbo paves the way towards controllable and high-quality symbolic music composition based on LDM, offering valuable contributions to both music amateurs and practitioners.