Search papers, labs, and topics across Lattice.
This study introduces a method for learning discrete representations of continuous pitch movements in music, specifically targeting Korean traditional music. By employing a VQ-VAE model to quantize local pitch-contour patterns from unlabeled audio, the authors achieve stable tokenization that aligns with expert-defined musical categories. The key finding is that these learned tokens effectively capture the nuances of traditional music analysis, facilitating corpus-level studies without the need for supervised labeling.
Unsupervised learning of pitch-contour tokens reveals hidden structures in Korean traditional music, aligning with expert categories and enhancing analysis.
Computational analysis of music often relies on discrete representations, yet many musical traditions are organized around continuous pitch movement that resists segmentation into note-like units. For such traditions, the discrete units that analysis would build on are not given in advance. We address this gap by learning a vocabulary of local pitch-contour patterns directly from unlabeled audio, using a VQ-VAE that quantizes fixed-length contour segments into a finite codebook. To make the learned tokens stable across segmentation positions and small variations in timing and pitch range, we train the model with a reconstruction objective evaluated under the best alignment among a set of candidate temporal and pitch-domain transformations. Applied to Korean traditional music, the learned tokens recover information about expert-defined sigimsae categories without supervision, and in pansori individual tokens align with the two principal modes, Gyemyeonjo and Ujo, supporting their use as units for corpus-level analysis of contour-centric traditions.