Search papers, labs, and topics across Lattice.
This paper introduces MIDI-RAE-JEPA, a novel framework for hierarchical representation learning in symbolic music that leverages a pitch- and time-shift equivariance objective alongside a Swin Transformer V2 encoder. The model is trained using self-supervised techniques, achieving a remarkable reconstruction F1 score of 0.995 and generating musically coherent outputs that align with the conditioning excerpts. The findings demonstrate that equivariance-based self-supervised learning can yield semantically rich representations that significantly outperform traditional methods in downstream tasks like emotion classification.
Equivariance-based self-supervised learning can unlock semantically rich representations in symbolic music that outperform traditional methods by a wide margin.
Rich internal representations of musical structure are essential for music understanding tasks such as machine-assisted music co-writing, yet self-supervised approaches for symbolic music representation remain underexplored, particularly those that encode the hierarchical multiscale nature of musical structures. We present MIDI-RAE-JEPA, combining a pitch- and time-shift equivariance objective with LeJEPA and a Swin Transformer V2 encoder to learn such hierarchical representations of symbolic music encoded as piano roll images. The time-shift equivariance objective encourages the model to internalize temporal musical relationships. The encoder is trained purely on self-supervised objectives -- including a masked embedding predictor (MEP) -- with collapse prevented via SIGReg. A separate decoder trained on the frozen encoder embeddings achieves reconstruction F1 of 0.995, and a flow matching generative model conditioned on those embeddings produces generations that closely match the pitch register and rhythmic density of the conditioning excerpt, while mismatched conditioning yields unrelated but musically plausible output. Learned representations outperform a Haar scattering transform baseline on a downstream emotion classification task, and embedding distances increase monotonically with pitch and time shift magnitude, confirming measurable equivariance. These results suggest that equivariance-based SSL objectives, combined with sufficient fine-level encoder capacity, provide a viable path toward semantically rich, generatively useful representations of symbolic music.