Search papers, labs, and topics across Lattice.
MAJEPPA is a self-supervised framework designed to learn comprehensive piano performance representations across a spectrum of skill levels, from beginners to virtuosos. The framework utilizes a curated dataset of approximately 4,000 annotated recordings and employs a pre-trained MIDI autoregressive model to achieve next-token prediction and align score and performance representations through InfoNCE and supervised contrastive losses. Evaluation through the newly introduced EVPMR benchmark shows significant advancements in quality assessment and classification tasks, indicating the model's potential for real-world applications in piano performance analysis.
A unified framework that not only generates but also interprets piano performances reveals critical insights across all skill levels, challenging traditional assessment methods.
We present MAJEPPA, a self-supervised framework to learn piano performance representations that span the full skill spectrum, from beginner practice sessions to virtuoso concert recordings. We curate the MAJEPPA dataset, comprising ~4,000 annotated recordings across six expertise levels and six recording contexts. We adapt a single pre-trained MIDI autoregressive model with a joint objective: next-token prediction learns score-conditioned performance generation at various skill levels, while InfoNCE and supervised contrastive losses align abstract score and performance representations in a joint embedding space. The proposed model both generates and understands performances in a unified framework. By introducing the EVPMR benchmark, a suite of downstream tasks spanning quality assessment, competition ranking, mistake and technique classification, we evaluate the learnt representations, demonstrating progress towards a real-world model for the piano performance space.