Search papers, labs, and topics across Lattice.
This study introduces a dual evaluation framework for automatic music transcription systems, emphasizing the need for complementary assessments of notation and playback similarity. By analyzing a diverse dataset of 230 piano recordings and leveraging a listening study with over 100 participants, the authors identify that the CLEWS metric not only aligns closely with human judgments but is also cost-effective. The findings reveal that different transcription systems excel in either notation or playback similarity, highlighting the importance of evaluating both dimensions to optimize performance across various models.
The most effective playback similarity metric, CLEWS, is also the least expensive to implement, revolutionizing evaluation strategies in music transcription.
Automatic music transcription systems produce sheet music that can be read and played back. We argue that these two targets call for complementary evaluations of notation similarity to a reference score and playback similarity to the original performance, respectively. Our study considers notation similarity metrics from the optical music recognition literature and a wide range of playback-similarity methods validated through a listening study across over 100 participants and 230 piano recordings covering 23 works, 30 performers, and six composers. We find, fortuitously, that the playback similarity metric that correlates best with human judgments, CLEWS, is also the cheapest to run. We also find that the two evaluation dimensions favor different systems among a collection of 24 pipelines formed by pairing eight audio-to-MIDI models with three MIDI-to-score converters, with the latter component systematically determining the favored objective. The complementarity between metrics also holds when adding to the pool Rubato, a new end-to-end system that offers substantially improved notation similarity while remaining competitive, though not the best, on playback similarity.