Search papers, labs, and topics across Lattice.
This paper introduces a score-agnostic method for grouping automatically transcribed music performances based on structural similarity, addressing the challenge of performance variations in large-scale, noisy datasets. They use sequence-to-sequence alignment and hierarchical clustering on pairwise transcriptions, leveraging alignment cost and sequence length dissimilarity to resolve structural mismatches. The method effectively groups performances with similar structural interpretations, enabling more meaningful performance analysis in the absence of ground-truth scores.
Automatically transcribed music datasets can be cleaned and organized for performance analysis without relying on ground truth scores, using only structural similarities between performances.
In recent years, thanks to advances in automatic music transcription (AMT), several large-scale datasets of automatically transcribed piano solo music have been released. While these datasets undoubtedly offer extensive material for performance studies, they vary substantially in quality. In the case of classical music, performances often differ not only in expressive aspects such as tempo, but also in their structural interpretation of the score (including repeat patterns and edition-specific variants). To meaningfully use large-scale transcribed datasets for performance research, transcriptions of the same piece must be grouped according to their underlying structural realisation to support valid comparison. We address this by applying sequence-to-sequence alignment followed by hierarchical clustering: we create pairwise alignments for all pairs of transcriptions of a given piece, and use the alignment cost and (dis)similarity of performed sequence lengths to resolve structural mismatches as features for grouping. We propose this approach as a first step towards automatically evaluating large-scale transcribed datasets that lack ground-truth score and/or audio, shifting the evaluation criterion from truth-based accuracy to musical coherence and plausibility. We demonstrate our score-agnostic approach on around 1,500 transcriptions of 88 compositions from a recently published large-scale transcribed piano performance dataset.