Search papers, labs, and topics across Lattice.
This study introduces DiVers, a novel dataset for musical version identification (VI) that includes over 1.1 million versions, addressing the limitations of existing datasets that focus on professionally recorded tracks. By incorporating a diverse range of amateur and user-generated content, DiVers enables more robust training of VI systems in real-world conditions. The evaluation reveals that models trained on DiVers significantly enhance performance in noisy environments while retaining effectiveness on cleaner audio, demonstrating the dataset's utility for improving VI robustness.
Models trained on the new DiVers dataset show remarkable resilience to noisy and acoustically diverse inputs, outperforming traditional benchmarks.
Existing datasets for musical version identification (VI) are primarily derived from curated metadata sources such as SecondHandSongs and Discogs, and are therefore dominated by professionally recorded tracks. This leads to a domain mismatch with real-world scenarios, where amateur and user-generated content is prevalent. To address this limitation, we introduce DiVers, a large-scale VI dataset comprising over 1.1 million musical versions, with train-validation-test splits compatible with established datasets such as Discogs-VI-YT, SHS100K, and Da-TACOS. In addition to standard version-level annotations, DiVers provides automatically assigned tags (e.g., instrumental, live) and segment-level predictions indicating the presence or absence of music. We evaluate the proposed dataset by training state-of-the-art VI systems. Our results show that models trained on DiVers achieve substantially improved robustness to acoustically diverse and noisy inputs, while maintaining a stable performance on cleaner, studio-quality benchmarks. We release the dataset metadata, code for its construction, and all experimental pipelines to support reproducibility.