Search papers, labs, and topics across Lattice.
This paper investigates the potential for version identification (VI) to subsume track identification (TI) in music identification tasks, proposing a unified benchmark to evaluate the accuracy and robustness of both tasks. By comparing seven existing models, the authors reveal that none achieve satisfactory performance across both dimensions, prompting the development of a baseline model that successfully integrates TI and VI with only 10 seconds of TI queries. The study also identifies two key constraints that hinder TI performance, paving the way for future advancements in music identification systems.
A unified music identification system can achieve robust performance across track and version identification tasks with just 10 seconds of audio input.
Given a music database, track identification (TI) retrieves the exact track matching an audio excerpt, whereas version identification (VI) retrieves its musical versions. Traditionally, the two tasks have been addressed separately. However, as every track is its own closest version, we investigate whether VI can subsume TI. This requires VI systems to be robust to both signal manipulation and audio degradation. We therefore propose a unified benchmark that evaluates accuracy and robustness on each task. Comparing seven existing models on this benchmark, we show that none of them are both accurate and robust on both tasks. We then train a baseline model targeting both tasks and show that a unified system is possible with 10 s TI queries. Lastly, we characterize the two retrieval constraints that limit our model's TI performance. We envision extending this unification to other music identification tasks.