Search papers, labs, and topics across Lattice.
This paper introduces DTT-BSR+, a two-stage cascade system for music source restoration (MSR) that separates the processes of distribution fitting and signal reconstruction. By employing a generative DTT-BSR separator in the first stage and a modified Demucs network in the second, the method achieves significant improvements in multi-mel signal-to-noise ratio (MMSNR) while maintaining semantic consistency. Notably, DTT-BSR+ outperforms the state-of-the-art X-LANCE system across multiple stems, highlighting an implicit trade-off between reconstruction accuracy and semantic fitting.
Achieving a breakthrough in music source restoration, DTT-BSR+ significantly enhances signal quality while preserving semantic integrity, outperforming existing methods.
Music source restoration (MSR) requires jointly addressing source unmixing and the inversion of non-linear production effects. Current methods struggle to achieve accurate target signal reconstruction while maintaining semantic consistency. To address this limitation, we propose DTT-BSR+, a two-stage cascade MSR system that decouples distribution fitting from signal reconstruction into separate stages. A generative DTT-BSR separator in the first stage produces stems matching the prior of clean sources, and a modified Demucs network in the second stage enhances the first stage output using time-domain and multi-resolution spectral losses. DTT-BSR+ improves multi-mel signal-to-noise ratio (MMSNR) over the single-stage DTT-BSR across all stems, and surpasses the state-of-the-art X-LANCE MSR system on five stems. We also reveal through Fr茅chet Audio Distance (FAD) decomposition an implicit trade-off between signal reconstruction accuracy and semantic distribution fitting across stems.