Search papers, labs, and topics across Lattice.
This paper introduces DSME, a novel music enhancement model that utilizes dual time-frequency spectral representations to improve the quality of non-professional music recordings. By integrating short-time Fourier transform (STFT) for generation and constant-Q transform (CQT) for discrimination within a generative adversarial framework, the model effectively reconstructs clean audio from degraded inputs while maintaining harmonic consistency through a chroma-spectrum loss. Experimental results demonstrate that DSME significantly outperforms existing baselines in both objective and subjective quality assessments, highlighting the advantages of the dual-spectrum methodology in music enhancement.
Dual time-frequency spectral representations can dramatically enhance the quality of degraded music recordings, outperforming traditional methods in both objective metrics and listener satisfaction.
Non-professional music recordings shared online often suffer from background noise and reverberation, degrading perceived quality and limiting reuse. This paper proposes DSME, a music enhancement model based on dual time-frequency spectral representations. Within a generative adversarial framework, DSME uses short-time Fourier transform (STFT) spectra for generation and constant-Q transform (CQT) spectra for discrimination. Leveraging STFT's fixed window, invertibility, and predictability, the generator estimates clean amplitude-phase spectra from degraded inputs and reconstructs waveforms via inverse STFT. Exploiting CQT's log-frequency, variable-window structure aligned with musical octaves, we design an octave-segmented CQT discriminator. We also introduce a chroma-spectrum loss to emphasize pitch and harmonic consistency. Experiments show DSME outperforms baselines in objective and subjective tests, validating the effectiveness of the dual-spectrum approach.