Search papers, labs, and topics across Lattice.
This paper presents a comprehensive framework for evaluating six language model-based generative speech enhancement paradigms using latent space features from neural audio codecs, including both discrete and continuous approaches. The authors conduct a rigorous comparative analysis of these paradigms in a unified experimental setup, revealing that continuous non-autoregressive speech enhancement (CNAR) outperforms its discrete counterparts across various metrics. Additionally, they introduce a fine-tuning strategy that leverages auxiliary losses on reconstructed speech, significantly enhancing performance on intrusive and non-intrusive evaluation metrics such as DNSMOS and PESQ.
Continuous non-autoregressive models outperform discrete methods in speech enhancement, revealing a critical shift in paradigm effectiveness.
Language model (LM)-based speech enhancement (SE) has recently emerged rapidly using latent space features of neural audio codecs (NACs). In this paper, first, we present a unified framework covering six popular LM-based generative SE modeling paradigms based on discrete/continuous latent NAC features: discrete or continuous autoregressive (D/CAR) SE, discrete or continuous non-autoregressive (D/CNAR) SE, discrete diffusion (DDiff) SE, and continuous flow matching (CFM) SE. Second, we are the first to compare their performance in a unified experimental setup and synopsis with diverse intrusive and non-intrusive metrics, enabling a fair and comprehensive evaluation. Third, we propose a fine-tuning strategy with auxiliary losses on reconstructed speech to improve both intrusive and non-intrusive metrics. Trained and evaluated on URGENT 2025 Speech Enhancement Challenge data splits, all continuous-domain paradigms excel their discrete-domain counterparts. The overall best approach turns out to be CNAR. We further show that our proposed auxiliary loss fine-tuning strategy helps to improve DNSMOS, NISQA, PESQ, and POLQA consistently in all six paradigms.