Search papers, labs, and topics across Lattice.
This paper introduces FocalSE, a novel speech enhancement method designed to improve the performance of neural speech codecs in noisy environments by employing feature denoising, noise feature separation, and noise recognition. The method utilizes focal modulation for effective compression and decompression, enabling the generation of focal masks that enhance clean feature embeddings while separating noise embeddings from noisy inputs. Experimental results on LibriTTS and ESC50 datasets show that FocalSE significantly outperforms existing state-of-the-art techniques in low-bitrate and low-SNR conditions, highlighting its robustness in real-world applications.
FocalSE achieves superior speech enhancement by effectively separating noise features, outperforming existing methods in challenging acoustic conditions.
Neural speech codec has attracted extensive attention for high-quality reconstruction at low-bitrate. However, real-world noise severely degrades its performance and hinders high-quality clean speech reconstruction. To tackle this problem, we propose FocalSE, a novel speech enhancement method that performs feature denoising, noise feature separation and noise recognition in the continuous embedding space of neural speech codecs. Specifically, we develop focal modulation-based compression and decompression to capture global context and local mutual information, and generate focal masks to recover clean feature embeddings. We then separate noise embeddings from noisy embeddings to improve denoising performance. Finally, we use ResNet1D-18 to recognize noise categories for better separation effectiveness. Extensive experiments on two standard datasets, LibriTTS and ESC50, demonstrate that our method outperforms state-of-the-art approaches under low-bitrate and low-SNR conditions.