Search papers, labs, and topics across Lattice.
This paper introduces Close-to-Distant microphone Projection (C2D projection), a novel method for generating training targets for speech enhancement by transforming close-microphone recordings into clean reference signals aligned with distant-microphone recordings. The approach addresses the critical issue of data mismatch between simulated and real recordings, which has hindered the accuracy of neural networks in distant speech scenarios. Experimental results reveal that neural networks trained with C2D-projected data significantly outperform existing state-of-the-art methods, such as Guided Source Separation, on the challenging CHiME6 dinner party ASR task when enhanced outputs are used as auxiliary inputs.
Neural networks trained with C2D-projected data achieve superior performance in real-world speech enhancement, outperforming state-of-the-art methods by effectively bridging the gap between close and distant microphone recordings.
Training neural networks (NNs) for speech enhancement (SE) in distant speech-capturing scenarios requires paired distorted and clean reference speech signals. While such data are often generated through simulation, the mismatch between simulated and real recordings significantly limits SE accuracy. To address this issue, we propose Close-to-Distant microphone Projection (C2D projection), a method that generates paired data from real recordings captured by close and distant microphones. C2D projection estimates an optimal projection matrix that transforms close-microphone inputs into clean reference signals aligned with distant-microphone recordings, while simultaneously performing denoising. We show this projection can be effectively realized using a variant of the Parametric Multichannel Wiener Filter (PMWF). Experimental results demonstrate that an NN trained with C2D-projected data outperforms the state-of-the-art Guided Source Separation (GSS) on the challenging CHiME6 dinner party ASR task under oracle diarization, when using the enhanced output from GSS as an auxiliary input to the NN.