Search papers, labs, and topics across Lattice.
This study introduces a novel training set synthesis approach for bioacoustic denoising, specifically targeting the ultrasonic vocalizations (USVs) of house mice, which are often obscured by ambient noise. By employing a ridge-guided loss function that emphasizes frequency contours, the developed supervised model significantly enhances the preservation of vocalization details during the denoising process. The results demonstrate improved classification performance on noisy recordings and a substantial increase in scale-invariant signal-to-distortion ratio, indicating the method's effectiveness and broader applicability to other bioacoustic signals.
Denoising bioacoustic signals using ridge-guided training synthesis can dramatically enhance the clarity of vocalizations, leading to better classification outcomes in noisy environments.
Bioacoustic recordings are often degraded by ambient noise, which complicates the analysis of weak or noise-overlapped vocalizations. Convolutional neural networks, particularly U-Net architectures, have shown strong denoising performance in speech and music processing. However, their application for denoising bioacoustic signals is limited by the scarcity of clean training data. To address this, we propose a training set synthesis approach and develop a supervised denoising model that predicts a complex ratio mask in the time-frequency domain. The model leverages ridges, or frequency contours, that represent the fundamental frequency together with one or more harmonic partial components of vocalizations. These ridges are used both for the synthesis of training set and for the design of a loss function that assigns higher weights to the ridge regions (ridge-guided loss function). This weighting step helps the network better preserve vocalization details during denoising. As a case study, we evaluate our approach using ultrasonic vocalization (USV) recordings of house mice, which are widely studied in behavioral biology and neuroscience. In field recordings, the proposed method enhances tracking of fundamental and harmonic partial ridges compared to previous signal-processing approaches. In addition, a classifier trained on denoised data improves USV classification on out-of-sample, noisy recordings from wild and domesticated mice compared to classifiers trained on noisy recordings. The proposed method also substantially improves the scale-invariant signal-to-distortion ratio on synthetic testing data across a wide range of input signal-to-noise ratios. While focused on USVs, the proposed approach is broadly applicable to other bioacoustic signals with trackable ridges, and thus enables ridge-based training set synthesis and denoising.