Search papers, labs, and topics across Lattice.
This study systematically evaluates three speech enhancement systems鈥擠enoiser, PASE, and RE-USE鈥攐n real-time magnetic resonance imaging (rtMRI) audio to assess their impact on signal quality and downstream tasks. The findings reveal that enhancement effects are highly dependent on the specific evaluation endpoint, with no single system outperforming others across all metrics. Notably, while RE-USE generally improved automatic speech recognition (ASR) performance, Denoiser showed mixed results, emphasizing the need for task-specific considerations in audio enhancement for rtMRI applications.
Enhancement systems can significantly alter ASR outcomes, but the best choice varies by task and context, challenging the notion of a one-size-fits-all solution.
Audio recorded during real-time magnetic resonance imaging (rtMRI) is heavily contaminated by scanner noise, but it remains unclear whether general-purpose speech enhancement improves the signal for speech research and downstream processing. Three off-the-shelf systems---Denoiser, PASE, and RE-USE---are evaluated across five rtMRI corpora using naturally recorded inputs, a clean-input probe, and an archived paired additive-noise probe. The multi-task evaluation spans learned quality predictors, speaker and phone representations, reference-based intelligibility and quality measures, acoustic--phonetic probes, automatic speech recognition (ASR), and paralinguistic tasks. The central result is that enhancement effects are endpoint dependent: higher predicted-quality scores do not reliably imply better ASR performance or greater source fidelity. Across 15 corpus--recognizer comparisons using corpus-provided processed inputs, RE-USE yielded lower word-error-rate point estimates in 11, whereas Denoiser yielded higher estimates in 13. In the paired additive-noise probe, PASE and RE-USE improved recognized-phone agreement, intelligibility, and perceptual-quality point estimates. Denoiser improved recognized-phone agreement and short-time objective intelligibility (STOI) but reduced speaker-embedding similarity. No system was uniformly best across corpora, recognizers, and endpoints. Enhanced rtMRI audio should therefore be treated as a task-specific transformed derivative rather than a universally improved replacement for the original or DSP-processed waveform.