Search papers, labs, and topics across Lattice.
This paper introduces a single-utterance test-time adaptation (TTA) method that enhances speech quality by leveraging an autoregressive prior derived from clean speech latent representations. By minimizing the Kullback-Leibler divergence between the enhanced speech distribution and the autoregressive clean speech prior, the approach effectively addresses the challenges posed by mismatched acoustic conditions without the need for labeled data. Experimental results across various noisy speech datasets demonstrate significant improvements in speech quality, particularly in scenarios where training and testing noise conditions differ.
Test-time adaptation can dramatically enhance speech quality under mismatched acoustic conditions without requiring labeled data.
Test-time adaptation (TTA) offers a promising direction for improving speech enhancement models under mismatched acoustic conditions, without requiring access to labeled target data. In this work, we propose a single-utterance TTA method that regularizes a pretrained speech enhancement model using an autoregressive prior trained on clean speech latent representations extracted from a neural audio codec. Adaptation is performed by minimizing the Kullback-Leibler divergence between the enhanced speech distribution and the clean speech prior. Experiments across multiple noisy speech datasets show consistent improvements in speech quality, particularly under training-testing noise mismatch conditions. Code and audio examples are available online.