Search papers, labs, and topics across Lattice.
5
0
7
5
A single-layer speech enhancement model outperforms naive architectures and achieves competitive quality with a significant speedup through progressive knowledge distillation.
A single universal speech enhancement model can effectively adapt to multiple latency requirements without sacrificing performance, challenging the need for specialized models.
LLMs can effectively aggregate diverse speech quality metrics, even outperforming specialized models when labeled data is scarce.
Text-only LLMs already contain surprisingly diverse levels of auditory knowledge, and this pre-existing knowledge strongly predicts their performance when adapted for audio-language tasks.
Time-shifted anechoic speech beats early reflections as a training target for universal speech enhancement, leading to better perceptual quality and ASR performance.