Search papers, labs, and topics across Lattice.
Affiliation:
5
0
6
3
Enhancement systems can significantly alter ASR outcomes, but the best choice varies by task and context, challenging the notion of a one-size-fits-all solution.
A single-layer speech enhancement model outperforms naive architectures and achieves competitive quality with a significant speedup through progressive knowledge distillation.
A single universal speech enhancement model can effectively adapt to multiple latency requirements without sacrificing performance, challenging the need for specialized models.
Gaze is a surprisingly effective cue for resolving the cocktail party problem, boosting audio-visual speech enhancement by over 23% in SI-SDR.
Time-shifted anechoic speech beats early reflections as a training target for universal speech enhancement, leading to better perceptual quality and ASR performance.