Search papers, labs, and topics across Lattice.
5
0
4
5
FireRedAudio achieves leading performance in audio understanding and speech generation by decoupling input representations, marking a significant advancement in unified audio-language modeling.
DiaScriber achieves unprecedented accuracy in multi-speaker scenarios, overcoming the challenges of overlapping speech and rapid transitions.
Dialect recognition in ASR can be significantly improved without sacrificing Mandarin accuracy, thanks to a novel self-distillation approach.
Real-world conversational dynamics significantly challenge target speaker extraction, revealing that even advanced systems struggle with natural overlap and noise.
Current spoken dialogue systems struggle with the nuances of human conversation, but a new benchmark offers a path to more natural interactions by focusing on handling interruptions and overlapping speech.