Search papers, labs, and topics across Lattice.
Affiliation:
4
0
5
Heavy speaker overlap drastically hinders speech recognition accuracy in smart glasses, revealing critical limitations in current audio-language models.
Multi-speaker conversational understanding is critically under-evaluated, with MSU-Bench revealing that even leading models struggle with complex speaker grounding tasks.
Achieving high accuracy in multi-speaker transcription, SoulX-Transcriber outperforms existing models by effectively addressing speaker overlap and rapid turn-taking.