Search papers, labs, and topics across Lattice.
4
0
6
6
Heavy speaker overlap drastically hinders speech recognition accuracy in smart glasses, revealing critical limitations in current audio-language models.
Achieving high accuracy in multi-speaker transcription, SoulX-Transcriber outperforms existing models by effectively addressing speaker overlap and rapid turn-taking.
3D reconstruction models can self-improve at test time by "hallucinating" better versions of themselves from more complete viewpoints, sidestepping the need for costly 3D ground truth.
Achieve human-like full-duplex voice interactions with SoulX-Duplug, a plug-and-play module that slashes latency and improves turn management by acting as a semantic VAD.