Search papers, labs, and topics across Lattice.
1
0
2
AV-STE is proposed, a modular streaming audio-visual front-end that restores corrupted semantic speech tokens from noisy audio and lip video before they reach the speech LLM, and remains entirely frozen, preserving its pretrained conversational capabilities.