Search papers, labs, and topics across Lattice.
5
0
4
AV-STE is proposed, a modular streaming audio-visual front-end that restores corrupted semantic speech tokens from noisy audio and lip video before they reach the speech LLM, and remains entirely frozen, preserving its pretrained conversational capabilities.
Autoregressive coordinate generation fundamentally breaks down on dense satellite imagery, but framing change localization as LLM selection over discrete candidate region tokens solves this bottleneck.
View-invariant video representations can be achieved without sacrificing the richness of view-variant semantics, leading to unprecedented performance in cross-view tasks.
Ditch the left-to-right constraint: diffusion models can now transcribe speech by iteratively refining predictions with bidirectional context, achieving state-of-the-art results.
LLMs can iteratively refine their reasoning and reduce errors by recursively evaluating and improving their own confidence, leading to more stable and faster inference.