Search papers, labs, and topics across Lattice.
6
0
7
6
RefCaptioner not only outperforms existing models in video captioning but also enables precise grounding of visual elements to multiple reference images, enhancing factual accuracy.
MVP-Nav achieves state-of-the-art performance in zero-shot object navigation by effectively integrating semantic reasoning with physical constraints, even in the absence of depth information.
Achieving superior zero-shot success rates, RelAfford6D redefines robotic manipulation by seamlessly linking abstract instructions to precise physical actions without the need for extensive training.
Current audio-visual generation models struggle to maintain coherence and alignment when scaling to minute-long content, a problem exposed by the new LongAV-Compass benchmark.
Ditching text-based chain-of-thought unlocks better audio-visual reasoning by interleaving textual steps with a unified latent space that preserves dense sensory information.
GUI agents struggle with long tasks not because they mis-click, but because they forget what they were doing, and a new "anchored memory" method can fix it.