Search papers, labs, and topics across Lattice.
12
9
14
18
VideoChat3 achieves unprecedented generalization in video understanding while maintaining high efficiency, outperforming larger models with just 4 billion parameters.
LIFT accelerates learning in vision-language-action policies by injecting reactive force feedback, achieving superior performance in contact-rich tasks.
Existing MLLMs struggle with procedural reasoning, but a new self-skill-exploration agent achieves state-of-the-art performance by discovering effective strategies for everyday tasks.
UAV swarms can now navigate complex indoor environments with mission-specific precision, even when communication is lost.
Generating scenarios that directly minimize operational costs can cut grid dispatch expenses by over 2% compared to conventional methods.
Imagine fixing your robot's mistakes *before* it even makes them: RoboPocket lets you train robots twice as efficiently using just your smartphone and AR.
Automatically generating personas from VR app store reviews can efficiently foster empathy and uncover hidden accessibility needs in VR development.
Achieve state-of-the-art small object detection by explicitly preserving fine-grained structural details and modeling global relations, even in complex backgrounds.
Current video LLMs falter when faced with the demands of real-time interaction, a gap RIVER Bench directly addresses by providing a challenging new evaluation framework.
InterFormer tackles the "interaction illusion" in egocentric hand-object parsing, achieving state-of-the-art results by explicitly modeling hand-object co-occurrence and spatial dynamics.
Escape the bottleneck of translating product intent into ranking system hypotheses: GEARS offers an agentic framework that autonomously discovers and validates superior ranking policies.
Current LLMs and VLMs struggle with multi-step reasoning in long videos, often failing to maintain temporal coherence and procedural validity, as revealed by a new benchmark of hour-long narratives.