Search papers, labs, and topics across Lattice.
11
2
11
18
World Tokens achieves state-of-the-art performance in embodied action tasks without the heavy inference burden of traditional world models.
NeuPAT recovers nearly all lost language capabilities in multimodal LLMs while ensuring robust performance across diverse tasks.
AtlasVLA outperforms multi-view baselines by over 17% in long-horizon tasks, showcasing the power of proactive reasoning in embodied AI.
TimeThink revolutionizes video reasoning by enabling models to pinpoint relevant temporal evidence with unprecedented accuracy, outperforming existing approaches.
SurveilNav achieves state-of-the-art navigation success rates by seamlessly integrating robot perception with multi-view surveillance, transforming how robots can navigate complex environments.
NavWM redefines navigation by combining perception, generation, and control into a single framework, leading to unprecedented improvements in planning and foresight.
Action verification can now be reliably performed in VLA models, reducing the risk of grasp failures and task errors in real-world robotic applications.
LongSpace reveals that integrating explicit spatial memory into MLLMs can dramatically enhance their performance on long-horizon tasks, a critical advancement for applications in autonomous navigation.
Today's best multimodal LLMs still struggle to grasp fine-grained details and reason across multiple entities in images, even with access to external knowledge.
Achieve near-dense Video-LLM performance on long videos with up to 57% fewer FLOPs by adaptively selecting which video cubes and tokens to process.
Forget slow visual token concatenation: LaVi modulates LLM features directly with visual context, slashing FLOPs by 94% while boosting speed and accuracy.