Search papers, labs, and topics across Lattice.
Galbot
9
70
5
10
WAM-TTT allows robot models to adapt to new tasks using only raw human videos, eliminating the need for additional demonstrations or fine-tuning.
AllDayNav achieves nearly 100% success in lifelong navigation by using a self-evolving memory system that outperforms traditional mapping techniques.
Cosmos 3 sets a new benchmark for omnimodal models, outperforming existing state-of-the-art in both Text-to-Image and Image-to-Video tasks.
By cleverly anchoring diffusion sampling near plausible solutions and adding a lightweight residual correction, AnchorVLA achieves robust mobile manipulation with significantly reduced inference costs.
A single spatial token, learned via occupancy prediction on a massive dataset, is surprisingly effective at injecting crucial spatial awareness into vision-language navigation, leading to state-of-the-art performance.
Forget hand-crafted heuristics: this new dynamics-aware policy learns to exploit contact forces in cluttered environments, outperforming traditional methods by 25% in simulation and showing impressive sim-to-real transfer.
Achieve more realistic and physically plausible scene reconstructions from video by explicitly optimizing viewpoints for object generation and synthesizing scene graphs within a 3D simulator.
Training a robot foundation model on 30,000 hours of heterogeneous embodied data lets it outperform prior methods by up to 48% on complex manipulation tasks and even benefit from low-quality data.
Forget painstakingly labeled real-world data – GraspVLA proves you can train a surprisingly capable grasping foundation model on a billion frames of purely synthetic action data.