Search papers, labs, and topics across Lattice.
8
9
5
16
VisCo not only compresses visual tokens more effectively than previous methods but also enhances model performance through innovative use of memory tokens.
Tactile dynamics are crucial for contact-rich manipulation, and VT-WAM outperforms existing models by 26.67% to 35.84% by effectively integrating visual and tactile cues.
TacForeSight enables robots to anticipate contact changes in real-time, outperforming traditional methods in dynamic manipulation tasks.
Robots can now perform complex, contact-rich tasks with significantly smoother and more continuous motions by learning high-frequency action chunks in a latent space.
Achieve robust robot manipulation across diverse viewpoints without camera calibration by synthesizing novel views with a geometry-aware video diffusion model.
Pocket-sized VLA models can now achieve state-of-the-art robot manipulation performance by pre-training on a curated multimodal dataset and injecting manipulation-relevant representations into the action space.
Zero-shot RL agents can now learn better representations by focusing on dynamics-relevant image regions, leading to state-of-the-art generalization performance.
Imagine training robots to manipulate objects in the real world, but entirely within a high-fidelity, diffusion-based dream.