Search papers, labs, and topics across Lattice.
4
0
5
4
Fine-tuning vision-language models with latent actions can dramatically improve robotic manipulation performance, revealing critical design choices that matter most.
LAFP achieves up to 15% higher success rates in imitation learning by preserving the multimodal structure of latent actions, challenging the limitations of traditional behavior cloning.
Robots can now navigate complex outdoor environments using only high-level human instructions and readily available GPS/map data, bypassing the need for expensive HD maps or limited short-horizon policies.
Latent reasoning can beat explicit Chain-of-Thought – but only if you force it to learn causal dynamics via a visual world model, not just language.