Search papers, labs, and topics across Lattice.
2
0
4
Egocentric human video can outperform traditional teleoperated robot data, achieving superior performance in embodied model pretraining with lower costs and greater diversity.
Camera pose, largely ignored in video LLMs, unlocks significant gains in spatial reasoning and even improves general video QA when used as a lightweight supervisory signal.