Search papers, labs, and topics across Lattice.
3
0
6
Video generative priors only translate into robust physical control when paired with explicit world-to-action information routing and synchronized joint denoising rather than standard monolithic fine-tuning.
Models with similar success rates can exhibit vastly different strengths and weaknesses, revealing the hidden complexities of mobile manipulation capabilities.
A surprisingly simple VLA model, StarVLA-$\alpha$, beats more complex systems on real-world robotics tasks, suggesting that VLM backbones are more critical than intricate architectures.