Search papers, labs, and topics across Lattice.
Affiliation:
3
0
7
IMLE-VLA is introduced, which replaces the iterative action head with a single-step conditional generator trained via conditional Implicit Maximum Likelihood Estimation (cIMLE), and promotes multimodal action coverage, avoiding the mode collapse of naive regression heads while eliminating multi-step sampling entirely.
Robots equipped with scene graphs can significantly outperform traditional imitation learning methods in complex, partially observed environments.
Get up to 1.79x faster ViT inference on high-resolution images without sacrificing accuracy by surgically replacing full-attention blocks with cheaper alternatives *after* pre-training.