Search papers, labs, and topics across Lattice.
4
0
6
1
OPD-Aha is introduced, which reconstructs the distillation target directly from this isolated visual preference rather than relying on the fragile teacher-student discrepancy, yielding broad and consistent improvements across diverse fine-grained perception and complex multimodal reasoning benchmarks.
Vid2WAM achieves superior task generalization and data efficiency by leveraging video diffusion priors, even with minimal expert demonstrations.
$\tau_0$-WM outperforms traditional models by seamlessly integrating action prediction and evaluation, leading to superior performance in complex robotic tasks.
Robots get a crucial boost in robustness by learning to "see" and predict how objects will move, not just react to the current frame.