Search papers, labs, and topics across Lattice.
This study conducts a comprehensive empirical analysis of Latent Action Models (LAMs) for robotic manipulation, systematically evaluating 41 design choices across latent action modeling, learning objectives, and integration strategies. By unifying various LAM methods within a common autoencoding framework, the research identifies key factors that significantly influence downstream performance. The findings reveal that fine-tuning vision-language model backbones with latent actions enhances initialization for policy learning, validated through extensive experiments on established benchmarks and real-world tasks.
Fine-tuning vision-language models with latent actions can dramatically improve robotic manipulation performance, revealing critical design choices that matter most.
Latent Action Models (LAMs) have emerged as a promising paradigm for enabling robot learning to leverage large-scale unlabeled videos through latent actions that serve as compact surrogates for physical actions. Despite rapid progress, research on LAM remains highly fragmented, with existing methods evaluating different design choices in isolation under inconsistent experimental settings, making it difficult to identify the factors that truly determine downstream robotic manipulation performance. In this work, we present the first comprehensive empirical study of latent action learning for robotic manipulation. We unify representative LAM methods within a common autoencoding framework and systematically investigate 41 LAM design choices across three dimensions, including latent action modeling paradigms, learning objectives and regularization methods, and latent action integration strategies. We further examine four proxy metrics for evaluating latent action quality and assess their ability to reliably predict downstream robotic manipulation performance. Extensive experiments on three widely used benchmarks provide strong empirical evidence that fine-tuning vision-language model (VLM) backbones with latent actions provides a stronger initialization for downstream policy learning, with further validation on real-world robot manipulation tasks.