Search papers, labs, and topics across Lattice.
This study systematically investigates zero-shot generalization in vision-language-action (VLA) models for tabletop manipulation by establishing a controlled benchmark that differentiates between strict and pretrain-exposed zero-shot transfer. The authors analyze the impact of various factors, including state-action representations and pretraining embodiment diversity, on transfer performance across 14 held-out target embodiments. Key findings reveal that improvements in cross-embodiment transfer can be achieved through specific representations and minimal target-embodiment data during pretraining, emphasizing the need for distinct reporting of transfer types.
Adding just 5% of target-embodiment data during pretraining can boost transfer performance by over 13 percentage points, highlighting the critical role of embodiment exposure in VLA models.
Zero-shot generalization to unseen embodiments is important for generalizable vision-language-action (VLA) models as robot hardware evolves and task-specific data collection remains costly. However, a systematic understanding of this problem remains limited, in part because the literature lacks a unified zero-shot transfer definition and controlled evaluation settings that isolate embodiment changes from differences in tasks, scenes, or protocols. To address this gap, we first distinguish strict zero-shot transfer, where the target embodiment is absent from all training data, from pretrain-exposed zero-shot transfer, where it appears only during pretraining. We then introduce a controlled benchmark spanning 14 held-out target embodiments across simulation and real-world validation. Within this framework, we conduct a controlled analysis of four factors: state-action representations, pretraining embodiment diversity, auxiliary co-training objectives, and target-embodiment exposure. Experimental results show that local end-effector (EEF) state-action representations, the source embodiment diversity, and auxiliary co-training improve cross-embodiment transfer by around 15, 18, and 7 percentage points, respectively. We further find that adding only 5% target-embodiment data during pretraining improves average target-embodiment progress by 13.4 percentage points, showing that strict and pretrain-exposed zero-shot transfer are distinct and should be reported separately. Together, these findings provide practical guidance for evaluating and improving cross-embodiment VLA transfer in stationary tabletop manipulation with two-finger grippers, while motivating future investigation of broader settings including mobile-base control, dexterous hands, and long-horizon tasks.