Search papers, labs, and topics across Lattice.
This study investigates the initial-pose dependence of vision-language-action (VLA) policies in humanoid dual-arm manipulation, revealing that task success can mask significant pose-specific failures and hand selection biases. By introducing metrics such as HandPriorScore and examining interactions across 17 initial configurations, the authors demonstrate that specific arm poses can drastically alter policy performance and hand preference. The research shows that enhancing initial-pose coverage in training datasets can significantly improve the robustness of these policies against initial configuration variations.
Initial arm configurations can dramatically skew hand selection in humanoid manipulation, revealing a hidden layer of complexity in VLA policy performance.
Vision-language-action (VLA) policies are expected to operate robustly across variations in the robot's initial configuration, yet aggregate task success can conceal pose-specific failures and inappropriate hand selection. This work investigates initial-pose dependence in VLA-based humanoid dual-arm manipulation. We characterize the initial-condition-dependent early hand preference as a policy-induced hand prior and quantify it using HandPriorScore, residual hand bias, and target responsiveness. Evaluations across multiple policies and 17 initial configurations reveal strong initial-pose--policy interactions: the same pose produces substantially different success rates across policies, while a single policy exhibits large performance variation across poses. Specific initial arm configurations can suppress or induce an asymmetric hand preference, with the resulting effect varying in direction and strength across policies. Wrist-camera observations also influence hand selection and task performance. Expanding initial-pose coverage in the training dataset substantially improves robustness, while targeted augmentation around a low-performing configuration increases its success rate. Comparisons across training configurations show that sufficient exposure to the target simulation task is beneficial, whereas the effect of real or auxiliary data depends on pose coverage, simulation ratio, and observation availability. These findings characterize a pose-conditioned hand prior, identify a localized initial arm configuration as a causal handle on hand-selection behavior, and demonstrate how data coverage and training composition affect initial-pose robustness.