Search papers, labs, and topics across Lattice.
This paper introduces EATR-Stereo, an embodiment-aware token-routing framework designed to enhance humanoid vision-language-action (VLA) control by effectively utilizing paired stereo evidence. By retaining primary-view tokens and constructing Cross-View Auxiliary Tokens (CVATs) that adapt to the robot's configuration, the framework significantly improves the integration of visual and proprioceptive information during action generation. The results reveal a 60% full-task success rate and a remarkable 100% grasp success rate, showcasing the framework's efficacy in overcoming challenges like severe occlusion in complex tasks.
Selectively routed stereo evidence boosts humanoid VLA control success rates, achieving 100% grasp success even under severe occlusion.
Long-horizon humanoid vision--language--action (VLA) control with head-mounted stereo cameras requires visual interfaces that can exploit complementary views while maintaining compatibility with pretrained representations. Existing interfaces often discard complementary stereo evidence or fuse additional observations without preserving the native primary-view pathway and adapting auxiliary information to robot embodiment. We present EATR-Stereo, an embodiment-aware token-routing framework that retains primary-view tokens and constructs primary-aligned Cross-View Auxiliary Tokens (CVATs) by querying the synchronized auxiliary-view token sequence. A body-segmented proprioceptive encoder further conditions token-wise auxiliary usage on robot configuration history, enabling selective incorporation of stereo evidence during action generation. The routed auxiliary stream augments the language and primary-visual context of a pretrained VLA while keeping its vision--language model frozen. On a 33-DoF physical humanoid with a 37-D proprioceptive state, we evaluate nine configurations in over-100-s search--approach--grasp--place--return tasks. EATR-Stereo achieves 60.0% full-task success, 100.0% grasp success, and 80.0% stage success. Under severe asymmetric occlusion, it improves recovery to 80% compared with 30% for CVAT alone. Ablation studies further show the importance of preserving primary tokens and combining cross-view auxiliary features with structured proprioceptive routing. These results demonstrate that selectively routed paired stereo evidence improves spatial grounding for reliable long-horizon humanoid VLA control.