Search papers, labs, and topics across Lattice.
This paper introduces UCAG-P, a novel camera-centric unified action formulation that aligns heterogeneous embodied datasets into a shared geometric action space, enabling a single policy to effectively manage diverse robot morphologies and action spaces. By representing manipulation through camera-observable anchor motion, UCAG-P allows for the decoupling of policy learning from embodiment-specific commands, enhancing the transferability of learned actions across different robotic platforms. The approach achieves impressive performance metrics, including 98.3% on LIBERO and 82.0% zero-shot on LIBERO-Plus, demonstrating its efficacy without the need for benchmark-specific fine-tuning.
A unified policy can achieve near-perfect performance across diverse robotic embodiments without requiring extensive fine-tuning for specific tasks.
Scaling generalist vision-language-action (VLA) policies is severely bottlenecked by the inherent heterogeneity of embodied data, which spans diverse robot morphologies, camera configurations, and low-level action spaces. Existing paradigms typically address this mismatch through explicit action retargeting, human-to-robot video synthesis, or dataset-specific adaptation branches, fundamentally hindering the joint learning of a unified policy. We introduce UCAG-P, a camera-centric unified action formulation that structurally aligns heterogeneous embodied datasets into a shared geometric action space. Rather than treating robot-specific commands as the shared policy target, UCAG-P represents manipulation through camera-observable anchor motion in image and camera-frame coordinates, treating robot arms, humanoids, and human hands as different embodiments of a common action schema. A geometry-conditioned action translator combines predicted motion with target-embodiment kinematics to produce executable controls. The resulting decoupled architecture allows a shared VLA policy to learn transferable manipulation geometry while retaining embodiment-specific controllability. UCAG-P is trained on 4.03K hours of robot and simulation data and 2.34K hours of human demonstrations. A single checkpoint reaches 98.3% on LIBERO, 88.7% and 89.2% on RoboTwin Easy and Hard, 82.0% zero-shot on LIBERO-Plus, and 62.0% on RoboCasa GR-1, without benchmark-specific fine-tuning.