Search papers, labs, and topics across Lattice.
To overcome the severe data bottleneck in robot learning, RoboTok constructs an internet-scale data engine that retrieves task-relevant human manipulation demonstrations from raw web videos using query clips. The system learns a compact, viewpoint- and appearance-invariant latent motion space derived from 3D hand trajectories projected into estimated actor-centered reference frames. Benchmarks show that RoboTok achieves superior demonstration retrieval accuracy over existing baselines and directly improves downstream task success rates for dexterous manipulation policies.
Internet video can finally solve the physical robot data bottleneck once manipulation behaviors are indexed by actor-centric 3D hand trajectories rather than fragile visual pixels.
Robot learning increasingly depends on broad and diverse demonstrations, yet collecting robot data remains expensive and poorly suited to covering the long tail of real-world tasks. To address this bottleneck, we introduce RoboTok, an internet-scale data engine that, given a query human manipulation video, retrieves manipulation-relevant human demonstrations from web videos for training dexterous robot policies. Specifically, we learn a latent motion space from 3D hand trajectories expressed in estimated actor-centered reference frames. This representation enables manipulation behaviors to be compared across variations in camera viewpoint, scene appearance, and actor occlusions, while remaining compact enough for efficient search and continual indexing over internet-scale video collections. We evaluate RoboTok against existing robot-data retrieval approaches on retrieval benchmarks and downstream robot policy performance. Our results show that RoboTok retrieves more relevant manipulation demonstrations and improves downstream task success, establishing hand-pose trajectory-aware retrieval as a way to make web video a scalable and continuously growing source of supervision for robot learning.