Search papers, labs, and topics across Lattice.
SiMDex introduces a similarity-based data mining framework that optimizes the selection of ego-centric human videos for robot dexterous manipulation by framing it as a recommendation problem. Utilizing a three-layer recall-ranking-re-ranking pipeline, SiMDex efficiently extracts relevant subsets from a vast pool of ~32 million samples without requiring modifications to the existing VLA architecture. The approach significantly enhances performance, improving the success rate from 47.7% to 61.1% using only ~1.49 million curated samples, demonstrating the effectiveness of selective data curation over random sampling.
Selective curation of ego-centric video data can boost robot manipulation success rates by over 28% with less than 5% of the original dataset.
Recent years have witnessed an explosive trend of scaling ego-centric human videos for robot manipulation, yet it remains unclear which data actually benefits dexterous manipulation. We present SiMDex, a similarity-based data mining framework that casts human data selection for VLA post-training in dexterous manipulation as a recommendation problem. For each robot demonstration, SiMDex employs a three-layer recall-ranking-re-ranking pipeline to extract task-relevant subsets from a pool of ~32M egocentric human samples, operating in a morphology-agnostic action space that requires no changes to VLA architecture or training. Against a strong baseline trained with an equal amount of randomly sampled human data, SiMDex uses only ~1.49M mined samples (<5% of the pool) yet improves the overall success rate from 47.7% to 61.1%, showing that selective curation outperforms indiscriminate data mixing.