Search papers, labs, and topics across Lattice.
This work introduces XR-2, a vision-language-action model trained on 1,500 hours of diverse bimanual manipulation demonstrations aimed at enhancing household task performance. By employing a high throughput data pipeline and a multi-stage training approach, XR-2 achieves impressive manipulation success rates while maintaining efficient training and data utilization. The study reveals that both the quantity of expert demonstrations and the incorporation of real-time human corrections significantly improve task success rates, underscoring the model's robust learning capacity and the dataset's scalability potential.
XR-2 shows that scaling demonstration data and incorporating real-time corrections can dramatically enhance bimanual manipulation success rates in household tasks.
Learning generalist policies for robust bimanual manipulation is bottlenecked by the scarcity of high quality large scale human demonstration data. In this work, we release 1,500 hours of diverse bimanual manipulation demonstrations covering everyday household tasks, and use this comprehensive corpus to train XR-2, a powerful vision-language-action (VLA) model. Enabled by a purpose built high throughput data pipeline and a carefully designed multi stage training paradigm, XR-2 attains strong manipulation performance in our systematic experiments while retaining favorable training efficiency and high data utilization. We further study two critical scaling axes: varying the amount of expert demonstration data, and post training on DAgger correction data from real time human interventions. In both settings, task success rate improves steadily over the data ranges we probe, exhibiting a clear consistent scaling trend at our current data scale. These results validate both the learning capacity of XR-2 and the promising scaling properties of the released dataset, which we open source to support reproducible research on bimanual robot manipulation learning.