Search papers, labs, and topics across Lattice.
This paper introduces ChronoFlow, a temporally unified representation that integrates past, current, and future interaction dynamics in visuomotor policy learning using sparse 3D keypoints. By leveraging this representation, the authors propose ChronoFlow-Policy, a diffusion-based approach that jointly learns interaction dynamics and action sequences through a co-training objective. Experimental results across 14 simulated tasks and 5 real-world manipulation tasks show that ChronoFlow-Policy significantly outperforms existing diffusion-policy baselines, particularly in long-horizon and non-Markovian scenarios.
ChronoFlow-Policy outperforms traditional methods by effectively unifying past and future interaction dynamics, enhancing performance in complex manipulation tasks.
Visual signals play a crucial role in policy learning by enabling models to capture object motion and interaction dynamics. Just as humans reason about actions using both past experience and anticipated outcomes, effective policies should integrate past interactions with future predictions. However, existing visuomotor policies typically model either historical context or future dynamics in isolation, lacking a unified temporal representation of interaction dynamics. In this work, we introduce \textbf{ChronoFlow}, a temporally unified representation that captures \textbf{past, current, and future} interaction dynamics through sparse 3D keypoints of both objects and the gripper. Based on this representation, we propose \textbf{ChronoFlow-Policy}, a diffusion-based visuomotor policy that jointly learns ChronoFlow and action sequences through a co-training objective. Experiments on 14 simulated tasks and 5 real-world manipulation tasks demonstrate that ChronoFlow-Policy consistently outperforms strong diffusion-policy baselines and improves robustness in long-horizon and non-Markovian manipulation scenarios.