Search papers, labs, and topics across Lattice.
This paper introduces MatchingPolicy, a correspondence-driven framework designed to enhance in-context imitation learning by decoupling demonstration-to-scene matching from policy learning. By employing a correspondence-aware diffusion policy that directly conditions robotic actions on dense semantic correspondences, the method effectively addresses the challenges of maintaining performance on unseen objects and novel scenarios. Extensive evaluations on RLBench and real-world manipulation tasks demonstrate that MatchingPolicy significantly improves few-shot performance and generalization across diverse object instances and semantic categories.
Robust out-of-distribution transfer is achieved by decoupling demonstration matching from policy learning, allowing robots to generalize effectively to unseen objects.
In-context imitation learning enables few-shot policy generalization but struggles to maintain performance on unseen objects and novel scenarios. To address this, we introduce MatchingPolicy, a correspondence-driven framework that explicitly decouples demonstration-to-scene matching from policy learning. Central to our method is a correspondence-aware diffusion policy that conditions robotic actions directly on dense semantic correspondences. This architectural separation resolves the inherent conflict between correspondence identification and action adaptation, enabling robust out-of-distribution transfer. Our framework integrates vision foundation models with a novel two-stage matching algorithm to dynamically establish reliable correspondences. Extensive evaluations on RLBench and real-world manipulation tasks confirm that MatchingPolicy achieves superior few-shot performance, generalizing reliably across unseen object instances and semantic categories.