Search papers, labs, and topics across Lattice.
This study investigates the impact of open-loop execution in robotic manipulation, revealing that while it aids in imitating non-Markovian demonstrations, it ultimately reduces reactivity in policies. Through experiments across various tasks, the authors demonstrate that expert non-Markovianity significantly influences task success, overshadowing the previously emphasized role of compounding errors. The findings advocate for the adoption of long-context, reactive policies over traditional long open-loop execution, suggesting a shift in how imitation learning is approached in robotics.
Long open-loop execution may hinder reactivity in robotic policies, revealing that expert non-Markovianity is a more critical factor for task success than previously thought.
Action chunking --- the practice of predicting a sequence of actions and executing a prefix open-loop --- has emerged as a key enabler of recent progress in imitation learning for robotic manipulation. However, executing long open-loop prefixes reduces reactivity, limiting policies' ability to correct for errors. Further, the mechanisms underlying these performance benefits remain poorly understood: prior works cite mitigating compounding errors, absorbing inference latency, or smoothing motions, but provide limited controlled evidence or guidance for preserving reactivity. In this work, we argue that long open-loop execution primarily helps short-context policies imitate "non-Markovian demonstrations". Across four simulation and two real-world tasks, we show that expert non-Markovianity strongly shapes the relationship between task success and open-loop execution horizon. Further, we investigate the impact of compounding errors --- the prevailing explanation for long open-loop execution in prior work --- and find that while they matter, expert non-Markovianity has a much stronger impact in our experimental setting. Finally, we show that when policies are provided with a sufficiently long context, open-loop execution is no longer beneficial and the most reactive, closed-loop policies perform best. While imitation learning has seen great success using long open-loop execution, our findings motivate long-context, reactive policies as a more principled and performant paradigm.