Search papers, labs, and topics across Lattice.
This paper addresses the challenges of robotic manipulation by introducing an agentic reinforcement learning framework that enhances execution stability in the face of uncertainty and long-horizon tasks. The authors develop two metrics to assess execution quality in real-time and implement a high-level decision-making policy that enables the robot to recover from execution deviations. Their approach shows significant improvements in success rates, with up to 39.2% enhancement under disturbance conditions on the LIBERO benchmark, highlighting its effectiveness in maintaining robust execution.
Robots can now recover from execution failures with up to 39.2% better success rates by leveraging high-level decision-making in uncertain environments.
Robotic manipulation poses fundamental challenges due to uncertainty, long-horizon execution, and compounding errors, which can easily destabilize execution and lead to task failure. Although recent vision-language-action (VLA) models exhibit strong generalization, they typically lack explicit mechanisms to assess execution stability and to recover when execution deviates from its nominal behavior. In this paper, we propose: (1) two complementary metrics to assess execution quality at runtime, and (2) an agentic reinforcement learning framework that learns to restore effective execution through high-level decision-making rather than directly learning low-level actions. In this framework, an agentic policy reasons over recent execution history and selects among a small set of execution modes to regulate the execution process. Under execution degradation, it triggers appropriate recovery mechanisms to restore the robot to previously visited nominal states, enabling the task to continue. We evaluate the proposed method on the LIBERO benchmark, achieving up to a 13.7% improvement in success rate under standard settings and up to a 39.2% improvement under disturbance settings, demonstrating substantially enhanced execution robustness.