Search papers, labs, and topics across Lattice.
This paper introduces a latency-aware reinforcement learning framework, Asynchronous RL with Intermediate Information (ARLI), designed to improve generalist robot policies despite the challenges posed by inference latency. By interleaving action generation with execution and incorporating state augmentations, ARLI restores the near-Markovian structure necessary for effective RL, enabling robots to finetune their policies in real-time environments. Experimental results demonstrate that ARLI not only overcomes the limitations of standard RL under latency but can also match or exceed performance in ideal conditions.
Robots can now learn and adapt in real-time, overcoming inference delays that previously stymied reinforcement learning effectiveness.
While reinforcement learning (RL) allows generalist robot policies to continually improve during deployment, the large model size of modern generalist policies, such as VLAs, poses a fundamental obstacle to effective RL improvement. In particular, their severe inference latency---which can lead to pauses or jerky movements---can alter the effective environment dynamics and, if not correctly accounted for, break the Markov assumption that RL relies on, causing standard RL algorithms to fail completely. In this work, we introduce a latency-aware framework, Asynchronous RL with Intermediate Information (ARLI), that enables RL-based improvement of generalist policies under inference delays. Our framework builds on asynchronous inference approaches, which interleave action generation with execution to hide latency, and addresses its incompatibility with RL by providing a low-latency RL policy design that maximizes reactivity within the inference window through two contributions: state augmentations that restore near-Markovian structure by incorporating committed actions and a mid-inference observation. We evaluate our approach across simulated and real-world manipulation tasks, and find that it enables effective finetuning under inference delays where standard RL fails entirely, even matching or exceeding the performance of standard RL in idealized no-latency settings.