Search papers, labs, and topics across Lattice.
This paper introduces ChainVLA, a 1.2B-parameter vision-language-action (VLA) policy designed to enhance long-horizon manipulation by maintaining a unified execution state that integrates both task progress and motion continuity. By employing a combination of a recurrent Working State and sparse event memory, ChainVLA effectively retains knowledge of prior actions while adapting to new observations, significantly improving performance on benchmark tasks. The model achieves an average success rate of 62.8% on RMBench and 98.8% across four LIBERO suites, highlighting the critical role of its dual components in facilitating effective action prediction and execution.
ChainVLA achieves a remarkable 62.8% success rate on long-horizon manipulation tasks by seamlessly chaining vision-language-action queries through a unified execution state.
Humans perform long-horizon manipulation by retaining knowledge of what earlier actions have established while continuously adapting the motion underway. By contrast, action-chunked vision-language-action (VLA) policies repeatedly replan from the current input at each query. Existing methods preserve either long-term task evidence through memory or short-term motion through action reuse and ensembling, leaving the cross-query handoff incomplete. We introduce ChainVLA, a 1.2B-parameter VLA policy that chains successive queries through a joint and revisable execution state. Progress Context combines a recurrent Working State with sparse event memory to carry observation-derived task progress, while Motion Tail feeds the preceding prediction's unexecuted continuation into state construction and action generation. Together, the two components condition a decoder that regenerates each action horizon under the latest observation, allowing the carried state to guide the next prediction without fixing it. ChainVLA reaches 62.8% average success on RMBench and 98.8% across four LIBERO suites, while removing Motion Tail or Progress Context reduces RMBench success to 11.2% and 3.0%, respectively. These asymmetric ablations are consistent with motion continuity helping preserve the observation stream from which task progress is inferred.