Tsinghua AIGigaAIApr 9, 2026arXiv:2604.08168

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning

Jindi Lv, Hao Li, Jie Li, Yifei Nie, Fankun Kong, Yang Wang, Xiaofeng Wang, Zheng Zhu, Chaojun Ni, Qiuping Deng, Hengtao Li, Jiancheng Lv, Guan Huang

AI Summary

This paper introduces ViVa, a video-generative value model for robot reinforcement learning that leverages a pretrained video generator to predict future proprioception and a scalar value. By grounding value estimation in anticipated embodiment dynamics, ViVa overcomes limitations of existing VLM-based value models in capturing temporal dynamics for long-horizon tasks. Experiments integrating ViVa into RECAP demonstrate substantial improvements on real-world box assembly and generalization to novel objects, showcasing the potential of video-generative models for value estimation.

Key Contribution

Robots can now better assemble boxes in the real world thanks to a video-generative value model that anticipates future states, moving beyond static snapshots for more reliable task progress assessment.

Abstract

Vision-language-action (VLA) models have advanced robot manipulation through large-scale pretraining, but real-world deployment remains challenging due to partial observability and delayed feedback. Reinforcement learning addresses this via value functions, which assess task progress and guide policy improvement. However, existing value models built on vision-language models (VLMs) struggle to capture temporal dynamics, undermining reliable value estimation in long-horizon tasks. In this paper, we propose ViVa, a video-generative value model that repurposes a pretrained video generator for value estimation. Taking the current observation and robot proprioception as input, ViVa jointly predicts future proprioception and a scalar value for the current state. By leveraging the spatiotemporal priors of a pretrained video generator, our approach grounds value estimation in anticipated embodiment dynamics, moving beyond static snapshots to intrinsically couple value with foresight. Integrated into RECAP, ViVa delivers substantial improvements on real-world box assembly. Qualitative analysis across all three tasks confirms that ViVa produces more reliable value signals, accurately reflecting task progress. By leveraging spatiotemporal priors from video corpora, ViVa also generalizes to novel objects, highlighting the promise of video-generative models for value estimation.

Multimodal Models Robotics & Embodied AI World Models & Planning

Citation Metrics

Citations0

Influential citations0

References0

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

ViVa: A Video-Generative Value Model for Robot Reinforcement Learning

Related Papers