Shanghai AI LabFeb 24, 2026arXiv:2602.20739

PyVision-RL: Forging Open Agentic Vision Models via RL

Shitian Zhao, Shitian Zhao, Shaoheng Lin, Shaoheng Lin, Ming Li, Haoquan Zhang, Wenshuo Peng, Kaipeng Zhang, Kaipeng Zhang

AI Summary

The paper introduces PyVision-RL, a reinforcement learning framework designed to mitigate interaction collapse in agentic multimodal models by encouraging sustained tool usage and multi-turn reasoning. They achieve this by employing an oversampling-filtering-ranking rollout strategy combined with an accumulative tool reward. The framework is used to develop PyVision-Image and PyVision-Video, demonstrating improved performance and efficiency through on-demand context construction for video reasoning, which reduces visual token usage.

Key Contribution

Reinforcement learning for multimodal agents doesn't have to collapse into uselessness: PyVision-RL shows how to stabilize training and encourage multi-turn tool use.

Abstract

Reinforcement learning for agentic multimodal models often suffers from interaction collapse, where models learn to reduce tool usage and multi-turn reasoning, limiting the benefits of agentic behavior. We introduce PyVision-RL, a reinforcement learning framework for open-weight multimodal models that stabilizes training and sustains interaction. Our approach combines an oversampling-filtering-ranking rollout strategy with an accumulative tool reward to prevent collapse and encourage multi-turn tool use. Using a unified training pipeline, we develop PyVision-Image and PyVision-Video for image and video understanding. For video reasoning, PyVision-Video employs on-demand context construction, selectively sampling task-relevant frames during reasoning to significantly reduce visual token usage. Experiments show strong performance and improved efficiency, demonstrating that sustained interaction and on-demand visual processing are critical for scalable multimodal agents.

Multimodal Models RLHF & Preference Learning Tool Use & Agents

Citation Metrics

Citations0

Influential citations0

References64

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

PyVision-RL: Forging Open Agentic Vision Models via RL

Related Papers