Tsinghua AIMar 17, 2026arXiv:2603.17024

HopChain: Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning

Shenzhi Wang, Shixuan Liu, Chang Gao, Binghai Wang, An Yang, Shiji Song, Bowen Yu, Junyang Lin

AI Summary

The paper introduces HopChain, a framework for synthesizing multi-hop vision-language reasoning data to address the limitations of existing datasets in exposing complex reasoning failures in VLMs. HopChain generates logically dependent chains of instance-grounded hops, where each hop builds upon previous ones, culminating in a verifiable numerical answer. Training Qwen3.5 models with HopChain-synthesized data improves performance across 20 out of 24 diverse benchmarks, demonstrating broad and generalizable gains, particularly in long-context reasoning.

Key Contribution

Multi-hop data synthesis using HopChain boosts VLM performance across a wide range of tasks, with gains of over 50 points in accuracy for ultra-long-context reasoning.

Abstract

VLMs show strong multimodal capabilities, but they still struggle with fine-grained vision-language reasoning. We find that long CoT reasoning exposes diverse failure modes, including perception, reasoning, knowledge, and hallucination errors, which can compound across intermediate steps. However, most existing vision-language data used for RLVR does not involve complex reasoning chains that rely on visual evidence throughout, leaving these weaknesses largely unexposed. We therefore propose HopChain, a scalable framework for synthesizing multi-hop vision-language reasoning data specifically for RLVR training of VLMs. Each synthesized multi-hop query forms a logically dependent chain of instance-grounded hops, where earlier hops establish the instances, sets, or conditions needed for later hops, while the final answer remains a specific, unambiguous number suitable for verifiable rewards. We add the multi-hop data synthesized by HopChain to the original RLVR data used to train Qwen3.5-35B-A3B and Qwen3.5-397B-A17B, and compare against RLVR on the original RLVR data alone across 24 benchmarks spanning STEM and Puzzle, General VQA, Text Recognition and Document Understanding, and Video Understanding. Although this multi-hop data is not synthesized to target any specific benchmark, adding it improves 20 out of 24 benchmarks on both models, indicating broad and generalizable gains. To demonstrate that full chained queries are important, we replace them with half-multi-hop or single-hop variants, reducing the 24-benchmark average accuracy by 5.3 and 7.0 points, respectively. Multi-hop training also strengthens long-CoT vision-language reasoning, with gains peaking at more than 50 accuracy points in the ultra-long-CoT regime. These experiments establish HopChain as an effective, scalable framework for synthesizing multi-hop data that improves generalizable vision-language reasoning.

Data Curation & Synthetic Data Multimodal Models Reasoning & Chain-of-Thought

Citation Metrics

Citations0

Influential citations0

References91

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

HopChain: Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning

Related Papers