Search papers, labs, and topics across Lattice.
This paper introduces OPLD (On-Policy Latent Distillation), a novel framework that enhances multimodal reasoning by transferring reasoning capabilities from interleaved multimodal Chain-of-Thought (CoT) into latent representations. By focusing on supervising latent representations at the reasoning-process level rather than relying on feature-level alignment, OPLD effectively internalizes abstract reasoning processes. The extensive experiments demonstrate that OPLD outperforms existing methods and achieves state-of-the-art results across multiple multimodal benchmarks, highlighting its significance in advancing visual reasoning capabilities.
Supervising latent representations at the reasoning-process level leads to a breakthrough in multimodal reasoning, outperforming traditional methods by a significant margin.
Interleaved multimodal Chain-of-Thought (CoT) improves visual reasoning by incorporating auxiliary visual evidence into intermediate reasoning. However, existing approaches remain constrained by externally defined reasoning traces and visual operations, limiting their ability to develop flexible and abstract visual thinking. Reasoning with latent has recently offered a promising direction by internalizing intermediate computation into continuous representations. Nevertheless, existing visual-latent methods mainly supervise latent states through alignment with compressed auxiliary visual features, treating them as proxies for visual observations rather than active reasoning states. Consequently, they capture the provided evidence but fail to fully internalize the abstract reasoning process induced by multimodal CoT. In this paper, we propose OPLD (On-Policy Latent Distillation), a simple framework that transfers the reasoning capability induced by privileged multimodal CoT into latent reasoning representations. Extensive experiments on diverse multimodal benchmarks demonstrate that OPLD consistently outperforms existing latent reasoning methods and achieves state-of-the-art performance on multiple benchmarks. The results suggest that supervising latent representations at the reasoning-process level provides a more effective paradigm for multimodal latent reasoning than conventional feature-level alignment.