HUSTFeb 26, 2026arXiv:2602.23306

ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding

Yiran Guan, Sifan Tu, Sifan Tu, Dingkang Liang, Linghao Zhu, Linghao Zhu, Jianzhong Ju, Zhenbo Luo, Zhenbo Luo, Jian Luan, Yuliang Liu

AI Summary

The paper introduces ThinkOmni, a training-free framework that enhances the reasoning capabilities of Omni-modal Large Language Models (OLLMs) by leveraging Large Reasoning Models (LRMs) to guide the decoding process. ThinkOmni employs "LRM-as-a-Guide" and "Stepwise Contrastive Scaling" to balance perception and reasoning signals adaptively. Experiments on six multi-modal reasoning benchmarks show consistent performance improvements, achieving 70.2 on MathVista and 75.5 on MMAU, demonstrating the framework's effectiveness in lifting textual reasoning to omni-modal scenarios.

Key Contribution

Unleashing powerful reasoning in OLLMs doesn't require expensive training data or compute – just clever guidance from existing Large Reasoning Models.

Abstract

Omni-modal reasoning is essential for intelligent systems to understand and draw inferences from diverse data sources. While existing omni-modal large language models (OLLM) excel at perceiving diverse modalities, they lack the complex reasoning abilities of recent large reasoning models (LRM). However, enhancing the reasoning ability of OLLMs through additional training presents significant challenges, including the need for high-quality data, task-specific adaptation, and substantial computational costs. To address these limitations, we propose ThinkOmni, a training-free and data-free framework that lifts textual reasoning to omni-modal scenarios. ThinkOmni introduces two key components: 1) LRM-as-a-Guide, which leverages off-the-shelf LRMs to guide the OLLM decoding process; 2) Stepwise Contrastive Scaling, which adaptively balances perception and reasoning signals without manual hyperparameter tuning. Experiments on six multi-modal reasoning benchmarks demonstrate that ThinkOmni consistently delivers performance improvements, with main results achieving 70.2 on MathVista and 75.5 on MMAU. Overall, ThinkOmni offers a flexible and generalizable solution for omni-modal reasoning and provides new insights into the generalization and application of reasoning capabilities.

Multimodal Models Reasoning & Chain-of-Thought Tool Use & Agents

Citation Metrics

Citations0

Influential citations0

References47

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding

Related Papers