Search papers, labs, and topics across Lattice.
This paper introduces SmartMage, a unified Multimodal Large Language Model (MLLM) designed to enhance 3D scene understanding by dynamically orchestrating heterogeneous modalities based on the specific needs of each query. By implementing a Semantic-guided Modality Adaptive Routing (SMART) module and a Modality-Aware Gating Expert (MAGE) module, SmartMage effectively selects and activates task-relevant modalities, reducing semantic noise and improving computational efficiency. The model achieves state-of-the-art performance across five benchmarks for 3D scene understanding and demonstrates competitive results in RGB-only video understanding tasks.
SmartMage redefines 3D scene understanding by dynamically selecting modalities, achieving state-of-the-art results while minimizing irrelevant computations.
Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, including visual and geometric cues. However, the relevance of these modalities often varies across queries. Existing Multimodal Large Language Models (MLLMs) typically rely on fixed modality combinations, overlooking query-dependent modality needs. Such a rigid design can introduce semantic noise from irrelevant modalities while underutilizing more informative ones, leading to wasted computation and diluted reasoning. To address these challenges, this paper proposes SmartMage, a unified MLLM that dynamically orchestrates heterogeneous modalities for semantic-aware 3D scene understanding. Specifically, SmartMage incorporates: (1) a Semantic-guided Modality Adaptive RouTng (SMART) module that selects task-relevant modalities using semantic priors, text-modality alignment, and modality quality; and (2) a Modality-Aware Gating Expert (MAGE) module that leverages modality priors to guide expert activation, fostering adaptive specialization in multimodal reasoning. Empirically, SmartMage achieves state-of-the-art performance across five 3D scene understanding benchmarks, and attains competitive results on RGB-only video understanding benchmarks. In our diagnostic benchmark ScanFacet, tasks are divided into fine-grained semantic categories, enabling analysis of modality combinations preferred by each semantic type. The observed modality-semantic patterns provide further evidence of SmartMage's effectiveness. Project page: https://yuecheong.github.io/SmartMage/.