Search papers, labs, and topics across Lattice.
This paper introduces DualG-MRAG, a novel framework that decouples macro-reasoning from micro-matching in Multimodal Retrieval-Augmented Generation (MM-RAG) to enhance performance on complex multi-hop reasoning tasks. By constructing separate Macro and Micro Graphs, the approach effectively mitigates retrieval noise while preserving critical local evidence, leading to improved accuracy in question answering. Extensive experiments reveal that DualG-MRAG significantly surpasses existing methods in both evidence recall and QA accuracy, underscoring its effectiveness in multimodal scenarios.
Isolating global reasoning from local evidence in multimodal retrieval can dramatically boost QA accuracy and evidence recall.
While Multimodal Retrieval-Augmented Generation (MM-RAG) has shown promising results, it still struggles with complex multi-hop reasoning tasks. Existing methods primarily focus on independent instance-level matching, which often fails to capture explicit relationships across modalities and documents. Although Graph-enhanced methods introduce structural modeling, they face a fundamental challenge in multimodal scenarios: incorporating fine-grained visual features leads to rapid graph expansion and retrieval noise, whereas coarse-grained representations cause the discarding of critical local evidence. To address this dilemma, we propose DualG-MRAG, a Dual-tier framework that introduces a decoupled architecture comprising Macro-reasoning and Micro-matching Graphs for Multimodal RAG. Specifically, to suppress retrieval noise by isolating global structural reasoning from fine-grained evidence matching, we construct a Macro Graph for global topological routing and a Micro Graph for precise local verification. Subsequently, to enable dynamic relevance propagation across heterogeneous evidence sources, we formulate retrieval as a query-driven message passing process via a GNN Retriever. Furthermore, to provide the generative model with coherent structural guidance, we introduce a dynamic programming decoding mechanism that extracts explicit reasoning paths directly from the GNN's forward pass, replacing the standard input of isolated document chunks. Extensive experiments demonstrate that DualG-MRAG outperforms baselines in both evidence recall and complex QA accuracy.