Search papers, labs, and topics across Lattice.
This paper introduces CM2, a multi-agent framework designed to enhance multimodal cultural reasoning in Large Language Models (LLMs) by mimicking human cognitive processes. By integrating multimodal perception, retrieval-augmented generation, networked reasoning, gated fusion, and reward-driven feedback, CM2 outperforms traditional reasoning paradigms in interdisciplinary contexts. Experiments demonstrate significant improvements across various MLLM backbones, with ablation studies confirming the effectiveness of each component and conflict analyses validating cross-modal arbitration capabilities.
CM2 achieves unprecedented gains in cultural reasoning for MLLMs, revealing the power of a multi-agent approach to interdisciplinary challenges.
Multimodal Large Language Models (MLLMs) have shown remarkable success in STEM domains, where progress is often driven by vertical, step-by-step deduction under relatively stable symbol systems. Their horizontal, interdisciplinary cultural reasoning, however, remains underexplored.We propose CM2, a multi-agent framework grounded in the cognitive pathway of human cultural interpretation. CM2 integrates multimodal perception, retrieval-augmented generation, networked reasoning, gated fusion, and reward-driven feedback.Experiments on CM2D across multiple MLLM backbones show consistent gains over CoT and typical reasoning paradigms; ablations validate each module's contribution, and conflict analyses confirm genuine cross-modal arbitration.