Search papers, labs, and topics across Lattice.
This paper systematically evaluates the capabilities and limitations of multimodal large language models (MLLMs) in the context of remote sensing image scene understanding (RSISU). It reveals that while domain-specific RS-MLLMs excel in certain tasks like visual grounding and high-resolution question answering, general-purpose computer vision MLLMs often outperform them without requiring specialized fine-tuning. The study highlights significant challenges in spatial reasoning and generalization, suggesting a need for improved evaluation and development strategies for MLLMs in remote sensing applications.
General-purpose computer vision MLLMs can outperform specialized remote sensing models on key tasks, challenging the notion of domain-specific superiority.
The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language interaction with remote sensing imagery. However, a systematic understanding of the capability boundaries, cross-task generalization, and task-specific limitations of existing remote sensing MLLMs (RS-MLLMs) is still lacking. This paper presents a systematic survey and diagnostic evaluation of MLLMs for RSISU. We review the technical evolution of RS-MLLMs, focusing on model design, multimodal learning, training data, and downstream capabilities. We further compare RS-MLLMs with general-purpose computer vision MLLMs (CV-MLLMs) across diverse RSISU tasks and benchmarks. RS-MLLMs remain competitive in domain-specific settings, particularly remote sensing visual grounding and high-resolution visual question answering. More notably, general-purpose CV-MLLMs can match or even outperform these specialized models on several RSISU tasks without remote sensing-specific fine-tuning. These findings demonstrate the strong transferability of general-purpose CV-MLLMs and show that current RS-MLLMs do not consistently outperform them across diverse RSISU tasks. Current MLLMs also face limitations in spatial and relational reasoning, fine-grained visual understanding, instruction diversity, and generalization across heterogeneous task formats. Based on these findings, we outline future directions toward reliable evaluation, multimodal and high-resolution reasoning, efficient deployment, and tool-augmented remote sensing agents. This survey provides a systematic reference for developing robust, generalizable, and practical MLLMs for RSISU.