Search papers, labs, and topics across Lattice.
This paper investigates the impact of situational illusions鈥攄iscrepancies between real-world appearances and their underlying physical states鈥攐n the performance of multimodal large language models (MLLMs). By developing a comprehensive taxonomy and the MSIBench benchmark, the authors evaluate 27 model configurations, revealing significant vulnerabilities and six common failure modes in MLLMs related to visual observation, grounding, and reasoning. To address these issues, the authors propose two mitigation strategies that enhance model performance by up to 20%, paving the way for more reliable multimodal understanding in complex environments.
Current multimodal large language models are highly susceptible to situational illusions, revealing critical vulnerabilities in their reasoning and perception capabilities.
Real-world situation appearances can deviate from their underlying physical states, challenging the reliability of multimodal large language models (MLLMs) in practical applications. In this paper, we term this phenomenon situational illusions and investigate: (1) how MLLMs perform under such illusions, and (2) how to mitigate the limitations. We first develop a comprehensive where-what-how taxonomy that characterizes where situational illusions occur, what targets they take, and how they arise. Building on this taxonomy, we introduce MSIBench, a benchmark designed to assess the discrimination, understanding, and reasoning capabilities of MLLMs under situational illusions. Evaluations of 27 model configurations reveal that current MLLMs are highly vulnerable to these illusions and exhibit 6 typical failure modes related to visual observation, grounding, and reasoning. To mitigate the limitations, we build on the core idea of systematically inspecting and reasoning over visual evidence for contextual understanding, developing prompting for closed-source models and supervised fine-tuning for open-source models, respectively. These two simple yet effective methods improve model performances by 20% at most, suggesting a practical path toward more reliable multimodal perception and reasoning in complex real-world environments.