Search papers, labs, and topics across Lattice.
This paper introduces Implicit Social Context Analysis (MoCA), a systematic framework for studying implicit social communication across dimensions of affection, intent, and stance. By constructing a benchmark of 3,108 multimodal instances with detailed cognitive annotations, the authors reveal that existing multimodal large language models struggle with implicit cues due to their reliance on explicit signals. The proposed Conflict-Driven Abductive Reasoning (CoDAR) framework significantly enhances model performance but still leaves a notable gap compared to human reasoning, underscoring the complexity of implicit social understanding.
State-of-the-art multimodal models falter in interpreting implicit social cues, revealing a critical gap in AI's understanding of human communication.
Human social communication, such as affection and intent, is often conveyed in highly implicit ways, where underlying meanings are expressed through indirect, socially and culturally grounded signals rather than explicit statements. Such implicit social contexts are pervasive in real-world interactions, yet there remains a lack of a formal and systematic framework for studying them. In this paper, we introduce Implicit Social Context Analysis (MoCA), a novel task that systematically models implicit social scenarios along three key dimensions: affection, intent, and stance. We construct a high-quality benchmark containing 3,108 multimodal instances collected from real-world sources, with fine-grained cognitive annotations revealing who expresses what toward whom, as well as how and why it is conveyed. Using the MoCA dataset, we show that state-of-the-art multimodal large language models struggle significantly with this task because of their reliance on explicit cues and limited ability to reason over latent social contexts. To address this challenge, we propose Conflict-Driven Abductive Reasoning (CoDAR), a novel framework that models the discrepancy between observed expressions and expected truthful behavior as cognitive conflict, thereby enabling the inference of hidden mental states. Extensive experiments demonstrate that CoDAR substantially improves model performance. Nevertheless, a large gap from human reasoning remains, highlighting the fundamental difficulty of implicit social understanding.