Search papers, labs, and topics across Lattice.
This paper introduces MeetingToM, a benchmark designed to evaluate the Theory of Mind (ToM) reasoning capabilities of Multimodal Large Language Models (MLLMs) in the context of multi-party meetings. It addresses the inadequacies of existing benchmarks by focusing on complex social behaviors, such as pseudo-consensus, and organizes evaluations across different levels of social granularity. Systematic analyses of representative MLLMs reveal significant limitations in their ability to integrate non-verbal cues and accurately infer hidden attitudes, underscoring the need for improved ToM in multimodal contexts.
MLLMs struggle to discern genuine consensus from pseudo-consensus in multi-party meetings, revealing critical gaps in their Theory of Mind reasoning abilities.
Theory of Mind (ToM), the ability to infer other's beliefs, intentions, and states of knowledge, is central to social interaction, yet remains challenging for current Multimodal Large Language Models (MLLMs), especially in multi-party meetings where cues are distributed across speech and behavior. Existing multimodal ToM benchmarks mainly focus on video-grounded question answering over overt, externally verifiable signals, and provide limited coverage of latent social states and group dynamics. We introduce MeetingToM, a benchmark for complex social behavior reasoning in naturalistic multi-party meetings. MeetingToM targets meeting-specific phenomena such as \textbf{pseudo-consensus}, where apparent agreement masks private dissent under social pressure. The benchmark is hierarchically organized to evaluate ToM at increasing levels of social granularity, including (i) subject-level mental state prediction, (ii) dyadic-level addressee understanding, and (iii) group-level consensus reasoning. We provide a unified evaluation protocol and conduct systematic analyses of representative MLLMs, revealing persistent limitations in integrating non-verbal cues, inferring hidden attitudes, and distinguishing genuine consensus from pseudo-consensus. Our results highlight key challenges and establish MeetingToM as a testbed for advancing meeting-grounded ToM in multimodal models.