Search papers, labs, and topics across Lattice.
This paper introduces MMHBench, a novel multimodal benchmark designed to enhance mental health understanding in long-form videos by incorporating nuanced reasoning across observable behaviors and psychological states. The benchmark includes 268 videos and 2,184 questions, organized into third-person assessments and first-person perspective-taking tasks, which reveal the limitations of existing models that often rely on superficial correlations. Evaluation of 22 multimodal large language models (MLLMs) shows that accurately interpreting mental health in this context remains a significant challenge, highlighting the need for more sophisticated approaches in AI understanding of psychological phenomena.
Long-form video analysis reveals that even advanced MLLMs struggle with nuanced mental health interpretations, underscoring a critical gap in AI understanding.
Mental health understanding in long-form videos requires nuanced reasoning over observable behavior, interpersonal context, and latent psychological states. Existing benchmarks largely reduce this task to coarse-grained classification, providing limited insight into whether models truly understand psychological phenomena or rely on superficial correlations. To address this limitation, we introduce MMHBench, a comprehensive multimodal benchmark for multi-perspective mental health understanding, comprising 268 long-form videos and 2,184 carefully curated questions. MMHBench organizes the evaluation into two complementary settings: (1) third-person assessment, consisting of 605 questions that focus on the interpretation of observable behaviors and multimodal evidence, and (2) first-person perspective-taking, comprising 1,579 questions that require perspective-conditioned reasoning to identify the interpretation of the mental state supported by the available multimodal evidence. We propose a Multi-Agent Question Generation (MAQG) framework that simulates diverse social roles to synthesize questions from multiple perspectives. The generated questions are refined through multi-role feedback and iterative optimization, followed by expert-guided verification to ensure quality and validity. Extensive evaluation of 22 representative multimodal large language models (MLLMs), spanning both open-source and leading closed-source models, demonstrates that long-form video mental health understanding remains highly challenging.