Search papers, labs, and topics across Lattice.
This paper introduces E$^3$mo-Bench, a comprehensive benchmark designed to evaluate multimodal large language models (MLLMs) on both evoked and expressed emotions through 12,314 question-answer pairs across 2,524 videos. The benchmark employs three tasks鈥攅motion perception, open-vocabulary recognition, and valence-arousal-dominance (VAD) assessment鈥攚hile utilizing a novel Bayesian Pairwise Alignment method to generate reliable continuous annotations from sparse judgments. The results reveal significant performance disparities between evoked and expressed emotion understanding, highlighting critical areas for improvement in MLLMs' emotional intelligence capabilities.
MLLMs show a striking performance gap in understanding evoked versus expressed emotions, revealing a crucial blind spot in current AI emotional intelligence.
Understanding both expressed and evoked emotions is critical for multimodal large language models (MLLMs) to achieve comprehensive affect-aware interactions. However, existing benchmarks typically examine expressed and evoked emotions in isolation or are constrained to coarse-grained and incomplete affective characterizations. To bridge this gap, we introduce E$^3$mo-Bench, a scalable benchmark comprising $12{,}314$ question-answer pairs across $2{,}524$ videos with predefined affective perspectives. It evaluates evoked and expressed emotion understanding via $3$ complementary tasks: emotion perception, open-vocabulary recognition, and valence-arousal-dominance (VAD) assessment. To efficiently scale reliable continuous annotations, we propose Bayesian Pairwise Alignment, which aggregates sparse, low-burden pairwise judgments into anchor-referenced VAD estimates. Furthermore, we develop E$^3$mo-Score, a training-free agent that aggregates complementary judgments from a five-model committee to improve VAD estimation. Extensive experiments validate the effectiveness of our framework and expose a pronounced performance skew between evoked and expressed emotion paradigms. These findings, coupled with MLLMs' persistent deficits in fine-grained recognition and dimensional assessment, chart a clear course for advancing multimodal emotional intelligence.