Search papers, labs, and topics across Lattice.
This paper introduces SpEmoC, a comprehensive multimodal emotion benchmark comprising 306,544 clips from English movies and TV series, curated to ensure class balance and modality synchronization. The dataset features a hybrid annotation pipeline that combines pretrained models with human validation, focusing on seven distinct emotions, including often-overlooked categories like Fear and Disgust. Rigorous evaluations demonstrate that balanced data and strategic dataset partitioning significantly enhance model generalization and stability across various emotion recognition tasks.
Balanced datasets can dramatically improve emotion recognition performance, especially for minority emotions often neglected in existing benchmarks.
Understanding human emotions in spoken conversations is a key challenge in affective computing, with applications in empathetic AI, human computer interaction, and mental health monitoring. However, existing datasets vary in scale, emotion distribution, modality alignment, and data partitioning strategies, which can influence reliable cross-dataset generalization and minority-emotion modeling. We introduce SpEmoC a Speaking segment Emotion for Conversations comprising 306,544 raw clips from 3,100 English language movies and TV series. From these, 30,000 high quality, class balanced clips are curated, featuring synchronized visual, audio, and textual modalities annotated for seven emotions through a hybrid pipeline that integrates pretrained models with human validation. SpEmoC uses strict movie- and series-level splits to prevent content overlap between split sets, allowing more reliable evaluation of model generalization. The dataset also maintains a near-balanced distribution across seven emotions, including minority classes such as Fear and Disgust, which supports more balanced learning across categories. Extensive experiments, including in-domain benchmarking, cross-dataset transfer, low-data training, class-imbalance analysis, and modality transfer show that balanced data and careful splitting lead to more stable performance across emotions when models are evaluated on other datasets. These results highlight the importance of dataset design for robust and transferable multimodal emotion recognition.