Search papers, labs, and topics across Lattice.
This paper introduces InCarEmo, a novel multimodal dataset designed for in-cabin emotion recognition and driver state monitoring, addressing the limitations of existing datasets that primarily focus on visual modalities. By integrating RGB and infrared video, in-cabin audio, and dialogue text from scripted scenarios, InCarEmo enables comprehensive analysis of driver emotions under varied conditions. Experimental results highlight the advantages of multimodal fusion for tasks such as emotion recognition, fatigue detection, and distraction monitoring, while also identifying challenges posed by real-world noise and low-light environments.
Multimodal fusion significantly enhances emotion recognition accuracy in vehicles, revealing critical insights into driver state monitoring.
Understanding driver emotion and state is critical for the next generation of intelligent in-cabin systems that ensure safety and enhance human-vehicle interaction. However, existing public datasets for in-cabin affective computing are largely limited to visual modalities and rarely include conversational information, making it difficult to capture the linguistic and interactive cues underlying driver emotion. To address these gaps, we introduce InCarEmo, a multimodal dataset for in-cabin emotion recognition and driver state monitoring. InCarEmo integrates RGB and infrared video, in-cabin audio, and dialogue text collected from scripted in-cabin scenarios designed to simulate realistic driver behaviors, covering diverse lighting conditions and driving contexts. The dataset supports three primary tasks: 1) multimodal emotion recognition, 2) fatigue detection, and 3) distraction monitoring. In addition to the original Chinese data, we construct an auxiliary English benchmark to support preliminary cross-lingual evaluation. We provide a unified benchmark with extensive baseline results across unimodal and multimodal methods, including analyses under modality-missing and noise conditions. Experimental results demonstrate the benefits of multimodal fusion and reveal remaining challenges under real-world noise and low-light conditions. By releasing InCarEmo, we aim to establish a comprehensive foundation for robust, interpretable, and human-centric in-cabin affective understanding, promoting safer and more empathetic driver-vehicle interaction.