Search papers, labs, and topics across Lattice.
This study tackles the challenge of emotion recognition in virtual reality (VR) environments where head-mounted displays (HMDs) occlude the upper face, limiting traditional facial expression analysis. By fusing lower-face video with upper-face electromyography (EMG) data, the authors classify seven emotional categories, achieving a macro-F1 score of 51% with their late-fusion architecture. This approach not only surpasses the performance of image-only and EMG-only methods but also provides a robust solution for real-time affective assessment in VR applications.
Fusing lower-face video with upper-face EMG boosts emotion recognition accuracy in VR, achieving a 51% macro-F1 score despite facial occlusion.
Head-mounted displays (HMDs) fundamentally limit emotion recognition in virtual reality (VR): by occluding the upper face, they render conventional image-based facial expression analysis incomplete, particularly for applications requiring real-time affective assessment. We address this challenge by fusing lower-face video with facial electromyography (EMG) from the occluded upper face to classify seven emotional categories (six basic emotions plus neutral). We introduce a synchronized multimodal dataset from 20 participants, pairing lower-face video with seven-channel upper-face EMG elicited by validated emotion stimuli. Under subject-independent test, our proposed late-fusion architecture merging convolutional visual embeddings with RBF-kernel EMG representations achieves 51% macro-F1, outperforming both image-only (41%) and EMG-only (43%) baselines. These results demonstrate that upper-face EMG provides robust complementary information under HMD-induced visual occlusion and establish a foundation for multimodal emotion recognition in naturalistic VR environments. This approach facilitates affect-adaptive applications, including communication training and therapeutic interventions. The dataset will be shared upon request under an ethical-use agreement.