Search papers, labs, and topics across Lattice.
The authors extend the Self-Distillation Transformer for multimodal Emotion Recognition in Conversations (ERC) by integrating geometric facial features alongside visual appearance, class-wise adaptive modality gating, and a continuous valence-arousal transition prior. Addressing the failure of conventional architectures to capture subtle affective shifts and non-uniform modality importance across discrete emotion classes, the framework systematically models conversational affective dynamics. Evaluating on MELD and IEMOCAP, the addition of geometric visual representations yields up to a 4.36-point weighted F1 gain, while the valence-arousal prior improves classification accuracy on emotionally dynamic conversational turns by up to 0.74 points without degrading performance on stable turns.
Standard multimodal fusion treats modality weights uniformly across emotional states, but coupling facial geometry with class-adaptive gating and valence-arousal transition priors drives up to a 4.36 F1 gain on conversational emotion tracking.
Emotion Recognition in Conversations (ERC) requires integrating heterogeneous textual, audio, and visual cues while accounting for conversational context and emotional dynamics. We extend the Self-Distillation Transformer architecture for ERC with appearance+geometry visual representations, class-wise adaptive modality fusion, and a valence-arousal prior for affective transitions. On the MELD and IEMOCAP datasets, geometry-enhanced visual representations improve weighted F1 by 0.27 and 4.36 points over appearance-only features, respectively, while class-wise adaptive fusion provides further gains of 0.17 and 0.25 points over the original softmax gate. The valence-arousal prior yields targeted improvements of 0.30 and 0.74 accuracy points on emotionally shifted utterances while preserving performance on stable turns. These results indicate that structured facial cues, emotion-dependent modality weighting, and affective geometry provide complementary benefits for multimodal ERC.