Search papers, labs, and topics across Lattice.
This paper introduces a novel text-centered multimodal approach for recognizing ambivalence and hesitancy in video data, addressing the challenges posed by inconsistent expressions across linguistic, acoustic, and facial modalities. By employing a Text Residual Fusion model that treats text as the anchor modality, the authors achieve significant performance improvements, with an average Macro F1-score of 78.24% on the Private Test subset. The findings highlight that integrating multimodal features can enhance recognition accuracy while avoiding the computational burdens associated with ensemble methods.
Text-centered multimodal fusion boosts ambivalence recognition accuracy by over 4% compared to traditional text-only models, challenging the need for complex ensembles.
Automatic recognition of ambivalence and hesitancy is challenging because these states may be expressed through inconsistent linguistic, acoustic, facial, and contextual patterns, while top-performing systems often rely on computationally expensive ensembles. We present a single text-centered multimodal approach for video-level ambivalence and hesitancy recognition for the 11th Affective&Behavior Analysis in-the-Wild (ABAW) Challenge. The proposed approach combines linguistic, acoustic, facial, and scene features using text-centered multimodal fusion model. Text Residual Fusion treats text as the anchor modality and applies gated residual adjustments based on the other modalities. Experiments on the Behavioural Ambivalence/Hesitancy (BAH) corpus confirm that text is the strongest unimodal modality. The Text Residual Fusion model achieves an average Macro F1-score (MF1) of 75.14% across the Development and Public Test subsets. On the Private Test subset, it reaches an MF1 of 78.24%, outperforming the text model by 4.03%. These results demonstrate that complementary multimodal information can improve recognition performance without requiring a large model ensemble.