Search papers, labs, and topics across Lattice.
This study investigates the presence of emotion-sensitive neurons (ESNs) in three multimodal foundation models (MFMs) to determine whether they utilize shared mechanisms for recognizing emotions across speech and facial expressions. By employing both speech and facial emotion recognition tasks, the authors demonstrate that visual ESNs are causally significant, as their deactivation impairs emotion recognition while their activation enhances it. The results reveal a structural alignment between acoustic and visual ESNs, indicating that these models leverage overlapping affective representations across modalities, with cross-modal interventions showing bidirectional causal transfer of emotion-specific effects.
Emotion-sensitive neurons in multimodal models reveal shared mechanisms for recognizing emotions across speech and faces, with implications for enhancing emotion recognition capabilities.
Modern multimodal foundation models (MFMs) have made rapid progress on tasks requiring integrated perception across speech, vision, and language, including emotion recognition. However, it remains unclear whether they recognize speech and facial emotion through shared affective functional units or modality-specific pathways. We explore emotion-sensitive neurons (ESNs), sparse decoder neurons selectively associated with emotion categories, in three MFMs: Gemma-4-12B-it, MiniCPM-o-4.5, and Qwen2.5-Omni-7B. Using speech emotion recognition and facial expression recognition as complementary probes, we identify acoustic and visual ESNs. Visual ESNs are causally meaningful: deactivating them selectively impairs recognition of the associated facial emotion, whereas steering their activations selectively enhances recognition of that emotion relative to other emotion categories. Acoustic and visual ESNs further show emotion-matched overlap and similar layer-wise distributions, indicating partial structural alignment between affective representations across speech and faces. Finally, cross-modal interventions reveal bidirectional causal transfer: ESNs identified from one modality produce emotion-specific effects when applied to the other. Our findings provide one of the first cross-modality activation-level analyses of affective functional units in MFMs, suggesting that speech and facial emotion recognition partially converge onto sparse decoder-level components that can be localized and manipulated without training.