Search papers, labs, and topics across Lattice.
This study investigates how large audio-language models (LALMs) encode emotional expressions across different languages by introducing the concept of Multilingual Emotion Neurons (MLENs), which exhibit stable emotional selectivity and causal effects. Using Consistency-Regularized Fusion (CR-Fusion), the authors demonstrate that emotion-sensitive neurons identified in individual languages show minimal overlap, highlighting the necessity of pooling cross-lingual evidence for effective neuron identification. The findings reveal that MLENs provide superior affective control in zero-shot and low-resource scenarios compared to traditional monolingual approaches, emphasizing the importance of multilingual neuron identification for understanding emotional communication in LALMs.
Emotion-sensitive neurons in LALMs are language-specific, but pooling cross-lingual evidence reveals powerful, transferable Multilingual Emotion Neurons that enhance affective control.
Emotion is central to human communication, and its expression varies across languages. Large audio-language models (LALMs) achieve strong performance on multilingual speech tasks, yet it remains unclear whether they encode emotion through language-specific correlations or language-agnostic representations. We present the first neuron-level interpretability study of this question. We define Multilingual Emotion Neurons (MLENs) as functional units exhibiting stable emotional selectivity and aligned causal effects across languages, and introduce Consistency-Regularized Fusion (CR-Fusion) to identify them. Across four modern LALMs and 12 typologically diverse languages, emotion-sensitive neurons identified independently per language show minimal overlap, and additional monolingual identification data saturates quickly without isolating more transferable units, motivating identification from pooled cross-lingual evidence. Causal interventions demonstrate that MLENs identified by CR-Fusion provide more precise and transferable affective control than monolingual neuron sets in both zero-shot and low-resource settings. Leave-one-out ablations further reveal asymmetric transfer: individual identification languages, including low-resource ones, contribute non-redundant evidence, while several low-resource languages benefit most from the resulting cross-lingual transfer. Together, our findings provide the first causal, neuron-level account of how LALMs encode emotion across languages, and establish multilingual neuron identification as an effective mechanism for understanding cross-lingual affective behavior.