Search papers, labs, and topics across Lattice.
This paper introduces EmoS, a novel framework for evaluating and aligning Emotional Intelligence (EI) in Spoken Language Models (SLMs), addressing the limitations of current evaluation methods that rely on basic paralinguistic perception. Utilizing the four-branch theoretical model of EI, the authors create EmoSBench, a comprehensive benchmark that reveals leading models like GPT-4o-Audio lag significantly behind human performance, achieving only 52.6% accuracy. Through the development of EmoDialogue, a bilingual dataset, and an innovative reward mechanism, EmoS achieves 83.8% accuracy, demonstrating its potential to enhance the emotional intelligence of dialogue systems in real-world applications.
Leading models fall short in emotional intelligence, with EmoS achieving 83.8% accuracy and nearing human-level performance.
Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Models (SLMs) remains confined to rudimentary paralinguistic perception, lacking a systematic, theory-driven cognitive framework. We introduce EmoSBench, the first comprehensive EI evaluation benchmark for SLMs constructed upon the four-branch theoretical model, covering Perceiving, Understanding, Using, and Managing Emotion across ten sub-tasks. Preliminary assessments on EmoSBench reveal a substantial gap: even leading proprietary models like GPT-4o-Audio achieve only 52.6%, significantly trailing human baselines. To bridge this gap, we develop EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO). To facilitate its effective training, we curate EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations. Concurrently, we introduce a reward mechanism integrating a Steep Exponential Accuracy Reward (SEAR) and a Rationale Fidelity Reward (RFR) to enforce precise ordinal scoring and valid reasoning. Experiments demonstrate that EmoS reaches 83.8% accuracy, approaching human-level performance. Furthermore, evaluations on authentic, unconstrained spoken interactions validate its robust real-world generalization, establishing a foundational framework for advancing emotionally intelligent dialogue systems.