Search papers, labs, and topics across Lattice.
This paper introduces CHIARO, a novel benchmark for contrastive emotion inference that captures the complexity of simultaneous opposing emotions arising from a single event, grounded in appraisal theory. The dataset comprises 1,000 human-annotated sentences reflecting a ten-class taxonomy, facilitating the evaluation of seven state-of-the-art LLMs and four emotion classifiers. Results show that while the best-performing LLM achieves a macro-F1 score of 67.3, it still falls short of human agreement, highlighting the challenges in accurately modeling nuanced emotional responses and demonstrating that CHIARO can enhance existing emotion classifiers when used as a training signal.
The strongest LLM only achieves a macro-F1 score of 67.3 on a benchmark designed to capture the complexity of opposing emotions, revealing significant gaps in current emotion recognition systems.
Emotion recognition benchmarks often predict one emotion per text, missing many real-world scenarios where two people arrive at opposing emotions from a single shared event. For example, a child kicks the seat in front of her in excitement while the passenger ahead grows angry. We introduce CHIARO, a 1,000 human-annotated sentence benchmark for contrastive emotion inference grounded in appraisal theory. Each scene describes one causal trigger eliciting a positive emotion in one person and a negative emotion in the other, drawn from a ten-class taxonomy. We benchmark seven frontier LLMs and four off-the-shelf emotion classifiers. The strongest LLM reaches 67.3 macro-F1, well below human agreement, while existing emotion classifiers score near chance. Beyond evaluation, CHIARO also serves as a training signal. When combined with an existing emotion corpus, the resulting downstream classifier improves on CHIARO itself and on six of ten external emotion benchmarks, which positions our dataset as a complementary signal for emotion recognition.