Search papers, labs, and topics across Lattice.
This study investigates the role of semantic primes from the Natural Semantic Metalanguage (NSM) as explanatory tools for understanding emotions in large language models (LLMs). By analyzing four instruction-tuned LLMs, the authors demonstrate that NSM primes serve as recoverable internal elements that significantly enhance emotional control鈥攕howing three times stronger and twice more selective influence compared to traditional appraisal-based methods. The findings suggest that NSM primes provide a more effective framework for explaining emotional responses in LLMs, challenging existing paradigms of emotional representation in AI.
NSM primes outperform traditional appraisal methods in controlling emotional responses in LLMs, revealing a more effective explanatory framework for AI emotion.
Progresses have been made on understanding emotion mechanisms of large language models (LLMs). However, how to explain emotion in LLMs, or even what constitutes good explanations, are less clear. Emotion representations, components, circuits are widely recoverable, but as explanations of a model's own computation they are circular; the emotion space dimensions tend to be arbitrary and non-terminating. A pressing question to ask is whether a more primitive set of internal variables does the work: the semantic primes of the Natural Semantic Metalanguage (NSM). Across four instruction-tuned LLMs (Llama-1B, Gemma-2B, Gemma-9B, OLMo-7B), experiments show that the NSM primes are (1) recoverable internal elements; and (2) on the reference model, intervening with a prime based direction controls emotion about three times as strongly, and twice as selectively, as the best appraisal based direction; and (3) the model treats a prime based explication as interchangeable with the corresponding emotion. These evidences suggest that NSM primes seem to be better explanans for emotion in LLMs than many alternative options according to scientific explanations criteria.