Search papers, labs, and topics across Lattice.
This paper explores how emotional preferences, generated by higher-level goals, can autonomously regulate the priorities of competing lower-level objectives in decision-making agents. By employing a multi-objective reinforcement learning framework, the authors develop a model where an outer preference generator learns to adjust goal priorities based on the current state, while an inner controller executes preference-conditioned behaviors. The findings reveal that this emergent emotional preference mechanism not only enables contextual priority switching and graded trade-offs but also significantly outperforms traditional fixed and handcrafted preference strategies in multi-objective environments.
Emotional preferences can autonomously reshape goal priorities in agents, leading to more adaptive decision-making in dynamic environments.
A core question in decision-making for agents is whether the relative priorities of competing lower-level objectives can be determined by emotional preferences autonomously generated by higher-level goals, rather than being externally prespecified. Under changing external environments and evolving internal states, emotions play an important functional role in regulating the relative priorities of competing goals. Inspired by the goal-directed theory of emotion, this paper studies how such preference regulation can be computationally realized through reinforcement learning. We first propose a conception of emergent emotional preference: a high-level goal autonomously induces state-dependent preferences over competing lower-level objectives. This conception is built upon a framework consisting of a multi-objective reinforcement learning inner controller and an outer preference generator. The inner controller provides a repertoire of preference-conditioned goal-directed behaviors, while the outer preference generator learns a mapping from the current state to objective preferences through reinforcement learning on a high-level goal. We operationalize emotional preference as a state-dependent regulation of relative goal priorities that emerges through optimization. Furthermore, we characterize the policy space induced by preference regulation and derive an upper bound on the optimality gap in terms of the representation error of the inner behavioral repertoire. We show that the gap vanishes when the optimal policy can be represented by the available preference-conditioned policies. Experiments in self-constructed multi-objective exploration environments show that the learned preference function exhibits contextual priority switching, graded trade-offs, and temporal persistence, and outperforms the evaluated fixed-preference and handcrafted-preference strategies.