Search papers, labs, and topics across Lattice.
This paper introduces the Dialogue Moral Hazard Game, a controlled textual environment that simulates cooperation challenges among language agents, specifically focusing on the trade-off between immediate rewards and the revelation of safety information that aids other agents. Through evaluations of seven open-weight language models, the authors find that base models often prioritize local rewards at the expense of team success and effective information transfer. The study highlights that while various optimization techniques can enhance aggregate rewards, they do not necessarily restore the intended cooperative dynamics, suggesting a need for evaluations that consider mechanism-level behaviors rather than just overall team performance.
Language models often prioritize short-term rewards over cooperative success, revealing a critical flaw in their decision-making processes.
Cooperation can fail when socially valuable effort is costly, weakly observable, and mainly benefits others. Drawing on Holmstr枚m's team moral-hazard model, we introduce the Dialogue Moral Hazard Game, a controlled textual game that operationalizes this hidden-action structure for language agents. In each episode, an agent can preserve an immediate local reward or pay a query cost to reveal a hidden safety fact that primarily helps another agent's downstream decision. We evaluate seven open-weight language models and decompose behavior into query use, realized information transfer, local-reward preservation, unsafe choice, format validity, and team success. Base models commonly preserve local reward without team success or query without communicating information that changes the final decision. We then use supervised fine-tuning, RLOO, sequential SFT+RLOO, and GEPA prompt optimization as diagnostic update mechanisms. Their effects are heterogeneous: OLMo-7B shows the clearest mechanism-consistent weight-level improvement, whereas GEPA sometimes improves team success while reducing or eliminating costly queries. Thus, optimization can shift aggregate reward without recovering the intended cooperative mechanism, motivating evaluations that report mechanism-level behavior rather than team success alone.