Search papers, labs, and topics across Lattice.
This study explores the calibration of guilt signals from human neural data to enhance reward shaping in cooperative multi-agent reinforcement learning. By analyzing fMRI data from 40 participants, the authors derive a guilt weight that is then integrated into a Social Lottery environment, where agents are trained using Proximal Policy Optimization under various reward shaping conditions. The results show that agents using the neurally calibrated guilt signal closely match human decision-making patterns, outperforming traditional reward shaping methods significantly in terms of alignment with human behavior.
Calibrating guilt from human neural data enables AI agents to mimic human prosocial behavior more accurately than conventional reward shaping methods.
Cooperative multi-agent reinforcement learning often adds social terms to individual rewards, yet the scale of those terms is usually chosen by hand. We ask whether a guilt signal can instead be calibrated from human neural and behavioural data and transferred to artificial agents. Using the public SoDec responsibility fMRI dataset (40 participants), we fit a subject-fixed-effects regression of momentary-happiness changes on outcome-type counts and recover a guilt weight as the Partner-negative minus Social-negative contrast ($\hat{w}=1.118$, Cohen's $d=0.214$). We embed this weight in a two-agent Social Lottery environment and train independent Proximal Policy Optimization actor-critics under four shaping regimes: neurally calibrated, uniform constant, zero (selfish), and a unit-coefficient oracle. Across 1{,}000 evaluation episodes per condition, the calibrated agents track the human Social safe-choice rate most closely ($0.459$ vs.\ human $0.484$; $\mathrm{KL}=0.0012$), while the other three conditions deviate by one to three orders of magnitude in KL. Human neurobehavioural priors can therefore act as quantitative constraints on prosocial reward shaping.