Search papers, labs, and topics across Lattice.
The authors investigated whether five LLMs that mirror human evaluations of a whistleblower's moral character also replicate the latent psychological motive attributions that drive those judgments, comparing model outputs against two human cohorts ($N = 125$ and $N = 742$). Although models accurately reproduced human rankings of overall moral character, they systematically idealized agents by overestimating helpfulness and underestimating self-interest and hostility, while failing to link competitive motives to moral judgments. This reveals that surface-level behavioral parity between humans and models can mask fundamentally misaligned underlying reasoning and contextual insensitivity.
LLMs can perfectly mirror human moral verdicts while harboring an entirely un-human theory of mind, systematically whitewashing agents as far more altruistic and less hostile than humans perceive them to be.
Large language models (LLMs) are used to simulate human participants in psychological research. We asked whether LLMs that reproduce human evaluations of a whistleblower's moral character also reproduce the motive attributions that accompany them. Five LLMs and two human samples (N = 125 and N = 742) evaluated a physician who either remained silent about fraudulent billing or reported it to a hospital, regulator, or newspaper. Models reproduced the human ranking of the physician's moral character but portrayed whistleblowers as more helpful, less self-interested, and less hostile. In four of five models, competitive motives were less strongly associated with moral-character judgments. Model ratings changed little when prompts reproduced the narratives and demographic profiles of both human samples, although this comparison cannot isolate a perspective effect. Thus, agreement in average ratings can conceal differences in attributed motives, relationships among judgments, and sensitivity to context. Validating LLMs as simulated participants therefore requires testing psychologically informative response patterns, not average agreement alone.