Search papers, labs, and topics across Lattice.
This study evaluates the ability of six large language models (LLMs) to simulate human belief updates in controlled environments, specifically by comparing their outputs to data from 391 UK participants who adjusted their stances on discussion topics after reading Reddit comments. The results reveal that while some models, like Qwen3-32B and GPT-5-Mini, can replicate human post-stance distributions when provided with actual initial stances, they consistently fail to generate accurate initial stances or produce reliable belief updates from self-generated stances. Systematic biases were identified across all models, including an overrepresentation of neutral positions and inadequate ranking of comment convincingness, highlighting the limitations of LLMs in accurately modeling human belief dynamics without realistic starting conditions.
LLMs can mimic human belief updates鈥攂ut only if they start from the right initial conditions, revealing critical limitations in their use as proxies for human participants.
LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been tested directly. We test whether six LLMs can simulate individual human belief updates, comparing LLM outputs 1-to-1 against ground truth data from 391 UK participants on Prolific, who updated their stances on three discussion topics after reading Reddit comments. Each participant was simulated by an LLM conditioned on a persona derived from their demographic and personality trait data. We find that some LLMs (Qwen3-32B and GPT-5-Mini) can match the human post-stance distribution, but only when given participants'actual initial stances. All six models fail to simulate initial stances themselves and to produce faithful belief updates from self-generated stances. Three systematic biases emerge across all models: overrepresentation of neutral positions, more frequent but smaller belief shifts than humans, and a failure to rank comments by convincingness. Demographic and personality trait personas had no consistent effect on fidelity. LLM simulations of human belief dynamics are only reliable when grounded in realistic starting conditions, that current multi-round social media simulations rarely provide.