Search papers, labs, and topics across Lattice.
This study investigates the divergent capabilities of base and post-trained language models in simulating human opinions, identifying two distinct tasks: emulation and estimation. By evaluating six matched models against the Pew American Trends Panel, the authors reveal that base models excel at emulation, closely aligning with human response distributions and maintaining demographic sensitivity, while post-trained models outperform in direct estimation tasks. The findings suggest that the choice of model should be informed by the specific requirements of the task at hand, whether it involves generating text or predicting population distributions.
Base models outperform post-trained models in emulating human opinions, revealing a crucial distinction in how we should approach opinion simulation tasks.
Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, while others find persona collapse and weak demographic sensitivity. We show that much of this conflict stems from conflating two distinct tasks. We call the first task emulation, in which models generate individual responses that aggregate into a population distribution. We call the second task estimation, in which models directly predict the population distribution. Evaluating six matched base and post-trained models on the Pew American Trends Panel, we find that base models are stronger emulators: they produce response distributions closer to human ground truth and better preserve demographic structure. Post-trained models are stronger estimators, producing more accurate distributional predictions when asked directly. We propose that model selection for human simulation should be guided by whether the task requires generating text or predicting distributions.