Search papers, labs, and topics across Lattice.
This study conducts a controlled empirical comparison between surveys generated by GPT and established human-designed survey instruments across three social domains: climate change, immigration, and diversity, equity, and inclusion (DEI). By utilizing a fixed prompting framework to produce GPT-generated surveys and comparing them with validated human baselines, the researchers collected responses from U.S. participants to analyze differences in response distributions and clustering behavior. The key finding indicates that while GPT-generated surveys align with dominant attitudinal divisions found in human-designed surveys, they differ in the resolution of belief structures, suggesting their potential for exploratory analyses in social research.
GPT-generated surveys can effectively capture social attitudes, matching human designs in key areas while revealing unique insights into belief structures.
Understanding human beliefs and social attitudes often relies on carefully designed survey instruments. Recent work has suggested that large language models (LLMs) could automate parts of this process by generating surveys at scale, raising questions about the comparability of such instruments to literature-grounded, human-designed surveys. We present a controlled empirical comparison between GPT-generated surveys and established survey baselines across three social domains: climate change, immigration, and diversity, equity, and inclusion (DEI). GPT-generated surveys were produced using a fixed prompting framework enforcing a 3x3 structure over beliefs, perceptions, and behaviors, while human baselines were assembled from validated instruments to match survey length and construct coverage. We collected responses from U.S.-based participants, who completed both survey types, allowing direct within-subject comparison. We analyze differences in response distributions, clustering behavior, and alignment with self-identified stances. Our results show that GPT-generated surveys capture the same dominant attitudinal divisions as human-designed instruments, while exhibiting differences in the resolution of belief structure and group separation. These findings suggest that LLM-generated surveys are suited for exploratory and large-scale analyses, and can be used to complement expert-designed instruments.