Search papers, labs, and topics across Lattice.
This study investigates whether large language models (LLMs) exhibit genuine decision-making structures or simply mimic rationales when faced with binary choices based on graded attributes. By employing a behavioral model to analyze the relationship between model outputs and the attributes influencing their decisions, the authors find that while LLMs demonstrate systematic behavior linked to visible attributes, their self-reported reasons only partially align with the inferred drivers of their choices. The findings suggest that LLMs operate under a framework of "superficial belief," where their decision-making is structured yet lacks full verbal articulation of the underlying priorities guiding their choices.
LLMs may seem rational, but they operate on "superficial beliefs," revealing a gap between decision-making and self-reported reasoning.
We ask whether large language models (LLMs) merely imitate rationales when choosing between two options, or whether their choices reflect a systematic underlying decision structure. Using synthetic binary decision settings in which models choose between profiles defined by graded attributes, we compare the attribute a model says mattered most with the attribute that best explains its choice under a behavioural model fit to prior decisions. The behavioural model predicts held-out choices well, showing that model behaviour is systematically related to the visible attributes rather than being random. However, direct self-reports and a separate score-based judge recover the behaviourally inferred driver only partially. The resulting picture is neither one of arbitrary behaviour nor one of fully articulated belief - outputs are structured enough to support prediction, but explicit reasons track the recovered driver only imperfectly. This qualitative pattern persists across prompt-order and sampling perturbations, alternative behavioural models, targeted occlusion analyses, and structurally varied decision settings. We interpret this as evidence for ``superficial belief'' in LLM decision-making: models behave as if guided by probabilistic local priorities over attributes, while having only limited verbal access to the attributes that drive their decisions.