Search papers, labs, and topics across Lattice.
This study investigates how demographic group identity is represented within the Mistral-7B language model, revealing that traditional last-token read-outs significantly underestimate the model's capabilities. Through representational similarity analysis and causal interventions, the authors find that attention-head read-outs provide a more accurate reflection of inter-group opinion structures, with a single attention head demonstrating consistent fidelity across various demographic types. Notably, the research uncovers that the model's causal use of demographic information does not align with its fidelity, highlighting the complexity of how LLMs simulate population dynamics.
A single attention head in the Mistral-7B model captures demographic identity with surprising fidelity, yet its causal use reveals a disconnect that complicates LLMs' ability to simulate real-world populations.
Large language models are widely used to simulate survey respondents, yet their answers are homogeneous and unfaithful to real inter-group differences. We ask where demographic group identity lives inside an LLM, how faithfully its geometry mirrors real inter-group opinion structure, and whether it uses what it encodes. Using representational similarity analysis against Pew ground truth over 169 demographic cells, we score 1,089 read-out locations in Mistral-7B and intervene causally across six attribute types. Four results. (1) The standard last-token residual read-out understates the model: attention-head read-outs dominate it in five of six types, with selection-corrected fidelity up to rho=0.63 -- roughly 70% of the measurement-reliability ceiling -- surviving a lexical-similarity control. (2) A single head (L11 H16) is significantly faithful in all six types as a fixed location, while race-based types stay weak and prompt-fragile. Both phenomena replicate -- the analogous head significant in five of six types, weakest on the same race type -- across three checkpoints of a second model family, where ten billion training tokens barely move the map. (3) Causal use does not follow fidelity: the clearest causal pathway sits in one of the least faithful types (p=0.002, cluster-robust, fixed depth), the most faithful type shows no correction-surviving single-layer effect, and replacing the entire identity moves predictions by under 2% of their error. (4) A 128-dimensional probe of the single head lands 21-31% closer to survey truth than the model's own answers -- yet recovers almost none of the per-question group ordering, no better than the answers themselves. Readable, faithfully arranged, and causally used are three dissociable properties of the same model; treating them as one claim is what keeps the"can LLMs simulate populations"debate unresolved.