Search papers, labs, and topics across Lattice.
This study investigates the internal representations of logical validity in five open-weight transformer models by analyzing their performance on matched premise-claim pairs across various inference families and semantic domains. Despite models demonstrating near-chance performance in behavioral tasks, the research reveals that logical validity can be almost perfectly decoded from hidden states, even in cases where the models produce incorrect outputs. The findings indicate a dissociation between the representation of validity and its expression in model behavior, highlighting that strong decodability does not equate to reliable output performance.
Logical validity can be accurately decoded from hidden states of language models, even when their outputs are incorrect, revealing a critical gap between representation and expression.
Large language models can look capable of logical reasoning, but correct or incorrect answers alone tell us little about what the model represents internally. We study logical verification in five open-weight transformer models using matched valid--invalid premise--claim pairs that vary across inference families, semantic domains, templates, and difficulty levels. Despite near-chance behavioral performance, logical validity is often almost perfectly decodable from hidden states and remains strongly decodable under held-out templates, domains, and inference families. Validity also remains highly decodable on behaviorally incorrect examples in the conditions where correctness-conditioned evaluation is well defined. At the same time, exhaustive leave-one-out tests reveal clear limits to this generalization, and interventions along probe-derived validity directions have only weak, nonspecific effects compared with random controls. Our results suggest that representing validity, expressing it in behavior, and using it causally are distinct. Validity related information can be strongly decodable from a model's hidden states without being reliably expressed in its output.