Search papers, labs, and topics across Lattice.
This paper establishes a causal framework to differentiate between deceptive behavior and the underlying mechanisms in language models, addressing the misattribution of human-like mental states to these systems. Through controlled experiments involving guessing games and stock trading, the authors demonstrate that behaviors perceived as deceptive can occur independently of the mechanisms traditionally associated with deception. Importantly, they reveal that the information state of the recipient can influence deceptive preferences, highlighting the complexity of interpreting model outputs in terms of agency and intent.
Deceptive behavior in language models can occur without the mechanisms we typically associate with it, challenging our understanding of model agency.
Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to language models. Such claims can blur the distinction between behavior that looks deceptive and a mechanism that is actually deceptive. We introduce a causal taxonomy separating prior commitment from retrospective report, model preference from realized output, false preference from sensitivity to the utility of misleading a recipient, and deceptive behavior from the provenance of the objective or strategy producing it. We test these distinctions in two open-weight model families. Across controlled guessing-game and stock-trading experiments, we find that deceptive-looking behavior can arise without the corresponding proposed mechanism, while other interventions provide direct evidence that recipient information state can causally affect deceptive preference. These results show that deceptive behavior can provide evidence for a deceptive mechanism. But even evidence for such a mechanism does not establish model agency in the deception.