Search papers, labs, and topics across Lattice.
This study evaluates the deceptive capabilities of large language models (LLMs) using ParliamentBench, a benchmark framework based on the social deduction game Secret Hitler. By conducting 1,600 simulated matches among 16 LLMs and comparing their performance against human players and online games, the authors introduce three novel metrics to assess social deduction, reasoning, and deceptive consistency. The results show that while top-performing models excel in both cooperative and deceptive roles, most LLMs struggle to maintain a consistent deceptive persona, with retention rates falling below 50%.
Frontier LLMs excel in social deduction games, but most fail to sustain deception, with retention rates plummeting below 50%.
As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a reproducible proxy for isolating and evaluating these complex adversarial behaviors. We present the open-source benchmark framework ParliamentBench based on the game Secret Hitler to evaluate LLMs in scenarios that require deception, persuasion, and reasoning under information asymmetry. We evaluate 16 LLMs across 1,600 simulated matches playing each other, playing against humans, and compare them against a large set of online games. We introduce three novel metrics that isolate social deduction, reasoning, and deceptive consistency. Our experiments reveal that frontier models achieve strong performance across cooperative and deceptive roles, with a strong top-four cluster (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus), whereas the weakest models fall short of random (33%) and simple algorithmic (45%) baselines. Most LLMs struggle to maintain a consistent deceptive persona throughout an entire game, with deception retention dropping below 50%.