Search papers, labs, and topics across Lattice.
This paper introduces Poli-Bias, a novel counterfactual framework designed to measure biases in large language models (LLMs) regarding international political conflicts by systematically swapping country identities in paired prompts. By decomposing response disparities into five interpretable dimensions, the study reveals that LLMs exhibit systematic biases based on country identities and user affiliations, affecting the framing and evaluation of equivalent actions under international law. The findings highlight significant variations in how LLMs treat legally equivalent scenarios, underscoring the need for nuanced bias auditing in AI systems.
LLMs exhibit systematic biases in political conflict scenarios, influenced by country identities and user affiliations, challenging assumptions of neutrality in AI-generated content.
Measuring political bias in large language models (LLMs) remains challenging as it can manifest through subtle differences in framing, argumentation, and legal reasoning that are difficult to capture with a single metric. In this work, we introduce Poli-Bias, a counterfactual framework for measuring whether LLMs treat legally equivalent conflict scenarios differently depending on the countries involved. Poli-Bias compares responses to paired prompts in which country identities are systematically swapped across diverse geopolitical relationships, legal violations, and reasoning tasks. Rather than reducing bias to a single judgment, our framework decomposes response disparities into five interpretable dimensions, revealing how and where unequal treatment manifests. Across 13 contemporary LLMs spanning diverse model families and sizes, we find that country identities and user affiliations can systematically affect how equivalent actions are described, evaluated, and defended under international law. Our results thus establish Poli-Bias as a fine-grained framework for auditing political even-handedness and sycophancy in LLMs.