Search papers, labs, and topics across Lattice.
The authors investigate how state political threat models manifest in LLM guardrails by auditing ten Chinese and Western models across three languages against political dissent and mobilization prompts. They find that Chinese systems selectively suppress collective-action coordination (even pro-government mobilization) directed at China rather than pure dissent, yet this censorship is remarkably brittle to adversarial paraphrasing. Consequently, superficial refusal rates create a false impression of strict information control, with Western frontier models actually exhibiting far greater robustness against jailbreak evasion.
Chinese model guardrails target coordination rather than ideology鈥攄eclining even to organize pro-government rallies鈥攜et their sky-high refusal rates collapse under basic adversarial paraphrasing.
As large language models become the front door to political information, what they refuse to discuss becomes a new instrument of information control. We argue that a model's guardrail encodes not a universal notion of harm but the political threat model of the state that governs its developer, and we derive the expected structure of that control from the comparative study of how authoritarian regimes censor. Across ten models and three languages, Chinese guardrails carry its signatures: they answer to the developer's own regime, refusing identical collective-action prompts far more when a prompt names China than a foreign state; within politics they target the capacity to coordinate rather than dissent, declining even to help organize pro-government mobilization; and their strictness is porous, collapsing under adversarial paraphrase, so that the models most resistant to attack are Western frontier systems, not the strictest refusers. Machine censorship thus reproduces the friction-based logic of prior-era information control while, lacking a censor's case-by-case judgment, proving blunter than the bureaucracy it resembles---so that audits which measure refusal directly overstate how controlled a model actually is.