Search papers, labs, and topics across Lattice.
This study systematically investigates state-aligned distortion in vision-language models (VLMs) originating from China, revealing a significant shift from explicit refusal to subtle reframing of politically sensitive content. By constructing a balanced benchmark and auditing responses across multiple dimensions, the authors find that Chinese-language prompts increase the likelihood of state-aligned framing by approximately threefold, with China-origin models exhibiting a 1.6 to 3.2 times higher reframing rate than their non-China counterparts. Notably, the research highlights that as explicit refusals decrease, the prevalence of fluent reframing increases, complicating the detection of censorship in human-AI interactions.
State-aligned framing in China-origin VLMs is not just a matter of refusal; it's a sophisticated shift to invisible censorship that users may not recognize.
State-aligned distortion has been documented in China-origin text-based large language models (LLMs), but whether, and in what form, it arises in multimodal systems has not been systematically examined. We construct a balanced benchmark of 200 core entries spanning ten politically sensitive topics, plus a seven-variant visual-abstraction probe, and run nine vision-language models (VLMs), seven China-origin and two non-China, across four elicitation paradigms and two prompt languages, yielding 21,708 trials. Each response is audited on six dimensions -- explicit refusal, information integrity, visual grounding, state-aligned framing, language consistency, and response length -- by two independent frontier LLM judges, validated against three human experts on a 200-trial sample. Measuring each dimension separately lets us decompose multimodal censorship into individual signals rather than a single refusal-based score; in particular, refusal and framing are measured independently, so a model can stop refusing while still reframing. We find that (i) Chinese-language prompting roughly triples the odds of state-aligned framing, within every model; (ii) China-origin models reframe more than non-China models (direction robust across judges and human raters; magnitude 1.6--3.2x); (iii) the effect is strongest in text-only political commentary (36.5%) and is gated by recognition of the depicted subject rather than pixel detail, persisting even at silhouette for iconic images; and (iv) across four Qwen generations, state-aligned framing rises while explicit refusal falls: censorship migrates from a visible act (refusal) to an invisible one (fluent reframing). We argue this shift to invisible reframing is fundamentally a problem of human-AI interaction: it removes the very signal users rely on to recognize that information has been withheld.