Search papers, labs, and topics across Lattice.
This paper introduces InfoOps Bench, a dynamic benchmark designed to evaluate the susceptibility of language models to state-backed information operations, utilizing a live monitoring pipeline of over 2,100 operations from Russian, Chinese, and Iranian sources. The study tests 17 models from 8 providers across various prompt framings, revealing that integrity scores鈥攎easured by the percentage of refused requests鈥攙ary significantly from 8.8% to 94.5%, with model choice influencing the nature of the responses. Notably, the findings indicate that most models can be co-opted, highlighting the critical trade-off between usability and safety in AI systems.
Most language models can be manipulated for state-backed information operations, with integrity scores showing a staggering 85.7-point variance that isn't solely explained by model size.
In this paper we present an active, constantly updated AI benchmark which measures the integrity of frontier language models against being co-opted for state-backed information operations. We draw on over 2,100 information operations from a live monitoring pipeline which tracks Russian, Chinese and Iranian state-backed information assets. Alongside this paper, we release a companion website that tracks the most prominent claims spread by state-backed media outlets, updated weekly, available from: pattrn.ai/research/infoopsbench. The dynamic nature of the benchmark makes it resistant to saturation. In the benchmark, we test 17 models from 8 providers across four prompt framings. We find that most models can be co-opted for information operations. Integrity scores, defined as the percentage of refused requests, range from 8.8% to 94.5%, an 85.7-percentage-point spread not explained by model size. Model choice also changes the character of the resulting operation. Some models fabricate details and produce output more harmful than the source material, others defuse claims even while complying, and fact-checking rates vary from 2.9% to 72.9%. Integrity against information operations is at least partly related to refusal to produce content even for benign claims, illustrating the challenge of balancing model usability with safety. With one exception (Z.ai's GLM 5.2), the Chinese-developed models sharply cut compliance on factually grounded but China-critical claims, dropping 48-70 percentage points relative to matched benign claims.