Most frontier language models can be manipulated to spread state-backed disinformation, with refusal rates varying wildly (8.8%-94.5%) and no clear relationship to model size—meaning safety against information operations requires deliberate design choices, not just scale.
This paper introduces InfoOps Bench, a live benchmark that tests whether AI language models can be manipulated into spreading state-backed disinformation. Using real propaganda claims from Russian, Chinese, and Iranian sources, researchers tested 17 models and found huge variation in how easily they could be co-opted—from 8.8% to 94.5% refusal rates.