Safety alignment in diffusion-based language models is mechanistically similar to autoregressive models and can be bypassed by identifying and manipulating specific safety neurons, raising urgent questions about deploying DLLMs in production systems.
This paper reveals critical safety vulnerabilities in Diffusion Large Language Models (DLLMs)—models that generate text through iterative denoising rather than predicting one token at a time. The researchers show that safety mechanisms in DLLMs are sparse and can be transferred between models, and introduce a jailbreak method that exploits these vulnerabilities with minimal computational cost.