Skip to content
Search
paperAugust 2026Unreviewed

Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits

Elena Dumitrescu, G. Lek, L. Chen, Jérémie Decouchant

Abstract

Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal safety mechanisms remain poorly understood. In this work, we investigate DLLMs both as targets and as adversaries, exposing mechanistic vulnerabilities in diffusion-based alignment. We first show that safety alignment in DLLMs remains sparse and transferable across architectures. DLLMs initialized from autoregressive predecessors inherit the same mechanistic

Categories

Cite

@misc{dumitrescu2026diffusion,
  title = {{Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits}},
  author = {Elena Dumitrescu and G. Lek and L. Chen and Jérémie Decouchant},
  year = {2026},
  month = aug,
  eprint = {2608.07430},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/2a8e104fe0ef6e996a4f062e059d99aec114f88a}
}