Skip to content
Search
paperMay 2026Unreviewed

BadDLM: Backdooring Diffusion Language Models with Diverse Targets

Shengfang Zhai, Xiaoyang Ji, Yuling Shi, Haoran Gao, Fanyu Meng, Yan Zeng, Yuejian Fang, Yinpeng Dong, Jiaheng Zhang

Abstract

Diffusion language models (DLMs) have recently emerged as an alternative modeling paradigm to autoregressive (AR) language models, enabling parallel generation and bidirectional context modeling. Yet their security implications, particularly their vulnerability to backdoor attacks, remain underexplored. We propose BadDLM, a unified framework for studying backdoor attacks against DLMs with diverse targets. We introduce a trigger-aware training objective that emphasizes target-relevant positions i

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{zhai2026baddlm,
  title = {{BadDLM: Backdooring Diffusion Language Models with Diverse Targets}},
  author = {Shengfang Zhai and Xiaoyang Ji and Yuling Shi and Haoran Gao and Fanyu Meng and Yan Zeng and Yuejian Fang and Yinpeng Dong and Jiaheng Zhang},
  year = {2026},
  month = may,
  eprint = {2605.09397},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.09397}
}