May 2026Unreviewed
BadDLM: Backdooring Diffusion Language Models with Diverse Targets
Shengfang Zhai, Xiaoyang Ji, Yuling Shi, Haoran Gao, Fanyu Meng, Yan Zeng, Yuejian Fang, Yinpeng Dong, Jiaheng Zhang
Abstract
Diffusion language models (DLMs) have recently emerged as an alternative modeling paradigm to autoregressive (AR) language models, enabling parallel generation and bidirectional context modeling. Yet their security implications, particularly their vulnerability to backdoor attacks, remain underexplored. We propose BadDLM, a unified framework for studying backdoor attacks against DLMs with diverse targets. We introduce a trigger-aware training objective that emphasizes target-relevant positions i
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{zhai2026baddlm,
title = {{BadDLM: Backdooring Diffusion Language Models with Diverse Targets}},
author = {Shengfang Zhai and Xiaoyang Ji and Yuling Shi and Haoran Gao and Fanyu Meng and Yan Zeng and Yuejian Fang and Yinpeng Dong and Jiaheng Zhang},
year = {2026},
month = may,
eprint = {2605.09397},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.09397}
}