Skip to content
Search
paperAugust 2026Unreviewed

Why Are LLM Backdoor Defenses Fragmented? A Feature-Level Explanation with Sparse Autoencoders

Yizhe Zeng, Chenxu Niu, Wei Zhang, Hao Huang, Yunpeng Li, Dongxu Han, Dan Du, Cheng Hong, Hequn Xian, Yuling Liu

Abstract

Backdoor attacks pose a serious threat to large language models (LLMs), but existing defenses remain fragmented, failing to pro?vide unified defense against both dirty-label and clean-label attacks. To investigate why such fragmentation arises, we present the first systematic feature-level mechanistic analysis of LLM backdoors using sparse autoencoders (SAEs). Starting from a 2 x 2 comparison of clean and poisoned models on clean and triggered inputs, we trace backdoor-induced logit shifts to hi

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{zeng2026why,
  title = {{Why Are LLM Backdoor Defenses Fragmented? A Feature-Level Explanation with Sparse Autoencoders}},
  author = {Yizhe Zeng and Chenxu Niu and Wei Zhang and Hao Huang and Yunpeng Li and Dongxu Han and Dan Du and Cheng Hong and Hequn Xian and Yuling Liu},
  year = {2026},
  month = aug,
  eprint = {2608.30403},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.30403}
}