August 2026Unreviewed
Why Are LLM Backdoor Defenses Fragmented? A Feature-Level Explanation with Sparse Autoencoders
Yizhe Zeng, Chenxu Niu, Wei Zhang, Hao Huang, Yunpeng Li, Dongxu Han, Dan Du, Cheng Hong, Hequn Xian, Yuling Liu
Abstract
Backdoor attacks pose a serious threat to large language models (LLMs), but existing defenses remain fragmented, failing to pro?vide unified defense against both dirty-label and clean-label attacks. To investigate why such fragmentation arises, we present the first systematic feature-level mechanistic analysis of LLM backdoors using sparse autoencoders (SAEs). Starting from a 2 x 2 comparison of clean and poisoned models on clean and triggered inputs, we trace backdoor-induced logit shifts to hi
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{zeng2026why,
title = {{Why Are LLM Backdoor Defenses Fragmented? A Feature-Level Explanation with Sparse Autoencoders}},
author = {Yizhe Zeng and Chenxu Niu and Wei Zhang and Hao Huang and Yunpeng Li and Dongxu Han and Dan Du and Cheng Hong and Hequn Xian and Yuling Liu},
year = {2026},
month = aug,
eprint = {2608.30403},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.30403}
}