April 2026Unreviewed
MirageBackdoor: A Stealthy Attack that Induces Think-Well-Answer-Wrong Reasoning
Yizhe Zeng, Wei Zhang, Yunpeng Li, Juxin Xiao, Xiao Wang, Yuling Liu
Abstract
While Chain-of-Thought (CoT) prompting has become a standard paradigm for eliciting complex reasoning capabilities in Large Language Models, it inadvertently exposes a new attack surface for backdoor attacks. Existing CoT backdoor attacks typically manipulate the intermediate reasoning steps to steer the model toward incorrect answers. However, these corrupted reasoning traces are readily detected by prevalent process-monitoring defenses. To address this limitation, we introduce MirageBackdoor(M
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{zeng2026miragebackdoor,
title = {{MirageBackdoor: A Stealthy Attack that Induces Think-Well-Answer-Wrong Reasoning}},
author = {Yizhe Zeng and Wei Zhang and Yunpeng Li and Juxin Xiao and Xiao Wang and Yuling Liu},
year = {2026},
month = apr,
eprint = {2604.06840},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2604.06840}
}