← Back to search
paper llmsec-2026-00150
MirageBackdoor: A Stealthy Attack that Induces Think-Well-Answer-Wrong Reasoning
Yizhe Zeng, Wei Zhang, Yunpeng Li, Juxin Xiao, Xiao Wang, Yuling Liu
2026-04
Abstract
While Chain-of-Thought (CoT) prompting has become a standard paradigm for eliciting complex reasoning capabilities in Large Language Models, it inadvertently exposes a new attack surface for backdoor attacks. Existing CoT backdoor attacks typically manipulate the intermediate reasoning steps to steer the model toward incorrect answers. However, these corrupted reasoning traces are readily detected by prevalent process-monitoring defenses. To address this limitation, we introduce MirageBackdoor(M
Categories
Cite This Resource
@article{llmsec202600150,
title = {MirageBackdoor: A Stealthy Attack that Induces Think-Well-Answer-Wrong Reasoning},
author = {Yizhe Zeng and Wei Zhang and Yunpeng Li and Juxin Xiao and Xiao Wang and Yuling Liu},
year = {2026},
url = {https://arxiv.org/abs/2604.06840},
} Metadata
- Added
- 2026-05-17
- Added by
- automation
- Source
- arxiv
- arxiv_id
- 2604.06840