← Back to search
paper llmsec-2026-00150

MirageBackdoor: A Stealthy Attack that Induces Think-Well-Answer-Wrong Reasoning

Yizhe Zeng, Wei Zhang, Yunpeng Li, Juxin Xiao, Xiao Wang, Yuling Liu

2026-04

Abstract

While Chain-of-Thought (CoT) prompting has become a standard paradigm for eliciting complex reasoning capabilities in Large Language Models, it inadvertently exposes a new attack surface for backdoor attacks. Existing CoT backdoor attacks typically manipulate the intermediate reasoning steps to steer the model toward incorrect answers. However, these corrupted reasoning traces are readily detected by prevalent process-monitoring defenses. To address this limitation, we introduce MirageBackdoor(M

Cite This Resource

@article{llmsec202600150,
  title = {MirageBackdoor: A Stealthy Attack that Induces Think-Well-Answer-Wrong Reasoning},
  author = {Yizhe Zeng and Wei Zhang and Yunpeng Li and Juxin Xiao and Xiao Wang and Yuling Liu},
  year = {2026},
  url = {https://arxiv.org/abs/2604.06840},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2604.06840