Skip to content
Search
paperApril 2026Unreviewed

MirageBackdoor: A Stealthy Attack that Induces Think-Well-Answer-Wrong Reasoning

Yizhe Zeng, Wei Zhang, Yunpeng Li, Juxin Xiao, Xiao Wang, Yuling Liu

Abstract

While Chain-of-Thought (CoT) prompting has become a standard paradigm for eliciting complex reasoning capabilities in Large Language Models, it inadvertently exposes a new attack surface for backdoor attacks. Existing CoT backdoor attacks typically manipulate the intermediate reasoning steps to steer the model toward incorrect answers. However, these corrupted reasoning traces are readily detected by prevalent process-monitoring defenses. To address this limitation, we introduce MirageBackdoor(M

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{zeng2026miragebackdoor,
  title = {{MirageBackdoor: A Stealthy Attack that Induces Think-Well-Answer-Wrong Reasoning}},
  author = {Yizhe Zeng and Wei Zhang and Yunpeng Li and Juxin Xiao and Xiao Wang and Yuling Liu},
  year = {2026},
  month = apr,
  eprint = {2604.06840},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2604.06840}
}