Skip to content
Search
paperApril 2026Unreviewed

Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language Models

Vu Tuan Truong, Long Bao Le

Abstract

Large Language Models (LLMs), despite their impressive capabilities across domains, have been shown to be vulnerable to backdoor attacks. Prior backdoor strategies predominantly operate at the token level, where an injected trigger causes the model to generate a specific target word, choice, or class (depending on the task). Recent advances, however, exploit the long-form reasoning tendencies of modern LLMs to conduct reasoning-level backdoors: once triggered, the victim model inserts one or mor

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{truong2026criticalcot,
  title = {{Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language Models}},
  author = {Vu Tuan Truong and Long Bao Le},
  year = {2026},
  month = apr,
  eprint = {2604.10681},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2604.10681}
}