September 2026Unreviewed
AKRASIA: Stealthy Backdoor Attack on Reasoning-based Code LLMs
Chua Jin Chou, Sarang Nambiar, Murali Srinivasan, Ezekiel Soremekun
Abstract
We present AKRASIA, a stealthy, inference-time backdoor attack against reasoning-based Code LLMs. AKRASIA aims to achieve a backdoor target (e.g., malicious code execution) in reasoning LLMs while evading automated defenses and human inspection. To achieve this, AKRASIA probes the victim LLM to construct a code-level backdoor trigger. It then employs in-context learning for backdoor learning, and model unfaithfulness to conceal the backdoor trigger, and generate plausible reasoning. We evaluate
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{chou2026akrasia,
title = {{AKRASIA: Stealthy Backdoor Attack on Reasoning-based Code LLMs}},
author = {Chua Jin Chou and Sarang Nambiar and Murali Srinivasan and Ezekiel Soremekun},
year = {2026},
month = sep,
eprint = {2609.01023},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.01023}
}