August 2026Unreviewed
Backdoor Decontamination Dynamics in LLM Agents
Gabriel Huang, Abhay Puri, L'eo Boisvert, Alexandre Drouin, Perouz Taslakian, Spandana Gella, Christopher Pal
Abstract
Open-weight LLM agents are vulnerable to backdoors installed during fine-tuning, which may be undetectable if the trigger conditions are never met during testing. Assuming defenders do not know the existing trigger, they cannot unlearn it directly. One decontamination strategy is to install a known backdoor (defensive poisoning) then to unlearn it, hoping that the original unknown backdoor is removed as a side effect. However, this procedure has uncertain outcomes: the original backdoor may pers
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{huang2026backdoor,
title = {{Backdoor Decontamination Dynamics in LLM Agents}},
author = {Gabriel Huang and Abhay Puri and L'eo Boisvert and Alexandre Drouin and Perouz Taslakian and Spandana Gella and Christopher Pal},
year = {2026},
month = aug,
eprint = {2608.11295},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/d5229f30a2f7494f971c21683cde87e9e29b1cbf}
}