Skip to content
Search
paperAugust 2026Unreviewed

Backdoor Decontamination Dynamics in LLM Agents

Gabriel Huang, Abhay Puri, L'eo Boisvert, Alexandre Drouin, Perouz Taslakian, Spandana Gella, Christopher Pal

Abstract

Open-weight LLM agents are vulnerable to backdoors installed during fine-tuning, which may be undetectable if the trigger conditions are never met during testing. Assuming defenders do not know the existing trigger, they cannot unlearn it directly. One decontamination strategy is to install a known backdoor (defensive poisoning) then to unlearn it, hoping that the original unknown backdoor is removed as a side effect. However, this procedure has uncertain outcomes: the original backdoor may pers

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{huang2026backdoor,
  title = {{Backdoor Decontamination Dynamics in LLM Agents}},
  author = {Gabriel Huang and Abhay Puri and L'eo Boisvert and Alexandre Drouin and Perouz Taslakian and Spandana Gella and Christopher Pal},
  year = {2026},
  month = aug,
  eprint = {2608.11295},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/d5229f30a2f7494f971c21683cde87e9e29b1cbf}
}