June 2026Unreviewed
Curvature-Guided Module Localization for Low-Rank Detoxification of Backdoored Large Language Models
Arash Raftari, Mehrdad Mahdavi, Nathan Blackthorn, Andrew Arash Mahyari
Abstract
Backdoor attacks pose a serious threat to large language models (LLMs) by causing otherwise benign systems to produce attacker-specified malicious behavior when a hidden trigger is present. In this work, we study post hoc detoxification of backdoored LLMs in a practical setting where the defender has access to the poisoned model but does not wish to retrain the full network from scratch. We propose a mechanistically guided weight-space repair framework that first localizes modules involved in pr
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{raftari2026curvatureguided,
title = {{Curvature-Guided Module Localization for Low-Rank Detoxification of Backdoored Large Language Models}},
author = {Arash Raftari and Mehrdad Mahdavi and Nathan Blackthorn and Andrew Arash Mahyari},
year = {2026},
month = jun,
eprint = {2606.30899},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.30899}
}