Skip to content
Search
paperJune 2026Unreviewed

Curvature-Guided Module Localization for Low-Rank Detoxification of Backdoored Large Language Models

Arash Raftari, Mehrdad Mahdavi, Nathan Blackthorn, Andrew Arash Mahyari

Abstract

Backdoor attacks pose a serious threat to large language models (LLMs) by causing otherwise benign systems to produce attacker-specified malicious behavior when a hidden trigger is present. In this work, we study post hoc detoxification of backdoored LLMs in a practical setting where the defender has access to the poisoned model but does not wish to retrain the full network from scratch. We propose a mechanistically guided weight-space repair framework that first localizes modules involved in pr

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{raftari2026curvatureguided,
  title = {{Curvature-Guided Module Localization for Low-Rank Detoxification of Backdoored Large Language Models}},
  author = {Arash Raftari and Mehrdad Mahdavi and Nathan Blackthorn and Andrew Arash Mahyari},
  year = {2026},
  month = jun,
  eprint = {2606.30899},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.30899}
}