Skip to content
Search
paperMay 2026Unreviewed

Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks

John T. Halloran, Noopur S. Bhatt

Abstract

Large language models (LLMs) are highly susceptible to backdoor attacks (BAs), wherein training samples are poisoned using trigger-based harmful content. Furthermore, existing defenses have proven ineffective when extensively tested across BA patterns. To better combat BAs, we explore the use of LLM rewriting as a proactive defense against data poisoning. First, we theoretically show that when LLM rewriting utilizes open-book benign samples--termed open-book benign rewriting (OBBR)--the probabil

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{halloran2026be,
  title = {{Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks}},
  author = {John T. Halloran and Noopur S. Bhatt},
  year = {2026},
  month = may,
  eprint = {2605.19147},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.19147}
}