May 2026Unreviewed
Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks
John T. Halloran, Noopur S. Bhatt
Abstract
Large language models (LLMs) are highly susceptible to backdoor attacks (BAs), wherein training samples are poisoned using trigger-based harmful content. Furthermore, existing defenses have proven ineffective when extensively tested across BA patterns. To better combat BAs, we explore the use of LLM rewriting as a proactive defense against data poisoning. First, we theoretically show that when LLM rewriting utilizes open-book benign samples--termed open-book benign rewriting (OBBR)--the probabil
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{halloran2026be,
title = {{Be Kind, Rewrite: Benign Projections via Rewriting Defend Against LLM Data Poisoning Attacks}},
author = {John T. Halloran and Noopur S. Bhatt},
year = {2026},
month = may,
eprint = {2605.19147},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.19147}
}