← Back to search
paper llmsec-2026-00131

CleanBase: Detecting Malicious Documents in RAG Knowledge Databases

Weifei Jin, Xilong Wang, Wei Zou, Jinyuan Jia, Neil Gong

2026-05

Abstract

Retrieval-augmented generation (RAG) is vulnerable to prompt injection attacks, in which an adversary inserts malicious documents containing carefully crafted injected prompts into the knowledge database. When a user issues a question targeted by the attack, the RAG system may retrieve these malicious documents, whose injected prompts mislead it into generating attacker-specified answers, thereby compromising the integrity of the RAG system. In this work, we propose CleanBase, a method to detect

Cite This Resource

@article{llmsec202600131,
  title = {CleanBase: Detecting Malicious Documents in RAG Knowledge Databases},
  author = {Weifei Jin and Xilong Wang and Wei Zou and Jinyuan Jia and Neil Gong},
  year = {2026},
  url = {https://arxiv.org/abs/2605.00460},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2605.00460