August 2026Unreviewed
Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents
Mehrdad Ghassabi
Abstract
Retrieval-Augmented Generation (RAG) has significantly enhanced the performance of large language models (LLMs), yet these systems remain vulnerable to knowledge-poisoning attacks, in which misinformation in retrieved documents can influence the model's final outputs. Notably, an LLM may correctly detect that a document contains incorrect information while nevertheless being influenced by it. Prior work has addressed this vulnerability through the Cordon Principle, which prevents models responsi
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{ghassabi2026safer,
title = {{Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents}},
author = {Mehrdad Ghassabi},
year = {2026},
month = aug,
eprint = {2608.17153},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/9ad90be31be191016980abf888e48404a08457d7}
}