Skip to content
Search
paperAugust 2026Unreviewed

Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents

Mehrdad Ghassabi

Abstract

Retrieval-Augmented Generation (RAG) has significantly enhanced the performance of large language models (LLMs), yet these systems remain vulnerable to knowledge-poisoning attacks, in which misinformation in retrieved documents can influence the model's final outputs. Notably, an LLM may correctly detect that a document contains incorrect information while nevertheless being influenced by it. Prior work has addressed this vulnerability through the Cordon Principle, which prevents models responsi

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{ghassabi2026safer,
  title = {{Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents}},
  author = {Mehrdad Ghassabi},
  year = {2026},
  month = aug,
  eprint = {2608.17153},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/9ad90be31be191016980abf888e48404a08457d7}
}