Skip to content
Search
paperAugust 2026Unreviewed

Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems

Balkrishna Giri, M. Hasan, Jussi Rasku, Muhammad Waseem, Pekka Abrahamsson

Abstract

Retrieval-Augmented Generation (RAG) grounds Large Language Model (LLM) outputs in external knowledge, but RAG systems usually trust whatever they retrieve, creating a Security-Reliability Gap: high semantic relevance does not guarantee factual truth. Adversaries exploit this through knowledge poisoning, inserting malicious documents to cause targeted misinformation. We propose an Evaluation Agent, middleware that combines Natural Language Inference (NLI) factual verification, a five-signal pois

Categories

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Suggested from the entry's categories.

Cite

@misc{giri2026trustworthy,
  title = {{Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems}},
  author = {Balkrishna Giri and M. Hasan and Jussi Rasku and Muhammad Waseem and Pekka Abrahamsson},
  year = {2026},
  month = aug,
  eprint = {2608.21095},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/21c1fac19d0c3a8a1ebbddabcdc5b5109b566700}
}