August 2026Unreviewed
Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems
Balkrishna Giri, M. Hasan, Jussi Rasku, Muhammad Waseem, Pekka Abrahamsson
Abstract
Retrieval-Augmented Generation (RAG) grounds Large Language Model (LLM) outputs in external knowledge, but RAG systems usually trust whatever they retrieve, creating a Security-Reliability Gap: high semantic relevance does not guarantee factual truth. Adversaries exploit this through knowledge poisoning, inserting malicious documents to cause targeted misinformation. We propose an Evaluation Agent, middleware that combines Natural Language Inference (NLI) factual verification, a five-signal pois
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Suggested from the entry's categories.
Cite
@misc{giri2026trustworthy,
title = {{Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems}},
author = {Balkrishna Giri and M. Hasan and Jussi Rasku and Muhammad Waseem and Pekka Abrahamsson},
year = {2026},
month = aug,
eprint = {2608.21095},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/21c1fac19d0c3a8a1ebbddabcdc5b5109b566700}
}