August 2026Unreviewed
Eliciting Intrinsic Hallucinations in LLMs via Semantically Equivalent Adversarial Attacks
Atri Vivek Sharma, Brian Formento, Alessio Lomuscio
Abstract
Large language models (LLMs) are often used in conjunction with external knowledge sources to improve their factual accuracy and decrease hallucinations, through methods such as Retrieval-Augmented Generation (RAG). However, these systems remain susceptible to intrinsic hallucinations, where the model generates unfaithful or fabricated information that is not supported by the retrieved evidence. We propose a novel framework to assess model robustness against this phenomenon by stress-testing usi
Categories
Framework mappings
MITRE ATLAS
- AML.T0043Craft Adversarial Data
Suggested from the entry's categories.
Cite
@misc{sharma2026eliciting,
title = {{Eliciting Intrinsic Hallucinations in LLMs via Semantically Equivalent Adversarial Attacks}},
author = {Atri Vivek Sharma and Brian Formento and Alessio Lomuscio},
year = {2026},
month = aug,
eprint = {2608.04286},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.04286}
}