Skip to content
Search
paperAugust 2026Unreviewed

Eliciting Intrinsic Hallucinations in LLMs via Semantically Equivalent Adversarial Attacks

Atri Vivek Sharma, Brian Formento, Alessio Lomuscio

Abstract

Large language models (LLMs) are often used in conjunction with external knowledge sources to improve their factual accuracy and decrease hallucinations, through methods such as Retrieval-Augmented Generation (RAG). However, these systems remain susceptible to intrinsic hallucinations, where the model generates unfaithful or fabricated information that is not supported by the retrieved evidence. We propose a novel framework to assess model robustness against this phenomenon by stress-testing usi

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data

Suggested from the entry's categories.

Cite

@misc{sharma2026eliciting,
  title = {{Eliciting Intrinsic Hallucinations in LLMs via Semantically Equivalent Adversarial Attacks}},
  author = {Atri Vivek Sharma and Brian Formento and Alessio Lomuscio},
  year = {2026},
  month = aug,
  eprint = {2608.04286},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.04286}
}