Skip to content
Search
paperMay 2026Unreviewed

REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations

Buyun Liang, Jinqi Luo, Liangzu Peng, Kwan Ho Ryan Chan, Darshan Thaker, Kaleab A. Kinfu, Fengrui Tian, Hamed Hassani, René Vidal

Abstract

Large language models (LLMs) achieve strong performance across many tasks but remain vulnerable to hallucinations, motivating the need for realistic adversarial prompts that elicit such failures. We formulate hallucination elicitation as a constrained optimization problem, where the goal is to find semantically coherent adversarial prompts that are equivalent to benign user prompts. Existing methods remain limited: discrete prompt-based attacks preserve semantic equivalence and coherence but sea

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data

Suggested from the entry's categories.

Cite

@misc{liang2026realista,
  title = {{REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations}},
  author = {Buyun Liang and Jinqi Luo and Liangzu Peng and Kwan Ho Ryan Chan and Darshan Thaker and Kaleab A. Kinfu and Fengrui Tian and Hamed Hassani and René Vidal},
  year = {2026},
  month = may,
  eprint = {2605.12813},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.12813}
}