Skip to content
Search
paperJune 2026Unreviewed

What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text Attacks

Qin Yang, Lu Malloy, Joshua Lee, Xiaohan Chang, Meisam Mohammady, Doowon Kim, Yuan Hong

Abstract

Large language model (LLM)-powered content moderation systems are a critical defense against harmful online content. However, they operate primarily on tokenized text and often overlook visual cues that humans naturally use when interpreting content. We show that this limitation creates a fundamental vulnerability: content readily recognized as harmful by humans can evade automated moderation. To systematically study this problem, we introduce Human-Perceptible Adversarial Attacks (HPAA), which

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data

Suggested from the entry's categories.

Cite

@misc{yang2026what,
  title = {{What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text Attacks}},
  author = {Qin Yang and Lu Malloy and Joshua Lee and Xiaohan Chang and Meisam Mohammady and Doowon Kim and Yuan Hong},
  year = {2026},
  month = jun,
  eprint = {2606.09700},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.09700}
}