June 2026Unreviewed
What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text Attacks
Qin Yang, Lu Malloy, Joshua Lee, Xiaohan Chang, Meisam Mohammady, Doowon Kim, Yuan Hong
Abstract
Large language model (LLM)-powered content moderation systems are a critical defense against harmful online content. However, they operate primarily on tokenized text and often overlook visual cues that humans naturally use when interpreting content. We show that this limitation creates a fundamental vulnerability: content readily recognized as harmful by humans can evade automated moderation. To systematically study this problem, we introduce Human-Perceptible Adversarial Attacks (HPAA), which
Categories
Framework mappings
MITRE ATLAS
- AML.T0043Craft Adversarial Data
Suggested from the entry's categories.
Cite
@misc{yang2026what,
title = {{What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text Attacks}},
author = {Qin Yang and Lu Malloy and Joshua Lee and Xiaohan Chang and Meisam Mohammady and Doowon Kim and Yuan Hong},
year = {2026},
month = jun,
eprint = {2606.09700},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.09700}
}