Skip to content
Search
paperAugust 2026Unreviewed

Adversarial Attacks on Deep OCR Systems

Wenbo Sun, Hong-Zong Li, Yanyun Wang, Jia-Hao Ma, Shuxin Zhuang, Rong Feng, Shi-Qin Tang, Zi Liang

Abstract

Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-context OCR at low token cost. However, its increased complexity may introduce new security vulnerabilities. In this paper, we present, to the best of our knowledge, the first pure black-box adversarial attack against a generative OCR vision-language model, where only the decoded string can be queried and no gradients, logits, or model internals are available. We

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data

Suggested from the entry's categories.

Cite

@misc{sun2026adversarial,
  title = {{Adversarial Attacks on Deep OCR Systems}},
  author = {Wenbo Sun and Hong-Zong Li and Yanyun Wang and Jia-Hao Ma and Shuxin Zhuang and Rong Feng and Shi-Qin Tang and Zi Liang},
  year = {2026},
  month = aug,
  eprint = {2608.07636},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/949627af3c2bcd16e81a3c0d9d82bf097b7b3bce}
}