August 2026Unreviewed
Adversarial Attacks on Deep OCR Systems
Wenbo Sun, Hong-Zong Li, Yanyun Wang, Jia-Hao Ma, Shuxin Zhuang, Rong Feng, Shi-Qin Tang, Zi Liang
Abstract
Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-context OCR at low token cost. However, its increased complexity may introduce new security vulnerabilities. In this paper, we present, to the best of our knowledge, the first pure black-box adversarial attack against a generative OCR vision-language model, where only the decoded string can be queried and no gradients, logits, or model internals are available. We
Categories
Framework mappings
MITRE ATLAS
- AML.T0043Craft Adversarial Data
Suggested from the entry's categories.
Cite
@misc{sun2026adversarial,
title = {{Adversarial Attacks on Deep OCR Systems}},
author = {Wenbo Sun and Hong-Zong Li and Yanyun Wang and Jia-Hao Ma and Shuxin Zhuang and Rong Feng and Shi-Qin Tang and Zi Liang},
year = {2026},
month = aug,
eprint = {2608.07636},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/949627af3c2bcd16e81a3c0d9d82bf097b7b3bce}
}