Skip to content
Search
paperAugust 2026Unreviewed

HFAVL: Hard-Label Fusion Attack Toward Vision-Language Model

Hongbo Cao, Yongqi Sun, Li Duan, Yifan Sui, Xisu Wang

IEEE Internet of Things Journal

Abstract

With the advancement of artificial intelligence, vision-language models (VLMs) that integrate text and image modalities have become central to multimodal learning and are increasingly deployed in Internet of Things (IoT) environments such as smart surveillance, autonomous driving, and industrial monitoring. However, VLMs are also highly susceptible to adversarial attacks. Most previous black-box attack methods toward the VLMs rely on either the soft label or substitute models. However, most of t

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data

Suggested from the entry's categories.

Cite

@article{cao2026hfavl,
  title = {{HFAVL: Hard-Label Fusion Attack Toward Vision-Language Model}},
  author = {Hongbo Cao and Yongqi Sun and Li Duan and Yifan Sui and Xisu Wang},
  year = {2026},
  month = aug,
  journal = {IEEE Internet of Things Journal},
  doi = {10.1109/JIOT.2026.3694181},
  url = {https://www.semanticscholar.org/paper/2e81df9a879be060866efffa59cd5fec618d8afe}
}