August 2026Unreviewed
HFAVL: Hard-Label Fusion Attack Toward Vision-Language Model
Hongbo Cao, Yongqi Sun, Li Duan, Yifan Sui, Xisu Wang
IEEE Internet of Things Journal
Abstract
With the advancement of artificial intelligence, vision-language models (VLMs) that integrate text and image modalities have become central to multimodal learning and are increasingly deployed in Internet of Things (IoT) environments such as smart surveillance, autonomous driving, and industrial monitoring. However, VLMs are also highly susceptible to adversarial attacks. Most previous black-box attack methods toward the VLMs rely on either the soft label or substitute models. However, most of t
Categories
Framework mappings
MITRE ATLAS
- AML.T0043Craft Adversarial Data
Suggested from the entry's categories.
Cite
@article{cao2026hfavl,
title = {{HFAVL: Hard-Label Fusion Attack Toward Vision-Language Model}},
author = {Hongbo Cao and Yongqi Sun and Li Duan and Yifan Sui and Xisu Wang},
year = {2026},
month = aug,
journal = {IEEE Internet of Things Journal},
doi = {10.1109/JIOT.2026.3694181},
url = {https://www.semanticscholar.org/paper/2e81df9a879be060866efffa59cd5fec618d8afe}
}