September 2026Unreviewed
Beyond Single-Pair Attacks: Disrupting Vision-Language Pre-Training Models With Dual-Semantic Frequency Stealth
Hai-Qi Zhang, Zi-Qiang Li, Hao Tang, Ze-Chao Li
IEEE Transactions on Dependable and Secure Computing
Abstract
Vision-Language Pre-training (VLP) models are highly capable in multimodal tasks but are critically vulnerable to adversarial attacks. Existing methods for creating transferable adversarial examples typically operate by modifying semantics within isolated image-text pairs. This strategy often yields poorly generalizable perturbations that insufficiently disrupt cross-modal alignment, leading to limited effectiveness in black-box settings. An additional limitation is the frequent oversight of adv
Categories
Framework mappings
MITRE ATLAS
- AML.T0043Craft Adversarial Data
Suggested from the entry's categories.
Cite
@article{zhang2026beyond,
title = {{Beyond Single-Pair Attacks: Disrupting Vision-Language Pre-Training Models With Dual-Semantic Frequency Stealth}},
author = {Hai-Qi Zhang and Zi-Qiang Li and Hao Tang and Ze-Chao Li},
year = {2026},
month = sep,
journal = {IEEE Transactions on Dependable and Secure Computing},
doi = {10.1109/TDSC.2026.3713210},
url = {https://www.semanticscholar.org/paper/94fb165ddf17a75721511d02381aa3da9e76fd1d}
}