Skip to content
Search
paperSeptember 2026Unreviewed

Beyond Single-Pair Attacks: Disrupting Vision-Language Pre-Training Models With Dual-Semantic Frequency Stealth

Hai-Qi Zhang, Zi-Qiang Li, Hao Tang, Ze-Chao Li

IEEE Transactions on Dependable and Secure Computing

Abstract

Vision-Language Pre-training (VLP) models are highly capable in multimodal tasks but are critically vulnerable to adversarial attacks. Existing methods for creating transferable adversarial examples typically operate by modifying semantics within isolated image-text pairs. This strategy often yields poorly generalizable perturbations that insufficiently disrupt cross-modal alignment, leading to limited effectiveness in black-box settings. An additional limitation is the frequent oversight of adv

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data

Suggested from the entry's categories.

Cite

@article{zhang2026beyond,
  title = {{Beyond Single-Pair Attacks: Disrupting Vision-Language Pre-Training Models With Dual-Semantic Frequency Stealth}},
  author = {Hai-Qi Zhang and Zi-Qiang Li and Hao Tang and Ze-Chao Li},
  year = {2026},
  month = sep,
  journal = {IEEE Transactions on Dependable and Secure Computing},
  doi = {10.1109/TDSC.2026.3713210},
  url = {https://www.semanticscholar.org/paper/94fb165ddf17a75721511d02381aa3da9e76fd1d}
}