Skip to content
Search
paperAugust 2026Unreviewed

Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models

Wenyun Li, Guiping Cao, Xiangyuan Lan, Zheng Zhang

Abstract

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language interaction, yet their safety alignment remains vulnerable to jailbreak attacks. A key challenge is that safety behavior learned in the textual space does not reliably transfer to fused cross-modal representations, leaving multimodal inputs exploitable through latent semantic cues. We propose Text-Anchored Semantic Perturbation Attack (TA-SPA), a black-box jailbreak framework that optimizes transferable

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{li2026textanchored,
  title = {{Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models}},
  author = {Wenyun Li and Guiping Cao and Xiangyuan Lan and Zheng Zhang},
  year = {2026},
  month = aug,
  eprint = {2608.22312},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.22312}
}