August 2026Unreviewed
Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models
Wenyun Li, Guiping Cao, Xiangyuan Lan, Zheng Zhang
Abstract
Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language interaction, yet their safety alignment remains vulnerable to jailbreak attacks. A key challenge is that safety behavior learned in the textual space does not reliably transfer to fused cross-modal representations, leaving multimodal inputs exploitable through latent semantic cues. We propose Text-Anchored Semantic Perturbation Attack (TA-SPA), a black-box jailbreak framework that optimizes transferable
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{li2026textanchored,
title = {{Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models}},
author = {Wenyun Li and Guiping Cao and Xiangyuan Lan and Zheng Zhang},
year = {2026},
month = aug,
eprint = {2608.22312},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.22312}
}