Skip to content
Search
paperSeptember 2026Unreviewed

MPDA: Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models

Xingkai Peng, Jun Jiang, Meng Tong, Shuai Li, Weiming Zhang, Neng H. Yu, Kejiang Chen

IEEE Transactions on Dependable and Secure Computing

Abstract

Text-to-image (T2I) models have been widely applied in generating high-fidelity images across various domains. However, these models may also be abused to produce Not-Safe-for-Work (NSFW) content via jailbreak attacks. Existing jailbreak methods primarily manipulate the textual prompt, leaving potential vulnerabilities in image-based inputs largely unexplored. Moreover, text-based methods face challenges in bypassing the model’s safety filters. In response to these limitations, we propose the Mu

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@article{peng2026mpda,
  title = {{MPDA: Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models}},
  author = {Xingkai Peng and Jun Jiang and Meng Tong and Shuai Li and Weiming Zhang and Neng H. Yu and Kejiang Chen},
  year = {2026},
  month = sep,
  journal = {IEEE Transactions on Dependable and Secure Computing},
  doi = {10.1109/TDSC.2026.3696145},
  url = {https://www.semanticscholar.org/paper/d71b2075655a52c19b88c09dfa9c50250927ade5}
}