September 2026Unreviewed
MPDA: Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models
Xingkai Peng, Jun Jiang, Meng Tong, Shuai Li, Weiming Zhang, Neng H. Yu, Kejiang Chen
IEEE Transactions on Dependable and Secure Computing
Abstract
Text-to-image (T2I) models have been widely applied in generating high-fidelity images across various domains. However, these models may also be abused to produce Not-Safe-for-Work (NSFW) content via jailbreak attacks. Existing jailbreak methods primarily manipulate the textual prompt, leaving potential vulnerabilities in image-based inputs largely unexplored. Moreover, text-based methods face challenges in bypassing the model’s safety filters. In response to these limitations, we propose the Mu
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@article{peng2026mpda,
title = {{MPDA: Multimodal Prompt Decoupling Attack on the Safety Filters in Text-to-Image Models}},
author = {Xingkai Peng and Jun Jiang and Meng Tong and Shuai Li and Weiming Zhang and Neng H. Yu and Kejiang Chen},
year = {2026},
month = sep,
journal = {IEEE Transactions on Dependable and Secure Computing},
doi = {10.1109/TDSC.2026.3696145},
url = {https://www.semanticscholar.org/paper/d71b2075655a52c19b88c09dfa9c50250927ade5}
}