August 2026Unreviewed
MMJailBench: A Factorized Benchmark for Disentangling Multimodal Jailbreak Vulnerabilities
Tianshi Wang, Jingsong Wang, Yafei Huang, Fengling Li, Xin Li, Lei Zhu
Abstract
Multimodal Large Language Models (MLLMs) are increasingly deployed in real-world applications, yet how different factors shape their jailbreak vulnerabilities remains poorly understood. Existing benchmarks often couple harmful intent, prompt framing, visual semantics, and instruction carrier within individual jailbreak instances, obscuring the specific sources of observed vulnerabilities. To address this limitation, we introduce MMJailBench, a factorized benchmark that systematically varies and
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{wang2026mmjailbench,
title = {{MMJailBench: A Factorized Benchmark for Disentangling Multimodal Jailbreak Vulnerabilities}},
author = {Tianshi Wang and Jingsong Wang and Yafei Huang and Fengling Li and Xin Li and Lei Zhu},
year = {2026},
month = aug,
eprint = {2608.25490},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.25490}
}