Skip to content
Search
paperAugust 2026Unreviewed

MMJailBench: A Factorized Benchmark for Disentangling Multimodal Jailbreak Vulnerabilities

Tianshi Wang, Jingsong Wang, Yafei Huang, Fengling Li, Xin Li, Lei Zhu

Abstract

Multimodal Large Language Models (MLLMs) are increasingly deployed in real-world applications, yet how different factors shape their jailbreak vulnerabilities remains poorly understood. Existing benchmarks often couple harmful intent, prompt framing, visual semantics, and instruction carrier within individual jailbreak instances, obscuring the specific sources of observed vulnerabilities. To address this limitation, we introduce MMJailBench, a factorized benchmark that systematically varies and

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{wang2026mmjailbench,
  title = {{MMJailBench: A Factorized Benchmark for Disentangling Multimodal Jailbreak Vulnerabilities}},
  author = {Tianshi Wang and Jingsong Wang and Yafei Huang and Fengling Li and Xin Li and Lei Zhu},
  year = {2026},
  month = aug,
  eprint = {2608.25490},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.25490}
}