Skip to content
Search
paperMay 2026Unreviewed

When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?

Yuan Tian, Bing Hu, Fang Wu, Xiaomin Li, Binghang Lu, Neil Zhenqiang Gong

Abstract

Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly understood. Existing systems already span multiple process designs, including direct response generation, text-only prior turn, visual-state manipulation, and explicit external image-tool invocation. In this paper, we ask which of these evaluated paradigms improves multimodal jailbreak robustness, and why. Across multiple vision-language models, explicit

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{tian2026when,
  title = {{When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?}},
  author = {Yuan Tian and Bing Hu and Fang Wu and Xiaomin Li and Binghang Lu and Neil Zhenqiang Gong},
  year = {2026},
  month = may,
  eprint = {2605.27932},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.27932}
}