May 2026Unreviewed
When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?
Yuan Tian, Bing Hu, Fang Wu, Xiaomin Li, Binghang Lu, Neil Zhenqiang Gong
Abstract
Think-with-image reasoning is emerging as a new inference paradigm for large vision-language models, but its safety implications remain poorly understood. Existing systems already span multiple process designs, including direct response generation, text-only prior turn, visual-state manipulation, and explicit external image-tool invocation. In this paper, we ask which of these evaluated paradigms improves multimodal jailbreak robustness, and why. Across multiple vision-language models, explicit
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{tian2026when,
title = {{When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?}},
author = {Yuan Tian and Bing Hu and Fang Wu and Xiaomin Li and Binghang Lu and Neil Zhenqiang Gong},
year = {2026},
month = may,
eprint = {2605.27932},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.27932}
}