Skip to content
Search
paperAugust 2026Unreviewed

SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning

Caoyuan Ma, Wen-Pu Liu, Weichu Xie, Tian Gu, Shilei Zhao, Lingxi Min, Shuai Dong, Yu-Qi Xu, Ji Zhao, Ziyue Wang, Wenzheng Chang, Taiqiang Wu, Yong-Fu Zhu, Wen-Qi Shao, Yinqiang Zheng

Abstract

Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones. We propose SafeCap, a reinforcement-learning framework that aligns LVLMs through learned self-captioning. SafeCap trains a policy model to first generate a safety-relevant image caption and then produce a final answer; the caption is further optimized by whether it enables a frozen LLM to reach a safety-aligned decision. This c

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{ma2026safecap,
  title = {{SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning}},
  author = {Caoyuan Ma and Wen-Pu Liu and Weichu Xie and Tian Gu and Shilei Zhao and Lingxi Min and Shuai Dong and Yu-Qi Xu and Ji Zhao and Ziyue Wang and Wenzheng Chang and Taiqiang Wu and Yong-Fu Zhu and Wen-Qi Shao and Yinqiang Zheng},
  year = {2026},
  month = aug,
  eprint = {2608.10513},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/e4ddb4f43295dae00e3e9f478c20020294f9157c}
}