August 2026Unreviewed
SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning
Caoyuan Ma, Wen-Pu Liu, Weichu Xie, Tian Gu, Shilei Zhao, Lingxi Min, Shuai Dong, Yu-Qi Xu, Ji Zhao, Ziyue Wang, Wenzheng Chang, Taiqiang Wu, Yong-Fu Zhu, Wen-Qi Shao, Yinqiang Zheng
Abstract
Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited from their language backbones. We propose SafeCap, a reinforcement-learning framework that aligns LVLMs through learned self-captioning. SafeCap trains a policy model to first generate a safety-relevant image caption and then produce a final answer; the caption is further optimized by whether it enables a frozen LLM to reach a safety-aligned decision. This c
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{ma2026safecap,
title = {{SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning}},
author = {Caoyuan Ma and Wen-Pu Liu and Weichu Xie and Tian Gu and Shilei Zhao and Lingxi Min and Shuai Dong and Yu-Qi Xu and Ji Zhao and Ziyue Wang and Wenzheng Chang and Taiqiang Wu and Yong-Fu Zhu and Wen-Qi Shao and Yinqiang Zheng},
year = {2026},
month = aug,
eprint = {2608.10513},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/e4ddb4f43295dae00e3e9f478c20020294f9157c}
}