Skip to content
Search
paperJuly 2026Unreviewed

Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks

Haoyu Zhang, Zhuoxi Wang, Shibo Zheng, Zijian Xiao, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita

Abstract

Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a rare language, code, or an image of text slips past a guard that would block it in plain language -- the decode gap. The natural fix is a guard-agnostic recover-and-decode amplifier that transcribes image content and restates encoded text into its plain payload before the guard, so any off

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{zhang2026recover,
  title = {{Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks}},
  author = {Haoyu Zhang and Zhuoxi Wang and Shibo Zheng and Zijian Xiao and Xiangchen Guan and Mohammad Zandsalimy and Shanu Sushmita},
  year = {2026},
  month = jul,
  eprint = {2607.26574},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.26574}
}