Skip to content
Search
paperAugust 2026Unreviewed

Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks

Haoyu Zhang, Xiangchen Guan, Shibo Zheng, Mohammad Zandsalimy, Shanu Sushmita

Abstract

We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrelated decoy image can sharply lower attack success rate (ASR). The operative change is in the defense pipeline, not in the image. Across five frontier VLMs, two encoded-attack families, and three black-box defenses, a caption-mediated defense (ECSO) that leaves ASR essentially unchanged on text-only encoded input drops i

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{zhang2026decoy,
  title = {{Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks}},
  author = {Haoyu Zhang and Xiangchen Guan and Shibo Zheng and Mohammad Zandsalimy and Shanu Sushmita},
  year = {2026},
  month = aug,
  eprint = {2608.01043},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.01043}
}