August 2026Unreviewed
Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks
Haoyu Zhang, Xiangchen Guan, Shibo Zheng, Mohammad Zandsalimy, Shanu Sushmita
Abstract
We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing an encoded jailbreak prompt with an unrelated decoy image can sharply lower attack success rate (ASR). The operative change is in the defense pipeline, not in the image. Across five frontier VLMs, two encoded-attack families, and three black-box defenses, a caption-mediated defense (ECSO) that leaves ASR essentially unchanged on text-only encoded input drops i
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{zhang2026decoy,
title = {{Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks}},
author = {Haoyu Zhang and Xiangchen Guan and Shibo Zheng and Mohammad Zandsalimy and Shanu Sushmita},
year = {2026},
month = aug,
eprint = {2608.01043},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.01043}
}