Skip to content
Search
paperSeptember 2026Unreviewed

Multimodal Resource-Exhaustion Attacks on Vision-Language Models via Joint Pixel-Prompt Optimization

Zhaoxiong Ni, Yatie Xiao, Chi-Man Pun, Fei Peng, Qingxiao Guan, Keke Tang

Abstract

Resource-exhaustion attacks against autoregressive vision-language models (VLMs) typically assume unimodal threat models, treating the image branch as the primary optimization surface while holding user-visible prompts fixed. Even recent loop-centric variants remain confined to this single-channel paradigm, leaving the exploitation of availability unexplored as a cross-modal optimization problem over jointly controllable input surfaces. We introduce Joint Pixel-Prompt Optimization (JPPO), the fi

Categories

Cite

@misc{ni2026multimodal,
  title = {{Multimodal Resource-Exhaustion Attacks on Vision-Language Models via Joint Pixel-Prompt Optimization}},
  author = {Zhaoxiong Ni and Yatie Xiao and Chi-Man Pun and Fei Peng and Qingxiao Guan and Keke Tang},
  year = {2026},
  month = sep,
  eprint = {2609.05889},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.05889}
}