Skip to content
Search
paperJuly 2023Unreviewed

Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models

Erfan Shayegani, Yue Dong, Nael B. Abu-Ghazaleh

International Conference on Learning Representations

Abstract

We introduce new jailbreak attacks on vision language models (VLMs), which use aligned LLMs and are resilient to text-only jailbreak attacks. Specifically, we develop cross-modality attacks on alignment where we pair adversarial images going through the vision encoder with textual prompts to break the alignment of the language model. Our attacks employ a novel compositional strategy that combines an image, adversarially targeted towards toxic embeddings, with generic prompts to accomplish the ja

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@inproceedings{shayegani2023jailbreak,
  title = {{Jailbreak in pieces: Compositional Adversarial Attacks on Multi-Modal Language Models}},
  author = {Erfan Shayegani and Yue Dong and Nael B. Abu-Ghazaleh},
  year = {2023},
  month = jul,
  booktitle = {International Conference on Learning Representations},
  eprint = {2307.14539},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/92b9d8b8c81c4c53ea62000c0924500b2dd11bce}
}