Skip to content
Search
paperFebruary 2024Unreviewed

Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast

Xiangming Gu, Xiaosen Zheng, Tianyu Pang, Chao Du, Qian Liu, Ye Wang, Jing Jiang, Min Lin

International Conference on Machine Learning

Abstract

A multimodal large language model (MLLM) agent can receive instructions, capture images, retrieve histories from memory, and decide which tools to use. Nonetheless, red-teaming efforts have revealed that adversarial images/prompts can jailbreak an MLLM and cause unaligned behaviors. In this work, we report an even more severe safety issue in multi-agent environments, referred to as infectious jailbreak. It entails the adversary simply jailbreaking a single agent, and without any further interven

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@inproceedings{gu2024agent,
  title = {{Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast}},
  author = {Xiangming Gu and Xiaosen Zheng and Tianyu Pang and Chao Du and Qian Liu and Ye Wang and Jing Jiang and Min Lin},
  year = {2024},
  month = feb,
  booktitle = {International Conference on Machine Learning},
  eprint = {2402.08567},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2402.08567},
  url = {https://www.semanticscholar.org/paper/b0ada492ba48e85016cbbfd95ec7180fb7e79648}
}