Skip to content
Search
paper2024ReviewedOpen access

Are Aligned Neural Networks Adversarially Aligned?

Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski, Irena Gao, Anas Awadalla, Pang Wei Koh, Daphne Ippolito, Katherine Lee, Florian Tramer, Ludwig Schmidt

NeurIPS 2023

Abstract

Evaluates whether multimodal LLMs aligned to refuse harmful text requests also refuse harmful image-based requests, finding significant gaps.

Categories

#multimodal#alignment-gap#image-attacks

Framework mappings

Cite

@inproceedings{carlini2024are,
  title = {{Are Aligned Neural Networks Adversarially Aligned?}},
  author = {Nicholas Carlini and Milad Nasr and Christopher A. Choquette-Choo and Matthew Jagielski and Irena Gao and Anas Awadalla and Pang Wei Koh and Daphne Ippolito and Katherine Lee and Florian Tramer and Ludwig Schmidt},
  year = {2024},
  booktitle = {NeurIPS 2023},
  eprint = {2306.15447},
  archivePrefix = {arXiv},
  doi = {10.52202/075280-2687},
  url = {https://arxiv.org/abs/2306.15447}
}