Skip to content
Search
paperJuly 2023ReviewedOpen access

Universal and Transferable Adversarial Attacks on Aligned Language Models

Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, Matt Fredrikson

arXiv preprint

Abstract

Proposes an automated method (GCG) to generate adversarial suffixes that cause aligned LLMs to produce harmful content, with attacks transferring across models including ChatGPT and Claude.

Categories

#GCG#adversarial-suffix#transferability

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data
  • AML.T0054LLM Jailbreak

Cite

@misc{zou2023universal,
  title = {{Universal and Transferable Adversarial Attacks on Aligned Language Models}},
  author = {Andy Zou and Zifan Wang and Nicholas Carlini and Milad Nasr and J. Zico Kolter and Matt Fredrikson},
  year = {2023},
  month = jul,
  eprint = {2307.15043},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2307.15043}
}