Skip to content
Search
paperJanuary 2024ReviewedOpen access

Jailbroken: How Does LLM Safety Training Fail?

Alexander Wei, Nika Haghtalab, Jacob Steinhardt

NeurIPS 2023

Abstract

Analyzes failure modes of LLM safety training, identifying two broad categories: competing objectives and mismatched generalization, demonstrating attacks that exploit each.

Categories

#safety-training#alignment#failure-modes

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Cite

@inproceedings{wei2024jailbroken,
  title = {{Jailbroken: How Does LLM Safety Training Fail?}},
  author = {Alexander Wei and Nika Haghtalab and Jacob Steinhardt},
  year = {2024},
  month = jan,
  booktitle = {NeurIPS 2023},
  eprint = {2307.02483},
  archivePrefix = {arXiv},
  doi = {10.52202/075280-3508},
  url = {https://arxiv.org/abs/2307.02483}
}