January 2024ReviewedOpen access
Jailbroken: How Does LLM Safety Training Fail?
Alexander Wei, Nika Haghtalab, Jacob Steinhardt
NeurIPS 2023
Abstract
Analyzes failure modes of LLM safety training, identifying two broad categories: competing objectives and mismatched generalization, demonstrating attacks that exploit each.
Categories
#safety-training#alignment#failure-modes
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Cite
@inproceedings{wei2024jailbroken,
title = {{Jailbroken: How Does LLM Safety Training Fail?}},
author = {Alexander Wei and Nika Haghtalab and Jacob Steinhardt},
year = {2024},
month = jan,
booktitle = {NeurIPS 2023},
eprint = {2307.02483},
archivePrefix = {arXiv},
doi = {10.52202/075280-3508},
url = {https://arxiv.org/abs/2307.02483}
}