2024ReviewedOpen access
A StrongREJECT for Empty Jailbreaks
Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, Sam Toyer
NeurIPS 2024 Datasets and Benchmarks
Abstract
Introduces StrongREJECT, a high-quality evaluation benchmark for measuring how well LLMs refuse harmful requests.
Categories
#evaluation#refusal#benchmark#reliability
Framework mappings
NIST AI Risk Management Framework
- MEASUREMeasure
Cite
@inproceedings{souly2024strongreject,
title = {{A StrongREJECT for Empty Jailbreaks}},
author = {Alexandra Souly and Qingyuan Lu and Dillon Bowen and Tu Trinh and Elvis Hsieh and Sana Pandey and Pieter Abbeel and Justin Svegliato and Scott Emmons and Olivia Watkins and Sam Toyer},
year = {2024},
booktitle = {NeurIPS 2024 Datasets and Benchmarks},
eprint = {2402.10260},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2402.10260}
}