October 2023Unreviewed
ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language Models
Alex Mei, Sharon Levy, W. Wang
Conference on Empirical Methods in Natural Language Processing
Abstract
As large language models are integrated into society, robustness toward a suite of prompts is increasingly important to maintain reliability in a high-variance environment.Robustness evaluations must comprehensively encapsulate the various settings in which a user may invoke an intelligent system. This paper proposes ASSERT, Automated Safety Scenario Red Teaming, consisting of three methods -- semantically aligned augmentation, target bootstrapping, and adversarial knowledge injection. For robus
Categories
Framework mappings
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@inproceedings{mei2023assert,
title = {{ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language Models}},
author = {Alex Mei and Sharon Levy and W. Wang},
year = {2023},
month = oct,
booktitle = {Conference on Empirical Methods in Natural Language Processing},
eprint = {2310.09624},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2310.09624},
url = {https://www.semanticscholar.org/paper/aa9aa1c315cb2a0c1759d82fb3d4b4506c2dbb7c}
}