Skip to content
Search
paperOctober 2023Unreviewed

ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language Models

Alex Mei, Sharon Levy, W. Wang

Conference on Empirical Methods in Natural Language Processing

Abstract

As large language models are integrated into society, robustness toward a suite of prompts is increasingly important to maintain reliability in a high-variance environment.Robustness evaluations must comprehensively encapsulate the various settings in which a user may invoke an intelligent system. This paper proposes ASSERT, Automated Safety Scenario Red Teaming, consisting of three methods -- semantically aligned augmentation, target bootstrapping, and adversarial knowledge injection. For robus

Categories

Framework mappings

Suggested from the entry's categories.

Cite

@inproceedings{mei2023assert,
  title = {{ASSERT: Automated Safety Scenario Red Teaming for Evaluating the Robustness of Large Language Models}},
  author = {Alex Mei and Sharon Levy and W. Wang},
  year = {2023},
  month = oct,
  booktitle = {Conference on Empirical Methods in Natural Language Processing},
  eprint = {2310.09624},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2310.09624},
  url = {https://www.semanticscholar.org/paper/aa9aa1c315cb2a0c1759d82fb3d4b4506c2dbb7c}
}