June 2025Unreviewed
Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models
Ren-Jian Wang, Ke Xue, Zeyu Qin, Ziniu Li, Sheng Tang, Hao-Tian Li, Shengcai Liu, Chao Qian
arXiv.org
Abstract
Ensuring safety of large language models (LLMs) is important. Red teaming--a systematic approach to identifying adversarial prompts that elicit harmful responses from target LLMs--has emerged as a crucial safety evaluation method. Within this framework, the diversity of adversarial prompts is essential for comprehensive safety assessments. We find that previous approaches to red-teaming may suffer from two key limitations. First, they often pursue diversity through simplistic metrics like word f
Categories
Framework mappings
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{wang2025qualitydiversity,
title = {{Quality-Diversity Red-Teaming: Automated Generation of High-Quality and Diverse Attackers for Large Language Models}},
author = {Ren-Jian Wang and Ke Xue and Zeyu Qin and Ziniu Li and Sheng Tang and Hao-Tian Li and Shengcai Liu and Chao Qian},
year = {2025},
month = jun,
eprint = {2506.07121},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2506.07121},
url = {https://www.semanticscholar.org/paper/da14e71f31e98612ebf871be742dddc55da509d1}
}