Skip to content
Search
paperSeptember 2025Unreviewed

PersonaTeaming: Exploring How Introducing Personas Can Improve Automated AI Red-Teaming

Wesley Hanwen Deng, Sunnie S. Y. Kim, Akshita Jha, Kenneth Holstein, Motahhare Eslami, L. Wilcox, Leon A. Gatys

arXiv.org

Abstract

Recent developments in AI governance and safety research have called for red-teaming methods that can effectively surface potential risks posed by AI models. Many of these calls have emphasized how the identities and backgrounds of red-teamers can shape their red-teaming strategies, and thus the kinds of risks they are likely to uncover. While automated red-teaming approaches promise to complement human red-teaming by enabling larger-scale exploration of model behavior, current approaches do not

Categories

Framework mappings

Suggested from the entry's categories.

Cite

@misc{deng2025personateaming,
  title = {{PersonaTeaming: Exploring How Introducing Personas Can Improve Automated AI Red-Teaming}},
  author = {Wesley Hanwen Deng and Sunnie S. Y. Kim and Akshita Jha and Kenneth Holstein and Motahhare Eslami and L. Wilcox and Leon A. Gatys},
  year = {2025},
  month = sep,
  eprint = {2509.03728},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2509.03728},
  url = {https://www.semanticscholar.org/paper/d9490d1ce9f58c5d545941246854e41831e2a74f}
}