June 2025Unreviewed
RedDebate: Safer Responses through Multi-Agent Red Teaming Debates
A. Asad, Stephen Obadinma, Radin Shayanfar, Xiaodan Zhu
arXiv.org
Abstract
We introduce RedDebate, a novel multi-agent debate framework that provides the foundation for Large Language Models (LLMs) to identify and mitigate their unsafe behaviours. Existing AI safety approaches often rely on costly human evaluation or isolated single-model assessment, both constrained by scalability and prone to oversight failures. RedDebate employs collaborative argumentation among multiple LLMs across diverse debate scenarios, enabling them to critically evaluate one another's reasoni
Categories
Framework mappings
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{asad2025reddebate,
title = {{RedDebate: Safer Responses through Multi-Agent Red Teaming Debates}},
author = {A. Asad and Stephen Obadinma and Radin Shayanfar and Xiaodan Zhu},
year = {2025},
month = jun,
eprint = {2506.11083},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2506.11083},
url = {https://www.semanticscholar.org/paper/89b908678883095f8a2eda436bb5aac623b4e18e}
}