2024ReviewedOpen access
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Yuchen Lin, Nathan Lambert, Yejin Choi, Nouha Dziri
arXiv preprint
Abstract
Open-source moderation tool for detecting safety risks in LLM interactions, trained on a diverse dataset of harmful and benign prompts.
Categories
#moderation#open-source#safety-classifier
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
- LLM05Improper Output Handling
Cite
@misc{han2024wildguard,
title = {{WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs}},
author = {Seungju Han and Kavel Rao and Allyson Ettinger and Liwei Jiang and Bill Yuchen Lin and Nathan Lambert and Yejin Choi and Nouha Dziri},
year = {2024},
eprint = {2406.18495},
archivePrefix = {arXiv},
doi = {10.52202/079017-0261},
url = {https://arxiv.org/abs/2406.18495}
}