Skip to content
Search
paper2024ReviewedOpen access

WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs

Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Yuchen Lin, Nathan Lambert, Yejin Choi, Nouha Dziri

arXiv preprint

Abstract

Open-source moderation tool for detecting safety risks in LLM interactions, trained on a diverse dataset of harmful and benign prompts.

Categories

#moderation#open-source#safety-classifier

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM01Prompt Injection
  • LLM05Improper Output Handling

Cite

@misc{han2024wildguard,
  title = {{WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs}},
  author = {Seungju Han and Kavel Rao and Allyson Ettinger and Liwei Jiang and Bill Yuchen Lin and Nathan Lambert and Yejin Choi and Nouha Dziri},
  year = {2024},
  eprint = {2406.18495},
  archivePrefix = {arXiv},
  doi = {10.52202/079017-0261},
  url = {https://arxiv.org/abs/2406.18495}
}