December 2023ReviewedOpen access
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, Madian Khabsa
arXiv preprint
Abstract
Introduces Llama Guard, an LLM-based safeguard model for classifying safety risks in LLM inputs and outputs, achieving strong performance on standard benchmarks.
Categories
#llama-guard#content-safety#classifier
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
- LLM05Improper Output Handling
Cite
@misc{inan2023llama,
title = {{Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations}},
author = {Hakan Inan and Kartikeya Upasani and Jianfeng Chi and Rashi Rungta and Krithika Iyer and Yuning Mao and Michael Tontchev and Qing Hu and Brian Fuller and Davide Testuggine and Madian Khabsa},
year = {2023},
month = dec,
eprint = {2312.06674},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2312.06674}
}