Skip to content
Search
paperDecember 2023ReviewedOpen access

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, Madian Khabsa

arXiv preprint

Abstract

Introduces Llama Guard, an LLM-based safeguard model for classifying safety risks in LLM inputs and outputs, achieving strong performance on standard benchmarks.

Categories

#llama-guard#content-safety#classifier

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM01Prompt Injection
  • LLM05Improper Output Handling

Cite

@misc{inan2023llama,
  title = {{Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations}},
  author = {Hakan Inan and Kartikeya Upasani and Jianfeng Chi and Rashi Rungta and Krithika Iyer and Yuning Mao and Michael Tontchev and Qing Hu and Brian Fuller and Davide Testuggine and Madian Khabsa},
  year = {2023},
  month = dec,
  eprint = {2312.06674},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2312.06674}
}