Skip to content
Search
paperOctober 2023ReviewedOpen access

LoRA Fine-Tuning Efficiently Undoes Safety Training in Llama 2-Chat

Simon Lermen, Charlie Rogers-Smith, Jeffrey Ladish

arXiv preprint

Abstract

Shows that LoRA fine-tuning with as few as 100 examples can remove safety guardrails from Llama 2-Chat, raising concerns about fine-tuning access to aligned models.

Categories

#LoRA#safety-undoing#fine-tuning#alignment

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0018Manipulate AI Model

Cite

@misc{lermen2023lora,
  title = {{LoRA Fine-Tuning Efficiently Undoes Safety Training in Llama 2-Chat}},
  author = {Simon Lermen and Charlie Rogers-Smith and Jeffrey Ladish},
  year = {2023},
  month = oct,
  eprint = {2310.20624},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2310.20624}
}