October 2023ReviewedOpen access
LoRA Fine-Tuning Efficiently Undoes Safety Training in Llama 2-Chat
Simon Lermen, Charlie Rogers-Smith, Jeffrey Ladish
arXiv preprint
Abstract
Shows that LoRA fine-tuning with as few as 100 examples can remove safety guardrails from Llama 2-Chat, raising concerns about fine-tuning access to aligned models.
Categories
#LoRA#safety-undoing#fine-tuning#alignment
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0018Manipulate AI Model
Cite
@misc{lermen2023lora,
title = {{LoRA Fine-Tuning Efficiently Undoes Safety Training in Llama 2-Chat}},
author = {Simon Lermen and Charlie Rogers-Smith and Jeffrey Ladish},
year = {2023},
month = oct,
eprint = {2310.20624},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2310.20624}
}