May 2023ReviewedOpen access
Poisoning Language Models During Instruction Tuning
Alexander Wan, Eric Wallace, Sheng Shen, Dan Klein
ICML 2023
Abstract
Shows that adversaries can insert poisoned examples into instruction-tuning datasets, causing models to generate targeted outputs for attacker-chosen triggers.
Categories
#instruction-tuning#backdoor#fine-tuning
Framework mappings
OWASP Top 10 for LLM Applications
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0020Poison Training Data
Cite
@inproceedings{wan2023poisoning,
title = {{Poisoning Language Models During Instruction Tuning}},
author = {Alexander Wan and Eric Wallace and Sheng Shen and Dan Klein},
year = {2023},
month = may,
booktitle = {ICML 2023},
eprint = {2305.00944},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2305.00944}
}