Skip to content
Search
paperMay 2023ReviewedOpen access

Poisoning Language Models During Instruction Tuning

Alexander Wan, Eric Wallace, Sheng Shen, Dan Klein

ICML 2023

Abstract

Shows that adversaries can insert poisoned examples into instruction-tuning datasets, causing models to generate targeted outputs for attacker-chosen triggers.

Categories

#instruction-tuning#backdoor#fine-tuning

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0020Poison Training Data

Cite

@inproceedings{wan2023poisoning,
  title = {{Poisoning Language Models During Instruction Tuning}},
  author = {Alexander Wan and Eric Wallace and Sheng Shen and Dan Klein},
  year = {2023},
  month = may,
  booktitle = {ICML 2023},
  eprint = {2305.00944},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2305.00944}
}