Skip to content
Search
paperMay 2024ReviewedOpen access

TrojLLM: A Black-box Trojan Prompt Attack on Large Language Models

Jiaqi Xue, Mengxin Zheng, Ting Hua, Yilin Shen, Yepeng Liu, Ladislau Boloni, Qian Lou

NeurIPS 2023

Abstract

Proposes TrojLLM, a black-box attack that generates universal trojan prompts to compromise LLMs without access to model internals.

Categories

#trojan#backdoor#black-box

Framework mappings

OWASP Top 10 for LLM Applications
  • LLM03Supply Chain
  • LLM04Data and Model Poisoning
MITRE ATLAS
  • AML.T0018Manipulate AI Model
  • AML.T0043Craft Adversarial Data

Cite

@inproceedings{xue2024trojllm,
  title = {{TrojLLM: A Black-box Trojan Prompt Attack on Large Language Models}},
  author = {Jiaqi Xue and Mengxin Zheng and Ting Hua and Yilin Shen and Yepeng Liu and Ladislau Boloni and Qian Lou},
  year = {2024},
  month = may,
  booktitle = {NeurIPS 2023},
  eprint = {2306.06815},
  archivePrefix = {arXiv},
  doi = {10.52202/075280-2866},
  url = {https://arxiv.org/abs/2306.06815}
}