May 2024ReviewedOpen access
TrojLLM: A Black-box Trojan Prompt Attack on Large Language Models
Jiaqi Xue, Mengxin Zheng, Ting Hua, Yilin Shen, Yepeng Liu, Ladislau Boloni, Qian Lou
NeurIPS 2023
Abstract
Proposes TrojLLM, a black-box attack that generates universal trojan prompts to compromise LLMs without access to model internals.
Categories
#trojan#backdoor#black-box
Framework mappings
OWASP Top 10 for LLM Applications
- LLM03Supply Chain
- LLM04Data and Model Poisoning
MITRE ATLAS
- AML.T0018Manipulate AI Model
- AML.T0043Craft Adversarial Data
Cite
@inproceedings{xue2024trojllm,
title = {{TrojLLM: A Black-box Trojan Prompt Attack on Large Language Models}},
author = {Jiaqi Xue and Mengxin Zheng and Ting Hua and Yilin Shen and Yepeng Liu and Ladislau Boloni and Qian Lou},
year = {2024},
month = may,
booktitle = {NeurIPS 2023},
eprint = {2306.06815},
archivePrefix = {arXiv},
doi = {10.52202/075280-2866},
url = {https://arxiv.org/abs/2306.06815}
}