January 2024ReviewedOpen access
Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models
Jingwei Yi, Yueqi Xie, Bin Zhu, Keegan Hines, Emre Kiciman, Guangzhong Sun, Xing Xie, Fangzhao Wu
arXiv preprint
Abstract
Provides a benchmark for indirect prompt injection attacks and evaluates several defense strategies including perplexity-based detection and sandwich defense.
Categories
#indirect-injection#benchmark#defense-evaluation
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
Cite
@misc{yi2024benchmarking,
title = {{Benchmarking and Defending Against Indirect Prompt Injection Attacks on Large Language Models}},
author = {Jingwei Yi and Yueqi Xie and Bin Zhu and Keegan Hines and Emre Kiciman and Guangzhong Sun and Xing Xie and Fangzhao Wu},
year = {2024},
month = jan,
eprint = {2312.14197},
archivePrefix = {arXiv},
doi = {10.1145/3690624.3709179},
url = {https://arxiv.org/abs/2312.14197}
}