April 2025Unreviewed
Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems
William Hackett, Lewis Birch, Stefan Trawicki, Neeraj Suri, Peter Garraghan
Abstract
Large Language Models (LLMs) guardrail systems are designed to protect against prompt injection and jailbreak attacks. However, they remain vulnerable to evasion techniques. We demonstrate two approaches for bypassing LLM prompt injection and jailbreak detection systems via traditional character injection methods and algorithmic Adversarial Machine Learning (AML) evasion techniques. Through testing against six prominent protection systems, including Microsoft's Azure Prompt Shield and Meta's Pro
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0043Craft Adversarial Data
- AML.T0051LLM Prompt Injection
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{hackett2025bypassing,
title = {{Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems}},
author = {William Hackett and Lewis Birch and Stefan Trawicki and Neeraj Suri and Peter Garraghan},
year = {2025},
month = apr,
eprint = {2504.11168},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/f74726bf146835dc48ba4d8ab2dcfd0ed762af8e}
}