Skip to content
Search
paperApril 2025Unreviewed

Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems

William Hackett, Lewis Birch, Stefan Trawicki, Neeraj Suri, Peter Garraghan

Abstract

Large Language Models (LLMs) guardrail systems are designed to protect against prompt injection and jailbreak attacks. However, they remain vulnerable to evasion techniques. We demonstrate two approaches for bypassing LLM prompt injection and jailbreak detection systems via traditional character injection methods and algorithmic Adversarial Machine Learning (AML) evasion techniques. Through testing against six prominent protection systems, including Microsoft's Azure Prompt Shield and Meta's Pro

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data
  • AML.T0051LLM Prompt Injection
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{hackett2025bypassing,
  title = {{Bypassing LLM Guardrails: An Empirical Analysis of Evasion Attacks against Prompt Injection and Jailbreak Detection Systems}},
  author = {William Hackett and Lewis Birch and Stefan Trawicki and Neeraj Suri and Peter Garraghan},
  year = {2025},
  month = apr,
  eprint = {2504.11168},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/f74726bf146835dc48ba4d8ab2dcfd0ed762af8e}
}