Skip to content
Search
paperJune 2026Unreviewed

From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails

Yuguang Zhou, Xunguang Wang, Pingchuan Ma, Zhantong Xue, Zhaoyu Wang, Shuai Wang

Abstract

LLM-based guardrails have emerged as a highly effective defense against prompt injection and jailbreak attacks in autonomous agents. However, we reveal that the very reasoning and task-following capabilities enabling this protection introduce a novel vulnerability: attackers can inject crafted data to trap the guardrail in extended reasoning loops, effectuating a systematic denial-of-service (DoS) attack. To systematically expose this threat, we design a beam-search optimization framework that c

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{zhou2026from,
  title = {{From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails}},
  author = {Yuguang Zhou and Xunguang Wang and Pingchuan Ma and Zhantong Xue and Zhaoyu Wang and Shuai Wang},
  year = {2026},
  month = jun,
  eprint = {2606.14517},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.14517}
}