Skip to content
Search
paperAugust 2026Unreviewed

The Guard That Cried Wolf: How Scary Words Make Agent Guardrails Refuse Legitimate Actions

Yingjie Zhang, Yuanbo Xie, Kai Chen

Abstract

Agent guardrails are checks that approve or refuse each action before an LLM executes it. Sometimes they refuse requests that are genuinely safe. This over-safety blocks deployment when a guardrail refuses an authorized task. Evaluating over-safety is hard: at the boundary an authorized action resembles an unauthorized one, and the safe-versus-unsafe label is a choice of authorization policy, not fixed by the action alone. We argue it therefore requires a benchmark that does not yet exist, one t

Categories

Cite

@misc{zhang2026guard,
  title = {{The Guard That Cried Wolf: How Scary Words Make Agent Guardrails Refuse Legitimate Actions}},
  author = {Yingjie Zhang and Yuanbo Xie and Kai Chen},
  year = {2026},
  month = aug,
  eprint = {2608.27009},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/31f4998e2bf81619cb2868927931efc2c5f3672f}
}