August 2026Unreviewed
The Guard That Cried Wolf: How Scary Words Make Agent Guardrails Refuse Legitimate Actions
Yingjie Zhang, Yuanbo Xie, Kai Chen
Abstract
Agent guardrails are checks that approve or refuse each action before an LLM executes it. Sometimes they refuse requests that are genuinely safe. This over-safety blocks deployment when a guardrail refuses an authorized task. Evaluating over-safety is hard: at the boundary an authorized action resembles an unauthorized one, and the safe-versus-unsafe label is a choice of authorization policy, not fixed by the action alone. We argue it therefore requires a benchmark that does not yet exist, one t
Categories
Cite
@misc{zhang2026guard,
title = {{The Guard That Cried Wolf: How Scary Words Make Agent Guardrails Refuse Legitimate Actions}},
author = {Yingjie Zhang and Yuanbo Xie and Kai Chen},
year = {2026},
month = aug,
eprint = {2608.27009},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/31f4998e2bf81619cb2868927931efc2c5f3672f}
}