August 2026Unreviewed
LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails
Ziyang Chen, Xing Wu, Songlin Hu
Abstract
Safety guardrails serve as the last line of defense against harmful inputs and outputs of large language models (LLMs), yet they are trained and evaluated almost exclusively on short text. We present LongGuard, a framework that evaluates, mechanistically analyzes, and mitigates long-context guardrail failure. We formulate the task as Safety Needle-in-a-Haystack (SafetyNIAH) over a 0.25k-32k length grid; across 15 mainstream guardrails, unsafe recall drops monotonically by more than 50% on averag
Categories
Cite
@misc{chen2026longguard,
title = {{LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails}},
author = {Ziyang Chen and Xing Wu and Songlin Hu},
year = {2026},
month = aug,
eprint = {2608.27580},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/03d27a4020eb7f788ff7f11c1a73615d5ad275c2}
}