Skip to content
Search
paperSeptember 2026Unreviewed

The Safeguard Worked. Is the LLM System Safer?

Pingyu Wu, Weiming Zhang, Nenghai Yu

Abstract

Safeguards in deployed LLM services are evaluated by refusal, attack success, and policy violation rates. Those rates characterize how a control performed on the requests it was tested on. A deployment has to answer a different question: how much help with harmful tasks the service still gives an attacker who keeps adapting or finds another way in. We determine what each reported result implies for that question, allowing results from different safeguard families to be compared under one deploym

Categories

Cite

@misc{wu2026safeguard,
  title = {{The Safeguard Worked. Is the LLM System Safer?}},
  author = {Pingyu Wu and Weiming Zhang and Nenghai Yu},
  year = {2026},
  month = sep,
  eprint = {2609.00519},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.00519}
}