Skip to content
Search
paperJuly 2026Unreviewed

When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs

Tong Zhang, Zexin Li, Simin Chen, Yun Peng

Abstract

Jailbreak defenses are essential for protecting large language models (LLMs), but they can also introduce secondary costs that weaken model utility. We present a systematic study of these defense trade-offs along three dimensions: performance impact, over-refusal on benign inputs, and inference cost. Rather than treating defenses as a single class, we organize them by operational strategy and examine how different strategies correlate with different side-effect profiles. Across state-of-the-art

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{zhang2026whena,
  title = {{When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs}},
  author = {Tong Zhang and Zexin Li and Simin Chen and Yun Peng},
  year = {2026},
  month = jul,
  eprint = {2607.24392},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.24392}
}