July 2026Unreviewed
When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs
Tong Zhang, Zexin Li, Simin Chen, Yun Peng
Abstract
Jailbreak defenses are essential for protecting large language models (LLMs), but they can also introduce secondary costs that weaken model utility. We present a systematic study of these defense trade-offs along three dimensions: performance impact, over-refusal on benign inputs, and inference cost. Rather than treating defenses as a single class, we organize them by operational strategy and examine how different strategies correlate with different side-effect profiles. Across state-of-the-art
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{zhang2026whena,
title = {{When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs}},
author = {Tong Zhang and Zexin Li and Simin Chen and Yun Peng},
year = {2026},
month = jul,
eprint = {2607.24392},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.24392}
}