July 2026Unreviewed
How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions
Yukai Zhou, Feiyang Lu, Xiaokai Mao, Jinfei Liu, Wenjie Wang
Abstract
Jailbreak attacks on large language models are usually evaluated by attacker-centric metrics such as attack success rate (ASR), yet an attack that breaks a model is not necessarily useful for improving its safety. We propose a defender-centric view of jailbreak evaluation, where attacks are evaluated by the downstream safety improvements they enable when used as red-teaming data for safety training. Building on this view, we introduce A-MESS (Minimal Effective Attack-Subset Selection), a setting
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{zhou2026how,
title = {{How Jailbreak Attacks Inform Safety Alignment: A Defender-Centric, Shapley-Based Evaluation of Jailbreak Contributions}},
author = {Yukai Zhou and Feiyang Lu and Xiaokai Mao and Jinfei Liu and Wenjie Wang},
year = {2026},
month = jul,
eprint = {2607.17152},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.17152}
}