Skip to content
Search
paperAugust 2026Unreviewed

Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks

Mingyu Luo, Ming Deng, Zilang Qiu, Yiming Cheng, Ci Tao, Xue Tan, Sijin Sun, Yangfu Li, Ping Chen, Jun Dai, Xiaoyan Sun

Abstract

Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones. That separation is then read as evidence that the score will also catch the attacks that succeed. Harmful intent is a property of the prompt. Jailbreak success is an outcome produced later by a particular target model, decoding policy, and judge. A filter tuned on a score that measures the wrong quantity spends its false positive budget on attacks

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{luo2026measuring,
  title = {{Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks}},
  author = {Mingyu Luo and Ming Deng and Zilang Qiu and Yiming Cheng and Ci Tao and Xue Tan and Sijin Sun and Yangfu Li and Ping Chen and Jun Dai and Xiaoyan Sun},
  year = {2026},
  month = aug,
  eprint = {2608.09624},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.09624}
}