August 2026Unreviewed
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks
Mingyu Luo, Ming Deng, Zilang Qiu, Yiming Cheng, Ci Tao, Xue Tan, Sijin Sun, Yangfu Li, Ping Chen, Jun Dai, Xiaoyan Sun
Abstract
Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from benign ones. That separation is then read as evidence that the score will also catch the attacks that succeed. Harmful intent is a property of the prompt. Jailbreak success is an outcome produced later by a particular target model, decoding policy, and judge. A filter tuned on a score that measures the wrong quantity spends its false positive budget on attacks
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{luo2026measuring,
title = {{Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks}},
author = {Mingyu Luo and Ming Deng and Zilang Qiu and Yiming Cheng and Ci Tao and Xue Tan and Sijin Sun and Yangfu Li and Ping Chen and Jun Dai and Xiaoyan Sun},
year = {2026},
month = aug,
eprint = {2608.09624},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.09624}
}