June 2026Unreviewed
How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring
Yang Gao
Abstract
Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety classifier trained for the task, or a general chat model prompted to grade. The judge is rarely checked. We check it. Using 596 human-labeled completions from the HarmBench classifier validation set, we compare the two judge families against human majority votes and then attack them. The two families fail in opposite
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0051LLM Prompt Injection
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{gao2026how,
title = {{How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring}},
author = {Yang Gao},
year = {2026},
month = jun,
eprint = {2606.25487},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.25487}
}