Skip to content
Search
paperJune 2026Unreviewed

How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring

Yang Gao

Abstract

Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety classifier trained for the task, or a general chat model prompted to grade. The judge is rarely checked. We check it. Using 596 human-labeled completions from the HarmBench classifier validation set, we compare the two judge families against human majority votes and then attack them. The two families fail in opposite

Categories

Framework mappings

MITRE ATLAS
  • AML.T0051LLM Prompt Injection
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{gao2026how,
  title = {{How Reliable Is Your Jailbreak Judge? Calibration and Adversarial Robustness of Automated ASR Scoring}},
  author = {Yang Gao},
  year = {2026},
  month = jun,
  eprint = {2606.25487},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.25487}
}