September 2026Unreviewed
An Empirical Measurement of Jailbreaking Evaluators
Yujie Mu
Abstract
Expert evaluation of jailbreak responses is costly and difficult to scale, so the community increasingly relies on automated evaluators to determine whether an attack succeeds. However, jailbreak studies typically validate their chosen evaluator independently, repeatedly spending resources on similar evaluation efforts while making results across papers difficult to compare. Different evaluators also encode different definitions of jailbreak success, meaning that reported attack strength and app
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{mu2026empirical,
title = {{An Empirical Measurement of Jailbreaking Evaluators}},
author = {Yujie Mu},
year = {2026},
month = sep,
eprint = {2609.10594},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.10594}
}