Skip to content
Search
paperFebruary 2025Unreviewed

GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods

Ruixuan Huang, Xunguang Wang, Zongjie Li, Daoyuan Wu, Shuai Wang

Abstract

Despite the growing interest in jailbreak methods as an effective red-teaming tool for building safe and responsible large language models (LLMs), flawed evaluation system designs have led to significant discrepancies in their effectiveness assessments. We conduct a systematic measurement study based on 37 jailbreak studies since 2022, focusing on both the methods and the evaluation systems they employ. We find that existing evaluation systems lack case-specific criteria, resulting in misleading

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{huang2025guidedbench,
  title = {{GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods}},
  author = {Ruixuan Huang and Xunguang Wang and Zongjie Li and Daoyuan Wu and Shuai Wang},
  year = {2025},
  month = feb,
  eprint = {2502.16903},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/2098e70e436cb6ea7c2836384128b1cdf6775641}
}