December 2025Unreviewed
Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
Kai Hu, Abhinav Aggarwal, Mehran Khodabandeh, David W. Zhang, Eric Hsin, Li Chen, Ankit Jain, Matt Fredrikson, Akash Bharadwaj
arXiv.org
Abstract
This paper introduces Jailbreak-Zero, a novel red teaming methodology that shifts the paradigm of Large Language Model (LLM) safety evaluation from a constrained example-based approach to a more expansive and effective policy-based framework. By leveraging an attack LLM to generate a high volume of diverse adversarial prompts and then fine-tuning this attack model with a preference dataset, Jailbreak-Zero achieves Pareto optimality across the crucial objectives of policy coverage, attack strateg
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{hu2025jailbreakzero,
title = {{Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models}},
author = {Kai Hu and Abhinav Aggarwal and Mehran Khodabandeh and David W. Zhang and Eric Hsin and Li Chen and Ankit Jain and Matt Fredrikson and Akash Bharadwaj},
year = {2025},
month = dec,
eprint = {2601.03265},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2601.03265},
url = {https://www.semanticscholar.org/paper/0e54275afd64916c2a2137b9a81a8402c9e6faff}
}