Skip to content
Search
paperDecember 2025Unreviewed

Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models

Kai Hu, Abhinav Aggarwal, Mehran Khodabandeh, David W. Zhang, Eric Hsin, Li Chen, Ankit Jain, Matt Fredrikson, Akash Bharadwaj

arXiv.org

Abstract

This paper introduces Jailbreak-Zero, a novel red teaming methodology that shifts the paradigm of Large Language Model (LLM) safety evaluation from a constrained example-based approach to a more expansive and effective policy-based framework. By leveraging an attack LLM to generate a high volume of diverse adversarial prompts and then fine-tuning this attack model with a preference dataset, Jailbreak-Zero achieves Pareto optimality across the crucial objectives of policy coverage, attack strateg

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{hu2025jailbreakzero,
  title = {{Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models}},
  author = {Kai Hu and Abhinav Aggarwal and Mehran Khodabandeh and David W. Zhang and Eric Hsin and Li Chen and Ankit Jain and Matt Fredrikson and Akash Bharadwaj},
  year = {2025},
  month = dec,
  eprint = {2601.03265},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2601.03265},
  url = {https://www.semanticscholar.org/paper/0e54275afd64916c2a2137b9a81a8402c9e6faff}
}