July 2026Unreviewed
DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection
Weiwei Qi, Zefeng Wu, Zhilin Guo, Tianhang Zheng, Chaochao Lu, Liang He, Zhan Qin, Kui Ren
Abstract
Most existing LLM safety evaluation and defense methods follow a static formulation: jailbreak vulnerabilities are evaluated with fixed attack methods, and guardrails are trained on fixed malicious prompt datasets. However, real-world adversaries continuously evolve their capabilities and expand the attack space. To address this challenge, we propose DARWIN, an evolutionary attack-defense framework that formulates jailbreaking as an open-ended evolution process and continuously updates guardrail
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{qi2026darwin,
title = {{DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection}},
author = {Weiwei Qi and Zefeng Wu and Zhilin Guo and Tianhang Zheng and Chaochao Lu and Liang He and Zhan Qin and Kui Ren},
year = {2026},
month = jul,
eprint = {2607.19829},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.19829}
}