May 2026Unreviewed
Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution
Xiaozhe Zhang, Chaozhuo Li, Hui Liu, Shaocheng Yan, Bingyu Yan, Qiwei Ye, Haoliang Li
Abstract
Large language models remain vulnerable to adversarial prompts that elicit harmful outputs. Existing safety paradigms typically couple red-teaming and post-training in a closed, policy-centric loop, causing attack discovery to suffer from rapid saturation and limiting the exposure of novel failure modes, while leaving defenses inefficient, rigid, and difficult to transfer across victim models. To this end, we propose EvoSafety, an LLM safety framework built around persistent, inspectable, and re
Categories
Framework mappings
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{zhang2026modelagnostic,
title = {{Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution}},
author = {Xiaozhe Zhang and Chaozhuo Li and Hui Liu and Shaocheng Yan and Bingyu Yan and Qiwei Ye and Haoliang Li},
year = {2026},
month = may,
eprint = {2605.13411},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.13411}
}