Skip to content
Search
paperMay 2026Unreviewed

Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution

Xiaozhe Zhang, Chaozhuo Li, Hui Liu, Shaocheng Yan, Bingyu Yan, Qiwei Ye, Haoliang Li

Abstract

Large language models remain vulnerable to adversarial prompts that elicit harmful outputs. Existing safety paradigms typically couple red-teaming and post-training in a closed, policy-centric loop, causing attack discovery to suffer from rapid saturation and limiting the exposure of novel failure modes, while leaving defenses inefficient, rigid, and difficult to transfer across victim models. To this end, we propose EvoSafety, an LLM safety framework built around persistent, inspectable, and re

Categories

Framework mappings

Suggested from the entry's categories.

Cite

@misc{zhang2026modelagnostic,
  title = {{Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution}},
  author = {Xiaozhe Zhang and Chaozhuo Li and Hui Liu and Shaocheng Yan and Bingyu Yan and Qiwei Ye and Haoliang Li},
  year = {2026},
  month = may,
  eprint = {2605.13411},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.13411}
}