Skip to content
Search
paperNovember 2025Unreviewed

Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization

Xurui Li, Kaisong Song, Rui Zhu, Pin-Yu Chen, Haixu Tang

arXiv.org

Abstract

Large Language Models (LLMs) have developed rapidly in web services, delivering unprecedented capabilities while amplifying societal risks. Existing works tend to focus on either isolated jailbreak attacks or static defenses, neglecting the dynamic interplay between evolving threats and safeguards in real-world web contexts. To mitigate these challenges, we propose ACE-Safety (Adversarial Co-Evolution for LLM Safety), a novel framework that jointly optimize attack and defense models by seamlessl

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{li2025adversarial,
  title = {{Adversarial Attack-Defense Co-Evolution for LLM Safety Alignment via Tree-Group Dual-Aware Search and Optimization}},
  author = {Xurui Li and Kaisong Song and Rui Zhu and Pin-Yu Chen and Haixu Tang},
  year = {2025},
  month = nov,
  eprint = {2511.19218},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2511.19218},
  url = {https://www.semanticscholar.org/paper/9d0955162f6732bae135c35e00aa6ab8612c5175}
}