July 2025Unreviewed
CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations
Xiaohu Li, Yunfeng Ning, Zepeng Bao, Mayi Xu, Jianhao Chen, Tieyun Qian
Annual Meeting of the Association for Computational Linguistics
Abstract
Security alignment enables the Large Language Model (LLM) to gain the protection against malicious queries, but various jailbreak attack methods reveal the vulnerability of this security mechanism. Previous studies have isolated LLM jailbreak attacks and defenses. We analyze the security protection mechanism of the LLM, and propose a framework that combines attack and defense. Our method is based on the linearly separable property of LLM intermediate layer embedding, as well as the essence of ja
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0043Craft Adversarial Data
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@article{li2025cavgan,
title = {{CAVGAN: Unifying Jailbreak and Defense of LLMs via Generative Adversarial Attacks on their Internal Representations}},
author = {Xiaohu Li and Yunfeng Ning and Zepeng Bao and Mayi Xu and Jianhao Chen and Tieyun Qian},
year = {2025},
month = jul,
journal = {Annual Meeting of the Association for Computational Linguistics},
eprint = {2507.06043},
archivePrefix = {arXiv},
doi = {10.18653/v1/2025.findings-acl.346},
url = {https://www.semanticscholar.org/paper/ff5db18341d678f2a9528dcd07ab5b63c031331c}
}