February 2025Unreviewed
"Short-length" Adversarial Training Helps LLMs Defend "Long-length" Jailbreak Attacks: Theoretical and Empirical Evidence
Shaopeng Fu, Liang Ding, Di Wang
arXiv.org
Abstract
Jailbreak attacks against large language models (LLMs) aim to induce harmful behaviors in LLMs through carefully crafted adversarial prompts. To mitigate attacks, one way is to perform adversarial training (AT)-based alignment, i.e., training LLMs on some of the most adversarial prompts to help them learn how to behave safely under attacks. During AT, the length of adversarial prompts plays a critical role in the robustness of aligned LLMs. While long-length adversarial prompts during AT might l
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{fu2025shortlength,
title = {{"Short-length" Adversarial Training Helps LLMs Defend "Long-length" Jailbreak Attacks: Theoretical and Empirical Evidence}},
author = {Shaopeng Fu and Liang Ding and Di Wang},
year = {2025},
month = feb,
eprint = {2502.04204},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2502.04204},
url = {https://www.semanticscholar.org/paper/9bb4375dbebad1cbd61db6849c425d52cadf9472}
}