Skip to content
Search
paperFebruary 2025Unreviewed

"Short-length" Adversarial Training Helps LLMs Defend "Long-length" Jailbreak Attacks: Theoretical and Empirical Evidence

Shaopeng Fu, Liang Ding, Di Wang

arXiv.org

Abstract

Jailbreak attacks against large language models (LLMs) aim to induce harmful behaviors in LLMs through carefully crafted adversarial prompts. To mitigate attacks, one way is to perform adversarial training (AT)-based alignment, i.e., training LLMs on some of the most adversarial prompts to help them learn how to behave safely under attacks. During AT, the length of adversarial prompts plays a critical role in the robustness of aligned LLMs. While long-length adversarial prompts during AT might l

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{fu2025shortlength,
  title = {{"Short-length" Adversarial Training Helps LLMs Defend "Long-length" Jailbreak Attacks: Theoretical and Empirical Evidence}},
  author = {Shaopeng Fu and Liang Ding and Di Wang},
  year = {2025},
  month = feb,
  eprint = {2502.04204},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2502.04204},
  url = {https://www.semanticscholar.org/paper/9bb4375dbebad1cbd61db6849c425d52cadf9472}
}