Skip to content
Search
paperApril 2026Unreviewed

Understanding and Improving Continuous Adversarial Training for LLMs via In-context Learning Theory

Shaopeng Fu, Di Wang

Abstract

Adversarial training (AT) is an effective defense for large language models (LLMs) against jailbreak attacks, but performing AT on LLMs is costly. To improve the efficiency of AT for LLMs, recent studies propose continuous AT (CAT) that searches for adversarial inputs within the continuous embedding space of LLMs during AT. While CAT has achieved empirical success, its underlying mechanism, i.e., why adversarial perturbations in the embedding space can help LLMs defend against jailbreak prompts

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{fu2026understanding,
  title = {{Understanding and Improving Continuous Adversarial Training for LLMs via In-context Learning Theory}},
  author = {Shaopeng Fu and Di Wang},
  year = {2026},
  month = apr,
  eprint = {2604.12817},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2604.12817}
}