April 2026Unreviewed
Understanding and Improving Continuous Adversarial Training for LLMs via In-context Learning Theory
Shaopeng Fu, Di Wang
Abstract
Adversarial training (AT) is an effective defense for large language models (LLMs) against jailbreak attacks, but performing AT on LLMs is costly. To improve the efficiency of AT for LLMs, recent studies propose continuous AT (CAT) that searches for adversarial inputs within the continuous embedding space of LLMs during AT. While CAT has achieved empirical success, its underlying mechanism, i.e., why adversarial perturbations in the embedding space can help LLMs defend against jailbreak prompts
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0043Craft Adversarial Data
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{fu2026understanding,
title = {{Understanding and Improving Continuous Adversarial Training for LLMs via In-context Learning Theory}},
author = {Shaopeng Fu and Di Wang},
year = {2026},
month = apr,
eprint = {2604.12817},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2604.12817}
}