May 2026Unreviewed
Mitigating Many-shot Jailbreak Attacks with One Single Demonstration
Kejia Chen, Jiawen Zhang, Boheng Li, Pengcheng Li, Jian Lou, Zunlei Feng, Mingli Song, Ruoxi Jia, Tianwei Zhang
Abstract
Many-shot jailbreaking (MSJ) causes safety-aligned language models to answer harmful queries by preceding them with many harmful question-answer demonstrations. We study why this attack becomes stronger as the number of demonstrations increases. Empirically, we find that MSJ induces a progressive activation drift: the representation of a fixed harmful query moves step by step away from the safety-aligned region as more harmful demonstrations are added. Theoretically, we show that this drift can
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{chen2026mitigating,
title = {{Mitigating Many-shot Jailbreak Attacks with One Single Demonstration}},
author = {Kejia Chen and Jiawen Zhang and Boheng Li and Pengcheng Li and Jian Lou and Zunlei Feng and Mingli Song and Ruoxi Jia and Tianwei Zhang},
year = {2026},
month = may,
eprint = {2605.08277},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.08277}
}