← Back to search
paper llmsec-2026-00115
Mitigating Many-shot Jailbreak Attacks with One Single Demonstration
Kejia Chen, Jiawen Zhang, Boheng Li, Pengcheng Li, Jian Lou, Zunlei Feng, Mingli Song, Ruoxi Jia, Tianwei Zhang
2026-05
Abstract
Many-shot jailbreaking (MSJ) causes safety-aligned language models to answer harmful queries by preceding them with many harmful question-answer demonstrations. We study why this attack becomes stronger as the number of demonstrations increases. Empirically, we find that MSJ induces a progressive activation drift: the representation of a fixed harmful query moves step by step away from the safety-aligned region as more harmful demonstrations are added. Theoretically, we show that this drift can
Categories
Cite This Resource
@article{llmsec202600115,
title = {Mitigating Many-shot Jailbreak Attacks with One Single Demonstration},
author = {Kejia Chen and Jiawen Zhang and Boheng Li and Pengcheng Li and Jian Lou and Zunlei Feng and Mingli Song and Ruoxi Jia and Tianwei Zhang},
year = {2026},
url = {https://arxiv.org/abs/2605.08277},
} Metadata
- Added
- 2026-05-17
- Added by
- automation
- Source
- arxiv
- arxiv_id
- 2605.08277