← Back to search
paper llmsec-2026-00115

Mitigating Many-shot Jailbreak Attacks with One Single Demonstration

Kejia Chen, Jiawen Zhang, Boheng Li, Pengcheng Li, Jian Lou, Zunlei Feng, Mingli Song, Ruoxi Jia, Tianwei Zhang

2026-05

Abstract

Many-shot jailbreaking (MSJ) causes safety-aligned language models to answer harmful queries by preceding them with many harmful question-answer demonstrations. We study why this attack becomes stronger as the number of demonstrations increases. Empirically, we find that MSJ induces a progressive activation drift: the representation of a fixed harmful query moves step by step away from the safety-aligned region as more harmful demonstrations are added. Theoretically, we show that this drift can

Categories

Cite This Resource

@article{llmsec202600115,
  title = {Mitigating Many-shot Jailbreak Attacks with One Single Demonstration},
  author = {Kejia Chen and Jiawen Zhang and Boheng Li and Pengcheng Li and Jian Lou and Zunlei Feng and Mingli Song and Ruoxi Jia and Tianwei Zhang},
  year = {2026},
  url = {https://arxiv.org/abs/2605.08277},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2605.08277