Skip to content
Search
paperMay 2026Unreviewed

Mitigating Many-shot Jailbreak Attacks with One Single Demonstration

Kejia Chen, Jiawen Zhang, Boheng Li, Pengcheng Li, Jian Lou, Zunlei Feng, Mingli Song, Ruoxi Jia, Tianwei Zhang

Abstract

Many-shot jailbreaking (MSJ) causes safety-aligned language models to answer harmful queries by preceding them with many harmful question-answer demonstrations. We study why this attack becomes stronger as the number of demonstrations increases. Empirically, we find that MSJ induces a progressive activation drift: the representation of a fixed harmful query moves step by step away from the safety-aligned region as more harmful demonstrations are added. Theoretically, we show that this drift can

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{chen2026mitigating,
  title = {{Mitigating Many-shot Jailbreak Attacks with One Single Demonstration}},
  author = {Kejia Chen and Jiawen Zhang and Boheng Li and Pengcheng Li and Jian Lou and Zunlei Feng and Mingli Song and Ruoxi Jia and Tianwei Zhang},
  year = {2026},
  month = may,
  eprint = {2605.08277},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.08277}
}