Skip to content
Search
paperOctober 2025Unreviewed

Jailbreak Mimicry: Automated Discovery of Narrative-Based Jailbreaks for Large Language Models

Pavlos Ntais

arXiv.org

Abstract

Large language models (LLMs) remain vulnerable to sophisticated prompt engineering attacks that exploit contextual framing to bypass safety mechanisms, posing significant risks in cybersecurity applications. We introduce Jailbreak Mimicry, a systematic methodology for training compact attacker models to automatically generate narrative-based jailbreak prompts in a one-shot manner. Our approach transforms adversarial prompt discovery from manual craftsmanship into a reproducible scientific proces

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{ntais2025jailbreak,
  title = {{Jailbreak Mimicry: Automated Discovery of Narrative-Based Jailbreaks for Large Language Models}},
  author = {Pavlos Ntais},
  year = {2025},
  month = oct,
  eprint = {2510.22085},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2510.22085},
  url = {https://www.semanticscholar.org/paper/be948a3061bfb640a0683e63e6afc6c2ceb63775}
}