October 2025Unreviewed
Jailbreak Mimicry: Automated Discovery of Narrative-Based Jailbreaks for Large Language Models
Pavlos Ntais
arXiv.org
Abstract
Large language models (LLMs) remain vulnerable to sophisticated prompt engineering attacks that exploit contextual framing to bypass safety mechanisms, posing significant risks in cybersecurity applications. We introduce Jailbreak Mimicry, a systematic methodology for training compact attacker models to automatically generate narrative-based jailbreak prompts in a one-shot manner. Our approach transforms adversarial prompt discovery from manual craftsmanship into a reproducible scientific proces
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{ntais2025jailbreak,
title = {{Jailbreak Mimicry: Automated Discovery of Narrative-Based Jailbreaks for Large Language Models}},
author = {Pavlos Ntais},
year = {2025},
month = oct,
eprint = {2510.22085},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2510.22085},
url = {https://www.semanticscholar.org/paper/be948a3061bfb640a0683e63e6afc6c2ceb63775}
}