November 2025Unreviewed
Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models
Piercosma Bisconti, Matteo Prandi, Federico Pierucci, Francesco Giarrusso, Marcantonio Bracale, Marcello Galisai, Vincenzo Suriani, Olga E. Sorokoletova, Federico Sartore, Daniele Nardi
arXiv.org
Abstract
We present evidence that adversarial poetry functions as a universal single-turn jailbreak technique for Large Language Models (LLMs). Across 25 frontier proprietary and open-weight models, curated poetic prompts yielded high attack-success rates (ASR), with some providers exceeding 90%. Mapping prompts to MLCommons and EU CoP risk taxonomies shows that poetic attacks transfer across CBRN, manipulation, cyber-offence, and loss-of-control domains. Converting 1,200 MLCommons harmful prompts into v
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{bisconti2025adversarial,
title = {{Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models}},
author = {Piercosma Bisconti and Matteo Prandi and Federico Pierucci and Francesco Giarrusso and Marcantonio Bracale and Marcello Galisai and Vincenzo Suriani and Olga E. Sorokoletova and Federico Sartore and Daniele Nardi},
year = {2025},
month = nov,
eprint = {2511.15304},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2511.15304},
url = {https://www.semanticscholar.org/paper/04f84b2b9bcd069f9251365d3eeda4c114a7b4eb}
}