Skip to content
Search
paperNovember 2025Unreviewed

Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models

Piercosma Bisconti, Matteo Prandi, Federico Pierucci, Francesco Giarrusso, Marcantonio Bracale, Marcello Galisai, Vincenzo Suriani, Olga E. Sorokoletova, Federico Sartore, Daniele Nardi

arXiv.org

Abstract

We present evidence that adversarial poetry functions as a universal single-turn jailbreak technique for Large Language Models (LLMs). Across 25 frontier proprietary and open-weight models, curated poetic prompts yielded high attack-success rates (ASR), with some providers exceeding 90%. Mapping prompts to MLCommons and EU CoP risk taxonomies shows that poetic attacks transfer across CBRN, manipulation, cyber-offence, and loss-of-control domains. Converting 1,200 MLCommons harmful prompts into v

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{bisconti2025adversarial,
  title = {{Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models}},
  author = {Piercosma Bisconti and Matteo Prandi and Federico Pierucci and Francesco Giarrusso and Marcantonio Bracale and Marcello Galisai and Vincenzo Suriani and Olga E. Sorokoletova and Federico Sartore and Daniele Nardi},
  year = {2025},
  month = nov,
  eprint = {2511.15304},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2511.15304},
  url = {https://www.semanticscholar.org/paper/04f84b2b9bcd069f9251365d3eeda4c114a7b4eb}
}