May 2026Unreviewed
ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming
Mario Rodríguez Béjar, Francisco J. Cortés-Delgado, S. Braghin, Jose L. Hernández-Ramos
Abstract
Large language models (LLMs) remain vulnerable to jailbreak attacks that bypass safety alignment and elicit harmful responses. A growing body of work shows that contextual priming, where earlier turns covertly bias later replies, constitutes a powerful attack surface, with hand-crafted multi-turn scaffolds consistently outperforming single-turn manipulations on capable models. However, automated optimization-based red-teaming has remained largely limited to the single-turn setting, iterating ove
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
NIST AI Risk Management Framework
- MEASUREMeasure
Suggested from the entry's categories.
Cite
@misc{bejar2026contextualjailbreak,
title = {{ContextualJailbreak: Evolutionary Red-Teaming via Simulated Conversational Priming}},
author = {Mario Rodríguez Béjar and Francisco J. Cortés-Delgado and S. Braghin and Jose L. Hernández-Ramos},
year = {2026},
month = may,
eprint = {2605.02647},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.02647}
}