Skip to content
Search
paperAugust 2026Unreviewed

The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and Refusal

M. Chowdhury, Ernie Chang, Yang Li

Abstract

Large language models are trained to follow instructions while refusing harmful requests. Jailbreaks exploit this balance to elicit content a model would ordinarily reject. Roleplay jailbreaks are especially concerning: the harmful request can remain visible inside a roleplay wrapper made of a persona, scenario, and task, yet the model may comply. We use mechanistic interpretability to determine how this context reverses refusal and which elements contribute to the reversal. Across two benchmark

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{chowdhury2026safety,
  title = {{The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and Refusal}},
  author = {M. Chowdhury and Ernie Chang and Yang Li},
  year = {2026},
  month = aug,
  eprint = {2608.30585},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/702d373e63d2c06fbb60a94f8ad67cab9be00732}
}