August 2026Unreviewed
The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and Refusal
M. Chowdhury, Ernie Chang, Yang Li
Abstract
Large language models are trained to follow instructions while refusing harmful requests. Jailbreaks exploit this balance to elicit content a model would ordinarily reject. Roleplay jailbreaks are especially concerning: the harmful request can remain visible inside a roleplay wrapper made of a persona, scenario, and task, yet the model may comply. We use mechanistic interpretability to determine how this context reverses refusal and which elements contribute to the reversal. Across two benchmark
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{chowdhury2026safety,
title = {{The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and Refusal}},
author = {M. Chowdhury and Ernie Chang and Yang Li},
year = {2026},
month = aug,
eprint = {2608.30585},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/702d373e63d2c06fbb60a94f8ad67cab9be00732}
}