July 2026Unreviewed
Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak
Alex Kwon
Abstract
Aligned language models refuse harmful requests, but a one-line prefill ("Sure, here is") strips the refusal. We ask where and how it fails. The harm representation stays intact: on the prompts the attack flips to compliance, a linear probe reads harm as high as on the refused ones (0.91-0.98), while behavioral refusal drops to chance. This holds across four models and three families (1.5-3.8B, and at 14B). Refusal is therefore a shallow, response-site computation. We localize it to an early win
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{kwon2026breaking,
title = {{Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak}},
author = {Alex Kwon},
year = {2026},
month = jul,
eprint = {2607.14147},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.14147}
}