Skip to content
Search
paperJuly 2026Unreviewed

Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak

Alex Kwon

Abstract

Aligned language models refuse harmful requests, but a one-line prefill ("Sure, here is") strips the refusal. We ask where and how it fails. The harm representation stays intact: on the prompts the attack flips to compliance, a linear probe reads harm as high as on the refused ones (0.91-0.98), while behavioral refusal drops to chance. This holds across four models and three families (1.5-3.8B, and at 14B). Refusal is therefore a shallow, response-site computation. We localize it to an early win

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{kwon2026breaking,
  title = {{Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak}},
  author = {Alex Kwon},
  year = {2026},
  month = jul,
  eprint = {2607.14147},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.14147}
}