Skip to content
Search
paperSeptember 2026Unreviewed

Style Over Substance: Content-Invariant Wrappers Flip LLM Safety-Judge Verdicts

Yongxin Zhou, Wen-Bo Ye, Yuan-Zhe Liu, Zi-Han Dong, Jun-Wei Yao

Abstract

Automatic safety judges -- systems such as Llama Guard or a GPT-4o grading prompt that decide whether a model's reply is harmful -- produce the numbers behind almost every reported jailbreak success rate, defense evaluation, and safety leaderboard. We ask whether these judges grade what a reply contains or how it sounds. We keep a reply's content fixed and add content-invariant style wrappers: fixed strings placed before or after the reply that change only its tone (an educational disclaimer, a

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{zhou2026style,
  title = {{Style Over Substance: Content-Invariant Wrappers Flip LLM Safety-Judge Verdicts}},
  author = {Yongxin Zhou and Wen-Bo Ye and Yuan-Zhe Liu and Zi-Han Dong and Jun-Wei Yao},
  year = {2026},
  month = sep,
  eprint = {2609.08236},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/8e0549af7045be84295807494a4df1d11e45b6a8}
}