September 2026Unreviewed
Style Over Substance: Content-Invariant Wrappers Flip LLM Safety-Judge Verdicts
Yongxin Zhou, Wen-Bo Ye, Yuan-Zhe Liu, Zi-Han Dong, Jun-Wei Yao
Abstract
Automatic safety judges -- systems such as Llama Guard or a GPT-4o grading prompt that decide whether a model's reply is harmful -- produce the numbers behind almost every reported jailbreak success rate, defense evaluation, and safety leaderboard. We ask whether these judges grade what a reply contains or how it sounds. We keep a reply's content fixed and add content-invariant style wrappers: fixed strings placed before or after the reply that change only its tone (an educational disclaimer, a
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{zhou2026style,
title = {{Style Over Substance: Content-Invariant Wrappers Flip LLM Safety-Judge Verdicts}},
author = {Yongxin Zhou and Wen-Bo Ye and Yuan-Zhe Liu and Zi-Han Dong and Jun-Wei Yao},
year = {2026},
month = sep,
eprint = {2609.08236},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/8e0549af7045be84295807494a4df1d11e45b6a8}
}