Skip to content
Search
paperJuly 2026Unreviewed

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization

Yu Cui, Ruiqing Yue, Tingyu Li, Sicheng Pan, Zhuoyu Sun, Xufeng Zhang, Baohan Huang, Haibin Zhang, Cong Zuo

Abstract

Safety defenses for large language models (LLMs) have been extensively studied, with existing approaches focusing on attack detection and refusal mechanisms. Such fixed-form direct refusal strategies may introduce the risk of prefix injection attacks. Recent work has explored a new direction that leverages humor as an indirect refusal mechanism to mitigate over-refusal in jailbreak scenarios and reduce prefix injection risks. However, this approach implicitly assumes that humorous responses are

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{cui2026refusal,
  title = {{Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization}},
  author = {Yu Cui and Ruiqing Yue and Tingyu Li and Sicheng Pan and Zhuoyu Sun and Xufeng Zhang and Baohan Huang and Haibin Zhang and Cong Zuo},
  year = {2026},
  month = jul,
  eprint = {2607.15977},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.15977}
}