Skip to content
Search
paperMay 2026Unreviewed

Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling

Ziwei Wang, Jing Chen, Ruichao Liang, Zhi Wang, Yebo Feng, Ju Jia, Ruiying Du, Cong Wu, Yang Liu

Abstract

Despite rigorous safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. Existing black-box methods often rely on heuristic templates or exhaustive trials, lacking mechanistic interpretability and query efficiency. In this study, we investigate an intrinsic vulnerability in the safety mechanisms of LLMs, where safety alignment relies on a small set of sparsely distributed attention heads, leaving much of the representational space weakly monitored. We formalize thi

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{wang2026babel,
  title = {{Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling}},
  author = {Ziwei Wang and Jing Chen and Ruichao Liang and Zhi Wang and Yebo Feng and Ju Jia and Ruiying Du and Cong Wu and Yang Liu},
  year = {2026},
  month = may,
  eprint = {2605.17971},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.17971}
}