May 2026Unreviewed
Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling
Ziwei Wang, Jing Chen, Ruichao Liang, Zhi Wang, Yebo Feng, Ju Jia, Ruiying Du, Cong Wu, Yang Liu
Abstract
Despite rigorous safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. Existing black-box methods often rely on heuristic templates or exhaustive trials, lacking mechanistic interpretability and query efficiency. In this study, we investigate an intrinsic vulnerability in the safety mechanisms of LLMs, where safety alignment relies on a small set of sparsely distributed attention heads, leaving much of the representational space weakly monitored. We formalize thi
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{wang2026babel,
title = {{Babel: Jailbreaking Safety Attention via Obfuscation Distribution Optimized Sampling}},
author = {Ziwei Wang and Jing Chen and Ruichao Liang and Zhi Wang and Yebo Feng and Ju Jia and Ruiying Du and Cong Wu and Yang Liu},
year = {2026},
month = may,
eprint = {2605.17971},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.17971}
}