Skip to content
Search
paperJune 2026Unreviewed

Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack

Long P. Hoang, Hai V. Le, Shaoyang Xu, Wei Lu, Wenxuan Zhang

Abstract

Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and recognize unsafe content. In this work, we reveal that this advanced safety awareness inadvertently introduces a fatal vulnerability. We introduce Posterior Attack, a single-query jailbreak that bypasses guardrails by prompting the model to generate the exact harmful response its internal classifier would normally flag as unsafe. Through extensive

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{hoang2026safety,
  title = {{Safety Paradox: How Enhanced Safety Awareness Leaves LLMs Vulnerable to Posterior Attack}},
  author = {Long P. Hoang and Hai V. Le and Shaoyang Xu and Wei Lu and Wenxuan Zhang},
  year = {2026},
  month = jun,
  eprint = {2606.05614},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.05614}
}