Skip to content
Search
paperMay 2026Unreviewed

Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation Simulation

Luoyu Chen, Weiqi Wang, Zhiyi Tian, Chenhan Zhang, Feng Wu, Jianhuan Huang, Ahmed Asiri, Shui Yu

Abstract

Jailbreak prompts can trigger harmful completions on aligned LLMs, In accordance, safety steering has been proposed: test-time activation interventions that steer jailbreak activations to trigger refusal while preserving benign utility. However, existing steering methods are fundamentally supervised and tied to a static, limited training set, whereas real jailbreaks evolve and are often out-of-distributed from the training set, leading to failures on unseen attacks. In this paper, we tackle the

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{chen2026steering,
  title = {{Steering Beyond the Support: Adversarial Training on Unsupervised Jailbroken Activation Simulation}},
  author = {Luoyu Chen and Weiqi Wang and Zhiyi Tian and Chenhan Zhang and Feng Wu and Jianhuan Huang and Ahmed Asiri and Shui Yu},
  year = {2026},
  month = may,
  eprint = {2605.24535},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2605.24535}
}