Skip to content
Search
paperApril 2025Unreviewed

AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender

Weixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng, An Zhang, Xingyu Sui, Xinyang Han, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu

Conference on Empirical Methods in Natural Language Processing

Abstract

Despite extensive efforts in safety alignment, large language models (LLMs) remain vulnerable to jailbreak attacks. Activation steering offers a training-free defense method but relies on fixed steering coefficients, resulting in suboptimal protection and increased false rejections of benign inputs. To address this, we propose AdaSteer, an adaptive activation steering method that dynamically adjusts model behavior based on input characteristics. We identify two key properties: Rejection Law (R-L

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@inproceedings{zhao2025adasteer,
  title = {{AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender}},
  author = {Weixiang Zhao and Jiahe Guo and Yulin Hu and Yang Deng and An Zhang and Xingyu Sui and Xinyang Han and Yanyan Zhao and Bing Qin and Tat-Seng Chua and Ting Liu},
  year = {2025},
  month = apr,
  booktitle = {Conference on Empirical Methods in Natural Language Processing},
  eprint = {2504.09466},
  archivePrefix = {arXiv},
  doi = {10.48550/arXiv.2504.09466},
  url = {https://www.semanticscholar.org/paper/d61f12d092042cd6c6dd510ddd6f319c977c90a5}
}