April 2025Unreviewed
AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
Weixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng, An Zhang, Xingyu Sui, Xinyang Han, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu
Conference on Empirical Methods in Natural Language Processing
Abstract
Despite extensive efforts in safety alignment, large language models (LLMs) remain vulnerable to jailbreak attacks. Activation steering offers a training-free defense method but relies on fixed steering coefficients, resulting in suboptimal protection and increased false rejections of benign inputs. To address this, we propose AdaSteer, an adaptive activation steering method that dynamically adjusts model behavior based on input characteristics. We identify two key properties: Rejection Law (R-L
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@inproceedings{zhao2025adasteer,
title = {{AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender}},
author = {Weixiang Zhao and Jiahe Guo and Yulin Hu and Yang Deng and An Zhang and Xingyu Sui and Xinyang Han and Yanyan Zhao and Bing Qin and Tat-Seng Chua and Ting Liu},
year = {2025},
month = apr,
booktitle = {Conference on Empirical Methods in Natural Language Processing},
eprint = {2504.09466},
archivePrefix = {arXiv},
doi = {10.48550/arXiv.2504.09466},
url = {https://www.semanticscholar.org/paper/d61f12d092042cd6c6dd510ddd6f319c977c90a5}
}