May 2026Unreviewed
Adaptive Probe-based Steering for Robust LLM Jailbreaking
Junxi Chen, Junhao Dong, Xiaohua Xie
Abstract
Recent work has demonstrated the potential of contrastive steering for jailbreaking Large Language Models (LLMs). However, existing methods rely on limited and inherently biased contrastive prompts and require laborious manual tuning of steering strength, limiting their robustness and effectiveness. In this paper, we leverage the idea of model extraction to guide the learned steering vectors to approximate the ideal one and propose tuning the steering strength adaptively based on contrastive act
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0024.002Extract AI Model
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{chen2026adaptive,
title = {{Adaptive Probe-based Steering for Robust LLM Jailbreaking}},
author = {Junxi Chen and Junhao Dong and Xiaohua Xie},
year = {2026},
month = may,
eprint = {2605.20286},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.20286}
}