June 2026Unreviewed
DoubtProbe: Black-Box Jailbreak Defense via Structural Verification and Semantic Auditing
Xuanyu Yin, Yilin Jiang, Jun Zhou, Kai Chen, Zhengfu Cao, Xiaolei Dong
Abstract
As large language models (LLMs) are increasingly deployed in user-facing systems, black-box jailbreak defense has become an important practical problem. Existing defenses often rely on known-attack coverage, prompt-level semantic judgment, or local runtime control, yet these paths can become unstable under evolving prompt packaging, expression rewriting, and structure manipulation. We observe that many black-box jailbreaks do not remove the harmful goal, but reorganize the information needed to
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{yin2026doubtprobe,
title = {{DoubtProbe: Black-Box Jailbreak Defense via Structural Verification and Semantic Auditing}},
author = {Xuanyu Yin and Yilin Jiang and Jun Zhou and Kai Chen and Zhengfu Cao and Xiaolei Dong},
year = {2026},
month = jun,
eprint = {2606.16527},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2606.16527}
}