Skip to content
Search
paperJune 2026Unreviewed

DoubtProbe: Black-Box Jailbreak Defense via Structural Verification and Semantic Auditing

Xuanyu Yin, Yilin Jiang, Jun Zhou, Kai Chen, Zhengfu Cao, Xiaolei Dong

Abstract

As large language models (LLMs) are increasingly deployed in user-facing systems, black-box jailbreak defense has become an important practical problem. Existing defenses often rely on known-attack coverage, prompt-level semantic judgment, or local runtime control, yet these paths can become unstable under evolving prompt packaging, expression rewriting, and structure manipulation. We observe that many black-box jailbreaks do not remove the harmful goal, but reorganize the information needed to

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{yin2026doubtprobe,
  title = {{DoubtProbe: Black-Box Jailbreak Defense via Structural Verification and Semantic Auditing}},
  author = {Xuanyu Yin and Yilin Jiang and Jun Zhou and Kai Chen and Zhengfu Cao and Xiaolei Dong},
  year = {2026},
  month = jun,
  eprint = {2606.16527},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.16527}
}