May 2026Unreviewed
BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning
Xuan Luo, Yue Wang, Geng Tu, Jing Li, Ruifeng Xu
Abstract
In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that approaches malicious goals through internal disclosure. BAIT first asks the model to identify the protection boundary, then requires it to refine that boundary, and finally requests a detailed example. By expanding each step upon the model's previous responses, BAIT turns the model's own reasoning and consistency tendency into a disclosure pathway. Experiments on AdvBench, JailbreakBench, AIR-Ben
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{luo2026bait,
title = {{BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning}},
author = {Xuan Luo and Yue Wang and Geng Tu and Jing Li and Ruifeng Xu},
year = {2026},
month = may,
eprint = {2605.27110},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.27110}
}