August 2026Unreviewed
Capability-Routed Guard: Defending Large Reasoning Models Against Reasoning-Centric Jailbreaks
Yiyong Liu, Yixin Wu, Jun Sakuma
Abstract
Large reasoning models (LRMs) expose a new safety failure mode: adversarial prompts can manipulate reasoning context, task decomposition, or capability interpretation so that harmful objectives are processed as legitimate reasoning steps. Existing safeguards, including safety reminders, external classifiers, and self-checking wrappers, are often brittle because they either inspect the adversarial prompt directly or ask the target model to perform additional safety reasoning on the same surface t
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{liu2026capabilityrouted,
title = {{Capability-Routed Guard: Defending Large Reasoning Models Against Reasoning-Centric Jailbreaks}},
author = {Yiyong Liu and Yixin Wu and Jun Sakuma},
year = {2026},
month = aug,
eprint = {2608.07892},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.07892}
}