Skip to content
Search
paperAugust 2026Unreviewed

Capability-Routed Guard: Defending Large Reasoning Models Against Reasoning-Centric Jailbreaks

Yiyong Liu, Yixin Wu, Jun Sakuma

Abstract

Large reasoning models (LRMs) expose a new safety failure mode: adversarial prompts can manipulate reasoning context, task decomposition, or capability interpretation so that harmful objectives are processed as legitimate reasoning steps. Existing safeguards, including safety reminders, external classifiers, and self-checking wrappers, are often brittle because they either inspect the adversarial prompt directly or ask the target model to perform additional safety reasoning on the same surface t

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{liu2026capabilityrouted,
  title = {{Capability-Routed Guard: Defending Large Reasoning Models Against Reasoning-Centric Jailbreaks}},
  author = {Yiyong Liu and Yixin Wu and Jun Sakuma},
  year = {2026},
  month = aug,
  eprint = {2608.07892},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.07892}
}