Skip to content
Search
paperJune 2026Unreviewed

Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics

Hangtao Zhang, Yucheng Zhao, Sishun Liu, Ziqi Zhou, Zeyu Ye, Wei Wan, Minghui Li, Shengshan Hu, Yanjun Zhang, Yi Liu, Leo Yu Zhang

Abstract

Jailbreak prompts can bypass alignment guardrails in large language models (LLMs) and elicit unsafe outputs, making reliable deployment-time detection critical. Prior detection approaches largely rely on a fixed metric space, e.g., raw inputs, gradients, or hidden features, in which benign and jailbreak prompts are linearly separable. We show this assumption breaks under (i) pseudo-malicious prompts that are benign by intent but contain safety-related keywords, and (ii) adaptive attacks that exp

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{zhang2026defending,
  title = {{Defending Jailbreak Attacks on Large Language Models via Manifold Trajectory Kinetics}},
  author = {Hangtao Zhang and Yucheng Zhao and Sishun Liu and Ziqi Zhou and Zeyu Ye and Wei Wan and Minghui Li and Shengshan Hu and Yanjun Zhang and Yi Liu and Leo Yu Zhang},
  year = {2026},
  month = jun,
  eprint = {2606.07335},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2606.07335}
}