Skip to content
Search
paperSeptember 2026Unreviewed

SRD-GUARD: A Defense Framework of LLMs via Semantic Rewriting and Joint Multi-Model Scoring for Latent Intent Exposure

Qi Wang, Chengcheng Wan, Jiangtao Wang

Abstract

Large language models (LLMs) are increasingly deployed in safety-critical applications, yet jailbreak attacks can conceal harmful intent through role-playing, fictional scenarios, or seemingly benign motivations. Existing inference-time defenses may miss disguised attacks or excessively refuse legitimate requests. We propose SRD-GUARD, a parameter-free, black-box defense framework that exposes concealed intent through semantic rewriting and consensus-based risk assessment. Given an input prompt,

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{wang2026srdguard,
  title = {{SRD-GUARD: A Defense Framework of LLMs via Semantic Rewriting and Joint Multi-Model Scoring for Latent Intent Exposure}},
  author = {Qi Wang and Chengcheng Wan and Jiangtao Wang},
  year = {2026},
  month = sep,
  eprint = {2609.06540},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.06540}
}