Skip to content
Search
paperAugust 2026UnreviewedOpen access

RefusalGuard-M: a scalable human–machine framework for multi-turn LLM jailbreak evaluation via semantic refusal manifold modeling

Michael Tchuindjang, Nathan Duran, Phil Legg, Faiza Medjek

Cybersecurity

Abstract

Existing multi-turn jailbreak evaluation methods increasingly rely on large language models (LLMs) as automated judges to reduce the cost and scalability limitations of human assessment. However, recent studies show that LLM-based evaluators can diverge from human judgments under adversarial strategies involving subtle linguistic and semantic variations, raising reliability concerns in safety-critical domains such as cybersecurity. To address this challenge, we propose Refusal Manifold Guard (Re

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@article{tchuindjang2026refusalguardm,
  title = {{RefusalGuard-M: a scalable human–machine framework for multi-turn LLM jailbreak evaluation via semantic refusal manifold modeling}},
  author = {Michael Tchuindjang and Nathan Duran and Phil Legg and Faiza Medjek},
  year = {2026},
  month = aug,
  journal = {Cybersecurity},
  doi = {10.1186/s42400-026-00633-z},
  url = {https://www.semanticscholar.org/paper/da9ee63b62208fdbb2199fda4de57efe4a4a4ee7}
}