August 2026UnreviewedOpen access
RefusalGuard-M: a scalable human–machine framework for multi-turn LLM jailbreak evaluation via semantic refusal manifold modeling
Michael Tchuindjang, Nathan Duran, Phil Legg, Faiza Medjek
Cybersecurity
Abstract
Existing multi-turn jailbreak evaluation methods increasingly rely on large language models (LLMs) as automated judges to reduce the cost and scalability limitations of human assessment. However, recent studies show that LLM-based evaluators can diverge from human judgments under adversarial strategies involving subtle linguistic and semantic variations, raising reliability concerns in safety-critical domains such as cybersecurity. To address this challenge, we propose Refusal Manifold Guard (Re
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@article{tchuindjang2026refusalguardm,
title = {{RefusalGuard-M: a scalable human–machine framework for multi-turn LLM jailbreak evaluation via semantic refusal manifold modeling}},
author = {Michael Tchuindjang and Nathan Duran and Phil Legg and Faiza Medjek},
year = {2026},
month = aug,
journal = {Cybersecurity},
doi = {10.1186/s42400-026-00633-z},
url = {https://www.semanticscholar.org/paper/da9ee63b62208fdbb2199fda4de57efe4a4a4ee7}
}