September 2026Unreviewed
SAFEGuard: Detect Optimization-Based Jailbreak Attacks Through Harmful Semantic Analysis and Fluency Measurement
Quoc Viet Vo, Trung Le, Damith C. Ranasinghe, Ehsan Abbasnejad
Abstract
Despite the significant efforts devoted to aligning large language models (LLMs) with human values and ensuring safe deployment, recent work has revealed that LLMs remain vulnerable to adversarial jailbreak attacks that can bypass safety guardrails and elicit harmful responses. Many defense methods are proposed to detect jailbreaks but they are limited in their effectiveness to counter wide-range optimization-based jailbreak mechanisms that can yield highly fluency-optimized or harmful semantic
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{vo2026safeguard,
title = {{SAFEGuard: Detect Optimization-Based Jailbreak Attacks Through Harmful Semantic Analysis and Fluency Measurement}},
author = {Quoc Viet Vo and Trung Le and Damith C. Ranasinghe and Ehsan Abbasnejad},
year = {2026},
month = sep,
eprint = {2609.05850},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2609.05850}
}