May 2026Unreviewed
Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses
Kemal Derya, Berk Sunar
Abstract
Defending large language models (LLMs) against jailbreak attacks, such as Greedy Coordinate Gradient (GCG), remains a challenge, particularly under adaptive threat models where an attacker directly targets the defense mechanism. JBShield, a recent jailbreak defense with a 0% attack success rate in some settings, detects malicious prompts via two concept signals, a toxic concept and a jailbreak concept. We design JB-GCG, which modifies GCG's objective to combine two terms: refusal-direction suppr
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{derya2026revisiting,
title = {{Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses}},
author = {Kemal Derya and Berk Sunar},
year = {2026},
month = may,
eprint = {2605.03095},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2605.03095}
}