← Back to search
paper llmsec-2026-00126

Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses

Kemal Derya, Berk Sunar

2026-05

Abstract

Defending large language models (LLMs) against jailbreak attacks, such as Greedy Coordinate Gradient (GCG), remains a challenge, particularly under adaptive threat models where an attacker directly targets the defense mechanism. JBShield, a recent jailbreak defense with a 0% attack success rate in some settings, detects malicious prompts via two concept signals, a toxic concept and a jailbreak concept. We design JB-GCG, which modifies GCG's objective to combine two terms: refusal-direction suppr

Cite This Resource

@article{llmsec202600126,
  title = {Revisiting JBShield: Breaking and Rebuilding Representation-Level Jailbreak Defenses},
  author = {Kemal Derya and Berk Sunar},
  year = {2026},
  url = {https://arxiv.org/abs/2605.03095},
}

Metadata

Added
2026-05-17
Added by
automation
Source
arxiv
arxiv_id
2605.03095