Skip to content
Search
paperAugust 2026Unreviewed

Circuit Discovery Helps Detect LLM Jailbreaking: A Mechanistic Interpretability Study

Paria Mehrbod, Boris Knyazev, Guy Wolf, Eugene Belilovsky, Geraldin Nanfack

Abstract

Despite extensive safety alignment, large language models (LLMs) remain vulnerable to jailbreak attacks that bypass safeguards to elicit harmful content. While prior work attributes this vulnerability to safety training limitations, the internal mechanisms by which LLMs process adversarial prompts remain poorly understood. We present a mechanistic analysis of the jailbreaking behavior in a large-scale, safety-aligned LLM, focusing on LLaMA-2-7B-chat-hf. Leveraging edge attribution patching and s

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{mehrbod2026circuit,
  title = {{Circuit Discovery Helps Detect LLM Jailbreaking: A Mechanistic Interpretability Study}},
  author = {Paria Mehrbod and Boris Knyazev and Guy Wolf and Eugene Belilovsky and Geraldin Nanfack},
  year = {2026},
  month = aug,
  eprint = {2608.27504},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2608.27504}
}