August 2026Unreviewed
Circuit Discovery Helps Detect LLM Jailbreaking: A Mechanistic Interpretability Study
Paria Mehrbod, Boris Knyazev, Guy Wolf, Eugene Belilovsky, Geraldin Nanfack
Abstract
Despite extensive safety alignment, large language models (LLMs) remain vulnerable to jailbreak attacks that bypass safeguards to elicit harmful content. While prior work attributes this vulnerability to safety training limitations, the internal mechanisms by which LLMs process adversarial prompts remain poorly understood. We present a mechanistic analysis of the jailbreaking behavior in a large-scale, safety-aligned LLM, focusing on LLaMA-2-7B-chat-hf. Leveraging edge attribution patching and s
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{mehrbod2026circuit,
title = {{Circuit Discovery Helps Detect LLM Jailbreaking: A Mechanistic Interpretability Study}},
author = {Paria Mehrbod and Boris Knyazev and Guy Wolf and Eugene Belilovsky and Geraldin Nanfack},
year = {2026},
month = aug,
eprint = {2608.27504},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2608.27504}
}