July 2026Unreviewed
Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs
Anupam Wagle, Ifrat Ikhtear Uddin, Chaowei Zhang, Longwei Wang
Abstract
Large language models (LLMs) exhibit remarkable capabilities but remain highly vulnerable to adversarial prompts and jailbreak attacks. Existing approaches primarily analyze these failures through input-output behaviors or attribution methods, offering limited insight into how adversarial perturbations alter the model's internal reasoning. Consequently, the mechanisms underlying unsafe or incorrect behaviors remain poorly understood. We introduce a mechanistic framework for diagnosing LLM vulner
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0043Craft Adversarial Data
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{wagle2026mechanistic,
title = {{Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs}},
author = {Anupam Wagle and Ifrat Ikhtear Uddin and Chaowei Zhang and Longwei Wang},
year = {2026},
month = jul,
eprint = {2607.07903},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.07903}
}