Skip to content
Search
paperJuly 2026Unreviewed

Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs

Anupam Wagle, Ifrat Ikhtear Uddin, Chaowei Zhang, Longwei Wang

Abstract

Large language models (LLMs) exhibit remarkable capabilities but remain highly vulnerable to adversarial prompts and jailbreak attacks. Existing approaches primarily analyze these failures through input-output behaviors or attribution methods, offering limited insight into how adversarial perturbations alter the model's internal reasoning. Consequently, the mechanisms underlying unsafe or incorrect behaviors remain poorly understood. We introduce a mechanistic framework for diagnosing LLM vulner

Categories

Framework mappings

MITRE ATLAS
  • AML.T0043Craft Adversarial Data
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{wagle2026mechanistic,
  title = {{Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs}},
  author = {Anupam Wagle and Ifrat Ikhtear Uddin and Chaowei Zhang and Longwei Wang},
  year = {2026},
  month = jul,
  eprint = {2607.07903},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2607.07903}
}