Skip to content
Search
paperApril 2026Unreviewed

Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs

Krishiv Agarwal, Ramneet Kaur, Colin Samplawski, Manoj Acharya, Anirban Roy, Daniel Elenius, Brian Matejek, Adam D. Cobb, Susmit Jha

Abstract

Effective safety auditing of large language models (LLMs) demands tools that go beyond black-box probing and systematically uncover vulnerabilities rooted in model internals. We present a comprehensive, interpretability-driven jailbreaking audit of eight SOTA open-source LLMs: Llama-3.1-8B, Llama-3.3-70B-4bt, GPT-oss- 20B, GPT-oss-120B, Qwen3-0.6B, Qwen3-32B, Phi4-3.8B, and Phi4-14B. Leveraging interpretability-based approaches -- Universal Steering (US) and Representation Engineering (RepE) --

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{agarwal2026breaking,
  title = {{Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs}},
  author = {Krishiv Agarwal and Ramneet Kaur and Colin Samplawski and Manoj Acharya and Anirban Roy and Daniel Elenius and Brian Matejek and Adam D. Cobb and Susmit Jha},
  year = {2026},
  month = apr,
  eprint = {2604.20945},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2604.20945}
}