April 2026Unreviewed
Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs
Krishiv Agarwal, Ramneet Kaur, Colin Samplawski, Manoj Acharya, Anirban Roy, Daniel Elenius, Brian Matejek, Adam D. Cobb, Susmit Jha
Abstract
Effective safety auditing of large language models (LLMs) demands tools that go beyond black-box probing and systematically uncover vulnerabilities rooted in model internals. We present a comprehensive, interpretability-driven jailbreaking audit of eight SOTA open-source LLMs: Llama-3.1-8B, Llama-3.3-70B-4bt, GPT-oss- 20B, GPT-oss-120B, Qwen3-0.6B, Qwen3-32B, Phi4-3.8B, and Phi4-14B. Leveraging interpretability-based approaches -- Universal Steering (US) and Representation Engineering (RepE) --
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{agarwal2026breaking,
title = {{Breaking Bad: Interpretability-Based Safety Audits of State-of-the-Art LLMs}},
author = {Krishiv Agarwal and Ramneet Kaur and Colin Samplawski and Manoj Acharya and Anirban Roy and Daniel Elenius and Brian Matejek and Adam D. Cobb and Susmit Jha},
year = {2026},
month = apr,
eprint = {2604.20945},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2604.20945}
}