April 2026Unreviewed
Towards Understanding the Robustness of Sparse Autoencoders
Ahson Saiyed, Sabrina Sadiekh, Chirag Agarwal
Abstract
Large Language Models (LLMs) remain vulnerable to optimization-based jailbreak attacks that exploit internal gradient structure. While Sparse Autoencoders (SAEs) are widely used for interpretability, their robustness implications remain underexplored. We present a study of integrating pretrained SAEs into transformer residual streams at inference time, without modifying model weights or blocking gradients. Across four model families (Gemma, LLaMA, Mistral, Qwen) and two strong white-box attacks
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{saiyed2026understanding,
title = {{Towards Understanding the Robustness of Sparse Autoencoders}},
author = {Ahson Saiyed and Sabrina Sadiekh and Chirag Agarwal},
year = {2026},
month = apr,
eprint = {2604.18756},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2604.18756}
}