Skip to content
Search
paperApril 2026Unreviewed

Towards Understanding the Robustness of Sparse Autoencoders

Ahson Saiyed, Sabrina Sadiekh, Chirag Agarwal

Abstract

Large Language Models (LLMs) remain vulnerable to optimization-based jailbreak attacks that exploit internal gradient structure. While Sparse Autoencoders (SAEs) are widely used for interpretability, their robustness implications remain underexplored. We present a study of integrating pretrained SAEs into transformer residual streams at inference time, without modifying model weights or blocking gradients. Across four model families (Gemma, LLaMA, Mistral, Qwen) and two strong white-box attacks

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{saiyed2026understanding,
  title = {{Towards Understanding the Robustness of Sparse Autoencoders}},
  author = {Ahson Saiyed and Sabrina Sadiekh and Chirag Agarwal},
  year = {2026},
  month = apr,
  eprint = {2604.18756},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2604.18756}
}