Skip to content
Search
paperSeptember 2026Unreviewed

SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?

Yuqiao Tan, Shizhu He, Jun Zhao, Kang Liu

Abstract

While research on recursive self-improvement (RSI) has predominantly automated model training pipelines, reliable autonomous development demands a missing pillar: post-hoc monitoring and auditing to understand what models learn and ensure safe alignment. Mechanistic interpretability tools are essential to bridge this gap, among which Sparse Autoencoders (SAEs) serve as a cornerstone by isolating interpretable features for model inspection and steering. In this paper, we introduce SAEScientist-Be

Categories

Cite

@misc{tan2026saescientistbench,
  title = {{SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?}},
  author = {Yuqiao Tan and Shizhu He and Jun Zhao and Kang Liu},
  year = {2026},
  month = sep,
  eprint = {2609.09113},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2609.09113}
}