Skip to content
Search
paperAugust 2026Unreviewed

Measuring Semantic Abstractness of SAE Features via Nonlocality

Chuqi Lin, S. Sondhi, Xiao-Liang Qi

Abstract

Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the corresponding task-relevant and causally effective features. To evaluate such mechanistic explanations, downstream studies must distinguish surface lexical features from genuinely high-level ones. However, neither an autointerp-based semantic description nor causal steering utility fully resolves the abstraction level of a feature. To this end, we

Categories

Framework mappings

MITRE ATLAS
  • AML.T0054LLM Jailbreak

Suggested from the entry's categories.

Cite

@misc{lin2026measuring,
  title = {{Measuring Semantic Abstractness of SAE Features via Nonlocality}},
  author = {Chuqi Lin and S. Sondhi and Xiao-Liang Qi},
  year = {2026},
  month = aug,
  eprint = {2608.10537},
  archivePrefix = {arXiv},
  url = {https://www.semanticscholar.org/paper/736ae9f2e1f4b71b506575926c7799e9ca2901df}
}