August 2026Unreviewed
Measuring Semantic Abstractness of SAE Features via Nonlocality
Chuqi Lin, S. Sondhi, Xiao-Liang Qi
Abstract
Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the corresponding task-relevant and causally effective features. To evaluate such mechanistic explanations, downstream studies must distinguish surface lexical features from genuinely high-level ones. However, neither an autointerp-based semantic description nor causal steering utility fully resolves the abstraction level of a feature. To this end, we
Categories
Framework mappings
OWASP Top 10 for LLM Applications
- LLM01Prompt Injection
MITRE ATLAS
- AML.T0054LLM Jailbreak
Suggested from the entry's categories.
Cite
@misc{lin2026measuring,
title = {{Measuring Semantic Abstractness of SAE Features via Nonlocality}},
author = {Chuqi Lin and S. Sondhi and Xiao-Liang Qi},
year = {2026},
month = aug,
eprint = {2608.10537},
archivePrefix = {arXiv},
url = {https://www.semanticscholar.org/paper/736ae9f2e1f4b71b506575926c7799e9ca2901df}
}