Feature Nonlocality: New Metric for Measuring Semantic Abstractness in SAE Features
A recent research article published on arXiv (2608.10537) presents a new metric called Feature Nonlocality (FNL), aimed at assessing the semantic abstractness of features derived from sparse autoencoders (SAEs) utilizing large language models (LLMs). This study tackles a significant issue in mechanistic interpretability: the differentiation between superficial lexical features and truly high-level, context-sensitive ones. FNL is defined as the entropy of the normalized influence per position on the activation of an SAE feature. The findings indicate that FNL aligns with current LLM-based proxy metrics for semantic abstractness and effectively distinguishes context-dependent reasoning features from those driven by tokens, showing higher FNL for contextual features in 73–84% of randomly selected pairs. The authors, whose names are not mentioned, contribute to AI interpretability by providing a quantitative method for evaluating feature abstraction, crucial for validating mechanistic explanations of LLM behaviors like reasoning and jailbreaking.
Key facts
- Paper on arXiv: 2608.10537
- Introduces Feature Nonlocality (FNL) metric
- FNL defined as entropy of normalized per-position influence on SAE feature activation
- FNL correlates with LLM-based proxy metrics of semantic abstractness
- Successfully distinguishes context-dependent reasoning features from token-driven ones
- Correctly assigns higher FNL to contextual feature in 73–84% of randomly drawn pairs
- Aims to evaluate mechanistic explanations of LLM behaviors
- Relevant to reasoning and jailbreaking behaviors
Entities
Institutions
- arXiv