ARTFEED — Contemporary Art Intelligence

Feature Nonlocality: New Metric for Measuring Semantic Abstractness in SAE Features

ai-technology · 2026-08-13

A recent research article published on arXiv (2608.10537) presents a new metric called Feature Nonlocality (FNL), aimed at assessing the semantic abstractness of features derived from sparse autoencoders (SAEs) utilizing large language models (LLMs). This study tackles a significant issue in mechanistic interpretability: the differentiation between superficial lexical features and truly high-level, context-sensitive ones. FNL is defined as the entropy of the normalized influence per position on the activation of an SAE feature. The findings indicate that FNL aligns with current LLM-based proxy metrics for semantic abstractness and effectively distinguishes context-dependent reasoning features from those driven by tokens, showing higher FNL for contextual features in 73–84% of randomly selected pairs. The authors, whose names are not mentioned, contribute to AI interpretability by providing a quantitative method for evaluating feature abstraction, crucial for validating mechanistic explanations of LLM behaviors like reasoning and jailbreaking.

Key facts

  • Paper on arXiv: 2608.10537
  • Introduces Feature Nonlocality (FNL) metric
  • FNL defined as entropy of normalized per-position influence on SAE feature activation
  • FNL correlates with LLM-based proxy metrics of semantic abstractness
  • Successfully distinguishes context-dependent reasoning features from token-driven ones
  • Correctly assigns higher FNL to contextual feature in 73–84% of randomly drawn pairs
  • Aims to evaluate mechanistic explanations of LLM behaviors
  • Relevant to reasoning and jailbreaking behaviors

Entities

Institutions

  • arXiv

Sources