Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects

· Source: Computation and Language · Field: Technology & Digital — Artificial Intelligence & Machine Learning · Depth: Expert, quick

Summary

Sparse Autoencoders (SAEs) are widely used for interpretability, but their utility is hampered by inconsistent links between features and model behavior. Features with clear descriptions often have weak or unexpected causal effects, and steering can vary or oppose intentions. To address this, researchers introduce Feature-Effect Geometry Analysis (FEGA), an unsupervised framework that removes active SAE features across contexts and analyzes the resulting cloud of logit changes. FEGA reveals that consistent one-dimensional effects are rare, meaning few features behave like reusable directions. The study distinguishes value-like features, tied to static information, from pointer-like features, associated with context-dependent operations. Value-like features exhibit structured, low-dimensional effects, while pointer-like features predominantly show diffuse effects. This indicates a feature can be interpretable and causally relevant without providing a stable direction for steering.

Key takeaway

For AI scientists and research scientists focused on model interpretability and steering, you should re-evaluate assumptions about how Sparse Autoencoder features influence model behavior. Recognize that a feature's interpretability does not automatically imply a stable, one-dimensional steering direction. Instead, consider the context-dependent nature of feature effects, distinguishing between value-like and pointer-like features to better predict and control model outputs.

Key insights

Sparse Autoencoder features' causal effects are complex and context-dependent, rarely providing stable, one-dimensional steering directions.

Principles

Method

FEGA is an unsupervised framework that removes the same active SAE feature across contexts and analyzes the resulting cloud of logit changes.

In practice

Topics

Best for: AI Scientist, Research Scientist

Related on AIssential

Open in AIssential →

Editorial summary, takeaway, and curation by AIssential. Original article published by Computation and Language.