Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects
Summary
Sparse Autoencoders (SAEs) are widely used for interpretability, but their utility is hampered by inconsistent links between features and model behavior. Features with clear descriptions often have weak or unexpected causal effects, and steering can vary or oppose intentions. To address this, researchers introduce Feature-Effect Geometry Analysis (FEGA), an unsupervised framework that removes active SAE features across contexts and analyzes the resulting cloud of logit changes. FEGA reveals that consistent one-dimensional effects are rare, meaning few features behave like reusable directions. The study distinguishes value-like features, tied to static information, from pointer-like features, associated with context-dependent operations. Value-like features exhibit structured, low-dimensional effects, while pointer-like features predominantly show diffuse effects. This indicates a feature can be interpretable and causally relevant without providing a stable direction for steering.
Key takeaway
For AI scientists and research scientists focused on model interpretability and steering, you should re-evaluate assumptions about how Sparse Autoencoder features influence model behavior. Recognize that a feature's interpretability does not automatically imply a stable, one-dimensional steering direction. Instead, consider the context-dependent nature of feature effects, distinguishing between value-like and pointer-like features to better predict and control model outputs.
Key insights
Sparse Autoencoder features' causal effects are complex and context-dependent, rarely providing stable, one-dimensional steering directions.
Principles
- SAE features rarely provide consistent one-dimensional steering directions.
- Feature effects can be value-like (static info) or pointer-like (context-dependent).
- Interpretability and causal relevance do not guarantee stable steering.
Method
FEGA is an unsupervised framework that removes the same active SAE feature across contexts and analyzes the resulting cloud of logit changes.
In practice
- Distinguish value-like from pointer-like features in analysis.
- Analyze logit changes from feature interventions to understand effects.
Topics
- Sparse Autoencoders
- Model Interpretability
- Feature-Effect Geometry Analysis
- Feature Steering
- Logit Analysis
- Value-like Features
Best for: AI Scientist, Research Scientist
Related on AIssential
See Counsel's argued verdicts on the open AI decisions leaders are weighing →
Editorial summary, takeaway, and curation by AIssential. Original article published by Computation and Language.