Skip to main content
Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
Train sparse autoencoders to interpret model features.

Skill metadata

Reference: full SKILL.md

The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.

SAELens: Sparse Autoencoders for Mechanistic Interpretability

SAELens is the primary library for training and analyzing Sparse Autoencoders (SAEs) - a technique for decomposing polysemantic neural network activations into sparse, interpretable features. Based on Anthropic’s groundbreaking research on monosemanticity. GitHub: jbloomAus/SAELens (1,100+ stars)

The Problem: Polysemanticity & Superposition

Individual neurons in neural networks are polysemantic - they activate in multiple, semantically distinct contexts. This happens because models use superposition to represent more features than they have neurons, making interpretability difficult. SAEs solve this by decomposing dense activations into sparse, monosemantic features - typically only a small number of features activate for any given input, and each feature corresponds to an interpretable concept.

When to Use SAELens

Use SAELens when you need to:
  • Discover interpretable features in model activations
  • Understand what concepts a model has learned
  • Study superposition and feature geometry
  • Perform feature-based steering or ablation
  • Analyze safety-relevant features (deception, bias, harmful content)
Consider alternatives when:
  • You need basic activation analysis → Use TransformerLens directly
  • You want causal intervention experiments → Use pyvene or TransformerLens
  • You need production steering → Consider direct activation engineering

Installation

Requirements: Python 3.10+, transformer-lens>=2.0.0

Core Concepts

What SAEs Learn

SAEs are trained to reconstruct model activations through a sparse bottleneck:
Loss Function: MSE(original, reconstructed) + L1_coefficient × L1(features)

Key Validation (Anthropic Research)

In “Towards Monosemanticity”, human evaluators found 70% of SAE features genuinely interpretable. Features discovered include:
  • DNA sequences, legal language, HTTP requests
  • Hebrew text, nutrition statements, code syntax
  • Sentiment, named entities, grammatical structures

Workflow 1: Loading and Analyzing Pre-trained SAEs

Step-by-Step

Available Pre-trained SAEs

Checklist

  • Load model with TransformerLens
  • Load matching SAE for target layer
  • Encode activations to sparse features
  • Identify top-activating features per token
  • Validate reconstruction quality

Workflow 2: Training a Custom SAE

Step-by-Step

v6 migration note: For other SAE types swap the sae= sub-config — GatedTrainingSAEConfig, TopKTrainingSAEConfig (set k directly), or JumpReLUTrainingSAEConfig (uses l0_coefficient). Legacy flat options (architecture, expansion_factor, hook_layer, activation_fn/activation_fn_kwargs, use_ghost_grads, ghost grads, b_dec/decoder init options) were removed in v6.

Key Hyperparameters

Evaluation Metrics

Checklist

  • Choose target layer and hook point
  • Set expansion factor (d_sae = 4-16× d_model)
  • Tune L1 coefficient for desired sparsity
  • Enable L1 warm-up to prevent dead features
  • Monitor metrics during training (W&B)
  • Validate L0 and CE loss recovery
  • Check dead feature ratio

Workflow 3: Feature Analysis and Steering

Analyzing Individual Features

Feature Steering

Feature Attribution

Common Issues & Solutions

All examples below use the v6 nested config: SAE-specific options go in the sae= sub-config (StandardTrainingSAEConfig / TopKTrainingSAEConfig / etc.), training knobs stay on the top-level LanguageModelSAERunnerConfig.

Issue: High dead feature ratio

Issue: Poor reconstruction (low CE recovery)

Issue: Features not interpretable

Issue: Memory errors during training

Integration with Neuronpedia

Browse pre-trained SAE features at neuronpedia.org:

Key Classes Reference

Reference Documentation

For detailed API documentation, tutorials, and advanced usage, see the references/ folder:

External Resources

Tutorials

Papers

Official Documentation

SAE Architectures