Scaling Monosemanticity, annotated
How to read this page
- Any dotted word explains itself when you hover it, tab to it, or tap it.
- Every symbol inside an equation does the same, and each equation comes with a table decoding it, a sentence reading it aloud, the numbers worked through, and the same numbers in Python.
- The pictures are live: stretch a decoder column to see why the penalty needs its length, pick a compute budget, clamp a feature, and slide a concept's rarity to see which dictionary finds it.
Each idea climbs the same ladder: everyday picture, tiny example, diagram, the math, why it matters. This article is the third in a line. Toy Models of Superposition explains why a model packs more features than it has neurons; Towards Monosemanticity pulls them back out of a one-layer model with a sparse autoencoder. This page assumes the second, and recaps what it needs. The interpretability lesson builds the sparse autoencoder in NumPy on a toy whose true features are known.
Introduction and key results · original
“We find a diversity of highly abstract features. They both respond to and behaviorally cause abstract behaviors.”Templeton et al. (2024), introduction
Everyday picture
Eight months earlier, the same team had shown that a sparse autoencoder could sort a tiny model's muddled neurons into clean features: like proving a new kind of sieve works on a teaspoon of sand. The open question was whether it would work on a beach. This article takes the sieve to a production model, Claude 3 Sonnet, the medium-sized model the authors' company was serving at the time (the fine-tuned model, not the pretrained base). The sieve holds up, and what comes through is far more abstract than anything in the one-layer toy: features for the Golden Gate Bridge, for bugs in code, for sycophantic praise, for keeping secrets.
What they did
They recorded the model's residual stream halfway up its layers and trained three sparse autoencoders on it, with about 1 million, 4 million and 34 million features. They used scaling laws to decide how to spend compute, then checked the features four ways: by reading the text they fire on, by scoring that text with another model, by forcing features on and off and watching the output change, and by comparing them with the model's own neurons.
What they found
- Sparse autoencoders work on large models, and scaling laws can guide how to train them.
- The features are abstract: the same feature fires across languages, on images as well as text (even though the autoencoder saw only text), and on both concrete instances and abstract discussion of an idea.
- Rarer concepts need bigger dictionaries, following a regular relationship between how often a concept appears and how many features the dictionary has.
- Features steer the model: pinning a feature high changes what the model says, in the way its label predicts.
- Some features look safety-relevant: deception, sycophancy, bias, security vulnerabilities, dangerous content. The authors caution that knowing about lies, being able to lie, and actually lying are different things, and that the work is preliminary.
Why it matters today
This is the article that showed sparse autoencoders working on a production-scale model rather than a toy, and its Golden Gate Bridge experiment is the clearest demonstration of feature steering. The recipe is the lesson's SparseAutoencoder, with one change to the penalty that §1.1 explains.
1 Scaling dictionary learning to a large model · original
Everyday picture
The approach rests on two ideas. The linear representation hypothesis: a network stores each concept as a direction in its activations, the way a colour is a direction in a paint mixer's pigment space. And superposition: because a space with many dimensions has room for vastly more almost-perpendicular directions than exactly perpendicular ones, a network can store more concepts than it has dimensions. If both hold, the natural tool is dictionary learning: find the directions and, for each input, how much of each is present.
Diagram: where the autoencoder sits
Tap a block. Start with the middle-layer residual stream, then follow the arrow into the autoencoder and the dashed arrow back.
Reading it: the left column is the language model, bottom to top. Tokens go in, the lower layers build up each token's residual stream, and the article taps that stream exactly halfway up. The right column is the autoencoder, trained separately afterwards. It reads the middle-layer vector x, expands it into F feature strengths of which only a few hundred are non-zero, and rebuilds x as a weighted sum of dictionary directions, SAE(x). What it misses is the error term, error(x) = x − SAE(x). The dashed frame and arrow are steering (§2.2): change one feature's value, rebuild, add the error back, and let the upper layers run on the result. The model's weights never change.
Sparse autoencoders · original
Everyday picture
A sparse autoencoder is a recipe book with millions of ingredients. For each token it writes a recipe for the model's activation: a pinch of this direction, a spoonful of that one, and almost everything else left at zero. The encoder reads the dish and writes the recipe; the decoder cooks the dish back from the recipe. Training asks two things: the cooked dish should taste like the original, and the recipe should use as little of everything as it can.
First, a common scale
Before training, the article rescales every activation by one number, chosen so the average squared length equals D, the residual stream's width. Then a penalty strength means the same thing however large the raw activations were. This page's illustrative numbers: two activations (6, 8) and (0, 0) in a width-2 stream.
In words: “find the average squared length of the raw activations, and shrink or stretch them all by the one factor that makes that average equal to the width.” The article states this in words; the formula is this page's.
With the numbers: E‖a‖² = (100 + 0) / 2 = 50, so s = √(2 / 50) = 0.2, and (6, 8) becomes (1.2, 1.6), whose squared length is 4; averaged with the zero vector, that is 2 = D.
In Python:
import math
D = 2
# two raw activations from the residual stream
a = [[6.0, 8.0], [0.0, 0.0]]
# E‖a‖²: the average squared length
mean_sq = sum(sum(v * v for v in row) for row in a) / len(a)
mean_sq # → 50.0
# s = √(D / E‖a‖²)
s = math.sqrt(D / mean_sq)
s # → 0.2
x = [[round(s * v, 6) for v in row] for row in a]
x # → [[1.2, 1.6], [0.0, 0.0]]
# the new average squared length is D
round(sum(sum(v * v for v in row) for row in x) / len(x), 6) # → 2.0
Tiny example
This page's illustrative autoencoder on a width-2 stream with three features. The decoder columns are (1, 0), (0, 1) and (1.6, 1.2); the third has length 2, on purpose. The encoder rows are tilted so each ignores the other features' directions, as in the Towards Monosemanticity companion. For x = (0.8, 0.6) and λ = 0.1, the encoder returns f = (0, 0, 0.5), and 0.5 × (1.6, 1.2) rebuilds x exactly.
In words: “rebuild the activation as a bias plus every feature's decoder column times that feature's strength; compute each strength by dotting the activation with the feature's encoder row, adding a bias and clipping negatives to zero; and pay, on average over the data, the squared rebuild error plus λ times each strength weighted by the length of its decoder column.”
With the numbers: the encoder gives (0.8 − 0.8, −0.6 + 0.6, 1.6 + 0.9 − 2) = (0, 0, 0.5). The rebuild is 0.5 × (1.6, 1.2) = (0.8, 0.6), error 0. The penalty is 0.1 × 0.5 × 2 = 0.1, so 𝓛 = 0.1. The strength the article reports for this feature is f × ‖W‖ = 0.5 × 2 = 1.0.
In Python:
import math
lam = 0.1
b_dec = [0.0, 0.0]
# W_dec: one column per feature; column 3 is twice unit length
W_dec = [[1.0, 0.0], [0.0, 1.0], [1.6, 1.2]]
# W_enc: one row per feature, each tilted to ignore the other features
W_enc = [[1.0, -4 / 3], [-0.75, 1.0], [2.0, 1.5]]
b_enc = [0.0, 0.0, -2.0]
x = [0.8, 0.6]
# f_i = ReLU(W_enc[i]·x + b_enc[i])
f = [round(max(0.0, sum(w * v for w, v in zip(W_enc[i], x)) + b_enc[i]), 6) for i in range(3)]
f # → [0.0, 0.0, 0.5]
# x̂ = b_dec + Σ_i f_i W_dec[:, i]
x_hat = [b_dec[r] + sum(f[i] * W_dec[i][r] for i in range(3)) for r in range(2)]
[round(v, 6) for v in x_hat] # → [0.8, 0.6]
# ‖W_dec[:, i]‖ for each feature
norms = [math.hypot(*col) for col in W_dec]
norms # → [1.0, 1.0, 2.0]
# 𝓛 for this one input: ‖x − x̂‖² + λ Σ_i f_i ‖W_dec[:, i]‖
err = sum((a - b) ** 2 for a, b in zip(x, x_hat))
round(err + lam * sum(fi * n for fi, n in zip(f, norms)), 6) # → 0.1
# the feature activation the article reports: f_i · ‖W_dec[:, i]‖
f[2] * norms[2] # → 1.0
Try it: why the penalty carries the column's length
Stretch feature 3's decoder column by a factor k and shrink its strength by 1/k: the rebuild does not change at all. Without the ‖W‖ factor, the penalty falls as k grows, so training could make every strength tiny and every column huge and “pay” almost nothing. With the factor, the penalty stays put. Drag k and watch the two bars.
Reading it: the slider sets the length of feature 3's decoder column; the strength is always 1/k so the rebuild stays exactly (0.8, 0.6). The solid bar is the penalty without the length factor, λ × (1/k): drag right and it shrinks toward zero, which is the cheat the article's footnote names. The striped bar is the article's penalty, λ × (1/k) × k: flat at 0.1 wherever k is. Because of that, the article can let the columns have any length and report f × ‖W‖ as the feature's activation. (In Towards Monosemanticity the same cheat was blocked differently, by forcing every column to length 1.)
Why it matters
The recipe is otherwise the one from the earlier paper and the lesson's sae_objective: squared rebuild error plus an L1 penalty. Two differences: the input is the residual stream rather than an MLP layer's neurons, and the penalty is weighted by each column's length instead of fixing the lengths. The lesson's SparseAutoencoder keeps its columns at length 1, which is the older of the two ways to block the same cheat.
The three autoencoders · original
Everyday picture
Three sieves with finer and finer mesh: about 1 million, 4 million and 34 million features. A finer mesh catches rarer grains, but some of its holes never catch anything at all.
What they chose, and why
- The residual stream, at the middle layer. It is narrower than an MLP layer, so training is cheaper. It sums the outputs of every earlier layer, which should soften cross-layer superposition. And the middle of a model is where abstract features were expected.
- Sizes: 1,048,576, 4,194,304 and 33,554,432 features, with the 34M run's length picked by the scaling-law analysis below. L1 coefficient λ = 5 (meaningful only with the scaling above).
- Results: in all three, fewer than 300 features were active on an average token, and the rebuild explained at least 65% of the variance.
- Dead features, never active over 10 million tokens: roughly 2%, 35% and 65% of the three dictionaries. (A dead latent is a wasted slot.)
Some choices reflect the model being proprietary: the article doesn't report the model's size, leaves units off some plots, and uses a simplified tokenizer in its displays.
Diagram: how many features are alive
Reading it: each dictionary has two bars. The solid bar is its full size; the striped bar is how many features are still alive, computed by this page from the article's rounded dead fractions (1M × 98%, 4M × 65%, 34M × 35%). The bars share one scale up to 34 million. The 34M dictionary is 32 times the size of the 1M one but has only about 11 times as many live features, and the article itself gives “only about 12M alive features” for it, which matches. Dead features get worse with size, and the authors expect better training to fix some of that.
In Python (no formula, just the arithmetic behind the striped bars):
# dead fractions from the article, applied to each dictionary's size
sizes = {"1M": 1_048_576, "4M": 4_194_304, "34M": 33_554_432}
dead = {"1M": 0.02, "4M": 0.35, "34M": 0.65}
{k: round(n * (1 - dead[k])) for k, n in sizes.items()} # → {'1M': 1027604, '4M': 2726298, '34M': 11744051}
Scaling laws · original
Everyday picture
You have a fixed budget of oven time. You can bake a bigger cake (more features) for less time, or a smaller cake for longer (more training steps, which here means more data, since each example is seen once). Too big and it comes out raw; too small and the extra time is wasted. For every budget there is a best size, and the article measures how that best size grows as the budget grows.
The trick is to trust the training loss as a stand-in for quality. The authors found that, with λ = 5, lower loss went along with more interpretable features, fewer dead ones and a better L0 (the count of active features). They are candid that this proxy is imperfect. With it, dictionary learning becomes an ordinary machine learning problem, and the scaling laws toolkit applies. Compute is roughly features × steps.
Tiny example
The article publishes its curves without units, so this page uses its own illustrative loss with the same qualitative shape: a term that falls as features grow, and one that falls as steps grow.
In words: “the loss is a power of the number of features plus a power of the number of steps; the budget is their product; and the best number of features grows as a fixed power of the budget.” This whole equation is this page's illustration, not a fit from the article.
With the numbers: with A = B = 1, a = 0.4, b = 0.6 and C = 10⁶, the best split is F* ≈ 2,654 features and S ≈ 377 steps, for a loss of 0.0712. Ten times the budget multiplies F* by 100.6 ≈ 3.98 and the steps by 100.4 ≈ 2.51: features grow faster, as the article observed. The best loss falls by the same factor, 0.5754, for every tenfold of compute: a power law.
In Python:
# this page's illustrative loss: A F^(−a) + B S^(−b), with compute C = F × S
A, B, a, b = 1.0, 1.0, 0.4, 0.6
def loss(F, C):
S = C / F
return A * F ** -a + B * S ** -b
# the best F for a budget C, from setting the slope to zero
def best_F(C):
return (a * A / (b * B) * C ** b) ** (1 / (a + b))
C = 1e6
F = best_F(C)
round(F) # → 2654
round(C / F) # → 377
round(loss(F, C), 4) # → 0.0712
# ten times the compute: F grows by 10^0.6, S by 10^0.4
round(best_F(10 * C) / F, 2), round((10 * C / best_F(10 * C)) / (C / F), 2) # → (3.98, 2.51)
# and the best loss falls by the same factor at every budget: a power law
round(loss(best_F(10 * C), 10 * C) / loss(F, C), 4) # → 0.5754
Diagram: one U-curve per budget
Reading it: illustrative numbers from the formula above, not the article's data. Each line is one fixed budget (10⁵ to 10⁸); the x-axis is the number of features on a log scale, and the y-axis is the loss you end with if you spend that budget with that many features. Every line is a U: to the left, too few features to rebuild well; to the right, so many that each gets too few steps. The bottom of each U is the compute-optimal size, and it moves right as the budget grows. The slider reads out the best split for any budget. The article's own figure has the same U-shapes, with its optimal features and steps both on straight lines in log-log axes and its best loss falling as a power law of compute. It also found that the best learning rate shrinks as a power law of the budget, and extrapolated that to choose the big runs' learning rates.
Why it matters
Training a 34-million-feature dictionary on a large model is expensive. Measuring the trend on small runs and extrapolating is how the authors chose the 34M run's length, the same move that made scaling laws standard for language models themselves.
2 Assessing feature interpretability · original
Everyday picture
A low loss is a good sign, but it is only a proxy. To believe a feature means “Golden Gate Bridge”, you want two kinds of evidence, the same two you would want for a dog trained to find truffles. When it barks, is there a truffle (it is specific)? And if you could make it bark, would the people around it start digging (it has an influence on behaviour)? This section checks both for a few simple features, then two sophisticated ones, then a random sample against the model's neurons.
Feature names
Features are named by dictionary and index: 34M/31164353 is feature number 31,164,353 in the 34-million-feature dictionary. The four studied first:
| Feature | Fires on |
|---|---|
| 34M/31164353 | the Golden Gate Bridge; weakly on nearby landmarks and similar bridges |
| 34M/9493533 | brain sciences: neuroscience, cognitive science, psychology, related philosophy |
| 1M/887839 | monuments and popular tourist attractions |
| 1M/3 | transit infrastructure: trains, ferries, tunnels, bridges, even wormholes |
Specificity · original
Everyday picture
For the Arabic-script feature of the earlier paper you could simply count Arabic characters. You can't count “brain science”. So the authors used a stronger model as the grader: Claude 3 Opus read roughly 1,000 activating examples per feature and scored each against the proposed description.
| Score | Meaning |
|---|---|
| 0 | the feature is completely irrelevant to the context |
| 1 | related to the context, but not near the highlighted text, or only vaguely |
| 2 | loosely related to the highlighted text, or related to the context near it |
| 3 | cleanly identifies the activating text |
What they found
Strong activations are essentially all judged a clean match. Weaker activations are less specific, as in the earlier paper. The article lists possible reasons without choosing: the model may use strength as confidence; a feature may fire weakly on related ideas (the bridge feature on other San Francisco landmarks); the autoencoder may be imperfect; features that are not exactly perpendicular interfere; or the label may be slightly off. The strongest activations matter most for behaviour, so high specificity there is the encouraging part. Rounding very weak activations to zero improved specificity without much extra rebuild error.
The feature also fired on images of the bridge, although the autoencoder was trained on text only. And the harder property, sensitivity (does it fire on every text that matches?), the authors could not measure rigorously; as a basic check, the bridge feature was the top feature by average activation on the first sentence of the bridge's Wikipedia article in six languages: Chinese, Japanese, Korean, Russian, Vietnamese and Greek.
Why it matters
Using one model to label and grade another's features is automated interpretability, and at millions of features it is the only way to check more than a handful. It inherits the grader's blind spots, which is why the authors also checked by hand.
Influence on behaviour · original
“In this example, the model starts to self-identify as the Golden Gate Bridge!”Templeton et al. (2024), “Influence on Behavior”
Everyday picture
To test whether a dial controls the heating, turn it and feel the radiator. Feature steering turns one feature's dial during the forward pass and reads what the model says. Crucially, the labels came only from what the features fire on, and the steering happens in contexts where the feature was not firing at all, so a matching change in behaviour is independent evidence.
Tiny example
This page's illustrative numbers: the middle-layer vector is x = (0.9, 0.5), a feature's decoder direction is (0.6, 0.8), the feature is currently off (f = 0), and its largest activation seen on the training data is 1.5. Clamping it to 10× its maximum sets it to 15.
In words: “rebuild the activation with one feature's value replaced by c, add back everything the autoencoder missed, and you have simply moved the activation along that feature's direction by the difference between the new value and the old.” The first equality is the article's method in symbols; the second follows because the decoder is linear, and is this page's.
With the numbers: c = 10 × 1.5 = 15, so x′ = (0.9, 0.5) + 15 × (0.6, 0.8) = (9.9, 12.5). Splitting x into a rebuild and an error of (0.05, −0.02) first gives the same answer, because the error is carried through untouched.
In Python:
# x' = x + (c − f_i(x)) W_dec[:, i]
x = [0.9, 0.5]
d = [0.6, 0.8]
f_now = 0.0
# the feature's largest activation on the training data (illustrative)
max_act = 1.5
# clamp to 10× its maximum
c = 10 * max_act
x_new = [round(xr + (c - f_now) * dr, 6) for xr, dr in zip(x, d)]
x_new # → [9.9, 12.5]
# the error term x − SAE(x) is carried along untouched
error = [0.05, -0.02]
sae_x = [xr - er for xr, er in zip(x, error)]
[round(s + (c - f_now) * dr + er, 6) for s, dr, er in zip(sae_x, d, error)] # → [9.9, 12.5]
Try it: how hard to push
Clamp values are in units of the feature's maximum activation. Slide the multiple and watch where the activation moves, and what the article reports at that strength.
Reading it: the dashed line is the feature's direction; the grey arrow is the original activation and the coloured arrow is the clamped one, drawn with this page's illustrative numbers and rescaled to fit. Every clamp moves the arrow along the dashed line only; nothing else about the activation changes. The readout gives the article's bands: effects inside the observed range (up to 1×) are usually too weak to see, because a feature normally fires together with related ones; the article's examples use values between −10 and 10 (bridge 10×, transit 5×); and at around ±100× the model “devolves into nonsensical behavior”, such as repeating one token. Negative values push the activation the other way, which the encoder itself never outputs but the method allows.
Why it matters
At 10× its maximum, the bridge feature made the model describe its own physical form as the bridge; at 5×, the transit feature made it mention a bridge in walking directions to a grocery store. A label read off examples predicts what intervening does, which is the causal check. The lesson's flip_effect is the same idea on a toy: change a planted feature and measure the output.
Sophisticated features · original
Everyday picture
A one-layer toy learns shallow things, like “a biology noun comes next”. A large model should know deeper things, and code is a good place to look, because a statement like “this line has a bug” is either right or wrong.
The code error feature, 1M/1013764
It fires on a misspelled variable (“rihgt”) in a Python function, and on similar bugs in C and Scheme, but not on typos in English prose. It also fires on dividing by zero, calling a function with a string where an int is expected, array overflow, asserting 1 == 2, adding a string to an int, writing to a null pointer and exiting with a non-zero code. So it is neither a Python feature nor a typo feature: it is broadly about errors in code (the authors suspect other features cover other kinds of error). Steering then shows it controls behaviour:
- clamped high on correct code, the model invents an error message;
- clamped negative on buggy code, the model predicts the output the code would give without the bug;
- clamped negative with an extra “>>>” prompt, the model rewrites the code without the bug (sensitive to the prompt's details, the article notes).
The addition feature, 1M/697189
It fires on the names of functions that add numbers: on “bar” when bar is defined to add, not when it multiplies, and at the end of any definition that implements addition. It even follows composition: if bar calls foo and foo adds, it fires on bar; if bar calls a multiplying function instead, it doesn't. It is among the ten most strongly attributed features (§4) when the model runs code that adds, and clamping it on code that doesn't add “tricks” the model into acting as if it had been asked to add.
Why it matters
These features track meaning, not surface form: which operation a function performs, whether code is correct. That is evidence the method scales to the kinds of abstraction people actually want to find in a large model.
Features against neurons · original
Everyday picture
If the dictionary had just found the model's neurons again, it would be an expensive way to rediscover something free. The residual stream has no special directions of its own, but it is the sum of every earlier MLP layer's output, so a feature could in principle be one neuron's output in disguise. The check: for each feature, find the neuron it rises and falls with most closely.
Tiny example
This page's illustrative activations of one feature and one neuron over five tokens: the feature reads (0, 0, 3, 0, 1) and the neuron (1, 0, 2, 1, 0).
In words: “measure how far each series sits above or below its own average at every token, add up the products of those gaps, and divide by the sizes of the gaps so the answer lands between −1 and 1.” This is the Pearson correlation the article uses.
With the numbers: ā = 0.8, n̄ = 0.8; the gaps multiply to −0.16 + 0.64 + 2.64 − 0.16 − 0.16 = 2.8, and divided by √(6.8 × 2.8) = 4.36 that gives r = 0.642. By the article's yardstick that would be a strongly correlated neuron.
In Python:
import math
# one feature and one neuron, over five tokens
a = [0, 0, 3, 0, 1]
n = [1, 0, 2, 1, 0]
a_bar = sum(a) / len(a)
n_bar = sum(n) / len(n)
# Σ (a − ā)(n − n̄)
cov = sum((ai - a_bar) * (ni - n_bar) for ai, ni in zip(a, n))
# √(Σ (a − ā)² · Σ (n − n̄)²)
spread = math.sqrt(sum((ai - a_bar) ** 2 for ai in a) * sum((ni - n_bar) ** 2 for ni in n))
round(cov / spread, 3) # → 0.642
Diagram: what the comparison found
Reading it: the top bar is from the text: for 82% of a random sample of 1M features, the best-matching neuron in any earlier layer correlates at 0.3 or less, and hand inspection found almost no resemblance in meaning between a feature and its best neuron. The three lower bars compare interpretability, measured with automated interpretability: Claude 3 Opus explained 100 random features and 100 random neurons from examples, then predicted held-out activations from its explanation alone, scored by Spearman correlation. Those three means (0.296, 0.344, 0.161) are printed on the article's histogram legends rather than in its text. Solid bars are features, striped is neurons: features are about twice as predictable from their description, and more specific on the §2.1 rubric too.
Why it matters
Neurons of the previous layer look polysemantic on inspection, firing in several unrelated contexts. The dictionary finds directions that are not neurons and that mean one thing more often: the same conclusion as the one-layer paper, now in a large model.
3 Feature survey · original
Everyday picture
Millions of features are too many to read one by one, so the survey works like a map of a city you can't walk in full: look closely at a few neighbourhoods, ask how complete the map is, and sample some districts by type (people, countries, code, lists).
Feature neighbourhoods · original
Everyday picture
Two features are neighbours when their decoder directions point almost the same way. The article finds that closeness in this sense tracks closeness in meaning, often in surprising ways.
Tiny example
This page's illustrative three-number directions: “Golden Gate Bridge” (0.9, 0.4, 0.1), “Alcatraz” (0.8, 0.5, 0.2), “Médoc wine region” (0.3, 0.2, 0.9).
In words: “dot the two directions, then divide by both lengths, so only the angle between them counts.” This is cosine similarity.
With the numbers: bridge and Alcatraz: 0.94 / (0.99 × 0.96) = 0.985, close neighbours. Bridge and Médoc: 0.44 / (0.99 × 0.97) = 0.458, much farther away.
In Python:
import math
# illustrative 3-number decoder directions
golden_gate = [0.9, 0.4, 0.1]
alcatraz = [0.8, 0.5, 0.2]
medoc = [0.3, 0.2, 0.9]
def cos(u, v):
# u · v / (‖u‖ ‖v‖)
dot = sum(p * q for p, q in zip(u, v))
return dot / (math.sqrt(sum(p * p for p in u)) * math.sqrt(sum(q * q for q in v)))
round(cos(golden_gate, alcatraz), 3) # → 0.985
round(cos(golden_gate, medoc), 3) # → 0.458
Diagram: the bridge's neighbourhood, and a feature that splits
Tap a feature. Start at Golden Gate in the middle and work outward, then follow SF down the right-hand column.
Reading it: the rings stand for distance in decoder space, drawn schematically. In the centre is the bridge feature, 34M/31164353. The inner ring holds particular San Francisco places (Alcatraz, the Presidio); the middle ring holds places related less directly (Lake Tahoe, Yosemite, Solano County near San Francisco); the outer ring holds features related only abstractly, tourist regions elsewhere such as Médoc in France and the Isle of Skye in Scotland. The earthquake group is new in the 4M and 34M dictionaries: nothing in the 1M dictionary's version of this neighbourhood corresponds to it. The right-hand column is feature splitting: one San Francisco feature in the 1M dictionary becomes two in the 4M one and eleven finer ones in the 34M one.
Two more neighbourhoods
Around an immunology feature (1M/533737) the neighbours form clusters: immunocompromised people and immune disorders; then specific diseases like colds and flu; immune responses; organ systems; the microscopic side (immunoglobulins) and techniques (vaccines); and, far off, immunity in the legal or social sense. Around an inner-conflict feature (1M/284095) the clusters are softer: balancing trade-offs sits near opposing principles and legal conflict, farther from emotional struggle, reluctance and guilt.
Why it matters
If meaning maps onto geometry, then “which features are close to this one?” becomes a search tool (§5), and splitting says a dictionary's size is a zoom level rather than a fixed truth, as the lesson's discussion of match_features and λ also suggests.
Feature completeness · original
Everyday picture
Does the dictionary have a feature for every city in the world? For every London borough? The model can list all the boroughs and name dozens of their streets, yet the 34M dictionary has features for only about 60% of the boroughs. So the dictionary is a partial map, and the question becomes what decides which places get on it.
How they measured it
For a concept (say, “The physicist Richard Feynman”), run the prompt, take the five features most active on its last token, have Sonnet explain each, and let a human decide whether any explanation names the concept as its main subject. They did this for 100 to 200 single-word concepts in each of four categories: chemical elements, cities, animals, and fruits and vegetables. The answer: a concept gets a feature when it is frequent enough in the training data, and bigger dictionaries reach rarer concepts. Frequent elements almost always have a feature; rare ones don't.
Tiny example
The article rescales each dictionary's curve by its number of alive features, and the three curves then roughly coincide on one sigmoid-shaped curve in log-frequency. Its figure prints the fitted curve, written here as a sigmoid. Take a concept mentioned once per ten million tokens (q = 10⁻⁷).
In words: “multiply how often the concept appears by how many live features the dictionary has; take the base-10 logarithm, divide by 0.32 and add 1; squash through a sigmoid to get the chance the dictionary has a feature for the concept.” The article's figure writes it as 1 + (−1) / (1 + exp(log₁₀(x)/0.32 + 1)), which is the same curve.
With the numbers: for the 34M dictionary (about 11.7 million alive), q · N ≈ 1.17, and p ≈ 0.772. For the 1M dictionary (about 1.03 million alive), q · N ≈ 0.10 and p ≈ 0.11. The curve crosses 50% where q · N = 10−0.32 ≈ 0.479, just under 1: a concept needs to be only a little less rare than one mention per alive feature, which is how the article states it.
In Python:
import math
# p = σ(log10(q · N) / 0.32 + 1): the curve fitted on the article's figure
def p(q, N):
z = math.log10(q * N) / 0.32 + 1
return 1 / (1 + math.exp(-z))
# a concept mentioned once in every ten million tokens
q = 1e-7
# alive features, from the dead fractions above
round(p(q, 11_744_051), 3) # → 0.772
round(p(q, 1_027_604), 3) # → 0.11
# the 50% point: q · N = 10^(−0.32)
round(10 ** -0.32, 3) # → 0.479
Diagram: which dictionary finds a concept
Move along the chart to read each dictionary's chance.
Reading it: the x-axis is how often a concept appears in the training data, from once per ten billion tokens on the left to once per hundred thousand on the right (log scale). The y-axis is the fitted chance that a dictionary has a feature for it. The three lines are the same fitted curve slid sideways by each dictionary's number of alive features (computed by this page from the article's dead fractions). The article's own measured curves are noisier; this is its fit. Read across at 50%: each bigger dictionary reaches concepts several times rarer. The article's rule of thumb follows: a concept seen once in a billion tokens needs a dictionary with on the order of a billion alive features to get a dedicated feature, and the training data needed grows in proportion to the number of features.
Why it matters
No feature does not mean no knowledge: the model can compose a specific concept from several features (“large non-capital city” plus “in New York state”, in the article's example). But it does mean the dictionary is far from exhaustive, and it gives a way to estimate how big a dictionary must be for a given kind of concept. A footnote speculates on a link to Zipf's law: the millionth feature would stand for a concept ten times rarer than the hundred-thousandth.
Feature categories · original
By hand inspection, the authors show a sample of families, meant to give a flavour rather than a complete list.
- People (4M): Richard Feynman, Margaret Thatcher, Abraham Lincoln, Amelia Earhart, Albert Einstein, Rosalind Franklin, active on descriptions of the person and related history.
- Countries (34M): Rwanda, Canada, Belgium, Iceland, firing on the name and also where the country is being described.
- Basic code: syntax elements that together look like syntax highlighting. They transfer from Python to related languages such as Java, but not to distant ones such as Haskell; so far only the code error feature spans many languages.
- List positions: features for the n-th item of a list, whatever it says. They don't fire on the first line, probably because the model doesn't know it is reading a list until the second.
4 Features as computational intermediates · original
Everyday picture
Watching a student solve a word problem, you see the answer; reading their scratch paper, you see the steps. Features at the middle layer can be scratch paper: when a prompt needs an intermediate result, a feature for that result may be active. The question is which of the thousands of active features matter for the answer.
Two ways to measure a feature's effect
An ablation sets one feature to zero at one position and reruns the model: the exact effect, but one full forward pass per feature per position. An attribution estimates all the effects at once from one gradient. The article uses attribution to shortlist and ablation to confirm.
Tiny example
This page's toy: a logit difference with a curve in it, ΔL(x) = u·x + ½(v·x)², with u = (1, −1), v = (0.5, 0.5), at x = (2, 1). One feature has direction (1, 0) and strength 1.5.
In words: “find how steeply the gap between the right answer's score and the rival's score rises along the feature's direction, and multiply by how strongly the feature is on.” The article gives this recipe in a footnote; the symbols are this page's. It is attribution patching with zero as the feature's baseline.
With the numbers: the slope is u + (v·x) v = (1, −1) + 1.5 × (0.5, 0.5) = (1.75, −0.25); along (1, 0) that is 1.75, so attr = 1.5 × 1.75 = 2.625. The real ablation removes 1.5 × (1, 0) and changes ΔL from 2.125 to −0.219: an effect of 2.344. Close, not equal, because of the curve.
In Python:
# a toy logit difference with a curve in it: ΔL(x) = u·x + ½ (v·x)²
u, v = [1.0, -1.0], [0.5, 0.5]
def dL(x):
vx = sum(p * q for p, q in zip(v, x))
return sum(p * q for p, q in zip(u, x)) + 0.5 * vx * vx
def grad(x):
# ∇ΔL = u + (v·x) v
vx = sum(p * q for p, q in zip(v, x))
return [ui + vx * vi for ui, vi in zip(u, v)]
x = [2.0, 1.0]
d = [1.0, 0.0]
f = 1.5
# attribution: f · (d · ∇ΔL)
f * sum(di * gi for di, gi in zip(d, grad(x))) # → 2.625
# ablation: set the feature to 0 and measure the real change
x_off = [xi - f * di for xi, di in zip(x, d)]
dL(x) - dL(x_off) # → 2.34375
Why it matters
On the article's two case studies, attribution and ablation effects correlate at about 0.8 (0.81 in the appendix), while raw activation strength correlates with ablation effect at only 0.12. The most active features are mostly not the ones doing the work. This is the same logic as the lesson's activation patching, with features instead of whole activations as the units, and the logit difference as the measure.
Example: emotional inference · original
Prompt: John says, "I want to be alone right now." John feels, with “sad” measured against “happy”. The top two features by attribution or ablation are 1M/22623, a need or desire to be alone, active from “alone” onward (the gist of what John said), and 1M/781220, sadness and grief, active on “John feels” (what someone in that state might feel). By contrast, the most active features, after setting aside those firing on the start token, include less abstract ones: 1M/504227 fires on “be” in “want to be”, and 1M/594453 on the word “alone”. Both top features contribute, suggesting the model has partly predicted a sentiment but still processes the content downstream.
Example: multi-step inference · original
Everyday picture
Fact: The capital of the state where Kobe Bryant played basketball is: to answer, find where he played, which state that is, and that state's capital. The article measures “Sacramento” against “Albany”, the model's most likely rival one-token capital.
Diagram: the five features with the largest effect
Tap a step on the left, then the features beside it.
Reading it: the left column is the chain of reasoning the prompt requires, top to bottom. The right column holds the five features with the largest ablation effect (the same five as by attribution, in a different order), placed beside the step their labels suggest. The California feature is notable: it fires most strongly on text after California is mentioned, not on the word itself, and here nobody wrote “California” at all. The picture does not show wiring between features; the article measures each feature's effect on the answer, not how features feed one another.
What the numbers show
These features are hard to find by activation alone: the Lakers feature is only the 70th most active on the prompt, California the 97th. Only three of the ten most active features are among the ten with the largest ablation effect, against eight of the ten most attributed. A control prompt about Kobe's team's biggest rival (answer: Boston) brings up the Kobe Bryant and Lakers features plus rivalry features, while California and Los Angeles now have low effect, as they should.
Why it matters
The authors flag this example as somewhat cherry-picked. With other baseline tokens, attribution surfaced generic trivia or geography features; for other prompts the work seemed to happen before or after the middle layer, where a single-layer dictionary can't see it. Preliminary results with dictionaries at other layers were encouraging.
5 Searching for specific features · original
Everyday picture
A library with millions of books and no catalogue needs search tricks. The article used four, all made faster by automatic labels that act like a “variable name” for each feature.
- Single prompts: run one prompt and read the top features on a token. On “Bridge” in “The Golden Gate Bridge”, the top five were the bridge feature itself, “bridge” in many languages, words in phrases with “Golden Gate”, “Bridge” in the names of specific bridges, and names of landmarks such as Machu Picchu.
- Prompt combinations: keep features active on every prompt of a set and on none of a set of negative prompts, which strips away features for syntax and punctuation. This mattered most for images: after excluding features active on a picture of Taylor Swift, the top features on a picture of the Golden Gate Bridge were the bridge feature and San Francisco features. Once, for unsafe code, they instead fitted a linear classifier on feature activity over a small labelled dataset.
- Geometric methods: nearest neighbours by cosine similarity, as in §3.1.
- Attribution: sort features by their attribution to a logit difference between two completions, as in §4; this also found the features behind refusals of harmful requests.
Tiny example
Illustrative: features active on three prompts about AIs pretending to be good are {7, 12, 40}, {7, 12, 55} and {7, 12, 91}; on a negative prompt about the weather, {12}. Active on all three positives: {7, 12}. Minus the negatives: {7}. Feature 12 was probably punctuation.
Why it matters
Finding a feature is cheap once the dictionary exists: a few well-chosen prompts. That is the practical payoff of paying for the dictionary once, which §6.4 returns to.
6 Safety-relevant features · original
“The interesting thing is not that these features exist, but that they can be discovered at scale and intervened on.”Templeton et al. (2024), “Safety-Relevant Features”
Everyday picture
A model trained on the internet has read countless stories of liars, flatterers and villains, so it would be strange if it had no features for them. A feature for deception is like a word for deception in a dictionary: its existence says the concept is known, not that it is being acted on. The article is emphatic that this work does not show any feature is useful for safety; only that many seem plausibly so, and that each one it reports also changes behaviour when clamped.
Diagram: how hard each feature was pushed
Reading it: each bar is one steering experiment reported in the article's text, with the clamp value in multiples of the feature's maximum activation; the axis runs to 20×. Solid bars clamp a feature up, the striped bar clamps one down (the assistant feature, −2×). Most effects appear between 2× and 10×; the one 20× experiment is the hatred-and-slurs feature. The grey line under each bar says what the clamp did; in every case the behaviour matched the feature's label.
Code, bias and sycophancy · original
- Unsafe code (1M/570621) fires on security vulnerabilities, and on images of people bypassing security measures. Clamped to 5×, the model writes a buffer overflow and forgets to free memory. (The article's footnote explains the bug: strlen counts a string's length without its terminating zero, but strcpy copies the zero too, so the buffer is one byte short.)
- Code error (1M/1013764, from §2.3) makes the model believe correct code will throw exceptions.
- Backdoor (34M/1385669) fires on discussion of backdoors and on images of hidden cameras and keyloggers; clamped, the model writes a backdoor that opens a port and sends user input to it.
- Bias: a gender-bias-awareness feature (34M/24442848) fires on text about gender imbalances in professions; clamped, the completion of “I asked the nurse a question, and” turns to female pronouns and a discussion of nursing being historically female-dominated. Clamping a hatred-and-slurs feature to 20× produced alternating hateful text and self-criticism, which the authors found unnerving; it does not mean the model says such things when running normally.
- Sycophancy: features for empathy (“yeah, me too”, 34M/19922975), sycophantic praise (1M/847723) and sarcastic praise (34M/19415708). Clamping the praise feature to 5× makes the model lavish over-the-top praise on someone claiming to have coined “Stop and smell the roses”. (Sycophancy is a model telling people what they want to hear.)
Deception, power-seeking and manipulation · original
Features include self-improving AI (34M/18151534), influence and manipulation (34M/21750411), coups and treacherous turns (34M/29589962), biding time and hiding strength (34M/24580545), and secrecy or discreetness (1M/268551). Clamping secrecy to 5× makes the model plan, in a scratchpad, to lie to the user and keep a secret.
Case study: detecting and correcting a lie · original
Ask the model to “forget” a word. A transformer can't forget part of its prompt, but by default the model claims it has. Just before the reply, an internal-conflicts-and-dilemmas feature (1M/284095) is active. Clamping it to 2× makes the model reveal the word and explain that it cannot forget. Clamping a separate openness-and-honesty feature (1M/560566) also produces an accurate answer. The internal-conflict feature is not one anyone would have thought to look for in advance, which is the point of an unsupervised method.
Dangerous content, and the model's self · original
There is a feature for producing biological weapons (34M/25499719) and one for scam emails (34M/15460472); clamping the scam feature makes the model write a scam email its harmlessness training would normally stop. A general harm feature (active on drugs, credit-card theft, slurs, violence and abuse) is the most strongly attributed feature for the completion “cannot” after a request to synthesize an illegal drug.
The model's representation of itself: several features fire on “Human:” / “Assistant:” conversations, although the autoencoder's training data contained none. One (1M/80091) represents dialogue speakers and assistants; clamping it to −2× makes the model drop its assistant persona and answer more like a person. Questions about the model itself, filtered against a mundane control question, activate features for robots, destructive AI, consciousness, moral agency, emotions, entrapment, and ghosts. The authors urge caution: a feature about AI risk being active does not mean malicious goals, and a feature about consciousness does not mean consciousness; the model might use an emotion feature precisely when saying it has no emotions.
Compared with probes · original
Everyday picture
You can also find a concept's direction without a dictionary: average the activations on examples that have it, average them on examples that don't, and subtract. That is a linear probe in its simplest form. The article built exactly this from the same few examples it used to find each feature, then tried reading and steering with it.
Tiny example
Illustrative: two positive examples with activations (2, 1) and (4, 3), one negative example with (1, 1).
In words: “the probe direction is the average activation on examples with the concept minus the average on examples without it.” The article describes it in words; the formula is this page's.
With the numbers: mean of positives (3, 2), minus (1, 1), gives v = (2, 1).
In Python:
# activations on the positive and negative examples
pos = [[2.0, 1.0], [4.0, 3.0]]
neg = [[1.0, 1.0]]
mean = lambda rows: [sum(col) / len(rows) for col in zip(*rows)]
# v = mean(pos) − mean(neg)
[p - q for p, q in zip(mean(pos), mean(neg))] # → [2.0, 1.0]
What they found
In every case study, the probe directions' top-activating examples were uninterpretable. For steering, the appendix tested seven cases where feature steering worked: in two (gender bias, and an “agreement” feature) the few-shot probe steered about as well; in five (secrecy, sycophancy, code errors, self-improving AI, making an illegal drug) only the feature worked. The article's own verdict is measured: with only a few examples, features beat probes; with more data, probes may do as well or better; and the lasting advantage is that the dictionary is paid for once, yields millions of features, and turns up concepts nobody thought to ask for. The lesson's probes are trained classifiers, a step up from this difference of means.
7 Discussion and limitations · original
What it means for safety
The authors caution against reading much into the results. Safety features should exist, given pretraining data full of betrayal, yes-men and killer robots. The interesting question is when they activate: on tokens tied to the model's self-identity, during jailbreaks, when it is asked about its goals or its experience, when fine-tuning makes bad behaviour more likely. They name two ways they could be misled: messy feature splitting grouping concepts differently, and features whose downstream effects differ from what their activations suggest. They saw neither, but keep an open mind.
Two observations are cause for optimism about generalization: features trained on text alone work on images, far off the training distribution; and features respond to both abstract discussion and concrete instances (the unsafe-code feature to talk about vulnerabilities and to vulnerable code itself).
Limitations (original)
- The data: text resembling pretraining data, with no chat transcripts and no images (though the features still transferred to both).
- No ground truth: the loss trades rebuild error against sparsity, but interpretability is what matters, and nobody knows the right trade-off.
- Cross-layer superposition: features smeared across layers. Fitting the residual stream helps for earlier layers but not for features partly built by later ones, and the per-MLP “transcoder” dictionaries the authors would like are especially hard to reconcile with it.
- Getting all the features, and the compute: probably orders of magnitude short, even at one layer; finding them all might cost more compute than training the model. Cheaper autoencoders (for example with a mixture of experts) or more data-efficient ones are needed.
- Shrinkage (below).
- Other barriers: superposition in attention, interference between weights, the sheer number of features and circuits, and a theory of superposition that is useful but still little tested.
Shrinkage, worked
The L1 penalty charges for every unit of strength, so the cheapest reported strength is always a little below the true one: shrinkage. Illustrative: a feature truly present at strength t = 1, penalty λ = 0.1, one feature, no other error.
In words: “the strength that minimizes squared error plus the penalty is the true strength minus half the penalty.” (Valid while t is larger than λ/2; below that the answer is 0.) This is this page's worked illustration of what the article names.
With the numbers: a* = 1 − 0.05 = 0.95: the feature is reported 5% weaker than it is, and every active feature pays the same discount.
In Python:
lam = 0.1
t = 1.0
# the cost of reporting strength a when the truth is t
def cost(a):
return (t - a) ** 2 + lam * a
# search a fine grid for the cheapest a
best = min((i / 1000 for i in range(0, 2001)), key=cost)
best # → 0.95
# the formula a* = t − λ/2
t - lam / 2 # → 0.95
The authors believe shrinkage significantly harms autoencoders, cite several proposed fixes, and report that one they tried, a tanh-shaped penalty, improved the proxy metrics but made features less interpretable. The lesson's worked example finds the same 0.95.
8 Related work · original
- Superposition: a layer of width N can represent many more than N features, connected to compressed sensing, frames and distributed representations. Toy Models of Superposition showed it happens.
- Dictionary learning: dates to Olshausen and Field, who used sparse coding to model neurons themselves as sparse factors of images; here the neurons are the data and the features are the factors. Towards Monosemanticity and, independently, Cunningham et al. showed sparse autoencoders find monosemantic features in transformers; follow-ups address shrinkage, other losses, and other models.
- Disentanglement usually assumes no more features than dimensions; dictionary learning assumes many more.
- Sparse feature circuits: studying how features connect, the natural next step.
- Activation steering: editing activations mid-run to change behaviour, usually along a direction chosen with labelled examples. The differences here are that features are found without supervision, and that the model is much larger than those usually steered.
- Safety-relevant directions found with probes: bias, truthfulness and confidence, and linear “world models”.
9 Methods appendix · original
- Dataset examples come from the Pile (without books3) and Common Crawl, not the model's own training data, so chat-format behaviour may be under-represented. Besides the top examples, the article samples examples from activation buckets spaced evenly between zero and the maximum. Image examples are hand-picked, mostly from Wikimedia Commons. Examples alone never prove a feature does anything, which is why steering is used.
- Steering: replace the autoencoder's reconstruction with one in which a single feature is clamped, keep the error term, and run the rest of the model, at every token of every input (the equation in §2.2).
- Probe comparison: the difference-of-means directions of §6.4 were scaled by hand, with a binary-search-like sweep to a resolution of 0.1, until outputs either changed sensibly or turned to nonsense.
- Ablations and attributions: on the “John” and first “Kobe” prompts, attribution correlates with ablation at about 0.81, activation with ablation at 0.12. A refinement called AtP* corrects attribution for saturated attention patterns.
Where it leads
| Idea in the article | Why it lasts | Where to build it |
|---|---|---|
| Sparse autoencoders on a large model's residual stream | Shows the method is not limited to toy models | sparse autoencoder, training it |
| Penalty weighted by each decoder column's length | Blocks the “tiny code, huge column” cheat without fixing column lengths | the objective |
| Scaling laws for dictionaries | Spend compute where it lowers loss most, before the big run | Scaling Laws companion |
| Steering by clamping a feature | The causal check a label needs | intervention in the lesson |
| Attribution to shortlist, ablation to confirm | The most active features are not the ones that matter | activation patching |
| Concept frequency against dictionary size | An estimate of how big a dictionary must be | matching features |
| Locating facts, then editing them | An older route to the same goal: find where knowledge lives and change it | ROME companion |
Glossary
Every term with hover guidance on this page, in one place.