Towards Monosemanticity, annotated
How to read this page
- Any dotted word explains itself when you hover it, tab to it, or tap it.
- Every symbol inside an equation does the same, and each equation comes with a table decoding it, a sentence reading it aloud, the numbers worked through, and the same numbers in Python.
- The pictures are live: feed inputs to a tiny autoencoder and watch why its encoder must not simply copy its decoder, sweep the sparsity penalty, and step a chain of features through an HTML tag.
Each idea climbs the same ladder: everyday picture, tiny example, diagram, the math, why it matters. This paper is the sequel to Toy Models of Superposition; read that companion first if “superposition” is new. The interpretability lesson builds the paper's tool, a sparse autoencoder, in NumPy and runs it on a toy whose true features are known.
Introduction and results · original
“Unfortunately, the most natural computational unit of the neural network – the neuron itself – turns out not to be a natural unit for human understanding.”Bricken et al. (2023), introduction
Everyday picture
You record a choir with 512 microphones hung from the ceiling. Every microphone picks up several singers, so listening to one microphone gives you a muddle. But each singer's voice reaches the microphones in a fixed pattern, and at any moment only a few singers are singing. If you could learn each singer's pattern, you could turn 512 muddled recordings into thousands of clean voices, one per singer. That is what this paper does to one layer of a language model: the microphones are neurons, the singers are features, and the tool that learns the patterns is a sparse autoencoder.
The problem it attacks: many neurons are polysemantic. In the small language model studied here, one neuron responds to academic citations, English dialogue, HTTP requests and Korean text. The suspected cause is superposition: the model stores more features than it has neurons, each feature a direction across many neurons. Toy Models proposed three ways out, and this paper reports what happened: building models without superposition didn't give clean neurons, classic dictionary learning overfit, and a deliberately weak method, the sparse autoencoder, worked.
What they did
They trained a one-layer transformer whose MLP layer has 512 neurons, recorded its MLP activations on 8 billion tokens, and trained sparse autoencoders to rewrite each activation as a sparse sum of learned feature directions, with between 512 (1×) and 131,072 (256×) features. Most of the analysis is on one run, called A/1, with 4,096 features.
What they found
- The features are far more monosemantic than neurons, by four lines of evidence: detailed studies of a few features, blind human scoring, and two kinds of automated interpretability.
- Some features are invisible among neurons: a Hebrew feature isn't among the top examples of any neuron.
- Features steer the model: pinning the base64 feature on makes the model write base64; pinning the Arabic feature makes it write Arabic script.
- Features are roughly universal: autoencoders on two different transformers find similar features, more similar to each other than either model's neurons are.
- Features split as the dictionary grows: one base64 feature becomes three.
- 512 neurons hold tens of thousands of features: bigger dictionaries keep finding new ones.
- Features chain into systems like finite-state machines, for example to write valid HTML.
Why it matters today
This is the paper that made sparse autoencoders a standard tool of the field. The next year, Scaling Monosemanticity (2024) applied the method to a production model, and features found this way have become a common way to name, monitor and steer what models represent. The lesson's SparseAutoencoder is this paper's architecture, small enough to read in one sitting.
1 Problem setup · original
“In some sense, this is the simplest language model we profoundly don't understand.”Bricken et al. (2023), Problem Setup
Everyday picture
To understand a machine, take it apart into pieces you can understand one at a time. Without such pieces the number of possible internal states is astronomically large (the curse of dimensionality). An attention-only transformer could be analysed without looking at hidden activations at all, but as soon as a model has an MLP layer with a ReLU, you have to decompose that layer's activations. A one-layer transformer with one MLP is the smallest such model, so it is the target.
Diagram: the setup
Tap a block. Start with MLP: 512 ReLU neurons on the left, then follow the arrow into the autoencoder.
Reading it: the left column is the transformer, bottom to top: tokens are embedded as 128 numbers, one attention layer mixes information between positions, the MLP's 512 ReLU neurons do the nonlinear work, and the unembedding turns the result into next-token scores. The right column is the autoencoder, trained afterwards and separately: it reads each token's 512 MLP activations, expands them into 4,096 feature strengths that are mostly zero, and rebuilds the 512 numbers from them. The dashed arrow is training: a loss made only of the rebuild error and the total feature strength adjusts the encoder and decoder. The transformer is never changed, and nothing about its output enters the autoencoder's training, which is what later makes the model's output a fair test of the features.
Features as a decomposition · original
Everyday picture
A paint colour can be written as a recipe: so much of this pigment, so much of that. The paper writes every activation as a recipe too: a fixed base, plus each feature's direction times how strongly that feature is present. The pigments are not the neurons; they can be any directions, and there are many more of them than neurons.
Tiny example
This page's illustrative numbers: a two-neuron activation x = (0.8, 0.6), no base (b = 0), and three feature directions, (1, 0), (0, 1) and (0.8, 0.6). With feature strengths (0, 0, 1), the recipe is 1 × (0.8, 0.6) = (0.8, 0.6): one feature explains the whole activation.
In words: “each activation is roughly a base vector plus a weighted sum of feature directions; a feature's weight on this input is computed by the encoder: shift the input, multiply by the encoder's weights, add a bias, and clip negatives to zero.” Equation (1) is numbered as in the paper.
With the numbers: b + 0 × (1, 0) + 0 × (0, 1) + 1 × (0.8, 0.6) = (0.8, 0.6) = x. The other recipe, 0.8 × (1, 0) + 0.6 × (0, 1), also gives (0.8, 0.6) but uses two features; §1.5 shows how the loss picks between them.
In Python:
b = [0.0, 0.0]
d = [[1, 0], [0, 1], [0.8, 0.6]]
f = [0, 0, 1]
# x ≈ b + Σ_i f_i d_i
[round(b[r] + sum(f_i * d_i[r] for f_i, d_i in zip(f, d)), 3) for r in range(2)] # → [0.8, 0.6]
Why it matters
This is just a matrix factorization, the kind dictionary learning has always used; the paper says so. What makes it a claim about the model is the expectation from Toy Models that the directions should be overcomplete (more directions than neurons) and the strengths sparse (few non-zero per input). The paper stays agnostic about whether features are the model's true parts or a convenient description, though universality (§4.3) suggests they are more than an artefact. The lesson's compose_features builds an activation from feature amounts exactly this way.
What makes a good decomposition · original
The paper asks three things of a decomposition. You can describe when each feature fires, in a way that holds up on varied and even synthetic inputs. You can describe what each feature does downstream, consistently with the first description. And together the features account for a large share of what the layer does, measured by the loss. With that, you could explain a layer's output on an example feature by feature, monitor a model for a feature, change behaviour predictably by editing features, and show that a model has learned, and is using, some property of its data.
Why not change the architecture? · original
Everyday picture
If superposition is the problem, why not train models that can't do it, for example by forcing only one neuron to fire at a time? The authors tried, down to one-hot activations. Superposition went away, and the neurons were still polysemantic. A neuron that means “A or B” can beat a neuron that means “A”, even when there is no squeezing at all.
Tiny example
One neuron that either fires or doesn't. Four features, A, B, C, D, exactly one present per example, each predicting its own next token. Monosemantic: fire on A only and predict A; otherwise spread the prediction over B, C, D. Polysemantic: fire on A or B and spread over those two; otherwise spread over C and D.
In words: “the one-feature neuron is perfect a quarter of the time and pays ln 3 the rest of the time; the two-feature neuron always pays ln 2; ln 2 is smaller, so the model prefers the polysemantic neuron.”
With the numbers: monosemantic: ¼ × 0 + ¾ × ln 3 = 0.75 × 1.0986 = 0.824. Polysemantic: ½ × ln 2 + ½ × ln 2 = 0.693.
In Python:
import math
# cross-entropy of a uniform guess over k tokens is ln k
# mono: A (prob 1/4) is predicted exactly; otherwise uniform over B, C, D
L_mono = 1/4 * 0 + 3/4 * math.log(3)
# poly: fires on A or B (uniform over two); otherwise uniform over C, D
L_poly = 1/2 * math.log(2) + 1/2 * math.log(2)
round(L_mono, 3), round(L_poly, 3) # → (0.824, 0.693)
Why it matters
Language models are trained with cross-entropy, and under cross-entropy the authors conclude that models will generally prefer more features represented ambiguously to fewer represented cleanly, even when sparsity makes superposition impossible. Models trained on mean squared error don't necessarily have this problem, which is why the autoencoder itself uses a squared-error loss: otherwise there could be superposition “all the way down”. So the paper decomposes a normal model after training rather than redesigning models.
Why a sparse autoencoder · original
Everyday picture
Recovering thousands of feature strengths from 512 numbers is like working out a long shopping list from the total on the receipt. In general it can't be done. It becomes possible only if you know the list is short: few items, many possible items. That is compressed sensing, and doing it exactly is NP-hard (no known algorithm solves every case efficiently).
Why weak on purpose
Classic dictionary learning methods optimize each code carefully or search greedily. The authors chose something simpler for two reasons: a sparse autoencoder scales to billions of examples, and a powerful method might be too strong, recovering features from the activations that the model itself could never read. A sparse autoencoder is built like an MLP layer, one matrix and a ReLU, so it should be about as able as the model to pull features out of superposition. Its training sees only the activations, never the model's outputs, so downstream effects become an independent check.
The autoencoder, precisely · original
Everyday picture
The autoencoder has two halves with different jobs. The decoder is the dictionary: one direction per feature, each kept at length 1. The encoder is the detector: for each feature, a direction to measure along, and a threshold. The loss asks for two things: rebuild the activation well, and use as little total feature strength as possible.
In words: “subtract the typical activation; measure every feature with the encoder and clip at zero; rebuild as a weighted sum of dictionary directions and add the typical activation back; and pay for the squared rebuild error plus λ times the total feature strength, averaged over the data.” The text adds two details the equation hides: the squared error is averaged over the vector's entries while the L1 penalty is summed, and every column of Wd has length 1, so the penalty can't be dodged by shrinking codes and stretching directions.
Tiny example
This page's own illustrative autoencoder, with the dictionary from §1.1: directions (1, 0), (0, 1), (0.8, 0.6), no bias to subtract, and λ = 0.1. Its encoder is not the decoder flipped: each detector is tilted so it ignores the other features' directions (see the widget below). For x = (0.8, 0.6), the encoder returns f = (0, 0, 1), the rebuild is exact, and the loss is just the penalty, 0.1 × 1 = 0.1. The two-feature recipe (0.8, 0.6, 0) would also rebuild exactly but pay 0.1 × 1.4 = 0.14: the L1 penalty prefers the sparser explanation.
With the numbers: Wex̄ + be = (0.8 − 0.8, −0.6 + 0.6, 3.2 + 1.8 − 4) = (0, 0, 1); f = (0, 0, 1); x̂ = (0.8, 0.6); 𝓛 = 0 + 0.1 × 1 = 0.1.
In Python:
lam = 0.1
b_d = [0.0, 0.0]
# decoder W_d: one unit-length direction per feature (stored here as a list of columns)
W_d = [[1, 0], [0, 1], [0.8, 0.6]]
# encoder W_e: one detector per feature, tilted away from the other features' directions
W_e = [[1, -4/3], [-0.75, 1], [4, 3]]
b_e = [0, 0, -4]
x = [0.8, 0.6]
# x̄ = x − b_d
x_bar = [x[r] - b_d[r] for r in range(2)]
# f = ReLU(W_e x̄ + b_e)
f = [round(max(0.0, sum(w * v for w, v in zip(W_e[i], x_bar)) + b_e[i]), 6) for i in range(3)]
f # → [0.0, 0.0, 1.0]
# x̂ = W_d f + b_d
x_hat = [sum(f[i] * W_d[i][r] for i in range(3)) + b_d[r] for r in range(2)]
# 𝓛 for this one example: ‖x − x̂‖² + λ ‖f‖₁
round(sum((a - b) ** 2 for a, b in zip(x, x_hat)) + lam * sum(abs(v) for v in f), 3) # → 0.1
Try it: why the encoder isn't tied to the decoder
Many autoencoders reuse the decoder, flipped, as the encoder. The paper found that its trained encoders drift away from the decoders on purpose: similar features get detectors tilted apart so they don't set each other off. Feed the tiny autoencoder an input and compare a tied encoder (decoder flipped, no bias) with the tilted one above. Try each feature alone, then two at once.
Reading it: the sliders set the two-neuron activation x. Each group of bars shows the three feature strengths an encoder produces for it, solid for the tied encoder and striped for the tilted one, and the readout gives each encoder's rebuild and its squared error. With x = (0.8, 0.6), exactly feature 3's direction, the tied encoder lights up all three features (0.8, 0.6 and 1) because the directions overlap, and its rebuild overshoots to (1.6, 1.2). The tilted encoder answers (0, 0, 1) and rebuilds exactly. The same holds for features 1 and 2 alone. Now set x = (0.5, 0.5), features 1 and 2 on together: the tilted encoder, built to answer one feature at a time, misses most of it. That is the bargain of superposition again: detectors tuned to ignore each other's typical leaks work when few features are on at once.
Sweeping the penalty
How strong should λ be? The paper watches the trade-off between how many features fire per token (the L0 norm, which it generally keeps below 10 or 20) and how much of the model's behaviour the rebuild preserves. The same trade-off shows up on the lesson's toy, where the answer is known: five planted features in two dimensions, one on at a time in most examples.
Hover or use the arrow keys on the chart to read each penalty.
Reading it: across is the penalty λ, on a log scale; each point is one autoencoder with 5 latents, trained by the lesson's train_sae on 20,000 hidden states of the pentagon toy (each feature on 5% of the time). The solid line is the worst match between a planted feature and its nearest latent (1 means found exactly, measured by match_features); the dashed line is the share of variance the rebuild misses; the dotted line is the average number of latents active on an example with any feature on. With a tiny penalty the rebuild is perfect but about two latents fire per example and the worst match is only 0.8 to 0.9: many dictionaries rebuild a two-dimensional space, and nothing picks the true one. At λ = 0.3 about one latent fires per example, as in the data, and every planted feature is matched at 0.99, for 11% of variance left unexplained. Push harder and the penalty crushes the codes. The paper faces the same dial with no answer key, which is why its judgment of when a run is working leans on many signals at once.
Resampling dead latents
During training some latents stop firing altogether: dead latents (the paper says dead neurons). The paper revives them: at steps 25,000, 50,000, 75,000 and 100,000 it finds latents that haven't fired in the previous 12,500 steps, scores a random batch of inputs by the autoencoder's loss, and re-aims each dead latent at an input picked with probability proportional to its loss squared, so new features go where the dictionary is doing worst. The new encoder direction is scaled to 0.2 times the average live encoder length, with bias zero, so the revived latent fires only weakly at first.
In words: “pick input k with probability proportional to the square of the autoencoder's loss on it.” This page's symbols for the article's sentence.
With the numbers: three inputs with losses 0.1, 0.2 and 0.7 (illustrative): squares 0.01, 0.04, 0.49, total 0.54, so probabilities 0.019, 0.074, 0.907. Squaring makes the worst input dominate.
In Python:
L = [0.1, 0.2, 0.7]
# p_k = L_k² / Σ_j L_j²
[round(L_k ** 2 / sum(L_j ** 2 for L_j in L), 3) for L_k in L] # → [0.019, 0.074, 0.907]
Other advice from the appendix: tie the pre-encoder bias to the decoder bias and start it at the geometric median of the data; train long after the loss looks flat; prefer lower learning rates; sample activations without repeating them; and, when keeping decoder columns at unit length, remove the part of each gradient that would change their length rather than renormalizing after every step.
Why it matters
These details decide whether an autoencoder finds features or garbage, and the lesson's implementation follows the same design: a decoder bias subtracted before encoding, separate encoder and decoder weights, and unit-length decoder columns (see SparseAutoencoder.loss_and_grads, with every gradient derived by hand).
Is it working? · original
Most of machine learning has a test loss to watch. Here there is none that settles the question. The authors tried an information-based score and found it didn't track interpretability, so they steer by several signals: reading features by hand; feature density, the fraction of tokens each feature fires on (the number of live features and how rare they get); the rebuild loss; and, early on, toy models where the true features are known. Scale mattered most: training the autoencoder on more data made features sharper, and in the end it saw 8 billion activations, in batches of 8,192 over a million steps.
The one-layer model · original
The transformer has one attention block and one MLP block, a 128-number residual stream, 512 ReLU neurons, and was trained on 100 billion tokens of the Pile (a public text dataset, chosen so others can reproduce the work). One layer brings three advantages: fewer true features than a big model, so a moderate dictionary may cover them; cheap overtraining, which may make its superposition cleaner; and effects on the output that are nearly linear in the features, so they can be read off directly. A second transformer, “B”, identical except for its random seed, exists to test universality. Features are named run/number: A/1/3450 is feature 3450 of run 1 on transformer A. Runs A/0 to A/5 share one penalty and grow the dictionary.
2 Four features up close · original
“The most important claim of our paper is that dictionary learning can extract features that are significantly more monosemantic than neurons.”Bricken et al. (2023), Detailed Investigations of Individual Features
Everyday picture
Before trusting a new instrument, test it on something you can check by other means. The authors pick four features whose meaning can be checked with a simple rule: text in Arabic script, DNA sequences, base64 strings and Hebrew text. For each they ask five questions. When it fires, is the context there (specificity)? When the context is there, does it fire (sensitivity)? Does it cause the right behaviour downstream? Is it just a neuron in disguise? Does another model have it too? The features are cherry-picked for being easy to check; §3 deals with the typical one.
The proxies
For each context the authors write a scoring rule, a proxy: how much more likely a string is under the hypothesis than in ordinary text, as a log ratio. Ordinary text uses each token's overall frequency. Base64 treats every base64 character as a 1-in-64 draw and anything else as nearly impossible (10−10); DNA treats A, T, C, G as 1-in-4 draws; for Arabic and Hebrew, whether every character lies in the script's Unicode block.
In words: “score a string by how many times more likely it is as random base64 than as ordinary text, on a log scale: positive means it looks like base64.”
With the numbers: the token “Qg” (two base64 characters) with an illustrative overall frequency of one in a million: ln((1/64)²) − ln(10−6) = −8.32 + 13.82 = 5.5, strongly base64. The token “ the” (a space, then three letters) with frequency 0.03: the space isn't a base64 character, so ln(10−10 × (1/64)³) − ln 0.03 = −35.5 + 3.5 = −32.0.
In Python:
import math
import string
BASE64 = set(string.ascii_letters + string.digits + "+/")
def proxy(token, unigram_p):
# log P(s | base64): each base64 character is a 1-in-64 draw, anything else 1e-10
log_p_ctx = sum(math.log(1 / 64 if c in BASE64 else 1e-10) for c in token)
# log P(s): how common the token is overall (one token here, so one factor)
return round(log_p_ctx - math.log(unigram_p), 1)
# the unigram frequencies are illustrative
proxy("Qg", 1e-6) # → 5.5
proxy(" the", 0.03) # → -32.0
The Arabic script feature, A/1/3450 · original
Specificity
Arabic-script text is rare in the data, 0.13% of training tokens, yet it is 81% of the tokens on which this feature fires. That share rises with the activation: 25% when the feature is barely on, 98% when its activation is above 5. So, weighted by how much it fires, the feature is almost entirely about Arabic script. (It fires on the last token of characters that the tokenizer splits into two; the weak activations on other text may be proxy errors, the tiny model's own mistakes, or unrecovered features leaking in.)
In words: “how many times more common Arabic script is among the tokens where the feature fires than among tokens in general.” This page's own summary of the article's two percentages.
With the numbers: 0.81 / 0.0013 = 623: when this feature is on, Arabic script is over six hundred times more likely than usual.
In Python:
P_on = 0.81
P_base = 0.0013
# lift = P(Arabic | feature on) / P(Arabic)
round(P_on / P_base) # → 623
Sensitivity
The feature misses some Arabic tokens, such as the prefix “ال” (“al-”, the word for “the”); another Arabic feature, A/1/3134, fires exactly there, and A/1/3399 takes the first half of split characters. Several features share the script. Even so, the feature's activity correlates with the proxy at 0.74 (Pearson, over 40 million tokens).
Downstream effects
Three checks that the feature is part of the model's machinery, not just a pattern in the data. First, its logit weights: push the feature's direction through the rest of the model and see which next tokens it favours. Their distribution has a big mode at zero and a small separate mode of Arabic characters and the byte tokens that begin them. Second, ablation: set the feature to zero in context, and the model's predictions of Arabic tokens get worse. Third, pinning: set the feature to its maximum while sampling after the prompt “1,2,3,4,5,6,7,8,9,10”, and the model produces Arabic script.
In words: “take the feature's direction among the neurons, write it into the residual stream with the MLP's output weights, apply a linear stand-in for the final layer normalization (remove the mean, then rescale), and score every vocabulary token with the unembedding; then shift so the median token scores zero.”
With the numbers (illustrative sizes: 2 neurons, a 2-number residual stream, 3 tokens): d = (1, 0.5) → d Wdown = (1, 2) → minus the mean 1.5 gives (−0.5, 0.5) → scaled by 2 gives (−1, 1) → logits (−1, 1, 2) → minus the median 1 gives (−2, 0, 1): the feature favours token 3 and disfavours token 1.
In Python:
import statistics
d = [1, 0.5]
W_down = [[1, 1], [0, 2]]
W_unembed = [[1, 0, -1], [0, 1, 1]]
scale = [2, 2]
# d W_down: the feature written into the residual stream
r = [sum(d[n] * W_down[n][k] for n in range(2)) for k in range(2)]
# π: remove the mean; L: the layer norm's scaling, as a diagonal
mean = sum(r) / len(r)
r = [(v - mean) * s for v, s in zip(r, scale)]
# W_unembed: one score per token
logits = [sum(r[k] * W_unembed[k][t] for k in range(2)) for t in range(3)]
logits # → [-1.0, 1.0, 2.0]
# shift so the median logit weight is zero
[v - statistics.median(logits) for v in logits] # → [-2.0, 0.0, 1.0]
Not a neuron, and universal
Only one neuron has any Arabic among its top 20 examples, and only one such example. The feature's direction is spread over neurons: its three largest coefficients are all negative, and 27 neurons have coefficients of at least 0.1 in size. The neuron most correlated with it, A/neurons/489, fires on a mixture of non-English languages, a sensible superposition since languages rarely co-occur. On transformer B, the most correlated feature, B/1/1334, matches at 0.91 and is, if anything, cleaner.
DNA, base64 and Hebrew · original
The same five checks, repeated. The numbers below are collected from the article's text.
| Feature | Correlation with its proxy | Best match on transformer B | Closest neuron |
|---|---|---|---|
| Arabic script, A/1/3450 | 0.74 | 0.91 | a mix of non-English languages |
| DNA, A/1/2937 | 0.8 (proxy made yes/no) | 0.92 | DNA is a tiny sliver of its examples |
| base64, A/1/2357 | 0.38 | 0.85 | correlation 0.18; also code, HTML, URLs |
| Hebrew, A/1/416 | 0.55 | 0.92 | correlation 0.1; no neuron shows Hebrew in its top examples |
DNA (original): the only feature devoted to DNA. Where it and the proxy disagree, the feature is almost always right: DNA with spaces (“TGG AGT”), or the “5'-” that announces a sequence is coming. By the end of a long sequence it is the only feature active. Its top logit weights are nucleotide combinations such as AGT and GCC.
base64 (original): a feature the team had seen before as a neuron in an earlier model, a hint that it is universal. The low proxy correlation is mostly the proxy's fault: hexadecimal strings score as base64 but have a feature of their own, A/1/3817. Its logit weights shade gradually into base64 tokens, because many ordinary tokens (“fr”, say) also occur inside base64 strings.
Hebrew (original): the feature that is essentially invisible among neurons. A partner feature, A/1/1016, fires on the byte that begins most Hebrew characters and predicts the byte that completes them.
Why it matters
Each of these is an existence proof: a unit of the model's computation, specific, causal and reproducible across models, that no single neuron shows. That is what superposition predicted, and what reading neurons one by one would never find.
3 Global analysis · original
Everyday picture
Four good features could be luck. The next question is about the typical feature: pick one at random, and is it interpretable? Of A/1's 4,096 features, 168 are dead (they never fire on 100 million examples) and 292 are “ultralow density” (firing on fewer than one in a million and behaving oddly); both groups are set aside.
How interpretable is a typical feature? · original
Three ways to score
- A human, blind (original). One of the authors scored features and neurons without knowing which was which, on a rubric (below). To avoid the trap of only looking at top examples, where many polysemantic neurons look clean, examples were drawn evenly from 11 intervals across each unit's activation range. In all, 412 intervals across 162 features and neurons.
- A language model explaining activations (original). Claude writes an explanation from examples, then, seeing only the explanation, predicts activations on 60 new nine-token examples (540 predictions); the score is the Spearman correlation between predicted and true activations.
- A language model predicting effects (original). Given the explanation, decide whether a token is one the unit pushes up; half the tokens are its top logit weights and half random, so guessing scores 50%.
The rubric
The human's score adds up: confidence in the interpretation (0 to 3), how consistent the strongly activating tokens are with it (0 to 5), how consistent the pushed-up tokens are (0 to 3), a point if the inconsistent ones are clearly weaker (0 or 1), and how specific the interpretation is (0 to 3). A perfect logit score makes the extra point impossible, so the maximum is 14. Very roughly, the authors found a feature quite interpretable above 8.
Diagram: features against neurons
Reading it: each pair of bars compares A/1's features (solid) with the transformer's neurons (striped) on one measure, with numbers from the article's text; each bar's length is its share of that measure's full scale, given in the label. On the human rubric, the median feature interval scores 12 out of 14, a confident, specific interpretation consistent with what the feature pushes up; the median neuron scores 0: the annotator could not even form a hypothesis. On predicting pushed-up tokens, where chance is 50%, features average 74% against neurons' 58%. On finding the same unit in the second transformer, the median feature's best match correlates at 0.72, the median neuron's at 0.46. Every measure points the same way.
Caveats the paper raises
Most activations are small and fall in the weaker, less interpretable intervals; some features stay consistent across their whole range, others only in the top 60% or so, which may mean a learned direction sits at a slight angle to the true one. Most of a feature's effect comes from its large activations, which are the interpretable ones. The appendix adds a subtlety about automated scoring: weighting examples by how often they really occur (importance scoring) rewards explanations that say what a feature does not fire on, because features are so rare that predicting zero is almost always right.
How much of the model does this explain? · original
Everyday picture
Swap the MLP's real activations for the autoencoder's rebuild and see how much worse the model predicts text. Compare that with the damage of deleting the MLP entirely (setting its output to zero). The share of the damage the rebuild avoids is the share of the MLP's work the features capture.
In words: “of the loss the MLP saves compared with no MLP at all, how much is still saved when the MLP's activations are replaced by the autoencoder's rebuild.” This page's symbols for the article's description.
With the numbers (illustrative losses chosen to give A/1's result): model 3.00, rebuild 3.21, no MLP 4.00: (4.00 − 3.21)/(4.00 − 3.00) = 0.79.
In Python:
# illustrative losses (nats per token)
L_model, L_SAE, L_zero = 3.00, 3.21, 4.00
# (L_zero − L_SAE) / (L_zero − L_model)
round((L_zero - L_SAE) / (L_zero - L_model), 2) # → 0.79
For A/1, 79% of the MLP's log-likelihood improvement survives. More features or a smaller penalty recover more: A/5, with 131,072 features and a penalty of 0.004, recovers 94.5%. The authors ask for these numbers to be taken with a grain of salt: there is probably a long tail of features, so each extra percent costs more, the baseline of deleting the MLP is harsh (making the percentage an overestimate), and some polysemanticity may hide in weak activations. The ablation and pinning experiments are what persuade them the 79% measures something real.
Do features describe the model or the data? · original
Activations reflect both the text and what the model does with it, so interpretable features might just be patterns in the text passing through. Two answers. First, the downstream effects (logit weights, ablations, pinned sampling) are never seen by the autoencoder, so if they make sense, that is a fact about the model. Second, a control: shuffle every weight matrix of the trained transformer, keeping the same numbers but destroying their structure, and run dictionary learning on that. Its features are mostly single tokens (“span”, “file”, “.”) plus features for arbitrary mixtures of contexts that the authors couldn't interpret. Training creates richer structure than the tokens alone. (Automated scoring actually rates the random model's features higher on median, because single-token features are trivially easy to predict; set those aside and they score like polysemantic neurons.)
4 Phenomenology · original
Everyday picture
With a working microscope, the next step is to look around and note what keeps turning up. Four patterns stand out.
Feature motifs · original
Most features are context features (DNA, base64) or token-in-context features: “the” in mathematics (A/0/341), “<” in HTML (A/0/20). In A/4 there are more than a hundred features for the token “the” in different contexts. There are trigram-like features, such as one that predicts the “19” in “COVID-19”, and features for long specific phrases. In a one-layer model every feature is also an action: the base64 feature both detects base64 and makes base64 more likely next. The action view explains oddities: A/0/341 fires on “the” in maths, but also on “special” and “this”, because all of them are followed by a noun phrase it predicts, such as “denominator”.
Feature splitting · original
Everyday picture
A map at three zoom levels: at the coarsest you see “city”, zoom in and it becomes districts, zoom again and streets. Dictionaries of different sizes are zoom levels on the same model. As the dictionary grows from 512 features (A/0) to 4,096 (A/1) to 16,384 (A/2), the number of base64 features goes from 1 to 3 to many more. This is feature splitting.
Diagram: one base64 feature becomes three
Tap a feature. Start with A/0/45 at the top, then each of its three children.
Reading it: top row, the small dictionary; bottom row, the larger one; arrows show which fine features jointly cover the coarse one's activations. One feature for all base64 becomes a letters feature, a digits feature and a feature for base64 that encodes readable ASCII text. The digits feature exists because of the tokenizer: two digits in a row would have been merged into one token, so a lone digit token tells the model the next token is not a digit, and this feature's logit weights for digits are much lower. The third was a surprise nobody went looking for: the clue was the token “ICAgICAg”, which is six spaces in base64.
A theory of splitting
The authors conjecture an ideal set of features that an unlimited dictionary would find, often in clusters of similar features the model packs tightly together. A limited dictionary returns features covering roughly the same territory, less specifically. Similar features have similar directions because they cause similar behaviour: several features that fire on periods would all predict a space and a capital letter next. This explains two apparent bugs. Single-token features in small dictionaries (one fires on every “ P”) are collapsed families that split into context-specific versions when given room. Several features for one context (three for base64) are genuine splits. It also means the “right” number of features matters less than feared: small dictionaries summarize, large ones refine. The splitting isn't a clean tree: in the maths and physics example, features at one level both split and merge at the next.
Why it matters
For a big model, the recipe becomes: read a coarse dictionary to map the territory, then a fine one to study the details. The lesson warns of the same effect: a bigger dictionary can split one feature into several finer ones, so feature counts depend on the dictionary's size and penalty.
Universality · original
Everyday picture
If two people independently map the same coastline and draw the same bays, the bays are probably real. The authors compare features across the two transformers in two ways: activation similarity (do they fire on the same tokens?) and logit weight similarity (do they push up the same next tokens?).
What they found
For each feature in A/1, the best-matching feature in B/1 correlates at a median of 0.72; for neurons across the two transformers, 0.46. Logit weights agree less, because most of each feature's logit weights are small noise-like “interference” that the model doesn't need to control; only the important tokens agree. The starkest case: two PLOS ONE citation features (they fire on “pone” in citation strings and predict the “.” after it) correlate at 0.98 in activation but negatively in logit weights, with only the “.” token strongly weighted by both. So the authors combine the two views into an attribution score: a feature's activation on a token times its logit weight for the token that actually comes next.
In words: “how much feature i, at token j, pushed up the token that really came next: its activation times its logit weight for that next token.” It approximates the classic gradient-times-activation attribution, ignoring the softmax's and layer norm's denominators.
With the numbers (illustrative): the PLOS feature fires at 2.0 on “pone”, the next token is “.”, and its logit weight for “.” is 1.5: attribution 2.0 × 1.5 = 3.0. On a token where it doesn't fire, 0 × anything = 0.
In Python:
# illustrative activations and logit weights
f = {"pone": 2.0, "the": 0.0}
v = {".": 1.5, "cat": 0.4}
# a = f_i(t_j) · v_(i, t_(j+1)) for two (token, next token) pairs
[f["pone"] * v["."], f["the"] * v["cat"]] # → [3.0, 0.0]
Stacking these scores over many tokens and correlating them across models gives an attribution similarity, which tracks activation similarity closely: features that fire together across models also help predict the same tokens. The authors also find features resembling ones reported elsewhere: base64, hexadecimal and all-caps neurons from their earlier SoLU models; German and title-case detectors from other dictionary-learning work; region features (Australia, Canada, Africa) echoing “region neurons” in a vision model. Person detectors exist but are much narrower, and nothing like that model's emotion neurons turned up.
Why it matters
Universality is evidence that features aren't artefacts of one training run or of the autoencoder, and it is what lets lessons from one model carry to the next.
“Finite state automata” · original
Everyday picture
A relay race: each runner finishes their leg and hands the baton to the next, who was waiting for exactly that. Features do this through the text itself: one feature makes certain tokens likely, the model writes one of them, and that token switches on the next feature. The model never learned the chain as a unit; it emerges from patterns in the data. The simplest chain is a loop of one: a base64 feature predicts base64 tokens, which keep it on.
Diagram: writing HTML
Tap a feature, or step through the text below. Start with A/0/20 at the top left.
Reading it: each box is a feature from the small dictionary A/0, labelled with the token it fires on and what it predicts; each arrow is a token the model writes, which wakes the next feature. Step through and the loop writes the article's prototypical sample, “<div>⏎ tabs <span>”: open a tag, name it, close it, indent, open another. It is a finite-state machine drawn in features, and a crude one: this small dictionary has no state for what happens when the name is followed by an attribute such as “href”, which in A/1 is a richer system.
Other chains
Two-feature loops are common where the tokenizer splits characters in two, as with Tamil or Chinese, one feature for the first byte and one for the second (for Chinese, one feature covers both whole characters and second halves, since both can be followed by the same things). Snake-case identifiers like ARRAY_MAX_VALUE alternate a capitals feature and an underscore feature. IRC chat logs have their own chain. And a chain in a larger dictionary, A/4, reproduces the licence boilerplate “MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE”, a small mechanism for memorization found in a model with only 512 neurons.
Why it matters
Features are not just labels; they compose into behaviour. Here they compose through the generated text, a different kind of circuit, and a hint of what to look for in models trained to act in the world, whose own outputs feed back into their inputs.
5 Related work · original
Three threads. Superposition: since Toy Models, sparse probes have found likely examples in real models; the Othello exchange between Li et al. and Nanda et al. strengthened the features-as-directions view; and follow-up work found small datasets memorized in superposition, with a sharp transition to generalization. Disentanglement and architecture: disentanglement aims for as many factors as dimensions, where superposition expects more; attempts to force clean neurons, such as the SoLU activation, made some neurons cleaner and others worse. Dictionary learning on networks: sparse coding of word embeddings came first; closest in time, Sharkey, Cunningham and colleagues trained sparse autoencoders on transformers in parallel, published as interim reports and a manuscript with many of the same conclusions, and the authors credit them for focusing on the sparse autoencoder approach.
6 Discussion · original
“This work has persuaded us that our previous model was missing something crucial.”Bricken et al. (2023), Theories of Superposition
Superposition, revised
Toy Models suggested an isotropic picture: features as separate one-dimensional directions repelling each other into an even spread. The real features clump. Correlated activation is one reason; the authors suspect a bigger one is shared actions: the base64 digits feature predicts almost the same tokens as the letters feature, so their directions sit close together. Features might not even be one-dimensional; there may be continuous families, which would explain why splitting never seems to stop. Still, the results leave them more confident that some version of superposition and the linear representation hypothesis is true: the number of interpretable features, activations that behave like intensity or confidence, sensible logit weights, and interference weights are all what superposition predicts.
Are hundreds of “the” features real? (original)
A separate feature for every (token, context) pair is a local code; a compositional code would have one “the” feature and one “physics” feature and add them. Either the model is compositional and the L1 penalty pushes the autoencoder towards the sparser local code, or the model really is partly local. The authors lean towards the second, at least in part: with a compositional code, the prediction after “the” in physics would be the sum of “what follows the” and “what appears in physics”, and a local code allows a sharper prediction than that sum.
Future work (original)
The biggest open problem is scale. An autoencoder with a 100× expansion on one MLP layer of width 10,000 has about 20 billion parameters, and rare features may need a large share of the model's own training data, so the autoencoder could cost more than the model.
In words: “an encoder and a decoder, each connecting the layer's d numbers to E times as many features.” This page's own count behind the article's figure, ignoring the small biases.
With the numbers: 2 × 10,000 × 1,000,000 = 20,000,000,000, about 20 billion. The paper's A/1: 2 × 512 × 4,096 = 4,194,304.
In Python:
def params(d, E):
# encoder (E·d × d) plus decoder (d × E·d)
return 2 * d * (E * d)
params(10_000, 100) # → 20000000000
params(512, 8) # → 4194304
Beyond scale: how the ideal expansion and data needs grow with model size (scaling laws for dictionary learning); a trustworthy way to tell good features from bad; turning many microscopic facts into an understanding of a whole model, which automated interpretability might help with; better algorithms than a plain L1 penalty; and whether attention layers hold superposition of their own.
7 Comments and replications · original
Neel Nanda replicated the core result on an open one-layer model with a GELU MLP: a large fraction of the learned features were interpretable. His decoder directions were mostly spread across many neurons (92% dense; only 4% well explained by a single neuron). He found no dead features, but more than half formed an ultralow-frequency cluster that turned out to be nearly one repeated encoder direction, which reappeared across random seeds and which he couldn't interpret. And each feature's encoder and decoder directions differed (median cosine similarity 0.5), supporting untied weights: the encoder's job is detecting, avoiding interference from similar features, while the decoder's is representing the feature's true direction.
Where it leads
| Idea in the paper | Why it lasts | Where to build it |
|---|---|---|
| Decompose activations into many sparse directions | Turned “what does this neuron do?” into “which features are active?” | sparse autoencoder |
| Rebuild error plus an L1 penalty, unit-length dictionary | Still the baseline recipe that later variants modify | the objective, training it |
| Check features by their effects, not just their activations | Ablation and pinning are the causal test a label needs | intervention in the lesson |
| Recovered loss as a coverage score | The standard way to report how much a dictionary explains | the lesson's limits section |
| Feature splitting and universality | Dictionaries are zoom levels; the same features recur across models | matching features |
| Scaling to production models | Done the next year in Scaling Monosemanticity | the papers index |
Glossary
Every term with hover guidance on this page, in one place.