Toy Models of Superposition, annotated
How to read this page
- Any dotted word explains itself when you hover it, tab to it, or tap it.
- Every symbol inside an equation does the same, and each equation comes with a table decoding it, a sentence reading it aloud, the numbers worked through, and the same numbers in Python.
- The pictures are live: slide the sparsity and watch features squeeze into two dimensions, switch features on and off to see them leak into each other, and move through the phase diagram.
Each idea climbs the same ladder: everyday picture, tiny example, diagram, the math, why it matters. The interpretability lesson builds the paper's toy model in NumPy, trains it, and then recovers its features with a sparse autoencoder. Most tiny examples here reuse its five-features-in-two-neurons pentagon, so you can run them there.
Introduction and key results · original
“When features are sparse, superposition allows compression beyond what a linear model would do, at the cost of "interference" that requires nonlinear filtering.”Elhage et al. (2022), introduction
Everyday picture
Five instruments play through two speakers. A sound engineer pans each instrument to its own spot around the room. When one plays alone you can tell which it is from where the sound comes from; when two play at once the positions blur, and you might hear a third instrument that isn't there. The trick works only because, in this band, most of the time only one instrument plays. A network faces the same squeeze: far more things worth tracking than it has neurons, but in any one input almost all of those things are absent.
It would be convenient if every neuron meant one thing: one fires for a left-facing curve, another for a dog's snout. Some neurons do. Many, especially in large language models, respond to a jumble of unrelated inputs: they are polysemantic. This paper asks why, and answers with tiny networks trained on made-up data where the true features are known exactly.
Tiny example
Squeeze five features into two numbers and ask the network to get all five back out. A purely linear network can't do better than keeping the two most important and throwing three away. Put a ReLU at the end and make the features rare, and the network keeps all five, pointing them 72° apart like the corners of a pentagon. Slide the density below to see it happen. Every arrow was learned by the lesson's own training code.
Reading it: on the left, each numbered arrow is one feature's learned direction in the model's two-number hidden space (its column of ); a long arrow is a feature the model stores, an arrow shrunk to the centre is one it gave up. On the right, the bar is that arrow's length: solid when the feature points somewhere no other feature does, striped when it shares its direction with others (in superposition), missing when dropped. Feature 0 matters most and feature 4 least (importance 0.8i). At density 1 every feature is present in every example, and the model does what a linear model would: two perpendicular arrows for features 0 and 1, the rest dropped. As features get rarer, the dropped ones come back, first as antipodal pairs (at density 0.2, feature 3 points exactly opposite feature 0 and feature 2 opposite feature 1), then, from density 0.1 down, all five as a pentagon. Each setting is the best of six runs of the lesson's train_superposition (3,000 steps), rotated so feature 0 points right.
What the paper shows
- Superposition is real: small ReLU networks do store more features than dimensions, when the features are sparse.
- Both kinds of neuron form: monosemantic neurons (one feature each) and polysemantic ones, sometimes in the same model.
- Some computation works in superposition: a network can compute absolute values of more features than it has room for.
- A phase change decides it: whether a feature is dropped, packed in, or given its own dimension switches sharply as sparsity and importance vary.
- The packing has geometry: features arrange themselves as pairs, triangles, pentagons and tetrahedra.
Why it matters today
This paper gave interpretability its central problem statement. If models pack features into shared directions, then reading neurons one at a time will never be enough, and the job becomes finding the hidden features again. That is exactly what the follow-up, Towards Monosemanticity, does with a sparse autoencoder, and what the lesson's SparseAutoencoder does on this very toy.
1 Features, directions and superposition · original
“In our work, we often think of neural networks as having features of the input represented as directions in activation space.”Elhage et al. (2022), Definitions and Motivation
Everyday picture
Think of a mixing desk with a slider for every sound. What you hear is the sum of each sound at its slider's level. The paper's working assumption is that a network's activations are like that: a sum of features, each feature a fixed direction, each scaled by how strongly it is present. Two properties are packed in there. Decomposability: the activation splits into features that can be understood one at a time. Linearity: each feature is a direction, so its strength can be read with a dot product. Together these are the linear representation hypothesis.
The evidence the paper leans on is empirical: word embeddings in which king − man + woman lands near queen; interpretable directions in image generators; neurons that reliably detect curves or dog heads; the same kinds of neurons turning up across different networks; and, on the other side, many polysemantic neurons.
What is a feature? · original
Everyday picture
“Cat” and “car” feel like features; “cat plus car” feels like a mixture of two. The paper weighs three working definitions. Features as any function of the input is too loose: it can't tell “cat” from “cat plus car”. Features as human-understandable concepts is too narrow: a protein model might find an important structure nobody has named yet. The paper settles, loosely, on a third: a feature is a property that a large enough network would reliably give a neuron of its own. Curve detectors count, because they appear across many vision models above some size.
Why it matters
The definition is circular on purpose, and the paper says not to hold it too tightly. It sets up the question the rest of the paper answers: if a feature would get its own neuron in a huge model, what happens to it in a small one?
Features as directions · original
Tiny example
A two-number activation space. Feature “is a question” has direction (1, 0); feature “mentions weather” has direction (0.6, 0.8). An input that is mildly a question (strength 0.5) and strongly about weather (strength 2) produces 0.5 × (1, 0) + 2 × (0.6, 0.8) = (1.7, 1.6). The features themselves may be complicated functions of the input; only the step from feature strengths to activations is linear.
In words: “the activation is each feature's direction, scaled by how strongly that feature is present, all added up.”
With the numbers: a = 0.5 × (1, 0) + 2 × (0.6, 0.8) = (0.5 + 1.2, 0 + 1.6) = (1.7, 1.6).
In Python:
W_question = [1.0, 0.0]
W_weather = [0.6, 0.8]
x_question, x_weather = 0.5, 2.0
# a = x_question W_question + x_weather W_weather, one number per dimension
[round(x_question * q + x_weather * w, 2) for q, w in zip(W_question, W_weather)] # → [1.7, 1.6]
Why directions?
The paper gives three reasons networks would favour this format. A neuron that matches a weight template fires more the better the match, which already produces a direction. A feature stored as a direction is linearly accessible: the next layer can pick it out with one weighted sum, in one step. And directions let a model generalize beyond the neighbourhood of examples it has seen. Most of a network's arithmetic is matrix multiplies; directions are the format those operations read and write natively.
A naive count says a space with m numbers holds at most m such directions. The whole paper is about why that count is wrong.
Privileged and non-privileged bases · original
Everyday picture
A map has north and east, but nothing about the land cares which way the grid lines run; you could rotate the grid and every distance would stay the same. A chessboard is different: the moves are defined along its rows and columns, so those directions are special. A word embedding is like the map: apply any rotation to it and the undoing rotation to the next weights, and the model is unchanged, so nothing makes “dimension 7” more meaningful than any other direction. A layer followed by an activation function such as ReLU is like the chessboard: the ReLU acts on each number separately, which makes the axes special. The paper calls this a privileged basis, and reserves the word neuron for its axes.
Why it matters
Asking “what does this neuron mean?” only makes sense in a privileged basis. And even there, nothing forces features to line up with neurons; the paper will show they often don't. In the residual stream of a transformer, which has no activation function of its own, there is no reason to expect them to line up at all.
The superposition hypothesis · original
“Roughly, the idea of superposition is that neural networks "want to represent more features than they have neurons", so they exploit a property of high-dimensional spaces to simulate a model with many more neurons.”Elhage et al. (2022), The Superposition Hypothesis
Everyday picture
In a flat, two-dimensional world only two directions can be exactly perpendicular. But “almost perpendicular” is much roomier, and the roominess grows explosively with the number of dimensions: by the Johnson–Lindenstrauss lemma, a space with n dimensions holds about exp(n) directions that are all nearly perpendicular. And compressed sensing shows that a long vector squashed into a few numbers can still be recovered, if you know it is sparse (mostly zeros).
Tiny example
Five features in two dimensions, 72° apart (the pentagon). Feature 0 on alone gives the hidden state (1, 0). Reading every feature back with a dot product gives 1 for feature 0, but also 0.309 for its two neighbours and −0.809 for the two across from it. That leak is interference. It is tolerable when features are rarely on together, and a ReLU can throw away small leaks. Section 2 does this arithmetic properly.
Why it matters
If the hypothesis holds, a small network is a noisy simulation of a much larger, sparser one in which every feature has its own neuron. Neurons in the small network are then mixtures by construction, which would explain polysemanticity. Superposition doesn't need a privileged basis: it just means more features than dimensions, so it can happen in an embedding too.
A hierarchy of feature properties · original
Diagram
Tap a box. Start at the outside, decomposable, and step inwards.
Reading it: each box is a stricter property than the one around it. The outer two, decomposable and linear, are what the paper expects of most representations. Inside the linear box a representation either packs more features than dimensions (superposition, left) or doesn't (right). Only on the right can a representation also be basis-aligned, with every feature on exactly one neuron, and that needs a privileged basis. The paper expects the outer two to be widespread, and the two on the right, no superposition and basis-aligned, to hold only sometimes.
Tiny example
Put three features in two dimensions, 120° apart. Every feature's direction dotted with every other's gives the 3 × 3 table : 1 on the diagonal, −0.5 everywhere else. Its determinant is 0: the table can't be inverted, because three directions in a two-dimensional space are bound to be tangled. Two features at right angles would give the identity table instead, which inverts fine.
In words: “a linear representation is in superposition exactly when the table of dot products between its feature directions has no inverse.”
With the numbers: for the triangle, det(WᵀW) = 1(1 − 0.25) − (−0.5)(−0.5 − 0.25) + (−0.5)(0.25 + 0.5) = 0.75 − 0.375 − 0.375 = 0.
In Python:
import math
# three features in two dimensions, 120° apart: each row is one feature's direction W_i
W = [[math.cos(math.radians(120 * k)), math.sin(math.radians(120 * k))] for k in range(3)]
# WᵀW: every direction dotted with every other
WtW = [[round(sum(p * q for p, q in zip(W[i], W[j])), 3) for j in range(3)] for i in range(3)]
WtW # → [[1.0, -0.5, -0.5], [-0.5, 1.0, -0.5], [-0.5, -0.5, 1.0]]
# the determinant of a 3 × 3 table; zero means it has no inverse
(a, b, c), (d, e, f), (g, h, i) = WtW
abs(round(a * (e * i - f * h) - b * (d * i - f * g) + c * (d * h - e * g), 6)) # → 0.0
Why it matters
The definition turns a vague idea into a check on the weights. It also tells you what you give up: a table with no inverse means some patterns of features (here, all three on at once in the right proportions) produce the same activation as no features at all, so no reader can tell them apart.
2 Demonstrating superposition · original
“But we'll see that adding just a slight nonlinearity can make models behave in a radically different way!”Elhage et al. (2022), Demonstrating Superposition
Everyday picture
Imagine the huge, tidy network in which every feature has its own neuron. Its activations are the vector x. The experiment squeezes x through a narrow pipe and asks a small network to rebuild it on the far side. If the small network gets back more features than the pipe is wide, it is using superposition. Because the data is synthetic, the right answer is known exactly, which is what real models never give us.
The setup · original
The features
The paper builds its data on three assumptions about real features. They are sparse: most images contain no dog head, most tokens don't mention any particular person. There are more features than neurons: a model could usefully know every person ever written about. And they vary in importance: some reduce the loss a lot, some barely. So each feature xi is zero with probability S (its sparsity) and otherwise a random number between 0 and 1, and each carries an importance weight Ii.
Tiny example
Twenty features, each with S = 0.9. On average 20 × 0.1 = 2 are on in any example, and a single feature's average value is 0.1 × 0.5 = 0.05: mostly zeros, with a random value between 0 and 1 in about one example in ten.
In words: “each feature is off with probability S; when it is on, its strength is equally likely to be anything between 0 and 1.” This page's own symbols for the paper's sentence defining the data.
With the numbers: with S = 0.9 and 20 features, the expected number on is 20 × (1 − 0.9) = 2, and each feature's mean is (1 − 0.9) × 0.5 = 0.05.
In Python:
n, S = 20, 0.9
# expected number of features on: each is on with probability 1 − S
round(n * (1 - S), 6) # → 2.0
# a feature's mean: on with probability 1 − S, and then 0.5 on average
round((1 - S) * 0.5, 6) # → 0.05
The two models
Both squeeze with a weight matrix W (m rows, n columns, m < n) and unsqueeze with its transpose. Column Wi is feature i's direction in the narrow space. Both add a bias. The only difference is the last step: the linear model stops there, the ReLU output model passes the result through a ReLU.
In words: “squeeze the n features into m numbers with W; spread them back out with the same weights, flipped; add a bias; and, in the second model, set anything negative to zero.”
With the numbers: the pentagon (n = 5, m = 2), feature 0 alone: h = (1, 0), and Wᵀh = (1, 0.309, −0.809, −0.809, 0.309). The linear model with no bias returns exactly that, leaks and all. The ReLU model with bias −0.31 on every feature returns ReLU(1 − 0.31, 0.309 − 0.31, …) = (0.69, 0, 0, 0, 0): the leaks are gone.
In Python:
import math
# the pentagon: feature k's direction is (cos 72k°, sin 72k°)
W = [[math.cos(math.radians(72 * k)), math.sin(math.radians(72 * k))] for k in range(5)]
def model(x, b, relu):
# h = W x: two numbers
h = [sum(W[k][r] * x[k] for k in range(5)) for r in range(2)]
# Wᵀh + b: one number per feature
z = [W[k][0] * h[0] + W[k][1] * h[1] + b for k in range(5)]
return [round(max(0.0, v) if relu else v, 3) for v in z]
# the linear model, no bias: the leaks come straight through
model([1, 0, 0, 0, 0], b=0, relu=False) # → [1.0, 0.309, -0.809, -0.809, 0.309]
# the ReLU output model with a small negative bias filters them
model([1, 0, 0, 0, 0], b=-0.31, relu=True) # → [0.69, 0.0, 0.0, 0.0, 0.0]
Diagram
Tap a block. Start with features x at the top and follow the pipe down.
Reading it: five numbers go in at the top, mostly zero. The squeeze turns them into just two, which is where any superposition has to happen. The unsqueeze uses the same weights flipped, so a feature's direction on the way in is the direction it is read along on the way out; that keeps the question “which direction is feature i?” unambiguous. The bias and the ReLU are the only nonlinear part, and they turn out to decide everything: take the ReLU away and you have the linear model, which never packs features.
The loss
The model is scored by squared error, weighted by importance, so an error on an important feature costs more.
In words: “over every example and every feature, square the gap between the true and rebuilt strength, weight it by the feature's importance, and add everything up.”
With the numbers: one example with two features, importances (1, 0.7), truth (1, 0.5), rebuilt (0.69, 0): L = 1 × 0.31² + 0.7 × 0.5² = 0.0961 + 0.175 = 0.2711.
In Python:
I = [1.0, 0.7]
x = [1.0, 0.5]
x_rebuilt = [0.69, 0.0]
# Σ_i I_i (x_i − x'_i)² for this one example
round(sum(I_i * (x_i - r_i) ** 2 for I_i, x_i, r_i in zip(I, x, x_rebuilt)), 4) # → 0.2711
Why it matters
Every piece has a job. Sparsity is the condition under which superposition pays. Importance gives the model a reason to choose which features to keep. The shared weights make each feature's direction well defined. The bias lets the model output a feature's average when it doesn't store it, and, more importantly, subtract small leaks. The lesson implements exactly this loss, with its gradient derived by hand, in superposition_loss_and_grads.
Basic results · original
Everyday picture
To read a trained model, the paper asks two questions of every feature. Is it stored at all? That is how long its arrow is. Does it share its space? That is how much every other feature's arrow points along it.
Tiny example
In the ideal pentagon every arrow has length 1, so every feature is stored. Project the other four onto feature 0's direction: the neighbours contribute 0.309 each and the far pair −0.809 each. Squared and added: 2 × 0.0955 + 2 × 0.6545 = 1.5. Anything at 1 or more means other features, together, can light up feature 0's direction as strongly as feature 0 itself.
In words: “how long is feature i's arrow; and how much do all the other arrows, squared, point along it.”
With the numbers: pentagon, feature 0: ‖W₀‖ = 1, and Σ = 0.309² + (−0.809)² + (−0.809)² + 0.309² = 1.5. In the dense solution, feature 0 has the same length and a sum of 0: it is alone.
In Python:
import math
W = [[math.cos(math.radians(72 * k)), math.sin(math.radians(72 * k))] for k in range(5)]
i = 0
# ‖W_i‖: the length of feature i's arrow
norm = math.sqrt(sum(v * v for v in W[i]))
round(norm, 3) # → 1.0
# Ŵ_i: the same arrow at length 1
W_hat = [v / norm for v in W[i]]
# Σ_(j≠i) (Ŵ_i · W_j)²
round(sum(sum(p * q for p, q in zip(W_hat, W[j])) ** 2 for j in range(5) if j != i), 3) # → 1.5
What the paper found
With 20 features, 5 hidden dimensions and importance 0.7i, the linear model always keeps the 5 most important features, like the top components of PCA. The ReLU model does the same on dense data. As sparsity rises, superposition appears, starting with the least important features and working up, first as antipodal pairs and then as other shapes. With 80 features in 20 dimensions the picture is the same, rescaled. The slider at the top of this page shows the same story on the lesson's five-feature toy.
Why it matters
This is the paper's first result, and the simplest: a network with a single nonlinearity at its output packs more features than it has room for, and the amount it packs is set by sparsity. The lesson's features_a_neuron_responds_to reads the consequence off the pentagon: each hidden number responds to three features.
Why the ReLU changes everything · original
Everyday picture
Two forces pull on every feature. Feature benefit: storing a feature lets the model rebuild it, which lowers the loss. Interference: storing it in a shared direction makes it leak into other features' readings, which raises the loss. In a linear model interference always wins once the space is full. A ReLU tilts the contest: any leak that comes out negative is clipped to zero, and so costs nothing.
Tiny example
Three features in two dimensions, importances 1, 0.7 and 0.49. Option A: store features 0 and 1 at right angles and drop feature 2. Option B: store all three as a triangle, 120° apart, so every pair dots to −0.5. The linear model scores A at 0.49 (the missing feature) and B at 1.095 (all that interference), so it drops a feature. For examples with a single feature on, the ReLU model scores B at 0: every leak is −0.5, and ReLU clips it. It keeps all three.
In words: “the linear model's loss behaves like a penalty for every feature whose arrow isn't length 1 (lost benefit), plus a penalty for every pair of arrows that aren't perpendicular (interference).” The paper derives this shape from Saxe et al.'s analysis of linear networks, with Gaussian inputs; ∼ means “behaves like”, not an exact equality.
With the numbers: A: benefit term 0.49 × (1 − 0)² = 0.49, interference 0, total 0.49. B: benefit 0; each ordered pair contributes Ij × 0.25 and each feature is the j of two pairs, so 0.25 × 2 × (1 + 0.7 + 0.49) = 1.095.
In Python:
import math
I = [1.0, 0.7, 0.49]
def linear_loss(W):
dot = lambda u, v: sum(p * q for p, q in zip(u, v))
# feature benefit: Σ_i I_i (1 − ‖W_i‖²)²
benefit = sum(I[i] * (1 - dot(W[i], W[i])) ** 2 for i in range(3))
# interference: Σ_(i≠j) I_j (W_j · W_i)²
interference = sum(I[j] * dot(W[j], W[i]) ** 2 for i in range(3) for j in range(3) if i != j)
return round(benefit + interference, 3)
# A: features 0 and 1 at right angles, feature 2 dropped
linear_loss([[1, 0], [0, 1], [0, 0]]) # → 0.49
# B: all three as a triangle
linear_loss([[math.cos(math.radians(120 * k)), math.sin(math.radians(120 * k))] for k in range(3)]) # → 1.095
The ReLU model, one feature at a time
For the ReLU model the paper splits the loss by how many features are on. With n features the chance that exactly k particular ones are on is (1 − S)k Sn−k, so the loss is a weighted sum of Lk, the loss on examples with k features on.
In words: “the loss is the loss on each pattern of active features, weighted by how likely that pattern is; when features are sparse, patterns with zero or one feature on dominate.”
With the numbers: five features at S = 0.95. No feature on: 0.95⁵ = 0.774 of examples. Exactly one on: 5 × 0.05 × 0.95⁴ = 0.204. Exactly two: 10 × 0.05² × 0.95³ = 0.021. So 98% of the loss comes from the empty and one-feature examples.
In Python:
import math
n, S = 5, 0.95
# share of examples with exactly k features on: (n choose k) (1 − S)^k S^(n − k)
[round(math.comb(n, k) * (1 - S) ** k * S ** (n - k), 3) for k in range(3)] # → [0.774, 0.204, 0.021]
L0 only punishes positive biases. The interesting term is L1. Written for a feature at full strength, xi = 1, it looks just like the linear loss with a ReLU wrapped around each part:
In words: “when only feature i is on, you lose whatever of it the model fails to return, and you pay for every other feature j that it wrongly switches on; but a leak that comes out negative after the bias is clipped to zero and costs nothing.”
With the numbers: the triangle with zero biases: benefit terms ReLU(1 + 0) = 1, so 0; every leak is ReLU(−0.5) = 0; L1 = 0. Option A still pays 0.49 for dropping feature 2. The ReLU model prefers the triangle.
In Python:
import math
I = [1.0, 0.7, 0.49]
relu = lambda v: max(0.0, v)
dot = lambda u, v: sum(p * q for p, q in zip(u, v))
def L1(W, b):
# benefit: Σ_i I_i (1 − ReLU(‖W_i‖² + b_i))²
benefit = sum(I[i] * (1 - relu(dot(W[i], W[i]) + b[i])) ** 2 for i in range(3))
# interference: Σ_(i≠j) I_j ReLU(W_j · W_i + b_j)²
interference = sum(I[j] * relu(dot(W[j], W[i]) + b[j]) ** 2 for i in range(3) for j in range(3) if i != j)
return round(benefit + interference, 3)
triangle = [[math.cos(math.radians(120 * k)), math.sin(math.radians(120 * k))] for k in range(3)]
L1(triangle, b=[0, 0, 0]) # → 0.0
L1([[1, 0], [0, 1], [0, 0]], b=[0, 0, 0]) # → 0.49
Try it: interference and the bias
The pentagon is harder than the triangle: its neighbours dot to +0.309, a positive leak that the ReLU lets through. A small negative bias fixes it by pushing that leak below zero. Switch features on and move the bias. Watch feature 0's neighbours when feature 0 is on alone, and then switch two neighbours on together.
Reading it: each row is one feature: the striped bar is its true strength (1 when switched on, else nothing), the solid bar is what the ReLU model reads back from the two hidden numbers. With the bias at 0 and feature 0 alone, features 1 and 4 read 0.309: false alarms. At −0.31 they vanish, and feature 0 reads 0.69 instead of 1, which training fixes by making the arrows a little longer (the trained pentagon's arrows are about 1.13 long, so a lone feature reads 1.13² − 0.24 ≈ 1.04). Switch on features 0 and 1 together and each reads about 1.0 instead of 0.69, and the far features stay silent. Switch on 0 and 2 instead and both vanish, while feature 1, which lies between them, lights up: two features combine into what looks like a third. That collision is the price, and sparsity is the bet that it is rare.
Why it matters
The asymmetry of the ReLU, where negative interference is free and positive interference costs, explains the shapes the paper finds: solutions arrange features so most leaks come out negative, and use a negative bias to push small positive ones below zero. The paper also notes the L1 term resembles the Thomson problem of spacing charges on a sphere, which §4 turns into a prediction.
3 Superposition as a phase change · original
“The optimal weight configuration discontinuously changes in magnitude and superposition.”Elhage et al. (2022), Superposition as a Phase Change
Everyday picture
Water doesn't get gradually more solid as it cools; at 0 °C it switches to ice. The paper finds the same kind of switch for features. A feature is either not learned, learned in superposition, or given a dimension of its own, and the transitions between these outcomes are sharp. It calls this a phase change, in the loose sense of a discontinuous jump.
Tiny example
The smallest possible case: two features, one hidden dimension. Feature 1 has importance 1; the extra feature 2 has importance I. There are three natural ways to use the single dimension. Drop the extra feature, W = [1, 0]. Give the extra feature the dimension instead (dedicated), W = [0, 1]. Or store both, pointing opposite ways (antipodal, superposition), W = [1, −1], which works whenever only one is on and fails when both are. The paper trains many models over a grid of I (0.1 to 10) and density (1 down to 0.01) and also computes the three solutions' losses in closed form; the two diagrams agree, and the boundaries between the regions are sharp.
This page redoes the closed-form part with its own simple choices, since the paper's derivation lives in a notebook: the dropped feature gets the best constant bias (its average, (1 − S)/2), the antipodal pair gets no bias. When the pair collides, both on with strengths u and v, the model outputs the difference for the larger and zero for the smaller, so each feature is off by the smaller strength: an error min(u, v)², whose average over two uniform strengths is 1/6.
In words: “dropping a feature costs its importance times its variance; the antipodal pair costs nothing unless both features are on at once, which happens with probability d², and then costs a sixth of their combined importance on average.” These three formulas are this page's own derivation, checked by simulation; the paper publishes its version in a notebook.
With the numbers: I = 1 and d = 0.5: dropping costs 0.5/3 − 0.25/4 = 0.104, dedicating the same, antipodal 2 × 0.25/6 = 0.083, so superposition wins. At d = 0.8 dropping costs 0.107 and antipodal 0.213: no superposition. With I = 1, superposition wins whenever d < 4/7 ≈ 0.57.
In Python:
def losses(I, d):
# a feature left out, replaced by its average: importance × variance
unstored = d / 3 - d ** 2 / 4
drop = I * unstored
dedicated = 1 * unstored
# both on with probability d², then each is off by the smaller strength: E[min(u, v)²] = 1/6
antipodal = (1 + I) * d ** 2 / 6
return round(drop, 3), round(dedicated, 3), round(antipodal, 3)
losses(1, 0.5) # → (0.104, 0.104, 0.083)
losses(1, 0.8) # → (0.107, 0.107, 0.213)
Diagram: the phase diagram
Reading it: across is the extra feature's importance, from a tenth of feature 1's to ten times it (log scale); up is the density of both features, from present in every example at the top to one example in a hundred at the bottom (log scale). Each cell is coloured by whichever of the three solutions has the lowest loss, and each region is labelled. Top left, dense and unimportant: the extra feature is dropped. Top right, dense and important: it takes the dimension and feature 1 is dropped. The whole bottom band, sparse: both are stored, antipodally. The boundaries are where two loss formulas cross, and there the best solution jumps rather than blends; the paper checks that the derivative of the best loss is discontinuous, the signature of a first-order phase change. This computed diagram shows what the paper's text reports of its own: sparsity is needed for superposition, and relative importance decides the rest.
Three features in two dimensions
The paper repeats the experiment with three features in two dimensions, varying the third feature's importance. Now four natural solutions compete, named by the direction W ignores: drop the extra feature (W ⟂ [0, 0, 1]); drop another (W ⟂ [1, 0, 0]); pair the extra feature antipodally with one of the others (W ⟂ [0, 1, 1]); or pair the other two and give the extra feature its own dimension (W ⟂ [1, 1, 0]). Again the empirical and theoretical diagrams show sharp boundaries.
Why it matters
A phase change is hopeful news for interpretability: it means there is a region with no superposition at all, rather than superposition that shrinks forever without vanishing. If training could be pushed into that region, the problem would disappear rather than merely shrink. (A comment on the paper by Tom McGrath solves the two-feature case exactly and finds a phase the paper hadn't considered, which he calls the “confused feature” regime: both weights near 1/√2, the features similar rather than opposed, when sparsity is low and both features are important.)
4 The geometry of superposition · original
“In some ways, the structure described in this section seems "too elegant to be true" and we think there's a good chance it's at least partly idiosyncratic to the toy model we're investigating.”Elhage et al. (2022), The Geometry of Superposition
Everyday picture
Seat guests around tables so that the ones who might argue sit as far apart as possible. With four guests at a round table you'd seat them in a square, with three in a triangle. Features in superposition do the same: to keep interference low, they spread out into regular shapes.
Uniform superposition · original
Tiny example
Start with the simplest world: every feature equally important and equally sparse. The paper uses 400 features in 30 dimensions, and counts how many features the model stores with the squared Frobenius norm of W, since a stored feature has ‖Wi‖² ≈ 1 and a dropped one ≈ 0. Its reciprocal, scaled, is “dimensions per feature”. The ideal pentagon stores 5 features in 2 dimensions: 2/5 = 0.4 dimensions each. The dense solution stores 2 features in 2 dimensions: 1 each.
In words: “the number of hidden dimensions divided by the number of features the model stores: how much room each stored feature gets.”
With the numbers: pentagon: ‖W‖F² = 5 × 1² = 5, so D* = 2/5 = 0.4. Antipodal pairs in 30 dimensions store 60 features: D* = 30/60 = 0.5.
In Python:
import math
W = [[math.cos(math.radians(72 * k)), math.sin(math.radians(72 * k))] for k in range(5)]
m = 2
# ‖W‖_F²: every weight squared, added up
frob = sum(v * v for column in W for v in column)
# D* = m / ‖W‖_F²
round(m / frob, 3) # → 0.4
What the paper found
Plotting D* against sparsity gives a curve that is sticky: it sits flat at 1 (every stored feature has its own dimension) and flat again at 1/2 (antipodal pairs, two features per dimension) over wide ranges of sparsity, before sliding lower. The paper notes the plateaus vaguely resemble those of the fractional quantum Hall effect in physics. Antipodal pairs are so cheap that the model prefers them over a wide range. Only the ratio matters: doubling the hidden dimensions doubles the features stored, and the number of input features is irrelevant as long as it is much larger than m.
Feature dimensionality · original
Everyday picture
D* is an average. To see what each feature gets, the paper defines a per-feature share: how much of a dimension is this feature's own? A feature alone in its direction gets 1. Two features sharing one direction, back to back, get half each.
In words: “how strongly feature i is stored, divided by how many features, itself included, are crowded along its direction.”
With the numbers: antipodal pair: 1 / (1 + 1) = 1/2. Triangle: 1 / (1 + 0.25 + 0.25) = 2/3. Pentagon: 1 / (1 + 2 × 0.0955 + 2 × 0.6545) = 1 / 2.5 = 2/5. The shares add up to the number of dimensions when the features are packed efficiently: 5 × 2/5 = 2.
In Python:
import math
def D(W, i):
dot = lambda u, v: sum(p * q for p, q in zip(u, v))
norm2 = dot(W[i], W[i])
W_hat = [v / math.sqrt(norm2) for v in W[i]]
# D_i = ‖W_i‖² / Σ_j (Ŵ_i · W_j)², where the sum includes j = i
return round(norm2 / sum(dot(W_hat, Wj) ** 2 for Wj in W), 3)
polygon = lambda k: [[math.cos(2 * math.pi * j / k), math.sin(2 * math.pi * j / k)] for j in range(k)]
# antipodal pair, triangle, pentagon
[D(polygon(k), 0) for k in (2, 3, 5)] # → [0.5, 0.667, 0.4]
# a tetrahedron in three dimensions: every pair dots to −1/3
s = 1 / math.sqrt(3)
D([[s, s, s], [s, -s, -s], [-s, s, -s], [-s, -s, s]], 0) # → 0.75
Try it
Reading it: the arrows are the features' directions, spread evenly round the circle; the bars on the right are each feature's dimensionality Di, all equal because every feature sits in the same position relative to the others. From three features up, each gets 2/k and the shares add up to exactly 2, the number of dimensions: an efficient packing. Two features back to back get 1/2 each and add up to only 1: an antipodal pair uses one dimension, leaving the other free for a second pair. Four features form exactly that, since a square is two antipodal pairs at right angles, and they get the same 1/2 each.
What the paper found
Plotting every feature's Di across sparsities, the points cluster on a few lines: ¾ (tetrahedron), ⅔ (triangle), ½ (antipodal pair), ⅖ (pentagon), ⅜ (square antiprism) and 0 (not learned). Each fraction is one geometric arrangement, and the model jumps between them as sparsity changes. Everything strictly between 0 and 1 is superposition, so superposition isn't one phase but many, the way ice has several crystal structures.
Why these shapes? · original
Everyday picture
Several of the shapes are solutions of the Thomson problem: where charges settle on a sphere. A stored feature is a point on a sphere in the hidden space, and interference plays the role of repulsion. The paper also notices that the lines appear for the uniform shapes, where every point sits in the same position, and split where a solution is non-uniform: the Thomson answer for five points in three dimensions, a triangular bipyramid, would give ⅗, but the models show points at ⅔ and ½ instead. That bipyramid is a triangle and an antipodal pair in perpendicular subspaces, a tegum product, and features in different perpendicular pieces never interfere.
The aside: shapes are superposition strategies
Any matrix of the form W⊤W is a table of dot products between the columns of W, which are n points in m dimensions. So every polytope is a way to superpose features, and every way is a polytope: three features in two dimensions are a triangle. Going the other way, pick the direction W should ignore. Ignore (1, 1, 1) in three dimensions by projecting each axis onto the plane perpendicular to it: axis 1 becomes (2/3, −1/3, −1/3), of length 0.816, and any two projected axes have cosine −0.5. That is an equilateral triangle. The fully dense direction (1, 1, …, 1) is the cheapest one to give up when features are equal and sparse, because all of them being on at once is the rarest pattern.
Why it matters
If anything like this holds in real models, features in superposition would come in small, independent, predictable groups, which would make them far easier to find. The paper is candid that this part may be special to the toy.
Non-uniform superposition · original
Everyday picture
Real features aren't equal. Make one guest at the round table talk more, and the others shuffle away to give them room; make one nearly silent and the others crowd in. Past some point the seating snaps into a different arrangement entirely.
Tiny example
The paper takes the uniform pentagon (5 features, 2 dimensions, density 0.05) and changes one feature's density. Denser, and the other four move away from it; sparser, and they move towards it. Sparse enough, and the pentagon collapses into two antipodal pairs with the rare feature dropped, at the point where the two arrangements' loss curves cross: a first-order phase change seen directly. The slider at the top of the page shows the reverse journey in the lesson's run: two antipodal pairs at density 0.2, a pentagon from 0.1.
Pentagon solutions also sit off the unit circle. The model sets a small negative bias to cut off noise and lengthens the arrows to compensate. In the lesson's run at density 0.05 the bias is about −0.24 and the arrows are about 1.13 long, so a lone feature at strength 1 reads 1.13² − 0.24 = 1.04, close to 1. The paper writes the compensation as ‖Wi‖ = 1/(1 − bi); with the bias counted as a cut of 0.23 to 0.25, the run's numbers fit that formula for the squared length (1.28 against 1/(1 − 0.23) = 1.30) better than for the length, so read it as the direction of the effect: the harder the bias cuts, the longer the arrows.
Correlated and anticorrelated features · original
Everyday picture
Fur, ears and eyes tend to show up together (a correlated bundle); a word being English and being German rarely do (anticorrelated). Features that show up together would collide constantly if they shared space, so the model keeps them apart; features that never co-occur can share space safely.
What the paper found
- Correlated features prefer to be perpendicular, in separate tegum factors. In larger models this produces local almost-orthogonal bases: the model as a whole is in superposition, but any one correlated set, taken alone, barely is. If real models do this, methods such as PCA might be safe to use on narrow slices of the data.
- When correlated features can't be perpendicular, they sit side by side, preferring positive interference with each other to negative.
- Anticorrelated features prefer to share, ideally as antipodal pairs: they are never on together, so the collision never happens.
- When there isn't room for a correlated pair, the model collapses them into their principal component.
Tiny example: collapse into a principal component
Two features a and b always switch on together (their strengths still vary independently). With room for only one direction, the model stores the direction halfway between them, reading the sum. If a = 0.6 and b = 0.8, that stored number is (0.6 + 0.8)/√2 = 0.99; the difference between them is lost.
In words: “when two features always arrive together and only one direction is available, store their normalized sum and give up their difference.”
With the numbers: (0.6 + 0.8)/1.414 = 0.99; the ignored second component is (0.6 − 0.8)/1.414 = −0.141.
In Python:
import math
a, b = 0.6, 0.8
# the kept first principal component, (a + b)/√2
round((a + b) / math.sqrt(2), 3) # → 0.99
# the ignored second one, (a − b)/√2
round((a - b) / math.sqrt(2), 3) # → -0.141
The paper's demonstration uses six features in three correlated pairs: very sparse, they form a hexagon with each pair side by side; less sparse, the pairs progressively collapse into their principal components, until in the dense regime the solution is just PCA.
Why it matters
PCA and superposition look like two ends of one trade-off: correlation favours PCA, sparsity favours superposition, and real data, both sparse and correlated, probably gets a mixture. It also answers a practical question: polysemantic neurons won't group features at random, so comparing which features share a neuron across models won't be a clean test for superposition.
5 Superposition and learning dynamics · original
Everyday picture
An electron in an atom doesn't drift between energies; it jumps from one allowed level to another. The paper sees features do something similar during training: a feature's dimensionality sits at one value, then jumps to another, often trading places with another feature, and the loss drops suddenly at the same moment.
What the paper found
- Energy-level jumps (original). In a uniform model whose features end up as antipodal pairs, the dimensionality of individual features jumps between values during training, and each jump lines up with a sudden drop in the loss (a small one at the first jump, a larger one at the second). The authors suspect that the smooth loss curves of large models are made of many such small jumps, in the same family as the sudden changes seen when induction heads form and in grokking.
- Learning as geometric moves (original). With six features in three dimensions, as two correlated sets of three, training passes through distinct regimes visible in the loss curve, each a simple geometric transformation of the arrangement, ending as an octahedron with features from different sets in antipodal pairs. The network first learns something close to the linear, PCA-like solution, then moves to a better nonlinear one.
Why it matters
If loss curves are sums of discrete rearrangements, then “the model learned X at step Y” can be a precise statement. This part of the paper is short and exploratory; the authors expect it to generalize less than the basic superposition results.
6 Relationship to adversarial robustness · original
“Note that this may remain true even in the infinite data limit: the optimal behavior of the model fit to sparse infinite data is to use superposition to represent more features, leaving it vulnerable to attack.”Elhage et al. (2022), Relationship to Adversarial Robustness
Everyday picture
If five instruments share two speakers, someone who knows the panning can fake the sound of the violin by playing the flute and the drum at the right volumes. An adversarial example does the same to a network: a small, deliberately chosen change to the input that makes the model see something that isn't there. Superposition hands the attacker the recipe: every feature's reading already leaks from others.
Tiny example
Row 0 of W⊤W says how much each input feature moves feature 0's reading. Without superposition it is (1, 0, 0, …): only feature 0 moves it. In the pentagon it is (1, 0.309, −0.809, −0.809, 0.309). The paper's attack on feature i pushes the input a fixed small distance λ along that row. The same push moves feature 0's reading by λ × the row's length: 1 × λ without superposition, 1.58 × λ in the pentagon.
In words: “without superposition only a feature moves its own reading; with it, every other feature nudges it a little, and the most damaging small push points along exactly those nudges.”
With the numbers: the pentagon's row has length √(1 + 2 × 0.309² + 2 × 0.809²) = √2.5 = 1.581. A push of size λ = 0.1 along it moves feature 0's reading by 0.1 × 1.581 = 0.158, against 0.1 without superposition.
In Python:
import math
W = [[math.cos(math.radians(72 * k)), math.sin(math.radians(72 * k))] for k in range(5)]
# (WᵀW)_0: how much each input feature moves feature 0's reading
row = [sum(p * q for p, q in zip(W[0], W[j])) for j in range(5)]
[round(v, 3) for v in row] # → [1.0, 0.309, -0.809, -0.809, 0.309]
lam = 0.1
# δ = λ (WᵀW)_0 / ‖(WᵀW)_0‖
row_norm = math.sqrt(sum(v * v for v in row))
delta = [lam * v / row_norm for v in row]
# how far feature 0's reading moves: (WᵀW)_0 · δ
round(sum(r * d for r, d in zip(row, delta)), 3) # → 0.158
What the paper found
Gradient-based attacks struggled on very sparse models, where ReLUs are off almost all the time (the gradient gives little signal), so the authors used the analytic attack above for each feature, keeping the most harmful, with a budget of 0.1 of the average input's length. Vulnerability rose sharply, more than threefold, as superposition formed, and tracked the number of features per dimension. Adversarial training did reduce superposition, but only with unreasonably large attacks (80% of the input's length) did it remove it.
Why it matters
The authors are careful not to claim superposition causes adversarial examples in real models; other theories exist. But the view makes predictions: robust models should do worse on the task (they give up packed features), be more interpretable, and share attacks when their superposition is shaped by the same correlations.
7 Superposition in a privileged basis · original
Everyday picture
So far the hidden space had no special axes: spin it any way and nothing changes, like the map from §1. Real layers with neurons are more like the chessboard. The paper adds a ReLU to the hidden layer so the axes become neurons, and asks: do features line up with neurons, and when do neurons become polysemantic?
Tiny example: why the axes stop being arbitrary
Without a hidden ReLU, rotating the hidden space by any rotation O changes nothing, because the model only ever uses W⊤W. With it, a rotation can wreck the model. Feature 0 sits on neuron 1 at (1, 0). Rotate by 180° and it sits at (−1, 0); the ReLU turns that into (0, 0), and the feature is gone.
In words: “rotating the hidden space leaves the table of dot products untouched, since a rotation followed by its undoing is nothing, so without a hidden ReLU no direction is special; once a ReLU sits on the hidden numbers, each neuron's own axis matters.”
With the numbers: the pentagon's W⊤W entry for features 0 and 1 is 0.309 before and after a 90° rotation. Feature 0 alone: ReLU(W x) = ReLU(1, 0) = (1, 0) before, and ReLU(−1, 0) = (0, 0) after a 180° rotation.
In Python:
import math
W = [[math.cos(math.radians(72 * k)), math.sin(math.radians(72 * k))] for k in range(5)]
def rotate(v, degrees):
c, s = math.cos(math.radians(degrees)), math.sin(math.radians(degrees))
return [c * v[0] - s * v[1], s * v[0] + c * v[1]]
dot = lambda u, v: sum(p * q for p, q in zip(u, v))
# (OW)ᵀ(OW) = WᵀW: the entry for features 0 and 1, before and after a 90° turn
round(dot(W[0], W[1]), 3), round(dot(rotate(W[0], 90), rotate(W[1], 90)), 3) # → (0.309, 0.309)
# with a hidden ReLU the turn matters: feature 0 alone, before and after 180°
[round(max(0.0, v), 3) for v in W[0]], [round(max(0.0, v), 3) for v in rotate(W[0], 180)] # → ([1.0, 0.0], [0.0, 0.0])
Diagram: neurons, one feature or many
Reading it: this is the paper's stack plot, drawn for this page's own run of its setup (10 features, 5 neurons, importance 0.75i, hidden ReLU, best of eight trainings at each density). Each column is one neuron; each block in it is one feature's weight into that neuron, labelled with the feature's number, its height the weight's size, above the line if positive and below if negative. Weights smaller than 0.2 are left out. Dense data (density 1): five columns of one block each, five monosemantic neurons for the five most important features. At density 0.3, three neurons still hold one feature each, while features 2, 4 and 5 share two neurons: both read feature 2 positively, one together with 5 and against 4, the other together with 4 and against 5. At density 0.1 only feature 0 keeps a neuron to itself, and four neurons share features 1 to 6. The most important features are the last to lose their own neuron.
What the paper found, and a weakness
The paper trains 10 features on 5 neurons and, keeping the best of 1,000 runs per setting, finds the same shift from monosemantic to polysemantic neurons as features get sparser, with both kinds in one model at intermediate sparsity and a neuron-level phase change mirroring the feature-level one. Polysemantic solutions are surprisingly structured: features map to small sets of neurons. But this model has a flaw the authors point out: the hidden ReLU does nothing useful for the task, so the model evades it whenever it can (with a hidden bias it would push every neuron into the positive, linear range). The next section fixes that by giving the ReLU real work.
Why it matters
This explains why looking at neurons works sometimes (early vision layers are full of clean detectors) and fails other times: whether a neuron is clean depends on how sparse and important its features are, not on luck.
8 Computation in superposition · original
“So far, we've shown that neural networks can store sparse features in superposition and then recover them. But we actually believe superposition is more powerful than this – we think that neural networks can perform computation entirely in superposition rather than just using it as storage.”Elhage et al. (2022), Computation in Superposition
Everyday picture
Storing things compressed is one skill; working on them without unpacking them first is another. The paper picks the simplest nonlinear job, taking an absolute value, and asks a small network to do it for more features than it has room for.
Tiny example
Absolute value has a neat two-neuron recipe: one neuron passes the positive side, one the negative side flipped, and adding them gives the size. For x = −0.6: ReLU(−0.6) = 0, ReLU(0.6) = 0.6, sum 0.6. So each feature needs two neurons if it is to have its own; with 3 features and 6 neurons the trained model does exactly that. Now the inputs are between −1 and 1 when on, and the model's two layers have separate weights.
In words: “the size of a number is its positive part plus its flipped negative part; the network computes such pieces in a hidden layer of neurons and adds them up in the output layer, now with two separately learned weight matrices.”
With the numbers: abs(−0.6) = ReLU(−0.6) + ReLU(0.6) = 0 + 0.6 = 0.6; abs(0.3) = 0.3 + 0 = 0.3.
In Python:
relu = lambda v: max(0.0, v)
# abs(x) = ReLU(x) + ReLU(−x), a "positive side" neuron plus a "negative side" neuron
[relu(x) + relu(-x) for x in (-0.6, 0.3)] # → [0.6, 0.3]
What the paper found
With 100 features, 40 neurons and importance 0.8i, dense data gives one clean feature per neuron, readable directly. Sparse enough, and the model computes absolute value for more features than 40 neurons could handle one-to-one. At intermediate sparsity it shows two patterns known from real networks: most neurons stay pure while a few become highly polysemantic (and the neurons for the most important features stay pure); and many neurons have one primary feature with a large weight plus some secondary features with small weights. Such a neuron looks clean if you only examine its strongest activations, but its weaker activations are a mix, which is exactly what researchers see in language models.
Diagram: the asymmetric superposition motif
Inside one trained model the authors find a recurring two-neuron trick. The first neuron stores two features at unequal sizes, such as weights 2 and −½, and reads them out with reciprocal weights ½ and 2. Feature b's leak into a's output is then small, but a's leak into b's output is large, so a second neuron fires on a and strongly inhibits b's output, turning a harmful positive leak into a harmless negative one that the output ReLU clips. The weights below are this page's own illustration of that pattern, using the paper's example magnitudes.
Tap a part. Start at x_a, then follow its wires to both outputs.
Reading it: inputs on the left, outputs on the right, neurons in the middle; each wire is labelled with its weight, and the dashed wire is the inhibition. Neuron n₁ computes ReLU(2xa − ½xb), so it fires for positive a and for negative b, two features in one neuron. Its output weights, ½ to ya and 2 to yb, undo the input sizes on the right path. Follow xb = −1: n₁ = 0.5, so yb gets 1 (correct) and ya gets only 0.25 (a small leak). Follow xa = 1: n₁ = 2, so ya gets 1 (correct) but yb would get 4, a disaster. That is why n₂ = ReLU(xa) exists: it adds −4 to yb, cancelling the leak exactly here, and the output ReLU keeps yb at zero if a is ever larger.
In Python:
relu = lambda v: max(0.0, v)
def motif(x_a, x_b):
# n1 stores a large and b small; n2 fires on a alone
n1 = relu(2 * x_a - 0.5 * x_b)
n2 = relu(x_a)
# outputs read n1 with reciprocal weights; n2 inhibits y_b
return relu(0.5 * n1), relu(2 * n1 - 4 * n2)
# b alone (negative side): y_b is right, y_a gets a small leak
motif(0.0, -1.0) # → (0.25, 1.0)
# a alone: y_a is right, and the inhibition cancels the big leak into y_b
motif(1.0, 0.0) # → (1.0, 0.0)
Why it matters
Computation in superposition is what makes the phenomenon serious: a network isn't just storing packed features, it is running circuits on them. Any account of what a model computes has to handle circuits whose inputs, neurons and outputs are all shared.
9 The strategic picture of superposition · original
“The ability to have a universal quantifier over the fundamental units of neural network computation is a significant step towards saying that certain types of circuits don't exist.”Elhage et al. (2022), Safety, Interpretability, & "Solving Superposition"
Everyday picture
To promise a building has no gas leaks, you need to be able to check every pipe. The authors want the same for models: to say a model never does something, such as deceive, one strong tool would be the ability to list all its features and check each. Superposition is what stops that: when features aren't neurons, there is no list to go through. They call any method that recovers the list, or equivalently unfolds the activations into those of a larger non-superposed model, a solution to superposition. Even simple checks depend on it: under superposition, a high cosine similarity between two activations can come from unrelated features that happen to share directions.
Three ways out
- Build models without superposition (original). In the toys it is easy: add an L1 penalty on the hidden activations. The catch is cost: superposition makes a model effectively larger, and giving it up may cost a lot. But the true budget is compute, not neurons, and a mixture of experts changes that relation by switching most neurons off on each input, so it eats the same gap between feature sparsity and neuron sparsity that superposition exploits.
- Find an overcomplete basis after the fact (original). Take a trained model's activations and search for more directions than dimensions such that every activation is a sparse combination of them: dictionary learning. Training is untouched, but new problems arrive: how many features to look for, a huge engineering job, and no training pressure helping you.
- Hybrids (original): train models with a little less superposition, or with architectures that make the after-the-fact search easier. At the optimum, the loss is flat in the amount of superposition, so a little can be removed almost for free.
Tiny example: the two formulas behind the first two routes
Route 1 adds a penalty for total activation. A hidden state h = (0.5, −0.2, 0) with λ = 0.1 adds 0.1 × (0.5 + 0.2 + 0) = 0.07 to the loss, so the model is pushed to use fewer, smaller activations. Route 2 writes a table of activations as a sparse table times a dictionary: with two stimuli, three features and a two-dimensional hidden space, the rows of B are the feature directions and each row of A says which features are on.
In words: “route 1: add to the loss a multiple of the total size of the hidden activations; route 2: explain the table of recorded activations as a mostly-zero table of feature strengths times a table of feature directions, with more features than dimensions.”
With the numbers: the penalty is 0.1 × 0.7 = 0.07. With directions B = [(1, 0), (0, 1), (−0.6, 0.8)] and sparse codes A = [(0, 0, 1), (2, 0, 0)], the rebuilt activations are H = [(−0.6, 0.8), (2, 0)]: each stimulus uses one feature, and there are three features in a two-dimensional space.
In Python:
lam = 0.1
h = [0.5, -0.2, 0.0]
# λ ‖h‖₁: λ times the sum of absolute values
round(lam * sum(abs(v) for v in h), 3) # → 0.07
# H ≈ A B: each row of A says how much of each feature direction (row of B) a stimulus uses
B = [[1, 0], [0, 1], [-0.6, 0.8]]
A = [[0, 0, 1], [2, 0, 0]]
[[round(sum(A[s][f] * B[f][r] for f in range(3)), 2) for r in range(2)] for s in range(2)] # → [[-0.6, 0.8], [2.0, 0.0]]
Why it matters today
Route 2 is the one the field took. Within a year, the same group trained sparse autoencoders on a transformer's activations (Towards Monosemanticity), and the lesson does it on this toy: train_sae learns five latents from the pentagon's hidden states and match_features checks them against the planted directions. The paper's other hopes remain: phase changes mean superposition can, in principle, be switched off entirely, and even an underperforming superposition-free model would be a valuable research tool, because it would provide the ground truth that real models lack. Local bases, the authors add, won't be enough on their own for safety claims, which need to cover the whole input distribution.
10 Discussion · original
Does this happen in real models?
The toys are only worth studying if their predictions match what is seen in real networks, and the authors list four matches. Polysemantic neurons exist. Clean and polysemantic neurons sit side by side in the same layer. In the InceptionV1 vision model, later layers have more polysemantic neurons; later layers detect rarer, higher-level things (a floppy ear rather than an edge), and the toy predicts more superposition for sparser features. And the first MLP layer of transformer language models is extremely polysemantic, as expected if its job is to tell apart meanings of the same token across languages (“die” in English, German, Dutch, Afrikaans), each very rare.
They expect superposition, the monosemantic/polysemantic split and perhaps the adversarial link to carry over to real models, and are much less sure about the geometry and learning dynamics.
Open questions (original)
Among the questions the paper leaves open: is there a statistical test for superposition? How can training control it (L1 on activations, adversarial training, other activation functions)? Can the importance and sparsity curves of a real model be estimated? Does superposition fade with scale, stay a constant fraction, or grow? How much computation can happen in superposition, and does it require sparse structure? Do models use nonlinear, rather than linear, compression? An appendix gives an example of such a scheme and argues it is inefficient for networks.
11 Related work · original
Everyday picture
The idea didn't start here. Arora and colleagues suggested that a word with several senses has an embedding that is a sum of one vector per sense, and extended it to many sparse “atoms of discourse”. Neuroscience has long debated local codes (one neuron per thing) against distributed ones. Research on disentanglement tries to force a model's axes to match factors like rotation or lighting. What is closest mathematically is compressed sensing and its sibling, sparse coding.
Compressed sensing: how much can fit? (original)
Compressed sensing recovers a long vector with at most k non-zero entries from a shorter projection of it, using an optimization algorithm. It needs the projection to have at least on the order of k log(n/k) dimensions. The toy model is a much weaker decoder (one matrix and a ReLU), so that lower bound becomes an upper bound on how much superposition the toy can achieve. Plugging in the paper's sparsity, where a typical example has k ≈ (1 − S)n features on, gives a bound linear in n.
In words: “recovering n features, k of them on, needs a hidden size that grows at least like k times the log of n over k; with sparsity S that is n times the density times the surprise of a feature being on.”
With the numbers: 400 features at density 0.05, so k = 20: k ln(n/k) = 20 × ln 20 = 59.9, and −n(1 − S) ln(1 − S) = −400 × 0.05 × ln 0.05 = 59.9, the same number: the hidden size must grow at least in proportion to this (the Ω hides a constant the paper doesn't give).
In Python:
import math
n, density = 400, 0.05
k = n * density
# k log(n/k)
round(k * math.log(n / k), 1) # → 59.9
# −n (1 − S) log(1 − S), the same quantity written with the density 1 − S
round(-n * density * math.log(density), 1) # → 59.9
The two expressions agree because k/n is the density. The authors note that the number of features that fit grows linearly with m, modulated by sparsity, and point out a suspicious parallel: compressed sensing also has sharp phase transitions between recoverable and not.
Sparse coding and dictionary learning (original)
Sparse coding is compressed sensing where the projection matrix is unknown too: find both the directions and the sparse codes. Neuroscience uses it to describe what neurons do to their inputs; the authors propose running it the other way, to find features in superposition over neurons. It is, they say, the most natural mathematical form of “solving superposition”, and that sentence is the seed of the sparse autoencoder work that followed.
Why it matters
The compressed-sensing view gives the toy a ceiling: however clever the network, it cannot pack more features than a best-possible decoder could recover. And it names the tool, dictionary learning, that the field would pick up.
12 Comments and replications · original
The article invited outside researchers to comment. Several (at Redwood Research, DeepMind, OpenAI, and independently) replicated the basic results. Redwood found the phase diagrams change noticeably with the activation function. Tom McGrath solved the two-feature, one-dimension case exactly and found the extra “confused feature” phase noted in §3. Jermyn, Hubinger and Schiefer found training and initialization methods that make these toys monosemantic, suggesting that in some limits there may be no price for it. And Sharkey, Braun and Millidge trained a sparse autoencoder with an L1 penalty on toy data built from known features and found it recovered features almost identical to the ground truth. That last comment is the bridge to Towards Monosemanticity. Neel Nanda describes a test of the linear representation hypothesis: an Othello-playing model whose board state looked nonlinear turned out to be linear once the features were “my colour” and “their colour” rather than black and white.
Where it leads
| Idea in the paper | Why it lasts | Where to build it |
|---|---|---|
| Features are directions, and sparse features can outnumber dimensions | The working model of representations behind almost all current interpretability | the pentagon |
| A ReLU with a negative bias filters interference | Explains why the output nonlinearity, not the squeeze, makes superposition possible | superposition readout |
| Sparsity and importance decide, with sharp transitions | Predicts when neurons will be clean and when polysemantic | training the toy |
| Polysemantic neurons are a consequence, not an accident | Why reading single neurons fails in language models | a neuron's features |
| Solve superposition by finding an overcomplete basis | Became sparse autoencoders, the main tool of the field | sparse autoencoder, Towards Monosemanticity |
| Enumerating features as a route to safety claims | Still the aspiration; coverage of a real model's features remains incomplete | the interpretability lesson's limits section |
Glossary
Every term with hover guidance on this page, in one place.