An annotated companion · AI Primer

Locating and Editing Factual Associations in GPT, annotated

Paper on arXiv PDF
About this page. This is a companion, not a copy. It follows the paper section by section, quotes only a sentence or two per section (clearly marked), and explains everything in its own words. The paper is distributed on arXiv under its standard non-exclusive licence, which grants no permission to reproduce figures, so every picture here is drawn from scratch: numbers in pictures either come from the paper's text and tables (a few selected rows, credited in each caption) or are labelled illustrative. Equations are reproduced with every symbol decoded; numbers invented to keep an example small are labelled illustrative. Read the original alongside: every section heading links to it.

How to read this page

  • Any dotted word explains itself when you hover it, tab to it, or tap it.
  • Every symbol inside an equation does the same, and each equation comes with a table decoding it, a sentence reading it aloud, the numbers worked through, and the same numbers in Python.
  • The pictures are live: hover the causal-trace map and switch between states, MLPs and attention; move a new key and watch how much a naive edit and ROME disturb the old memories; drag three scores and see why the paper averages them the way it does.

Each idea climbs the same ladder: everyday picture, tiny example, diagram, the math, why it matters. The interpretability lesson builds causal tracing, there called activation patching, on a hand-built model where the answer is known.

1 Introduction · original

“Where does a large language model store its facts?”Meng et al. (2022), §1

Everyday picture

Give GPT “The Space Needle is located in the city of” and it says “Seattle”. Somewhere in its billions of weights is the fact. Finding where is like finding which wire in a house controls one light: you can flick switches at random, or you can cut power everywhere and restore it one wire at a time until the light comes back on. The paper does the second, then proves it found the right wire by rewiring it: after one small change to one layer's weights, the model says the Space Needle is in Paris, however the question is phrased, while leaving other landmarks where they were.

What they did

  1. Causal tracing: run the model on a factual prompt; run it again with the subject's words blurred by noise; then rerun the blurred version many times, each time restoring one hidden state to its clean value, and measure which restorations bring the right answer back.
  2. ROME, Rank-One Model Editing: treat one middle-layer MLP's output matrix as a lookup table from subject to facts, and insert a new entry with a single rank-one update.
  3. Two tests: an existing benchmark (zsRE) and a harder new one, CounterFact, built from counterfactual facts, which checks that an edit generalizes to paraphrases without spilling onto similar subjects.

What they found

  • In GPT-2 XL, factual recall runs through MLP modules at middle layers, while the model is reading the subject's last token; attention at late layers then carries the result to the end of the prompt.
  • Editing one MLP layer's weights there works: ROME matches other editors on zsRE and, on CounterFact, is the only method that is both general and specific.
  • Editing works best exactly where tracing pointed, which ties the two halves together.

Why it matters today

Clean run, corrupted run, one restored activation: that recipe is the one the lesson builds as activation patching, and its patching_map is the loop. The idea that an MLP layer is a key–value memory you can write to directly is the root of model editing, which the authors' own follow-up scaled to many facts at once, and of the sobering point the paper makes itself: an editable model can be edited to say false things.

2 Interventions on activations for tracing information flow · original

Everyday picture

A fact is a triple: a subject, a relation and an object, like (Space Needle, is located in, Seattle). The prompt states the subject and relation; the model supplies the object. Inside, the model keeps one running vector per token, updated layer by layer: think of a grid of notebooks, one row per word and one column per layer, where each layer adds two notes to every notebook. One note comes from attention, which reads earlier words' notebooks; the other from the MLP, which looks only at this word's notebook.

Tiny example

This page's illustrative numbers: a two-number hidden state h = (1, 0) at the previous layer, an attention contribution a = (0.5, 0.5), and an MLP with three hidden neurons.

In words: “a token's new state is its old state plus what attention adds plus what the MLP adds; the MLP normalizes attention's output plus the old state, multiplies by one matrix, clips negatives to zero, and multiplies by a second matrix.” This is the paper's equation (1), with attention and the MLP in sequence, as in GPT-2 and GPT-3; a footnote says the methods also apply to models such as GPT-J that compute them in parallel.

With the numbers: a + h = (1.5, 0.5); normalized, γ gives (1, −1). Wfc turns that into (1, −1, 0); the ReLU keeps (1, 0, 0). Wproj maps it to m = (0.5, 0). New state: (1, 0) + (0.5, 0.5) + (0.5, 0) = (2, 0.5). Keep an eye on the middle vector (1, 0, 0): §3.1 will call it the key.

In Python:

import statistics
h_prev = [1.0, 0.0]
a = [0.5, 0.5]
# γ: normalize a + h (subtract the mean, divide by the spread)
v = [p + q for p, q in zip(a, h_prev)]
mu, sd = statistics.mean(v), statistics.pstdev(v)
g = [(x - mu) / sd for x in v]
g  # → [1.0, -1.0]
# σ(W_fc γ(...)): the MLP's hidden layer, one number per neuron
W_fc = [[1, 0], [0, 1], [1, 1]]
k = [max(0.0, sum(w * x for w, x in zip(row, g))) for row in W_fc]
k  # → [1.0, 0.0, 0.0]
# m = W_proj k
W_proj = [[0.5, 0, 0], [0, 0, 1]]
m = [sum(w * x for w, x in zip(row, k)) for row in W_proj]
m  # → [0.5, 0.0]
# h = h_prev + a + m
[p + q + r for p, q, r in zip(h_prev, a, m)]  # → [2.0, 0.5]

Why it matters

Written this way, every token at every layer has three separate places information could live: the state h, the attention output a, and the MLP output m. The whole grid is a causal graph, with arrows only from earlier layers and (through attention) from earlier tokens. Causal tracing asks which nodes in that graph the fact must pass through. The lesson's residual_stream collects the same h at every layer of a tiny GPT.

2.1 Causal tracing of factual associations · original

Everyday picture

This is causal mediation analysis, and it needs three runs. The clean run answers correctly, and every hidden state is saved. The corrupted run adds noise to the subject's word vectors, so the model can no longer tell which landmark it is reading about, and usually gets the answer wrong. The corrupted-with-restoration run repeats the corrupted run but forces one hidden state, at one token and one layer, back to its clean value, and lets everything after it run normally. If that single state brings the answer back, it carries the fact.

How much noise

Every subject token's embedding gets Gaussian noise added, with a spread three times the standard deviation of real token embeddings. This page's illustrative sample of embedding entries: 0.1, −0.1, 0.3, −0.3, 0.

In words: “for every token of the subject, replace its embedding by itself plus random noise whose spread is three times the typical spread of embeddings.” The third part is from Appendix B.1.

With the numbers: the sample's standard deviation is σt = 0.2, so the noise has spread ν = 0.6. The paper repeats each corrupted run ten times with different noise.

In Python:

import statistics
# a few entries of token embeddings sampled from text (illustrative)
samples = [0.1, -0.1, 0.3, -0.3, 0.0]
sigma_t = statistics.pstdev(samples)
round(sigma_t, 3)  # → 0.2
# ν = 3 σ_t: the noise added to every subject embedding
round(3 * sigma_t, 3)  # → 0.6

Diagram: the three runs

token embed layers 1 → L The* Space* Need* le* is in downtown early site late site Seattle?P[o] copy in one clean state,then rerun what follows

Tap a part. Start with the noisy embeddings (the starred subject tokens), then the dashed restore arrow, then follow the flow to Seattle?

The paper's Figure 1 (a to d), redrawn and simplified: five layer columns stand for GPT-2 XL's 48. Based on Meng et al. (2022), Figure 1.

Reading it: rows are the tokens of “The Space Needle is in downtown” (the tokenizer splits “Needle” into “Need” and “le”); columns are layers, left to right. The four starred rows are the subject, and their embeddings (the small boxes on the left) get noise, so every later state in those rows, and every state that attends to them, is corrupted. The dashed arrow restores one clean state, here at the subject's last token in a middle layer, and everything after it is recomputed. The pill-shaped “early site” and “late site” mark the two places the paper finds restoration works: middle layers at the last subject token, and late layers at the last token. The curved arrow is the example flow: once the early site is clean again, attention carries the fact down to “downtown”, and the output recovers.

Tiny example: the two effects

This page's illustrative single prompt: the clean run gives “Seattle” probability 0.90, the corrupted run 0.10, and restoring one middle-layer state at the subject's last token brings it back to 0.60.

In words: “the total effect is how much the noise costs: clean probability minus corrupted probability. The indirect effect of one state is how much restoring just that state gives back: restored probability minus corrupted probability.” Averaged over many prompts they become the ATE and the AIE.

With the numbers: TE = 0.90 − 0.10 = 0.8; IE = 0.60 − 0.10 = 0.5. This one state recovers 0.5 / 0.8 = 62.5% of what the noise took away.

In Python:

# illustrative probabilities of "Seattle" for one prompt
P_clean = 0.90
P_corrupt = 0.10
# restoring one clean state: at the last subject token, a middle layer
P_restored = 0.60
# TE = P[o] − P*[o]
round(P_clean - P_corrupt, 2)  # → 0.8
# IE = P*,clean h[o] − P*[o]
round(P_restored - P_corrupt, 2)  # → 0.5
# as a share of the total effect
round((P_restored - P_corrupt) / (P_clean - P_corrupt), 3)  # → 0.625

Why it matters

This is exactly the lesson's “fraction restored”, computed by fraction_restored, with probabilities instead of a logit difference and noise instead of a swapped word. The paper found a third quantity, the direct effect, too noisy to be useful. It also found the method more informative than gradient-based saliency (Appendix B).

2.2 Causal tracing results · original

Everyday picture

Repeat the restoration for every token and every layer, average over 1,000 facts GPT-2 XL knows, and you get a heat map of where facts live. It shows two bright spots, and one of them was a surprise.

The numbers from the text

  • Over the 1,000 facts, the average total effect is 18.6% (in Appendix B.1, the correct token's probability averages 27.0% clean and 8.47% corrupted).
  • At the last subject token, single states reach an AIE of 8.7% at layer 15: the early site, the new discovery.
  • At the last token of the prompt, late layers are strongly causal: the late site, unsurprising, since that is where the prediction is read out.
  • Split by module: MLP contributions peak at an AIE of 6.6% at the early site, while attention there reaches only 1.6%; attention matters at the late site instead. Because restoring one MLP or attention output alone had little effect, these module traces restore a window of ten consecutive layers at once.

Diagram: the average trace, three ways

Try it: switch between restoring whole states, MLP outputs and attention outputs, and hover (or use the arrow keys on) the map to read a value. Watch where the bright band sits in each view.

Rows, top to bottom: first subject token, middle subject tokens, last subject token, first token after the subject, further tokens, last token. Columns: layers 0 to 47.

Hover the map.

Reading it: this is an illustrative map shaped to the paper's description, not its data: the three peak values the text reports are placed where the text says (states 8.7% at layer 15 and MLPs 6.6% around layer 17, both at the last subject token; attention 1.6% there), and the rest is smooth filler. Each cell is the average indirect effect, in percentage points, of restoring that layer at that position; darker red is a bigger effect. In the states view there are two bands: the early site across middle layers of the last-subject-token row, and the late site in the last row at high layers, which rises to the whole total effect because restoring the final state restores everything. The MLP view keeps only the early band; the attention view keeps only the late one. That split is the paper's main finding: MLPs do the recalling while the subject is being read, attention moves the result to where it is needed.

Severing the MLPs

To test that the early site really depends on the MLPs, the paper modifies the graph: restore a clean state as before, but freeze every MLP output at that token in its corrupted value, so the clean state cannot pass through later MLPs. The lowest layers then lose their effect, while higher layers barely change. So a clean state at an early layer matters because the MLPs after it use it; by the upper middle layers the MLP work is done. Freezing attention instead shows no such gap. The paper's summary is its hypothesis: a localized middle-layer MLP key–value mapping recalls facts about the subject.

Why it matters

The same picture appears in GPT-J and GPT-NeoX (Appendix B.3), and in smaller GPT-2s, though the peak layers move from model to model. It is a claim about where a computation happens, found by intervening rather than by looking, the rung the lesson calls intervention.

2.3 The localized factual association hypothesis · original

“We conjecture that any fact could be equivalently stored in any one of the middle MLP layers.”Meng et al. (2022), §2.3

Everyday picture

A receptionist hears a visitor's name, looks it up in a card file, and writes the relevant facts on a sticky note; later, a clerk at the end of the corridor reads the note to answer a question. The card file is the middle-layer MLPs, consulted while the model reads the subject's last token; the clerk is attention at the high layers.

subject tokensThe Space Need le middle-layer MLPsat the last subject token:look up the subject's properties high-layer attentioncopies them to the last token last token “downtown”predicts Seattle localized by module, layer and token

Tap a step, top to bottom.

The paper's hypothesis from §2.3, drawn by this page.

Reading it: read top to bottom. The hypothesis pins factual recall down three ways: in the MLP modules (not attention), at a range of middle layers (not early or late), and at the subject's last token (not the other tokens). Each middle MLP adds a little, and the sum accumulates in the last subject token's residual stream; high attention layers then copy the accumulated information to the end. The paper goes further: since reordering a transformer's middle layers changes little (a finding it cites), it conjectures that any single middle MLP could hold any fact. That is testable: pick one layer and try to write a new fact into it.

Why it matters

It connects two earlier observations: that MLP layers behave like memories, and that attention's job is largely copying information between positions. And it turns an observation about activations into a prediction about weights, which §3 tests.

3 Interventions on weights for understanding factual association storage · original

Everyday picture

Finding the right wire is one thing; proving you understand it is rewiring it and having the light switch correctly from every switch in the house, and no other light change. The test: replace a true fact (Space Needle, located in, Seattle) with a false one (Space Needle, located in, Paris), so the change works for any phrasing (generalization) but leaves other subjects alone (specificity). Earlier work observed that an MLP's first layer acts like a set of keys and its second like the values they retrieve; this paper treats the second layer as a single linear memory rather than a set of separate neurons.

3.1 Rank-one model editing: the MLP as an associative memory · original

Everyday picture

Any matrix can act as a lookup table. Multiply it by a key vector and out comes a value vector. With many key–value pairs, the best matrix is the one that gets all of them as right as possible: least squares, solved by the pseudoinverse. Now add one new pair that must be stored exactly. You could change the matrix any number of ways; the right way changes it least on the keys it already stores, and especially on the keys it sees most often. ROME uses the average statistics of keys, their uncentred covariance C, to know which directions are crowded.

Tiny example

This page's illustrative memory in two dimensions. It stores key k₁ = (1, 0) three times and k₂ = (0, 1) once, each as itself, so W is the identity. The new association: key k* = (1, 1) should give v* = (2, 0).

In words: “keep the memory as good as possible on all the old keys, subject to storing the new pair exactly; the answer is the old matrix plus one column-times-row table: the column is the new value's current error, scaled; the row is the new key reshaped by the inverse key covariance, which steers the change away from directions where common keys live.” This is the paper's equation (2) with the two definitions that follow it.

With the numbers: C = [[3, 0], [0, 1]], so u = C⁻¹k* = (1/3, 1) and uᵀk* = 4/3. The new value's error is v* − Wk* = (2, 0) − (1, 1) = (1, −1), so Λ = (3/4, −3/4), and Ŵ = [[5/4, 3/4], [−1/4, 1/4]]. Check: Ŵk* = (2, 0). The old memories move by a total squared amount of 1.5; a plain rank-one edit that ignores C moves them by 2.0, because it disturbs the thrice-stored k₁ four times as much.

In Python:

from fractions import Fraction as Fr
# old keys: k1 = (1, 0) stored three times, k2 = (0, 1) once; W = I stores each as itself
K = [(1, 0), (1, 0), (1, 0), (0, 1)]
W = [[Fr(1), Fr(0)], [Fr(0), Fr(1)]]
# C = K Kᵀ: add up k kᵀ over the stored keys
C = [[sum(Fr(k[r] * k[c]) for k in K) for c in range(2)] for r in range(2)]
C  # → [[Fraction(3, 1), Fraction(0, 1)], [Fraction(0, 1), Fraction(1, 1)]]
det = C[0][0] * C[1][1] - C[0][1] * C[1][0]
C_inv = [[C[1][1] / det, -C[0][1] / det], [-C[1][0] / det, C[0][0] / det]]
# the new association: key k* = (1, 1) should now give v* = (2, 0)
k_star, v_star = [1, 1], [2, 0]
# u = C⁻¹ k*
u = [sum(C_inv[r][c] * k_star[c] for c in range(2)) for r in range(2)]
[str(x) for x in u]  # → ['1/3', '1']
# Λ = (v* − W k*) / (uᵀ k*)
Wk = [sum(W[r][c] * k_star[c] for c in range(2)) for r in range(2)]
Lam = [(v_star[r] - Wk[r]) / sum(u[c] * k_star[c] for c in range(2)) for r in range(2)]
[str(x) for x in Lam]  # → ['3/4', '-3/4']
# Ŵ = W + Λ uᵀ
W_hat = [[W[r][c] + Lam[r] * u[c] for c in range(2)] for r in range(2)]
[[str(x) for x in row] for row in W_hat]  # → [['5/4', '3/4'], ['-1/4', '1/4']]
# the new memory is stored exactly: Ŵ k* = v*
[sum(W_hat[r][c] * k_star[c] for c in range(2)) for r in range(2)]  # → [Fraction(2, 1), Fraction(0, 1)]
# damage to the old memories: Σ ‖Ŵ k − W k‖² over the stored keys
def damage(M):
    return sum(sum((sum((M[r][c] - W[r][c]) * k[c] for c in range(2))) ** 2 for r in range(2)) for k in K)
float(damage(W_hat))  # → 1.5
# a plain rank-one edit that ignores C (u = k* / ‖k*‖²) damages more
u0 = [Fr(1, 2), Fr(1, 2)]
Lam0 = [v_star[r] - Wk[r] for r in range(2)]
W0 = [[W[r][c] + Lam0[r] * u0[c] for c in range(2)] for r in range(2)]
float(damage(W0))  # → 2.0

Try it: crowded keys

Change how often the old key k₁ was stored and the direction of the new key, and compare how much each edit disturbs the old memories. Both edits store the new pair exactly.

3 45°

Reading it: the old memory stores k₁ = (1, 0) as many times as the first slider says and k₂ = (0, 1) once; the new key k* has length √2 at the chosen angle from k₁ (45° is the worked example) and must map to (2, 0). The solid bar is the total squared disturbance to the old memories from a plain rank-one edit along k*; the striped bar is ROME's. With k₁ stored once the two memories are equally common, C is the identity, and the edits coincide. The more copies of k₁, the more ROME tilts its change away from k₁'s direction and the bigger its advantage. At 0°, the new key points exactly along k₁, and no edit can avoid changing what k₁ retrieves: the bars meet again. At 90° it points along k₂, and both edits leave k₁ untouched. The readout shows ROME's trade: it disturbs the rarely stored k₂ more than the plain edit does, in exchange for protecting the common k₁, and comes out ahead overall.

Step 1: the key, from the subject

The key is the MLP's hidden activation (after the nonlinearity) at the subject's last token, in the chosen layer: the vector (1, 0, 0) of §2. It depends a little on the words before the subject, so ROME averages it over several random prefixes. Illustrative: three prefixes give keys (1.0, 0.2), (0.8, 0.4) and (1.2, 0.0).

In words: “run several texts that end with the subject, read the MLP's hidden activation at the subject's last token each time, and average them.”

With the numbers: ((1.0 + 0.8 + 1.2) / 3, (0.2 + 0.4 + 0.0) / 3) = (1.0, 0.2).

In Python:

# the key at the subject's last token, after three different prefixes (illustrative)
keys = [[1.0, 0.2], [0.8, 0.4], [1.2, 0.0]]
N = len(keys)
# k* = (1/N) Σ_j k(x_j + s)
[round(sum(k[i] for k in keys) / N, 6) for i in range(2)]  # → [1.0, 0.2]

Step 2: the value, by optimization

The value v* is found, not written by hand: search for a vector z that, placed as the MLP's output at the subject's last token, makes the model predict the new object after each prefix, while keeping what the model says about the subject's nature (“{subject} is a”) close to before. The second term prevents essence drift. Illustrative: after two prefixes the new object has probability 0.5 and 0.25; the “is a” predictions are (0.7, 0.2, 0.1) edited and (0.6, 0.3, 0.1) originally.

In words: “make the new object likely after every prefix (the first term), and keep the model's idea of what the subject is unchanged (the second term, a KL divergence); the z that minimizes both is v*.” The optimization changes no weights; it only finds the vector the weights should produce.

With the numbers: (−ln 0.5 − ln 0.25) / 2 = (0.693 + 1.386) / 2 = 1.04; the KL term is 0.7 ln(0.7/0.6) + 0.2 ln(0.2/0.3) = 0.0268; 𝓛 = 1.067.

In Python:

import math
# (a): probability of the new object o* after each prefix, with z written into the MLP output
p_new = [0.5, 0.25]
nll = sum(-math.log(p) for p in p_new) / len(p_new)
round(nll, 3)  # → 1.04
# (b): KL between next-token predictions for "{subject} is a", edited against original
P_edit = [0.7, 0.2, 0.1]
P_orig = [0.6, 0.3, 0.1]
kl = sum(p * math.log(p / q) for p, q in zip(P_edit, P_orig))
round(kl, 4)  # → 0.0268
round(nll + kl, 3)  # → 1.067

Step 3: insert

With k* and v* in hand, equation (2) gives the new Wproj in closed form. In the paper's setup (Appendix E.5): layer 18 of GPT-2 XL; C estimated from 100,000 keys collected on Wikipedia text; v* found with Adam in at most 20 steps; the whole edit about 2 seconds on one GPU.

Why it matters

The edit is one outer product added to one matrix, derived rather than trained, which is what makes it a test of understanding: if the key–value picture were wrong, a single rank-one change would not generalize to paraphrases. Nothing in this primer's code edits weights this way; the closest neighbour is the lesson's LandmarkModel, whose middle MLP is a hand-written lookup from landmark to city.

3.2 Evaluating ROME on zero-shot relation extraction · original

The first test uses zsRE, as used by earlier editors: 10,000 records, each with a fact to insert, a paraphrase of it, and an unrelated fact. The baselines are fine-tuning one layer (FT), fine-tuning with a limit on how far any weight may move (FT+L), and two hypernetworks that learn to predict weight edits, KE and MEND, also retrained on zsRE itself (the “-zsRE” rows).

Selected rows of Table 1 (GPT-2 XL, accuracy %, ± 95% interval), Meng et al. (2022)
EditorEfficacyParaphraseSpecificity
GPT-2 XL, unedited22.2 (±0.5)21.3 (±0.5)24.2 (±0.5)
FT99.6 (±0.1)82.1 (±0.6)23.2 (±0.5)
MEND-zsRE99.4 (±0.1)99.3 (±0.1)24.1 (±0.5)
ROME99.8 (±0.0)88.1 (±0.5)24.2 (±0.5)

ROME, which needs no training at all, is competitive; the hypernetwork trained on zsRE's own data generalizes to paraphrases better. The authors note that zsRE's specificity score barely moves for any method, because its unrelated facts are random and edits bleed mostly into similar subjects. That motivates the next benchmark.

3.3 Evaluating ROME on CounterFact · original

Everyday picture

Teaching a model a fact it half-believes already is easy. CounterFact uses counterfactuals the model starts out rating unlikely, and measures three things per edit: did it take (efficacy), does it hold under rewording (paraphrase), and did it leave similar subjects alone (neighbourhood)? For the Eiffel Tower moved to a new city, the neighbours are other things in Paris, such as the Louvre, found through Wikidata.

The scores

  • ES, efficacy score: the share of edits after which P[o*] > P[oc]; EM is the mean gap.
  • PS, paraphrase score: the same on reworded prompts; PM its mean gap.
  • NS, neighbourhood score: the share of neighbouring subjects that still prefer the true object; NM its mean gap.
  • RS, consistency of generated text with the new fact (a cosine similarity of word statistics against reference texts), and GE, fluency (below).

Tiny example: one score from three

The overall score S is the harmonic mean of ES, PS and NS. The numbers are ROME's and FT's on GPT-2 XL, from Table 4.

In words: “average the reciprocals of the three scores and take the reciprocal of that”: a mean that is dragged down hard by whichever score is worst. The paper names it; the formula is the standard definition.

With the numbers: ROME: 3 / (1/100 + 1/96.4 + 1/75.4) = 89.2, as printed. FT: 3 / (1/100 + 1/87.9 + 1/40.4) = 65.0; the table prints 65.1, presumably computed before its inputs were rounded. An ordinary average would give FT 76.1 and hide how badly it damages neighbours.

In Python:

# S = harmonic mean of ES, PS, NS (Table 4, GPT-2 XL)
def S(es, ps, ns):
    return 3 / (1 / es + 1 / ps + 1 / ns)
round(S(100.0, 96.4, 75.4), 1)  # → 89.2
# FT: the table prints 65.1, computed before its inputs were rounded
round(S(100.0, 87.9, 40.4), 1)  # → 65.0
# the ordinary average would hide FT's failure on neighbours
round((100.0 + 87.9 + 40.4) / 3, 1)  # → 76.1

Try it: why a harmonic mean

Drag the three scores. Watch how far a single low score pulls the harmonic mean compared with the ordinary average.

100 88 40

Reading it: the sliders are the three success rates of one editor, starting near FT's (100, 88, 40). The solid bar is their ordinary average, the striped bar the harmonic mean S. Drop any one slider toward zero and S follows it down, however high the other two are, while the average barely notices. An editor can only score well on S by being effective, general and specific at once, which is the point of the benchmark: the paper's two failure modes, overfitting to one phrasing and bleeding into neighbours, each sink one of the three.

Fluency, measured by repetition

GE is a weighted average of bigram and trigram entropy of generated text: it drops when an edit makes the model repeat itself. Illustrative: five repetitions of one word, against five different words.

In words: “for each distinct n-gram, multiply its share of all n-grams by the base-2 log of that share, add up, and flip the sign”: zero when one n-gram repeats, larger when many different ones appear.

With the numbers: “medicine” five times gives one bigram with share 1 and entropy 0; “he studied medicine in paris” gives four distinct bigrams, each ¼, and entropy 2 bits.

In Python:

import math
from collections import Counter
def bigram_entropy(words):
    # f(k): how often each bigram k occurs, as a share
    grams = Counter(zip(words, words[1:]))
    total = sum(grams.values())
    # −Σ_k f(k) log₂ f(k); adding 0.0 turns a −0.0 into 0.0
    return -sum(c / total * math.log2(c / total) for c in grams.values()) + 0.0
bigram_entropy("medicine medicine medicine medicine medicine".split())  # → 0.0
bigram_entropy("he studied medicine in paris".split())  # → 2.0

Why it matters

The dataset has 21,919 records, with two paraphrase prompts, ten neighbourhood prompts and several generation prompts each, built from the ParaRel dataset. Its design, and especially its neighbourhood prompts, is what exposes the failures the next section shows.

3.4 Confirming the importance of decisive states · original

Everyday picture

If tracing found the right place, editing there should work best. So the authors ran ROME at every layer and every token position of GPT-2 XL. Edits at the subject's last token do best, with both generalization and specificity peaking at the middle layers, and generalization peaking at layer 18, inside the early site. Edits at other tokens generalize or stay specific poorly. (Appendix I adds that editing the late attention layers instead makes the model repeat the new fact only in the trained wording.)

Diagram: the editors compared

Reading it: selected rows of the paper's Table 4, for GPT-2 XL on 7,500 CounterFact records; the axis runs from 0 to 100%. For each editor, the solid bar is efficacy (ES), the striped bar paraphrase (PS), and the grey striped bar neighbourhood (NS). Look for the editor whose three bars are all long. FT has perfect efficacy but keeps only 40.4% of neighbours correct: it bleeds. FT+L protects neighbours but generalizes to under half the paraphrases: it overfits one wording. KE and MEND show both problems at once. ROME's three bars are all long, which is why its S is 89.2 against the next best, FT+L's 66.9. The unedited model's NS of 78.1 is the ceiling to compare neighbourhood scores with: ROME's 75.4 is close. On GPT-J, with 2,000 records, ROME scores S = 91.5.

The paper names the two failure modes (F1) overfitting to the counterfactual statement and failing to generalize, and (F2) underfitting and predicting the new object for unrelated subjects. Every method but ROME shows one or both. Knowledge Neurons (KN), which edits the MLP rows of neurons picked out by gradients, fails both ways on this benchmark.

3.5 and 3.6 Generated text, and people's ratings · original

Insert “Pierre Curie's area of work is medicine” (he was a physicist) and let each edited model write. FT and ROME describe him as a physician under every wording; FT+L, KE and MEND switch between medicine and physics depending on the phrasing, and KE repeats “medicine” to the point of nonsense. On specificity, GPT-2 XL already (wrongly) called Robert Millikan an astronomer; after the Curie edit, FT+L calls him a biologist and KE and MEND a medical scientist, while ROME leaves him unchanged. Appendix G adds cases where edits compose with other knowledge (after moving Liberty Island to Scotland, a ROME-edited model connects it to Loch Lomond) and cases of essence drift.

Human evaluation: 15 volunteers compared generated text on 50 inserted facts. They were 1.8 times more likely to judge ROME's text more consistent with the new fact than FT+L's, but 1.3 times less likely to judge it more fluent: a loss of fluency the automatic entropy score did not catch.

3.7 Limitations · original

“it only edits a single fact at a time, and it is not intended as a practical method for large-scale model training.”Meng et al. (2022), §3.7
  • One fact per edit, and one direction: “the Space Needle is in Seattle” and “Seattle's iconic landmark is the Space Needle” are stored separately and need two edits. The authors point to their follow-up work for editing many facts at once.
  • Only factual associations: logical, spatial or numerical knowledge were not studied, and the structure of the vector spaces involved is not well understood.
  • Plausible guessing: even after a successful edit, the model will invent plausible new facts with no basis, which limits its use as a source of facts.

4 Related work · original

Three lines of prior work

  • What representations encode. Probing classifiers read properties from hidden states but can be dissociated from what the network actually does. Causal methods avoid that: earlier work used causal mediation analysis on individual neurons (gender bias, syntactic agreement) and erased information to measure its effect. Causal tracing adds paired interventions that measure a single hidden state's indirect effect.
  • How much knowledge models hold: fill-in-the-blank prompts, better and more varied prompts, and ParaRel, the paraphrase dataset CounterFact is built from. This paper asks how knowledge is recalled rather than how much can be extracted.
  • Localizing and editing knowledge: MLP layers as key–value memories; Knowledge Neurons, which writes the object's embedding into selected MLP rows; KE and MEND, hypernetworks that predict weight updates. ROME is compared with all of them.

This primer's probes and patching sit on exactly this reading-versus-intervening divide.

5 Conclusion and ethical considerations · original

“we stress that large language models should not be used as an authoritative source of factual knowledge in critical settings.”Meng et al. (2022), §6

The paper clarifies how information flows when a GPT recalls a fact, and turns that into a simple editor. It frames ROME as a tool for testing where knowledge is stored, not a product. The ethics section is two-sided: understanding a model's organization improves transparency and makes fixing errors cheaper; but the same ability can insert misinformation, bias or other adversarial content, which, together with the guessing behaviour above, is why the authors warn against treating a language model as an authoritative source.

Appendix A: solving for Λ algebraically · original

Everyday picture

This is the classic problem of least squares with one equality constraint. The old matrix W already solves the unconstrained problem. Adding the constraint with a Lagrange multiplier Λ and setting the slope to zero gives two equations, one for the old optimum and one for the new; subtracting them, almost everything cancels.

Tiny example

Reuse §3.1's numbers: C = [[3, 0], [0, 1]], k* = (1, 1), Λ = (3/4, −3/4), and the change Ŵ − W = Λuᵀ with u = (1/3, 1).

In words: “the old matrix satisfies the normal equations; the new one satisfies them plus a term from the constraint; subtract, and the change times C equals Λ times the new key laid on its side; multiply by C's inverse to get the rank-one update.” Then substituting into Ŵk* = v* gives Λ = (v* − Wk*) / ((C⁻¹k*)ᵀk*), equation (26).

With the numbers: Λuᵀ = [[1/4, 3/4], [−1/4, −3/4]]; times C, that is [[3/4, 3/4], [−3/4, −3/4]], which is exactly Λk*ᵀ.

In Python:

from fractions import Fraction as Fr
# the numbers from §3.1: C, k*, Λ and the change Ŵ − W = Λ uᵀ
C = [[Fr(3), Fr(0)], [Fr(0), Fr(1)]]
k_star = [1, 1]
Lam = [Fr(3, 4), Fr(-3, 4)]
u = [Fr(1, 3), Fr(1)]
dW = [[Lam[r] * u[c] for c in range(2)] for r in range(2)]
# left side of (12): (Ŵ − W) C
left = [[sum(dW[r][j] * C[j][c] for j in range(2)) for c in range(2)] for r in range(2)]
# right side of (12): Λ k*ᵀ
right = [[Lam[r] * k_star[c] for c in range(2)] for r in range(2)]
left == right  # → True
[[str(x) for x in row] for row in left]  # → [['3/4', '3/4'], ['-3/4', '-3/4']]

Why it matters

The derivation assumes C is invertible and that W was the least-squares memory for the keys C summarizes. In practice C is estimated from keys on Wikipedia text, which is why it stands for “the keys this layer usually sees”. The Frobenius norm in the paper's version is just the squared error summed over every entry of the matrix.

Appendices B and I: robustness checks · original

Is the trace an artefact of the noise?

The authors vary the corruption: also corrupting the token after the subject, drawing noise from a Gaussian matched to the embeddings' mean and covariance, or from a uniform distribution. The early site at the last subject token persists every time. The noise must be large, though: with a spread of just one standard deviation, the total effect becomes too small to read indirect effects from. Gradient saliency (integrated gradients) on the same prompts gives scattered maps that show neither the last subject token's importance nor the middle MLPs.

Is the last subject token always the one?

Not always. For “Windows Media Player”, the first word “Windows” triggers the decisive lookup; for “Mitsubishi Electric”, “Electric” doesn't matter; for “Madame de Montesson”, the title “Madame” does most of the work. Some low-confidence facts show no decisive MLP lookup inside the subject at all.

Other models

GPT-J (6B, 28 layers) and GPT-NeoX (20B, 44 layers), with noise rescaled to their embeddings, show the same early and late sites, with attention playing a larger role at the first layers of the last subject token, perhaps because fewer layers force the work into fewer places. Across GPT-2 Medium, GPT-2 Large and NeoX, early-site MLPs keep large effects, though the peak layer differs.

Editing attention instead (Appendix I)

Constrained fine-tuning of the query, key and value weights of the attention at layer 33 (the peak of the attention trace) makes the model repeat the new fact when given the exact training prompt, but not under paraphrase: word prediction rather than recall. ROME at the MLP generalizes. It supports the division of labour in the hypothesis.

Where it leads

Idea in the paperWhy it lastsWhere to build it
Clean run, corrupted run, restore one activationThe basic causal experiment of interpretabilitypatching every site, a run that accepts patches
Indirect effect as the measureSeparates what a site carries from what reaches the outputfraction restored
MLPs recall, attention movesA reusable picture of how facts flow through a transformerthe landmark model, the transformer lesson
A linear layer as a key–value memory, edited in closed formModel editing without retrainingNo code in this primer; the worked example above
Test edits on paraphrases and on neighboursGeneralization and specificity, reported togetherevaluations
Features as the unit of interventionA later route from “where” to “what”Scaling Monosemanticity companion

Glossary

Every term with hover guidance on this page, in one place.