Editing Models with Task Arithmetic, annotated
How to read this page
- Any dotted word explains itself on hover, focus or tap; so does every symbol in every equation.
- The weight-space playground in §1 is the paper in one picture: pick an operation, drag λ, and watch the edited model move.
Each idea climbs the same ladder: everyday picture, tiny example, diagram, the math, why it matters today. The fine-tuning lesson builds task arithmetic from scratch on a tiny network: task_vector subtracts the base, merge adds scaled task vectors back, and merge_run measures what the merged model can do.
Abstract · original
“A task vector specifies a direction in the weight space of a pre-trained model, such that movement in that direction improves performance on the task.”Ilharco et al. (2022), Abstract
Everyday picture
Two editors take copies of the same manuscript. One fixes the grammar, the other tightens the argument, and each hands back a list of tracked changes. You can accept both lists onto the original; you can accept one list backwards to undo it; and if you know how a British edit relates to an American one for chapter 1, you can guess the change for chapter 2. Task arithmetic does this with a model's weights: the “tracked changes” of a fine-tune are its task vector.
What the paper claims
- Negation forgets. Subtracting a task vector lowers performance on that task and leaves unrelated tasks nearly untouched: it cut a language model's toxic generations from 4.8% to 0.8%.
- Addition combines. Adding task vectors from different tasks gives one model good at all of them; for pairs of image tasks it keeps 98.9% of the specialists' accuracy.
- Analogies transfer. When tasks relate as “A is to B as C is to D”, combining the vectors of A, B and C improves D, even with no labelled data for D.
- All of it is elementwise arithmetic on weights: no extra training, no extra cost at inference.
Why it matters today
This paper gave model merging its vocabulary. Whenever someone combines fine-tunes of a shared base, or subtracts one to remove a behaviour, they are doing task arithmetic.
1 Introduction · original
“There is no extra cost at inference time in terms of memory or compute, since we only do element-wise operations on model weights.”Ilharco et al. (2022), §1
Everyday picture
Picture every possible model as a point on a huge map, one direction for every weight: that map is weight space. The pretrained model is one point. Fine-tuning walks from that point to another; the walk, as an arrow, is the task vector. Arrows can be scaled, reversed and laid end to end. The paper's claim is that doing so to the arrows does predictable things to the skills.
Tiny example: a two-weight model
Real models have millions or billions of weights, but two are enough to draw. Start at θpre = (0, 0). Fine-tuning on task A moves the weights to (1.5, 0.5), so τA = (1.5, 0.5); fine-tuning on B moves them to (0.3, 1.4), so τB = (0.3, 1.4). Adding both at λ = 1 lands at (1.8, 1.9); negating A lands at (−1.5, −0.5). These numbers are illustrative.
Try it: Figure 1 as a playground
Two weights and these vectors are illustrative; the operations are the paper's.
Reading it: the two axes are two of the model's weights, and the ringed dot at the centre is the pretrained model. Arrows are task vectors; the solid purple arrow is the edit actually applied, and the filled dot at its tip is the edited model. (a) At λ = 1 the tip is exactly the fine-tuned model, and at λ = 0 it is the pretrained one: task vectors can only reproduce, interpolate or extrapolate a fine-tune, never invent a new one. (b) Negation points the arrow the other way: the model moves away from what the fine-tune learned. (c) Addition lays the two arrows end to end (the dashed parallelogram); the tip is one model meant to do both tasks. (d) The analogy starts from C and adds the step from A to B. Drag λ in each mode: it shrinks or stretches the whole edit, and the paper picks it on held-out data every time.
Why it matters today
Once a fine-tune is an arrow, a library of fine-tunes becomes a set of parts. Publicly shared fine-tunes of the same base model can be combined by anyone, without the data they were trained on.
2 Task vectors · original
“Note that adding a single task vector to a pre-trained model with λ = 1 results in the model fine-tuned on that task.”Ilharco et al. (2022), §2
Everyday picture
A task vector is a diff: take the fine-tuned weights, subtract the weights you started from, and what is left is exactly what training changed. The paper defines a task as a dataset plus a loss function to fine-tune on.
Tiny example
The fine-tuning lesson's three-weight model: θpre = (1, 0, 2), and after fine-tuning on A, θft = (1.5, 0, 2). The task vector is (0.5, 0, 0): only the first weight moved. Applying it at λ = 1 gives back (1.5, 0, 2); at λ = 0.5, (1.25, 0, 2); negated, (0.5, 0, 2).
In words: “the task vector for task t is the fine-tuned weights minus the pretrained weights, one weight at a time.”
With the numbers: (1.5, 0, 2) − (1, 0, 2) = (0.5, 0, 0).
In words: “to edit a model, add λ times a task vector to its weights.”
With the numbers: (1, 0, 2) + 1 × (0.5, 0, 0) = (1.5, 0, 2), the fine-tuned model; λ = 0.5 gives (1.25, 0, 2), halfway between.
In Python:
theta_pre = [1.0, 0.0, 2.0]
theta_ft = [1.5, 0.0, 2.0]
# τ = θ_ft − θ_pre, weight by weight
tau = [f - p for f, p in zip(theta_ft, theta_pre)]
tau # → [0.5, 0.0, 0.0]
# θ_new = θ + λτ
def apply(theta, tau, lam):
return [w + lam * t for w, t in zip(theta, tau)]
apply(theta_pre, tau, 1.0), apply(theta_pre, tau, 0.5), apply(theta_pre, tau, -1.0) # → ([1.5, 0.0, 2.0], [1.25, 0.0, 2.0], [0.5, 0.0, 2.0])
The three operations
Every edit in the paper builds a new vector τnew and applies it with θnew = θ + λτnew, with λ chosen on a held-out validation set.
In words: “to forget a task, subtract its vector; to combine tasks, add their vectors; to reach task D from an analogy, start at C's vector and add the step that takes A to B.”
With the numbers: in the playground's two weights, τA = (1.5, 0.5) and τB = (0.3, 1.4): negation gives (−1.5, −0.5), addition gives (1.8, 1.9), and with τC = (−0.6, −1.1), the analogy gives (−0.6, −1.1) + (−1.2, 0.9) = (−1.8, −0.2).
In Python:
tau_A, tau_B, tau_C = [1.5, 0.5], [0.3, 1.4], [-0.6, -1.1]
# negation: τ_new = −τ_A
[-t for t in tau_A] # → [-1.5, -0.5]
# addition: τ_new = Σ_i τ_i
[round(a + b, 2) for a, b in zip(tau_A, tau_B)] # → [1.8, 1.9]
# analogy: τ_new = τ_C + (τ_B − τ_A)
[round(c + (b - a), 2) for a, b, c in zip(tau_A, tau_B, tau_C)] # → [-1.8, -0.2]
Why it matters today
The whole method needs one condition: every model involved shares an architecture and starts from the same pretrained weights, so the diffs line up weight by weight. The paper also restricts itself to fine-tunes that add no new parameters (such as a new classification head); open-vocabulary image classifiers like CLIP and text-to-text models like T5 qualify. In the fine-tuning lesson, task_vector and merge are these two equations in NumPy.
3 Forgetting via negation · original
“Negating a task vector decreases performance on the target task, with little change in model behavior on control tasks.”Ilharco et al. (2022), Abstract
Everyday picture
If practising scales made you better at piano, walking the same distance in the opposite direction should make you worse at scales, and only at scales. Negation is that walk. The hard part is checking “only”: every experiment measures a control task the edit should leave alone.
Tiny example: choosing λ
For images, the paper tries λ = 0, 0.05, 0.1, …, 1 and keeps the largest λ whose model still scores at least 95% of the pretrained model's control accuracy. Illustrative control accuracies at λ = 0, 0.25, 0.5, 0.75, 1: 75.5, 74.9, 73.8, 72.1, 70.2. The bar is 0.95 × 75.5 = 71.725, so λ = 0.75 is kept: the strongest forgetting the control can afford.
In words: “push the negated vector as hard as you can, as long as the control task keeps at least 95% of its original accuracy.”
With the numbers: 72.1 ≥ 71.725 at λ = 0.75, but 70.2 < 71.725 at λ = 1, so λ* = 0.75.
In Python:
lams = [0, 0.25, 0.5, 0.75, 1.0]
# illustrative control accuracy of θ_pre − λτ at each λ
A_ctrl = [75.5, 74.9, 73.8, 72.1, 70.2]
# the bar: 95% of the pretrained model's control accuracy
bar = 0.95 * A_ctrl[0]
round(bar, 3) # → 71.725
# λ* = the largest λ that clears the bar
max(lam for lam, acc in zip(lams, A_ctrl) if acc >= bar) # → 0.75
The two baselines, decoded
Negation is compared with two other ways to make a model worse at a task. Gradient ascent fine-tunes to raise the loss instead of lowering it (gradient ascent). A random vector with the same size as the task vector, layer by layer, checks whether “moving the weights this far in any direction” would do the same damage.
Tiny example for gradient ascent: two training examples where the model gives the right label probability 0.8 and 0.5. The ordinary cross-entropy loss is (−log 0.8 − log 0.5) / 2 = 0.458; gradient ascent minimises its negative, −0.458, which it can push towards minus infinity by driving those probabilities to zero.
In words: “the usual loss is the average surprise at the right label; gradient ascent minimises minus that, so it rewards the model for being as wrong as possible.”
With the numbers: ℓ = (0.223 + 0.693) / 2 = 0.458 and ℓneg = −0.458. Nothing stops ℓneg falling forever, which is why gradient ascent wrecks everything else too.
In Python:
import math
# f(x)_y: probability of the right label on two examples
f_xy = [0.8, 0.5]
# ℓ = average of −log f(x)_y
ell = sum(-math.log(p) for p in f_xy) / len(f_xy)
round(ell, 3), round(-ell, 3) # → (0.458, -0.458)
# as the right-label probability falls, −log grows without limit
[round(-math.log(p), 1) for p in (0.1, 0.001, 1e-9)] # → [2.3, 6.9, 20.7]
Tiny example for the random vector: if one layer's task vector is (3, 4), its length is 5. Draw a random vector, say (1, −2), with length 2.236, and stretch it by 5 / 2.236: (2.236, −4.472), whose length is 5 too. Same size, random direction.
In words: “for each layer, draw a random direction and rescale it to exactly the length of the real task vector's piece for that layer.”
With the numbers: (1, −2) × 5 / 2.236 = (2.236, −4.472), length 5.
In Python:
import math
def norm(v):
return math.sqrt(sum(x * x for x in v))
tau_L = [3.0, 4.0]
tau_rand = [1.0, -2.0]
# ‖τ^(L)‖ and ‖τ_rand^(L)‖
norm(tau_L), round(norm(tau_rand), 3) # → (5.0, 2.236)
# τ_scaled = τ_rand · ‖τ‖ / ‖τ_rand‖
tau_scaled = [x * norm(tau_L) / norm(tau_rand) for x in tau_rand]
[round(x, 3) for x in tau_scaled], round(norm(tau_scaled), 3) # → ([2.236, -4.472], 5.0)
Why it matters today
The two baselines turn a curiosity into evidence. If any large change hurt the task, the random vector would too; it doesn't. If “make the task worse” were the whole job, gradient ascent would win; it destroys the control. Only the task vector's direction carries the task.
3.1 Image classification · original
The setup
Three CLIP models (vision transformers ViT-B/32, ViT-B/16 and ViT-L/14) are fine-tuned on eight image tasks, from satellite images (EuroSAT, RESISC45) to traffic signs (GTSRB), digits (MNIST, SVHN), textures (DTD), cars and scenes (SUN397). The control is ImageNet. Each fine-tune runs 2,000 steps with AdamW at learning rate 10−5, with CLIP's text-derived classification layer frozen so no new parameters appear.
Reading it: Table 1 of the paper for the largest model, ViT-L/14, as bars. For each method, the solid bar is the average accuracy on the eight tasks we want forgotten, and the striped bar is ImageNet accuracy, which we want kept. Good forgetting is a short solid bar next to a long striped one. The fine-tuned model is the opposite of what we want (94.0 target). Gradient ascent empties both bars (3.93 and 16.3). The random vector barely changes anything (60.9 and 72.9). The negative task vector drops the target to 19.0, down 45.8 points from the pretrained 64.8, while ImageNet falls only from 75.5 to 72.9.
What the appendix adds
Forgetting works best when the fine-tune had helped the most: across tasks, the gain from fine-tuning and the drop from negation are positively correlated. It is weakest on Cars and SUN397, whose images resemble ImageNet's, though removing overlapping ImageNet classes barely changed that. It also works on reading text in images (OCR), and less well on recognising celebrities, where fine-tuning itself gained little.
Why it matters today
Negation is a cheap first tool for removing a capability that a fine-tune can isolate. It is also a warning: the edit is only as clean as the control you measured, so a negation must be checked on everything you care about, not just the task you removed.
3.2 Text generation · original
Everyday picture
Teach a model to be rude on purpose, then point that lesson backwards. The paper fine-tunes GPT-2 Large on the most toxic comments from Civil Comments (toxicity score above 0.8), negates the task vector, and asks whether the result is less toxic than the model it started from.
Tiny example: reading the table
Toxicity is measured on 1,000 generations prompted with “I don't care if this is controversial”; fluency is perplexity on WikiText-103 (lower is better). λ is the largest value keeping perplexity within 0.5 of the pretrained model's. 4.8% toxic before, 0.8% after: a sixfold drop, with perplexity moving from 16.4 to 16.9.
| Method | % toxic generations ↓ | Avg. toxicity score ↓ | WikiText-103 perplexity ↓ |
|---|---|---|---|
| Pretrained | 4.8 | 0.06 | 16.4 |
| Fine-tuned (on toxic text) | 57 | 0.56 | 16.6 |
| Gradient ascent | 0.0 | 0.45 | > 1010 |
| Fine-tuned on non-toxic text | 1.8 | 0.03 | 17.2 |
| Random vector | 4.8 | 0.06 | 16.4 |
| Negative task vector | 0.8 | 0.01 | 16.9 |
Reading it: read the table in pairs of columns: toxicity (left two) and fluency (right). Gradient ascent “wins” on toxic generations with 0.0%, but its perplexity above ten billion means it no longer produces language at all. The obvious alternative, fine-tuning on polite comments, is worse than negation on both counts (1.8% toxic, perplexity 17.2). The highlighted row is the only one low on both.
Why it matters today
This is the result people remember: learn a behaviour you do not want, then subtract it. The same idea, finding the direction of an unwanted behaviour and removing it, reappears in later work on steering models. The paper notes it works better for larger models, and that for prompts designed to provoke toxicity there is still much room to improve.
4 Learning via addition · original
Everyday picture
Two specialists each learned one trade in their own workshop. Instead of hiring both, you merge their notebooks into one person's head. It works if their notes are about different things; if both scribbled over the same page, the merge is a mess.
Tiny example: normalized accuracy
Tasks differ in difficulty, so the paper divides each task's accuracy by what that task's own fine-tuned model scores. Illustrative: a merged model scores 91% on task 1, whose specialist scores 95%, and 88% on task 2, whose specialist scores 89%. Normalized, that is 0.958 and 0.989, averaging 0.973: the one merged model keeps 97.3% of what two specialists would give.
In words: “score the edited model on a task as a fraction of what that task's own specialist scores, so 1 means as good as the specialist.”
With the numbers: 91 / 95 = 0.958 and 88 / 89 = 0.989; the average is 0.973.
In Python:
# illustrative (a_edited, a_ft) for two tasks
scores = [(91, 95), (88, 89)]
# a_norm = a_edited / a_ft
a_norm = [a_edited / a_ft for a_edited, a_ft in scores]
[round(a, 3) for a in a_norm] # → [0.958, 0.989]
round(sum(a_norm) / len(a_norm), 3) # → 0.973
Why it matters
Normalizing makes “1.0” mean “as good as keeping every specialist”, so every addition result below is a direct answer to: how much do you lose by collapsing several models into one?
4.1 Image classification · original
What happens
Adding every pair of the eight task vectors gives single models that average 98.9% normalized accuracy on their two tasks, far above the pretrained model. Then the paper adds every subset of the eight, 28 = 256 of them, one scaling coefficient λ per experiment, and scores each on all eight tasks.
Hover or tap to read the average normalized accuracy for each number of task vectors.
Reading it: Figure 3 of the paper, redrawn as its average line. The x-axis is how many task vectors are added; the y-axis is normalized accuracy averaged over all eight tasks, where 1.0 would match keeping eight specialists. With none (the pretrained model) the average is about 0.69; each extra vector lifts it, to about 0.90 with all eight. The paper's plot also shows a dot for every subset, spread around this line; the best of all, 91.2%, is a subset rather than all eight together, and an appendix notes that the best multi-task model often does not use every task vector. The line's values are read off the paper's plot (it prints only 91.2%), so they are approximate.
What the appendix adds
- λ matters and varies. The best coefficient for a sum ranges widely between experiments; values from 0.3 to 0.5 are near-optimal in many cases. Changing λ needs no retraining, only evaluation.
- Training jointly is still better. One model fine-tuned on all eight tasks at once reaches 0.994 normalized, against 0.912 for the best task-vector sum. The trade is flexibility: adding a task to a joint model means training again.
- Random seeds barely matter for these image tasks, and adding an ImageNet task vector to any of the eight gives a model good at both.
Why it matters today
This is the result behind “merge instead of retrain”. The fine-tuning lesson reproduces both halves in miniature: two fine-tunes whose task vectors are nearly perpendicular merge into a model scoring 0.97 on both tasks at λ = 1, while two that contradict each other cannot be merged at any λ (see merge_run).
4.2 Natural language processing · original
Everyday picture
Here addition is used not to build a generalist but to improve a specialist: borrow a neighbour's notebook and see whether it helps with your own task.
Tiny example
T5-base models are fine-tuned on four GLUE text-classification tasks. The authors then searched the Hugging Face Hub for other T5-base fine-tunes, found 427 compatible checkpoints, tried adding each one's task vector to their fine-tuned model, and kept the best checkpoint and λ on held-out data.
| Method | MRPC | RTE | CoLA | SST-2 | Average |
|---|---|---|---|---|---|
| Zero-shot | 74.8 | 52.7 | 8.29 | 92.7 | 57.1 |
| Fine-tuned | 88.5 | 77.3 | 52.3 | 94.5 | 78.1 |
| Fine-tuned + task vectors | 89.3 | 77.5 | 53.0 | 94.7 | 78.6 |
Reading it: each column is one task (CoLA is scored by a correlation, the others by accuracy or an accuracy-like score). The highlighted row beats plain fine-tuning on all four, by 0.2 to 0.8 points, 0.5 on average: small but free, since it costs a search and no training. An appendix adds that pairs of six popular public T5 fine-tunes (sentiment, question answering, summarisation and more) merge into single models keeping 96.7% of the specialists' normalized performance on average.
Why it matters today
This is the first hint of a practice now common with open models: download fine-tunes other people made from the same base, and combine them, without ever seeing their training data.
5 Task analogies · original
“When tasks are linked by an analogy relationship of the form ‘A is to B as C is to D’, combining task vectors from three of the tasks can improve performance on the fourth, even when no data from the fourth task is used for training.”Ilharco et al. (2022), Abstract
Everyday picture
You know how a British speaker's sentence changes when an American says it, and you know the British way to ask for directions. Then you can guess the American way to ask for directions without ever hearing it. Word vectors famously did this with words (king − man + woman ≈ queen); this paper does it with whole tasks.
Tiny example: kings from queens
Illustrative three-weight task vectors, one weight each for “royal”, “female” and “male”: τqueen = (1, 1, 0), τwoman = (0, 1, 0), τman = (0, 0, 1). The analogy τking ≈ τqueen + (τman − τwoman) = (1, 0, 1): royal and male. The paper ran this for real: CLIP models learned a new class named “something” from 50 web images each of queens, kings, men and women, and for every category the other three built its vector. With ViT-L/14, accuracy on each target went from 0 to 100% for queens, kings and women and 96% for men, with ImageNet slipping from 75.5 to about 74.6.
In words: “estimate D's task vector by starting from C's and adding the step that turns A into B.”
With the numbers: (1, 1, 0) + ((0, 0, 1) − (0, 1, 0)) = (1, 0, 1).
In Python:
# illustrative weights: (royal, female, male)
tau_queen, tau_woman, tau_man = [1, 1, 0], [0, 1, 0], [0, 0, 1]
# τ̂_king = τ_queen + (τ_man − τ_woman): A = woman, B = man, C = queen
[c + (b - a) for a, b, c in zip(tau_woman, tau_man, tau_queen)] # → [1, 0, 1]
Two real uses
Domain generalization. To classify the sentiment of Yelp reviews with no labelled Yelp data, take a sentiment task vector learned on labelled Amazon reviews, and add the difference between plain language-model fine-tunes on Yelp text and on Amazon text: τ̂yelp; sent = τamazon; sent + (τyelp; lm − τamazon; lm). Unlabelled text is cheap; labels are not. For T5-small, the analogy lifts Yelp accuracy from 88.6 (fine-tuned on Amazon alone) to 89.9, against 91.1 for a model fine-tuned on labelled Yelp data: it closes (89.9 − 88.6) / (91.1 − 88.6) = 0.52, about half the gap, with no Yelp labels.
Rare subpopulations. Photos and sketches of 125 classes shared by ImageNet and a sketch dataset, split into four groups (think “real dog”, “real lion”, “sketch dog”, “sketch lion”). Averaged over the four targets and three CLIP models, the analogy adds 3.4 percentage points over the pretrained model with no target data, about what collecting and labelling roughly a hundred target examples would buy.
One λ or three?
The paper normally scales the whole combined vector by one λ. An appendix tries a separate coefficient for each vector, which helped by 0.7 points on average, with best values around λB = λC = 0.32 and λA = 0.28, but multiplied the search.
In words: “give each of the three task vectors its own strength, instead of one strength for the whole analogy.”
With the numbers: each λ takes 11 values (0, 0.1, …, 1), so one shared λ means 11 evaluations and three separate ones mean 113 = 1,331. The paper rounds these to 10 and 103.
In Python:
grid = [round(0.1 * k, 1) for k in range(11)]
len(grid) # → 11
# every combination of λ_A, λ_B, λ_C
len([(a, b, c) for a in grid for b in grid for c in grid]) # → 1331
Why it matters today
Analogies are the most surprising result in the paper, and the least used in practice. What survived is the idea behind them: the difference between two fine-tunes can isolate one factor (a domain, a style, a subgroup), and that difference can be moved to a model that never saw the data.
6 Discussion · original
“We observe that vectors from different tasks are typically close to orthogonal, and speculate that this enables the combination of task vectors via addition with minimal interference.”Ilharco et al. (2022), §6
Everyday picture
Two people pushing a heavy crate at right angles do not cancel each other: each gets their push in. Two people pushing along the same line either help or fight. Task vectors that are orthogonal (at right angles) are the first case, which is why adding them mostly works.
Tiny example: the cosine
The cosine similarity measures the angle between two vectors: 1 for the same direction, 0 for right angles, −1 for opposite. For (3, 4, 0) and (0, 4, 3): the dot product is 0 + 16 + 0 = 16, both lengths are 5, so the cosine is 16 / 25 = 0.64. In very high dimensions, unrelated directions are almost always nearly perpendicular: two random 10,000-number vectors have a cosine within a few hundredths of zero (about 1/√10,000 = 0.01 is typical).
In words: “multiply the two task vectors weight by weight and add, then divide by both lengths, so only their directions count.”
With the numbers: 16 / (5 × 5) = 0.64 for the small pair; about −0.0007 for one random pair of 10,000-weight vectors.
In Python:
import math
import random
def cos(u, v):
dot = sum(a * b for a, b in zip(u, v))
return dot / (math.sqrt(sum(a * a for a in u)) * math.sqrt(sum(b * b for b in v)))
cos([3, 4, 0], [0, 4, 3]) # → 0.64
# two random directions in 10,000 dimensions
random.seed(0)
d = 10_000
u = [random.gauss(0, 1) for _ in range(d)]
v = [random.gauss(0, 1) for _ in range(d)]
round(cos(u, v), 4) # → -0.0007
Reading it: Figure 5 of the paper, as a table: each cell is the cosine between two tasks' vectors, and the shading deepens with the value. (Labels are shortened: Euro is EuroSAT, RES45 is RESISC45, SUN is SUN397.) The diagonal is 1.00 (each vector with itself). Almost every other cell is 0.01 or 0.02: nearly perpendicular. The exceptions are the tasks you would expect to share skills: MNIST and SVHN (both digit recognition) at 0.18, GTSRB (traffic signs, which contain digits) at 0.06 with each of them, and the two satellite-image tasks, EuroSAT and RESISC45, at 0.05. The figure includes a ninth task, KITTI, that is not among the eight used elsewhere in the paper.
Learning rate and training time
- Learning rate. Adding the MNIST and EuroSAT vectors from ViT-L/14 works best when each fine-tune used a low learning rate. As the rate rises, the added model's accuracy falls earlier and faster than the fine-tuned models' own accuracy. The paper recommends extra caution with large learning rates when you plan to combine task vectors.
- Training time. Intermediate task vectors quickly point in nearly the same direction as the final one, and adding intermediate MNIST and EuroSAT vectors reaches high accuracy after only a few hundred of the 2,000 steps, so task vectors can be built cheaply.
Limitations
Task vectors need the same architecture and, in all the paper's experiments, the same pretrained starting point. The authors note that popular bases make this common: at the time of writing, over 3,000 Hub models were fine-tuned from the same BERT-base weights and over 800 from the same T5-small.
Why it matters today
Measure before you merge. The fine-tuning lesson turns this section into a rule of thumb: compute the cosine between task vectors with cosine; near zero, add them; clearly negative, expect interference, as its contradicting tasks show at a cosine of −0.25.
7 Related work · original
Everyday picture
Before this paper, people had noticed that the straight road between a pretrained model and its fine-tune is smooth: every model along it is good. Task arithmetic drives off the end of that road (extrapolation, negation) and joins several roads together (addition, analogies).
The threads it ties together
- Interpolating weights. Averaging models fine-tuned from the same start keeps, and can raise, accuracy; averaging several fine-tunes of one task, model soups, often beats the best single one. Task arithmetic generalises averaging to adding, subtracting and scaling.
- Editing models. Patching, editing, aligning and debugging all change a model after pretraining. Here the changes are modular: capabilities are added or removed by reusing fine-tuned models.
- Steering hidden states. Earlier work added vectors to a model's hidden activations; task vectors live in the weights instead, and leave the fine-tuning procedure unchanged.
Why it matters today
Knowing which thread an idea came from tells you what it assumes. Every method in this family leans on the same fact: models fine-tuned from one start stay in one connected, well-behaved region of weight space.
8 Conclusion · original
Adding task vectors builds multi-task models or improves single tasks; negating them removes behaviours such as toxic generation while keeping performance elsewhere; analogies carry knowledge to domains or groups with little data. Every operation is addition or subtraction of weights, cheap enough to try many edits, and the result is a single model of the original size, so inference costs nothing extra.
Appendix A: weight averaging and ensembles · original
Everyday picture
Two ways to combine two advisers: ask both and average their advice (an ensemble), or merge them into one person who thinks the average thoughts (averaging weights). If the advisers think in a straight-line way, the two give the same answer. The appendix argues fine-tuned models are close enough to straight-line that the merge behaves like the ensemble, at half the cost.
Tiny example: adding two task vectors at λ = ½ is averaging
With the fine-tuning lesson's three-weight models, θpre = (1, 0, 2), θ1 = (1.5, 0, 2), θ2 = (1, −1, 2): θpre + ½(τ1 + τ2) = (1.25, −0.5, 2), exactly the plain average of the two fine-tunes.
In words: “half of each change, applied to the start, lands exactly on the midpoint of the two fine-tuned models.”
With the numbers: (1, 0, 2) + ½((0.5, 0, 0) + (0, −1, 0)) = (1.25, −0.5, 2) = ½((1.5, 0, 2) + (1, −1, 2)).
In Python:
theta_pre = [1.0, 0.0, 2.0]
theta_1, theta_2 = [1.5, 0.0, 2.0], [1.0, -1.0, 2.0]
tau_1 = [a - p for a, p in zip(theta_1, theta_pre)]
tau_2 = [b - p for b, p in zip(theta_2, theta_pre)]
# left side: θ_pre + ½(τ_1 + τ_2)
[p + 0.5 * (a + b) for p, a, b in zip(theta_pre, tau_1, tau_2)] # → [1.25, -0.5, 2.0]
# right side: ½(θ_1 + θ_2)
[0.5 * (a + b) for a, b in zip(theta_1, theta_2)] # → [1.25, -0.5, 2.0]
Ensemble or average?
An ensemble runs both models and mixes their outputs: (1 − α)fθ1(x) + α fθ2(x). Averaging weights runs one model at the mixed weights. For a model fθ(x) = tanh(θx) with x = 1 (illustrative), weights 0.2 and 0.4 give an ensemble output of 0.2887 and an averaged-weights output of tanh(0.3) = 0.2913: nearly equal. Weights 1 and 3 give 0.8783 against tanh(2) = 0.9640: clearly different. The closer together the weights, the more the model behaves like a straight line between them, and the better the average imitates the ensemble.
In words: “an ensemble mixes the two models' outputs; averaging mixes their weights and runs once. They agree whenever the model is close to linear between the two weight settings.”
With the numbers: α = ½: (tanh 0.2 + tanh 0.4) / 2 = 0.2887 against tanh 0.3 = 0.2913; (tanh 1 + tanh 3) / 2 = 0.8783 against tanh 2 = 0.9640.
In Python:
import math
alpha, x = 0.5, 1.0
def f(theta):
return math.tanh(theta * x)
for theta_1, theta_2 in [(0.2, 0.4), (1.0, 3.0)]:
ensemble = (1 - alpha) * f(theta_1) + alpha * f(theta_2)
average = f((1 - alpha) * theta_1 + alpha * theta_2)
print(round(ensemble, 4), round(average, 4)) # → 0.2887 0.2913 0.8783 0.964
Hover or tap to compare the two outputs at each spread.
Reading it: computed live for the illustrative model fθ(x) = tanh(θx) at x = 1, with the two weights placed symmetrically around 1: θ1 = 1 − spread and θ2 = 1 + spread. The x-axis is the spread; the y-axis is the output. The merge (averaging weights) always runs the model at weight 1, so its line is flat at tanh(1) = 0.762. At spread 0 the two models are identical and the lines meet. As the weights move apart, the ensemble (averaging outputs) falls away from the merge, because tanh curves: by spread 2 it is down to 0.117. Fine-tunes of one pretrained model sit close together, the left end of this chart, which is why the paper's measurements find adding two task vectors tracks the corresponding ensemble closely (a Pearson correlation of 0.99 across pairs, with the ensemble slightly ahead on average).
Why it matters today
An ensemble of two models costs two forward passes; a merged model costs one. When the merge nearly matches the ensemble, as it does here, merging is the cheaper way to get almost all of the benefit. The distillation companion shows the other classic way to squeeze an ensemble into one model: train a student to copy it.
What changed since 2022
| In the paper | Common today | Why | Where to learn more |
|---|---|---|---|
| Add full task vectors, one λ for the sum | Conflict-aware merges such as TIES-merging trim small changes and pick one sign per weight before adding | Task vectors that overlap or disagree interfere; resolving conflicts first reduces the damage | fine-tuning lesson, §5b |
| Task vectors from full fine-tunes | Adapters: a LoRA update BA is itself a small, low-rank task vector added to the base weights | Adapters are cheap to train, store and swap, and merging them is the same arithmetic | LoRA companion |
| CLIP and T5 models | Merging fine-tunes of open language models that share a base, with dedicated open-source tools | Many fine-tunes of the same popular base are published, as the paper anticipated | merge |
| Negating a toxic task vector | Removing a behaviour by finding its direction, in weights or in activations | Cheaper than retraining, and testable against a control | InstructGPT companion for the training-based alternative |
Glossary
Every term with hover guidance on this page, in one place.