CLIP, annotated
How to read this page
- Any dotted word explains itself when you hover it, tab to it, or tap it.
- Every symbol inside an equation does the same.
- The similarity matrix in §2.3 and the classifier in §3.1 compute live. Drag the temperature.
Each idea climbs the same ladder: everyday picture, tiny example, diagram, the math, why it matters. The contrastive training lesson implements CLIP's symmetric loss in NumPy.
Abstract · original
“We demonstrate that the simple pre-training task of predicting which caption goes with which image is an efficient and scalable way to learn SOTA image representations from scratch on a dataset of 400 million (image, text) pairs collected from the internet.”Radford et al. (2021), Abstract
Everyday picture
Traditional image models learn from a fixed list of labels: “this is one of 1,000 categories”. To recognize something new you need new labelled photos. CLIP learns from the captions people already write next to images on the web, and ends up able to recognize anything you can describe in words. You classify a photo by asking which of several descriptions matches it best.
What the paper claims
- Matching images to captions, trained on 400 million web pairs, learns strong image representations from scratch.
- With no task-specific training at all (“zero-shot”), it matches the original ResNet-50 on ImageNet, 76.2%, without using any of ImageNet's 1.28 million labelled training images.
- It transfers to over 30 datasets: text in images, actions in videos, locations, fine-grained categories.
Why it matters today
CLIP-style shared image-text spaces power image search by description, many multimodal assistants' vision encoders, filtering of training data, and text-to-image generators that need to know what a caption “looks like”.
1 Introduction · original
Everyday picture
Language models improved dramatically by learning from raw web text instead of hand-labelled datasets. Vision was still trained mostly on crowd-labelled sets like ImageNet. The question: can learning from raw web text do for vision what it did for language?
Earlier attempts to learn from captions existed, but performed far below labelled training; one reached 11.5% zero-shot on ImageNet. The paper's answer is mostly scale (far more data) plus a more efficient training objective.
2 Approach · original
2.1 Learning from captions · original
Learning from natural language has two advantages over fixed labels. It scales, because nobody has to label anything: the captions already exist. And the model learns to connect images to language, so at test time you can describe any category in words instead of retraining.
2.2 A 400-million-pair dataset · original
Existing caption datasets were either small (about 100,000 photos) or noisy (a 100-million-photo set shrank to 15 million after keeping only images with real English descriptions). So the authors built WIT (WebImageText): 400 million image-text pairs, gathered by searching for pairs whose text contained one of 500,000 queries, with at most 20,000 pairs per query to keep it balanced. Its total word count is similar to the text dataset used to train GPT-2.
2.3 The contrastive objective · original
Everyday picture
First they tried to make the model write each image's caption word for word. That is very hard, because the same photo could be captioned a thousand ways. The easier game is matching: shuffle a stack of photos and a stack of captions, and learn to pair them up. That turned out to be far more efficient. Predicting only the bag of words, ignoring order, learned 3× faster than writing the caption, and switching to matching gave another 4×.
Hover or tap a cell of the matrix, or an encoder. The green diagonal holds the true pairs.
Reading it: a batch of N image-caption pairs goes in, here N = 4. The image encoder turns each image into a vector I₁ to I₄ (the rows); the text encoder turns each caption into a vector T₁ to T₄ (the columns). Every cell is the similarity of one image with one caption: all N × N combinations. The N cells on the green diagonal are the real pairs; the other N² − N are mismatches. Training pushes diagonal cells up and everything else down. In the real paper N is 32,768, so each image is contrasted with over 32,000 wrong captions at once.
Tiny example: two pairs
Two images and their two captions. After normalizing every vector to length 1, the similarities are: image 1 with caption 1 = 0.9, with caption 2 = 0.1; image 2 with caption 1 = 0.2, with caption 2 = 0.8. Read row 1 as a two-way multiple-choice question: softmax(0.9, 0.1) gives the right caption probability e0.9 / (e0.9 + e0.1) = 2.46 / 3.57 = 0.690, a loss of −ln 0.690 = 0.371. Row 2 gives 0.646 and 0.437. The columns ask the reverse question, “which image matches this caption?”, and give 0.403 each. The average of the two directions is 0.404.
The math
In words: “score every image against every caption with a scaled dot product of unit vectors; for each image, the right caption should win a softmax over all captions, and for each caption, the right image should win a softmax over all images; average the two losses.”
With the numbers: with the scale at 1, the two-pair example gives image-to-text losses 0.371 and 0.437 (mean 0.404) and text-to-image losses 0.403 and 0.403 (mean 0.403); the total is ½(0.404 + 0.403) = 0.404. At CLIP's starting scale of 1/0.07 ≈ 14.3, the same similarities give a loss below 0.001, because scaling stretches the gap between right and wrong.
In Python:
import math
# I_i · T_j: row i is image i, column j is caption j
sims = [[0.9, 0.1],
[0.2, 0.8]]
def loss(scale):
# ℓ_ij = e^t · (I_i · T_j)
l = [[scale * s for s in row] for row in sims]
N = len(l)
img = [-math.log(math.exp(l[i][i]) / sum(math.exp(l[i][j]) for j in range(N))) for i in range(N)]
txt = [-math.log(math.exp(l[j][j]) / sum(math.exp(l[i][j]) for i in range(N))) for j in range(N)]
# ½ (image-to-text + text-to-image)
return img, txt, (sum(img) / N + sum(txt) / N) / 2
img, txt, L = loss(1)
[round(x, 3) for x in img], [round(x, 3) for x in txt] # → ([0.371, 0.437], [0.403, 0.403])
round(sum(img) / 2, 3), round(sum(txt) / 2, 3), round(L, 3) # → (0.404, 0.403, 0.404)
# CLIP's starting scale, 1/0.07 ≈ 14.3
loss(1 / 0.07)[2] < 0.001 # → True
The paper notes this is the same batch trick as the multi-class N-pair loss and InfoNCE, earlier used for medical images: the other items in the batch serve as free negatives. See the losses lesson.
Try it: the similarity matrix, live
Illustrative: four coloured squares stand in for photos, and all eight embeddings are hand-made 5-number vectors (animal, dog, cat, vehicle, food). The arithmetic is CLIP's.
Reading it: rows are images, columns are captions, and each cell shows the probability that this caption is the right one for this image (a softmax along the row). Darker cells mean higher probability, and the green-bordered diagonal holds the correct pairs. At the default τ = 0.07 each row is nearly certain, even though the dog and cat photos are fairly similar to each other's captions (both are animals). Drag τ towards 1 and the rows go soft: the same cosine similarities now give much less decisive probabilities. CLIP learns τ during training, starting at 0.07 and capped so the scale 1/τ never exceeds 100. Hover a cell for its raw cosine similarity.
Why it matters
Because the loss only cares about ranking right pairs above wrong ones within a batch, bigger batches mean more, harder competition. That is why CLIP used a batch of 32,768, and why contrastive training is hungry for scale. The contrastive lesson shows how adding hard negatives sharpens this further.
2.4–2.5 Models and training · original
- Image encoders: five ResNets (modified, with attention pooling at the end) and three Vision Transformers, which cut the image into patches and treat each patch as a token.
- Text encoder: a Transformer of 63 million parameters, 12 layers, 512 wide, 8 heads, reading lower-cased byte-pair-encoded text with a 49,152-token vocabulary, capped at 76 tokens. The vector at the end-of-text token becomes the caption's embedding.
- Simple projections: each encoder's output is mapped into the shared space by one linear layer and normalized to length 1.
- Training: 32 epochs, batch 32,768, Adam with decoupled weight decay (AdamW) and cosine learning-rate decay. The largest ResNet took 18 days on 592 V100 GPUs; the largest Vision Transformer 12 days on 256.
- Learned temperature: τ is a trained parameter, initialized at 0.07 and clipped so logits are never scaled by more than 100, which the authors found necessary for stability.
The best model, ViT-L/14 fine-tuned briefly at 336-pixel resolution, is the one reported as “CLIP” in most results.
3.1 Zero-shot transfer · original
Classifying with text · original
Everyday picture
To sort photos of pets into breeds, you don't retrain CLIP. You write one caption per breed, “a photo of a beagle”, “a photo of a pug”, and so on, and for each photo ask: which caption fits best? The captions become the classifier.
In words: “embed the image and one caption per class; the probability of each class is a softmax over the image's cosine similarity with each caption, sharpened by the temperature.”
With the numbers: with two classes, cosines 0.30 (dog) and 0.25 (cat), and τ = 0.01, the logits are 30 and 25, so p(dog) = 1 / (1 + e−5) = 0.993. A 0.05 difference in cosine becomes a confident answer.
In Python:
import math
# cos(I, T_k) for each class caption
cos = {"dog": 0.30, "cat": 0.25}
tau = 0.01
logits = {k: c / tau for k, c in cos.items()}
{k: round(z) for k, z in logits.items()} # → {'dog': 30, 'cat': 25}
# Σ_j exp(cos(I, T_j) / τ)
total = sum(math.exp(z) for z in logits.values())
# p(y = dog | image)
round(math.exp(logits["dog"]) / total, 3) # → 0.993
# the same thing, written with the gap of 5
round(1 / (1 + math.exp(-5)), 3) # → 0.993
The paper points out that this is a linear classifier whose weights are written by the text encoder from the class names. That is why no training examples are needed.
Illustrative: hand-made embeddings, CLIP's arithmetic, τ = 0.07.
Reading it: pick a stand-in photo. Each bar is the probability CLIP's zero-shot rule gives to one class, computed from the cosine similarity between the photo's embedding and the embedding of “a photo of a {class}”. Press “add a class” to add “bird”: the probabilities are recomputed over the new set, and nothing is retrained. Notice that the dog photo gives the cat caption, and then the bird caption, a little probability, far more than car or pizza: they share the “animal” direction.
Against the earlier attempt · original
| Model | aYahoo | ImageNet | SUN |
|---|---|---|---|
| Visual N-Grams (2017) | 72.4 | 11.5 | 23.0 |
| CLIP | 98.4 | 76.2 | 58.5 |
76.2% top-1 matches the original fully supervised ResNet-50; CLIP's top-5 accuracy reached 95%. The authors stress this is context, not a controlled comparison: CLIP used about 10× more data and far more compute.
Prompt engineering and ensembling · original
Everyday picture
A single word is ambiguous. “Crane” could be a bird or a machine; “boxer” a dog or an athlete. And web captions are rarely one word. Wrapping the label in a sentence tells CLIP what kind of thing you mean.
- “A photo of a {label}.” alone improved ImageNet accuracy by 1.3%.
- Task hints helped more: “a photo of a {label}, a type of pet.” on pet breeds; “a satellite photo of a {label}.” for satellite images; quotes around text for reading digits.
- Ensembling: average the text embeddings of 80 different prompts (“a photo of a big {label}”, “a photo of a small {label}”, …) into one per class, which costs nothing extra at prediction time. That added 3.5% on ImageNet.
Together, almost 5 points on ImageNet, which the paper likens to the gain from 4× more compute, for free.
Why it matters today
This was an early, clean demonstration that the wording of what you ask a model changes the answer, for vision as much as for language. Averaging prompt embeddings is still a common trick.
How good is zero-shot? · original
- Against a fully supervised linear probe on ResNet-50 features, zero-shot CLIP won on 16 of 27 datasets.
- It did especially well on everyday objects and scenes, and poorly on specialized or abstract tasks: satellite images, tumour detection, counting objects, traffic signs, distance estimation.
- Zero-shot CLIP matched a 4-shot linear probe on its own features: describing a class in words was worth about four labelled examples per class.
3.2 Representation learning · original
Setting zero-shot aside, how good are CLIP's image embeddings when you do train a linear classifier on them? A linear probe on the best CLIP model beat one on the best ImageNet model at the time (Noisy Student EfficientNet-L2) on 21 of 27 datasets, with the biggest gains on reading text in images, geolocation, scene recognition and actions in video. Small CLIP models were not better than similarly sized ImageNet models; the advantage appears with scale.
3.3 Robustness to distribution shift · original
Everyday picture
A student who memorized one textbook's photos of bananas may fail on a sketch of a banana, or a banana photographed in a kitchen. ImageNet models are like that: accuracy drops sharply on new kinds of images of the same things (sketches, renditions, adversarially chosen photos, differently collected sets). A ResNet-101 makes 5 times as many mistakes on these natural shifts as on ImageNet itself.
- Zero-shot CLIP closed the gap between ImageNet accuracy and accuracy under shift by up to 75%.
- Training a classifier on CLIP's features specifically for ImageNet raised ImageNet accuracy by 9.2% but gave essentially no gain on the shifted sets: the gain was ImageNet-specific.
Why it matters
A benchmark score is a measurement on one distribution. Models trained on broad data and evaluated zero-shot tend to hold up better when the world shifts, which is the practical situation.
4 Comparison to humans · original
Five people classified the 3,669 test images of the Oxford-IIIT Pets dataset into 37 breeds. With no examples they averaged about 54%; shown one example per breed, about 76%, with the gain almost entirely on images they had been unsure about. A second example added little. Humans “know what they don't know” and learn from one example; CLIP's few-shot learning does not use prior knowledge this efficiently. The breeds hardest for CLIP were also hardest for people.
5 Data overlap · original
With 400 million web images, some test images were probably seen in training (contamination). The authors built a duplicate detector and checked 35 datasets: 9 had no detected overlap, overlap was usually in the single digits of percent, and only a handful of datasets showed a statistically significant accuracy change from it, two of them for the worse. Overall inflation was small.
6 Limitations · original
“However, CLIP only achieves 88% accuracy on the handwritten digits of MNIST.”Radford et al. (2021), §6
- Truly unfamiliar data: handwritten digits, rare online, trip it up; logistic regression on raw pixels does better.
- Fine-grained and abstract tasks: car models, flower species, aircraft variants, counting objects.
- Only chooses, never describes: it can pick among the labels you give it, but can't write a caption itself.
- Compute: the authors estimate about a 1000× increase in compute would be needed for zero-shot CLIP to reach state-of-the-art overall.
- Methodology: development repeatedly used full validation sets, which isn't truly zero-shot.
7 Broader impacts · original
Because anyone can write new classes, CLIP can be pointed at tasks with serious social consequences, from surveillance to labelling people. The paper's bias probes found that class design and thresholds strongly change what harmful associations appear, and that uncurated web data carries social biases into the model. The authors call for evaluating such models in the context of each specific deployment.
9 Conclusion · original
The paper's answer to its opening question is yes: the recipe of web-scale, task-agnostic pre-training transfers from language to vision. To be good at matching, CLIP learns many visual tasks along the way, and natural-language prompts unlock them zero-shot, competitively with supervised models at sufficient scale, with plenty of room left.
What changed since
| In the paper | Common today | Lesson or companion |
|---|---|---|
| Image-text matching on 400M pairs | The same objective on billions of pairs, and in text-text embedding models | contrastive |
| Classify by picking a caption | Search photos by description; vision encoders inside multimodal chat models | retrieval |
| Softmax over the whole batch | Variants that score each pair on its own, making huge batches cheaper | losses |
| One fixed embedding size | Embeddings that can be truncated to smaller sizes | Matryoshka companion |
Glossary
Every term with hover guidance on this page, in one place.