Sentence-BERT, annotated
How to read this page
Any dotted word explains itself when you hover it, tab to it or tap it; so does every symbol in every equation. Diagrams and demos are live. Each idea climbs the same ladder: an everyday picture, a tiny example you can check by hand, a diagram or demo, the math, and why it still matters. The contrastive training lesson trains a small bi-encoder from scratch, and the retrieval lesson puts one to work next to a cross-encoder.
Abstract · original
“Finding the most similar pair in a collection of 10,000 sentences requires about 50 million inference computations (~65 hours) with BERT.”Reimers and Gurevych (2019), Abstract
Everyday picture
You have 10,000 job applicants and want the two most similar CVs. Method one: sit every possible pair of CVs side by side and read them together, 50 million readings. Method two: write a short summary card for each CV once, 10,000 summaries, then compare the cards. BERT, as used in 2019, could only do method one. Sentence-BERT teaches it method two, without losing much accuracy.
What the paper claims
- A small change to BERT, a pooling step plus siamese fine-tuning, produces sentence embeddings you can compare with cosine similarity.
- Finding the most similar pair among 10,000 sentences drops from about 65 hours to about 5 seconds.
- Better than earlier sentence embedding methods (InferSent, Universal Sentence Encoder) on standard similarity benchmarks.
Why it matters today
Almost every semantic search system, RAG pipeline and embedding API follows this recipe: encode each text once into a vector, store the vectors, and compare vectors at query time. This paper is where BERT-quality models first became usable that way.
1 Introduction: 65 hours versus 5 seconds · original
Everyday picture
A cross-encoder is a judge who must hear both sides at once: to score two sentences, BERT reads them together in one pass. That is accurate, because every word of one sentence can attend to every word of the other, but the work grows with the number of pairs. A bi-encoder reads each sentence alone and writes down a vector; comparing two vectors is a single dot product.
Tiny example
With 5 sentences there are 5 × 4 / 2 = 10 pairs; with 10,000 sentences there are 10,000 × 9,999 / 2 = 49,995,000 pairs. The number of pairs grows with the square of the collection.
In words: “each of the n sentences can be paired with the n − 1 others, and each pair is counted twice that way, so halve it.”
With the numbers: n = 10,000 gives 49,995,000 pairs. At the paper's reported 65 hours on a V100 GPU, that is 49,995,000 / 234,000 seconds ≈ 214 pairs per second (our arithmetic from the paper's figures).
In Python:
n = 10_000
# pairs = n (n − 1) / 2
pairs = n * (n - 1) // 2
pairs # → 49995000
# the paper's 65 hours
seconds = 65 * 60 * 60
seconds # → 234000
# pairs scored per second
round(pairs / seconds) # → 214
Reading it: the left card is BERT as a cross-encoder, the right card is SBERT as a bi-encoder, both using rates derived from the paper's own numbers (about 214 cross-encoded pairs per second; about 2,000 sentences embedded per second; about 5 billion cosine comparisons per second). Slide n to the right. In pair mode the BERT time grows with n² and passes a year of GPU time before n reaches a million, while SBERT's stays in minutes. In search mode, with the collection already embedded, SBERT answers a query in a fraction of a second; at the paper's Quora example of 40 million questions BERT needs over 50 hours per query. Adding an approximate index such as HNSW brings the search time down to milliseconds.
The surprise in the introduction
People had tried the obvious shortcut: feed one sentence into plain BERT and average its output vectors, or take the output of its special [CLS] token. The paper shows this gives bad sentence embeddings for similarity, often worse than averaging old-fashioned word vectors (GloVe). BERT was never trained to make its vectors comparable by cosine; SBERT's fine-tuning is what fixes that.
Why it matters today
The “retrieve with a bi-encoder, rerank the top few with a cross-encoder” pipeline used in nearly every production search system is this trade-off put to work. See the retrieval lesson.
2 Related work · original
Everyday picture
Before SBERT there were two camps. One had accurate but slow judges (BERT reading pairs). The other had fast summary-writers that were less accurate: InferSent, a BiLSTM trained on natural language inference data, and Google's Universal Sentence Encoder, a transformer trained on several tasks. SBERT's idea is to take the best judge and teach it to write summaries.
- BERT for pairs: the two sentences are joined with a [SEP] token and read together, with a small regression head on top. It set the state of the art on the STS benchmark.
- Poly-encoders (Humeau et al.) speed up scoring against precomputed candidates, but their score is not symmetric and they still need O(n²) comparisons for clustering.
- Starting point: unlike earlier sentence encoders trained from scratch, SBERT starts from pre-trained BERT and only fine-tunes it, in under 20 minutes.
3 Model · original
“SBERT adds a pooling operation to the output of BERT / RoBERTa to derive a fixed sized sentence embedding.”Reimers and Gurevych (2019), §3
Pooling: many token vectors to one sentence vector
Everyday picture
BERT gives you one vector per token, like one opinion per committee member. You need a single vector for the whole sentence, a committee decision. Pooling is the voting rule: take the average opinion (MEAN), the strongest opinion on each question (MAX), or just ask the chair (CLS, the special first token).
Try it
Illustrative token vectors with 4 numbers each; real BERT-base vectors have 768.
Reading it: each row is one token's output vector from BERT; each column is one of its numbers. The bold row at the bottom is the sentence vector. MEAN averages each column over all tokens; MAX keeps each column's largest value (the highlighted cells show where each came from); CLS simply copies the first row. Switch between them and notice that MAX builds its answer from different tokens in different columns, while CLS ignores every other token entirely. The paper's ablation (section 6) found MEAN best overall.
In words: “the sentence vector is the average of all the token vectors BERT produced.”
With the numbers: for the first column of the demo, (0.2 + 0.9 + 0.7 + 0.1 + 0.6) / 5 = 0.50.
In Python:
# the first number of each token vector h_i
h = [0.2, 0.9, 0.7, 0.1, 0.6]
# L = 5 tokens
L = len(h)
# u = (1/L) Σ h_i
u = sum(h) / L
print(f"{u:.2f}") # → 0.50
The siamese network
Everyday picture
Two identical twins, trained identically, each read one sentence and write a summary card. Because they are the same person in effect (every weight is shared), their cards are written in the same language and can be compared line by line. During training a coach looks at both cards and says how the sentences relate; at test time you just compare cards.
Hover or tap any block. Start at the bottom with the two sentences.
Reading it: read each side from the bottom up. Each sentence passes through BERT on its own (the two BERT boxes are one network with shared weights, marked “tied”), then through pooling, giving vectors u and v. On the left, during training, u, v and their element-wise difference |u − v| are glued together and fed to a small classifier that must say whether sentence A entails, is neutral towards, or contradicts sentence B; its errors are what tune BERT. On the right, at inference, the classifier is thrown away: you only compute cosine(u, v). The key design point is visible in the shape: nothing ever joins the two sides until the very end, so u can be computed once and stored.
Three training objectives
What the network is trained to do depends on the labels you have. The paper uses three set-ups.
1. Classification (for NLI data)
Everyday picture: a quiz where each question is two sentences and the answer is one of three labels. To answer well, the summary cards must capture meaning precisely enough that their differences reveal agreement or contradiction.
In words: “glue the two sentence vectors and their element-wise distance into one long vector, multiply by a learned matrix to get one score per label, and turn the scores into probabilities.”
With the numbers: with 2-number vectors u = (1, 2) and v = (2, 1), |u − v| = (1, 1) and the glued vector is (1, 2, 2, 1, 1, 1), six numbers. For BERT-base, n = 768 and k = 3 labels, so Wt holds 3 × 768 × 3 = 6,912 learned numbers, tiny next to BERT's 110 million. The loss is cross-entropy on the correct label.
In Python:
u, v = [1, 2], [2, 1]
# |u − v|
diff = [abs(ui - vi) for ui, vi in zip(u, v)]
diff # → [1, 1]
# (u, v, |u − v|) glued into one list
u + v + diff # → [1, 2, 2, 1, 1, 1]
# BERT-base width, number of labels
n, k = 768, 3
# W_t turns 3n numbers into k scores
3 * n * k # → 6912
2. Regression (for similarity scores)
Everyday picture: show the model two sentences and a human's similarity rating, and train it so the cosine of the two cards matches the rating.
In words: “the squared gap between the predicted similarity and the human rating” (mean squared error when averaged over pairs).
With the numbers (illustrative): cos((1, 2), (2, 1)) = 4 / (√5 × √5) = 0.8. If the rating, rescaled to the same range, is 0.9, the loss is (0.8 − 0.9)² = 0.01.
In Python:
import math
u, v = [1, 2], [2, 1]
# 1×2 + 2×1 = 4
dot = sum(ui * vi for ui, vi in zip(u, v))
# 4 / (√5 × √5)
cos = dot / (math.hypot(*u) * math.hypot(*v))
round(cos, 2) # → 0.8
# the human rating, rescaled
y = 0.9
# L = (cos(u, v) − y)²
round((cos - y) ** 2, 2) # → 0.01
3. Triplet (for anchor, positive, negative)
Everyday picture: a sentence (the anchor), one that should be near it (the positive) and one that should be far (the negative). The rule is not “positive close, negative far” in absolute terms, but “the positive must be closer than the negative by at least a safety margin”.
In words: “if the positive is not at least ε closer to the anchor than the negative, pay the shortfall; otherwise pay nothing.”
With the numbers: with ε = 1 (the paper's value), d(a, p) = 0.5 and d(a, n) = 1.2 give 0.5 − 1.2 + 1 = 0.3 of loss. Move the negative out to 1.6 and the loss is max(−0.1, 0) = 0: nothing left to learn from this triplet.
In Python:
# ε, the safety margin
eps = 1
# ‖s_a − s_p‖ and ‖s_a − s_n‖
d_ap, d_an = 0.5, 1.2
round(max(d_ap - d_an + eps, 0), 1) # → 0.3
# move the negative further out
d_an = 1.6
round(d_ap - d_an + eps, 1) # → -0.1
# nothing left to learn
max(d_ap - d_an + eps, 0) # → 0
Reading it: the blue dot is the anchor sentence, the green dot the positive and the red dot the negative; drag the green and red dots. The inner dashed circle has radius d(a, p), the anchor-to-positive distance. The outer circle is that radius plus the margin ε = 1: the negative must sit outside it for the loss to be zero. While the red dot is inside the outer ring, the readout shows the loss and the two forces training would apply: pull the positive in, push the negative out. Once it is outside, the triplet teaches nothing, which is why later systems deliberately mine hard triplets whose negatives sit inside the ring.
Why it matters today
All three objectives are still in use; today's embedding models are mostly trained with a fourth, the in-batch softmax loss popularised for retrieval by DPR, where every other example in the batch is a negative. The losses lesson covers the family.
3.1 Training details · original
- Data: SNLI (570,000 sentence pairs) plus MultiNLI (430,000 pairs), each labelled contradiction, entailment or neutral.
- Objective: the 3-way softmax classifier, for one epoch.
- Settings: batch size 16, the Adam optimizer with learning rate 2 × 10−5, linear warm-up over the first 10% of training, MEAN pooling.
Everyday picture: this is a short refresher course, not a degree. BERT already knows the language from pre-training; one pass over a million labelled pairs teaches it to lay out its knowledge so that cosine similarity means something.
4 Evaluation: semantic textual similarity · original
Everyday picture
Humans rated thousands of sentence pairs from 0 (unrelated) to 5 (same meaning). A good embedding model should put the pairs in the same order as the humans. We do not care whether its cosine for a “5” pair is 0.9 or 0.7, only that the “5” pairs come out ahead of the “3” pairs.
Tiny example: Spearman rank correlation
Four pairs with human scores 4.8, 3.9, 2.0 and 0.5 (ranks 1, 2, 3, 4) get cosines 0.91, 0.72, 0.80 and 0.30 (ranks 1, 3, 2, 4). The rank differences are 0, −1, 1 and 0.
In words: “1 minus a penalty that grows with how much the two rankings disagree; 1 means identical order, 0 means no relation, −1 means reversed.”
With the numbers: Σd² = 0 + 1 + 1 + 0 = 2, m = 4, so ρ = 1 − 12 / 60 = 0.8. The paper reports ρ × 100, so this would be “80”.
In Python:
# from the scores 4.8, 3.9, 2.0, 0.5
human_rank = [1, 2, 3, 4]
cosines = [0.91, 0.72, 0.80, 0.30]
model_rank = [sorted(cosines, reverse=True).index(c) + 1 for c in cosines]
model_rank # → [1, 3, 2, 4]
d = [h - m for h, m in zip(human_rank, model_rank)]
d # → [0, -1, 1, 0]
m = len(d)
# 1 − 12 / 60
rho = 1 - 6 * sum(di**2 for di in d) / (m * (m**2 - 1))
rho # → 0.8
4.1 Without any similarity training
Reading it: each bar is a method's average Spearman correlation (× 100) over seven STS datasets, none of which the models were trained on. The grey bars are the obvious shortcuts: plain BERT's [CLS] vector scores 29.2 and averaged BERT token vectors 54.8, both below averaged GloVe word vectors (61.3). The coloured bars are purpose-built sentence encoders; SBERT (74.9 base, 76.6 large) beats InferSent (65.0) and Universal Sentence Encoder (71.2). Selected averages from Table 1 of Reimers and Gurevych (2019).
4.2 With similarity training: the cross-encoder still wins
Fine-tuned on the STS benchmark's own training data (after NLI), the BERT cross-encoder reaches 88.3 (base) against SBERT's 85.4. Reading both sentences together is genuinely more accurate. The trade is roughly three points of correlation for an enormous speed-up, and it explains the modern two-stage design: a bi-encoder retrieves a shortlist fast, a cross-encoder reranks it carefully.
4.3 A harder test: argument similarity
On the Argument Facet Similarity corpus (arguments about gun control, gay marriage and the death penalty), SBERT nearly matches BERT when trained and tested on the same topics. Trained on two topics and tested on the third, it falls about 7 points behind (50.7 against 57.2 Spearman for the base models). The authors' explanation: BERT can compare two arguments word by word with attention; SBERT must place an argument on an unseen topic in the right spot in vector space on its own, which is harder.
4.4 Wikipedia sections
Using the triplet objective on about 1.8 million triplets (anchor and positive from the same Wikipedia section, negative from another section of the same article), SBERT reached 80.4% accuracy against 74% for the earlier BiLSTM approach.
Why it matters today
“Measure the ranking, not the raw score” is the right habit for any similarity system. Cosine values from different models are not comparable (see anisotropy in the similarity lesson); orderings are.
5 Evaluation: SentEval · original
“Cosine-similarity treats all dimensions equally.”Reimers and Gurevych (2019), §5
Everyday picture
SentEval freezes the sentence vectors and trains a small logistic regression classifier on top for tasks such as sentiment. That is a different exam from cosine similarity: a classifier can learn to pay attention to the few dimensions that matter and ignore the rest.
Tiny example
Two review sentences have vectors (0.1, 5) and (−0.1, 5): the second number is a big “topic: film” signal and the first is a small sentiment signal. Their cosine is (−0.01 + 25) / (5.001 × 5.001) = 0.9992: nearly identical. A classifier with weights (10, 0) gives them scores +1 and −1: perfectly separated. Averaged BERT vectors do reasonably well on SentEval (84.9 average) while failing at cosine similarity for exactly this reason.
Result
SBERT-NLI-large averages 87.7 over seven transfer tasks, about 2 points ahead of InferSent (85.6) and Universal Sentence Encoder (85.1), with large gains on the sentiment tasks. The authors stress that transfer learning is not SBERT's purpose; fine-tuning all of BERT is better for that.
6 Ablation study · original
Everyday picture
Take the recipe apart one ingredient at a time and taste the result. Which pooling rule matters? Which pieces does the training classifier need to see?
Reading it: each bar is a Spearman score (× 100) on the STS benchmark's development set. In what the classifier sees, feeding it only (u, v) scores 66.0; adding the element-wise difference |u − v| jumps to 80.8, the best. The difference vector matters most, because it directly tells the classifier how far apart the two sentences are in each dimension, which is what pulls similar sentences together during training. Adding the element-wise product u × v, used by InferSent, slightly hurts here. In pooling rule, the choice barely matters for NLI training, but with regression training MAX pooling collapses to 69.9. Values from Table 6 of Reimers and Gurevych (2019); concatenation results use MEAN pooling.
Note that the concatenation choice only affects training. At inference time only u, v and cosine are used.
7 Computational efficiency · original
Everyday picture
A GPU processes a batch of sentences as one rectangle of numbers, so short sentences are padded to the length of the longest one in the batch. If a batch mixes a 5-word sentence with a 60-word one, the short one wastes 55 slots. Smart batching sorts sentences by length first, so each batch holds similar lengths and little is wasted.
Reading it: each bar is one sentence (illustrative lengths), grouped into batches of four (the separated blocks). Solid blue is real tokens; hatched grey is padding, the slots the GPU computes on but throws away. In arrival order, one long sentence forces its whole batch to its length. Sorted, each batch is nearly rectangular. The readout counts computed slots. The paper measured smart batching speeding SBERT up by 89% on CPU and 48% on GPU, reaching about 2,042 sentences per second on a V100 (against 1,876 for InferSent and 1,318 for Universal Sentence Encoder). On CPU, the simpler InferSent is about 65% faster.
What changed since 2019
| In the paper | Today | Why | Where to learn it |
|---|---|---|---|
| Fine-tune on NLI labels with a softmax classifier | Contrastive training on huge numbers of (query, passage) pairs with in-batch negatives | Retrieval data is plentiful and the loss matches the use | DPR companion, contrastive lesson |
| BERT-base, 768 dimensions | Many sizes; some models let you truncate the vector | Storage and speed trade-offs | Matryoshka companion |
| Brute-force cosine over all pairs | Approximate nearest-neighbour indexes | Millions to billions of vectors | HNSW companion |
| One vector per sentence | Also multi-vector models (one vector per token) | Finer matching without cross-encoding | ColBERT companion |
What survived intact: the bi-encoder shape, pooling a transformer's token vectors into one embedding, cosine comparison, and the idea of fine-tuning a pre-trained model specifically so that its vectors are comparable.
Glossary
Every term with hover guidance on this page, in one place.