primer.ml.embeddings.similarity

Similarity: how close are two meanings?

Run: python -m primer.ml.embeddings.similarity

New to the notation (vectors, Σ, ‖x‖)? Start with primer.notation, which builds every symbol used here from zero.

Level 1: The practitioner's guide

In one sentence. A similarity metric turns two embedding vectors into one number that says how alike their meanings are, and the three in common use (cosine, dot product, Euclidean distance) agree with each other only under conditions you have to arrange.

When you need it. Every time you compare embeddings: semantic search, retrieval for RAG, near-duplicate detection, a semantic cache that answers a question from a stored one, clustering, recommendations. You choose a metric whether you mean to or not, because every vector database and every search library has a default, and the wrong one ranks silently wrong. The tells: you are about to write ORDER BY against a vector column and have not checked which operator the embedding model expects; or your team argues about whether "0.8 similar" is a lot. This lesson's demo shows why the second argument can't be settled by intuition: for one model, paraphrases score between 0.76 and 0.87 and unrelated pairs between 0.72 and 0.78, so "0.8" is either a near-duplicate or a stranger. You don't need any of this for exact-match lookups (an id, a hash) or for keyword search, which scores words, not vectors.

Your options. The ways to score a pair, from the cheapest to the most deliberate:

Option What it does What it guarantees What it costs Where it lives
Dot product on normalized vectors Normalize every vector to length 1 once, at write time, then multiply and add The same ranking as cosine and as Euclidean distance, at the lowest price One multiply-add per dimension per pair; a normalization step on insert Your ingestion code, then the database's inner-product operator
Cosine similarity Dot product divided by both lengths Direction only; lengths can't skew the order, whatever came in Two extra norms per comparison, unless the store caches them The database's cosine operator, or a library's default
Raw dot product Direction and length together Keeps signal a model put in the length, such as popularity Rankings that a mix of long and short vectors can dominate Recommenders and models trained with it
Euclidean distance Straight-line gap between the tips The natural metric for clustering and for many index structures Disagrees with cosine unless vectors are normalized k-means, some indexes
Mean-centred scores Subtract the collection's average vector before comparing Unrelated pairs land near 0, similar ones stand clear A pass over the collection, and a re-run when it changes Your code, before scoring
A calibrated threshold Label a few hundred pairs, sweep the cut-off, keep the best F1 A yes/no that matches your data and your model A labelling session, repeated at every model change A number stored with the model version

How to choose. Start from how the model was trained, not from which metric sounds best.

  • The model's documentation says cosine, or says its vectors are unit length (most text embedding APIs): normalize on write and rank by dot product. Same order as cosine, fewer operations.
  • The model was trained with a raw dot product and encodes something in the length (recommenders, some retrieval models): keep raw vectors and use the dot product. Normalizing "to be safe" throws the signal away.
  • You need an ordered list (search, RAG): rank, and never threshold. A shared offset in every score leaves the order alone.
  • You need a yes/no (duplicates, a cache hit, "no good answer found"): calibrate the threshold on labelled pairs from your own data, for this model, and store it with the model version.
  • Whatever you pick, use one metric, one model and one normalization policy for every vector in the index. A mixture ranks wrong without an error.

What it costs. A comparison is arithmetic over the vector's dimensions: real models produce 384 to 3,072 numbers per text, so a dot product is a few thousand multiply-adds, and cosine adds two square roots and a division unless the lengths are precomputed. Exact search is one comparison per stored vector; the sentence-transformers documentation puts the practical limit of that brute-force loop at about a million entries, after which you move to an approximate index (primer.ml.embeddings.ann). Storage is 4 bytes per dimension as float32 (pgvector charges 4 × dimensions + 8 bytes per vector, half that as 16-bit floats), which primer.ml.embeddings.compression shrinks. Normalizing is one pass at write time, once. Calibration costs labelling: a few hundred pairs, scored and swept, and the demo shows the return on it: a guessed threshold of 0.80 reaches F1 0.868 on the demo's pairs, the calibrated 0.776 reaches 0.992.

What breaks.

  • Mixed normalization. Some vectors normalized, some not, and the ranking quietly changes: in the demo, raw dot product and cosine order 200 documents differently, and agree only after every vector is normalized. Normalize in one place, on the write path.
  • A threshold from folklore. A cut-off copied from a blog post or from the previous model is meaningless: F1 swings from 0.99 to 0.87 across four hundredths of cosine. Measure it on labelled pairs, per model.
  • Every score looks like 0.8. Transformer embeddings crowd into a narrow cone (Ethayarajh, 2019, found them "not isotropic in any layer"). Rank instead of judging absolute scores, or mean-centre: after centring, the demo's unrelated pairs sit at 0.00 and paraphrases at 0.30.
  • A model upgrade with an old index. A new model is a new space: old vectors, old thresholds and old scores don't carry over. Re-embed the whole collection and recalibrate (primer.ml.embeddings.operations).
  • The operator is a distance, not a similarity. Databases often expose cosine distance (1 − cosine) and the negative inner product so that ascending order is best-first, as pgvector does. Read the sign before you sort.
  • Queries and documents embedded differently. Some providers train separate treatments for the query and for the passage (Cohere's input_type of search_query and search_document); embed both sides the way the model expects or the scores are off.
  • Nearest means little on unstructured data. With 1,000 random points, the nearest is 0.7% as far as the farthest in 2 dimensions and 90% as far in 1,000: the curse of dimensionality. Real embeddings are structured, so search works, but a score gap that small is a warning that the vectors carry little.

In the wild. pgvector exposes one operator per metric on a Postgres column: <-> for Euclidean distance, <#> for the negative inner product, <=> for cosine distance, plus Hamming and Jaccard on bit vectors, up to 16,000 dimensions. sentence-transformers' semantic search defaults to cosine and offers dot-product scoring for normalized vectors. OpenAI's embedding guide states that its embeddings are normalized to length 1, recommends cosine, and notes that it can then be computed as a plain dot product; Cohere asks for an input_type per side of the search. Faiss's flat indexes (IndexFlatIP, IndexFlatL2) are the exact inner-product and Euclidean searches every approximate index is measured against. Ethayarajh (2019) measured anisotropy in contextual models and Su et al. (2021) showed that centring and whitening spreads the vectors back out.

Go deeper. Level 2 computes all three metrics by hand on three 3-dimensional vectors, proves in one line why normalizing makes them agree (the squared distance is 2 − 2 × cosine), builds the popularity example where cosine and dot product disagree on purpose, measures the curse of dimensionality with random points, and manufactures anisotropic embeddings to show a calibrated threshold beating a guessed one. If you only needed to choose, you are done.

Level 2: How it works, from scratch.

Level 2: How it works, from scratch

An embedding gives every piece of text a set of coordinates on a map of meaning. Texts about similar things live on the same street; unrelated ones are across town. Draw an arrow from the centre of the map to each text's spot, and "how similar are these two texts?" becomes "how alike are these two arrows?".

There are three natural ways to compare two arrows:

  • Cosine similarity compares the direction they point and ignores how long they are, like two people pointing at the same mountain, one with a longer arm.
  • Euclidean distance is the straight-line distance between the two arrow tips, like measuring with a ruler.
  • Dot product mixes both: pointing the same way and being long both make it bigger.

A tiny worked example

Three texts, each placed in a 3-dimensional space (real models use 384 to 3,072 dimensions, meaning numbers per vector; the arithmetic is the same):

a = (1, 2, 2), b = (2, 1, 2), c = (2, −1, −2)

  1. Dot product: multiply position by position, then add. a·b = 1·2 + 2·1 + 2·2 = 8. a·c = 1·2 + 2·(−1) + 2·(−2) = −4.
  2. Length (the norm, written ‖a‖): square each number, add, and take the square root. ‖a‖ = √(1 + 4 + 4) = 3, and b and c also have length 3.
  3. Cosine: dot product divided by both lengths. cos(a, b) = 8 / (3·3) = 0.89. cos(a, c) = −4 / 9 = −0.44.
  4. Euclidean distance: subtract, square, add, square-root. ‖a − b‖ = √(1 + 1 + 0) = 1.41. ‖a − c‖ = √(1 + 9 + 16) = 5.10.
Pair Dot product Lengths multiplied Cosine Distance Reading
a, b 8 9 0.89 1.41 Very similar
a, c −4 9 −0.44 5.10 Pointing away

Cosine runs from −1 (opposite directions) through 0 (unrelated, at right angles) to 1 (same direction).

flowchart LR A["a = (1, 2, 2)"] --> D["dot product<br/>a·b = 8"] B["b = (2, 1, 2)"] --> D A --> NA["length ‖a‖ = 3"] B --> NB["length ‖b‖ = 3"] D --> COS["cosine = 8 / (3 × 3) = 0.89"] NA --> COS NB --> COS A --> SUB["difference a − b = (−1, 1, 0)"] B --> SUB SUB --> L2["distance = √2 = 1.41"]

Reading it: the two input vectors feed three calculations. The top path is the dot product. The middle paths measure each vector's length, and cosine is simply the dot product divided by those lengths, which is how it comes to "ignore length". The bottom path subtracts the vectors and measures what's left: that's distance. Everything later in this lesson is these three boxes.

In code: worked_example computes every number in the table above from a, b and c, so you can check your hand arithmetic against it.

The math

Level 3: the formula and its symbols

$$ \cos(a, b) = \frac{a \cdot b}{\lVert a \rVert \, \lVert b \rVert} \qquad a \cdot b = \sum_{i=1}^{d} a_i b_i \qquad \lVert a - b \rVert = \sqrt{\sum_{i=1}^{d} (a_i - b_i)^2} $$

Symbols

Symbol Meaning here Shape / range
a, b the two embedding vectors being compared d numbers each
d number of dimensions 3 in the example; 384 to 3,072 in real models
i which position (dimension) we're looking at 1 to d
aᵢ the i-th number in a a real number
Σ "add up", over i from 1 to d
a · b dot product: multiply matching positions and add one number, any sign
‖a‖ length (norm) of a: √(Σ aᵢ²) one number ≥ 0
√ square root
cos(a, b) cosine of the angle between a and b −1 to 1
‖a − b‖ Euclidean distance between the tips ≥ 0; 0 means identical

In words: the cosine of a and b is their dot product divided by the product of their lengths; the dot product adds up the products of matching positions; the distance is the square root of the summed squared differences.

On the example: cos(a, b) = (1·2 + 2·1 + 2·2) / (3 · 3) = 8 / 9 = 0.89; ‖a − b‖ = √((1−2)² + (2−1)² + (2−2)²) = √2 = 1.41.

Level 3: in Python

In Python:

import math
a, b = [1, 2, 2], [2, 1, 2]
# a · b = Σ a_i b_i
dot = sum(a_i * b_i for a_i, b_i in zip(a, b))
dot  # → 8
# ‖a‖
norm_a = math.sqrt(sum(a_i ** 2 for a_i in a))
# ‖b‖
norm_b = math.sqrt(sum(b_i ** 2 for b_i in b))
norm_a, norm_b  # → (3.0, 3.0)
# cos(a, b)
round(dot / (norm_a * norm_b), 2)  # → 0.89
# ‖a - b‖
round(math.sqrt(sum((a_i - b_i) ** 2 for a_i, b_i in zip(a, b))), 2)  # → 1.41

The code is cosine, dot and euclidean, each a line of NumPy.

In code: rank sorts a whole collection of documents from best to worst match under whichever of the three metrics you name.

Why it matters: every vector database, semantic search box and RAG system does exactly this comparison, millions of times per second. Knowing which of the three a system uses, and why, is the difference between correct and silently wrong rankings.

Normalize, and the three agree

Everyday picture: trim every arrow to the same length, one unit. Now the only thing that can differ is direction, so "tips are close together" and "arrows point the same way" become the same statement.

Tiny example: a = (1, 0) and b = (0.6, 0.8) both have length 1 (check: 0.6² + 0.8² = 0.36 + 0.64 = 1). Their cosine is 1·0.6 + 0·0.8 = 0.6. The squared distance is (1 − 0.6)² + (0 − 0.8)² = 0.16 + 0.64 = 0.8, and 2 − 2·0.6 = 0.8. Same number.

Level 3: the formula and its symbols

$$ \lVert a - b \rVert^2 = \lVert a \rVert^2 + \lVert b \rVert^2 - 2\,a\cdot b = 2 - 2\cos(a, b) \quad \text{when } \lVert a \rVert = \lVert b \rVert = 1 $$

Symbols

Symbol Meaning here Shape / range
‖a − b‖² squared distance between the tips 0 to 4 for unit vectors
‖a‖², ‖b‖² squared lengths, both 1 after normalizing 1
a · b dot product, equal to cosine for unit vectors −1 to 1
cos(a, b) cosine similarity −1 to 1

In words: for unit-length vectors, the squared distance is two minus twice the cosine, so the higher the cosine, the smaller the distance, always.

On the example: ‖a − b‖² = 1 + 1 − 2·0.6 = 0.8 = 2 − 2·0.6.

Level 3: in Python

In Python:

# both length 1
a, b = [1, 0], [0.6, 0.8]
# for unit vectors, a · b is the cosine
cos = sum(a_i * b_i for a_i, b_i in zip(a, b))
# ‖a - b‖²
dist_sq = sum((a_i - b_i) ** 2 for a_i, b_i in zip(a, b))
# the two sides of the identity
round(dist_sq, 2), round(2 - 2 * cos, 2)  # → (0.8, 0.8)

L2-normalizing (dividing a vector by its own length) is what trims every arrow to length 1. After it, the dot product is the cosine, and distance is a decreasing function of the cosine. So once vectors are L2-normalized, cosine, dot product and Euclidean distance give the same ranking.

Try it: pick "Same direction, different lengths", then drag B's length. The dot product and the distance change, but the cosine stays at 1, because it only sees the angle. Drag an angle instead and watch the cosine slide from 1 through 0 to −1. Then switch on "Normalize to length 1": both arrows shrink to the unit circle, the dot product becomes the cosine, and the squared distance becomes 2 − 2 × cosine.

All 400 random unit-vector pairs fall exactly on the line squared distance = 2 - 2 x cosine, so higher cosine always means smaller distance

Reading it: each dot is one pair of random unit vectors. The horizontal axis is their cosine and the vertical axis is their squared Euclidean distance. Every dot sits exactly on the line y = 2 − 2x. Because the line only goes down, "higher cosine" and "smaller distance" always mean the same thing, so the two metrics can never disagree about which neighbour is nearer.

flowchart TD S{What was the model<br/>trained with?} -->|cosine| N[L2-normalize every vector<br/>once, at write time] S -->|dot product| D{Does the vector length<br/>carry signal?} D -->|yes, e.g. popularity| RAW[Keep raw vectors,<br/>rank by dot product] D -->|no| N N --> FAST[Rank by dot product:<br/>same order as cosine and L2,<br/>and the cheapest to compute]

Reading it: start at the top with the question that matters, which is how the model was trained, not which metric sounds best. Most text embedding models are trained with cosine, so you normalize once when you write vectors, then use the plain dot product at query time. It gives the same order as cosine and costs less. Only when length is deliberate signal do you skip normalizing.

In code: l2_normalize trims every row to length 1; rankings_agree_after_normalizing ranks random documents by all three metrics before and after normalizing and reports which rankings match, and l2_identity_check returns both sides of the identity for two random vectors.

Why it matters: vector databases typically normalize on insert and then use the dot product, the cheapest of the three. If you mix normalized and unnormalized vectors, rankings quietly change. Use the metric the embedding model was trained with.

When length carries signal

Everyday picture: on the map of meaning, some shops are big and popular. A recommender may deliberately make popular items' arrows longer, so they win ties. Cosine throws that away; the dot product keeps it.

Tiny example (popularity_in_the_norm): the query points along (1, 0). A niche item (0.99, 0.10) is almost perfectly aligned. A popular item (0.95, 0.30) × 3 = (2.85, 0.90) is slightly off but three times longer.

Item Cosine with query Dot with query
niche (0.99, 0.10) 0.99 / 0.995 = 0.995 0.99
popular (2.85, 0.90) 2.85 / 2.99 = 0.954 2.85

Cosine prefers the niche item; the dot product prefers the popular one.

The short item points almost along the query and wins on cosine (0.99 vs 0.95); the item three times longer wins on dot product (2.85 vs 0.99)

Reading it: the black arrow is the query. The blue arrow (niche item) points almost exactly the same way but is short. The orange arrow (popular item) leans a little away but is three times longer. Cosine only looks at the angle between each arrow and the query, so blue wins. The dot product also rewards length, so orange wins. Neither is wrong: it depends on what the model was trained to encode.

Why it matters: some retrieval and recommender models are trained with the raw dot product and put useful signal in the length. Normalizing their vectors "to be safe" throws that signal away.

The curse of dimensionality

Everyday picture: in a small village some neighbours live next door and some live across the valley. Now imagine a city with thousands of streets running in thousands of directions: pick random strangers and they all turn out to live about equally far from you. "Nearest" stops meaning much.

Tiny example: scatter 1,000 random points in a cube and measure from a random spot. In 2 dimensions the nearest point is about 1/100th as far away as the farthest. In 1,000 dimensions it's about 90% as far: everyone is nearly equidistant.

Level 3: the formula and its symbols

$$ \text{ratio} = \frac{\min_j \lVert q - x_j \rVert}{\max_j \lVert q - x_j \rVert} $$

Symbols

Symbol Meaning here Shape / range
q the query point d numbers
xⱼ the j-th of the random points d numbers each
j which random point 1 to 1,000
min, max over j the smallest / largest value across all the points
ratio nearest distance divided by farthest distance 0 to 1; near 1 means "all equally far"

In words: the ratio is the distance to the closest point divided by the distance to the farthest point.

On the example: by hand first: a query at (0, 0) and three points at (1, 0), (0, 2) and (3, 4) are 1, 2 and 5 away, so the ratio is 1 / 5 = 0.2. With 1,000 random points the demo table shows about 0.01 at d = 2 and about 0.9 at d = 1,000. The Python below draws its own random points, so it lands close by rather than on the same digits: 0.02 and 0.9.

Level 3: in Python

In Python:

import math, random
def ratio(q, xs):
    # ‖q - x_j‖ for every j
    dists = [math.dist(q, x_j) for x_j in xs]
    # min over j ÷ max over j
    return min(dists) / max(dists)
ratio((0, 0), [(1, 0), (0, 2), (3, 4)])  # → 0.2
random.seed(0)
for d in (2, 1000):
    q = [random.random() for _ in range(d)]
    xs = [[random.random() for _ in range(d)] for _ in range(1000)]
    # one significant figure
    print(d, f"{ratio(q, xs):.1g}")  # → 2 0.02 1000 0.9

The nearest-to-farthest distance ratio rises from about 0 in 2 dimensions to 0.8 at 200 and 0.92 at 2,000: random points become almost equidistant

Reading it: the horizontal axis is the number of dimensions (log scale); the vertical axis is the nearest-to-farthest ratio above. In 2-D the ratio is near 0: the nearest point really is much nearer. By a few hundred dimensions the curve flattens toward 1, so every random point is roughly equally far away, and "nearest" tells you almost nothing about random data.

In code: curse_of_dimensionality scatters the random points, measures from a random query, and returns the nearest, farthest and ratio for each dimension.

Why it matters: exact nearest-neighbour search in high dimensions gets slow and fragile, which is one reason vector search uses approximate indexes (primer.ml.embeddings.ann). Real embeddings aren't random: meaning puts them on a much lower-dimensional structure, so nearest-neighbour search still works well in practice.

Anisotropy: why every score looks like 0.8

Everyday picture: a stadium crowd all facing the stage. Ask "are these two people facing the same way?" and the answer is "roughly yes" for everyone, so the question tells you little. Many embedding models behave like that crowd: their vectors bunch in a narrow cone, so even unrelated texts get cosines of 0.7 or more. This lopsidedness is called anisotropy ("not the same in every direction").

Tiny example: build each vector as √3·m + n, where m is one direction everybody shares (the stage) and n is a unit-length direction of each text's own. For two unrelated texts the n parts are at right angles, so their dot product is (√3)² = 3 and each length is √(3 + 1) = 2, and the cosine is 3 / (2 · 2) = 0.75. For paraphrases, whose own parts agree with cosine ρ, the cosine becomes 0.75 + 0.25·ρ: about 0.79 to 0.86. All the signal is squeezed into a few hundredths.

Level 3: the formula and its symbols

$$ v = \alpha\, m + \beta\, n, \qquad \cos(v_1, v_2) = \frac{\alpha^2 + \beta^2 \rho}{\alpha^2 + \beta^2} $$

Symbols

Symbol Meaning here Shape / range
v, v₁, v₂ an embedding vector d numbers (768 here)
m the shared direction every vector leans toward unit vector
n the text's own direction unit vector
α (alpha) how strongly every vector leans toward m √3 here
β (beta) how strongly each keeps its own direction 1 here
ρ (rho) cosine between the two texts' own directions 0 unrelated; 0.15 to 0.45 for paraphrases

In words: every vector is a big shared part plus a small personal part, so the cosine is mostly the shared part's weight, nudged up a little by how much the personal parts agree.

On the example: unrelated, ρ = 0: (3 + 0) / (3 + 1) = 0.75. A paraphrase with ρ = 0.4: (3 + 0.4) / 4 = 0.85.

Level 3: in Python

In Python:

import math
alpha, beta = math.sqrt(3), 1
# shared direction; two unrelated own directions
m, n1, n2 = (1, 0, 0), (0, 1, 0), (0, 0, 1)
# v = α m + β n
v1 = [alpha * m_i + beta * n_i for m_i, n_i in zip(m, n1)]
v2 = [alpha * m_i + beta * n_i for m_i, n_i in zip(m, n2)]
dot = sum(a * b for a, b in zip(v1, v2))
# the cosine, measured directly
round(dot / (math.hypot(*v1) * math.hypot(*v2)), 2)  # → 0.75
def cos_formula(rho):
    return (alpha ** 2 + beta ** 2 * rho) / (alpha ** 2 + beta ** 2)
# unrelated, and a paraphrase
round(cos_formula(0), 2), round(cos_formula(0.4), 2)  # → (0.75, 0.85)

Raw cosines crowd paraphrases (0.76 to 0.87) and unrelated pairs (0.72 to 0.78) together; mean-centred, unrelated sit near 0 and paraphrases near 0.3

Reading it: the left panel shows raw cosine scores for 300 paraphrase pairs (one colour) and 300 unrelated pairs (the other). Both piles sit between 0.7 and 0.9, so a score of "0.8" alone says little. The right panel shows the same pairs after mean-centering, which means subtracting the average vector of the whole collection (that removes the shared "stage" direction). Unrelated pairs drop to about 0, paraphrases sit clearly above, and the gap is easy to see. The information was there all along, squeezed into a narrow band.

In code: anisotropic_embeddings builds the pairs as α·m + β·n, pair_cosines scores each pair, and mean_center subtracts the average vector to remove the shared direction.

Picking a threshold from data, not folklore

Sometimes you need a yes/no cut-off: "are these duplicates?", "is this cached answer close enough?". Pick it by measuring. Label some pairs, score them, and choose the threshold with the best F1, the balance of precision (of the pairs I called similar, how many were?) and recall (of the truly similar pairs, how many did I catch?):

Level 3: the formula and its symbols

$$ F_1 = \frac{2PR}{P + R} $$

Symbols

Symbol Meaning here Range
P precision: true matches ÷ everything called a match 0 to 1
R recall: true matches found ÷ all true matches 0 to 1
F₁ their harmonic mean; high only when both are high 0 to 1

In words: F1 is twice the product of precision and recall, divided by their sum.

On a hand example: scores 0.9, 0.8, 0.7, 0.6 where the first two are true matches. Threshold 0.8 calls exactly those two matches: P = 2/2 = 1, R = 2/2 = 1, F₁ = 2·1·1 / 2 = 1. best_threshold finds 0.8.

Level 3: in Python

In Python:

scores = [0.9, 0.8, 0.7, 0.6]
# the labels
is_match = [True, True, False, False]
# threshold 0.8
called = [s >= 0.8 for s in scores]
true_matches = sum(c and y for c, y in zip(called, is_match))
# precision
P = true_matches / sum(called)
# recall
R = true_matches / sum(is_match)
# F₁
2 * P * R / (P + R)  # → 1.0
flowchart LR L[Sample a few hundred pairs<br/>from your own data] --> H[Label each:<br/>same meaning or not] H --> E[Score every pair<br/>with THIS model] E --> W[Sweep thresholds,<br/>compute precision, recall, F1] W --> P[Pick the threshold that fits<br/>your cost of each error] P --> V[Store it with the<br/>model version] V -->|model upgraded| E

Reading it: a threshold is a property of one model on one kind of data, so it is measured, not guessed. Label pairs, score them, sweep, and pick. The loop back from "model upgraded" is the step people forget: a new model has a different score distribution, so the old threshold is meaningless.

F1 on raw scores swings within a few hundredths: the calibrated threshold 0.776 reaches F1 0.99, while the guessed 0.80 drops to 0.87

Reading it: the curve is F1 at each candidate threshold on the raw anisotropic scores. The dashed line is the "obvious" guess of 0.8; the solid line is the calibrated threshold. F1 changes sharply over a few hundredths of cosine, which is exactly why a threshold carried over from a different model (or a blog post) can be badly wrong.

In code: f1_curve computes F1 at every candidate threshold for this plot; best_threshold returns the single best one with its precision and recall.

Why it matters:

  • Absolute scores mean different things for different models. "0.8" is not a universal "similar".
  • Rank, don't threshold, whenever you can. A shared offset doesn't change which document is on top.
  • When you must threshold (dedup, semantic caching, "no good answer found"), calibrate on labeled pairs for that model, and redo it when the model changes.

In 20 seconds

  • Cosine compares direction, the dot product compares direction and length, Euclidean distance measures the gap between the tips.
  • On L2-normalized vectors all three give the same ranking, because ‖a − b‖² = 2 − 2cos.
  • Use the metric the model was trained with.
  • Scores aren't comparable across models (anisotropy). Rank when you can; calibrate thresholds on labeled pairs when you can't.

Self-test questions

Q: When do cosine, dot product and Euclidean distance give the same ranking? When all vectors are L2-normalized: the dot product equals the cosine, and the squared distance equals 2 − 2cos, which always moves opposite to cosine.

Q: When would you deliberately use the raw dot product on unnormalized vectors? When the model was trained that way and encodes useful signal in the length (e.g. popularity or quality in some recommender and retrieval models). Normalizing would throw that signal away.

Q: Similarity scores all sit between 0.75 and 0.85. What's going on and how do you set a threshold? The model is anisotropic: its vectors share a dominant direction, so every pair looks similar and the real signal lives in a narrow band. Ranking still works. For a threshold, label a few hundred pairs from your own data, sweep the threshold, and pick the one that maximizes F1 (or matches your precision/recall needs). Optionally mean-center the vectors. Redo it for every model version.

Q: What is the curse of dimensionality for nearest-neighbour search? For random high-dimensional data, the nearest and farthest points are almost the same distance away, so "nearest" carries little information and exact search gets expensive. Real embeddings have structure, and approximate indexes trade a little recall for large speedups.

The papers behind this lesson

  • Beyer, Goldstein, Ramakrishnan and Shaft, When Is "Nearest Neighbor" Meaningful? (1999): https://doi.org/10.1007/3-540-49257-7_15. Proved that for broad families of random data, the nearest and farthest neighbours become equally far as dimension grows: the curse of dimensionality measured in this lesson.
  • Ethayarajh, How Contextual are Contextualized Word Representations? (2019): https://arxiv.org/abs/1909.00512. Showed that transformer embeddings occupy a narrow cone (anisotropy), which is why raw cosine scores bunch together.
  • Su, Cao, Liu and Ou, Whitening Sentence Representations for Better Semantics and Faster Retrieval (2021): https://arxiv.org/abs/2103.15316. Showed that centering and whitening sentence embeddings spreads them back out and improves similarity quality.

Further reading

on GitHub
   1r"""
   2# Similarity: how close are two meanings?
   3
   4Run: `python -m primer.ml.embeddings.similarity`
   5
   6New to the notation (vectors, Σ, ‖x‖)? Start with `primer.notation`, which
   7builds every symbol used here from zero.
   8
   9## Level 1: The practitioner's guide
  10
  11**In one sentence.** A similarity metric turns two embedding vectors into
  12one number that says how alike their meanings are, and the three in common
  13use (cosine, dot product, Euclidean distance) agree with each other only
  14under conditions you have to arrange.
  15
  16**When you need it.** Every time you compare embeddings: semantic search,
  17retrieval for RAG, near-duplicate detection, a semantic cache that answers a
  18question from a stored one, clustering, recommendations. You choose a metric
  19whether you mean to or not, because every vector database and every search
  20library has a default, and the wrong one ranks silently wrong. The tells:
  21you are about to write `ORDER BY` against a vector column and have not
  22checked which operator the embedding model expects; or your team argues
  23about whether "0.8 similar" is a lot. This lesson's demo shows why the
  24second argument can't be settled by intuition: for one model, paraphrases
  25score between 0.76 and 0.87 and unrelated pairs between 0.72 and 0.78, so
  26"0.8" is either a near-duplicate or a stranger. You don't need any of this
  27for exact-match lookups (an id, a hash) or for keyword search, which scores
  28words, not vectors.
  29
  30**Your options.** The ways to score a pair, from the cheapest to the most
  31deliberate:
  32
  33| Option | What it does | What it guarantees | What it costs | Where it lives |
  34|---|---|---|---|---|
  35| Dot product on normalized vectors | Normalize every vector to length 1 once, at write time, then multiply and add | The same ranking as cosine and as Euclidean distance, at the lowest price | One multiply-add per dimension per pair; a normalization step on insert | Your ingestion code, then the database's inner-product operator |
  36| Cosine similarity | Dot product divided by both lengths | Direction only; lengths can't skew the order, whatever came in | Two extra norms per comparison, unless the store caches them | The database's cosine operator, or a library's default |
  37| Raw dot product | Direction and length together | Keeps signal a model put in the length, such as popularity | Rankings that a mix of long and short vectors can dominate | Recommenders and models trained with it |
  38| Euclidean distance | Straight-line gap between the tips | The natural metric for clustering and for many index structures | Disagrees with cosine unless vectors are normalized | k-means, some indexes |
  39| Mean-centred scores | Subtract the collection's average vector before comparing | Unrelated pairs land near 0, similar ones stand clear | A pass over the collection, and a re-run when it changes | Your code, before scoring |
  40| A calibrated threshold | Label a few hundred pairs, sweep the cut-off, keep the best F1 | A yes/no that matches your data and your model | A labelling session, repeated at every model change | A number stored with the model version |
  41
  42**How to choose.** Start from how the model was trained, not from which
  43metric sounds best.
  44
  45- The model's documentation says cosine, or says its vectors are unit
  46  length (most text embedding APIs): normalize on write and rank by dot
  47  product. Same order as cosine, fewer operations.
  48- The model was trained with a raw dot product and encodes something in the
  49  length (recommenders, some retrieval models): keep raw vectors and use the
  50  dot product. Normalizing "to be safe" throws the signal away.
  51- You need an ordered list (search, RAG): rank, and never threshold. A shared
  52  offset in every score leaves the order alone.
  53- You need a yes/no (duplicates, a cache hit, "no good answer found"):
  54  calibrate the threshold on labelled pairs from your own data, for this
  55  model, and store it with the model version.
  56- Whatever you pick, use one metric, one model and one normalization policy
  57  for every vector in the index. A mixture ranks wrong without an error.
  58
  59**What it costs.** A comparison is arithmetic over the vector's dimensions:
  60real models produce 384 to 3,072 numbers per text, so a dot product is a few
  61thousand multiply-adds, and cosine adds two square roots and a division
  62unless the lengths are precomputed. Exact search is one comparison per
  63stored vector; the sentence-transformers documentation puts the practical
  64limit of that brute-force loop at about a million entries, after which you
  65move to an approximate index (`primer.ml.embeddings.ann`). Storage is 4 bytes
  66per dimension as float32 (pgvector charges 4 × dimensions + 8 bytes per
  67vector, half that as 16-bit floats), which `primer.ml.embeddings.compression`
  68shrinks. Normalizing is one pass at write time, once. Calibration costs
  69labelling: a few hundred pairs, scored and swept, and the demo shows the
  70return on it: a guessed threshold of 0.80 reaches F1 0.868 on the demo's
  71pairs, the calibrated 0.776 reaches 0.992.
  72
  73**What breaks.**
  74
  75- **Mixed normalization.** Some vectors normalized, some not, and the
  76  ranking quietly changes: in the demo, raw dot product and cosine order 200
  77  documents differently, and agree only after every vector is normalized.
  78  Normalize in one place, on the write path.
  79- **A threshold from folklore.** A cut-off copied from a blog post or from
  80  the previous model is meaningless: F1 swings from 0.99 to 0.87 across four
  81  hundredths of cosine. Measure it on labelled pairs, per model.
  82- **Every score looks like 0.8.** Transformer embeddings crowd into a narrow
  83  cone (Ethayarajh, 2019, found them "not isotropic in any layer"). Rank
  84  instead of judging absolute scores, or mean-centre: after centring, the
  85  demo's unrelated pairs sit at 0.00 and paraphrases at 0.30.
  86- **A model upgrade with an old index.** A new model is a new space: old
  87  vectors, old thresholds and old scores don't carry over. Re-embed the whole
  88  collection and recalibrate (`primer.ml.embeddings.operations`).
  89- **The operator is a distance, not a similarity.** Databases often expose
  90  cosine *distance* (1 − cosine) and the *negative* inner product so that
  91  ascending order is best-first, as pgvector does. Read the sign before you
  92  sort.
  93- **Queries and documents embedded differently.** Some providers train
  94  separate treatments for the query and for the passage (Cohere's
  95  `input_type` of search_query and search_document); embed both sides the way
  96  the model expects or the scores are off.
  97- **Nearest means little on unstructured data.** With 1,000 random points,
  98  the nearest is 0.7% as far as the farthest in 2 dimensions and 90% as far
  99  in 1,000: the curse of dimensionality. Real embeddings are structured, so
 100  search works, but a score gap that small is a warning that the vectors
 101  carry little.
 102
 103**In the wild.** pgvector exposes one operator per metric on a Postgres
 104column: `<->` for Euclidean distance, `<#>` for the negative inner product,
 105`<=>` for cosine distance, plus Hamming and Jaccard on bit vectors, up to
 10616,000 dimensions. sentence-transformers' semantic search defaults to cosine
 107and offers dot-product scoring for normalized vectors. OpenAI's embedding
 108guide states that its embeddings are normalized to length 1, recommends
 109cosine, and notes that it can then be computed as a plain dot product; Cohere
 110asks for an `input_type` per side of the search. Faiss's flat indexes
 111(IndexFlatIP, IndexFlatL2) are the exact inner-product and Euclidean
 112searches every approximate index is measured against. Ethayarajh (2019)
 113measured anisotropy in contextual models and Su et al. (2021) showed that
 114centring and whitening spreads the vectors back out.
 115
 116**Go deeper.** Level 2 computes all three metrics by hand on three
 1173-dimensional vectors, proves in one line why normalizing makes them agree
 118(the squared distance is 2 − 2 × cosine), builds the popularity example where
 119cosine and dot product disagree on purpose, measures the curse of
 120dimensionality with random points, and manufactures anisotropic embeddings
 121to show a calibrated threshold beating a guessed one. If you only needed to
 122choose, you are done.
 123
 124## Level 2: How it works, from scratch
 125
 126An **embedding** gives every piece of text a set of coordinates on a map of
 127meaning. Texts about similar things live on the same street; unrelated ones
 128are across town. Draw an arrow from the centre of the map to each text's
 129spot, and "how similar are these two texts?" becomes "how alike are these two
 130arrows?".
 131
 132There are three natural ways to compare two arrows:
 133
 134- **Cosine similarity** compares the *direction* they point and ignores how
 135  long they are, like two people pointing at the same mountain, one with a
 136  longer arm.
 137- **Euclidean distance** is the straight-line distance between the two arrow
 138  tips, like measuring with a ruler.
 139- **Dot product** mixes both: pointing the same way *and* being long both
 140  make it bigger.
 141
 142## A tiny worked example
 143
 144Three texts, each placed in a 3-dimensional space (real models use 384 to
 1453,072 **dimensions**, meaning numbers per vector; the arithmetic is the same):
 146
 147a = (1, 2, 2), b = (2, 1, 2), c = (2, −1, −2)
 148
 1491. **Dot product**: multiply position by position, then add.
 150   a·b = 1·2 + 2·1 + 2·2 = **8**. a·c = 1·2 + 2·(−1) + 2·(−2) = **−4**.
 1512. **Length** (the **norm**, written ‖a‖): square each number, add, and take
 152   the square root. ‖a‖ = √(1 + 4 + 4) = **3**, and b and c also have length 3.
 1533. **Cosine**: dot product divided by both lengths.
 154   cos(a, b) = 8 / (3·3) = **0.89**. cos(a, c) = −4 / 9 = **−0.44**.
 1554. **Euclidean distance**: subtract, square, add, square-root.
 156   ‖a − b‖ = √(1 + 1 + 0) = **1.41**. ‖a − c‖ = √(1 + 9 + 16) = **5.10**.
 157
 158| Pair | Dot product | Lengths multiplied | Cosine | Distance | Reading |
 159|---|---|---|---|---|---|
 160| a, b | 8 | 9 | 0.89 | 1.41 | Very similar |
 161| a, c | −4 | 9 | −0.44 | 5.10 | Pointing away |
 162
 163Cosine runs from −1 (opposite directions) through 0 (unrelated, at right
 164angles) to 1 (same direction).
 165
 166```mermaid
 167flowchart LR
 168  A["a = (1, 2, 2)"] --> D["dot product<br/>a·b = 8"]
 169  B["b = (2, 1, 2)"] --> D
 170  A --> NA["length ‖a‖ = 3"]
 171  B --> NB["length ‖b‖ = 3"]
 172  D --> COS["cosine = 8 / (3 × 3) = 0.89"]
 173  NA --> COS
 174  NB --> COS
 175  A --> SUB["difference a − b = (−1, 1, 0)"]
 176  B --> SUB
 177  SUB --> L2["distance = √2 = 1.41"]
 178```
 179
 180**Reading it:** the two input vectors feed three calculations. The top path
 181is the dot product. The middle paths measure each vector's length, and
 182cosine is simply the dot product divided by those lengths, which is how it
 183comes to "ignore length". The bottom path subtracts the vectors and measures
 184what's left: that's distance. Everything later in this lesson is these three
 185boxes.
 186
 187**In code:** `worked_example` computes every number in the table above from
 188a, b and c, so you can check your hand arithmetic against it.
 189
 190## The math
 191
 192$$
 193\cos(a, b) = \frac{a \cdot b}{\lVert a \rVert \, \lVert b \rVert}
 194\qquad
 195a \cdot b = \sum_{i=1}^{d} a_i b_i
 196\qquad
 197\lVert a - b \rVert = \sqrt{\sum_{i=1}^{d} (a_i - b_i)^2}
 198$$
 199
 200**Symbols**
 201
 202| Symbol | Meaning here | Shape / range |
 203|---|---|---|
 204| a, b | the two embedding vectors being compared | d numbers each |
 205| d | number of dimensions | 3 in the example; 384 to 3,072 in real models |
 206| i | which position (dimension) we're looking at | 1 to d |
 207| aᵢ | the i-th number in a | a real number |
 208| Σ | "add up", over i from 1 to d | |
 209| a · b | dot product: multiply matching positions and add | one number, any sign |
 210| ‖a‖ | length (norm) of a: √(Σ aᵢ²) | one number ≥ 0 |
 211| √ | square root | |
 212| cos(a, b) | cosine of the angle between a and b | −1 to 1 |
 213| ‖a − b‖ | Euclidean distance between the tips | ≥ 0; 0 means identical |
 214
 215**In words:** the cosine of a and b is their dot product divided by the
 216product of their lengths; the dot product adds up the products of matching
 217positions; the distance is the square root of the summed squared
 218differences.
 219
 220**On the example:** cos(a, b) = (1·2 + 2·1 + 2·2) / (3 · 3) = 8 / 9 = 0.89;
 221‖a − b‖ = √((1−2)² + (2−1)² + (2−2)²) = √2 = 1.41.
 222
 223**In Python:**
 224
 225```python
 226import math
 227a, b = [1, 2, 2], [2, 1, 2]
 228# a · b = Σ a_i b_i
 229dot = sum(a_i * b_i for a_i, b_i in zip(a, b))
 230dot  # → 8
 231# ‖a‖
 232norm_a = math.sqrt(sum(a_i ** 2 for a_i in a))
 233# ‖b‖
 234norm_b = math.sqrt(sum(b_i ** 2 for b_i in b))
 235norm_a, norm_b  # → (3.0, 3.0)
 236# cos(a, b)
 237round(dot / (norm_a * norm_b), 2)  # → 0.89
 238# ‖a - b‖
 239round(math.sqrt(sum((a_i - b_i) ** 2 for a_i, b_i in zip(a, b))), 2)  # → 1.41
 240```
 241
 242The code is `cosine`, `dot` and `euclidean`, each a line of NumPy.
 243
 244**In code:** `rank` sorts a whole collection of documents from best to worst
 245match under whichever of the three metrics you name.
 246
 247**Why it matters:** every vector database, semantic search box and RAG
 248system does exactly this comparison, millions of times per second. Knowing
 249which of the three a system uses, and why, is the difference between
 250correct and silently wrong rankings.
 251
 252## Normalize, and the three agree
 253
 254**Everyday picture:** trim every arrow to the same length, one unit. Now the
 255only thing that can differ is direction, so "tips are close together" and
 256"arrows point the same way" become the same statement.
 257
 258**Tiny example:** a = (1, 0) and b = (0.6, 0.8) both have length 1 (check:
 2590.6² + 0.8² = 0.36 + 0.64 = 1). Their cosine is 1·0.6 + 0·0.8 = 0.6. The squared
 260distance is (1 − 0.6)² + (0 − 0.8)² = 0.16 + 0.64 = 0.8, and 2 − 2·0.6 = 0.8.
 261Same number.
 262
 263$$
 264\lVert a - b \rVert^2 = \lVert a \rVert^2 + \lVert b \rVert^2 - 2\,a\cdot b = 2 - 2\cos(a, b)
 265\quad \text{when } \lVert a \rVert = \lVert b \rVert = 1
 266$$
 267
 268**Symbols**
 269
 270| Symbol | Meaning here | Shape / range |
 271|---|---|---|
 272| ‖a − b‖² | squared distance between the tips | 0 to 4 for unit vectors |
 273| ‖a‖², ‖b‖² | squared lengths, both 1 after normalizing | 1 |
 274| a · b | dot product, equal to cosine for unit vectors | −1 to 1 |
 275| cos(a, b) | cosine similarity | −1 to 1 |
 276
 277**In words:** for unit-length vectors, the squared distance is two minus
 278twice the cosine, so the higher the cosine, the smaller the distance, always.
 279
 280**On the example:** ‖a − b‖² = 1 + 1 − 2·0.6 = 0.8 = 2 − 2·0.6.
 281
 282**In Python:**
 283
 284```python
 285# both length 1
 286a, b = [1, 0], [0.6, 0.8]
 287# for unit vectors, a · b is the cosine
 288cos = sum(a_i * b_i for a_i, b_i in zip(a, b))
 289# ‖a - b‖²
 290dist_sq = sum((a_i - b_i) ** 2 for a_i, b_i in zip(a, b))
 291# the two sides of the identity
 292round(dist_sq, 2), round(2 - 2 * cos, 2)  # → (0.8, 0.8)
 293```
 294
 295**L2-normalizing** (dividing a vector by its own length) is what trims every
 296arrow to length 1. After it, the dot product *is* the cosine, and distance is
 297a decreasing function of the cosine. So **once vectors are L2-normalized,
 298cosine, dot product and Euclidean distance give the same ranking.**
 299
 300**Try it:** pick "Same direction, different lengths", then drag B's length.
 301The dot product and the distance change, but the cosine stays at 1, because
 302it only sees the angle. Drag an angle instead and watch the cosine slide from
 3031 through 0 to −1. Then switch on "Normalize to length 1": both arrows shrink
 304to the unit circle, the dot product becomes the cosine, and the squared
 305distance becomes 2 − 2 × cosine.
 306
 307<div class="viz" data-viz="cosine-similarity" aria-label="Cosine similarity of two arrows"></div>
 308
 309![All 400 random unit-vector pairs fall exactly on the line squared distance = 2 - 2 x cosine, so higher cosine always means smaller distance](figures/primer.ml.embeddings.similarity.identity.svg)
 310
 311**Reading it:** each dot is one pair of random unit vectors. The horizontal
 312axis is their cosine and the vertical axis is their squared Euclidean
 313distance. Every dot sits exactly on the line y = 2 − 2x. Because the line only
 314goes down, "higher cosine" and "smaller distance" always mean the same thing,
 315so the two metrics can never disagree about which neighbour is nearer.
 316
 317```mermaid
 318flowchart TD
 319  S{What was the model<br/>trained with?} -->|cosine| N[L2-normalize every vector<br/>once, at write time]
 320  S -->|dot product| D{Does the vector length<br/>carry signal?}
 321  D -->|yes, e.g. popularity| RAW[Keep raw vectors,<br/>rank by dot product]
 322  D -->|no| N
 323  N --> FAST[Rank by dot product:<br/>same order as cosine and L2,<br/>and the cheapest to compute]
 324```
 325
 326**Reading it:** start at the top with the question that matters, which is
 327how the model was trained, not which metric sounds best. Most text
 328embedding models are trained with cosine, so you normalize once when you
 329write vectors, then use the plain dot product at query time. It gives the
 330same order as cosine and costs less. Only when length is deliberate signal
 331do you skip normalizing.
 332
 333**In code:** `l2_normalize` trims every row to length 1;
 334`rankings_agree_after_normalizing` ranks random documents by all three metrics
 335before and after normalizing and reports which rankings match, and
 336`l2_identity_check` returns both sides of the identity for two random vectors.
 337
 338**Why it matters:** vector databases typically normalize on insert and
 339then use the dot product, the cheapest of the three. If you mix normalized
 340and unnormalized vectors, rankings quietly change. **Use the metric the
 341embedding model was trained with.**
 342
 343## When length carries signal
 344
 345**Everyday picture:** on the map of meaning, some shops are big and popular.
 346A recommender may deliberately make popular items' arrows longer, so they
 347win ties. Cosine throws that away; the dot product keeps it.
 348
 349**Tiny example** (`popularity_in_the_norm`): the query points along
 350(1, 0). A niche item (0.99, 0.10) is almost perfectly aligned. A popular item
 351(0.95, 0.30) × 3 = (2.85, 0.90) is slightly off but three times longer.
 352
 353| Item | Cosine with query | Dot with query |
 354|---|---|---|
 355| niche (0.99, 0.10) | 0.99 / 0.995 = 0.995 | 0.99 |
 356| popular (2.85, 0.90) | 2.85 / 2.99 = 0.954 | 2.85 |
 357
 358Cosine prefers the niche item; the dot product prefers the popular one.
 359
 360![The short item points almost along the query and wins on cosine (0.99 vs 0.95); the item three times longer wins on dot product (2.85 vs 0.99)](figures/primer.ml.embeddings.similarity.arrows.svg)
 361
 362**Reading it:** the black arrow is the query. The blue arrow (niche item)
 363points almost exactly the same way but is short. The orange arrow (popular
 364item) leans a little away but is three times longer. Cosine only looks at
 365the angle between each arrow and the query, so blue wins. The dot product
 366also rewards length, so orange wins. Neither is wrong: it depends on what
 367the model was trained to encode.
 368
 369**Why it matters:** some retrieval and recommender models are trained with
 370the raw dot product and put useful signal in the length. Normalizing their
 371vectors "to be safe" throws that signal away.
 372
 373## The curse of dimensionality
 374
 375**Everyday picture:** in a small village some neighbours live next door and
 376some live across the valley. Now imagine a city with thousands of streets
 377running in thousands of directions: pick random strangers and they all turn
 378out to live about equally far from you. "Nearest" stops meaning much.
 379
 380**Tiny example:** scatter 1,000 random points in a cube and measure from a
 381random spot. In 2 dimensions the nearest point is about 1/100th as far away
 382as the farthest. In 1,000 dimensions it's about 90% as far: everyone is
 383nearly equidistant.
 384
 385$$
 386\text{ratio} = \frac{\min_j \lVert q - x_j \rVert}{\max_j \lVert q - x_j \rVert}
 387$$
 388
 389**Symbols**
 390
 391| Symbol | Meaning here | Shape / range |
 392|---|---|---|
 393| q | the query point | d numbers |
 394| xⱼ | the j-th of the random points | d numbers each |
 395| j | which random point | 1 to 1,000 |
 396| min, max over j | the smallest / largest value across all the points | |
 397| ratio | nearest distance divided by farthest distance | 0 to 1; near 1 means "all equally far" |
 398
 399**In words:** the ratio is the distance to the closest point divided by the
 400distance to the farthest point.
 401
 402**On the example:** by hand first: a query at (0, 0) and three points at
 403(1, 0), (0, 2) and (3, 4) are 1, 2 and 5 away, so the ratio is 1 / 5 = 0.2.
 404With 1,000 random points the demo table shows about 0.01 at d = 2 and about
 4050.9 at d = 1,000. The Python below draws its own random points, so it lands
 406close by rather than on the same digits: 0.02 and 0.9.
 407
 408**In Python:**
 409
 410```python
 411import math, random
 412def ratio(q, xs):
 413    # ‖q - x_j‖ for every j
 414    dists = [math.dist(q, x_j) for x_j in xs]
 415    # min over j ÷ max over j
 416    return min(dists) / max(dists)
 417ratio((0, 0), [(1, 0), (0, 2), (3, 4)])  # → 0.2
 418random.seed(0)
 419for d in (2, 1000):
 420    q = [random.random() for _ in range(d)]
 421    xs = [[random.random() for _ in range(d)] for _ in range(1000)]
 422    # one significant figure
 423    print(d, f"{ratio(q, xs):.1g}")  # → 2 0.02 1000 0.9
 424```
 425
 426![The nearest-to-farthest distance ratio rises from about 0 in 2 dimensions to 0.8 at 200 and 0.92 at 2,000: random points become almost equidistant](figures/primer.ml.embeddings.similarity.curse.svg)
 427
 428**Reading it:** the horizontal axis is the number of dimensions (log scale);
 429the vertical axis is the nearest-to-farthest ratio above. In 2-D the ratio is
 430near 0: the nearest point really is much nearer. By a few hundred dimensions
 431the curve flattens toward 1, so every random point is roughly equally far
 432away, and "nearest" tells you almost nothing about random data.
 433
 434**In code:** `curse_of_dimensionality` scatters the random points, measures
 435from a random query, and returns the nearest, farthest and ratio for each
 436dimension.
 437
 438**Why it matters:** exact nearest-neighbour search in high dimensions gets
 439slow and fragile, which is one reason vector search uses approximate indexes
 440(`primer.ml.embeddings.ann`). Real embeddings aren't random: meaning puts
 441them on a much lower-dimensional structure, so nearest-neighbour search still
 442works well in practice.
 443
 444## Anisotropy: why every score looks like 0.8
 445
 446**Everyday picture:** a stadium crowd all facing the stage. Ask "are these
 447two people facing the same way?" and the answer is "roughly yes" for
 448everyone, so the question tells you little. Many embedding models behave
 449like that crowd: their vectors bunch in a narrow cone, so even unrelated
 450texts get cosines of 0.7 or more. This lopsidedness is called
 451**anisotropy** ("not the same in every direction").
 452
 453**Tiny example:** build each vector as √3·m + n, where m is one direction
 454everybody shares (the stage) and n is a unit-length direction of each
 455text's own. For two unrelated texts the n parts are at right angles, so
 456their dot product is (√3)² = 3 and each length is √(3 + 1) = 2, and the
 457cosine is 3 / (2 · 2) = **0.75**. For paraphrases, whose own parts agree with
 458cosine ρ, the cosine becomes 0.75 + 0.25·ρ: about 0.79 to 0.86. All the signal
 459is squeezed into a few hundredths.
 460
 461$$
 462v = \alpha\, m + \beta\, n, \qquad
 463\cos(v_1, v_2) = \frac{\alpha^2 + \beta^2 \rho}{\alpha^2 + \beta^2}
 464$$
 465
 466**Symbols**
 467
 468| Symbol | Meaning here | Shape / range |
 469|---|---|---|
 470| v, v₁, v₂ | an embedding vector | d numbers (768 here) |
 471| m | the shared direction every vector leans toward | unit vector |
 472| n | the text's own direction | unit vector |
 473| α (alpha) | how strongly every vector leans toward m | √3 here |
 474| β (beta) | how strongly each keeps its own direction | 1 here |
 475| ρ (rho) | cosine between the two texts' own directions | 0 unrelated; 0.15 to 0.45 for paraphrases |
 476
 477**In words:** every vector is a big shared part plus a small personal part,
 478so the cosine is mostly the shared part's weight, nudged up a little by how
 479much the personal parts agree.
 480
 481**On the example:** unrelated, ρ = 0: (3 + 0) / (3 + 1) = 0.75. A paraphrase with
 482ρ = 0.4: (3 + 0.4) / 4 = 0.85.
 483
 484**In Python:**
 485
 486```python
 487import math
 488alpha, beta = math.sqrt(3), 1
 489# shared direction; two unrelated own directions
 490m, n1, n2 = (1, 0, 0), (0, 1, 0), (0, 0, 1)
 491# v = α m + β n
 492v1 = [alpha * m_i + beta * n_i for m_i, n_i in zip(m, n1)]
 493v2 = [alpha * m_i + beta * n_i for m_i, n_i in zip(m, n2)]
 494dot = sum(a * b for a, b in zip(v1, v2))
 495# the cosine, measured directly
 496round(dot / (math.hypot(*v1) * math.hypot(*v2)), 2)  # → 0.75
 497def cos_formula(rho):
 498    return (alpha ** 2 + beta ** 2 * rho) / (alpha ** 2 + beta ** 2)
 499# unrelated, and a paraphrase
 500round(cos_formula(0), 2), round(cos_formula(0.4), 2)  # → (0.75, 0.85)
 501```
 502
 503![Raw cosines crowd paraphrases (0.76 to 0.87) and unrelated pairs (0.72 to 0.78) together; mean-centred, unrelated sit near 0 and paraphrases near 0.3](figures/primer.ml.embeddings.similarity.anisotropy.svg)
 504
 505**Reading it:** the left panel shows raw cosine scores for 300 paraphrase
 506pairs (one colour) and 300 unrelated pairs (the other). Both piles sit
 507between 0.7 and 0.9, so a score of "0.8" alone says little. The right panel
 508shows the same pairs after **mean-centering**, which means subtracting the
 509average vector of the whole collection (that removes the shared "stage"
 510direction). Unrelated pairs drop to about 0, paraphrases sit clearly above,
 511and the gap is easy to see. The information was there all along, squeezed
 512into a narrow band.
 513
 514**In code:** `anisotropic_embeddings` builds the pairs as α·m + β·n,
 515`pair_cosines` scores each pair, and `mean_center` subtracts the average
 516vector to remove the shared direction.
 517
 518### Picking a threshold from data, not folklore
 519
 520Sometimes you need a yes/no cut-off: "are these duplicates?", "is this
 521cached answer close enough?". Pick it by measuring. Label some pairs,
 522score them, and choose the threshold with the best **F1**, the balance of
 523**precision** (of the pairs I called similar, how many were?) and **recall**
 524(of the truly similar pairs, how many did I catch?):
 525
 526$$
 527F_1 = \frac{2PR}{P + R}
 528$$
 529
 530**Symbols**
 531
 532| Symbol | Meaning here | Range |
 533|---|---|---|
 534| P | precision: true matches ÷ everything called a match | 0 to 1 |
 535| R | recall: true matches found ÷ all true matches | 0 to 1 |
 536| F₁ | their harmonic mean; high only when both are high | 0 to 1 |
 537
 538**In words:** F1 is twice the product of precision and recall, divided by
 539their sum.
 540
 541**On a hand example:** scores 0.9, 0.8, 0.7, 0.6 where the first two are true
 542matches. Threshold 0.8 calls exactly those two matches: P = 2/2 = 1,
 543R = 2/2 = 1, F₁ = 2·1·1 / 2 = 1. `best_threshold` finds 0.8.
 544
 545**In Python:**
 546
 547```python
 548scores = [0.9, 0.8, 0.7, 0.6]
 549# the labels
 550is_match = [True, True, False, False]
 551# threshold 0.8
 552called = [s >= 0.8 for s in scores]
 553true_matches = sum(c and y for c, y in zip(called, is_match))
 554# precision
 555P = true_matches / sum(called)
 556# recall
 557R = true_matches / sum(is_match)
 558# F₁
 5592 * P * R / (P + R)  # → 1.0
 560```
 561
 562```mermaid
 563flowchart LR
 564  L[Sample a few hundred pairs<br/>from your own data] --> H[Label each:<br/>same meaning or not]
 565  H --> E[Score every pair<br/>with THIS model]
 566  E --> W[Sweep thresholds,<br/>compute precision, recall, F1]
 567  W --> P[Pick the threshold that fits<br/>your cost of each error]
 568  P --> V[Store it with the<br/>model version]
 569  V -->|model upgraded| E
 570```
 571
 572**Reading it:** a threshold is a property of one model on one kind of data,
 573so it is *measured*, not guessed. Label pairs, score them, sweep, and pick.
 574The loop back from "model upgraded" is the step people forget: a new model
 575has a different score distribution, so the old threshold is meaningless.
 576
 577![F1 on raw scores swings within a few hundredths: the calibrated threshold 0.776 reaches F1 0.99, while the guessed 0.80 drops to 0.87](figures/primer.ml.embeddings.similarity.threshold.svg)
 578
 579**Reading it:** the curve is F1 at each candidate threshold on the raw
 580anisotropic scores. The dashed line is the "obvious" guess of 0.8; the
 581solid line is the calibrated threshold. F1 changes sharply over a few
 582hundredths of cosine, which is exactly why a threshold carried over from a
 583different model (or a blog post) can be badly wrong.
 584
 585**In code:** `f1_curve` computes F1 at every candidate threshold for this
 586plot; `best_threshold` returns the single best one with its precision and
 587recall.
 588
 589**Why it matters:**
 590
 591* Absolute scores mean different things for different models. "0.8" is
 592  not a universal "similar".
 593* **Rank, don't threshold, whenever you can.** A shared offset doesn't change
 594  which document is on top.
 595* When you *must* threshold (dedup, semantic caching, "no good answer
 596  found"), calibrate on labeled pairs for that model, and redo it when the
 597  model changes.
 598
 599## In 20 seconds
 600- Cosine compares direction, the dot product compares direction and length,
 601  Euclidean distance measures the gap between the tips.
 602- On L2-normalized vectors all three give the same ranking, because
 603  ‖a − b‖² = 2 − 2cos.
 604- Use the metric the model was trained with.
 605- Scores aren't comparable across models (anisotropy). Rank when you can;
 606  calibrate thresholds on labeled pairs when you can't.
 607
 608## Self-test questions
 609
 610**Q: When do cosine, dot product and Euclidean distance give the same ranking?**
 611When all vectors are L2-normalized: the dot product equals the cosine, and
 612the squared distance equals 2 − 2cos, which always moves opposite to cosine.
 613
 614**Q: When would you deliberately use the raw dot product on unnormalized vectors?**
 615When the model was trained that way and encodes useful signal in the length
 616(e.g. popularity or quality in some recommender and retrieval models).
 617Normalizing would throw that signal away.
 618
 619**Q: Similarity scores all sit between 0.75 and 0.85. What's going on and how do you set a threshold?**
 620The model is anisotropic: its vectors share a dominant direction, so every
 621pair looks similar and the real signal lives in a narrow band. Ranking still
 622works. For a threshold, label a few hundred pairs from your own data, sweep
 623the threshold, and pick the one that maximizes F1 (or matches your
 624precision/recall needs). Optionally mean-center the vectors. Redo it for
 625every model version.
 626
 627**Q: What is the curse of dimensionality for nearest-neighbour search?**
 628For random high-dimensional data, the nearest and farthest points are almost
 629the same distance away, so "nearest" carries little information and exact
 630search gets expensive. Real embeddings have structure, and approximate
 631indexes trade a little recall for large speedups.
 632
 633## The papers behind this lesson
 634
 635- **Beyer, Goldstein, Ramakrishnan and Shaft, *When Is "Nearest Neighbor" Meaningful?* (1999)**: https://doi.org/10.1007/3-540-49257-7_15.
 636  Proved that for broad families of random data, the nearest and farthest neighbours become equally far as dimension grows: the curse of dimensionality measured in this lesson.
 637- **Ethayarajh, *How Contextual are Contextualized Word Representations?* (2019)**: https://arxiv.org/abs/1909.00512.
 638  Showed that transformer embeddings occupy a narrow cone (anisotropy), which is why raw cosine scores bunch together.
 639- **Su, Cao, Liu and Ou, *Whitening Sentence Representations for Better Semantics and Faster Retrieval* (2021)**: https://arxiv.org/abs/2103.15316.
 640  Showed that centering and whitening sentence embeddings spreads them back out and improves similarity quality.
 641
 642## Further reading
 643- Cosine similarity (Wikipedia): https://en.wikipedia.org/wiki/Cosine_similarity
 644- Beyer et al., *When Is "Nearest Neighbor" Meaningful?* (1999): https://doi.org/10.1007/3-540-49257-7_15
 645- Ethayarajh, *How Contextual are Contextualized Word Representations?* (anisotropy, 2019): https://arxiv.org/abs/1909.00512
 646- Su et al., *Whitening Sentence Representations* (2021): https://arxiv.org/abs/2103.15316
 647- sentence-transformers, semantic search and score functions: https://www.sbert.net/examples/applications/semantic-search/README.html
 648"""
 649
 650from __future__ import annotations
 651
 652import numpy as np
 653
 654from primer._show import banner, say, table, takeaway
 655
 656# ---------------------------------------------------------------------------
 657# 1. The three measures
 658# ---------------------------------------------------------------------------
 659
 660
 661def l2_normalize(X: np.ndarray, axis: int = -1, eps: float = 1e-12) -> np.ndarray:
 662    """Scale each vector (row, by default) to unit length.
 663
 664    `eps` avoids dividing by zero for an all-zero vector (e.g. empty text).
 665    """
 666    return X / np.maximum(np.linalg.norm(X, axis=axis, keepdims=True), eps)
 667
 668
 669def dot(a: np.ndarray, b: np.ndarray) -> float:
 670    """Sum of element-wise products. Sensitive to both angle and length."""
 671    return float(np.dot(a, b))
 672
 673
 674def cosine(a: np.ndarray, b: np.ndarray) -> float:
 675    """cos(angle) = a·b / (|a||b|). In [-1, 1]; ignores length."""
 676    return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))
 677
 678
 679def euclidean(a: np.ndarray, b: np.ndarray) -> float:
 680    """Straight-line distance. Smaller means closer."""
 681    return float(np.linalg.norm(a - b))
 682
 683
 684def rank(query: np.ndarray, docs: np.ndarray, metric: str) -> np.ndarray:
 685    """Doc indices from best to worst match under `metric` (cosine | dot | l2)."""
 686    if metric == "dot":
 687        scores = docs @ query
 688    elif metric == "cosine":
 689        scores = (docs @ query) / (np.linalg.norm(docs, axis=1) * np.linalg.norm(query))
 690    elif metric == "l2":
 691        scores = -np.linalg.norm(docs - query, axis=1)  # negate: closer = higher
 692    else:
 693        raise ValueError(metric)
 694    # Stable sort so ties break the same way for every metric.
 695    return np.argsort(-scores, kind="stable")
 696
 697
 698# ---------------------------------------------------------------------------
 699# 2. Worked example
 700# ---------------------------------------------------------------------------
 701
 702A = np.array([1.0, 2.0, 2.0])
 703B = np.array([2.0, 1.0, 2.0])
 704C = np.array([2.0, -1.0, -2.0])
 705
 706
 707def worked_example() -> dict[str, float]:
 708    """a, b, c from the table above. Returns the numbers you should be able to compute by hand."""
 709    return {
 710        "len_a": float(np.linalg.norm(A)),
 711        "dot_ab": dot(A, B),
 712        "dot_ac": dot(A, C),
 713        "cos_ab": cosine(A, B),
 714        "cos_ac": cosine(A, C),
 715        "l2_ab": euclidean(A, B),
 716        "l2_ac": euclidean(A, C),
 717    }
 718
 719
 720# ---------------------------------------------------------------------------
 721# 3. Normalization makes the metrics agree; without it they can disagree
 722# ---------------------------------------------------------------------------
 723
 724
 725def rankings_agree_after_normalizing(n_docs: int = 200, dim: int = 32, seed: int = 0) -> dict[str, bool]:
 726    """Random docs with wildly different lengths: do the three rankings agree?
 727
 728    Before normalizing: dot product favors long vectors, so it disagrees.
 729    After normalizing: all three rankings are identical, as the algebra says.
 730    """
 731    rng = np.random.default_rng(seed)
 732    docs = rng.standard_normal((n_docs, dim)) * rng.uniform(0.2, 5.0, (n_docs, 1))  # varied lengths
 733    q = rng.standard_normal(dim)
 734
 735    raw = {m: rank(q, docs, m) for m in ("cosine", "dot", "l2")}
 736    dn, qn = l2_normalize(docs), l2_normalize(q)
 737    norm = {m: rank(qn, dn, m) for m in ("cosine", "dot", "l2")}
 738    return {
 739        "raw_dot_equals_cosine": bool(np.array_equal(raw["dot"], raw["cosine"])),
 740        "raw_l2_equals_cosine": bool(np.array_equal(raw["l2"], raw["cosine"])),
 741        "norm_dot_equals_cosine": bool(np.array_equal(norm["dot"], norm["cosine"])),
 742        "norm_l2_equals_cosine": bool(np.array_equal(norm["l2"], norm["cosine"])),
 743    }
 744
 745
 746def l2_identity_check(seed: int = 0) -> tuple[float, float]:
 747    """Return (||a-b||², 2 - 2cos(a,b)) for two random unit vectors. They're equal."""
 748    rng = np.random.default_rng(seed)
 749    a, b = l2_normalize(rng.standard_normal(16)), l2_normalize(rng.standard_normal(16))
 750    return float(np.sum((a - b) ** 2)), 2 - 2 * cosine(a, b)
 751
 752
 753def popularity_in_the_norm() -> list[tuple[str, float, float]]:
 754    """When magnitude carries signal, dot product and cosine rank differently.
 755
 756    Two items point in almost the same direction as the query. The "popular"
 757    one has a long vector (a model trained with dot product can learn to do
 758    this). Cosine ignores that and prefers the slightly better-aligned niche
 759    item; dot product prefers the popular one.
 760    """
 761    q = np.array([1.0, 0.0])
 762    items = {
 763        "niche item (well aligned, short)": np.array([0.99, 0.10]) * 1.0,
 764        "popular item (a bit off, long)": np.array([0.95, 0.30]) * 3.0,
 765    }
 766    return [(name, cosine(q, v), dot(q, v)) for name, v in items.items()]
 767
 768
 769# ---------------------------------------------------------------------------
 770# 4. Curse of dimensionality
 771# ---------------------------------------------------------------------------
 772
 773
 774def curse_of_dimensionality(dims=(2, 10, 100, 1000), n_points: int = 1000, seed: int = 0) -> list[dict]:
 775    """For random points, how much nearer is the nearest than the farthest?
 776
 777    ratio = nearest / farthest distance from a random query. Near 0 means
 778    "nearest" is meaningful; near 1 means every point is about equally far.
 779    """
 780    rng = np.random.default_rng(seed)
 781    rows = []
 782    for d in dims:
 783        X = rng.uniform(0, 1, (n_points, d))
 784        q = rng.uniform(0, 1, d)
 785        dist = np.linalg.norm(X - q, axis=1)
 786        rows.append({"dim": d, "nearest": float(dist.min()), "farthest": float(dist.max()), "ratio": float(dist.min() / dist.max())})
 787    return rows
 788
 789
 790# ---------------------------------------------------------------------------
 791# 5. Anisotropy, and thresholds chosen from labeled pairs
 792# ---------------------------------------------------------------------------
 793
 794
 795def anisotropic_embeddings(n_pairs: int = 300, dim: int = 768, seed: int = 0):
 796    """Simulate a model whose vectors all live in a narrow cone.
 797
 798    Each vector = sqrt(3)·m + 1·n, where m is one shared unit direction and n
 799    is a per-text unit direction. For unrelated texts the n parts are nearly
 800    orthogonal, so cos ≈ 3 / (3 + 1) = 0.75. For paraphrases we make the n parts
 801    correlated (rho in [0.15, 0.45]), so cos ≈ 0.75 + 0.25·rho, about 0.79 to 0.86.
 802    Everything lands in the 0.75-0.85 band even though the underlying
 803    signal (rho) is clear.
 804
 805    Returns (left, right, is_paraphrase): left[i] and right[i] form pair i.
 806    """
 807    rng = np.random.default_rng(seed)
 808    m = l2_normalize(rng.standard_normal(dim))
 809    alpha, beta = np.sqrt(3.0), 1.0
 810
 811    def unit(k):
 812        return l2_normalize(rng.standard_normal((k, dim)))
 813
 814    n_left = unit(2 * n_pairs)
 815    labels = np.array([True] * n_pairs + [False] * n_pairs)
 816    rho = np.where(labels, rng.uniform(0.15, 0.45, 2 * n_pairs), 0.0)
 817    # Correlated partner: rho·n + sqrt(1 - rho²)·fresh, which is unit length with cos(n, partner) ≈ rho.
 818    n_right = l2_normalize(rho[:, None] * n_left + np.sqrt(1 - rho[:, None] ** 2) * unit(2 * n_pairs))
 819    left = alpha * m + beta * n_left
 820    right = alpha * m + beta * n_right
 821    return left, right, labels
 822
 823
 824def pair_cosines(left: np.ndarray, right: np.ndarray) -> np.ndarray:
 825    return np.sum(l2_normalize(left) * l2_normalize(right), axis=1)
 826
 827
 828def mean_center(*arrays: np.ndarray) -> list[np.ndarray]:
 829    """Subtract the mean vector of all given vectors (the shared "cone" direction).
 830
 831    In production you'd estimate the mean once from a sample of your corpus
 832    and store it with the index; queries get the same shift.
 833    """
 834    mu = np.concatenate(arrays).mean(axis=0)
 835    return [a - mu for a in arrays]
 836
 837
 838def best_threshold(scores: np.ndarray, labels: np.ndarray) -> dict[str, float]:
 839    """Sweep every candidate threshold; return the one with the best F1.
 840
 841    Predict "similar" when score >= threshold. This is how you pick a dedup
 842    or cache threshold for a specific model, from a few hundred labeled pairs,
 843    instead of guessing "0.8".
 844    """
 845    best = {"threshold": 0.0, "f1": -1.0, "precision": 0.0, "recall": 0.0}
 846    for t in np.unique(scores):
 847        pred = scores >= t
 848        tp = np.sum(pred & labels)
 849        if tp == 0:
 850            continue
 851        precision, recall = tp / pred.sum(), tp / labels.sum()
 852        f1 = 2 * precision * recall / (precision + recall)
 853        if f1 > best["f1"]:
 854            best = {"threshold": float(t), "f1": float(f1), "precision": float(precision), "recall": float(recall)}
 855    return best
 856
 857
 858def f1_curve(scores: np.ndarray, labels: np.ndarray, thresholds: np.ndarray) -> np.ndarray:
 859    """F1 of the rule "score >= t" for each t in `thresholds` (for plotting)."""
 860    out = []
 861    for t in thresholds:
 862        pred = scores >= t
 863        tp = np.sum(pred & labels)
 864        out.append(0.0 if tp == 0 else 2 * tp / (pred.sum() + labels.sum()))
 865    return np.array(out)
 866
 867
 868# ---------------------------------------------------------------------------
 869# 6. Figures (rendered into docs/figures by `make figures`) and the widget's data
 870# ---------------------------------------------------------------------------
 871
 872
 873def viz_data() -> dict:
 874    """The starting arrows for the site's interactive cosine-similarity widget."""
 875    # Each arrow is an angle in degrees (0 points right, 90 up) and a length.
 876    # The widget recomputes dot, cosine and distance with the same formulas as
 877    # dot, cosine and euclidean above; tests check the two agree.
 878    return {
 879        "cosine-similarity": {
 880            "presets": [
 881                {"name": "Same direction, different lengths", "a": {"angle": 30, "length": 1.0}, "b": {"angle": 30, "length": 3.0}},
 882                {"name": "Close in direction", "a": {"angle": 20, "length": 2.0}, "b": {"angle": 50, "length": 1.5}},
 883                {"name": "At right angles", "a": {"angle": 0, "length": 2.0}, "b": {"angle": 90, "length": 2.0}},
 884                {"name": "Opposite", "a": {"angle": 30, "length": 1.5}, "b": {"angle": 210, "length": 1.5}},
 885            ],
 886        }
 887    }
 888
 889
 890def figures() -> dict:
 891    """Plots computed from this module's own functions. Keys match the docstring's image names."""
 892    import matplotlib
 893
 894    matplotlib.use("Agg")
 895    import matplotlib.pyplot as plt
 896
 897    figs = {}
 898
 899    # identity: ||a-b||² = 2 - 2cos for unit vectors
 900    rng = np.random.default_rng(0)
 901    a, b = l2_normalize(rng.standard_normal((400, 8))), l2_normalize(rng.standard_normal((400, 8)))
 902    cos = np.sum(a * b, axis=1)
 903    sq = np.sum((a - b) ** 2, axis=1)
 904    fig, ax = plt.subplots(figsize=(5.5, 4))
 905    xs = np.linspace(-1, 1, 50)
 906    ax.plot(xs, 2 - 2 * xs, color="0.6", lw=2, label="y = 2 - 2x")
 907    ax.scatter(cos, sq, s=10, alpha=0.7, label="random unit-vector pairs")
 908    ax.set(xlabel="cosine similarity", ylabel="squared Euclidean distance", title="On unit vectors, distance is a function of cosine")
 909    ax.legend()
 910    figs["identity"] = fig
 911
 912    # arrows: niche vs popular item against a query
 913    fig, ax = plt.subplots(figsize=(5.5, 3.5))
 914    q = np.array([1.0, 0.0])
 915    niche, popular = np.array([0.99, 0.10]), np.array([0.95, 0.30]) * 3.0
 916    for v, color, label in ((q, "black", "query"), (niche, "C0", "niche item (aligned, short)"), (popular, "C1", "popular item (off-angle, long)")):
 917        ax.annotate("", xy=v, xytext=(0, 0), arrowprops=dict(arrowstyle="->", color=color, lw=2))
 918        ax.plot([], [], color=color, lw=2, label=label)
 919    ax.set(xlim=(-0.1, 3.1), ylim=(-0.1, 1.1), xlabel="dimension 1", ylabel="dimension 2", title="Cosine picks blue (angle); dot product picks orange (length)")
 920    ax.set_aspect("equal")
 921    ax.legend(loc="upper left")
 922    figs["arrows"] = fig
 923
 924    # curse: nearest/farthest ratio vs dimension
 925    rows = curse_of_dimensionality(dims=(1, 2, 5, 10, 20, 50, 100, 200, 500, 1000, 2000))
 926    fig, ax = plt.subplots(figsize=(5.5, 4))
 927    ax.plot([r["dim"] for r in rows], [r["ratio"] for r in rows], marker="o")
 928    ax.set_xscale("log")
 929    ax.set_ylim(0, 1)
 930    ax.set(xlabel="dimensions (log scale)", ylabel="nearest distance / farthest distance", title="Random points become equidistant as dimension grows")
 931    figs["curse"] = fig
 932
 933    # anisotropy: raw vs centered score histograms
 934    left, right, related = anisotropic_embeddings()
 935    raw = pair_cosines(left, right)
 936    lc, rc = mean_center(left, right)
 937    centered = pair_cosines(lc, rc)
 938    fig, axes = plt.subplots(1, 2, figsize=(10, 4))
 939    for ax, s, title in ((axes[0], raw, "raw cosine: everything is 0.7 to 0.9"), (axes[1], centered, "after mean-centering: the gap opens")):
 940        bins = np.linspace(s.min(), s.max(), 40)
 941        ax.hist(s[related], bins=bins, alpha=0.7, label="paraphrase pairs")
 942        ax.hist(s[~related], bins=bins, alpha=0.7, label="unrelated pairs")
 943        ax.set(xlabel="cosine similarity", ylabel="number of pairs", title=title)
 944        ax.legend()
 945    figs["anisotropy"] = fig
 946
 947    # threshold: F1 sweep, guessed vs calibrated
 948    ts = np.linspace(0.72, 0.88, 200)
 949    best = best_threshold(raw, related)
 950    fig, ax = plt.subplots(figsize=(5.5, 4))
 951    ax.plot(ts, f1_curve(raw, related, ts))
 952    ax.axvline(0.8, ls="--", color="0.4", label="guessed 0.80")
 953    ax.axvline(best["threshold"], color="C3", label=f"calibrated {best['threshold']:.3f}")
 954    ax.set(xlabel="threshold on raw cosine", ylabel="F1", title="Threshold choice on an anisotropic model")
 955    ax.legend()
 956    figs["threshold"] = fig
 957
 958    for f in figs.values():
 959        f.tight_layout()
 960    return figs
 961
 962
 963# ---------------------------------------------------------------------------
 964# 7. Narrated walkthrough
 965# ---------------------------------------------------------------------------
 966
 967
 968def demo() -> None:
 969    banner("1. Worked example: a=(1,2,2), b=(2,1,2), c=(2,-1,-2)")
 970    w = worked_example()
 971    table(
 972        ["pair", "dot", "|x||y|", "cosine", "L2 distance"],
 973        [("a, b", w["dot_ab"], 9, w["cos_ab"], w["l2_ab"]), ("a, c", w["dot_ac"], 9, w["cos_ac"], w["l2_ac"])],
 974        floatfmt=".2f",
 975    )
 976    say("Every vector has length 3, so cosine = dot / 9: 8/9 = 0.89 and -4/9 = -0.44.")
 977
 978    banner("2. Normalize and the three metrics agree")
 979    sq, formula = l2_identity_check()
 980    say(f"For two random unit vectors: ||a-b||² = {sq:.6f} and 2 - 2cos = {formula:.6f}.")
 981    agree = rankings_agree_after_normalizing()
 982    table(["check (200 docs, varied lengths)", "same ranking?"], [(k, v) for k, v in agree.items()])
 983    takeaway("On L2-normalized vectors, cosine, dot product and Euclidean distance rank identically.")
 984
 985    banner("3. When they disagree: signal stored in the vector length")
 986    table(["item", "cosine", "dot"], popularity_in_the_norm(), floatfmt=".3f")
 987    say(
 988        """
 989        Cosine prefers the better-aligned niche item; dot product prefers the
 990        long "popular" vector. Neither is wrong. Use whatever the model was
 991        trained with.
 992        """
 993    )
 994
 995    banner("4. Curse of dimensionality (random points)")
 996    table(["dim", "nearest", "farthest", "nearest / farthest"], [(r["dim"], r["nearest"], r["farthest"], r["ratio"]) for r in curse_of_dimensionality()], floatfmt=".3f")
 997    say("As dimension grows the ratio climbs toward 1: every random point is about equally far away.")
 998
 999    banner("5. Anisotropy: every score sits in 0.75-0.85")
1000    left, right, labels = anisotropic_embeddings()
1001    raw = pair_cosines(left, right)
1002    table(
1003        ["pairs", "min cos", "mean cos", "max cos"],
1004        [
1005            ("paraphrases", raw[labels].min(), raw[labels].mean(), raw[labels].max()),
1006            ("unrelated", raw[~labels].min(), raw[~labels].mean(), raw[~labels].max()),
1007        ],
1008        floatfmt=".3f",
1009    )
1010    naive = raw >= 0.8
1011    naive_f1 = 2 * np.sum(naive & labels) / (naive.sum() + labels.sum())
1012    best = best_threshold(raw, labels)
1013    say(
1014        f"""
1015        A "magic" threshold of 0.80 gives F1 = {naive_f1:.3f}. Calibrating on
1016        labeled pairs picks {best['threshold']:.3f} with F1 = {best['f1']:.3f}.
1017        The signal is there, squeezed into a few hundredths.
1018        """
1019    )
1020    lc, rc = mean_center(left, right)
1021    centered = pair_cosines(lc, rc)
1022    table(
1023        ["after mean-centering", "mean cos"],
1024        [("paraphrases", centered[labels].mean()), ("unrelated", centered[~labels].mean())],
1025        floatfmt=".3f",
1026    )
1027    takeaway(
1028        "Raw scores aren't comparable across models. Rank when you can; when you "
1029        "must threshold, calibrate on labeled pairs from your own data, per model."
1030    )
1031
1032
1033if __name__ == "__main__":
1034    demo()
Level 3: the code, function by function.
def l2_normalize(X: numpy.ndarray, axis: int = -1, eps: float = 1e-12) -> numpy.ndarray: on GitHub
662def l2_normalize(X: np.ndarray, axis: int = -1, eps: float = 1e-12) -> np.ndarray:
663    """Scale each vector (row, by default) to unit length.
664
665    `eps` avoids dividing by zero for an all-zero vector (e.g. empty text).
666    """
667    return X / np.maximum(np.linalg.norm(X, axis=axis, keepdims=True), eps)

Scale each vector (row, by default) to unit length.

eps avoids dividing by zero for an all-zero vector (e.g. empty text).

def dot(a: numpy.ndarray, b: numpy.ndarray) -> float: on GitHub
670def dot(a: np.ndarray, b: np.ndarray) -> float:
671    """Sum of element-wise products. Sensitive to both angle and length."""
672    return float(np.dot(a, b))

Sum of element-wise products. Sensitive to both angle and length.

def cosine(a: numpy.ndarray, b: numpy.ndarray) -> float: on GitHub
675def cosine(a: np.ndarray, b: np.ndarray) -> float:
676    """cos(angle) = a·b / (|a||b|). In [-1, 1]; ignores length."""
677    return float(np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b)))

cos(angle) = a·b / (|a||b|). In [-1, 1]; ignores length.

def euclidean(a: numpy.ndarray, b: numpy.ndarray) -> float: on GitHub
680def euclidean(a: np.ndarray, b: np.ndarray) -> float:
681    """Straight-line distance. Smaller means closer."""
682    return float(np.linalg.norm(a - b))

Straight-line distance. Smaller means closer.

def rank(query: numpy.ndarray, docs: numpy.ndarray, metric: str) -> numpy.ndarray: on GitHub
685def rank(query: np.ndarray, docs: np.ndarray, metric: str) -> np.ndarray:
686    """Doc indices from best to worst match under `metric` (cosine | dot | l2)."""
687    if metric == "dot":
688        scores = docs @ query
689    elif metric == "cosine":
690        scores = (docs @ query) / (np.linalg.norm(docs, axis=1) * np.linalg.norm(query))
691    elif metric == "l2":
692        scores = -np.linalg.norm(docs - query, axis=1)  # negate: closer = higher
693    else:
694        raise ValueError(metric)
695    # Stable sort so ties break the same way for every metric.
696    return np.argsort(-scores, kind="stable")

Doc indices from best to worst match under metric (cosine | dot | l2).

A = array([1., 2., 2.])
B = array([2., 1., 2.])
C = array([ 2., -1., -2.])
def worked_example() -> dict[str, float]: on GitHub
708def worked_example() -> dict[str, float]:
709    """a, b, c from the table above. Returns the numbers you should be able to compute by hand."""
710    return {
711        "len_a": float(np.linalg.norm(A)),
712        "dot_ab": dot(A, B),
713        "dot_ac": dot(A, C),
714        "cos_ab": cosine(A, B),
715        "cos_ac": cosine(A, C),
716        "l2_ab": euclidean(A, B),
717        "l2_ac": euclidean(A, C),
718    }

a, b, c from the table above. Returns the numbers you should be able to compute by hand.

def rankings_agree_after_normalizing(n_docs: int = 200, dim: int = 32, seed: int = 0) -> dict[str, bool]: on GitHub
726def rankings_agree_after_normalizing(n_docs: int = 200, dim: int = 32, seed: int = 0) -> dict[str, bool]:
727    """Random docs with wildly different lengths: do the three rankings agree?
728
729    Before normalizing: dot product favors long vectors, so it disagrees.
730    After normalizing: all three rankings are identical, as the algebra says.
731    """
732    rng = np.random.default_rng(seed)
733    docs = rng.standard_normal((n_docs, dim)) * rng.uniform(0.2, 5.0, (n_docs, 1))  # varied lengths
734    q = rng.standard_normal(dim)
735
736    raw = {m: rank(q, docs, m) for m in ("cosine", "dot", "l2")}
737    dn, qn = l2_normalize(docs), l2_normalize(q)
738    norm = {m: rank(qn, dn, m) for m in ("cosine", "dot", "l2")}
739    return {
740        "raw_dot_equals_cosine": bool(np.array_equal(raw["dot"], raw["cosine"])),
741        "raw_l2_equals_cosine": bool(np.array_equal(raw["l2"], raw["cosine"])),
742        "norm_dot_equals_cosine": bool(np.array_equal(norm["dot"], norm["cosine"])),
743        "norm_l2_equals_cosine": bool(np.array_equal(norm["l2"], norm["cosine"])),
744    }

Random docs with wildly different lengths: do the three rankings agree?

Before normalizing: dot product favors long vectors, so it disagrees. After normalizing: all three rankings are identical, as the algebra says.

def l2_identity_check(seed: int = 0) -> tuple[float, float]: on GitHub
747def l2_identity_check(seed: int = 0) -> tuple[float, float]:
748    """Return (||a-b||², 2 - 2cos(a,b)) for two random unit vectors. They're equal."""
749    rng = np.random.default_rng(seed)
750    a, b = l2_normalize(rng.standard_normal(16)), l2_normalize(rng.standard_normal(16))
751    return float(np.sum((a - b) ** 2)), 2 - 2 * cosine(a, b)

Return (||a-b||², 2 - 2cos(a,b)) for two random unit vectors. They're equal.

def popularity_in_the_norm() -> list[tuple[str, float, float]]: on GitHub
754def popularity_in_the_norm() -> list[tuple[str, float, float]]:
755    """When magnitude carries signal, dot product and cosine rank differently.
756
757    Two items point in almost the same direction as the query. The "popular"
758    one has a long vector (a model trained with dot product can learn to do
759    this). Cosine ignores that and prefers the slightly better-aligned niche
760    item; dot product prefers the popular one.
761    """
762    q = np.array([1.0, 0.0])
763    items = {
764        "niche item (well aligned, short)": np.array([0.99, 0.10]) * 1.0,
765        "popular item (a bit off, long)": np.array([0.95, 0.30]) * 3.0,
766    }
767    return [(name, cosine(q, v), dot(q, v)) for name, v in items.items()]

When magnitude carries signal, dot product and cosine rank differently.

Two items point in almost the same direction as the query. The "popular" one has a long vector (a model trained with dot product can learn to do this). Cosine ignores that and prefers the slightly better-aligned niche item; dot product prefers the popular one.

def curse_of_dimensionality( dims=(2, 10, 100, 1000), n_points: int = 1000, seed: int = 0) -> list[dict]: on GitHub
775def curse_of_dimensionality(dims=(2, 10, 100, 1000), n_points: int = 1000, seed: int = 0) -> list[dict]:
776    """For random points, how much nearer is the nearest than the farthest?
777
778    ratio = nearest / farthest distance from a random query. Near 0 means
779    "nearest" is meaningful; near 1 means every point is about equally far.
780    """
781    rng = np.random.default_rng(seed)
782    rows = []
783    for d in dims:
784        X = rng.uniform(0, 1, (n_points, d))
785        q = rng.uniform(0, 1, d)
786        dist = np.linalg.norm(X - q, axis=1)
787        rows.append({"dim": d, "nearest": float(dist.min()), "farthest": float(dist.max()), "ratio": float(dist.min() / dist.max())})
788    return rows

For random points, how much nearer is the nearest than the farthest?

ratio = nearest / farthest distance from a random query. Near 0 means "nearest" is meaningful; near 1 means every point is about equally far.

def anisotropic_embeddings(n_pairs: int = 300, dim: int = 768, seed: int = 0): on GitHub
796def anisotropic_embeddings(n_pairs: int = 300, dim: int = 768, seed: int = 0):
797    """Simulate a model whose vectors all live in a narrow cone.
798
799    Each vector = sqrt(3)·m + 1·n, where m is one shared unit direction and n
800    is a per-text unit direction. For unrelated texts the n parts are nearly
801    orthogonal, so cos ≈ 3 / (3 + 1) = 0.75. For paraphrases we make the n parts
802    correlated (rho in [0.15, 0.45]), so cos ≈ 0.75 + 0.25·rho, about 0.79 to 0.86.
803    Everything lands in the 0.75-0.85 band even though the underlying
804    signal (rho) is clear.
805
806    Returns (left, right, is_paraphrase): left[i] and right[i] form pair i.
807    """
808    rng = np.random.default_rng(seed)
809    m = l2_normalize(rng.standard_normal(dim))
810    alpha, beta = np.sqrt(3.0), 1.0
811
812    def unit(k):
813        return l2_normalize(rng.standard_normal((k, dim)))
814
815    n_left = unit(2 * n_pairs)
816    labels = np.array([True] * n_pairs + [False] * n_pairs)
817    rho = np.where(labels, rng.uniform(0.15, 0.45, 2 * n_pairs), 0.0)
818    # Correlated partner: rho·n + sqrt(1 - rho²)·fresh, which is unit length with cos(n, partner) ≈ rho.
819    n_right = l2_normalize(rho[:, None] * n_left + np.sqrt(1 - rho[:, None] ** 2) * unit(2 * n_pairs))
820    left = alpha * m + beta * n_left
821    right = alpha * m + beta * n_right
822    return left, right, labels

Simulate a model whose vectors all live in a narrow cone.

Each vector = sqrt(3)·m + 1·n, where m is one shared unit direction and n is a per-text unit direction. For unrelated texts the n parts are nearly orthogonal, so cos ≈ 3 / (3 + 1) = 0.75. For paraphrases we make the n parts correlated (rho in [0.15, 0.45]), so cos ≈ 0.75 + 0.25·rho, about 0.79 to 0.86. Everything lands in the 0.75-0.85 band even though the underlying signal (rho) is clear.

Returns (left, right, is_paraphrase): left[i] and right[i] form pair i.

def pair_cosines(left: numpy.ndarray, right: numpy.ndarray) -> numpy.ndarray: on GitHub
825def pair_cosines(left: np.ndarray, right: np.ndarray) -> np.ndarray:
826    return np.sum(l2_normalize(left) * l2_normalize(right), axis=1)
def mean_center(*arrays: numpy.ndarray) -> list[numpy.ndarray]: on GitHub
829def mean_center(*arrays: np.ndarray) -> list[np.ndarray]:
830    """Subtract the mean vector of all given vectors (the shared "cone" direction).
831
832    In production you'd estimate the mean once from a sample of your corpus
833    and store it with the index; queries get the same shift.
834    """
835    mu = np.concatenate(arrays).mean(axis=0)
836    return [a - mu for a in arrays]

Subtract the mean vector of all given vectors (the shared "cone" direction).

In production you'd estimate the mean once from a sample of your corpus and store it with the index; queries get the same shift.

def best_threshold(scores: numpy.ndarray, labels: numpy.ndarray) -> dict[str, float]: on GitHub
839def best_threshold(scores: np.ndarray, labels: np.ndarray) -> dict[str, float]:
840    """Sweep every candidate threshold; return the one with the best F1.
841
842    Predict "similar" when score >= threshold. This is how you pick a dedup
843    or cache threshold for a specific model, from a few hundred labeled pairs,
844    instead of guessing "0.8".
845    """
846    best = {"threshold": 0.0, "f1": -1.0, "precision": 0.0, "recall": 0.0}
847    for t in np.unique(scores):
848        pred = scores >= t
849        tp = np.sum(pred & labels)
850        if tp == 0:
851            continue
852        precision, recall = tp / pred.sum(), tp / labels.sum()
853        f1 = 2 * precision * recall / (precision + recall)
854        if f1 > best["f1"]:
855            best = {"threshold": float(t), "f1": float(f1), "precision": float(precision), "recall": float(recall)}
856    return best

Sweep every candidate threshold; return the one with the best F1.

Predict "similar" when score >= threshold. This is how you pick a dedup or cache threshold for a specific model, from a few hundred labeled pairs, instead of guessing "0.8".

def f1_curve( scores: numpy.ndarray, labels: numpy.ndarray, thresholds: numpy.ndarray) -> numpy.ndarray: on GitHub
859def f1_curve(scores: np.ndarray, labels: np.ndarray, thresholds: np.ndarray) -> np.ndarray:
860    """F1 of the rule "score >= t" for each t in `thresholds` (for plotting)."""
861    out = []
862    for t in thresholds:
863        pred = scores >= t
864        tp = np.sum(pred & labels)
865        out.append(0.0 if tp == 0 else 2 * tp / (pred.sum() + labels.sum()))
866    return np.array(out)

F1 of the rule "score >= t" for each t in thresholds (for plotting).

def viz_data() -> dict: on GitHub
874def viz_data() -> dict:
875    """The starting arrows for the site's interactive cosine-similarity widget."""
876    # Each arrow is an angle in degrees (0 points right, 90 up) and a length.
877    # The widget recomputes dot, cosine and distance with the same formulas as
878    # dot, cosine and euclidean above; tests check the two agree.
879    return {
880        "cosine-similarity": {
881            "presets": [
882                {"name": "Same direction, different lengths", "a": {"angle": 30, "length": 1.0}, "b": {"angle": 30, "length": 3.0}},
883                {"name": "Close in direction", "a": {"angle": 20, "length": 2.0}, "b": {"angle": 50, "length": 1.5}},
884                {"name": "At right angles", "a": {"angle": 0, "length": 2.0}, "b": {"angle": 90, "length": 2.0}},
885                {"name": "Opposite", "a": {"angle": 30, "length": 1.5}, "b": {"angle": 210, "length": 1.5}},
886            ],
887        }
888    }

The starting arrows for the site's interactive cosine-similarity widget.

def figures() -> dict: on GitHub
891def figures() -> dict:
892    """Plots computed from this module's own functions. Keys match the docstring's image names."""
893    import matplotlib
894
895    matplotlib.use("Agg")
896    import matplotlib.pyplot as plt
897
898    figs = {}
899
900    # identity: ||a-b||² = 2 - 2cos for unit vectors
901    rng = np.random.default_rng(0)
902    a, b = l2_normalize(rng.standard_normal((400, 8))), l2_normalize(rng.standard_normal((400, 8)))
903    cos = np.sum(a * b, axis=1)
904    sq = np.sum((a - b) ** 2, axis=1)
905    fig, ax = plt.subplots(figsize=(5.5, 4))
906    xs = np.linspace(-1, 1, 50)
907    ax.plot(xs, 2 - 2 * xs, color="0.6", lw=2, label="y = 2 - 2x")
908    ax.scatter(cos, sq, s=10, alpha=0.7, label="random unit-vector pairs")
909    ax.set(xlabel="cosine similarity", ylabel="squared Euclidean distance", title="On unit vectors, distance is a function of cosine")
910    ax.legend()
911    figs["identity"] = fig
912
913    # arrows: niche vs popular item against a query
914    fig, ax = plt.subplots(figsize=(5.5, 3.5))
915    q = np.array([1.0, 0.0])
916    niche, popular = np.array([0.99, 0.10]), np.array([0.95, 0.30]) * 3.0
917    for v, color, label in ((q, "black", "query"), (niche, "C0", "niche item (aligned, short)"), (popular, "C1", "popular item (off-angle, long)")):
918        ax.annotate("", xy=v, xytext=(0, 0), arrowprops=dict(arrowstyle="->", color=color, lw=2))
919        ax.plot([], [], color=color, lw=2, label=label)
920    ax.set(xlim=(-0.1, 3.1), ylim=(-0.1, 1.1), xlabel="dimension 1", ylabel="dimension 2", title="Cosine picks blue (angle); dot product picks orange (length)")
921    ax.set_aspect("equal")
922    ax.legend(loc="upper left")
923    figs["arrows"] = fig
924
925    # curse: nearest/farthest ratio vs dimension
926    rows = curse_of_dimensionality(dims=(1, 2, 5, 10, 20, 50, 100, 200, 500, 1000, 2000))
927    fig, ax = plt.subplots(figsize=(5.5, 4))
928    ax.plot([r["dim"] for r in rows], [r["ratio"] for r in rows], marker="o")
929    ax.set_xscale("log")
930    ax.set_ylim(0, 1)
931    ax.set(xlabel="dimensions (log scale)", ylabel="nearest distance / farthest distance", title="Random points become equidistant as dimension grows")
932    figs["curse"] = fig
933
934    # anisotropy: raw vs centered score histograms
935    left, right, related = anisotropic_embeddings()
936    raw = pair_cosines(left, right)
937    lc, rc = mean_center(left, right)
938    centered = pair_cosines(lc, rc)
939    fig, axes = plt.subplots(1, 2, figsize=(10, 4))
940    for ax, s, title in ((axes[0], raw, "raw cosine: everything is 0.7 to 0.9"), (axes[1], centered, "after mean-centering: the gap opens")):
941        bins = np.linspace(s.min(), s.max(), 40)
942        ax.hist(s[related], bins=bins, alpha=0.7, label="paraphrase pairs")
943        ax.hist(s[~related], bins=bins, alpha=0.7, label="unrelated pairs")
944        ax.set(xlabel="cosine similarity", ylabel="number of pairs", title=title)
945        ax.legend()
946    figs["anisotropy"] = fig
947
948    # threshold: F1 sweep, guessed vs calibrated
949    ts = np.linspace(0.72, 0.88, 200)
950    best = best_threshold(raw, related)
951    fig, ax = plt.subplots(figsize=(5.5, 4))
952    ax.plot(ts, f1_curve(raw, related, ts))
953    ax.axvline(0.8, ls="--", color="0.4", label="guessed 0.80")
954    ax.axvline(best["threshold"], color="C3", label=f"calibrated {best['threshold']:.3f}")
955    ax.set(xlabel="threshold on raw cosine", ylabel="F1", title="Threshold choice on an anisotropic model")
956    ax.legend()
957    figs["threshold"] = fig
958
959    for f in figs.values():
960        f.tight_layout()
961    return figs

Plots computed from this module's own functions. Keys match the docstring's image names.

def demo() -> None: on GitHub
 969def demo() -> None:
 970    banner("1. Worked example: a=(1,2,2), b=(2,1,2), c=(2,-1,-2)")
 971    w = worked_example()
 972    table(
 973        ["pair", "dot", "|x||y|", "cosine", "L2 distance"],
 974        [("a, b", w["dot_ab"], 9, w["cos_ab"], w["l2_ab"]), ("a, c", w["dot_ac"], 9, w["cos_ac"], w["l2_ac"])],
 975        floatfmt=".2f",
 976    )
 977    say("Every vector has length 3, so cosine = dot / 9: 8/9 = 0.89 and -4/9 = -0.44.")
 978
 979    banner("2. Normalize and the three metrics agree")
 980    sq, formula = l2_identity_check()
 981    say(f"For two random unit vectors: ||a-b||² = {sq:.6f} and 2 - 2cos = {formula:.6f}.")
 982    agree = rankings_agree_after_normalizing()
 983    table(["check (200 docs, varied lengths)", "same ranking?"], [(k, v) for k, v in agree.items()])
 984    takeaway("On L2-normalized vectors, cosine, dot product and Euclidean distance rank identically.")
 985
 986    banner("3. When they disagree: signal stored in the vector length")
 987    table(["item", "cosine", "dot"], popularity_in_the_norm(), floatfmt=".3f")
 988    say(
 989        """
 990        Cosine prefers the better-aligned niche item; dot product prefers the
 991        long "popular" vector. Neither is wrong. Use whatever the model was
 992        trained with.
 993        """
 994    )
 995
 996    banner("4. Curse of dimensionality (random points)")
 997    table(["dim", "nearest", "farthest", "nearest / farthest"], [(r["dim"], r["nearest"], r["farthest"], r["ratio"]) for r in curse_of_dimensionality()], floatfmt=".3f")
 998    say("As dimension grows the ratio climbs toward 1: every random point is about equally far away.")
 999
1000    banner("5. Anisotropy: every score sits in 0.75-0.85")
1001    left, right, labels = anisotropic_embeddings()
1002    raw = pair_cosines(left, right)
1003    table(
1004        ["pairs", "min cos", "mean cos", "max cos"],
1005        [
1006            ("paraphrases", raw[labels].min(), raw[labels].mean(), raw[labels].max()),
1007            ("unrelated", raw[~labels].min(), raw[~labels].mean(), raw[~labels].max()),
1008        ],
1009        floatfmt=".3f",
1010    )
1011    naive = raw >= 0.8
1012    naive_f1 = 2 * np.sum(naive & labels) / (naive.sum() + labels.sum())
1013    best = best_threshold(raw, labels)
1014    say(
1015        f"""
1016        A "magic" threshold of 0.80 gives F1 = {naive_f1:.3f}. Calibrating on
1017        labeled pairs picks {best['threshold']:.3f} with F1 = {best['f1']:.3f}.
1018        The signal is there, squeezed into a few hundredths.
1019        """
1020    )
1021    lc, rc = mean_center(left, right)
1022    centered = pair_cosines(lc, rc)
1023    table(
1024        ["after mean-centering", "mean cos"],
1025        [("paraphrases", centered[labels].mean()), ("unrelated", centered[~labels].mean())],
1026        floatfmt=".3f",
1027    )
1028    takeaway(
1029        "Raw scores aren't comparable across models. Rank when you can; when you "
1030        "must threshold, calibrate on labeled pairs from your own data, per model."
1031    )