Scaling laws and Chinchilla, annotated
How to read this page
- Any dotted word explains itself on hover, focus or tap, and so does every symbol in every equation.
- The two papers disagree, so read Part I as “the first answer” and Part II as “the corrected answer”. The budget-splitting tool at the end puts their advice side by side.
Every idea climbs the ladder: everyday picture, tiny example, diagram or chart, the math, why it matters today. The training pipeline these laws plan is described in the training stages lesson.
The question both papers answer
Everyday picture
You have a fixed budget for building a library of knowledge in someone's head. You can hire a bigger brain (more parameters, N) or buy more books for it to read (more training tokens, D). A bigger brain costs more per book read, so with a fixed budget, every extra neuron means fewer books. What is the best split?
Tiny example
Training cost is roughly 6 × N × D operations (six per parameter per token). So a budget of 6 × 1021 operations buys a 1-billion-parameter model reading 1 trillion tokens, or a 10-billion-parameter model reading 100 billion tokens, or a 100-billion-parameter model reading only 10 billion. Same bill, very different models. Which one ends up smartest?
In words: “training compute is about six operations for every parameter, for every token read.”
With the numbers: 6 × 109 × 1012 = 6 × 1021, and 6 × 1010 × 1011 is the same 6 × 1021.
In Python:
# 1 billion parameters, 1 trillion tokens
N, D = 10**9, 10**12
# C ≈ 6 N D
C = 6 * N * D
print(f"{C:.0e}") # → 6e+21
# 10× the model, a tenth of the data: same bill
6 * 10**10 * 10**11 == C # → True
Why it matters
Training runs cost millions of dollars and happen once. Getting this split wrong wastes most of the budget, and the two papers gave very different answers: Kaplan's said “mostly spend it on size”, Chinchilla's said “split it evenly”.
Part I: Kaplan et al. (2020), Scaling Laws for Neural Language Models
Abstract · original
“The loss scales as a power-law with model size, dataset size, and the amount of compute used for training, with some trends spanning more than seven orders of magnitude.”Kaplan et al. (2020), Abstract
Everyday picture
Children's height follows growth charts: predictable curves, not surprises. This paper found that a language model's loss (how wrong its next-word guesses are, measured with cross-entropy) follows equally predictable curves as you add parameters, data or compute. That makes it possible to forecast how good a giant model will be from a handful of small experiments.
Why it matters today
This is the paper that made “just scale it up” a plan instead of a hope. It directly shaped how large GPT-3 was made (the GPT-3 companion shares authors with it).
1.1 Scale, not shape · original
“Within reasonable limits, performance depends very weakly on other architectural hyperparameters such as depth vs. width.”Kaplan et al. (2020), §1.1
Everyday picture
For a cake, the amount of flour matters far more than the shape of the tin. The paper finds that for a Transformer, how many parameters it has matters far more than how they are arranged: many thin layers or fewer wide ones, many attention heads or few, all land close to the same loss for the same total size.
Its other headline findings, each explained below: smooth power laws in N, D and compute; overfitting is predictable from the ratio of model size to data; and bigger models learn more per example, reaching the same loss after reading fewer tokens.
1.2 Three power laws · original
Everyday picture
A power law is a rule of the form “every time you double X, Y shrinks by the same percentage”. Double the model and the loss drops by about 5%, double it again and it drops another 5%. Each doubling helps a bit less in absolute terms, but it never stops helping, at least over the range they tested.
Tiny example
With the paper's fitted constants, a 100-million-parameter model (trained to convergence on plenty of data) has a predicted loss of 2.83; 1 billion parameters gives 2.38; 10 billion gives 1.99. Each 10× in size multiplies the loss by about 0.84.
In words: “when only one resource is holding the model back, loss is a fixed constant divided by that resource, raised to a small power.”
With the numbers: L(109) = (8.8 × 1013 / 109)0.076 = 88,0000.076 = 2.38 nats. Doubling N multiplies the loss by 2−0.076 = 0.95. For data: L(1010 tokens) = 5,4000.095 = 2.26.
In Python:
# the fitted constants of L(N)
N_c, alpha_N = 8.8e13, 0.076
N = 1e9
round(N_c / N) # → 88000
# L(N) = (N_c / N)^α_N, in nats
round((N_c / N) ** alpha_N, 2) # → 2.38
# doubling N multiplies the loss by this
round(2 ** -alpha_N, 2) # → 0.95
# the fitted constants of L(D)
D_c, alpha_D = 5.4e13, 0.095
# L(D) at D = 10^10 tokens
round((D_c / 1e10) ** alpha_D, 2) # → 2.26
Reading it: the x-axis is model size on a log scale, from a million to a trillion parameters; the y-axis is the predicted loss from the L(N) law, computed live from the paper's constants. Drag the slider to move the marker and read the loss. Notice there is no cliff and no plateau anywhere in the range: each tenfold increase buys a similar-looking step down. The authors stress that the curve must flatten eventually, since text is never perfectly predictable, but they saw no sign of it at the scales they tested.
Why it matters today
The shape of these curves is why labs could justify bigger and bigger training runs: the next model's quality was forecastable. The exact constants depend on the dataset and tokenizer, but the power-law shape has held up widely.
1.3 and 2.1 Counting N and C · original
Everyday picture
To compare models fairly you need a common currency. The paper counts only the parameters doing the real work (excluding the embedding table) and measures compute in petaflop/s-days.
In words: “a standard Transformer block holds about twelve times its width squared in parameters; training compute is six operations per parameter per token, and the tokens are the batch size times the number of steps.”
With the numbers: a 24-layer model of width 1,024 has about 12 × 24 × 1,0242 ≈ 302 million non-embedding parameters. Training it for 100,000 steps at 500,000 tokens per batch (50 billion tokens) costs about 6 × 3.02 × 108 × 5 × 1010 ≈ 9.1 × 1019 operations, about 1 petaflop/s-day.
In Python:
n_layer, d_model = 24, 1024
# N ≈ 12 n_layer d_model²
N = 12 * n_layer * d_model**2
# millions of parameters
round(N / 1e6) # → 302
# tokens per batch, training steps
B, S = 500_000, 100_000
# C ≈ 6 N B S
C = 6 * N * B * S
print(f"{C:.1e}") # → 9.1e+19
# petaflop/s-days: 10^15 operations a second for a day
round(C / 8.64e19, 1) # → 1.0
The 12 comes from four d × d attention matrices (query, key, value, output) plus a feed-forward network that widens to 4d and back (8d2). The 6 comes from about 2 operations per parameter per token going forward and about 4 going backward. The transformer lesson counts both exactly.
4 Overfitting: L(N, D) · original
Everyday picture
A big brain with too few books starts memorizing the books instead of learning the subject: overfitting. The paper finds a single formula for both limits at once, and a rule for how much more reading a bigger brain needs.
In words: “add a term for being too small and a term for having too little data, then apply the same small power to the total; whichever term is larger dominates the loss.”
With the numbers: the overfitting penalty depends on N0.74/D, so making the model 8 times larger needs 80.74 ≈ 4.7 times more data to keep the penalty the same: the paper's “8× the model, about 5× the data”.
In Python:
# α_N / α_D from the single-resource laws
round(0.076 / 0.095, 2) # → 0.8
# the exponent the paper measured for N^0.74 / D
ratio = 0.74
# 8× the model needs this many times the data
round(8 ** ratio, 1) # → 4.7
Why it matters
This rule says data needs grow more slowly than model size. That conclusion is exactly what Part II overturns.
6 Spending a compute budget · original
“Larger models are significantly more sample-efficient, such that optimally compute-efficient training involves training very large models on a relatively modest amount of data and stopping significantly before convergence.”Kaplan et al. (2020), Abstract
Everyday picture
Kaplan's advice: when the budget grows, spend almost all of it on a bigger brain, add a little more reading, and don't bother finishing the book; stop training long before the model has squeezed everything out of its data.
In words: “as compute grows, the best model size grows almost as fast as compute itself; the batch grows a bit; the number of training steps barely grows at all.”
With the numbers: 10× more compute means 100.73 ≈ 5.4× bigger model, and since C ≈ 6ND the data only grows by 10 / 5.4 ≈ 1.9×. The paper's own fitted rule is Nopt ≈ 1.3 × 109 × C0.73 parameters, with C in petaflop/s-days.
In Python:
# 10× more compute
growth = 10
# N_opt ∝ C_min^0.73
N_growth = growth ** 0.73
round(N_growth, 1) # → 5.4
# C ≈ 6 N D, so the data grows by what is left
round(growth / N_growth, 1) # → 1.9
§6.3 then notices a paradox: extrapolated far enough, this recipe predicts a loss lower than the data-limited law allows, so the laws must break somewhere. The authors guess that point is where Transformers hit their ceiling.
Why it matters today
This advice explains the models of 2020 and 2021: GPT-3 (175B parameters on 300B tokens), Gopher (280B on 300B), Megatron-Turing NLG (530B on about 300B). All huge, all trained on roughly the same amount of text. Part II shows that was the wrong way round.
Part II: Hoffmann et al. (2022), Training Compute-Optimal Large Language Models
Abstract · original
“We find that current large language models are significantly undertrained, a consequence of the recent focus on scaling language models whilst keeping the amount of training data constant.”Hoffmann et al. (2022), Abstract
Everyday picture
The field had been building ever bigger brains and giving each the same stack of books. DeepMind trained over 400 models, from 70 million to 16 billion parameters on 5 to 500 billion tokens, and found the stack of books should grow as fast as the brain. To prove it, they trained Chinchilla: 4 times smaller than their own Gopher, fed 4 times more data, same compute bill. Chinchilla won almost everywhere.
2 Where Kaplan went wrong · original
Everyday picture
Imagine judging runners by timing everyone at the 5 km mark of a race, even the ones who paced themselves for 10 km. The 10 km runners look slow at 5 km, not because they are slow but because they were pacing for a longer race. Kaplan et al. used one fixed learning-rate schedule for all runs and read off losses part-way through training, so shorter training looked worse than it really is.
Hoffmann et al. instead set each run's cosine decay to finish exactly when that run's data ran out. With every run properly paced, training on more data turned out to be worth much more than Kaplan's fits suggested.
Why it matters
It is a lesson far beyond this paper: an experimental design choice that looks like a detail (the learning-rate schedule) quietly changed the conclusion that guided billions of dollars of training. The optimizers lesson shows what warmup and cosine decay do.
3 Three approaches, one answer · original
Everyday picture
To be sure of a surprising answer, measure it three different ways. Approach 1 trains models of fixed sizes for four different lengths and, for every compute budget, keeps whichever run reached the lowest loss. Approach 2 fixes nine compute budgets and, for each, trains a range of model sizes (so bigger models read fewer tokens) to find the sweet spot: an “IsoFLOP” profile. Approach 3 fits one formula to every run and solves it with algebra.
Hover or tap to read the predicted loss at each model size.
Reading it: each curve is one fixed compute budget, drawn from the paper's fitted loss formula (§3.3) with D = C / (6N). Moving right along a curve means a bigger model that, to stay on budget, reads fewer tokens. Every curve is a valley: too small a model can't learn enough, too big a model doesn't get to read enough. The bottom of each valley is the compute-optimal size, and as the budget grows (lower curves) the bottom moves right, but only about 3× per 10× of compute. These are predictions from the fit, not the paper's measured runs.
3.3 The fitted loss · original
Everyday picture
Split the model's error into three bills. One is unavoidable: language is partly unpredictable, and even a perfect model pays it. One is the penalty for a small brain. One is the penalty for too little reading. Shrink the second by growing N, shrink the third by growing D.
In words: “predicted loss is an unbeatable floor, plus a penalty that shrinks as the model grows, plus a penalty that shrinks as the data grows.”
With the numbers: for Chinchilla (70 billion parameters, 1.4 trillion tokens): 406.4 / (7 × 1010)0.34 ≈ 0.083 and 410.7 / (1.4 × 1012)0.28 ≈ 0.163, so L̂ ≈ 1.69 + 0.083 + 0.163 = 1.94. For Gopher (280 billion parameters, 300 billion tokens): 1.69 + 0.052 + 0.251 = 1.99. Gopher's model penalty is smaller, but its data penalty is much bigger.
In Python:
E, A, B, alpha, beta = 1.69, 406.4, 410.7, 0.34, 0.28
def L_hat(N, D):
# L̂ = E + A / N^α + B / D^β
return E + A / N**alpha + B / D**beta
# Chinchilla's two penalties
round(A / 7e10**alpha, 3), round(B / 1.4e12**beta, 3) # → (0.083, 0.163)
# Chinchilla: 70B parameters, 1.4T tokens
round(L_hat(7e10, 1.4e12), 2) # → 1.94
# Gopher's two penalties
round(A / 2.8e11**alpha, 3), round(B / 3e11**beta, 3) # → (0.052, 0.251)
# Gopher: 280B parameters, 300B tokens
round(L_hat(2.8e11, 3e11), 2) # → 1.99
Minimizing this formula under the budget C = 6ND has a closed-form answer:
In words: “the best model size and the best data size are both powers of the budget; the two powers add up to 1, and they are set by how fast each penalty shrinks.”
With the numbers: a = 0.28 / 0.62 ≈ 0.45 and b = 0.34 / 0.62 ≈ 0.55: close to half and half, and close to the 0.46 and 0.54 the paper reports for this approach. Compare Kaplan's 0.73 and 0.27.
In Python:
alpha, beta = 0.34, 0.28
# the power on the budget for N_opt
a = beta / (alpha + beta)
# the power on the budget for D_opt
b = alpha / (alpha + beta)
round(a, 2), round(b, 2) # → (0.45, 0.55)
# the two powers always add up to 1
round(a + b, 6) # → 1.0
3.4 The answer: scale both equally · original
“…for compute-optimal training, the model size and the number of training tokens should be scaled equally: for every doubling of model size the number of training tokens should also be doubled.”Hoffmann et al. (2022), Abstract
| Method | a, where Nopt ∝ Ca | b, where Dopt ∝ Cb |
|---|---|---|
| Approach 1: best of fixed-size runs | 0.50 | 0.50 |
| Approach 2: IsoFLOP profiles | 0.49 | 0.51 |
| Approach 3: fitted loss | 0.46 | 0.54 |
| Kaplan et al. (2020) | 0.73 | 0.27 |
| Parameters | Training tokens | Tokens per parameter |
|---|---|---|
| 400 million | 8.0 billion | 20 |
| 1 billion | 20.2 billion | 20 |
| 10 billion | 205.1 billion | 21 |
| 67 billion | 1.5 trillion | 22 |
| 175 billion | 3.7 trillion | 21 |
| 280 billion | 5.9 trillion | 21 |
The famous rule of thumb: the tokens-per-parameter column hovers around 20. The paper does not state it as a rule, but it falls straight out of this table, and it is how people usually quote Chinchilla. By that yardstick GPT-3, at 300 billion tokens, should have read about 3.5 trillion.
4 Chinchilla against Gopher · original
Everyday picture
The proof of a prediction is a head-to-head test. Same compute budget, same data source, same architecture family: one model built the old way, one built the new way.
| Gopher | Chinchilla | |
|---|---|---|
| Parameters | 280 billion | 70 billion |
| Training tokens | 300 billion | 1.4 trillion |
| Layers · width | 80 · 16,384 | 80 · 8,192 |
| Training compute | the same budget (about 5.76 × 1023 FLOPs) | |
| MMLU (57 exam subjects, 5-shot) | 60.0% | 67.5% |
| LAMBADA (zero-shot) | 74.5% | 77.4% |
Chinchilla also beat GPT-3 (175B), Jurassic-1 (178B) and Megatron-Turing NLG (530B) on a large range of tasks. (The abstract reports 67.5% on MMLU; the detailed table rounds it to 67.6%.) And because it is 4 times smaller, it is also about 4 times cheaper to run for every question afterwards.
Why it matters today
Being smaller for the same quality is a double win: the training bill is the same, but every one of the millions of later requests is cheaper, a point developed in the inference lesson.
5 Discussion · original
“Speculatively, we expect that scaling to larger and larger datasets is only beneficial when the data is high-quality.”Hoffmann et al. (2022), §5
The authors flag their limits honestly: only two truly large runs (Chinchilla and Gopher), an assumption that the best-size curve is a straight power law (with hints of curvature at the top that could mean even smaller models are optimal), and all runs using less than one pass over the data. Their conclusion shifts attention to the other half of the budget: if data must grow as fast as models, collecting enough good text becomes the bottleneck.
Try it: split a compute budget
| Advice | Parameters N | Tokens D | Tokens per parameter | Predicted loss |
|---|
Reading it: drag the budget (in FLOPs, on a log scale). The first row applies Kaplan et al.'s fitted rule, N ≈ 1.3 × 109 × C0.73 with C in petaflop/s-days, and gives the rest of the budget to data. The second applies Chinchilla's roughly 20 tokens per parameter, which with C = 6ND means N = √(C / 120). The last column scores both with Chinchilla's fitted loss formula (§3.3), so it tells you which split that formula prefers. The starting position is Gopher's budget: Kaplan's rule asks for a model of about 800 billion parameters reading only about 120 billion tokens, while Chinchilla's asks for about 69 billion parameters on 1.4 trillion tokens, and the fit predicts a clearly lower loss for the second. Kaplan's rule is applied naively here (its compute and parameter definitions differ slightly), so treat its row as indicative.
What happened next
| Development | Why |
|---|---|
| “Chinchilla-optimal” became the standard starting point for planning training runs | Balanced N and D gives the best model for a fixed training bill |
| Many popular models are now trained well past 20 tokens per parameter | Chinchilla optimizes the training bill only. If a model will answer billions of questions, a smaller model trained much longer costs more to train but far less to run: see the inference lesson |
| Data quality and deduplication became central research topics | If D must grow with N, the supply of good text becomes the bottleneck |
| Scaling laws are used for more than size | Teams fit small-scale curves for architectures, data mixes and hyperparameters, then extrapolate before committing to a big run |
Glossary
Every term with hover guidance on this page, in one place.