Language Models are Few-Shot Learners (GPT-3), annotated
How to read this page
- Any dotted word explains itself on hover, focus or tap. So does every symbol in every equation.
- The prompt builder in §2 lets you assemble zero-, one- and few-shot prompts yourself. It is the single most useful idea in the paper, and it is how everyone still talks to language models.
Each idea climbs the ladder: everyday picture, tiny example, diagram, the math, why it matters today. The code that walks a prompt through a model end to end is in the big picture lesson.
Abstract · original
“Here we show that scaling up language models greatly improves task-agnostic, few-shot performance, sometimes even reaching competitiveness with prior state-of-the-art fine-tuning approaches.”Brown et al. (2020), Abstract
Everyday picture
Before GPT-3, teaching a model a new job was like hiring a temp and sending them on a week-long training course for every new task: thousands of labelled examples and a round of fine-tuning. GPT-3 is more like a very well-read new hire who, shown two or three filled-in forms, can fill in the next one straight away. No retraining, no change to a single weight: just examples in the message you send.
What the paper claims
- GPT-3 has 175 billion parameters, 10 times more than any earlier non-sparse language model.
- With no fine-tuning at all, only examples in the prompt, it is strong on translation, question answering and fill-in-the-blank tasks, and sometimes matches fine-tuned systems.
- It can do some on-the-fly tasks it was never trained for, such as 3-digit arithmetic and unscrambling words.
- People could barely tell its short news articles from human-written ones.
- It still fails on some tasks, and training on a web-scale corpus raises real methodological problems.
Why it matters today
This is the paper that turned “prompting” into the main way people use AI. Every time you paste an example into a chat and say “do it like this”, you are doing what this paper named few-shot learning.
1 Introduction · original
“We use the term “in-context learning” to describe the inner loop of this process, which occurs within the forward-pass upon each sequence.”Brown et al. (2020), Figure 1.1 caption
Everyday picture
There are two kinds of learning in this paper, running on two very different clocks. The outer loop is pre-training: months of reading the internet, slowly adjusting billions of weights. The inner loop happens in a fraction of a second, inside a single prompt: the model notices the pattern in the examples you gave (“English word, arrow, French word”) and continues it. Nothing is stored; close the chat and the “learning” is gone. That inner loop is in-context learning.
Tiny example
Somewhere in its training data the model has seen pages that list word translations, pages of worked sums, and pages of quiz questions followed by answers. So a prompt that looks like the start of such a page is naturally continued in the same style. Show it “cat → chat, dog → chien, cheese →” and the most likely next token is “fromage”, because that is what such a list would say next.
Why it matters
The introduction's argument is practical: collecting thousands of labelled examples for every task is expensive, fine-tuned models can overfit narrow datasets, and people learn most tasks from a sentence of instruction and a couple of examples. If a model could too, it would be vastly more useful. The bet is that scale makes this possible.
2 Approach · original
Zero, one and few shots
“…the model is given a few demonstrations of the task at inference time as conditioning […], but no weight updates are allowed.”Brown et al. (2020), §2, defining few-shot
Everyday picture
Imagine asking a colleague to translate a word. Zero-shot: you just describe the job (“translate English to French: cheese”). One-shot: you describe it and show one worked example. Few-shot: you show several. Fine-tuning is the old way: send them on a course and permanently change what they know.
Hover or tap a block. Compare the left column with the other three.
Reading it: the left column is fine-tuning. Each example triggers a gradient step, thousands of them, and the result is a new copy of the model for that one task. The three columns on the right all feed a single frozen model (the dashed box): the only difference between them is how much text is placed in front of the question. Zero-shot has only a description; one-shot adds one worked example; few-shot adds several (the paper uses 10 to 100, as many as fit). The model simply continues the text, and its continuation is the answer.
Try it: build a prompt
Reading it: the grey box is exactly the text the model would receive. Drag K from 0 up: at K = 0 it is a zero-shot prompt, at 1 one-shot, above that few-shot. The last line always ends where the model must continue, and the expected continuation is shown in the readout. The token count is estimated at about 4 characters per token and compared with GPT-3's 2,048-token context window, which is what capped K in the paper. The example words and numbers are new ones chosen for this page.
The math: what the model is doing
Underneath, GPT-3 does one thing: it gives a probability for the next token given everything before it. The probability of a whole piece of text is the product of those next-token probabilities.
In words: “the chance of a whole text is the chance of its first token, times the chance of the second given the first, and so on to the end.”
With the numbers: if the model gives “the” 0.1, then “cat” 0.3 given “the”, then “sat” 0.5 given “the cat”, the three-token text has probability 0.1 × 0.3 × 0.5 = 0.015. A few-shot prompt is simply x<t: the demonstrations change which next token looks likely.
In Python:
# P(the), P(cat | the), P(sat | the cat)
P_next = [0.1, 0.3, 0.5]
P_x = 1
# Π over t
for p in P_next:
P_x *= p
round(P_x, 3) # → 0.015
Why it matters today
Prompting is still how most applications steer a model: a system prompt describes the task and a few examples pin down the format. Modern chat models are additionally trained to follow instructions (see InstructGPT), so zero-shot works far better than it did for GPT-3, but few-shot examples remain the most reliable way to show an exact output format.
2.1 Model and architectures · original
Everyday picture
GPT-3 is not a new design. It is GPT-2 (a decoder-only Transformer) scaled up, trained in eight sizes from 125 million to 175 billion parameters so the authors could watch how abilities change with size. The one architectural tweak: layers alternate between full attention and a cheaper, locally banded sparse attention.
Tiny example: counting the parameters
Almost all of a Transformer's weights live in the attention and feed-forward matrices, and a standard block has about 12·dmodel2 of them. Pick a size to check the paper's parameter count against this rule of thumb and to see the training compute.
Reading it: each position of the slider is one of the paper's eight models (layer count, width and heads restated from Table 2.1). The first computed row applies the 12 × layers × width² rule; the second adds the embedding table (50,257 tokens × width), and together they land close to the paper's stated size. The last rows estimate training compute with the 6 × parameters × tokens rule, since every model saw 300 billion tokens. For the 175B model this gives 3.15 × 1023 FLOPs, matching the paper's own figure of 3.14 × 1023.
In words: “a Transformer's size is about twelve times the number of layers times the width squared; training costs about six arithmetic operations per parameter for every token of data.”
With the numbers: 12 × 96 × 12,2882 ≈ 174 billion parameters. Then 6 × 175 × 109 × 300 × 109 = 3.15 × 1023 operations, about 3,646 petaflop/s-days (the paper's own table, from its exact counts, lists 3,640): a machine doing 1015 operations per second would need about 10 years.
In Python:
n_layer, d_model = 96, 12_288
N = 12 * n_layer * d_model ** 2
# billions of parameters
round(N / 1e9) # → 174
# parameters, training tokens
N, D = 175e9, 300e9
C = 6 * N * D
f"{C:.3g}" # → '3.15e+23'
# petaflop/s-days
round(C / (1e15 * 24 * 3600)) # → 3646
# years at 10^15 operations per second
round(C / 1e15 / (365 * 24 * 3600)) # → 10
Why it matters today
These two rules of thumb (12·L·d² and 6·N·D) are how practitioners estimate model size and training cost on the back of an envelope. The transformer lesson counts parameters exactly and checks the rule; the scaling laws companion asks how to split a compute budget between N and D.
2.2 Training dataset · original
“This essentially accepts a small amount of overfitting in exchange for higher quality training data.”Brown et al. (2020), §2.2
Everyday picture
A reading list for a model. Most of the pages come from a filtered web crawl, but the authors deliberately hand the model the better books more often than their size alone would suggest, like a student who skims the newspaper once but rereads the good textbooks three times.
Reading it: each row is one source. The bar is its share of what the model actually read during training; the right-hand column shows how big the source is and how many times it was read in full over the 300 billion training tokens. Common Crawl makes up 60% of training but was not even read once in full (0.44 epochs), while Wikipedia, a small 3 billion tokens, was read 3.4 times. Numbers restated from Brown et al. (2020), Table 2.2.
Why it matters today
Data curation, weighting and deduplication are still among the most important and least visible parts of training a language model. The paper also removed near-duplicate documents so the held-out set stayed honest, a theme that returns in §4.
2.3 Training process · original
Larger models used larger batches and smaller learning rates: the 175B model used 3.2 million tokens per batch and a peak learning rate of 0.6 × 10−4, against 0.5 million tokens and 6.0 × 10−4 for the smallest. To fit in memory, each model was split across many GPUs, both within each matrix multiply and across layers, on a Microsoft cluster of V100 GPUs.
Everyday picture: a huge model takes careful, small steps (low learning rate) but averages each step over far more examples (big batch), like a large ship that turns slowly but steadily. The optimizers lesson shows why step size matters.
2.4 Evaluation · original
Everyday picture
For a multiple-choice question, you don't ask the model to write an answer; you ask how likely it finds each option as a continuation, and pick the most likely. One trap: a long answer is a product of many probabilities below 1, so it always looks less likely than a short one. The fix is to compare the average per token.
In words: “for each candidate answer, add up the log-probabilities of its tokens and divide by how many tokens it has; pick the candidate with the highest average.”
With the numbers: answer A is two tokens with probabilities 0.5 and 0.4: ln 0.5 + ln 0.4 = −0.693 − 0.916 = −1.609, average −0.805. Answer B is one token with probability 0.3: ln 0.3 = −1.204. By total, B would win (−1.204 > −1.609); by average, A wins (−0.805 > −1.204). Length normalization stops long answers being penalized just for being long.
In Python:
import math
# c: the candidate's token probabilities
def score(c):
# (1/|c|) Σ_t log P(c_t | ...)
return sum(math.log(p) for p in c) / len(c)
# totals
round(math.log(0.5) + math.log(0.4), 3), round(math.log(0.3), 3) # → (-1.609, -1.204)
# averages
round(score([0.5, 0.4]), 3), round(score([0.3]), 3) # → (-0.805, -1.204)
The paper uses per-token averaging for most multiple-choice tasks, draws the K demonstrations at random from each task's training set, and uses K between 10 and 100 depending on how many fit in 2,048 tokens.
Why it matters today
Scoring options by likelihood is still how many benchmarks evaluate base models. The loss functions lesson explains log-probabilities and why they are summed rather than multiplied.
3 Results · original
The results span more than two dozen benchmarks. Four stories stand out.
3.1 LAMBADA and completion tasks · original
Everyday picture
LAMBADA is a fill-in-the-last-word test where you must read a whole paragraph to know the word: a test of long-range understanding. Few-shot examples help in a surprising way: they show the model that the answer must be exactly one word, which a plain continuation might not respect.
| Setting | Accuracy |
|---|---|
| Previous best (a 17B-parameter model) | 68.0% |
| GPT-3 zero-shot | 76.2% |
| GPT-3 one-shot | 72.5% |
| GPT-3 few-shot | 86.4% |
An 18-point jump over the previous state of the art, from examples in a prompt, with no training on the task.
3.2 Closed-book question answering · original
Everyday picture
“Open book” systems search documents for the answer first (that is retrieval-augmented generation); “closed book” means answering from memory alone. On TriviaQA, GPT-3 few-shot scored 71.2% closed-book, above the 68.0% of a fine-tuned open-book RAG system at the time (Table 3.3). On Natural Questions, which asks about more obscure facts, it scored 29.9% against RAG's 44.5%.
Why it matters today
This split still holds: large models remember popular facts well and obscure or recent facts badly, which is exactly why production systems retrieve documents instead of trusting memory. See the RAG lesson.
3.9.1 Arithmetic · original
Everyday picture
Nobody taught GPT-3 to add. The test asks questions in plain English (“Q: What is 23 plus 58? A:”) to see whether a model that only predicts text has picked up arithmetic along the way.
Reading it: each bar is GPT-3 175B's accuracy on one kind of problem, restated from Brown et al. (2020), Table 3.9. Toggle between settings. Two-digit addition is essentially solved few-shot (100%), three-digit addition is right about 80% of the time, and accuracy falls steeply with more digits (about 9% at five digits) and for multiplication (29%). Examples in the prompt help a great deal: two-digit addition goes from 77% zero-shot to 100% few-shot. The pattern looks like partial skill rather than a memorized table, but not like a reliable calculator either.
Why it matters today
Arithmetic is where the tokenizer quirks from the tokenization lesson bite: numbers split into irregular chunks. Modern systems either train heavily on maths, write out intermediate steps, or simply call a calculator tool.
3.9.4 News article generation · original
Everyday picture
People were shown short news articles (about 200 words), some real and some written by models of each size, and asked which were machine-written. If they could tell perfectly they would score 100%; pure guessing scores 50%.
Hover or tap the chart to read each model's result.
Reading it: the x-axis is model size on a log scale, from 125 million to 175 billion parameters; the y-axis is how often people correctly spotted the machine-written article (restated from Brown et al., 2020, Table 3.11). The flat line is chance. As models grow, detection falls towards the line, reaching 52% for GPT-3 175B: barely better than a coin toss. A deliberately bad control model was spotted 86% of the time.
Why it matters today
This result drove the paper's long discussion of misuse, and it is why detecting machine-written text remains an unsolved problem rather than a solved one.
4 Measuring and preventing memorization of benchmarks · original
Everyday picture
If a student has seen the exam paper before the exam, a high score proves little. A model trained on a large slice of the internet may well have read the test questions. That is benchmark contamination.
The authors tried to remove every overlap with their benchmarks from the training data, but a bug left some in, and the model was too expensive to retrain. So they built “clean” versions of each benchmark, dropping every question that shared a 13-word sequence with the training data, and compared scores. For most benchmarks the score barely moved, though a few were flagged as possibly inflated.
Why it matters today
Contamination is now a standing concern for every published benchmark score. The regularization lesson covers data leakage, its smaller cousin.
5 Limitations · original
“GPT-3 samples still sometimes repeat themselves semantically at the document level, start to lose coherence over sufficiently long passages, contradict themselves, and occasionally contain non-sequitur sentences or paragraphs.”Brown et al. (2020), §5
The authors list further weaknesses: poor “common-sense physics” (the paper's example asks whether cheese put in a fridge will melt), near-chance performance on some comparison tasks, a training objective that weights every token equally whether it matters or not, no grounding in the physical world, poor sample efficiency during pre-training, and uncertainty about whether few-shot learning is genuinely learning a new task or recognizing one seen in training.
Why it matters today: several of these limitations directly motivated the next wave of work, especially the one about the objective. Training a model to follow instructions and to prefer helpful answers, rather than just to predict web text, is the subject of the InstructGPT companion.
6 Broader impacts and energy · original
“…even with the full GPT-3 175B, generating 100 pages of content from a trained model can cost on the order of 0.4 kW-hr, or only a few cents in energy costs.”Brown et al. (2020), §6.3
The paper devotes a full section to misuse (disinformation, spam, phishing), to bias in how the model writes about gender, race and religion, and to energy. The energy argument is about amortization: training is enormously expensive once (thousands of petaflop/s-days), but each use afterwards is cheap, so the cost should be judged over the model's lifetime. It also points to distillation as a way to make smaller, cheaper versions.
What happened next
| In the paper | What changed | Where to learn more |
|---|---|---|
| Next-token prediction only | Instruction tuning and preference training (SFT, RLHF) made zero-shot requests work well | training stages lesson, InstructGPT companion |
| 175B parameters on 300B tokens | Later work found this model was under-trained for its size: smaller models on far more tokens do better per unit of compute | scaling laws companion |
| 2,048-token context | Context windows grew by hundreds of times, thanks to efficient attention and position schemes | FlashAttention, RoFormer |
| Closed-book knowledge | Retrieval supplies facts the model does not reliably remember | RAG lesson |
| Few-shot examples in the prompt | Still the standard way to pin down an output format; now usually combined with a system prompt | context engineering lesson |
Glossary
Every term with hover guidance on this page, in one place.