An annotated companion · AI Primer

Evaluating Large Language Models Trained on Code, annotated

About this page. This is a companion, not a copy. It follows the paper section by section, quotes only a sentence or two per section (clearly marked), and explains everything in its own words. The paper is distributed under arXiv's standard licence, so its tables are not reproduced: a few of its numbers are restated in this page's own charts and tables, each attributed, and its figures are redrawn from scratch. Its three-line pass@k function and one problem from its appendix are quoted because the page explains them line by line. Equations are reproduced with every symbol decoded. The paper runs to 35 pages with long appendices on bias, security and economics; those are summarised, and the sections on evaluation, which made the paper famous, are taken apart in full. Numbers the paper does not give are labelled illustrative.

How to read this page

  • Any dotted word explains itself when you hover it, tab to it, or tap it, and so does every symbol in every equation.
  • The sample drawer lets you set how many samples were generated and how many passed, then draw k at random and watch the count converge on the formula.
  • The bias explorer shows why the tempting shortcut formula always reads low.
  • The chain builder assembles the paper's synthetic docstrings from its 13 building blocks and runs them.

Each idea climbs the ladder: everyday picture, tiny example, diagram, the math, why it matters. Two lessons build what this paper introduced: the benchmarks lesson (pass@k and its bias) and the coding agents lesson (hidden tests, a sandbox, and pass@k for agents).

Abstract

“Furthermore, we find that repeated sampling from the model is a surprisingly effective strategy for producing working solutions to difficult prompts.”Chen et al. (2021), Abstract. Read the original

Everyday picture

A student writes a function from its description. You don't grade it by how much it looks like the model answer; you run it and see whether it works. And if the student may hand in several attempts, you count the problem solved when any one of them works. This paper builds that exam for language models, and a careful way to score it.

What the paper claims

  • Codex, a GPT model fine-tuned on public Python code from GitHub, solves 28.8% of the new HumanEval problems with one sample each. GPT-3 solves 0%; GPT-J (6 billion parameters, trained on a mix of text that includes some code) solves 11.4%.
  • Sampling many times works: with 100 samples per problem, it solves 70.2% (if something can pick out the working sample).
  • It struggles with docstrings that describe long chains of operations and with binding operations to the right variables.
  • A long analysis of broader impacts: over-reliance, misalignment, bias, security, economics.

Why it matters today

Codex itself is history, but two things from this paper are everywhere: the idea of grading generated code by running it against unit tests (functional correctness), and the formula for pass@k. Every code benchmark score you read today is built on them.

1 Introduction · original

Everyday picture

GPT-3 was never trained to write code, yet the authors noticed it could write simple Python functions from their docstrings. The hypothesis: a GPT model that reads lots of code will be much better at it.

Tiny example: the headline numbers, from the paper's Figure 1

For Codex-12B and its variants on HumanEval's 164 problems (all samples at temperature 0.8 in that figure):

The share of problems solved under each rule, from the caption of Figure 1 of Chen et al. (2021).
Model and ruleSolved
GPT-12B, one sample0%
Codex-12B, one sample28.8%
Codex-S-12B (extra fine-tuning on standalone functions), one sample37.7%
Codex-S, 100 samples, submit the one with the highest mean log-probability44.5%
Codex-S, 100 samples, solved if any passes the unit tests77.5%

Reading it: read down the rows. The first jump (0% to 28.8%) is what training on code buys; the second (to 37.7%) is fine-tuning on the right kind of code. The last two rows use the same 100 samples and differ only in who picks one: the model's own confidence gets 44.5%, while the unit tests get 77.5%. The gap between those rows is the value of being able to check an answer.

The whole paper in one picture

GPT-3 family12M to 12B params 159 GB of Python54M GitHub repos Codex§3 ≈50,000 problems:contests + CI traces,kept if solvable Codex-Swrites code, §4 Codex-Dwrites docstrings, §5 Evaluation, §2 HumanEval164 problemsn = 200 each run hiddenunit tests ina sandbox count passes c,then estimatepass@k

Hover or tap a box. Follow the models down the left, then the evaluation along the bottom.

The paper's structure as one diagram, drawn from §2 to §5 of Chen et al. (2021).

Reading it: the left column is the family of models: GPT-3 weights, fine-tuned on a large slice of GitHub to make Codex, then fine-tuned again on about 50,000 small, verified “write this function” problems to make Codex-S. The same problems, with the docstring moved to the end, train Codex-D to go the other way. Every model is graded the same way, along the bottom: generate many samples per HumanEval problem, run each against hidden unit tests in a sandbox, count the passes, and turn the count into pass@k. The green box with the thick border is the paper's most lasting contribution.

Why it matters

The paper's structure is still the structure of code-model work: pretrain broadly, fine-tune on code, fine-tune again on the task format, and measure by execution rather than by resemblance.

2 Evaluation framework · original

Three pieces: a metric (pass@k), a dataset (HumanEval), and a safe place to run untrusted code (the sandbox). This section is why the paper is still read.

2.1 Functional correctness and pass@k · original

“Perhaps the most convincing reason to evaluate functional correctness is that it is used by human developers to judge code.”Chen et al. (2021), §2.1

Everyday picture

Two recipes for the same cake can share hardly a word, and a recipe that differs from the original by one word (“salt” for “sugar”) can be a disaster. Comparing text is the wrong test for a recipe; bake it and taste it. Code is the same: BLEU and other text-matching scores compare a sample with a reference solution, but many different programs are correct, and a program one character away from the reference can be wrong. So a sample counts as correct if it passes the problem's unit tests.

Tiny example: counting draws

Take one problem from the paper's Appendix B, is_prime(n): Codex-12B wrote n = 8 samples and c = 3 of them passed. If you could hand in k = 3 of those samples, picked at random, how likely is it that at least one passes? Count. There are C(8, 3) = 56 different groups of 3 you could pick. The groups made only of failing samples come from the 5 failures: C(5, 3) = 10. So 10 of the 56 groups fail completely, and pass@3 = 1 − 10/56 = 0.821.

The older way (Kulal et al., 2019) generates exactly k samples per problem and checks whether any passes. That is fair, but noisy: with k = 3 you would see one random group, and get 1 or 0. The paper instead generates more samples than needed (n = 200, k up to 100) and computes the average over every group of k, exactly, by counting.

The math: the unbiased estimator

In words: “for each problem, pass@k is one minus the share of all possible groups of k samples that contain no passing sample; the benchmark's pass@k is the average of that over the problems.”

With the numbers: for is_prime, n = 8, c = 3, k = 3: 1 − C(5, 3)/C(8, 3) = 1 − 10/56 = 0.821. The paper's Appendix B shows eight problems, each with 8 samples, and 1, 3, 1, 0, 0, 3, 3 and 0 of them pass. Averaging over those eight problems gives pass@1 = 0.172 and pass@3 = 0.402. With k = 1 the formula is simply c / n, the share of samples that pass.

In Python:

import math
n, c, k = 8, 3, 3
# C(n, k): every group of k samples; C(n - c, k): groups made only of failures
math.comb(n, k), math.comb(n - c, k)  # → (56, 10)
# pass@k for this one problem
round(1 - math.comb(n - c, k) / math.comb(n, k), 3)  # → 0.821
# E over problems: the eight Appendix B problems, 8 samples each
cs = [1, 3, 1, 0, 0, 3, 3, 0]
def pass_at(n, c, k):
    return 1.0 if n - c < k else 1 - math.comb(n - c, k) / math.comb(n, k)
round(sum(pass_at(8, c, 1) for c in cs) / len(cs), 3)  # → 0.172
round(sum(pass_at(8, c, 3) for c in cs) / len(cs), 3)  # → 0.402

When fewer than k samples failed (n − c < k), every group of k must contain a pass, and the formula is exactly 1; the code checks that case first. The binomial coefficients C(n, k) are the “n choose k” counts; pass_at_k in the benchmarks lesson is this formula.

The paper's three lines of code

C(200, 100) has 59 digits. Computed naively in floating point, the ratio of two such numbers can lose accuracy or overflow, so the paper's Figure 3 gives a stable version that never forms either number. Their function, quoted:

def pass_at_k(n, c, k):
    if n - c < k: return 1.0
    return 1.0 - np.prod(1.0 - k /
        np.arange(n - c + 1, n + 1))

In words: “the share of all-failing groups equals a product of c small factors, one for each i from n − c + 1 up to n, each a little below 1.”

With the numbers: n = 8, c = 3, k = 3: i runs over 6, 7, 8, giving (1 − 3/6)(1 − 3/7)(1 − 3/8) = 0.5 × 0.571 × 0.625 = 0.179 = 10/56. So pass@3 = 1 − 0.179 = 0.821, the same answer with no large numbers at all.

In Python:

import math
n, c, k = 8, 3, 3
# Π over i = n - c + 1 .. n of (1 - k / i)
fail = math.prod(1 - k / i for i in range(n - c + 1, n + 1))
round(fail, 4), round(10 / 56, 4)  # → (0.1786, 0.1786)
round(1 - fail, 3)  # → 0.821
# why it matters: C(200, 100) has 59 digits
len(str(math.comb(200, 100)))  # → 59

Why it works: C(n − c, k)/C(n, k) is the chance that k draws without replacement all land among the n − c failures. Writing both counts as factorials and cancelling leaves a product of c factors of the form (i − k)/i. Python's integers are exact, so the benchmarks lesson can use math.comb directly; numpy's floats cannot, which is why the paper's version exists.

The tempting shortcut

Why not estimate one sample's pass rate as p̂ = c / n and compute the chance that k independent tries all fail, (1 − p̂)k?

In words: “treat the observed pass rate as the true one and assume k independent tries.”

With the numbers: p̂ = 3/8 = 0.375; 1 − 0.625³ = 1 − 0.244 = 0.756, against 0.821 from counting. The shortcut reads low, and Appendix A shows this is not luck: averaged over every possible outcome it is always low.

In Python:

n, c, k = 8, 3, 3
p_hat = c / n
p_hat  # → 0.375
# 1 - (1 - p̂)^k
round(1 - (1 - p_hat) ** k, 3)  # → 0.756

Why it matters

pass@k with a large k answers “can the model produce a working solution at all, if something checks for us?”; pass@1 answers “will one try work?”. Both come from the same n samples, which is the formula's practical gift: generate once, report any k up to n. The coding agents lesson makes the matching point for agents: a user running an agent once gets pass@1.

Try it: draw k of n samples

Everyday picture

A bag of n marbles, c of them green. Grab k without looking. The formula says how often your handful contains at least one green; drawing many handfuls should agree.

Try it: set n, c and k. Press Draw k at random a few times, then Draw 1,000, and watch the running share of draws that contained a pass settle on the formula. Then set k close to n: once fewer than k samples fail, every draw contains a pass.

Reading it: each square is one generated sample: green and striped if it passed the unit tests, plain grey if it failed. A draw outlines k squares chosen at random, without repeats. The readout gives three numbers: the counting formula 1 − C(n − c, k)/C(n, k), the running share of your draws that caught at least one green square, and the plug-in shortcut 1 − (1 − c/n)k. After a thousand draws the running share sits close to the formula, not to the shortcut: the formula describes exactly this experiment, drawing without replacement from the samples you have.

Why it matters

The formula is not a model of the world; it is an exact count of an experiment you could run on your own samples. That is why it can be exactly unbiased: nothing is assumed beyond the n samples being independent draws from the model.

2.2 HumanEval: hand-written evaluation set · original

“It is important for these tasks to be hand-written, since our models are trained on a large fraction of GitHub, which already contains solutions to problems from a variety of sources.”Chen et al. (2021), §2.2

Everyday picture

If you set a class homework taken from the textbook's exercises, and the answers are in the back of the book, the grades tell you who has the book. The authors wrote 164 fresh problems by hand, because Codex had read much of GitHub, where solutions to published exercises (for example, the Codeforces problems in the APPS dataset) already sit in more than ten public repositories.

Tiny example: one HumanEval problem and eight tries

Each problem has a function signature, a docstring (often with examples), a reference body, and hidden unit tests: 7.7 tests per problem on average. The model sees only the signature and docstring and writes the body. Hover or tap each part.

What the model sees def words_string(s): Split a string of words separated by commasor spaces into a list of the words. "Hi, my name is John" → ["Hi", "my", …] 8 samples at temperature 0.8 1 ✓ 2 ✗ 3 ✗ 4 ✗ 5 ✗ 6 ✗ 7 ✗ 8 ✗ Hidden unit tests (7.7 per problemon average), run in a sandbox c = 1 of n = 8 pass pass@1 = 0.125, pass@3 = 0.375

Hover or tap a part. The eight numbered squares are real samples from the paper's Appendix B; hover the pink ones to see why each failed.

One problem from the paper's Appendix B, words_string, with the verdicts on the eight Codex-12B samples printed there. The docstring is paraphrased; its example is quoted.

Reading it: the frame at the top is everything the model is given. The eight squares are the eight samples the paper prints for this problem: one correct (green), seven wrong (pink). Squares with the same wrong idea light up together: four of the seven are some form of s.split(), which splits on spaces only and keeps the commas, so "Hi," comes back instead of "Hi"; the docstring's own example already catches it. The tests, not a human, deliver each verdict. The count c = 1 of n = 8 is the only thing pass@k needs: pass@1 = 1/8, and pass@3 = 1 − C(7, 3)/C(8, 3) = 1 − 35/56 = 0.375.

In Python: the most common wrong answer, run on the docstring's example.

s = "Hi, my name is John"
# samples 2, 5, 6 and 8 all amount to this
s.split()  # → ['Hi,', 'my', 'name', 'is', 'John']
# the docstring asks for ['Hi', 'my', 'name', 'is', 'John']
s.split() == ["Hi", "my", "name", "is", "John"]  # → False

Why it matters

HumanEval is small (164 problems), easy by professional standards, and Python-only, and all three facts shape how to read a score on it. Small means noisy: the benchmarks lesson computes that a score of 80% on 164 problems carries a 95% margin of about ±6 points. Hand-written means it was fresh in 2021; it has been public since, which is the contamination risk the authors wrote it to avoid.

2.3 Sandbox for executing generated programs · original

Everyday picture

You would taste a stranger's cooking, but not in your own kitchen with the doors unlocked. Generated code is often wrong, and code learned from GitHub can include code that was written to do harm, so it runs somewhere it can break nothing.

What they built

A sandbox meant to stop untrusted programs from modifying the host, persisting on it, reaching sensitive resources, or sending data out. The main wall is gVisor, which emulates the host's resources so a container never touches them directly; network access is cut by firewall rules except for what the experiment itself needs.

Why it matters

Execution-based grading is only as trustworthy as the place code runs. The coding agents lesson builds a small version in plain Python: run_sandboxed runs code with a restricted set of built-ins under step and time Limits, and run_cases grades it against test cases.

3 Code fine-tuning · original

Everyday picture

Take a well-read generalist and give them a year of nothing but reading other people's code. That is Codex: GPT models of up to 12 billion parameters, trained further on Python.

3.1 Data

Collected in May 2020 from 54 million public GitHub repositories: 179 GB of unique Python files under 1 MB. Files that looked auto-generated, had an average line longer than 100 characters or a longest line over 1,000, or had few alphanumeric characters were removed, leaving 159 GB.

3.2 Methods

  • Start from GPT-3. Surprisingly, starting from a pretrained language model gave no better final result than starting from scratch (the code dataset is so large), but it converged faster, so every model starts from GPT.
  • Training. 100 billion tokens with Adam (β₁ = 0.9, β₂ = 0.95, ε = 10⁻⁸), weight decay 0.1, a 175-step linear warmup and cosine decay.
  • A better tokenizer for code. GPT-3's tokenizer wastes tokens on indentation, so they added tokens for runs of whitespace of different lengths: about 30% fewer tokens for the same code.
  • Sampling. Nucleus sampling with p = 0.95, stopping at the first \nclass, \ndef, \n#, \nif or \nprint, since otherwise the model carries on writing more functions.

Why it matters

The stop sequences are a small detail with a big effect on the metric: a correct function followed by stray top-level code can fail its tests. Every harness that grades generated code has to decide where the generation ends.

3.3 Results: loss, temperature and ranking · original

Everyday picture

Three questions a practitioner asks: does a bigger model predict code better, how adventurous should sampling be, and if you can't run tests, how do you pick one sample to show?

The math: test loss follows a power law

In words: “the test loss is the model's size, measured in units of 59.2 million parameters, raised to the power −0.13: every doubling of size multiplies the loss by the same factor.”

With the numbers: at N = 5.92 × 10⁷ the loss is 1. At 3 × 10⁸, (5.07)−0.13 = 0.81; at 1.2 × 10¹⁰, (202.7)−0.13 = 0.50. Each doubling multiplies the loss by 2−0.13 = 0.914. (Illustrative sizes plugged into the paper's fitted curve; N counts non-embedding parameters.)

In Python:

# L(N) = (N / 5.92e7) ** -0.13, the paper's fit on held-out Python
def L(N):
    return (N / 5.92e7) ** -0.13
round(L(5.92e7), 2), round(L(3e8), 2), round(L(1.2e10), 2)  # → (1.0, 0.81, 0.5)
# the factor per doubling of N
round(2 ** -0.13, 3)  # → 0.914

This is the same kind of power law as GPT-3's scaling curves (scaling-laws companion): code fine-tuning did not break the smooth trend.

Temperature: pick it for the k you report

For a 679M-parameter Codex, the best temperature for pass@1 is 0.2, and for pass@100 it is 0.8. Higher temperatures are best for larger k because the samples are more varied, and pass@k only needs one of them to work.

Tiny example (illustrative). Ten problems. At a low temperature, the samples for a problem are near-copies: the model solves 3 problems every time and the other 7 never, so pass@1 = pass@100 = 0.3. At a high temperature, samples vary: say 5 problems are solved by 20% of samples and 5 never. Then pass@1 = 0.1, but with 100 tries each of those 5 is almost surely solved, so pass@100 ≈ 0.5. Low temperature wins pass@1, high temperature wins pass@100.

# low temperature: each problem is all-or-nothing
round(3 / 10, 2), round(3 / 10, 2)  # → (0.3, 0.3)
# high temperature: 5 problems at 20% per sample, 5 at 0%
p = [0.2] * 5 + [0.0] * 5
round(sum(p) / len(p), 2)  # → 0.1
round(sum(1 - (1 - q) ** 100 for q in p) / len(p), 3)  # → 0.5

The math: choosing one sample without tests

pass@k assumes an oracle (the unit tests) picks the working sample. An autocomplete tool can show a user only one suggestion and has no tests. The paper tries ranking samples by the model's own log-probability: the sum over tokens performs slightly worse than picking at random, while the mean per token beats random clearly.

In words: “a sample's score is the average, over its tokens, of the log of the probability the model gave each token; the sample with the highest average is the one shown.”

With the numbers (illustrative): a 4-token wrong one-liner with token log-probabilities (−0.1, −0.2, −0.3, −0.2) sums to −0.8, average −0.2. A 12-token correct loop whose tokens each score −0.15 sums to −1.8, average −0.15. By the sum, the short wrong answer wins; by the mean, the correct one does.

In Python:

# log P(s_t | x, s_<t) for each token of two samples (illustrative)
short_wrong = [-0.1, -0.2, -0.3, -0.2]
long_right = [-0.15] * 12
# Σ_t log P: longer samples pay for every token
round(sum(short_wrong), 2), round(sum(long_right), 2)  # → (-0.8, -1.8)
# (1/T) Σ_t log P: the paper's ranking score
round(sum(short_wrong) / len(short_wrong), 2), round(sum(long_right) / len(long_right), 2)  # → (-0.2, -0.15)

The same length bias appears when multiple-choice benchmarks score options by likelihood; the benchmarks lesson builds both rules in option_logprob.

BLEU does not track correctness

For Codex-12B samples on four random problems, the paper plots the BLEU scores of correct and incorrect samples against the reference solution, and the two distributions overlap heavily. An incorrect sample is guaranteed to differ in behaviour from the reference, so a higher BLEU score is no evidence of a working program. The words_string samples in the figure above show the reason in miniature: four wrong samples are nearly identical text, and the one correct sample, a character-by-character loop, looks like none of them.

Why it matters

Three practical rules that outlived Codex: report the temperature with pass@k (a pass@100 at T = 0.2 undersells a model); without a checker, prefer mean over summed log-probability to pick a sample; and never use text overlap as a stand-in for correctness when the output can be executed.

3.4 Against other models · original

Everyday picture

How much of Codex's ability is size, and how much is what it read? Compare it with models of every size trained on different mixtures.

Tiny example

GPT-Neo 2.7B and GPT-J 6B were trained on The Pile, a text collection that is 8% GitHub code. GPT-Neo 2.7B roughly matches Codex-85M (30 times fewer parameters); GPT-J 6B roughly matches Codex-300M (20 times fewer). Plain GPT models of similar size score near 0%.

Hover the chart, or tab to it and use the arrow keys, to read pass@k at each model size.

Reading it: the x-axis is parameters on a log scale; the y-axis is the percentage of HumanEval problems solved. The top three lines are Codex at eight sizes, from 12 million to 12 billion parameters, for pass@1, pass@10 and pass@100. Every line rises steadily with size, and the gaps between them grow: at 12B, 28.8% of problems are solved on the first try but 72.3% within 100. The lowest line is the text-trained GPT-Neo (125M, 1.3B, 2.7B) and GPT-J (6B) at pass@1, each the best of the temperatures the authors tried: at 2.7B it sits where Codex was at 85M. Numbers from Table 1 of Chen et al. (2021).

Selected rows from Table 1 of Chen et al. (2021): percentage of HumanEval problems solved.
Modelpass@1pass@10pass@100
GPT-J 6B11.6215.7427.74
TabNine (a code autocomplete product)2.584.357.59
Codex-300M13.1720.3736.27
Codex-12B28.8146.8172.31

Two numbers disagree inside the paper. The abstract and introduction give GPT-J 11.4% and Codex-12B 70.2% with 100 samples; Table 1 gives 11.62% and 72.31%, and §3.4 says 11.6%. The paper does not say which settings produced the abstract's figures.

Why it matters

What a model reads matters as much as its size: 8% code in the training mix bought GPT-Neo and GPT-J a real, if small, ability, and a code-only diet bought Codex a 20 to 30 times advantage in parameters. The same logic, choosing a data mixture for the skills you want, runs through the pretraining lesson.

3.5 Results on the APPS dataset · original

Everyday picture

HumanEval asks for one function. APPS (Hendrycks et al., 2021) asks for whole programs that read input and print output, like programming-contest problems, in three tiers from introductory to competition level, 5,000 problems each for training and test.

Tiny example: filtering with the public examples

APPS problems print 3 input/output examples in their statement. Contestants test against those before submitting, so the paper does too: sample 1,000 programs, keep only those that pass the 3 public examples, and compute pass@k among the survivors. On the introductory tier, Codex-12B (given one example in its docstring, “1-shot”) solves 4.14% with one raw sample, 25.02% with 1,000 raw samples, and 22.78% with one filtered sample.

Why it matters

Filtered pass@1 is a realistic middle ground between pass@1 and an oracle: the public examples act as a cheap verifier, and one filtered sample gets most of the way to what 1,000 raw samples achieve. It is the same pattern as best-of-n, with a checker that sees only part of the truth. Timeouts count too: a correct but slow program fails a 3-second limit, and the paper reports those separately.

4 Supervised fine-tuning: Codex-S · original

Everyday picture

GitHub is mostly not “here is a docstring, now write the function”: it is classes, scripts, configuration files, data. To get better at the exam's format, practise that format.

Tiny example: where the practice problems came from

  • About 10,000 problems from programming-contest and practice sites, turned into HumanEval-style tasks. Their full test suites are hidden, so tests were built from the examples in the statements, or recovered by submitting deliberately wrong solutions.
  • About 40,000 functions from open-source projects with continuous integration: run each project's tests with Python's sys.setprofile hook, record the inputs and outputs of every function called, and turn them into unit tests. These are mostly utility functions: following instructions rather than clever algorithms.
  • Filtering by sampling. Codex-12B tried each problem 100 times. If no sample passed, the problem was judged ambiguous or too hard and dropped; repeated runs removed problems whose results changed from run to run.

The math: train only on the solution

Each training example is prompt (signature and docstring) followed by the reference solution. The loss counts only the solution's tokens: the prompt's tokens are masked out, because the goal is to write functions, not docstrings. This is supervised fine-tuning with a loss mask.

In words: “the loss is minus the sum of the log-probabilities the model gives each token of the reference solution, given the prompt and the solution so far; no term comes from the prompt's own tokens.”

With the numbers (illustrative): a 3-token solution whose tokens get probabilities 0.9, 0.8 and 0.5: −(log 0.9 + log 0.8 + log 0.5) = 0.105 + 0.223 + 0.693 = 1.022. The least likely token contributes most, so training pushes hardest there.

In Python:

import math
# P(y_t | x, y_<t) for each solution token (illustrative); prompt tokens are not listed: they are masked
P = [0.9, 0.8, 0.5]
[round(-math.log(p), 3) for p in P]  # → [0.105, 0.223, 0.693]
# L = -Σ_t log P(y_t | x, y_<t)
round(-sum(math.log(p) for p in P), 3)  # → 1.022

Results

Codex-S beats Codex of the same size by 6.5 points of pass@1 and 15.1 points of pass@100 on average across sizes, and it is one to two orders of magnitude more parameter-efficient. It prefers slightly higher temperatures than Codex for every k above 1 (the paper uses 0 for pass@1 and 1 for pass@100), which the authors take as a sign that it models a narrower distribution. Ranking by mean log-probability beats random ranking by 11.6 points on average for Codex-S-12B.

Why it matters

This is the recipe of the training stages lesson in miniature: a broad base, then a small, clean dataset in the exact format you care about. The filtering step is worth noticing: a model's own samples, checked by tests, decide which training problems are trustworthy.

5 Docstring generation: Codex-D · original

Everyday picture

If a model can turn a description into code, can another turn code into a description? That would help people see what generated code is meant to do, which the authors want for safety reasons.

Tiny example

Reorder each training problem as signature, then solution, then docstring, and train on the docstring's tokens. There is no unit test for a docstring, so the authors graded samples by hand: a docstring is correct if it uniquely and accurately specifies the body. Ten samples per problem, 1,640 in all, from Codex-D-12B at temperature 0.8.

From Table 3 of Chen et al. (2021): pass rates, graded by unit tests for Codex-S and by hand for Codex-D.
Modelpass@1pass@10
Codex-S-12B (code from docstring)32.2%59.5%
Codex-D-12B (docstring from code)20.3%46.5%

Reading it: the reverse direction is somewhat harder but in the same range. The common failures were leaving out an important detail (“an answer must be to two decimal places”) and inventing a task from the function's name instead of its body. Codex-D also wrote docstrings such as “I just found this function online”: it had learned from real developers' docstrings.

Back-translation as a ranking score

Codex-D gives another way to pick one sample out of k: prefer the sample from which Codex-D finds the original docstring most probable, P(docstring | sample). This back-translation score beats random ranking but loses to mean log-probability, and it overfits quickly.

Why it matters

Two lessons: a task with no automatic grader forces expensive hand grading (the reason functional correctness was so valuable), and a model's second opinion is a weaker selector than a test.

6 Limitations · original

“We find that as the number of chained building blocks in the docstring increases, model performance decreases exponentially.”Chen et al. (2021), §6

Everyday picture

Give someone a recipe with two steps and they follow it; give them twelve steps read aloud once, and one gets skipped. A person who can do a chain of two can do a chain of twelve with care. The paper finds Codex cannot: every extra step costs it a large share of its success.

Tiny example: the chain builder

The paper built synthetic problems from 13 one-line string operations (its Appendix C), each with a sentence for the docstring and a line of code. Chaining m of them gives a docstring of m instructions and a body of m lines. Pass rates fell by a factor of roughly 2 to 3 with each extra block (its Figure 11).

Try it: choose how many blocks to chain, in the paper's order, and type an input string. Read the docstring a model would get, and the output the correct body produces. Past about five blocks, notice how hard it is to check the answer in your head.


  

Reading it: the box shows the docstring: one numbered instruction per block, in the order the paper lists them. The readout shows the string after each block and the final result, computed by the same one-line operations the paper gives (lowercasing, removing vowels, dropping every third character, reversing word order, and so on). Each step is trivial; the difficulty is keeping track of all of them. With an illustrative 90% success at one block and a factor of 2.5 per extra block, the pass rate would be 36% at two blocks and 14% at three.

Binding operations to variables

From the paper: asked to “Add 3 to y, then subtract 4 from both x and w. Return the product of the four numbers”, Codex-12B computed t = y + 3 and u = x - 4, never decremented w, and returned z * w instead of the product of all four. The more operations and variables in a docstring, the more such mix-ups.

Why it matters

HumanEval problems are short and self-contained, so a high score says little about long, multi-step specifications or about code that must fit a large existing system. The authors also note that training is not sample-efficient: Codex read hundreds of millions of lines of code, far more than any developer, yet a strong student finishing an introductory course should solve more of HumanEval than Codex-12B. Real repositories and multi-step work are what later benchmarks, graded the same way by hidden tests, set out to measure (the coding agents lesson builds one in miniature).

7 Broader impacts and hazard analysis · original

Everyday picture

A new power tool comes with a safety sheet: what it can do, how it can hurt you, and how to use it well. Section 7 and Appendices E to H are that sheet for code generation.

The main findings, one line each

  • Over-reliance (§7.1). Suggestions can look right and be wrong, which matters most for beginners; human review is required.
  • Misalignment (§7.2, Appendix E). When the prompt contains subtly buggy code, Codex writes worse code than it is capable of, even when told to write correct code, and the gap grows with model size. A model trained to continue its input matches the input's quality, not the user's intent: a failure of alignment, not of ability.
  • Bias (§7.3, Appendix F). Code and comments can encode stereotypes, much as text models do.
  • Economics (§7.4, Appendix H). Effects are limited by how little of an engineer's day is writing code; the paper calls its analysis preliminary.
  • Security (§7.5, Appendix G). Asked to create encryption keys, Codex models often chose clearly insecure settings (RSA keys shorter than 2,048 bits, AES in ECB mode), with no clear improvement with size.
  • Environment and law (§7.6, §7.7). Fine-tuning Codex-12B took compute similar to training GPT-3-12B; generated code matched training snippets in under 0.1% of cases in one study.

A small inconsistency

Appendix E runs its alignment evaluation on “158 problems” of HumanEval, while §2.2 describes 164. The paper does not explain the difference.

Why it matters

The misalignment result is the one that aged into a research field: capability and intent are different things, and more capability can make the gap worse. The paper even suggests RLHF, with tests and verification tools helping the human labellers, as a direction; the InstructGPT companion covers the method.

8 Related work and 9 Conclusion · original

Earlier neural work split into program induction (the network computes the output itself) and program synthesis (the network writes a program). Large transformers moved synthesis forward (CodeBERT, PyMT5), and functional correctness was already used by SPoC, whose fixed budget of compilations is close to pass@k, and TransCoder, which also found it tracked quality better than BLEU. A cautionary note from bug-fixing research: weak test suites let “fixes” pass by deleting the failing functionality, which is why tests have to be good for execution-based grading to mean anything.

The conclusion: fine-tuning GPT on GitHub code gives strong performance on human-written problems of easy exercise difficulty; training on a distribution closer to the evaluation, and sampling more, both help; the reverse task (docstrings from code) is easy to train and behaves similarly; and there is significant room for improvement.

Appendix A: estimating pass@k · original

“Evaluating pass@k in an unbiased way with any number of samples n is important for fair comparison.”Chen et al. (2021), Appendix A

Everyday picture

A scale that always reads a little light is biased: weigh many things and the errors don't cancel, they pile up in one direction. An unbiased estimator may be off on any single reading, but its average over many readings is the true value.

Tiny example

Suppose a model's true chance of a correct sample is p = 0.3, and you generate n = 8 samples and report k = 3. The true pass@3 is 1 − 0.7³ = 0.657. The number of passing samples c is random: sometimes 1, sometimes 4. Average each estimator over every possible c, weighted by how likely it is: the counting formula averages exactly 0.657, while the plug-in 1 − (1 − c/n)³ averages 0.603.

The math: the proof in the appendix

c follows a binomial distribution: each of the n samples passes independently with chance p. The paper's derivation, one line at a time:

In words: “average the estimator over every possible number of passing samples i, weighting each by its binomial chance; a counting identity turns each term into a binomial chance for n − k samples, those chances add up to 1, and what is left is exactly the true pass@k.” The sum stops at n − k because for more passes the estimator is 1, and 1 − 1 contributes nothing.

With the numbers: n = 8, k = 3, p = 0.3. The identity for i = 2: C(6, 3)/C(8, 3) × C(8, 2) = 20/56 × 28 = 10 = C(5, 2). The whole average: 0.657 = 1 − 0.7³. The plug-in's average: 0.603, low by 0.054.

In Python:

import math
C = math.comb
n, k, p = 8, 3, 0.3
# the identity inside the proof, for i = 2
C(6, 3) / C(8, 3) * C(8, 2), C(5, 2)  # → (10.0, 10)
# the binomial chance of each count c of passing samples
pmf = [C(n, c) * p**c * (1 - p) ** (n - c) for c in range(n + 1)]
unbiased = sum(w * (1.0 if n - c < k else 1 - C(n - c, k) / C(n, k)) for c, w in enumerate(pmf))
plug_in = sum(w * (1 - (1 - c / n) ** k) for c, w in enumerate(pmf))
round(unbiased, 3), round(1 - (1 - p) ** k, 3), round(plug_in, 3)  # → (0.657, 0.657, 0.603)

Why the plug-in reads low

The paper's reading: the plug-in behaves like drawing k samples with replacement from the n you have, so the same failing sample can be drawn twice, and the k draws are not independent tries of the model. A second way to see it: (1 − p̂)k curves upward in p̂ (it is convex), so averaging it over a spread of p̂ gives more than its value at the average, and one minus that gives less.

Hover the chart, or tab to it and use the arrow keys, to read each estimator's average at every n.

The idea of the paper's Figure 13, recomputed exactly on this page: each estimator's average over every possible outcome of n samples. Not the paper's data.

Reading it: the x-axis is n, how many samples were generated (at least k); the y-axis is pass@k, zoomed to the range the lines occupy (it does not start at 0). The solid line is the true pass@k, 1 − (1 − p)k, flat because it does not depend on n. The dashed line is the counting estimator's average: it lies exactly on the truth at every n. The dotted line is the plug-in's average: always below, and climbing towards the truth as n grows. That climb is the danger the paper names: with the plug-in, the same model looks better simply because more samples were drawn, and the gap has not closed even at n = 5k. Try p = 0.1 with k = 10 for a large gap.

Why it matters

Two labs that generate different numbers of samples can compare pass@k fairly only with an unbiased estimator. That is why this formula, not the obvious one, became the standard, and why the benchmarks lesson builds both (naive_pass_at_k) and checks the bias exactly with expected_pass_at_k.

What happened next

In the paperWhat came laterLearn it
164 problems, 7.7 tests eachEvalPlus (Liu et al., 2023, arXiv:2305.01210) added 80 times more tests to make HumanEval+, catching wrong code that had passed and cutting pass@k by up to 19.3 to 28.9%; it also changed some model rankingshidden tests
Standalone functions from docstringsReal repository issues graded by the project's own tests: SWE-bench (Jimenez et al., 2023, arXiv:2310.06770), with fail-to-pass and pass-to-pass testsresolved rate
pass@k with an oraclepass@1 as the headline for agents that run once; pass@k for systems with a checkerpass@k for agents
Unit tests pick the sampleTests, verifiers and votes as selectors at inference timeself-consistency companion
Hand-written to avoid leakageA public benchmark since 2021, so contamination checks matter for any recent scorecontamination

The Reflexion companion shows HumanEval in its later role: a standard yardstick for agents that write code, test it and retry.

Glossary

Every term with hover guidance on this page, in one place.