primer.ml.big_picture

The big picture: what happens when you send a prompt

Run: python -m primer.ml.big_picture

New to the notation (vectors, sums, logarithms)? Every symbol is decoded where it appears, and primer.notation teaches them all from zero.

Level 1: The practitioner's guide

In one sentence. A language model is a next-token scorer run in a loop: the prompt is cut into tokens, each token becomes a vector, the vectors pass through a stack of transformer blocks, the last position is scored against the whole vocabulary, one token is drawn, appended, and the loop runs again until a stop token or a length limit.

When you need it. You need this picture whenever you set a parameter you can't explain (temperature, top_p, max_tokens), read a bill that counts input and output tokens differently, or debug an answer that was cut off, repeats itself, or comes out as gibberish. The tell: you are adjusting a sampling knob by trial and error, or estimating cost in words when the meter counts tokens. This lesson's toy tokenizer turns "Reset your password" into 5 tokens, and " password" is one of them while "Reset" is two, so a word count is only an estimate. You don't need this lesson to write a good prompt, and you don't need it to pick a model by its benchmark scores; you need it the first time a model's behaviour has to be predicted rather than observed.

Your options. The loop has a few knobs a caller can turn, from the cheapest to the most certain:

Option What it does What it guarantees What it costs Where it lives
Leave sampling at its defaults Draws each token in proportion to the model's probabilities (temperature 1) The variety the model was trained to produce Run-to-run variation The API's defaults
Lower the temperature, down to greedy Divides the scores before softmax; at 0 it takes the single most likely token The likely answer more often: at temperature 0.5 this lesson's top token rises from 0.63 to 0.84 Blander text, and repetition at 0 (Holtzman et al., 2019); still not identical runs One request parameter
Cut the tail (top-k, top-p) Keeps only the k most likely tokens, or the smallest set whose probabilities reach p No draws from the long tail of near-zero tokens A knob some hosted APIs have withdrawn; open-model servers keep it One request parameter
Bound the length (max_tokens, stop sequences) Ends the loop at a token limit or at a string you name A ceiling on cost and latency per call Truncated answers if the limit is too low; check the stop reason One request parameter
Reuse the prefix (KV cache, prompt caching) Keeps the work done on tokens that haven't changed Each new token costs one position instead of a re-read of everything: 219 positions against 23,900 for 200 tokens after a 20-token prompt Server memory; caching rules and prices that differ by vendor The serving stack
Constrain the output (schema, grammar) Forbids, at every step, any token that cannot lead to a valid shape A parseable answer by construction A compiled schema and a small check per token The model server (primer.ml.structured_output)

How to choose. Start from who reads the answer and how long it is.

  • Extraction, classification, tool calls: low temperature, a schema where the server offers one, and a max_tokens sized to the answer.
  • Writing, brainstorming, dialogue: the default temperature, and a length bound that stops runaway output rather than shaping it.
  • Long documents in the prompt: the input is read in one parallel pass, so it costs money more than time; cache the unchanged prefix when you send it repeatedly.
  • Agents and multi-turn systems: every turn re-sends the whole history, so the prompt grows with the conversation; bound each answer and watch the stop reason, because a truncated tool call is a broken one.
  • Whatever you pick, measure tokens in and out on your own traffic. Output tokens are produced one loop iteration at a time, so the cheapest way to make a call faster is to ask for less.

What it costs. Two meters run. Input tokens go through the model in one parallel pass, so a long prompt costs money and memory more than time. Output tokens are produced one per trip round the loop, so latency is proportional to length: this lesson's naive loop reads the prompt plus everything generated so far at every step (14 positions to produce 4 tokens from a 2-token prompt), and a KV cache turns that into one position per token. Context is a hard ceiling: the position table has a fixed number of rows (64 in this toy; max_position_embeddings in a Hugging Face config, where LlamaConfig defaults to 2048, and 4k in the Llama 2 paper), and prompt plus answer must fit inside it. Quality is measured on the same loop: the training loss is the average of −ln p for the real next token, and perplexity is e raised to it. This lesson's untrained model scores 5.97 against 5.99 for a blind guess over 400 tokens (perplexity 393), and training over trillions of tokens (2.0T for Llama 2) is what pushes that number down.

What breaks.

  • Gibberish. Random weights give a nearly flat guess over the vocabulary, and the demo's untrained model continues "The cat" with byte fragments. In practice the same symptom comes from a tokenizer that doesn't match the model, so ids fetch the wrong rows of the table. Check the pairing before the weights.
  • Cut-off answers. The loop stopped at max_tokens, not at a stop token. Claude's API reports end_turn when the model finished on its own and another stop reason when your limit or stop sequence ended it; treat anything but a natural end as incomplete.
  • Temperature 0 that still varies. Greedy picks the argmax, but Claude's reference says results are not fully deterministic even at 0.0, and the arithmetic on real serving hardware is why.
  • Repetition. Always taking the most likely token produces loops of the same phrase; Holtzman et al. (2019) showed it and proposed nucleus sampling as the fix.
  • A bill that surprised you. Cost is counted in tokens, not words, and a conversation re-sends its whole history every turn.
  • A context error. Prompt plus requested output exceeded the position table. Shorten the prompt or the max_tokens, or retrieve less.

In the wild. The pipeline is the decoder-only recipe of GPT-2 (Radford et al., 2019: byte-level BPE, learned positions, tied embeddings) that GPT-3 (Brown et al., 2020) scaled until instructions in the prompt were enough. Hugging Face's generate() exposes the loop's knobs in a GenerationConfig (do_sample, otherwise greedy; temperature 1.0, top_k 50, top_p 1.0, max_new_tokens, repetition_penalty, num_beams), and its LlamaForCausalLM returns logits of shape (batch, sequence, vocabulary) with a logits_to_keep option because, as its docs put it, only the last token's logits are needed for generation. Claude's Messages API offers temperature (0.0 to 1.0, default 1.0; models released after Claude Opus 4.6 accept only 1.0), max_tokens, stop_sequences and a stop_reason on every response, and lets you set max_tokens to 0 to warm the prompt cache without generating. Karpathy's nanoGPT is this lesson at full size, in a few hundred lines. The papers are linked at the end of the lesson.

Go deeper. Level 2 traces "Reset your password" through every box with its shapes, works temperature by hand on three scores, counts the loop's positions, and measures the loss of a model that has learned nothing. If you only needed to set the knobs, you are done.

Level 2: How it works, from scratch.

Level 2: How it works, from scratch

This lesson wires the real pieces from the other lessons into one working pipeline: the tokenizer from primer.ml.tokenization, the transformer from primer.ml.transformer, and a sampling loop. Keep this one picture in your head; every other lesson zooms into one box of it.

flowchart LR A[Prompt text] --> B[Tokenizer<br/>text to IDs] B --> C[Embedding lookup<br/>IDs to vectors] C --> D[Add position info] D --> E[Transformer blocks<br/>repeated N times] E --> F[Output layer<br/>score per vocab token] F --> G[Softmax + sampling] G --> H[Next token] H -->|append and repeat| E

Reading it: read left to right, then follow the loop back. The prompt is cut into token ids, each id picks a vector from a table, and position information is mixed in so order counts. The vectors pass through the transformer blocks, where tokens exchange information. The output layer turns the last token's final vector into a score for every token in the vocabulary; softmax makes those scores probabilities and one token is drawn. That token is appended and the loop runs again, until the model emits a stop token or hits a length limit. Training uses the same boxes with one addition, a loss, at the end (section 6).

1. Text to ids: the coat check

Everyday picture. A coat check. You hand over a coat (a piece of text) and get back a numbered ticket (a token id). The model only ever handles the tickets.

Tiny worked example. With this module's toy tokenizer, "Reset your password" becomes 5 tickets: Re set y our password → [346, 377, 309, 328, 291]. The common word " password" got a single ticket; the rarer pieces got several.

flowchart LR T["'Reset your password'"] --> TK["tokenizer<br/>(learned kit of 400 pieces)"] TK --> I["[346, 377, 309, 328, 291]"]

Reading it: one box, text in and integers out. Everything to the right of this box works only with integers and vectors.

The code. tok.encode(prompt); how the kit is learned is the whole of primer.ml.tokenization.

In code: build_pipeline trains the toy tokenizer (primer.ml.tokenization.ByteBPE) and builds an untrained primer.ml.transformer.TinyGPT sized to its vocabulary.

Why it matters. Prompt length, price and context limits are all counted in these tickets.

2. Ids to vectors: a lookup, not a computation

Everyday picture. A dictionary where ticket number 291 opens to page 291, and each page holds a list of numbers describing that token. A second dictionary, indexed by seat number, describes where the token sits.

Tiny worked example. The table has 400 rows (one per token) of 32 numbers. Id 291 fetches row 291. The 5 ids fetch 5 rows: a 5 × 32 grid. Row i of the position table (i = 0 to 4) is added to row i of that grid.

flowchart LR I["ids (5)"] --> E["token table<br/>400 × 32"] E --> X["5 × 32: what each token is"] P["positions 0..4"] --> PT["position table<br/>64 × 32"] PT --> Y["5 × 32: where each token is"] X --> ADD(("+")) Y --> ADD ADD --> OUT["5 × 32 input to the blocks"]

Reading it: two lookups, one add. Nothing is multiplied here, rows are simply fetched, which is why this step is nearly free. The tables themselves are learned during training, so similar tokens end up with similar rows (see primer.ml.embeddings).

The math and the code.

Level 3: the formula and its symbols

$$ x_i = E_{t_i} + P_i $$

Symbols

Symbol Meaning here In the example
$i$ a token's position in the prompt, from 0 4 (the last token)
$t_i$ the token id at position $i$ $t_4$ = 291
$E$ the token embedding table 400 × 32
$E_{t_i}$ row $t_i$ of that table row 291: 32 numbers
$P_i$ row $i$ of the position table row 4: 32 numbers
$x_i$ what the first block receives for position $i$ 32 numbers

In words: "each token's input vector is its token's row plus its position's row."

With the numbers: $x_4 = E_{291} + P_4$, one 32-number list plus another. The code is model.wte[ids] + model.wpe[:len(ids)]. In miniature, with 3-number rows instead of 32: $E_{291}$ = (0.2, −0.1, 0.5) and $P_4$ = (0.1, 0.3, −0.2) give $x_4$ = (0.3, 0.2, 0.3).

Level 3: in Python

In Python:

# row 291 of the token table (3 numbers, not 32)
E_291 = [0.2, -0.1, 0.5]
# row 4 of the position table
P_4 = [0.1, 0.3, -0.2]
# E_(t_i) + P_i, number by number
x_4 = [round(e + p, 2) for e, p in zip(E_291, P_4)]
x_4  # → [0.3, 0.2, 0.3]

In code: trace does both lookups and the add, and keeps every stage's result so you can inspect the grid before and after positions are mixed in.

Why it matters. This is the only place a token's identity enters the model; every later step works on these vectors.

3. The transformer blocks: rounds of meeting and desk work

Everyday picture. The team from primer.ml.transformer: each round is a meeting where every token listens to the others (attention), then desk work where each token thinks alone (feed-forward). This toy runs 2 rounds; large models run dozens.

Tiny worked example. The 5 × 32 grid goes into block 1 and comes out 5 × 32; the same through block 2. By the end, the vector at the last position ("password") has absorbed information from "Re", "set", "y" and "our".

flowchart LR X["5 × 32"] --> B1["block 1<br/>meeting + desk work"] --> B2["block 2"] --> LN["final norm"] --> H["5 × 32<br/>context-aware vectors"]

Reading it: the shape never changes, only the contents. Each block edits every token's vector by adding what it learned from the others.

The math and the code.

Level 3: the formula and its symbols

$$ h = \text{LN}\big(\text{Block}_N(\cdots\text{Block}_2(\text{Block}_1(x))\cdots)\big) $$

Symbols

Symbol Meaning here In the example
$x$ the 5 × 32 input grid from section 2
$\text{Block}_k$ the $k$-th transformer block $N$ = 2
$\cdots$ "and so on, for every block in between"
$\text{LN}$ the final layer norm
$h$ the context-aware vectors 5 × 32

In words: "run the input through every block in turn, then normalize."

With the numbers: here N = 2, so h = LN(Block₂(Block₁(x))). The code is the for block in model.blocks loop in trace. In miniature, with one 3-number vector and two stand-in blocks that each add an edit: x = (1, 2, 3), Block₁ adds (0, 1, 0) and Block₂ adds (0, 0, 2), giving (1, 3, 5). The final norm subtracts the mean (3) and divides by the spread (1.63), so h = (−1.22, 0, 1.22).

Level 3: in Python

In Python:

import statistics
x = [1.0, 2.0, 3.0]
# a stand-in edit
def block_1(v): return [v_i + d for v_i, d in zip(v, [0, 1, 0])]
def block_2(v): return [v_i + d for v_i, d in zip(v, [0, 0, 2])]
def LN(v):
    mu, sigma = statistics.fmean(v), statistics.pstdev(v)
    # centre, then rescale
    return [round((v_i - mu) / sigma, 2) for v_i in v]
h = x
# Block_1 first, then Block_2 ... up to Block_N
for block in [block_1, block_2]:
    h = block(h)
h  # → [1.0, 3.0, 5.0]
LN(h)  # → [-1.22, 0.0, 1.22]

Why it matters. This is where nearly all the compute and all the "understanding" happen.

4. Scores, softmax and temperature

Everyday picture. A scoreboard with one line for every token in the vocabulary. Softmax turns the scores into shares of a pie. Temperature sets how adventurous the pick is: low temperature almost always takes the biggest slice; high temperature gives the small slices a real chance.

Tiny worked example. Three candidate tokens scored 2.0, 1.0 and 0.5.

temperature p(A) p(B) p(C)
0 (greedy) 1.000 0.000 0.000
0.5 0.844 0.114 0.042
1 0.629 0.231 0.140
2 0.481 0.292 0.227
flowchart LR H["last row of h<br/>32 numbers"] --> S["× token tableᵀ<br/>400 scores"] S --> T["÷ temperature"] T --> SM["softmax<br/>400 probabilities"] SM --> PICK["draw one token"]

Reading it: only the last position's vector is used to choose the next token. It is scored against every token's embedding, the scores are divided by the temperature, softmax turns them into probabilities that sum to 1, and one token is drawn at random in proportion to them.

The math and the code. Softmax raises e (≈ 2.718) to the power of each score, so bigger scores get disproportionately bigger shares, then divides by the total so the shares add up to 1:

Level 3: the formula and its symbols

$$ p_i = \frac{e^{z_i / T}}{\sum_{j=1}^{V} e^{z_j / T}} $$

Symbols

Symbol Meaning here In the example
$z_i$ the score of candidate token $i$ $z_A$ = 2.0
$T$ the temperature 0.5
$e^{\cdot}$ e ≈ 2.718 raised to that power $e^{4}$ = 54.6
$V$ number of candidates (the vocabulary size) 3 here, 400 in the model
$\sum_{j=1}^{V}$ add up over every candidate $j$
$p_i$ probability that token $i$ comes next 0.844

In words: "the chance of token i is e to the power of its score over the temperature, divided by the same quantity summed over all tokens."

With the numbers: T = 0.5 doubles every score to (4, 2, 1): e⁴ = 54.6, e² = 7.39, e¹ = 2.72, total 64.7, so p(A) = 54.6 / 64.7 = 0.844 (next_token_probs). Temperature 0 is the limit case: all probability on the top score.

Level 3: in Python

In Python:

import math
z = [2.0, 1.0, 0.5]
T = 0.5
# e^(z_i / T) for each candidate
exps = [math.exp(z_i / T) for z_i in z]
[round(e, 2) for e in exps]  # → [54.6, 7.39, 2.72]
# Σ_j e^(z_j / T)
total = sum(exps)
round(total, 1)  # → 64.7
# p_i: each share of the total
[round(e / total, 3) for e in exps]  # → [0.844, 0.114, 0.042]

At T = 0.25 token A takes 98% of the probability; at T = 2 the three tokens share it 48%, 29% and 23%

Reading it: three groups of bars, one per candidate token; within each group the bars run from low temperature (left) to high (right). For token A the bars fall as temperature rises, and for tokens B and C they rise. At T = 0.25, A takes almost everything; at T = 2 the three are much closer.

The untrained model's 12 favourites are random byte fragments, each only 1.25 to 1.4 times the uniform 1-in-400 share

Reading it: the bars are the 12 most probable next tokens after "The cat sat on the", according to our untrained model; the dashed line is a uniform guess, 1 in 400 (0.25%). Even these favourites clear the line only modestly: the top one gets about 0.35%, 1.4 times the uniform share, and the twelfth about 1.25 times. Across all 400 tokens every probability stays between 0.74 and 1.39 times uniform, and the "favourites" are random byte fragments. That is exactly what random weights should give: a nearly flat guess with small random bumps. The machinery works, but nothing has been learned yet.

Why it matters. Use low temperature for extraction and tool calls, where you want the most likely answer, and higher temperature for creative writing. Temperature 0 reduces randomness but doesn't guarantee identical outputs on real serving hardware.

5. The loop: append and repeat

Everyday picture. Writing a sentence one word at a time, and re-reading everything you've written before choosing each new word.

Tiny worked example. A 2-token prompt, 4 new tokens. Step 1 reads 2 tokens, step 2 reads 3, step 3 reads 4, step 4 reads 5: 14 token positions to produce 4 tokens. With a 20-token prompt and 200 new tokens, the naive loop reads 23,900 positions; a KV cache reads 219.

flowchart LR S["ids so far"] --> M["full forward pass<br/>over ALL ids"] M --> P["probabilities for the next id"] P --> D["draw one id"] D --> A["append it"] A -->|"not done"| S A -->|"stop token or length limit"| OUT["decode ids to text"]

Reading it: the loop box is the whole of generation. The expensive arrow is "full forward pass over ALL ids": every step re-reads text that hasn't changed. Because of the causal mask, earlier tokens' internal vectors can't change when new tokens arrive, so that repeated work is pure waste. The KV cache (primer.ml.inference) stores it once.

The math. Total positions processed without a cache:

Level 3: the formula and its symbols

$$ W = \sum_{t=0}^{n-1} (p + t) = n\,p + \frac{n(n-1)}{2} $$

Symbols

Symbol Meaning here In the example
$p$ prompt length in tokens 2
$n$ tokens to generate 4
$t$ the step counter, from 0 to $n-1$ 0, 1, 2, 3
$p + t$ tokens re-read at step $t$ 2, 3, 4, 5
$W$ total positions run through the model 14

In words: "each step re-reads the prompt plus everything generated so far; add that up over all steps."

With the numbers: 4 × 2 + (4 × 3) / 2 = 8 + 6 = 14, the count generate reports as positions_processed.

Level 3: in Python

In Python:

p, n = 2, 4
# Σ over t = 0 .. n-1 of (p + t): 2 + 3 + 4 + 5
sum(p + t for t in range(n))  # → 14
# the closed form gives the same count
n * p + n * (n - 1) // 2  # → 14

Without a cache work grows with the square of output length, 23,900 positions at 200 tokens against 219 with a KV cache

Reading it: the x-axis is how many tokens have been generated after a 20-token prompt; the y-axis is total work in token positions. The red curve bends upward: it grows with the square of the output length. The blue line (with a KV cache) grows by exactly one per token. At 200 tokens the gap is more than a hundredfold.

In code: Generation holds the generated ids, the decoded text and the positions count; naive_vs_cached_work counts the positions processed with and without a KV cache for the figure.

Why it matters. Output length drives latency and cost; this is why every serving system caches keys and values.

6. Training: the same forward pass, plus a loss

Everyday picture. A guessing game with instant feedback. Cover the next word, guess it, uncover it, and note how surprised you were. Training nudges every weight to make the surprise smaller next time.

Tiny worked example. If the model gives the right next token probability 0.5, the loss is −ln 0.5 = 0.693. Our untrained model averages 5.97 on a real sentence, close to ln 400 = 5.99, the score of a blind uniform guess.

flowchart LR T["training text"] --> F["forward pass<br/>(this whole lesson)"] F --> P["probabilities at every position"] P --> L["loss: −ln p(actual next token)<br/>averaged over positions"] L --> B["backpropagation<br/>gradient for every weight"] B --> U["optimizer nudges weights"] U -->|next batch| F

Reading it: the first two boxes are the forward pass you just traced. Training adds the loss (how surprised the model was by the real next tokens), then backpropagation (primer.ml.neural_net) works out how each weight contributed, and the optimizer (primer.ml.optimizers) adjusts them. One pass over n tokens gives n − 1 guesses at once, in parallel, which is a big reason transformers train fast.

The math and the code. A logarithm answers "to what power must I raise e to get this number?" For probabilities between 0 and 1 it is negative, so we flip the sign; ln 1 = 0 (no surprise) and ln of a tiny number is very negative (huge surprise).

Level 3: the formula and its symbols

$$ \mathcal{L} = -\frac{1}{n-1}\sum_{i=1}^{n-1} \ln p\big(t_{i+1} \mid t_1, \ldots, t_i\big) $$

Symbols

Symbol Meaning here In the example
$n$ tokens in the training text 12 ("The model reads tokens, not words.")
$t_i$ the token at position $i$
$p(t_{i+1} \mid t_1,\ldots,t_i)$ probability the model gave the actual next token, having seen everything before it; the bar $\mid$ reads "given" ≈ 1/400 when untrained
$\ln$ natural logarithm ln(1/400) = −5.99
$\frac{1}{n-1}\sum$ the average over all $n - 1$ guesses
$\mathcal{L}$ the loss training pushes down 5.97

In words: "for every position, take the log of the probability the model gave the true next token, average them, and flip the sign."

With the numbers: a uniform guess over 400 tokens gives every true token p = 1/400, so each term is −ln(1/400) = ln 400 = 5.99. Our untrained model scores 5.97, so it is still essentially guessing. Perplexity, e raised to the loss, is 393: "as unsure as choosing among 393 equally likely tokens" (next_token_loss, cross_entropy).

Level 3: in Python

In Python:

import math
# one guess that gave the true token p = 0.5
round(-math.log(0.5), 3)  # → 0.693
n = 12
# a uniform guess gives every true token 1/400
p = [1 / 400] * (n - 1)
# -(1/(n-1)) Σ ln p
L = -sum(math.log(p_i) for p_i in p) / (n - 1)
round(L, 2)  # → 5.99

Why it matters. Pretraining is exactly this, over trillions of tokens: predicting the next token well forces the model to absorb grammar, facts and reasoning patterns. Random weights produce gibberish, and training is what turns the same machinery into a useful model.

In 20 seconds

  • Tokenize the prompt, look up a vector per token, add position, run the transformer blocks, score every vocabulary entry from the last position, softmax, sample, append, repeat.
  • Temperature divides the scores before softmax: low is predictable, high is varied.
  • The naive loop re-reads everything each step; the KV cache makes each step cost one token. Training is the same forward pass plus a next-token loss.

Self-test questions

What happens, step by step, when you send a prompt? Tokenizer turns text into ids; each id looks up an embedding; position information is added; the vectors pass through N transformer blocks; the last position's vector is scored against the whole vocabulary; softmax and sampling pick a token; it's appended and the loop repeats until a stop token.

Why does only the last position matter when generating? Its vector has attended to every earlier token and is the one trained to predict what comes next; earlier rows predict tokens we already have.

What does temperature do to the scores (2, 1, 0.5) at T = 0.5? Doubles them to (4, 2, 1) before softmax, sharpening the distribution: the top token rises from 0.63 to 0.84.

How many token positions does the naive loop process for a 2-token prompt and 4 new tokens? 2 + 3 + 4 + 5 = 14. A KV cache avoids re-reading the unchanged prefix.

What loss does an untrained model get over a 400-token vocabulary, and why? About ln 400 ≈ 5.99 (perplexity about 400), because with random, tiny weights its predictions are close to uniform.

How is training different from inference? Same forward pass, plus a loss comparing predictions to the real next tokens, then backpropagation and a weight update. Inference only runs the forward pass.

The papers behind this lesson

Further reading

on GitHub
  1r"""
  2# The big picture: what happens when you send a prompt
  3
  4Run: `python -m primer.ml.big_picture`
  5
  6New to the notation (vectors, sums, logarithms)? Every symbol is decoded
  7where it appears, and `primer.notation` teaches them all from zero.
  8
  9## Level 1: The practitioner's guide
 10
 11**In one sentence.** A language model is a next-token scorer run in a loop:
 12the prompt is cut into tokens, each token becomes a vector, the vectors
 13pass through a stack of transformer blocks, the last position is scored
 14against the whole vocabulary, one token is drawn, appended, and the loop
 15runs again until a stop token or a length limit.
 16
 17**When you need it.** You need this picture whenever you set a parameter
 18you can't explain (`temperature`, `top_p`, `max_tokens`), read a bill that
 19counts input and output tokens differently, or debug an answer that was cut
 20off, repeats itself, or comes out as gibberish. The tell: you are adjusting
 21a sampling knob by trial and error, or estimating cost in words when the
 22meter counts tokens. This lesson's toy tokenizer turns "Reset your password"
 23into 5 tokens, and " password" is one of them while "Reset" is two, so a
 24word count is only an estimate. You don't need this lesson to write a good
 25prompt, and you don't need it to pick a model by its benchmark scores; you
 26need it the first time a model's behaviour has to be predicted rather than
 27observed.
 28
 29**Your options.** The loop has a few knobs a caller can turn, from the
 30cheapest to the most certain:
 31
 32| Option | What it does | What it guarantees | What it costs | Where it lives |
 33|---|---|---|---|---|
 34| Leave sampling at its defaults | Draws each token in proportion to the model's probabilities (temperature 1) | The variety the model was trained to produce | Run-to-run variation | The API's defaults |
 35| Lower the temperature, down to greedy | Divides the scores before softmax; at 0 it takes the single most likely token | The likely answer more often: at temperature 0.5 this lesson's top token rises from 0.63 to 0.84 | Blander text, and repetition at 0 (Holtzman et al., 2019); still not identical runs | One request parameter |
 36| Cut the tail (top-k, top-p) | Keeps only the k most likely tokens, or the smallest set whose probabilities reach p | No draws from the long tail of near-zero tokens | A knob some hosted APIs have withdrawn; open-model servers keep it | One request parameter |
 37| Bound the length (`max_tokens`, stop sequences) | Ends the loop at a token limit or at a string you name | A ceiling on cost and latency per call | Truncated answers if the limit is too low; check the stop reason | One request parameter |
 38| Reuse the prefix (KV cache, prompt caching) | Keeps the work done on tokens that haven't changed | Each new token costs one position instead of a re-read of everything: 219 positions against 23,900 for 200 tokens after a 20-token prompt | Server memory; caching rules and prices that differ by vendor | The serving stack |
 39| Constrain the output (schema, grammar) | Forbids, at every step, any token that cannot lead to a valid shape | A parseable answer by construction | A compiled schema and a small check per token | The model server (`primer.ml.structured_output`) |
 40
 41**How to choose.** Start from who reads the answer and how long it is.
 42
 43- Extraction, classification, tool calls: low temperature, a schema where
 44  the server offers one, and a `max_tokens` sized to the answer.
 45- Writing, brainstorming, dialogue: the default temperature, and a length
 46  bound that stops runaway output rather than shaping it.
 47- Long documents in the prompt: the input is read in one parallel pass, so
 48  it costs money more than time; cache the unchanged prefix when you send
 49  it repeatedly.
 50- Agents and multi-turn systems: every turn re-sends the whole history, so
 51  the prompt grows with the conversation; bound each answer and watch the
 52  stop reason, because a truncated tool call is a broken one.
 53- Whatever you pick, measure tokens in and out on your own traffic. Output
 54  tokens are produced one loop iteration at a time, so the cheapest way to
 55  make a call faster is to ask for less.
 56
 57**What it costs.** Two meters run. Input tokens go through the model in
 58one parallel pass, so a long prompt costs money and memory more than
 59time. Output tokens are produced one per trip round the loop, so latency is
 60proportional to length: this lesson's naive loop reads the prompt plus
 61everything generated so far at every step (14 positions to produce 4 tokens
 62from a 2-token prompt), and a KV cache turns that into one position per
 63token. Context is a hard ceiling: the position table has a fixed number of
 64rows (64 in this toy; `max_position_embeddings` in a Hugging Face config,
 65where `LlamaConfig` defaults to 2048, and 4k in the Llama 2 paper), and
 66prompt plus answer must fit inside it. Quality is measured on the same loop: the training loss is
 67the average of −ln p for the real next token, and perplexity is *e* raised
 68to it. This lesson's untrained model scores 5.97 against 5.99 for a blind
 69guess over 400 tokens (perplexity 393), and training over trillions of
 70tokens (2.0T for Llama 2) is what pushes that number down.
 71
 72**What breaks.**
 73
 74- **Gibberish.** Random weights give a nearly flat guess over the
 75  vocabulary, and the demo's untrained model continues "The cat" with byte
 76  fragments. In practice the same symptom comes from a tokenizer that
 77  doesn't match the model, so ids fetch the wrong rows of the table. Check
 78  the pairing before the weights.
 79- **Cut-off answers.** The loop stopped at `max_tokens`, not at a stop
 80  token. Claude's API reports `end_turn` when the model finished on its own
 81  and another stop reason when your limit or stop sequence ended it; treat
 82  anything but a natural end as incomplete.
 83- **Temperature 0 that still varies.** Greedy picks the argmax, but Claude's
 84  reference says results are not fully deterministic even at 0.0, and the
 85  arithmetic on real serving hardware is why.
 86- **Repetition.** Always taking the most likely token produces loops of the
 87  same phrase; Holtzman et al. (2019) showed it and proposed nucleus
 88  sampling as the fix.
 89- **A bill that surprised you.** Cost is counted in tokens, not words, and
 90  a conversation re-sends its whole history every turn.
 91- **A context error.** Prompt plus requested output exceeded the position
 92  table. Shorten the prompt or the `max_tokens`, or retrieve less.
 93
 94**In the wild.** The pipeline is the decoder-only recipe of GPT-2 (Radford
 95et al., 2019: byte-level BPE, learned positions, tied embeddings) that
 96GPT-3 (Brown et al., 2020) scaled until instructions in the prompt were
 97enough. Hugging Face's `generate()` exposes the loop's knobs in a
 98`GenerationConfig` (`do_sample`, otherwise greedy; `temperature` 1.0,
 99`top_k` 50, `top_p` 1.0, `max_new_tokens`, `repetition_penalty`,
100`num_beams`), and its `LlamaForCausalLM` returns logits of shape (batch,
101sequence, vocabulary) with a `logits_to_keep` option because, as its docs
102put it, only the last token's logits are needed for generation. Claude's
103Messages API offers `temperature` (0.0 to 1.0, default 1.0; models
104released after Claude Opus 4.6 accept only 1.0), `max_tokens`,
105`stop_sequences` and a `stop_reason` on every response, and lets you set
106`max_tokens` to 0 to warm the prompt cache without generating. Karpathy's
107nanoGPT is this lesson at full size, in a few hundred lines. The papers are
108linked at the end of the lesson.
109
110**Go deeper.** Level 2 traces "Reset your password" through every box with
111its shapes, works temperature by hand on three scores, counts the loop's
112positions, and measures the loss of a model that has learned nothing. If you
113only needed to set the knobs, you are done.
114
115## Level 2: How it works, from scratch
116
117This lesson wires the real pieces from the other lessons into one working
118pipeline: the tokenizer from `primer.ml.tokenization`, the transformer from
119`primer.ml.transformer`, and a sampling loop. Keep this one picture in your
120head; every other lesson zooms into one box of it.
121
122```mermaid
123flowchart LR
124  A[Prompt text] --> B[Tokenizer<br/>text to IDs]
125  B --> C[Embedding lookup<br/>IDs to vectors]
126  C --> D[Add position info]
127  D --> E[Transformer blocks<br/>repeated N times]
128  E --> F[Output layer<br/>score per vocab token]
129  F --> G[Softmax + sampling]
130  G --> H[Next token]
131  H -->|append and repeat| E
132```
133
134**Reading it:** read left to right, then follow the loop back. The prompt is
135cut into token ids, each id picks a vector from a table, and position
136information is mixed in so order counts. The vectors pass through the
137transformer blocks, where tokens exchange information. The output layer turns
138the *last* token's final vector into a score for every token in the
139vocabulary; softmax makes those scores probabilities and one token is
140drawn. That token is appended and the loop runs again, until the model emits
141a stop token or hits a length limit. Training uses the same boxes with one
142addition, a loss, at the end (section 6).
143
144## 1. Text to ids: the coat check
145
146**Everyday picture.** A coat check. You hand over a coat (a piece of text)
147and get back a numbered ticket (a token id). The model only ever handles the
148tickets.
149
150**Tiny worked example.** With this module's toy tokenizer, "Reset your
151password" becomes 5 tickets: `Re` `set` ` y` `our` ` password` →
152**[346, 377, 309, 328, 291]**. The common word " password" got a single
153ticket; the rarer pieces got several.
154
155```mermaid
156flowchart LR
157  T["'Reset your password'"] --> TK["tokenizer<br/>(learned kit of 400 pieces)"]
158  TK --> I["[346, 377, 309, 328, 291]"]
159```
160
161**Reading it:** one box, text in and integers out. Everything to the right
162of this box works only with integers and vectors.
163
164**The code.** `tok.encode(prompt)`; how the kit is learned is the whole of
165`primer.ml.tokenization`.
166
167**In code:** `build_pipeline` trains the toy tokenizer (`primer.ml.tokenization.ByteBPE`) and builds an untrained `primer.ml.transformer.TinyGPT` sized to its vocabulary.
168
169**Why it matters.** Prompt length, price and context limits are all counted
170in these tickets.
171
172## 2. Ids to vectors: a lookup, not a computation
173
174**Everyday picture.** A dictionary where ticket number 291 opens to page 291,
175and each page holds a list of numbers describing that token. A second
176dictionary, indexed by seat number, describes *where* the token sits.
177
178**Tiny worked example.** The table has 400 rows (one per token) of 32
179numbers. Id 291 fetches row 291. The 5 ids fetch 5 rows: a 5 × 32 grid. Row
180i of the position table (i = 0 to 4) is added to row i of that grid.
181
182```mermaid
183flowchart LR
184  I["ids (5)"] --> E["token table<br/>400 × 32"]
185  E --> X["5 × 32: what each token is"]
186  P["positions 0..4"] --> PT["position table<br/>64 × 32"]
187  PT --> Y["5 × 32: where each token is"]
188  X --> ADD(("+"))
189  Y --> ADD
190  ADD --> OUT["5 × 32 input to the blocks"]
191```
192
193**Reading it:** two lookups, one add. Nothing is multiplied here, rows are
194simply fetched, which is why this step is nearly free. The tables themselves
195are learned during training, so similar tokens end up with similar rows (see
196`primer.ml.embeddings`).
197
198**The math and the code.**
199
200$$
201x_i = E_{t_i} + P_i
202$$
203
204**Symbols**
205
206| Symbol | Meaning here | In the example |
207|---|---|---|
208| $i$ | a token's position in the prompt, from 0 | 4 (the last token) |
209| $t_i$ | the token id at position $i$ | $t_4$ = 291 |
210| $E$ | the token embedding table | 400 × 32 |
211| $E_{t_i}$ | row $t_i$ of that table | row 291: 32 numbers |
212| $P_i$ | row $i$ of the position table | row 4: 32 numbers |
213| $x_i$ | what the first block receives for position $i$ | 32 numbers |
214
215**In words:** "each token's input vector is its token's row plus its
216position's row."
217
218**With the numbers:** $x_4 = E_{291} + P_4$, one 32-number list plus
219another. The code is `model.wte[ids] + model.wpe[:len(ids)]`. In miniature,
220with 3-number rows instead of 32: $E_{291}$ = (0.2, −0.1, 0.5) and $P_4$ =
221(0.1, 0.3, −0.2) give $x_4$ = (0.3, 0.2, 0.3).
222
223**In Python:**
224
225```python
226# row 291 of the token table (3 numbers, not 32)
227E_291 = [0.2, -0.1, 0.5]
228# row 4 of the position table
229P_4 = [0.1, 0.3, -0.2]
230# E_(t_i) + P_i, number by number
231x_4 = [round(e + p, 2) for e, p in zip(E_291, P_4)]
232x_4  # → [0.3, 0.2, 0.3]
233```
234
235**In code:** `trace` does both lookups and the add, and keeps every stage's result so you can inspect the grid before and after positions are mixed in.
236
237**Why it matters.** This is the only place a token's identity enters the
238model; every later step works on these vectors.
239
240## 3. The transformer blocks: rounds of meeting and desk work
241
242**Everyday picture.** The team from `primer.ml.transformer`: each round is a
243meeting where every token listens to the others (attention), then desk work
244where each token thinks alone (feed-forward). This toy runs 2 rounds; large
245models run dozens.
246
247**Tiny worked example.** The 5 × 32 grid goes into block 1 and comes out
2485 × 32; the same through block 2. By the end, the vector at the last position
249("password") has absorbed information from "Re", "set", "y" and "our".
250
251```mermaid
252flowchart LR
253  X["5 × 32"] --> B1["block 1<br/>meeting + desk work"] --> B2["block 2"] --> LN["final norm"] --> H["5 × 32<br/>context-aware vectors"]
254```
255
256**Reading it:** the shape never changes, only the contents. Each block
257edits every token's vector by adding what it learned from the others.
258
259**The math and the code.**
260
261$$
262h = \text{LN}\big(\text{Block}_N(\cdots\text{Block}_2(\text{Block}_1(x))\cdots)\big)
263$$
264
265**Symbols**
266
267| Symbol | Meaning here | In the example |
268|---|---|---|
269| $x$ | the 5 × 32 input grid from section 2 | |
270| $\text{Block}_k$ | the $k$-th transformer block | $N$ = 2 |
271| $\cdots$ | "and so on, for every block in between" | |
272| $\text{LN}$ | the final layer norm | |
273| $h$ | the context-aware vectors | 5 × 32 |
274
275**In words:** "run the input through every block in turn, then normalize."
276
277**With the numbers:** here N = 2, so h = LN(Block₂(Block₁(x))). The code is
278the `for block in model.blocks` loop in `trace`. In miniature, with one
2793-number vector and two stand-in blocks that each add an edit: x = (1, 2, 3),
280Block₁ adds (0, 1, 0) and Block₂ adds (0, 0, 2), giving (1, 3, 5). The final
281norm subtracts the mean (3) and divides by the spread (1.63), so
282h = (−1.22, 0, 1.22).
283
284**In Python:**
285
286```python
287import statistics
288x = [1.0, 2.0, 3.0]
289# a stand-in edit
290def block_1(v): return [v_i + d for v_i, d in zip(v, [0, 1, 0])]
291def block_2(v): return [v_i + d for v_i, d in zip(v, [0, 0, 2])]
292def LN(v):
293    mu, sigma = statistics.fmean(v), statistics.pstdev(v)
294    # centre, then rescale
295    return [round((v_i - mu) / sigma, 2) for v_i in v]
296h = x
297# Block_1 first, then Block_2 ... up to Block_N
298for block in [block_1, block_2]:
299    h = block(h)
300h  # → [1.0, 3.0, 5.0]
301LN(h)  # → [-1.22, 0.0, 1.22]
302```
303
304**Why it matters.** This is where nearly all the compute and all the
305"understanding" happen.
306
307## 4. Scores, softmax and temperature
308
309**Everyday picture.** A scoreboard with one line for every token in the
310vocabulary. Softmax turns the scores into shares of a pie. **Temperature**
311sets how adventurous the pick is: low temperature almost always takes the
312biggest slice; high temperature gives the small slices a real chance.
313
314**Tiny worked example.** Three candidate tokens scored 2.0, 1.0 and 0.5.
315
316| temperature | p(A) | p(B) | p(C) |
317|---|---|---|---|
318| 0 (greedy) | 1.000 | 0.000 | 0.000 |
319| 0.5 | 0.844 | 0.114 | 0.042 |
320| 1 | 0.629 | 0.231 | 0.140 |
321| 2 | 0.481 | 0.292 | 0.227 |
322
323```mermaid
324flowchart LR
325  H["last row of h<br/>32 numbers"] --> S["× token tableᵀ<br/>400 scores"]
326  S --> T["÷ temperature"]
327  T --> SM["softmax<br/>400 probabilities"]
328  SM --> PICK["draw one token"]
329```
330
331**Reading it:** only the last position's vector is used to choose the next
332token. It is scored against every token's embedding, the scores are divided
333by the temperature, softmax turns them into probabilities that sum to 1, and
334one token is drawn at random in proportion to them.
335
336**The math and the code.** **Softmax** raises *e* (≈ 2.718) to the power of
337each score, so bigger scores get disproportionately bigger shares, then
338divides by the total so the shares add up to 1:
339
340$$
341p_i = \frac{e^{z_i / T}}{\sum_{j=1}^{V} e^{z_j / T}}
342$$
343
344**Symbols**
345
346| Symbol | Meaning here | In the example |
347|---|---|---|
348| $z_i$ | the score of candidate token $i$ | $z_A$ = 2.0 |
349| $T$ | the temperature | 0.5 |
350| $e^{\cdot}$ | *e* ≈ 2.718 raised to that power | $e^{4}$ = 54.6 |
351| $V$ | number of candidates (the vocabulary size) | 3 here, 400 in the model |
352| $\sum_{j=1}^{V}$ | add up over every candidate $j$ | |
353| $p_i$ | probability that token $i$ comes next | 0.844 |
354
355**In words:** "the chance of token *i* is *e* to the power of its score over
356the temperature, divided by the same quantity summed over all tokens."
357
358**With the numbers:** T = 0.5 doubles every score to (4, 2, 1): e⁴ = 54.6,
359e² = 7.39, e¹ = 2.72, total 64.7, so p(A) = 54.6 / 64.7 = **0.844**
360(`next_token_probs`). Temperature 0 is the limit case: all probability on the
361top score.
362
363**In Python:**
364
365```python
366import math
367z = [2.0, 1.0, 0.5]
368T = 0.5
369# e^(z_i / T) for each candidate
370exps = [math.exp(z_i / T) for z_i in z]
371[round(e, 2) for e in exps]  # → [54.6, 7.39, 2.72]
372# Σ_j e^(z_j / T)
373total = sum(exps)
374round(total, 1)  # → 64.7
375# p_i: each share of the total
376[round(e / total, 3) for e in exps]  # → [0.844, 0.114, 0.042]
377```
378
379![At T = 0.25 token A takes 98% of the probability; at T = 2 the three tokens share it 48%, 29% and 23%](figures/primer.ml.big_picture.temperature.svg)
380
381**Reading it:** three groups of bars, one per candidate token; within each
382group the bars run from low temperature (left) to high (right). For token A
383the bars fall as temperature rises, and for tokens B and C they rise. At
384T = 0.25, A takes almost everything; at T = 2 the three are much closer.
385
386![The untrained model's 12 favourites are random byte fragments, each only 1.25 to 1.4 times the uniform 1-in-400 share](figures/primer.ml.big_picture.untrained_next_token.svg)
387
388**Reading it:** the bars are the 12 most probable next tokens after "The cat
389sat on the", according to our untrained model; the dashed line is a uniform
390guess, 1 in 400 (0.25%). Even these favourites clear the line only
391modestly: the top one gets about 0.35%, 1.4 times the uniform share, and the
392twelfth about 1.25 times. Across all 400 tokens every probability stays
393between 0.74 and 1.39 times uniform, and the "favourites" are random byte
394fragments. That is exactly what random weights should give: a nearly flat
395guess with small random bumps. The machinery works, but nothing has been
396learned yet.
397
398**Why it matters.** Use low temperature for extraction and tool calls,
399where you want the most likely answer, and higher temperature for creative
400writing. Temperature 0 reduces randomness but doesn't guarantee identical
401outputs on real serving hardware.
402
403## 5. The loop: append and repeat
404
405**Everyday picture.** Writing a sentence one word at a time, and re-reading
406everything you've written before choosing each new word.
407
408**Tiny worked example.** A 2-token prompt, 4 new tokens. Step 1 reads 2
409tokens, step 2 reads 3, step 3 reads 4, step 4 reads 5: **14** token
410positions to produce 4 tokens. With a 20-token prompt and 200 new tokens, the
411naive loop reads 23,900 positions; a KV cache reads 219.
412
413```mermaid
414flowchart LR
415  S["ids so far"] --> M["full forward pass<br/>over ALL ids"]
416  M --> P["probabilities for the next id"]
417  P --> D["draw one id"]
418  D --> A["append it"]
419  A -->|"not done"| S
420  A -->|"stop token or length limit"| OUT["decode ids to text"]
421```
422
423**Reading it:** the loop box is the whole of generation. The expensive
424arrow is "full forward pass over ALL ids": every step re-reads text that
425hasn't changed. Because of the causal mask, earlier tokens' internal vectors
426can't change when new tokens arrive, so that repeated work is pure waste.
427The KV cache (`primer.ml.inference`) stores it once.
428
429**The math.** Total positions processed without a cache:
430
431$$
432W = \sum_{t=0}^{n-1} (p + t) = n\,p + \frac{n(n-1)}{2}
433$$
434
435**Symbols**
436
437| Symbol | Meaning here | In the example |
438|---|---|---|
439| $p$ | prompt length in tokens | 2 |
440| $n$ | tokens to generate | 4 |
441| $t$ | the step counter, from 0 to $n-1$ | 0, 1, 2, 3 |
442| $p + t$ | tokens re-read at step $t$ | 2, 3, 4, 5 |
443| $W$ | total positions run through the model | 14 |
444
445**In words:** "each step re-reads the prompt plus everything generated so
446far; add that up over all steps."
447
448**With the numbers:** 4 × 2 + (4 × 3) / 2 = 8 + 6 = **14**, the count
449`generate` reports as `positions_processed`.
450
451**In Python:**
452
453```python
454p, n = 2, 4
455# Σ over t = 0 .. n-1 of (p + t): 2 + 3 + 4 + 5
456sum(p + t for t in range(n))  # → 14
457# the closed form gives the same count
458n * p + n * (n - 1) // 2  # → 14
459```
460
461![Without a cache work grows with the square of output length, 23,900 positions at 200 tokens against 219 with a KV cache](figures/primer.ml.big_picture.loop_cost.svg)
462
463**Reading it:** the x-axis is how many tokens have been generated after a
46420-token prompt; the y-axis is total work in token positions. The red curve
465bends upward: it grows with the square of the output length. The blue line
466(with a KV cache) grows by exactly one per token. At 200 tokens the gap is
467more than a hundredfold.
468
469**In code:** `Generation` holds the generated ids, the decoded text and the positions count; `naive_vs_cached_work` counts the positions processed with and without a KV cache for the figure.
470
471**Why it matters.** Output length drives latency and cost; this is why
472every serving system caches keys and values.
473
474## 6. Training: the same forward pass, plus a loss
475
476**Everyday picture.** A guessing game with instant feedback. Cover the next
477word, guess it, uncover it, and note how surprised you were. Training nudges
478every weight to make the surprise smaller next time.
479
480**Tiny worked example.** If the model gives the right next token probability
4810.5, the loss is −ln 0.5 = **0.693**. Our untrained model averages **5.97** on
482a real sentence, close to ln 400 = 5.99, the score of a blind uniform guess.
483
484```mermaid
485flowchart LR
486  T["training text"] --> F["forward pass<br/>(this whole lesson)"]
487  F --> P["probabilities at every position"]
488  P --> L["loss: −ln p(actual next token)<br/>averaged over positions"]
489  L --> B["backpropagation<br/>gradient for every weight"]
490  B --> U["optimizer nudges weights"]
491  U -->|next batch| F
492```
493
494**Reading it:** the first two boxes are the forward pass you just traced.
495Training adds the loss (how surprised the model was by the real next
496tokens), then backpropagation (`primer.ml.neural_net`) works out how each
497weight contributed, and the optimizer (`primer.ml.optimizers`) adjusts them.
498One pass over n tokens gives n − 1 guesses at once, in parallel, which is a
499big reason transformers train fast.
500
501**The math and the code.** A **logarithm** answers "to what power must I
502raise *e* to get this number?" For probabilities between 0 and 1 it is
503negative, so we flip the sign; ln 1 = 0 (no surprise) and ln of a tiny
504number is very negative (huge surprise).
505
506$$
507\mathcal{L} = -\frac{1}{n-1}\sum_{i=1}^{n-1} \ln p\big(t_{i+1} \mid t_1, \ldots, t_i\big)
508$$
509
510**Symbols**
511
512| Symbol | Meaning here | In the example |
513|---|---|---|
514| $n$ | tokens in the training text | 12 ("The model reads tokens, not words.") |
515| $t_i$ | the token at position $i$ | |
516| $p(t_{i+1} \mid t_1,\ldots,t_i)$ | probability the model gave the *actual* next token, having seen everything before it; the bar $\mid$ reads "given" | ≈ 1/400 when untrained |
517| $\ln$ | natural logarithm | ln(1/400) = −5.99 |
518| $\frac{1}{n-1}\sum$ | the average over all $n - 1$ guesses | |
519| $\mathcal{L}$ | the loss training pushes down | 5.97 |
520
521**In words:** "for every position, take the log of the probability the model
522gave the true next token, average them, and flip the sign."
523
524**With the numbers:** a uniform guess over 400 tokens gives every true token
525p = 1/400, so each term is −ln(1/400) = ln 400 = **5.99**. Our untrained model
526scores 5.97, so it is still essentially guessing. **Perplexity**, e raised to
527the loss, is 393: "as unsure as choosing among 393 equally likely tokens"
528(`next_token_loss`, `cross_entropy`).
529
530**In Python:**
531
532```python
533import math
534# one guess that gave the true token p = 0.5
535round(-math.log(0.5), 3)  # → 0.693
536n = 12
537# a uniform guess gives every true token 1/400
538p = [1 / 400] * (n - 1)
539# -(1/(n-1)) Σ ln p
540L = -sum(math.log(p_i) for p_i in p) / (n - 1)
541round(L, 2)  # → 5.99
542```
543
544**Why it matters.** Pretraining is exactly this, over trillions of tokens:
545predicting the next token well forces the model to absorb grammar, facts and
546reasoning patterns. Random weights produce gibberish, and training is what
547turns the same machinery into a useful model.
548
549## In 20 seconds
550- Tokenize the prompt, look up a vector per token, add position, run the
551  transformer blocks, score every vocabulary entry from the last position,
552  softmax, sample, append, repeat.
553- Temperature divides the scores before softmax: low is predictable, high is
554  varied.
555- The naive loop re-reads everything each step; the KV cache makes each step
556  cost one token. Training is the same forward pass plus a next-token loss.
557
558## Self-test questions
559
560**What happens, step by step, when you send a prompt?**
561Tokenizer turns text into ids; each id looks up an embedding; position
562information is added; the vectors pass through N transformer blocks; the last
563position's vector is scored against the whole vocabulary; softmax and sampling
564pick a token; it's appended and the loop repeats until a stop token.
565
566**Why does only the last position matter when generating?**
567Its vector has attended to every earlier token and is the one trained to
568predict what comes next; earlier rows predict tokens we already have.
569
570**What does temperature do to the scores (2, 1, 0.5) at T = 0.5?**
571Doubles them to (4, 2, 1) before softmax, sharpening the distribution: the
572top token rises from 0.63 to 0.84.
573
574**How many token positions does the naive loop process for a 2-token prompt and 4 new tokens?**
5752 + 3 + 4 + 5 = 14. A KV cache avoids re-reading the unchanged prefix.
576
577**What loss does an untrained model get over a 400-token vocabulary, and why?**
578About ln 400 ≈ 5.99 (perplexity about 400), because with random, tiny weights
579its predictions are close to uniform.
580
581**How is training different from inference?**
582Same forward pass, plus a loss comparing predictions to the real next tokens,
583then backpropagation and a weight update. Inference only runs the forward pass.
584
585## The papers behind this lesson
586
587- **Vaswani et al. (2017), *Attention Is All You Need*.** https://arxiv.org/abs/1706.03762.
588  The architecture inside the "transformer blocks" box.
589  [annotated companion](../../papers/attention-is-all-you-need.html)
590- **Radford et al. (2019), *Language Models are Unsupervised Multitask Learners* (GPT-2).**
591  https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf.
592  The decoder-only, next-token-prediction recipe this pipeline follows,
593  including learned positions, byte-level BPE and tied embeddings.
594- **Brown et al. (2020), *Language Models are Few-Shot Learners* (GPT-3).**
595  https://arxiv.org/abs/2005.14165. Showed that scaling this same loop up
596  produces models that follow instructions from examples in the prompt.
597  [annotated companion](../../papers/gpt-3.html)
598- **Holtzman et al. (2019), *The Curious Case of Neural Text Degeneration*.**
599  https://arxiv.org/abs/1904.09751. Introduced nucleus (top-p) sampling and
600  explained why pure greedy decoding produces repetitive text.
601
602## Further reading
603- Andrej Karpathy, *Let's build GPT* (video): https://www.youtube.com/watch?v=kCc8FmEb1nY
604- Karpathy's `nanoGPT`: https://github.com/karpathy/nanoGPT
605- Jay Alammar, *The Illustrated GPT-2*: https://jalammar.github.io/illustrated-gpt2/
606- 3Blue1Brown, *But what is a GPT?* (video): https://www.youtube.com/watch?v=wjZofJX0v4M
607- Hugging Face, *How to generate text* (decoding strategies): https://huggingface.co/blog/how-to-generate
608"""
609
610from __future__ import annotations
611
612from dataclasses import dataclass
613
614import numpy as np
615
616from primer._show import banner, say, table, takeaway
617from primer.ml.tokenization import ByteBPE, trained_tokenizer
618from primer.ml.transformer import TinyGPT, layer_norm
619
620# Model size for the walkthrough: small enough to run instantly in NumPy.
621D_MODEL, N_LAYERS, N_HEADS, MAX_LEN = 32, 2, 4, 64
622
623
624def build_pipeline(seed: int = 0) -> tuple[ByteBPE, TinyGPT]:
625    """A trained toy tokenizer plus an *untrained* TinyGPT sized to its vocabulary."""
626    tok = trained_tokenizer()
627    model = TinyGPT(tok.vocab_size, D_MODEL, N_LAYERS, N_HEADS, MAX_LEN, seed=seed)
628    return tok, model
629
630
631# ---------------------------------------------------------------------------
632# 1. One forward pass, stage by stage
633# ---------------------------------------------------------------------------
634
635
636def next_token_probs(logits: np.ndarray, temperature: float = 1.0) -> np.ndarray:
637    """Turn one row of scores into next-token probabilities.
638
639    Temperature divides the scores before softmax: below 1 sharpens the
640    distribution (more predictable), above 1 flattens it (more varied).
641    Temperature 0 is the limit: all probability on the top score (greedy).
642    """
643    if temperature == 0:
644        p = np.zeros_like(logits, dtype=float)
645        p[np.argmax(logits)] = 1.0
646        return p
647    z = logits / temperature
648    e = np.exp(z - z.max())  # subtract the max so exp never overflows
649    return e / e.sum()
650
651
652def trace(prompt: str, tok: ByteBPE, model: TinyGPT) -> dict[str, np.ndarray]:
653    """Run one prompt through every stage and keep each intermediate result."""
654    ids = np.array(tok.encode(prompt))  # (n,)       text -> integers
655    emb = model.wte[ids]  # (n, d)     integers -> vectors (a table lookup)
656    x = emb + model.wpe[: len(ids)]  # (n, d)     add "where" to "what"
657    for block in model.blocks:  # (n, d)     N rounds of meeting + desk work
658        x = block(x)
659    h = layer_norm(x, model.lnf_g, model.lnf_b)
660    logits = h @ model.wte.T  # (n, vocab) a score for every possible next token
661    return {
662        "ids": ids,
663        "embeddings": emb,
664        "with_positions": emb + model.wpe[: len(ids)],
665        "after_blocks": x,
666        "logits": logits,
667        "next_token_probs": next_token_probs(logits[-1]),  # only the last row predicts what comes next
668    }
669
670
671# ---------------------------------------------------------------------------
672# 2. The generation loop: predict, sample, append, repeat
673# ---------------------------------------------------------------------------
674
675
676@dataclass
677class Generation:
678    ids: list[int]
679    text: str
680    positions_processed: int  # total token positions run through the model
681
682
683def generate(
684    prompt: str, n_new: int, tok: ByteBPE, model: TinyGPT, temperature: float = 1.0, seed: int = 0
685) -> Generation:
686    """The naive autoregressive loop: re-run the whole sequence for every new token.
687
688    Step t processes all (prompt + t) tokens again, so total work grows with
689    the square of the output length. The KV cache (primer.ml.inference)
690    stores each token's keys and values so every step only processes the one
691    new token. A real model also stops at an end-of-text token; this toy just
692    stops after n_new.
693    """
694    rng = np.random.default_rng(seed)
695    ids = tok.encode(prompt)
696    processed = 0
697    for _ in range(n_new):
698        logits = model(np.array(ids))
699        processed += len(ids)
700        p = next_token_probs(logits[-1], temperature)
701        ids.append(int(rng.choice(len(p), p=p)))
702    return Generation(ids=ids, text=tok.decode(ids), positions_processed=processed)
703
704
705# ---------------------------------------------------------------------------
706# 3. Training: the same forward pass, plus a loss
707# ---------------------------------------------------------------------------
708
709
710def cross_entropy(probs: np.ndarray, target: int) -> float:
711    """−ln(probability given to the right answer). Big when confidently wrong."""
712    return float(-np.log(probs[target]))
713
714
715def next_token_loss(text: str, tok: ByteBPE, model: TinyGPT) -> float:
716    """Average next-token cross-entropy over `text`, the number training pushes down.
717
718    Every position predicts the token after it, so one forward pass over n
719    tokens gives n-1 training examples at once, in parallel. That is a big
720    reason transformers train so much faster than RNNs.
721    """
722    ids = tok.encode(text)
723    logits = model(np.array(ids))
724    losses = [cross_entropy(next_token_probs(logits[i]), ids[i + 1]) for i in range(len(ids) - 1)]
725    return float(np.mean(losses))
726
727
728# ---------------------------------------------------------------------------
729# 4. Figures (rendered to docs/figures by `make figures`)
730# ---------------------------------------------------------------------------
731
732TEMPERATURE_SCORES = np.array([2.0, 1.0, 0.5])
733
734
735def naive_vs_cached_work(prompt_len: int, max_new: int) -> list[tuple[int, int, int]]:
736    """(new tokens, positions processed without a cache, with a KV cache).
737
738    Without a cache step t re-runs prompt_len + t positions. With a cache the
739    first step runs the prompt once (prefill), and every later step runs just
740    the one new token.
741    """
742    rows = []
743    for n in range(1, max_new + 1):
744        naive = sum(prompt_len + t for t in range(n))
745        cached = prompt_len + (n - 1)
746        rows.append((n, naive, cached))
747    return rows
748
749
750def figures() -> dict:
751    """Plot this lesson's data. matplotlib is imported here, and only here."""
752    import matplotlib
753
754    matplotlib.use("Agg")
755    import matplotlib.pyplot as plt
756
757    BLUE, RED, MUTED = "#2563eb", "#dc2626", "#9ca3af"
758    figs = {}
759    tok, model = build_pipeline()
760
761    p = trace("The cat sat on the", tok, model)["next_token_probs"]
762    top = np.argsort(-p)[:12]
763    fig, ax = plt.subplots(figsize=(6.5, 3.6))
764    ax.bar([repr(tok.decode([int(i)])) for i in top], p[top], color=BLUE)
765    ax.axhline(1 / len(p), color=RED, ls="--", label=f"uniform guess 1/{len(p)}")
766    ax.set_ylabel("probability of being next")
767    ax.set_title("Untrained model: 'The cat sat on the' → ?")
768    ax.tick_params(axis="x", rotation=45)
769    ax.legend(frameon=False)
770    fig.tight_layout()
771    figs["untrained_next_token"] = fig
772
773    temps = (0.25, 0.5, 1.0, 2.0)
774    fig, ax = plt.subplots(figsize=(6, 3.6))
775    width = 0.2
776    for j, T in enumerate(temps):
777        ax.bar(np.arange(3) + (j - 1.5) * width, next_token_probs(TEMPERATURE_SCORES, T), width, label=f"T = {T}")
778    ax.set_xticks(range(3), ["token A (score 2.0)", "token B (1.0)", "token C (0.5)"])
779    ax.set_ylabel("probability")
780    ax.set_title("Temperature: low sharpens, high flattens")
781    ax.legend(frameon=False)
782    fig.tight_layout()
783    figs["temperature"] = fig
784
785    rows = naive_vs_cached_work(prompt_len=20, max_new=200)
786    fig, ax = plt.subplots(figsize=(6, 3.6))
787    ax.plot([r[0] for r in rows], [r[1] for r in rows], color=RED, label="no cache: re-read everything")
788    ax.plot([r[0] for r in rows], [r[2] for r in rows], color=BLUE, label="KV cache: one new token per step")
789    ax.set_xlabel("tokens generated (after a 20-token prompt)")
790    ax.set_ylabel("token positions run through the model")
791    ax.set_title("Why the naive loop is too slow")
792    ax.legend(frameon=False)
793    fig.tight_layout()
794    figs["loop_cost"] = fig
795    return figs
796
797
798# ---------------------------------------------------------------------------
799# 5. Narrated walkthrough
800# ---------------------------------------------------------------------------
801
802
803def demo() -> None:
804    tok, model = build_pipeline()
805    prompt = "Reset your password"
806
807    banner("1. Text -> ids -> vectors -> blocks -> scores")
808    st = trace(prompt, tok, model)
809    table(
810        ["stage", "shape", "what it holds"],
811        [
812            ("ids", st["ids"].shape, f"{st['ids'].tolist()} = {tok.tokens(prompt)}"),
813            ("embeddings", st["embeddings"].shape, "one row of the token table per id"),
814            ("with positions", st["with_positions"].shape, "plus one row of the position table"),
815            ("after blocks", st["after_blocks"].shape, f"{N_LAYERS} rounds of attention + feed-forward"),
816            ("logits", st["logits"].shape, f"a score for all {tok.vocab_size} tokens, per position"),
817            ("next-token probs", st["next_token_probs"].shape, "softmax of the LAST row only"),
818        ],
819    )
820    takeaway("Only the last position's scores decide the next token; the rest matter during training.")
821
822    banner("2. Temperature on the scores (2.0, 1.0, 0.5)")
823    table(
824        ["temperature", "p(A)", "p(B)", "p(C)"],
825        [(T, *next_token_probs(TEMPERATURE_SCORES, T)) for T in (0.0, 0.5, 1.0, 2.0)],
826        floatfmt=".3f",
827    )
828
829    banner("3. The loop: predict, sample, append, repeat")
830    g = generate("The cat", 12, tok, model, seed=0)
831    say(
832        f"""
833        Continuation of 'The cat': {g.text!r}. It's gibberish, and it should be:
834        the weights are random, so every next token is close to a uniform guess
835        over {tok.vocab_size} tokens. The machinery is identical to a real model's;
836        only the numbers inside the matrices are missing. They come from training.
837        This loop ran {g.positions_processed} token positions to produce 12 tokens,
838        because it re-reads the whole sequence every step (see primer.ml.inference
839        for the KV cache).
840        """
841    )
842
843    banner("4. Training = the same forward pass + a loss")
844    loss = next_token_loss("The model reads tokens, not words.", tok, model)
845    say(
846        f"""
847        Average next-token loss of the untrained model: {loss:.2f}. A uniform guess
848        over {tok.vocab_size} tokens scores ln {tok.vocab_size} = {np.log(tok.vocab_size):.2f}.
849        Perplexity e^loss = {np.exp(loss):.0f}: the model is as unsure as picking among
850        ~{np.exp(loss):.0f} equally likely tokens. Training nudges every weight to push
851        this number down, across trillions of tokens.
852        """
853    )
854    takeaway("A language model is a next-token scorer run in a loop; training makes its scores good.")
855
856
857if __name__ == "__main__":
858    demo()
Level 3: the code, function by function.
625def build_pipeline(seed: int = 0) -> tuple[ByteBPE, TinyGPT]:
626    """A trained toy tokenizer plus an *untrained* TinyGPT sized to its vocabulary."""
627    tok = trained_tokenizer()
628    model = TinyGPT(tok.vocab_size, D_MODEL, N_LAYERS, N_HEADS, MAX_LEN, seed=seed)
629    return tok, model

A trained toy tokenizer plus an untrained TinyGPT sized to its vocabulary.

def next_token_probs(logits: numpy.ndarray, temperature: float = 1.0) -> numpy.ndarray: on GitHub
637def next_token_probs(logits: np.ndarray, temperature: float = 1.0) -> np.ndarray:
638    """Turn one row of scores into next-token probabilities.
639
640    Temperature divides the scores before softmax: below 1 sharpens the
641    distribution (more predictable), above 1 flattens it (more varied).
642    Temperature 0 is the limit: all probability on the top score (greedy).
643    """
644    if temperature == 0:
645        p = np.zeros_like(logits, dtype=float)
646        p[np.argmax(logits)] = 1.0
647        return p
648    z = logits / temperature
649    e = np.exp(z - z.max())  # subtract the max so exp never overflows
650    return e / e.sum()

Turn one row of scores into next-token probabilities.

Temperature divides the scores before softmax: below 1 sharpens the distribution (more predictable), above 1 flattens it (more varied). Temperature 0 is the limit: all probability on the top score (greedy).

def trace( prompt: str, tok: primer.ml.tokenization.ByteBPE, model: primer.ml.transformer.TinyGPT) -> dict[str, numpy.ndarray]: on GitHub
653def trace(prompt: str, tok: ByteBPE, model: TinyGPT) -> dict[str, np.ndarray]:
654    """Run one prompt through every stage and keep each intermediate result."""
655    ids = np.array(tok.encode(prompt))  # (n,)       text -> integers
656    emb = model.wte[ids]  # (n, d)     integers -> vectors (a table lookup)
657    x = emb + model.wpe[: len(ids)]  # (n, d)     add "where" to "what"
658    for block in model.blocks:  # (n, d)     N rounds of meeting + desk work
659        x = block(x)
660    h = layer_norm(x, model.lnf_g, model.lnf_b)
661    logits = h @ model.wte.T  # (n, vocab) a score for every possible next token
662    return {
663        "ids": ids,
664        "embeddings": emb,
665        "with_positions": emb + model.wpe[: len(ids)],
666        "after_blocks": x,
667        "logits": logits,
668        "next_token_probs": next_token_probs(logits[-1]),  # only the last row predicts what comes next
669    }

Run one prompt through every stage and keep each intermediate result.

@dataclass
class Generation: on GitHub
677@dataclass
678class Generation:
679    ids: list[int]
680    text: str
681    positions_processed: int  # total token positions run through the model
Generation(ids: list[int], text: str, positions_processed: int)
ids: list[int]
text: str
def generate( prompt: str, n_new: int, tok: primer.ml.tokenization.ByteBPE, model: primer.ml.transformer.TinyGPT, temperature: float = 1.0, seed: int = 0) -> Generation: on GitHub
684def generate(
685    prompt: str, n_new: int, tok: ByteBPE, model: TinyGPT, temperature: float = 1.0, seed: int = 0
686) -> Generation:
687    """The naive autoregressive loop: re-run the whole sequence for every new token.
688
689    Step t processes all (prompt + t) tokens again, so total work grows with
690    the square of the output length. The KV cache (primer.ml.inference)
691    stores each token's keys and values so every step only processes the one
692    new token. A real model also stops at an end-of-text token; this toy just
693    stops after n_new.
694    """
695    rng = np.random.default_rng(seed)
696    ids = tok.encode(prompt)
697    processed = 0
698    for _ in range(n_new):
699        logits = model(np.array(ids))
700        processed += len(ids)
701        p = next_token_probs(logits[-1], temperature)
702        ids.append(int(rng.choice(len(p), p=p)))
703    return Generation(ids=ids, text=tok.decode(ids), positions_processed=processed)

The naive autoregressive loop: re-run the whole sequence for every new token.

Step t processes all (prompt + t) tokens again, so total work grows with the square of the output length. The KV cache (primer.ml.inference) stores each token's keys and values so every step only processes the one new token. A real model also stops at an end-of-text token; this toy just stops after n_new.

def cross_entropy(probs: numpy.ndarray, target: int) -> float: on GitHub
711def cross_entropy(probs: np.ndarray, target: int) -> float:
712    """−ln(probability given to the right answer). Big when confidently wrong."""
713    return float(-np.log(probs[target]))

−ln(probability given to the right answer). Big when confidently wrong.

def next_token_loss( text: str, tok: primer.ml.tokenization.ByteBPE, model: primer.ml.transformer.TinyGPT) -> float: on GitHub
716def next_token_loss(text: str, tok: ByteBPE, model: TinyGPT) -> float:
717    """Average next-token cross-entropy over `text`, the number training pushes down.
718
719    Every position predicts the token after it, so one forward pass over n
720    tokens gives n-1 training examples at once, in parallel. That is a big
721    reason transformers train so much faster than RNNs.
722    """
723    ids = tok.encode(text)
724    logits = model(np.array(ids))
725    losses = [cross_entropy(next_token_probs(logits[i]), ids[i + 1]) for i in range(len(ids) - 1)]
726    return float(np.mean(losses))

Average next-token cross-entropy over text, the number training pushes down.

Every position predicts the token after it, so one forward pass over n tokens gives n-1 training examples at once, in parallel. That is a big reason transformers train so much faster than RNNs.

TEMPERATURE_SCORES = array([2. , 1. , 0.5])
def naive_vs_cached_work(prompt_len: int, max_new: int) -> list[tuple[int, int, int]]: on GitHub
736def naive_vs_cached_work(prompt_len: int, max_new: int) -> list[tuple[int, int, int]]:
737    """(new tokens, positions processed without a cache, with a KV cache).
738
739    Without a cache step t re-runs prompt_len + t positions. With a cache the
740    first step runs the prompt once (prefill), and every later step runs just
741    the one new token.
742    """
743    rows = []
744    for n in range(1, max_new + 1):
745        naive = sum(prompt_len + t for t in range(n))
746        cached = prompt_len + (n - 1)
747        rows.append((n, naive, cached))
748    return rows

(new tokens, positions processed without a cache, with a KV cache).

Without a cache step t re-runs prompt_len + t positions. With a cache the first step runs the prompt once (prefill), and every later step runs just the one new token.

def figures() -> dict: on GitHub
751def figures() -> dict:
752    """Plot this lesson's data. matplotlib is imported here, and only here."""
753    import matplotlib
754
755    matplotlib.use("Agg")
756    import matplotlib.pyplot as plt
757
758    BLUE, RED, MUTED = "#2563eb", "#dc2626", "#9ca3af"
759    figs = {}
760    tok, model = build_pipeline()
761
762    p = trace("The cat sat on the", tok, model)["next_token_probs"]
763    top = np.argsort(-p)[:12]
764    fig, ax = plt.subplots(figsize=(6.5, 3.6))
765    ax.bar([repr(tok.decode([int(i)])) for i in top], p[top], color=BLUE)
766    ax.axhline(1 / len(p), color=RED, ls="--", label=f"uniform guess 1/{len(p)}")
767    ax.set_ylabel("probability of being next")
768    ax.set_title("Untrained model: 'The cat sat on the' → ?")
769    ax.tick_params(axis="x", rotation=45)
770    ax.legend(frameon=False)
771    fig.tight_layout()
772    figs["untrained_next_token"] = fig
773
774    temps = (0.25, 0.5, 1.0, 2.0)
775    fig, ax = plt.subplots(figsize=(6, 3.6))
776    width = 0.2
777    for j, T in enumerate(temps):
778        ax.bar(np.arange(3) + (j - 1.5) * width, next_token_probs(TEMPERATURE_SCORES, T), width, label=f"T = {T}")
779    ax.set_xticks(range(3), ["token A (score 2.0)", "token B (1.0)", "token C (0.5)"])
780    ax.set_ylabel("probability")
781    ax.set_title("Temperature: low sharpens, high flattens")
782    ax.legend(frameon=False)
783    fig.tight_layout()
784    figs["temperature"] = fig
785
786    rows = naive_vs_cached_work(prompt_len=20, max_new=200)
787    fig, ax = plt.subplots(figsize=(6, 3.6))
788    ax.plot([r[0] for r in rows], [r[1] for r in rows], color=RED, label="no cache: re-read everything")
789    ax.plot([r[0] for r in rows], [r[2] for r in rows], color=BLUE, label="KV cache: one new token per step")
790    ax.set_xlabel("tokens generated (after a 20-token prompt)")
791    ax.set_ylabel("token positions run through the model")
792    ax.set_title("Why the naive loop is too slow")
793    ax.legend(frameon=False)
794    fig.tight_layout()
795    figs["loop_cost"] = fig
796    return figs

Plot this lesson's data. matplotlib is imported here, and only here.

def demo() -> None: on GitHub
804def demo() -> None:
805    tok, model = build_pipeline()
806    prompt = "Reset your password"
807
808    banner("1. Text -> ids -> vectors -> blocks -> scores")
809    st = trace(prompt, tok, model)
810    table(
811        ["stage", "shape", "what it holds"],
812        [
813            ("ids", st["ids"].shape, f"{st['ids'].tolist()} = {tok.tokens(prompt)}"),
814            ("embeddings", st["embeddings"].shape, "one row of the token table per id"),
815            ("with positions", st["with_positions"].shape, "plus one row of the position table"),
816            ("after blocks", st["after_blocks"].shape, f"{N_LAYERS} rounds of attention + feed-forward"),
817            ("logits", st["logits"].shape, f"a score for all {tok.vocab_size} tokens, per position"),
818            ("next-token probs", st["next_token_probs"].shape, "softmax of the LAST row only"),
819        ],
820    )
821    takeaway("Only the last position's scores decide the next token; the rest matter during training.")
822
823    banner("2. Temperature on the scores (2.0, 1.0, 0.5)")
824    table(
825        ["temperature", "p(A)", "p(B)", "p(C)"],
826        [(T, *next_token_probs(TEMPERATURE_SCORES, T)) for T in (0.0, 0.5, 1.0, 2.0)],
827        floatfmt=".3f",
828    )
829
830    banner("3. The loop: predict, sample, append, repeat")
831    g = generate("The cat", 12, tok, model, seed=0)
832    say(
833        f"""
834        Continuation of 'The cat': {g.text!r}. It's gibberish, and it should be:
835        the weights are random, so every next token is close to a uniform guess
836        over {tok.vocab_size} tokens. The machinery is identical to a real model's;
837        only the numbers inside the matrices are missing. They come from training.
838        This loop ran {g.positions_processed} token positions to produce 12 tokens,
839        because it re-reads the whole sequence every step (see primer.ml.inference
840        for the KV cache).
841        """
842    )
843
844    banner("4. Training = the same forward pass + a loss")
845    loss = next_token_loss("The model reads tokens, not words.", tok, model)
846    say(
847        f"""
848        Average next-token loss of the untrained model: {loss:.2f}. A uniform guess
849        over {tok.vocab_size} tokens scores ln {tok.vocab_size} = {np.log(tok.vocab_size):.2f}.
850        Perplexity e^loss = {np.exp(loss):.0f}: the model is as unsure as picking among
851        ~{np.exp(loss):.0f} equally likely tokens. Training nudges every weight to push
852        this number down, across trillions of tokens.
853        """
854    )
855    takeaway("A language model is a next-token scorer run in a loop; training makes its scores good.")