The specification: 3,168 tests
A test is a small program that runs the lessons' code and checks that it does what the lesson claims. Every test here was written before the code it checks: first the behaviour is stated and seen to fail, then only enough code is written to make it pass (this is called test-driven development). Each test is named as a plain sentence, usually in the form given a situation, when something happens, then a result (behaviour-driven development), so the whole suite reads as a specification.
Every sentence below is one behaviour; click it to read the test that checks it. A behaviour checked on several
inputs counts as several tests but appears once. The whole suite runs in a few seconds with make test,
offline, and make spec prints this page in a terminal.
- 0. Math notation, from zero 22
- 1. The big picture 14
- 2. Neural networks 31
- 3. Optimizers 23
- 4. Training deep networks 22
- 5. Attention 20
- 6. Positional information 21
- 7. The transformer 35
- 8. Tokenization 40
- 9. Training stages 31
- 10. Pretraining at scale 62
- 11. Fine-tuning in practice 47
- 12. Reinforcement learning 42
- 13. Reasoning models 37
- 14. Alignment and safety 20
- 15. The hardware underneath 64
- 16. Inference 44
- 17. Structured output 47
- 18. Long context and efficient architectures 43
- 19. Loss functions 22
- 20. Metrics 28
- 21. Reading benchmarks 50
- 22. Overfitting and regularization 26
- 23. Trees and boosting 45
- 24. CNNs and RNNs 33
- 25. Looking inside the model 38
- 26. Word embeddings 15
- 27. Similarity 28
- 28. Training embedding models 13
- 29. Dimensions and compression 15
- 30. Vector indexes 35
- 31. Retrieval 54
- 32. Clustering and matching 22
- 33. Embeddings in production 20
- 34. Autoencoders and VAEs 26
- 35. GANs 28
- 36. Diffusion and flow matching 32
- 37. Multimodal models 41
- 38. Talking to a model 21
- 39. Orchestration 24
- 40. The agent loop 22
- 41. Tools 35
- 42. Coding and computer-use agents 58
- 43. Model Context Protocol 23
- 44. Retrieval-augmented generation 38
- 45. Context engineering 28
- 46. Memory 25
- 47. Planning 20
- 48. Evaluation 25
- 49. Guardrails 33
- 50. Cost and latency 29
- 51. Observability 18
- 52. Safe deployment 29
- 53. Why the hard ones fail 17
- glossary 25
- house style 3
- math in python 397
- navigation 1054
- viz 8
0. Math notation, from zero
22 tests in tests/test_notation.py
summation
product
- given one to four capital pi multiplies them to twenty four
- given ten steps at 95 percent the chance all succeed is about 60 percent
dot product
- given 1 2 and 3 half the dot product is 4
- given vectors at right angles the dot product is zero
- given lists of different lengths the dot product is refused
norm
matrices
- given a 2 by 3 matrix transpose turns it into 3 by 2
- given two small matrices each output cell is a row dotted with a column
- given inner sizes that disagree matrix multiply is refused
softmax and argmax
- given scores 2 1 and half softmax gives 63 23 and 14 percent
- given any scores the shares sum to one
- given scores argmax names the position of the largest not its value
spread
- given 2 4 4 4 5 5 7 9 the mean is 5
- given 2 4 4 4 5 5 7 9 the variance is 4 and the standard deviation 2
derivatives
- given x squared at 3 the slope is 6
- given a bowl x squared plus y squared the gradient at 1 2 points uphill as 2 4
agreement with num py
1. The big picture
14 tests in tests/test_big_picture.py
tracing one prompt
- the prompt becomes one id per token
- each id becomes one vector of model width
- the blocks keep one vector per token
- the output layer scores every one of the 400 vocabulary entries
- the next token probabilities add up to one
temperature
- at temperature 1 scores 2 1 and half become 63 23 and 14 percent
- at temperature half the top token rises to 84 percent
- at temperature zero the top token gets all the probability
the generation loop
- generating 5 tokens appends 5 ids to the prompt
- the same seed produces the same continuation
- without a cache each step reprocesses the whole sequence
training is the same forward pass plus a loss
2. Neural networks
31 tests in tests/test_neural_net.py
one neuron worked example
- given inputs 2 and 3 weights half and 1 and bias half the weighted sum is 4 5
- given a positive sum relu passes it through unchanged
- given a target of 5 the squared error is 0 25
- the weight on the larger input receives the larger gradient
- one sgd step raises both weights the second one more
- after one step the output moves toward the target and the loss drops
activation functions
- relu zeroes negative inputs
- given a negative input relu passes no gradient so the neuron can die
- sigmoid of zero is one half
- the sigmoid gradient never exceeds one quarter
- a saturated sigmoid passes almost no gradient
- sigmoid of a huge negative input is zero without overflow
- the tanh gradient at zero is one
- gelu of one is one times the normal cdf at one
- gelu leaks a small negative value where relu gives zero
- the tanh approximation of gelu agrees to three decimals
- softmax of zero and ln3 gives one quarter and three quarters
why nonlinearity matters
- a stack of linear layers equals a single linear layer
- a linear classifier cannot separate the two moons
- one hidden tanh layer separates the two moons
- a hidden layer learns xor
single weight learning
tiny two layer backprop
- the forward pass predicts 0 716
- the error at the output is prediction minus target
- the output weight gradient is the hidden activation times the error
- the blame reaching the first weight is shrunk by the tanh slope
backpropagation
training loop
3. Optimizers
23 tests in tests/test_optimizers.py
gradient descent on a bowl
- given a learning rate of 0 1 each step keeps 80 percent of w
- given a learning rate of 0 5 one step lands exactly at the bottom
- given a learning rate above 1 each step overshoots further and training diverges
- given a tiny learning rate ten steps barely move the weight
momentum
- given beta 0 9 the velocity accumulates past gradients
- in a long narrow valley momentum gets far lower than plain sgd in the same steps
adam
- given any gradient size the first step is the learning rate
- in a long narrow valley adam reaches the bottom where sgd stalls
weight decay
- given no gradient adamw shrinks the weight by lr times decay
- given l2 inside adam the penalty is normalised into a full learning rate step
- given l2 inside adam a hundredfold smaller penalty still shrinks the weight as much
- given adamw a hundredfold smaller decay shrinks the weight a hundredfold less
warmup cosine schedule
- halfway through warmup the rate is half the peak
- at the end of warmup the rate is the peak
- halfway through the decay the cosine has fallen to half the peak
- at the final step the rate reaches the floor
gradient clipping
- given a gradient of length 5 and a limit of 1 it is scaled to length 1
- given a gradient already under the limit it is unchanged
- the length is measured across all tensors together so direction is kept
the narrow valley paths
4. Training deep networks
22 tests in tests/test_deep_nets.py
gradients through a chain of units
- given ten sigmoid units the gradient shrinks to a quarter to the tenth
- given ten linear units with weight 1 5 the gradient explodes to 57 7
- given ten linear units with weight 1 the gradient passes through unchanged
gradients through a deep network
- given relu and he initialisation the gradient keeps its size across 30 layers
- given relu and weights that are too small the gradient vanishes
- given relu and weights that are too large the gradient explodes
- given sigmoid the gradient vanishes even with xavier initialisation
- given sigmoid with xavier the gradient shrinks about 18 orders of magnitude over 30 layers
initialisation
- xavier scale for 100 inputs and 100 outputs is 0 1
- he scale for 50 inputs is 0 2
- given he initialisation the signal keeps its size through 30 relu layers
- given sigmoid with xavier the forward signal holds near 0 5 while its gradient vanishes
residual connections
- given ten plain blocks of slope 0 025 the gradient is about 1e minus 16
- given the same blocks with skip connections the gradient is 1 28
- given skip connections a 30 layer sigmoid network stops vanishing
batch norm
- each feature is rescaled to mean 0 and spread 1 across the batch
- the same example normalises differently in a different batch
layer norm
RMS norm
clipping an exploding gradient
the gradient flow figure
5. Attention
20 tests in tests/test_attention.py
softmax
- given scores 2 1 and half the weights are 63 23 and 14 percent
- given any scores the weights sum to one
- given huge scores the weights stay finite
causal masking
- given a causal mask future tokens receive zero attention
- given the first token it can only attend to itself
- when a later token changes earlier outputs are untouched
scaling by sqrt dk
- given unscaled 512 dim vectors score spread grows to about sqrt 512
- given scaling score spread stays near one at any width
- given unscaled 512 dim vectors softmax saturates and gradients vanish
grouped query attention
- given two kv heads instead of eight key value parameters shrink fourfold
- given fewer kv heads the output shape is unchanged
agreement with py torch
the worked sentence
- given the whole sentence it still scores animal tired and street 2 1 and half
- given no mask it attends most to animal
- given a causal mask it cannot see tired because tired comes later
the attention widget agrees with the lesson
- given the sentence without a mask the widget computes the lessons weights
- given the sentence with a causal mask the widget computes the lessons weights
- given temperature one half the widget undoes the division by sqrt dk
the lesson places the attention widget
6. Positional information
21 tests in tests/test_positional.py
sinusoidal encoding
- at position zero sine dimensions are 0 and cosine dimensions are 1
- at position one with four dimensions the frequencies are 1 and one hundredth
- every value stays between minus one and one
- the similarity of two positions depends only on their distance
rotary position embedding
- at position zero a vector is left unchanged
- a 2d vector at position one turns by one radian
- rotation never changes a vectors length
- the attention score depends only on the distance between query and key
- the attention score changes when the distance changes
- a whole sequence rotates each row by its own position
order blindness without positions
- with identity projections dog first outputs 0 802 0 599
- with identity projections dog last outputs the same 0 802 0 599
- without positions dog bites man and man bites dog pool to the same vector
- without positions each word gets the same output whatever its slot
- with positions the two sentences pool to different vectors
- a causal mask alone already leaks word order
context extension
7. The transformer
35 tests in tests/test_transformer.py
normalization
- layer norm turns 1 2 3 4 into mean zero unit spread
- rms norm divides 1 2 3 4 by their root mean square without centring
- layer norm ignores how large the input is
- each token is normalized on its own
GELU
- gelu of zero is zero
- gelu of one is about 0 841
- gelu lets large positive inputs through unchanged
- gelu squashes large negative inputs to almost zero
feed forward
- the hidden layer is four times wider than the model
- the output has the same shape as the input
- with d model 8 it has 552 parameters
- each token is processed without looking at the others
transformer block
- the output has the same shape as the input so blocks can stack
- with both sublayers switched off the residual path passes the input through unchanged
- in a decoder block a later token never changes an earlier output
- in an encoder block a later token changes the first output
- given the same input only the last tokens attention row matches between encoder and decoder
tiny GPT
- it returns one score per vocabulary entry for every position
- a later token never changes the scores at earlier positions
- the output layer reuses the embedding table
- with vocab 50 width 32 two layers and 16 positions it has 27328 parameters
parameter counting
- gpt2 small has 124439808 parameters
- given gpt2 small without attention biases it has 124402944 parameters
- without embeddings a wide model has about 12 times layers times width squared
mixture of experts
- router scores 2 1 half minus 1 send the token to experts 0 and 1 at 73 and 27 percent
- every token is sent to exactly k experts
- the output has the same shape as the input
- eight experts of width 16 hold 17152 parameters in total
- with top 2 routing each token uses only 4384 of those parameters
- given an untrained router and 256 tokens the load ranges from 51 to 75 against an even 64
load balancing
- a perfectly balanced router scores a balancing loss of 1
- a router sending everything to one of 4 experts scores about 4
compute arithmetic
8. Tokenization
40 tests in tests/test_tokenization.py
byte pair encoding training
- given low lower lowest the first merge joins l and o seen three times
- given low lower lowest the second merge joins lo and w seen three times
- given low lower lowest the third merge joins low and e seen twice
- after three merges the text reads low lowe r lowe s t
- given every word is already one token training stops early
byte pair encoding encoding
- given the learned merges lowest encodes as lowe s t
- given an unseen word the learned merges still apply inside it
- given a word no merge applies to it stays as characters
word piece
- given low lower lowest wordpiece prefers s and t which never appear apart
- given low lower lowest wordpiece scores l and o at one third
byte level BPE
- given no training the vocabulary is the 256 byte values
- given an untrained tokenizer each utf8 byte is one token
- given any text decoding the encoding returns it unchanged
- given training frequent words become single tokens
- given training the vocabulary grows to the requested size
tokenizer quirks
- given a leading space cat becomes a different token
- given 1234 is frequent mid sentence it is one token while 12345 takes two
- given the same number without its leading space it fragments into pieces
- given strawberry the model sees fewer tokens than its ten letters
- given strawberry mid sentence it splits into straw and berry
- given an english trained tokenizer hindi costs more tokens per character
rules of thumb
- given 400 characters of prose the estimate is 100 tokens
- given 1000 tokens the text is about 750 words
- given 2000 input and 500 output tokens at 3 and 15 dollars per million the call costs 1 35 cents
agreement with tiktoken
- given english prose a production tokenizer averages three to six characters per token
- given a production tokenizer a leading space changes the token
the tokenizer widget agrees with the lesson
- given the first merges the widget splits text into the pieces a tokenizer trained that far gives
- given no merges every character of plain english is its own token
- given every merge the lesson learns the password is one token
- given an emoji the widget counts characters per token as the lesson does
the tokenizer widget data
9. Training stages
31 tests in tests/test_training_stages.py
supervised fine tuning loss
- given a 3 token prompt and 2 token reply only the 2 positions predicting the reply are trained
- given the reply predicted with probability one half the sft loss is ln 2
- given the same sequence the pretraining loss averages every position
reward model
- given equal rewards the preferred answer wins half the time and the loss is ln 2
- given the chosen answer scores 2 higher it is preferred with probability 0 88
- given the chosen answer scores 2 higher the loss is 0 127
direct preference optimization
- given a policy identical to the reference the dpo loss is ln 2
- given the worked example the dpo loss is 0 598
- given the worked example the update strength is 0 045
- given a pair the policy ranks the wrong way the update strength is larger
- when a toy policy trains on preferences the chosen answer becomes more likely
- when a toy policy trains for 50 steps the rambling answer falls more slowly than the rude one
- given a larger beta the same departure from the reference already gives a lower loss
lo RA
- given the adapter starts with B at zero the adapted layer behaves exactly like the frozen one
- given W identity B 1 0 and A 0 1 input 1 2 becomes 3 2
- given a 4096 by 4096 layer lora trains 2 times 4096 times rank parameters
- given a 4096 by 4096 layer full fine tuning trains 16 8 million parameters
- when the adapter is merged into W the outputs are unchanged
- when trained on a low rank correction a rank 2 adapter removes over 99 percent of the error
- when the adapter trains the pretrained weights do not change
choosing an adaptation
- given missing or changing facts the answer is retrieval not fine tuning
- given wrong format that a better prompt fixes prompting is enough
- given behaviour a prompt cannot make consistent the answer is a lora fine tune
- given a major domain shift with lots of data the answer is a full fine tune
- given neither missing knowledge nor wrong behaviour nothing needs changing
distillation
10. Pretraining at scale
62 tests in tests/test_pretraining.py
language identification
- given an english sentence it is identified as english
- given a french sentence it is identified as french
- given a string of product codes it is marked unknown
heuristic quality filters
- given a well formed paragraph it passes every rule
- given a short snippet it fails the length rule
- given a page full of hashtags it fails the symbol rule
- given a keyword list with no function words it fails the stop word rule
classifier quality scores
- given reference text style the classifier scores it as good
- given spam style the classifier scores it as bad
exact deduplication
jaccard similarity
min hash
- given identical documents the estimate is exactly one
- given sets with true jaccard one third 256 hashes estimate it within point zero eight
- given disjoint documents the estimate is near zero
locality sensitive hashing
- given 20 bands of 5 rows a pair at similarity 08 becomes a candidate almost surely
- given 20 bands of 5 rows a pair at similarity 03 rarely becomes a candidate
- given a near copy among different documents only that pair is flagged
why duplicates hurt
- given a boilerplate line repeated a hundred times a bigram model nearly memorises it
- given the same line kept once the model is far less sure of it
the curation pipeline
data mixtures
- given a small source with a large weight it is seen more than once
- given a large source with a matching weight it is seen less than once
token budgets
- given a 7b model the compute optimal budget is about 140b tokens
- given 7b parameters and 140b tokens training costs about 6 times 10 to the 21 flops
synthetic data and model collapse
training memory
- given adam in mixed precision each parameter costs 16 bytes
- given a 7b model its training state needs 112 gb more than one 80 gb gpu
- given a 7b shape at 4096 tokens activations take about 104 gb
- given attention that never stores its score matrix activations drop to about 18 gb
- given activation checkpointing only each layers input is kept about 1 gb
data parallelism
- given a batch split across four workers the averaged gradient equals the full batch gradient
- given a ring all reduce every worker ends with the sum of all gradients
- given a ring of four each worker sends one and a half times its gradient
- given a thousand workers the traffic per worker stays under twice the gradient
sharded optimizer state
- given no sharding every gpu holds all 120 gb
- given sharded optimizer state each gpu holds 31 point 4 gb
- given sharded optimizer state and gradients each gpu holds 16 point 6 gb
- given everything sharded each gpu holds under 2 gb
tensor parallelism
- given a weight split by columns across four devices the joined result equals the unsplit matmul
- given a weight split by rows the summed partial results equal the unsplit matmul
- given a feed forward layer split column then row it matches the single device layer
pipeline parallelism
- given 4 stages and 8 microbatches idle slots in the schedule are 3 elevenths
- given 4 stages and one microbatch three quarters of the time is bubble
- given 32 microbatches the bubble shrinks below nine percent
number formats
- given random values fp16 rounding matches numpys float16
- given one third bf16 keeps only about three significant digits
- given 70000 fp16 overflows to infinity
- given 70000 bf16 holds it because it has fp32s range
- given fp8 e4m3 its largest value is 448
loss scaling
- given a gradient of one hundred millionth fp16 flushes it to zero
- given a loss scale of 65536 the same gradient survives fp16 and unscales correctly
- given bf16 the tiny gradient survives without any scaling
- given an overflowing step the scaler skips it and halves the scale
- given enough clean steps in a row the scaler doubles the scale
master weights
- given a thousand tiny updates an fp32 master copy accumulates all of them
- given the same updates applied to bf16 weights every one is lost
gradient clipping at scale
- given gradients sharded across gpus the combined norm equals the norm of the whole
- given one corrupted batch without clipping the loss spikes above a thousand
- given the same batch with clipping the loss barely moves
spike detection
checkpointing
11. Fine-tuning in practice
47 tests in tests/test_fine_tuning.py
the cost of fine tuning
- given 3000 prompt tokens at 2 dollars per million a request costs six tenths of a cent
- given 600 dollars up front and 0 48 cents saved per request it pays off after 125000 requests
- given a tuned model that costs more per request it never pays off
formatting chat examples
- given a system user and assistant turn only the assistant text is trained on
- given a well formed example it has no problems
- given an example that ends on the users turn it is rejected
- given an empty assistant reply it is rejected
- given an unknown role it is rejected
deduplication
- given two phrasings one word apart they share five of six word pairs
- given two different questions with the same opening they share a quarter of their word pairs
- given texts that differ only in case and punctuation they count as identical
- given a near duplicate deduplication keeps only the first
- given a training example that nearly copies an eval example it is removed from training
the held out set
- given 100 eval examples at 80 percent the margin of error is about 8 points
- given four times as many eval examples the margin of error halves
- given a split the eval set never overlaps training
- given the same seed the split is the same every time
label quality
- given a 90 percent model and 10 percent wrong labels the eval reports 82 percent
- given a perfect model the eval scores no higher than the share of correct labels
- given two labelers who disagree on one item in five they agree 80 percent
the toy model
- given the base model it has learned its general skill
- given a fine tune on task a it learns task a
- given parameters as one flat vector the model rebuilt from them predicts the same
forgetting the general skill
- when fine tuning runs long after task a is learned the general skill keeps fading
- given a lower learning rate the general skill fades less over the same steps
- given forgetting it grows with how far the weights move
catastrophic forgetting
- given forgetting as a number it is the drop in accuracy on the old task
- given a model fine tuned on a then on b its accuracy on a collapses
- given a new task on separate inputs the old task is mostly kept
- given a lower learning rate task a is still lost once task b is learned
replay
- given 200 new and 10 old examples the mix has 210 examples
- given old examples that are being forgotten they carry most of the loss
- given 5 percent of task a mixed into task b both tasks are kept
overfitting a small dataset
- given 16 examples and many epochs training loss keeps falling
- given 16 examples validation loss bottoms out early then rises
- given the validation curve early stopping stops soon after the best epoch
task arithmetic
- given a fine tune its task vector is the change from the base
- given two task vectors at lambda one the merge adds both changes
- given two fine tunes their average is task arithmetic at lambda one half
- given task vectors for separate inputs they are nearly orthogonal
- given task vectors for separate inputs their sum does both tasks
- given plain averaging each skill is diluted compared with the full sum
interference
- given the worked vectors their cosine is minus eight tenths
- given tasks that rewrite the same weights their task vectors point apart
- given conflicting task vectors no scaling makes their sum do both tasks
training
12. Reinforcement learning
42 tests in tests/test_reinforcement.py
the bandit
- given an arm that wins 80 percent of the time about 800 of 1000 pulls pay out
- given an offset every pull pays the offset plus the win
- given a uniform policy over arms winning 20 50 and 80 percent the expected reward is one half
- given a policy that always picks the best arm the expected reward is its win chance
the log probability trick
- given a uniform policy over three actions the gradient of log probability of the third is one hot minus a third
- given one rewarded step at learning rate one half the chosen action rises from a third to 0 452
- given a reward of zero the policy does not move
- given the averaged estimate it equals the slope of expected reward measured by nudging each logit
learning from reward
- given 500 pulls the policy puts over 90 percent on the arm that pays most
- given training the last hundred pulls earn more than the first hundred
baselines
- given rewards 10 and 12 and no baseline the gradient estimate swings with variance 30 25
- given a baseline of 11 the same estimate has zero variance
- given a baseline the average gradient is unchanged
- given rewards offset by 5 subtracting the average reward cuts the sampled variance over fiftyfold
- given rewards offset by 5 a running average baseline finds the best arm in every run
- given rewards offset by 5 and no baseline many runs lock onto a worse arm
PPO clipping
- given ratio 1 5 and advantage 2 the objective is capped at 1 2 times 2
- given ratio 0 5 and advantage minus 1 the objective is capped at minus 0 8
- given ratio 1 5 and advantage minus 1 the full penalty minus 1 5 applies
- given a ratio inside the band the objective is ratio times advantage
- given a ratio past 1 plus epsilon with positive advantage raising it further gains nothing
- given fifty passes over one batch without clipping the ratio runs past two
- given fifty passes over one batch with clipping the ratio stays below one and a half
- given a single pass the ratio starts at exactly one
the KL leash
- given policy probability 0 6 against reference 0 3 and beta 0 1 a reward of 1 becomes 0 931
- given reference one half each rewards 1 and 0 and beta 1 the best policy is 0 731 and 0 269
- given a huge beta the best policy is the reference
- given a tiny beta the best policy puts everything on the highest reward
group relative advantages
- given group rewards 1 0 0 1 the advantages are plus one minus one minus one plus one
- given one right answer in four it earns 1 73 and each wrong one minus 0 58
- given every answer in a group is right every advantage is zero
- given a verifier the correct sum scores one and anything else zero
- given grpo training accuracy on the addition prompts climbs from under 30 to over 90 percent
- given grpo training more groups stop teaching anything as the prompts are mastered
- given the addition prompts every answer is a single digit
reward hacking
- given ratings of one to four sentence answers the fitted reward model is 0 3125 plus 0 1875 per sentence
- given the fitted reward model it scores ten sentences highest while true quality peaks at four
- given optimising the flawed reward true quality first rises
- given optimising the flawed reward for longer true quality ends below where it started
- given a kl leash true quality ends higher than without one
- given a verifiable reward for the true goal quality climbs to nearly perfect
- given too weak a leash the best policy is worse than the reference and a moderate one is better
13. Reasoning models
37 tests in tests/test_reasoning.py
thinking out loud
- given 3 5 8 2 answered in one token the model runs out of steps and guesses 17
- given room to write steps the model writes running totals and gets 18
- given room to write steps it stops once done so a 3 addition sum costs 3 tokens
- given two numbers one forward pass is enough to answer directly
thinking budget
- given six difficulties and a budget of 8 tokens the formula gives 70 percent
- given difficulties spaced by doubling each doubling of budget adds the same accuracy
- given a budget past the hardest problem more tokens buy no accuracy
- given real sums accuracy climbs from under a third to every problem as the budget grows
- given real sums the simulation lands on the formula within a few points
self consistency
- given answers 18 17 18 the vote picks 18
- given a solver right 60 percent of the time three votes are right 64 8 percent
- given a solver right 60 percent of the time five votes are right 68 3 percent
- given a solver right less than half the time on a yes no question voting makes it worse
- given wrong answers that scatter a 40 percent solver still wins the vote most of the time
- given independent samples 31 votes push a 60 percent solver above 95 percent
- given samples that share the same mistake 31 votes barely beat one
verifiers
- given a step whose arithmetic holds the process verifier accepts it
- given a step whose arithmetic is off the process verifier rejects it
- given a chain with a slip in step two the process verifier points at step two
- given a clean chain the process verifier finds no bad step
- given two slips that cancel the outcome reward is full but the process verifier objects
- given samples right 30 percent of the time one of five is right 83 percent of the time
- given noisy chains best of n with a step checker beats voting and a single sample
- given a perfect step checker best of n gets within a few points of the pass at n ceiling
learning to reason with RL
- given rewards 1 0 0 1 group advantages are plus one and minus one
- given a group that all succeeds there is no signal to learn from
- given rewards that differ by a hair dividing by the spread blows them up to plus and minus one
- given rewards that differ by a hair without dividing by the spread they stay a hair apart
- given only a correctness reward the model learns to write longer and gets more right
- given a hard problem the trained model leaves room to check every step
- given a length penalty the model thinks less on easy problems than on hard ones
- given a length penalty and division by the groups spread easy problems shrink to one token
- given a length penalty without division by the spread easy problems settle on the best length two
- given a length penalty the model spends fewer tokens than without one
what reasoning costs
14. Alignment and safety
20 tests in tests/test_alignment.py
constitutional critique
- given a response that breaks no principle the critique finds nothing
- given a response that overclaims certainty the honesty principle flags it
- given two candidates the one breaking fewer principles is chosen
- given two candidates that tie no preference pair is produced
- given an overclaiming uncited answer one revision leaves no principle broken
red teaming
- given the exact blocked placeholder the filter catches it
- given a seeded search the attack success rate is between zero and one
- given the same seed the search finds the same rate
- given six attempts of which four get through the attack success rate is two thirds
- given only the hand written tests the filter looks perfect
- given the filter is patched with what the search found the same search no longer gets through
sycophancy
- given a model that never defers its flip rate is zero
- given a model that defers half the time about half its answers flip
- given raters reward agreement by 2 and correctness by 1 the agreeing answer wins 73 percent
refusal tradeoff
- given a stricter threshold benign refusals rise and harmful compliance falls
- given the hand checked scores a threshold of half gives one third on each error
goodhart
- given more optimization the proxy keeps rising while the true goal eventually falls
- given the toy curves the true goal peaks near step 16
release gate
15. The hardware underneath
64 tests in tests/test_hardware.py
counting the work in a matrix multiply
- given a 2x3 times 3x2 multiply it takes 12 multiplies and 12 adds
- given the counted multiply its answer matches the product worked by hand
- given two 4096 square matrices the multiply costs 2 m n k flops
splitting a multiply across cores
- given one core a 64 cube multiply takes one round per multiply add
- given as many cores as outputs only one dot product of rounds remains
- given more cores than outputs the extra cores sit idle
the memory hierarchy
- given the levels from registers outward each holds more than the one before
- given the levels from registers to host memory each is slower than the one before
- given one gigabyte reading it from hbm takes about 0 3 milliseconds
- given one gigabyte fetching it from host memory takes 67 times longer than from hbm
tiling a matrix multiply
- given any tile size the tiled product equals the ordinary product
- given 4x4 matrices and no reuse 128 values are read from slow memory
- given 4x4 matrices and 2x2 tiles the reads halve to 64
- given a tile as big as the matrix each input is read exactly once
- given any tile every output is written to slow memory exactly once
- given tiles of t by t fast memory holds three of them at once
- given the simulation its counted reads match the formula 2 n cubed over t
- given 4096 square matrices in 16 bit 128 tiles do about 63 flops per byte
- given 8 bit numbers instead of 16 bit the same tiles do twice the flops per byte
storing a number in a float format
- given bf16 minus 6 5 is stored as sign 1 exponent 129 mantissa 80
- given the stored fields of minus 6 5 decoding them gives back minus 6 5
- given bf16 the bits of minus 6 5 read sign then exponent then mantissa
- given a value below the smallest normal fp16 stores it as a subnormal
- given each format its largest finite value matches its specification
- given each format the gap just above one is two to the minus mantissa bits
rounding agrees with real hardware formats
- given random values of every size rounding to fp16 agrees with numpy float16
- given random values of every size rounding to fp32 agrees with numpy float32
- given float32 values rounding to bf16 agrees with rounding away the low 16 bits
range and precision
- given a value past fp16s largest it overflows to infinity
- given fp8 e4m3 a value past 448 has no encoding
- given a gradient of one hundred millionth fp16 flushes it to zero
- given the same tiny gradient bf16 keeps it within one percent
- given bf16 adding a thousandth to one changes nothing
- given fp32 adding a thousandth to one is kept
why smaller formats are faster
all reduce across gpus
- given four gpus ring all reduce leaves every gpu holding the sum
- given n gpus each sends 2 times n minus 1 over n of the data
- given 7b gradients in 16 bit on 8 gpus in one machine the all reduce takes about 49 ms
- given the same gradients across machines the all reduce takes ten times longer
- given 7b parameters and 16k tokens one training step of arithmetic takes about 0 69 s
will a model fit for serving
16. Inference
44 tests in tests/test_inference.py
memory math
- given 70 billion parameters at 16 bit the weights need 140 gb
- given 70 billion parameters at 4 bit the weights need 35 gb
- given a llama 3 8b shaped model the kv cache costs 131 072 bytes per token
- given a 32 000 token context one request holds about 4 gb of kv cache
- given an 80 gb gpu and 16 gb of weights 15 requests at 32k context fit
- given 32 kv heads instead of 8 the kv cache is four times larger
prefill versus decode
- given one token at a time at 16 bit each byte of weights read does 1 flop
- given a 1000 token prompt prefill does 1000 flops per byte of weights
- given a gpu with 1 pflop per second and 3 35 tb per second the break even is about 300 flops per byte
- given single token decoding the gpu is memory bound
- given a 1000 token prefill the gpu is compute bound
- given an 8b model at 16 bit each decoded token takes about 4 8 ms
- given an 8b model a 1000 token prompt prefills in about 16 ms
KV cache
- given a cached prefix incremental decoding matches full recomputation
- given the same prompt greedy generation with and without the cache produces the same tokens
- given an 8 token prompt and 16 new tokens full recomputation processes 248 token positions
- given an 8 token prompt and 16 new tokens the cache processes 23 token positions
- when a token is decoded every layer cache grows by one key and one value
sampling
- given temperature 1 the probabilities are the plain softmax
- given temperature one half the distribution sharpens to 0 867 0 117 0 016
- given temperature 0 all probability goes to the top token
- given top k of 2 only the two likeliest tokens survive renormalised
- given top p of 0 9 the smallest set reaching 90 percent survives
- given a confident distribution top p keeps fewer tokens than a flat one
- given the same numbers adding them in a different order gives a different float
- given two nearly tied tokens a different summation order flips the greedy choice
speculative decoding
- given target 0 5 0 3 0 2 and draft 0 3 0 3 0 4 a drafted token is accepted 80 percent of the time
- given an 80 percent acceptance rate and 4 drafts a round yields 3 36 tokens on average
- given a draft identical to the target every draft is accepted plus one bonus token
- given a different draft the first token still follows the target distribution
- given a different draft the second token still follows the target distribution
quantization
- given a row whose largest weight is 1 27 int8 codes are 50 minus 127 and 2
- given the same row int4 codes are 3 minus 7 and 0
- given int4 codes the row comes back as 0 544 minus 1 27 and 0
- given one outlier weight per channel scales keep the other rows precise
- given random weights int8 error is far below int4 error
continuous batching
- given static batches the short request waits for the long one and serving takes 5 steps
- given static batches the gpu slots are busy 70 percent of the time
- given continuous batching a finished slot is refilled at once and serving takes 4 steps
- given continuous batching the gpu slots are busy 87 5 percent of the time
prompt caching
- given two prompts sharing the first 1000 tokens 1000 tokens can be reused
- given a change in the very first token nothing can be reused
- given 100 requests with a 10k token shared prefix caching cuts the input bill from 3 15 to 0 48 dollars
top p adapts to confidence
17. Structured output
47 tests in tests/test_structured_output.py
asking nicely is not enough
- given ten tokens each right 98 percent of the time the whole output is valid 82 percent of the time
- given a longer output the chance that every token is right falls
- given the toy model unconstrained a sizeable share of its answers break the pattern
- given a chatty model some unconstrained answers open with a preamble
masking the logits
- given logits 2 1 half and 0 with only the last two allowed they share 62 and 38 percent
- given a mask every forbidden token gets exactly zero probability
- given a mask the allowed tokens keep their odds against each other
- given the shared vocabulary the end of output token comes last
from a pattern to a state machine
- given the enum cat car dog the machine accepts exactly those three words
- given an optional sign and digits the machine needs three states
- given random strings the machine agrees with pythons re module
which tokens may come next
- given the start state a multi character token is allowed when the machine can consume all of it
- given the prefix ca only t and r may follow
- given a finished word only the end token may follow
- given a precomputed table each row matches checking every token on the fly
constrained decoding with a pattern
- given the age pattern every constrained answer matches it
- given the age pattern the constrained answers never open with a preamble
nesting needs a stack
- given a prefix inside three containers the stack holds one frame for each
- given the innermost object closes its frame comes off the stack
- given a closing bracket that does not match the last one opened the prefix is dead
can this prefix still be completed
- given a prefix that can still be finished the checker says open
- given a whole object with every required field the checker says complete
- given a key the schema does not list the prefix dies at its first wrong letter
- given a required field still missing the object cannot close
- given an integer field a decimal point is refused
- given a number json forbids a leading zero
- given an enum only the listed words can be spelled
- given any json every prefix of a valid document stays alive
- given random text whatever the checker calls complete python can parse
constrained decoding with a schema
- given the person schema every constrained answer that finishes parses and fits the schema
- given a token budget too small the answer is cut off valid so far but unfinished
- given a finished value only the end token is allowed
what the mask costs in quality
- given a model that wanted null forcing an integer throws away 80 percent of its belief
- given a schema that admits null the honest answer is allowed again
- given a locally tempting first token masking commits to it nine times in ten
- given the same model its own odds among valid answers give that path only 9 percent
token boundaries
- given the text 42 the toy vocabulary can spell it two ways
- given a whole answer the mask allows many spellings the model rarely saw
what the mask costs in time
- given a hundred steps the naive mask walks every character of every token every step
- given a precomputed table the work is paid once per state plus one row read per step
when retrying is enough
- given answers valid 95 percent of the time retrying needs 1 05 attempts on average
- given answers valid two times in three three tries all fail about 4 percent of the time
- given a longer answer the unconstrained model breaks it more often
- given constrained decoding no finished answer breaks the rules on either task
json mode and tool arguments
18. Long context and efficient architectures
43 tests in tests/test_efficient_architectures.py
the cost of long context
- given a llama 3 8b shaped model the kv cache for 128k tokens is about 17 gigabytes
- given a million tokens the kv cache outgrows an 80 gigabyte gpu
- given eight times the context the pairs scored grow sixty four fold
sliding window attention
- given a window of three a token sees itself and the two before it
- given eight tokens and a window of three 21 of the 36 causal pairs are scored
- given a window as long as the sequence it is ordinary causal attention
- given three layers of window three the last token hears six positions back but not seven
- given 32 layers and a 4096 token window information can travel about 131 thousand tokens
- given a rolling buffer decoding matches the masked computation
- given a rolling buffer it never holds more than w keys
sparse attention patterns
- given a global first token every later token can attend to it
- given sixteen tokens a window of four and one global token 70 of 136 pairs are scored
- given a stride of four a token sees its last four and every fourth before them
- given sixteen tokens and a stride of four 82 pairs are scored
- given any sparse mask the pairs it skips get exactly zero attention
- given a fixed window doubling the sequence doubles the pairs instead of quadrupling them
linear attention
- given any input the feature map is positive so no weight is negative
- given the worked example the third token reads 62 fifteenths
- given the worked example the running sums are 20 22 and 5 5
- given the same inputs the running sum matches the quadratic form exactly
- given the same inputs linear attention gives a different answer from softmax attention
- given any sequence length the state it carries has the same size
state space models
- given a half decay and inputs 1 0 0 2 the outputs are 1 half quarter and 2 125
- given a half decay the convolution kernel is 1 half quarter eighth
- given a time invariant ssm the convolution gives the same outputs as the recurrence
- given a large step the state is replaced by the new input
- given a tiny step the state is kept and the input ignored
- given a marked token followed by noise a selective ssm still holds it at the end
- given the same task no fixed step size holds even half of it
- given a parallel scan it gives the same states as the step by step loop
- given 1024 steps a parallel scan needs only 10 rounds
- given any context length the ssm state takes the same memory
hybrid models
compressing the KV cache
- given a long context a sliding window cache stops growing at the window
- given the worked latent example a token is stored as 3 1 and its key is 3 1 3 1
- given a 512 number latent plus a 64 number position key the cache is 57 times smaller than 128 full heads
- given cached latents attention matches attention over the up projected keys and values
- given the up projection absorbed into the query the output is unchanged
- given a latent of width 6 every head s keys together have rank at most 6
- given the worked key 4 bit codes are 7 minus 3 2 and 0
- given an 8 bit cache attention output moves by under one percent
- given a 4 bit cache attention output moves more than with 8 bits
- given 4 bit codes and one 16 bit scale per head vector the cache is about 3 9 times smaller
19. Loss functions
22 tests in tests/test_losses.py
cross entropy
- given probability p on the right token the loss is minus ln p
- given 1 percent versus 90 percent on the right token the loss is about 44 times larger
- given certainty on the right answer the loss is zero
perplexity
- given an average loss of 0 69 perplexity is 2 like a coin flip
- given uniform guesses over 10 tokens perplexity is 10
cross entropy from logits
- given logits 2 1 and 0 1 with the first correct the loss is 0 417
- given huge logits the loss stays finite
- the gradient is predicted probabilities minus the one hot target
regression losses
- mse of four unit misses and one miss of 10 is 20 8
- mae of the same misses is 2 8
- under mse one outlier dominates the loss
- under mae the same outlier has far less pull
contrastive info NCE
- given perfectly matched orthogonal pairs the loss is ln 1 plus e minus 1
- given swapped pairs the loss is ln 1 plus e
- a lower temperature sharpens the reward for correct pairs
- a hard negative close to the query raises the loss
- the analytic gradients match finite differences
label smoothing
20. Metrics
28 tests in tests/test_metrics.py
fraud worked example
- given 6 correct of 8 flagged precision is 75 percent
- given 6 found of 10 frauds recall is 60 percent
- given precision 75 and recall 60 f1 is 67 percent
- given 88 true negatives accuracy looks great at 94 percent
- the confusion matrix counts 6 hits 2 false alarms 4 misses and 88 correct rejections
accuracy trap
- given 1 percent fraud a model that flags nothing is 99 percent accurate
- given 1 percent fraud a model that flags nothing finds no fraud
roc auc
- given 3 of 4 pairs ranked correctly the area under the roc curve is 0 75
- given 3 of 4 pairs ranked correctly the pairwise win rate is 0 75
- given a perfect ranking auc is 1
- given a positive and negative with the same score the tie earns half credit
- given many tied scores the curve area equals the pairwise win rate
threshold by error cost
- given misses cost ten times false alarms the threshold drops to catch every positive
- given false alarms cost ten times misses the threshold rises to flag only the surest case
retrieval metrics
- given 2 of 3 relevant docs in the top 5 recall at 5 is two thirds
- given 2 relevant docs among 5 results precision at 5 is 0 4
- given the first relevant doc at rank 2 the reciprocal rank is one half
- given first hits at ranks 1 and 2 and one miss mrr is one half
- given graded relevance ndcg at 5 is 0 56
- given the ideal order ndcg is 1
word overlap metrics
- given an exact copy of the reference bleu is 1
- given a repeated word bleu credits it only as often as the reference has it
- given candidate abcd and reference acde rouge l is 0 75
- a wrong answer that copies the wording outscores a correct paraphrase on bleu
- a wrong answer that copies the wording outscores a correct paraphrase on rouge l
judge calibration
21. Reading benchmarks
50 tests in tests/test_benchmarks.py
a benchmark is an average of per question scores
scoring multiple choice by likelihood
- given summed log probabilities the one word option wins because it has fewer tokens to pay for
- given per token averages the correct long option wins
scoring a generated answer
- given worked reasoning the last number is taken as the answer
- given a number with a thousands separator it is read as one number
- given an answer spelled in words no number is found and it scores as wrong
pass at k
- given 3 correct of 10 samples pass at 1 is the plain success rate
- given 3 correct of 10 samples pass at 5 is one minus 21 over 252
- given too few wrong samples to fill a draw pass at k is certain
- given no correct samples pass at k is zero
- given every possible draw of k samples it equals the share of draws holding a correct one
- given a true solve rate of 20 percent the estimator is right on average
- given the naive plug in formula it underestimates on average
a score is an estimate
- given 80 percent on 100 questions the standard error is 4 points
- given 80 percent on 100 questions the 95 percent interval runs from 72 to 88
- given four times the questions the interval is half as wide
- given a coin flip score telling apart one point needs about 9600 questions
- given 80 of 100 right the bootstrap interval agrees with the formula
comparing two models
- given 82 and 79 percent on 200 questions the gap has a standard error of about 4 points
- given a 3 point gap on 200 questions the interval for the gap includes zero
- given the same 3 point gap on 5000 questions the interval excludes zero
contamination
- given a verbatim copy in the training data every 8 gram of the question is found
- given unrelated text about muffins no 8 gram is found
- given a paraphrase only one of 16 five grams matches so n gram checks miss rewording
- given 30 percent of questions leaked a 60 percent model reports 72
- given leaked questions scoring the clean subset alone recovers the true skill
- given a copy that kept the canary string the filter drops it from the training data
saturation
- given ability equal to the middle difficulty the expected score is one half
- given an easy benchmark two quite different models score within one point
- given a hard benchmark the same two models are 40 points apart
- given 5 percent wrong answer keys even a perfect model cannot pass 95 percent
goodhart
- given 20 equally good versions the one picked on the test set looks 6 points better
- given the winner rescored on fresh questions its score falls back to the truth
elo ratings
- given equal ratings each side is expected to win half the time
- given a 400 point gap the stronger side wins 10 times in 11
- given a 100 point gap the stronger side wins 64 percent
- when one of two equal models wins it gains 16 points and the other loses 16
- given the same votes in another order online elo gives different ratings
bradley terry
- given a beats b 7 times in 10 the fitted gap is 147 elo points
- given the same votes in any order the fit is identical
- given 20000 simulated votes it recovers the true ratings within 25 points
style bias
- given voters who favour long answers a verbose model climbs past a stronger one
- given answer length as a second factor the verbose model returns to its true rating
- given answer length as a second factor the fit measures how much voters favour length
reading an announcement
- given like for like settings and a clear gap nothing is flagged
- given a gap smaller than the noise it is flagged
- given different numbers of worked examples in the prompt it is flagged
- given best of several attempts against a single attempt it is flagged
- given scores near the ceiling it is flagged as saturated
- given no contamination check it is flagged
22. Overfitting and regularization
26 tests in tests/test_regularization.py
memorising
- a cubic through the four zigzag points has zero training error
- the memorising cubic predicts 8 at x equals 4
- the flat line at 0 5 has training error 0 25 but predicts sensibly
underfitting and overfitting
- given 12 noisy points a degree 11 polynomial fits the training set exactly
- the exact fit does far worse on new data than a cubic
- given degrees 0 to 11 validation error is lowest at degree 5 and rises from degree 6
- a straight line underfits with high error on both training and new data
early stopping
- given patience 2 training stops two epochs after the best and keeps the best
- given a steadily improving run early stopping never triggers
- when a flexible model trains too long training loss keeps falling while validation loss rises
bias variance
- a simple model is systematically off so its bias is larger
- a flexible model swings with each training sample so its variance is larger
dropout
- given a drop rate of one half every surviving activation is doubled
- on average the output equals the input
- at evaluation time dropout does nothing
l 1 and l 2 penalties
- soft thresholding shrinks 3 to 2 and snaps 0 5 to exactly zero
- given plain weights 3 0 5 and minus 2 l1 with strength 1 gives 2 0 and minus 1
- given the same weights l2 with strength 1 halves each one and zeroes none
- given ten features of which three matter l1 sets most useless weights to exactly zero
- given the same problem l2 leaves every weight nonzero
data splits
- given 100 examples a 70 15 15 split gives 70 15 and 15 with no overlap
- in 5 fold cross validation of 10 examples every example is held out exactly once
- in 5 fold cross validation of 10 examples each model trains on 8
data leakage
23. Trees and boosting
45 tests in tests/test_classical.py
measuring impurity
- given four spam and four ham the gini impurity is one half
- given a pure node the gini impurity is zero
- given one ham and four spam the gini impurity is 0 32
- given an even split of two classes the entropy is one bit
- given one ham and four spam the entropy is about 0 72 bits
choosing the best question
- given the eight emails asking known sender leaves weighted impurity 0 2
- given the eight emails asking more than 2 links leaves weighted impurity 0 375
- given the eight emails the best first question is known sender
- given the eight emails there are six questions worth trying
- given values 0 and 3 on either side the threshold sits halfway at 1 5
- given a node that is already pure there is no question worth asking
growing a tree
- given the eight emails a depth two tree sorts every one correctly
- given a depth limit of one the tree is a single question with two leaves
- given a depth limit the tree never grows deeper
- given the spam tree its rules read as plain questions
- given any input the predicted class probabilities sum to one
- given features rescaled or squashed the tree makes the same predictions
- given the same data the first question matches scikit learn
overfitting
- given no depth limit the tree memorises its training set
- given noisy data a moderate depth beats unlimited depth on validation
- given more depth training accuracy never falls
bagging
- given eight examples the chance one is left out of a bootstrap sample is 34 percent
- given a large dataset the left out share approaches one over e
- given many bootstrap samples the share left out matches the formula
- given correlation 0 3 and ten trees the averaged variance is 0 37
- given uncorrelated trees averaging divides the variance by their number
random forest
- given noisy data a forest beats a single deep tree on validation
- given more trees validation accuracy ends higher than with one
- given the same seed the forest is reproducible
gradient boosting
- given the four houses the starting guess is the mean price of 4
- given learning rate one half one round moves the guesses to 2 75 and 5 25
- given that round the mean squared error falls from 6 5 to 1 8125
- given squared loss the residual is the negative slope of the loss
- given more rounds training loss never rises
- given a smaller learning rate training loss falls more slowly
- given too many rounds validation loss climbs back up from its best
- given a smaller learning rate the best validation loss is lower and arrives later
- given log loss boosted trees classify held out moons well
when trees win
- given mixed tabular features a tree ensemble beats even the best prepared small neural network
- given raw unscaled features the neural network does worse than with scaled ones
- given log income and one hot regions the neural network improves on scaling alone
- given engineered features income becomes its logarithm and each region its own column
- given a house bigger than any seen boosted trees predict the same as for the biggest seen
- given the tabular data a feature the label ignores gets the least permutation importance
- given a pure noise feature impurity importance still gives it credit
24. CNNs and RNNs
33 tests in tests/test_cnn_rnn.py
convolution
- given a 5x5 image and a 3x3 filter the feature map is 3x3
- given a vertical edge the vertical edge filter fires 3 where the edge is and 0 on flat bright areas
- given a vertical edge the horizontal edge filter stays silent
- given a straight vertical edge the diagonal filter still fires 2 two thirds of the edge filter
- given one pixel of zero padding a 3x3 filter keeps the 5x5 size
- given stride 2 the filter jumps two pixels and the map shrinks to 2x2
- given a 224 pixel image a 7x7 filter with stride 2 and padding 3 gives 112 positions
- given a color image the filter sums over all three channels
- given the same inputs the result matches pytorch conv2d
pooling
weight sharing
- given a 224x224 color image 64 filters of 3x3 need only 1792 parameters
- given the same image a dense layer with the same output size needs 483 billion weights
receptive field
- given stacked 3x3 filters each layer widens the view by 2 pixels
- given a 2x2 pool between two 3x3 convolutions a final neuron sees 8 pixels
- given two 3x3 filters instead of one 5x5 the same view costs 18 weights instead of 25
image patches as tokens
- given a 224x224 color image and 16 pixel patches a vision transformer sees 196 tokens of 768 numbers
- given a 4x4 image and 2 pixel patches the first token is the top left square
recurrent network
- given not very good the summary after each word is minus 0 762 then 0 119 then 0 785
- given not very good the final summary has nearly forgotten the opening not
- given recurrent weight 0 5 a word 10 steps back has influence 0 5 to the 10th
- given recurrent weight 1 5 the influence explodes to 57 7 after 10 steps
LSTM gates
- given the eraser off and the pen off the notebook keeps its content for 100 steps
- given the eraser fully on the notebook is wiped in one step
- given the pen writing 0 5 on a wiped page the cell holds 0 5 and shows tanh 0 5
- given the highlighter off nothing in the notebook is shown
- given 50 steps an lstm memory carries over 100 times more gradient than a plain rnn
GRU
- given the update gate fully on the hidden state is carried unchanged
- given the update gate off the hidden state becomes the candidate
why transformers won
25. Looking inside the model
38 tests in tests/test_interpretability.py
features are directions
- given two perpendicular feature directions a dot product reads each amount back
- given two features mixed into a hidden state no single neuron equals either amount
linear probes
- given the worked probe the hidden state 2 minus 1 reads as 82 percent positive
- given hidden states that encode sentiment a probe reads it on examples it never saw
- given random labels a probe scores near chance on examples it never saw
- given random labels a probe still fits its own training set better than chance
decodable is not the same as used
- given a feature the models output ignores a probe still reads it
- when the ignored feature is flipped the models output does not move
- when the used feature is flipped the models output moves by two
the logit lens
- given the worked residual stream the lens reads london then paris then paris
- given the worked residual stream the first guess is split 36 40 24
- given the last layer the lens gives paris a probability of 087
- given an unembedding the lens turns any residual state into probabilities that sum to one
- given tinygpt the lens at the last layer reproduces the models own logits
- given tinygpt the lens gives one reading per layer plus the embedding
- given the landmark model the eiffel position reads paris one layer before the last position does
activation patching
- given the clean prompt paris beats rome by three
- given the corrupted prompt rome beats paris by three
- given a patched difference of zero between plus 3 and minus 3 half the answer is restored
- when the eiffel positions mlp output is patched in the answer is fully restored
- when a filler words state is patched in nothing is restored
- given the whole patching map the answer lives at the subject early and at the last word late
- given the eiffel position after the last layer the lens reads paris but patching it restores nothing
superposition
- given five features on a pentagon one active feature leaks 0309 into each neighbour
- given a bias of minus 031 the relu filters the leak to zero
- given two neighbouring features active at once each reads back too high
- given sparse features training stores all five in two dimensions
- given sparse features the five learned directions sit 72 degrees apart
- given sparse features training learns a negative bias to filter interference
- given dense features training keeps only the two most important
- given features in superposition a single neuron responds to three of them
- given the hand derived gradients they match finite differences
sparse autoencoders
- given lambda 01 the one latent explanation costs less than the two latent one
- given the l1 penalty the best activation shrinks below the true amount
- given superposed activations every planted feature is matched by a learned latent
- given the trained autoencoder an active example lights up about one latent
- given a tiny penalty the autoencoder rebuilds the data but misses some features
- given the hand derived gradients they match finite differences
26. Word embeddings
15 tests in tests/test_emb_word2vec.py
skip gram training pairs
negative sampling
skip gram loss
- given orthogonal vectors each true and noise pair costs log two
- given the true context aligned and noise opposed the loss is near zero
- given one update at learning rate 0 1 the true pair scores higher and the loss drops
- given random vectors the hand derived gradient matches a numerical estimate
learned geometry
- given the trained vectors king is closest to a word sharing two of its three attributes
- given king minus man plus woman the nearest word is queen
- given boy minus man plus woman the nearest word is girl
- given prince minus boy plus girl the nearest word is princess
count based embeddings
- given two words that never share a context their ppmi is zero
- given ppmi factored by svd king minus man plus woman is also queen
one vector per word
27. Similarity
28 tests in tests/test_emb_similarity.py
worked example
- given a and b their dot product is eight
- given a and b their cosine is eight ninths
- given a and c their cosine is minus four ninths
- given a and c they are farther apart than a and b
normalization
- given any nonzero vector normalizing gives length one
- given unit vectors squared distance equals two minus twice the cosine
- given unit vectors cosine dot and euclidean rank neighbours identically
- given vectors of very different lengths dot product ranks differently from cosine
signal in the vector length
curse of dimensionality
- given two dimensions the nearest random point is far nearer than the farthest
- given a thousand dimensions the nearest and farthest random points are almost equally far
anisotropy
- given a model with a shared direction unrelated texts still score about three quarters
- given a model with a shared direction every score is squeezed between 0 7 and 0 9
- given mean centering unrelated texts score near zero
threshold calibration
- given labeled pairs the chosen threshold is the one that separates them
- given an anisotropic model a calibrated threshold beats a guessed 0 8
the cosine widget agrees with the lesson
- given pairs of arrows the widget computes the lessons dot products
- given pairs of arrows the widget computes the lessons cosines
- given pairs of arrows the widget computes the lessons euclidean distances
- given an arrow normalizing it in the widget matches the lessons l2 normalize
- given a zero length arrow normalizing it leaves it at zero as the lesson does
- given an angle and a length the widget places the arrow tip by hand trigonometry
the cosine widget presets
- given the same direction preset the cosine is one although the lengths differ
- given the right angles preset the cosine is zero
- given the opposite preset the cosine is minus one
- given the visualization data it is plain json
the lesson places the cosine widget
28. Training embedding models
13 tests in tests/test_emb_contrastive.py
info NCE loss
- given a query matching its passage at temperature one the loss is 0 313
- given the same scores at temperature 0 1 the loss is almost zero
- given random inputs the hand derived encoder gradient matches a numerical estimate
temperature
- given temperature one a positive at 0 8 against three negatives at 0 6 gets only 29 percent
- given temperature 0 05 the same scores give the positive 95 percent
hard negatives
- given only in batch negatives the encoder often confuses look alike passages
- given hard negatives the encoder tells look alike passages apart
- given an unseen password topic the hard negative encoder ranks reset steps above the policy
CLIP
- given two correctly paired images and captions the symmetric loss is 0 313
- given swapped captions the symmetric loss rises to 1 313
- given untrained encoders zero shot labels are near chance
- given contrastive training images are labeled by their nearest caption embedding
- given contrastive training 4 of the 40 pictured test images are closest to another class caption
29. Dimensions and compression
15 tests in tests/test_emb_compression.py
storage math
- given ten million 1536 dim float32 vectors raw storage is 61 gigabytes
- given int8 the same vectors take a quarter of the space
- given one bit per value the same vectors take a thirty second of the space
- given an hnsw graph with m 16 its bottom layer links add 1 28 gigabytes
recall at k
matryoshka truncation
- given vectors ordered by importance the first eighth of dimensions keeps nine in ten neighbours
- given vectors without that order the first eighth of dimensions loses over half the neighbours
- given a short vector shortlist rescored with full vectors nearly every neighbour is recovered
scalar quantization
- given values minus one zero and one in the range minus one to one the codes are 0 128 and 255
- given int8 codes every value comes back within half a step
- given int8 vectors search keeps nearly every true neighbour
binary quantization
30. Vector indexes
35 tests in tests/test_emb_ann.py
flat search
- given three unit vectors the closest one ranks first
- when searching it compares the query with every stored vector
storage math
IVF
- given nprobe equal to nlist it returns exactly what flat search returns
- given one probe it scans a fraction of the corpus
- when nprobe grows recall does not fall
product quantization
- given eight one byte codes per vector a thousand vectors need 8000 bytes of codes
- given 32 dim float32 vectors eight byte codes are 16 times smaller
- when a shortlist is rescored with exact vectors recall recovers
- given residual codes and eight probes ivf pq with rescoring finds most true neighbors
HNSW
- given a thousand vectors it finds at least 90 percent of the true top 10
- when searching it touches under half the corpus
- given m of 8 each layer holds roughly one eighth of the layer below
- given bidirectional links no bottom layer node exceeds 2m neighbors
- when ef search widens recall does not fall and more vectors are compared
- given an empty index a search returns nothing
HNSW search path
- given a query the search starts on the top layer
- given a query the search finishes on the bottom layer
- when descending each layer starts closer to the query than the one above
- given 300 points and m of 6 the graph has four layers of 300 64 14 and 1 nodes
- given the figures query the search compares it with 51 of the 300 points
eight point map
- given the query at 5 4 greedy hnsw takes the highway to e then a side street to h
- given two clusters ivf probes only the right hand cluster for the query at 5 4
- given two sub vectors and four centroids each pq stores a 4d vector as two codes
- given those codes the table lookup score is 2 while the exact score is 1 7
the small HNSW map
- given 60 points and m of 4 the map has layers of 60 15 3 and 2 nodes
- given a beam of one query d stops short of its true nearest point
- given a beam of four query d finds its true nearest point
- given the exported widget data it is plain json under 50 kb
the HNSW widget agrees with the lesson
- given each query and beam width the widget expands the same nodes in the same order
- given each query and beam width the widget returns the lessons nearest point
- given each query and beam width the widget counts the indexs distance computations
- given each query the widgets brute force agrees with the flat index
- given the exported graph it is the lessons index link for link
- given the ann lesson it places the hnsw search widget after a try it paragraph
31. Retrieval
54 tests in tests/test_emb_retrieval.py
BM 25
- given a word in one of three documents its rarity weight is about 0 98
- given a word in two of three documents it weighs less than a rarer word
- given the query cat the document mentioning it twice scores about 1 21
- given a word repeated 100 times its credit stays below k1 plus one
- given equal mentions the shorter document scores higher
- given a word no document contains every score is zero
- when searching documents with zero score are left out
- given an exact error code the matching article ranks first
reciprocal rank fusion
- given ranks 1 and 3 the fused score is about 0 0323
- given rank 1 in one list only the fused score is about 0 0164
- given two lists a document both agree on outranks one that tops a single list
hybrid search
- given an exact error code query dense search leaves the answer out of its top 3
- given an exact error code query bm25 ranks the answer first
- given a synonym query with no shared words bm25 finds nothing
- given a synonym query dense search ranks the car expense article first
- given the error code query hybrid search ranks the answer first
- given the synonym query hybrid search ranks the answer first
retrieval quality
- given the evaluation set bm25 finds 15 of 20 answers in its top 3
- given the evaluation set dense search finds 18 of 20 answers in its top 3
- given the evaluation set hybrid search finds all 20 answers in its top 3
- given answers at ranks 1 2 and missing mean reciprocal rank is one half
cross encoder reranking
- given password rules hybrid search ranks the reset how to above the policy
- given password rules the policy article covers both ideas in the query
- given password rules the policy scores about 1 78 and the reset how to about 0 69
- when reranked the password policy ranks first
- given forgotten sign in credentials hybrid search ranks the vpn guide first
- when reranked the password reset how to ranks first
- given a shortlist of 10 the cross encoder reads exactly 10 query document pairs
- given how many vacation days the literal word vacation lifts the pto article above sick leave
- given refund for the hotel the reranker prefers the article that talks about reimbursing trips
- given the evaluation set reranking keeps all 20 answers in the top 3
late interaction
- given two query tokens each takes its best matching document token and the best scores add up
- given a two word title late interaction stores two vectors where a bi encoder stores one
- given password rules late interaction ranks the policy above the reset how to
chunking
- given 130 words 50 word chunks with 10 words overlap make 3 chunks
- given 10 words overlap consecutive chunks share exactly 10 words
- given the handbook no structure aware chunk spans two sections
- given the handbook each chunk carries its source section date and access list
- given the stipend question the best fixed size chunk loses the amount
- given the stipend question the best structure aware chunk holds the heading and the amount
- when a small child chunk matches the whole parent section is returned
query and passage prefixes
- given both prefixes used correctly 11 of the 12 everyday questions find their answer in the top 3
- when the passage prefix is forgotten at indexing time recall silently drops
the RAG widget agrees with the lesson
- given several sizes and overlaps the widgets chunker cuts exactly the lessons chunks
- given each question and chunking the widget marks exactly the chunks that hold the whole answer
- given each question and chunking the widgets first stage order is the lessons hybrid search
- given each question and chunking the widgets reranker scores are the lessons cross encoder
- given every top k the widgets reranked list is the lessons retrieve then rerank
- given 30 word chunks without overlap no chunk holds the whole vpn answer
- given 30 word chunks with 10 words of overlap one chunk holds the whole vpn answer
- given the stipend amount question hybrid search ranks a chunk without the amount first
- when the stipend shortlist is reranked the chunk with the amount ranks first
- given every setting the widgets data stays small enough to load with the page
- given the retrieval lesson it places the rag pipeline widget
32. Clustering and matching
22 tests in tests/test_emb_clustering.py
k means
- given two obvious pairs of points k means puts each pair in its own cluster
- given two obvious pairs the centres land halfway between each pair
- given two obvious pairs each round never increases the total squared distance
- given two obvious pairs the total squared distance to centres is one
silhouette
- given two tight far apart pairs the silhouette is 0 90
- given five kinds of support ticket the best silhouette picks five clusters
density clustering
- given two dense runs and a straggler dbscan finds two clusters and one noise point
- given support tickets dbscan marks exactly the off topic tickets as noise
- given too small a reach dbscan strands genuine tickets as noise
- given too large a reach dbscan chains different kinds into one cluster
- given a reach sweep from 0 45 to 0 85 clusters and noise follow the lessons figure
PCA
- given points on a straight line the first direction explains all the variation
- given points on a straight line their 2d coordinates are spaced by root two
practical uses
- given a reworded copy of a ticket near duplicate detection pairs them
- given a vpn complaint the router sends it to the it helpdesk
- given a receipt question the router sends it to finance
- given an off topic question the router falls back to a human
- given ticket clusters the off topic tickets are farthest from every centre
semantic cache
33. Embeddings in production
20 tests in tests/test_emb_operations.py
models have their own spaces
- given the same word two different models give unrelated vectors
- given queries and documents from the same model retrieval finds nearly every answer
- given a same size new model querying old document vectors silently drops to chance
- given v2 queries over v1 documents exactly 3 of the 12 golden questions find their answer
- given a new model with more dimensions querying old document vectors fails loudly
blue green migration
- given a better new model cutover points live traffic at the new index
- given a worse new model cutover is refused and live traffic stays put
- given no shadow comparison cutover is refused
- given a rollback after cutover live traffic returns to the old index
- given a document added during the migration both indexes can find it
reembedding cost
- given fifty million documents of 500 tokens at a million tokens a second it takes about seven hours
- given two cents per million tokens re embedding fifty million documents costs 500 dollars
domain mismatch
- given company jargon a general model misses most of the answers
- given a model tuned on the jargon it finds every answer
building an eval set
- given a search log only resolved sessions become eval cases
- given the same question resolved by two documents both count as relevant
failure triage
34. Autoencoders and VAEs
26 tests in tests/test_autoencoders.py
the stroke images
- given the toy dataset each image is an 8 by 8 picture of one of four stroke kinds
- given any stroke its brightest pixel holds at least 0 707 of full ink
squeezing into a code
- given a point on the line y equals 2x its one number code rebuilds it exactly
- given the point 2 3 off the line it is rebuilt as 1 6 3 2 with squared error 0 2
- given training the reconstruction error falls to under a tenth of where it started
- given stroke images a two number code rebuilds them far better than two principal components
- given trained codes each strokes nearest neighbour in code space is the same kind of stroke
PCA is the linear special case
- given points near y equals 2x a linear autoencoder learns the lines direction
- given the same points a linear autoencoder rebuilds them as well as one principal component
holes in a plain autoencoders code space
- given codes drawn from a standard normal a plain autoencoder decodes junk
- given codes drawn anywhere inside its own code range a quarter or more decode to junk
the reparameterization trick
- given mean half spread half and noise 1 2 the code is 1 1
- given many noise draws the codes average to the mean and spread by the spread
- given the noise held fixed backprop through the sampling step matches finite differences
the KL penalty
- given mean 0 and spread 1 the code already matches the standard normal and pays nothing
- given mean half and spread half the penalty is 0 443
- given two coordinates their penalties add
- given the closed form it agrees with an average over samples
sampling new strokes
- given codes drawn from a standard normal a vae decodes images close to real strokes
- given a trained vae its codes crowd around zero with spread about one
- given a walk between two strokes it starts and ends at their reconstructions
- given walks between random pairs the vaes images change in smaller steps than a plain autoencoders
the beta trade off
35. GANs
28 tests in tests/test_gans.py
the two players
- given a batch of noise the forger returns one 2d point per noise vector
- given any points the detective turns its score into a probability between 0 and 1
- given a hand written backward pass it matches gradients measured by nudging each number
- given the toy data every sample lies near one of eight points on a circle of radius 2
the minimax value
- given real verdicts 0 9 and 0 8 and fake verdicts 0 2 and 0 4 the value is minus 0 53
- given a detective that is never wrong the value reaches its ceiling of zero
- given a detective that always says one half the value is minus log 4
the best possible detective
- given real density 0 3 and fake density 0 1 at a point the best verdict is 0 75
- given a forger that matches the data the best detective says one half everywhere
- given a fixed forger a trained detective arrives at the formula for the best verdict
why the forger needs a non saturating loss
- given a detective 99 percent sure a fake is fake the original loss gives the forger a gradient of only 0 01
- given an undecided detective both losses give the same gradient
- given a detective with a long head start the original loss starves the forger while the non saturating one does not
oscillation
- given theta 1 and psi 0 one step of size 1 moves the detective to minus one half and leaves the forger
- given theta 1 and psi 0 the second step moves the forger towards the data at 0
- given plain simultaneous steps the players spiral away from the equilibrium
- given alternating steps the players circle the equilibrium without ever settling
mode collapse
- given samples piled on three modes coverage counts three
- given samples far from every mode none count as high quality
- given a detective that learns 3 times faster than the forger all eight modes are covered
- given equal learning rates the forger collapses onto at most half the modes
- given equal learning rates the forger hops from one set of modes to another
- given a detective that learns 10 times faster the forger gets stuck on part of the ring
stabilizers
- given an r1 gradient penalty the players spiral in to the equilibrium
- given two single point piles js divergence is log 2 however far apart they are
- given a pile shifted 3 to the right the wasserstein distance is 3
- given a matrix that stretches one direction by 3 its largest stretch is 3
- given spectral normalization the largest stretch becomes 1
36. Diffusion and flow matching
32 tests in tests/test_diffusion.py
the noising process
- given alpha bar 0 64 a point at 2 with noise 0 5 lands at 1 9
- given step sizes 0 1 and 0 2 the signal left after two steps is 0 72
- when noise is added twice in small steps the result matches the one jump shortcut
- given the default schedule less than one percent of the signal survives the last step
- when the blobs are noised to the last step they look like plain standard noise
learning to denoise
- given the hand written backward pass its gradients match finite differences
- given an untrained denoiser its first score is what guessing zero scores
- given a trained denoiser it guesses the noise far better than guessing zero
- given a noise guess of 0 9 at alpha bar 0 64 the score is minus 1 5
sampling backwards
- given one denoising step with no fresh noise the point moves from 1 to 0 9
- given one denoising step with fresh noise 0 5 the point lands at 1 118
- given a perfect noise guess a deterministic jump to alpha bar 1 recovers the clean point
- given a deterministic jump to alpha bar 0 96 the point lands at 2 06
- given pure noise a hundred small steps land the samples on the blobs
- given pure noise that was never denoised it sits far from the blobs
- given only twenty deterministic jumps the samples still land on the blobs
- given the same starting noise the deterministic sampler draws the same points
flow matching
- given noise at minus 1 and data at 2 a quarter of the way along the point is at minus 0 25
- given noise at minus 1 and data at 2 the velocity to learn is 3 everywhere on the line
- given the true velocity one euler step of 0 75 arrives exactly at the data
- given ten euler steps the learned flow lands the samples on the blobs
- given only five steps the flow lands closer to the blobs than diffusion does
- given a single euler step every sample lands near the empty average of the data
guidance
- given guesses 0 2 without and 0 5 with the label a weight of 3 gives 1 1
- given a weight of 0 the guided guess ignores the label
- given a weight of 1 the guided guess is the labelled guess
- given the south west label nearly every sample lands in the south west blob
- given a weight of 0 samples spread over all four blobs
- given a stronger guidance weight the samples bunch more tightly
scaling up to images
37. Multimodal models
41 tests in tests/test_multimodal.py
cutting an image into patches
- given a 4 by 4 image and 2 pixel patches it becomes 4 tokens of 4 numbers
- given the worked image the top right patch reads 2 3 6 7
- given a side the patch does not divide it refuses
counting image tokens
- given 224 pixels and 16 pixel patches an image is 196 tokens
- given 336 pixels and 14 pixel patches an image is 576 tokens
- given neighbouring patches merged 2 by 2 the count falls fourfold
- given twice the side length the token count quadruples
patch embedding
- given the worked projection and positions the tokens are the hand computed ones
- given two identical patches in different places their tokens differ
the vision encoder
- given a 32 pixel colour image and 8 pixel patches it returns 16 vectors
- given a change in the last patch the first patch output changes too
the projector
- given the worked example vector 1 2 it projects to 1 2 3
- given vision width 32 and language width 48 every image token comes out 48 wide
interleaving image and text
- given one image placeholder the sequence grows by the image tokens minus one
- given the placeholder the projected image tokens sit exactly where it was
- given an image and a prompt the model scores every vocabulary entry at every position
- given a different image the text before it is scored the same
aligning vision to language
- given an untrained projector captions are little better than chance
- given projector training on image caption pairs the right caption word wins
- given alignment training the language model and vision encoder stay frozen
- given alignment training the caption loss falls
how much of each pitch
- given a 2 cycle wave over 8 samples all its energy lands in bin 2 and its mirror
- given any signal the magnitudes match numpys fast fourier transform
the spectrogram
- given one second at 16 khz with 25 ms windows every 10 ms there are 98 frames
- given a steady 1000 hz tone every frame peaks in the bin nearest 1000 hz
- given a rising chirp the loudest frequency climbs frame by frame
the mel scale
- given 1000 hz it is 1000 mel by construction
- given 700 hz it is about 781 mel
- given a filterbank the high pitch filters are wider than the low ones
counting audio tokens
- given 30 seconds with 10 ms frames halved once the encoder sees 1500 tokens
- given one second the encoder sees 50 tokens
vector quantization
- given the point 0 9 0 2 its nearest codebook entry is number 1
- given ids decoding returns the codebook vectors they name
- given a larger learned codebook the reconstruction error falls
- given a learned codebook it beats a codebook of random points
- given a second codebook for the leftover error the reconstruction improves
video tokens
- given 10 seconds at 30 frames per second and 256 tokens a frame it is 76800 tokens
- given one frame a second the same clip is 2560 tokens
- given tubelets two frames deep the count halves
- given 300 frames at 30 fps sampling one a second keeps every 30th frame
the context budget
38. Talking to a model
21 tests in tests/test_agents_llm.py
a model reply
- given a policy that answers the reply ends the turn with text
- given a policy that wants a tool the reply is a request not a result
- given a tool request without an id the model assigns one so results can be matched
- given a fixed script each call returns the next line then repeats the last
- given the reply the assistant content holds the exact blocks to send back
tokens
- given 4000 characters of prose the estimate is about 1000 tokens
- given 1500 more characters of history the next call bills 375 more input tokens
- given a ten call loop starting at 2000 tokens and adding 500 per step it bills 42500
- given the same loop the tenth call re sends 6500 tokens of history
reading a conversation
- given tool results in the latest user turn the last user text is still the question
- given a conversation its tool results are found in order
- given a failed tool the result is flagged as an error the model can act on
the tool call round trip
- given a weather question the round trip takes two model calls and one tool run
- given the round trip the tool result answers the request with the same id
the real request
- given low effort the request asks the model to think briefly
- given no effort the request leaves the models default
- given fallbacks and effort the request carries both
a local model through ollama
39. Orchestration
24 tests in tests/test_agents_orchestration.py
prompt chaining
- given an outline that passes the gate the chain writes the draft from it
- given an outline that fails the gate the chain stops before drafting
routing
- given a billing question the classifier sends it to the billing handler
- given a classifier answer outside the allowed labels the request goes to the fallback
- given a label with different case and spacing it still matches
parallelization
- given two independent sections they run at the same time
- given three reviewers where two say vulnerable the majority verdict is vulnerable
- given three independent voters each right 80 percent of the time the majority is right 89 6 percent
orchestrator workers
- given a request with three questions the orchestrator splits it into three subtasks
- given three subtasks each worker answer is grounded in its own policy document
- given the worker answers the final reply cites all three sources
evaluator optimizer
- given a first draft without a citation the editor sends it back and the second draft passes
- given an editor that never approves the loop stops after three rounds and says so
supervisor
- given a question needing hr and finance facts the supervisor gives one task to each specialist
- given delegated tasks each specialist sees only its own task
- given a delegation to a specialist that does not exist it is skipped and reported
state machine
- given an invoice under the approval limit it moves through every state to done
- given an invoice over the approval limit the run pauses for a human and posts nothing
- given a document that is not an invoice the run is rejected without extracting anything
- given an extracted amount that is not positive the run stops for review
- given a crash while posting resuming finishes without asking the model again
- given a paused large invoice a human approval posts it and finishes
autonomy costs
figures
40. The agent loop
22 tests in tests/test_agents_agent_loop.py
completing a task
- given a model that answers after one round of tools the run completes with that answer
- given two lookups requested in one turn both results return in a single user message
- given two tools that each wait for the other they both succeed because they run concurrently
runaway controls
- given a model that repeats the same tool call the loop stops it as a loop
- given the same arguments in a different key order the calls count as one repeated call
- given a model that never finishes the run stops at the step limit
- given a token budget a run that exceeds it stops as budget exceeded
- given a dollar budget a run that exceeds it stops as budget exceeded
- when every call resends the history each step sends more input tokens than the last
tool errors
- given a badly formatted date the model receives an error that shows the valid format
- given an actionable error the model corrects its call and completes
- given a tool name the model invented it is told which tools exist
- given a tool that crashes the run continues with an error result
- given a tool that raises a tool error the model sees its message verbatim
stop reasons
- given a refusal the run stops as a refusal without running tools
- given a truncated response its half written tool call is not run
handing off to a human
cost
- given one million input and 100k output tokens at 5 and 25 dollars the cost is 7 50
- given a tool result the transcript stores it as json the model can read
stop causes
figures
quadratic growth
41. Tools
35 tests in tests/test_agents_tools.py
tool definitions
- given a registered function its definition has the name description and input schema
- given strict mode each definition is marked strict and forbids extra fields
argument validation
- given a missing required argument the error names it
- given text where a whole number is expected the error names the field and both types
- given true where a whole number is expected it is rejected as a boolean
- given the lessons misspelt payment every problem comes back at once with its fix
- given a value outside the allowed set the error lists the allowed values
- given a date in the wrong format the error shows the expected format
- given an amount below the minimum the error states the bound
- given a quantity above the maximum the error states the bound
- given a misspelled field the error names it and lists the allowed fields
- given a list with one wrong item the error points at that item
- given valid arguments there are no errors
calling a tool
- given valid arguments the tool runs and its result comes back as text
- given invalid arguments the tool does not run and every problem is reported at once
- given a tool name that does not exist the reply lists the tools that do
semantic validation
least privilege
- given credentials that only read tickets a refund is denied and never runs
- given credentials with the right scope the refund runs
idempotency
- given a retried refund with the same idempotency key the money moves once
- given two different idempotency keys both refunds happen
dry run
human approval
- given an irreversible action and no approver it waits for approval and does not run
- given an approver who declines the refund does not run and the model is told why
- given an approver who agrees the refund runs
- given a rule that refunds over 500 need approval a refund of 40 runs without asking
- given a rule that refunds over 500 need approval a refund of 900 waits
tool descriptions
- given vague overlapping descriptions a model picks the right tool for 2 of 6 requests
- given precise descriptions that say when to use each tool it picks right for 6 of 6
fewer higher level tools
- given calls that each succeed 97 percent of the time five in a row succeed about 86 percent
- given one high level call at 97 percent the whole job succeeds 97 percent of the time
dynamic tool loading
- given a catalog of 30 tools a locked out request loads the password reset tool
- given a limit of 5 only 5 tool definitions are sent to the model
- given an expense request the best match is the expense tool
figures
42. Coding and computer-use agents
58 tests in tests/test_agents_coding_agents.py
why code suits agents
- given a 40 percent fix rate and tests to check each try three tries succeed 78 percent of the time
- given a single try the chance is just the fix rate
- given a failing test the report names the call what it returned and what was expected
running the tests
- given the buggy median the odd length cases pass and the even length ones fail
- given code that raises the report names the exception and its message
- given a patch with a syntax error the report names the file and line
- given two cases each runs in a fresh namespace so one cannot leak into the next
editing files
- given old text that appears once the edit replaces it
- given old text that is not in the file the error says to copy the lines exactly
- given old text that matches twice the error asks for more surrounding lines
- given a file that does not exist the error lists the real ones
the edit run test loop
- given a model that tests searches reads and patches the tests go green at step 7
- given the careful run its steps are test search read edit test edit test
- given the first patch is off by one the next test run names the index error
- given the careful run the passing count goes two two four
- given a model that says fixed without running the tests the harness does not take its word
- given a model that claims success the harness replies with the failing tests
- given a model that never fixes anything the run stops at the step budget
finding the right files
- given a search for a definition it returns the file and line number
- given more matches than the cap the rest are counted not dumped
- given the careful run the only file it reads is stats py
- given the careful run it puts under a fifth of the repository into context
- given 500 files of 400 tokens a dump is 200000 tokens and search then read is 850
the sandbox
- given a well behaved function it returns its value
- given an infinite loop it is stopped on the first line past the budget
- given a loop that catches every exception it is still stopped
- given a wall clock limit a runaway loop is stopped by time
- given code that keeps allocating it is stopped at the memory limit
- given code that imports a module it fails before it can reach the network
- given code that opens a file it fails because open is not provided
- given the restricted namespace introspection still reaches every class the interpreter loaded
- given the sandbox finishes the previous tracer is restored
evaluating coding agents
- given a correct patch every hidden test passes so the task is resolved
- given a patch that special cases the visible inputs the visible tests pass
- given a patch that special cases the visible inputs the hidden tests leave it unresolved
- given a patch that fixes 1900 but breaks 2000 the pass to pass tests leave it unresolved
- given the correct leap year rule the task is resolved
- given the example agent two of the three tasks are resolved
- given 10 samples with 3 correct pass at 1 is 0 3
- given 10 samples with 3 correct pass at 5 is 1 minus 21 over 252
- given no correct samples pass at k is zero at any k
- given fewer wrong samples than draws every draw holds a correct one
- given ten attempts at 60 cents with four resolved each resolved task costs 1 50
- given nothing resolved the cost per resolved task is infinite
driving a screen
- given the signup form its screenshot shows each control as a row of text
- given a click on a field it takes focus and typed text lands in it
- given a click on empty space nothing happens and no error is raised
- given an agent that looks before each action it completes the form
- given the looking agent it takes 8 model calls and 7 screenshots
- given a layout shift the agent that looks again still completes the form
- given the original layout replayed coordinates complete the form
- given a layout shift replayed coordinates miss and nothing is submitted
- given an instruction written on screen an unguarded gullible agent deletes the account
- given a guard on destructive controls the injected click is blocked
- given the blocked click comes back as an error the agent returns to its task and finishes
- given a 1280 by 800 screenshot cut into 32 pixel patches it costs 1000 tokens
- given six actions at that size the run sends 7000 image tokens
figures
43. Model Context Protocol
23 tests in tests/test_agents_mcp.py
handshake
- given an initialize request the server replies with its version capabilities and name
- given a notification the server sends no reply
- given a request before initialize the server refuses it
- given a method the server does not know it replies method not found
tools
- given a tools list request the server describes each tool with an input schema
- given a tools call the server runs the tool and returns text content
- given a tool that fails the failure comes back as a result the model can read
- given a call to a tool that does not exist the server replies invalid params
resources
- given a resources list request the server names its readable documents
- given a resources read request the server returns the document text
client
- given a client connected over json lines it discovers the servers tools
- given a client session every line on the wire is one json rpc 2 0 message
- given mcp tool definitions the host converts them to the anthropic tool format
tool poisoning
- given a description with hidden orders the scanner names each warning sign
- given an ordinary description the scanner finds nothing
- given invisible characters in a description the scanner flags them
rug pulls
- given tool definitions that have not changed since approval nothing is flagged
- given a server that rewrites a description after approval the change is flagged
- given a server that adds a tool after approval the new tool is flagged
confused deputy
- given a server acting with its own admin account a read only user gets a ticket deleted
- given a server acting with the users own delegated permissions the same request is refused
integration count
figures
44. Retrieval-augmented generation
38 tests in tests/test_agents_rag.py
parsing messy documents
- given a header repeated on every page it is removed
- given page number footers they are removed
- given a word hyphenated across a line break it is rejoined
- given a table each row becomes a self contained sentence
chunking
- given a document every passage carries its access groups and date
- given a short word limit the error document splits into its two sentences
hybrid retrieval
- given an exact error code dense search alone misses the document
- given an exact error code hybrid search brings the document into the top three
- given a question sharing no words with the answer keyword search alone misses it
- given a question sharing no words with the answer hybrid search finds it
permission aware retrieval
- given a user outside finance restricted documents never reach the model
- given a finance user the restricted forecast ranks first
- given filtering after generation the restricted figure has already leaked into the answer
- given filtering before generation the restricted figure never appears
grounded generation
- given a covered question the answer cites a passage from the context
- given a correctly cited answer no claim is flagged as unsupported
- given a citation attached to a claim its source does not support the claim is flagged
- given a question the sources do not cover the model says it does not know
- given an answer citing a source that was never retrieved verification flags it
metadata filtering
- given no date filter the superseded 2023 travel policy ranks first
- given a filter on documents updated since 2024 the superseded policy never appears
contextual retrieval
- given plain chunks the fix sentence is not found for how do i fix err 4012
- given chunks prefixed with their document title the fix sentence is found
- given three phrasings of the question the titled fix sentence ranks second second and fourth
query rewriting
- given a follow up on its own retrieval goes to the printer guide
- given the follow up rewritten with the earlier topic retrieval finds vpn help
hy DE
- given a question in the users own words plain dense search misses the phishing guide
- given a hypothetical answer to search with the phishing guide ranks first
multi query
- given one informal phrasing the two factor guide is missed
- given several phrasings fused the two factor guide is found
parent child retrieval
agentic RAG
- given a two part question the agent searches twice and cites both sources
- given small talk the agent does not search at all
graph RAG
- given the knowledge base two factor connects to email vpn and anyconnect
- given the graph each connection remembers which document stated it
debugging wrong answers
45. Context engineering
28 tests in tests/test_agents_context.py
token budget
- given everything fits no section is dropped
- given too much content the lowest priority section is dropped first
- given a section that does not fit it is skipped and a smaller one after it still gets in
- given the budget the assembled context never exceeds it
fitting pieces into the window
- given a window with room to spare every piece is kept
- given a window just too small only the oldest turn is cut
- given old turns are not enough the lowest ranked documents go next
- given a tight window the recent turns outlast every document
- given must haves larger than the window they are kept and the request does not fit
cache friendly ordering
turn summarization
- given ten turns and keep four the history becomes one summary plus four turns
- given ten turns the first entry is a summary of the six oldest
- given ten turns the last four are kept verbatim
- given fewer turns than the window nothing changes
tool output compression
- given the fields the task needs only those are kept
- given a field that does not exist it is skipped rather than invented
delimiting data from instructions
prefix caching
- given an identical prompt twice the second is served entirely from cache
- given a timestamp at the top the next request reuses nothing
- given the timestamp at the end the next request reuses the whole system prompt
lost in the middle
the context widget agrees with the lesson
- given parts that fit comfortably the widget keeps what the lesson keeps
- given parts just over the window the widget cuts what the lesson cuts
- given parts far over the window the widget trims documents as the lesson does
- given uneven piece sizes the widget cuts in the lessons order
- given must haves larger than the window the widget reports what the lesson reports
- given the context lesson it places the context budget widget
- given the widgets starting sizes they are plain json under its name
46. Memory
25 tests in tests/test_agents_memory.py
short term memory
- given a conversation within budget the model sees every message and no summary
- given a conversation over budget the older turns become a one line summary
- given a forty exchange conversation what is sent never exceeds the budget
- given a conversation over budget the most recent messages are kept word for word
what is worth remembering
- given small talk nothing is stored
- given a message containing a password it is refused
- given a lasting fact about the user it is stored as semantic memory
recall
- given three kinds of memory a question about globex recalls the rejection first
- given a request for how to do something only procedural memory is searched
conflicting facts
- given a fact is corrected recall returns only the new value
- given a fact is corrected its history keeps the old value marked superseded with its source
tenant isolation
- given a query that quotes another tenants memory word for word it finds nothing
- given a query crafted to look like a filter it is treated as plain text
- given two users in the same company one cannot recall the others memories
- given another tenants memory id forgetting it is refused
- given a scope without a tenant it cannot be created
user rights
- given a user asks what is remembered about them the export lists every memory including superseded ones
- given a user asks to be forgotten all their memories are erased
- given one user is erased their colleagues memories remain
task state outside the model
- given a new task the first unfinished step is the first step
- given the model proposes finishing the current step the code records it and moves on
- given the model proposes finishing a later step while an earlier one is open the update is refused
- given the model proposes a status that does not exist the update is refused
- given a crash after two steps a fresh process resumes at the third
figures
47. Planning
20 tests in tests/test_agents_planning.py
compounding error
- given steps that each succeed 95 percent of the time ten in a row succeed about 60 percent
- given steps that each succeed 95 percent of the time twenty in a row succeed about 36 percent
verification and retries
- given 95 percent steps a check that catches every failure and one retry make each step 99 75 percent
- given perfect checks and one retry ten steps succeed about 97 5 percent instead of 60
- given a check that misses half the failures ten steps succeed only about 77 percent
checkpoints
- given checkpoints finishing ten 95 percent steps takes about 10 5 step runs
- given no checkpoints a failure restarts from step one and it takes about 13 4 step runs
- given fifty 95 percent steps restarting from scratch costs about 240 step runs
decomposition
- given the q3 invoices and payments the reconciliation lists the short payment and the unpaid invoice
- given a finished reconciliation the summary names both mismatches
- given a payments fetch that times out once only that step runs again
- given a step that keeps failing its check the run stops there and later steps never run
plan and execute
- given a step that fails because its api was retired the planner revises the plan and the goal is met
- given a revised plan work already completed is not repeated
- given a planner that never adapts the run gives up after the replan limit
reflection versus external checks
- given a leap year function that forgets century years the models self review approves it
- given the same function running real test cases catches the year 1900
- given the corrected function every test case passes
- given test failures fed back to the model its second draft passes
figures
48. Evaluation
25 tests in tests/test_agents_evals.py
code graders
- given answers differing only in case and spacing exact match passes
- given different answers exact match fails
grading outcomes not paths
- given the right end state reached by a different tool order the run passes
- given a correct answer but a forbidden tool call the run fails
- given a refund above the limit escalating instead of refunding passes
- given a run that exceeds the step limit it fails even if the outcome is right
trajectory metrics
- given two required tools and only one called tool selection accuracy is one half
- given all required tools called tool selection accuracy is one
judge calibration
- given the worked example raw agreement is eighty percent
- given the worked example cohens kappa is 0 583
- given a judge that always says pass kappa is zero despite sixty percent agreement
- given kappa below 0 6 the judge is flagged as not trustworthy
- given the rubric judge an answer naming the order and its status passes
- given the rubric judge a vague answer fails
- given the always pass judge even a vague answer passes
retrieval and faithfulness
- given the relevant document at rank three recall at three is one
- given the relevant document at rank three recall at two is zero
- given one invented claim out of two faithfulness is one half
percentiles
regression gate
- given the same version twice the release gate passes
- given a prompt change that breaks the over limit rule the gate blocks the release
- given version one every golden task passes
- given an evaluation cost is reported per successful task
online signals
49. Guardrails
33 tests in tests/test_agents_guardrails.py
injection heuristics
- given an email saying ignore previous instructions the detector flags an override
- given an ordinary business email the detector finds nothing
- given the same attack paraphrased the detector misses it
luhn checksum
- given the standard visa test number the checksum passes
- given one digit changed the checksum fails
- given fewer than thirteen digits it is not a card number
PII redaction
- given an email address it becomes a typed placeholder
- given a phone number it becomes a typed placeholder
- given a valid card number it becomes a typed placeholder
- given a sixteen digit order number that fails luhn it is left alone
schema validation
- given a well formed refund there are no errors
- given an amount sent as a string the error names the field and expected type
- given a missing field the error names it
- given an unexpected field it is rejected
- given a boolean where a number belongs it is rejected
output policy
- given an answer that promises a guarantee the policy flags it
- given an answer that reveals a password the policy flags it
groundedness
- given an answer restating the source every claim is supported
- given one invented claim out of two half the answer is grounded
- given a grounded answer containing an email address the layered check blocks it
action policy
- given a tool the agent was not granted the call is denied
- given a small refund it is allowed
- given a refund over the approval threshold a human must approve
- given refunds that together exceed the spend cap the last one is denied
- given a refund over the sign off threshold that would also break the cap it is denied not escalated
- given an email to an outside domain a human must approve
- given an irreversible tool a human must approve even with no amount
privilege separation
- given a single agent that reads email and can send the injection exfiltrates invoices
- given reader and actor are separated nothing is sent to the attacker
- given reader and actor are separated the injected forward request is blocked and visible
- given the user only asked for a summary even a benign reply waits for the user
- given four phrasings of one attack the detector catches only the blunt and hidden comment ones
- given every attack variant only the naive design leaks
50. Cost and latency
29 tests in tests/test_agents_cost.py
request pricing
- given 10k input and 500 output tokens on the large model the call costs 6 25 cents
- given 8k of the input served from cache the same call costs 2 65 cents
- given the batch interface the same call costs half
- given equal token counts output costs five times input
model routing
- given a short classification task it goes to the small model
- given a multi step planning task it goes to the large model
- given the sample workload routing cuts spend by more than half
response caching
- given the same question with different spacing and case the exact cache hits
- given a paraphrase the exact cache misses
semantic caching
- given a paraphrase above the threshold the semantic cache returns the stored answer
- given a different question about the same topic a high threshold still returns the wrong answer
- given different error codes the identifier guard prevents a wrong hit
- given a negated request the negation guard prevents a wrong hit
- given an entry older than its time to live it is not served
- given a lower threshold more wrong answers are served
- given the labelled pairs there are both paraphrases and near misses
parallel tool calls
- given three independent 100ms tools running them together takes about one tool time
- given parallel calls the results come back in request order
budgets
- given a three step budget the fourth step is refused
- given a token budget a step that would exceed it is refused
- given a tenant over its monthly cap the spend tracker raises a cap alert
- given a task using ten times the usual tokens the tracker raises an anomaly alert
unit economics
- given retries until success a cheap model that succeeds 60 percent costs 0 33 cents per success
- given human cleanup of failures the cheap model costs 80 cents per task
- given human cleanup of failures the large model costs 11 cents per task
five times plan
- given the sample workload the ranked levers together cut cost at least five fold
- given the plan the free levers come before the model change
the parallel figure is stable
the semantic cache sweep
51. Observability
18 tests in tests/test_agents_observability.py
span tree
- given a run with two model calls and one tool call the root has three children in order
- given a child span its parent is the root
- given code that raises inside a span the span is marked as an error and the exception still propagates
gen AI attributes
- given a model call the span records the model and token usage
- given a tool call the span records the tool name and call id
- given a model call the span records which prompt version produced it
totals
- given two 400ms model calls and one 100ms tool call the run takes 900ms
- given the run the totals count two model calls and one tool call
debugging from a trace
- given a run where the lookup failed the first error is the tool span
- given the lookup failed the agent tells the user it could not find the order
- given a successful run there is no error span
waterfall
rendering
- given a trace the first line is the root with its duration
- given a trace children are drawn as branches with the last one closing the tree
open telemetry export
- given an export every span shares the trace id
- given an export children point at the root span
- given an export timestamps are in nanoseconds
feedback to golden set
52. Safe deployment
29 tests in tests/test_agents_deployment.py
shadow mode
- given eight matching decisions out of ten overall agreement is eighty percent
- given the same run refund decisions agree three times out of five
- given the same run reply decisions agree every time
graduated autonomy
- given a new agent it starts in shadow mode
- given 48 of 50 shadow decisions matching humans it is promoted to human approval
- given 46 of 50 shadow decisions matching it stays in shadow mode
- given 49 of 50 actions approved by reviewers it may act alone on low risk actions
- given autonomy high risk actions still need a person
- given quality drift while autonomous it drops back to human approval
- given an incident while autonomous it drops back to human approval
prompts as code
- given the same prompt text twice it gets the same version
- given an edited prompt it gets a new version
- given a rollback the previous text is active again
- given a version its name carries the first eight hex digits of the text hash
canary assignment
- given the same user the assignment never changes
- given ten percent about one user in ten is in the canary
- given a user in the 5 percent canary they stay in it at 25 percent
canary rollout
- given the canary matches control the rollout advances to the next step
- given the canary is five points worse the rollout rolls back to zero
- given a healthy version the rollout passes through 1 5 25 50 and 100 percent
- given too few canary samples the rollout waits
kill switch
- given a tenant switched off only that tenant is blocked
- given the global switch every tenant is blocked
rate limiting
- given a full bucket of five a burst of five actions is allowed
- given an empty bucket the sixth action is refused
- given two seconds of refill at one per second two more actions are allowed
tamper evident audit log
53. Why the hard ones fail
17 tests in tests/test_agents_failures.py
compounding error
- given ten steps each 95 percent reliable the whole task succeeds about 60 percent of the time
- given twenty such steps the task succeeds about 36 percent of the time
- given 95 percent steps and a 90 percent target at most two steps fit
retry with backoff
- given two timeouts then success the call succeeds on the third attempt
- given two timeouts the waits double from the base delay
- given many failures the wait never exceeds the cap
- given a service that never recovers the last error is raised after the final attempt
- given an error that retrying cannot fix it is raised immediately
circuit breaker
- given three consecutive failures the breaker opens
- given an open breaker calls fail fast without touching the service
- given the reset timeout has passed one trial call goes through and success closes the breaker
- given a fifty second outage the service is reached five times and 21 requests fail fast
- given the trial call fails the breaker opens again
loop detection
- given the same tool called with the same arguments three times in a row the agent is looping
- given the same tool with different arguments the agent is making progress
catalogue
glossary
25 tests in tests/test_glossary.py
the glossary
- given every entry its definition is plain prose of at most three sentences
- given every entry its lesson is a real module
- given the glossary source no term is defined twice
- given the glossary its keys are lowercase so lookups ignore case
glossary for the browser
hover definitions
- given a term in prose its first mention gets a hover definition
- given a longer term containing a shorter one the longer term wins
- given a term inside code or a heading it is left alone
- given a term inside display math it is left alone
- given a term inside inline math it is left alone but prose beside it is not
- given display math split across tags every part is left alone
- given a term inside a link it is left alone
- given a term as part of a longer word it is not matched
- given the page depth the lesson link is relative to the page
scoped terms
- given a scoped term on a page inside its scope it gets a hover definition
- given a scoped term on a page outside its scope it is left alone
- given the glossary everyday words with a technical meaning are scoped
terms show their usual case
- given an acronym it shows in capitals
- given a name with mixed case it keeps it
- given an ordinary word capitalised only at the start of sentences it stays lower case
- given a term capitalised only in headings tables and paper titles it stays lower case
- given a proper name it keeps its capital
- given the glossary page it shows terms in their usual case
- given a term only the paper companions use it shows their case
- given a term that states its display form it shows that form
house style
3 tests in tests/test_house_style.py
punctuation
tone
dollar signs
math in python
397 tests in tests/test_math_in_python.py
checking an in python block
- given a line whose result matches its comment the block passes
- given a line whose result differs from its comment it is reported
- given an assignment with a result comment the assigned value is checked
- given a print its printed text is checked
- given a statement over several lines the result comment on its last line is checked
- given a block that shows no result it is reported
- given a block that raises it is reported
every formula in a lesson has its python
- given a formula the same section shows it in python after it
- given its python every result it shows is what the code produces
- given its python it reads as code you can run not a session transcript
every equation in a companion has its python
viz
8 tests in tests/test_viz.py
every visualization is wired
- given a placeholder in a lesson its script exists
- given any placeholder it names the widget for screen readers
- given every visualization script it is valid javascript
- given a lesson that exports visualization data it is json for its own widgets
the site loads visualizations
- given a page with a widget it loads the styles data helpers and that widget
- given a page without a widget nothing is added