The specification: 3,168 tests

A test is a small program that runs the lessons' code and checks that it does what the lesson claims. Every test here was written before the code it checks: first the behaviour is stated and seen to fail, then only enough code is written to make it pass (this is called test-driven development). Each test is named as a plain sentence, usually in the form given a situation, when something happens, then a result (behaviour-driven development), so the whole suite reads as a specification.

Every sentence below is one behaviour; click it to read the test that checks it. A behaviour checked on several inputs counts as several tests but appears once. The whole suite runs in a few seconds with make test, offline, and make spec prints this page in a terminal.

0. Math notation, from zero

22 tests in tests/test_notation.py

summation

product

dot product

norm

matrices

softmax and argmax

spread

derivatives

agreement with num py

1. The big picture

14 tests in tests/test_big_picture.py

tracing one prompt

temperature

the generation loop

training is the same forward pass plus a loss

2. Neural networks

31 tests in tests/test_neural_net.py

one neuron worked example

activation functions

why nonlinearity matters

single weight learning

tiny two layer backprop

backpropagation

training loop

3. Optimizers

23 tests in tests/test_optimizers.py

gradient descent on a bowl

momentum

adam

weight decay

warmup cosine schedule

gradient clipping

the narrow valley paths

4. Training deep networks

22 tests in tests/test_deep_nets.py

gradients through a chain of units

gradients through a deep network

initialisation

residual connections

batch norm

layer norm

RMS norm

clipping an exploding gradient

the gradient flow figure

5. Attention

20 tests in tests/test_attention.py

softmax

causal masking

scaling by sqrt dk

grouped query attention

agreement with py torch

the worked sentence

the attention widget agrees with the lesson

the lesson places the attention widget

6. Positional information

21 tests in tests/test_positional.py

sinusoidal encoding

rotary position embedding

order blindness without positions

context extension

7. The transformer

35 tests in tests/test_transformer.py

normalization

GELU

feed forward

transformer block

tiny GPT

parameter counting

mixture of experts

load balancing

compute arithmetic

8. Tokenization

40 tests in tests/test_tokenization.py

byte pair encoding training

byte pair encoding encoding

word piece

byte level BPE

tokenizer quirks

rules of thumb

agreement with tiktoken

the tokenizer widget agrees with the lesson

the tokenizer widget data

9. Training stages

31 tests in tests/test_training_stages.py

supervised fine tuning loss

reward model

direct preference optimization

lo RA

choosing an adaptation

distillation

10. Pretraining at scale

62 tests in tests/test_pretraining.py

language identification

heuristic quality filters

classifier quality scores

exact deduplication

jaccard similarity

min hash

locality sensitive hashing

why duplicates hurt

the curation pipeline

data mixtures

token budgets

synthetic data and model collapse

training memory

data parallelism

sharded optimizer state

tensor parallelism

pipeline parallelism

number formats

loss scaling

master weights

gradient clipping at scale

spike detection

checkpointing

11. Fine-tuning in practice

47 tests in tests/test_fine_tuning.py

the cost of fine tuning

formatting chat examples

deduplication

the held out set

label quality

the toy model

forgetting the general skill

catastrophic forgetting

replay

overfitting a small dataset

task arithmetic

interference

training

12. Reinforcement learning

42 tests in tests/test_reinforcement.py

the bandit

the log probability trick

learning from reward

baselines

PPO clipping

the KL leash

group relative advantages

reward hacking

13. Reasoning models

37 tests in tests/test_reasoning.py

thinking out loud

thinking budget

self consistency

verifiers

learning to reason with RL

what reasoning costs

14. Alignment and safety

20 tests in tests/test_alignment.py

constitutional critique

red teaming

sycophancy

refusal tradeoff

goodhart

release gate

15. The hardware underneath

64 tests in tests/test_hardware.py

counting the work in a matrix multiply

splitting a multiply across cores

the memory hierarchy

tiling a matrix multiply

storing a number in a float format

rounding agrees with real hardware formats

range and precision

why smaller formats are faster

all reduce across gpus

will a model fit for serving

16. Inference

44 tests in tests/test_inference.py

memory math

prefill versus decode

KV cache

sampling

speculative decoding

quantization

continuous batching

prompt caching

top p adapts to confidence

17. Structured output

47 tests in tests/test_structured_output.py

asking nicely is not enough

masking the logits

from a pattern to a state machine

which tokens may come next

constrained decoding with a pattern

nesting needs a stack

can this prefix still be completed

constrained decoding with a schema

what the mask costs in quality

token boundaries

what the mask costs in time

when retrying is enough

json mode and tool arguments

18. Long context and efficient architectures

43 tests in tests/test_efficient_architectures.py

the cost of long context

sliding window attention

sparse attention patterns

linear attention

state space models

hybrid models

compressing the KV cache

19. Loss functions

22 tests in tests/test_losses.py

cross entropy

perplexity

cross entropy from logits

regression losses

contrastive info NCE

label smoothing

20. Metrics

28 tests in tests/test_metrics.py

fraud worked example

accuracy trap

roc auc

threshold by error cost

retrieval metrics

word overlap metrics

judge calibration

21. Reading benchmarks

50 tests in tests/test_benchmarks.py

a benchmark is an average of per question scores

scoring multiple choice by likelihood

scoring a generated answer

pass at k

a score is an estimate

comparing two models

contamination

saturation

goodhart

elo ratings

bradley terry

style bias

reading an announcement

22. Overfitting and regularization

26 tests in tests/test_regularization.py

memorising

underfitting and overfitting

early stopping

bias variance

dropout

l 1 and l 2 penalties

data splits

data leakage

23. Trees and boosting

45 tests in tests/test_classical.py

measuring impurity

choosing the best question

growing a tree

overfitting

bagging

random forest

gradient boosting

when trees win

24. CNNs and RNNs

33 tests in tests/test_cnn_rnn.py

convolution

pooling

weight sharing

receptive field

image patches as tokens

recurrent network

LSTM gates

GRU

why transformers won

25. Looking inside the model

38 tests in tests/test_interpretability.py

features are directions

linear probes

decodable is not the same as used

the logit lens

activation patching

superposition

sparse autoencoders

26. Word embeddings

15 tests in tests/test_emb_word2vec.py

skip gram training pairs

negative sampling

skip gram loss

learned geometry

count based embeddings

one vector per word

27. Similarity

28 tests in tests/test_emb_similarity.py

worked example

normalization

signal in the vector length

curse of dimensionality

anisotropy

threshold calibration

the cosine widget agrees with the lesson

the cosine widget presets

the lesson places the cosine widget

28. Training embedding models

13 tests in tests/test_emb_contrastive.py

info NCE loss

temperature

hard negatives

CLIP

29. Dimensions and compression

15 tests in tests/test_emb_compression.py

storage math

recall at k

matryoshka truncation

scalar quantization

binary quantization

30. Vector indexes

35 tests in tests/test_emb_ann.py

flat search

storage math

IVF

product quantization

HNSW

HNSW search path

eight point map

the small HNSW map

the HNSW widget agrees with the lesson

31. Retrieval

54 tests in tests/test_emb_retrieval.py

BM 25

reciprocal rank fusion

hybrid search

retrieval quality

cross encoder reranking

late interaction

chunking

query and passage prefixes

the RAG widget agrees with the lesson

32. Clustering and matching

22 tests in tests/test_emb_clustering.py

k means

silhouette

density clustering

PCA

practical uses

semantic cache

33. Embeddings in production

20 tests in tests/test_emb_operations.py

models have their own spaces

blue green migration

reembedding cost

domain mismatch

building an eval set

failure triage

34. Autoencoders and VAEs

26 tests in tests/test_autoencoders.py

the stroke images

squeezing into a code

PCA is the linear special case

holes in a plain autoencoders code space

the reparameterization trick

the KL penalty

sampling new strokes

the beta trade off

35. GANs

28 tests in tests/test_gans.py

the two players

the minimax value

the best possible detective

why the forger needs a non saturating loss

oscillation

mode collapse

stabilizers

36. Diffusion and flow matching

32 tests in tests/test_diffusion.py

the noising process

learning to denoise

sampling backwards

flow matching

guidance

scaling up to images

37. Multimodal models

41 tests in tests/test_multimodal.py

cutting an image into patches

counting image tokens

patch embedding

the vision encoder

the projector

interleaving image and text

aligning vision to language

how much of each pitch

the spectrogram

the mel scale

counting audio tokens

vector quantization

video tokens

the context budget

38. Talking to a model

21 tests in tests/test_agents_llm.py

a model reply

tokens

reading a conversation

the tool call round trip

the real request

a local model through ollama

39. Orchestration

24 tests in tests/test_agents_orchestration.py

prompt chaining

routing

parallelization

orchestrator workers

evaluator optimizer

supervisor

state machine

autonomy costs

figures

40. The agent loop

22 tests in tests/test_agents_agent_loop.py

completing a task

runaway controls

tool errors

stop reasons

handing off to a human

cost

stop causes

figures

quadratic growth

41. Tools

35 tests in tests/test_agents_tools.py

tool definitions

argument validation

calling a tool

semantic validation

least privilege

idempotency

dry run

human approval

tool descriptions

fewer higher level tools

dynamic tool loading

figures

42. Coding and computer-use agents

58 tests in tests/test_agents_coding_agents.py

why code suits agents

running the tests

editing files

the edit run test loop

finding the right files

the sandbox

evaluating coding agents

driving a screen

figures

43. Model Context Protocol

23 tests in tests/test_agents_mcp.py

handshake

tools

resources

client

tool poisoning

rug pulls

confused deputy

integration count

figures

44. Retrieval-augmented generation

38 tests in tests/test_agents_rag.py

parsing messy documents

chunking

hybrid retrieval

permission aware retrieval

grounded generation

metadata filtering

contextual retrieval

query rewriting

hy DE

multi query

parent child retrieval

agentic RAG

graph RAG

debugging wrong answers

45. Context engineering

28 tests in tests/test_agents_context.py

token budget

fitting pieces into the window

cache friendly ordering

turn summarization

tool output compression

delimiting data from instructions

prefix caching

lost in the middle

the context widget agrees with the lesson

46. Memory

25 tests in tests/test_agents_memory.py

short term memory

what is worth remembering

recall

conflicting facts

tenant isolation

user rights

task state outside the model

figures

47. Planning

20 tests in tests/test_agents_planning.py

compounding error

verification and retries

checkpoints

decomposition

plan and execute

reflection versus external checks

figures

48. Evaluation

25 tests in tests/test_agents_evals.py

code graders

grading outcomes not paths

trajectory metrics

judge calibration

retrieval and faithfulness

percentiles

regression gate

online signals

49. Guardrails

33 tests in tests/test_agents_guardrails.py

injection heuristics

luhn checksum

PII redaction

schema validation

output policy

groundedness

action policy

privilege separation

50. Cost and latency

29 tests in tests/test_agents_cost.py

request pricing

model routing

response caching

semantic caching

parallel tool calls

budgets

unit economics

five times plan

the parallel figure is stable

the semantic cache sweep

51. Observability

18 tests in tests/test_agents_observability.py

span tree

gen AI attributes

totals

debugging from a trace

waterfall

rendering

open telemetry export

feedback to golden set

52. Safe deployment

29 tests in tests/test_agents_deployment.py

shadow mode

graduated autonomy

prompts as code

canary assignment

canary rollout

kill switch

rate limiting

tamper evident audit log

53. Why the hard ones fail

17 tests in tests/test_agents_failures.py

compounding error

retry with backoff

circuit breaker

loop detection

catalogue

glossary

25 tests in tests/test_glossary.py

the glossary

glossary for the browser

hover definitions

scoped terms

terms show their usual case

house style

3 tests in tests/test_house_style.py

punctuation

tone

dollar signs

math in python

397 tests in tests/test_math_in_python.py

checking an in python block

every formula in a lesson has its python

every equation in a companion has its python

navigation

1054 tests in tests/test_navigation.py

the curriculum

next and previous

the readme

every lesson meets the standard

the site

the self test book

big questions

where code links point

links into the code

checking links into this repository

pages that are not lessons

one theme everywhere

no page repeats the home page

the home page order

module names on git hub

no code mention is left unlinked

every anchor exists

names written without their module

no link goes nowhere

the home page keeps each parts introduction

the specification is published

nothing renders broken

every way of naming code is linked

the same thing looks the same everywhere

every page can be reached

companion links

the home page header

diagrams draw once

every page fits a phone

every page works with a keyboard and a screen reader

every page runs its scripts cleanly

the title

the catalog reads its lessons

learning paths

a page can be seen as a phone shows it

landing a branch

figure words stay readable

every lesson starts at the top

a reader descends by choice

viz

8 tests in tests/test_viz.py

every visualization is wired

the site loads visualizations

the KV cache widget agrees with the lesson