AI Primer

Modern AI from first principles: every concept explained, implemented and tested.

54 lessons, numbered 0 to 53. Every lesson has two levels, and you choose how deep to go. Level 1, the practitioner's guide, is the level of a good professional book: what the thing is, when you need it, your options and their trade-offs, what it costs, what breaks, and who does it in the wild. Most readers stop there, well equipped. Level 2, how it works from scratch, builds the same thing in plain Python, drawn and explained, with every formula decoded symbol by symbol; that decoding, and the code itself, fold away as Level 3. A switch at the top of each lesson picks the depth, and the paths below open every lesson at the depth their reader wants. Hover over any underlined term for a plain-English definition.

Every lesson is also an executable specification: don't take its word for how attention, retrieval or tool calling works, run it. All of it is pinned down by a suite of 3,168 tests: small programs that run the lessons' code and check it does what the lessons claim. They were written before the code (test-driven), and each is named as a plain sentence in the form given a situation, then a result (behaviour-driven), such as given a causal mask, future tokens receive zero attention (the attention lesson's tests). Read together, they are a precise specification of what every lesson teaches: read the specification.

Every lesson also runs on its own in a terminal as a narrated walkthrough (python -m primer.ml.attention), and ends with links to the primary sources. The shared toy data and stand-in embedder the lessons use live in primer.common.

The code: https://github.com/that-mathevs/ai-primer

Every map on this site (this page, the reading order, each lesson's previous and next) is generated from one file, primer.curriculum, so none of them can go stale.

Where to start

That's a lot of lessons, so pick the path that fits you. Each skips lessons but never jumps backwards.

Just explain LLMs to me

The shortest route to understanding what happens when you send a prompt.

  1. 1The big picture
  2. 5Attention
  3. 7The transformer
  4. 8Tokenization
  5. 9Training stages
  6. 13Reasoning models
  7. 16Inference

Curious about images, audio and video

How models generate pictures and sound, and how they see and hear.

  1. 2Neural networks
  2. 24CNNs and RNNs
  3. 28Training embedding models
  4. 34Autoencoders and VAEs
  5. 35GANs
  6. 36Diffusion and flow matching
  7. 37Multimodal models

Before you begin

The notation every formula in this primer uses, decoded as short loops.

  1. 0Math notation, from zeroEvery symbol in an ML formula, as a short loop

Part 1: how the model works inside

From a single neuron to a working transformer, and how models are trained and served.

Embeddings, the centerpiece, have their own package: primer.ml.embeddings.

  1. 1The big pictureWhat happens, end to end, when you send a prompt
  2. 2Neural networksNeurons, activations, the forward pass, backprop by hand
  3. 3OptimizersSGD, momentum, Adam/AdamW, learning-rate warmup and decay
  4. 4Training deep networksVanishing/exploding gradients, residuals, normalization, initialization
  5. 5AttentionQueries, keys, values, softmax, masking, multi-head, GQA, O(n²)
  6. 6Positional informationWhy order must be added, sinusoids and RoPE
  7. 7The transformerThe block, a tiny GPT, parameter counts, mixture of experts
  8. 8TokenizationBPE from scratch, byte-level tokens, why models miscount letters
  9. 9Training stagesPretraining, SFT, RLHF and DPO, LoRA, fine-tuning vs. RAG
  10. 10Pretraining at scaleData curation and deduplication, parallelism across GPUs, mixed precision
  11. 11Fine-tuning in practicePreparing data, forgetting old skills, merging models
  12. 12Reinforcement learningPolicy gradients from scratch, PPO, GRPO, reward hacking
  13. 13Reasoning modelsChain of thought, test-time compute, verifiers, learning to reason with RL
  14. 14Alignment and safetyConstitutional AI, red-teaming, sycophancy, refusals
  15. 15The hardware underneathGPUs, the memory hierarchy, FLOPs vs. bandwidth, number formats
  16. 16InferencePrefill vs. decode, the KV cache, sampling, speculative decoding, memory math
  17. 17Structured outputConstrained decoding: grammars and JSON schemas that guarantee valid output
  18. 18Long context and efficient architecturesSliding-window and sparse attention, state-space models, KV-cache compression
  19. 19Loss functionsCross-entropy, perplexity, MSE/MAE, contrastive losses
  20. 20MetricsPrecision/recall/F1, ROC-AUC, recall@k, MRR, nDCG, BLEU/ROUGE
  21. 21Reading benchmarksWhat benchmarks measure, contamination, leaderboards and arenas
  22. 22Overfitting and regularizationOverfitting, early stopping, dropout, L1/L2, leakage
  23. 23Trees and boostingDecision trees, random forests, gradient boosting, and when they still win
  24. 24CNNs and RNNsHow convolutions see and recurrent nets remember, and why transformers won
  25. 25Looking inside the modelProbes, the logit lens, activation patching, superposition, sparse autoencoders

Embeddings, the centerpiece

Vectors that capture meaning, and the search systems built on them.

An embedding is a learned mapping from an input (a word, sentence, image or piece of code) to a list of numbers, a vector, arranged so that closeness in space means similarity in meaning. Search, retrieval-augmented generation, routing, deduplication, clustering and semantic caching all run on them.

  1. 26Word embeddingsWhere embeddings came from, analogies, the "bank" problem
  2. 27SimilarityCosine vs. dot vs. distance, normalization, anisotropy, thresholds
  3. 28Training embedding modelsContrastive learning, hard negatives, CLIP
  4. 29Dimensions and compressionStorage math, Matryoshka truncation, int8 and binary quantization
  5. 30Vector indexesFlat, IVF, PQ and HNSW from scratch, recall vs. latency
  6. 31RetrievalBM25, hybrid search with RRF, rerankers, ColBERT, chunking
  7. 32Clustering and matchingk-means, density clustering, dedup, routing, semantic caching
  8. 33Embeddings in productionModel migrations, domain mismatch, measuring retrieval on its own

Generating images, audio and video

Autoencoders, GANs, diffusion, and the multimodal models that connect them to language.

Language models generate text one token at a time. Pictures, sound and video are made differently: an autoencoder squeezes them into a small code and back, a GAN pits a forger against a detective, and a diffusion model learns to turn pure noise into an image by removing a little noise at a time. Multimodal models then connect all of these to a language model, so one system can read an image, hear speech and answer in words.

  1. 34Autoencoders and VAEsSqueezing data into a code and back, and sampling new data from it
  2. 35GANsA forger against a detective: adversarial training, and why it is unstable
  3. 36Diffusion and flow matchingTurning noise into images one small step at a time
  4. 37Multimodal modelsImages, audio and video into a language model

Part 2: building systems people rely on

Agents, tools, retrieval, memory, evaluation, safety, cost and deployment.

Everything here runs offline against primer.agents.llm.ScriptedLLM, a deterministic stand-in for a model, so every failure mode can be reproduced on purpose. Swap in primer.agents.llm.ClaudeLLM to run the same code against a real model.

  1. 38Talking to a modelThe message format, and what tool calling really is
  2. 39OrchestrationWorkflows vs. agents, and the named patterns
  3. 40The agent loopA production agent loop: budgets, loop detection, recovery
  4. 41ToolsTool design, validation, idempotency, approvals, least privilege
  5. 42Coding and computer-use agentsEdit, run, test, repeat; sandboxes; driving a screen
  6. 43Model Context ProtocolMCP on the wire, and its security risks
  7. 44Retrieval-augmented generationRAG end to end, with citations and access control
  8. 45Context engineeringWhat goes in the window, compression, cache-friendly layout
  9. 46MemoryShort- and long-term memory, tenant isolation, forgetting
  10. 47PlanningPlan-and-execute, decomposition, reflection, compounding error
  11. 48EvaluationGolden sets, graders, LLM-as-judge calibration
  12. 49GuardrailsPrompt injection and privilege separation, PII, output checks
  13. 50Cost and latencyRouting, caching, batching, budgets, cost per successful task
  14. 51ObservabilityTraces, OpenTelemetry GenAI attributes, the improvement loop
  15. 52Safe deploymentShadow mode, graduated autonomy, canaries, kill switches, audit logs
  16. 53Why the hard ones failThe common failure modes, and the fix for each

Big questions

The lessons build the field from the bottom up. These questions give the top-down view: open one to see the short version, and the lessons that tell the full story, in order.

What happens, step by step, when I send a prompt to a language model?

Route: Math notation, from zero → The big picture → Tokenization → Attention → Positional information → The transformer → Inference

  1. Tokenizer: text becomes subword IDs; cost and context limits are counted in tokens.
  2. Embedding lookup turns each ID into a vector; position information is mixed in.
  3. Dozens of transformer blocks: attention mixes information across tokens, the feed-forward layer processes each token.
  4. The last position's vector becomes a score for every vocabulary token; softmax turns scores into probabilities.
  5. Sampling (temperature, top-p) picks one token, which is appended; the loop repeats until a stop token.
  6. Prefill processes the prompt in parallel; decode generates one token at a time, made cheap by the KV cache.
How does a neural network actually learn?

Route: Neural networks → Loss functions → Optimizers → Training deep networks → Overfitting and regularization

  1. A forward pass makes a prediction; a loss turns 'how wrong' into one number.
  2. Backpropagation applies the chain rule to find every weight's gradient.
  3. An optimizer (SGD, Adam/AdamW) steps each weight against its gradient; the learning rate sets the step size.
  4. Depth brings vanishing and exploding gradients; residual connections, normalization and good initialization fix them.
  5. Watch validation loss: when it rises while training loss falls, the model is overfitting; regularize or stop early.
How does attention work, and why did transformers replace RNNs?

Route: Attention → Positional information → The transformer → CNNs and RNNs

  1. Each token forms a query, key and value; query-key dot products score relevance; softmax turns scores into weights; the output blends values.
  2. Scores are divided by the square root of d_k so softmax doesn't saturate and gradients keep flowing.
  3. A causal mask hides future tokens, which makes next-token training honest and generation cacheable.
  4. Multi-head attention runs several attentions in parallel; grouped-query attention shares keys and values to shrink the KV cache.
  5. RNNs pass everything through one hidden state, one step at a time; attention gives every pair of tokens a direct path and trains in parallel.
  6. The price is O(n²) cost in sequence length, which FlashAttention, sparse attention and state-space models attack.
How are large language models trained, and when should I fine-tune instead of using RAG?

Route: Training stages → Fine-tuning in practice → Tokenization → Loss functions → Embeddings in production → Retrieval-augmented generation

  1. Pretraining: next-token prediction over trillions of tokens produces a knowledgeable base model.
  2. Supervised fine-tuning teaches the assistant format; preference tuning (RLHF or DPO) shapes helpfulness and safety.
  3. Adaptation, cheapest first: prompting, then RAG, then LoRA, then (rarely) a full fine-tune.
  4. Fine-tuning changes behavior; RAG supplies knowledge that changes or must be cited.
  5. Distillation trains a small model to imitate a large one, often the biggest production cost win.
What makes serving a model fast and affordable?

Route: The hardware underneath → Inference → Long context and efficient architectures → Attention → Cost and latency → Context engineering

  1. Prefill is compute-bound and sets time to first token; decode is memory-bound and sets tokens per second.
  2. The KV cache trades GPU memory for speed; its size is 2 × layers × KV heads × head dimension × bytes, per token.
  3. Memory math: weights = parameters × bytes per parameter (70B at 16-bit is about 140 GB).
  4. Speedups: quantization, continuous batching, speculative decoding, grouped-query attention, prompt caching.
  5. At the system level: route easy work to small models, cache stable prefixes, trim tokens, batch offline work.
What is an embedding, and how is an embedding model trained?

Route: Word embeddings → Training embedding models → Similarity → Loss functions

  1. An embedding is a learned vector where closeness means similar meaning.
  2. word2vec learned one vector per word from co-occurrence; contextual models give each token a vector that depends on its sentence.
  3. Sentence embeddings pool token vectors; models trained for similarity beat plain pooled encoders.
  4. Contrastive training pulls matching pairs together and pushes others apart, using in-batch negatives (InfoNCE).
  5. Hard negatives (right topic, wrong answer) are the biggest driver of retrieval quality.
  6. CLIP applies the same idea across images and text, putting both in one space.
How do you search millions of vectors quickly, and what does it cost?

Route: Similarity → Dimensions and compression → Vector indexes

  1. On normalized vectors, cosine, dot product and Euclidean distance give the same ranking; use what the model was trained with.
  2. Storage math: vectors × dimensions × 4 bytes (10M × 1536 is about 61 GB) before index overhead.
  3. Exact search is too slow at scale; approximate indexes trade a little recall for a lot of speed.
  4. HNSW: layered graph, long jumps on top, local search at the bottom; M, efConstruction and efSearch are the knobs.
  5. IVF searches only the nearest clusters (nprobe); PQ compresses vectors into codes.
  6. Matryoshka truncation and scalar or binary quantization shrink memory; rescoring the shortlist recovers accuracy.
How do you build retrieval that returns the right passages?

Route: Retrieval → Clustering and matching → Embeddings in production → Metrics → Retrieval-augmented generation

  1. Chunk on document structure, with overlap and metadata; chunking often matters more than the model.
  2. Dense search finds meaning; BM25 finds exact IDs and rare terms; hybrid search fuses both with reciprocal rank fusion.
  3. Retrieve wide with a bi-encoder, then rerank the shortlist with a cross-encoder.
  4. Measure retrieval on its own with recall@k, MRR and nDCG on a labeled set before tuning prompts.
  5. Operations: new embedding model means re-embedding everything; version indexes and switch traffic behind an alias.
How do you know if a model or an agent is any good?

Route: Metrics → Reading benchmarks → Loss functions → Overfitting and regularization → Evaluation

  1. Pick metrics by the cost of each error: precision vs. recall; accuracy misleads on imbalanced data.
  2. Keep training, validation and test data apart, and watch for leakage and benchmark contamination.
  3. For agents: a golden set of real tasks, graded by code wherever possible (end state, schema, tests).
  4. For open-ended output: an LLM judge with an explicit rubric, calibrated against human labels.
  5. Track trajectory, cost and latency beside quality; run the suite on every change and feed production failures back in.
When should you build an agent, and how does one work?

Route: Talking to a model → Orchestration → The agent loop → Tools → Coding and computer-use agents → Model Context Protocol → Planning

  1. Use the least autonomy that solves the problem: fixed workflow, then router, then agent loop, then multiple agents.
  2. Tool calling: the model emits a structured request; your code validates it, runs it and returns the result. The model executes nothing.
  3. The loop: think, call a tool, observe, decide, with step limits, token budgets and loop detection.
  4. Tool design is prompt design: few, high-level tools with precise descriptions, validated arguments and actionable errors.
  5. Long tasks fail by compounding error (0.95¹⁰ ≈ 0.60), so plan, verify each step externally, and checkpoint.
  6. MCP standardizes how apps connect to tools, and brings its own risks: tool poisoning, rug pulls, broad permissions.
How do you give an AI system the right context and memory?

Route: Context engineering → Memory → Retrieval-augmented generation

  1. Context engineering: the smallest set of high-signal content, structured with clear delimiters.
  2. More context is not better: models use the middle of long inputs least reliably, and quality rots as sessions grow.
  3. Put stable content first so prompt caching can reuse it; summarize or drop old turns; compress tool outputs.
  4. Long-term memory is retrieval over the system's own history: episodic, semantic and procedural.
  5. Memory must be isolated per tenant and user, updatable, and deletable on request.
How do you make an AI system safe to put in front of real users?

Route: Alignment and safety → Guardrails → Tools → Safe deployment → Observability

  1. Layer guardrails on inputs, outputs and actions; no single check is reliable alone.
  2. Treat retrieved content, emails and tool outputs as untrusted: no prompt wording fully prevents injection.
  3. Privilege separation: the part that reads untrusted content holds no dangerous tools; actions pass a policy or human check.
  4. Graduate autonomy with evidence: shadow mode, then approval per action, then autonomy for low-risk actions.
  5. Trace every run, keep tamper-evident audit logs, rate-limit actions and keep a kill switch and a one-step rollback.
How do you cut cost and latency without hurting quality?

Route: Cost and latency → Context engineering → Inference → Clustering and matching

  1. Measure cost per successful task, not per call.
  2. Route each step to the cheapest model that handles it; this is usually the biggest lever.
  3. Cache: prompt caching for stable prefixes, response and semantic caches for repeated questions.
  4. Trim tokens: tight prompts, compressed tool outputs, only the top reranked chunks.
  5. Run independent tool calls in parallel, stream output, and move offline work to batch APIs.
  6. Enforce per-task and per-tenant budgets with anomaly alerts.
Why do AI systems fail in production, and how do you fix them?

Route: Why the hard ones fail → Structured output → Planning → Retrieval-augmented generation → Evaluation → Observability

  1. Compounding error over long tasks: shorten paths, verify steps, checkpoint.
  2. Bad retrieval behind confident wrong answers: hybrid search, reranking, retrieval evals.
  3. Ambiguous tools, loops and runaway cost: better tool design, budgets, loop detection.
  4. Prompt injection and messy enterprise data: untrusted-content boundaries, permission-aware retrieval, investment in parsing.
  5. No evals and no traces: regressions ship silently; build the loop from production failure to trace to test case to fix.
What does it take to pretrain a large model?

Route: Training stages → Tokenization → Pretraining at scale → Optimizers → The hardware underneath

  1. Most of a web crawl is thrown away: language ID, quality rules and classifiers, and exact and near-duplicate removal.
  2. Sources are mixed by weight, not size; about 20 tokens per parameter is compute-optimal, but models meant for heavy use train far longer.
  3. Adam in mixed precision needs about 16 bytes per parameter before activations, so one GPU can't hold a large model.
  4. Data parallelism shares gradients, ZeRO/FSDP shards the training state, tensor parallelism splits each matrix multiply, and pipeline parallelism splits the layers.
  5. The maths runs in bf16 or fp8 with scaling, while the master weights stay in fp32.
  6. Warmup, gradient clipping, spike rollback and regular checkpoints keep a months-long run alive.
How does a model learn from rewards instead of examples?

Route: Reinforcement learning → Training stages → Reasoning models → Alignment and safety

  1. Reinforcement learning samples an action, scores it, and makes high-scoring actions more likely.
  2. A baseline turns rewards into advantages (better or worse than usual), which cuts noise without bias.
  3. PPO reuses each batch for several steps, clips how far the policy moves, and leashes it to a reference with a KL penalty.
  4. GRPO drops the value network by comparing rewards within a group of answers to the same prompt.
  5. A checker as the reward (the right answer, passing tests) is how reasoning models are trained.
  6. The policy optimises the reward you wrote, not the goal you meant; verifiable rewards, a KL leash and held-out checks defend against that.
How do reasoning models think, and when is extra thinking worth it?

Route: Inference → Reinforcement learning → Reasoning models → Planning → Cost and latency

  1. Every written token is another forward pass, so a chain of thought buys serial computation, and the text is the model's working memory.
  2. Test-time compute can go into one longer chain, or into many chains with a vote or a verifier picking one answer.
  3. Voting helps only when the right answer is the most common one and the samples' mistakes are independent.
  4. Checking each step catches errors that checking only the final answer misses.
  5. Training with verifiable rewards makes longer, self-checking reasoning emerge.
  6. Thinking is paid for per token and slips compound over long chains, so route easy tasks to little thinking and measure cost per successful task.
How do models handle very long contexts?

Route: Attention → Positional information → Inference → Long context and efficient architectures

  1. Attention scores every pair of tokens, and the KV cache grows with every token, per layer, per conversation.
  2. Sliding windows and sparse patterns score fewer pairs; stacked layers still carry information far.
  3. Linear attention and state-space models keep a fixed-size summary: linear time and constant memory, but blurrier recall.
  4. Mamba makes the summary selective: each token decides how much to keep and how much to write.
  5. Hybrids keep a few attention layers for exact lookup.
  6. The cache shrinks by sharing key/value heads, caching a small latent, or storing fewer bits.
What is going on inside a trained model, and how can we tell?

Route: The transformer → Word embeddings → Looking inside the model → Alignment and safety

  1. Models store features as directions across many neurons, not one feature per neuron.
  2. Probes and the logit lens read what is present; they show correlation, not use.
  3. Activation patching changes one activation and watches the output: the causal test.
  4. Sparse features get packed in superposition, which makes individual neurons respond to several things.
  5. Sparse autoencoders unpack superposition into interpretable features, at the cost of some unexplained activity.
  6. These tools give evidence, not proof; full explanations exist only for narrow behaviours.
When is a neural network the wrong tool?

Route: Trees and boosting → Neural networks → Overfitting and regularization → Metrics

  1. On tables whose columns each mean something alone, gradient-boosted trees or a random forest are the model to beat.
  2. A tree asks one column at a time whether it's above a threshold, so it needs no feature scaling and handles categories natively.
  3. A single deep tree overfits; forests average many decorrelated trees, and boosting adds small trees fit to the remaining errors.
  4. Neural networks win when meaning lives in arrangements of raw values (images, audio, text), when data is huge, or when a pretrained model can be reused.
  5. Trees can't extrapolate beyond the values they trained on, and impurity importances credit noise, so check importances on held-out data.
How do AI models generate images, audio and video?

Route: Autoencoders and VAEs → GANs → Diffusion and flow matching → Multimodal models

  1. A generator learns a whole distribution, so it can sample new examples; predicting the average gives blur.
  2. An autoencoder squeezes data into a small code and back; a VAE shapes that code so random codes decode to new data.
  3. A GAN trains a generator against a discriminator: sharp, one-pass samples, but unstable training and mode collapse.
  4. Diffusion adds noise on purpose and learns to remove it, generating from pure noise in many small steps.
  5. Flow matching learns straight paths from noise to data, so it needs fewer steps; guidance trades variety for following the prompt.
  6. Real systems denoise an autoencoder's latent with a transformer that reads the prompt; video and audio are the same idea with more tokens.
How do AI models see images and hear audio?

Route: CNNs and RNNs → Training embedding models → Multimodal models

  1. Every modality becomes a sequence of vectors a transformer attends over.
  2. Images become patch tokens, (H/P)·(W/P) of them, so cost grows with the square of the resolution.
  3. A small projector connects a vision or audio encoder to a language model; it is trained first, with both models frozen.
  4. Audio becomes a log-mel spectrogram, then about 50 tokens a second.
  5. Video multiplies image tokens by time, so frames are sampled; text stays the densest input.

The papers behind the lessons

Annotated, interactive companions: hover over any term or equation symbol. All papers.