primer.curriculum

Curriculum: the map of every lesson

The single source of truth for lesson order, titles and what each lesson teaches. Everything navigational is generated from CURRICULUM:

  • the reading-order tables in README.md (make readme; a test fails when they drift);
  • the reading lists in each package's page (primer.ml, primer.agents, …);
  • the site's home page, breadcrumbs and previous/next links (make docs).

To add a lesson, add one Lesson here, in reading order. Nothing else needs editing by hand.

Before you begin

The notation every formula in this primer uses, decoded as short loops.

  1. primer.notation: Math notation, from zero. Every symbol in an ML formula, as a short loop.

Part 1: how the model works inside

From a single neuron to a working transformer, and how models are trained and served.

  1. primer.ml.big_picture: The big picture. What happens, end to end, when you send a prompt.
  2. primer.ml.neural_net: Neural networks. Neurons, activations, the forward pass, backprop by hand.
  3. primer.ml.optimizers: Optimizers. SGD, momentum, Adam/AdamW, learning-rate warmup and decay.
  4. primer.ml.deep_nets: Training deep networks. Vanishing/exploding gradients, residuals, normalization, initialization.
  5. primer.ml.attention: Attention. Queries, keys, values, softmax, masking, multi-head, GQA, O(n²).
  6. primer.ml.positional: Positional information. Why order must be added, sinusoids and RoPE.
  7. primer.ml.transformer: The transformer. The block, a tiny GPT, parameter counts, mixture of experts.
  8. primer.ml.tokenization: Tokenization. BPE from scratch, byte-level tokens, why models miscount letters.
  9. primer.ml.training_stages: Training stages. Pretraining, SFT, RLHF and DPO, LoRA, fine-tuning vs. RAG.
  10. primer.ml.pretraining: Pretraining at scale. Data curation and deduplication, parallelism across GPUs, mixed precision.
  11. primer.ml.fine_tuning: Fine-tuning in practice. Preparing data, forgetting old skills, merging models.
  12. primer.ml.reinforcement: Reinforcement learning. Policy gradients from scratch, PPO, GRPO, reward hacking.
  13. primer.ml.reasoning: Reasoning models. Chain of thought, test-time compute, verifiers, learning to reason with RL.
  14. primer.ml.alignment: Alignment and safety. Constitutional AI, red-teaming, sycophancy, refusals.
  15. primer.ml.hardware: The hardware underneath. GPUs, the memory hierarchy, FLOPs vs. bandwidth, number formats.
  16. primer.ml.inference: Inference. Prefill vs. decode, the KV cache, sampling, speculative decoding, memory math.
  17. primer.ml.structured_output: Structured output. Constrained decoding: grammars and JSON schemas that guarantee valid output.
  18. primer.ml.efficient_architectures: Long context and efficient architectures. Sliding-window and sparse attention, state-space models, KV-cache compression.
  19. primer.ml.losses: Loss functions. Cross-entropy, perplexity, MSE/MAE, contrastive losses.
  20. primer.ml.metrics: Metrics. Precision/recall/F1, ROC-AUC, recall@k, MRR, nDCG, BLEU/ROUGE.
  21. primer.ml.benchmarks: Reading benchmarks. What benchmarks measure, contamination, leaderboards and arenas.
  22. primer.ml.regularization: Overfitting and regularization. Overfitting, early stopping, dropout, L1/L2, leakage.
  23. primer.ml.classical: Trees and boosting. Decision trees, random forests, gradient boosting, and when they still win.
  24. primer.ml.cnn_rnn: CNNs and RNNs. How convolutions see and recurrent nets remember, and why transformers won.
  25. primer.ml.interpretability: Looking inside the model. Probes, the logit lens, activation patching, superposition, sparse autoencoders.

Embeddings, the centerpiece

Vectors that capture meaning, and the search systems built on them.

  1. primer.ml.embeddings.word2vec: Word embeddings. Where embeddings came from, analogies, the "bank" problem.
  2. primer.ml.embeddings.similarity: Similarity. Cosine vs. dot vs. distance, normalization, anisotropy, thresholds.
  3. primer.ml.embeddings.contrastive: Training embedding models. Contrastive learning, hard negatives, CLIP.
  4. primer.ml.embeddings.compression: Dimensions and compression. Storage math, Matryoshka truncation, int8 and binary quantization.
  5. primer.ml.embeddings.ann: Vector indexes. Flat, IVF, PQ and HNSW from scratch, recall vs. latency.
  6. primer.ml.embeddings.retrieval: Retrieval. BM25, hybrid search with RRF, rerankers, ColBERT, chunking.
  7. primer.ml.embeddings.clustering: Clustering and matching. k-means, density clustering, dedup, routing, semantic caching.
  8. primer.ml.embeddings.operations: Embeddings in production. Model migrations, domain mismatch, measuring retrieval on its own.

Generating images, audio and video

Autoencoders, GANs, diffusion, and the multimodal models that connect them to language.

  1. primer.ml.generative.autoencoders: Autoencoders and VAEs. Squeezing data into a code and back, and sampling new data from it.
  2. primer.ml.generative.gans: GANs. A forger against a detective: adversarial training, and why it is unstable.
  3. primer.ml.generative.diffusion: Diffusion and flow matching. Turning noise into images one small step at a time.
  4. primer.ml.generative.multimodal: Multimodal models. Images, audio and video into a language model.

Part 2: building systems people rely on

Agents, tools, retrieval, memory, evaluation, safety, cost and deployment.

  1. primer.agents.llm: Talking to a model. The message format, and what tool calling really is.
  2. primer.agents.orchestration: Orchestration. Workflows vs. agents, and the named patterns.
  3. primer.agents.agent_loop: The agent loop. A production agent loop: budgets, loop detection, recovery.
  4. primer.agents.tools: Tools. Tool design, validation, idempotency, approvals, least privilege.
  5. primer.agents.coding_agents: Coding and computer-use agents. Edit, run, test, repeat; sandboxes; driving a screen.
  6. primer.agents.mcp: Model Context Protocol. MCP on the wire, and its security risks.
  7. primer.agents.rag: Retrieval-augmented generation. RAG end to end, with citations and access control.
  8. primer.agents.context: Context engineering. What goes in the window, compression, cache-friendly layout.
  9. primer.agents.memory: Memory. Short- and long-term memory, tenant isolation, forgetting.
  10. primer.agents.planning: Planning. Plan-and-execute, decomposition, reflection, compounding error.
  11. primer.agents.evals: Evaluation. Golden sets, graders, LLM-as-judge calibration.
  12. primer.agents.guardrails: Guardrails. Prompt injection and privilege separation, PII, output checks.
  13. primer.agents.cost: Cost and latency. Routing, caching, batching, budgets, cost per successful task.
  14. primer.agents.observability: Observability. Traces, OpenTelemetry GenAI attributes, the improvement loop.
  15. primer.agents.deployment: Safe deployment. Shadow mode, graduated autonomy, canaries, kill switches, audit logs.
  16. primer.agents.failures: Why the hard ones fail. The common failure modes, and the fix for each.
on GitHub
  1"""
  2# Curriculum: the map of every lesson
  3
  4The single source of truth for lesson order, titles and what each lesson
  5teaches. Everything navigational is generated from `CURRICULUM`:
  6
  7* the reading-order tables in `README.md` (`make readme`; a test fails when
  8  they drift);
  9* the reading lists in each package's page (`primer.ml`, `primer.agents`, …);
 10* the site's home page, breadcrumbs and previous/next links (`make docs`).
 11
 12To add a lesson, add one `Lesson` here, in reading order. Nothing else needs
 13editing by hand.
 14"""
 15
 16from __future__ import annotations
 17
 18from dataclasses import dataclass
 19
 20# Where the project lives. The README needs absolute links to the site (the
 21# built HTML isn't committed). The site builder reads the repository's address
 22# from git or GitHub Actions instead, so a fork's site links to the fork; this
 23# constant is only its fallback.
 24REPO_URL = "https://github.com/that-mathevs/ai-primer"
 25SITE_URL = "https://that-mathevs.github.io/ai-primer/"
 26BRANCH = "main"
 27
 28
 29def source_path(module: str) -> str:
 30    """The file a module lives in, relative to the repository root."""
 31    return module.replace(".", "/") + ".py"
 32
 33
 34def link_code_references(markdown: str, from_dir: str) -> str:
 35    """Make every code span that names something in this repository a link GitHub can follow.
 36
 37    GitHub renders Markdown but not docstrings, so in a .md file a reference to
 38    code only helps a reader if it's a relative link. Linked, when written as code:
 39
 40    * a module, or a name inside one: `primer.agents.llm`, `primer.agents.llm.ClaudeLLM`;
 41    * a command that runs a lesson: `python -m primer.ml.attention`;
 42    * a committed file or folder, written from the repository root (`primer/glossary.py`,
 43      `CLAUDE.md`) or from the Markdown file's own folder;
 44    * a class defined in exactly one module: `ScriptedLLM`.
 45
 46    Anything else (build output such as `docs/html/index.html`, commands, other
 47    projects' names) stays plain. Existing links and fenced code blocks are untouched.
 48
 49    Args:
 50        markdown: the text to link.
 51        from_dir: the directory the Markdown file sits in, relative to the repository root.
 52    """
 53    import os
 54    import re
 55    from pathlib import Path
 56
 57    root = Path(__file__).resolve().parent.parent
 58    here = root / from_dir
 59
 60    def rel(target: Path) -> str:
 61        return os.path.relpath(target, here).replace(os.sep, "/")
 62
 63    def target_for(code: str) -> Path | None:
 64        run = re.fullmatch(r"python -m (primer(?:\.\w+)+)", code)
 65        dotted = run.group(1) if run else code
 66        if re.fullmatch(r"primer(?:\.\w+)+", dotted):
 67            return _module_file(dotted)
 68        if re.fullmatch(r"[\w.-]+(?:/[\w.-]+)*/?", code) and ("/" in code or "." in code):
 69            for base in (root, here):
 70                if _is_committed(base / code):
 71                    return base / code
 72            return None
 73        if re.fullmatch(r"[A-Z]\w+(?:\(\))?", code):
 74            homes = _class_homes().get(code.removesuffix("()"), ())
 75            return _module_file(homes[0]) if len(homes) == 1 else None
 76        return None
 77
 78    def link(m: re.Match) -> str:
 79        target = target_for(m.group(1))
 80        return f"[`{m.group(1)}`]({rel(target)})" if target else m.group(0)
 81
 82    # A code span that is already a link's text is followed by "]"; one inside [..] is preceded by "[".
 83    span = re.compile(r"(?<!\[)`([^`\n]+)`(?!\])")
 84    pieces = re.split(r"(```.*?```)", markdown, flags=re.S)
 85    return "".join(piece if i % 2 else span.sub(link, piece) for i, piece in enumerate(pieces))
 86
 87
 88def _module_file(name: str):
 89    """The file that defines a dotted name: primer.agents.llm.ClaudeLLM -> primer/agents/llm.py."""
 90    from pathlib import Path
 91
 92    root = Path(__file__).resolve().parent.parent
 93    parts = name.split(".")
 94    for i in range(len(parts), 0, -1):
 95        base = root.joinpath(*parts[:i])
 96        if base.with_suffix(".py").is_file():
 97            return base.with_suffix(".py")
 98        if (base / "__init__.py").is_file():
 99            return base / "__init__.py"
100    return None
101
102
103_COMMITTED: set[str] | None = None
104
105
106def _is_committed(path) -> bool:
107    """Whether GitHub will have this file or folder: it is tracked by git (or exists, outside a checkout)."""
108    import subprocess
109    from pathlib import Path
110
111    global _COMMITTED
112    root = Path(__file__).resolve().parent.parent
113    if _COMMITTED is None:
114        out = subprocess.run(["git", "ls-files"], cwd=root, capture_output=True, text=True)
115        _COMMITTED = set(out.stdout.split("\n")) if out.returncode == 0 else set()
116    try:
117        rel = Path(path).resolve().relative_to(root).as_posix()
118    except ValueError:
119        return False
120    if not _COMMITTED:
121        return Path(path).exists()
122    return rel in _COMMITTED or any(f.startswith(rel.rstrip("/") + "/") for f in _COMMITTED)
123
124
125_CLASS_HOMES: dict[str, list[str]] | None = None
126
127
128def _class_homes() -> dict[str, list[str]]:
129    """Every public class in the package, mapped to the modules that define one by that name."""
130    import importlib
131    import inspect
132    import pkgutil
133
134    import primer
135
136    global _CLASS_HOMES
137    if _CLASS_HOMES is None:
138        _CLASS_HOMES = {}
139        for info in pkgutil.walk_packages(primer.__path__, "primer."):
140            if info.name.rsplit(".", 1)[-1].startswith("_"):
141                continue
142            mod = importlib.import_module(info.name)
143            for n, o in vars(mod).items():
144                if inspect.isclass(o) and o.__module__ == info.name and not n.startswith("_"):
145                    _CLASS_HOMES.setdefault(n, []).append(info.name)
146    return _CLASS_HOMES
147
148
149def tests_for(module: str) -> str:
150    """The test file that specifies a lesson (tests/test_<name>.py, test_emb_ or test_agents_)."""
151    name = module.rsplit(".", 1)[-1]
152    prefix = "emb_" if module.startswith("primer.ml.embeddings.") else "agents_" if module.startswith("primer.agents.") else ""
153    return f"tests/test_{prefix}{name}.py"
154
155
156@dataclass(frozen=True)
157class Part:
158    key: str
159    title: str
160    blurb: str
161
162
163@dataclass(frozen=True)
164class Lesson:
165    module: str
166    title: str
167    outcome: str  # "What you'll be able to explain"
168    part: str  # Part.key
169
170
171PARTS: list[Part] = [
172    Part("start", "Before you begin", "The notation every formula in this primer uses, decoded as short loops."),
173    Part("ml", "Part 1: how the model works inside", "From a single neuron to a working transformer, and how models are trained and served."),
174    Part("embeddings", "Embeddings, the centerpiece", "Vectors that capture meaning, and the search systems built on them."),
175    Part("generative", "Generating images, audio and video", "Autoencoders, GANs, diffusion, and the multimodal models that connect them to language."),
176    Part("agents", "Part 2: building systems people rely on", "Agents, tools, retrieval, memory, evaluation, safety, cost and deployment."),
177]
178
179CURRICULUM: list[Lesson] = [
180    Lesson("primer.notation", "Math notation, from zero", "Every symbol in an ML formula, as a short loop", "start"),
181    Lesson("primer.ml.big_picture", "The big picture", "What happens, end to end, when you send a prompt", "ml"),
182    Lesson("primer.ml.neural_net", "Neural networks", "Neurons, activations, the forward pass, backprop by hand", "ml"),
183    Lesson("primer.ml.optimizers", "Optimizers", "SGD, momentum, Adam/AdamW, learning-rate warmup and decay", "ml"),
184    Lesson("primer.ml.deep_nets", "Training deep networks", "Vanishing/exploding gradients, residuals, normalization, initialization", "ml"),
185    Lesson("primer.ml.attention", "Attention", "Queries, keys, values, softmax, masking, multi-head, GQA, O(n²)", "ml"),
186    Lesson("primer.ml.positional", "Positional information", "Why order must be added, sinusoids and RoPE", "ml"),
187    Lesson("primer.ml.transformer", "The transformer", "The block, a tiny GPT, parameter counts, mixture of experts", "ml"),
188    Lesson("primer.ml.tokenization", "Tokenization", "BPE from scratch, byte-level tokens, why models miscount letters", "ml"),
189    Lesson("primer.ml.training_stages", "Training stages", "Pretraining, SFT, RLHF and DPO, LoRA, fine-tuning vs. RAG", "ml"),
190    Lesson("primer.ml.pretraining", "Pretraining at scale", "Data curation and deduplication, parallelism across GPUs, mixed precision", "ml"),
191    Lesson("primer.ml.fine_tuning", "Fine-tuning in practice", "Preparing data, forgetting old skills, merging models", "ml"),
192    Lesson("primer.ml.reinforcement", "Reinforcement learning", "Policy gradients from scratch, PPO, GRPO, reward hacking", "ml"),
193    Lesson("primer.ml.reasoning", "Reasoning models", "Chain of thought, test-time compute, verifiers, learning to reason with RL", "ml"),
194    Lesson("primer.ml.alignment", "Alignment and safety", "Constitutional AI, red-teaming, sycophancy, refusals", "ml"),
195    Lesson("primer.ml.hardware", "The hardware underneath", "GPUs, the memory hierarchy, FLOPs vs. bandwidth, number formats", "ml"),
196    Lesson("primer.ml.inference", "Inference", "Prefill vs. decode, the KV cache, sampling, speculative decoding, memory math", "ml"),
197    Lesson("primer.ml.structured_output", "Structured output", "Constrained decoding: grammars and JSON schemas that guarantee valid output", "ml"),
198    Lesson("primer.ml.efficient_architectures", "Long context and efficient architectures", "Sliding-window and sparse attention, state-space models, KV-cache compression", "ml"),
199    Lesson("primer.ml.losses", "Loss functions", "Cross-entropy, perplexity, MSE/MAE, contrastive losses", "ml"),
200    Lesson("primer.ml.metrics", "Metrics", "Precision/recall/F1, ROC-AUC, recall@k, MRR, nDCG, BLEU/ROUGE", "ml"),
201    Lesson("primer.ml.benchmarks", "Reading benchmarks", "What benchmarks measure, contamination, leaderboards and arenas", "ml"),
202    Lesson("primer.ml.regularization", "Overfitting and regularization", "Overfitting, early stopping, dropout, L1/L2, leakage", "ml"),
203    Lesson("primer.ml.classical", "Trees and boosting", "Decision trees, random forests, gradient boosting, and when they still win", "ml"),
204    Lesson("primer.ml.cnn_rnn", "CNNs and RNNs", "How convolutions see and recurrent nets remember, and why transformers won", "ml"),
205    Lesson("primer.ml.interpretability", "Looking inside the model", "Probes, the logit lens, activation patching, superposition, sparse autoencoders", "ml"),
206    Lesson("primer.ml.embeddings.word2vec", "Word embeddings", "Where embeddings came from, analogies, the \"bank\" problem", "embeddings"),
207    Lesson("primer.ml.embeddings.similarity", "Similarity", "Cosine vs. dot vs. distance, normalization, anisotropy, thresholds", "embeddings"),
208    Lesson("primer.ml.embeddings.contrastive", "Training embedding models", "Contrastive learning, hard negatives, CLIP", "embeddings"),
209    Lesson("primer.ml.embeddings.compression", "Dimensions and compression", "Storage math, Matryoshka truncation, int8 and binary quantization", "embeddings"),
210    Lesson("primer.ml.embeddings.ann", "Vector indexes", "Flat, IVF, PQ and HNSW from scratch, recall vs. latency", "embeddings"),
211    Lesson("primer.ml.embeddings.retrieval", "Retrieval", "BM25, hybrid search with RRF, rerankers, ColBERT, chunking", "embeddings"),
212    Lesson("primer.ml.embeddings.clustering", "Clustering and matching", "k-means, density clustering, dedup, routing, semantic caching", "embeddings"),
213    Lesson("primer.ml.embeddings.operations", "Embeddings in production", "Model migrations, domain mismatch, measuring retrieval on its own", "embeddings"),
214    Lesson("primer.ml.generative.autoencoders", "Autoencoders and VAEs", "Squeezing data into a code and back, and sampling new data from it", "generative"),
215    Lesson("primer.ml.generative.gans", "GANs", "A forger against a detective: adversarial training, and why it is unstable", "generative"),
216    Lesson("primer.ml.generative.diffusion", "Diffusion and flow matching", "Turning noise into images one small step at a time", "generative"),
217    Lesson("primer.ml.generative.multimodal", "Multimodal models", "Images, audio and video into a language model", "generative"),
218    Lesson("primer.agents.llm", "Talking to a model", "The message format, and what tool calling really is", "agents"),
219    Lesson("primer.agents.orchestration", "Orchestration", "Workflows vs. agents, and the named patterns", "agents"),
220    Lesson("primer.agents.agent_loop", "The agent loop", "A production agent loop: budgets, loop detection, recovery", "agents"),
221    Lesson("primer.agents.tools", "Tools", "Tool design, validation, idempotency, approvals, least privilege", "agents"),
222    Lesson("primer.agents.coding_agents", "Coding and computer-use agents", "Edit, run, test, repeat; sandboxes; driving a screen", "agents"),
223    Lesson("primer.agents.mcp", "Model Context Protocol", "MCP on the wire, and its security risks", "agents"),
224    Lesson("primer.agents.rag", "Retrieval-augmented generation", "RAG end to end, with citations and access control", "agents"),
225    Lesson("primer.agents.context", "Context engineering", "What goes in the window, compression, cache-friendly layout", "agents"),
226    Lesson("primer.agents.memory", "Memory", "Short- and long-term memory, tenant isolation, forgetting", "agents"),
227    Lesson("primer.agents.planning", "Planning", "Plan-and-execute, decomposition, reflection, compounding error", "agents"),
228    Lesson("primer.agents.evals", "Evaluation", "Golden sets, graders, LLM-as-judge calibration", "agents"),
229    Lesson("primer.agents.guardrails", "Guardrails", "Prompt injection and privilege separation, PII, output checks", "agents"),
230    Lesson("primer.agents.cost", "Cost and latency", "Routing, caching, batching, budgets, cost per successful task", "agents"),
231    Lesson("primer.agents.observability", "Observability", "Traces, OpenTelemetry GenAI attributes, the improvement loop", "agents"),
232    Lesson("primer.agents.deployment", "Safe deployment", "Shadow mode, graduated autonomy, canaries, kill switches, audit logs", "agents"),
233    Lesson("primer.agents.failures", "Why the hard ones fail", "The common failure modes, and the fix for each", "agents"),
234]
235
236_BY_MODULE = {l.module: i for i, l in enumerate(CURRICULUM)}
237
238
239def neighbours(module: str) -> tuple[Lesson | None, Lesson | None]:
240    """(previous lesson, next lesson) in reading order; None at either end."""
241    i = _BY_MODULE[module]
242    return (CURRICULUM[i - 1] if i > 0 else None, CURRICULUM[i + 1] if i + 1 < len(CURRICULUM) else None)
243
244
245def lessons_in(part_key: str) -> list[tuple[int, Lesson]]:
246    """(number, lesson) pairs for one part, numbered across the whole curriculum."""
247    return [(i, l) for i, l in enumerate(CURRICULUM) if l.part == part_key]
248
249
250def readme_section() -> str:
251    """The README's reading-order tables. Regenerate with `make readme`."""
252    out = []
253    for part in PARTS:
254        out.append(f"### {part.title}\n\n{part.blurb}\n\n| # | Lesson | What you'll be able to explain | Read |\n|---|---|---|---|")
255        out += [
256            f"| {i} | [{l.title}]({source_path(l.module)}) | {l.outcome} | "
257            f"[page]({SITE_URL}{l.module.replace('.', '/')}.html) · [tests]({tests_for(l.module)}) |"
258            for i, l in lessons_in(part.key)
259        ]
260        out.append("")
261    return link_code_references("\n".join(out) + "\n", '.')
262
263
264def reading_list(package: str) -> str:
265    """A markdown reading list of the lessons inside `package`, for its page."""
266    rows = [
267        f"{i}. `{l.module}`: **{l.title}.** {l.outcome}."
268        for i, l in enumerate(CURRICULUM)
269        if l.module.startswith(package + ".") and "." not in l.module[len(package) + 1 :]
270    ]
271    return "\n## Reading order\n\n" + "\n".join(rows) + "\n\nThe full map is in `primer.curriculum`.\n"
272
273
274# ---------------------------------------------------------------------------
275# Every lesson opens with "## Level 1: The practitioner's guide" and continues with
276# "## Level 2: How it works, from scratch". A lesson added without its guide goes here
277# until it has one (tests/test_navigation.py holds the bar); the set is empty now.
278# ---------------------------------------------------------------------------
279
280LEVELS_PENDING: frozenset[str] = frozenset()
281
282
283# ---------------------------------------------------------------------------
284# Learning paths: a short route through the lessons for each kind of reader.
285# A path skips lessons but never jumps backwards, so prerequisites come first.
286# ---------------------------------------------------------------------------
287
288
289@dataclass(frozen=True)
290class LearningPath:
291    who: str  # the reader it's for
292    why: str  # what they get out of it
293    route: tuple[str, ...]  # lesson modules, in reading order
294    # How deep this reader goes by default: 1 the practitioner's guide, 2 the mechanism built
295    # from scratch, 3 its math and code as well. The site opens each lesson at this depth.
296    depth: int = 2
297
298
299_M, _E, _G, _A = "primer.ml.", "primer.ml.embeddings.", "primer.ml.generative.", "primer.agents."
300
301LEARNING_PATHS: list[LearningPath] = [
302    LearningPath(
303        "Software engineer new to AI",
304        "How a language model works, then how to build on one.",
305        depth=2, route=
306        ("primer.notation", _M + "big_picture", _M + "neural_net", _M + "attention", _M + "transformer", _M + "tokenization",
307         _M + "inference", _E + "similarity", _E + "retrieval", _A + "llm", _A + "agent_loop", _A + "rag", _A + "evals"),
308    ),
309    LearningPath(
310        "AI application engineer",
311        "Agents, retrieval and tools, and keeping them reliable, safe and affordable.",
312        depth=1, route=
313        (_M + "structured_output", _E + "retrieval", _A + "llm", _A + "orchestration", _A + "agent_loop", _A + "tools",
314         _A + "coding_agents", _A + "mcp", _A + "rag", _A + "context", _A + "memory", _A + "evals", _A + "guardrails",
315         _A + "cost", _A + "observability", _A + "deployment", _A + "failures"),
316    ),
317    LearningPath(
318        "ML engineer",
319        "The model itself: training, scaling, serving and looking inside.",
320        depth=3, route=
321        ("primer.notation", _M + "neural_net", _M + "optimizers", _M + "deep_nets", _M + "attention", _M + "positional",
322         _M + "transformer", _M + "training_stages", _M + "pretraining", _M + "fine_tuning", _M + "reinforcement",
323         _M + "hardware", _M + "inference", _M + "efficient_architectures", _M + "losses", _M + "metrics",
324         _M + "benchmarks", _M + "regularization", _M + "interpretability"),
325    ),
326    LearningPath(
327        "Engineering manager or architect",
328        "What these systems can do, what they cost, and how they fail.",
329        depth=1, route=
330        (_M + "big_picture", _M + "training_stages", _M + "reasoning", _M + "alignment", _M + "inference",
331         _M + "benchmarks", _A + "orchestration", _A + "rag", _A + "evals", _A + "cost", _A + "deployment", _A + "failures"),
332    ),
333    LearningPath(
334        "Just explain LLMs to me",
335        "The shortest route to understanding what happens when you send a prompt.",
336        depth=2, route=
337        (_M + "big_picture", _M + "attention", _M + "transformer", _M + "tokenization", _M + "training_stages",
338         _M + "reasoning", _M + "inference"),
339    ),
340    LearningPath(
341        "Curious about images, audio and video",
342        "How models generate pictures and sound, and how they see and hear.",
343        depth=2, route=
344        (_M + "neural_net", _M + "cnn_rnn", _E + "contrastive", _G + "autoencoders", _G + "gans", _G + "diffusion",
345         _G + "multimodal"),
346    ),
347]
348
349
350def learning_paths_table() -> str:
351    """The README's learning paths. Regenerate with `make readme`."""
352    number = {l.module: i for i, l in enumerate(CURRICULUM)}
353    rows = ["| If you are… | You'll learn | Lessons, in order |", "|---|---|---|"]
354    for p in LEARNING_PATHS:
355        lessons = " → ".join(f"[{number[m]}]({source_path(m)})" for m in p.route)
356        rows.append(f"| **{p.who}** | {p.why} | {lessons} |")
357    return "\n".join(rows) + "\n"
358
359
360# ---------------------------------------------------------------------------
361# Big questions: the macro map. Lessons are organized bottom-up; real
362# conversations about AI systems start top-down with questions like these.
363# Each one lists the lessons that answer it, in order, and the short version:
364# the main ideas, in the order that builds understanding.
365# ---------------------------------------------------------------------------
366
367
368@dataclass(frozen=True)
369class BigQuestion:
370    question: str
371    route: tuple[str, ...]  # lesson modules, in the order to read them
372    in_brief: tuple[str, ...]  # the short version: the main ideas, in the order that builds them
373
374
375_ML, _EMB, _AG = "primer.ml.", "primer.ml.embeddings.", "primer.agents."
376_GEN = "primer.ml.generative."
377
378BIG_QUESTIONS: list[BigQuestion] = [
379    BigQuestion(
380        "What happens, step by step, when I send a prompt to a language model?",
381        ("primer.notation", _ML + "big_picture", _ML + "tokenization", _ML + "attention", _ML + "positional", _ML + "transformer", _ML + "inference"),
382        (
383            "Tokenizer: text becomes subword IDs; cost and context limits are counted in tokens.",
384            "Embedding lookup turns each ID into a vector; position information is mixed in.",
385            "Dozens of transformer blocks: attention mixes information across tokens, the feed-forward layer processes each token.",
386            "The last position's vector becomes a score for every vocabulary token; softmax turns scores into probabilities.",
387            "Sampling (temperature, top-p) picks one token, which is appended; the loop repeats until a stop token.",
388            "Prefill processes the prompt in parallel; decode generates one token at a time, made cheap by the KV cache.",
389        ),
390    ),
391    BigQuestion(
392        "How does a neural network actually learn?",
393        (_ML + "neural_net", _ML + "losses", _ML + "optimizers", _ML + "deep_nets", _ML + "regularization"),
394        (
395            "A forward pass makes a prediction; a loss turns 'how wrong' into one number.",
396            "Backpropagation applies the chain rule to find every weight's gradient.",
397            "An optimizer (SGD, Adam/AdamW) steps each weight against its gradient; the learning rate sets the step size.",
398            "Depth brings vanishing and exploding gradients; residual connections, normalization and good initialization fix them.",
399            "Watch validation loss: when it rises while training loss falls, the model is overfitting; regularize or stop early.",
400        ),
401    ),
402    BigQuestion(
403        "How does attention work, and why did transformers replace RNNs?",
404        (_ML + "attention", _ML + "positional", _ML + "transformer", _ML + "cnn_rnn"),
405        (
406            "Each token forms a query, key and value; query-key dot products score relevance; softmax turns scores into weights; the output blends values.",
407            "Scores are divided by the square root of d_k so softmax doesn't saturate and gradients keep flowing.",
408            "A causal mask hides future tokens, which makes next-token training honest and generation cacheable.",
409            "Multi-head attention runs several attentions in parallel; grouped-query attention shares keys and values to shrink the KV cache.",
410            "RNNs pass everything through one hidden state, one step at a time; attention gives every pair of tokens a direct path and trains in parallel.",
411            "The price is O(n²) cost in sequence length, which FlashAttention, sparse attention and state-space models attack.",
412        ),
413    ),
414    BigQuestion(
415        "How are large language models trained, and when should I fine-tune instead of using RAG?",
416        (_ML + "training_stages", _ML + "fine_tuning", _ML + "tokenization", _ML + "losses", _EMB + "operations", _AG + "rag"),
417        (
418            "Pretraining: next-token prediction over trillions of tokens produces a knowledgeable base model.",
419            "Supervised fine-tuning teaches the assistant format; preference tuning (RLHF or DPO) shapes helpfulness and safety.",
420            "Adaptation, cheapest first: prompting, then RAG, then LoRA, then (rarely) a full fine-tune.",
421            "Fine-tuning changes behavior; RAG supplies knowledge that changes or must be cited.",
422            "Distillation trains a small model to imitate a large one, often the biggest production cost win.",
423        ),
424    ),
425    BigQuestion(
426        "What makes serving a model fast and affordable?",
427        (_ML + "hardware", _ML + "inference", _ML + "efficient_architectures", _ML + "attention", _AG + "cost", _AG + "context"),
428        (
429            "Prefill is compute-bound and sets time to first token; decode is memory-bound and sets tokens per second.",
430            "The KV cache trades GPU memory for speed; its size is 2 × layers × KV heads × head dimension × bytes, per token.",
431            "Memory math: weights = parameters × bytes per parameter (70B at 16-bit is about 140 GB).",
432            "Speedups: quantization, continuous batching, speculative decoding, grouped-query attention, prompt caching.",
433            "At the system level: route easy work to small models, cache stable prefixes, trim tokens, batch offline work.",
434        ),
435    ),
436    BigQuestion(
437        "What is an embedding, and how is an embedding model trained?",
438        (_EMB + "word2vec", _EMB + "contrastive", _EMB + "similarity", _ML + "losses"),
439        (
440            "An embedding is a learned vector where closeness means similar meaning.",
441            "word2vec learned one vector per word from co-occurrence; contextual models give each token a vector that depends on its sentence.",
442            "Sentence embeddings pool token vectors; models trained for similarity beat plain pooled encoders.",
443            "Contrastive training pulls matching pairs together and pushes others apart, using in-batch negatives (InfoNCE).",
444            "Hard negatives (right topic, wrong answer) are the biggest driver of retrieval quality.",
445            "CLIP applies the same idea across images and text, putting both in one space.",
446        ),
447    ),
448    BigQuestion(
449        "How do you search millions of vectors quickly, and what does it cost?",
450        (_EMB + "similarity", _EMB + "compression", _EMB + "ann"),
451        (
452            "On normalized vectors, cosine, dot product and Euclidean distance give the same ranking; use what the model was trained with.",
453            "Storage math: vectors × dimensions × 4 bytes (10M × 1536 is about 61 GB) before index overhead.",
454            "Exact search is too slow at scale; approximate indexes trade a little recall for a lot of speed.",
455            "HNSW: layered graph, long jumps on top, local search at the bottom; M, efConstruction and efSearch are the knobs.",
456            "IVF searches only the nearest clusters (nprobe); PQ compresses vectors into codes.",
457            "Matryoshka truncation and scalar or binary quantization shrink memory; rescoring the shortlist recovers accuracy.",
458        ),
459    ),
460    BigQuestion(
461        "How do you build retrieval that returns the right passages?",
462        (_EMB + "retrieval", _EMB + "clustering", _EMB + "operations", _ML + "metrics", _AG + "rag"),
463        (
464            "Chunk on document structure, with overlap and metadata; chunking often matters more than the model.",
465            "Dense search finds meaning; BM25 finds exact IDs and rare terms; hybrid search fuses both with reciprocal rank fusion.",
466            "Retrieve wide with a bi-encoder, then rerank the shortlist with a cross-encoder.",
467            "Measure retrieval on its own with recall@k, MRR and nDCG on a labeled set before tuning prompts.",
468            "Operations: new embedding model means re-embedding everything; version indexes and switch traffic behind an alias.",
469        ),
470    ),
471    BigQuestion(
472        "How do you know if a model or an agent is any good?",
473        (_ML + "metrics", _ML + "benchmarks", _ML + "losses", _ML + "regularization", _AG + "evals"),
474        (
475            "Pick metrics by the cost of each error: precision vs. recall; accuracy misleads on imbalanced data.",
476            "Keep training, validation and test data apart, and watch for leakage and benchmark contamination.",
477            "For agents: a golden set of real tasks, graded by code wherever possible (end state, schema, tests).",
478            "For open-ended output: an LLM judge with an explicit rubric, calibrated against human labels.",
479            "Track trajectory, cost and latency beside quality; run the suite on every change and feed production failures back in.",
480        ),
481    ),
482    BigQuestion(
483        "When should you build an agent, and how does one work?",
484        (_AG + "llm", _AG + "orchestration", _AG + "agent_loop", _AG + "tools", _AG + "coding_agents", _AG + "mcp", _AG + "planning"),
485        (
486            "Use the least autonomy that solves the problem: fixed workflow, then router, then agent loop, then multiple agents.",
487            "Tool calling: the model emits a structured request; your code validates it, runs it and returns the result. The model executes nothing.",
488            "The loop: think, call a tool, observe, decide, with step limits, token budgets and loop detection.",
489            "Tool design is prompt design: few, high-level tools with precise descriptions, validated arguments and actionable errors.",
490            "Long tasks fail by compounding error (0.95¹⁰ ≈ 0.60), so plan, verify each step externally, and checkpoint.",
491            "MCP standardizes how apps connect to tools, and brings its own risks: tool poisoning, rug pulls, broad permissions.",
492        ),
493    ),
494    BigQuestion(
495        "How do you give an AI system the right context and memory?",
496        (_AG + "context", _AG + "memory", _AG + "rag"),
497        (
498            "Context engineering: the smallest set of high-signal content, structured with clear delimiters.",
499            "More context is not better: models use the middle of long inputs least reliably, and quality rots as sessions grow.",
500            "Put stable content first so prompt caching can reuse it; summarize or drop old turns; compress tool outputs.",
501            "Long-term memory is retrieval over the system's own history: episodic, semantic and procedural.",
502            "Memory must be isolated per tenant and user, updatable, and deletable on request.",
503        ),
504    ),
505    BigQuestion(
506        "How do you make an AI system safe to put in front of real users?",
507        (_ML + "alignment", _AG + "guardrails", _AG + "tools", _AG + "deployment", _AG + "observability"),
508        (
509            "Layer guardrails on inputs, outputs and actions; no single check is reliable alone.",
510            "Treat retrieved content, emails and tool outputs as untrusted: no prompt wording fully prevents injection.",
511            "Privilege separation: the part that reads untrusted content holds no dangerous tools; actions pass a policy or human check.",
512            "Graduate autonomy with evidence: shadow mode, then approval per action, then autonomy for low-risk actions.",
513            "Trace every run, keep tamper-evident audit logs, rate-limit actions and keep a kill switch and a one-step rollback.",
514        ),
515    ),
516    BigQuestion(
517        "How do you cut cost and latency without hurting quality?",
518        (_AG + "cost", _AG + "context", _ML + "inference", _EMB + "clustering"),
519        (
520            "Measure cost per successful task, not per call.",
521            "Route each step to the cheapest model that handles it; this is usually the biggest lever.",
522            "Cache: prompt caching for stable prefixes, response and semantic caches for repeated questions.",
523            "Trim tokens: tight prompts, compressed tool outputs, only the top reranked chunks.",
524            "Run independent tool calls in parallel, stream output, and move offline work to batch APIs.",
525            "Enforce per-task and per-tenant budgets with anomaly alerts.",
526        ),
527    ),
528    BigQuestion(
529        "Why do AI systems fail in production, and how do you fix them?",
530        (_AG + "failures", _ML + "structured_output", _AG + "planning", _AG + "rag", _AG + "evals", _AG + "observability"),
531        (
532            "Compounding error over long tasks: shorten paths, verify steps, checkpoint.",
533            "Bad retrieval behind confident wrong answers: hybrid search, reranking, retrieval evals.",
534            "Ambiguous tools, loops and runaway cost: better tool design, budgets, loop detection.",
535            "Prompt injection and messy enterprise data: untrusted-content boundaries, permission-aware retrieval, investment in parsing.",
536            "No evals and no traces: regressions ship silently; build the loop from production failure to trace to test case to fix.",
537        ),
538    ),
539    BigQuestion(
540        "What does it take to pretrain a large model?",
541        (_ML + "training_stages", _ML + "tokenization", _ML + "pretraining", _ML + "optimizers", _ML + "hardware"),
542        (
543            "Most of a web crawl is thrown away: language ID, quality rules and classifiers, and exact and near-duplicate removal.",
544            "Sources are mixed by weight, not size; about 20 tokens per parameter is compute-optimal, but models meant for heavy use train far longer.",
545            "Adam in mixed precision needs about 16 bytes per parameter before activations, so one GPU can't hold a large model.",
546            "Data parallelism shares gradients, ZeRO/FSDP shards the training state, tensor parallelism splits each matrix multiply, and pipeline parallelism splits the layers.",
547            "The maths runs in bf16 or fp8 with scaling, while the master weights stay in fp32.",
548            "Warmup, gradient clipping, spike rollback and regular checkpoints keep a months-long run alive.",
549        ),
550    ),
551    BigQuestion(
552        "How does a model learn from rewards instead of examples?",
553        (_ML + "reinforcement", _ML + "training_stages", _ML + "reasoning", _ML + "alignment"),
554        (
555            "Reinforcement learning samples an action, scores it, and makes high-scoring actions more likely.",
556            "A baseline turns rewards into advantages (better or worse than usual), which cuts noise without bias.",
557            "PPO reuses each batch for several steps, clips how far the policy moves, and leashes it to a reference with a KL penalty.",
558            "GRPO drops the value network by comparing rewards within a group of answers to the same prompt.",
559            "A checker as the reward (the right answer, passing tests) is how reasoning models are trained.",
560            "The policy optimises the reward you wrote, not the goal you meant; verifiable rewards, a KL leash and held-out checks defend against that.",
561        ),
562    ),
563    BigQuestion(
564        "How do reasoning models think, and when is extra thinking worth it?",
565        (_ML + "inference", _ML + "reinforcement", _ML + "reasoning", _AG + "planning", _AG + "cost"),
566        (
567            "Every written token is another forward pass, so a chain of thought buys serial computation, and the text is the model's working memory.",
568            "Test-time compute can go into one longer chain, or into many chains with a vote or a verifier picking one answer.",
569            "Voting helps only when the right answer is the most common one and the samples' mistakes are independent.",
570            "Checking each step catches errors that checking only the final answer misses.",
571            "Training with verifiable rewards makes longer, self-checking reasoning emerge.",
572            "Thinking is paid for per token and slips compound over long chains, so route easy tasks to little thinking and measure cost per successful task.",
573        ),
574    ),
575    BigQuestion(
576        "How do models handle very long contexts?",
577        (_ML + "attention", _ML + "positional", _ML + "inference", _ML + "efficient_architectures"),
578        (
579            "Attention scores every pair of tokens, and the KV cache grows with every token, per layer, per conversation.",
580            "Sliding windows and sparse patterns score fewer pairs; stacked layers still carry information far.",
581            "Linear attention and state-space models keep a fixed-size summary: linear time and constant memory, but blurrier recall.",
582            "Mamba makes the summary selective: each token decides how much to keep and how much to write.",
583            "Hybrids keep a few attention layers for exact lookup.",
584            "The cache shrinks by sharing key/value heads, caching a small latent, or storing fewer bits.",
585        ),
586    ),
587    BigQuestion(
588        "What is going on inside a trained model, and how can we tell?",
589        (_ML + "transformer", _EMB + "word2vec", _ML + "interpretability", _ML + "alignment"),
590        (
591            "Models store features as directions across many neurons, not one feature per neuron.",
592            "Probes and the logit lens read what is present; they show correlation, not use.",
593            "Activation patching changes one activation and watches the output: the causal test.",
594            "Sparse features get packed in superposition, which makes individual neurons respond to several things.",
595            "Sparse autoencoders unpack superposition into interpretable features, at the cost of some unexplained activity.",
596            "These tools give evidence, not proof; full explanations exist only for narrow behaviours.",
597        ),
598    ),
599    BigQuestion(
600        "When is a neural network the wrong tool?",
601        (_ML + "classical", _ML + "neural_net", _ML + "regularization", _ML + "metrics"),
602        (
603            "On tables whose columns each mean something alone, gradient-boosted trees or a random forest are the model to beat.",
604            "A tree asks one column at a time whether it's above a threshold, so it needs no feature scaling and handles categories natively.",
605            "A single deep tree overfits; forests average many decorrelated trees, and boosting adds small trees fit to the remaining errors.",
606            "Neural networks win when meaning lives in arrangements of raw values (images, audio, text), when data is huge, or when a pretrained model can be reused.",
607            "Trees can't extrapolate beyond the values they trained on, and impurity importances credit noise, so check importances on held-out data.",
608        ),
609    ),
610    BigQuestion(
611        "How do AI models generate images, audio and video?",
612        (_GEN + "autoencoders", _GEN + "gans", _GEN + "diffusion", _GEN + "multimodal"),
613        (
614            "A generator learns a whole distribution, so it can sample new examples; predicting the average gives blur.",
615            "An autoencoder squeezes data into a small code and back; a VAE shapes that code so random codes decode to new data.",
616            "A GAN trains a generator against a discriminator: sharp, one-pass samples, but unstable training and mode collapse.",
617            "Diffusion adds noise on purpose and learns to remove it, generating from pure noise in many small steps.",
618            "Flow matching learns straight paths from noise to data, so it needs fewer steps; guidance trades variety for following the prompt.",
619            "Real systems denoise an autoencoder's latent with a transformer that reads the prompt; video and audio are the same idea with more tokens.",
620        ),
621    ),
622    BigQuestion(
623        "How do AI models see images and hear audio?",
624        (_ML + "cnn_rnn", _EMB + "contrastive", _GEN + "multimodal"),
625        (
626            "Every modality becomes a sequence of vectors a transformer attends over.",
627            "Images become patch tokens, (H/P)·(W/P) of them, so cost grows with the square of the resolution.",
628            "A small projector connects a vision or audio encoder to a language model; it is trained first, with both models frozen.",
629            "Audio becomes a log-mel spectrogram, then about 50 tokens a second.",
630            "Video multiplies image tokens by time, so frames are sampled; text stays the densest input.",
631        ),
632    ),
633]
634
635
636def _lesson(module: str) -> "Lesson":
637    return CURRICULUM[_BY_MODULE[module]]
638
639
640def big_questions_table() -> str:
641    """The README's compact map: each big question and its route. Regenerate with `make readme`."""
642    rows = ["| Big question | Lessons that answer it, in order |", "|---|---|"]
643    for q in BIG_QUESTIONS:
644        route = ", ".join(f"[{_lesson(m).title}]({m.replace('.', '/')}.py)" for m in q.route)
645        rows.append(f"| {q.question} | {route} |")
646    return link_code_references("\n".join(rows) + "\n", '.')
647
648
649def big_questions_page() -> str:
650    """docs/BIG_QUESTIONS.md: every big question with its route and the short version."""
651    out = [
652        "# Big questions: the map from the top down\n",
653        "The lessons build the field from the bottom up. Real conversations about AI systems start from the top, "
654        "with questions like these. For each one: the lessons that answer it, in order, and the short version, "
655        "the main ideas in the order that builds understanding. Read the short version first, then open the "
656        "lessons wherever you want the full story.\n",
657        "Generated from `primer/curriculum.py` by `make readme`; edit it there.\n",
658    ]
659    for n, q in enumerate(BIG_QUESTIONS, 1):
660        route = " → ".join(f"[{_lesson(m).title}](../{m.replace('.', '/')}.py)" for m in q.route)
661        in_brief = "\n".join(f"{i}. {point}" for i, point in enumerate(q.in_brief, 1))
662        out.append(f"\n## {n}. {q.question}\n\n**Route:** {route}\n\n**In brief:**\n\n{in_brief}\n")
663    return link_code_references("\n".join(out), 'docs')
664
665
666def self_test_book() -> str:
667    """Every lesson's self-test questions, in reading order, as one markdown page.
668
669    Generated from the lesson docstrings (the single source of truth) into
670    docs/SELF_TEST.md by `make readme`.
671    """
672    import importlib
673    import re
674
675    out = [
676        "# Self-test: every question in the primer\n",
677        "Generated from each lesson's `## Self-test questions` section by `make readme`; "
678        "edit the lesson, not this file. Answer each question out loud before reading the answer.\n",
679    ]
680    for part in PARTS:
681        out.append(f"\n## {part.title}\n")
682        for i, lesson in lessons_in(part.key):
683            try:
684                doc = importlib.import_module(lesson.module).__doc__ or ""
685            except ModuleNotFoundError:
686                continue
687            m = re.search(r"^## Self-test questions\s*\n(.*?)(?=^## |\Z)", doc, re.S | re.M)
688            if not m:
689                continue
690            body = re.sub(r"^(#+) ", lambda h: "#" * (len(h.group(1)) + 2) + " ", m.group(1).strip(), flags=re.M)
691            link = lesson.module.replace(".", "/") + ".py"
692            out.append(f"\n### {i}. {lesson.title}\n\nFrom [`{lesson.module}`](../{link}).\n\n{body}\n")
693    return link_code_references("\n".join(out), "docs")
694
695
696def _render_doc() -> str:
697    parts = []
698    for part in PARTS:
699        parts.append(f"\n## {part.title}\n\n{part.blurb}\n")
700        parts += [f"{i}. `{l.module}`: **{l.title}.** {l.outcome}." for i, l in lessons_in(part.key)]
701    return __doc__ + "\n".join(parts) + "\n"
702
703
704__doc__ = _render_doc()
705
706
707if __name__ == "__main__":
708    # `make readme`: rewrite the generated section of README.md in place.
709    import re
710    import sys
711    from pathlib import Path
712
713    readme = Path(__file__).resolve().parent.parent / "README.md"
714    text = readme.read_text()
715    new, n = re.subn(
716        r"(<!-- BEGIN curriculum -->\n).*?(<!-- END curriculum -->)",
717        lambda m: m.group(1) + readme_section() + m.group(2),
718        text,
719        flags=re.S,
720    )
721    if n != 1:
722        sys.exit("README.md needs exactly one <!-- BEGIN curriculum --> ... <!-- END curriculum --> block")
723    new, n = re.subn(
724        r"(<!-- BEGIN big-questions -->\n).*?(<!-- END big-questions -->)",
725        lambda m: m.group(1) + big_questions_table() + m.group(2),
726        new,
727        flags=re.S,
728    )
729    if n != 1:
730        sys.exit("README.md needs exactly one <!-- BEGIN big-questions --> ... <!-- END big-questions --> block")
731    new, n = re.subn(
732        r"(<!-- BEGIN paths -->\n).*?(<!-- END paths -->)",
733        lambda m: m.group(1) + learning_paths_table() + m.group(2),
734        new,
735        flags=re.S,
736    )
737    if n != 1:
738        sys.exit("README.md needs exactly one <!-- BEGIN paths --> ... <!-- END paths --> block")
739    readme.write_text(new)
740    (readme.parent / "docs" / "SELF_TEST.md").write_text(self_test_book())
741    (readme.parent / "docs" / "BIG_QUESTIONS.md").write_text(big_questions_page())
742    print("README.md, docs/SELF_TEST.md and docs/BIG_QUESTIONS.md regenerated from primer/curriculum.py")
REPO_URL = 'https://github.com/that-mathevs/ai-primer'
SITE_URL = 'https://that-mathevs.github.io/ai-primer/'
BRANCH = 'main'
def source_path(module: str) -> str: on GitHub
30def source_path(module: str) -> str:
31    """The file a module lives in, relative to the repository root."""
32    return module.replace(".", "/") + ".py"

The file a module lives in, relative to the repository root.

def tests_for(module: str) -> str: on GitHub
150def tests_for(module: str) -> str:
151    """The test file that specifies a lesson (tests/test_<name>.py, test_emb_ or test_agents_)."""
152    name = module.rsplit(".", 1)[-1]
153    prefix = "emb_" if module.startswith("primer.ml.embeddings.") else "agents_" if module.startswith("primer.agents.") else ""
154    return f"tests/test_{prefix}{name}.py"

The test file that specifies a lesson (tests/test_.py, test_emb_ or test_agents_).

@dataclass(frozen=True)
class Part: on GitHub
157@dataclass(frozen=True)
158class Part:
159    key: str
160    title: str
161    blurb: str
Part(key: str, title: str, blurb: str)
key: str
title: str
blurb: str
@dataclass(frozen=True)
class Lesson: on GitHub
164@dataclass(frozen=True)
165class Lesson:
166    module: str
167    title: str
168    outcome: str  # "What you'll be able to explain"
169    part: str  # Part.key
Lesson(module: str, title: str, outcome: str, part: str)
module: str
title: str
outcome: str
part: str
PARTS: list[Part] = [Part(key='start', title='Before you begin', blurb='The notation every formula in this primer uses, decoded as short loops.'), Part(key='ml', title='Part 1: how the model works inside', blurb='From a single neuron to a working transformer, and how models are trained and served.'), Part(key='embeddings', title='Embeddings, the centerpiece', blurb='Vectors that capture meaning, and the search systems built on them.'), Part(key='generative', title='Generating images, audio and video', blurb='Autoencoders, GANs, diffusion, and the multimodal models that connect them to language.'), Part(key='agents', title='Part 2: building systems people rely on', blurb='Agents, tools, retrieval, memory, evaluation, safety, cost and deployment.')]
CURRICULUM: list[Lesson] = [Lesson(module='primer.notation', title='Math notation, from zero', outcome='Every symbol in an ML formula, as a short loop', part='start'), Lesson(module='primer.ml.big_picture', title='The big picture', outcome='What happens, end to end, when you send a prompt', part='ml'), Lesson(module='primer.ml.neural_net', title='Neural networks', outcome='Neurons, activations, the forward pass, backprop by hand', part='ml'), Lesson(module='primer.ml.optimizers', title='Optimizers', outcome='SGD, momentum, Adam/AdamW, learning-rate warmup and decay', part='ml'), Lesson(module='primer.ml.deep_nets', title='Training deep networks', outcome='Vanishing/exploding gradients, residuals, normalization, initialization', part='ml'), Lesson(module='primer.ml.attention', title='Attention', outcome='Queries, keys, values, softmax, masking, multi-head, GQA, O(n²)', part='ml'), Lesson(module='primer.ml.positional', title='Positional information', outcome='Why order must be added, sinusoids and RoPE', part='ml'), Lesson(module='primer.ml.transformer', title='The transformer', outcome='The block, a tiny GPT, parameter counts, mixture of experts', part='ml'), Lesson(module='primer.ml.tokenization', title='Tokenization', outcome='BPE from scratch, byte-level tokens, why models miscount letters', part='ml'), Lesson(module='primer.ml.training_stages', title='Training stages', outcome='Pretraining, SFT, RLHF and DPO, LoRA, fine-tuning vs. RAG', part='ml'), Lesson(module='primer.ml.pretraining', title='Pretraining at scale', outcome='Data curation and deduplication, parallelism across GPUs, mixed precision', part='ml'), Lesson(module='primer.ml.fine_tuning', title='Fine-tuning in practice', outcome='Preparing data, forgetting old skills, merging models', part='ml'), Lesson(module='primer.ml.reinforcement', title='Reinforcement learning', outcome='Policy gradients from scratch, PPO, GRPO, reward hacking', part='ml'), Lesson(module='primer.ml.reasoning', title='Reasoning models', outcome='Chain of thought, test-time compute, verifiers, learning to reason with RL', part='ml'), Lesson(module='primer.ml.alignment', title='Alignment and safety', outcome='Constitutional AI, red-teaming, sycophancy, refusals', part='ml'), Lesson(module='primer.ml.hardware', title='The hardware underneath', outcome='GPUs, the memory hierarchy, FLOPs vs. bandwidth, number formats', part='ml'), Lesson(module='primer.ml.inference', title='Inference', outcome='Prefill vs. decode, the KV cache, sampling, speculative decoding, memory math', part='ml'), Lesson(module='primer.ml.structured_output', title='Structured output', outcome='Constrained decoding: grammars and JSON schemas that guarantee valid output', part='ml'), Lesson(module='primer.ml.efficient_architectures', title='Long context and efficient architectures', outcome='Sliding-window and sparse attention, state-space models, KV-cache compression', part='ml'), Lesson(module='primer.ml.losses', title='Loss functions', outcome='Cross-entropy, perplexity, MSE/MAE, contrastive losses', part='ml'), Lesson(module='primer.ml.metrics', title='Metrics', outcome='Precision/recall/F1, ROC-AUC, recall@k, MRR, nDCG, BLEU/ROUGE', part='ml'), Lesson(module='primer.ml.benchmarks', title='Reading benchmarks', outcome='What benchmarks measure, contamination, leaderboards and arenas', part='ml'), Lesson(module='primer.ml.regularization', title='Overfitting and regularization', outcome='Overfitting, early stopping, dropout, L1/L2, leakage', part='ml'), Lesson(module='primer.ml.classical', title='Trees and boosting', outcome='Decision trees, random forests, gradient boosting, and when they still win', part='ml'), Lesson(module='primer.ml.cnn_rnn', title='CNNs and RNNs', outcome='How convolutions see and recurrent nets remember, and why transformers won', part='ml'), Lesson(module='primer.ml.interpretability', title='Looking inside the model', outcome='Probes, the logit lens, activation patching, superposition, sparse autoencoders', part='ml'), Lesson(module='primer.ml.embeddings.word2vec', title='Word embeddings', outcome='Where embeddings came from, analogies, the "bank" problem', part='embeddings'), Lesson(module='primer.ml.embeddings.similarity', title='Similarity', outcome='Cosine vs. dot vs. distance, normalization, anisotropy, thresholds', part='embeddings'), Lesson(module='primer.ml.embeddings.contrastive', title='Training embedding models', outcome='Contrastive learning, hard negatives, CLIP', part='embeddings'), Lesson(module='primer.ml.embeddings.compression', title='Dimensions and compression', outcome='Storage math, Matryoshka truncation, int8 and binary quantization', part='embeddings'), Lesson(module='primer.ml.embeddings.ann', title='Vector indexes', outcome='Flat, IVF, PQ and HNSW from scratch, recall vs. latency', part='embeddings'), Lesson(module='primer.ml.embeddings.retrieval', title='Retrieval', outcome='BM25, hybrid search with RRF, rerankers, ColBERT, chunking', part='embeddings'), Lesson(module='primer.ml.embeddings.clustering', title='Clustering and matching', outcome='k-means, density clustering, dedup, routing, semantic caching', part='embeddings'), Lesson(module='primer.ml.embeddings.operations', title='Embeddings in production', outcome='Model migrations, domain mismatch, measuring retrieval on its own', part='embeddings'), Lesson(module='primer.ml.generative.autoencoders', title='Autoencoders and VAEs', outcome='Squeezing data into a code and back, and sampling new data from it', part='generative'), Lesson(module='primer.ml.generative.gans', title='GANs', outcome='A forger against a detective: adversarial training, and why it is unstable', part='generative'), Lesson(module='primer.ml.generative.diffusion', title='Diffusion and flow matching', outcome='Turning noise into images one small step at a time', part='generative'), Lesson(module='primer.ml.generative.multimodal', title='Multimodal models', outcome='Images, audio and video into a language model', part='generative'), Lesson(module='primer.agents.llm', title='Talking to a model', outcome='The message format, and what tool calling really is', part='agents'), Lesson(module='primer.agents.orchestration', title='Orchestration', outcome='Workflows vs. agents, and the named patterns', part='agents'), Lesson(module='primer.agents.agent_loop', title='The agent loop', outcome='A production agent loop: budgets, loop detection, recovery', part='agents'), Lesson(module='primer.agents.tools', title='Tools', outcome='Tool design, validation, idempotency, approvals, least privilege', part='agents'), Lesson(module='primer.agents.coding_agents', title='Coding and computer-use agents', outcome='Edit, run, test, repeat; sandboxes; driving a screen', part='agents'), Lesson(module='primer.agents.mcp', title='Model Context Protocol', outcome='MCP on the wire, and its security risks', part='agents'), Lesson(module='primer.agents.rag', title='Retrieval-augmented generation', outcome='RAG end to end, with citations and access control', part='agents'), Lesson(module='primer.agents.context', title='Context engineering', outcome='What goes in the window, compression, cache-friendly layout', part='agents'), Lesson(module='primer.agents.memory', title='Memory', outcome='Short- and long-term memory, tenant isolation, forgetting', part='agents'), Lesson(module='primer.agents.planning', title='Planning', outcome='Plan-and-execute, decomposition, reflection, compounding error', part='agents'), Lesson(module='primer.agents.evals', title='Evaluation', outcome='Golden sets, graders, LLM-as-judge calibration', part='agents'), Lesson(module='primer.agents.guardrails', title='Guardrails', outcome='Prompt injection and privilege separation, PII, output checks', part='agents'), Lesson(module='primer.agents.cost', title='Cost and latency', outcome='Routing, caching, batching, budgets, cost per successful task', part='agents'), Lesson(module='primer.agents.observability', title='Observability', outcome='Traces, OpenTelemetry GenAI attributes, the improvement loop', part='agents'), Lesson(module='primer.agents.deployment', title='Safe deployment', outcome='Shadow mode, graduated autonomy, canaries, kill switches, audit logs', part='agents'), Lesson(module='primer.agents.failures', title='Why the hard ones fail', outcome='The common failure modes, and the fix for each', part='agents')]
def neighbours( module: str) -> tuple[Lesson | None, Lesson | None]: on GitHub
240def neighbours(module: str) -> tuple[Lesson | None, Lesson | None]:
241    """(previous lesson, next lesson) in reading order; None at either end."""
242    i = _BY_MODULE[module]
243    return (CURRICULUM[i - 1] if i > 0 else None, CURRICULUM[i + 1] if i + 1 < len(CURRICULUM) else None)

(previous lesson, next lesson) in reading order; None at either end.

def lessons_in(part_key: str) -> list[tuple[int, Lesson]]: on GitHub
246def lessons_in(part_key: str) -> list[tuple[int, Lesson]]:
247    """(number, lesson) pairs for one part, numbered across the whole curriculum."""
248    return [(i, l) for i, l in enumerate(CURRICULUM) if l.part == part_key]

(number, lesson) pairs for one part, numbered across the whole curriculum.

def readme_section() -> str: on GitHub
251def readme_section() -> str:
252    """The README's reading-order tables. Regenerate with `make readme`."""
253    out = []
254    for part in PARTS:
255        out.append(f"### {part.title}\n\n{part.blurb}\n\n| # | Lesson | What you'll be able to explain | Read |\n|---|---|---|---|")
256        out += [
257            f"| {i} | [{l.title}]({source_path(l.module)}) | {l.outcome} | "
258            f"[page]({SITE_URL}{l.module.replace('.', '/')}.html) · [tests]({tests_for(l.module)}) |"
259            for i, l in lessons_in(part.key)
260        ]
261        out.append("")
262    return link_code_references("\n".join(out) + "\n", '.')

The README's reading-order tables. Regenerate with make readme.

def reading_list(package: str) -> str: on GitHub
265def reading_list(package: str) -> str:
266    """A markdown reading list of the lessons inside `package`, for its page."""
267    rows = [
268        f"{i}. `{l.module}`: **{l.title}.** {l.outcome}."
269        for i, l in enumerate(CURRICULUM)
270        if l.module.startswith(package + ".") and "." not in l.module[len(package) + 1 :]
271    ]
272    return "\n## Reading order\n\n" + "\n".join(rows) + "\n\nThe full map is in `primer.curriculum`.\n"

A markdown reading list of the lessons inside package, for its page.

LEVELS_PENDING: frozenset[str] = frozenset()
@dataclass(frozen=True)
class LearningPath: on GitHub
290@dataclass(frozen=True)
291class LearningPath:
292    who: str  # the reader it's for
293    why: str  # what they get out of it
294    route: tuple[str, ...]  # lesson modules, in reading order
295    # How deep this reader goes by default: 1 the practitioner's guide, 2 the mechanism built
296    # from scratch, 3 its math and code as well. The site opens each lesson at this depth.
297    depth: int = 2
LearningPath(who: str, why: str, route: tuple[str, ...], depth: int = 2)
who: str
why: str
route: tuple[str, ...]
depth: int = 2
LEARNING_PATHS: list[LearningPath] = [LearningPath(who='Software engineer new to AI', why='How a language model works, then how to build on one.', route=('primer.notation', 'primer.ml.big_picture', 'primer.ml.neural_net', 'primer.ml.attention', 'primer.ml.transformer', 'primer.ml.tokenization', 'primer.ml.inference', 'primer.ml.embeddings.similarity', 'primer.ml.embeddings.retrieval', 'primer.agents.llm', 'primer.agents.agent_loop', 'primer.agents.rag', 'primer.agents.evals'), depth=2), LearningPath(who='AI application engineer', why='Agents, retrieval and tools, and keeping them reliable, safe and affordable.', route=('primer.ml.structured_output', 'primer.ml.embeddings.retrieval', 'primer.agents.llm', 'primer.agents.orchestration', 'primer.agents.agent_loop', 'primer.agents.tools', 'primer.agents.coding_agents', 'primer.agents.mcp', 'primer.agents.rag', 'primer.agents.context', 'primer.agents.memory', 'primer.agents.evals', 'primer.agents.guardrails', 'primer.agents.cost', 'primer.agents.observability', 'primer.agents.deployment', 'primer.agents.failures'), depth=1), LearningPath(who='ML engineer', why='The model itself: training, scaling, serving and looking inside.', route=('primer.notation', 'primer.ml.neural_net', 'primer.ml.optimizers', 'primer.ml.deep_nets', 'primer.ml.attention', 'primer.ml.positional', 'primer.ml.transformer', 'primer.ml.training_stages', 'primer.ml.pretraining', 'primer.ml.fine_tuning', 'primer.ml.reinforcement', 'primer.ml.hardware', 'primer.ml.inference', 'primer.ml.efficient_architectures', 'primer.ml.losses', 'primer.ml.metrics', 'primer.ml.benchmarks', 'primer.ml.regularization', 'primer.ml.interpretability'), depth=3), LearningPath(who='Engineering manager or architect', why='What these systems can do, what they cost, and how they fail.', route=('primer.ml.big_picture', 'primer.ml.training_stages', 'primer.ml.reasoning', 'primer.ml.alignment', 'primer.ml.inference', 'primer.ml.benchmarks', 'primer.agents.orchestration', 'primer.agents.rag', 'primer.agents.evals', 'primer.agents.cost', 'primer.agents.deployment', 'primer.agents.failures'), depth=1), LearningPath(who='Just explain LLMs to me', why='The shortest route to understanding what happens when you send a prompt.', route=('primer.ml.big_picture', 'primer.ml.attention', 'primer.ml.transformer', 'primer.ml.tokenization', 'primer.ml.training_stages', 'primer.ml.reasoning', 'primer.ml.inference'), depth=2), LearningPath(who='Curious about images, audio and video', why='How models generate pictures and sound, and how they see and hear.', route=('primer.ml.neural_net', 'primer.ml.cnn_rnn', 'primer.ml.embeddings.contrastive', 'primer.ml.generative.autoencoders', 'primer.ml.generative.gans', 'primer.ml.generative.diffusion', 'primer.ml.generative.multimodal'), depth=2)]
def learning_paths_table() -> str: on GitHub
351def learning_paths_table() -> str:
352    """The README's learning paths. Regenerate with `make readme`."""
353    number = {l.module: i for i, l in enumerate(CURRICULUM)}
354    rows = ["| If you are… | You'll learn | Lessons, in order |", "|---|---|---|"]
355    for p in LEARNING_PATHS:
356        lessons = " → ".join(f"[{number[m]}]({source_path(m)})" for m in p.route)
357        rows.append(f"| **{p.who}** | {p.why} | {lessons} |")
358    return "\n".join(rows) + "\n"

The README's learning paths. Regenerate with make readme.

@dataclass(frozen=True)
class BigQuestion: on GitHub
369@dataclass(frozen=True)
370class BigQuestion:
371    question: str
372    route: tuple[str, ...]  # lesson modules, in the order to read them
373    in_brief: tuple[str, ...]  # the short version: the main ideas, in the order that builds them
BigQuestion(question: str, route: tuple[str, ...], in_brief: tuple[str, ...])
question: str
route: tuple[str, ...]
in_brief: tuple[str, ...]
BIG_QUESTIONS: list[BigQuestion] = [BigQuestion(question='What happens, step by step, when I send a prompt to a language model?', route=('primer.notation', 'primer.ml.big_picture', 'primer.ml.tokenization', 'primer.ml.attention', 'primer.ml.positional', 'primer.ml.transformer', 'primer.ml.inference'), in_brief=('Tokenizer: text becomes subword IDs; cost and context limits are counted in tokens.', 'Embedding lookup turns each ID into a vector; position information is mixed in.', 'Dozens of transformer blocks: attention mixes information across tokens, the feed-forward layer processes each token.', "The last position's vector becomes a score for every vocabulary token; softmax turns scores into probabilities.", 'Sampling (temperature, top-p) picks one token, which is appended; the loop repeats until a stop token.', 'Prefill processes the prompt in parallel; decode generates one token at a time, made cheap by the KV cache.')), BigQuestion(question='How does a neural network actually learn?', route=('primer.ml.neural_net', 'primer.ml.losses', 'primer.ml.optimizers', 'primer.ml.deep_nets', 'primer.ml.regularization'), in_brief=("A forward pass makes a prediction; a loss turns 'how wrong' into one number.", "Backpropagation applies the chain rule to find every weight's gradient.", 'An optimizer (SGD, Adam/AdamW) steps each weight against its gradient; the learning rate sets the step size.', 'Depth brings vanishing and exploding gradients; residual connections, normalization and good initialization fix them.', 'Watch validation loss: when it rises while training loss falls, the model is overfitting; regularize or stop early.')), BigQuestion(question='How does attention work, and why did transformers replace RNNs?', route=('primer.ml.attention', 'primer.ml.positional', 'primer.ml.transformer', 'primer.ml.cnn_rnn'), in_brief=('Each token forms a query, key and value; query-key dot products score relevance; softmax turns scores into weights; the output blends values.', "Scores are divided by the square root of d_k so softmax doesn't saturate and gradients keep flowing.", 'A causal mask hides future tokens, which makes next-token training honest and generation cacheable.', 'Multi-head attention runs several attentions in parallel; grouped-query attention shares keys and values to shrink the KV cache.', 'RNNs pass everything through one hidden state, one step at a time; attention gives every pair of tokens a direct path and trains in parallel.', 'The price is O(n²) cost in sequence length, which FlashAttention, sparse attention and state-space models attack.')), BigQuestion(question='How are large language models trained, and when should I fine-tune instead of using RAG?', route=('primer.ml.training_stages', 'primer.ml.fine_tuning', 'primer.ml.tokenization', 'primer.ml.losses', 'primer.ml.embeddings.operations', 'primer.agents.rag'), in_brief=('Pretraining: next-token prediction over trillions of tokens produces a knowledgeable base model.', 'Supervised fine-tuning teaches the assistant format; preference tuning (RLHF or DPO) shapes helpfulness and safety.', 'Adaptation, cheapest first: prompting, then RAG, then LoRA, then (rarely) a full fine-tune.', 'Fine-tuning changes behavior; RAG supplies knowledge that changes or must be cited.', 'Distillation trains a small model to imitate a large one, often the biggest production cost win.')), BigQuestion(question='What makes serving a model fast and affordable?', route=('primer.ml.hardware', 'primer.ml.inference', 'primer.ml.efficient_architectures', 'primer.ml.attention', 'primer.agents.cost', 'primer.agents.context'), in_brief=('Prefill is compute-bound and sets time to first token; decode is memory-bound and sets tokens per second.', 'The KV cache trades GPU memory for speed; its size is 2 × layers × KV heads × head dimension × bytes, per token.', 'Memory math: weights = parameters × bytes per parameter (70B at 16-bit is about 140 GB).', 'Speedups: quantization, continuous batching, speculative decoding, grouped-query attention, prompt caching.', 'At the system level: route easy work to small models, cache stable prefixes, trim tokens, batch offline work.')), BigQuestion(question='What is an embedding, and how is an embedding model trained?', route=('primer.ml.embeddings.word2vec', 'primer.ml.embeddings.contrastive', 'primer.ml.embeddings.similarity', 'primer.ml.losses'), in_brief=('An embedding is a learned vector where closeness means similar meaning.', 'word2vec learned one vector per word from co-occurrence; contextual models give each token a vector that depends on its sentence.', 'Sentence embeddings pool token vectors; models trained for similarity beat plain pooled encoders.', 'Contrastive training pulls matching pairs together and pushes others apart, using in-batch negatives (InfoNCE).', 'Hard negatives (right topic, wrong answer) are the biggest driver of retrieval quality.', 'CLIP applies the same idea across images and text, putting both in one space.')), BigQuestion(question='How do you search millions of vectors quickly, and what does it cost?', route=('primer.ml.embeddings.similarity', 'primer.ml.embeddings.compression', 'primer.ml.embeddings.ann'), in_brief=('On normalized vectors, cosine, dot product and Euclidean distance give the same ranking; use what the model was trained with.', 'Storage math: vectors × dimensions × 4 bytes (10M × 1536 is about 61 GB) before index overhead.', 'Exact search is too slow at scale; approximate indexes trade a little recall for a lot of speed.', 'HNSW: layered graph, long jumps on top, local search at the bottom; M, efConstruction and efSearch are the knobs.', 'IVF searches only the nearest clusters (nprobe); PQ compresses vectors into codes.', 'Matryoshka truncation and scalar or binary quantization shrink memory; rescoring the shortlist recovers accuracy.')), BigQuestion(question='How do you build retrieval that returns the right passages?', route=('primer.ml.embeddings.retrieval', 'primer.ml.embeddings.clustering', 'primer.ml.embeddings.operations', 'primer.ml.metrics', 'primer.agents.rag'), in_brief=('Chunk on document structure, with overlap and metadata; chunking often matters more than the model.', 'Dense search finds meaning; BM25 finds exact IDs and rare terms; hybrid search fuses both with reciprocal rank fusion.', 'Retrieve wide with a bi-encoder, then rerank the shortlist with a cross-encoder.', 'Measure retrieval on its own with recall@k, MRR and nDCG on a labeled set before tuning prompts.', 'Operations: new embedding model means re-embedding everything; version indexes and switch traffic behind an alias.')), BigQuestion(question='How do you know if a model or an agent is any good?', route=('primer.ml.metrics', 'primer.ml.benchmarks', 'primer.ml.losses', 'primer.ml.regularization', 'primer.agents.evals'), in_brief=('Pick metrics by the cost of each error: precision vs. recall; accuracy misleads on imbalanced data.', 'Keep training, validation and test data apart, and watch for leakage and benchmark contamination.', 'For agents: a golden set of real tasks, graded by code wherever possible (end state, schema, tests).', 'For open-ended output: an LLM judge with an explicit rubric, calibrated against human labels.', 'Track trajectory, cost and latency beside quality; run the suite on every change and feed production failures back in.')), BigQuestion(question='When should you build an agent, and how does one work?', route=('primer.agents.llm', 'primer.agents.orchestration', 'primer.agents.agent_loop', 'primer.agents.tools', 'primer.agents.coding_agents', 'primer.agents.mcp', 'primer.agents.planning'), in_brief=('Use the least autonomy that solves the problem: fixed workflow, then router, then agent loop, then multiple agents.', 'Tool calling: the model emits a structured request; your code validates it, runs it and returns the result. The model executes nothing.', 'The loop: think, call a tool, observe, decide, with step limits, token budgets and loop detection.', 'Tool design is prompt design: few, high-level tools with precise descriptions, validated arguments and actionable errors.', 'Long tasks fail by compounding error (0.95¹⁰ ≈ 0.60), so plan, verify each step externally, and checkpoint.', 'MCP standardizes how apps connect to tools, and brings its own risks: tool poisoning, rug pulls, broad permissions.')), BigQuestion(question='How do you give an AI system the right context and memory?', route=('primer.agents.context', 'primer.agents.memory', 'primer.agents.rag'), in_brief=('Context engineering: the smallest set of high-signal content, structured with clear delimiters.', 'More context is not better: models use the middle of long inputs least reliably, and quality rots as sessions grow.', 'Put stable content first so prompt caching can reuse it; summarize or drop old turns; compress tool outputs.', "Long-term memory is retrieval over the system's own history: episodic, semantic and procedural.", 'Memory must be isolated per tenant and user, updatable, and deletable on request.')), BigQuestion(question='How do you make an AI system safe to put in front of real users?', route=('primer.ml.alignment', 'primer.agents.guardrails', 'primer.agents.tools', 'primer.agents.deployment', 'primer.agents.observability'), in_brief=('Layer guardrails on inputs, outputs and actions; no single check is reliable alone.', 'Treat retrieved content, emails and tool outputs as untrusted: no prompt wording fully prevents injection.', 'Privilege separation: the part that reads untrusted content holds no dangerous tools; actions pass a policy or human check.', 'Graduate autonomy with evidence: shadow mode, then approval per action, then autonomy for low-risk actions.', 'Trace every run, keep tamper-evident audit logs, rate-limit actions and keep a kill switch and a one-step rollback.')), BigQuestion(question='How do you cut cost and latency without hurting quality?', route=('primer.agents.cost', 'primer.agents.context', 'primer.ml.inference', 'primer.ml.embeddings.clustering'), in_brief=('Measure cost per successful task, not per call.', 'Route each step to the cheapest model that handles it; this is usually the biggest lever.', 'Cache: prompt caching for stable prefixes, response and semantic caches for repeated questions.', 'Trim tokens: tight prompts, compressed tool outputs, only the top reranked chunks.', 'Run independent tool calls in parallel, stream output, and move offline work to batch APIs.', 'Enforce per-task and per-tenant budgets with anomaly alerts.')), BigQuestion(question='Why do AI systems fail in production, and how do you fix them?', route=('primer.agents.failures', 'primer.ml.structured_output', 'primer.agents.planning', 'primer.agents.rag', 'primer.agents.evals', 'primer.agents.observability'), in_brief=('Compounding error over long tasks: shorten paths, verify steps, checkpoint.', 'Bad retrieval behind confident wrong answers: hybrid search, reranking, retrieval evals.', 'Ambiguous tools, loops and runaway cost: better tool design, budgets, loop detection.', 'Prompt injection and messy enterprise data: untrusted-content boundaries, permission-aware retrieval, investment in parsing.', 'No evals and no traces: regressions ship silently; build the loop from production failure to trace to test case to fix.')), BigQuestion(question='What does it take to pretrain a large model?', route=('primer.ml.training_stages', 'primer.ml.tokenization', 'primer.ml.pretraining', 'primer.ml.optimizers', 'primer.ml.hardware'), in_brief=('Most of a web crawl is thrown away: language ID, quality rules and classifiers, and exact and near-duplicate removal.', 'Sources are mixed by weight, not size; about 20 tokens per parameter is compute-optimal, but models meant for heavy use train far longer.', "Adam in mixed precision needs about 16 bytes per parameter before activations, so one GPU can't hold a large model.", 'Data parallelism shares gradients, ZeRO/FSDP shards the training state, tensor parallelism splits each matrix multiply, and pipeline parallelism splits the layers.', 'The maths runs in bf16 or fp8 with scaling, while the master weights stay in fp32.', 'Warmup, gradient clipping, spike rollback and regular checkpoints keep a months-long run alive.')), BigQuestion(question='How does a model learn from rewards instead of examples?', route=('primer.ml.reinforcement', 'primer.ml.training_stages', 'primer.ml.reasoning', 'primer.ml.alignment'), in_brief=('Reinforcement learning samples an action, scores it, and makes high-scoring actions more likely.', 'A baseline turns rewards into advantages (better or worse than usual), which cuts noise without bias.', 'PPO reuses each batch for several steps, clips how far the policy moves, and leashes it to a reference with a KL penalty.', 'GRPO drops the value network by comparing rewards within a group of answers to the same prompt.', 'A checker as the reward (the right answer, passing tests) is how reasoning models are trained.', 'The policy optimises the reward you wrote, not the goal you meant; verifiable rewards, a KL leash and held-out checks defend against that.')), BigQuestion(question='How do reasoning models think, and when is extra thinking worth it?', route=('primer.ml.inference', 'primer.ml.reinforcement', 'primer.ml.reasoning', 'primer.agents.planning', 'primer.agents.cost'), in_brief=("Every written token is another forward pass, so a chain of thought buys serial computation, and the text is the model's working memory.", 'Test-time compute can go into one longer chain, or into many chains with a vote or a verifier picking one answer.', "Voting helps only when the right answer is the most common one and the samples' mistakes are independent.", 'Checking each step catches errors that checking only the final answer misses.', 'Training with verifiable rewards makes longer, self-checking reasoning emerge.', 'Thinking is paid for per token and slips compound over long chains, so route easy tasks to little thinking and measure cost per successful task.')), BigQuestion(question='How do models handle very long contexts?', route=('primer.ml.attention', 'primer.ml.positional', 'primer.ml.inference', 'primer.ml.efficient_architectures'), in_brief=('Attention scores every pair of tokens, and the KV cache grows with every token, per layer, per conversation.', 'Sliding windows and sparse patterns score fewer pairs; stacked layers still carry information far.', 'Linear attention and state-space models keep a fixed-size summary: linear time and constant memory, but blurrier recall.', 'Mamba makes the summary selective: each token decides how much to keep and how much to write.', 'Hybrids keep a few attention layers for exact lookup.', 'The cache shrinks by sharing key/value heads, caching a small latent, or storing fewer bits.')), BigQuestion(question='What is going on inside a trained model, and how can we tell?', route=('primer.ml.transformer', 'primer.ml.embeddings.word2vec', 'primer.ml.interpretability', 'primer.ml.alignment'), in_brief=('Models store features as directions across many neurons, not one feature per neuron.', 'Probes and the logit lens read what is present; they show correlation, not use.', 'Activation patching changes one activation and watches the output: the causal test.', 'Sparse features get packed in superposition, which makes individual neurons respond to several things.', 'Sparse autoencoders unpack superposition into interpretable features, at the cost of some unexplained activity.', 'These tools give evidence, not proof; full explanations exist only for narrow behaviours.')), BigQuestion(question='When is a neural network the wrong tool?', route=('primer.ml.classical', 'primer.ml.neural_net', 'primer.ml.regularization', 'primer.ml.metrics'), in_brief=('On tables whose columns each mean something alone, gradient-boosted trees or a random forest are the model to beat.', "A tree asks one column at a time whether it's above a threshold, so it needs no feature scaling and handles categories natively.", 'A single deep tree overfits; forests average many decorrelated trees, and boosting adds small trees fit to the remaining errors.', 'Neural networks win when meaning lives in arrangements of raw values (images, audio, text), when data is huge, or when a pretrained model can be reused.', "Trees can't extrapolate beyond the values they trained on, and impurity importances credit noise, so check importances on held-out data.")), BigQuestion(question='How do AI models generate images, audio and video?', route=('primer.ml.generative.autoencoders', 'primer.ml.generative.gans', 'primer.ml.generative.diffusion', 'primer.ml.generative.multimodal'), in_brief=('A generator learns a whole distribution, so it can sample new examples; predicting the average gives blur.', 'An autoencoder squeezes data into a small code and back; a VAE shapes that code so random codes decode to new data.', 'A GAN trains a generator against a discriminator: sharp, one-pass samples, but unstable training and mode collapse.', 'Diffusion adds noise on purpose and learns to remove it, generating from pure noise in many small steps.', 'Flow matching learns straight paths from noise to data, so it needs fewer steps; guidance trades variety for following the prompt.', "Real systems denoise an autoencoder's latent with a transformer that reads the prompt; video and audio are the same idea with more tokens.")), BigQuestion(question='How do AI models see images and hear audio?', route=('primer.ml.cnn_rnn', 'primer.ml.embeddings.contrastive', 'primer.ml.generative.multimodal'), in_brief=('Every modality becomes a sequence of vectors a transformer attends over.', 'Images become patch tokens, (H/P)·(W/P) of them, so cost grows with the square of the resolution.', 'A small projector connects a vision or audio encoder to a language model; it is trained first, with both models frozen.', 'Audio becomes a log-mel spectrogram, then about 50 tokens a second.', 'Video multiplies image tokens by time, so frames are sampled; text stays the densest input.'))]
def big_questions_table() -> str: on GitHub
641def big_questions_table() -> str:
642    """The README's compact map: each big question and its route. Regenerate with `make readme`."""
643    rows = ["| Big question | Lessons that answer it, in order |", "|---|---|"]
644    for q in BIG_QUESTIONS:
645        route = ", ".join(f"[{_lesson(m).title}]({m.replace('.', '/')}.py)" for m in q.route)
646        rows.append(f"| {q.question} | {route} |")
647    return link_code_references("\n".join(rows) + "\n", '.')

The README's compact map: each big question and its route. Regenerate with make readme.

def big_questions_page() -> str: on GitHub
650def big_questions_page() -> str:
651    """docs/BIG_QUESTIONS.md: every big question with its route and the short version."""
652    out = [
653        "# Big questions: the map from the top down\n",
654        "The lessons build the field from the bottom up. Real conversations about AI systems start from the top, "
655        "with questions like these. For each one: the lessons that answer it, in order, and the short version, "
656        "the main ideas in the order that builds understanding. Read the short version first, then open the "
657        "lessons wherever you want the full story.\n",
658        "Generated from `primer/curriculum.py` by `make readme`; edit it there.\n",
659    ]
660    for n, q in enumerate(BIG_QUESTIONS, 1):
661        route = " → ".join(f"[{_lesson(m).title}](../{m.replace('.', '/')}.py)" for m in q.route)
662        in_brief = "\n".join(f"{i}. {point}" for i, point in enumerate(q.in_brief, 1))
663        out.append(f"\n## {n}. {q.question}\n\n**Route:** {route}\n\n**In brief:**\n\n{in_brief}\n")
664    return link_code_references("\n".join(out), 'docs')

docs/BIG_QUESTIONS.md: every big question with its route and the short version.

def self_test_book() -> str: on GitHub
667def self_test_book() -> str:
668    """Every lesson's self-test questions, in reading order, as one markdown page.
669
670    Generated from the lesson docstrings (the single source of truth) into
671    docs/SELF_TEST.md by `make readme`.
672    """
673    import importlib
674    import re
675
676    out = [
677        "# Self-test: every question in the primer\n",
678        "Generated from each lesson's `## Self-test questions` section by `make readme`; "
679        "edit the lesson, not this file. Answer each question out loud before reading the answer.\n",
680    ]
681    for part in PARTS:
682        out.append(f"\n## {part.title}\n")
683        for i, lesson in lessons_in(part.key):
684            try:
685                doc = importlib.import_module(lesson.module).__doc__ or ""
686            except ModuleNotFoundError:
687                continue
688            m = re.search(r"^## Self-test questions\s*\n(.*?)(?=^## |\Z)", doc, re.S | re.M)
689            if not m:
690                continue
691            body = re.sub(r"^(#+) ", lambda h: "#" * (len(h.group(1)) + 2) + " ", m.group(1).strip(), flags=re.M)
692            link = lesson.module.replace(".", "/") + ".py"
693            out.append(f"\n### {i}. {lesson.title}\n\nFrom [`{lesson.module}`](../{link}).\n\n{body}\n")
694    return link_code_references("\n".join(out), "docs")

Every lesson's self-test questions, in reading order, as one markdown page.

Generated from the lesson docstrings (the single source of truth) into docs/SELF_TEST.md by make readme.