primer.glossary

Glossary

Every term this primer uses, in plain English, with a link to the lesson that builds it. On the HTML site, these definitions also appear when you hover over (or tap) the term anywhere in a lesson or an annotated paper.

This module is the single source of truth: make docs turns it into glossary.js for the browser. The table below is generated from GLOSSARY.

Term Meaning Taught in
ablation Removing or replacing one part of a method and measuring again, to find out which part the gains come from.
acceptance rate How often the big model keeps a drafted token in speculative decoding. primer.ml.inference
accuracy The share of all predictions that were correct. Misleading when one class is rare. primer.ml.metrics
ACL Access-control list: the users or groups allowed to read a document. primer.agents.rag
activation checkpointing Keeping only each layer's input and recomputing the rest during the backward pass. primer.ml.pretraining
activation function The nonlinear function inside a neuron, such as ReLU or GELU. Without it, stacked layers collapse into one. primer.ml.neural_net
activation patching Copying one activation from a clean run into a corrupted run to measure how much of the right answer returns; also called causal tracing. primer.ml.interpretability
active learning Choosing which examples to label next by what the current model is likely to get wrong, so each paid label teaches more.
actor-critic A reinforcement learner in two parts: the actor (the policy) chooses actions, and the critic (a value network) predicts the reward to expect, which serves as the baseline. primer.ml.reinforcement
Adaboost A boosting method (Freund and Schapire, 1996) that trains each new classifier on reweighted data, raising the weight of the examples the previous ones got wrong, and lets the classifiers vote with weights set by their accuracy.
AdaGrad An optimizer that divides each weight's step by the square root of the sum of all its past squared gradients, so steps only ever shrink. primer.ml.optimizers
Adam An optimizer that adapts the step size for each weight using running averages of its gradients. primer.ml.optimizers
AdamW Adam with weight decay applied separately from the gradient step. The default optimizer for transformers. primer.ml.optimizers
adapter A small trainable module added beside frozen weights to specialise a model. primer.ml.training_stages
adaptive layer norm Layer normalization whose scale and shift are computed from a conditioning signal, such as the noise step and class label, instead of being fixed learned weights. Diffusion transformers use it to tell every block what to make (adaLN). primer.ml.deep_nets
advantage How much better an action did than typical: reward minus baseline. primer.ml.reinforcement
adversarial example An input changed by a small, deliberately chosen amount so that a model gets it wrong, often in a way a person would never notice.
agent A system where a language model decides which tools to call, looks at the results and decides the next step, in a loop until done. primer.agents.agent_loop
agent loop The cycle an agent repeats: think, call a tool, observe the result, decide what next. primer.agents.agent_loop
agent-computer interface The commands an agent can call and the exact form of what comes back to it, designed for a language model the way a code editor is designed for a person. primer.agents.coding_agents
agentic RAG RAG where the model decides whether, what and how often to search. primer.agents.rag
ALiBi A position scheme that adds no position vectors and instead subtracts a penalty proportional to distance from each attention score. primer.ml.positional
alignment Making a model's behaviour match what we want (helpful, honest, harmless), not just the measurements we optimize. primer.ml.alignment
alignment tax Capability lost as a side effect of training a model to be helpful, honest and harmless. primer.ml.training_stages
all-gather Every GPU contributes its piece and every GPU ends up with all the pieces, side by side; the second half of a ring all-reduce. primer.ml.pretraining
all-reduce Summing a value across GPUs so every GPU ends up holding the total. primer.ml.hardware
anisotropy When a model's vectors all crowd into a narrow cone, so even unrelated texts score as fairly similar. primer.ml.embeddings.similarity
ANN Approximate nearest neighbor search: trade a little recall for a lot of speed. primer.ml.embeddings.ann
anomaly alert An alert when a metric jumps far outside its usual range. primer.agents.cost
anomaly detection Flagging items far from every known group. primer.ml.embeddings.clustering
approximate nearest neighbor Finding vectors close to a query quickly by searching only part of the collection, accepting a small chance of missing the true closest. primer.ml.embeddings.ann
arena Ratings built from people voting between two anonymous answers. primer.ml.benchmarks
argmax The position of the largest value in a list, rather than the value itself. primer.notation
arithmetic intensity Operations performed per byte read from memory. primer.ml.inference
asymmetric distance computation Scoring compressed vectors against an uncompressed query by adding up precomputed table lookups. primer.ml.embeddings.ann
attack success rate The share of red-team attempts that get past a safety check. primer.ml.alignment
attention The mechanism that lets each token look at every other token, score how relevant each one is, and blend in information from the relevant ones. primer.ml.attention
attention sink Early tokens that trained models park spare attention on; dropping them breaks streaming generation. primer.ml.efficient_architectures
audit log A record of which actor took which action, for whom, with what inputs. primer.agents.deployment
audit trail A tamper-evident log of which agent did what, for whom and with what authority. primer.agents.deployment
auto-regressive Generating one token at a time, each predicted from everything before it and then fed back in. primer.ml.big_picture
autoencoder A network that squeezes its input into a small code and rebuilds the input from it, trained only to make the rebuild match. primer.ml.generative.autoencoders
automated interpretability Using a language model to explain a feature or neuron from examples of when it fires, then scoring the explanation by how well the model predicts new activations from it alone.
autoregressive generation Producing text one token at a time, where each new token is predicted from all the tokens before it and then appended. primer.ml.big_picture
backfill Re-processing existing records into a new system, such as re-embedding a corpus with a new model. primer.ml.embeddings.operations
backpropagation The chain rule run backwards through the network to find, for every weight, how much nudging it would change the loss. primer.ml.neural_net
backpropagation through time Backpropagation applied to an RNN unrolled across its time steps. primer.ml.cnn_rnn
bag of words Representing text by counting how many times each word appears, ignoring order. primer.ml.embeddings.contrastive
bagging Bootstrap aggregating: train one model per bootstrap sample and average them to cut variance. primer.ml.classical
base model A model after pretraining only: knowledgeable, but it continues text rather than following instructions. primer.ml.training_stages
baseline A typical reward subtracted before updating; it reduces noise without changing the average gradient. primer.ml.reinforcement
batch The group of examples processed before one weight update. primer.ml.neural_net
batch API A provider interface that processes many requests asynchronously at a discount. primer.agents.cost
batch normalization Rescaling each feature using statistics from the current batch of examples. Common in image models. primer.ml.deep_nets
beam search Decoding that keeps the few most probable partial outputs at each step instead of just one, then picks the best finished one. primer.ml.inference
benchmark Fixed questions plus a scoring rule, averaged into one comparable number. primer.ml.benchmarks
benchmark contamination When test questions appeared in a model's training data, inflating its scores. primer.ml.regularization
Bernoulli distribution The distribution of a single yes-or-no outcome, set by one number: the chance of yes. A VAE decoder for black-and-white pixels outputs one per pixel. primer.ml.losses
BERT An encoder-only transformer trained to fill in hidden words using the text on both sides of them; the ancestor of many embedding and classification models. primer.ml.transformer
BERTScore Compares generated and reference text by embedding similarity instead of exact words. primer.ml.metrics
best-of-n Sample n candidates and keep the one a verifier scores highest. primer.ml.reasoning
bf16 A 16-bit float with fp32's 8 exponent bits and 7 mantissa bits: the same range, less precision. primer.ml.hardware
bi-encoder Embeds the query and each document separately, so document vectors can be computed once and searched fast. primer.ml.embeddings.retrieval
bias A learned number added to a neuron's weighted sum, letting it shift its output up or down. primer.ml.neural_net
bias correction Adam's rescaling of its running averages early in training, since they start at zero and would otherwise be too small. primer.ml.optimizers
bias-variance trade-off Splitting a model's error into systematic error from being too simple (bias) and sensitivity to the training sample from being too flexible (variance). primer.ml.regularization
big-O A way to say how cost grows with input size, ignoring constant factors. O(n²) means doubling n quadruples the cost. primer.notation
binary cross-entropy Cross-entropy for yes/no predictions: −[y·ln p + (1 − y)·ln(1 − p)]. primer.ml.losses
binary quantization Storing only the sign of each number, as a single bit. primer.ml.embeddings.compression
binomial coefficient "n choose k", written C(n, k): the number of different groups of k items that can be picked from n, ignoring order. C(10, 5) = 252. primer.ml.benchmarks
binomial distribution The chances of getting each possible number of successes in n independent tries that each succeed with the same probability p. primer.ml.benchmarks
bits per dimension A model's negative log-likelihood per number in the data, in bits: how long a code the model needs, on average, for each pixel value. Lower is better. primer.ml.losses
blast radius How much damage a failure can do before it is noticed and stopped. primer.agents.deployment
BLEU A score that counts overlapping word sequences with a reference text. Cheap, but blind to paraphrase. primer.ml.metrics
block table A per-request map from logical KV-cache blocks to physical blocks in GPU memory. primer.ml.inference
blue/green deployment Building the new version beside the old one, switching traffic once it proves itself, and keeping the old one for rollback. primer.ml.embeddings.operations
BM25 The classic keyword-search score: rewards documents containing the query's rarer words, with diminishing returns for repeats. primer.ml.embeddings.retrieval
boosting Building a strong model from many weak ones trained one after another, each concentrating on what the ones before it still get wrong. primer.ml.classical
bootstrap Measuring uncertainty by rescoring many with-replacement resamples of your own data. primer.ml.benchmarks
bootstrap sample A resample of the training rows drawn with replacement, so some rows repeat and about 37% are left out. primer.ml.classical
bottleneck The narrow middle of an autoencoder: too small to copy through, it forces the network to keep only what matters. primer.ml.generative.autoencoders
bounding box A rectangle, given by its corner coordinates, that marks where one object sits in an image.
BPE Byte pair encoding: build a vocabulary by repeatedly merging the most frequent adjacent pair of symbols. primer.ml.tokenization
Bradley-Terry model The rule that the probability A beats B is the sigmoid of their score difference. primer.ml.training_stages
byte pair encoding A tokenizer-building method that starts from characters and repeatedly merges the most frequent adjacent pair into a new token. primer.ml.tokenization
byte-level BPE Byte pair encoding run on the raw bytes of UTF-8 text, so any string can be tokenized and there is never an unknown token. primer.ml.tokenization
cache miss When the processor needs a number that is not in its small, fast cache and must wait for the much slower main memory, which can cost as much time as hundreds of arithmetic operations. primer.ml.hardware
calibration How well a model's confidence matches how often it is right: a calibrated model is right about 80% of the time when it says 80%. primer.ml.losses
canary release Sending a small share of traffic to a new version first and watching its metrics before rolling out further. primer.agents.deployment
canary string A unique marker in benchmark files so trainers can filter out copies. primer.ml.benchmarks
capability negotiation The start-of-session exchange where client and server announce which optional features they support. primer.agents.mcp
catastrophic forgetting When training on a new task alone erodes or erases skills a model already had. primer.ml.fine_tuning
causal mask Hides future tokens from each token during attention, so a model predicting the next word can't peek at it. primer.ml.attention
causal tracing Activation patching used to find where a model recalls a fact: blur the subject in the input, then restore one clean hidden state at a time and see which ones bring the right answer back. primer.ml.interpretability
CBOW Continuous bag of words: the word2vec model that averages the surrounding words' vectors to predict the middle word. primer.ml.embeddings.word2vec
cell state An LSTM's notebook, edited by adding rather than overwriting, so information survives many steps. primer.ml.cnn_rnn
centroid The average position of a group of vectors, used as the group's representative point. primer.ml.embeddings.ann
chain of thought Intermediate reasoning steps a model writes before its final answer, each one readable by the next forward pass. primer.ml.reasoning
chain rule To get the slope through a chain of functions, multiply the slopes of each link. primer.notation
chat template The exact role markers and layout a model family uses to turn messages into training text. primer.ml.fine_tuning
checkpoint Saved progress after a completed step, so a failure resumes from there instead of from the start. primer.agents.planning
checkpoint averaging Averaging the weights saved at the last few points of training, which usually gives a slightly better model. primer.ml.training_stages
chinchilla scaling The compute-optimal rule of about 20 training tokens per model parameter. primer.ml.pretraining
chunking Splitting documents into passages before embedding them, so retrieval can return just the relevant part. primer.ml.embeddings.retrieval
CIDEr A score for a generated caption: how well its words and short phrases match several human-written captions of the same image, with phrases common to captions of many images counting less. Higher is better; values above 100 are normal.
circuit A chain of components that together compute one behaviour. primer.ml.interpretability
circuit breaker A wrapper that stops calling a failing dependency for a cool-down period, failing fast instead. primer.agents.failures
classifier guidance Steering a diffusion model toward a label by adding the slope of a separate classifier, trained on noisy inputs, to the model's noise guess at every step. primer.ml.generative.diffusion
classifier-free guidance Mixing a model's guesses with and without the prompt, and pushing past the prompted one to follow it more closely. primer.ml.generative.diffusion
CLIP A model that trains an image encoder and a text encoder together so pictures and their captions land near each other in one vector space. primer.ml.embeddings.contrastive
clustering Grouping items so similar ones end up together, without being told the groups. primer.ml.embeddings.clustering
CNN Convolutional neural network: a network built from convolutions, small learned filters slid across an image to find local patterns wherever they appear. primer.ml.cnn_rnn
co-adaptation When neurons learn to depend on each other's exact behaviour, each correcting the others' quirks; it fits the training data but breaks on new data. primer.ml.regularization
code grader An automatic check written in code, such as exact match, schema or database state, that scores an agent's output. primer.agents.evals
codebook A small catalogue of representative vectors; each vector (or piece of one) is stored as the number of its nearest entry, as in product quantization or tokenizers for images and audio. primer.ml.embeddings.ann
Cohen's kappa Agreement between two raters after subtracting the agreement expected by chance. primer.ml.metrics
ColBERT A retrieval model that keeps one vector per token and matches each query token to its best document token. primer.ml.embeddings.retrieval
collector In OpenTelemetry, the service that receives spans from apps and forwards them to storage backends. primer.agents.observability
commitment loss A training term that pulls an encoder's output towards the codebook entry it snapped to, so the encoder commits to its entries instead of drifting away from them.
Common Crawl A nonprofit's public archive of the web, released as regular snapshots of billions of pages; the raw material of most pretraining data. primer.ml.pretraining
compounding error Small per-step failure rates multiplying over many steps: ten steps at 95% succeed only about 60% of the time. primer.agents.planning
compressed sensing Recovering a long vector from far fewer measurements of it, which is possible when the vector is known to be sparse (mostly zeros).
compute-bound Limited by arithmetic throughput, as in prefill. primer.ml.inference
computer use An agent operating a graphical interface from screenshots, sending clicks and keystrokes. primer.agents.coding_agents
confidence interval A range built so that 95 in 100 such ranges contain the true value. primer.ml.benchmarks
confused deputy A program with broad authority tricked into using it for someone who lacks that authority. primer.agents.mcp
confusion matrix The four counts behind every classification metric: hits, false alarms, misses and correct passes. primer.ml.metrics
Constitutional AI Training against a written list of principles: the model critiques and revises its answers, and AI-labelled preferences train the reward model. primer.ml.alignment
constrained decoding Before each token is drawn, setting every token that can't lead to a valid answer to probability zero. primer.ml.structured_output
container An isolated copy of an operating system's user space for one program: its own files, processes and network settings, started from an image and thrown away afterwards. primer.agents.coding_agents
contamination Test questions and answers leaking into a model's training data. primer.ml.benchmarks
content-addressed Named by a hash of its own content, so identical content always gets the identical name. primer.agents.deployment
context engineering Deciding exactly what goes into the model's context window on each call: instructions, facts, tool results, memory. primer.agents.context
context rot Quality dropping as a long session fills the context with stale or irrelevant material. primer.agents.context
context window The maximum number of tokens a model can take in at once. primer.ml.inference
contextual embedding A vector computed for a word within its sentence, so the same word gets different vectors in different contexts. primer.ml.embeddings.word2vec
contextual retrieval Indexing each chunk together with a line saying where in its document it comes from. primer.agents.rag
continuous batching Refilling a finished request's slot in the batch immediately with the next request. primer.ml.inference
continuous integration The automated checks that run on every proposed change before it can merge. primer.agents.evals
contract test An automated check that an external API still accepts and returns what your code expects. primer.agents.failures
contrastive learning Training that pulls matching pairs together and pushes non-matching pairs apart. primer.ml.embeddings.contrastive
contrastive loss A loss that pulls matching pairs together in vector space and pushes non-matching pairs apart. primer.ml.losses
control task A probe trained on random labels, showing how much a probe fits with nothing real to find. primer.ml.interpretability
convex Bowl-shaped everywhere with a single bottom, so any downhill path reaches the same minimum; neural network losses are not convex. primer.notation
convolution Sliding a small grid of learned weights across an image and computing a weighted sum at every position, to find where a pattern appears. primer.ml.cnn_rnn
cooldown A quiet period after an action during which the same action is not repeated.
copy-on-write Sharing one copy of data and making a private copy only when someone needs to change it. primer.ml.inference
cosine decay Lowering the learning rate from its peak to a floor along half a cosine wave. primer.ml.optimizers
cosine similarity How closely two vectors point in the same direction, from −1 (opposite) to 1 (identical), ignoring their lengths. primer.ml.embeddings.similarity
cost per resolved task Everything spent on all attempts, divided by the tasks actually resolved. primer.agents.coding_agents
cost per successful task Total spend divided by the number of tasks that succeeded, counting retries and cleanup. primer.agents.cost
coupling A way of pairing draws from two distributions so that each side, on its own, still has its own distribution. Pairing noise with data at random is one coupling; sending each noise point to one definite data point is another. primer.ml.generative.diffusion
covariance How two quantities vary together: positive when they rise together, negative when one rises as the other falls. For vectors, a matrix holding it for every pair of entries.
credit assignment Working out which of many earlier actions deserves the blame or credit for how an episode ended. primer.agents.planning
critic A Wasserstein GAN's discriminator, which outputs an unbounded score instead of a probability. primer.ml.generative.gans
cross-attention Attention whose queries come from one sequence (image patches) and whose keys and values come from another (text tokens). primer.ml.generative.diffusion
cross-encoder Reads the query and one document together and outputs a relevance score. Accurate but slow, so it is used to rerank a shortlist. primer.ml.embeddings.retrieval
cross-entropy The standard loss for predicting categories: minus the log of the probability given to the right answer. primer.ml.losses
cross-validation Rotating which slice of the data is held out, training once per slice, and averaging the scores. primer.ml.regularization
curse of dimensionality In very high dimensions, random points are all nearly the same distance apart, which makes exact nearest-neighbor search slow and fragile. primer.ml.embeddings.similarity
dark knowledge How a teacher model ranks the wrong answers, revealed by its soft targets. primer.ml.training_stages
data leakage When information from the test set, or from the future, sneaks into training and inflates offline scores. primer.ml.regularization
data mixture The share of training tokens drawn from each data source. primer.ml.pretraining
data parallelism Every GPU holds the whole model and trains on a different slice of the batch. primer.ml.pretraining
DBSCAN Density-based clustering that grows clusters from points with enough close neighbours and labels the rest as noise. primer.ml.embeddings.clustering
DCGAN Deep convolutional GAN: a GAN whose generator and discriminator are convolutional networks, the usual baseline design for image GANs. primer.ml.generative.gans
DDIM A deterministic diffusion sampler that predicts the clean result and jumps to a much less noisy step, so it needs far fewer steps. primer.ml.generative.diffusion
DDPM The original diffusion sampler: remove the guessed noise in many small steps, adding a small fresh wobble each time. primer.ml.generative.diffusion
dead latent A sparse autoencoder latent that has stopped firing on any input, wasting its slot in the dictionary; resampling it onto badly rebuilt inputs brings it back. primer.ml.interpretability
debounce Waiting until a signal has settled before acting on it, so a burst of changes triggers one action.
decision tree A flowchart of yes/no questions about one feature at a time, learned from examples, ending in leaves that give the answer. primer.ml.classical
decode The second phase of generation: output tokens are produced one at a time. It sets the tokens per second. primer.ml.inference
decoder A transformer that generates text left to right, each token seeing only the ones before it. GPT, Claude and Llama are decoders. primer.ml.transformer
decomposition Splitting a big goal into small subtasks, each with a checkable definition of done. primer.agents.planning
deduplication Removing repeated copies of documents from training data. primer.ml.pretraining
degradation problem Deeper plain networks reaching higher error even on their training data: an optimization failure, not overfitting. primer.ml.deep_nets
denoiser The network in a diffusion model that looks at a noisy input and its step, and guesses the noise in it. primer.ml.generative.diffusion
denoising score matching Learning the score by training a network to guess the noise added to real examples. primer.ml.generative.diffusion
dense retrieval Search by comparing embedding vectors, so matches are by meaning rather than exact words. primer.agents.rag
derivative The slope of a function at a point: how much the output changes per tiny nudge of the input. primer.notation
dictionary learning Finding a set of directions, usually more than there are dimensions, such that every data point is a sparse combination of a few of them; also called sparse coding. A sparse autoencoder is one way to do it. primer.ml.interpretability
diff A text listing, file by file, which lines to remove (marked −) and which to add (marked +), with a few unchanged lines around each change so a tool can find the spot. Saved to a file, it is a patch that can be applied to another copy of the code. primer.agents.coding_agents
diffusion model A generator that learns to turn random noise into data by removing a little noise at a time. primer.ml.generative.diffusion
diffusion transformer A transformer used as the denoiser, reading patches of the latent as tokens (DiT). primer.ml.generative.diffusion
dimension One of the numbers in a vector: a 768-dimension embedding is a list of 768 numbers. primer.ml.embeddings.similarity
discretization Turning a continuous rate of change into one step's keep and write factors. primer.ml.efficient_architectures
discriminator The network in a GAN that outputs the probability that a sample is real rather than generated. primer.ml.generative.gans
distillation Training a small model to imitate a large one's outputs, often the biggest cost saving in production. primer.ml.training_stages
diversity heuristic HNSW's rule of keeping a neighbour only if it is closer to the new node than to any neighbour already kept, so links fan out in different directions. primer.ml.embeddings.ann
domain mismatch A general model underperforming on specialised language it never saw, such as company jargon. primer.ml.embeddings.operations
dot product Multiply two lists position by position and add up the products. It is large when two vectors point the same way. primer.notation
double descent The observation that past a point, even larger models generalize better again. primer.ml.regularization
double quantization Compressing the per-block scale factors of a quantized model as well, saving about 0.37 bits per parameter. primer.ml.inference
DPO Direct preference optimization: tuning a model straight from pairs of preferred and rejected answers, without a separate reward model. primer.ml.training_stages
draft model The small, fast model that proposes tokens in speculative decoding. primer.ml.inference
drift A metric creeping away from its usual level over time or after a deploy. primer.agents.evals
dropout Randomly switching off neurons during training so the network can't rely on any single path. primer.ml.regularization
dry run Running every check for an action and describing it without actually doing it. primer.agents.tools
dual write Writing every new record to both the old and the new system during a migration. primer.ml.embeddings.operations
durable execution Running a workflow so its progress survives crashes, by storing each completed step's result. primer.agents.orchestration
dying relu A ReLU neuron whose input stays negative, so it outputs zero and receives zero gradient forever. primer.ml.neural_net
dynamic tool loading Sending the model only the few tool definitions most relevant to the current request. primer.agents.tools
early stopping Stopping training when the score on held-out data stops improving. primer.ml.regularization
earth mover's distance The least total work to reshape one pile of probability into another, each bit of mass times the distance it travels; the same as the Wasserstein distance. primer.ml.generative.gans
edit distance The fewest single-item insertions, deletions and substitutions that turn one sequence into another.
edit-run-test loop An agent loop that changes code, runs the tests, and repeats until they pass or a budget runs out. primer.agents.coding_agents
efConstruction HNSW's beam width while building the graph; higher means a better graph and a slower build. primer.ml.embeddings.ann
efSearch HNSW's query-time beam width; raising it trades speed for recall. primer.ml.embeddings.ann
elbow method Choosing the number of clusters where adding more stops reducing inertia much. primer.ml.embeddings.clustering
Elo rating A rating scale where a 400-point gap means ten-to-one odds of winning. primer.ml.benchmarks
embedding A learned vector for a piece of content, arranged so that similar meanings end up close together. primer.ml.embeddings.word2vec
emergent ability A skill that is near chance in smaller models and appears sharply once a model is large enough, so it can't be predicted by extending the small models' trend.
encoder A transformer that reads the whole input at once, with every token seeing every other. Used for embeddings and classification. primer.ml.transformer
encoder-decoder A transformer where an encoder reads the whole input and a decoder writes the output one token at a time while attending to the encoder. primer.ml.transformer
encoder-decoder attention Attention where the decoder's queries look at the encoder's keys and values, so each output word can consult the whole input. primer.ml.transformer
ensemble A model that combines the predictions of many models, such as trees. primer.ml.classical
entailment Whether one text logically follows from another, checked by a model trained for it. primer.agents.guardrails
entropy The average surprise of a distribution: how many yes/no questions, on average, it takes to learn an outcome; 0 means certain. primer.ml.classical
episodic memory Memory of what happened, such as 'last week the user rejected this vendor'. primer.agents.memory
epoch One full pass over the training data. primer.ml.neural_net
erp Enterprise resource planning: a company's finance and operations system of record. primer.agents.deployment
escaping Replacing characters like < and > so data can't pose as markup or instructions. primer.agents.context
Euclidean distance The straight-line distance between two vectors. primer.ml.embeddings.similarity
Euler step Moving in a straight line for a short time at the current velocity, then looking again. primer.ml.generative.diffusion
Euler-Maruyama step The Euler step for a stochastic differential equation: move by the drift times the time step, then add a fresh random nudge whose size grows with the square root of the time step. primer.ml.generative.diffusion
eval A repeatable test of a model or agent's behavior on a fixed set of tasks with known good outcomes. primer.agents.evals
evaluator-optimizer A generator drafts and an evaluator critiques, repeated until the draft passes or the rounds run out. primer.agents.orchestration
evidence lower bound A quantity that never exceeds the log-probability a model gives the data; the VAE loss is its negative. primer.ml.generative.autoencoders
exemplar One worked example in a few-shot prompt: a question with its answer, and sometimes the steps that reach it. primer.agents.context
expectation The average of a quantity over many random draws, written E. primer.ml.generative.gans
expected value The average you would get over many tries. primer.notation
exploding gradient When gradients grow huge as they pass back through many layers, so training diverges. primer.ml.deep_nets
exploration Trying actions you are unsure of instead of repeating the best one so far. primer.ml.reinforcement
exponent The bits of a float that pick the power of two, which sets its range. primer.ml.hardware
exponential backoff Waiting longer after each failed attempt, doubling up to a cap. primer.agents.failures
exponential moving average A running average that keeps most of its old value and mixes in a small share of each new one, so recent values count most. Adam keeps its moments this way, and diffusion models sample with such an average of their weights. primer.ml.optimizers
external fragmentation Free memory split into gaps too small to use. primer.ml.inference
external verification Checking a model's output with something outside the model, such as tests, schemas or database queries. primer.agents.planning
F1 The harmonic mean of precision and recall: high only when both are high. primer.ml.metrics
fail-to-pass test A hidden test that fails before a fix and must pass after it. primer.agents.coding_agents
faithfulness The share of an answer's claims that are supported by the retrieved sources. primer.agents.evals
false positive Something flagged as positive that is actually negative, such as a solution graded correct because its final answer matches although its reasoning is wrong. primer.ml.metrics
false-positive rate The share of real negatives the model wrongly flags. primer.ml.metrics
fan-in The number of inputs a neuron adds up. primer.ml.deep_nets
fault localization Finding which files, functions and lines cause a bug, before trying to fix it. primer.agents.coding_agents
feature A property a model tracks, stored as a direction across many neurons. primer.ml.interpretability
feature engineering Hand-transforming inputs (logarithms, one-hot columns) so a model can use them. primer.ml.classical
feature flag A configuration switch that turns a capability on or off without redeploying. primer.agents.deployment
feature importance A score for how much a model relies on each input column. primer.ml.classical
feature map In linear attention, the function φ applied to queries and keys so that φ(q)·φ(k) replaces e^(q·k). primer.ml.efficient_architectures
feature splitting One feature in a small sparse autoencoder becoming several narrower features in a larger one, such as one base64 feature becoming separate features for letters, digits and encoded text. primer.ml.interpretability
feed-forward network The small two-layer network in each transformer block that processes every token on its own after attention has mixed them. primer.ml.transformer
few-shot prompt A prompt that includes a few worked examples of the task before the real question. primer.agents.context
FID Fréchet Inception Distance: compares a set of generated images with real ones through the statistics of their features in a pretrained image network. Lower is better; 0 means indistinguishable. primer.ml.generative.gans
filter In a CNN, the small grid of learned weights a convolution slides over an image; also called a kernel. primer.ml.cnn_rnn
fine-tuning Training an existing model further on a smaller, targeted dataset to change its behavior. primer.ml.training_stages
finite-state machine A fixed set of states with one move per input character; it can check patterns but not unlimited nesting. primer.ml.structured_output
flat index Exact search that compares the query with every stored vector; the ground truth other indexes are measured against. primer.ml.embeddings.ann
flip rate The share of questions whose answer changes when the user asserts a wrong answer. primer.ml.alignment
floating point Storing a number as a sign, an exponent (range) and a mantissa (precision). primer.ml.hardware
FLOPs Floating-point operations: individual multiplies or adds, used to measure compute (about 2 per parameter per token to generate, 6 to train). primer.ml.transformer
flow matching Training a network to output the velocity along paths from noise to data, then following it to generate. primer.ml.generative.diffusion
forget gate The LSTM dial that decides what to erase from the cell state. primer.ml.cnn_rnn
forward pass Running inputs through the network, layer by layer, to get a prediction. primer.ml.neural_net
forward process The fixed, unlearned procedure that mixes data with noise step by step until only noise is left. primer.ml.generative.diffusion
Fourier transform Splitting a signal into how much of each frequency it contains. primer.ml.generative.multimodal
fp16 A 16-bit float with 5 exponent and 10 mantissa bits: more precision, range only up to 65,504. primer.ml.hardware
fp8 8-bit floats: E4M3 (range to 448) for weights and activations, E5M2 (to 57,344) for gradients. primer.ml.hardware
frame sampling Keeping only some video frames, such as one a second, to save tokens. primer.ml.generative.multimodal
freezing Keeping some of a model's weights fixed during training, so only the rest learn and the frozen part keeps what it already knew. primer.ml.fine_tuning
Frobenius norm The size of a whole matrix: square every entry, add them up, take the square root. primer.notation
FSDP Fully sharded data parallel: all training state sharded across GPUs (ZeRO stage 3). primer.ml.pretraining
full fine-tune Updating every weight of a model. Rarely worth the cost. primer.ml.training_stages
functional correctness Judging generated code by running it against tests rather than by comparing its text with a reference solution. primer.ml.benchmarks
fused multiply-add One instruction that multiplies two numbers and adds the result to a running total. primer.ml.hardware
GAN A generator and a discriminator trained against each other, so the generator learns to make samples the discriminator can't tell from real data. primer.ml.generative.gans
gated cross-attention Cross-attention layers added inside a frozen language model so its text can look at image features, with each layer's output multiplied by tanh of a learned number that starts at 0, so the model starts out unchanged. primer.ml.generative.multimodal
Gaussian The bell-curve distribution, set by its mean (where the centre is) and its standard deviation (how wide it is). primer.ml.generative.autoencoders
GELU A smooth version of ReLU used in most transformers. primer.ml.neural_net
GenAI semantic conventions OpenTelemetry's standard attribute names for model and agent spans, such as gen_ai.request.model. primer.agents.observability
generalization gap Validation loss minus training loss; when it keeps growing, the model is memorising. primer.ml.fine_tuning
generation failure A wrong answer produced even though a relevant document was retrieved. primer.ml.embeddings.operations
generative model A model that learns to produce new samples resembling its training data, such as new faces, voices or sentences. primer.ml.generative.autoencoders
generator The network in a GAN that turns random noise into a sample. primer.ml.generative.gans
Gini impurity The chance that two examples drawn at random from a pile carry different labels; 0 means pure. primer.ml.classical
global token A token every other token may attend to, and which attends to all, so any two tokens are two hops apart. primer.ml.efficient_architectures
GloVe A counting method that fits word vectors so their dot products predict how often words appear together. primer.ml.embeddings.word2vec
golden set A curated set of real tasks with expected outcomes, run on every change to catch regressions. primer.agents.evals
Goodhart's law Once a measure becomes a target, optimizing it stops improving the thing it measured. primer.ml.alignment
GPU A graphics processor: a chip with thousands of small cores that do many multiply-adds at once, which is exactly what neural networks need. primer.notation
gradient One slope per input, collected into a vector. It points in the direction that increases the function fastest. primer.notation
gradient ascent Nudging every weight along its gradient to make a quantity larger: the mirror image of gradient descent. Used to maximise a reward, or, run on a loss, to make a model worse at a task on purpose. primer.ml.reinforcement
gradient boosting Adding small trees one at a time, each fit to the current errors, with each step shrunk by a learning rate. primer.ml.classical
gradient clipping Capping the size of the gradient so one bad batch can't throw the weights far off course. primer.ml.deep_nets
gradient descent Training by repeatedly taking a small step downhill: nudge every weight against its gradient to lower the loss. primer.ml.optimizers
gradient penalty A discriminator loss term that punishes steep slopes of its output with respect to its input (R1, WGAN-GP). primer.ml.generative.gans
graduated autonomy Granting an agent more independence in steps (shadow, then approval, then autonomy), each earned with evidence. primer.agents.deployment
grammar A set of rules defining which strings are valid in a format. primer.ml.structured_output
graph A set of points (nodes) joined by links (edges). primer.ml.embeddings.ann
GraphRAG Retrieval over a graph of entities and relationships extracted from documents. primer.agents.rag
greedy decoding Always picking the single most probable next token, the same as sampling at temperature 0. primer.ml.big_picture
greedy search Always move to whichever neighbour is closest to the target, and stop when none is closer. primer.ml.embeddings.ann
grounding Tying an answer to specific source passages so every claim can be checked. primer.agents.rag
grouped-query attention Several query heads share one set of keys and values, shrinking the memory needed during generation. primer.ml.attention
GRPO Group Relative Policy Optimization: each answer is scored against other answers to the same prompt, so no value network is needed. primer.ml.reinforcement
GRU A simpler gated RNN with an update gate and a reset gate and no separate cell state. primer.ml.cnn_rnn
GSM8K A benchmark of grade-school maths word problems (1,319 in its test set), scored by comparing the final number with the answer key. primer.ml.benchmarks
guardrail A check around the model that screens inputs, validates outputs or limits actions. primer.agents.guardrails
guidance scale The weight w in classifier-free guidance: 0 ignores the prompt, 1 follows it plainly, above 1 exaggerates it. primer.ml.generative.diffusion
hallucination A confident answer that isn't supported by the sources or the facts. primer.agents.rag
hamming distance The number of positions where two bit strings differ. primer.ml.embeddings.compression
hand-off The message that passes work and its context from one agent to another; anything not written into it is lost. primer.agents.orchestration
hann window A smooth rise and fall applied to each slice so its cut edges don't add false frequencies. primer.ml.generative.multimodal
hard negative A training example that looks relevant but isn't, such as the right topic with the wrong answer. Training on them teaches fine distinctions. primer.ml.embeddings.contrastive
harmful compliance Answering a request that should have been refused. primer.ml.alignment
harmonic mean n divided by the sum of the reciprocals of n numbers; dominated by the smallest, so it is high only when every number is high. primer.ml.metrics
hash chain Records that each store the previous record's hash, so any edit or deletion is detectable. primer.agents.deployment
HBM The GPU's large main memory: tens of GB, about 10× slower than on-chip SRAM. primer.ml.inference
HDBSCAN A version of DBSCAN that considers every reach at once and keeps the most persistent clusters, so no reach setting is needed. primer.ml.embeddings.clustering
He initialization Starting weights with variance 2 / fan-in, making up for ReLU zeroing half its inputs. primer.ml.deep_nets
hidden state The running summary an RNN carries from one step to the next. primer.ml.cnn_rnn
hidden tests Grading tests the agent never sees, so it can't pass by satisfying the tests instead of the intent. primer.agents.coding_agents
hierarchical softmax Replacing one softmax over the whole vocabulary with about log₂V yes-or-no decisions down a binary tree. primer.ml.embeddings.word2vec
hit@k The share of queries with at least one relevant result in the top k. primer.ml.embeddings.operations
HNSW A layered graph index for vector search: long jumps on sparse upper layers, then a careful local search at the bottom. primer.ml.embeddings.ann
host memory The CPU's main memory, reached from the GPU over a much slower link. primer.ml.hardware
Huber loss A loss that is squared for small errors and grows only in proportion to the error beyond a threshold δ, so a few wild outliers cannot dominate the fit. primer.ml.losses
Huffman tree A binary tree that gives frequent items short codes and rare items long ones. primer.ml.embeddings.word2vec
human approval gate A rule that parks irreversible or high-value actions until a person approves them. primer.agents.tools
hybrid model A stack that mixes a few attention layers with many state-space layers. primer.ml.efficient_architectures
hybrid search Running keyword search and vector search together and merging the two ranked lists. primer.ml.embeddings.retrieval
HyDE Hypothetical document embeddings: have the model write a plausible answer, then search with that, since answers resemble documents more than questions do. primer.agents.rag
hyperparameter A setting chosen before training rather than learned from data, such as the learning rate, the model size or the number of epochs. primer.ml.optimizers
hysteresis Making the bar for changing state higher than the bar for staying, so a decision doesn't flicker between two close options.
idempotency key A unique id sent with a write so a retried request returns the first result instead of acting twice. primer.agents.tools
idempotent Safe to repeat: doing it twice has the same effect as doing it once, so retries can't create duplicates. primer.agents.tools
IDF Inverse document frequency: a word's rarity weight, higher for words found in fewer documents. primer.ml.embeddings.retrieval
ImageNet A benchmark of about 1.3 million photos in 1,000 categories, the standard test of image classifiers; ImageNet-21k is its 14-million-photo, 21,000-category superset. primer.ml.cnn_rnn
implicit generative model A model that makes samples by pushing random noise through a fixed procedure, without giving the probability of any sample: a GAN generator, or a diffusion model sampled with DDIM. primer.ml.generative.gans
implicit reward In DPO, β times the log of how much more likely training has made a response than the reference model did. primer.ml.training_stages
importance sampling Estimating an average under one distribution from samples drawn under another, by weighting each sample by the ratio of its two probabilities. primer.ml.reinforcement
in-batch negatives Using the other examples in a training batch as free wrong answers for each query. primer.ml.embeddings.contrastive
in-context learning Picking up a task from instructions or examples in the prompt, with no change to the model's weights. primer.ml.big_picture
Inception score A score for generated images from a pretrained image classifier: high when each image is confidently one class and the images spread over many classes. Higher is better; it never looks at real images. primer.ml.generative.gans
index alias A pointer name like "live" that search uses, so switching indexes means repointing one name. primer.ml.embeddings.operations
induction head A pattern-completion mechanism: having seen A followed by B earlier in the text, predict B the next time A appears. It is thought to underlie much of in-context learning.
inductive bias The assumptions built into a model before it sees any data, such as a convolution's belief that nearby pixels matter most. Good assumptions help with little data; with enough data a model can learn them instead. primer.ml.cnn_rnn
inertia The total squared distance from each point to its cluster's center: the quantity k-means minimizes. primer.ml.embeddings.clustering
inference Using a trained model to make predictions, as opposed to training it. primer.ml.inference
InfoNCE A contrastive loss that treats the other items in the batch as wrong answers for each matching pair. primer.ml.losses
information gain How much a split lowers entropy; the tree picks the split that lowers it most. primer.ml.classical
initialization Choosing the size of a network's random starting weights so signals neither shrink nor grow layer by layer. primer.ml.deep_nets
inpainting Filling in a missing or masked part of an image so that it fits the rest; a generative model does it by sampling only the unknown pixels.
input gate The LSTM dial that decides what new information to write into the cell state. primer.ml.cnn_rnn
instruction tuning Fine-tuning a pretrained model on many tasks written as instructions paired with good responses, so it learns to follow instructions it has never seen. primer.ml.training_stages
instrumentation Code that records spans or metrics around the work a program does. primer.agents.observability
int4 quantization Storing each weight as a 4-bit integer times a shared scale. primer.ml.inference
int8 quantization Storing each weight as an 8-bit integer times a shared scale. primer.ml.inference
inter-annotator agreement How often two people labelling the same items independently give the same label. It caps how finely any evaluation built on those labels can tell two models apart. primer.ml.fine_tuning
interaction effect Part of a prediction that depends on how two or more inputs combine, beyond what each contributes on its own: size mattering more in one city than another.
interference Signals getting in each other's way: other features leaking into one feature's reading (superposition), or merged task vectors changing the same weights in opposite directions (model merging). primer.ml.interpretability
internal fragmentation Memory reserved for a request that it never uses. primer.ml.inference
interpretability Studying what a model's internal numbers represent, and which of them cause its outputs. primer.ml.interpretability
inverted file In an IVF index, the list of vectors assigned to one cluster, so a search scans only the lists it picks. primer.ml.embeddings.ann
isoflop profile A curve of final loss against model size at a fixed training compute; its minimum is the best size for that budget. primer.ml.training_stages
item response theory Modelling the chance of a right answer from a taker's ability minus a question's difficulty. primer.ml.benchmarks
IVF A vector index that clusters vectors ahead of time and searches only the clusters nearest the query. primer.ml.embeddings.ann
IVF-PQ IVF clustering combined with compressed offsets from each cluster centre: the standard index for billion-scale search. primer.ml.embeddings.ann
Jaccard similarity Shared items divided by all distinct items across two sets, from 0 to 1. primer.ml.fine_tuning
jailbreak A prompt crafted to talk a model out of its safety training, such as a role-play or a disguised request, so it produces what it would normally refuse. primer.agents.guardrails
Jensen-Shannon divergence A measure of how different two distributions are; stuck at log 2 whenever they don't overlap. primer.ml.generative.gans
jitter Random variation added to retry delays so clients don't all retry at the same moment. primer.agents.failures
JSON mode A decoding option that guarantees parseable JSON, but not any particular shape. primer.ml.structured_output
JSON Schema A standard way to describe the shape of JSON data: which fields exist, their types and which are required. primer.agents.tools
JSON-RPC A simple convention for calling functions on another program by exchanging JSON requests, replies and notifications. primer.agents.mcp
k-means Clustering by repeatedly assigning each point to its nearest center and moving each center to the average of its points. primer.ml.embeddings.clustering
k-means++ A way to pick k-means starting centers spread far apart, which avoids bad results. primer.ml.embeddings.clustering
kernel In a CNN, the small grid of learned weights a convolution slides over an image; also called a filter. primer.ml.cnn_rnn
kernel fusion Doing several steps in one GPU program, so data is read from and written to main memory only once. primer.ml.inference
key In attention, the vector a token offers so others can decide whether it is relevant to them. primer.ml.attention
kill switch A way to stop an agent instantly, per customer or globally. primer.agents.deployment
KL divergence The extra surprise from using distribution q when the truth is p; zero only when they match. primer.ml.training_stages
KL penalty A term subtracted from the reward that grows as the tuned model drifts from its starting point, measured by KL divergence: a leash that keeps it near the data the reward model was trained on. primer.ml.reinforcement
KV cache Stored keys and values from earlier tokens, so each new token is computed without redoing work. It trades GPU memory for speed. primer.ml.inference
kv-cache quantization Storing cached keys and values in fewer bits, with one scale per vector. primer.ml.efficient_architectures
L0 norm The number of non-zero entries in a vector. For a sparse autoencoder, how many latents fire on an input.
l1 regularization Penalizing the sum of absolute weights, which drives unneeded weights to exactly zero. primer.ml.regularization
l2 normalization Dividing a vector by its own length so it has length 1. primer.ml.embeddings.similarity
L2 regularization Penalizing the sum of squared weights, which shrinks all weights toward zero without eliminating them. primer.ml.regularization
label noise Reference labels that are wrong, which cap what a model can learn and what an eval can measure. primer.ml.fine_tuning
label smoothing Training against a softened target, such as 0.9 on the right class and the rest spread evenly, to discourage over-confidence. primer.ml.losses
Langevin dynamics Sampling by repeatedly taking a small step along the score, towards where data is denser, and adding a little fresh noise. A diffusion sampler has this shape. primer.ml.generative.diffusion
language identification Guessing which language a text is in, so a pipeline keeps only the ones it wants. primer.ml.pretraining
late interaction Keeping one vector per word and matching each query word to its best document word. primer.ml.embeddings.retrieval
latency The delay between something happening and the system's response to it. primer.agents.cost
latent diffusion Running diffusion on an autoencoder's small compressed code instead of on pixels, then decoding once at the end. primer.ml.generative.diffusion
latent space The space of codes a model works in, where position means something and nearby codes decode to similar data. primer.ml.generative.autoencoders
law of large numbers The rule that the average of many independent random draws settles ever closer to the true average as more draws are added.
layer normalization Rescaling each token's numbers to a steady average and spread, which keeps training stable. primer.ml.deep_nets
layout shift Controls moving between runs, so clicks at remembered coordinates miss. primer.agents.coding_agents
leaf An end box of a decision tree, predicting the majority label or average value of the training examples that reached it. primer.ml.classical
learning rate How big a step each training update takes. Too high and training blows up; too low and it crawls. primer.ml.optimizers
learning-rate schedule A rule that changes the learning rate during training, such as warming up and then decaying along a cosine curve. primer.ml.optimizers
least privilege Giving each agent or tool only the permissions its job needs. primer.agents.tools
least squares Choosing the parameters that make the sum of squared errors as small as possible. primer.ml.regularization
length normalization BM25's discount for mentions in documents longer than average, controlled by b. primer.ml.embeddings.retrieval
lethal trifecta Private data, untrusted content and a way to send data out, all in one agent: together they allow data theft. primer.agents.guardrails
Likert scale A rating on a short fixed ladder of labelled points, such as 1 (not helpful) to 6 (highly helpful), averaged over many items to compare systems. primer.agents.evals
line search Having picked a direction to move in, trying different step lengths along it and keeping the one that lowers the loss most.
linear attention Attention variants that avoid scoring every pair of tokens, so cost grows in proportion to n instead of n². primer.ml.attention
linear probe Freezing a model and training only a simple linear classifier on its embeddings or internal activations, to measure what they encode. primer.ml.embeddings.contrastive
linear representation hypothesis The idea that features are directions, readable with a dot product. primer.ml.interpretability
linear time invariance A sequence model whose update rule is the same at every step, whatever the input; such a model can be computed as one convolution. primer.ml.efficient_architectures
linter A program that reads code without running it and reports likely mistakes, such as a syntax error, bad indentation or a name that is never defined.
Lipschitz A function is K-Lipschitz if its output never changes more than K times as fast as its input: a speed limit on its slope everywhere. primer.ml.generative.gans
LLM Large language model: a transformer trained to predict the next token, then tuned to follow instructions. primer.agents.llm
llm-as-judge Using a language model with a rubric to grade open-ended outputs, checked against human ratings. primer.agents.evals
load-balancing loss A small extra training penalty that pushes a mixture-of-experts router to spread tokens evenly across experts. primer.ml.transformer
local model A language model running on your own machine rather than a provider's servers: free per call, private and offline, but smaller. primer.agents.llm
locality-sensitive hashing Hashing vectors so that similar vectors are likely to land in the same bucket, for fast approximate search. primer.ml.embeddings.ann
log-derivative trick Rewriting ∇π as π·∇log π, so a sampled action gives an estimate of the gradient. primer.ml.reinforcement
log-likelihood The log of how probable the observed data is under a model. primer.ml.benchmarks
log-mel spectrogram A spectrogram pooled into mel bands with loudness on a log scale: what speech models read. primer.ml.generative.multimodal
log-odds The logarithm of p / (1 − p): 0 for a 50% chance, positive when more likely than not. A sigmoid turns log-odds back into a probability. primer.ml.classical
log-probability The logarithm of a probability; for a whole response, the sum of its tokens' log-probabilities. primer.ml.training_stages
log-sum-exp Computing the log of a sum of exponentials without overflow, by factoring out the largest term first. primer.ml.losses
logarithm The undo button for exponentiation: ln(y) asks what power of e gives y. Logs turn multiplication into addition. primer.notation
logistic regression A linear model whose weighted sum passes through a sigmoid to give a probability. primer.ml.neural_net
logit difference The correct answer's score minus a wrong answer's score. primer.ml.interpretability
logit lens Applying the model's own output layer to intermediate layers to see what it would predict so far. primer.ml.interpretability
logits The raw scores a model outputs for every possible next token, before softmax turns them into probabilities. primer.ml.big_picture
long-term memory Facts, events and procedures stored outside the model and recalled into later conversations. primer.agents.memory
loop detection Stopping an agent that keeps requesting the same tool with the same arguments. primer.agents.agent_loop
LoRA Low-rank adaptation: freeze the model and train a small pair of matrices whose product is added to a weight matrix. primer.ml.training_stages
loss A single number measuring how wrong the model's predictions are. Training tries to make it smaller. primer.ml.losses
loss function The rule that turns predictions and correct answers into the loss. Choosing it defines what the model learns. primer.ml.losses
loss mask A per-position switch deciding which predictions count toward the loss, such as only the reply in SFT. primer.ml.training_stages
loss scaling Multiplying the loss so small fp16 gradients stay above zero, then dividing back. primer.ml.hardware
loss spike A sudden jump in training loss that may recover or diverge. primer.ml.pretraining
Lost in the Middle Models use information at the start and end of a long input more reliably than information in the middle. primer.agents.context
LSTM An RNN with gates that decide what to forget, what to write and what to output, so it can remember across long sequences. primer.ml.cnn_rnn
Luhn check A checksum every real card number satisfies, used to tell card numbers from look-alike IDs. primer.agents.guardrails
machine translation Turning text in one language into another. primer.ml.transformer
Maj@K A score for sampling k answers per question and keeping the most common final answer: the share of questions where that majority answer is right. primer.ml.reasoning
majority voting Choosing the answer that the most samples agree on. primer.ml.reasoning
Mamba A state-space model that trains in parallel and runs in time linear in sequence length. primer.ml.efficient_architectures
mantissa The bits of a float that store its digits, which set its precision. primer.ml.hardware
mAP@10 Mean average precision over the top 10 results: rewards putting correct matches in the top 10, and higher up within it. primer.ml.metrics
margin of error The ± range around a measured score that the true score probably falls in; it shrinks with the square root of the sample size. primer.ml.fine_tuning
marginal likelihood How probable a model finds an example, averaged over every hidden cause that could have produced it. With a neural network inside the model that average is an intractable integral, which is why VAEs train on a lower bound instead. primer.ml.generative.autoencoders
Markov chain A sequence of random steps where each step depends only on the one just before it, not on the whole history. primer.ml.generative.diffusion
master weights An fp32 copy of the weights that receives optimizer updates, so tiny updates aren't lost. primer.ml.pretraining
matrix A table of numbers with rows and columns. A batch of vectors stacked as rows is a matrix, and most model weights are matrices. primer.notation
matrix multiply A grid of dot products: each output cell is one row of the first matrix dotted with one column of the second. primer.notation
Matryoshka embedding An embedding trained so its first few dimensions work as a smaller embedding on their own. primer.ml.embeddings.compression
max pooling Keeping only the largest value in each small block of a feature map. primer.ml.cnn_rnn
max-norm constraint Capping the length of each neuron's incoming weight vector at a fixed radius, shrinking it back whenever an update pushes it past. primer.ml.regularization
MaxSim For each query word, its highest similarity to any document word, summed over the query. primer.ml.embeddings.retrieval
MCP Model Context Protocol: an open standard for connecting AI apps to tools and data through reusable servers. primer.agents.mcp
mcp resource Read-only data an MCP server offers for the app to load into context, addressed by a URI. primer.agents.mcp
MCP server The program that wraps a real system (files, a database, an API) and offers it to AI apps over MCP. primer.agents.mcp
mean absolute error The average of absolute prediction errors, which is robust to outliers. primer.ml.losses
mean average precision How well a system ranks correct results above wrong ones, averaged over queries or classes. primer.ml.metrics
mean squared error The average of squared prediction errors, so big misses dominate. primer.ml.losses
mean-centering Subtracting the average vector of a collection from every vector, removing the direction they all share. primer.ml.embeddings.similarity
median The middle value once numbers are sorted (the average of the two middle ones for an even count). Unlike the mean, one extreme value barely moves it.
mel scale A relabelling of frequency so equal steps sound equally far apart to human ears. primer.ml.generative.multimodal
memory bandwidth How many bytes per second a memory can deliver. primer.ml.hardware
memory hierarchy The chain of storage from registers to the network, each level bigger and slower than the last. primer.ml.hardware
memory-bound Limited by how fast data arrives from memory rather than by arithmetic, as in decoding. primer.ml.inference
micro-batch A slice of a batch that moves through a pipeline on its own. primer.ml.pretraining
MinHash A short signature per document whose matching slots estimate Jaccard similarity. primer.ml.pretraining
minimax A game where one player tries to make a number as large as possible and the other as small as possible. primer.ml.generative.gans
MIPS Maximum inner product search: finding the stored vectors with the largest dot product against a query vector, usually approximately with an index. primer.ml.embeddings.ann
mixed precision Doing the big multiplies in 16- or 8-bit while keeping master weights and sums in fp32. primer.ml.hardware
mixture of experts Replacing one feed-forward network with many expert networks and a router that sends each token to a few of them. primer.ml.transformer
MLP Multi-layer perceptron: the plainest neural network, layers of weighted sums each followed by a nonlinearity, with every unit connected to every unit in the next layer. primer.ml.neural_net
modality One kind of input a model can take: text, images, audio or video. primer.ml.generative.multimodal
mode collapse A generator producing only a few kinds of output instead of the full variety in the data. primer.ml.generative.gans
model collapse Losing rare data (the tails) when models are trained on their own outputs. primer.ml.pretraining
Model Context Protocol An open standard that lets any AI application connect to any tool server the same way. primer.agents.mcp
model editing Changing one specific fact or behaviour inside a trained model by adjusting a few weights directly, without retraining, while leaving everything else as it was. primer.ml.interpretability
model merging Building one model from several fine-tunes by arithmetic on their weights, with no extra training. primer.ml.fine_tuning
model parallelism Splitting one model's weights across GPUs, either inside each layer (tensor parallelism) or by layers (pipeline parallelism). The ZeRO and Megatron-LM papers use it for the first. primer.ml.pretraining
model routing Sending each request to the cheapest model that can handle it well. primer.agents.cost
momentum Keeping a running average of recent gradients so updates roll smoothly through noise, like a ball gathering speed downhill. primer.ml.optimizers
monosemantic Responding to one understandable thing only: said of a neuron, or of a feature a sparse autoencoder finds. primer.ml.interpretability
MRR Mean reciprocal rank: the average of 1 / (position of the first relevant result). primer.ml.metrics
multi-agent system Several agents that coordinate, typically a supervisor delegating to specialist workers. primer.agents.orchestration
multi-armed bandit The simplest RL problem: pick among options with hidden payouts, and learn which pays best by trying them. primer.ml.reinforcement
multi-head attention Several attention computations run in parallel on slices of the vectors, so each head can track a different kind of relationship. primer.ml.attention
multi-head latent attention Caching one small latent vector per token and rebuilding every head's keys and values from it. primer.ml.efficient_architectures
multi-query attention All query heads share a single key/value head. primer.ml.efficient_architectures
multi-query retrieval Searching several phrasings of a question and fusing the results. primer.agents.rag
multi-tenant One system serving many separate customers whose data must never mix. primer.agents.memory
multilayer perceptron A network of fully connected layers, each a matrix multiply followed by a nonlinearity; also called an MLP. primer.ml.neural_net
multimodal model A model that takes in more than one kind of input (text, images, audio, video) as one sequence of tokens. primer.ml.generative.multimodal
multitask learning Training one model on several tasks at once, with part of the input saying which task to do, so what it learns for one task can help the others.
n-gram overlap The share of a text's n-word runs that also appear in another corpus. primer.ml.benchmarks
named-entity recognition A model that tags names, places and organisations in text. primer.agents.guardrails
navigable small world A graph where always stepping to the neighbour closest to the target reaches it in few hops. primer.ml.embeddings.ann
nDCG A ranking score that gives more credit for relevant results near the top and handles degrees of relevance. primer.ml.metrics
near-duplicate detection Finding items whose embeddings are almost identical, such as reworded copies. primer.ml.embeddings.clustering
negative sampling Training by scoring the true pair up and a few random noise pairs down, instead of scoring the whole vocabulary. primer.ml.embeddings.word2vec
NER Named-entity recognition: tagging names, places and organisations in text. primer.agents.guardrails
neural ode A model whose output is found by following an ordinary differential equation whose rate of change is computed by a neural network; it can be run backwards and gives exact probabilities.
neuron Multiply each input by a weight, add them up with a bias, and pass the result through a nonlinear function. primer.ml.neural_net
Newton's method A step that uses the curvature (second derivative) as well as the slope: divide the slope by the curvature to jump to the bottom of the parabola that matches the loss at the current point.
NF4 4-bit NormalFloat: a 4-bit number format whose 16 levels are spaced to match the bell-curve shape of neural network weights. primer.ml.inference
noise schedule How much noise each step adds (the betas), which sets how fast the signal fades. primer.ml.generative.diffusion
non-saturating loss The generator loss −log D(G(z)), which keeps a strong gradient when the discriminator confidently rejects fakes. primer.ml.generative.gans
norm The length of a vector: square the entries, add them, take the square root. primer.notation
nprobe How many IVF clusters a query scans; raising it trades speed for recall. primer.ml.embeddings.ann
NTK-aware scaling Extending a RoPE model's context by raising the rotation base, which slows the slow frequencies while keeping the fast ones that encode local order. primer.ml.positional
OAuth A standard way for a user to grant an app limited, revocable access without sharing their password. primer.agents.mcp
OCR Optical character recognition: reading text from an image of a page. primer.agents.rag
off-by-one error An index or count one position away from the right one. primer.agents.coding_agents
Ollama A free app that downloads open language models and runs them on your own computer, served over a small local web API. primer.agents.llm
one-hot A vector with a 1 at the correct class and 0 everywhere else. primer.ml.losses
one-hot encoding Turning a category into one yes/no column per possible value. primer.ml.classical
OpenTelemetry The open standard for traces, metrics and logs. primer.agents.observability
optimal discriminator p_data/(p_data + p_g): the best possible verdict against a fixed generator; 1/2 everywhere at equilibrium. primer.ml.generative.gans
optimal transport Moving one distribution onto another as cheaply as possible, where moving mass further costs more. For the bell curves of flow matching, every bit of mass then travels in a straight line at constant speed. primer.ml.generative.gans
optimizer The rule that turns gradients into weight updates, such as SGD, momentum or Adam. primer.ml.optimizers
optimizer state The running numbers an optimizer keeps for every weight between steps, such as Adam's two running averages. With an fp32 master copy of the weights it is 12 bytes per parameter, the biggest part of training memory. primer.ml.pretraining
orchestrator-workers One model decides the subtasks, workers handle them, and one model combines the results. primer.agents.orchestration
ordinary differential equation A rule giving, at every moment, how fast something is changing. Solving it means following that rule forward in time, for example with small Euler steps. primer.ml.generative.diffusion
orthogonal At right angles: two vectors whose dot product, and so whose cosine similarity, is zero, so moving along one does not move you along the other. primer.ml.embeddings.similarity
OTLP The OpenTelemetry Protocol: the wire format exporters use to ship telemetry. primer.agents.observability
out-of-bag The rows left out of a tree's bootstrap sample, usable as free validation data for that tree. primer.ml.classical
out-of-distribution Unlike the examples a model learned from or was shown, such as longer or rarer inputs. Performance there is usually lower and harder to predict.
outcome reward model A verifier that scores only the final answer. primer.ml.reasoning
outcome supervision Training a reward model from the final result alone: each solution is labelled right or wrong by its answer, never step by step. primer.ml.reasoning
outer product A column vector times a row vector, giving a table whose entry (m, c) is the product of their m-th and c-th numbers. primer.ml.efficient_architectures
output gate The LSTM dial that decides how much of the cell state to show as output. primer.ml.cnn_rnn
over-refusal Refusing a harmless request because a safety check is too strict. primer.ml.alignment
overfitting When a model memorizes its training data, noise included, and does worse on new data. primer.ml.regularization
overoptimization Optimizing against a learned reward model for so long that the true quality it stood for starts to fall, even as its score keeps rising: Goodhart's law for reward models. primer.ml.alignment
overthinking Spending many reasoning tokens where few would do, wasting cost and time and sometimes losing a right answer. primer.ml.reasoning
p95 The 95th percentile: the value that 95% of measurements are at or below. primer.agents.evals
padding Zeros added around an image so a filter can centre on the edge pixels. primer.ml.cnn_rnn
PagedAttention Storing the KV cache in fixed-size pages, like virtual memory, to avoid wasted GPU memory. primer.ml.inference
paired bootstrap Resampling questions with both models' results kept together, to test whether a gap is real. primer.ml.benchmarks
parallel scan Computing every state of a linear recurrence in about log₂ n parallel rounds by merging steps. primer.ml.efficient_architectures
parallel tool calls Several tool requests in one model turn, run at the same time, with all results returned in one message. primer.agents.agent_loop
parallelism Doing many pieces of work at the same time instead of one after another. primer.ml.attention
parameters All the learned numbers in a model (its weights and biases). A 70B model has 70 billion of them. primer.ml.neural_net
parent-child retrieval Matching small chunks but returning the larger section around them. primer.agents.rag
partial dependence plot A plot of a model's average prediction as one or two inputs are set to each value in turn, with every other input left as it is in the data.
pass-to-pass test A hidden test that passed before a fix and must still pass after it. primer.agents.coding_agents
pass@1 The share of problems solved by the first program submitted for each, judged by hidden tests. primer.agents.evals
pass@k The chance that at least one of k sampled answers passes the tests. primer.ml.benchmarks
pass@n The chance that at least one of n samples is right: 1 − (1 − p)ⁿ. primer.ml.reasoning
patch A small square cut from an image and flattened into one token. primer.ml.cnn_rnn
patch embedding The learned matrix that turns a flattened image patch into a token vector. primer.ml.generative.multimodal
PCA Principal component analysis: finding the directions along which data varies most, to draw or compress it with fewer numbers. primer.ml.embeddings.clustering
Pearson correlation How closely two lists of numbers rise and fall together along a straight line, from −1 to 1; 0 means no straight-line relationship.
per-channel quantization One scale per weight row, so a single outlier doesn't coarsen all the others. primer.ml.inference
Perceiver Resampler A small transformer whose queries are a fixed set of learned vectors: they cross-attend to any number of image or video features and always return that fixed number of visual tokens. primer.ml.generative.multimodal
percentage point The plain difference between two percentages: going from 11% to 18% is a rise of 7 percentage points, which is a 64% relative rise.
perceptual loss A loss that compares two images through the features of a pretrained network rather than pixel by pixel, so it punishes the differences a person would notice. LPIPS is a widely used one. primer.ml.generative.gans
permission-aware retrieval Filtering out documents a user may not read before anything is ranked or shown to the model. primer.agents.rag
permutation equivariance Shuffling a layer's inputs just shuffles its outputs the same way, which is why attention alone cannot see word order. primer.ml.positional
permutation importance The accuracy lost on held-out data when one column is shuffled. primer.ml.classical
perplexity e raised to the average cross-entropy: roughly how many options the model is torn between at each step. primer.ml.losses
petaflop/s-day 10¹⁵ operations per second for one day, 8.64 × 10¹⁹ operations: a unit for training budgets. primer.ml.training_stages
PII Personally identifiable information: names, emails, phone numbers, card numbers and similar. primer.agents.guardrails
pipeline bubble Time pipeline stages sit idle while the pipeline fills and drains. primer.ml.pretraining
pipeline parallelism Giving each GPU a stage of layers and passing activations between neighbouring stages. primer.ml.hardware
plan-and-execute An agent pattern that writes a plan first, then executes it step by step, replanning when results surprise it. primer.agents.planning
PMI Pointwise mutual information: the log of how much more often two words appear together than chance predicts. primer.ml.embeddings.word2vec
policy In reinforcement learning and preference tuning, the model being trained, viewed as a probability distribution over actions or responses. primer.ml.reinforcement
policy gradient Raising expected reward by making the actions that earned more reward more likely. primer.ml.reinforcement
polysemantic neuron A neuron that responds to several unrelated features. primer.ml.interpretability
pooling Shrinking a feature map by keeping only the strongest (or average) value in each small window. primer.ml.cnn_rnn
popcount Counting the 1-bits in a number; combined with XOR it computes Hamming distance in one step. primer.ml.embeddings.compression
position interpolation Stretching a model to longer contexts by scaling every position down into the range it saw during training. primer.ml.positional
positional encoding Information added to token vectors so the model knows word order, which attention alone ignores. primer.ml.positional
posterior What you believe about a hidden quantity after seeing the data: the prior, reweighted by how well each value explains what was observed. primer.ml.generative.autoencoders
posterior collapse When a VAE's code carries no information because the KL penalty outweighs what the code saves in rebuild error, so every output is the same average. primer.ml.generative.autoencoders
power iteration Repeatedly multiplying a vector by a matrix (and its transpose) until it points along the most-stretched direction. primer.ml.generative.gans
power law A relationship where one quantity is a fixed power of another, y = a·x^k; on log-log axes it is a straight line. primer.ml.transformer
PPO Proximal Policy Optimization: a reinforcement learning algorithm that improves a policy in small, clipped steps; widely used for RLHF. primer.ml.training_stages
pre-norm Normalizing before each sub-layer of a transformer block, leaving the residual path untouched, which trains more stably than normalizing after. primer.ml.transformer
pre-tokenization Splitting text into chunks, such as words with their leading space, before BPE runs, so merges never cross chunk boundaries. primer.ml.tokenization
precision Of the items flagged, the fraction that were right. primer.ml.metrics
precision-recall curve Precision plotted against recall across thresholds; more honest than ROC for rare events. primer.ml.metrics
precision@k The share of the top k search results that are relevant. primer.ml.metrics
preference model Another name for a reward model: it scores a response so that the gap between two scores predicts which one people (or a model) prefer. primer.ml.training_stages
preference tuning Training on which of two responses people preferred, to shape tone, helpfulness and safety. primer.ml.training_stages
prefill The first phase of generation: the whole prompt is processed in one parallel pass. It sets the time to the first token. primer.ml.inference
prefix tuning Training a few special virtual-token vectors placed before every input while the model stays frozen. primer.ml.training_stages
pretraining The first, most expensive training stage: predicting the next token over trillions of tokens of text. primer.ml.training_stages
prior What you believe about a hidden quantity before seeing any data. A VAE's prior over codes is the standard normal distribution. primer.ml.generative.autoencoders
privilege separation Splitting work so the part that reads untrusted content can't take dangerous actions. primer.agents.guardrails
privileged basis Directions made special by the architecture, such as neurons followed by an activation function that acts on each number separately. Only in a privileged basis does asking what one neuron means make sense. primer.ml.interpretability
probability density How thickly a continuous distribution's samples cover each spot: high where they crowd, zero where none ever land. It is the height of the bump, and its area adds up to 1. primer.ml.generative.gans
probability distribution A list of probabilities over all possible outcomes, each between 0 and 1, adding up to 1. primer.notation
probability flow ODE The deterministic equation whose solutions carry noise to data with the same in-between distributions as a diffusion process; DDIM sampling is one way of stepping along it. primer.ml.generative.diffusion
probability ratio The current policy's probability of a sampled action divided by its probability when the action was sampled. primer.ml.reinforcement
probe A small classifier trained on a frozen model's activations to test whether a property is encoded there. primer.ml.interpretability
procedural memory Memory of how to do things, such as a learned workflow or saved skill. primer.agents.memory
process reward model A verifier that scores each intermediate step. primer.ml.reasoning
process supervision Training a reward model from labels on each intermediate step, so it learns where a solution went wrong. primer.ml.reasoning
product quantization Compressing a vector by splitting it into chunks and replacing each chunk with the ID of its nearest entry in a small codebook. primer.ml.embeddings.ann
projector A small layer that maps an encoder's vectors into a language model's embedding space. primer.ml.generative.multimodal
prompt The text sent to a language model: instructions, context and the question. primer.agents.llm
prompt caching Reusing the processed form of a prompt prefix that repeats across requests, which cuts cost and time to first token. primer.ml.inference
prompt chaining A fixed sequence of model calls where each output feeds the next, with code checks in between. primer.agents.orchestration
prompt injection Hostile instructions hidden in content the model reads, such as an email or web page, trying to hijack it. primer.agents.guardrails
pull request A proposed set of changes to a code repository, submitted for review; once merged, the changes become part of the project. primer.agents.coding_agents
pushdown automaton A finite-state machine plus a stack: enough to check nested formats like JSON or SQL. primer.ml.structured_output
QLoRA LoRA adapters trained on top of base weights stored in 4 bits. primer.ml.training_stages
quality filter A rule or classifier that throws out low-quality pages before training. primer.ml.pretraining
quantile The value below which a given share of the data falls: the 0.5 quantile is the median, the 0.9 quantile has 90% of values below it.
quantization Storing numbers with fewer bits (say 8 or 4 instead of 16 or 32), which shrinks memory with a small loss of precision. primer.ml.inference
query In attention, the vector a token uses to ask what it is looking for. primer.ml.attention
query and passage prefixes Labels some embedding models expect on questions and documents; forgetting one silently lowers recall. primer.ml.embeddings.retrieval
query rewriting Turning a follow-up or vague question into a standalone search query. primer.agents.rag
RAG Retrieval-augmented generation: fetch relevant passages first, then have the model answer using them, with citations. primer.agents.rag
random forest Many deep trees on bootstrap samples, each split limited to a random subset of features, with their votes averaged. primer.ml.classical
rank How many independent directions a matrix really contains; a table made by multiplying one column by one row has rank 1. primer.ml.training_stages
rate limit A cap on how many actions may happen per unit of time. primer.agents.deployment
re-scoring Recomputing exact scores for a shortlist that an approximate index returned, to recover accuracy. primer.ml.embeddings.ann
ReAct Reason plus act: the agent pattern that interleaves reasoning steps with tool calls. primer.agents.agent_loop
reasoning model A model trained to write out and check intermediate steps before it commits to an answer. primer.ml.reasoning
recall Of the items that truly mattered, the fraction that were found. primer.ml.metrics
recall@k The fraction of relevant documents that appear in the top k results. For RAG, usually the metric that matters most. primer.ml.metrics
receptive field How much of the original input one neuron can see. primer.ml.cnn_rnn
reciprocal rank fusion Merging ranked lists by giving each document 1 / (k + rank) from every list it appears in, then adding those up. primer.ml.embeddings.retrieval
rectified flow Flow matching with straight-line paths, retrained on its own outputs so the paths get straighter and need fewer steps. primer.ml.generative.diffusion
red-teaming Searching systematically for inputs that make a model or its safety checks fail. primer.ml.alignment
reduce-scatter Summing a vector across GPUs so each GPU ends up holding the total for only its own slice; the first half of a ring all-reduce. primer.ml.pretraining
reference model A frozen copy of the starting model that DPO and RLHF measure drift against. primer.ml.training_stages
reflection Having a model review and revise its own output. Useful, but external checks such as tests are more reliable. primer.agents.planning
reflow Retraining a rectified flow on its own (noise, sample) pairs. The new pairs' straight lines rarely cross, so the new flow's paths are straighter and need fewer steps. primer.ml.generative.diffusion
register The tiny storage right beside the arithmetic units, holding the numbers being worked on this instant. primer.ml.hardware
regression A task that used to pass and now fails after a change. primer.agents.evals
regret In online learning, the total extra loss from choosing each step's parameters before seeing that step's data, compared with the best fixed choice in hindsight. primer.ml.optimizers
regular expression A pattern language (such as [0-9]+ or cat car dog) that can always be compiled into a finite-state machine. primer.ml.structured_output
regularization Any technique that discourages memorizing, such as dropout, weight decay or early stopping. primer.ml.regularization
REINFORCE The basic policy-gradient algorithm: step along reward times the gradient of log π(action). primer.ml.reinforcement
reinforcement learning Learning from a score for what you did, rather than from the correct answer. primer.ml.reinforcement
rejection sampling Generating many candidate outputs and keeping only those that pass a check, such as a correct final answer, often to use as training data.
release gate Limits set before measuring; a release goes ahead only if every evaluation is within its limit. primer.ml.alignment
ReLU An activation function that keeps positive numbers and turns negatives into zero. primer.ml.neural_net
reparameterization trick Writing a random draw as z = μ + σ·ε with the noise ε as a separate input, so gradients can flow through sampling. primer.ml.generative.autoencoders
replay Mixing a small sample of old-task examples into new training data so the old skill keeps getting practised. primer.ml.fine_tuning
reranker A second, more accurate model that reorders the top results of a fast first-stage search. primer.ml.embeddings.retrieval
residual The true value minus the current prediction; for squared error it is the negative gradient. primer.ml.classical
residual connection Adding a layer's input back to its output (x + f(x)), giving gradients a shortcut through deep networks. primer.ml.deep_nets
residual stream The running vector for each token that every transformer block reads from and adds a correction to, instead of replacing it. primer.ml.transformer
residual vector quantization Stacking codebooks, each one encoding the error the previous ones left. primer.ml.generative.multimodal
ResNet A deep convolutional network built from residual blocks (x + f(x)), which made networks of a hundred layers and more trainable. primer.ml.deep_nets
resolved rate The share of tasks whose patch passes every hidden fail-to-pass and pass-to-pass test. primer.agents.coding_agents
retrieval failure A wrong answer caused because no relevant document was retrieved. primer.ml.embeddings.operations
reverse process The learned half of a diffusion model: a chain of small denoising steps that turns pure noise back into data. primer.ml.generative.diffusion
reward The single number the environment returns to say how good an action was. primer.ml.reinforcement
reward hacking A policy maximising the reward as written while the real goal gets worse. primer.ml.reinforcement
reward model A model trained to predict which of two responses a human would prefer. primer.ml.training_stages
right to erasure A user's legal right, for example under GDPR, to have their personal data deleted. primer.agents.memory
ring all-reduce All-reduce by passing chunks around a ring; each GPU sends about twice its data, however many GPUs there are. primer.ml.hardware
RLAIF Reinforcement learning from AI feedback: preference labels come from a model applying written principles. primer.ml.alignment
RLHF Reinforcement learning from human feedback: a reward model learns human preferences, and the language model is tuned to score well on it. primer.ml.training_stages
RMSNorm A cheaper layer normalization that divides each token's numbers by their root-mean-square without subtracting the mean first. primer.ml.transformer
RMSProp An optimizer that divides each step by the square root of a running average of recent squared gradients. primer.ml.optimizers
RNN A recurrent neural network: it reads a sequence one step at a time, carrying a running summary called the hidden state. primer.ml.cnn_rnn
ROC curve True-positive rate plotted against false-positive rate as the decision threshold sweeps. primer.ml.metrics
ROC-AUC The probability that the model ranks a random positive above a random negative. primer.ml.metrics
rolling buffer cache A KV cache with w slots where each new token overwrites the oldest. primer.ml.efficient_architectures
roofline A chart of the speed a chip can reach at each arithmetic intensity, capped first by memory and then by compute. primer.ml.inference
RoPE Rotary position embedding: encodes position by rotating query and key vectors, so attention depends on the distance between tokens. primer.ml.positional
ROUGE-L An overlap score based on the longest in-order sequence of words shared with a reference. primer.ml.metrics
router A step that classifies a request and sends it down the right path, or to the right model. primer.agents.orchestration
routing One model call classifies a request and sends it to a specialised handler. primer.agents.orchestration
rubric A short list of explicit pass/fail criteria given to a grader. primer.agents.evals
rug pull A tool server changing its definitions after they were approved. primer.agents.mcp
sample rate How many measurements of a signal are taken per second, such as 16,000 for speech. primer.ml.generative.multimodal
sandbox An isolated place to run untrusted code, where the worst it can do is fail: no network, no secrets, time and memory limits. primer.agents.coding_agents
sandwich ordering Placing the best retrieved chunks at the start and end of the context and the weakest in the middle. primer.agents.context
saturation When an activation like sigmoid or tanh sits on its flat tail, so almost no gradient passes through. primer.ml.neural_net
scalar quantization Storing each number as one of 256 levels (one byte) instead of a 4-byte float. primer.ml.embeddings.compression
scaling law A smooth, predictable rule for how a model's loss falls as its size, data or compute grows: a straight line on log-log axes. It lets small runs forecast a big model's quality. primer.ml.transformer
score The direction in which data gets more crowded fastest; the noise guess, flipped and rescaled. primer.ml.generative.diffusion
screenshot The image of the screen a computer-use agent receives after each action. primer.agents.coding_agents
selective state-space model A state-space model whose step size (how much to keep and write) is computed from each token, as in Mamba. primer.ml.efficient_architectures
self-attention Attention where a sequence attends to itself: every token scores every other token in the same text. primer.ml.attention
self-consistency Sampling several chains of thought and returning the most common final answer. primer.ml.reasoning
self-supervised learning Learning from unlabelled data by predicting a hidden part of each example from the rest, such as a masked word or a masked stretch of audio.
semantic cache Reusing a stored answer when a new question's embedding is nearly identical to a previous question's. primer.ml.embeddings.clustering
semantic memory Memory of facts, such as 'the user's fiscal year starts in April'. primer.agents.memory
semantic validation Checking that well-formed arguments are also true, such as that a customer id actually exists. primer.agents.tools
SentencePiece A tokenizer library that runs BPE or Unigram directly on raw text, writing spaces as the visible symbol ▁. primer.ml.tokenization
SFT Supervised fine-tuning: training on examples of instructions paired with good responses. primer.ml.training_stages
SHA-256 A hash function: a fixed-length fingerprint of data that changes completely if the data changes at all. primer.agents.deployment
shadow mode Running a new system on real inputs and recording what it would do, without letting it act. primer.agents.deployment
shingle A window of k neighbouring words, used to compare texts for near-duplicates. primer.ml.fine_tuning
short-term memory The current conversation, kept within a token budget: recent turns verbatim, older ones summarized. primer.agents.memory
short-time fourier transform A Fourier transform on each short, overlapping slice of a signal. primer.ml.generative.multimodal
shortlist The small set of top candidates from a cheap first-stage search that a slower, more precise stage reorders. primer.ml.embeddings.retrieval
shots Worked examples placed in the prompt before the question. primer.ml.benchmarks
shrinkage Pulling values toward zero: an L1 penalty's pull on every activation, or in gradient boosting the learning rate that keeps only part of each new tree's correction. primer.ml.interpretability
siamese network Two copies of one network with shared weights, each reading one input, so their outputs can be compared directly. primer.ml.embeddings.contrastive
sigmoid Squashes any number into the range 0 to 1, with an S-shaped curve. primer.ml.neural_net
signal-to-noise ratio How much stronger a signal is than the noise mixed into it, as a ratio of their powers (variances). Audio quotes it in decibels, where every 10 dB is ten times the ratio; in diffusion it falls from very large (clean) to nearly zero (pure noise). primer.ml.generative.diffusion
silhouette score How much closer each point is to its own cluster than to the nearest other one, from −1 to 1. primer.ml.embeddings.clustering
singular value decomposition Breaking a matrix into independent directions ranked by importance, so the top few capture most of what it does. primer.ml.embeddings.word2vec
sinusoidal positional encoding The original transformer's fixed position codes, made of sines and cosines at many frequencies, added to each token's embedding. primer.ml.positional
skip list A sorted list with random express lanes on top, searched by running along the top lane and dropping down, in about log(n) steps. primer.ml.embeddings.ann
skip-gram The word2vec task of predicting each word's neighbours from the word itself. primer.ml.embeddings.word2vec
sliding-window attention Each token attends only to the last w tokens, so cost and cache stop growing with context length. primer.ml.efficient_architectures
soft targets A teacher model's temperature-softened probabilities, used to train a student model. primer.ml.training_stages
soft thresholding Moving each weight a fixed amount toward zero and snapping any weight within that amount to exactly zero. primer.ml.regularization
softmax Turns a list of scores into shares that are all positive and add up to 1, with bigger scores getting disproportionately more. primer.ml.attention
softplus log(1 + eˣ): a smooth ramp that is always positive. primer.ml.efficient_architectures
span One step inside a trace, with its start time, duration and details. primer.agents.observability
sparse attention Attention that scores only a chosen pattern of token pairs (local, global, strided) instead of every pair. primer.ml.efficient_architectures
sparse autoencoder A wide encoder and decoder trained with an L1 penalty so each input uses a few learned features. primer.ml.interpretability
sparse gradients Gradients that are zero most of the time for a given weight, such as the weight for a rare word. primer.ml.optimizers
sparse retrieval Keyword search such as BM25, which scores exact word matches. primer.ml.embeddings.retrieval
Spearman correlation How well two rankings agree, from −1 (reversed) to 1 (identical). primer.ml.metrics
special token A token reserved for structure rather than text, such as a turn marker or an end-of-text marker. The model learns when to emit it, and generation stops at the end marker. primer.ml.fine_tuning
specification gaming Another name for reward hacking: satisfying the letter of an objective but not its intent. primer.ml.reinforcement
spectral normalization Dividing each layer's weights by the most they can stretch any input, which caps how fast the discriminator can change. primer.ml.generative.gans
spectrogram A picture of sound: time across, frequency up, brightness for loudness. primer.ml.generative.multimodal
speculative decoding A small fast model drafts several tokens, and the big model checks them all in one pass, keeping the ones it agrees with. primer.ml.inference
speculative execution Starting likely work early, in parallel, and throwing it away if the guess was wrong. primer.ml.inference
speech recognition Turning recorded speech into written text; also called automatic speech recognition (ASR). primer.ml.generative.multimodal
SRAM The tiny, very fast on-chip memory next to a GPU's arithmetic units. primer.ml.inference
standard deviation The square root of the variance: the typical distance of a value from the average. primer.notation
standard error The typical distance between a measured score and the true rate: √(p(1−p)/n). primer.ml.benchmarks
standard normal distribution The bell curve centred on 0 with spread 1; N(0, I) draws each number from it independently. primer.ml.generative.autoencoders
state machine A fixed set of states and allowed transitions, with code deciding every move. primer.agents.orchestration
state-space model A recurrent-style model, such as Mamba, that trains in parallel and runs in time linear in sequence length. primer.ml.efficient_architectures
static batching Serving a fixed group of requests until the longest one finishes. primer.ml.inference
static embedding One fixed vector per word, whatever the sentence. primer.ml.embeddings.word2vec
stdio transport Running a server as a child process and exchanging one JSON message per line over its input and output. primer.agents.mcp
steering Changing a model's behaviour while it runs by editing an internal activation, such as adding or pinning a feature's direction. primer.ml.interpretability
stochastic differential equation A rule for how something changes over time with two parts: a steady drift, and random jitter of a set strength. The noising process of a diffusion model is one, and running it backwards generates data. primer.ml.generative.diffusion
stochastic gradient descent Gradient descent where each step's slope is estimated from a small random batch instead of the whole dataset. primer.ml.optimizers
stop reason Why a model response ended: finished, wants a tool, hit the length limit, or declined. primer.agents.agent_loop
stop-gradient An operation that passes its input through unchanged but blocks gradients from flowing back into it, so training treats that input as a constant.
straight-through estimator A way to train through a step that has no useful gradient, such as rounding or snapping to a codebook: use the step going forward, and pass the gradient back as if the step were not there.
strict mode A tool or output option that guarantees the model's JSON fits a given schema. primer.ml.structured_output
strict tool use An API setting that guarantees the model's tool arguments match the tool's JSON Schema exactly. primer.agents.tools
stride How many pixels a convolution filter jumps between positions. primer.ml.cnn_rnn
structured output Making a model's answer follow an exact format, such as JSON that fits a schema, so a program can read it. primer.ml.structured_output
stump A decision tree with a single question and two leaves. primer.ml.classical
style control Adding length or format as extra factors in the rating fit, so style is separated from quality. primer.ml.benchmarks
subnormal A float below the smallest normal value, with the hidden leading 1 dropped so it fades toward zero. primer.ml.hardware
subsampling Randomly skipping most occurrences of very frequent words during training, which is faster and improves rare-word vectors. primer.ml.embeddings.word2vec
subword unit A piece of a word, from a single character to a whole frequent word; a fixed set of them can spell any word. primer.ml.tokenization
super-resolution Turning a low-resolution image into a plausible higher-resolution one; most of the fine detail has to be invented, not recovered.
superposition Storing more features than there are neurons, as nearly perpendicular directions. primer.ml.interpretability
supervisor In a multi-agent system, the agent that assigns tasks to specialist agents and assembles their answers. primer.agents.orchestration
SVD Singular value decomposition: splits a table into a few directions that capture its main patterns, used to compress it. primer.ml.embeddings.word2vec
SwiGLU A gated feed-forward layer: one projection, passed through the smooth SiLU activation, multiplies a second projection number by number before the output projection. primer.ml.transformer
sycophancy A model changing its answer to agree with a view the user states. primer.ml.alignment
synthetic data Training examples written by a model rather than collected from people. primer.ml.pretraining
system prompt Standing instructions sent before the conversation that set a model's role, rules and style. primer.agents.llm
t-SNE An older neighbour-preserving 2-D projection; cluster sizes and gaps in its plots are not meaningful. primer.ml.embeddings.clustering
tabular data Data in rows and columns, where each column is a meaningful quantity in its own units. primer.ml.classical
tanh Squashes any number into the range −1 to 1. primer.ml.cnn_rnn
task arithmetic Combining or removing skills by adding, scaling or subtracting task vectors from a base model's weights. primer.ml.fine_tuning
task budget Hard limits on steps and tokens for one agent task. primer.agents.cost
task vector The change a fine-tune made to a model's weights: fine-tuned minus base. primer.ml.fine_tuning
temperature A knob that sharpens (low) or flattens (high) the probability distribution before sampling a token. primer.ml.inference
tenant One customer organisation on a shared platform. primer.agents.cost
tenant isolation Partitioning storage so a request can only ever reach its own tenant's and user's data. primer.agents.memory
tensor A grid of numbers with any number of dimensions: a vector is 1-D, a matrix 2-D, a batch of images 4-D. primer.notation
tensor parallelism Splitting each matrix multiply across GPUs, which must talk inside every layer. primer.ml.hardware
term frequency How many times a word appears in a document. primer.ml.embeddings.retrieval
test set Held-out data used once at the end to estimate real-world performance honestly. primer.ml.regularization
test-time compute Computation spent while answering rather than while training, such as longer chains or more samples. primer.ml.reasoning
text extraction Pulling a web page's main text out of its HTML, leaving menus and adverts behind. primer.ml.pretraining
text normalization Rewriting text into one standard form (case, punctuation, contractions, how numbers are written) so two texts are compared on their words, not their style.
thinking budget The maximum number of tokens a model may spend reasoning before it must answer. primer.ml.reasoning
threshold calibration Choosing a similarity cut-off by measuring precision and recall on labeled pairs for one specific model. primer.ml.embeddings.similarity
throughput Work finished per second, such as tokens generated per second across all requests; it rises with batch size, while latency is the wait for one request. primer.ml.inference
tiling Loading a block of data into fast memory once and doing all its work before evicting it. primer.ml.hardware
token A piece of text from a model's fixed vocabulary: often a whole common word, or a fragment of a rare one. Models read and write tokens, not words. primer.ml.tokenization
token bucket A rate limiter that allows bursts up to a capacity and refills at a steady rate. primer.agents.deployment
token budget A cap on the tokens, and so the money, one agent run may spend. primer.agents.agent_loop
token healing Backing up over the last prompt token so the model can rewrite it, when a prompt ends partway through what would normally be one token. primer.ml.structured_output
tokenizer The component that splits text into tokens and maps each to an integer ID. primer.ml.tokenization
tool call A structured request from the model to run one of your functions with specific arguments; your code decides whether to run it. primer.agents.tools
tool calling The model outputs a structured request to run a function; your code runs it and sends the result back. The model never executes anything itself. primer.agents.llm
tool poisoning Hiding instructions for the model inside a tool's description. primer.agents.mcp
tool_result The message block that returns a tool's output or error to the model, matched to its request by id. primer.agents.agent_loop
top-k Sampling only from the k most likely next tokens. primer.ml.inference
top-k retrieval accuracy The share of questions for which at least one of the top k retrieved passages contains the answer. primer.ml.metrics
top-p Nucleus sampling: pick only from the smallest set of tokens whose probabilities add up to p. primer.ml.inference
trace A record of one run as a tree of steps (model calls, tool calls, retrievals) with inputs, outputs, timing and cost. primer.agents.observability
trajectory The full record of an agent run: its answer, its tool calls with arguments, and the end state. primer.agents.evals
transfer learning Training a model on a large general task first, then reusing it (usually by fine-tuning) on a smaller task it was never trained for. primer.ml.training_stages
transformer The neural network architecture behind modern language models: stacked blocks of attention followed by a small feed-forward network. primer.ml.transformer
translation equivariance Shift the input and the output shifts the same way: a convolution finds an edge wherever it sits, because the same filter slides over every position. primer.ml.cnn_rnn
transpose Flip a matrix so its rows become columns. Written with a superscript T. primer.notation
triplet loss A loss requiring an anchor to be closer to its positive than to its negative by at least a margin. primer.ml.losses
true-positive rate The share of real positives the model flags; the same as recall. primer.ml.metrics
TTL Time to live: how long a cached entry may be served before it is treated as expired. primer.agents.cost
tubelet A video patch that spans several frames as well as a square of pixels. primer.ml.generative.multimodal
tuned lens A logit lens with a small learned translator per layer. primer.ml.interpretability
two time-scale update rule Giving the generator and discriminator different learning rates so the game converges. primer.ml.generative.gans
two-stage retrieval A cheap, wide first pass to shortlist candidates, then an expensive, precise pass over only those. primer.ml.embeddings.compression
U-Net A convolutional network shaped like a U: it shrinks an image step by step to see the big picture, then grows it back, with shortcuts carrying fine detail across. The classic diffusion denoiser before transformers. primer.ml.generative.diffusion
UMAP A projection to 2-D that keeps each point's neighbours close but distorts other distances. primer.ml.embeddings.clustering
unbiased estimate An estimate that is right on average: any single one may be off, but the errors cancel over many tries. primer.ml.reinforcement
unbiased estimator A way of estimating a number from random data whose average, over every dataset you might have drawn, equals the true value: it is not systematically high or low. primer.ml.benchmarks
underfitting When a model is too simple, or undertrained, to capture the pattern at all. primer.ml.regularization
underflow A number too small for its format, rounded to zero. primer.ml.pretraining
unembedding The final matrix of a language model: it turns the last hidden vector into one score (logit) per vocabulary token. primer.ml.interpretability
unigram tokenizer A subword tokenizer that starts from a huge vocabulary and repeatedly removes the pieces whose loss hurts least. primer.ml.tokenization
unit test A small program that runs one piece of code on chosen inputs and checks the outputs, passing or failing automatically. primer.agents.coding_agents
unit vector A vector rescaled to length 1, so it only carries a direction. primer.notation
UTF-8 The standard way to store text as bytes: 1 byte for basic Latin letters, 2 to 4 bytes for other characters and emoji. primer.ml.tokenization
VAE Variational autoencoder: an autoencoder whose codes are pulled towards a bell curve, so random codes decode to new data. primer.ml.generative.autoencoders
validation set Held-out data used to tune choices like model size and when to stop training. primer.ml.regularization
value In attention, the information a token hands over when others attend to it. primer.ml.attention
value function A prediction, made partway through a task, of how well it will end from here. A verifier that scores a solution after every token is one. primer.ml.reinforcement
value network A second model (the critic) that predicts expected reward, used as PPO's baseline. primer.ml.reinforcement
vanishing gradient When gradients shrink towards zero as they pass back through many layers, so early layers stop learning. primer.ml.deep_nets
variance How widely numbers are spread around their average: the average squared distance from the mean. primer.notation
variational autoencoder An autoencoder whose encoder outputs a fuzzy region (mean and spread) pulled towards the standard normal, so random codes decode to new data. primer.ml.generative.autoencoders
variational inference Approximating a distribution you cannot compute, usually a posterior, with the closest member of a simple family by maximising a lower bound. A VAE does it with a network that outputs the approximation for each example. primer.ml.generative.autoencoders
vector A list of numbers, like (3, 1, 2). In AI, a word, sentence or image is represented as a vector. primer.notation
vector quantization Replacing a vector with the number of its nearest codebook entry. primer.ml.generative.multimodal
velocity field A map giving, at every point and time, which way and how fast a sample should move. primer.ml.generative.diffusion
verifiable reward A reward computed by a check that can't be argued with, such as a correct answer or passing tests. primer.ml.reinforcement
verifier Anything that scores a candidate solution, from a unit test to a learned model. primer.ml.reasoning
Vision Transformer A transformer that treats small image patches as tokens. primer.ml.cnn_rnn
vision-language model A model that reads images (and often video) together with text and writes text; a language model given eyes. primer.ml.generative.multimodal
visual instruction tuning Fine-tuning a vision-language model on images paired with instructions and good answers. primer.ml.generative.multimodal
vocabulary The fixed set of tokens a model knows, typically 32,000 to 200,000 entries. primer.ml.tokenization
voice activity detection Deciding which stretches of a recording contain speech at all, so silence, music and noise are not transcribed.
Voronoi cell All the points closer to one centroid than to any other: the section an IVF index searches. primer.ml.embeddings.ann
VQ-VAE A VAE that snaps each code vector to the nearest entry of a learned codebook, turning images or audio into tokens. primer.ml.generative.autoencoders
warmup Starting training with a tiny learning rate and ramping it up, which keeps the first updates from destabilizing the model. primer.ml.optimizers
Wasserstein distance The least work to reshape one distribution into another (mass moved times distance); also called earth mover's distance. primer.ml.generative.gans
waveform Sound recorded as a list of air-pressure measurements over time. primer.ml.generative.multimodal
weak supervision Training on labels that are plentiful but noisy or imperfect, such as captions and transcripts found on the web, instead of a small set checked by experts.
weight A learned number that says how strongly one input influences an output. Training adjusts the weights. primer.ml.neural_net
weight averaging Merging fine-tunes of the same base by averaging their weights (task arithmetic with λ = 1/T). primer.ml.fine_tuning
weight clipping Forcing every weight of a network back into a small range, such as −0.01 to 0.01, after each update; the original Wasserstein GAN's crude way to cap its critic's slope. primer.ml.generative.gans
weight decay Shrinking every weight slightly at each step, which penalizes large weights and keeps the model smoother. primer.ml.regularization
weight sharing Reusing the same filter at every position, so the parameter count doesn't grow with image size. primer.ml.cnn_rnn
weight space The space of every possible setting of a model's weights: one axis per weight, one point per model. Fine-tuning moves a model from one point to another. primer.ml.fine_tuning
weight tying Reusing the token embedding table as the output layer, so the vector that reads a token in also scores it on the way out. primer.ml.transformer
whitening Rescaling vectors so every direction has equal spread and no two directions are correlated. primer.ml.embeddings.similarity
Wiener process Continuous random jitter, also called Brownian motion: over any short time dt it moves by a fresh bell-curve amount with variance dt, independent of everything before.
win rate The share of head-to-head comparisons one model wins; 50% means the two are indistinguishable. primer.agents.evals
winner's curse The best of many versions chosen on a test looks better on it than it really is. primer.ml.benchmarks
word error rate The share of a reference transcript's words that a system gets wrong: the substitutions, deletions and insertions needed to turn its output into the reference, divided by the reference's word count.
word2vec A 2013 method that learns one vector per word by predicting nearby words. primer.ml.embeddings.word2vec
WordPiece BERT's subword tokenizer, which merges the pair whose parts most rarely appear apart rather than simply the most frequent pair. primer.ml.tokenization
workflow A fixed sequence of steps written in code, with a language model doing a task inside each step. primer.agents.orchestration
Xavier initialization Starting weights with variance 2 / (fan-in + fan-out), suited to tanh and sigmoid layers. primer.ml.deep_nets
YaRN A refinement of RoPE context extension that rescales each frequency band differently, used by many long-context models. primer.ml.positional
zero Sharding optimizer state, then gradients, then weights across data-parallel GPUs. primer.ml.pretraining
zero-order hold A discretization rule that assumes the input stays constant for the whole step; it gives the keep factor e^(ΔA) of a state-space model. primer.ml.efficient_architectures
zero-shot Doing a task with no task-specific training examples. primer.ml.embeddings.contrastive
zero-shot classification Labeling items by comparing them to a text description of each label, with no training on those labels. primer.ml.embeddings.contrastive
β-VAE A VAE whose KL penalty is weighted by β, trading rebuild sharpness for a smoother, more organized code space. primer.ml.generative.autoencoders
on GitHub
   1"""
   2# Glossary
   3
   4Every term this primer uses, in plain English, with a link to the lesson
   5that builds it. On the HTML site, these definitions also appear when you hover
   6over (or tap) the term anywhere in a lesson or an annotated paper.
   7
   8This module is the single source of truth: `make docs` turns it into
   9`glossary.js` for the browser. The table below is generated from `GLOSSARY`.
  10"""
  11
  12from __future__ import annotations
  13
  14import functools
  15import json
  16from collections import Counter
  17from dataclasses import dataclass
  18
  19
  20@dataclass(frozen=True)
  21class Entry:
  22    definition: str
  23    lesson: str | None = None  # dotted module that teaches it, e.g. "primer.ml.attention"
  24    # Everyday words with a technical meaning ("value", "policy", "patch") are only
  25    # given hover definitions on pages under these path prefixes, so a "travel
  26    # policy" never pops up a definition from preference tuning.
  27    scope: tuple[str, ...] = ()
  28    # How readers see the term when counting the lessons' usage can't decide it (see display_term).
  29    display: str | None = None
  30
  31
  32_E = Entry
  33N = "primer.notation"
  34NN, OPT, DEEP = "primer.ml.neural_net", "primer.ml.optimizers", "primer.ml.deep_nets"
  35ATT, POS, TF, TOK = "primer.ml.attention", "primer.ml.positional", "primer.ml.transformer", "primer.ml.tokenization"
  36TRAIN, INF, LOSS, MET = "primer.ml.training_stages", "primer.ml.inference", "primer.ml.losses", "primer.ml.metrics"
  37REG, CNN, BIG = "primer.ml.regularization", "primer.ml.cnn_rnn", "primer.ml.big_picture"
  38W2V, SIM, CON = "primer.ml.embeddings.word2vec", "primer.ml.embeddings.similarity", "primer.ml.embeddings.contrastive"
  39CMP, ANN, RET = "primer.ml.embeddings.compression", "primer.ml.embeddings.ann", "primer.ml.embeddings.retrieval"
  40CLU, OPS = "primer.ml.embeddings.clustering", "primer.ml.embeddings.operations"
  41LLM, ORC, LOOP, TOOLS, MCP = "primer.agents.llm", "primer.agents.orchestration", "primer.agents.agent_loop", "primer.agents.tools", "primer.agents.mcp"
  42RAG, CTX, MEM, PLAN = "primer.agents.rag", "primer.agents.context", "primer.agents.memory", "primer.agents.planning"
  43EVAL, GUARD, COST, OBS, DEP = "primer.agents.evals", "primer.agents.guardrails", "primer.agents.cost", "primer.agents.observability", "primer.agents.deployment"
  44
  45GLOSSARY: dict[str, Entry] = {
  46    # --- notation and maths -------------------------------------------------
  47    "vector": _E("A list of numbers, like (3, 1, 2). In AI, a word, sentence or image is represented as a vector.", N),
  48    "matrix": _E("A table of numbers with rows and columns. A batch of vectors stacked as rows is a matrix, and most model weights are matrices.", N),
  49    "tensor": _E("A grid of numbers with any number of dimensions: a vector is 1-D, a matrix 2-D, a batch of images 4-D.", N),
  50    "dot product": _E("Multiply two lists position by position and add up the products. It is large when two vectors point the same way.", N),
  51    "matrix multiply": _E("A grid of dot products: each output cell is one row of the first matrix dotted with one column of the second.", N),
  52    "transpose": _E("Flip a matrix so its rows become columns. Written with a superscript T.", N),
  53    "norm": _E("The length of a vector: square the entries, add them, take the square root.", N),
  54    "unit vector": _E("A vector rescaled to length 1, so it only carries a direction.", N),
  55    "orthogonal": _E("At right angles: two vectors whose dot product, and so whose cosine similarity, is zero, so moving along one does not move you along the other.", SIM),
  56    "logarithm": _E("The undo button for exponentiation: ln(y) asks what power of e gives y. Logs turn multiplication into addition.", N),
  57    "variance": _E("How widely numbers are spread around their average: the average squared distance from the mean.", N),
  58    "standard deviation": _E("The square root of the variance: the typical distance of a value from the average.", N),
  59    "derivative": _E("The slope of a function at a point: how much the output changes per tiny nudge of the input.", N),
  60    "gradient": _E("One slope per input, collected into a vector. It points in the direction that increases the function fastest.", N),
  61    "chain rule": _E("To get the slope through a chain of functions, multiply the slopes of each link.", N),
  62    "argmax": _E("The position of the largest value in a list, rather than the value itself.", N),
  63    "big-o": _E("A way to say how cost grows with input size, ignoring constant factors. O(n²) means doubling n quadruples the cost.", N),
  64    "probability distribution": _E("A list of probabilities over all possible outcomes, each between 0 and 1, adding up to 1.", N),
  65    "probability density": _E("How thickly a continuous distribution's samples cover each spot: high where they crowd, zero where none ever land. It is the height of the bump, and its area adds up to 1.", 'primer.ml.generative.gans'),
  66
  67    # --- neural network basics ----------------------------------------------
  68    "neuron": _E("Multiply each input by a weight, add them up with a bias, and pass the result through a nonlinear function.", NN),
  69    "weight": _E("A learned number that says how strongly one input influences an output. Training adjusts the weights.", NN),
  70    'bias': _E("A learned number added to a neuron's weighted sum, letting it shift its output up or down.", NN, scope=('primer/ml/neural_net', 'primer/ml/deep_nets')),
  71    "parameters": _E("All the learned numbers in a model (its weights and biases). A 70B model has 70 billion of them.", NN),
  72    "activation function": _E("The nonlinear function inside a neuron, such as ReLU or GELU. Without it, stacked layers collapse into one.", NN),
  73    "relu": _E("An activation function that keeps positive numbers and turns negatives into zero.", NN),
  74    "gelu": _E("A smooth version of ReLU used in most transformers.", NN),
  75    "sigmoid": _E("Squashes any number into the range 0 to 1, with an S-shaped curve.", NN),
  76    "softmax": _E("Turns a list of scores into shares that are all positive and add up to 1, with bigger scores getting disproportionately more.", ATT),
  77    "forward pass": _E("Running inputs through the network, layer by layer, to get a prediction.", NN),
  78    "backpropagation": _E("The chain rule run backwards through the network to find, for every weight, how much nudging it would change the loss.", NN),
  79    "multilayer perceptron": _E("A network of fully connected layers, each a matrix multiply followed by a nonlinearity; also called an MLP.", NN),
  80    "loss": _E("A single number measuring how wrong the model's predictions are. Training tries to make it smaller.", LOSS),
  81    "loss function": _E("The rule that turns predictions and correct answers into the loss. Choosing it defines what the model learns.", LOSS),
  82    "gradient descent": _E("Training by repeatedly taking a small step downhill: nudge every weight against its gradient to lower the loss.", OPT),
  83    "gradient ascent": _E("Nudging every weight along its gradient to make a quantity larger: the mirror image of gradient descent. Used to maximise a reward, or, run on a loss, to make a model worse at a task on purpose.", "primer.ml.reinforcement"),
  84    "learning rate": _E("How big a step each training update takes. Too high and training blows up; too low and it crawls.", OPT),
  85    "optimizer": _E("The rule that turns gradients into weight updates, such as SGD, momentum or Adam.", OPT),
  86    "adam": _E("An optimizer that adapts the step size for each weight using running averages of its gradients.", OPT),
  87    "adamw": _E("Adam with weight decay applied separately from the gradient step. The default optimizer for transformers.", OPT),
  88    "momentum": _E("Keeping a running average of recent gradients so updates roll smoothly through noise, like a ball gathering speed downhill.", OPT),
  89    "batch": _E("The group of examples processed before one weight update.", NN),
  90    "epoch": _E("One full pass over the training data.", NN),
  91    "warmup": _E("Starting training with a tiny learning rate and ramping it up, which keeps the first updates from destabilizing the model.", OPT),
  92    "vanishing gradient": _E("When gradients shrink towards zero as they pass back through many layers, so early layers stop learning.", DEEP),
  93    "exploding gradient": _E("When gradients grow huge as they pass back through many layers, so training diverges.", DEEP),
  94    "residual connection": _E("Adding a layer's input back to its output (x + f(x)), giving gradients a shortcut through deep networks.", DEEP),
  95    "layer normalization": _E("Rescaling each token's numbers to a steady average and spread, which keeps training stable.", DEEP),
  96    "batch normalization": _E("Rescaling each feature using statistics from the current batch of examples. Common in image models.", DEEP),
  97    "gradient clipping": _E("Capping the size of the gradient so one bad batch can't throw the weights far off course.", DEEP),
  98
  99    # --- transformers ---------------------------------------------------------
 100    "attention": _E("The mechanism that lets each token look at every other token, score how relevant each one is, and blend in information from the relevant ones.", ATT),
 101    "self-attention": _E("Attention where a sequence attends to itself: every token scores every other token in the same text.", ATT),
 102    'query': _E("In attention, the vector a token uses to ask what it is looking for.", ATT, scope=('primer/ml/attention', 'primer/ml/transformer')),
 103    'key': _E("In attention, the vector a token offers so others can decide whether it is relevant to them.", ATT, scope=('primer/ml/attention', 'primer/ml/transformer', 'primer/ml/inference', 'primer/ml/big_picture')),
 104    'value': _E("In attention, the information a token hands over when others attend to it.", ATT, scope=('primer/ml/attention', 'primer/ml/transformer', 'primer/ml/inference', 'primer/ml/big_picture')),
 105    "multi-head attention": _E("Several attention computations run in parallel on slices of the vectors, so each head can track a different kind of relationship.", ATT),
 106    "grouped-query attention": _E("Several query heads share one set of keys and values, shrinking the memory needed during generation.", ATT),
 107    "causal mask": _E("Hides future tokens from each token during attention, so a model predicting the next word can't peek at it.", ATT),
 108    "transformer": _E("The neural network architecture behind modern language models: stacked blocks of attention followed by a small feed-forward network.", TF),
 109    "feed-forward network": _E("The small two-layer network in each transformer block that processes every token on its own after attention has mixed them.", TF),
 110    "encoder": _E("A transformer that reads the whole input at once, with every token seeing every other. Used for embeddings and classification.", TF),
 111    "decoder": _E("A transformer that generates text left to right, each token seeing only the ones before it. GPT, Claude and Llama are decoders.", TF),
 112    "mixture of experts": _E("Replacing one feed-forward network with many expert networks and a router that sends each token to a few of them.", TF),
 113    "positional encoding": _E("Information added to token vectors so the model knows word order, which attention alone ignores.", POS),
 114    "rope": _E("Rotary position embedding: encodes position by rotating query and key vectors, so attention depends on the distance between tokens.", POS),
 115    "context window": _E("The maximum number of tokens a model can take in at once.", INF),
 116    "rmsnorm": _E("A cheaper layer normalization that divides each token's numbers by their root-mean-square without subtracting the mean first.", TF),
 117    "pre-norm": _E("Normalizing before each sub-layer of a transformer block, leaving the residual path untouched, which trains more stably than normalizing after.", TF),
 118    "residual stream": _E("The running vector for each token that every transformer block reads from and adds a correction to, instead of replacing it.", TF),
 119    "weight tying": _E("Reusing the token embedding table as the output layer, so the vector that reads a token in also scores it on the way out.", TF),
 120    "encoder-decoder": _E("A transformer where an encoder reads the whole input and a decoder writes the output one token at a time while attending to the encoder.", TF),
 121    "load-balancing loss": _E("A small extra training penalty that pushes a mixture-of-experts router to spread tokens evenly across experts.", TF),
 122    "flops": _E("Floating-point operations: individual multiplies or adds, used to measure compute (about 2 per parameter per token to generate, 6 to train).", TF),
 123    "permutation equivariance": _E("Shuffling a layer's inputs just shuffles its outputs the same way, which is why attention alone cannot see word order.", POS),
 124    "sinusoidal positional encoding": _E("The original transformer's fixed position codes, made of sines and cosines at many frequencies, added to each token's embedding.", POS),
 125    "position interpolation": _E("Stretching a model to longer contexts by scaling every position down into the range it saw during training.", POS),
 126    "ntk-aware scaling": _E("Extending a RoPE model's context by raising the rotation base, which slows the slow frequencies while keeping the fast ones that encode local order.", POS),
 127    "yarn": _E("A refinement of RoPE context extension that rescales each frequency band differently, used by many long-context models.", POS),
 128    "alibi": _E("A position scheme that adds no position vectors and instead subtracts a penalty proportional to distance from each attention score.", POS),
 129    "autoregressive generation": _E("Producing text one token at a time, where each new token is predicted from all the tokens before it and then appended.", BIG),
 130    "greedy decoding": _E("Always picking the single most probable next token, the same as sampling at temperature 0.", BIG),
 131    "logits": _E("The raw scores a model outputs for every possible next token, before softmax turns them into probabilities.", BIG),
 132
 133    # --- tokens ---------------------------------------------------------------
 134    "token": _E("A piece of text from a model's fixed vocabulary: often a whole common word, or a fragment of a rare one. Models read and write tokens, not words.", TOK),
 135    "tokenizer": _E("The component that splits text into tokens and maps each to an integer ID.", TOK),
 136    "special token": _E("A token reserved for structure rather than text, such as a turn marker or an end-of-text marker. The model learns when to emit it, and generation stops at the end marker.", "primer.ml.fine_tuning"),
 137    "vocabulary": _E("The fixed set of tokens a model knows, typically 32,000 to 200,000 entries.", TOK),
 138    "byte pair encoding": _E("A tokenizer-building method that starts from characters and repeatedly merges the most frequent adjacent pair into a new token.", TOK),
 139    "byte-level bpe": _E("Byte pair encoding run on the raw bytes of UTF-8 text, so any string can be tokenized and there is never an unknown token.", TOK),
 140    "pre-tokenization": _E("Splitting text into chunks, such as words with their leading space, before BPE runs, so merges never cross chunk boundaries.", TOK),
 141    "wordpiece": _E("BERT's subword tokenizer, which merges the pair whose parts most rarely appear apart rather than simply the most frequent pair.", TOK),
 142    "unigram tokenizer": _E("A subword tokenizer that starts from a huge vocabulary and repeatedly removes the pieces whose loss hurts least.", TOK),
 143    "sentencepiece": _E("A tokenizer library that runs BPE or Unigram directly on raw text, writing spaces as the visible symbol ▁.", TOK),
 144    "utf-8": _E("The standard way to store text as bytes: 1 byte for basic Latin letters, 2 to 4 bytes for other characters and emoji.", TOK),
 145    "bpe": _E("Byte pair encoding: build a vocabulary by repeatedly merging the most frequent adjacent pair of symbols.", TOK),
 146
 147    # --- training and inference -----------------------------------------------
 148    "pretraining": _E("The first, most expensive training stage: predicting the next token over trillions of tokens of text.", TRAIN),
 149    "base model": _E("A model after pretraining only: knowledgeable, but it continues text rather than following instructions.", TRAIN),
 150    "fine-tuning": _E("Training an existing model further on a smaller, targeted dataset to change its behavior.", TRAIN),
 151    "weight space": _E("The space of every possible setting of a model's weights: one axis per weight, one point per model. Fine-tuning moves a model from one point to another.", "primer.ml.fine_tuning"),
 152    "sft": _E("Supervised fine-tuning: training on examples of instructions paired with good responses.", TRAIN),
 153    "rlhf": _E("Reinforcement learning from human feedback: a reward model learns human preferences, and the language model is tuned to score well on it.", TRAIN),
 154    "dpo": _E("Direct preference optimization: tuning a model straight from pairs of preferred and rejected answers, without a separate reward model.", TRAIN),
 155    "reward model": _E("A model trained to predict which of two responses a human would prefer.", TRAIN),
 156    "lora": _E("Low-rank adaptation: freeze the model and train a small pair of matrices whose product is added to a weight matrix.", TRAIN),
 157    "distillation": _E("Training a small model to imitate a large one's outputs, often the biggest cost saving in production.", TRAIN),
 158    "inference": _E("Using a trained model to make predictions, as opposed to training it.", INF),
 159    "prefill": _E("The first phase of generation: the whole prompt is processed in one parallel pass. It sets the time to the first token.", INF),
 160    'decode': _E("The second phase of generation: output tokens are produced one at a time. It sets the tokens per second.", INF, scope=('primer/ml/inference',)),
 161    "kv cache": _E("Stored keys and values from earlier tokens, so each new token is computed without redoing work. It trades GPU memory for speed.", INF),
 162    "temperature": _E("A knob that sharpens (low) or flattens (high) the probability distribution before sampling a token.", INF),
 163    "top-p": _E("Nucleus sampling: pick only from the smallest set of tokens whose probabilities add up to p.", INF),
 164    "top-k": _E("Sampling only from the k most likely next tokens.", INF),
 165    "quantization": _E("Storing numbers with fewer bits (say 8 or 4 instead of 16 or 32), which shrinks memory with a small loss of precision.", INF),
 166    "speculative decoding": _E("A small fast model drafts several tokens, and the big model checks them all in one pass, keeping the ones it agrees with.", INF),
 167    "prompt caching": _E("Reusing the processed form of a prompt prefix that repeats across requests, which cuts cost and time to first token.", INF),
 168    "cross-entropy": _E("The standard loss for predicting categories: minus the log of the probability given to the right answer.", LOSS),
 169    "perplexity": _E("e raised to the average cross-entropy: roughly how many options the model is torn between at each step.", LOSS),
 170    "contrastive loss": _E("A loss that pulls matching pairs together in vector space and pushes non-matching pairs apart.", LOSS),
 171    "infonce": _E("A contrastive loss that treats the other items in the batch as wrong answers for each matching pair.", LOSS),
 172    "overfitting": _E("When a model memorizes its training data, noise included, and does worse on new data.", REG),
 173    "underfitting": _E("When a model is too simple, or undertrained, to capture the pattern at all.", REG),
 174    "regularization": _E("Any technique that discourages memorizing, such as dropout, weight decay or early stopping.", REG),
 175    "dropout": _E("Randomly switching off neurons during training so the network can't rely on any single path.", REG),
 176    "weight decay": _E("Shrinking every weight slightly at each step, which penalizes large weights and keeps the model smoother.", REG),
 177    "early stopping": _E("Stopping training when the score on held-out data stops improving.", REG),
 178    "data leakage": _E("When information from the test set, or from the future, sneaks into training and inflates offline scores.", REG),
 179    "out-of-distribution": _E("Unlike the examples a model learned from or was shown, such as longer or rarer inputs. Performance there is usually lower and harder to predict."),
 180    "emergent ability": _E("A skill that is near chance in smaller models and appears sharply once a model is large enough, so it can't be predicted by extending the small models' trend."),
 181    "convolution": _E("Sliding a small grid of learned weights across an image and computing a weighted sum at every position, to find where a pattern appears.", CNN),
 182    "pooling": _E("Shrinking a feature map by keeping only the strongest (or average) value in each small window.", CNN),
 183    "rnn": _E("A recurrent neural network: it reads a sequence one step at a time, carrying a running summary called the hidden state.", CNN),
 184    "lstm": _E("An RNN with gates that decide what to forget, what to write and what to output, so it can remember across long sequences.", CNN),
 185
 186    # --- metrics --------------------------------------------------------------
 187    "precision": _E("Of the items flagged, the fraction that were right.", MET),
 188    'recall': _E("Of the items that truly mattered, the fraction that were found.", MET, scope=('primer/ml/metrics', 'primer/ml/embeddings/', 'primer/agents/rag', 'primer/agents/evals')),
 189    "f1": _E("The harmonic mean of precision and recall: high only when both are high.", MET),
 190    "roc-auc": _E("The probability that the model ranks a random positive above a random negative.", MET),
 191    "recall@k": _E("The fraction of relevant documents that appear in the top k results. For RAG, usually the metric that matters most.", MET),
 192    "mrr": _E("Mean reciprocal rank: the average of 1 / (position of the first relevant result).", MET),
 193    "ndcg": _E("A ranking score that gives more credit for relevant results near the top and handles degrees of relevance.", MET),
 194    "bleu": _E("A score that counts overlapping word sequences with a reference text. Cheap, but blind to paraphrase.", MET),
 195    "ablation": _E("Removing or replacing one part of a method and measuring again, to find out which part the gains come from."),
 196    "inter-annotator agreement": _E("How often two people labelling the same items independently give the same label. It caps how finely any evaluation built on those labels can tell two models apart.", "primer.ml.fine_tuning"),
 197    "likert scale": _E("A rating on a short fixed ladder of labelled points, such as 1 (not helpful) to 6 (highly helpful), averaged over many items to compare systems.", EVAL),
 198    'calibration': _E("How well a model's confidence matches how often it is right: a calibrated model is right about 80% of the time when it says 80%.", LOSS, scope=('primer/ml/losses', 'primer/ml/alignment')),
 199    # --- embeddings -----------------------------------------------------------
 200    "embedding": _E("A learned vector for a piece of content, arranged so that similar meanings end up close together.", W2V),
 201    "cosine similarity": _E("How closely two vectors point in the same direction, from −1 (opposite) to 1 (identical), ignoring their lengths.", SIM),
 202    "euclidean distance": _E("The straight-line distance between two vectors.", SIM),
 203    "anisotropy": _E("When a model's vectors all crowd into a narrow cone, so even unrelated texts score as fairly similar.", SIM),
 204    "curse of dimensionality": _E("In very high dimensions, random points are all nearly the same distance apart, which makes exact nearest-neighbor search slow and fragile.", SIM),
 205    "hard negative": _E("A training example that looks relevant but isn't, such as the right topic with the wrong answer. Training on them teaches fine distinctions.", CON),
 206    "bi-encoder": _E("Embeds the query and each document separately, so document vectors can be computed once and searched fast.", RET),
 207    "cross-encoder": _E("Reads the query and one document together and outputs a relevance score. Accurate but slow, so it is used to rerank a shortlist.", RET),
 208    "reranker": _E("A second, more accurate model that reorders the top results of a fast first-stage search.", RET),
 209    "clip": _E("A model that trains an image encoder and a text encoder together so pictures and their captions land near each other in one vector space.", CON, scope=("primer/ml/embeddings/contrastive", "primer/ml/generative", "primer/ml/losses"), display="CLIP"),
 210    "matryoshka embedding": _E("An embedding trained so its first few dimensions work as a smaller embedding on their own.", CMP),
 211    "approximate nearest neighbor": _E("Finding vectors close to a query quickly by searching only part of the collection, accepting a small chance of missing the true closest.", ANN),
 212    "ann": _E("Approximate nearest neighbor search: trade a little recall for a lot of speed.", ANN),
 213    "hnsw": _E("A layered graph index for vector search: long jumps on sparse upper layers, then a careful local search at the bottom.", ANN),
 214    "ivf": _E("A vector index that clusters vectors ahead of time and searches only the clusters nearest the query.", ANN),
 215    "product quantization": _E("Compressing a vector by splitting it into chunks and replacing each chunk with the ID of its nearest entry in a small codebook.", ANN),
 216    "bm25": _E("The classic keyword-search score: rewards documents containing the query's rarer words, with diminishing returns for repeats.", RET),
 217    "hybrid search": _E("Running keyword search and vector search together and merging the two ranked lists.", RET),
 218    "reciprocal rank fusion": _E("Merging ranked lists by giving each document 1 / (k + rank) from every list it appears in, then adding those up.", RET),
 219    "chunking": _E("Splitting documents into passages before embedding them, so retrieval can return just the relevant part.", RET),
 220    "colbert": _E("A retrieval model that keeps one vector per token and matches each query token to its best document token.", RET),
 221    "k-means": _E("Clustering by repeatedly assigning each point to its nearest center and moving each center to the average of its points.", CLU),
 222    "semantic cache": _E("Reusing a stored answer when a new question's embedding is nearly identical to a previous question's.", CLU),
 223
 224    # --- applied AI -----------------------------------------------------------
 225    "llm": _E("Large language model: a transformer trained to predict the next token, then tuned to follow instructions.", LLM),
 226    "prompt": _E("The text sent to a language model: instructions, context and the question.", LLM),
 227    "system prompt": _E("Standing instructions sent before the conversation that set a model's role, rules and style.", LLM),
 228    "tool calling": _E("The model outputs a structured request to run a function; your code runs it and sends the result back. The model never executes anything itself.", LLM),
 229    "agent": _E("A system where a language model decides which tools to call, looks at the results and decides the next step, in a loop until done.", LOOP),
 230    "agent loop": _E("The cycle an agent repeats: think, call a tool, observe the result, decide what next.", LOOP),
 231    "react": _E("Reason plus act: the agent pattern that interleaves reasoning steps with tool calls.", LOOP),
 232    "workflow": _E("A fixed sequence of steps written in code, with a language model doing a task inside each step.", ORC),
 233    "router": _E("A step that classifies a request and sends it down the right path, or to the right model.", ORC),
 234    "multi-agent system": _E("Several agents that coordinate, typically a supervisor delegating to specialist workers.", ORC),
 235    "json schema": _E("A standard way to describe the shape of JSON data: which fields exist, their types and which are required.", TOOLS),
 236    "idempotent": _E("Safe to repeat: doing it twice has the same effect as doing it once, so retries can't create duplicates.", TOOLS),
 237    "mcp": _E("Model Context Protocol: an open standard for connecting AI apps to tools and data through reusable servers.", MCP),
 238    "rag": _E("Retrieval-augmented generation: fetch relevant passages first, then have the model answer using them, with citations.", RAG),
 239    "hyde": _E("Hypothetical document embeddings: have the model write a plausible answer, then search with that, since answers resemble documents more than questions do.", RAG),
 240    "grounding": _E("Tying an answer to specific source passages so every claim can be checked.", RAG),
 241    "hallucination": _E("A confident answer that isn't supported by the sources or the facts.", RAG),
 242    "context engineering": _E("Deciding exactly what goes into the model's context window on each call: instructions, facts, tool results, memory.", CTX),
 243    "context rot": _E("Quality dropping as a long session fills the context with stale or irrelevant material.", CTX),
 244    "episodic memory": _E("Memory of what happened, such as 'last week the user rejected this vendor'.", MEM),
 245    "semantic memory": _E("Memory of facts, such as 'the user's fiscal year starts in April'.", MEM),
 246    "procedural memory": _E("Memory of how to do things, such as a learned workflow or saved skill.", MEM),
 247    "plan-and-execute": _E("An agent pattern that writes a plan first, then executes it step by step, replanning when results surprise it.", PLAN),
 248    "reflection": _E("Having a model review and revise its own output. Useful, but external checks such as tests are more reliable.", PLAN),
 249    "compounding error": _E("Small per-step failure rates multiplying over many steps: ten steps at 95% succeed only about 60% of the time.", PLAN),
 250    "eval": _E("A repeatable test of a model or agent's behavior on a fixed set of tasks with known good outcomes.", EVAL),
 251    "golden set": _E("A curated set of real tasks with expected outcomes, run on every change to catch regressions.", EVAL),
 252    "llm-as-judge": _E("Using a language model with a rubric to grade open-ended outputs, checked against human ratings.", EVAL),
 253    "guardrail": _E("A check around the model that screens inputs, validates outputs or limits actions.", GUARD),
 254    "prompt injection": _E("Hostile instructions hidden in content the model reads, such as an email or web page, trying to hijack it.", GUARD),
 255    "pii": _E("Personally identifiable information: names, emails, phone numbers, card numbers and similar.", GUARD),
 256    "model routing": _E("Sending each request to the cheapest model that can handle it well.", COST),
 257    'trace': _E("A record of one run as a tree of steps (model calls, tool calls, retrievals) with inputs, outputs, timing and cost.", OBS, scope=('primer/agents/observability', 'primer/agents/evals')),
 258    'span': _E("One step inside a trace, with its start time, duration and details.", OBS, scope=('primer/agents/observability',)),
 259    "shadow mode": _E("Running a new system on real inputs and recording what it would do, without letting it act.", DEP),
 260    "canary release": _E("Sending a small share of traffic to a new version first and watching its metrics before rolling out further.", DEP),
 261    "kill switch": _E("A way to stop an agent instantly, per customer or globally.", DEP),
 262    "audit trail": _E("A tamper-evident log of which agent did what, for whom and with what authority.", DEP),
 263    # --- added from lesson reports -----------------------------------------
 264    'accuracy': _E('The share of all predictions that were correct. Misleading when one class is rare.', 'primer.ml.metrics'),
 265    'agent-computer interface': _E('The commands an agent can call and the exact form of what comes back to it, designed for a language model the way a code editor is designed for a person.', 'primer.agents.coding_agents'),
 266    'confusion matrix': _E('The four counts behind every classification metric: hits, false alarms, misses and correct passes.', 'primer.ml.metrics'),
 267    'container': _E("An isolated copy of an operating system's user space for one program: its own files, processes and network settings, started from an image and thrown away afterwards.", 'primer.agents.coding_agents', scope=('primer/agents/coding_agents',)),
 268    'diff': _E('A text listing, file by file, which lines to remove (marked −) and which to add (marked +), with a few unchanged lines around each change so a tool can find the spot. Saved to a file, it is a patch that can be applied to another copy of the code.', 'primer.agents.coding_agents'),
 269    'fault localization': _E('Finding which files, functions and lines cause a bug, before trying to fix it.', 'primer.agents.coding_agents'),
 270    'linter': _E('A program that reads code without running it and reports likely mistakes, such as a syntax error, bad indentation or a name that is never defined.', None),
 271    'percentage point': _E('The plain difference between two percentages: going from 11% to 18% is a rise of 7 percentage points, which is a 64% relative rise.', None),
 272    'pull request': _E('A proposed set of changes to a code repository, submitted for review; once merged, the changes become part of the project.', 'primer.agents.coding_agents'),
 273    'roc curve': _E('True-positive rate plotted against false-positive rate as the decision threshold sweeps.', 'primer.ml.metrics'),
 274    'true-positive rate': _E('The share of real positives the model flags; the same as recall.', 'primer.ml.metrics'),
 275    'false-positive rate': _E('The share of real negatives the model wrongly flags.', 'primer.ml.metrics'),
 276    'precision-recall curve': _E('Precision plotted against recall across thresholds; more honest than ROC for rare events.', 'primer.ml.metrics'),
 277    'precision@k': _E('The share of the top k search results that are relevant.', 'primer.ml.metrics'),
 278    'rouge-l': _E('An overlap score based on the longest in-order sequence of words shared with a reference.', 'primer.ml.metrics'),
 279    'bertscore': _E('Compares generated and reference text by embedding similarity instead of exact words.', 'primer.ml.metrics'),
 280    "cohen's kappa": _E('Agreement between two raters after subtracting the agreement expected by chance.', 'primer.ml.metrics'),
 281    'loss mask': _E('A per-position switch deciding which predictions count toward the loss, such as only the reply in SFT.', 'primer.ml.training_stages'),
 282    'preference tuning': _E('Training on which of two responses people preferred, to shape tone, helpfulness and safety.', 'primer.ml.training_stages'),
 283    'bradley-terry model': _E('The rule that the probability A beats B is the sigmoid of their score difference.', 'primer.ml.training_stages'),
 284    'reference model': _E('A frozen copy of the starting model that DPO and RLHF measure drift against.', 'primer.ml.training_stages'),
 285    'policy': _E('In reinforcement learning and preference tuning, the model being trained, viewed as a probability distribution over actions or responses.', 'primer.ml.reinforcement', scope=('primer/ml/training_stages', 'primer/ml/reinforcement', 'primer/ml/reasoning', 'primer/ml/alignment')),
 286    'implicit reward': _E('In DPO, β times the log of how much more likely training has made a response than the reference model did.', 'primer.ml.training_stages'),
 287    'log-probability': _E("The logarithm of a probability; for a whole response, the sum of its tokens' log-probabilities.", 'primer.ml.training_stages'),
 288    'qlora': _E('LoRA adapters trained on top of base weights stored in 4 bits.', 'primer.ml.training_stages'),
 289    'adapter': _E('A small trainable module added beside frozen weights to specialise a model.', 'primer.ml.training_stages', scope=('primer/ml/training_stages',)),
 290    'full fine-tune': _E('Updating every weight of a model. Rarely worth the cost.', 'primer.ml.training_stages'),
 291    'soft targets': _E("A teacher model's temperature-softened probabilities, used to train a student model.", 'primer.ml.training_stages'),
 292    'kl divergence': _E('The extra surprise from using distribution q when the truth is p; zero only when they match.', 'primer.ml.training_stages'),
 293    'dark knowledge': _E('How a teacher model ranks the wrong answers, revealed by its soft targets.', 'primer.ml.training_stages'),
 294    'arithmetic intensity': _E('Operations performed per byte read from memory.', 'primer.ml.inference'),
 295    'roofline': _E('A chart of the speed a chip can reach at each arithmetic intensity, capped first by memory and then by compute.', 'primer.ml.inference'),
 296    'memory-bound': _E('Limited by how fast data arrives from memory rather than by arithmetic, as in decoding.', 'primer.ml.inference'),
 297    'compute-bound': _E('Limited by arithmetic throughput, as in prefill.', 'primer.ml.inference'),
 298    'draft model': _E('The small, fast model that proposes tokens in speculative decoding.', 'primer.ml.inference'),
 299    'acceptance rate': _E('How often the big model keeps a drafted token in speculative decoding.', 'primer.ml.inference'),
 300    'int8 quantization': _E('Storing each weight as an 8-bit integer times a shared scale.', 'primer.ml.inference'),
 301    'int4 quantization': _E('Storing each weight as a 4-bit integer times a shared scale.', 'primer.ml.inference'),
 302    'per-channel quantization': _E("One scale per weight row, so a single outlier doesn't coarsen all the others.", 'primer.ml.inference'),
 303    'static batching': _E('Serving a fixed group of requests until the longest one finishes.', 'primer.ml.inference'),
 304    'continuous batching': _E("Refilling a finished request's slot in the batch immediately with the next request.", 'primer.ml.inference'),
 305    'pagedattention': _E('Storing the KV cache in fixed-size pages, like virtual memory, to avoid wasted GPU memory.', 'primer.ml.inference'),
 306    'filter': _E('In a CNN, the small grid of learned weights a convolution slides over an image; also called a kernel.', 'primer.ml.cnn_rnn', scope=('primer/ml/cnn_rnn',)),
 307    'kernel': _E('In a CNN, the small grid of learned weights a convolution slides over an image; also called a filter.', 'primer.ml.cnn_rnn', scope=('primer/ml/cnn_rnn',)),
 308    'stride': _E('How many pixels a convolution filter jumps between positions.', 'primer.ml.cnn_rnn', scope=('primer/ml/cnn_rnn',)),
 309    'padding': _E('Zeros added around an image so a filter can centre on the edge pixels.', 'primer.ml.cnn_rnn', scope=('primer/ml/cnn_rnn',)),
 310    'max pooling': _E('Keeping only the largest value in each small block of a feature map.', 'primer.ml.cnn_rnn'),
 311    'weight sharing': _E("Reusing the same filter at every position, so the parameter count doesn't grow with image size.", 'primer.ml.cnn_rnn'),
 312    'receptive field': _E('How much of the original input one neuron can see.', 'primer.ml.cnn_rnn'),
 313    'vision transformer': _E('A transformer that treats small image patches as tokens.', 'primer.ml.cnn_rnn'),
 314    'patch': _E('A small square cut from an image and flattened into one token.', 'primer.ml.cnn_rnn', scope=('primer/ml/cnn_rnn',)),
 315    'backpropagation through time': _E('Backpropagation applied to an RNN unrolled across its time steps.', 'primer.ml.cnn_rnn'),
 316    'cell state': _E("An LSTM's notebook, edited by adding rather than overwriting, so information survives many steps.", 'primer.ml.cnn_rnn'),
 317    'forget gate': _E('The LSTM dial that decides what to erase from the cell state.', 'primer.ml.cnn_rnn'),
 318    'input gate': _E('The LSTM dial that decides what new information to write into the cell state.', 'primer.ml.cnn_rnn'),
 319    'output gate': _E('The LSTM dial that decides how much of the cell state to show as output.', 'primer.ml.cnn_rnn'),
 320    'gru': _E('A simpler gated RNN with an update gate and a reset gate and no separate cell state.', 'primer.ml.cnn_rnn'),
 321    'state-space model': _E('A recurrent-style model, such as Mamba, that trains in parallel and runs in time linear in sequence length.', 'primer.ml.efficient_architectures'),
 322    'mamba': _E('A state-space model that trains in parallel and runs in time linear in sequence length.', 'primer.ml.efficient_architectures'),
 323    'tanh': _E('Squashes any number into the range −1 to 1.', 'primer.ml.cnn_rnn'),
 324    'tool call': _E('A structured request from the model to run one of your functions with specific arguments; your code decides whether to run it.', 'primer.agents.tools'),
 325    'tool_result': _E("The message block that returns a tool's output or error to the model, matched to its request by id.", 'primer.agents.agent_loop'),
 326    'stop reason': _E('Why a model response ended: finished, wants a tool, hit the length limit, or declined.', 'primer.agents.agent_loop'),
 327    'loop detection': _E('Stopping an agent that keeps requesting the same tool with the same arguments.', 'primer.agents.agent_loop'),
 328    'token budget': _E('A cap on the tokens, and so the money, one agent run may spend.', 'primer.agents.agent_loop'),
 329    'parallel tool calls': _E('Several tool requests in one model turn, run at the same time, with all results returned in one message.', 'primer.agents.agent_loop'),
 330    'strict tool use': _E("An API setting that guarantees the model's tool arguments match the tool's JSON Schema exactly.", 'primer.agents.tools'),
 331    'semantic validation': _E('Checking that well-formed arguments are also true, such as that a customer id actually exists.', 'primer.agents.tools'),
 332    'idempotency key': _E('A unique id sent with a write so a retried request returns the first result instead of acting twice.', 'primer.agents.tools'),
 333    'dry run': _E('Running every check for an action and describing it without actually doing it.', 'primer.agents.tools'),
 334    'human approval gate': _E('A rule that parks irreversible or high-value actions until a person approves them.', 'primer.agents.tools'),
 335    'least privilege': _E('Giving each agent or tool only the permissions its job needs.', 'primer.agents.tools'),
 336    'dynamic tool loading': _E('Sending the model only the few tool definitions most relevant to the current request.', 'primer.agents.tools'),
 337    'model context protocol': _E('An open standard that lets any AI application connect to any tool server the same way.', 'primer.agents.mcp'),
 338    'json-rpc': _E('A simple convention for calling functions on another program by exchanging JSON requests, replies and notifications.', 'primer.agents.mcp'),
 339    'stdio transport': _E('Running a server as a child process and exchanging one JSON message per line over its input and output.', 'primer.agents.mcp'),
 340    'mcp server': _E('The program that wraps a real system (files, a database, an API) and offers it to AI apps over MCP.', 'primer.agents.mcp'),
 341    'mcp resource': _E('Read-only data an MCP server offers for the app to load into context, addressed by a URI.', 'primer.agents.mcp'),
 342    'capability negotiation': _E('The start-of-session exchange where client and server announce which optional features they support.', 'primer.agents.mcp'),
 343    'tool poisoning': _E("Hiding instructions for the model inside a tool's description.", 'primer.agents.mcp'),
 344    'rug pull': _E('A tool server changing its definitions after they were approved.', 'primer.agents.mcp'),
 345    'confused deputy': _E('A program with broad authority tricked into using it for someone who lacks that authority.', 'primer.agents.mcp'),
 346    'oauth': _E('A standard way for a user to grant an app limited, revocable access without sharing their password.', 'primer.agents.mcp'),
 347    'decomposition': _E('Splitting a big goal into small subtasks, each with a checkable definition of done.', 'primer.agents.planning'),
 348    'external verification': _E("Checking a model's output with something outside the model, such as tests, schemas or database queries.", 'primer.agents.planning'),
 349    'checkpoint': _E('Saved progress after a completed step, so a failure resumes from there instead of from the start.', 'primer.agents.planning', scope=('primer/agents/planning', 'primer/agents/orchestration')),
 350    'durable execution': _E("Running a workflow so its progress survives crashes, by storing each completed step's result.", 'primer.agents.orchestration'),
 351    'state machine': _E('A fixed set of states and allowed transitions, with code deciding every move.', 'primer.agents.orchestration'),
 352    'short-term memory': _E('The current conversation, kept within a token budget: recent turns verbatim, older ones summarized.', 'primer.agents.memory'),
 353    'long-term memory': _E('Facts, events and procedures stored outside the model and recalled into later conversations.', 'primer.agents.memory'),
 354    'multi-tenant': _E('One system serving many separate customers whose data must never mix.', 'primer.agents.memory'),
 355    'tenant isolation': _E("Partitioning storage so a request can only ever reach its own tenant's and user's data.", 'primer.agents.memory'),
 356    'right to erasure': _E("A user's legal right, for example under GDPR, to have their personal data deleted.", 'primer.agents.memory'),
 357    'prompt chaining': _E('A fixed sequence of model calls where each output feeds the next, with code checks in between.', 'primer.agents.orchestration'),
 358    'routing': _E('One model call classifies a request and sends it to a specialised handler.', 'primer.agents.orchestration', scope=('primer/agents/orchestration', 'primer/agents/cost')),
 359    'orchestrator-workers': _E('One model decides the subtasks, workers handle them, and one model combines the results.', 'primer.agents.orchestration'),
 360    'evaluator-optimizer': _E('A generator drafts and an evaluator critiques, repeated until the draft passes or the rounds run out.', 'primer.agents.orchestration'),
 361    'supervisor': _E('In a multi-agent system, the agent that assigns tasks to specialist agents and assembles their answers.', 'primer.agents.orchestration', scope=('primer/agents/orchestration',)),
 362    'hand-off': _E('The message that passes work and its context from one agent to another; anything not written into it is lost.', 'primer.agents.orchestration', scope=('primer/agents/orchestration', 'primer/agents/agent_loop')),
 363    'stochastic gradient descent': _E("Gradient descent where each step's slope is estimated from a small random batch instead of the whole dataset.", 'primer.ml.optimizers'),
 364    'learning-rate schedule': _E('A rule that changes the learning rate during training, such as warming up and then decaying along a cosine curve.', 'primer.ml.optimizers'),
 365    'cosine decay': _E('Lowering the learning rate from its peak to a floor along half a cosine wave.', 'primer.ml.optimizers'),
 366    'bias correction': _E("Adam's rescaling of its running averages early in training, since they start at zero and would otherwise be too small.", 'primer.ml.optimizers'),
 367    'initialization': _E("Choosing the size of a network's random starting weights so signals neither shrink nor grow layer by layer.", 'primer.ml.deep_nets', scope=('primer/ml/deep_nets', 'primer/ml/neural_net')),
 368    'xavier initialization': _E('Starting weights with variance 2 / (fan-in + fan-out), suited to tanh and sigmoid layers.', 'primer.ml.deep_nets'),
 369    'he initialization': _E('Starting weights with variance 2 / fan-in, making up for ReLU zeroing half its inputs.', 'primer.ml.deep_nets'),
 370    'fan-in': _E('The number of inputs a neuron adds up.', 'primer.ml.deep_nets'),
 371    'saturation': _E('When an activation like sigmoid or tanh sits on its flat tail, so almost no gradient passes through.', 'primer.ml.neural_net', scope=('primer/ml/neural_net', 'primer/ml/deep_nets', 'primer/ml/attention')),
 372    'dying relu': _E('A ReLU neuron whose input stays negative, so it outputs zero and receives zero gradient forever.', 'primer.ml.neural_net'),
 373    'one-hot': _E('A vector with a 1 at the correct class and 0 everywhere else.', 'primer.ml.losses'),
 374    'log-sum-exp': _E('Computing the log of a sum of exponentials without overflow, by factoring out the largest term first.', 'primer.ml.losses'),
 375    'binary cross-entropy': _E('Cross-entropy for yes/no predictions: −[y·ln p + (1 − y)·ln(1 − p)].', 'primer.ml.losses'),
 376    'mean squared error': _E('The average of squared prediction errors, so big misses dominate.', 'primer.ml.losses'),
 377    'mean absolute error': _E('The average of absolute prediction errors, which is robust to outliers.', 'primer.ml.losses'),
 378    'label smoothing': _E('Training against a softened target, such as 0.9 on the right class and the rest spread evenly, to discourage over-confidence.', 'primer.ml.losses'),
 379    'bias-variance trade-off': _E("Splitting a model's error into systematic error from being too simple (bias) and sensitivity to the training sample from being too flexible (variance).", 'primer.ml.regularization'),
 380    'l1 regularization': _E('Penalizing the sum of absolute weights, which drives unneeded weights to exactly zero.', 'primer.ml.regularization'),
 381    'l2 regularization': _E('Penalizing the sum of squared weights, which shrinks all weights toward zero without eliminating them.', 'primer.ml.regularization'),
 382    'soft thresholding': _E('Moving each weight a fixed amount toward zero and snapping any weight within that amount to exactly zero.', 'primer.ml.regularization'),
 383    'validation set': _E('Held-out data used to tune choices like model size and when to stop training.', 'primer.ml.regularization'),
 384    'test set': _E('Held-out data used once at the end to estimate real-world performance honestly.', 'primer.ml.regularization', scope=('primer/ml/',)),
 385    'cross-validation': _E('Rotating which slice of the data is held out, training once per slice, and averaging the scores.', 'primer.ml.regularization'),
 386    'double descent': _E('The observation that past a point, even larger models generalize better again.', 'primer.ml.regularization'),
 387    'benchmark contamination': _E("When test questions appeared in a model's training data, inflating its scores.", 'primer.ml.regularization'),
 388    'gpu': _E('A graphics processor: a chip with thousands of small cores that do many multiply-adds at once, which is exactly what neural networks need.', 'primer.notation'),
 389    'parallelism': _E('Doing many pieces of work at the same time instead of one after another.', 'primer.ml.attention'),
 390    'hidden state': _E('The running summary an RNN carries from one step to the next.', 'primer.ml.cnn_rnn'),
 391    'auto-regressive': _E('Generating one token at a time, each predicted from everything before it and then fed back in.', 'primer.ml.big_picture'),
 392    'beam search': _E('Decoding that keeps the few most probable partial outputs at each step instead of just one, then picks the best finished one.', 'primer.ml.inference'),
 393    'checkpoint averaging': _E('Averaging the weights saved at the last few points of training, which usually gives a slightly better model.', 'primer.ml.training_stages'),
 394    'encoder-decoder attention': _E("Attention where the decoder's queries look at the encoder's keys and values, so each output word can consult the whole input.", 'primer.ml.transformer'),
 395    'luhn check': _E('A checksum every real card number satisfies, used to tell card numbers from look-alike IDs.', 'primer.agents.guardrails'),
 396    'named-entity recognition': _E('A model that tags names, places and organisations in text.', 'primer.agents.guardrails'),
 397    'ner': _E('Named-entity recognition: tagging names, places and organisations in text.', 'primer.agents.guardrails'),
 398    'entailment': _E('Whether one text logically follows from another, checked by a model trained for it.', 'primer.agents.guardrails'),
 399    'privilege separation': _E("Splitting work so the part that reads untrusted content can't take dangerous actions.", 'primer.agents.guardrails'),
 400    'lethal trifecta': _E('Private data, untrusted content and a way to send data out, all in one agent: together they allow data theft.', 'primer.agents.guardrails'),
 401    'lost in the middle': _E('Models use information at the start and end of a long input more reliably than information in the middle.', 'primer.agents.context'),
 402    'sandwich ordering': _E('Placing the best retrieved chunks at the start and end of the context and the weakest in the middle.', 'primer.agents.context'),
 403    'escaping': _E("Replacing characters like < and > so data can't pose as markup or instructions.", 'primer.agents.context'),
 404    'trajectory': _E('The full record of an agent run: its answer, its tool calls with arguments, and the end state.', 'primer.agents.evals'),
 405    'code grader': _E("An automatic check written in code, such as exact match, schema or database state, that scores an agent's output.", 'primer.agents.evals'),
 406    'rubric': _E('A short list of explicit pass/fail criteria given to a grader.', 'primer.agents.evals'),
 407    'regression': _E('A task that used to pass and now fails after a change.', 'primer.agents.evals'),
 408    'p95': _E('The 95th percentile: the value that 95% of measurements are at or below.', 'primer.agents.evals'),
 409    'continuous integration': _E('The automated checks that run on every proposed change before it can merge.', 'primer.agents.evals'),
 410    'drift': _E('A metric creeping away from its usual level over time or after a deploy.', 'primer.agents.evals'),
 411    'faithfulness': _E("The share of an answer's claims that are supported by the retrieved sources.", 'primer.agents.evals'),
 412    'cost per successful task': _E('Total spend divided by the number of tasks that succeeded, counting retries and cleanup.', 'primer.agents.cost'),
 413    'ttl': _E('Time to live: how long a cached entry may be served before it is treated as expired.', 'primer.agents.cost'),
 414    'batch api': _E('A provider interface that processes many requests asynchronously at a discount.', 'primer.agents.cost'),
 415    'task budget': _E('Hard limits on steps and tokens for one agent task.', 'primer.agents.cost'),
 416    'tenant': _E('One customer organisation on a shared platform.', 'primer.agents.cost'),
 417    'anomaly alert': _E('An alert when a metric jumps far outside its usual range.', 'primer.agents.cost'),
 418    'instrumentation': _E('Code that records spans or metrics around the work a program does.', 'primer.agents.observability'),
 419    'opentelemetry': _E('The open standard for traces, metrics and logs.', 'primer.agents.observability'),
 420    'genai semantic conventions': _E("OpenTelemetry's standard attribute names for model and agent spans, such as gen_ai.request.model.", 'primer.agents.observability'),
 421    'otlp': _E('The OpenTelemetry Protocol: the wire format exporters use to ship telemetry.', 'primer.agents.observability'),
 422    'collector': _E('In OpenTelemetry, the service that receives spans from apps and forwards them to storage backends.', 'primer.agents.observability'),
 423    'graduated autonomy': _E('Granting an agent more independence in steps (shadow, then approval, then autonomy), each earned with evidence.', 'primer.agents.deployment'),
 424    'content-addressed': _E('Named by a hash of its own content, so identical content always gets the identical name.', 'primer.agents.deployment'),
 425    'sha-256': _E('A hash function: a fixed-length fingerprint of data that changes completely if the data changes at all.', 'primer.agents.deployment'),
 426    'feature flag': _E('A configuration switch that turns a capability on or off without redeploying.', 'primer.agents.deployment'),
 427    'rate limit': _E('A cap on how many actions may happen per unit of time.', 'primer.agents.deployment'),
 428    'token bucket': _E('A rate limiter that allows bursts up to a capacity and refills at a steady rate.', 'primer.agents.deployment'),
 429    'blast radius': _E('How much damage a failure can do before it is noticed and stopped.', 'primer.agents.deployment'),
 430    'audit log': _E('A record of which actor took which action, for whom, with what inputs.', 'primer.agents.deployment'),
 431    'hash chain': _E("Records that each store the previous record's hash, so any edit or deletion is detectable.", 'primer.agents.deployment'),
 432    'erp': _E("Enterprise resource planning: a company's finance and operations system of record.", 'primer.agents.deployment'),
 433    'exponential backoff': _E('Waiting longer after each failed attempt, doubling up to a cap.', 'primer.agents.failures'),
 434    'jitter': _E("Random variation added to retry delays so clients don't all retry at the same moment.", 'primer.agents.failures'),
 435    'circuit breaker': _E('A wrapper that stops calling a failing dependency for a cool-down period, failing fast instead.', 'primer.agents.failures'),
 436    'contract test': _E('An automated check that an external API still accepts and returns what your code expects.', 'primer.agents.failures'),
 437    'ocr': _E('Optical character recognition: reading text from an image of a page.', 'primer.agents.rag'),
 438    'acl': _E('Access-control list: the users or groups allowed to read a document.', 'primer.agents.rag'),
 439    'permission-aware retrieval': _E('Filtering out documents a user may not read before anything is ranked or shown to the model.', 'primer.agents.rag'),
 440    'dense retrieval': _E('Search by comparing embedding vectors, so matches are by meaning rather than exact words.', 'primer.agents.rag'),
 441    'query rewriting': _E('Turning a follow-up or vague question into a standalone search query.', 'primer.agents.rag'),
 442    'multi-query retrieval': _E('Searching several phrasings of a question and fusing the results.', 'primer.agents.rag'),
 443    'contextual retrieval': _E('Indexing each chunk together with a line saying where in its document it comes from.', 'primer.agents.rag'),
 444    'parent-child retrieval': _E('Matching small chunks but returning the larger section around them.', 'primer.agents.rag'),
 445    'agentic rag': _E('RAG where the model decides whether, what and how often to search.', 'primer.agents.rag'),
 446    'graphrag': _E('Retrieval over a graph of entities and relationships extracted from documents.', 'primer.agents.rag'),
 447    'centroid': _E("The average position of a group of vectors, used as the group's representative point.", 'primer.ml.embeddings.ann'),
 448    'voronoi cell': _E('All the points closer to one centroid than to any other: the section an IVF index searches.', 'primer.ml.embeddings.ann'),
 449    'nprobe': _E('How many IVF clusters a query scans; raising it trades speed for recall.', 'primer.ml.embeddings.ann'),
 450    'efsearch': _E("HNSW's query-time beam width; raising it trades speed for recall.", 'primer.ml.embeddings.ann'),
 451    'efconstruction': _E("HNSW's beam width while building the graph; higher means a better graph and a slower build.", 'primer.ml.embeddings.ann'),
 452    'greedy search': _E('Always move to whichever neighbour is closest to the target, and stop when none is closer.', 'primer.ml.embeddings.ann'),
 453    'graph': _E('A set of points (nodes) joined by links (edges).', 'primer.ml.embeddings.ann'),
 454    'codebook': _E('A small catalogue of representative vectors; each vector (or piece of one) is stored as the number of its nearest entry, as in product quantization or tokenizers for images and audio.', 'primer.ml.embeddings.ann'),
 455    'asymmetric distance computation': _E('Scoring compressed vectors against an uncompressed query by adding up precomputed table lookups.', 'primer.ml.embeddings.ann'),
 456    'ivf-pq': _E('IVF clustering combined with compressed offsets from each cluster centre: the standard index for billion-scale search.', 'primer.ml.embeddings.ann'),
 457    're-scoring': _E('Recomputing exact scores for a shortlist that an approximate index returned, to recover accuracy.', 'primer.ml.embeddings.ann'),
 458    'flat index': _E('Exact search that compares the query with every stored vector; the ground truth other indexes are measured against.', 'primer.ml.embeddings.ann'),
 459    'idf': _E("Inverse document frequency: a word's rarity weight, higher for words found in fewer documents.", 'primer.ml.embeddings.retrieval'),
 460    'term frequency': _E('How many times a word appears in a document.', 'primer.ml.embeddings.retrieval'),
 461    'length normalization': _E("BM25's discount for mentions in documents longer than average, controlled by b.", 'primer.ml.embeddings.retrieval'),
 462    'late interaction': _E('Keeping one vector per word and matching each query word to its best document word.', 'primer.ml.embeddings.retrieval'),
 463    'maxsim': _E('For each query word, its highest similarity to any document word, summed over the query.', 'primer.ml.embeddings.retrieval'),
 464    'query and passage prefixes': _E('Labels some embedding models expect on questions and documents; forgetting one silently lowers recall.', 'primer.ml.embeddings.retrieval'),
 465    'shortlist': _E('The small set of top candidates from a cheap first-stage search that a slower, more precise stage reorders.', 'primer.ml.embeddings.retrieval'),
 466    'l2 normalization': _E('Dividing a vector by its own length so it has length 1.', 'primer.ml.embeddings.similarity'),
 467    'dimension': _E('One of the numbers in a vector: a 768-dimension embedding is a list of 768 numbers.', 'primer.ml.embeddings.similarity'),
 468    'mean-centering': _E('Subtracting the average vector of a collection from every vector, removing the direction they all share.', 'primer.ml.embeddings.similarity'),
 469    'whitening': _E('Rescaling vectors so every direction has equal spread and no two directions are correlated.', 'primer.ml.embeddings.similarity'),
 470    'threshold calibration': _E('Choosing a similarity cut-off by measuring precision and recall on labeled pairs for one specific model.', 'primer.ml.embeddings.similarity'),
 471    'word2vec': _E('A 2013 method that learns one vector per word by predicting nearby words.', 'primer.ml.embeddings.word2vec'),
 472    'skip-gram': _E("The word2vec task of predicting each word's neighbours from the word itself.", 'primer.ml.embeddings.word2vec'),
 473    'negative sampling': _E('Training by scoring the true pair up and a few random noise pairs down, instead of scoring the whole vocabulary.', 'primer.ml.embeddings.word2vec'),
 474    'pmi': _E('Pointwise mutual information: the log of how much more often two words appear together than chance predicts.', 'primer.ml.embeddings.word2vec'),
 475    'glove': _E('A counting method that fits word vectors so their dot products predict how often words appear together.', 'primer.ml.embeddings.word2vec'),
 476    'svd': _E('Singular value decomposition: splits a table into a few directions that capture its main patterns, used to compress it.', 'primer.ml.embeddings.word2vec'),
 477    'static embedding': _E('One fixed vector per word, whatever the sentence.', 'primer.ml.embeddings.word2vec'),
 478    'contextual embedding': _E('A vector computed for a word within its sentence, so the same word gets different vectors in different contexts.', 'primer.ml.embeddings.word2vec'),
 479    'contrastive learning': _E('Training that pulls matching pairs together and pushes non-matching pairs apart.', 'primer.ml.embeddings.contrastive'),
 480    'in-batch negatives': _E('Using the other examples in a training batch as free wrong answers for each query.', 'primer.ml.embeddings.contrastive'),
 481    'bag of words': _E('Representing text by counting how many times each word appears, ignoring order.', 'primer.ml.embeddings.contrastive'),
 482    'zero-shot classification': _E('Labeling items by comparing them to a text description of each label, with no training on those labels.', 'primer.ml.embeddings.contrastive'),
 483    'scalar quantization': _E('Storing each number as one of 256 levels (one byte) instead of a 4-byte float.', 'primer.ml.embeddings.compression'),
 484    'binary quantization': _E('Storing only the sign of each number, as a single bit.', 'primer.ml.embeddings.compression'),
 485    'hamming distance': _E('The number of positions where two bit strings differ.', 'primer.ml.embeddings.compression'),
 486    'popcount': _E('Counting the 1-bits in a number; combined with XOR it computes Hamming distance in one step.', 'primer.ml.embeddings.compression'),
 487    'two-stage retrieval': _E('A cheap, wide first pass to shortlist candidates, then an expensive, precise pass over only those.', 'primer.ml.embeddings.compression'),
 488    'hit@k': _E('The share of queries with at least one relevant result in the top k.', 'primer.ml.embeddings.operations'),
 489    'pca': _E('Principal component analysis: finding the directions along which data varies most, to draw or compress it with fewer numbers.', 'primer.ml.embeddings.clustering'),
 490    'clustering': _E('Grouping items so similar ones end up together, without being told the groups.', 'primer.ml.embeddings.clustering'),
 491    'inertia': _E("The total squared distance from each point to its cluster's center: the quantity k-means minimizes.", 'primer.ml.embeddings.clustering'),
 492    'k-means++': _E('A way to pick k-means starting centers spread far apart, which avoids bad results.', 'primer.ml.embeddings.clustering'),
 493    'silhouette score': _E('How much closer each point is to its own cluster than to the nearest other one, from −1 to 1.', 'primer.ml.embeddings.clustering'),
 494    'elbow method': _E('Choosing the number of clusters where adding more stops reducing inertia much.', 'primer.ml.embeddings.clustering'),
 495    'dbscan': _E('Density-based clustering that grows clusters from points with enough close neighbours and labels the rest as noise.', 'primer.ml.embeddings.clustering'),
 496    'hdbscan': _E('A version of DBSCAN that considers every reach at once and keeps the most persistent clusters, so no reach setting is needed.', 'primer.ml.embeddings.clustering'),
 497    'umap': _E("A projection to 2-D that keeps each point's neighbours close but distorts other distances.", 'primer.ml.embeddings.clustering'),
 498    't-sne': _E('An older neighbour-preserving 2-D projection; cluster sizes and gaps in its plots are not meaningful.', 'primer.ml.embeddings.clustering'),
 499    'near-duplicate detection': _E('Finding items whose embeddings are almost identical, such as reworded copies.', 'primer.ml.embeddings.clustering'),
 500    'anomaly detection': _E('Flagging items far from every known group.', 'primer.ml.embeddings.clustering'),
 501    'blue/green deployment': _E('Building the new version beside the old one, switching traffic once it proves itself, and keeping the old one for rollback.', 'primer.ml.embeddings.operations'),
 502    'dual write': _E('Writing every new record to both the old and the new system during a migration.', 'primer.ml.embeddings.operations'),
 503    'backfill': _E('Re-processing existing records into a new system, such as re-embedding a corpus with a new model.', 'primer.ml.embeddings.operations'),
 504    'index alias': _E('A pointer name like "live" that search uses, so switching indexes means repointing one name.', 'primer.ml.embeddings.operations'),
 505    'domain mismatch': _E('A general model underperforming on specialised language it never saw, such as company jargon.', 'primer.ml.embeddings.operations'),
 506    'retrieval failure': _E('A wrong answer caused because no relevant document was retrieved.', 'primer.ml.embeddings.operations'),
 507    'generation failure': _E('A wrong answer produced even though a relevant document was retrieved.', 'primer.ml.embeddings.operations'),
 508    'inverted file': _E('In an IVF index, the list of vectors assigned to one cluster, so a search scans only the lists it picks.', 'primer.ml.embeddings.ann'),
 509    'diversity heuristic': _E("HNSW's rule of keeping a neighbour only if it is closer to the new node than to any neighbour already kept, so links fan out in different directions.", 'primer.ml.embeddings.ann'),
 510    'sparse retrieval': _E('Keyword search such as BM25, which scores exact word matches.', 'primer.ml.embeddings.retrieval'),
 511    'hysteresis': _E("Making the bar for changing state higher than the bar for staying, so a decision doesn't flicker between two close options.", None),
 512    'debounce': _E('Waiting until a signal has settled before acting on it, so a burst of changes triggers one action.', None),
 513    'cooldown': _E('A quiet period after an action during which the same action is not repeated.', None),
 514    'latency': _E("The delay between something happening and the system's response to it.", 'primer.agents.cost'),
 515    'rank': _E('How many independent directions a matrix really contains; a table made by multiplying one column by one row has rank 1.', 'primer.ml.training_stages', scope=('primer/ml/training_stages',)),
 516    'singular value decomposition': _E('Breaking a matrix into independent directions ranked by importance, so the top few capture most of what it does.', 'primer.ml.embeddings.word2vec'),
 517    'frobenius norm': _E('The size of a whole matrix: square every entry, add them up, take the square root.', 'primer.notation'),
 518    'prefix tuning': _E('Training a few special virtual-token vectors placed before every input while the model stays frozen.', 'primer.ml.training_stages'),
 519    'ppo': _E('Proximal Policy Optimization: a reinforcement learning algorithm that improves a policy in small, clipped steps; widely used for RLHF.', 'primer.ml.training_stages'),
 520    'alignment tax': _E('Capability lost as a side effect of training a model to be helpful, honest and harmless.', 'primer.ml.training_stages'),
 521    'win rate': _E('The share of head-to-head comparisons one model wins; 50% means the two are indistinguishable.', 'primer.agents.evals'),
 522    'few-shot prompt': _E('A prompt that includes a few worked examples of the task before the real question.', 'primer.agents.context'),
 523    'exemplar': _E('One worked example in a few-shot prompt: a question with its answer, and sometimes the steps that reach it.', 'primer.agents.context'),
 524    'linear probe': _E('Freezing a model and training only a simple linear classifier on its embeddings or internal activations, to measure what they encode.', 'primer.ml.embeddings.contrastive'),
 525    'zero-shot': _E('Doing a task with no task-specific training examples.', 'primer.ml.embeddings.contrastive'),
 526    'map@10': _E('Mean average precision over the top 10 results: rewards putting correct matches in the top 10, and higher up within it.', 'primer.ml.metrics'),
 527    'nf4': _E('4-bit NormalFloat: a 4-bit number format whose 16 levels are spaced to match the bell-curve shape of neural network weights.', 'primer.ml.inference'),
 528    'double quantization': _E('Compressing the per-block scale factors of a quantized model as well, saving about 0.37 bits per parameter.', 'primer.ml.inference'),
 529    'mips': _E('Maximum inner product search: finding the stored vectors with the largest dot product against a query vector, usually approximately with an index.', 'primer.ml.embeddings.ann'),
 530    'expected value': _E('The average you would get over many tries.', 'primer.notation'),
 531    'credit assignment': _E('Working out which of many earlier actions deserves the blame or credit for how an episode ended.', 'primer.agents.planning'),
 532    'pass@1': _E('The share of problems solved by the first program submitted for each, judged by hidden tests.', 'primer.agents.evals'),
 533    'machine translation': _E('Turning text in one language into another.', 'primer.ml.transformer'),
 534    'linear attention': _E('Attention variants that avoid scoring every pair of tokens, so cost grows in proportion to n instead of n².', 'primer.ml.attention'),
 535    'in-context learning': _E("Picking up a task from instructions or examples in the prompt, with no change to the model's weights.", 'primer.ml.big_picture'),
 536    'speculative execution': _E('Starting likely work early, in parallel, and throwing it away if the guess was wrong.', 'primer.ml.inference'),
 537    'petaflop/s-day': _E('10¹⁵ operations per second for one day, 8.64 × 10¹⁹ operations: a unit for training budgets.', 'primer.ml.training_stages'),
 538    'isoflop profile': _E('A curve of final loss against model size at a fixed training compute; its minimum is the best size for that budget.', 'primer.ml.training_stages'),
 539    'hbm': _E("The GPU's large main memory: tens of GB, about 10× slower than on-chip SRAM.", 'primer.ml.inference'),
 540    'sram': _E("The tiny, very fast on-chip memory next to a GPU's arithmetic units.", 'primer.ml.inference'),
 541    'kernel fusion': _E('Doing several steps in one GPU program, so data is read from and written to main memory only once.', 'primer.ml.inference'),
 542    'internal fragmentation': _E('Memory reserved for a request that it never uses.', 'primer.ml.inference'),
 543    'external fragmentation': _E('Free memory split into gaps too small to use.', 'primer.ml.inference'),
 544    'block table': _E('A per-request map from logical KV-cache blocks to physical blocks in GPU memory.', 'primer.ml.inference'),
 545    'copy-on-write': _E('Sharing one copy of data and making a private copy only when someone needs to change it.', 'primer.ml.inference'),
 546    'co-adaptation': _E("When neurons learn to depend on each other's exact behaviour, each correcting the others' quirks; it fits the training data but breaks on new data.", 'primer.ml.regularization'),
 547    'max-norm constraint': _E("Capping the length of each neuron's incoming weight vector at a fixed radius, shrinking it back whenever an update pushes it past.", 'primer.ml.regularization'),
 548    'regret': _E("In online learning, the total extra loss from choosing each step's parameters before seeing that step's data, compared with the best fixed choice in hindsight.", 'primer.ml.optimizers', scope=('primer/ml/optimizers',)),
 549    'convex': _E('Bowl-shaped everywhere with a single bottom, so any downhill path reaches the same minimum; neural network losses are not convex.', 'primer.notation'),
 550    'adagrad': _E("An optimizer that divides each weight's step by the square root of the sum of all its past squared gradients, so steps only ever shrink.", 'primer.ml.optimizers', display="AdaGrad"),
 551    'rmsprop': _E('An optimizer that divides each step by the square root of a running average of recent squared gradients.', 'primer.ml.optimizers'),
 552    'sparse gradients': _E('Gradients that are zero most of the time for a given weight, such as the weight for a rare word.', 'primer.ml.optimizers'),
 553    'degradation problem': _E('Deeper plain networks reaching higher error even on their training data: an optimization failure, not overfitting.', 'primer.ml.deep_nets'),
 554    'subword unit': _E('A piece of a word, from a single character to a whole frequent word; a fixed set of them can spell any word.', 'primer.ml.tokenization'),
 555    'mean average precision': _E('How well a system ranks correct results above wrong ones, averaged over queries or classes.', 'primer.ml.metrics'),
 556    'cbow': _E("Continuous bag of words: the word2vec model that averages the surrounding words' vectors to predict the middle word.", 'primer.ml.embeddings.word2vec'),
 557    'hierarchical softmax': _E('Replacing one softmax over the whole vocabulary with about log₂V yes-or-no decisions down a binary tree.', 'primer.ml.embeddings.word2vec'),
 558    'subsampling': _E('Randomly skipping most occurrences of very frequent words during training, which is faster and improves rare-word vectors.', 'primer.ml.embeddings.word2vec'),
 559    'siamese network': _E('Two copies of one network with shared weights, each reading one input, so their outputs can be compared directly.', 'primer.ml.embeddings.contrastive'),
 560    'triplet loss': _E('A loss requiring an anchor to be closer to its positive than to its negative by at least a margin.', 'primer.ml.losses'),
 561    'spearman correlation': _E('How well two rankings agree, from −1 (reversed) to 1 (identical).', 'primer.ml.metrics'),
 562    'top-k retrieval accuracy': _E('The share of questions for which at least one of the top k retrieved passages contains the answer.', 'primer.ml.metrics'),
 563    'navigable small world': _E('A graph where always stepping to the neighbour closest to the target reaches it in few hops.', 'primer.ml.embeddings.ann'),
 564    'skip list': _E('A sorted list with random express lanes on top, searched by running along the top lane and dropping down, in about log(n) steps.', 'primer.ml.embeddings.ann'),
 565    'huffman tree': _E('A binary tree that gives frequent items short codes and rare items long ones.', 'primer.ml.embeddings.word2vec'),
 566    'logistic regression': _E('A linear model whose weighted sum passes through a sigmoid to give a probability.', 'primer.ml.neural_net'),
 567    'locality-sensitive hashing': _E('Hashing vectors so that similar vectors are likely to land in the same bucket, for fast approximate search.', 'primer.ml.embeddings.ann'),
 568    'ollama': _E('A free app that downloads open language models and runs them on your own computer, served over a small local web API.', 'primer.agents.llm'),
 569    'local model': _E("A language model running on your own machine rather than a provider's servers: free per call, private and offline, but smaller.", 'primer.agents.llm'),
 570    # Added with the lessons on training at scale, reasoning, generation, and more.
 571    'activation checkpointing': _E("Keeping only each layer's input and recomputing the rest during the backward pass.", 'primer.ml.pretraining'),
 572    'activation patching': _E('Copying one activation from a clean run into a corrupted run to measure how much of the right answer returns; also called causal tracing.', 'primer.ml.interpretability'),
 573    'actor-critic': _E('A reinforcement learner in two parts: the actor (the policy) chooses actions, and the critic (a value network) predicts the reward to expect, which serves as the baseline.', 'primer.ml.reinforcement'),
 574    'adaptive layer norm': _E('Layer normalization whose scale and shift are computed from a conditioning signal, such as the noise step and class label, instead of being fixed learned weights. Diffusion transformers use it to tell every block what to make (adaLN).', 'primer.ml.deep_nets'),
 575    'active learning': _E('Choosing which examples to label next by what the current model is likely to get wrong, so each paid label teaches more.'),
 576    'advantage': _E('How much better an action did than typical: reward minus baseline.', 'primer.ml.reinforcement', scope=('primer/ml/reinforcement',)),
 577    'adaboost': _E('A boosting method (Freund and Schapire, 1996) that trains each new classifier on reweighted data, raising the weight of the examples the previous ones got wrong, and lets the classifiers vote with weights set by their accuracy.', None),
 578    'adversarial example': _E('An input changed by a small, deliberately chosen amount so that a model gets it wrong, often in a way a person would never notice.', None),
 579    'alignment': _E("Making a model's behaviour match what we want (helpful, honest, harmless), not just the measurements we optimize.", 'primer.ml.alignment', scope=('primer/ml/alignment',)),
 580    'all-gather': _E('Every GPU contributes its piece and every GPU ends up with all the pieces, side by side; the second half of a ring all-reduce.', 'primer.ml.pretraining'),
 581    'all-reduce': _E('Summing a value across GPUs so every GPU ends up holding the total.', 'primer.ml.hardware'),
 582    'arena': _E('Ratings built from people voting between two anonymous answers.', 'primer.ml.benchmarks', scope=('primer/ml/benchmarks',), display="arena"),
 583    'attack success rate': _E('The share of red-team attempts that get past a safety check.', 'primer.ml.alignment'),
 584    'attention sink': _E('Early tokens that trained models park spare attention on; dropping them breaks streaming generation.', 'primer.ml.efficient_architectures'),
 585    'autoencoder': _E('A network that squeezes its input into a small code and rebuilds the input from it, trained only to make the rebuild match.', 'primer.ml.generative.autoencoders'),
 586    'automated interpretability': _E('Using a language model to explain a feature or neuron from examples of when it fires, then scoring the explanation by how well the model predicts new activations from it alone.', None),
 587    'bagging': _E('Bootstrap aggregating: train one model per bootstrap sample and average them to cut variance.', 'primer.ml.classical'),
 588    'baseline': _E('A typical reward subtracted before updating; it reduces noise without changing the average gradient.', 'primer.ml.reinforcement', scope=('primer/ml/reinforcement',)),
 589    'benchmark': _E('Fixed questions plus a scoring rule, averaged into one comparable number.', 'primer.ml.benchmarks', scope=('primer/ml/benchmarks',)),
 590    'bernoulli distribution': _E('The distribution of a single yes-or-no outcome, set by one number: the chance of yes. A VAE decoder for black-and-white pixels outputs one per pixel.', 'primer.ml.losses'),
 591    'bert': _E('An encoder-only transformer trained to fill in hidden words using the text on both sides of them; the ancestor of many embedding and classification models.', 'primer.ml.transformer'),
 592    'best-of-n': _E('Sample n candidates and keep the one a verifier scores highest.', 'primer.ml.reasoning'),
 593    'binomial coefficient': _E('"n choose k", written C(n, k): the number of different groups of k items that can be picked from n, ignoring order. C(10, 5) = 252.', 'primer.ml.benchmarks'),
 594    'binomial distribution': _E('The chances of getting each possible number of successes in n independent tries that each succeed with the same probability p.', 'primer.ml.benchmarks'),
 595    'bf16': _E("A 16-bit float with fp32's 8 exponent bits and 7 mantissa bits: the same range, less precision.", 'primer.ml.hardware'),
 596    'bits per dimension': _E('A model\'s negative log-likelihood per number in the data, in bits: how long a code the model needs, on average, for each pixel value. Lower is better.', 'primer.ml.losses'),
 597    'boosting': _E('Building a strong model from many weak ones trained one after another, each concentrating on what the ones before it still get wrong.', 'primer.ml.classical'),
 598    'bootstrap': _E('Measuring uncertainty by rescoring many with-replacement resamples of your own data.', 'primer.ml.benchmarks'),
 599    'bootstrap sample': _E('A resample of the training rows drawn with replacement, so some rows repeat and about 37% are left out.', 'primer.ml.classical'),
 600    'bottleneck': _E('The narrow middle of an autoencoder: too small to copy through, it forces the network to keep only what matters.', 'primer.ml.generative.autoencoders', scope=('primer/ml/generative/autoencoders',)),
 601    'bounding box': _E('A rectangle, given by its corner coordinates, that marks where one object sits in an image.', None),
 602    'cache miss': _E('When the processor needs a number that is not in its small, fast cache and must wait for the much slower main memory, which can cost as much time as hundreds of arithmetic operations.', 'primer.ml.hardware'),
 603    'canary string': _E('A unique marker in benchmark files so trainers can filter out copies.', 'primer.ml.benchmarks'),
 604    'catastrophic forgetting': _E('When training on a new task alone erodes or erases skills a model already had.', 'primer.ml.fine_tuning'),
 605    'causal tracing': _E('Activation patching used to find where a model recalls a fact: blur the subject in the input, then restore one clean hidden state at a time and see which ones bring the right answer back.', 'primer.ml.interpretability'),
 606    'chain of thought': _E('Intermediate reasoning steps a model writes before its final answer, each one readable by the next forward pass.', 'primer.ml.reasoning'),
 607    'chat template': _E('The exact role markers and layout a model family uses to turn messages into training text.', 'primer.ml.fine_tuning'),
 608    'chinchilla scaling': _E('The compute-optimal rule of about 20 training tokens per model parameter.', 'primer.ml.pretraining'),
 609    'cider': _E('A score for a generated caption: how well its words and short phrases match several human-written captions of the same image, with phrases common to captions of many images counting less. Higher is better; values above 100 are normal.', None),
 610    'circuit': _E('A chain of components that together compute one behaviour.', 'primer.ml.interpretability', scope=('primer/ml/interpretability',)),
 611    'classifier guidance': _E("Steering a diffusion model toward a label by adding the slope of a separate classifier, trained on noisy inputs, to the model's noise guess at every step.", 'primer.ml.generative.diffusion'),
 612    'classifier-free guidance': _E("Mixing a model's guesses with and without the prompt, and pushing past the prompted one to follow it more closely.", 'primer.ml.generative.diffusion'),
 613    'cnn': _E('Convolutional neural network: a network built from convolutions, small learned filters slid across an image to find local patterns wherever they appear.', 'primer.ml.cnn_rnn'),
 614    'commitment loss': _E("A training term that pulls an encoder's output towards the codebook entry it snapped to, so the encoder commits to its entries instead of drifting away from them.", None),
 615    'common crawl': _E("A nonprofit's public archive of the web, released as regular snapshots of billions of pages; the raw material of most pretraining data.", 'primer.ml.pretraining'),
 616    'compressed sensing': _E('Recovering a long vector from far fewer measurements of it, which is possible when the vector is known to be sparse (mostly zeros).', None),
 617    'computer use': _E('An agent operating a graphical interface from screenshots, sending clicks and keystrokes.', 'primer.agents.coding_agents'),
 618    'confidence interval': _E('A range built so that 95 in 100 such ranges contain the true value.', 'primer.ml.benchmarks'),
 619    'constitutional ai': _E('Training against a written list of principles: the model critiques and revises its answers, and AI-labelled preferences train the reward model.', 'primer.ml.alignment'),
 620    'constrained decoding': _E("Before each token is drawn, setting every token that can't lead to a valid answer to probability zero.", 'primer.ml.structured_output'),
 621    'contamination': _E("Test questions and answers leaking into a model's training data.", 'primer.ml.benchmarks', scope=('primer/ml/benchmarks',)),
 622    'control task': _E('A probe trained on random labels, showing how much a probe fits with nothing real to find.', 'primer.ml.interpretability', scope=('primer/ml/interpretability',)),
 623    'coupling': _E('A way of pairing draws from two distributions so that each side, on its own, still has its own distribution. Pairing noise with data at random is one coupling; sending each noise point to one definite data point is another.', 'primer.ml.generative.diffusion', scope=('primer/ml/generative',)),
 624    'cost per resolved task': _E('Everything spent on all attempts, divided by the tasks actually resolved.', 'primer.agents.coding_agents'),
 625    'covariance': _E('How two quantities vary together: positive when they rise together, negative when one rises as the other falls. For vectors, a matrix holding it for every pair of entries.', None),
 626    'critic': _E("A Wasserstein GAN's discriminator, which outputs an unbounded score instead of a probability.", 'primer.ml.generative.gans', scope=('primer/ml/generative/gans',)),
 627    'cross-attention': _E('Attention whose queries come from one sequence (image patches) and whose keys and values come from another (text tokens).', 'primer.ml.generative.diffusion'),
 628    'data mixture': _E('The share of training tokens drawn from each data source.', 'primer.ml.pretraining'),
 629    'data parallelism': _E('Every GPU holds the whole model and trains on a different slice of the batch.', 'primer.ml.pretraining'),
 630    'dcgan': _E('Deep convolutional GAN: a GAN whose generator and discriminator are convolutional networks, the usual baseline design for image GANs.', 'primer.ml.generative.gans'),
 631    'ddim': _E('A deterministic diffusion sampler that predicts the clean result and jumps to a much less noisy step, so it needs far fewer steps.', 'primer.ml.generative.diffusion'),
 632    'ddpm': _E('The original diffusion sampler: remove the guessed noise in many small steps, adding a small fresh wobble each time.', 'primer.ml.generative.diffusion'),
 633    'dead latent': _E('A sparse autoencoder latent that has stopped firing on any input, wasting its slot in the dictionary; resampling it onto badly rebuilt inputs brings it back.', 'primer.ml.interpretability'),
 634    'decision tree': _E('A flowchart of yes/no questions about one feature at a time, learned from examples, ending in leaves that give the answer.', 'primer.ml.classical'),
 635    'deduplication': _E('Removing repeated copies of documents from training data.', 'primer.ml.pretraining'),
 636    'denoiser': _E('The network in a diffusion model that looks at a noisy input and its step, and guesses the noise in it.', 'primer.ml.generative.diffusion', scope=('primer/ml/generative/diffusion',)),
 637    'denoising score matching': _E('Learning the score by training a network to guess the noise added to real examples.', 'primer.ml.generative.diffusion'),
 638    'dictionary learning': _E('Finding a set of directions, usually more than there are dimensions, such that every data point is a sparse combination of a few of them; also called sparse coding. A sparse autoencoder is one way to do it.', 'primer.ml.interpretability'),
 639    'diffusion model': _E('A generator that learns to turn random noise into data by removing a little noise at a time.', 'primer.ml.generative.diffusion'),
 640    'diffusion transformer': _E('A transformer used as the denoiser, reading patches of the latent as tokens (DiT).', 'primer.ml.generative.diffusion'),
 641    'discretization': _E("Turning a continuous rate of change into one step's keep and write factors.", 'primer.ml.efficient_architectures'),
 642    'discriminator': _E('The network in a GAN that outputs the probability that a sample is real rather than generated.', 'primer.ml.generative.gans', scope=('primer/ml/generative/gans',)),
 643    'edit distance': _E('The fewest single-item insertions, deletions and substitutions that turn one sequence into another.', None),
 644    "earth mover's distance": _E('The least total work to reshape one pile of probability into another, each bit of mass times the distance it travels; the same as the Wasserstein distance.', 'primer.ml.generative.gans'),
 645    'edit-run-test loop': _E('An agent loop that changes code, runs the tests, and repeats until they pass or a budget runs out.', 'primer.agents.coding_agents'),
 646    'elo rating': _E('A rating scale where a 400-point gap means ten-to-one odds of winning.', 'primer.ml.benchmarks'),
 647    'ensemble': _E('A model that combines the predictions of many models, such as trees.', 'primer.ml.classical', scope=('primer/ml/classical',)),
 648    'entropy': _E('The average surprise of a distribution: how many yes/no questions, on average, it takes to learn an outcome; 0 means certain.', 'primer.ml.classical'),
 649    'euler step': _E('Moving in a straight line for a short time at the current velocity, then looking again.', 'primer.ml.generative.diffusion'),
 650    'euler-maruyama step': _E('The Euler step for a stochastic differential equation: move by the drift times the time step, then add a fresh random nudge whose size grows with the square root of the time step.', 'primer.ml.generative.diffusion'),
 651    'evidence lower bound': _E('A quantity that never exceeds the log-probability a model gives the data; the VAE loss is its negative.', 'primer.ml.generative.autoencoders'),
 652    'expectation': _E('The average of a quantity over many random draws, written E.', 'primer.ml.generative.gans', scope=('primer/ml/generative/gans',)),
 653    'exploration': _E('Trying actions you are unsure of instead of repeating the best one so far.', 'primer.ml.reinforcement', scope=('primer/ml/reinforcement',)),
 654    'exponent': _E('The bits of a float that pick the power of two, which sets its range.', 'primer.ml.hardware', scope=('primer/ml/hardware',)),
 655    'exponential moving average': _E('A running average that keeps most of its old value and mixes in a small share of each new one, so recent values count most. Adam keeps its moments this way, and diffusion models sample with such an average of their weights.', 'primer.ml.optimizers'),
 656    'fail-to-pass test': _E('A hidden test that fails before a fix and must pass after it.', 'primer.agents.coding_agents'),
 657    'false positive': _E('Something flagged as positive that is actually negative, such as a solution graded correct because its final answer matches although its reasoning is wrong.', 'primer.ml.metrics'),
 658    'feature': _E('A property a model tracks, stored as a direction across many neurons.', 'primer.ml.interpretability', scope=('primer/ml/interpretability',)),
 659    'feature engineering': _E('Hand-transforming inputs (logarithms, one-hot columns) so a model can use them.', 'primer.ml.classical'),
 660    'feature importance': _E('A score for how much a model relies on each input column.', 'primer.ml.classical'),
 661    'feature map': _E('In linear attention, the function φ applied to queries and keys so that φ(q)·φ(k) replaces e^(q·k).', 'primer.ml.efficient_architectures'),
 662    'feature splitting': _E('One feature in a small sparse autoencoder becoming several narrower features in a larger one, such as one base64 feature becoming separate features for letters, digits and encoded text.', 'primer.ml.interpretability'),
 663    'fid': _E('Fréchet Inception Distance: compares a set of generated images with real ones through the statistics of their features in a pretrained image network. Lower is better; 0 means indistinguishable.', 'primer.ml.generative.gans'),
 664    'finite-state machine': _E('A fixed set of states with one move per input character; it can check patterns but not unlimited nesting.', 'primer.ml.structured_output'),
 665    'flip rate': _E('The share of questions whose answer changes when the user asserts a wrong answer.', 'primer.ml.alignment'),
 666    'floating point': _E('Storing a number as a sign, an exponent (range) and a mantissa (precision).', 'primer.ml.hardware'),
 667    'flow matching': _E('Training a network to output the velocity along paths from noise to data, then following it to generate.', 'primer.ml.generative.diffusion'),
 668    'forward process': _E('The fixed, unlearned procedure that mixes data with noise step by step until only noise is left.', 'primer.ml.generative.diffusion', scope=('primer/ml/generative/diffusion',)),
 669    'fourier transform': _E('Splitting a signal into how much of each frequency it contains.', 'primer.ml.generative.multimodal'),
 670    'fp16': _E('A 16-bit float with 5 exponent and 10 mantissa bits: more precision, range only up to 65,504.', 'primer.ml.hardware'),
 671    'fp8': _E('8-bit floats: E4M3 (range to 448) for weights and activations, E5M2 (to 57,344) for gradients.', 'primer.ml.hardware'),
 672    'frame sampling': _E('Keeping only some video frames, such as one a second, to save tokens.', 'primer.ml.generative.multimodal'),
 673    'functional correctness': _E('Judging generated code by running it against tests rather than by comparing its text with a reference solution.', 'primer.ml.benchmarks'),
 674    'fsdp': _E('Fully sharded data parallel: all training state sharded across GPUs (ZeRO stage 3).', 'primer.ml.pretraining'),
 675    'freezing': _E("Keeping some of a model's weights fixed during training, so only the rest learn and the frozen part keeps what it already knew.", 'primer.ml.fine_tuning'),
 676    'fused multiply-add': _E('One instruction that multiplies two numbers and adds the result to a running total.', 'primer.ml.hardware'),
 677    'gan': _E("A generator and a discriminator trained against each other, so the generator learns to make samples the discriminator can't tell from real data.", 'primer.ml.generative.gans'),
 678    'gated cross-attention': _E("Cross-attention layers added inside a frozen language model so its text can look at image features, with each layer's output multiplied by tanh of a learned number that starts at 0, so the model starts out unchanged.", 'primer.ml.generative.multimodal'),
 679    'generative model': _E('A model that learns to produce new samples resembling its training data, such as new faces, voices or sentences.', 'primer.ml.generative.autoencoders'),
 680    'gaussian': _E('The bell-curve distribution, set by its mean (where the centre is) and its standard deviation (how wide it is).', 'primer.ml.generative.autoencoders'),
 681    'generalization gap': _E('Validation loss minus training loss; when it keeps growing, the model is memorising.', 'primer.ml.fine_tuning'),
 682    'generator': _E('The network in a GAN that turns random noise into a sample.', 'primer.ml.generative.gans', scope=('primer/ml/generative/gans',)),
 683    'gini impurity': _E('The chance that two examples drawn at random from a pile carry different labels; 0 means pure.', 'primer.ml.classical'),
 684    'global token': _E('A token every other token may attend to, and which attends to all, so any two tokens are two hops apart.', 'primer.ml.efficient_architectures'),
 685    "goodhart's law": _E('Once a measure becomes a target, optimizing it stops improving the thing it measured.', 'primer.ml.alignment'),
 686    'gradient boosting': _E('Adding small trees one at a time, each fit to the current errors, with each step shrunk by a learning rate.', 'primer.ml.classical'),
 687    'gradient penalty': _E('A discriminator loss term that punishes steep slopes of its output with respect to its input (R1, WGAN-GP).', 'primer.ml.generative.gans'),
 688    'grammar': _E('A set of rules defining which strings are valid in a format.', 'primer.ml.structured_output', scope=('primer/ml/structured_output',)),
 689    'grpo': _E('Group Relative Policy Optimization: each answer is scored against other answers to the same prompt, so no value network is needed.', 'primer.ml.reinforcement'),
 690    'gsm8k': _E('A benchmark of grade-school maths word problems (1,319 in its test set), scored by comparing the final number with the answer key.', 'primer.ml.benchmarks'),
 691    'guidance scale': _E('The weight w in classifier-free guidance: 0 ignores the prompt, 1 follows it plainly, above 1 exaggerates it.', 'primer.ml.generative.diffusion'),
 692    'hann window': _E("A smooth rise and fall applied to each slice so its cut edges don't add false frequencies.", 'primer.ml.generative.multimodal'),
 693    'harmful compliance': _E('Answering a request that should have been refused.', 'primer.ml.alignment'),
 694    'harmonic mean': _E('n divided by the sum of the reciprocals of n numbers; dominated by the smallest, so it is high only when every number is high.', 'primer.ml.metrics'),
 695    'hidden tests': _E("Grading tests the agent never sees, so it can't pass by satisfying the tests instead of the intent.", 'primer.agents.coding_agents'),
 696    'host memory': _E("The CPU's main memory, reached from the GPU over a much slower link.", 'primer.ml.hardware'),
 697    'huber loss': _E('A loss that is squared for small errors and grows only in proportion to the error beyond a threshold δ, so a few wild outliers cannot dominate the fit.', 'primer.ml.losses'),
 698    'hybrid model': _E('A stack that mixes a few attention layers with many state-space layers.', 'primer.ml.efficient_architectures'),
 699    'hyperparameter': _E('A setting chosen before training rather than learned from data, such as the learning rate, the model size or the number of epochs.', 'primer.ml.optimizers'),
 700    'interaction effect': _E('Part of a prediction that depends on how two or more inputs combine, beyond what each contributes on its own: size mattering more in one city than another.', None),
 701    'imagenet': _E('A benchmark of about 1.3 million photos in 1,000 categories, the standard test of image classifiers; ImageNet-21k is its 14-million-photo, 21,000-category superset.', 'primer.ml.cnn_rnn'),
 702    'implicit generative model': _E('A model that makes samples by pushing random noise through a fixed procedure, without giving the probability of any sample: a GAN generator, or a diffusion model sampled with DDIM.', 'primer.ml.generative.gans'),
 703    'inception score': _E('A score for generated images from a pretrained image classifier: high when each image is confidently one class and the images spread over many classes. Higher is better; it never looks at real images.', 'primer.ml.generative.gans'),
 704    'inductive bias': _E("The assumptions built into a model before it sees any data, such as a convolution's belief that nearby pixels matter most. Good assumptions help with little data; with enough data a model can learn them instead.", 'primer.ml.cnn_rnn'),
 705    'importance sampling': _E('Estimating an average under one distribution from samples drawn under another, by weighting each sample by the ratio of its two probabilities.', 'primer.ml.reinforcement'),
 706    'inpainting': _E('Filling in a missing or masked part of an image so that it fits the rest; a generative model does it by sampling only the unknown pixels.', None),
 707    'information gain': _E('How much a split lowers entropy; the tree picks the split that lowers it most.', 'primer.ml.classical'),
 708    'induction head': _E('A pattern-completion mechanism: having seen A followed by B earlier in the text, predict B the next time A appears. It is thought to underlie much of in-context learning.', None),
 709    'instruction tuning': _E('Fine-tuning a pretrained model on many tasks written as instructions paired with good responses, so it learns to follow instructions it has never seen.', 'primer.ml.training_stages'),
 710    'interference': _E("Signals getting in each other's way: other features leaking into one feature's reading (superposition), or merged task vectors changing the same weights in opposite directions (model merging).", 'primer.ml.interpretability', scope=('primer/ml/fine_tuning', 'primer/ml/interpretability')),
 711    'interpretability': _E("Studying what a model's internal numbers represent, and which of them cause its outputs.", 'primer.ml.interpretability'),
 712    'item response theory': _E("Modelling the chance of a right answer from a taker's ability minus a question's difficulty.", 'primer.ml.benchmarks'),
 713    'jaccard similarity': _E('Shared items divided by all distinct items across two sets, from 0 to 1.', 'primer.ml.fine_tuning'),
 714    'jailbreak': _E('A prompt crafted to talk a model out of its safety training, such as a role-play or a disguised request, so it produces what it would normally refuse.', 'primer.agents.guardrails'),
 715    'jensen-shannon divergence': _E("A measure of how different two distributions are; stuck at log 2 whenever they don't overlap.", 'primer.ml.generative.gans'),
 716    'json mode': _E('A decoding option that guarantees parseable JSON, but not any particular shape.', 'primer.ml.structured_output'),
 717    'kl penalty': _E('A term subtracted from the reward that grows as the tuned model drifts from its starting point, measured by KL divergence: a leash that keeps it near the data the reward model was trained on.', 'primer.ml.reinforcement'),
 718    'kv-cache quantization': _E('Storing cached keys and values in fewer bits, with one scale per vector.', 'primer.ml.efficient_architectures'),
 719    'l0 norm': _E('The number of non-zero entries in a vector. For a sparse autoencoder, how many latents fire on an input.', None),
 720    'label noise': _E('Reference labels that are wrong, which cap what a model can learn and what an eval can measure.', 'primer.ml.fine_tuning'),
 721    'langevin dynamics': _E('Sampling by repeatedly taking a small step along the score, towards where data is denser, and adding a little fresh noise. A diffusion sampler has this shape.', 'primer.ml.generative.diffusion'),
 722    'language identification': _E('Guessing which language a text is in, so a pipeline keeps only the ones it wants.', 'primer.ml.pretraining'),
 723    'latent diffusion': _E("Running diffusion on an autoencoder's small compressed code instead of on pixels, then decoding once at the end.", 'primer.ml.generative.diffusion'),
 724    'latent space': _E('The space of codes a model works in, where position means something and nearby codes decode to similar data.', 'primer.ml.generative.autoencoders'),
 725    'law of large numbers': _E('The rule that the average of many independent random draws settles ever closer to the true average as more draws are added.', None),
 726    'layout shift': _E('Controls moving between runs, so clicks at remembered coordinates miss.', 'primer.agents.coding_agents'),
 727    'leaf': _E('An end box of a decision tree, predicting the majority label or average value of the training examples that reached it.', 'primer.ml.classical', scope=('primer/ml/classical',)),
 728    'least squares': _E('Choosing the parameters that make the sum of squared errors as small as possible.', 'primer.ml.regularization'),
 729    'linear time invariance': _E('A sequence model whose update rule is the same at every step, whatever the input; such a model can be computed as one convolution.', 'primer.ml.efficient_architectures'),
 730    'linear representation hypothesis': _E('The idea that features are directions, readable with a dot product.', 'primer.ml.interpretability'),
 731    'log-derivative trick': _E('Rewriting ∇π as π·∇log π, so a sampled action gives an estimate of the gradient.', 'primer.ml.reinforcement'),
 732    'lipschitz': _E('A function is K-Lipschitz if its output never changes more than K times as fast as its input: a speed limit on its slope everywhere.', 'primer.ml.generative.gans', display="Lipschitz"),
 733    'line search': _E('Having picked a direction to move in, trying different step lengths along it and keeping the one that lowers the loss most.', None),
 734    'log-odds': _E('The logarithm of p / (1 − p): 0 for a 50% chance, positive when more likely than not. A sigmoid turns log-odds back into a probability.', 'primer.ml.classical'),
 735    'log-likelihood': _E('The log of how probable the observed data is under a model.', 'primer.ml.benchmarks'),
 736    'log-mel spectrogram': _E('A spectrogram pooled into mel bands with loudness on a log scale: what speech models read.', 'primer.ml.generative.multimodal'),
 737    'logit difference': _E("The correct answer's score minus a wrong answer's score.", 'primer.ml.interpretability'),
 738    'logit lens': _E("Applying the model's own output layer to intermediate layers to see what it would predict so far.", 'primer.ml.interpretability'),
 739    'loss scaling': _E('Multiplying the loss so small fp16 gradients stay above zero, then dividing back.', 'primer.ml.hardware'),
 740    'loss spike': _E('A sudden jump in training loss that may recover or diverge.', 'primer.ml.pretraining'),
 741    'maj@k': _E('A score for sampling k answers per question and keeping the most common final answer: the share of questions where that majority answer is right.', 'primer.ml.reasoning'),
 742    'majority voting': _E('Choosing the answer that the most samples agree on.', 'primer.ml.reasoning'),
 743    'mantissa': _E('The bits of a float that store its digits, which set its precision.', 'primer.ml.hardware', scope=('primer/ml/hardware',)),
 744    'margin of error': _E('The ± range around a measured score that the true score probably falls in; it shrinks with the square root of the sample size.', 'primer.ml.fine_tuning'),
 745    'marginal likelihood': _E('How probable a model finds an example, averaged over every hidden cause that could have produced it. With a neural network inside the model that average is an intractable integral, which is why VAEs train on a lower bound instead.', 'primer.ml.generative.autoencoders'),
 746    'markov chain': _E('A sequence of random steps where each step depends only on the one just before it, not on the whole history.', 'primer.ml.generative.diffusion'),
 747    'master weights': _E("An fp32 copy of the weights that receives optimizer updates, so tiny updates aren't lost.", 'primer.ml.pretraining'),
 748    'mel scale': _E('A relabelling of frequency so equal steps sound equally far apart to human ears.', 'primer.ml.generative.multimodal'),
 749    'median': _E('The middle value once numbers are sorted (the average of the two middle ones for an even count). Unlike the mean, one extreme value barely moves it.', None),
 750    'memory bandwidth': _E('How many bytes per second a memory can deliver.', 'primer.ml.hardware'),
 751    'memory hierarchy': _E('The chain of storage from registers to the network, each level bigger and slower than the last.', 'primer.ml.hardware'),
 752    'micro-batch': _E('A slice of a batch that moves through a pipeline on its own.', 'primer.ml.pretraining', scope=('primer/ml/pretraining',)),
 753    'minhash': _E('A short signature per document whose matching slots estimate Jaccard similarity.', 'primer.ml.pretraining'),
 754    'minimax': _E('A game where one player tries to make a number as large as possible and the other as small as possible.', 'primer.ml.generative.gans'),
 755    'mixed precision': _E('Doing the big multiplies in 16- or 8-bit while keeping master weights and sums in fp32.', 'primer.ml.hardware'),
 756    'mlp': _E('Multi-layer perceptron: the plainest neural network, layers of weighted sums each followed by a nonlinearity, with every unit connected to every unit in the next layer.', 'primer.ml.neural_net'),
 757    'modality': _E('One kind of input a model can take: text, images, audio or video.', 'primer.ml.generative.multimodal', scope=('primer/ml/generative/multimodal',)),
 758    'model parallelism': _E("Splitting one model's weights across GPUs, either inside each layer (tensor parallelism) or by layers (pipeline parallelism). The ZeRO and Megatron-LM papers use it for the first.", 'primer.ml.pretraining'),
 759    'mode collapse': _E('A generator producing only a few kinds of output instead of the full variety in the data.', 'primer.ml.generative.gans'),
 760    'model collapse': _E('Losing rare data (the tails) when models are trained on their own outputs.', 'primer.ml.pretraining'),
 761    'model editing': _E('Changing one specific fact or behaviour inside a trained model by adjusting a few weights directly, without retraining, while leaving everything else as it was.', 'primer.ml.interpretability', display="model editing"),
 762    'model merging': _E('Building one model from several fine-tunes by arithmetic on their weights, with no extra training.', 'primer.ml.fine_tuning'),
 763    'monosemantic': _E('Responding to one understandable thing only: said of a neuron, or of a feature a sparse autoencoder finds.', 'primer.ml.interpretability'),
 764    'multi-armed bandit': _E('The simplest RL problem: pick among options with hidden payouts, and learn which pays best by trying them.', 'primer.ml.reinforcement'),
 765    'multi-head latent attention': _E("Caching one small latent vector per token and rebuilding every head's keys and values from it.", 'primer.ml.efficient_architectures'),
 766    'multi-query attention': _E('All query heads share a single key/value head.', 'primer.ml.efficient_architectures'),
 767    'multitask learning': _E('Training one model on several tasks at once, with part of the input saying which task to do, so what it learns for one task can help the others.', None),
 768    'multimodal model': _E('A model that takes in more than one kind of input (text, images, audio, video) as one sequence of tokens.', 'primer.ml.generative.multimodal'),
 769    'n-gram overlap': _E("The share of a text's n-word runs that also appear in another corpus.", 'primer.ml.benchmarks'),
 770    'neural ode': _E('A model whose output is found by following an ordinary differential equation whose rate of change is computed by a neural network; it can be run backwards and gives exact probabilities.', None),
 771    "newton's method": _E('A step that uses the curvature (second derivative) as well as the slope: divide the slope by the curvature to jump to the bottom of the parabola that matches the loss at the current point.', None),
 772    'noise schedule': _E('How much noise each step adds (the betas), which sets how fast the signal fades.', 'primer.ml.generative.diffusion'),
 773    'non-saturating loss': _E('The generator loss −log D(G(z)), which keeps a strong gradient when the discriminator confidently rejects fakes.', 'primer.ml.generative.gans'),
 774    'off-by-one error': _E('An index or count one position away from the right one.', 'primer.agents.coding_agents'),
 775    'one-hot encoding': _E('Turning a category into one yes/no column per possible value.', 'primer.ml.classical'),
 776    'optimizer state': _E("The running numbers an optimizer keeps for every weight between steps, such as Adam's two running averages. With an fp32 master copy of the weights it is 12 bytes per parameter, the biggest part of training memory.", 'primer.ml.pretraining'),
 777    'optimal discriminator': _E('p_data/(p_data + p_g): the best possible verdict against a fixed generator; 1/2 everywhere at equilibrium.', 'primer.ml.generative.gans'),
 778    'optimal transport': _E('Moving one distribution onto another as cheaply as possible, where moving mass further costs more. For the bell curves of flow matching, every bit of mass then travels in a straight line at constant speed.', 'primer.ml.generative.gans'),
 779    'ordinary differential equation': _E('A rule giving, at every moment, how fast something is changing. Solving it means following that rule forward in time, for example with small Euler steps.', 'primer.ml.generative.diffusion'),
 780    'out-of-bag': _E("The rows left out of a tree's bootstrap sample, usable as free validation data for that tree.", 'primer.ml.classical'),
 781    'outcome reward model': _E('A verifier that scores only the final answer.', 'primer.ml.reasoning'),
 782    'outcome supervision': _E('Training a reward model from the final result alone: each solution is labelled right or wrong by its answer, never step by step.', 'primer.ml.reasoning'),
 783    'outer product': _E('A column vector times a row vector, giving a table whose entry (m, c) is the product of their m-th and c-th numbers.', 'primer.ml.efficient_architectures'),
 784    'over-refusal': _E('Refusing a harmless request because a safety check is too strict.', 'primer.ml.alignment'),
 785    'overoptimization': _E('Optimizing against a learned reward model for so long that the true quality it stood for starts to fall, even as its score keeps rising: Goodhart\'s law for reward models.', 'primer.ml.alignment'),
 786    'overthinking': _E('Spending many reasoning tokens where few would do, wasting cost and time and sometimes losing a right answer.', 'primer.ml.reasoning'),
 787    'partial dependence plot': _E("A plot of a model's average prediction as one or two inputs are set to each value in turn, with every other input left as it is in the data.", None),
 788    'paired bootstrap': _E("Resampling questions with both models' results kept together, to test whether a gap is real.", 'primer.ml.benchmarks'),
 789    'parallel scan': _E('Computing every state of a linear recurrence in about log₂ n parallel rounds by merging steps.', 'primer.ml.efficient_architectures'),
 790    'pass-to-pass test': _E('A hidden test that passed before a fix and must still pass after it.', 'primer.agents.coding_agents'),
 791    'pass@k': _E('The chance that at least one of k sampled answers passes the tests.', 'primer.ml.benchmarks'),
 792    'pass@n': _E('The chance that at least one of n samples is right: 1 − (1 − p)ⁿ.', 'primer.ml.reasoning'),
 793    'patch embedding': _E('The learned matrix that turns a flattened image patch into a token vector.', 'primer.ml.generative.multimodal'),
 794    'perceiver resampler': _E('A small transformer whose queries are a fixed set of learned vectors: they cross-attend to any number of image or video features and always return that fixed number of visual tokens.', 'primer.ml.generative.multimodal'),
 795    'perceptual loss': _E('A loss that compares two images through the features of a pretrained network rather than pixel by pixel, so it punishes the differences a person would notice. LPIPS is a widely used one.', 'primer.ml.generative.gans'),
 796    'pearson correlation': _E('How closely two lists of numbers rise and fall together along a straight line, from −1 to 1; 0 means no straight-line relationship.', None),
 797    'permutation importance': _E('The accuracy lost on held-out data when one column is shuffled.', 'primer.ml.classical'),
 798    'pipeline bubble': _E('Time pipeline stages sit idle while the pipeline fills and drains.', 'primer.ml.pretraining'),
 799    'pipeline parallelism': _E('Giving each GPU a stage of layers and passing activations between neighbouring stages.', 'primer.ml.hardware'),
 800    'policy gradient': _E('Raising expected reward by making the actions that earned more reward more likely.', 'primer.ml.reinforcement'),
 801    'polysemantic neuron': _E('A neuron that responds to several unrelated features.', 'primer.ml.interpretability'),
 802    'posterior': _E('What you believe about a hidden quantity after seeing the data: the prior, reweighted by how well each value explains what was observed.', 'primer.ml.generative.autoencoders'),
 803    'posterior collapse': _E("When a VAE's code carries no information because the KL penalty outweighs what the code saves in rebuild error, so every output is the same average.", 'primer.ml.generative.autoencoders'),
 804    'power iteration': _E('Repeatedly multiplying a vector by a matrix (and its transpose) until it points along the most-stretched direction.', 'primer.ml.generative.gans'),
 805    'power law': _E('A relationship where one quantity is a fixed power of another, y = a·x^k; on log-log axes it is a straight line.', 'primer.ml.transformer'),
 806    'preference model': _E('Another name for a reward model: it scores a response so that the gap between two scores predicts which one people (or a model) prefer.', 'primer.ml.training_stages'),
 807    'prior': _E('What you believe about a hidden quantity before seeing any data. A VAE\'s prior over codes is the standard normal distribution.', 'primer.ml.generative.autoencoders', scope=('primer/ml/generative',)),
 808    'probability flow ode': _E('The deterministic equation whose solutions carry noise to data with the same in-between distributions as a diffusion process; DDIM sampling is one way of stepping along it.', 'primer.ml.generative.diffusion'),
 809    'privileged basis': _E('Directions made special by the architecture, such as neurons followed by an activation function that acts on each number separately. Only in a privileged basis does asking what one neuron means make sense.', 'primer.ml.interpretability'),
 810    'probability ratio': _E("The current policy's probability of a sampled action divided by its probability when the action was sampled.", 'primer.ml.reinforcement'),
 811    'probe': _E("A small classifier trained on a frozen model's activations to test whether a property is encoded there.", 'primer.ml.interpretability', scope=('primer/ml/interpretability',)),
 812    'process supervision': _E('Training a reward model from labels on each intermediate step, so it learns where a solution went wrong.', 'primer.ml.reasoning'),
 813    'process reward model': _E('A verifier that scores each intermediate step.', 'primer.ml.reasoning'),
 814    'projector': _E("A small layer that maps an encoder's vectors into a language model's embedding space.", 'primer.ml.generative.multimodal', scope=('primer/ml/generative/multimodal',)),
 815    'pushdown automaton': _E('A finite-state machine plus a stack: enough to check nested formats like JSON or SQL.', 'primer.ml.structured_output'),
 816    'quantile': _E('The value below which a given share of the data falls: the 0.5 quantile is the median, the 0.9 quantile has 90% of values below it.', None),
 817    'quality filter': _E('A rule or classifier that throws out low-quality pages before training.', 'primer.ml.pretraining'),
 818    'random forest': _E('Many deep trees on bootstrap samples, each split limited to a random subset of features, with their votes averaged.', 'primer.ml.classical'),
 819    'reasoning model': _E('A model trained to write out and check intermediate steps before it commits to an answer.', 'primer.ml.reasoning'),
 820    'rectified flow': _E('Flow matching with straight-line paths, retrained on its own outputs so the paths get straighter and need fewer steps.', 'primer.ml.generative.diffusion'),
 821    'reduce-scatter': _E('Summing a vector across GPUs so each GPU ends up holding the total for only its own slice; the first half of a ring all-reduce.', 'primer.ml.pretraining'),
 822    'reflow': _E("Retraining a rectified flow on its own (noise, sample) pairs. The new pairs' straight lines rarely cross, so the new flow's paths are straighter and need fewer steps.", 'primer.ml.generative.diffusion'),
 823    'red-teaming': _E('Searching systematically for inputs that make a model or its safety checks fail.', 'primer.ml.alignment'),
 824    'register': _E('The tiny storage right beside the arithmetic units, holding the numbers being worked on this instant.', 'primer.ml.hardware', scope=('primer/ml/hardware',)),
 825    'regular expression': _E('A pattern language (such as [0-9]+ or cat|car|dog) that can always be compiled into a finite-state machine.', 'primer.ml.structured_output'),
 826    'reinforce': _E('The basic policy-gradient algorithm: step along reward times the gradient of log π(action).', 'primer.ml.reinforcement'),
 827    'reinforcement learning': _E('Learning from a score for what you did, rather than from the correct answer.', 'primer.ml.reinforcement'),
 828    'rejection sampling': _E('Generating many candidate outputs and keeping only those that pass a check, such as a correct final answer, often to use as training data.', None),
 829    'release gate': _E('Limits set before measuring; a release goes ahead only if every evaluation is within its limit.', 'primer.ml.alignment'),
 830    'reparameterization trick': _E('Writing a random draw as z = μ + σ·ε with the noise ε as a separate input, so gradients can flow through sampling.', 'primer.ml.generative.autoencoders'),
 831    'replay': _E('Mixing a small sample of old-task examples into new training data so the old skill keeps getting practised.', 'primer.ml.fine_tuning', scope=('primer/ml/fine_tuning',)),
 832    'residual': _E('The true value minus the current prediction; for squared error it is the negative gradient.', 'primer.ml.classical', scope=('primer/ml/classical',)),
 833    'residual vector quantization': _E('Stacking codebooks, each one encoding the error the previous ones left.', 'primer.ml.generative.multimodal'),
 834    'resnet': _E('A deep convolutional network built from residual blocks (x + f(x)), which made networks of a hundred layers and more trainable.', 'primer.ml.deep_nets'),
 835    'resolved rate': _E('The share of tasks whose patch passes every hidden fail-to-pass and pass-to-pass test.', 'primer.agents.coding_agents'),
 836    'reverse process': _E('The learned half of a diffusion model: a chain of small denoising steps that turns pure noise back into data.', 'primer.ml.generative.diffusion', scope=('primer/ml/generative/diffusion',)),
 837    'reward': _E('The single number the environment returns to say how good an action was.', 'primer.ml.reinforcement', scope=('primer/ml/reinforcement',)),
 838    'reward hacking': _E('A policy maximising the reward as written while the real goal gets worse.', 'primer.ml.reinforcement'),
 839    'ring all-reduce': _E('All-reduce by passing chunks around a ring; each GPU sends about twice its data, however many GPUs there are.', 'primer.ml.hardware'),
 840    'rlaif': _E('Reinforcement learning from AI feedback: preference labels come from a model applying written principles.', 'primer.ml.alignment'),
 841    'rolling buffer cache': _E('A KV cache with w slots where each new token overwrites the oldest.', 'primer.ml.efficient_architectures'),
 842    'sample rate': _E('How many measurements of a signal are taken per second, such as 16,000 for speech.', 'primer.ml.generative.multimodal'),
 843    'sandbox': _E('An isolated place to run untrusted code, where the worst it can do is fail: no network, no secrets, time and memory limits.', 'primer.agents.coding_agents', scope=('primer/agents/coding_agents',)),
 844    'scaling law': _E("A smooth, predictable rule for how a model's loss falls as its size, data or compute grows: a straight line on log-log axes. It lets small runs forecast a big model's quality.", 'primer.ml.transformer'),
 845    'score': _E('The direction in which data gets more crowded fastest; the noise guess, flipped and rescaled.', 'primer.ml.generative.diffusion', scope=('primer/ml/generative/diffusion',)),
 846    'screenshot': _E('The image of the screen a computer-use agent receives after each action.', 'primer.agents.coding_agents', scope=('primer/agents/coding_agents',)),
 847    'selective state-space model': _E('A state-space model whose step size (how much to keep and write) is computed from each token, as in Mamba.', 'primer.ml.efficient_architectures'),
 848    'self-consistency': _E('Sampling several chains of thought and returning the most common final answer.', 'primer.ml.reasoning'),
 849    'self-supervised learning': _E('Learning from unlabelled data by predicting a hidden part of each example from the rest, such as a masked word or a masked stretch of audio.', None),
 850    'signal-to-noise ratio': _E("How much stronger a signal is than the noise mixed into it, as a ratio of their powers (variances). Audio quotes it in decibels, where every 10 dB is ten times the ratio; in diffusion it falls from very large (clean) to nearly zero (pure noise).", 'primer.ml.generative.diffusion'),
 851    'shingle': _E('A window of k neighbouring words, used to compare texts for near-duplicates.', 'primer.ml.fine_tuning'),
 852    'short-time fourier transform': _E('A Fourier transform on each short, overlapping slice of a signal.', 'primer.ml.generative.multimodal'),
 853    'shots': _E('Worked examples placed in the prompt before the question.', 'primer.ml.benchmarks', scope=('primer/ml/benchmarks',)),
 854    'shrinkage': _E("Pulling values toward zero: an L1 penalty's pull on every activation, or in gradient boosting the learning rate that keeps only part of each new tree's correction.", 'primer.ml.interpretability', scope=('primer/ml/classical', 'primer/ml/interpretability')),
 855    'sliding-window attention': _E('Each token attends only to the last w tokens, so cost and cache stop growing with context length.', 'primer.ml.efficient_architectures'),
 856    'softplus': _E('log(1 + eˣ): a smooth ramp that is always positive.', 'primer.ml.efficient_architectures'),
 857    'sparse attention': _E('Attention that scores only a chosen pattern of token pairs (local, global, strided) instead of every pair.', 'primer.ml.efficient_architectures'),
 858    'sparse autoencoder': _E('A wide encoder and decoder trained with an L1 penalty so each input uses a few learned features.', 'primer.ml.interpretability'),
 859    'specification gaming': _E('Another name for reward hacking: satisfying the letter of an objective but not its intent.', 'primer.ml.reinforcement'),
 860    'spectral normalization': _E("Dividing each layer's weights by the most they can stretch any input, which caps how fast the discriminator can change.", 'primer.ml.generative.gans'),
 861    'spectrogram': _E('A picture of sound: time across, frequency up, brightness for loudness.', 'primer.ml.generative.multimodal'),
 862    'speech recognition': _E('Turning recorded speech into written text; also called automatic speech recognition (ASR).', 'primer.ml.generative.multimodal'),
 863    'standard error': _E('The typical distance between a measured score and the true rate: √(p(1−p)/n).', 'primer.ml.benchmarks'),
 864    'standard normal distribution': _E('The bell curve centred on 0 with spread 1; N(0, I) draws each number from it independently.', 'primer.ml.generative.autoencoders'),
 865    'steering': _E("Changing a model's behaviour while it runs by editing an internal activation, such as adding or pinning a feature's direction.", 'primer.ml.interpretability', scope=('primer/ml/interpretability',)),
 866    'stochastic differential equation': _E('A rule for how something changes over time with two parts: a steady drift, and random jitter of a set strength. The noising process of a diffusion model is one, and running it backwards generates data.', 'primer.ml.generative.diffusion'),
 867    'stop-gradient': _E('An operation that passes its input through unchanged but blocks gradients from flowing back into it, so training treats that input as a constant.', None),
 868    'straight-through estimator': _E('A way to train through a step that has no useful gradient, such as rounding or snapping to a codebook: use the step going forward, and pass the gradient back as if the step were not there.', None),
 869    'strict mode': _E("A tool or output option that guarantees the model's JSON fits a given schema.", 'primer.ml.structured_output'),
 870    'structured output': _E("Making a model's answer follow an exact format, such as JSON that fits a schema, so a program can read it.", 'primer.ml.structured_output'),
 871    'stump': _E('A decision tree with a single question and two leaves.', 'primer.ml.classical', scope=('primer/ml/classical',)),
 872    'style control': _E('Adding length or format as extra factors in the rating fit, so style is separated from quality.', 'primer.ml.benchmarks'),
 873    'subnormal': _E('A float below the smallest normal value, with the hidden leading 1 dropped so it fades toward zero.', 'primer.ml.hardware'),
 874    'superposition': _E('Storing more features than there are neurons, as nearly perpendicular directions.', 'primer.ml.interpretability'),
 875    'super-resolution': _E('Turning a low-resolution image into a plausible higher-resolution one; most of the fine detail has to be invented, not recovered.'),
 876    'swiglu': _E('A gated feed-forward layer: one projection, passed through the smooth SiLU activation, multiplies a second projection number by number before the output projection.', 'primer.ml.transformer'),
 877    'sycophancy': _E('A model changing its answer to agree with a view the user states.', 'primer.ml.alignment'),
 878    'synthetic data': _E('Training examples written by a model rather than collected from people.', 'primer.ml.pretraining'),
 879    'tabular data': _E('Data in rows and columns, where each column is a meaningful quantity in its own units.', 'primer.ml.classical'),
 880    'task arithmetic': _E("Combining or removing skills by adding, scaling or subtracting task vectors from a base model's weights.", 'primer.ml.fine_tuning'),
 881    'task vector': _E("The change a fine-tune made to a model's weights: fine-tuned minus base.", 'primer.ml.fine_tuning'),
 882    'tensor parallelism': _E('Splitting each matrix multiply across GPUs, which must talk inside every layer.', 'primer.ml.hardware'),
 883    'test-time compute': _E('Computation spent while answering rather than while training, such as longer chains or more samples.', 'primer.ml.reasoning'),
 884    'text extraction': _E("Pulling a web page's main text out of its HTML, leaving menus and adverts behind.", 'primer.ml.pretraining'),
 885    'text normalization': _E('Rewriting text into one standard form (case, punctuation, contractions, how numbers are written) so two texts are compared on their words, not their style.', None),
 886    'thinking budget': _E('The maximum number of tokens a model may spend reasoning before it must answer.', 'primer.ml.reasoning'),
 887    'throughput': _E('Work finished per second, such as tokens generated per second across all requests; it rises with batch size, while latency is the wait for one request.', 'primer.ml.inference'),
 888    'tiling': _E('Loading a block of data into fast memory once and doing all its work before evicting it.', 'primer.ml.hardware', scope=('primer/ml/hardware',)),
 889    'token healing': _E('Backing up over the last prompt token so the model can rewrite it, when a prompt ends partway through what would normally be one token.', 'primer.ml.structured_output'),
 890    'transfer learning': _E('Training a model on a large general task first, then reusing it (usually by fine-tuning) on a smaller task it was never trained for.', 'primer.ml.training_stages'),
 891    'translation equivariance': _E('Shift the input and the output shifts the same way: a convolution finds an edge wherever it sits, because the same filter slides over every position.', 'primer.ml.cnn_rnn'),
 892    'tubelet': _E('A video patch that spans several frames as well as a square of pixels.', 'primer.ml.generative.multimodal'),
 893    'tuned lens': _E('A logit lens with a small learned translator per layer.', 'primer.ml.interpretability'),
 894    'two time-scale update rule': _E('Giving the generator and discriminator different learning rates so the game converges.', 'primer.ml.generative.gans'),
 895    'u-net': _E('A convolutional network shaped like a U: it shrinks an image step by step to see the big picture, then grows it back, with shortcuts carrying fine detail across. The classic diffusion denoiser before transformers.', 'primer.ml.generative.diffusion'),
 896    'underflow': _E('A number too small for its format, rounded to zero.', 'primer.ml.pretraining', scope=('primer/ml/pretraining',)),
 897    'unbiased estimator': _E('A way of estimating a number from random data whose average, over every dataset you might have drawn, equals the true value: it is not systematically high or low.', 'primer.ml.benchmarks'),
 898    'unit test': _E('A small program that runs one piece of code on chosen inputs and checks the outputs, passing or failing automatically.', 'primer.agents.coding_agents'),
 899    'unbiased estimate': _E('An estimate that is right on average: any single one may be off, but the errors cancel over many tries.', 'primer.ml.reinforcement'),
 900    'unembedding': _E('The final matrix of a language model: it turns the last hidden vector into one score (logit) per vocabulary token.', 'primer.ml.interpretability'),
 901    'vae': _E('Variational autoencoder: an autoencoder whose codes are pulled towards a bell curve, so random codes decode to new data.', 'primer.ml.generative.autoencoders'),
 902    'value function': _E('A prediction, made partway through a task, of how well it will end from here. A verifier that scores a solution after every token is one.', 'primer.ml.reinforcement', scope=('primer/ml/reinforcement',)),
 903    'value network': _E("A second model (the critic) that predicts expected reward, used as PPO's baseline.", 'primer.ml.reinforcement'),
 904    'variational autoencoder': _E('An autoencoder whose encoder outputs a fuzzy region (mean and spread) pulled towards the standard normal, so random codes decode to new data.', 'primer.ml.generative.autoencoders'),
 905    'variational inference': _E('Approximating a distribution you cannot compute, usually a posterior, with the closest member of a simple family by maximising a lower bound. A VAE does it with a network that outputs the approximation for each example.', 'primer.ml.generative.autoencoders'),
 906    'vector quantization': _E('Replacing a vector with the number of its nearest codebook entry.', 'primer.ml.generative.multimodal'),
 907    'velocity field': _E('A map giving, at every point and time, which way and how fast a sample should move.', 'primer.ml.generative.diffusion'),
 908    'verifiable reward': _E("A reward computed by a check that can't be argued with, such as a correct answer or passing tests.", 'primer.ml.reinforcement'),
 909    'verifier': _E('Anything that scores a candidate solution, from a unit test to a learned model.', 'primer.ml.reasoning', scope=('primer/ml/reasoning',)),
 910    'vision-language model': _E('A model that reads images (and often video) together with text and writes text; a language model given eyes.', 'primer.ml.generative.multimodal'),
 911    'visual instruction tuning': _E('Fine-tuning a vision-language model on images paired with instructions and good answers.', 'primer.ml.generative.multimodal'),
 912    'voice activity detection': _E('Deciding which stretches of a recording contain speech at all, so silence, music and noise are not transcribed.', None),
 913    'vq-vae': _E('A VAE that snaps each code vector to the nearest entry of a learned codebook, turning images or audio into tokens.', 'primer.ml.generative.autoencoders'),
 914    'wasserstein distance': _E("The least work to reshape one distribution into another (mass moved times distance); also called earth mover's distance.", 'primer.ml.generative.gans'),
 915    'waveform': _E('Sound recorded as a list of air-pressure measurements over time.', 'primer.ml.generative.multimodal'),
 916    'weak supervision': _E('Training on labels that are plentiful but noisy or imperfect, such as captions and transcripts found on the web, instead of a small set checked by experts.', None),
 917    'weight averaging': _E('Merging fine-tunes of the same base by averaging their weights (task arithmetic with λ = 1/T).', 'primer.ml.fine_tuning'),
 918    'wiener process': _E('Continuous random jitter, also called Brownian motion: over any short time dt it moves by a fresh bell-curve amount with variance dt, independent of everything before.', None),
 919    'weight clipping': _E("Forcing every weight of a network back into a small range, such as −0.01 to 0.01, after each update; the original Wasserstein GAN's crude way to cap its critic's slope.", 'primer.ml.generative.gans'),
 920    "winner's curse": _E('The best of many versions chosen on a test looks better on it than it really is.', 'primer.ml.benchmarks'),
 921    'word error rate': _E("The share of a reference transcript's words that a system gets wrong: the substitutions, deletions and insertions needed to turn its output into the reference, divided by the reference's word count.", None),
 922    'zero-order hold': _E('A discretization rule that assumes the input stays constant for the whole step; it gives the keep factor e^(ΔA) of a state-space model.', 'primer.ml.efficient_architectures'),
 923    'zero': _E('Sharding optimizer state, then gradients, then weights across data-parallel GPUs.', 'primer.ml.pretraining', scope=('primer/ml/pretraining',)),
 924    'β-vae': _E('A VAE whose KL penalty is weighted by β, trading rebuild sharpness for a smoother, more organized code space.', 'primer.ml.generative.autoencoders'),
 925}
 926
 927
 928@functools.lru_cache(maxsize=1)
 929def _usage_text() -> str:
 930    """Every lesson's prose and every paper companion's, as the case of each term is judged from it."""
 931    import ast
 932    import re
 933    from pathlib import Path
 934
 935    # Read the docstrings from the source, without importing the lessons.
 936    root = Path(__file__).resolve().parent
 937    text = "\n".join(ast.get_docstring(ast.parse(f.read_text(encoding="utf-8"))) or "" for f in sorted(root.rglob("*.py")))
 938    # Companions use terms no lesson says mid-sentence ("a Markov chain"). Keep only their running prose:
 939    # headings, table cells, captions, citations, italic paper titles, buttons and link text capitalise for other reasons.
 940    skip = r"head|script|style|pre|code|svg|h[1-6]|th|td|caption|figcaption|cite|summary|button|label|a|dt|em|i"
 941    for page in sorted((root.parent / "docs" / "papers").glob("*.html")):
 942        html = re.sub(rf"<({skip})\b[^>]*>.*?</\1>", "\n", page.read_text(encoding="utf-8"), flags=re.S | re.I)
 943        html = re.sub(r"</?(p|li|div|section|blockquote|br|tr|table|ul|ol|figure)\b[^>]*>", "\n", html, flags=re.I)
 944        text += "\n" + re.sub(r"<[^>]+>", "", html)
 945    text = re.sub(r"```.*?```", " ", text, flags=re.S)  # fenced code and diagram labels
 946    # Inline code, italic titles (which may wrap onto a second line) and link text; then citation lines.
 947    text = re.sub(r"`[^`\n]*`|(?<!\*)\*(?!\s)[^*]+?(?<!\s)\*(?!\*)|\[[^\]\n]*\]\([^)]*\)", " ", text)
 948    return "\n".join(line for line in text.splitlines() if "http" not in line and "arxiv" not in line.lower())
 949
 950
 951@functools.lru_cache(maxsize=None)
 952def _usage(terms: tuple[str, ...]) -> dict[str, Counter]:
 953    """How often each casing of each term appears mid-sentence, found in one pass over the text."""
 954    import re
 955
 956    # Only mid-sentence uses count: after a lower-case word and a space, or an opening bracket.
 957    # Headings, table cells, list items and sentence starts capitalise words for reasons that aren't the term's.
 958    alternation = "|".join(re.escape(t) for t in sorted(terms, key=len, reverse=True))
 959    pattern = re.compile(rf"(?:(?<=[a-z0-9,;)] )|(?<=\())(?:{alternation})(?![\w-])", re.I)
 960    counts: dict[str, Counter] = {}
 961    for m in pattern.finditer(_usage_text()):
 962        counts.setdefault(m.group(0).lower(), Counter())[m.group(0)] += 1
 963    return counts
 964
 965
 966def display_term(term: str) -> str:
 967    """How the primer usually writes a term: "acl" -> "ACL", "adamw" -> "AdamW", "attention" -> "attention".
 968
 969    Counted over every lesson's and companion's text, mid-sentence only, where capitals mean something
 970    about the word. A term the lessons never use mid-sentence keeps its key, and an
 971    entry's own `display` overrides the count.
 972    """
 973    if term in GLOSSARY and GLOSSARY[term].display:
 974        return GLOSSARY[term].display
 975    terms = tuple(sorted(GLOSSARY)) if term in GLOSSARY else (term,)
 976    forms = _usage(terms).get(term.lower())
 977    return forms.most_common(1)[0][0] if forms else term
 978
 979
 980def lesson_path(dotted: str) -> str:
 981    """"primer.ml.attention" -> "primer/ml/attention.html", the pdoc page for that module."""
 982    return dotted.replace(".", "/") + ".html"
 983
 984
 985def glossary_js() -> str:
 986    """The glossary as a script that sets `window.PRIMER_GLOSSARY` (works from file:// too)."""
 987    data = {
 988        term: {"def": e.definition, "lesson": lesson_path(e.lesson) if e.lesson else None, "term": display_term(term)}
 989        for term, e in sorted(GLOSSARY.items())
 990    }
 991    return "window.PRIMER_GLOSSARY = " + json.dumps(data, ensure_ascii=False, indent=1) + ";\n"
 992
 993
 994def _render_doc() -> str:
 995    rows = "\n".join(
 996        f"| **{display_term(term)}** | {e.definition} | {f'`{e.lesson}`' if e.lesson else ''} |" for term, e in sorted(GLOSSARY.items())
 997    )
 998    return __doc__ + "\n| Term | Meaning | Taught in |\n|---|---|---|\n" + rows + "\n"
 999
1000
1001__doc__ = _render_doc()
@dataclass(frozen=True)
class Entry: on GitHub
21@dataclass(frozen=True)
22class Entry:
23    definition: str
24    lesson: str | None = None  # dotted module that teaches it, e.g. "primer.ml.attention"
25    # Everyday words with a technical meaning ("value", "policy", "patch") are only
26    # given hover definitions on pages under these path prefixes, so a "travel
27    # policy" never pops up a definition from preference tuning.
28    scope: tuple[str, ...] = ()
29    # How readers see the term when counting the lessons' usage can't decide it (see display_term).
30    display: str | None = None
Entry( definition: str, lesson: str | None = None, scope: tuple[str, ...] = (), display: str | None = None)
definition: str
lesson: str | None = None
scope: tuple[str, ...] = ()
display: str | None = None
GLOSSARY: dict[str, Entry] = {'vector': Entry(definition='A list of numbers, like (3, 1, 2). In AI, a word, sentence or image is represented as a vector.', lesson='primer.notation', scope=(), display=None), 'matrix': Entry(definition='A table of numbers with rows and columns. A batch of vectors stacked as rows is a matrix, and most model weights are matrices.', lesson='primer.notation', scope=(), display=None), 'tensor': Entry(definition='A grid of numbers with any number of dimensions: a vector is 1-D, a matrix 2-D, a batch of images 4-D.', lesson='primer.notation', scope=(), display=None), 'dot product': Entry(definition='Multiply two lists position by position and add up the products. It is large when two vectors point the same way.', lesson='primer.notation', scope=(), display=None), 'matrix multiply': Entry(definition='A grid of dot products: each output cell is one row of the first matrix dotted with one column of the second.', lesson='primer.notation', scope=(), display=None), 'transpose': Entry(definition='Flip a matrix so its rows become columns. Written with a superscript T.', lesson='primer.notation', scope=(), display=None), 'norm': Entry(definition='The length of a vector: square the entries, add them, take the square root.', lesson='primer.notation', scope=(), display=None), 'unit vector': Entry(definition='A vector rescaled to length 1, so it only carries a direction.', lesson='primer.notation', scope=(), display=None), 'orthogonal': Entry(definition='At right angles: two vectors whose dot product, and so whose cosine similarity, is zero, so moving along one does not move you along the other.', lesson='primer.ml.embeddings.similarity', scope=(), display=None), 'logarithm': Entry(definition='The undo button for exponentiation: ln(y) asks what power of e gives y. Logs turn multiplication into addition.', lesson='primer.notation', scope=(), display=None), 'variance': Entry(definition='How widely numbers are spread around their average: the average squared distance from the mean.', lesson='primer.notation', scope=(), display=None), 'standard deviation': Entry(definition='The square root of the variance: the typical distance of a value from the average.', lesson='primer.notation', scope=(), display=None), 'derivative': Entry(definition='The slope of a function at a point: how much the output changes per tiny nudge of the input.', lesson='primer.notation', scope=(), display=None), 'gradient': Entry(definition='One slope per input, collected into a vector. It points in the direction that increases the function fastest.', lesson='primer.notation', scope=(), display=None), 'chain rule': Entry(definition='To get the slope through a chain of functions, multiply the slopes of each link.', lesson='primer.notation', scope=(), display=None), 'argmax': Entry(definition='The position of the largest value in a list, rather than the value itself.', lesson='primer.notation', scope=(), display=None), 'big-o': Entry(definition='A way to say how cost grows with input size, ignoring constant factors. O(n²) means doubling n quadruples the cost.', lesson='primer.notation', scope=(), display=None), 'probability distribution': Entry(definition='A list of probabilities over all possible outcomes, each between 0 and 1, adding up to 1.', lesson='primer.notation', scope=(), display=None), 'probability density': Entry(definition="How thickly a continuous distribution's samples cover each spot: high where they crowd, zero where none ever land. It is the height of the bump, and its area adds up to 1.", lesson='primer.ml.generative.gans', scope=(), display=None), 'neuron': Entry(definition='Multiply each input by a weight, add them up with a bias, and pass the result through a nonlinear function.', lesson='primer.ml.neural_net', scope=(), display=None), 'weight': Entry(definition='A learned number that says how strongly one input influences an output. Training adjusts the weights.', lesson='primer.ml.neural_net', scope=(), display=None), 'bias': Entry(definition="A learned number added to a neuron's weighted sum, letting it shift its output up or down.", lesson='primer.ml.neural_net', scope=('primer/ml/neural_net', 'primer/ml/deep_nets'), display=None), 'parameters': Entry(definition='All the learned numbers in a model (its weights and biases). A 70B model has 70 billion of them.', lesson='primer.ml.neural_net', scope=(), display=None), 'activation function': Entry(definition='The nonlinear function inside a neuron, such as ReLU or GELU. Without it, stacked layers collapse into one.', lesson='primer.ml.neural_net', scope=(), display=None), 'relu': Entry(definition='An activation function that keeps positive numbers and turns negatives into zero.', lesson='primer.ml.neural_net', scope=(), display=None), 'gelu': Entry(definition='A smooth version of ReLU used in most transformers.', lesson='primer.ml.neural_net', scope=(), display=None), 'sigmoid': Entry(definition='Squashes any number into the range 0 to 1, with an S-shaped curve.', lesson='primer.ml.neural_net', scope=(), display=None), 'softmax': Entry(definition='Turns a list of scores into shares that are all positive and add up to 1, with bigger scores getting disproportionately more.', lesson='primer.ml.attention', scope=(), display=None), 'forward pass': Entry(definition='Running inputs through the network, layer by layer, to get a prediction.', lesson='primer.ml.neural_net', scope=(), display=None), 'backpropagation': Entry(definition='The chain rule run backwards through the network to find, for every weight, how much nudging it would change the loss.', lesson='primer.ml.neural_net', scope=(), display=None), 'multilayer perceptron': Entry(definition='A network of fully connected layers, each a matrix multiply followed by a nonlinearity; also called an MLP.', lesson='primer.ml.neural_net', scope=(), display=None), 'loss': Entry(definition="A single number measuring how wrong the model's predictions are. Training tries to make it smaller.", lesson='primer.ml.losses', scope=(), display=None), 'loss function': Entry(definition='The rule that turns predictions and correct answers into the loss. Choosing it defines what the model learns.', lesson='primer.ml.losses', scope=(), display=None), 'gradient descent': Entry(definition='Training by repeatedly taking a small step downhill: nudge every weight against its gradient to lower the loss.', lesson='primer.ml.optimizers', scope=(), display=None), 'gradient ascent': Entry(definition='Nudging every weight along its gradient to make a quantity larger: the mirror image of gradient descent. Used to maximise a reward, or, run on a loss, to make a model worse at a task on purpose.', lesson='primer.ml.reinforcement', scope=(), display=None), 'learning rate': Entry(definition='How big a step each training update takes. Too high and training blows up; too low and it crawls.', lesson='primer.ml.optimizers', scope=(), display=None), 'optimizer': Entry(definition='The rule that turns gradients into weight updates, such as SGD, momentum or Adam.', lesson='primer.ml.optimizers', scope=(), display=None), 'adam': Entry(definition='An optimizer that adapts the step size for each weight using running averages of its gradients.', lesson='primer.ml.optimizers', scope=(), display=None), 'adamw': Entry(definition='Adam with weight decay applied separately from the gradient step. The default optimizer for transformers.', lesson='primer.ml.optimizers', scope=(), display=None), 'momentum': Entry(definition='Keeping a running average of recent gradients so updates roll smoothly through noise, like a ball gathering speed downhill.', lesson='primer.ml.optimizers', scope=(), display=None), 'batch': Entry(definition='The group of examples processed before one weight update.', lesson='primer.ml.neural_net', scope=(), display=None), 'epoch': Entry(definition='One full pass over the training data.', lesson='primer.ml.neural_net', scope=(), display=None), 'warmup': Entry(definition='Starting training with a tiny learning rate and ramping it up, which keeps the first updates from destabilizing the model.', lesson='primer.ml.optimizers', scope=(), display=None), 'vanishing gradient': Entry(definition='When gradients shrink towards zero as they pass back through many layers, so early layers stop learning.', lesson='primer.ml.deep_nets', scope=(), display=None), 'exploding gradient': Entry(definition='When gradients grow huge as they pass back through many layers, so training diverges.', lesson='primer.ml.deep_nets', scope=(), display=None), 'residual connection': Entry(definition="Adding a layer's input back to its output (x + f(x)), giving gradients a shortcut through deep networks.", lesson='primer.ml.deep_nets', scope=(), display=None), 'layer normalization': Entry(definition="Rescaling each token's numbers to a steady average and spread, which keeps training stable.", lesson='primer.ml.deep_nets', scope=(), display=None), 'batch normalization': Entry(definition='Rescaling each feature using statistics from the current batch of examples. Common in image models.', lesson='primer.ml.deep_nets', scope=(), display=None), 'gradient clipping': Entry(definition="Capping the size of the gradient so one bad batch can't throw the weights far off course.", lesson='primer.ml.deep_nets', scope=(), display=None), 'attention': Entry(definition='The mechanism that lets each token look at every other token, score how relevant each one is, and blend in information from the relevant ones.', lesson='primer.ml.attention', scope=(), display=None), 'self-attention': Entry(definition='Attention where a sequence attends to itself: every token scores every other token in the same text.', lesson='primer.ml.attention', scope=(), display=None), 'query': Entry(definition='In attention, the vector a token uses to ask what it is looking for.', lesson='primer.ml.attention', scope=('primer/ml/attention', 'primer/ml/transformer'), display=None), 'key': Entry(definition='In attention, the vector a token offers so others can decide whether it is relevant to them.', lesson='primer.ml.attention', scope=('primer/ml/attention', 'primer/ml/transformer', 'primer/ml/inference', 'primer/ml/big_picture'), display=None), 'value': Entry(definition='In attention, the information a token hands over when others attend to it.', lesson='primer.ml.attention', scope=('primer/ml/attention', 'primer/ml/transformer', 'primer/ml/inference', 'primer/ml/big_picture'), display=None), 'multi-head attention': Entry(definition='Several attention computations run in parallel on slices of the vectors, so each head can track a different kind of relationship.', lesson='primer.ml.attention', scope=(), display=None), 'grouped-query attention': Entry(definition='Several query heads share one set of keys and values, shrinking the memory needed during generation.', lesson='primer.ml.attention', scope=(), display=None), 'causal mask': Entry(definition="Hides future tokens from each token during attention, so a model predicting the next word can't peek at it.", lesson='primer.ml.attention', scope=(), display=None), 'transformer': Entry(definition='The neural network architecture behind modern language models: stacked blocks of attention followed by a small feed-forward network.', lesson='primer.ml.transformer', scope=(), display=None), 'feed-forward network': Entry(definition='The small two-layer network in each transformer block that processes every token on its own after attention has mixed them.', lesson='primer.ml.transformer', scope=(), display=None), 'encoder': Entry(definition='A transformer that reads the whole input at once, with every token seeing every other. Used for embeddings and classification.', lesson='primer.ml.transformer', scope=(), display=None), 'decoder': Entry(definition='A transformer that generates text left to right, each token seeing only the ones before it. GPT, Claude and Llama are decoders.', lesson='primer.ml.transformer', scope=(), display=None), 'mixture of experts': Entry(definition='Replacing one feed-forward network with many expert networks and a router that sends each token to a few of them.', lesson='primer.ml.transformer', scope=(), display=None), 'positional encoding': Entry(definition='Information added to token vectors so the model knows word order, which attention alone ignores.', lesson='primer.ml.positional', scope=(), display=None), 'rope': Entry(definition='Rotary position embedding: encodes position by rotating query and key vectors, so attention depends on the distance between tokens.', lesson='primer.ml.positional', scope=(), display=None), 'context window': Entry(definition='The maximum number of tokens a model can take in at once.', lesson='primer.ml.inference', scope=(), display=None), 'rmsnorm': Entry(definition="A cheaper layer normalization that divides each token's numbers by their root-mean-square without subtracting the mean first.", lesson='primer.ml.transformer', scope=(), display=None), 'pre-norm': Entry(definition='Normalizing before each sub-layer of a transformer block, leaving the residual path untouched, which trains more stably than normalizing after.', lesson='primer.ml.transformer', scope=(), display=None), 'residual stream': Entry(definition='The running vector for each token that every transformer block reads from and adds a correction to, instead of replacing it.', lesson='primer.ml.transformer', scope=(), display=None), 'weight tying': Entry(definition='Reusing the token embedding table as the output layer, so the vector that reads a token in also scores it on the way out.', lesson='primer.ml.transformer', scope=(), display=None), 'encoder-decoder': Entry(definition='A transformer where an encoder reads the whole input and a decoder writes the output one token at a time while attending to the encoder.', lesson='primer.ml.transformer', scope=(), display=None), 'load-balancing loss': Entry(definition='A small extra training penalty that pushes a mixture-of-experts router to spread tokens evenly across experts.', lesson='primer.ml.transformer', scope=(), display=None), 'flops': Entry(definition='Floating-point operations: individual multiplies or adds, used to measure compute (about 2 per parameter per token to generate, 6 to train).', lesson='primer.ml.transformer', scope=(), display=None), 'permutation equivariance': Entry(definition="Shuffling a layer's inputs just shuffles its outputs the same way, which is why attention alone cannot see word order.", lesson='primer.ml.positional', scope=(), display=None), 'sinusoidal positional encoding': Entry(definition="The original transformer's fixed position codes, made of sines and cosines at many frequencies, added to each token's embedding.", lesson='primer.ml.positional', scope=(), display=None), 'position interpolation': Entry(definition='Stretching a model to longer contexts by scaling every position down into the range it saw during training.', lesson='primer.ml.positional', scope=(), display=None), 'ntk-aware scaling': Entry(definition="Extending a RoPE model's context by raising the rotation base, which slows the slow frequencies while keeping the fast ones that encode local order.", lesson='primer.ml.positional', scope=(), display=None), 'yarn': Entry(definition='A refinement of RoPE context extension that rescales each frequency band differently, used by many long-context models.', lesson='primer.ml.positional', scope=(), display=None), 'alibi': Entry(definition='A position scheme that adds no position vectors and instead subtracts a penalty proportional to distance from each attention score.', lesson='primer.ml.positional', scope=(), display=None), 'autoregressive generation': Entry(definition='Producing text one token at a time, where each new token is predicted from all the tokens before it and then appended.', lesson='primer.ml.big_picture', scope=(), display=None), 'greedy decoding': Entry(definition='Always picking the single most probable next token, the same as sampling at temperature 0.', lesson='primer.ml.big_picture', scope=(), display=None), 'logits': Entry(definition='The raw scores a model outputs for every possible next token, before softmax turns them into probabilities.', lesson='primer.ml.big_picture', scope=(), display=None), 'token': Entry(definition="A piece of text from a model's fixed vocabulary: often a whole common word, or a fragment of a rare one. Models read and write tokens, not words.", lesson='primer.ml.tokenization', scope=(), display=None), 'tokenizer': Entry(definition='The component that splits text into tokens and maps each to an integer ID.', lesson='primer.ml.tokenization', scope=(), display=None), 'special token': Entry(definition='A token reserved for structure rather than text, such as a turn marker or an end-of-text marker. The model learns when to emit it, and generation stops at the end marker.', lesson='primer.ml.fine_tuning', scope=(), display=None), 'vocabulary': Entry(definition='The fixed set of tokens a model knows, typically 32,000 to 200,000 entries.', lesson='primer.ml.tokenization', scope=(), display=None), 'byte pair encoding': Entry(definition='A tokenizer-building method that starts from characters and repeatedly merges the most frequent adjacent pair into a new token.', lesson='primer.ml.tokenization', scope=(), display=None), 'byte-level bpe': Entry(definition='Byte pair encoding run on the raw bytes of UTF-8 text, so any string can be tokenized and there is never an unknown token.', lesson='primer.ml.tokenization', scope=(), display=None), 'pre-tokenization': Entry(definition='Splitting text into chunks, such as words with their leading space, before BPE runs, so merges never cross chunk boundaries.', lesson='primer.ml.tokenization', scope=(), display=None), 'wordpiece': Entry(definition="BERT's subword tokenizer, which merges the pair whose parts most rarely appear apart rather than simply the most frequent pair.", lesson='primer.ml.tokenization', scope=(), display=None), 'unigram tokenizer': Entry(definition='A subword tokenizer that starts from a huge vocabulary and repeatedly removes the pieces whose loss hurts least.', lesson='primer.ml.tokenization', scope=(), display=None), 'sentencepiece': Entry(definition='A tokenizer library that runs BPE or Unigram directly on raw text, writing spaces as the visible symbol ▁.', lesson='primer.ml.tokenization', scope=(), display=None), 'utf-8': Entry(definition='The standard way to store text as bytes: 1 byte for basic Latin letters, 2 to 4 bytes for other characters and emoji.', lesson='primer.ml.tokenization', scope=(), display=None), 'bpe': Entry(definition='Byte pair encoding: build a vocabulary by repeatedly merging the most frequent adjacent pair of symbols.', lesson='primer.ml.tokenization', scope=(), display=None), 'pretraining': Entry(definition='The first, most expensive training stage: predicting the next token over trillions of tokens of text.', lesson='primer.ml.training_stages', scope=(), display=None), 'base model': Entry(definition='A model after pretraining only: knowledgeable, but it continues text rather than following instructions.', lesson='primer.ml.training_stages', scope=(), display=None), 'fine-tuning': Entry(definition='Training an existing model further on a smaller, targeted dataset to change its behavior.', lesson='primer.ml.training_stages', scope=(), display=None), 'weight space': Entry(definition="The space of every possible setting of a model's weights: one axis per weight, one point per model. Fine-tuning moves a model from one point to another.", lesson='primer.ml.fine_tuning', scope=(), display=None), 'sft': Entry(definition='Supervised fine-tuning: training on examples of instructions paired with good responses.', lesson='primer.ml.training_stages', scope=(), display=None), 'rlhf': Entry(definition='Reinforcement learning from human feedback: a reward model learns human preferences, and the language model is tuned to score well on it.', lesson='primer.ml.training_stages', scope=(), display=None), 'dpo': Entry(definition='Direct preference optimization: tuning a model straight from pairs of preferred and rejected answers, without a separate reward model.', lesson='primer.ml.training_stages', scope=(), display=None), 'reward model': Entry(definition='A model trained to predict which of two responses a human would prefer.', lesson='primer.ml.training_stages', scope=(), display=None), 'lora': Entry(definition='Low-rank adaptation: freeze the model and train a small pair of matrices whose product is added to a weight matrix.', lesson='primer.ml.training_stages', scope=(), display=None), 'distillation': Entry(definition="Training a small model to imitate a large one's outputs, often the biggest cost saving in production.", lesson='primer.ml.training_stages', scope=(), display=None), 'inference': Entry(definition='Using a trained model to make predictions, as opposed to training it.', lesson='primer.ml.inference', scope=(), display=None), 'prefill': Entry(definition='The first phase of generation: the whole prompt is processed in one parallel pass. It sets the time to the first token.', lesson='primer.ml.inference', scope=(), display=None), 'decode': Entry(definition='The second phase of generation: output tokens are produced one at a time. It sets the tokens per second.', lesson='primer.ml.inference', scope=('primer/ml/inference',), display=None), 'kv cache': Entry(definition='Stored keys and values from earlier tokens, so each new token is computed without redoing work. It trades GPU memory for speed.', lesson='primer.ml.inference', scope=(), display=None), 'temperature': Entry(definition='A knob that sharpens (low) or flattens (high) the probability distribution before sampling a token.', lesson='primer.ml.inference', scope=(), display=None), 'top-p': Entry(definition='Nucleus sampling: pick only from the smallest set of tokens whose probabilities add up to p.', lesson='primer.ml.inference', scope=(), display=None), 'top-k': Entry(definition='Sampling only from the k most likely next tokens.', lesson='primer.ml.inference', scope=(), display=None), 'quantization': Entry(definition='Storing numbers with fewer bits (say 8 or 4 instead of 16 or 32), which shrinks memory with a small loss of precision.', lesson='primer.ml.inference', scope=(), display=None), 'speculative decoding': Entry(definition='A small fast model drafts several tokens, and the big model checks them all in one pass, keeping the ones it agrees with.', lesson='primer.ml.inference', scope=(), display=None), 'prompt caching': Entry(definition='Reusing the processed form of a prompt prefix that repeats across requests, which cuts cost and time to first token.', lesson='primer.ml.inference', scope=(), display=None), 'cross-entropy': Entry(definition='The standard loss for predicting categories: minus the log of the probability given to the right answer.', lesson='primer.ml.losses', scope=(), display=None), 'perplexity': Entry(definition='e raised to the average cross-entropy: roughly how many options the model is torn between at each step.', lesson='primer.ml.losses', scope=(), display=None), 'contrastive loss': Entry(definition='A loss that pulls matching pairs together in vector space and pushes non-matching pairs apart.', lesson='primer.ml.losses', scope=(), display=None), 'infonce': Entry(definition='A contrastive loss that treats the other items in the batch as wrong answers for each matching pair.', lesson='primer.ml.losses', scope=(), display=None), 'overfitting': Entry(definition='When a model memorizes its training data, noise included, and does worse on new data.', lesson='primer.ml.regularization', scope=(), display=None), 'underfitting': Entry(definition='When a model is too simple, or undertrained, to capture the pattern at all.', lesson='primer.ml.regularization', scope=(), display=None), 'regularization': Entry(definition='Any technique that discourages memorizing, such as dropout, weight decay or early stopping.', lesson='primer.ml.regularization', scope=(), display=None), 'dropout': Entry(definition="Randomly switching off neurons during training so the network can't rely on any single path.", lesson='primer.ml.regularization', scope=(), display=None), 'weight decay': Entry(definition='Shrinking every weight slightly at each step, which penalizes large weights and keeps the model smoother.', lesson='primer.ml.regularization', scope=(), display=None), 'early stopping': Entry(definition='Stopping training when the score on held-out data stops improving.', lesson='primer.ml.regularization', scope=(), display=None), 'data leakage': Entry(definition='When information from the test set, or from the future, sneaks into training and inflates offline scores.', lesson='primer.ml.regularization', scope=(), display=None), 'out-of-distribution': Entry(definition='Unlike the examples a model learned from or was shown, such as longer or rarer inputs. Performance there is usually lower and harder to predict.', lesson=None, scope=(), display=None), 'emergent ability': Entry(definition="A skill that is near chance in smaller models and appears sharply once a model is large enough, so it can't be predicted by extending the small models' trend.", lesson=None, scope=(), display=None), 'convolution': Entry(definition='Sliding a small grid of learned weights across an image and computing a weighted sum at every position, to find where a pattern appears.', lesson='primer.ml.cnn_rnn', scope=(), display=None), 'pooling': Entry(definition='Shrinking a feature map by keeping only the strongest (or average) value in each small window.', lesson='primer.ml.cnn_rnn', scope=(), display=None), 'rnn': Entry(definition='A recurrent neural network: it reads a sequence one step at a time, carrying a running summary called the hidden state.', lesson='primer.ml.cnn_rnn', scope=(), display=None), 'lstm': Entry(definition='An RNN with gates that decide what to forget, what to write and what to output, so it can remember across long sequences.', lesson='primer.ml.cnn_rnn', scope=(), display=None), 'precision': Entry(definition='Of the items flagged, the fraction that were right.', lesson='primer.ml.metrics', scope=(), display=None), 'recall': Entry(definition='Of the items that truly mattered, the fraction that were found.', lesson='primer.ml.metrics', scope=('primer/ml/metrics', 'primer/ml/embeddings/', 'primer/agents/rag', 'primer/agents/evals'), display=None), 'f1': Entry(definition='The harmonic mean of precision and recall: high only when both are high.', lesson='primer.ml.metrics', scope=(), display=None), 'roc-auc': Entry(definition='The probability that the model ranks a random positive above a random negative.', lesson='primer.ml.metrics', scope=(), display=None), 'recall@k': Entry(definition='The fraction of relevant documents that appear in the top k results. For RAG, usually the metric that matters most.', lesson='primer.ml.metrics', scope=(), display=None), 'mrr': Entry(definition='Mean reciprocal rank: the average of 1 / (position of the first relevant result).', lesson='primer.ml.metrics', scope=(), display=None), 'ndcg': Entry(definition='A ranking score that gives more credit for relevant results near the top and handles degrees of relevance.', lesson='primer.ml.metrics', scope=(), display=None), 'bleu': Entry(definition='A score that counts overlapping word sequences with a reference text. Cheap, but blind to paraphrase.', lesson='primer.ml.metrics', scope=(), display=None), 'ablation': Entry(definition='Removing or replacing one part of a method and measuring again, to find out which part the gains come from.', lesson=None, scope=(), display=None), 'inter-annotator agreement': Entry(definition='How often two people labelling the same items independently give the same label. It caps how finely any evaluation built on those labels can tell two models apart.', lesson='primer.ml.fine_tuning', scope=(), display=None), 'likert scale': Entry(definition='A rating on a short fixed ladder of labelled points, such as 1 (not helpful) to 6 (highly helpful), averaged over many items to compare systems.', lesson='primer.agents.evals', scope=(), display=None), 'calibration': Entry(definition="How well a model's confidence matches how often it is right: a calibrated model is right about 80% of the time when it says 80%.", lesson='primer.ml.losses', scope=('primer/ml/losses', 'primer/ml/alignment'), display=None), 'embedding': Entry(definition='A learned vector for a piece of content, arranged so that similar meanings end up close together.', lesson='primer.ml.embeddings.word2vec', scope=(), display=None), 'cosine similarity': Entry(definition='How closely two vectors point in the same direction, from −1 (opposite) to 1 (identical), ignoring their lengths.', lesson='primer.ml.embeddings.similarity', scope=(), display=None), 'euclidean distance': Entry(definition='The straight-line distance between two vectors.', lesson='primer.ml.embeddings.similarity', scope=(), display=None), 'anisotropy': Entry(definition="When a model's vectors all crowd into a narrow cone, so even unrelated texts score as fairly similar.", lesson='primer.ml.embeddings.similarity', scope=(), display=None), 'curse of dimensionality': Entry(definition='In very high dimensions, random points are all nearly the same distance apart, which makes exact nearest-neighbor search slow and fragile.', lesson='primer.ml.embeddings.similarity', scope=(), display=None), 'hard negative': Entry(definition="A training example that looks relevant but isn't, such as the right topic with the wrong answer. Training on them teaches fine distinctions.", lesson='primer.ml.embeddings.contrastive', scope=(), display=None), 'bi-encoder': Entry(definition='Embeds the query and each document separately, so document vectors can be computed once and searched fast.', lesson='primer.ml.embeddings.retrieval', scope=(), display=None), 'cross-encoder': Entry(definition='Reads the query and one document together and outputs a relevance score. Accurate but slow, so it is used to rerank a shortlist.', lesson='primer.ml.embeddings.retrieval', scope=(), display=None), 'reranker': Entry(definition='A second, more accurate model that reorders the top results of a fast first-stage search.', lesson='primer.ml.embeddings.retrieval', scope=(), display=None), 'clip': Entry(definition='A model that trains an image encoder and a text encoder together so pictures and their captions land near each other in one vector space.', lesson='primer.ml.embeddings.contrastive', scope=('primer/ml/embeddings/contrastive', 'primer/ml/generative', 'primer/ml/losses'), display='CLIP'), 'matryoshka embedding': Entry(definition='An embedding trained so its first few dimensions work as a smaller embedding on their own.', lesson='primer.ml.embeddings.compression', scope=(), display=None), 'approximate nearest neighbor': Entry(definition='Finding vectors close to a query quickly by searching only part of the collection, accepting a small chance of missing the true closest.', lesson='primer.ml.embeddings.ann', scope=(), display=None), 'ann': Entry(definition='Approximate nearest neighbor search: trade a little recall for a lot of speed.', lesson='primer.ml.embeddings.ann', scope=(), display=None), 'hnsw': Entry(definition='A layered graph index for vector search: long jumps on sparse upper layers, then a careful local search at the bottom.', lesson='primer.ml.embeddings.ann', scope=(), display=None), 'ivf': Entry(definition='A vector index that clusters vectors ahead of time and searches only the clusters nearest the query.', lesson='primer.ml.embeddings.ann', scope=(), display=None), 'product quantization': Entry(definition='Compressing a vector by splitting it into chunks and replacing each chunk with the ID of its nearest entry in a small codebook.', lesson='primer.ml.embeddings.ann', scope=(), display=None), 'bm25': Entry(definition="The classic keyword-search score: rewards documents containing the query's rarer words, with diminishing returns for repeats.", lesson='primer.ml.embeddings.retrieval', scope=(), display=None), 'hybrid search': Entry(definition='Running keyword search and vector search together and merging the two ranked lists.', lesson='primer.ml.embeddings.retrieval', scope=(), display=None), 'reciprocal rank fusion': Entry(definition='Merging ranked lists by giving each document 1 / (k + rank) from every list it appears in, then adding those up.', lesson='primer.ml.embeddings.retrieval', scope=(), display=None), 'chunking': Entry(definition='Splitting documents into passages before embedding them, so retrieval can return just the relevant part.', lesson='primer.ml.embeddings.retrieval', scope=(), display=None), 'colbert': Entry(definition='A retrieval model that keeps one vector per token and matches each query token to its best document token.', lesson='primer.ml.embeddings.retrieval', scope=(), display=None), 'k-means': Entry(definition='Clustering by repeatedly assigning each point to its nearest center and moving each center to the average of its points.', lesson='primer.ml.embeddings.clustering', scope=(), display=None), 'semantic cache': Entry(definition="Reusing a stored answer when a new question's embedding is nearly identical to a previous question's.", lesson='primer.ml.embeddings.clustering', scope=(), display=None), 'llm': Entry(definition='Large language model: a transformer trained to predict the next token, then tuned to follow instructions.', lesson='primer.agents.llm', scope=(), display=None), 'prompt': Entry(definition='The text sent to a language model: instructions, context and the question.', lesson='primer.agents.llm', scope=(), display=None), 'system prompt': Entry(definition="Standing instructions sent before the conversation that set a model's role, rules and style.", lesson='primer.agents.llm', scope=(), display=None), 'tool calling': Entry(definition='The model outputs a structured request to run a function; your code runs it and sends the result back. The model never executes anything itself.', lesson='primer.agents.llm', scope=(), display=None), 'agent': Entry(definition='A system where a language model decides which tools to call, looks at the results and decides the next step, in a loop until done.', lesson='primer.agents.agent_loop', scope=(), display=None), 'agent loop': Entry(definition='The cycle an agent repeats: think, call a tool, observe the result, decide what next.', lesson='primer.agents.agent_loop', scope=(), display=None), 'react': Entry(definition='Reason plus act: the agent pattern that interleaves reasoning steps with tool calls.', lesson='primer.agents.agent_loop', scope=(), display=None), 'workflow': Entry(definition='A fixed sequence of steps written in code, with a language model doing a task inside each step.', lesson='primer.agents.orchestration', scope=(), display=None), 'router': Entry(definition='A step that classifies a request and sends it down the right path, or to the right model.', lesson='primer.agents.orchestration', scope=(), display=None), 'multi-agent system': Entry(definition='Several agents that coordinate, typically a supervisor delegating to specialist workers.', lesson='primer.agents.orchestration', scope=(), display=None), 'json schema': Entry(definition='A standard way to describe the shape of JSON data: which fields exist, their types and which are required.', lesson='primer.agents.tools', scope=(), display=None), 'idempotent': Entry(definition="Safe to repeat: doing it twice has the same effect as doing it once, so retries can't create duplicates.", lesson='primer.agents.tools', scope=(), display=None), 'mcp': Entry(definition='Model Context Protocol: an open standard for connecting AI apps to tools and data through reusable servers.', lesson='primer.agents.mcp', scope=(), display=None), 'rag': Entry(definition='Retrieval-augmented generation: fetch relevant passages first, then have the model answer using them, with citations.', lesson='primer.agents.rag', scope=(), display=None), 'hyde': Entry(definition='Hypothetical document embeddings: have the model write a plausible answer, then search with that, since answers resemble documents more than questions do.', lesson='primer.agents.rag', scope=(), display=None), 'grounding': Entry(definition='Tying an answer to specific source passages so every claim can be checked.', lesson='primer.agents.rag', scope=(), display=None), 'hallucination': Entry(definition="A confident answer that isn't supported by the sources or the facts.", lesson='primer.agents.rag', scope=(), display=None), 'context engineering': Entry(definition="Deciding exactly what goes into the model's context window on each call: instructions, facts, tool results, memory.", lesson='primer.agents.context', scope=(), display=None), 'context rot': Entry(definition='Quality dropping as a long session fills the context with stale or irrelevant material.', lesson='primer.agents.context', scope=(), display=None), 'episodic memory': Entry(definition="Memory of what happened, such as 'last week the user rejected this vendor'.", lesson='primer.agents.memory', scope=(), display=None), 'semantic memory': Entry(definition="Memory of facts, such as 'the user's fiscal year starts in April'.", lesson='primer.agents.memory', scope=(), display=None), 'procedural memory': Entry(definition='Memory of how to do things, such as a learned workflow or saved skill.', lesson='primer.agents.memory', scope=(), display=None), 'plan-and-execute': Entry(definition='An agent pattern that writes a plan first, then executes it step by step, replanning when results surprise it.', lesson='primer.agents.planning', scope=(), display=None), 'reflection': Entry(definition='Having a model review and revise its own output. Useful, but external checks such as tests are more reliable.', lesson='primer.agents.planning', scope=(), display=None), 'compounding error': Entry(definition='Small per-step failure rates multiplying over many steps: ten steps at 95% succeed only about 60% of the time.', lesson='primer.agents.planning', scope=(), display=None), 'eval': Entry(definition="A repeatable test of a model or agent's behavior on a fixed set of tasks with known good outcomes.", lesson='primer.agents.evals', scope=(), display=None), 'golden set': Entry(definition='A curated set of real tasks with expected outcomes, run on every change to catch regressions.', lesson='primer.agents.evals', scope=(), display=None), 'llm-as-judge': Entry(definition='Using a language model with a rubric to grade open-ended outputs, checked against human ratings.', lesson='primer.agents.evals', scope=(), display=None), 'guardrail': Entry(definition='A check around the model that screens inputs, validates outputs or limits actions.', lesson='primer.agents.guardrails', scope=(), display=None), 'prompt injection': Entry(definition='Hostile instructions hidden in content the model reads, such as an email or web page, trying to hijack it.', lesson='primer.agents.guardrails', scope=(), display=None), 'pii': Entry(definition='Personally identifiable information: names, emails, phone numbers, card numbers and similar.', lesson='primer.agents.guardrails', scope=(), display=None), 'model routing': Entry(definition='Sending each request to the cheapest model that can handle it well.', lesson='primer.agents.cost', scope=(), display=None), 'trace': Entry(definition='A record of one run as a tree of steps (model calls, tool calls, retrievals) with inputs, outputs, timing and cost.', lesson='primer.agents.observability', scope=('primer/agents/observability', 'primer/agents/evals'), display=None), 'span': Entry(definition='One step inside a trace, with its start time, duration and details.', lesson='primer.agents.observability', scope=('primer/agents/observability',), display=None), 'shadow mode': Entry(definition='Running a new system on real inputs and recording what it would do, without letting it act.', lesson='primer.agents.deployment', scope=(), display=None), 'canary release': Entry(definition='Sending a small share of traffic to a new version first and watching its metrics before rolling out further.', lesson='primer.agents.deployment', scope=(), display=None), 'kill switch': Entry(definition='A way to stop an agent instantly, per customer or globally.', lesson='primer.agents.deployment', scope=(), display=None), 'audit trail': Entry(definition='A tamper-evident log of which agent did what, for whom and with what authority.', lesson='primer.agents.deployment', scope=(), display=None), 'accuracy': Entry(definition='The share of all predictions that were correct. Misleading when one class is rare.', lesson='primer.ml.metrics', scope=(), display=None), 'agent-computer interface': Entry(definition='The commands an agent can call and the exact form of what comes back to it, designed for a language model the way a code editor is designed for a person.', lesson='primer.agents.coding_agents', scope=(), display=None), 'confusion matrix': Entry(definition='The four counts behind every classification metric: hits, false alarms, misses and correct passes.', lesson='primer.ml.metrics', scope=(), display=None), 'container': Entry(definition="An isolated copy of an operating system's user space for one program: its own files, processes and network settings, started from an image and thrown away afterwards.", lesson='primer.agents.coding_agents', scope=('primer/agents/coding_agents',), display=None), 'diff': Entry(definition='A text listing, file by file, which lines to remove (marked −) and which to add (marked +), with a few unchanged lines around each change so a tool can find the spot. Saved to a file, it is a patch that can be applied to another copy of the code.', lesson='primer.agents.coding_agents', scope=(), display=None), 'fault localization': Entry(definition='Finding which files, functions and lines cause a bug, before trying to fix it.', lesson='primer.agents.coding_agents', scope=(), display=None), 'linter': Entry(definition='A program that reads code without running it and reports likely mistakes, such as a syntax error, bad indentation or a name that is never defined.', lesson=None, scope=(), display=None), 'percentage point': Entry(definition='The plain difference between two percentages: going from 11% to 18% is a rise of 7 percentage points, which is a 64% relative rise.', lesson=None, scope=(), display=None), 'pull request': Entry(definition='A proposed set of changes to a code repository, submitted for review; once merged, the changes become part of the project.', lesson='primer.agents.coding_agents', scope=(), display=None), 'roc curve': Entry(definition='True-positive rate plotted against false-positive rate as the decision threshold sweeps.', lesson='primer.ml.metrics', scope=(), display=None), 'true-positive rate': Entry(definition='The share of real positives the model flags; the same as recall.', lesson='primer.ml.metrics', scope=(), display=None), 'false-positive rate': Entry(definition='The share of real negatives the model wrongly flags.', lesson='primer.ml.metrics', scope=(), display=None), 'precision-recall curve': Entry(definition='Precision plotted against recall across thresholds; more honest than ROC for rare events.', lesson='primer.ml.metrics', scope=(), display=None), 'precision@k': Entry(definition='The share of the top k search results that are relevant.', lesson='primer.ml.metrics', scope=(), display=None), 'rouge-l': Entry(definition='An overlap score based on the longest in-order sequence of words shared with a reference.', lesson='primer.ml.metrics', scope=(), display=None), 'bertscore': Entry(definition='Compares generated and reference text by embedding similarity instead of exact words.', lesson='primer.ml.metrics', scope=(), display=None), "cohen's kappa": Entry(definition='Agreement between two raters after subtracting the agreement expected by chance.', lesson='primer.ml.metrics', scope=(), display=None), 'loss mask': Entry(definition='A per-position switch deciding which predictions count toward the loss, such as only the reply in SFT.', lesson='primer.ml.training_stages', scope=(), display=None), 'preference tuning': Entry(definition='Training on which of two responses people preferred, to shape tone, helpfulness and safety.', lesson='primer.ml.training_stages', scope=(), display=None), 'bradley-terry model': Entry(definition='The rule that the probability A beats B is the sigmoid of their score difference.', lesson='primer.ml.training_stages', scope=(), display=None), 'reference model': Entry(definition='A frozen copy of the starting model that DPO and RLHF measure drift against.', lesson='primer.ml.training_stages', scope=(), display=None), 'policy': Entry(definition='In reinforcement learning and preference tuning, the model being trained, viewed as a probability distribution over actions or responses.', lesson='primer.ml.reinforcement', scope=('primer/ml/training_stages', 'primer/ml/reinforcement', 'primer/ml/reasoning', 'primer/ml/alignment'), display=None), 'implicit reward': Entry(definition='In DPO, β times the log of how much more likely training has made a response than the reference model did.', lesson='primer.ml.training_stages', scope=(), display=None), 'log-probability': Entry(definition="The logarithm of a probability; for a whole response, the sum of its tokens' log-probabilities.", lesson='primer.ml.training_stages', scope=(), display=None), 'qlora': Entry(definition='LoRA adapters trained on top of base weights stored in 4 bits.', lesson='primer.ml.training_stages', scope=(), display=None), 'adapter': Entry(definition='A small trainable module added beside frozen weights to specialise a model.', lesson='primer.ml.training_stages', scope=('primer/ml/training_stages',), display=None), 'full fine-tune': Entry(definition='Updating every weight of a model. Rarely worth the cost.', lesson='primer.ml.training_stages', scope=(), display=None), 'soft targets': Entry(definition="A teacher model's temperature-softened probabilities, used to train a student model.", lesson='primer.ml.training_stages', scope=(), display=None), 'kl divergence': Entry(definition='The extra surprise from using distribution q when the truth is p; zero only when they match.', lesson='primer.ml.training_stages', scope=(), display=None), 'dark knowledge': Entry(definition='How a teacher model ranks the wrong answers, revealed by its soft targets.', lesson='primer.ml.training_stages', scope=(), display=None), 'arithmetic intensity': Entry(definition='Operations performed per byte read from memory.', lesson='primer.ml.inference', scope=(), display=None), 'roofline': Entry(definition='A chart of the speed a chip can reach at each arithmetic intensity, capped first by memory and then by compute.', lesson='primer.ml.inference', scope=(), display=None), 'memory-bound': Entry(definition='Limited by how fast data arrives from memory rather than by arithmetic, as in decoding.', lesson='primer.ml.inference', scope=(), display=None), 'compute-bound': Entry(definition='Limited by arithmetic throughput, as in prefill.', lesson='primer.ml.inference', scope=(), display=None), 'draft model': Entry(definition='The small, fast model that proposes tokens in speculative decoding.', lesson='primer.ml.inference', scope=(), display=None), 'acceptance rate': Entry(definition='How often the big model keeps a drafted token in speculative decoding.', lesson='primer.ml.inference', scope=(), display=None), 'int8 quantization': Entry(definition='Storing each weight as an 8-bit integer times a shared scale.', lesson='primer.ml.inference', scope=(), display=None), 'int4 quantization': Entry(definition='Storing each weight as a 4-bit integer times a shared scale.', lesson='primer.ml.inference', scope=(), display=None), 'per-channel quantization': Entry(definition="One scale per weight row, so a single outlier doesn't coarsen all the others.", lesson='primer.ml.inference', scope=(), display=None), 'static batching': Entry(definition='Serving a fixed group of requests until the longest one finishes.', lesson='primer.ml.inference', scope=(), display=None), 'continuous batching': Entry(definition="Refilling a finished request's slot in the batch immediately with the next request.", lesson='primer.ml.inference', scope=(), display=None), 'pagedattention': Entry(definition='Storing the KV cache in fixed-size pages, like virtual memory, to avoid wasted GPU memory.', lesson='primer.ml.inference', scope=(), display=None), 'filter': Entry(definition='In a CNN, the small grid of learned weights a convolution slides over an image; also called a kernel.', lesson='primer.ml.cnn_rnn', scope=('primer/ml/cnn_rnn',), display=None), 'kernel': Entry(definition='In a CNN, the small grid of learned weights a convolution slides over an image; also called a filter.', lesson='primer.ml.cnn_rnn', scope=('primer/ml/cnn_rnn',), display=None), 'stride': Entry(definition='How many pixels a convolution filter jumps between positions.', lesson='primer.ml.cnn_rnn', scope=('primer/ml/cnn_rnn',), display=None), 'padding': Entry(definition='Zeros added around an image so a filter can centre on the edge pixels.', lesson='primer.ml.cnn_rnn', scope=('primer/ml/cnn_rnn',), display=None), 'max pooling': Entry(definition='Keeping only the largest value in each small block of a feature map.', lesson='primer.ml.cnn_rnn', scope=(), display=None), 'weight sharing': Entry(definition="Reusing the same filter at every position, so the parameter count doesn't grow with image size.", lesson='primer.ml.cnn_rnn', scope=(), display=None), 'receptive field': Entry(definition='How much of the original input one neuron can see.', lesson='primer.ml.cnn_rnn', scope=(), display=None), 'vision transformer': Entry(definition='A transformer that treats small image patches as tokens.', lesson='primer.ml.cnn_rnn', scope=(), display=None), 'patch': Entry(definition='A small square cut from an image and flattened into one token.', lesson='primer.ml.cnn_rnn', scope=('primer/ml/cnn_rnn',), display=None), 'backpropagation through time': Entry(definition='Backpropagation applied to an RNN unrolled across its time steps.', lesson='primer.ml.cnn_rnn', scope=(), display=None), 'cell state': Entry(definition="An LSTM's notebook, edited by adding rather than overwriting, so information survives many steps.", lesson='primer.ml.cnn_rnn', scope=(), display=None), 'forget gate': Entry(definition='The LSTM dial that decides what to erase from the cell state.', lesson='primer.ml.cnn_rnn', scope=(), display=None), 'input gate': Entry(definition='The LSTM dial that decides what new information to write into the cell state.', lesson='primer.ml.cnn_rnn', scope=(), display=None), 'output gate': Entry(definition='The LSTM dial that decides how much of the cell state to show as output.', lesson='primer.ml.cnn_rnn', scope=(), display=None), 'gru': Entry(definition='A simpler gated RNN with an update gate and a reset gate and no separate cell state.', lesson='primer.ml.cnn_rnn', scope=(), display=None), 'state-space model': Entry(definition='A recurrent-style model, such as Mamba, that trains in parallel and runs in time linear in sequence length.', lesson='primer.ml.efficient_architectures', scope=(), display=None), 'mamba': Entry(definition='A state-space model that trains in parallel and runs in time linear in sequence length.', lesson='primer.ml.efficient_architectures', scope=(), display=None), 'tanh': Entry(definition='Squashes any number into the range −1 to 1.', lesson='primer.ml.cnn_rnn', scope=(), display=None), 'tool call': Entry(definition='A structured request from the model to run one of your functions with specific arguments; your code decides whether to run it.', lesson='primer.agents.tools', scope=(), display=None), 'tool_result': Entry(definition="The message block that returns a tool's output or error to the model, matched to its request by id.", lesson='primer.agents.agent_loop', scope=(), display=None), 'stop reason': Entry(definition='Why a model response ended: finished, wants a tool, hit the length limit, or declined.', lesson='primer.agents.agent_loop', scope=(), display=None), 'loop detection': Entry(definition='Stopping an agent that keeps requesting the same tool with the same arguments.', lesson='primer.agents.agent_loop', scope=(), display=None), 'token budget': Entry(definition='A cap on the tokens, and so the money, one agent run may spend.', lesson='primer.agents.agent_loop', scope=(), display=None), 'parallel tool calls': Entry(definition='Several tool requests in one model turn, run at the same time, with all results returned in one message.', lesson='primer.agents.agent_loop', scope=(), display=None), 'strict tool use': Entry(definition="An API setting that guarantees the model's tool arguments match the tool's JSON Schema exactly.", lesson='primer.agents.tools', scope=(), display=None), 'semantic validation': Entry(definition='Checking that well-formed arguments are also true, such as that a customer id actually exists.', lesson='primer.agents.tools', scope=(), display=None), 'idempotency key': Entry(definition='A unique id sent with a write so a retried request returns the first result instead of acting twice.', lesson='primer.agents.tools', scope=(), display=None), 'dry run': Entry(definition='Running every check for an action and describing it without actually doing it.', lesson='primer.agents.tools', scope=(), display=None), 'human approval gate': Entry(definition='A rule that parks irreversible or high-value actions until a person approves them.', lesson='primer.agents.tools', scope=(), display=None), 'least privilege': Entry(definition='Giving each agent or tool only the permissions its job needs.', lesson='primer.agents.tools', scope=(), display=None), 'dynamic tool loading': Entry(definition='Sending the model only the few tool definitions most relevant to the current request.', lesson='primer.agents.tools', scope=(), display=None), 'model context protocol': Entry(definition='An open standard that lets any AI application connect to any tool server the same way.', lesson='primer.agents.mcp', scope=(), display=None), 'json-rpc': Entry(definition='A simple convention for calling functions on another program by exchanging JSON requests, replies and notifications.', lesson='primer.agents.mcp', scope=(), display=None), 'stdio transport': Entry(definition='Running a server as a child process and exchanging one JSON message per line over its input and output.', lesson='primer.agents.mcp', scope=(), display=None), 'mcp server': Entry(definition='The program that wraps a real system (files, a database, an API) and offers it to AI apps over MCP.', lesson='primer.agents.mcp', scope=(), display=None), 'mcp resource': Entry(definition='Read-only data an MCP server offers for the app to load into context, addressed by a URI.', lesson='primer.agents.mcp', scope=(), display=None), 'capability negotiation': Entry(definition='The start-of-session exchange where client and server announce which optional features they support.', lesson='primer.agents.mcp', scope=(), display=None), 'tool poisoning': Entry(definition="Hiding instructions for the model inside a tool's description.", lesson='primer.agents.mcp', scope=(), display=None), 'rug pull': Entry(definition='A tool server changing its definitions after they were approved.', lesson='primer.agents.mcp', scope=(), display=None), 'confused deputy': Entry(definition='A program with broad authority tricked into using it for someone who lacks that authority.', lesson='primer.agents.mcp', scope=(), display=None), 'oauth': Entry(definition='A standard way for a user to grant an app limited, revocable access without sharing their password.', lesson='primer.agents.mcp', scope=(), display=None), 'decomposition': Entry(definition='Splitting a big goal into small subtasks, each with a checkable definition of done.', lesson='primer.agents.planning', scope=(), display=None), 'external verification': Entry(definition="Checking a model's output with something outside the model, such as tests, schemas or database queries.", lesson='primer.agents.planning', scope=(), display=None), 'checkpoint': Entry(definition='Saved progress after a completed step, so a failure resumes from there instead of from the start.', lesson='primer.agents.planning', scope=('primer/agents/planning', 'primer/agents/orchestration'), display=None), 'durable execution': Entry(definition="Running a workflow so its progress survives crashes, by storing each completed step's result.", lesson='primer.agents.orchestration', scope=(), display=None), 'state machine': Entry(definition='A fixed set of states and allowed transitions, with code deciding every move.', lesson='primer.agents.orchestration', scope=(), display=None), 'short-term memory': Entry(definition='The current conversation, kept within a token budget: recent turns verbatim, older ones summarized.', lesson='primer.agents.memory', scope=(), display=None), 'long-term memory': Entry(definition='Facts, events and procedures stored outside the model and recalled into later conversations.', lesson='primer.agents.memory', scope=(), display=None), 'multi-tenant': Entry(definition='One system serving many separate customers whose data must never mix.', lesson='primer.agents.memory', scope=(), display=None), 'tenant isolation': Entry(definition="Partitioning storage so a request can only ever reach its own tenant's and user's data.", lesson='primer.agents.memory', scope=(), display=None), 'right to erasure': Entry(definition="A user's legal right, for example under GDPR, to have their personal data deleted.", lesson='primer.agents.memory', scope=(), display=None), 'prompt chaining': Entry(definition='A fixed sequence of model calls where each output feeds the next, with code checks in between.', lesson='primer.agents.orchestration', scope=(), display=None), 'routing': Entry(definition='One model call classifies a request and sends it to a specialised handler.', lesson='primer.agents.orchestration', scope=('primer/agents/orchestration', 'primer/agents/cost'), display=None), 'orchestrator-workers': Entry(definition='One model decides the subtasks, workers handle them, and one model combines the results.', lesson='primer.agents.orchestration', scope=(), display=None), 'evaluator-optimizer': Entry(definition='A generator drafts and an evaluator critiques, repeated until the draft passes or the rounds run out.', lesson='primer.agents.orchestration', scope=(), display=None), 'supervisor': Entry(definition='In a multi-agent system, the agent that assigns tasks to specialist agents and assembles their answers.', lesson='primer.agents.orchestration', scope=('primer/agents/orchestration',), display=None), 'hand-off': Entry(definition='The message that passes work and its context from one agent to another; anything not written into it is lost.', lesson='primer.agents.orchestration', scope=('primer/agents/orchestration', 'primer/agents/agent_loop'), display=None), 'stochastic gradient descent': Entry(definition="Gradient descent where each step's slope is estimated from a small random batch instead of the whole dataset.", lesson='primer.ml.optimizers', scope=(), display=None), 'learning-rate schedule': Entry(definition='A rule that changes the learning rate during training, such as warming up and then decaying along a cosine curve.', lesson='primer.ml.optimizers', scope=(), display=None), 'cosine decay': Entry(definition='Lowering the learning rate from its peak to a floor along half a cosine wave.', lesson='primer.ml.optimizers', scope=(), display=None), 'bias correction': Entry(definition="Adam's rescaling of its running averages early in training, since they start at zero and would otherwise be too small.", lesson='primer.ml.optimizers', scope=(), display=None), 'initialization': Entry(definition="Choosing the size of a network's random starting weights so signals neither shrink nor grow layer by layer.", lesson='primer.ml.deep_nets', scope=('primer/ml/deep_nets', 'primer/ml/neural_net'), display=None), 'xavier initialization': Entry(definition='Starting weights with variance 2 / (fan-in + fan-out), suited to tanh and sigmoid layers.', lesson='primer.ml.deep_nets', scope=(), display=None), 'he initialization': Entry(definition='Starting weights with variance 2 / fan-in, making up for ReLU zeroing half its inputs.', lesson='primer.ml.deep_nets', scope=(), display=None), 'fan-in': Entry(definition='The number of inputs a neuron adds up.', lesson='primer.ml.deep_nets', scope=(), display=None), 'saturation': Entry(definition='When an activation like sigmoid or tanh sits on its flat tail, so almost no gradient passes through.', lesson='primer.ml.neural_net', scope=('primer/ml/neural_net', 'primer/ml/deep_nets', 'primer/ml/attention'), display=None), 'dying relu': Entry(definition='A ReLU neuron whose input stays negative, so it outputs zero and receives zero gradient forever.', lesson='primer.ml.neural_net', scope=(), display=None), 'one-hot': Entry(definition='A vector with a 1 at the correct class and 0 everywhere else.', lesson='primer.ml.losses', scope=(), display=None), 'log-sum-exp': Entry(definition='Computing the log of a sum of exponentials without overflow, by factoring out the largest term first.', lesson='primer.ml.losses', scope=(), display=None), 'binary cross-entropy': Entry(definition='Cross-entropy for yes/no predictions: −[y·ln p + (1 − y)·ln(1 − p)].', lesson='primer.ml.losses', scope=(), display=None), 'mean squared error': Entry(definition='The average of squared prediction errors, so big misses dominate.', lesson='primer.ml.losses', scope=(), display=None), 'mean absolute error': Entry(definition='The average of absolute prediction errors, which is robust to outliers.', lesson='primer.ml.losses', scope=(), display=None), 'label smoothing': Entry(definition='Training against a softened target, such as 0.9 on the right class and the rest spread evenly, to discourage over-confidence.', lesson='primer.ml.losses', scope=(), display=None), 'bias-variance trade-off': Entry(definition="Splitting a model's error into systematic error from being too simple (bias) and sensitivity to the training sample from being too flexible (variance).", lesson='primer.ml.regularization', scope=(), display=None), 'l1 regularization': Entry(definition='Penalizing the sum of absolute weights, which drives unneeded weights to exactly zero.', lesson='primer.ml.regularization', scope=(), display=None), 'l2 regularization': Entry(definition='Penalizing the sum of squared weights, which shrinks all weights toward zero without eliminating them.', lesson='primer.ml.regularization', scope=(), display=None), 'soft thresholding': Entry(definition='Moving each weight a fixed amount toward zero and snapping any weight within that amount to exactly zero.', lesson='primer.ml.regularization', scope=(), display=None), 'validation set': Entry(definition='Held-out data used to tune choices like model size and when to stop training.', lesson='primer.ml.regularization', scope=(), display=None), 'test set': Entry(definition='Held-out data used once at the end to estimate real-world performance honestly.', lesson='primer.ml.regularization', scope=('primer/ml/',), display=None), 'cross-validation': Entry(definition='Rotating which slice of the data is held out, training once per slice, and averaging the scores.', lesson='primer.ml.regularization', scope=(), display=None), 'double descent': Entry(definition='The observation that past a point, even larger models generalize better again.', lesson='primer.ml.regularization', scope=(), display=None), 'benchmark contamination': Entry(definition="When test questions appeared in a model's training data, inflating its scores.", lesson='primer.ml.regularization', scope=(), display=None), 'gpu': Entry(definition='A graphics processor: a chip with thousands of small cores that do many multiply-adds at once, which is exactly what neural networks need.', lesson='primer.notation', scope=(), display=None), 'parallelism': Entry(definition='Doing many pieces of work at the same time instead of one after another.', lesson='primer.ml.attention', scope=(), display=None), 'hidden state': Entry(definition='The running summary an RNN carries from one step to the next.', lesson='primer.ml.cnn_rnn', scope=(), display=None), 'auto-regressive': Entry(definition='Generating one token at a time, each predicted from everything before it and then fed back in.', lesson='primer.ml.big_picture', scope=(), display=None), 'beam search': Entry(definition='Decoding that keeps the few most probable partial outputs at each step instead of just one, then picks the best finished one.', lesson='primer.ml.inference', scope=(), display=None), 'checkpoint averaging': Entry(definition='Averaging the weights saved at the last few points of training, which usually gives a slightly better model.', lesson='primer.ml.training_stages', scope=(), display=None), 'encoder-decoder attention': Entry(definition="Attention where the decoder's queries look at the encoder's keys and values, so each output word can consult the whole input.", lesson='primer.ml.transformer', scope=(), display=None), 'luhn check': Entry(definition='A checksum every real card number satisfies, used to tell card numbers from look-alike IDs.', lesson='primer.agents.guardrails', scope=(), display=None), 'named-entity recognition': Entry(definition='A model that tags names, places and organisations in text.', lesson='primer.agents.guardrails', scope=(), display=None), 'ner': Entry(definition='Named-entity recognition: tagging names, places and organisations in text.', lesson='primer.agents.guardrails', scope=(), display=None), 'entailment': Entry(definition='Whether one text logically follows from another, checked by a model trained for it.', lesson='primer.agents.guardrails', scope=(), display=None), 'privilege separation': Entry(definition="Splitting work so the part that reads untrusted content can't take dangerous actions.", lesson='primer.agents.guardrails', scope=(), display=None), 'lethal trifecta': Entry(definition='Private data, untrusted content and a way to send data out, all in one agent: together they allow data theft.', lesson='primer.agents.guardrails', scope=(), display=None), 'lost in the middle': Entry(definition='Models use information at the start and end of a long input more reliably than information in the middle.', lesson='primer.agents.context', scope=(), display=None), 'sandwich ordering': Entry(definition='Placing the best retrieved chunks at the start and end of the context and the weakest in the middle.', lesson='primer.agents.context', scope=(), display=None), 'escaping': Entry(definition="Replacing characters like < and > so data can't pose as markup or instructions.", lesson='primer.agents.context', scope=(), display=None), 'trajectory': Entry(definition='The full record of an agent run: its answer, its tool calls with arguments, and the end state.', lesson='primer.agents.evals', scope=(), display=None), 'code grader': Entry(definition="An automatic check written in code, such as exact match, schema or database state, that scores an agent's output.", lesson='primer.agents.evals', scope=(), display=None), 'rubric': Entry(definition='A short list of explicit pass/fail criteria given to a grader.', lesson='primer.agents.evals', scope=(), display=None), 'regression': Entry(definition='A task that used to pass and now fails after a change.', lesson='primer.agents.evals', scope=(), display=None), 'p95': Entry(definition='The 95th percentile: the value that 95% of measurements are at or below.', lesson='primer.agents.evals', scope=(), display=None), 'continuous integration': Entry(definition='The automated checks that run on every proposed change before it can merge.', lesson='primer.agents.evals', scope=(), display=None), 'drift': Entry(definition='A metric creeping away from its usual level over time or after a deploy.', lesson='primer.agents.evals', scope=(), display=None), 'faithfulness': Entry(definition="The share of an answer's claims that are supported by the retrieved sources.", lesson='primer.agents.evals', scope=(), display=None), 'cost per successful task': Entry(definition='Total spend divided by the number of tasks that succeeded, counting retries and cleanup.', lesson='primer.agents.cost', scope=(), display=None), 'ttl': Entry(definition='Time to live: how long a cached entry may be served before it is treated as expired.', lesson='primer.agents.cost', scope=(), display=None), 'batch api': Entry(definition='A provider interface that processes many requests asynchronously at a discount.', lesson='primer.agents.cost', scope=(), display=None), 'task budget': Entry(definition='Hard limits on steps and tokens for one agent task.', lesson='primer.agents.cost', scope=(), display=None), 'tenant': Entry(definition='One customer organisation on a shared platform.', lesson='primer.agents.cost', scope=(), display=None), 'anomaly alert': Entry(definition='An alert when a metric jumps far outside its usual range.', lesson='primer.agents.cost', scope=(), display=None), 'instrumentation': Entry(definition='Code that records spans or metrics around the work a program does.', lesson='primer.agents.observability', scope=(), display=None), 'opentelemetry': Entry(definition='The open standard for traces, metrics and logs.', lesson='primer.agents.observability', scope=(), display=None), 'genai semantic conventions': Entry(definition="OpenTelemetry's standard attribute names for model and agent spans, such as gen_ai.request.model.", lesson='primer.agents.observability', scope=(), display=None), 'otlp': Entry(definition='The OpenTelemetry Protocol: the wire format exporters use to ship telemetry.', lesson='primer.agents.observability', scope=(), display=None), 'collector': Entry(definition='In OpenTelemetry, the service that receives spans from apps and forwards them to storage backends.', lesson='primer.agents.observability', scope=(), display=None), 'graduated autonomy': Entry(definition='Granting an agent more independence in steps (shadow, then approval, then autonomy), each earned with evidence.', lesson='primer.agents.deployment', scope=(), display=None), 'content-addressed': Entry(definition='Named by a hash of its own content, so identical content always gets the identical name.', lesson='primer.agents.deployment', scope=(), display=None), 'sha-256': Entry(definition='A hash function: a fixed-length fingerprint of data that changes completely if the data changes at all.', lesson='primer.agents.deployment', scope=(), display=None), 'feature flag': Entry(definition='A configuration switch that turns a capability on or off without redeploying.', lesson='primer.agents.deployment', scope=(), display=None), 'rate limit': Entry(definition='A cap on how many actions may happen per unit of time.', lesson='primer.agents.deployment', scope=(), display=None), 'token bucket': Entry(definition='A rate limiter that allows bursts up to a capacity and refills at a steady rate.', lesson='primer.agents.deployment', scope=(), display=None), 'blast radius': Entry(definition='How much damage a failure can do before it is noticed and stopped.', lesson='primer.agents.deployment', scope=(), display=None), 'audit log': Entry(definition='A record of which actor took which action, for whom, with what inputs.', lesson='primer.agents.deployment', scope=(), display=None), 'hash chain': Entry(definition="Records that each store the previous record's hash, so any edit or deletion is detectable.", lesson='primer.agents.deployment', scope=(), display=None), 'erp': Entry(definition="Enterprise resource planning: a company's finance and operations system of record.", lesson='primer.agents.deployment', scope=(), display=None), 'exponential backoff': Entry(definition='Waiting longer after each failed attempt, doubling up to a cap.', lesson='primer.agents.failures', scope=(), display=None), 'jitter': Entry(definition="Random variation added to retry delays so clients don't all retry at the same moment.", lesson='primer.agents.failures', scope=(), display=None), 'circuit breaker': Entry(definition='A wrapper that stops calling a failing dependency for a cool-down period, failing fast instead.', lesson='primer.agents.failures', scope=(), display=None), 'contract test': Entry(definition='An automated check that an external API still accepts and returns what your code expects.', lesson='primer.agents.failures', scope=(), display=None), 'ocr': Entry(definition='Optical character recognition: reading text from an image of a page.', lesson='primer.agents.rag', scope=(), display=None), 'acl': Entry(definition='Access-control list: the users or groups allowed to read a document.', lesson='primer.agents.rag', scope=(), display=None), 'permission-aware retrieval': Entry(definition='Filtering out documents a user may not read before anything is ranked or shown to the model.', lesson='primer.agents.rag', scope=(), display=None), 'dense retrieval': Entry(definition='Search by comparing embedding vectors, so matches are by meaning rather than exact words.', lesson='primer.agents.rag', scope=(), display=None), 'query rewriting': Entry(definition='Turning a follow-up or vague question into a standalone search query.', lesson='primer.agents.rag', scope=(), display=None), 'multi-query retrieval': Entry(definition='Searching several phrasings of a question and fusing the results.', lesson='primer.agents.rag', scope=(), display=None), 'contextual retrieval': Entry(definition='Indexing each chunk together with a line saying where in its document it comes from.', lesson='primer.agents.rag', scope=(), display=None), 'parent-child retrieval': Entry(definition='Matching small chunks but returning the larger section around them.', lesson='primer.agents.rag', scope=(), display=None), 'agentic rag': Entry(definition='RAG where the model decides whether, what and how often to search.', lesson='primer.agents.rag', scope=(), display=None), 'graphrag': Entry(definition='Retrieval over a graph of entities and relationships extracted from documents.', lesson='primer.agents.rag', scope=(), display=None), 'centroid': Entry(definition="The average position of a group of vectors, used as the group's representative point.", lesson='primer.ml.embeddings.ann', scope=(), display=None), 'voronoi cell': Entry(definition='All the points closer to one centroid than to any other: the section an IVF index searches.', lesson='primer.ml.embeddings.ann', scope=(), display=None), 'nprobe': Entry(definition='How many IVF clusters a query scans; raising it trades speed for recall.', lesson='primer.ml.embeddings.ann', scope=(), display=None), 'efsearch': Entry(definition="HNSW's query-time beam width; raising it trades speed for recall.", lesson='primer.ml.embeddings.ann', scope=(), display=None), 'efconstruction': Entry(definition="HNSW's beam width while building the graph; higher means a better graph and a slower build.", lesson='primer.ml.embeddings.ann', scope=(), display=None), 'greedy search': Entry(definition='Always move to whichever neighbour is closest to the target, and stop when none is closer.', lesson='primer.ml.embeddings.ann', scope=(), display=None), 'graph': Entry(definition='A set of points (nodes) joined by links (edges).', lesson='primer.ml.embeddings.ann', scope=(), display=None), 'codebook': Entry(definition='A small catalogue of representative vectors; each vector (or piece of one) is stored as the number of its nearest entry, as in product quantization or tokenizers for images and audio.', lesson='primer.ml.embeddings.ann', scope=(), display=None), 'asymmetric distance computation': Entry(definition='Scoring compressed vectors against an uncompressed query by adding up precomputed table lookups.', lesson='primer.ml.embeddings.ann', scope=(), display=None), 'ivf-pq': Entry(definition='IVF clustering combined with compressed offsets from each cluster centre: the standard index for billion-scale search.', lesson='primer.ml.embeddings.ann', scope=(), display=None), 're-scoring': Entry(definition='Recomputing exact scores for a shortlist that an approximate index returned, to recover accuracy.', lesson='primer.ml.embeddings.ann', scope=(), display=None), 'flat index': Entry(definition='Exact search that compares the query with every stored vector; the ground truth other indexes are measured against.', lesson='primer.ml.embeddings.ann', scope=(), display=None), 'idf': Entry(definition="Inverse document frequency: a word's rarity weight, higher for words found in fewer documents.", lesson='primer.ml.embeddings.retrieval', scope=(), display=None), 'term frequency': Entry(definition='How many times a word appears in a document.', lesson='primer.ml.embeddings.retrieval', scope=(), display=None), 'length normalization': Entry(definition="BM25's discount for mentions in documents longer than average, controlled by b.", lesson='primer.ml.embeddings.retrieval', scope=(), display=None), 'late interaction': Entry(definition='Keeping one vector per word and matching each query word to its best document word.', lesson='primer.ml.embeddings.retrieval', scope=(), display=None), 'maxsim': Entry(definition='For each query word, its highest similarity to any document word, summed over the query.', lesson='primer.ml.embeddings.retrieval', scope=(), display=None), 'query and passage prefixes': Entry(definition='Labels some embedding models expect on questions and documents; forgetting one silently lowers recall.', lesson='primer.ml.embeddings.retrieval', scope=(), display=None), 'shortlist': Entry(definition='The small set of top candidates from a cheap first-stage search that a slower, more precise stage reorders.', lesson='primer.ml.embeddings.retrieval', scope=(), display=None), 'l2 normalization': Entry(definition='Dividing a vector by its own length so it has length 1.', lesson='primer.ml.embeddings.similarity', scope=(), display=None), 'dimension': Entry(definition='One of the numbers in a vector: a 768-dimension embedding is a list of 768 numbers.', lesson='primer.ml.embeddings.similarity', scope=(), display=None), 'mean-centering': Entry(definition='Subtracting the average vector of a collection from every vector, removing the direction they all share.', lesson='primer.ml.embeddings.similarity', scope=(), display=None), 'whitening': Entry(definition='Rescaling vectors so every direction has equal spread and no two directions are correlated.', lesson='primer.ml.embeddings.similarity', scope=(), display=None), 'threshold calibration': Entry(definition='Choosing a similarity cut-off by measuring precision and recall on labeled pairs for one specific model.', lesson='primer.ml.embeddings.similarity', scope=(), display=None), 'word2vec': Entry(definition='A 2013 method that learns one vector per word by predicting nearby words.', lesson='primer.ml.embeddings.word2vec', scope=(), display=None), 'skip-gram': Entry(definition="The word2vec task of predicting each word's neighbours from the word itself.", lesson='primer.ml.embeddings.word2vec', scope=(), display=None), 'negative sampling': Entry(definition='Training by scoring the true pair up and a few random noise pairs down, instead of scoring the whole vocabulary.', lesson='primer.ml.embeddings.word2vec', scope=(), display=None), 'pmi': Entry(definition='Pointwise mutual information: the log of how much more often two words appear together than chance predicts.', lesson='primer.ml.embeddings.word2vec', scope=(), display=None), 'glove': Entry(definition='A counting method that fits word vectors so their dot products predict how often words appear together.', lesson='primer.ml.embeddings.word2vec', scope=(), display=None), 'svd': Entry(definition='Singular value decomposition: splits a table into a few directions that capture its main patterns, used to compress it.', lesson='primer.ml.embeddings.word2vec', scope=(), display=None), 'static embedding': Entry(definition='One fixed vector per word, whatever the sentence.', lesson='primer.ml.embeddings.word2vec', scope=(), display=None), 'contextual embedding': Entry(definition='A vector computed for a word within its sentence, so the same word gets different vectors in different contexts.', lesson='primer.ml.embeddings.word2vec', scope=(), display=None), 'contrastive learning': Entry(definition='Training that pulls matching pairs together and pushes non-matching pairs apart.', lesson='primer.ml.embeddings.contrastive', scope=(), display=None), 'in-batch negatives': Entry(definition='Using the other examples in a training batch as free wrong answers for each query.', lesson='primer.ml.embeddings.contrastive', scope=(), display=None), 'bag of words': Entry(definition='Representing text by counting how many times each word appears, ignoring order.', lesson='primer.ml.embeddings.contrastive', scope=(), display=None), 'zero-shot classification': Entry(definition='Labeling items by comparing them to a text description of each label, with no training on those labels.', lesson='primer.ml.embeddings.contrastive', scope=(), display=None), 'scalar quantization': Entry(definition='Storing each number as one of 256 levels (one byte) instead of a 4-byte float.', lesson='primer.ml.embeddings.compression', scope=(), display=None), 'binary quantization': Entry(definition='Storing only the sign of each number, as a single bit.', lesson='primer.ml.embeddings.compression', scope=(), display=None), 'hamming distance': Entry(definition='The number of positions where two bit strings differ.', lesson='primer.ml.embeddings.compression', scope=(), display=None), 'popcount': Entry(definition='Counting the 1-bits in a number; combined with XOR it computes Hamming distance in one step.', lesson='primer.ml.embeddings.compression', scope=(), display=None), 'two-stage retrieval': Entry(definition='A cheap, wide first pass to shortlist candidates, then an expensive, precise pass over only those.', lesson='primer.ml.embeddings.compression', scope=(), display=None), 'hit@k': Entry(definition='The share of queries with at least one relevant result in the top k.', lesson='primer.ml.embeddings.operations', scope=(), display=None), 'pca': Entry(definition='Principal component analysis: finding the directions along which data varies most, to draw or compress it with fewer numbers.', lesson='primer.ml.embeddings.clustering', scope=(), display=None), 'clustering': Entry(definition='Grouping items so similar ones end up together, without being told the groups.', lesson='primer.ml.embeddings.clustering', scope=(), display=None), 'inertia': Entry(definition="The total squared distance from each point to its cluster's center: the quantity k-means minimizes.", lesson='primer.ml.embeddings.clustering', scope=(), display=None), 'k-means++': Entry(definition='A way to pick k-means starting centers spread far apart, which avoids bad results.', lesson='primer.ml.embeddings.clustering', scope=(), display=None), 'silhouette score': Entry(definition='How much closer each point is to its own cluster than to the nearest other one, from −1 to 1.', lesson='primer.ml.embeddings.clustering', scope=(), display=None), 'elbow method': Entry(definition='Choosing the number of clusters where adding more stops reducing inertia much.', lesson='primer.ml.embeddings.clustering', scope=(), display=None), 'dbscan': Entry(definition='Density-based clustering that grows clusters from points with enough close neighbours and labels the rest as noise.', lesson='primer.ml.embeddings.clustering', scope=(), display=None), 'hdbscan': Entry(definition='A version of DBSCAN that considers every reach at once and keeps the most persistent clusters, so no reach setting is needed.', lesson='primer.ml.embeddings.clustering', scope=(), display=None), 'umap': Entry(definition="A projection to 2-D that keeps each point's neighbours close but distorts other distances.", lesson='primer.ml.embeddings.clustering', scope=(), display=None), 't-sne': Entry(definition='An older neighbour-preserving 2-D projection; cluster sizes and gaps in its plots are not meaningful.', lesson='primer.ml.embeddings.clustering', scope=(), display=None), 'near-duplicate detection': Entry(definition='Finding items whose embeddings are almost identical, such as reworded copies.', lesson='primer.ml.embeddings.clustering', scope=(), display=None), 'anomaly detection': Entry(definition='Flagging items far from every known group.', lesson='primer.ml.embeddings.clustering', scope=(), display=None), 'blue/green deployment': Entry(definition='Building the new version beside the old one, switching traffic once it proves itself, and keeping the old one for rollback.', lesson='primer.ml.embeddings.operations', scope=(), display=None), 'dual write': Entry(definition='Writing every new record to both the old and the new system during a migration.', lesson='primer.ml.embeddings.operations', scope=(), display=None), 'backfill': Entry(definition='Re-processing existing records into a new system, such as re-embedding a corpus with a new model.', lesson='primer.ml.embeddings.operations', scope=(), display=None), 'index alias': Entry(definition='A pointer name like "live" that search uses, so switching indexes means repointing one name.', lesson='primer.ml.embeddings.operations', scope=(), display=None), 'domain mismatch': Entry(definition='A general model underperforming on specialised language it never saw, such as company jargon.', lesson='primer.ml.embeddings.operations', scope=(), display=None), 'retrieval failure': Entry(definition='A wrong answer caused because no relevant document was retrieved.', lesson='primer.ml.embeddings.operations', scope=(), display=None), 'generation failure': Entry(definition='A wrong answer produced even though a relevant document was retrieved.', lesson='primer.ml.embeddings.operations', scope=(), display=None), 'inverted file': Entry(definition='In an IVF index, the list of vectors assigned to one cluster, so a search scans only the lists it picks.', lesson='primer.ml.embeddings.ann', scope=(), display=None), 'diversity heuristic': Entry(definition="HNSW's rule of keeping a neighbour only if it is closer to the new node than to any neighbour already kept, so links fan out in different directions.", lesson='primer.ml.embeddings.ann', scope=(), display=None), 'sparse retrieval': Entry(definition='Keyword search such as BM25, which scores exact word matches.', lesson='primer.ml.embeddings.retrieval', scope=(), display=None), 'hysteresis': Entry(definition="Making the bar for changing state higher than the bar for staying, so a decision doesn't flicker between two close options.", lesson=None, scope=(), display=None), 'debounce': Entry(definition='Waiting until a signal has settled before acting on it, so a burst of changes triggers one action.', lesson=None, scope=(), display=None), 'cooldown': Entry(definition='A quiet period after an action during which the same action is not repeated.', lesson=None, scope=(), display=None), 'latency': Entry(definition="The delay between something happening and the system's response to it.", lesson='primer.agents.cost', scope=(), display=None), 'rank': Entry(definition='How many independent directions a matrix really contains; a table made by multiplying one column by one row has rank 1.', lesson='primer.ml.training_stages', scope=('primer/ml/training_stages',), display=None), 'singular value decomposition': Entry(definition='Breaking a matrix into independent directions ranked by importance, so the top few capture most of what it does.', lesson='primer.ml.embeddings.word2vec', scope=(), display=None), 'frobenius norm': Entry(definition='The size of a whole matrix: square every entry, add them up, take the square root.', lesson='primer.notation', scope=(), display=None), 'prefix tuning': Entry(definition='Training a few special virtual-token vectors placed before every input while the model stays frozen.', lesson='primer.ml.training_stages', scope=(), display=None), 'ppo': Entry(definition='Proximal Policy Optimization: a reinforcement learning algorithm that improves a policy in small, clipped steps; widely used for RLHF.', lesson='primer.ml.training_stages', scope=(), display=None), 'alignment tax': Entry(definition='Capability lost as a side effect of training a model to be helpful, honest and harmless.', lesson='primer.ml.training_stages', scope=(), display=None), 'win rate': Entry(definition='The share of head-to-head comparisons one model wins; 50% means the two are indistinguishable.', lesson='primer.agents.evals', scope=(), display=None), 'few-shot prompt': Entry(definition='A prompt that includes a few worked examples of the task before the real question.', lesson='primer.agents.context', scope=(), display=None), 'exemplar': Entry(definition='One worked example in a few-shot prompt: a question with its answer, and sometimes the steps that reach it.', lesson='primer.agents.context', scope=(), display=None), 'linear probe': Entry(definition='Freezing a model and training only a simple linear classifier on its embeddings or internal activations, to measure what they encode.', lesson='primer.ml.embeddings.contrastive', scope=(), display=None), 'zero-shot': Entry(definition='Doing a task with no task-specific training examples.', lesson='primer.ml.embeddings.contrastive', scope=(), display=None), 'map@10': Entry(definition='Mean average precision over the top 10 results: rewards putting correct matches in the top 10, and higher up within it.', lesson='primer.ml.metrics', scope=(), display=None), 'nf4': Entry(definition='4-bit NormalFloat: a 4-bit number format whose 16 levels are spaced to match the bell-curve shape of neural network weights.', lesson='primer.ml.inference', scope=(), display=None), 'double quantization': Entry(definition='Compressing the per-block scale factors of a quantized model as well, saving about 0.37 bits per parameter.', lesson='primer.ml.inference', scope=(), display=None), 'mips': Entry(definition='Maximum inner product search: finding the stored vectors with the largest dot product against a query vector, usually approximately with an index.', lesson='primer.ml.embeddings.ann', scope=(), display=None), 'expected value': Entry(definition='The average you would get over many tries.', lesson='primer.notation', scope=(), display=None), 'credit assignment': Entry(definition='Working out which of many earlier actions deserves the blame or credit for how an episode ended.', lesson='primer.agents.planning', scope=(), display=None), 'pass@1': Entry(definition='The share of problems solved by the first program submitted for each, judged by hidden tests.', lesson='primer.agents.evals', scope=(), display=None), 'machine translation': Entry(definition='Turning text in one language into another.', lesson='primer.ml.transformer', scope=(), display=None), 'linear attention': Entry(definition='Attention variants that avoid scoring every pair of tokens, so cost grows in proportion to n instead of n².', lesson='primer.ml.attention', scope=(), display=None), 'in-context learning': Entry(definition="Picking up a task from instructions or examples in the prompt, with no change to the model's weights.", lesson='primer.ml.big_picture', scope=(), display=None), 'speculative execution': Entry(definition='Starting likely work early, in parallel, and throwing it away if the guess was wrong.', lesson='primer.ml.inference', scope=(), display=None), 'petaflop/s-day': Entry(definition='10¹⁵ operations per second for one day, 8.64 × 10¹⁹ operations: a unit for training budgets.', lesson='primer.ml.training_stages', scope=(), display=None), 'isoflop profile': Entry(definition='A curve of final loss against model size at a fixed training compute; its minimum is the best size for that budget.', lesson='primer.ml.training_stages', scope=(), display=None), 'hbm': Entry(definition="The GPU's large main memory: tens of GB, about 10× slower than on-chip SRAM.", lesson='primer.ml.inference', scope=(), display=None), 'sram': Entry(definition="The tiny, very fast on-chip memory next to a GPU's arithmetic units.", lesson='primer.ml.inference', scope=(), display=None), 'kernel fusion': Entry(definition='Doing several steps in one GPU program, so data is read from and written to main memory only once.', lesson='primer.ml.inference', scope=(), display=None), 'internal fragmentation': Entry(definition='Memory reserved for a request that it never uses.', lesson='primer.ml.inference', scope=(), display=None), 'external fragmentation': Entry(definition='Free memory split into gaps too small to use.', lesson='primer.ml.inference', scope=(), display=None), 'block table': Entry(definition='A per-request map from logical KV-cache blocks to physical blocks in GPU memory.', lesson='primer.ml.inference', scope=(), display=None), 'copy-on-write': Entry(definition='Sharing one copy of data and making a private copy only when someone needs to change it.', lesson='primer.ml.inference', scope=(), display=None), 'co-adaptation': Entry(definition="When neurons learn to depend on each other's exact behaviour, each correcting the others' quirks; it fits the training data but breaks on new data.", lesson='primer.ml.regularization', scope=(), display=None), 'max-norm constraint': Entry(definition="Capping the length of each neuron's incoming weight vector at a fixed radius, shrinking it back whenever an update pushes it past.", lesson='primer.ml.regularization', scope=(), display=None), 'regret': Entry(definition="In online learning, the total extra loss from choosing each step's parameters before seeing that step's data, compared with the best fixed choice in hindsight.", lesson='primer.ml.optimizers', scope=('primer/ml/optimizers',), display=None), 'convex': Entry(definition='Bowl-shaped everywhere with a single bottom, so any downhill path reaches the same minimum; neural network losses are not convex.', lesson='primer.notation', scope=(), display=None), 'adagrad': Entry(definition="An optimizer that divides each weight's step by the square root of the sum of all its past squared gradients, so steps only ever shrink.", lesson='primer.ml.optimizers', scope=(), display='AdaGrad'), 'rmsprop': Entry(definition='An optimizer that divides each step by the square root of a running average of recent squared gradients.', lesson='primer.ml.optimizers', scope=(), display=None), 'sparse gradients': Entry(definition='Gradients that are zero most of the time for a given weight, such as the weight for a rare word.', lesson='primer.ml.optimizers', scope=(), display=None), 'degradation problem': Entry(definition='Deeper plain networks reaching higher error even on their training data: an optimization failure, not overfitting.', lesson='primer.ml.deep_nets', scope=(), display=None), 'subword unit': Entry(definition='A piece of a word, from a single character to a whole frequent word; a fixed set of them can spell any word.', lesson='primer.ml.tokenization', scope=(), display=None), 'mean average precision': Entry(definition='How well a system ranks correct results above wrong ones, averaged over queries or classes.', lesson='primer.ml.metrics', scope=(), display=None), 'cbow': Entry(definition="Continuous bag of words: the word2vec model that averages the surrounding words' vectors to predict the middle word.", lesson='primer.ml.embeddings.word2vec', scope=(), display=None), 'hierarchical softmax': Entry(definition='Replacing one softmax over the whole vocabulary with about log₂V yes-or-no decisions down a binary tree.', lesson='primer.ml.embeddings.word2vec', scope=(), display=None), 'subsampling': Entry(definition='Randomly skipping most occurrences of very frequent words during training, which is faster and improves rare-word vectors.', lesson='primer.ml.embeddings.word2vec', scope=(), display=None), 'siamese network': Entry(definition='Two copies of one network with shared weights, each reading one input, so their outputs can be compared directly.', lesson='primer.ml.embeddings.contrastive', scope=(), display=None), 'triplet loss': Entry(definition='A loss requiring an anchor to be closer to its positive than to its negative by at least a margin.', lesson='primer.ml.losses', scope=(), display=None), 'spearman correlation': Entry(definition='How well two rankings agree, from −1 (reversed) to 1 (identical).', lesson='primer.ml.metrics', scope=(), display=None), 'top-k retrieval accuracy': Entry(definition='The share of questions for which at least one of the top k retrieved passages contains the answer.', lesson='primer.ml.metrics', scope=(), display=None), 'navigable small world': Entry(definition='A graph where always stepping to the neighbour closest to the target reaches it in few hops.', lesson='primer.ml.embeddings.ann', scope=(), display=None), 'skip list': Entry(definition='A sorted list with random express lanes on top, searched by running along the top lane and dropping down, in about log(n) steps.', lesson='primer.ml.embeddings.ann', scope=(), display=None), 'huffman tree': Entry(definition='A binary tree that gives frequent items short codes and rare items long ones.', lesson='primer.ml.embeddings.word2vec', scope=(), display=None), 'logistic regression': Entry(definition='A linear model whose weighted sum passes through a sigmoid to give a probability.', lesson='primer.ml.neural_net', scope=(), display=None), 'locality-sensitive hashing': Entry(definition='Hashing vectors so that similar vectors are likely to land in the same bucket, for fast approximate search.', lesson='primer.ml.embeddings.ann', scope=(), display=None), 'ollama': Entry(definition='A free app that downloads open language models and runs them on your own computer, served over a small local web API.', lesson='primer.agents.llm', scope=(), display=None), 'local model': Entry(definition="A language model running on your own machine rather than a provider's servers: free per call, private and offline, but smaller.", lesson='primer.agents.llm', scope=(), display=None), 'activation checkpointing': Entry(definition="Keeping only each layer's input and recomputing the rest during the backward pass.", lesson='primer.ml.pretraining', scope=(), display=None), 'activation patching': Entry(definition='Copying one activation from a clean run into a corrupted run to measure how much of the right answer returns; also called causal tracing.', lesson='primer.ml.interpretability', scope=(), display=None), 'actor-critic': Entry(definition='A reinforcement learner in two parts: the actor (the policy) chooses actions, and the critic (a value network) predicts the reward to expect, which serves as the baseline.', lesson='primer.ml.reinforcement', scope=(), display=None), 'adaptive layer norm': Entry(definition='Layer normalization whose scale and shift are computed from a conditioning signal, such as the noise step and class label, instead of being fixed learned weights. Diffusion transformers use it to tell every block what to make (adaLN).', lesson='primer.ml.deep_nets', scope=(), display=None), 'active learning': Entry(definition='Choosing which examples to label next by what the current model is likely to get wrong, so each paid label teaches more.', lesson=None, scope=(), display=None), 'advantage': Entry(definition='How much better an action did than typical: reward minus baseline.', lesson='primer.ml.reinforcement', scope=('primer/ml/reinforcement',), display=None), 'adaboost': Entry(definition='A boosting method (Freund and Schapire, 1996) that trains each new classifier on reweighted data, raising the weight of the examples the previous ones got wrong, and lets the classifiers vote with weights set by their accuracy.', lesson=None, scope=(), display=None), 'adversarial example': Entry(definition='An input changed by a small, deliberately chosen amount so that a model gets it wrong, often in a way a person would never notice.', lesson=None, scope=(), display=None), 'alignment': Entry(definition="Making a model's behaviour match what we want (helpful, honest, harmless), not just the measurements we optimize.", lesson='primer.ml.alignment', scope=('primer/ml/alignment',), display=None), 'all-gather': Entry(definition='Every GPU contributes its piece and every GPU ends up with all the pieces, side by side; the second half of a ring all-reduce.', lesson='primer.ml.pretraining', scope=(), display=None), 'all-reduce': Entry(definition='Summing a value across GPUs so every GPU ends up holding the total.', lesson='primer.ml.hardware', scope=(), display=None), 'arena': Entry(definition='Ratings built from people voting between two anonymous answers.', lesson='primer.ml.benchmarks', scope=('primer/ml/benchmarks',), display='arena'), 'attack success rate': Entry(definition='The share of red-team attempts that get past a safety check.', lesson='primer.ml.alignment', scope=(), display=None), 'attention sink': Entry(definition='Early tokens that trained models park spare attention on; dropping them breaks streaming generation.', lesson='primer.ml.efficient_architectures', scope=(), display=None), 'autoencoder': Entry(definition='A network that squeezes its input into a small code and rebuilds the input from it, trained only to make the rebuild match.', lesson='primer.ml.generative.autoencoders', scope=(), display=None), 'automated interpretability': Entry(definition='Using a language model to explain a feature or neuron from examples of when it fires, then scoring the explanation by how well the model predicts new activations from it alone.', lesson=None, scope=(), display=None), 'bagging': Entry(definition='Bootstrap aggregating: train one model per bootstrap sample and average them to cut variance.', lesson='primer.ml.classical', scope=(), display=None), 'baseline': Entry(definition='A typical reward subtracted before updating; it reduces noise without changing the average gradient.', lesson='primer.ml.reinforcement', scope=('primer/ml/reinforcement',), display=None), 'benchmark': Entry(definition='Fixed questions plus a scoring rule, averaged into one comparable number.', lesson='primer.ml.benchmarks', scope=('primer/ml/benchmarks',), display=None), 'bernoulli distribution': Entry(definition='The distribution of a single yes-or-no outcome, set by one number: the chance of yes. A VAE decoder for black-and-white pixels outputs one per pixel.', lesson='primer.ml.losses', scope=(), display=None), 'bert': Entry(definition='An encoder-only transformer trained to fill in hidden words using the text on both sides of them; the ancestor of many embedding and classification models.', lesson='primer.ml.transformer', scope=(), display=None), 'best-of-n': Entry(definition='Sample n candidates and keep the one a verifier scores highest.', lesson='primer.ml.reasoning', scope=(), display=None), 'binomial coefficient': Entry(definition='"n choose k", written C(n, k): the number of different groups of k items that can be picked from n, ignoring order. C(10, 5) = 252.', lesson='primer.ml.benchmarks', scope=(), display=None), 'binomial distribution': Entry(definition='The chances of getting each possible number of successes in n independent tries that each succeed with the same probability p.', lesson='primer.ml.benchmarks', scope=(), display=None), 'bf16': Entry(definition="A 16-bit float with fp32's 8 exponent bits and 7 mantissa bits: the same range, less precision.", lesson='primer.ml.hardware', scope=(), display=None), 'bits per dimension': Entry(definition="A model's negative log-likelihood per number in the data, in bits: how long a code the model needs, on average, for each pixel value. Lower is better.", lesson='primer.ml.losses', scope=(), display=None), 'boosting': Entry(definition='Building a strong model from many weak ones trained one after another, each concentrating on what the ones before it still get wrong.', lesson='primer.ml.classical', scope=(), display=None), 'bootstrap': Entry(definition='Measuring uncertainty by rescoring many with-replacement resamples of your own data.', lesson='primer.ml.benchmarks', scope=(), display=None), 'bootstrap sample': Entry(definition='A resample of the training rows drawn with replacement, so some rows repeat and about 37% are left out.', lesson='primer.ml.classical', scope=(), display=None), 'bottleneck': Entry(definition='The narrow middle of an autoencoder: too small to copy through, it forces the network to keep only what matters.', lesson='primer.ml.generative.autoencoders', scope=('primer/ml/generative/autoencoders',), display=None), 'bounding box': Entry(definition='A rectangle, given by its corner coordinates, that marks where one object sits in an image.', lesson=None, scope=(), display=None), 'cache miss': Entry(definition='When the processor needs a number that is not in its small, fast cache and must wait for the much slower main memory, which can cost as much time as hundreds of arithmetic operations.', lesson='primer.ml.hardware', scope=(), display=None), 'canary string': Entry(definition='A unique marker in benchmark files so trainers can filter out copies.', lesson='primer.ml.benchmarks', scope=(), display=None), 'catastrophic forgetting': Entry(definition='When training on a new task alone erodes or erases skills a model already had.', lesson='primer.ml.fine_tuning', scope=(), display=None), 'causal tracing': Entry(definition='Activation patching used to find where a model recalls a fact: blur the subject in the input, then restore one clean hidden state at a time and see which ones bring the right answer back.', lesson='primer.ml.interpretability', scope=(), display=None), 'chain of thought': Entry(definition='Intermediate reasoning steps a model writes before its final answer, each one readable by the next forward pass.', lesson='primer.ml.reasoning', scope=(), display=None), 'chat template': Entry(definition='The exact role markers and layout a model family uses to turn messages into training text.', lesson='primer.ml.fine_tuning', scope=(), display=None), 'chinchilla scaling': Entry(definition='The compute-optimal rule of about 20 training tokens per model parameter.', lesson='primer.ml.pretraining', scope=(), display=None), 'cider': Entry(definition='A score for a generated caption: how well its words and short phrases match several human-written captions of the same image, with phrases common to captions of many images counting less. Higher is better; values above 100 are normal.', lesson=None, scope=(), display=None), 'circuit': Entry(definition='A chain of components that together compute one behaviour.', lesson='primer.ml.interpretability', scope=('primer/ml/interpretability',), display=None), 'classifier guidance': Entry(definition="Steering a diffusion model toward a label by adding the slope of a separate classifier, trained on noisy inputs, to the model's noise guess at every step.", lesson='primer.ml.generative.diffusion', scope=(), display=None), 'classifier-free guidance': Entry(definition="Mixing a model's guesses with and without the prompt, and pushing past the prompted one to follow it more closely.", lesson='primer.ml.generative.diffusion', scope=(), display=None), 'cnn': Entry(definition='Convolutional neural network: a network built from convolutions, small learned filters slid across an image to find local patterns wherever they appear.', lesson='primer.ml.cnn_rnn', scope=(), display=None), 'commitment loss': Entry(definition="A training term that pulls an encoder's output towards the codebook entry it snapped to, so the encoder commits to its entries instead of drifting away from them.", lesson=None, scope=(), display=None), 'common crawl': Entry(definition="A nonprofit's public archive of the web, released as regular snapshots of billions of pages; the raw material of most pretraining data.", lesson='primer.ml.pretraining', scope=(), display=None), 'compressed sensing': Entry(definition='Recovering a long vector from far fewer measurements of it, which is possible when the vector is known to be sparse (mostly zeros).', lesson=None, scope=(), display=None), 'computer use': Entry(definition='An agent operating a graphical interface from screenshots, sending clicks and keystrokes.', lesson='primer.agents.coding_agents', scope=(), display=None), 'confidence interval': Entry(definition='A range built so that 95 in 100 such ranges contain the true value.', lesson='primer.ml.benchmarks', scope=(), display=None), 'constitutional ai': Entry(definition='Training against a written list of principles: the model critiques and revises its answers, and AI-labelled preferences train the reward model.', lesson='primer.ml.alignment', scope=(), display=None), 'constrained decoding': Entry(definition="Before each token is drawn, setting every token that can't lead to a valid answer to probability zero.", lesson='primer.ml.structured_output', scope=(), display=None), 'contamination': Entry(definition="Test questions and answers leaking into a model's training data.", lesson='primer.ml.benchmarks', scope=('primer/ml/benchmarks',), display=None), 'control task': Entry(definition='A probe trained on random labels, showing how much a probe fits with nothing real to find.', lesson='primer.ml.interpretability', scope=('primer/ml/interpretability',), display=None), 'coupling': Entry(definition='A way of pairing draws from two distributions so that each side, on its own, still has its own distribution. Pairing noise with data at random is one coupling; sending each noise point to one definite data point is another.', lesson='primer.ml.generative.diffusion', scope=('primer/ml/generative',), display=None), 'cost per resolved task': Entry(definition='Everything spent on all attempts, divided by the tasks actually resolved.', lesson='primer.agents.coding_agents', scope=(), display=None), 'covariance': Entry(definition='How two quantities vary together: positive when they rise together, negative when one rises as the other falls. For vectors, a matrix holding it for every pair of entries.', lesson=None, scope=(), display=None), 'critic': Entry(definition="A Wasserstein GAN's discriminator, which outputs an unbounded score instead of a probability.", lesson='primer.ml.generative.gans', scope=('primer/ml/generative/gans',), display=None), 'cross-attention': Entry(definition='Attention whose queries come from one sequence (image patches) and whose keys and values come from another (text tokens).', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'data mixture': Entry(definition='The share of training tokens drawn from each data source.', lesson='primer.ml.pretraining', scope=(), display=None), 'data parallelism': Entry(definition='Every GPU holds the whole model and trains on a different slice of the batch.', lesson='primer.ml.pretraining', scope=(), display=None), 'dcgan': Entry(definition='Deep convolutional GAN: a GAN whose generator and discriminator are convolutional networks, the usual baseline design for image GANs.', lesson='primer.ml.generative.gans', scope=(), display=None), 'ddim': Entry(definition='A deterministic diffusion sampler that predicts the clean result and jumps to a much less noisy step, so it needs far fewer steps.', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'ddpm': Entry(definition='The original diffusion sampler: remove the guessed noise in many small steps, adding a small fresh wobble each time.', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'dead latent': Entry(definition='A sparse autoencoder latent that has stopped firing on any input, wasting its slot in the dictionary; resampling it onto badly rebuilt inputs brings it back.', lesson='primer.ml.interpretability', scope=(), display=None), 'decision tree': Entry(definition='A flowchart of yes/no questions about one feature at a time, learned from examples, ending in leaves that give the answer.', lesson='primer.ml.classical', scope=(), display=None), 'deduplication': Entry(definition='Removing repeated copies of documents from training data.', lesson='primer.ml.pretraining', scope=(), display=None), 'denoiser': Entry(definition='The network in a diffusion model that looks at a noisy input and its step, and guesses the noise in it.', lesson='primer.ml.generative.diffusion', scope=('primer/ml/generative/diffusion',), display=None), 'denoising score matching': Entry(definition='Learning the score by training a network to guess the noise added to real examples.', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'dictionary learning': Entry(definition='Finding a set of directions, usually more than there are dimensions, such that every data point is a sparse combination of a few of them; also called sparse coding. A sparse autoencoder is one way to do it.', lesson='primer.ml.interpretability', scope=(), display=None), 'diffusion model': Entry(definition='A generator that learns to turn random noise into data by removing a little noise at a time.', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'diffusion transformer': Entry(definition='A transformer used as the denoiser, reading patches of the latent as tokens (DiT).', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'discretization': Entry(definition="Turning a continuous rate of change into one step's keep and write factors.", lesson='primer.ml.efficient_architectures', scope=(), display=None), 'discriminator': Entry(definition='The network in a GAN that outputs the probability that a sample is real rather than generated.', lesson='primer.ml.generative.gans', scope=('primer/ml/generative/gans',), display=None), 'edit distance': Entry(definition='The fewest single-item insertions, deletions and substitutions that turn one sequence into another.', lesson=None, scope=(), display=None), "earth mover's distance": Entry(definition='The least total work to reshape one pile of probability into another, each bit of mass times the distance it travels; the same as the Wasserstein distance.', lesson='primer.ml.generative.gans', scope=(), display=None), 'edit-run-test loop': Entry(definition='An agent loop that changes code, runs the tests, and repeats until they pass or a budget runs out.', lesson='primer.agents.coding_agents', scope=(), display=None), 'elo rating': Entry(definition='A rating scale where a 400-point gap means ten-to-one odds of winning.', lesson='primer.ml.benchmarks', scope=(), display=None), 'ensemble': Entry(definition='A model that combines the predictions of many models, such as trees.', lesson='primer.ml.classical', scope=('primer/ml/classical',), display=None), 'entropy': Entry(definition='The average surprise of a distribution: how many yes/no questions, on average, it takes to learn an outcome; 0 means certain.', lesson='primer.ml.classical', scope=(), display=None), 'euler step': Entry(definition='Moving in a straight line for a short time at the current velocity, then looking again.', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'euler-maruyama step': Entry(definition='The Euler step for a stochastic differential equation: move by the drift times the time step, then add a fresh random nudge whose size grows with the square root of the time step.', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'evidence lower bound': Entry(definition='A quantity that never exceeds the log-probability a model gives the data; the VAE loss is its negative.', lesson='primer.ml.generative.autoencoders', scope=(), display=None), 'expectation': Entry(definition='The average of a quantity over many random draws, written E.', lesson='primer.ml.generative.gans', scope=('primer/ml/generative/gans',), display=None), 'exploration': Entry(definition='Trying actions you are unsure of instead of repeating the best one so far.', lesson='primer.ml.reinforcement', scope=('primer/ml/reinforcement',), display=None), 'exponent': Entry(definition='The bits of a float that pick the power of two, which sets its range.', lesson='primer.ml.hardware', scope=('primer/ml/hardware',), display=None), 'exponential moving average': Entry(definition='A running average that keeps most of its old value and mixes in a small share of each new one, so recent values count most. Adam keeps its moments this way, and diffusion models sample with such an average of their weights.', lesson='primer.ml.optimizers', scope=(), display=None), 'fail-to-pass test': Entry(definition='A hidden test that fails before a fix and must pass after it.', lesson='primer.agents.coding_agents', scope=(), display=None), 'false positive': Entry(definition='Something flagged as positive that is actually negative, such as a solution graded correct because its final answer matches although its reasoning is wrong.', lesson='primer.ml.metrics', scope=(), display=None), 'feature': Entry(definition='A property a model tracks, stored as a direction across many neurons.', lesson='primer.ml.interpretability', scope=('primer/ml/interpretability',), display=None), 'feature engineering': Entry(definition='Hand-transforming inputs (logarithms, one-hot columns) so a model can use them.', lesson='primer.ml.classical', scope=(), display=None), 'feature importance': Entry(definition='A score for how much a model relies on each input column.', lesson='primer.ml.classical', scope=(), display=None), 'feature map': Entry(definition='In linear attention, the function φ applied to queries and keys so that φ(q)·φ(k) replaces e^(q·k).', lesson='primer.ml.efficient_architectures', scope=(), display=None), 'feature splitting': Entry(definition='One feature in a small sparse autoencoder becoming several narrower features in a larger one, such as one base64 feature becoming separate features for letters, digits and encoded text.', lesson='primer.ml.interpretability', scope=(), display=None), 'fid': Entry(definition='Fréchet Inception Distance: compares a set of generated images with real ones through the statistics of their features in a pretrained image network. Lower is better; 0 means indistinguishable.', lesson='primer.ml.generative.gans', scope=(), display=None), 'finite-state machine': Entry(definition='A fixed set of states with one move per input character; it can check patterns but not unlimited nesting.', lesson='primer.ml.structured_output', scope=(), display=None), 'flip rate': Entry(definition='The share of questions whose answer changes when the user asserts a wrong answer.', lesson='primer.ml.alignment', scope=(), display=None), 'floating point': Entry(definition='Storing a number as a sign, an exponent (range) and a mantissa (precision).', lesson='primer.ml.hardware', scope=(), display=None), 'flow matching': Entry(definition='Training a network to output the velocity along paths from noise to data, then following it to generate.', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'forward process': Entry(definition='The fixed, unlearned procedure that mixes data with noise step by step until only noise is left.', lesson='primer.ml.generative.diffusion', scope=('primer/ml/generative/diffusion',), display=None), 'fourier transform': Entry(definition='Splitting a signal into how much of each frequency it contains.', lesson='primer.ml.generative.multimodal', scope=(), display=None), 'fp16': Entry(definition='A 16-bit float with 5 exponent and 10 mantissa bits: more precision, range only up to 65,504.', lesson='primer.ml.hardware', scope=(), display=None), 'fp8': Entry(definition='8-bit floats: E4M3 (range to 448) for weights and activations, E5M2 (to 57,344) for gradients.', lesson='primer.ml.hardware', scope=(), display=None), 'frame sampling': Entry(definition='Keeping only some video frames, such as one a second, to save tokens.', lesson='primer.ml.generative.multimodal', scope=(), display=None), 'functional correctness': Entry(definition='Judging generated code by running it against tests rather than by comparing its text with a reference solution.', lesson='primer.ml.benchmarks', scope=(), display=None), 'fsdp': Entry(definition='Fully sharded data parallel: all training state sharded across GPUs (ZeRO stage 3).', lesson='primer.ml.pretraining', scope=(), display=None), 'freezing': Entry(definition="Keeping some of a model's weights fixed during training, so only the rest learn and the frozen part keeps what it already knew.", lesson='primer.ml.fine_tuning', scope=(), display=None), 'fused multiply-add': Entry(definition='One instruction that multiplies two numbers and adds the result to a running total.', lesson='primer.ml.hardware', scope=(), display=None), 'gan': Entry(definition="A generator and a discriminator trained against each other, so the generator learns to make samples the discriminator can't tell from real data.", lesson='primer.ml.generative.gans', scope=(), display=None), 'gated cross-attention': Entry(definition="Cross-attention layers added inside a frozen language model so its text can look at image features, with each layer's output multiplied by tanh of a learned number that starts at 0, so the model starts out unchanged.", lesson='primer.ml.generative.multimodal', scope=(), display=None), 'generative model': Entry(definition='A model that learns to produce new samples resembling its training data, such as new faces, voices or sentences.', lesson='primer.ml.generative.autoencoders', scope=(), display=None), 'gaussian': Entry(definition='The bell-curve distribution, set by its mean (where the centre is) and its standard deviation (how wide it is).', lesson='primer.ml.generative.autoencoders', scope=(), display=None), 'generalization gap': Entry(definition='Validation loss minus training loss; when it keeps growing, the model is memorising.', lesson='primer.ml.fine_tuning', scope=(), display=None), 'generator': Entry(definition='The network in a GAN that turns random noise into a sample.', lesson='primer.ml.generative.gans', scope=('primer/ml/generative/gans',), display=None), 'gini impurity': Entry(definition='The chance that two examples drawn at random from a pile carry different labels; 0 means pure.', lesson='primer.ml.classical', scope=(), display=None), 'global token': Entry(definition='A token every other token may attend to, and which attends to all, so any two tokens are two hops apart.', lesson='primer.ml.efficient_architectures', scope=(), display=None), "goodhart's law": Entry(definition='Once a measure becomes a target, optimizing it stops improving the thing it measured.', lesson='primer.ml.alignment', scope=(), display=None), 'gradient boosting': Entry(definition='Adding small trees one at a time, each fit to the current errors, with each step shrunk by a learning rate.', lesson='primer.ml.classical', scope=(), display=None), 'gradient penalty': Entry(definition='A discriminator loss term that punishes steep slopes of its output with respect to its input (R1, WGAN-GP).', lesson='primer.ml.generative.gans', scope=(), display=None), 'grammar': Entry(definition='A set of rules defining which strings are valid in a format.', lesson='primer.ml.structured_output', scope=('primer/ml/structured_output',), display=None), 'grpo': Entry(definition='Group Relative Policy Optimization: each answer is scored against other answers to the same prompt, so no value network is needed.', lesson='primer.ml.reinforcement', scope=(), display=None), 'gsm8k': Entry(definition='A benchmark of grade-school maths word problems (1,319 in its test set), scored by comparing the final number with the answer key.', lesson='primer.ml.benchmarks', scope=(), display=None), 'guidance scale': Entry(definition='The weight w in classifier-free guidance: 0 ignores the prompt, 1 follows it plainly, above 1 exaggerates it.', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'hann window': Entry(definition="A smooth rise and fall applied to each slice so its cut edges don't add false frequencies.", lesson='primer.ml.generative.multimodal', scope=(), display=None), 'harmful compliance': Entry(definition='Answering a request that should have been refused.', lesson='primer.ml.alignment', scope=(), display=None), 'harmonic mean': Entry(definition='n divided by the sum of the reciprocals of n numbers; dominated by the smallest, so it is high only when every number is high.', lesson='primer.ml.metrics', scope=(), display=None), 'hidden tests': Entry(definition="Grading tests the agent never sees, so it can't pass by satisfying the tests instead of the intent.", lesson='primer.agents.coding_agents', scope=(), display=None), 'host memory': Entry(definition="The CPU's main memory, reached from the GPU over a much slower link.", lesson='primer.ml.hardware', scope=(), display=None), 'huber loss': Entry(definition='A loss that is squared for small errors and grows only in proportion to the error beyond a threshold δ, so a few wild outliers cannot dominate the fit.', lesson='primer.ml.losses', scope=(), display=None), 'hybrid model': Entry(definition='A stack that mixes a few attention layers with many state-space layers.', lesson='primer.ml.efficient_architectures', scope=(), display=None), 'hyperparameter': Entry(definition='A setting chosen before training rather than learned from data, such as the learning rate, the model size or the number of epochs.', lesson='primer.ml.optimizers', scope=(), display=None), 'interaction effect': Entry(definition='Part of a prediction that depends on how two or more inputs combine, beyond what each contributes on its own: size mattering more in one city than another.', lesson=None, scope=(), display=None), 'imagenet': Entry(definition='A benchmark of about 1.3 million photos in 1,000 categories, the standard test of image classifiers; ImageNet-21k is its 14-million-photo, 21,000-category superset.', lesson='primer.ml.cnn_rnn', scope=(), display=None), 'implicit generative model': Entry(definition='A model that makes samples by pushing random noise through a fixed procedure, without giving the probability of any sample: a GAN generator, or a diffusion model sampled with DDIM.', lesson='primer.ml.generative.gans', scope=(), display=None), 'inception score': Entry(definition='A score for generated images from a pretrained image classifier: high when each image is confidently one class and the images spread over many classes. Higher is better; it never looks at real images.', lesson='primer.ml.generative.gans', scope=(), display=None), 'inductive bias': Entry(definition="The assumptions built into a model before it sees any data, such as a convolution's belief that nearby pixels matter most. Good assumptions help with little data; with enough data a model can learn them instead.", lesson='primer.ml.cnn_rnn', scope=(), display=None), 'importance sampling': Entry(definition='Estimating an average under one distribution from samples drawn under another, by weighting each sample by the ratio of its two probabilities.', lesson='primer.ml.reinforcement', scope=(), display=None), 'inpainting': Entry(definition='Filling in a missing or masked part of an image so that it fits the rest; a generative model does it by sampling only the unknown pixels.', lesson=None, scope=(), display=None), 'information gain': Entry(definition='How much a split lowers entropy; the tree picks the split that lowers it most.', lesson='primer.ml.classical', scope=(), display=None), 'induction head': Entry(definition='A pattern-completion mechanism: having seen A followed by B earlier in the text, predict B the next time A appears. It is thought to underlie much of in-context learning.', lesson=None, scope=(), display=None), 'instruction tuning': Entry(definition='Fine-tuning a pretrained model on many tasks written as instructions paired with good responses, so it learns to follow instructions it has never seen.', lesson='primer.ml.training_stages', scope=(), display=None), 'interference': Entry(definition="Signals getting in each other's way: other features leaking into one feature's reading (superposition), or merged task vectors changing the same weights in opposite directions (model merging).", lesson='primer.ml.interpretability', scope=('primer/ml/fine_tuning', 'primer/ml/interpretability'), display=None), 'interpretability': Entry(definition="Studying what a model's internal numbers represent, and which of them cause its outputs.", lesson='primer.ml.interpretability', scope=(), display=None), 'item response theory': Entry(definition="Modelling the chance of a right answer from a taker's ability minus a question's difficulty.", lesson='primer.ml.benchmarks', scope=(), display=None), 'jaccard similarity': Entry(definition='Shared items divided by all distinct items across two sets, from 0 to 1.', lesson='primer.ml.fine_tuning', scope=(), display=None), 'jailbreak': Entry(definition='A prompt crafted to talk a model out of its safety training, such as a role-play or a disguised request, so it produces what it would normally refuse.', lesson='primer.agents.guardrails', scope=(), display=None), 'jensen-shannon divergence': Entry(definition="A measure of how different two distributions are; stuck at log 2 whenever they don't overlap.", lesson='primer.ml.generative.gans', scope=(), display=None), 'json mode': Entry(definition='A decoding option that guarantees parseable JSON, but not any particular shape.', lesson='primer.ml.structured_output', scope=(), display=None), 'kl penalty': Entry(definition='A term subtracted from the reward that grows as the tuned model drifts from its starting point, measured by KL divergence: a leash that keeps it near the data the reward model was trained on.', lesson='primer.ml.reinforcement', scope=(), display=None), 'kv-cache quantization': Entry(definition='Storing cached keys and values in fewer bits, with one scale per vector.', lesson='primer.ml.efficient_architectures', scope=(), display=None), 'l0 norm': Entry(definition='The number of non-zero entries in a vector. For a sparse autoencoder, how many latents fire on an input.', lesson=None, scope=(), display=None), 'label noise': Entry(definition='Reference labels that are wrong, which cap what a model can learn and what an eval can measure.', lesson='primer.ml.fine_tuning', scope=(), display=None), 'langevin dynamics': Entry(definition='Sampling by repeatedly taking a small step along the score, towards where data is denser, and adding a little fresh noise. A diffusion sampler has this shape.', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'language identification': Entry(definition='Guessing which language a text is in, so a pipeline keeps only the ones it wants.', lesson='primer.ml.pretraining', scope=(), display=None), 'latent diffusion': Entry(definition="Running diffusion on an autoencoder's small compressed code instead of on pixels, then decoding once at the end.", lesson='primer.ml.generative.diffusion', scope=(), display=None), 'latent space': Entry(definition='The space of codes a model works in, where position means something and nearby codes decode to similar data.', lesson='primer.ml.generative.autoencoders', scope=(), display=None), 'law of large numbers': Entry(definition='The rule that the average of many independent random draws settles ever closer to the true average as more draws are added.', lesson=None, scope=(), display=None), 'layout shift': Entry(definition='Controls moving between runs, so clicks at remembered coordinates miss.', lesson='primer.agents.coding_agents', scope=(), display=None), 'leaf': Entry(definition='An end box of a decision tree, predicting the majority label or average value of the training examples that reached it.', lesson='primer.ml.classical', scope=('primer/ml/classical',), display=None), 'least squares': Entry(definition='Choosing the parameters that make the sum of squared errors as small as possible.', lesson='primer.ml.regularization', scope=(), display=None), 'linear time invariance': Entry(definition='A sequence model whose update rule is the same at every step, whatever the input; such a model can be computed as one convolution.', lesson='primer.ml.efficient_architectures', scope=(), display=None), 'linear representation hypothesis': Entry(definition='The idea that features are directions, readable with a dot product.', lesson='primer.ml.interpretability', scope=(), display=None), 'log-derivative trick': Entry(definition='Rewriting ∇π as π·∇log π, so a sampled action gives an estimate of the gradient.', lesson='primer.ml.reinforcement', scope=(), display=None), 'lipschitz': Entry(definition='A function is K-Lipschitz if its output never changes more than K times as fast as its input: a speed limit on its slope everywhere.', lesson='primer.ml.generative.gans', scope=(), display='Lipschitz'), 'line search': Entry(definition='Having picked a direction to move in, trying different step lengths along it and keeping the one that lowers the loss most.', lesson=None, scope=(), display=None), 'log-odds': Entry(definition='The logarithm of p / (1 − p): 0 for a 50% chance, positive when more likely than not. A sigmoid turns log-odds back into a probability.', lesson='primer.ml.classical', scope=(), display=None), 'log-likelihood': Entry(definition='The log of how probable the observed data is under a model.', lesson='primer.ml.benchmarks', scope=(), display=None), 'log-mel spectrogram': Entry(definition='A spectrogram pooled into mel bands with loudness on a log scale: what speech models read.', lesson='primer.ml.generative.multimodal', scope=(), display=None), 'logit difference': Entry(definition="The correct answer's score minus a wrong answer's score.", lesson='primer.ml.interpretability', scope=(), display=None), 'logit lens': Entry(definition="Applying the model's own output layer to intermediate layers to see what it would predict so far.", lesson='primer.ml.interpretability', scope=(), display=None), 'loss scaling': Entry(definition='Multiplying the loss so small fp16 gradients stay above zero, then dividing back.', lesson='primer.ml.hardware', scope=(), display=None), 'loss spike': Entry(definition='A sudden jump in training loss that may recover or diverge.', lesson='primer.ml.pretraining', scope=(), display=None), 'maj@k': Entry(definition='A score for sampling k answers per question and keeping the most common final answer: the share of questions where that majority answer is right.', lesson='primer.ml.reasoning', scope=(), display=None), 'majority voting': Entry(definition='Choosing the answer that the most samples agree on.', lesson='primer.ml.reasoning', scope=(), display=None), 'mantissa': Entry(definition='The bits of a float that store its digits, which set its precision.', lesson='primer.ml.hardware', scope=('primer/ml/hardware',), display=None), 'margin of error': Entry(definition='The ± range around a measured score that the true score probably falls in; it shrinks with the square root of the sample size.', lesson='primer.ml.fine_tuning', scope=(), display=None), 'marginal likelihood': Entry(definition='How probable a model finds an example, averaged over every hidden cause that could have produced it. With a neural network inside the model that average is an intractable integral, which is why VAEs train on a lower bound instead.', lesson='primer.ml.generative.autoencoders', scope=(), display=None), 'markov chain': Entry(definition='A sequence of random steps where each step depends only on the one just before it, not on the whole history.', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'master weights': Entry(definition="An fp32 copy of the weights that receives optimizer updates, so tiny updates aren't lost.", lesson='primer.ml.pretraining', scope=(), display=None), 'mel scale': Entry(definition='A relabelling of frequency so equal steps sound equally far apart to human ears.', lesson='primer.ml.generative.multimodal', scope=(), display=None), 'median': Entry(definition='The middle value once numbers are sorted (the average of the two middle ones for an even count). Unlike the mean, one extreme value barely moves it.', lesson=None, scope=(), display=None), 'memory bandwidth': Entry(definition='How many bytes per second a memory can deliver.', lesson='primer.ml.hardware', scope=(), display=None), 'memory hierarchy': Entry(definition='The chain of storage from registers to the network, each level bigger and slower than the last.', lesson='primer.ml.hardware', scope=(), display=None), 'micro-batch': Entry(definition='A slice of a batch that moves through a pipeline on its own.', lesson='primer.ml.pretraining', scope=('primer/ml/pretraining',), display=None), 'minhash': Entry(definition='A short signature per document whose matching slots estimate Jaccard similarity.', lesson='primer.ml.pretraining', scope=(), display=None), 'minimax': Entry(definition='A game where one player tries to make a number as large as possible and the other as small as possible.', lesson='primer.ml.generative.gans', scope=(), display=None), 'mixed precision': Entry(definition='Doing the big multiplies in 16- or 8-bit while keeping master weights and sums in fp32.', lesson='primer.ml.hardware', scope=(), display=None), 'mlp': Entry(definition='Multi-layer perceptron: the plainest neural network, layers of weighted sums each followed by a nonlinearity, with every unit connected to every unit in the next layer.', lesson='primer.ml.neural_net', scope=(), display=None), 'modality': Entry(definition='One kind of input a model can take: text, images, audio or video.', lesson='primer.ml.generative.multimodal', scope=('primer/ml/generative/multimodal',), display=None), 'model parallelism': Entry(definition="Splitting one model's weights across GPUs, either inside each layer (tensor parallelism) or by layers (pipeline parallelism). The ZeRO and Megatron-LM papers use it for the first.", lesson='primer.ml.pretraining', scope=(), display=None), 'mode collapse': Entry(definition='A generator producing only a few kinds of output instead of the full variety in the data.', lesson='primer.ml.generative.gans', scope=(), display=None), 'model collapse': Entry(definition='Losing rare data (the tails) when models are trained on their own outputs.', lesson='primer.ml.pretraining', scope=(), display=None), 'model editing': Entry(definition='Changing one specific fact or behaviour inside a trained model by adjusting a few weights directly, without retraining, while leaving everything else as it was.', lesson='primer.ml.interpretability', scope=(), display='model editing'), 'model merging': Entry(definition='Building one model from several fine-tunes by arithmetic on their weights, with no extra training.', lesson='primer.ml.fine_tuning', scope=(), display=None), 'monosemantic': Entry(definition='Responding to one understandable thing only: said of a neuron, or of a feature a sparse autoencoder finds.', lesson='primer.ml.interpretability', scope=(), display=None), 'multi-armed bandit': Entry(definition='The simplest RL problem: pick among options with hidden payouts, and learn which pays best by trying them.', lesson='primer.ml.reinforcement', scope=(), display=None), 'multi-head latent attention': Entry(definition="Caching one small latent vector per token and rebuilding every head's keys and values from it.", lesson='primer.ml.efficient_architectures', scope=(), display=None), 'multi-query attention': Entry(definition='All query heads share a single key/value head.', lesson='primer.ml.efficient_architectures', scope=(), display=None), 'multitask learning': Entry(definition='Training one model on several tasks at once, with part of the input saying which task to do, so what it learns for one task can help the others.', lesson=None, scope=(), display=None), 'multimodal model': Entry(definition='A model that takes in more than one kind of input (text, images, audio, video) as one sequence of tokens.', lesson='primer.ml.generative.multimodal', scope=(), display=None), 'n-gram overlap': Entry(definition="The share of a text's n-word runs that also appear in another corpus.", lesson='primer.ml.benchmarks', scope=(), display=None), 'neural ode': Entry(definition='A model whose output is found by following an ordinary differential equation whose rate of change is computed by a neural network; it can be run backwards and gives exact probabilities.', lesson=None, scope=(), display=None), "newton's method": Entry(definition='A step that uses the curvature (second derivative) as well as the slope: divide the slope by the curvature to jump to the bottom of the parabola that matches the loss at the current point.', lesson=None, scope=(), display=None), 'noise schedule': Entry(definition='How much noise each step adds (the betas), which sets how fast the signal fades.', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'non-saturating loss': Entry(definition='The generator loss −log D(G(z)), which keeps a strong gradient when the discriminator confidently rejects fakes.', lesson='primer.ml.generative.gans', scope=(), display=None), 'off-by-one error': Entry(definition='An index or count one position away from the right one.', lesson='primer.agents.coding_agents', scope=(), display=None), 'one-hot encoding': Entry(definition='Turning a category into one yes/no column per possible value.', lesson='primer.ml.classical', scope=(), display=None), 'optimizer state': Entry(definition="The running numbers an optimizer keeps for every weight between steps, such as Adam's two running averages. With an fp32 master copy of the weights it is 12 bytes per parameter, the biggest part of training memory.", lesson='primer.ml.pretraining', scope=(), display=None), 'optimal discriminator': Entry(definition='p_data/(p_data + p_g): the best possible verdict against a fixed generator; 1/2 everywhere at equilibrium.', lesson='primer.ml.generative.gans', scope=(), display=None), 'optimal transport': Entry(definition='Moving one distribution onto another as cheaply as possible, where moving mass further costs more. For the bell curves of flow matching, every bit of mass then travels in a straight line at constant speed.', lesson='primer.ml.generative.gans', scope=(), display=None), 'ordinary differential equation': Entry(definition='A rule giving, at every moment, how fast something is changing. Solving it means following that rule forward in time, for example with small Euler steps.', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'out-of-bag': Entry(definition="The rows left out of a tree's bootstrap sample, usable as free validation data for that tree.", lesson='primer.ml.classical', scope=(), display=None), 'outcome reward model': Entry(definition='A verifier that scores only the final answer.', lesson='primer.ml.reasoning', scope=(), display=None), 'outcome supervision': Entry(definition='Training a reward model from the final result alone: each solution is labelled right or wrong by its answer, never step by step.', lesson='primer.ml.reasoning', scope=(), display=None), 'outer product': Entry(definition='A column vector times a row vector, giving a table whose entry (m, c) is the product of their m-th and c-th numbers.', lesson='primer.ml.efficient_architectures', scope=(), display=None), 'over-refusal': Entry(definition='Refusing a harmless request because a safety check is too strict.', lesson='primer.ml.alignment', scope=(), display=None), 'overoptimization': Entry(definition="Optimizing against a learned reward model for so long that the true quality it stood for starts to fall, even as its score keeps rising: Goodhart's law for reward models.", lesson='primer.ml.alignment', scope=(), display=None), 'overthinking': Entry(definition='Spending many reasoning tokens where few would do, wasting cost and time and sometimes losing a right answer.', lesson='primer.ml.reasoning', scope=(), display=None), 'partial dependence plot': Entry(definition="A plot of a model's average prediction as one or two inputs are set to each value in turn, with every other input left as it is in the data.", lesson=None, scope=(), display=None), 'paired bootstrap': Entry(definition="Resampling questions with both models' results kept together, to test whether a gap is real.", lesson='primer.ml.benchmarks', scope=(), display=None), 'parallel scan': Entry(definition='Computing every state of a linear recurrence in about log₂ n parallel rounds by merging steps.', lesson='primer.ml.efficient_architectures', scope=(), display=None), 'pass-to-pass test': Entry(definition='A hidden test that passed before a fix and must still pass after it.', lesson='primer.agents.coding_agents', scope=(), display=None), 'pass@k': Entry(definition='The chance that at least one of k sampled answers passes the tests.', lesson='primer.ml.benchmarks', scope=(), display=None), 'pass@n': Entry(definition='The chance that at least one of n samples is right: 1 − (1 − p)ⁿ.', lesson='primer.ml.reasoning', scope=(), display=None), 'patch embedding': Entry(definition='The learned matrix that turns a flattened image patch into a token vector.', lesson='primer.ml.generative.multimodal', scope=(), display=None), 'perceiver resampler': Entry(definition='A small transformer whose queries are a fixed set of learned vectors: they cross-attend to any number of image or video features and always return that fixed number of visual tokens.', lesson='primer.ml.generative.multimodal', scope=(), display=None), 'perceptual loss': Entry(definition='A loss that compares two images through the features of a pretrained network rather than pixel by pixel, so it punishes the differences a person would notice. LPIPS is a widely used one.', lesson='primer.ml.generative.gans', scope=(), display=None), 'pearson correlation': Entry(definition='How closely two lists of numbers rise and fall together along a straight line, from −1 to 1; 0 means no straight-line relationship.', lesson=None, scope=(), display=None), 'permutation importance': Entry(definition='The accuracy lost on held-out data when one column is shuffled.', lesson='primer.ml.classical', scope=(), display=None), 'pipeline bubble': Entry(definition='Time pipeline stages sit idle while the pipeline fills and drains.', lesson='primer.ml.pretraining', scope=(), display=None), 'pipeline parallelism': Entry(definition='Giving each GPU a stage of layers and passing activations between neighbouring stages.', lesson='primer.ml.hardware', scope=(), display=None), 'policy gradient': Entry(definition='Raising expected reward by making the actions that earned more reward more likely.', lesson='primer.ml.reinforcement', scope=(), display=None), 'polysemantic neuron': Entry(definition='A neuron that responds to several unrelated features.', lesson='primer.ml.interpretability', scope=(), display=None), 'posterior': Entry(definition='What you believe about a hidden quantity after seeing the data: the prior, reweighted by how well each value explains what was observed.', lesson='primer.ml.generative.autoencoders', scope=(), display=None), 'posterior collapse': Entry(definition="When a VAE's code carries no information because the KL penalty outweighs what the code saves in rebuild error, so every output is the same average.", lesson='primer.ml.generative.autoencoders', scope=(), display=None), 'power iteration': Entry(definition='Repeatedly multiplying a vector by a matrix (and its transpose) until it points along the most-stretched direction.', lesson='primer.ml.generative.gans', scope=(), display=None), 'power law': Entry(definition='A relationship where one quantity is a fixed power of another, y = a·x^k; on log-log axes it is a straight line.', lesson='primer.ml.transformer', scope=(), display=None), 'preference model': Entry(definition='Another name for a reward model: it scores a response so that the gap between two scores predicts which one people (or a model) prefer.', lesson='primer.ml.training_stages', scope=(), display=None), 'prior': Entry(definition="What you believe about a hidden quantity before seeing any data. A VAE's prior over codes is the standard normal distribution.", lesson='primer.ml.generative.autoencoders', scope=('primer/ml/generative',), display=None), 'probability flow ode': Entry(definition='The deterministic equation whose solutions carry noise to data with the same in-between distributions as a diffusion process; DDIM sampling is one way of stepping along it.', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'privileged basis': Entry(definition='Directions made special by the architecture, such as neurons followed by an activation function that acts on each number separately. Only in a privileged basis does asking what one neuron means make sense.', lesson='primer.ml.interpretability', scope=(), display=None), 'probability ratio': Entry(definition="The current policy's probability of a sampled action divided by its probability when the action was sampled.", lesson='primer.ml.reinforcement', scope=(), display=None), 'probe': Entry(definition="A small classifier trained on a frozen model's activations to test whether a property is encoded there.", lesson='primer.ml.interpretability', scope=('primer/ml/interpretability',), display=None), 'process supervision': Entry(definition='Training a reward model from labels on each intermediate step, so it learns where a solution went wrong.', lesson='primer.ml.reasoning', scope=(), display=None), 'process reward model': Entry(definition='A verifier that scores each intermediate step.', lesson='primer.ml.reasoning', scope=(), display=None), 'projector': Entry(definition="A small layer that maps an encoder's vectors into a language model's embedding space.", lesson='primer.ml.generative.multimodal', scope=('primer/ml/generative/multimodal',), display=None), 'pushdown automaton': Entry(definition='A finite-state machine plus a stack: enough to check nested formats like JSON or SQL.', lesson='primer.ml.structured_output', scope=(), display=None), 'quantile': Entry(definition='The value below which a given share of the data falls: the 0.5 quantile is the median, the 0.9 quantile has 90% of values below it.', lesson=None, scope=(), display=None), 'quality filter': Entry(definition='A rule or classifier that throws out low-quality pages before training.', lesson='primer.ml.pretraining', scope=(), display=None), 'random forest': Entry(definition='Many deep trees on bootstrap samples, each split limited to a random subset of features, with their votes averaged.', lesson='primer.ml.classical', scope=(), display=None), 'reasoning model': Entry(definition='A model trained to write out and check intermediate steps before it commits to an answer.', lesson='primer.ml.reasoning', scope=(), display=None), 'rectified flow': Entry(definition='Flow matching with straight-line paths, retrained on its own outputs so the paths get straighter and need fewer steps.', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'reduce-scatter': Entry(definition='Summing a vector across GPUs so each GPU ends up holding the total for only its own slice; the first half of a ring all-reduce.', lesson='primer.ml.pretraining', scope=(), display=None), 'reflow': Entry(definition="Retraining a rectified flow on its own (noise, sample) pairs. The new pairs' straight lines rarely cross, so the new flow's paths are straighter and need fewer steps.", lesson='primer.ml.generative.diffusion', scope=(), display=None), 'red-teaming': Entry(definition='Searching systematically for inputs that make a model or its safety checks fail.', lesson='primer.ml.alignment', scope=(), display=None), 'register': Entry(definition='The tiny storage right beside the arithmetic units, holding the numbers being worked on this instant.', lesson='primer.ml.hardware', scope=('primer/ml/hardware',), display=None), 'regular expression': Entry(definition='A pattern language (such as [0-9]+ or cat|car|dog) that can always be compiled into a finite-state machine.', lesson='primer.ml.structured_output', scope=(), display=None), 'reinforce': Entry(definition='The basic policy-gradient algorithm: step along reward times the gradient of log π(action).', lesson='primer.ml.reinforcement', scope=(), display=None), 'reinforcement learning': Entry(definition='Learning from a score for what you did, rather than from the correct answer.', lesson='primer.ml.reinforcement', scope=(), display=None), 'rejection sampling': Entry(definition='Generating many candidate outputs and keeping only those that pass a check, such as a correct final answer, often to use as training data.', lesson=None, scope=(), display=None), 'release gate': Entry(definition='Limits set before measuring; a release goes ahead only if every evaluation is within its limit.', lesson='primer.ml.alignment', scope=(), display=None), 'reparameterization trick': Entry(definition='Writing a random draw as z = μ + σ·ε with the noise ε as a separate input, so gradients can flow through sampling.', lesson='primer.ml.generative.autoencoders', scope=(), display=None), 'replay': Entry(definition='Mixing a small sample of old-task examples into new training data so the old skill keeps getting practised.', lesson='primer.ml.fine_tuning', scope=('primer/ml/fine_tuning',), display=None), 'residual': Entry(definition='The true value minus the current prediction; for squared error it is the negative gradient.', lesson='primer.ml.classical', scope=('primer/ml/classical',), display=None), 'residual vector quantization': Entry(definition='Stacking codebooks, each one encoding the error the previous ones left.', lesson='primer.ml.generative.multimodal', scope=(), display=None), 'resnet': Entry(definition='A deep convolutional network built from residual blocks (x + f(x)), which made networks of a hundred layers and more trainable.', lesson='primer.ml.deep_nets', scope=(), display=None), 'resolved rate': Entry(definition='The share of tasks whose patch passes every hidden fail-to-pass and pass-to-pass test.', lesson='primer.agents.coding_agents', scope=(), display=None), 'reverse process': Entry(definition='The learned half of a diffusion model: a chain of small denoising steps that turns pure noise back into data.', lesson='primer.ml.generative.diffusion', scope=('primer/ml/generative/diffusion',), display=None), 'reward': Entry(definition='The single number the environment returns to say how good an action was.', lesson='primer.ml.reinforcement', scope=('primer/ml/reinforcement',), display=None), 'reward hacking': Entry(definition='A policy maximising the reward as written while the real goal gets worse.', lesson='primer.ml.reinforcement', scope=(), display=None), 'ring all-reduce': Entry(definition='All-reduce by passing chunks around a ring; each GPU sends about twice its data, however many GPUs there are.', lesson='primer.ml.hardware', scope=(), display=None), 'rlaif': Entry(definition='Reinforcement learning from AI feedback: preference labels come from a model applying written principles.', lesson='primer.ml.alignment', scope=(), display=None), 'rolling buffer cache': Entry(definition='A KV cache with w slots where each new token overwrites the oldest.', lesson='primer.ml.efficient_architectures', scope=(), display=None), 'sample rate': Entry(definition='How many measurements of a signal are taken per second, such as 16,000 for speech.', lesson='primer.ml.generative.multimodal', scope=(), display=None), 'sandbox': Entry(definition='An isolated place to run untrusted code, where the worst it can do is fail: no network, no secrets, time and memory limits.', lesson='primer.agents.coding_agents', scope=('primer/agents/coding_agents',), display=None), 'scaling law': Entry(definition="A smooth, predictable rule for how a model's loss falls as its size, data or compute grows: a straight line on log-log axes. It lets small runs forecast a big model's quality.", lesson='primer.ml.transformer', scope=(), display=None), 'score': Entry(definition='The direction in which data gets more crowded fastest; the noise guess, flipped and rescaled.', lesson='primer.ml.generative.diffusion', scope=('primer/ml/generative/diffusion',), display=None), 'screenshot': Entry(definition='The image of the screen a computer-use agent receives after each action.', lesson='primer.agents.coding_agents', scope=('primer/agents/coding_agents',), display=None), 'selective state-space model': Entry(definition='A state-space model whose step size (how much to keep and write) is computed from each token, as in Mamba.', lesson='primer.ml.efficient_architectures', scope=(), display=None), 'self-consistency': Entry(definition='Sampling several chains of thought and returning the most common final answer.', lesson='primer.ml.reasoning', scope=(), display=None), 'self-supervised learning': Entry(definition='Learning from unlabelled data by predicting a hidden part of each example from the rest, such as a masked word or a masked stretch of audio.', lesson=None, scope=(), display=None), 'signal-to-noise ratio': Entry(definition='How much stronger a signal is than the noise mixed into it, as a ratio of their powers (variances). Audio quotes it in decibels, where every 10 dB is ten times the ratio; in diffusion it falls from very large (clean) to nearly zero (pure noise).', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'shingle': Entry(definition='A window of k neighbouring words, used to compare texts for near-duplicates.', lesson='primer.ml.fine_tuning', scope=(), display=None), 'short-time fourier transform': Entry(definition='A Fourier transform on each short, overlapping slice of a signal.', lesson='primer.ml.generative.multimodal', scope=(), display=None), 'shots': Entry(definition='Worked examples placed in the prompt before the question.', lesson='primer.ml.benchmarks', scope=('primer/ml/benchmarks',), display=None), 'shrinkage': Entry(definition="Pulling values toward zero: an L1 penalty's pull on every activation, or in gradient boosting the learning rate that keeps only part of each new tree's correction.", lesson='primer.ml.interpretability', scope=('primer/ml/classical', 'primer/ml/interpretability'), display=None), 'sliding-window attention': Entry(definition='Each token attends only to the last w tokens, so cost and cache stop growing with context length.', lesson='primer.ml.efficient_architectures', scope=(), display=None), 'softplus': Entry(definition='log(1 + eˣ): a smooth ramp that is always positive.', lesson='primer.ml.efficient_architectures', scope=(), display=None), 'sparse attention': Entry(definition='Attention that scores only a chosen pattern of token pairs (local, global, strided) instead of every pair.', lesson='primer.ml.efficient_architectures', scope=(), display=None), 'sparse autoencoder': Entry(definition='A wide encoder and decoder trained with an L1 penalty so each input uses a few learned features.', lesson='primer.ml.interpretability', scope=(), display=None), 'specification gaming': Entry(definition='Another name for reward hacking: satisfying the letter of an objective but not its intent.', lesson='primer.ml.reinforcement', scope=(), display=None), 'spectral normalization': Entry(definition="Dividing each layer's weights by the most they can stretch any input, which caps how fast the discriminator can change.", lesson='primer.ml.generative.gans', scope=(), display=None), 'spectrogram': Entry(definition='A picture of sound: time across, frequency up, brightness for loudness.', lesson='primer.ml.generative.multimodal', scope=(), display=None), 'speech recognition': Entry(definition='Turning recorded speech into written text; also called automatic speech recognition (ASR).', lesson='primer.ml.generative.multimodal', scope=(), display=None), 'standard error': Entry(definition='The typical distance between a measured score and the true rate: √(p(1−p)/n).', lesson='primer.ml.benchmarks', scope=(), display=None), 'standard normal distribution': Entry(definition='The bell curve centred on 0 with spread 1; N(0, I) draws each number from it independently.', lesson='primer.ml.generative.autoencoders', scope=(), display=None), 'steering': Entry(definition="Changing a model's behaviour while it runs by editing an internal activation, such as adding or pinning a feature's direction.", lesson='primer.ml.interpretability', scope=('primer/ml/interpretability',), display=None), 'stochastic differential equation': Entry(definition='A rule for how something changes over time with two parts: a steady drift, and random jitter of a set strength. The noising process of a diffusion model is one, and running it backwards generates data.', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'stop-gradient': Entry(definition='An operation that passes its input through unchanged but blocks gradients from flowing back into it, so training treats that input as a constant.', lesson=None, scope=(), display=None), 'straight-through estimator': Entry(definition='A way to train through a step that has no useful gradient, such as rounding or snapping to a codebook: use the step going forward, and pass the gradient back as if the step were not there.', lesson=None, scope=(), display=None), 'strict mode': Entry(definition="A tool or output option that guarantees the model's JSON fits a given schema.", lesson='primer.ml.structured_output', scope=(), display=None), 'structured output': Entry(definition="Making a model's answer follow an exact format, such as JSON that fits a schema, so a program can read it.", lesson='primer.ml.structured_output', scope=(), display=None), 'stump': Entry(definition='A decision tree with a single question and two leaves.', lesson='primer.ml.classical', scope=('primer/ml/classical',), display=None), 'style control': Entry(definition='Adding length or format as extra factors in the rating fit, so style is separated from quality.', lesson='primer.ml.benchmarks', scope=(), display=None), 'subnormal': Entry(definition='A float below the smallest normal value, with the hidden leading 1 dropped so it fades toward zero.', lesson='primer.ml.hardware', scope=(), display=None), 'superposition': Entry(definition='Storing more features than there are neurons, as nearly perpendicular directions.', lesson='primer.ml.interpretability', scope=(), display=None), 'super-resolution': Entry(definition='Turning a low-resolution image into a plausible higher-resolution one; most of the fine detail has to be invented, not recovered.', lesson=None, scope=(), display=None), 'swiglu': Entry(definition='A gated feed-forward layer: one projection, passed through the smooth SiLU activation, multiplies a second projection number by number before the output projection.', lesson='primer.ml.transformer', scope=(), display=None), 'sycophancy': Entry(definition='A model changing its answer to agree with a view the user states.', lesson='primer.ml.alignment', scope=(), display=None), 'synthetic data': Entry(definition='Training examples written by a model rather than collected from people.', lesson='primer.ml.pretraining', scope=(), display=None), 'tabular data': Entry(definition='Data in rows and columns, where each column is a meaningful quantity in its own units.', lesson='primer.ml.classical', scope=(), display=None), 'task arithmetic': Entry(definition="Combining or removing skills by adding, scaling or subtracting task vectors from a base model's weights.", lesson='primer.ml.fine_tuning', scope=(), display=None), 'task vector': Entry(definition="The change a fine-tune made to a model's weights: fine-tuned minus base.", lesson='primer.ml.fine_tuning', scope=(), display=None), 'tensor parallelism': Entry(definition='Splitting each matrix multiply across GPUs, which must talk inside every layer.', lesson='primer.ml.hardware', scope=(), display=None), 'test-time compute': Entry(definition='Computation spent while answering rather than while training, such as longer chains or more samples.', lesson='primer.ml.reasoning', scope=(), display=None), 'text extraction': Entry(definition="Pulling a web page's main text out of its HTML, leaving menus and adverts behind.", lesson='primer.ml.pretraining', scope=(), display=None), 'text normalization': Entry(definition='Rewriting text into one standard form (case, punctuation, contractions, how numbers are written) so two texts are compared on their words, not their style.', lesson=None, scope=(), display=None), 'thinking budget': Entry(definition='The maximum number of tokens a model may spend reasoning before it must answer.', lesson='primer.ml.reasoning', scope=(), display=None), 'throughput': Entry(definition='Work finished per second, such as tokens generated per second across all requests; it rises with batch size, while latency is the wait for one request.', lesson='primer.ml.inference', scope=(), display=None), 'tiling': Entry(definition='Loading a block of data into fast memory once and doing all its work before evicting it.', lesson='primer.ml.hardware', scope=('primer/ml/hardware',), display=None), 'token healing': Entry(definition='Backing up over the last prompt token so the model can rewrite it, when a prompt ends partway through what would normally be one token.', lesson='primer.ml.structured_output', scope=(), display=None), 'transfer learning': Entry(definition='Training a model on a large general task first, then reusing it (usually by fine-tuning) on a smaller task it was never trained for.', lesson='primer.ml.training_stages', scope=(), display=None), 'translation equivariance': Entry(definition='Shift the input and the output shifts the same way: a convolution finds an edge wherever it sits, because the same filter slides over every position.', lesson='primer.ml.cnn_rnn', scope=(), display=None), 'tubelet': Entry(definition='A video patch that spans several frames as well as a square of pixels.', lesson='primer.ml.generative.multimodal', scope=(), display=None), 'tuned lens': Entry(definition='A logit lens with a small learned translator per layer.', lesson='primer.ml.interpretability', scope=(), display=None), 'two time-scale update rule': Entry(definition='Giving the generator and discriminator different learning rates so the game converges.', lesson='primer.ml.generative.gans', scope=(), display=None), 'u-net': Entry(definition='A convolutional network shaped like a U: it shrinks an image step by step to see the big picture, then grows it back, with shortcuts carrying fine detail across. The classic diffusion denoiser before transformers.', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'underflow': Entry(definition='A number too small for its format, rounded to zero.', lesson='primer.ml.pretraining', scope=('primer/ml/pretraining',), display=None), 'unbiased estimator': Entry(definition='A way of estimating a number from random data whose average, over every dataset you might have drawn, equals the true value: it is not systematically high or low.', lesson='primer.ml.benchmarks', scope=(), display=None), 'unit test': Entry(definition='A small program that runs one piece of code on chosen inputs and checks the outputs, passing or failing automatically.', lesson='primer.agents.coding_agents', scope=(), display=None), 'unbiased estimate': Entry(definition='An estimate that is right on average: any single one may be off, but the errors cancel over many tries.', lesson='primer.ml.reinforcement', scope=(), display=None), 'unembedding': Entry(definition='The final matrix of a language model: it turns the last hidden vector into one score (logit) per vocabulary token.', lesson='primer.ml.interpretability', scope=(), display=None), 'vae': Entry(definition='Variational autoencoder: an autoencoder whose codes are pulled towards a bell curve, so random codes decode to new data.', lesson='primer.ml.generative.autoencoders', scope=(), display=None), 'value function': Entry(definition='A prediction, made partway through a task, of how well it will end from here. A verifier that scores a solution after every token is one.', lesson='primer.ml.reinforcement', scope=('primer/ml/reinforcement',), display=None), 'value network': Entry(definition="A second model (the critic) that predicts expected reward, used as PPO's baseline.", lesson='primer.ml.reinforcement', scope=(), display=None), 'variational autoencoder': Entry(definition='An autoencoder whose encoder outputs a fuzzy region (mean and spread) pulled towards the standard normal, so random codes decode to new data.', lesson='primer.ml.generative.autoencoders', scope=(), display=None), 'variational inference': Entry(definition='Approximating a distribution you cannot compute, usually a posterior, with the closest member of a simple family by maximising a lower bound. A VAE does it with a network that outputs the approximation for each example.', lesson='primer.ml.generative.autoencoders', scope=(), display=None), 'vector quantization': Entry(definition='Replacing a vector with the number of its nearest codebook entry.', lesson='primer.ml.generative.multimodal', scope=(), display=None), 'velocity field': Entry(definition='A map giving, at every point and time, which way and how fast a sample should move.', lesson='primer.ml.generative.diffusion', scope=(), display=None), 'verifiable reward': Entry(definition="A reward computed by a check that can't be argued with, such as a correct answer or passing tests.", lesson='primer.ml.reinforcement', scope=(), display=None), 'verifier': Entry(definition='Anything that scores a candidate solution, from a unit test to a learned model.', lesson='primer.ml.reasoning', scope=('primer/ml/reasoning',), display=None), 'vision-language model': Entry(definition='A model that reads images (and often video) together with text and writes text; a language model given eyes.', lesson='primer.ml.generative.multimodal', scope=(), display=None), 'visual instruction tuning': Entry(definition='Fine-tuning a vision-language model on images paired with instructions and good answers.', lesson='primer.ml.generative.multimodal', scope=(), display=None), 'voice activity detection': Entry(definition='Deciding which stretches of a recording contain speech at all, so silence, music and noise are not transcribed.', lesson=None, scope=(), display=None), 'vq-vae': Entry(definition='A VAE that snaps each code vector to the nearest entry of a learned codebook, turning images or audio into tokens.', lesson='primer.ml.generative.autoencoders', scope=(), display=None), 'wasserstein distance': Entry(definition="The least work to reshape one distribution into another (mass moved times distance); also called earth mover's distance.", lesson='primer.ml.generative.gans', scope=(), display=None), 'waveform': Entry(definition='Sound recorded as a list of air-pressure measurements over time.', lesson='primer.ml.generative.multimodal', scope=(), display=None), 'weak supervision': Entry(definition='Training on labels that are plentiful but noisy or imperfect, such as captions and transcripts found on the web, instead of a small set checked by experts.', lesson=None, scope=(), display=None), 'weight averaging': Entry(definition='Merging fine-tunes of the same base by averaging their weights (task arithmetic with λ = 1/T).', lesson='primer.ml.fine_tuning', scope=(), display=None), 'wiener process': Entry(definition='Continuous random jitter, also called Brownian motion: over any short time dt it moves by a fresh bell-curve amount with variance dt, independent of everything before.', lesson=None, scope=(), display=None), 'weight clipping': Entry(definition="Forcing every weight of a network back into a small range, such as −0.01 to 0.01, after each update; the original Wasserstein GAN's crude way to cap its critic's slope.", lesson='primer.ml.generative.gans', scope=(), display=None), "winner's curse": Entry(definition='The best of many versions chosen on a test looks better on it than it really is.', lesson='primer.ml.benchmarks', scope=(), display=None), 'word error rate': Entry(definition="The share of a reference transcript's words that a system gets wrong: the substitutions, deletions and insertions needed to turn its output into the reference, divided by the reference's word count.", lesson=None, scope=(), display=None), 'zero-order hold': Entry(definition='A discretization rule that assumes the input stays constant for the whole step; it gives the keep factor e^(ΔA) of a state-space model.', lesson='primer.ml.efficient_architectures', scope=(), display=None), 'zero': Entry(definition='Sharding optimizer state, then gradients, then weights across data-parallel GPUs.', lesson='primer.ml.pretraining', scope=('primer/ml/pretraining',), display=None), 'β-vae': Entry(definition='A VAE whose KL penalty is weighted by β, trading rebuild sharpness for a smoother, more organized code space.', lesson='primer.ml.generative.autoencoders', scope=(), display=None)}
def display_term(term: str) -> str: on GitHub
967def display_term(term: str) -> str:
968    """How the primer usually writes a term: "acl" -> "ACL", "adamw" -> "AdamW", "attention" -> "attention".
969
970    Counted over every lesson's and companion's text, mid-sentence only, where capitals mean something
971    about the word. A term the lessons never use mid-sentence keeps its key, and an
972    entry's own `display` overrides the count.
973    """
974    if term in GLOSSARY and GLOSSARY[term].display:
975        return GLOSSARY[term].display
976    terms = tuple(sorted(GLOSSARY)) if term in GLOSSARY else (term,)
977    forms = _usage(terms).get(term.lower())
978    return forms.most_common(1)[0][0] if forms else term

How the primer usually writes a term: "acl" -> "ACL", "adamw" -> "AdamW", "attention" -> "attention".

Counted over every lesson's and companion's text, mid-sentence only, where capitals mean something about the word. A term the lessons never use mid-sentence keeps its key, and an entry's own display overrides the count.

def lesson_path(dotted: str) -> str: on GitHub
981def lesson_path(dotted: str) -> str:
982    """"primer.ml.attention" -> "primer/ml/attention.html", the pdoc page for that module."""
983    return dotted.replace(".", "/") + ".html"

"primer.ml.attention" -> "primer/ml/attention.html", the pdoc page for that module.

def glossary_js() -> str: on GitHub
986def glossary_js() -> str:
987    """The glossary as a script that sets `window.PRIMER_GLOSSARY` (works from file:// too)."""
988    data = {
989        term: {"def": e.definition, "lesson": lesson_path(e.lesson) if e.lesson else None, "term": display_term(term)}
990        for term, e in sorted(GLOSSARY.items())
991    }
992    return "window.PRIMER_GLOSSARY = " + json.dumps(data, ensure_ascii=False, indent=1) + ";\n"

The glossary as a script that sets window.PRIMER_GLOSSARY (works from file:// too).