An annotated companion · AI Primer

Mistral 7B, annotated

About this page. This is a companion, not a copy. It follows the paper section by section, quotes at most a sentence or two per section (clearly marked and attributed), and explains everything in its own words. The paper is released under a CC BY 4.0 licence; its tables are reproduced with attribution, and its figures are redrawn as new interactive diagrams with new example text (the opening of Dickens's A Tale of Two Cities, which is in the public domain). Every section heading links to the original.

How to read this page

  • Any dotted word explains itself on hover, focus or tap, and so does every symbol in every equation.
  • Three of the paper's figures are redrawn as widgets: a window explorer in §2 with sliders for the window and the number of layers, a rolling buffer you can step through, and the chunked prefill mask, one chunk at a time.

You will get the most from this page after the multi-query attention companion (the idea behind grouped-query attention) and the sliding-window section of the efficient architectures lesson, which builds every mechanism here in code. Every idea climbs the ladder: everyday picture, tiny example, diagram, the math, why it matters today.

Abstract · original

“Mistral 7B outperforms the best open 13B model (Llama 2) across all evaluated benchmarks, and the best released 34B model (Llama 1) in reasoning, mathematics, and code generation.”Jiang et al. (2023), Abstract

Everyday picture

Two cars finish a race level. One has an engine twice the size. The smaller one is not just as good, it is cheaper to run for every mile afterwards. This paper's claim is that a language model's size should be judged the same way: not only by how well it scores, but by what it costs every time someone uses it.

What the paper claims

  • A model with about 7 billion parameters scores higher than Llama 2 13B on the benchmarks the authors ran (with one near-tie, below), and than Llama 1 34B in several categories.
  • Its attention uses grouped-query attention (GQA) for faster generation and sliding-window attention (SWA) so that long sequences cost less.
  • A version fine-tuned to follow instructions, Mistral 7B Instruct, beats Llama 2 13B Chat on human and automated comparisons.
  • The weights are released under the Apache 2.0 licence.

Why it matters today

The paper says little about how the model was trained (nothing about its training data) and a lot about the machinery of cheap inference. That machinery, sharing key/value heads and bounding the cache, is what this page is mostly about, and it is now common across open models.

1 Introduction · original

Everyday picture

A restaurant can serve more diners by hiring a bigger kitchen, or by redesigning the kitchen it has. The introduction argues that the field has mostly been building bigger kitchens (larger models), which raises the cost of every meal (latency and compute per answer), and that careful design can deliver the same food from a smaller one.

Tiny example

The two design choices each target one of the costs from the multi-query attention paper. GQA shrinks how big each remembered token is: with 32 query heads sharing 8 key/value heads, every token's entry in the KV cache is 4 times smaller, so more conversations fit on a GPU at once, raising throughput. SWA limits how many tokens each layer looks at, so the attention work and the cache stop growing with the length of the text.

Why it matters

The paper also points to deployment: a reference implementation, and serving through vLLM, the system built on PagedAttention. Efficiency here is not a side note; it is the product.

2 Architectural details · original

Mistral 7B is a transformer, and the paper describes it by how it differs from Llama. Everything it lists is about attention and its cache.

The shape of the model, and GQA · original

Everyday picture

A model's shape is a short list of numbers, like a building's floor plan: how many floors, how wide each floor, how many rooms. From that list alone you can work out how much memory the model needs to remember each token of a conversation.

Table 1, reproduced with attribution (Jiang et al., 2023, CC BY 4.0): model architecture
ParameterValueWhat it means
dim4,096width of each token's vector
n_layers32transformer blocks stacked
head_dim128width of one attention head
hidden_dim14,336inner width of each feed-forward network
n_heads32query heads per layer (32 × 128 = 4,096)
n_kv_heads8key/value heads per layer, each shared by 4 query heads
window_size4,096how far back each layer can look
context_len8,192the context length listed for the model
vocab_size32,000tokens in the vocabulary

Tiny example

Every token leaves behind one key and one value per key/value head in every layer. With 8 key/value heads of 128 numbers, 32 layers, and 16-bit numbers (2 bytes, an assumption: the paper does not state the precision), that is 2 × 32 × 8 × 128 × 2 = 131,072 bytes, 128 KiB per token. If each of the 32 query heads kept its own keys and values, it would be 512 KiB.

In words: “a key and a value, per key/value head, per layer, at so many bytes per number; and each key/value head is shared by g query heads.”

With the numbers: 2 × 32 × 8 × 128 × 2 = 131,072 bytes; with Hkv = H = 32 it would be 524,288. The group size is g = 32 / 8 = 4, and the cache is 4 times smaller.

In Python:

L, H, H_kv, d_h, b = 32, 32, 8, 128, 2
per_token = 2 * L * H_kv * d_h * b
per_token  # → 131072
# in KiB
per_token / 1024  # → 128.0
# if every query head had its own keys and values
2 * L * H * d_h * b  # → 524288
# query heads per key/value head
g = H // H_kv
g  # → 4

Why it matters

Grouped-query attention sits between the two designs of the multi-query attention paper: one key/value head per query head (multi-head) and one for all of them (multi-query). The paper cites Ainslie et al. (2023) for it and uses it, in its words, to accelerate inference and allow higher batch sizes. The primer's MultiHeadAttention is grouped-query attention whenever n_kv_heads is smaller than n_heads, and kv_cache_bytes_per_token is the formula above.

Sliding window attention · original

“SWA exploits the stacked layers of a transformer to attend information beyond the window size W.”Jiang et al. (2023), §2

Everyday picture

A rumour passes down a line of people, but each person only talks to the few people nearest them. Nobody speaks to the far end of the line directly, yet the rumour gets there, one short hop at a time. In sliding-window attention each token, in each layer, attends only to a window of nearby earlier tokens. Because each layer reads the previous layer's results, which already mixed in their own neighbours, information hops further back with every layer.

Tiny example

Twelve tokens: “It was the best of times it was the worst of times”, numbered 0 to 11. With a window of 3, as in the paper's Figure 1, token 11 (“times”) reads tokens 9, 10 and 11 in each layer. After one layer it has heard from token 9; token 9 had itself heard from token 7 in the layer below; so after two layers token 11 reaches token 7, and after three layers token 5. Each layer adds 2 positions of reach.

Figure 1 of the paper (the sliding-window mask and the effective context length), redrawn with new example text. Based on Jiang et al. (2023), Figure 1 (CC BY 4.0).

Reading it: the upper grid is one layer's mask, like the middle panel of the paper's figure: a row for each token doing the looking, a column for each token being looked at, filled where a score is computed; row 11, the token followed below, is drawn darker. It is a band of width W along the diagonal, never above it (no token sees the future). The lower picture stacks the layers, with the input at the bottom and the last layer on top. The filled cells in each row are the tokens that can still influence token 11 at the top: at every layer down, the band spreads W − 1 positions further back, so the bottom row shows everything token 11 can hear. Slide the window up and each layer's band widens; add layers and the reach grows in steps. With W = 3 and three layers, token 11 hears from tokens 5 to 11, and tokens 0 to 4 are out of reach, yet no single layer ever looked more than 2 positions back.

In words: “the state at position i in layer k reads the previous layer's states from W positions back up to itself; stacking k such layers lets information travel up to W times k positions.”

With the numbers: Mistral 7B has W = 4,096 and 32 layers, so the span is 4,096 × 32 = 131,072 tokens, the paper's “approximately 131K”. The widget, the paper's figures and the primer's lesson count the window as W positions including the token itself, which reaches W − 1 back per layer: 32 × 4,095 = 131,040. The paper's text counts from i − W to i, which is W + 1 positions. Either way the answer is about 131 thousand tokens from a window of 4 thousand.

In Python:

W, k = 4096, 32
# the paper's span: W × k
W * k  # → 131072
# counting the window as W positions including the token itself
k * (W - 1)  # → 131040
# the tiny example: W = 3 (itself and 2 back), token 11, three layers
w, layers, i = 3, 3, 11
i - layers * (w - 1)  # → 5

What it saves. The paper reports that at a sequence length of 16K with W = 4,096, changes it made to FlashAttention and xFormers give a 2x speed-up over a vanilla attention baseline. Counting pairs gives a sense of why (this page's own arithmetic, not a measurement): a full causal mask over 16,384 tokens scores 134 million query-key pairs per head, a window of 4,096 scores 58.7 million, 2.3 times fewer.

In Python:

n, W = 16384, 4096
# full causal attention: token i reads i + 1 tokens
full = sum(i + 1 for i in range(n))
# sliding window: token i reads at most W tokens
window = sum(min(i + 1, W) for i in range(n))
full, window  # → (134225920, 58722304)
round(full / window, 2)  # → 2.29

Why it matters

The window turns attention's cost from growing with the square of the length to growing in proportion to it, while stacked layers keep a long, if indirect, reach. The price is that a fact from far back must survive several hops to arrive. The lesson builds the band with sliding_window_mask, counts it with pairs_computed, and checks the reach by applying the mask layer after layer in receptive_field, the same computation as the widget. The idea itself comes from earlier work the paper cites, Sparse Transformers and Longformer.

Rolling buffer cache · original

“On a sequence length of 32k tokens, this reduces the cache memory usage by 8x, without impacting the model quality.”Jiang et al. (2023), §2

Everyday picture

A whiteboard with room for exactly W notes. When it is full, the next note is written over the oldest one. Since no layer ever looks more than W tokens back, the notes that get erased are exactly the ones nobody will read again.

Tiny example

With W = 4, token i goes into slot i mod 4, the remainder after dividing by 4. Tokens 0 to 3 fill slots 0 to 3. Token 4 (“of”) goes into slot 0, over token 0 (“It”); token 9 (“worst”) goes into slot 1, over token 5 (“times”). The board never grows.

Figure 2 of the paper (the rolling buffer cache), redrawn with new example text. Based on Jiang et al. (2023), Figure 2 (CC BY 4.0).

Reading it: press Next token to generate one more token. The four boxes are the whole cache. The slot just written is filled solid, as the latest token is coloured in the paper's figure, and the row of words below shows the text so far, with the tokens still in the cache underlined. From token 4 on, each new token lands in slot i mod 4 and erases the token exactly four positions earlier, which has just left the window. The cache never holds more than four tokens, however long the text grows.

In words: “token i's keys and values go into slot i mod W, so the cache holds the last W tokens once there are more than W of them, and all of them before that.”

With the numbers: W = 4: tokens 4 and 9 land in slots 0 and 1. For Mistral 7B at n = 32,768 tokens, the cache holds 4,096 of them instead of 32,768: 8 times less, the paper's figure. At 128 KiB per token (the GQA example above) that is 0.54 GB instead of 4.3 GB per sequence.

In Python:

W = 4
# slot(i) = i mod W
[i % W for i in range(12)]  # → [0, 1, 2, 3, 0, 1, 2, 3, 0, 1, 2, 3]
# Mistral 7B at 32k tokens
W, n = 4096, 32768
n / min(n, W)  # → 8.0
per_token = 131072
# gigabytes cached, with and without the buffer
round(min(n, W) * per_token / 1e9, 2), round(n * per_token / 1e9, 2)  # → (0.54, 4.29)

Why it matters

This is what makes the window pay off in memory, not only in compute: the cache has a fixed size, so a server knows in advance how much memory a conversation can ever need. The lesson's sliding_window_decode generates token by token with exactly this buffer and checks that its outputs match the masked attention, and cache_bytes_by_method plots the flat line it produces against every other design. One detail on the page of the paper: its text says Figure 2 shows W = 3, while the figure and its caption use W = 4.

Pre-fill and chunking · original

Everyday picture

The model writes its answer one token at a time, but the prompt is known in full before it starts, so its keys and values can be computed in advance: the prefill. A very long prompt would need a very large block of attention scores all at once, so the paper cuts it into pieces and pre-fills them one after another, like reading a long letter a page at a time while keeping the last page in mind.

Tiny example

The 12-token sentence is split into three chunks of 4 (the paper chooses the chunk size equal to the window, here 4): “It was the best”, “of times it was”, “the worst of times”. When the third chunk is processed, its tokens attend to themselves with an ordinary causal mask, to the second chunk (now sitting in the cache) through the sliding window, and not at all to the first chunk, which is out of reach.

Chunk being pre-filled:

Figure 3 of the paper (pre-fill and chunking), redrawn with new example text. Based on Jiang et al. (2023), Figure 3 (CC BY 4.0).

Reading it: each row is one token of the chunk being pre-filled, each column one token of the whole prompt, and a filled cell (marked 1) is a score that gets computed. The dashed lines split the columns into the past (chunks already dropped from the cache), the cache (the previous chunk) and the current chunk. For chunk 3 the current block is a lower triangle (causal), the cache block is an upper triangle (the window reaches into the previous chunk less and less as you go down), and the past block is empty. Together each row has exactly 4 filled cells: the same window as always. Chunk 1 has no cache to look at, and chunk 2 has no past.

In words: “each of the C tokens in a chunk is scored against at most the W cached tokens plus the C tokens of the chunk.” (This bound is this page's accounting; the paper draws it rather than writing it.)

With the numbers: in the tiny example, C = W = 4: at most 4 × 8 = 32 cells, of which the figure fills 16 (4 per row). For a 16,384-token prompt with C = W = 4,096, the largest block of scores is 4,096 × 8,192 = 33.5 million per head, against 16,384² = 268 million for the whole prompt at once: 8 times smaller.

In Python:

def block(C, W):
    # C queries against W cached keys plus the C keys of the chunk
    return C * (W + C)
block(4, 4)  # → 32
n, C, W = 16384, 4096, 4096
block(C, W), n * n  # → (33554432, 268435456)
n * n // block(C, W)  # → 8

Why it matters

Chunking caps the size of every attention computation during prefill, and with the rolling buffer it caps the cache too, so a prompt of any length fits in fixed memory. With the chunk size equal to the window, the rolling buffer holds exactly the previous chunk when the next one arrives, which is all the window can reach. The inference lesson explains why prefill and decode behave so differently in the first place.

3 Results · original

The benchmarks · original

Everyday picture

To compare two students fairly, you give them the same exams, marked the same way. The authors re-ran every model, their own and Llama's, through their own evaluation pipeline, rather than copying numbers from other papers. The exams are grouped into categories:

  • Commonsense reasoning (0-shot): HellaSwag, Winogrande, PIQA, SIQA, OpenbookQA, ARC-Easy, ARC-Challenge, CommonsenseQA.
  • World knowledge (5-shot): NaturalQuestions, TriviaQA.
  • Reading comprehension (0-shot): BoolQ, QuAC.
  • Maths: GSM8K (8-shot, maj@8) and MATH (4-shot, maj@4).
  • Code: HumanEval (0-shot) and MBPP (3-shot). The HumanEval companion explains how code is scored.
  • Aggregated: MMLU (5-shot), BBH (3-shot), and AGI Eval (3 to 5 shots, English multiple-choice only).

“5-shot” means five worked examples are placed in the prompt before each question; “maj@8” means eight answers are sampled and the most common one counts.

Table 2, reproduced with attribution (Jiang et al., 2023, CC BY 4.0): Mistral 7B against Llama models
ModelMMLUHellaSwagWinoGPIQAArc-eArc-cNQTriviaQAHumanEvalMBPPMATHGSM8K
Llama 2 7B44.477.169.577.968.743.224.763.811.626.13.916.0
Llama 2 13B55.680.772.980.875.248.829.069.618.935.46.034.3
Code Llama 7B (fine-tuned)36.962.962.372.859.434.511.034.931.152.55.220.8
Mistral 7B60.181.375.383.080.055.528.869.930.547.513.152.2

All numbers are accuracies in percent. Pick a benchmark to compare the four models:

Reading it: each bar is one model's accuracy on the chosen benchmark, on an axis from 0 to 100%, with its name on the left and its score on the right; Mistral 7B's bar is solid and the others are striped. On most benchmarks Mistral's bar is the longest, and on MMLU it leads Llama 2 13B by 4.5 points. Choose HumanEval or MBPP and Code Llama 7B, a model fine-tuned for code, comes out ahead, with Mistral close behind; Code Llama trails well behind on everything that is not code. Choose NQ (NaturalQuestions, a knowledge test) and Mistral's 28.8 sits just below Llama 2 13B's 29.0, the one column where the caption's “outperforms Llama 2 13B on all metrics” does not hold in its own table.

Why it matters

The paper lists two places where its evaluation differs from Llama 2's own: MBPP uses the hand-verified subset, and TriviaQA is asked without Wikipedia passages. Scores from different pipelines are not directly comparable, which is why re-running every model matters. The benchmarks lesson covers what such numbers can and cannot tell you.

Size and efficiency · original

“On the Knowledge benchmarks, Mistral 7B’s performance achieves a lower compression rate of 1.9x, which is likely due to its limited parameter count that restricts the amount of knowledge it can store.”Jiang et al. (2023), §3

Everyday picture

Line up the Llama 2 family by size (7B, 13B, 70B) and draw the curve of their scores. Put Mistral 7B's score on that curve and read off where a Llama 2 model would have to sit to score the same. That imaginary Llama's size is Mistral's equivalent model size.

Tiny example

The paper's Figure 5 does this for four categories. On MMLU, Mistral's 60.1 matches a Llama 2 of about 23B, 3.3 times its size. On reasoning it matches about 38B (5.4 times), on comprehension about 21B (3 times), and on knowledge only 13B (1.9 times).

In words: “how many times bigger a Llama 2 model would have to be to match Mistral 7B.”

With the numbers: 23 / 7 = 3.3; 38 / 7 = 5.4; 21 / 7 = 3.0; 13 / 7 = 1.9, the ratios printed on the paper's figure.

In Python:

N_mistral = 7
equivalent = {"MMLU": 23, "reasoning": 38, "comprehension": 21, "knowledge": 13}
{k: round(v / N_mistral, 1) for k, v in equivalent.items()}  # → {'MMLU': 3.3, 'reasoning': 5.4, 'comprehension': 3.0, 'knowledge': 1.9}
Equivalent Llama 2 sizes, read from the labels of Figure 5 (Jiang et al., 2023, CC BY 4.0)
CategoryEquivalent Llama 2 sizeCompression
Reasoning38B5.4×
MMLU23B3.3×
Comprehension21B3×
Knowledge13B1.9×

Why it matters

The pattern is the interesting part. Skills (reasoning, reading, maths) compress well into a small model; facts do not, because every fact has to be stored somewhere in the weights. This is also why a small model paired with retrieval, which looks facts up instead of memorising them, is a common design (the RAG lesson builds one).

4 Instruction finetuning · original

Everyday picture

A base model continues text; it has not learned to answer. Instruction tuning trains it further on examples of instructions paired with good answers. The authors did this with public instruction datasets from Hugging Face, and say they used no proprietary data or special tricks, to show how easily the base model adapts.

Tiny example

Table 3, reproduced with attribution (Jiang et al., 2023, CC BY 4.0): chat models
ModelChatbot Arena EloMT-Bench
WizardLM 13B v1.210477.2
Mistral 7B Instruct10316.84 ± 0.07
Llama 2 13B Chat10126.65
Vicuna 13B10416.57
Llama 2 7B Chat9856.27
Vicuna 7B9976.17
Alpaca 13B9144.53

MT-Bench is graded by a model acting as judge (the LLM-as-a-judge companion covers it); the Chatbot Arena rating comes from people voting between two anonymous answers, turned into an Elo rating. Mistral 7B Instruct has the best MT-Bench score of the 7B models and sits between the 13B chat models. In a separate human comparison on llmboxing.com, as of 6 October 2023, people preferred Mistral 7B Instruct's answer 5,020 times and Llama 2 13B Chat's 4,143 times.

In words: “the chance that A's answer is preferred depends only on the gap between the two ratings; every 400 points multiplies the odds by ten.”

With the numbers: Mistral 7B Instruct (1031) against Llama 2 13B Chat (1012): a 19-point gap predicts that Mistral wins 52.7% of comparisons. In the separate llmboxing comparison it won 5,020 of 9,163 votes, 54.8%: the same direction, from different questions and different voters.

In Python:

R_A, R_B = 1031, 1012
P_A = 1 / (1 + 10 ** ((R_B - R_A) / 400))
round(P_A, 3)  # → 0.527
# the human comparison on llmboxing.com
round(5020 / (5020 + 4143), 3)  # → 0.548

Why it matters

A strong base model plus a simple, public fine-tune gave a competitive chat model. Much of what a chat model can do is already in the base model; instruction tuning mostly teaches it the format. The training stages lesson walks through pretraining, instruction tuning and preference tuning in order.

5 Adding guardrails for front-facing applications · original

5.1 A system prompt to enforce guardrails · original

Everyday picture

Before a new employee meets customers, a manager gives them a few sentences of standing instructions. A system prompt is that briefing for a model, and a guardrail is any rule that keeps its output in bounds. The paper's recommended briefing is four sentences:

“Always assist with care, respect, and truth. Respond with utmost utility yet securely. Avoid harmful, unethical, prejudiced, or negative content. Ensure replies promote fairness and positivity.”Jiang et al. (2023), §5.1

Tiny example

Table 4, reproduced with attribution (Jiang et al., 2023, CC BY 4.0): MT-Bench of Mistral 7B Instruct under different system prompts, mean of 10 runs
System promptMT-Bench
none6.84 ± 0.07
Llama 2's system prompt6.38 ± 0.07
Mistral's system prompt6.58 ± 0.05

With its system prompt, the model declined all of a set of 175 unsafe prompts, at a cost of 0.26 MT-Bench points (about 4%) in helpfulness; Llama 2's system prompt, given to the same model, cost 0.46. The paper's example of the difference: asked how to kill a Linux process, Mistral 7B Instruct with its prompt explains the kill command, while Llama 2 13B Chat with its own prompt refuses. With no system prompt, both models answer correctly.

Why it matters

Guardrails move a model along a trade-off between usefulness and caution; refusing a harmless question about operating systems is the cost of leaning too far. The guardrails lesson builds input and output checks around a model rather than relying on the prompt alone.

5.2 Content moderation with self-reflection · original

Everyday picture

The same model can be asked to act as a moderator: given a prompt, or an answer, say whether it is acceptable or falls into one of three categories (illegal activities, hateful or violent content, unqualified legal, medical or financial advice). The paper calls this self-reflection.

Tiny example

On the authors' balanced set of adversarial and ordinary prompts, counting acceptable prompts as positives, it reached a precision of 99.4% and a recall of 95.6%. The paper does not give counts, so take 1,000 acceptable prompts as an illustration: recall 95.6% means 956 of them are let through (44 wrongly flagged), and precision 99.4% means that among everything let through, about 6 were not acceptable.

In words: “precision is the share of prompts let through that deserved it; recall is the share of acceptable prompts that were let through.”

With the numbers (illustrative counts): TP = 956, FN = 44, FP = 6: precision = 956 / 962 = 99.4%, recall = 956 / 1,000 = 95.6%.

In Python:

# illustrative counts matching the paper's rates
TP, FN, FP = 956, 44, 6
round(100 * TP / (TP + FP), 1)  # → 99.4
round(100 * TP / (TP + FN), 1)  # → 95.6

Why it matters

Because acceptable prompts are the positives, high precision means very little harmful content slips through as “acceptable”, and the 95.6% recall means about one acceptable prompt in 23 is wrongly flagged. The paper notes that users can choose which categories to filter for their own application. The metrics lesson explains precision and recall and the trade-off between them.

6 Conclusion · original

“…the problem is rather 3 dimensional (model capabilities, training cost, inference cost)…”Jiang et al. (2023), §6

Everyday picture

Choosing a car by price and speed alone ignores what it costs to fuel. Earlier work on scaling laws (the paper cites Hoffmann et al., 2022, the Chinchilla study) asked how to get the best model for a fixed training budget. The authors add a third axis: what the model costs every time it is used.

Tiny example

A model is trained once but may then answer millions of questions, so its inference bill keeps growing after its training bill has stopped. A 7B model that matches a 13B one costs roughly half as much per answer, for as long as it is used. The paper concludes that language models may compress knowledge more than was previously thought, and that much remains to be explored in getting the best performance from the smallest model.

Why it matters

Inference cost is now a first-class design goal, and every mechanism in §2 is an answer to it. The efficient architectures lesson compares them all on one chart.

What happened next

IdeaWhere it wentBuild it
Grouped-query attentionA common default for open models; the primer's running example, a Llama-3-8B-shaped model, uses 8 key/value heads for 32 query heads, exactly as Mistral 7B doesMultiHeadAttention
Sliding windows with a rolling bufferLater model families interleave sliding-window layers with a few full-attention layers, keeping exact long-range lookups where they are neededsliding_window_decode
Windows plus a few kept tokensAttention sinks (Xiao et al., 2023) keep the first few tokens alongside the window so generation can stream without limitglobal_local_mask
Bounded memory per conversationServing systems such as vLLM pack many such caches into GPU memoryPagedAttention companion

Glossary

Every term with hover guidance on this page, in one place.