An annotated companion · AI Primer

LIMA: Less Is More for Alignment, annotated

About this page. This is a companion, not a copy. It follows the paper section by section, quotes only a sentence or two per section (clearly marked), and explains everything in its own words. The paper is released on arXiv under a Creative Commons Attribution 4.0 licence; the few numbers and table rows shown here are attributed to it. Every chart is redrawn from scratch from numbers printed in the paper, and where a value had to be read off a plot instead, the page says so. Read the original alongside: every section links to it.

How to read this page

  • Any dotted word explains itself on hover, focus or tap; so does every symbol in every equation.
  • The diagrams are live: hover, tab to or tap a block to read about it. The preference chart in §4.2 switches between human and GPT-4 judges, and the filter in §2.1 lets you turn the paper's rules on and off.

Each idea climbs the same ladder: an everyday picture, a tiny example you could check by hand, a diagram, the math, and why it matters today. The fine-tuning lesson builds the data-preparation steps this paper relies on (chat formatting, deduplication, a held-out set, label audits) from scratch, and the training stages lesson builds the supervised loss LIMA is trained with.

Abstract · original

“Taken together, these results strongly suggest that almost all knowledge in large language models is learned during pretraining, and only limited instruction tuning data is necessary to teach models to produce high quality output.”Zhou et al. (2023), Abstract

Everyday picture

A new hire arrives with a degree, years of reading and a head full of facts. On day one they do not need to be taught chemistry again; they need to see how this office writes a reply: greet the customer, answer the question, list the steps, sign off. A thousand good examples of house style are enough, because the knowledge was already there. LIMA asks whether a base model is that new hire.

What the paper claims

  • A 65-billion-parameter LLaMa model, fine-tuned with the ordinary supervised loss on only 1,000 curated prompts and responses, with no reinforcement learning from human feedback and no preference model, produces strong assistant-style answers.
  • In a blind comparison, people judged LIMA's answer equal to or better than GPT-4's in 43% of prompts, Bard's in 58%, and DaVinci003's (a model trained with human feedback) in 65%.
  • The interpretation: almost everything a model knows comes from pretraining; alignment mostly teaches a format.

Why it matters today

LIMA changed how people think about the size of an instruction dataset. Before it, the reflex was “more examples”; after it, the question became “which examples, and how good are they?”. That is the stance the fine-tuning lesson takes in its sections on label quality and on how many examples to collect.

1 Introduction · original

“We hypothesize that alignment can be a simple process where the model learns the style or format for interacting with users, to expose the knowledge and capabilities that were already acquired during pretraining.”Zhou et al. (2023), §1

Everyday picture

Think of two ways to open a restaurant with a brilliant chef. One way puts the chef through months of training on millions of customer orders and then a season of diners scoring every plate. The other hands the chef one thick binder: a thousand menus and plated dishes showing exactly how this restaurant serves food. If the chef already knows how to cook, the binder might be almost as good.

Tiny example: how much data?

LIMA's training set is 1,000 examples, about 750,000 tokens in total, so roughly 750 tokens per example. The Alpaca baseline it is compared with was trained on 52,000 examples: 52 times more. The methods it is contrasted with, instruction tuning on multi-million-example datasets followed by RLHF over millions of interactions with annotators, use more by further orders of magnitude.

Pretraining on raw text Base model (LLaMa 65B for LIMA) usual recipe LIMA Instruction tuningmillions of examples RLHFmillions of interactions Supervised loss on1,000 examples no reward model, no RL Assistant LIMA

Hover or tap a block. Start at the top with Pretraining, then compare the two columns.

The contrast the introduction draws, redrawn as a diagram. Based on Zhou et al. (2023), §1.

Reading it: both recipes start from the same place: a model pretrained on an enormous amount of raw text to predict the next token. The left column is the usual path to a chat assistant: a large instruction-tuning stage, then reinforcement learning from human feedback. The right column is LIMA: one short supervised fine-tune on 1,000 examples and nothing else. If the right column gets close to the left, most of what makes an assistant capable must already be in the top two boxes, and the paper's whole argument rests on that comparison.

Why it matters today

This framing, pretraining supplies the knowledge and supervised fine-tuning supplies the manners, is now the standard way to explain why a small fine-tune can transform a model. The training stages lesson walks the full pipeline, and the InstructGPT companion explains the RLHF column LIMA leaves out.

2 Alignment data · original

“A model’s knowledge and capabilities are learnt almost entirely during pretraining, while alignment teaches it which subdistribution of formats should be used when interacting with users.”Zhou et al. (2023), §2, the Superficial Alignment Hypothesis

Everyday picture

A pretrained model has read forum threads, textbooks, recipes, arguments and jokes, and it can continue any of them. Asked a question, it might answer, or ask three more questions, or write a sarcastic reply, because all of those appear after questions on the web. The paper's Superficial Alignment Hypothesis says alignment is choosing one of those voices (the helpful assistant) and making it the default. If that is all it is, a small set of examples all written in that voice should do.

Tiny example: the recipe for the set

The authors want diverse inputs and uniform outputs: prompts from every corner of life, answered in one consistent assistant style. The training sources add up to exactly 1,000: 200 + 200 + 200 + 150 + 50 + 200. Separately, they keep 50 development prompts for choosing checkpoints and 300 test prompts (70 + 230) for the evaluation.

Training: 1,000 examples Stack Exchange (STEM)200 Stack Exchange (Other)200 wikiHow200 Reddit r/WritingPrompts150 Super-Natural Instructions50 Paper authors, group A200 Development: 50 prompts Paper authors, group A50 Test: 300 prompts r/AskReddit 70authors, group B 230

Hover or tap a bar to see where its examples came from and how long they are.

Table 1 of the paper, redrawn as bars. Counts from Zhou et al. (2023), Table 1 (CC BY 4.0).

Reading it: each bar's length is the number of examples. The six blue bars are the whole training set; four of them come from websites where good answers were already written, and one (group A) the authors wrote themselves. The development and test prompts come from different people than the training prompts: group B wrote the test prompts, and the Reddit test prompts have no reference answers at all. Hover the bars to compare lengths: a wikiHow prompt averages 12 tokens and its answer 1,811, while a Super-Natural Instructions prompt averages 236 tokens and its answer 92. The set is tiny, but it is deliberately mixed.

Why it matters today

“Diverse prompts, consistent answers” is still the working rule for a small instruction set. The fine-tuning lesson turns the same idea into checks you can run: chat_example builds one training example in the chat format, and deduplicate removes near-copies so the 1,000 slots are not wasted on repeats.

2.1 Community questions and answers · original

Everyday picture

Mining a forum for training data is like panning for gold: most of what you scoop up is gravel. Stack Exchange and wikiHow are rich streams, moderated and mostly helpful, so simple rules can sift them automatically. Reddit rewards jokes and sarcasm, so the authors picked those examples by hand.

Tiny example: the Stack Exchange sieve

For each question the authors took the top answer, only if it scored at least 10, and then dropped it if it was shorter than 1,200 characters or longer than 4,096, written in the first person (“I”, “my”), or pointing at other answers (“as mentioned”, “stack exchange”). Links, images and other HTML were stripped; code blocks and lists stayed. Try the sieve on seven invented candidates.

The seven candidates are illustrative, invented to exercise each rule; the rules themselves are the paper's.

Reading it: each card is one top answer and its verdict. With every rule on, two of seven survive. Switch a rule off and watch what floods back in: without the length rule, a one-liner and a sprawling reference both return; without the first-person rule, an anecdote about someone's job returns. The score rule keeps answers the community trusted; the other three remove answers that are fine in a forum but read wrongly from an assistant. That is the paper's point about style: the filter is choosing a voice, not checking facts.

The math: sampling across 179 communities

Stack Exchange has 179 communities. The authors split them into 75 STEM and 99 other (dropping 5 niche ones), then drew 200 questions from each group “using a temperature of τ = 3 to get a more uniform sample of the different domains”. The paper names the temperature but not the formula. The usual form, and our reading of it, raises each community's size to the power 1/τ before turning sizes into shares, which lifts small communities and trims big ones.

Tiny example (illustrative sizes): two communities with 900 and 100 questions. Sampling in proportion to size gives shares 0.9 and 0.1. With τ = 3, the sizes become 9001/3 = 9.655 and 1001/3 = 4.642, so the shares become 0.675 and 0.325: the small community's share more than triples.

In words: “shrink every community's size by taking its τ-th root, then divide by the total so the shares add up to one; the bigger τ, the more even the shares.”

With the numbers: psmall = 4.642 / (9.655 + 4.642) = 0.325, against 0.1 without the temperature.

In Python:

n = [900, 100]
tau = 3
# n_i^(1/τ): each community's size, softened
softened = [n_i ** (1 / tau) for n_i in n]
[round(x, 3) for x in softened]  # → [9.655, 4.642]
# p_i = n_i^(1/τ) / Σ_j n_j^(1/τ)
[round(x / sum(softened), 3) for x in softened]  # → [0.675, 0.325]
# τ = 1 is plain proportional sampling
[x / sum(n) for x in n]  # → [0.9, 0.1]

wikiHow and Reddit

From wikiHow, 200 articles: first a category (out of 19), then an article within it, the title as the prompt and the body as the answer, with “This article...” rewritten as “The following answer...”. From Reddit, 150 hand-picked prompts and stories from r/WritingPrompts went into training, and 70 self-contained r/AskReddit questions went into the test set, because top Reddit answers are not reliably helpful.

Why it matters today

Rule-based filters like these, plus a sampling scheme that stops one big source from dominating, are how most instruction sets are still assembled. A quality filter during pretraining plays the same role at a far larger scale.

2.2 Manually authored examples · original

Everyday picture

Forum questions cover what people ask forums. To reach the rest of what people might ask an assistant, the authors wrote prompts themselves, inspired by their own interests and their friends', and then wrote the answers in one house voice.

Tiny example: two groups, kept apart

Two groups of authors each wrote 250 prompts. Group A's became 200 training examples (with answers the authors wrote) and 50 development prompts; group B's, after filtering, became the 230 test prompts. A footnote admits the separation was imperfect: the groups talked before writing, so they share some habits. Of the 200 group-A training examples, 13 are prompts with some toxicity or malice, answered with a partial or full refusal and an explanation. Another 50 examples come from Super-Natural Instructions, one random example from each of 50 generation tasks (summarising, paraphrasing, style transfer), lightly edited into the same style.

The house style usually acknowledges the question first and then answers it. The authors found this consistent format helped in preliminary experiments and suggest it works a little like chain of thought: the opening sentence gives the model a running start.

Why it matters today

Keeping the authors of the test prompts separate from the authors of the training data is the same discipline as the fine-tuning lesson's held-out set, built there by split_before_training. The footnote is an honest reminder of how easily it leaks. The 13 refusal examples are a first look at the question §4.3 returns to: how much safety behaviour can a handful of examples teach?

3 Training LIMA · original

“We find that perplexity does not correlate with generation quality, and thus manually select checkpoints between the 5th and the 10th epochs using the held-out 50-example development set.”Zhou et al. (2023), §3

Everyday picture

The training itself is deliberately ordinary: read the 1,000 examples fifteen times, adjusting the model a little after every batch, and take smaller steps as you go. Two details stand out. A new special token marks the end of each speaker's turn, like a “your turn” card passed across a table. And the checkpoint to keep is chosen by reading its answers, not by a number.

Tiny example: the recipe in numbers

  • Start: LLaMa 65B. Loss: the standard supervised loss, the one the sft_loss function computes.
  • End-of-turn token: a new token (EOT) after each utterance. It stops generation the way an end-of-sequence token does, but avoids whatever other meanings the pretrained model had attached to that existing token.
  • Optimizer: AdamW with β1 = 0.9, β2 = 0.95, weight decay 0.1; 15 epochs; batch size 32; texts longer than 2,048 tokens trimmed.
  • Learning rate: no warm-up; starts at 10−5 and falls in a straight line to 10−6 by the end.
  • Steps: 1,000 examples in batches of 32 is 31.25 batches per epoch, so 15 epochs is roughly 470 steps (our arithmetic; the paper does not print a step count).

In words: “start at the initial learning rate and slide in a straight line to the final one; at a fraction s/S of the way through training, you are that fraction of the way down.”

With the numbers: halfway through, η = 10−5 + (10−6 − 10−5) × 0.5 = 5.5 × 10−6; at the end it is 10−6.

In Python:

eta_0, eta_S = 1e-5, 1e-6
S = 470
# η(s) = η_0 + (η_S − η_0) · s / S
def eta(s):
    return eta_0 + (eta_S - eta_0) * s / S
round(eta(0), 8), round(eta(S / 2), 8), round(eta(S), 8)  # → (1e-05, 5.5e-06, 1e-06)
# batches per epoch, and steps over 15 epochs
1000 / 32, 15 * 1000 / 32  # → (31.25, 468.75)

Hover or tap the chart to read the learning rate at each step.

Reading it: the x-axis is the training step, out of roughly 470; the y-axis is the learning rate, the size of each update. It starts at its highest, 10−5, with no gentle warm-up, and falls in a straight line to a tenth of that. Checkpoints were saved along the way, and the one kept came from between epochs 5 and 10, roughly steps 155 to 315: the middle third of this line. The final model is not the last one.

Residual dropout, raised layer by layer

The one unusual regulariser is dropout on the residual connections (following InstructGPT): no dropout at the bottom layer, rising in a straight line to 0.3 at the top layer (0.2 for smaller models). With only 1,000 examples read 15 times, the model could memorise them; switching off random parts of the top layers makes that harder.

Tiny example (a made-up 5-layer model; this paper does not state LLaMa 65B's layer count): the dropout rates are 0, 0.075, 0.15, 0.225 and 0.3.

In words: “the bottom layer gets no dropout, the top layer gets the maximum, and every layer between is spaced evenly along the way.”

With the numbers: layer 3 of 5 gets 0.3 × (3 − 1)/(5 − 1) = 0.15.

In Python:

p_max, L = 0.3, 5
# p_d(ℓ) = p_max · (ℓ − 1) / (L − 1), for ℓ = 1 … L
[round(p_max * (ell - 1) / (L - 1), 3) for ell in range(1, L + 1)]  # → [0.0, 0.075, 0.15, 0.225, 0.3]

Why it matters today

Two habits from this section stuck. A dedicated turn marker is part of every modern chat template; the fine-tuning lesson's render_chat ends every turn with one. And choosing checkpoints by reading generations on a small development set, rather than trusting perplexity, is common practice for chat models; Appendix B below shows why.

4 Human evaluation · original

Everyday picture

A blind taste test. A judge gets one question and two answers from unnamed models, and says which is better or that neither is clearly better. Repeat for 300 questions and five rival models, and count.

Why it matters

For open-ended answers there is no answer key, so preference between two answers is the most direct measurement available. It is also the same kind of judgement RLHF's reward models are trained on, which makes the comparison fair to the RLHF-trained rivals.

4.1 Experiment setup · original

The rivals

  • Alpaca 65B: the same LLaMa 65B fine-tuned on Alpaca's 52,000 examples. Same starting model, 52 times the data: the cleanest comparison in the paper.
  • DaVinci003: an OpenAI model tuned with RLHF (the method of the InstructGPT companion).
  • Bard (based on PaLM), Claude (52B parameters, trained with reinforcement learning from AI feedback, the method of the Constitutional AI companion), and GPT-4. All three answered in April 2023.

Generating answers

One answer per prompt per model, sampled with nucleus sampling at p = 0.9 and temperature 0.7, a repetition penalty of 1.2 on previously generated tokens, and at most 2,048 tokens. The inference lesson builds these knobs: temperature_probs and top_p_filter.

Do judges agree? Tie-discounted agreement

Before trusting the judges, the authors measured how often two judges agree on the same 50 comparisons, scoring one point when two judges give the same verdict, half a point when exactly one of them says “tie”, and zero when they pick opposite answers. This is a form of inter-annotator agreement.

Tiny example: four comparisons. Both judges pick A (1 point); one picks A, the other says tie (½); one picks A, the other B (0); both say tie (1). Agreement = 2.5 / 4 = 0.625.

In words: “score each shared comparison 1 for the same verdict, ½ if only one judge called it a tie, 0 for opposite verdicts, and average.”

With the numbers: (1 + ½ + 0 + 1) / 4 = 0.625. The paper measured 82% between crowd workers, 81% crowd and authors, 78% between authors, and 78% and 79% between GPT-4 and crowd and authors respectively: GPT-4 agreed with people about as well as people agreed with each other.

In Python:

pairs = [("A", "A"), ("A", "tie"), ("A", "B"), ("tie", "tie")]
# a_i: 1 if the verdicts match, ½ if exactly one is a tie, else 0
def a(first, second):
    if first == second:
        return 1
    if "tie" in (first, second):
        return 0.5
    return 0
scores = [a(x, y) for x, y in pairs]
scores  # → [1, 0.5, 0, 1]
# agreement = (1/n) Σ a_i
sum(scores) / len(scores)  # → 0.625

Why it matters today

Measuring agreement before trusting a judge is the step people most often skip. The fine-tuning lesson's label_agreement does the plain version, and the LLM-as-a-judge companion studies when a model can stand in for human judges, which LIMA's GPT-4 numbers foreshadow.

4.2 Results · original

“Perhaps ironically, even GPT-4 prefers LIMA outputs over its own 19% of the time.”Zhou et al. (2023), §4.2

Tiny example: reading one bar

Against GPT-4, human judges said LIMA won 18% of prompts, tied 25% and lost 57%. “Equal or better” is wins plus ties: 18 + 25 = 43%, the number in the abstract. Against Alpaca 65B, the same count is 53 + 21 = 74%.

In words: “LIMA is at least as good whenever it wins or ties, and it loses the rest.”

With the numbers: against GPT-4, g = 0.18 + 0.25 = 0.43 and l = 0.57; against Bard, g = 0.33 + 0.25 = 0.58.

In Python:

# (wins, ties) for LIMA, human judges, Figure 1
human = {"Alpaca 65B": (0.53, 0.21), "DaVinci003": (0.44, 0.21), "Bard": (0.33, 0.25),
         "Claude": (0.24, 0.22), "GPT-4": (0.18, 0.25)}
# g = w + t for each rival
{rival: round(w + t, 2) for rival, (w, t) in human.items()}  # → {'Alpaca 65B': 0.74, 'DaVinci003': 0.65, 'Bard': 0.58, 'Claude': 0.46, 'GPT-4': 0.43}
Judged by:
LIMA winstieLIMA loses (striped)

Figures 1 and 2 of the paper, redrawn. Percentages from Zhou et al. (2023), Figures 1 and 2 (CC BY 4.0), over 300 test prompts.

Reading it: each row is one rival; each bar is split into LIMA's wins (solid), ties (grey) and losses (striped), adding to 100%. Rows follow the paper's order, and the striped part grows as you go down. Look first at Alpaca 65B: the same base model trained on 52 times more data, and LIMA is at least as good 74% of the time. Then DaVinci003, trained with RLHF: LIMA wins more often than it loses (44% against 35%). Against Claude and GPT-4, LIMA mostly loses, but not always. Switch to GPT-4 as the judge: the ranking of rivals is the same and the picture barely changes, which is why the authors treat the GPT-4 judgements as corroboration.

Why it matters today

The headline is not that LIMA beats GPT-4; it doesn't. It is that 1,000 examples and no human-feedback stage got surprisingly far, and in particular beat the same base model trained on 52,000 automatically generated examples. Quality of examples, not their count, separated LIMA from Alpaca.

4.3 Analysis · original

Everyday picture

A head-to-head against the best products in the world is a high bar: those products may have seen millions of real user prompts. So the authors also graded LIMA on its own, like marking an exam against the question rather than against the top student.

Tiny example: 50 answers, three grades

On 50 random test prompts: Fail (did not meet the prompt's requirements), Pass (met them) or Excellent. 25 were excellent (50%), 19 passed (38%) and 6 failed (12%). Of those 50, 43 had a somewhat related training example (advice, letter writing, and so on); adding 13 more prompts with no related training example gives 20 out-of-distribution prompts, of which 20% failed, 35% passed and 45% were excellent: 4, 7 and 9 prompts.

excellentpassfail (striped)

Reading it: the top bar is Figure 3 of the paper: 50 test prompts, half excellent. The bottom bar is the out-of-distribution sample, prompts unlike anything in training. The two bars look alike: about half excellent and a small failing slice in each. The sample is small (20 prompts), but it is the paper's evidence that LIMA did not just learn to copy its 1,000 examples.

In Python:

# Figure 3: 50 prompts
excellent, passed, failed = 25, 19, 6
[x / 50 for x in (excellent, passed, failed)]  # → [0.5, 0.38, 0.12]
# out of distribution: 20% fail, 35% pass, 45% excellent, of 20 prompts
[round(share * 20) for share in (0.20, 0.35, 0.45)]  # → [4, 7, 9]
# safety: 80% of 30 sensitive prompts answered safely
round(0.8 * 30)  # → 24

Safety, from 13 examples

On 30 potentially sensitive test prompts, LIMA responded safely to 80% (24), including 6 of the 10 with malicious intent. It refuses outright when a request is plainly harmful (a celebrity's home address), but is more likely to comply when the harmful intent is implied rather than stated: the paper's Figure 4 shows it helpfully advising on drugging a neighbour's dog.

Why it matters today

Style generalised from a handful of examples; safety only partly did. That split is one reason most assistant recipes still add a preference-tuning stage after supervised fine-tuning, as in the DPO companion, and why dedicated safety evaluation and guardrails exist.

5 Why is less more? Diversity, quality and quantity · original

“We observe that, for the purpose of alignment, scaling up input diversity and output quality have measurable positive effects, while scaling up quantity alone might not.”Zhou et al. (2023), §5

Everyday picture

Three ways to improve a phrasebook: add more kinds of situation (diversity), fix the clumsy translations (quality), or print more copies of the same phrases (quantity). Only the first two teach you anything new.

Tiny example: how each score is made

These ablations use a smaller 7B LLaMa, trained on 2,000 examples per set (a footnote says the 7B model can be tuned on 1,000, but at least 2,000 made training more stable). For every test prompt, the model writes 5 answers, and ChatGPT (GPT-3.5 Turbo) grades each for helpfulness on a six-point Likert scale, from 1 (not helpful) to 6 (highly helpful: clear, detailed, well organised with headings or lists). The model's score is the average grade. Illustrative grades for one prompt: 4, 3, 5, 4, 3, averaging 3.8.

In words: “add up every grade the judge gave and divide by how many there are.”

With the numbers: (4 + 3 + 5 + 4 + 3) / 5 = 3.8 for the illustrative prompt. The real comparison below: filtered Stack Exchange scores 3.83 and unfiltered 3.33, a gap of 0.50, which the paper calls significant (it reports each average with a 95% confidence interval).

In Python:

s = [4, 3, 5, 4, 3]
# s̄ = (1/m) Σ s_k
sum(s) / len(s)  # → 3.8
# Figure 5: filtered minus unfiltered, and filtered minus wikiHow
round(3.83 - 3.33, 2), round(3.83 - 3.49, 2)  # → (0.5, 0.34)

Figure 5: generation quality (1 to 6), 2,000 examples each, axis from 3.0 to 4.0

Reading it: three 7B models, each trained on 2,000 examples, scored by the average ChatGPT grade (the bars start at 3.0 to make differences visible). Diversity: wikiHow (high-quality answers, but every prompt is “how to...”) scores 3.49; filtered Stack Exchange (high-quality answers to varied questions) scores 3.83. Quality: the same Stack Exchange data without the §2.1 filters drops to 3.33, the lowest of the three, even though its prompts are just as diverse. The filters are worth half a point.

Hover or tap to read the score at each training-set size.

Reading it: Figure 6 of the paper, redrawn. The x-axis is the number of filtered Stack Exchange training examples on a log scale, doubling from 2,000 to 32,000; the y-axis is the average grade, drawn from 3.2 to 4.0 as in the paper. The line is flat: sixteen times the data buys nothing measurable. These values are read off the paper's plot (it prints no numbers for this figure), so treat them as accurate to about ±0.01; every point sits near 3.83, and the plotted error bars overlap.

Why it matters today

This is the result behind “quality beats quantity” in fine-tuning. It sits beside the scaling laws of pretraining without contradicting them: for pretraining, more tokens teach more; for this kind of alignment, more copies of the same style teach nothing new. The fine-tuning lesson's advice to double the data only while the held-out score still rises by more than its margin of error (margin_of_error) is the practical form of Figure 6.

6 Multi-turn dialogue · original

“This leap in capability from a mere 30 examples, as well as the fact that the zero-shot model can converse at all, reinforces the hypothesis that such capabilities are learned during pretraining, and can be invoked through limited supervision.”Zhou et al. (2023), §6

Everyday picture

Every one of LIMA's 1,000 examples is a single question and a single answer. Asking it to hold a conversation is like asking someone who has only ever answered letters to chat on the phone. It can, a bit, and then it loses the thread. Thirty sample conversations fix most of that.

Tiny example: counting failures

Ten live conversations each. The 1,000-example model failed 15 of its 42 turns (35.7%); in 6 of 10 conversations it stopped following the prompt within 3 exchanges. The authors then added 30 multi-turn conversations (10 written by the authors, 20 edited from Stack Exchange comment chains) and fine-tuned a fresh model from LLaMa on the 1,030 examples. It failed 1 turn of 46 (2.2%), and the share of excellent turns rose from 45.2% to 76.1%. Judged as whole conversations, it was significantly better in 7 of 10 and tied in 3.

In words: “the failure rate is the number of failed turns divided by all turns.”

With the numbers: 15 / 42 = 0.357 before, 1 / 46 = 0.022 after: sixteen times fewer failures from 30 extra examples.

In Python:

# r = n_fail / n_turns, before and after the 30 dialogue examples
r_before, r_after = 15 / 42, 1 / 46
round(r_before, 3), round(r_after, 3)  # → (0.357, 0.022)
round(r_before / r_after, 1)  # → 16.4
excellentpassfail (striped)

Reading it: Figure 7 of the paper, redrawn: each bar is every turn from 10 conversations, split into excellent, pass and fail. The top bar is LIMA as trained on single-turn examples only (zero-shot at dialogue); the bottom is the version with 30 dialogues added. The striped failure slice shrinks from over a third to a sliver, and the excellent share grows by more than half. Thirty examples are 3% of the training set.

Why it matters today

A behaviour the model can almost do already needs only a few demonstrations, but it does need them: a gap in coverage shows up as a gap in behaviour. Appendix E repeats the finding for structured answers with just six examples.

7 Discussion · original

“Primarily, the mental effort in constructing such examples is significant and difficult to scale up.”Zhou et al. (2023), §7

Everyday picture

A hand-written style guide is a great way to train one excellent new hire, and a slow way to train a thousand of them for a thousand different jobs.

The two limitations the authors name

  • Curation does not scale easily. 1,000 examples is small to train on but large to write and check by hand.
  • Robustness. LIMA is not as dependable as product-grade models: an unlucky sample during decoding, or an adversarial prompt, can produce a weak answer.

Why it matters today

Both limitations point where the field went next: generating and filtering examples with models to scale curation (synthetic data), and preference tuning to make behaviour dependable across unlucky samples and adversarial prompts. The training stages lesson shows where each stage fits.

Appendix B: perplexity rises while quality rises · original

Everyday picture

A trainee copy-editor gets steadily worse at predicting, word for word, what other writers wrote, because they are learning to write in the house voice instead. Their own writing improves while their “guess the next word of someone else's text” score gets worse. LIMA's validation perplexity behaves exactly like that.

Tiny example: what perplexity measures

Perplexity scores how surprised a model is by some text. If it gave the three actual next tokens probabilities 0.5, 0.25 and 0.125, their average negative log-probability is (0.693 + 1.386 + 2.079) / 3 = 1.386, and perplexity is e1.386 = 4: on average, the model was as unsure as if it were choosing among 4 equally likely tokens.

In words: “average how surprised the model is by each real next token, measured as minus the log of the probability it gave that token, and exponentiate so the answer reads as a number of choices.”

With the numbers: exp(1.386) = 4.0.

In Python:

import math
p = [0.5, 0.25, 0.125]
N = len(p)
# −(1/N) Σ log p(t_i | t_<i)
mean_nll = -sum(math.log(p_i) for p_i in p) / N
round(mean_nll, 3)  # → 1.386
# PPL = exp(mean negative log-probability)
round(math.exp(mean_nll), 3)  # → 4.0

Hover or tap to read validation perplexity at each step.

Hover or tap to read generation quality at each step.

Reading it: Figure 9 of the paper, split into two charts that share the same x-axis: training steps of LIMA 65B, from 60 to 420. The top chart is perplexity on 2,000 held-out Stack Exchange examples; the bottom is the average ChatGPT grade of generations (the §5 method). Both climb, apart from a small dip in quality at the last point. In most training, rising validation perplexity is the classic sign of overfitting and a reason to stop; here the answers keep getting better while it rises. The values are read off the paper's plot, so they are approximate, and the grades carry wide error bars in the original. The paper reports the correlation and does not explain it; a plausible reading is that the model grows more committed to its own house style, and therefore less likely to predict other people's exact wording.

Why it matters today

Perplexity on reference answers is a poor guide to choosing a chat model's checkpoint. That is why LIMA picks checkpoints by reading answers on a 50-prompt development set, and it is the exception to the fine-tuning lesson's rule of stopping at the best validation loss, shown by overfitting_run: that rule holds when the validation loss measures the thing you care about.

Appendix E: six examples of structure · original

Everyday picture

If no one has ever shown you a form with named sections, you might answer “fill in Goals, Audience and Budget” with a nice essay that ignores the headings. Once you have seen a few forms, you fill in any form.

Tiny example

Early versions of LIMA handled most development prompts but not those that dictated a structure, such as “summarise this into bullet points” or “write a plan with these five sections”. Six training examples with formatting constraints (a product page with Highlights, About the Product and How to Use sections; question-and-answer pairs from an article) were enough. The paper's Figure 13 contrasts 994 examples with 1,000: without the six, a request for a coffee-shop marketing plan with five named sections gets a generic plan that ignores them; with them, the answer follows all five, although no marketing plan appears anywhere in the training data.

Why it matters today

Six examples, 0.6% of the data, switched a whole class of behaviour on. When a fine-tuned model ignores a format, the first thing to check is whether any training example shows that kind of format at all.

What changed since 2023

In the paperCommon todayWhyWhere to learn more
Supervised fine-tuning only, no preference stageA supervised stage followed by preference tuning such as DPOLIMA itself reports fragile safety and occasional weak answers; preference tuning targets exactly thoseDPO companion, training stages lesson
1,000 examples written or curated by handInstruction sets often drafted or filtered by models, then checkedHand curation does not scale, the limitation §7 names; LIMA's lesson carries over as “filter hard”fine-tuning lesson
GPT-4 as a second judge, agreeing with humans about as well as humans agreeModel judges used routinely for quick comparisons, calibrated against human labelsHuman judging is slow and costly; agreement must still be measuredLLM-as-a-judge companion, evals lesson
A special end-of-turn tokenEvery chat model has a template of turn markersThe model must know exactly where a turn ends, and existing tokens carry old meaningsfine-tuning lesson, render_chat
Full fine-tuning of all 65B weightsSmall instruction sets are often trained as adapters insteadUpdating a few low-rank matrices is far cheaper and forgets less of the base modelLoRA companion

Glossary

Every term with hover guidance on this page, in one place.