LoRA and QLoRA, annotated
How to read this page
Nothing on this page assumes you already know the jargon.
- Any dotted word explains itself when you hover it, tab to it, or tap it.
- Every symbol inside an equation does the same. Hover the B in the first equation and see.
- The diagrams and sliders are live. Drag the rank slider in §4.1 and watch the parameter count collapse.
Each idea climbs the same ladder: an everyday picture, a tiny example you can check by hand, a diagram, the math, and why it still matters. The training stages lesson builds LoRA from scratch in NumPy.
LoRA: Abstract · original
“We propose Low-Rank Adaptation, or LoRA, which freezes the pre-trained model weights and injects trainable rank decomposition matrices into each layer of the Transformer architecture, greatly reducing the number of trainable parameters for downstream tasks.”Hu et al. (2021), Abstract
Everyday picture
You own a 1,000-page textbook and want it to suit a new course. Reprinting the whole book for every course is absurd. Instead you keep the book exactly as it is and add a thin booklet of corrections for each course. Swap the booklet, and the same book serves a different class. LoRA is that booklet for a neural network: the huge weights stay frozen, and a tiny set of new numbers carries each task's changes.
What the paper claims
- Compared with full fine-tuning of GPT-3 175B with Adam, LoRA cuts the number of trainable parameters by 10,000× and the GPU memory needed by 3×.
- Quality is on par with or better than full fine-tuning on RoBERTa, DeBERTa, GPT-2 and GPT-3.
- No extra delay when the model runs, unlike adapter layers, because the booklet can be folded back into the book.
Why it matters today
LoRA is the default way to customize open-weight language models. When someone ships a “LoRA” for a model, they are shipping a few megabytes that turn a general model into a specialist.
1 Introduction · original
“We hypothesize that the change in weights during model adaptation also has a low ‘intrinsic rank’, leading to our proposed Low-Rank Adaptation (LoRA) approach.”Hu et al. (2021), §1
Everyday picture
Adapting a model that already knows English to write SQL, or to summarize conversations, probably doesn't require rewriting everything it knows. It needs a few focused adjustments. The authors bet that the change to each weight matrix can be described with very few independent “directions” of change. That count of directions is what rank means.
Tiny example: a big table from two short lists
Take a column of 4 numbers, B = (1, 0, 2, 0), and a row of 4 numbers, A = (0.5, 1, 0, −1). Multiply every entry of B by every entry of A and you get a full 4 × 4 table, 16 numbers from only 8:
| 0.5 | 1 | 0 | −1 | |
|---|---|---|---|---|
| 1 | 0.5 | 1 | 0 | −1 |
| 0 | 0 | 0 | 0 | 0 |
| 2 | 1 | 2 | 0 | −2 |
| 0 | 0 | 0 | 0 | 0 |
Every row is a multiple of (0.5, 1, 0, −1), so the table has rank 1. At the scale of a real model, a 4,096 × 4,096 weight matrix has 16.8 million entries, but a rank-1 change to it needs only 4,096 + 4,096 = 8,192 numbers. That ratio, 2,048 to 1, is the whole trick.
Four advantages the paper lists
- Many tasks, one model: keep one frozen copy and swap small A and B pairs per task.
- Cheaper training: no gradients or optimizer state for the frozen weights, which lowers the hardware needed by up to 3×.
- No inference latency: the update can be added into the weights before serving.
- Composable: it combines with other methods such as prefix tuning.
Why it matters
The claim is not only “cheaper”. It is that the useful part of fine-tuning really is low-dimensional, which §7 tests directly.
2 Problem statement · original
Everyday picture
Fine-tuning means showing the model many examples of a task, such as a question paired with its SQL, and nudging its weights so that the right answer becomes more likely, one token at a time.
In words: “adjust all the weights so that, across every training pair, the model gives as much probability as possible to each correct next token of the target, given the prompt and the target tokens so far.”
With the numbers: if the target SQL is 3 tokens long and the model gives them probabilities 0.5, 0.8 and 0.9, this pair contributes ln 0.5 + ln 0.8 + ln 0.9 = −0.69 − 0.22 − 0.11 = −1.02. Training pushes that number up towards 0.
In Python:
import math
# P_Φ(y_t | x, y_<t) for t = 1, 2, 3
P = [0.5, 0.8, 0.9]
[round(math.log(p), 2) for p in P] # → [-0.69, -0.22, -0.11]
# Σ over t of log P
round(sum(math.log(p) for p in P), 2) # → -1.02
The problem: full fine-tuning learns a change ΔΦ as large as the model itself. For GPT-3 that is about 175 billion numbers per task. LoRA instead describes the change with a much smaller set of numbers Θ; for GPT-3, Θ can be as small as 0.01% of the model.
Why it matters
Storage and switching cost scale with the size of what you train. Shrink that, and serving hundreds of custom models from one base becomes practical. See the training stages lesson for where fine-tuning sits among pretraining, SFT and preference tuning.
3 Why not adapters or prompts? · original
Everyday picture
Two earlier ways to avoid full fine-tuning each had a catch. Adapters insert small extra layers into the network: like adding extra stops on a train line, every journey now takes longer, and at a batch size of 1 the stops dominate. Prefix tuning learns special tokens that sit in the prompt: like reserving seats on the train for luggage, it takes up space the passengers (your actual input) could have used, and the paper found it hard to optimize.
The latency cost, measured
| Method | Latency | Extra |
|---|---|---|
| Fine-tuned model or LoRA (merged) | 19.8 | none |
| Adapter (Lin et al. design) | 23.9 | +20.7% |
| Adapter (Houlsby et al. design) | 25.8 | +30.3% |
With big batches the overhead shrinks (about 2 to 3% at batch 32), because the GPU is busy anyway. Online serving usually has small batches, where adapters hurt most.
Why it matters
The goal was an adaptation method with no serving penalty and no loss of usable context. That design constraint is what led to a change that can be merged into the existing weights.
4 The method · original
4.1 Low-rank update matrices · original
“During training, W₀ is frozen and does not receive gradient updates, while A and B contain trainable parameters.”Hu et al. (2021), §4.1
Everyday picture
A weight matrix is like a big mixing desk: it takes an input signal x and blends it into an output h. LoRA leaves the desk untouched and runs a thin side channel in parallel: squeeze x down to a handful of numbers (A), expand them back up (B), and add the result to the desk's output.
Hover or tap a part. Start at the bottom with x.
Reading it: the input x goes two ways. On the left, it passes through the original weight matrix W₀, which is frozen: training never touches it. On the right, it passes through the LoRA side path: A squeezes the d numbers of x down to just r numbers (the narrow waist), and B expands them back to d. The two results meet at the ⊕ and are added. Because B starts at zero, the side path contributes nothing on day one, so training begins from exactly the original model's behaviour. Only the two trapezoids (A and B) learn.
The math
In words: “the layer's output is what the frozen weights would have produced, plus a correction computed by squeezing the input through A and expanding it through B, scaled by α over r.”
With the numbers: with d = 4, r = 1, B = (1, 0, 2, 0) and A = (0.5, 1, 0, −1) from §1, and input x = (2, 1, 0, 1): A x = 0.5·2 + 1·1 + 0·0 − 1·1 = 1, a single number. B times that is (1, 0, 2, 0). So the correction adds 1 to the first output and 2 to the third, on top of W₀x (taking α = r, so the scale is 1).
In Python:
# d × r column, with r = 1
B = [1, 0, 2, 0]
# r × d row
A = [0.5, 1, 0, -1]
x = [2, 1, 0, 1]
# α = r, so the scale α/r is 1
alpha, r = 1, 1
# squeeze d numbers down to r = 1
Ax = sum(a_j * x_j for a_j, x_j in zip(A, x))
Ax # → 1.0
# expand: the correction (α/r)·BAx
[alpha / r * b_i * Ax for b_i in B] # → [1.0, 0.0, 2.0, 0.0]
Try it: how many numbers does LoRA train?
The heatmap is a 24 × 24 corner of the update BA at the chosen rank (random A and B, fixed seed). Hover it.
Reading it: the two bars compare how many numbers you would train to change one d × d weight matrix: all of them (full fine-tuning) or just A and B (LoRA, 2 × d × r). The bars use a log scale (a bar's length follows how many digits its number has), so the real gap is far bigger than it looks. At d = 4,096 and r = 8 LoRA trains 65,536 numbers instead of 16.8 million: 256× fewer. Now watch the heatmap: at r = 1 every row is a scaled copy of every other, so you see clean stripes. Raise r and the pattern gets richer, because more independent directions are mixed in. Low rank means “structured”, not “small values”.
Two details that make it work
- Initialization: A starts random (Gaussian) and B starts at zero, so ΔW = BA = 0 at the start.
- Scaling: the update is multiplied by α/r. The authors set α to the first r they try and never tune it, so changing r doesn't force re-tuning the learning rate.
No extra latency, by construction
Before serving, compute W = W₀ + BA once and store it. The network then has exactly the same shape and speed as the original. To switch tasks, subtract BA and add a different B′A′.
In words: “fold the correction into the weights once, and the model runs exactly as fast as before.”
With the numbers: the rank-1 table from §1 (16 entries) is added entry by entry into a 4 × 4 W₀; the merged matrix is still 4 × 4. Taking W₀ to be the 4 × 4 identity (ones on the diagonal, zeros elsewhere) to keep the numbers small, the merged W sends x = (2, 1, 0, 1) to (3, 1, 2, 1): exactly W₀x = (2, 1, 0, 1) plus the correction (1, 0, 2, 0) from above.
In Python:
B = [1, 0, 2, 0]
A = [0.5, 1, 0, -1]
# the identity
W0 = [[1 if i == j else 0 for j in range(4)] for i in range(4)]
# W = W₀ + BA
W = [[W0[i][j] + B[i] * A[j] for j in range(4)] for i in range(4)]
# still 4 × 4
len(W), len(W[0]) # → (4, 4)
x = [2, 1, 0, 1]
# W x
[sum(W[i][j] * x[j] for j in range(4)) for i in range(4)] # → [3.0, 1.0, 2.0, 1.0]
Why it matters today
“Merge the adapter” is a standard deployment step. The flip side, which the paper notes, is that if you merge, you can't mix different tasks' adapters in one batch; serving systems that keep adapters unmerged exist for exactly that. Build it yourself in the training stages lesson.
4.2 Applying LoRA to a Transformer · original
Everyday picture
A Transformer layer is full of weight matrices: four in attention (for queries, keys, values and the output) and two in the feed-forward network. The paper puts LoRA side paths only on the attention matrices, mostly Wq and Wv, and freezes everything else.
Tiny example: GPT-3's budget
GPT-3 has 96 layers of width 12,288. LoRA with r = 4 on Wq and Wv trains 96 layers × 2 matrices × (12,288 + 12,288) × 4 = 18.9 million numbers. In 16-bit that is about 38 MB, which the paper rounds to 35 MB, against 350 GB for the model.
| Quantity | Full fine-tuning | LoRA |
|---|---|---|
| GPU memory (VRAM) during training | 1.2 TB | 350 GB |
| Checkpoint size per task | 350 GB | 35 MB |
| 100 tasks stored | ≈ 35 TB | ≈ 354 GB |
| Training throughput per V100 | 32.5 tokens/s | 43.1 tokens/s |
The memory saving comes from not keeping Adam's extra per-weight state for frozen weights. The 25% speed-up comes from not computing gradients for them.
Why it matters today
This table is the business case: one base model in memory, hundreds of customers' adapters on disk, each loaded in milliseconds.
5 Experiments · original
Everyday picture
The test is simple: does the booklet teach as well as reprinting the book? The paper compares LoRA with full fine-tuning and other cheap methods on language understanding (the GLUE benchmark), text generation (E2E NLG) and three GPT-3 tasks: turning questions into SQL (WikiSQL), judging whether one sentence follows from another (MultiNLI) and summarizing chats (SAMSum).
| Model and method | Trainable params | Score |
|---|---|---|
| RoBERTa-base, full fine-tune | 125M | GLUE average 86.4 |
| RoBERTa-base, LoRA | 0.3M | GLUE average 87.2 |
| DeBERTa-XXL, full fine-tune | 1,500M | GLUE average 91.1 |
| DeBERTa-XXL, LoRA | 4.7M | GLUE average 91.3 |
| GPT-2 medium, full fine-tune | 354.9M | E2E BLEU 68.2 |
| GPT-2 medium, LoRA | 0.35M | E2E BLEU 70.4 |
| GPT-3 175B, full fine-tune | 175,255.8M | WikiSQL 73.8 · MNLI 89.5 |
| GPT-3 175B, LoRA | 4.7M | WikiSQL 73.4 · MNLI 91.7 |
| GPT-3 175B, LoRA | 37.7M | WikiSQL 74.0 · MNLI 91.6 |
Reading it: each row pair compares full fine-tuning with LoRA on the same model. The bars show trainable parameters on a log scale (a bar's length follows how many digits its number has), so even a modest gap in length means hundreds of times fewer parameters. LoRA trains between a hundred and tens of thousands of times fewer numbers and lands within a point of full fine-tuning, often above it. Scores come from different benchmarks, so compare rows within a pair, not across pairs.
One more finding: not every cheap method improves with more trainable parameters. Prefix methods got worse past a few dozen to a few hundred special tokens, while LoRA stayed stable.
Why it matters
“As good as full fine-tuning with a fraction of the parameters” is the result that made LoRA the default. The metrics lesson explains BLEU and ROUGE, two of the scores used here.
7 Understanding the low-rank updates · original
Having shown that LoRA works, the paper asks why. All three studies use GPT-3 175B.
7.1 Which matrices to adapt? · original
Everyday picture
With a fixed budget of about 18 million numbers, is it better to make deep changes to one kind of matrix, or shallow changes to several?
| Adapted matrices | Rank r | WikiSQL | MultiNLI |
|---|---|---|---|
| Wq only | 8 | 70.4 | 91.0 |
| Wk only | 8 | 70.0 | 90.8 |
| Wq and Wv | 4 | 73.7 | 91.3 |
| Wq, Wk, Wv, Wo | 2 | 73.7 | 91.7 |
Finding: spreading the budget wins. Even rank 2 on all four attention matrices beats rank 8 on one. A small rank captures enough per matrix; coverage matters more.
7.2 How small can r be? · original
“To our surprise, a rank as small as one suffices for adapting both Wq and Wv on these datasets while training Wq alone needs a larger r.”Hu et al. (2021), Table 6 caption
Tiny example
On WikiSQL, adapting Wq and Wv scored 73.4 at r = 1, 73.7 at r = 4, 73.8 at r = 8 and 73.5 at r = 64. Going from 1 to 64 directions of change bought essentially nothing.
How they checked it: comparing subspaces
They trained with r = 8 and r = 64, took the most important directions each learned (via singular value decomposition), and measured how much the two sets overlap.
In words: “take the top i directions from the rank-8 run and the top j from the rank-64 run, and score how much of the smaller set lies inside the larger one: 1 means completely, 0 means not at all.”
With the numbers: for the single top direction (i = j = 1), the two runs overlap with similarity above 0.5, while the other directions barely overlap. The important change lives in about one direction; the rest looks like noise. To see the formula itself at work, take i = j = 1 and one direction of length 1 from each run, u = (0.8, 0.6) and v = (1, 0): φ = (0.8 × 1 + 0.6 × 0)² / min(1, 1) = 0.64, above 0.5.
In Python:
# top direction of the r = 8 run (length 1)
u = [0.8, 0.6]
# top direction of the r = 64 run (length 1)
v = [1, 0]
i = j = 1
# Uᵀ U is a 1 × 1 table here
overlap = sum(u_k * v_k for u_k, v_k in zip(u, v))
# ‖·‖_F² / min(i, j)
round(overlap ** 2 / min(i, j), 2) # → 0.64
Why it matters
This is evidence for the paper's opening hypothesis: the task-specific change really is low rank. The authors caution that a very different task, such as a new language, would likely need a larger r.
7.3 How ΔW relates to W · original
Everyday picture
Does fine-tuning add brand-new knowledge, or turn up the volume on something the model already had? The paper projects W onto the directions of ΔW and compares sizes, using the Frobenius norm.
With the numbers: in layer 48, ‖Wq‖ = 61.95, and Wq projected onto ΔW's top 4 directions has size only 0.32, while ΔW itself has size 6.91. So along those directions the update is about 6.91 / 0.32 ≈ 21.6 times bigger than what W already had.
Finding: ΔW amplifies features that are present in W but not emphasized, rather than repeating W's strongest directions. Fine-tuning turns up quiet features that matter for the task.
8 Conclusion · original
The authors conclude that LoRA keeps quality while removing inference latency and cutting the input-length cost of prompt methods, and that it lets you switch tasks quickly on a shared model. Their open questions: combining LoRA with other methods, understanding the mechanism more deeply, choosing which matrices to adapt on a principled basis, and whether W itself is rank-deficient.
QLoRA (Dettmers et al., 2023) · original
“We present QLoRA, an efficient finetuning approach that reduces memory usage enough to finetune a 65B parameter model on a single 48GB GPU while preserving full 16-bit finetuning task performance.”Dettmers et al. (2023), Abstract
Everyday picture
LoRA shrank what you train, but you still had to hold the whole frozen model in GPU memory at 16 bits per number. QLoRA stores the frozen book in shorthand, 4 bits per number, and writes the corrections booklet in full precision. When a page is needed, it is expanded from shorthand on the fly.
Try it: how much GPU memory does fine-tuning need?
A rough rule of thumb for this page, not the paper's accounting: full 16-bit fine-tuning with Adam holds about 12 bytes per parameter (2 for the weight, 2 for its gradient, 8 for Adam's two running averages); LoRA in 16-bit holds about 2 bytes per frozen weight; QLoRA about 0.52 bytes (4 bits plus the scaling constants). Activations and the small adapters are left out.
Reading it: each bar is the memory for the model's weights and training state, on a log scale; the dashed marker is a single 48 GB GPU. At 65B parameters the rule of thumb gives 780 GB for full fine-tuning, matching the paper's “more than 780 GB”, and about 34 GB for QLoRA's frozen weights, which leaves room on one 48 GB card for the adapters and activations. Slide down to 7B and even 16-bit LoRA fits on one GPU.
QLoRA adds three tricks, described next: a 4-bit data type shaped for neural network weights, quantizing the quantization constants, and paging optimizer memory to the CPU.
4-bit NormalFloat (NF4) · original
Everyday picture
With 4 bits you get only 16 possible values. Suppose you must describe everyone's height with just 16 labels. Spacing the labels evenly from shortest to tallest wastes most of them on rare extremes, while crowds of average-height people share a few labels. Better: space the labels so each covers the same number of people. Neural network weights follow a bell curve, so NF4 puts its 16 levels where the weights actually are.
Tiny example
Take a block of weights (0.02, −0.10, 0.05, 0.31). Divide by the largest magnitude, 0.31, so they fit in [−1, 1]: (0.065, −0.323, 0.161, 1.0). Round each to the nearest NF4 level: 0.080, −0.284, 0.161, 1.0. Store the four level numbers (4 bits each) plus the one scale, 0.31. To use a weight, multiply back: 0.080 × 0.31 = 0.025, close to the original 0.02.
In words: “cut the bell curve into equal-probability slices and put each level halfway between two neighbouring cut points.”
With the numbers: with k = 4 there are 2⁴ = 16 levels and 2⁴ + 1 = 17 slices. Straight from the formula, q8 sits halfway between the cut points QX(8/17) = −0.074 and QX(9/17) = 0.074, at 0, and its neighbour q9 = 0.148 is close by; out in the tail, q14 = 1.058 and q15 = 1.376 are 0.318 apart. After a small fix so that zero is represented exactly, and rescaling into [−1, 1], the paper's NF4 levels run −1, −0.696, −0.525, −0.395, −0.284, −0.185, −0.091, 0, 0.080, 0.161, 0.246, 0.338, 0.441, 0.563, 0.723, 1: dense near zero, sparse at the tails.
In Python:
from statistics import NormalDist
# Q_X: the bell curve's quantile function
Q_X = NormalDist().inv_cdf
k = 4
# halfway between two neighbouring cut points
def q(i):
return (Q_X(i / (2**k + 1)) + Q_X((i + 1) / (2**k + 1))) / 2
round(Q_X(8 / 17), 3), round(Q_X(9 / 17), 3) # → (-0.074, 0.074)
# near zero: levels packed tight
round(q(8), 3), round(q(9), 3) # → (0.0, 0.148)
# in the tail: levels spread out
round(q(14), 3), round(q(15), 3) # → (1.058, 1.376)
Reading it: the grey bars are 25,600 weights drawn from a bell curve (computed live in your browser with a fixed seed), each block of 64 divided by its largest magnitude, which is how QLoRA normalizes them. The vertical lines are the 16 levels every weight must be rounded to. Toggle between NF4 and evenly spaced levels and read the error: NF4 packs levels where the bars are tall, so the average rounding error is lower and the 16 levels are used more evenly. With so few levels, where you put them matters.
| 4-bit data type | Mean perplexity |
|---|---|
| Int4 | 34.34 |
| Float4 (E2M1) | 31.07 |
| Float4 (E3M0) | 29.48 |
| NFloat4 + double quantization | 27.41 |
Perplexity measures how surprised a model is by real text. Same 4 bits, better placement, a less surprised model.
Double quantization · original
Everyday picture
Every block of 64 weights needs its own scale factor (the 0.31 in the tiny example), stored as a 32-bit number. That overhead adds up, so QLoRA compresses the scale factors too: a note about the notes.
In words: “one 32-bit scale per 64 weights costs half a bit per weight; storing those scales in 8 bits, with one 32-bit scale per 256 of them, cuts it to about an eighth of a bit.”
With the numbers: the saving is 0.5 − 0.127 = 0.373 bits per parameter. For a 65B model that is 65 × 10⁹ × 0.373 / 8 ≈ 3 GB, which the paper also reports.
In Python:
# one 32-bit scale per 64 weights
before = 32 / 64
# 8-bit scales, plus one 32-bit scale per 256 of them
after = 8 / 64 + 32 / (64 * 256)
before, round(after, 3) # → (0.5, 0.127)
# bits saved per parameter
round(before - after, 3) # → 0.373
# bytes saved on a 65B model, in GB
round(65e9 * 0.373 / 8 / 1e9, 1) # → 3.0
Paged optimizers · original
Everyday picture
A long training example can briefly spike memory use and crash the run. Paged optimizers keep the optimizer's state in memory that the GPU driver can move to ordinary CPU RAM when the GPU runs short, and page it back when needed, the way an operating system swaps memory to disk.
Why it matters
It turns rare out-of-memory crashes into a small slowdown, which is what makes “fine-tune a 65B model on one GPU” reliable rather than lucky.
QLoRA in one equation · original
In words: “expand the 4-bit frozen weights back to 16-bit on the fly, multiply the input by them, and add the LoRA correction computed in 16-bit.”
With the numbers: it is §4.1's equation with W₀ replaced by its decompressed 4-bit copy. In the tiny example, the stored level 0.080 and scale 0.31 are expanded to 0.025 just before the multiply. Gradients flow through that decompressed copy but are only kept for L₁ and L₂.
In Python:
# a W^NF4 entry and its block's scale
level, scale = 0.080, 0.31
# doubleDequant: back to a usable 16-bit weight
w = level * scale
round(w, 3) # → 0.025
Results · original
What the paper shows
- No quality loss from 4 bits: with NF4 and double quantization, QLoRA matched 16-bit full fine-tuning and 16-bit LoRA on GLUE and Super-NaturalInstructions, and matched 16-bit LoRA on 5-shot MMLU for LLaMA 7B to 65B (mean 53.1 versus 53.0).
- Adapter coverage beats rank: putting LoRA on every linear layer of each block was needed to match full fine-tuning; the rank r mattered little.
- 65B on one GPU: memory for fine-tuning a 65B model drops from more than 780 GB to under 48 GB.
- Guanaco: their best model reached 99.3% of ChatGPT's score on the Vicuna benchmark after 24 hours of fine-tuning on a single GPU.
- Data quality over quantity: a 9,000-example dataset (OASST1) beat a 450,000-example one (a subsample of FLAN v2) for chatbot quality.
| Data type | 7B | 13B | 33B | 65B |
|---|---|---|---|---|
| BFloat16 (16-bit) | 38.4 | 47.2 | 57.7 | 61.8 |
| Float4 | 37.2 | 47.3 | 55.9 | 61.3 |
| NFloat4 + double quantization | 39.0 | 47.5 | 57.3 | 61.8 |
The authors also warn that chatbot benchmarks judged by another model are not fully trustworthy, and they publish examples of where their model fails.
Why it matters today
QLoRA made fine-tuning large open models a single-GPU job, and “bigger model at lower precision beats smaller model at higher precision” became a common rule of thumb. The inference lesson builds int8 and int4 quantization from scratch.
What changed since
| In the papers | Common today | Lesson |
|---|---|---|
| LoRA on Wq and Wv only | LoRA on all linear layers (QLoRA's finding) | training stages |
| One adapter per deployed model | Servers that batch many unmerged adapters over one base model | inference |
| 16-bit or 4-bit base weights | 4-bit and 8-bit bases are routine for fine-tuning and serving | inference |
| Fine-tune to add behaviour | Still the rule: fine-tuning shapes behaviour, retrieval supplies knowledge | RAG |
Glossary
Every term with hover guidance on this page, in one place.