Visual Instruction Tuning, annotated
How to read this page
Nothing here assumes you already know the jargon. Three things help:
- Any dotted word explains itself when you hover it, tab to it, or tap it.
- Every symbol inside an equation does the same, and each equation is followed by a table of its symbols, a sentence reading it aloud, the numbers of a tiny example, and the same example in a few lines of Python.
- The pictures are live: hover or tap the parts of the architecture and the data pipeline, switch the architecture between its two training stages, and step through a training conversation to see which tokens the model is graded on.
The code that builds each piece from scratch lives in the multimodal lesson, whose toy vision-language model is a miniature LLaVA: a vision encoder, a projector, and a language model that reads the projected image as extra tokens.
Abstract
“We present the first attempt to use language-only GPT-4 to generate multimodal language-image instruction-following data.”Liu et al. (2023), Abstract. Read the original
Everyday picture
A chat model is a brilliant conversationalist who has never seen anything. A CLIP image encoder is a sharp-eyed observer who cannot hold a conversation. LLaVA puts them in the same room with a small interpreter between them, then gives the pair a few weeks of practice conversations about photos. The twist is where the practice conversations come from: a second chat model that also cannot see wrote them, working only from written descriptions of each photo.
What the paper claims
- Data: a text-only GPT-4, shown captions and object boxes instead of pixels, can write 158,000 useful image conversations, descriptions and reasoning questions.
- Model: a frozen CLIP vision encoder, one trainable matrix, and the Vicuna chat model, tuned end to end on that data, becomes a multimodal assistant: LLaVA.
- Results: on a synthetic benchmark, LLaVA's answers score 85.1% of the score a text-only GPT-4 gets when it is handed the ground-truth description; fine-tuned on ScienceQA, LLaVA combined with GPT-4 as a judge reaches 92.53%, a new best at the time.
- Openness: the data, the model, the code and two new benchmarks are all released.
Why it matters today
“Encode the image, project it, splice it into the prompt, then tune on image conversations” became the default recipe for open vision-language models. It is the design the multimodal lesson builds and trains.
1 Introduction · original
“In this paper, we present visual instruction-tuning, the first attempt to extend instruction-tuning to the language-image multimodal space, to pave the way towards building a general-purpose visual assistant.”Liu et al. (2023), §1
Everyday picture
Before this paper, vision models were like appliances with one button each: a classifier names objects, a captioner writes a caption, a detector draws boxes. Language models had just shown a different kind of interface: one text box where you say what you want. The trick that made them follow requests is instruction tuning: fine-tuning on many examples of an instruction followed by a good response. The paper asks whether the same trick works when some of the instruction is a picture.
Tiny example
The same photo, three instructions, three different jobs: “What vehicle is this?” wants a short fact; “Describe this photo in detail” wants a paragraph; “What challenges do these people face?” wants reasoning. A one-button captioner gives the same caption to all three. An instruction-tuned model reads the request and switches task.
The four contributions
- Data: a pipeline that turns ordinary image-caption pairs into instruction-following conversations, using ChatGPT or GPT-4 as the writer (§3).
- Model: a large multimodal model made by connecting CLIP's open-set image encoder to the Vicuna language decoder, fine-tuned end to end (§4).
- Benchmark: LLaVA-Bench, two small, hard test sets of images, questions and detailed annotations (§5.1).
- Release: data, code, checkpoints and a chat demo.
Why it matters
Instruction tuning is the step that turns a base language model into an assistant (see the training stages lesson and the InstructGPT companion). This paper is where that step first reached images in a cheap, open, reproducible form.
2 Related work · original
“Finally, note that visual instruction tuning is different from visual prompt tuning: the former aims to improve the model’s instruction-following abilities, while the latter aims to improve the parameter-efficiency in model adaptation.”Liu et al. (2023), §2
Everyday picture
There were two ways to build a visual assistant. One is to train a single model end to end for one kind of task, such as following directions through a house or editing a photo on request. The other is a switchboard: a language model that calls separate vision tools (a captioner here, a detector there) and stitches their outputs together, as Visual ChatGPT and MM-REACT did. LLaVA is the first kind, but general: one model trained end to end for many tasks.
What sits nearby
- Flamingo joined a frozen image encoder to a frozen language model through new cross-attention layers and showed strong in-context learning; the paper calls it “the GPT-3 moment in the multimodal domain”.
- BLIP-2 compressed each image into a few tokens with a small learned module (the Q-Former) before handing it to the language model.
- OpenFlamingo and LLaMA-Adapter gave the open LLaMA model image inputs.
The gap the paper names: none of these were tuned on vision-language instruction data, and their multimodal performance lagged their text-only performance.
Why it matters
The three families in the multimodal lesson map onto this section: splice projected tokens into the prompt (LLaVA), add cross-attention layers (Flamingo), or compress to a fixed few tokens first (BLIP-2).
3 GPT-assisted visual instruction data · original
“This symbolic representation allows us to encode the image as an LLM-recognizable sequence.”Liu et al. (2023), §3
Everyday picture
You want practice conversations about thousands of photos, and you have a brilliant writer who is blindfolded. So you read the writer a description of each photo: five one-line captions, plus a list of every object and the rectangle it sits in. Show the writer a few hand-made example conversations, and ask for more like them. The writer never sees a pixel, yet the conversations it writes are about the photo.
Tiny example
The paper's own example (its Table 1) is a photo of people packing luggage into a black SUV in an underground garage. What the writer is given:
- Captions, such as “People try to fit all of their luggage in an SUV.”
- Boxes, one line per object:
person: [0.681, 0.242, 0.774, 0.694],backpack: [0.384, 0.696, 0.485, 0.914], and so on. The four numbers are the box's corners as fractions of the image's width and height, so the writer can tell left from right and near from far.
From that, it writes three kinds of response: a conversation (“What type of vehicle is featured in the image?” “The image features a black sport utility vehicle (SUV).”), a detailed description, and a complex reasoning answer (“What challenges do these people face?”, answered with the difficulty of fitting all that luggage in).
The cheapest possible version, used later for stage 1, needs no writer at all: pick a question such as “Describe the image concisely.” and use the image's caption as the answer. The paper writes it as : question and image in, caption out. Cheap, but every answer is one caption long.
Hover or tap a part. Start with the image at the top left and notice which wire goes around GPT-4.
Reading it: start at the top. A COCO photo comes with human-written annotations (the dashed arrow): five captions and a list of object boxes. Those two texts, plus a few hand-written example conversations, go into a text-only GPT-4, the long box in the middle. GPT-4 writes one of three kinds of response (the three green boxes, with how many of each were kept). The photo itself travels down the left edge and never enters GPT-4; it only rejoins its text at the bottom, where each response is paired with its image to make a training example for LLaVA. The seed examples are the only human annotation in the whole pipeline.
The numbers
158K samples in total: 58K conversations, 23K detailed descriptions and 77K reasoning answers. Reasoning is the largest share, almost half. The paper also tried ChatGPT as the writer and found GPT-4 “consistently provides higher quality” data, especially for spatial reasoning.
In Python:
# samples written by GPT-4, in thousands
conversation, detail, reasoning = 58, 23, 77
conversation + detail + reasoning # → 158
# each type's share of the total
[round(100 * k / 158, 1) for k in (conversation, detail, reasoning)] # → [36.7, 14.6, 48.7]
Why it matters
This is distillation through a text bottleneck: a stronger model's language skill is poured into training data for a weaker one, and the image is described rather than shown. Synthetic data written by a stronger model is now a standard ingredient of instruction tuning. Its limit is visible too: the writer can only talk about what the captions and boxes mention, so whatever the annotations miss, the conversations miss.
4 Visual instruction tuning · original
Everyday picture
Two parts: how the picture gets into the language model (4.1, one matrix), and how the pair is taught (4.2, two stages, graded only on the assistant's replies).
4.1 Architecture · original
“We consider a simple linear layer to connect image features into the word embedding space.”Liu et al. (2023), §4.1
Everyday picture
The vision encoder writes its notes in its own shorthand; the language model reads only its own vocabulary of word vectors. The projector is a plug adapter between them: it reshapes each of the encoder's vectors into something the language model treats like a word. After that the language model reads the picture the way it reads a sentence.
Tiny example
Say the encoder describes one patch of the image with the two numbers (1, 2), and the language model's word vectors have three numbers. A 3 × 2 matrix W with rows (1, 0), (0, 1) and (1, 1) turns (1, 2) into (1, 2, 3): a vector of the right width, ready to sit in the prompt. The multimodal lesson works this same example (it writes the matrix the other way round, as a row vector times a 2 × 3 matrix; the numbers are the same).
Hover or tap a part. Start with the image at the bottom left and follow it up. Then switch stages below and watch which parts learn.
Reading it: two roads lead up into one language model. On the left, the image goes through CLIP's vision encoder g, whose grid of patch vectors Zv the projection W turns into visual tokens Hv (the four squares on the left). On the right, the instruction's words are looked up in the language model's own embedding table to give word tokens Hq (the three squares on the right). The language model reads both rows as one sequence and writes the response at the top. The small labels say which parts learn. Switch to stage 2 and the language model starts learning too; the vision encoder never does.
The math: one matrix
In words: “run the image through the frozen vision encoder to get one feature vector per patch, then multiply each by one learned matrix so it has the width of the language model's word vectors.” The paper writes W on the left, so each feature is a column; the lesson's code keeps features as rows and multiplies on the right. Same arithmetic.
With the numbers: W · (1, 2) = (1·1 + 0·2, 0·1 + 1·2, 1·1 + 1·2) = (1, 2, 3). At full size: CLIP's ViT-L/14 reads a 224 × 224 image (CLIP paper, Table 20) in 14-pixel patches, a 16 × 16 grid, so Zv holds 256 vectors, each 1,024 wide. Vicuna-13B is a fine-tuned LLaMA-13B, whose word vectors are 5,120 wide (LLaMA paper, Table 2). So W is 5,120 × 1,024: 5,242,880 learned numbers, turning one image into 256 visual tokens. (The paper names the encoder and the language model but not these widths; they come from the two models' own papers.)
In Python:
# Z_v: one patch feature, as a column
z = [1, 2]
# W: 3 rows (the language model's width), 2 columns (the encoder's width)
W = [[1, 0], [0, 1], [1, 1]]
# W · z: dot each row of W with z
[sum(W[r][c] * z[c] for c in range(2)) for r in range(3)] # → [1, 2, 3]
# full size: a 224-pixel image in 14-pixel patches
grid = (224 // 14) ** 2
grid # → 256
# W is d_LM × d_vision
print(f"{5120 * 1024:,}") # → 5,242,880
# as a share of Vicuna's 13 billion weights, in percent
round(100 * 5120 * 1024 / 13.0e9, 2) # → 0.04
In code: Projector is this matrix (and, given a hidden width, the two-layer version later LLaVAs switched to); worked_example_projection computes the (1, 2, 3) above; VisionEncoder plays the part of g; and VisionLanguageModel.splice puts the projected tokens where the image placeholder sits in the prompt.
Which features, exactly?
The encoder is CLIP's (see the CLIP companion), a Vision Transformer (see the ViT companion for how an image becomes a grid of patch vectors). LLaVA takes the grid of patch features, not CLIP's single summary vector, and the paper tries the grid from the last layer and from the layer before it. For ScienceQA, the layer before the last wins (§5.2): 90.92% against 89.96%. The authors' guess is that CLIP's last layer is trained to summarise the whole image to match a caption, while the layer before still keeps local detail.
Why it matters today
The authors chose the plainest possible connector on purpose, because it “allows us to iterate data centric experiments quickly”, and name richer options (Flamingo's gated cross-attention, BLIP-2's Q-Former) as future work. It turned out that the plain connector was enough: most open vision-language models still splice projected patch tokens into the prompt.
4.2 Training · original
“This stage can be understood as training a compatible visual tokenizer for the frozen LLM.”Liu et al. (2023), §4.2
Everyday picture
Every training example is a chat transcript with a photo taped into the first message. The model reads the whole transcript but is only marked on the assistant's lines, the way a language teacher reads your whole essay but only corrects the part you wrote. The course comes in two terms: first the interpreter alone learns the vocabulary, then interpreter and writer practise real conversations together.
Building the transcript
Tiny example: a three-turn conversation about the SUV photo has questions Xq1, Xq2, Xq3 and answers Xa1, Xa2, Xa3. The image goes into the first turn only, and a coin flip decides whether it comes before or after the first question.
In words: “the user's message in the first turn is the first question together with the image, in a random order; in every later turn it is just that turn's question.”
With the numbers: with T = 3 turns and the coin landing “image first”, the three user messages are [image, “What vehicle is this?”], [“How many people are there?”], [“What are they doing?”]. The image is attached once, and every later answer can still see it, because it sits earlier in the same sequence.
In Python:
questions = ["What vehicle is this?", "How many people are there?", "What are they doing?"]
T = len(questions)
# X_instruct^t: the image joins the first question, in an order set by a coin flip
def instruct(t, image_first):
if t == 1:
return ["<image>", questions[0]] if image_first else [questions[0], "<image>"]
return [questions[t - 1]]
[instruct(t, image_first=True) for t in range(1, T + 1)] # → [['<image>', 'What vehicle is this?'], ['How many people are there?'], ['What are they doing?']]
instruct(1, image_first=False) # → ['What vehicle is this?', '<image>']
The random order is there so the model does not learn to expect the picture in one fixed place. The turns are then laid out one after another, each user message and each answer followed by a stop marker (the paper uses ###), after a fixed system message borrowed from Vicuna.
Try it: the strip below is the paper's Table 2 as one flat sequence. Switch between a stage-1 example (one question, the caption as answer) and a stage-2 conversation, and tap any piece to see what it is. Highlighted pieces are the ones the loss counts.
Reading it: every piece is part of one long input that the language model reads left to right. The shaded piece labelled “image” is the image: 256 visual tokens in reality, one chip here. Plain pieces are context the model reads but is never graded on: the system message, the “Human:” and “Assistant:” labels, the user's questions and the image. Highlighted pieces are the targets: each answer and the ### that ends it, because the model must learn where to stop as well as what to say. In the stage-1 example the only target is the original caption; in the stage-2 conversation, every answer is.
The objective: predict the answers, token by token
Tiny example: the answer “black SUV” followed by the stop marker is three tokens. Say the model gives the right first token probability 0.6, the right second token 0.5 (given the first), and the stop marker 0.9 (given both). The probability of the whole answer is the product: 0.6 × 0.5 × 0.9 = 0.27.
In words: “the probability of all the answers is the product, over the answer tokens, of the probability the model gives each correct token, given the image, the instructions so far, and the answer tokens before it.” Training makes this number as large as possible, which is the same as making its negative logarithm (the cross-entropy) as small as possible. The product runs over positions 1 to L, but only the highlighted target tokens of Table 2 contribute a factor; the paper also leaves the system message and earlier stop markers out of the conditions “for better readability”.
With the numbers: 0.6 × 0.5 × 0.9 = 0.27. As a loss: −log 0.27 = 1.309, or 0.436 per answer token. A perfect model would give each token probability 1, a product of 1 and a loss of 0.
In Python:
import math
# p_θ(x_i | ...) for the three target tokens: "black", "SUV", "###"
p = [0.6, 0.5, 0.9]
# Π over the target tokens
round(math.prod(p), 2) # → 0.27
# the loss: −log of that product, in total and per token
round(-math.log(math.prod(p)), 3) # → 1.309
round(-sum(math.log(q) for q in p) / len(p), 3) # → 0.436
In code: the lesson's loss section computes exactly this average over the answer positions only (the loss mask), and cross_entropy_from_prob is the −log of one probability.
Two stages: what learns when
Everyday picture: first term, the interpreter studies alone while both experts keep doing exactly what they already do. Second term, the writer is allowed to change too. The photographer (the vision encoder) never changes.
In words: “in stage 1 only the projection matrix learns; in stage 2 the projection and every weight of the language model learn; the vision encoder is frozen throughout.”
With the numbers: stage 1 trains 5,242,880 numbers; stage 2 trains those plus Vicuna-13B's roughly 13.0 billion, a jump of about 2,500 times.
In Python:
# θ in stage 1: W only
W = 5120 * 1024
# φ: the language model's weights (LLaMA-13B size, LLaMA paper Table 2)
phi = 13.0e9
print(f"{W:,}") # → 5,242,880
# θ in stage 2: W and φ, as a multiple of stage 1
round((W + phi) / W) # → 2481
- Stage 1, pre-training for feature alignment. 595K image-caption pairs filtered from CC3M (Conceptual Captions, a set of about three million web images with captions; the filter is in Appendix E below), turned into single-turn examples with the cheap “describe briefly” recipe of §3. Frozen encoder, frozen language model, trainable W. The paper's phrase for it: “training a compatible visual tokenizer for the frozen LLM”.
- Stage 2, fine-tuning end to end. The encoder stays frozen; W and the language model learn. Two uses: a chatbot trained on the 158K samples of §3 (the three types sampled uniformly), and a ScienceQA model trained on that benchmark's questions, with the reasoning and answer as the target.
How long each stage is: stage 1 runs one pass over 595K pairs in batches of 128; stage 2 runs three passes over 158K samples in batches of 32 (§5). On eight A100 GPUs, stage 1 took under 4 hours and stage 2 under 10 (Appendix C).
In Python:
import math
# stage 1: 1 epoch over 595K pairs, batch 128
math.ceil(595_000 / 128) # → 4649
# stage 2: 3 epochs over 158K samples, batch 32
math.ceil(158_000 * 3 / 32) # → 14813
In code: align_projector is stage 1 in miniature: the vision encoder and language model of build_toy_vlm stay frozen and only the projector learns, from 80 toy pictures and their one-word captions; caption_accuracy measures the result. Stage 2 is ordinary supervised fine-tuning, as in the training stages lesson.
Why it matters today
Stage 1 is cheap and safe: it touches a few million weights and cannot damage what either big model knows. The paper's ablation (§5.2) shows it is not optional: skipping it costs 5.11 points on ScienceQA. The two-stage split, align a small connector first and then tune on instructions, survives in most vision-language recipes.
5 Experiments · original
Everyday picture
Two tests: can LLaVA hold a useful conversation about a picture (a chatbot, judged by another model), and can it answer science questions with pictures in them (ScienceQA, graded right or wrong)? All training used eight A100 GPUs with Vicuna's hyperparameters: stage 1 at a learning rate of 2e-3, stage 2 at 2e-5.
5.1 Multimodal chatbot · original
Everyday picture: the ironing man
The paper opens with a photo from the GPT-4 technical report: a man ironing clothes on a board fixed to the back of a yellow taxi in traffic. Asked “What is unusual about this image?”, LLaVA explains that ironing on a moving vehicle is not where one usually irons and is unsafe. BLIP-2 answers with a caption (“a man is sitting on the back of a yellow cab”); OpenFlamingo says the man is drying his clothes on the hood of his car. The difference is not eyesight but following the instruction: the other two describe, LLaVA answers the question. The paper notes LLaVA was trained on only about 80K unique images, and this photo is not among them.
How do you score an open-ended answer?
Everyday picture: you can't mark a free-form answer against an answer key, so you hire an examiner. Here the examiner is a text-only GPT-4. It gets the question, a written description of the image, and two answers: the candidate's (LLaVA's, which saw the image) and a reference written by text-only GPT-4 from the ground-truth description. It rates each answer from 1 to 10 for helpfulness, relevance, accuracy and detail, with an explanation. See the LLM-as-a-judge companion for what such examiners get right and wrong.
Tiny example: the judge gives the candidate 7 and the reference 8. The candidate's relative score is 7 / 8 = 87.5%.
In words: “the candidate's score as a percentage of the reference's score, both given by the same judge.” The paper reports “relative scores w.r.t. the text-only GPT-4 model” without printing a formula; this is the plain reading of that phrase, and the one that makes a headline like “85.1%” mean what it says.
With the numbers: 7 / 8 × 100% = 87.5%. Over a benchmark, the scores are totalled over all questions first.
In Python:
# the judge's ratings, 1 to 10
S_cand, S_ref = 7, 8
S_cand / S_ref * 100 # → 87.5
The reference is a strong one: it answers from the ground-truth annotations, so it cannot misread the picture. A relative score near 100% means “as good as a strong model that was told exactly what is in the image”.
LLaVA-Bench (COCO): which data matters?
30 COCO validation images, each with one question of each of the three types: 90 questions. The images are of the same kind as the training data, so this test isolates the effect of the training data mix.
| Training data | Conversation | Detail | Reasoning | All |
|---|---|---|---|---|
| Full data | 83.1 | 75.3 | 96.5 | 85.1 |
| Detail + Complex | 81.5 (−1.6) | 73.3 (−2.0) | 90.8 (−5.7) | 81.9 (−3.2) |
| Conv + 5% Detail + 10% Complex | 81.0 (−2.1) | 68.4 (−7.1) | 91.5 (−5.0) | 80.5 (−4.4) |
| Conversation | 76.5 (−6.6) | 59.8 (−16.2) | 84.9 (−12.4) | 73.8 (−11.3) |
| No instruction tuning | 22.0 (−61.1) | 24.0 (−51.3) | 18.5 (−78.0) | 21.5 (−63.6) |
Reading it: each bar is one training mix's overall relative score on the same 90 questions. The shortest bar is the model after stage 1 only: it can caption, but asked a question it scores 21.5%. Any instruction tuning adds over 50 points. Among the tuned models, conversation data alone reaches 73.8%; adding just a little description and reasoning data (5% and 10% of it) adds about 7 points, and even improves the conversation column, from 76.5 to 81.0. All three types together are best, at 85.1%. Reasoning questions score highest of the three columns: the judge rewards the long, careful answers GPT-4-written data teaches.
In Python:
# overall relative score, %
full, conv_only, conv_plus, untuned = 85.1, 73.8, 80.5, 21.5
# what instruction tuning adds over stage 1 alone
round(full - untuned, 1) # → 63.6
# what a little detail and reasoning data adds to conversation data
round(conv_plus - conv_only, 1) # → 6.7
LLaVA-Bench (In-the-Wild): harder pictures
24 new images (indoor and outdoor scenes, memes, paintings, sketches) with 60 questions and very detailed hand-written descriptions for the judge.
| Model | Conversation | Detail | Reasoning | All |
|---|---|---|---|---|
| OpenFlamingo | 19.3 ± 0.5 | 19.0 ± 0.5 | 19.1 ± 0.7 | 19.1 ± 0.4 |
| BLIP-2 | 54.6 ± 1.4 | 29.1 ± 1.2 | 32.9 ± 0.7 | 38.1 ± 1.0 |
| LLaVA | 57.3 ± 1.9 | 52.5 ± 6.3 | 81.7 ± 1.8 | 67.3 ± 2.0 |
| LLaVA† | 58.8 ± 0.6 | 49.2 ± 0.8 | 81.4 ± 0.3 | 66.7 ± 0.3 |
Reading it: overall relative scores on the harder set. LLaVA's bar is about 29 points longer than BLIP-2's and 48 longer than OpenFlamingo's; the paper writes these gaps as “+29%” and “+48%”, but they are differences in percentage points. The ± is the spread over three runs. For the first three rows, the three runs are three separate generations by the model; for LLaVA† (the striped bar), one set of LLaVA answers is judged three times, and its small spread (± 0.3) shows the judge is consistent with itself. The striking column is reasoning: 81.7% of the reference's score, while detailed description is the weakest, at 52.5%.
In Python:
llava, blip2, openflamingo = 67.3, 38.1, 19.1
# the gaps, in percentage points
round(llava - blip2, 1), round(llava - openflamingo, 1) # → (29.2, 48.2)
Limitations: a bag of patches
“This indicates that, at times, LLaVA perceives the image as a “bag of patches”, failing to grasp the complex semantics within the image.”Liu et al. (2023), §5.1
Two of the hard examples show where it breaks. Naming the restaurant from a photo of a ramen bowl needs broad knowledge and multilingual understanding. Reading the brand on a yogurt in a crowded fridge needs a high-resolution view, and LLaVA sees every photo at 224 pixels a side. And asked whether there is strawberry-flavoured yogurt in the fridge, LLaVA says yes, when the fridge holds yogurt and, separately, strawberries. Both words' patches are present; the relation between them is not. It is the image version of a bag of words, and a clean example of visual hallucination.
Why it matters today
The judge-based relative score became a common way to evaluate open-ended multimodal chat, with its known weaknesses (a judge that only reads a text description of the image, a single reference). The failures named here, fine detail at low resolution and relations between objects, are the ones later models attacked with higher resolutions and more varied data.
5.2 ScienceQA · original
“To the best of our knowledge, this is the first time that GPT-4 is used for model ensembling.”Liu et al. (2023), §5.2
Everyday picture
A school science exam with multiple-choice questions, some with a diagram or a photo, each with a worked explanation. ScienceQA has about 21,000 such questions across 3 subjects, 26 topics, 127 categories and 379 skills. LLaVA is fine-tuned on its training split and asked to write its reasoning first, then its answer: a trained-in chain of thought (see the chain-of-thought companion).
In Python:
# train, validation and test examples
splits = [12726, 4241, 4241]
sum(splits) # → 21208
The results
| Method | IMG | NO | Average |
|---|---|---|---|
| Human | 87.50 | 88.10 | 88.40 |
| GPT-3.5 with chain of thought | 67.43 | 79.93 | 75.17 |
| LLaMA-Adapter | 80.32 | 86.90 | 85.19 |
| MM-CoT (Large), previous best | 88.80 | 92.89 | 91.68 |
| GPT-4† (2-shot) | 70.75 | 90.73 | 82.69 |
| LLaVA | 88.00 | 90.66 | 90.92 |
| LLaVA + GPT-4† (complement) | 87.80 | 91.08 | 90.97 |
| LLaVA + GPT-4† (judge) | 88.99 | 93.52 | 92.53 |
LLaVA alone reaches 90.92%, close to the previous best of 91.68%. GPT-4 without images, shown two examples, reaches 82.69%; it often fails by saying it lacks the image. Its IMG column (70.75%) is well above chance, because some questions with an image do not actually need it. Two ways to combine the models:
- Complement: use GPT-4's answer, and fall back to LLaVA's whenever GPT-4 gives none. 90.97%: almost exactly LLaVA alone.
- GPT-4 as the judge: when the two answers differ, show GPT-4 the question and both answers (with their reasoning) and ask it for a final answer. 92.53%, a new best, and better than LLaVA in every question category.
Hover or tap a part. Start with the question at the top.
Reading it: both models answer the same question; only LLaVA sees the image. The circle compares their final choices. If they agree, that answer stands (the wire straight down). If they differ, GPT-4 is asked again, now with both answers and their reasoning in front of it, and its pick is final. The judge helps in two ways the paper describes: on questions that don't really need the image, GPT-4's knowledge can overrule a LLaVA mistake, and it can spot reasoning that doesn't hold up. An ensemble usually averages models; this one lets the stronger reasoner arbitrate.
In Python:
# the ensemble rule, with the judge as a function of both answers
def ensemble(llava, gpt4, judge):
return llava if llava == gpt4 else judge(llava, gpt4)
ensemble("A", "A", judge=lambda a, b: b) # → 'A'
ensemble("B", "A", judge=lambda a, b: b) # → 'A'
# the new best over the old one, and over LLaVA alone, in points
round(92.53 - 91.68, 2), round(92.53 - 90.92, 2) # → (0.85, 1.61)
The appendix's worked example (Table 10) is a rocking chair: LLaVA reasons that its seat is silk and answers (B); GPT-4, without the image, answers wood, (A); the judge finds silk implausible and picks (A), correctly. In its explanation the judge credits the right answer to “Assistant 1”, though it was Assistant 2 who said wood: a slip in GPT-4's own output that the paper prints as it came, and a small reminder that a judge's reasons deserve the same scrutiny as its verdicts.
Ablations
| Variant | Features before the last layer | Last-layer features |
|---|---|---|
| Best variant (reasoning, then answer) | 90.92 | 89.96 (−0.96) |
| Predict the answer first | n/a | 89.77 (−1.15) |
| Training from scratch (no stage 1) | 85.81 (−5.11) | n/a |
| 7B model size | 89.84 (−1.08) | n/a |
Reading it: each bar is how many points of ScienceQA accuracy a variant loses against the best 13B model (90.92%). Skipping stage 1 costs by far the most, 5.11 points: the alignment stage is doing real work. The other bars are each about a point: last-layer instead of second-to-last features (0.96), a 7B language model instead of 13B (1.08), and answering before reasoning (1.15). Look at the table before reading too much into that last bar: the answer-first run also used last-layer features, so its 1.15 mixes two changes. Against reasoning-first with the same last-layer features (89.96%), answering first costs only 0.19 points.
In Python:
best, last_layer, answer_first = 90.92, 89.96, 89.77
round(best - answer_first, 2) # → 1.15
# the answer-order effect alone, both runs on last-layer features
round(last_layer - answer_first, 2) # → 0.19
A sentence that doesn't match the table: describing this ablation, the text says “answer-first reports the best number 89.77% accuracy in 12 epochs, while reasoning-first can quickly reach 89.77% accuracy in 6 epochs, but no further improvement with more training.” Read literally, both orders stop at 89.77%. Table 8 has reasoning-first at 89.96% with the same features, and at 90.92% for the best variant. The likely meaning is that reasoning-first matched answer-first's best score in half the epochs; the paper's conclusion is that reasoning first “can largely improve convergence, but contributes relatively little to the final performance”. This page follows the table's numbers.
Why it matters today
Two lessons outlived the leaderboard. A tuned small open model plus a strong general model as arbiter beat either alone. And writing the reasoning before the answer mostly sped up learning here, rather than raising the ceiling: a useful counterpoint to the idea that chain of thought always helps.
6 Conclusion · original
“This paper is an initial step in visual instruction tuning, and mainly focuses on real-life tasks.”Liu et al. (2023), §6
Everyday picture
Take two models that already exist, connect them with the thinnest possible adapter, and spend the effort on the practice material instead. The paper's bet was that data, not architecture, was the missing piece for visual assistants.
Why it matters today
The follow-up the conclusion points to, “improved baselines with visual instruction tuning” (LLaVA-1.5), kept the recipe and changed only the parts named as weaknesses here: a two-layer MLP projector, a CLIP encoder at 336 pixels, and academic question-answering data added to the mix.
Appendices, briefly · original
What they contain
- A, broader impact: user text is filtered with a moderation API and images with a not-safe-for-work filter; the model can hallucinate and inherits biases from CLIP and LLaMA/Vicuna.
- B, more results: LLaVA writes HTML and JavaScript for a joke website from a hand-drawn sketch (after one small fix), explains memes, recognises the Mona Lisa, and reads text in images (OCR), which is “rarely covered” in its training data. It recognises Elon Musk, who appears in neither stage's training data; the paper's reading is that CLIP had seen him and the language model generalises to the new visual concept.
- C, training details: Adam with no weight decay, a cosine learning rate with 3% warmup, FSDP and gradient checkpointing to fit memory, BF16. Stage 1 under 4 hours, stage 2 under 10, ScienceQA under 4, all on 8 A100s.
- E, data: the question lists used for brief and detailed descriptions (such as “Describe the image concisely.”), and the CC3M filter, below.
- F, prompts: the system prompt and few-shot layout used to make GPT-4 write the conversations.
Appendix E: choosing 595K captions from CC3M · original
Everyday picture
You want a small caption set that still covers as many different things as possible. So count how often each noun phrase (“a red bus”, “the beach”) appears across all captions, ignore the very rare ones, and then go through the phrases from rarest to most common, adding captions that mention each one, but never more than a fixed number per phrase. Rare concepts get all their captions; common ones are capped.
Tiny example
The paper's rule, with its thresholds: skip phrases seen fewer than 3 times; take at most 100 captions per phrase, chosen at random. Below, the same rule on seven toy captions with the cap lowered to 3 so it bites: “dog” appears in 5 captions, “cat” in 3, “bus” in 1. “bus” is too rare to count. “cat” is rarer than “dog”, so it goes first and keeps all three of its captions. “dog” has five, more than the cap, so three are drawn at random. Caption 4 survives because it also mentions a dog; caption 3, a plain dog caption that was not drawn, is left out: 6 of 7 kept.
In Python:
import random
from collections import Counter
captions = {1: ["dog"], 2: ["dog", "cat"], 3: ["dog"], 4: ["dog", "bus"], 5: ["cat"], 6: ["cat"], 7: ["dog"]}
def filter_captions(captions, min_count=3, cap=100, seed=0):
rng = random.Random(seed)
counts = Counter(p for phrases in captions.values() for p in phrases)
pool = set()
# rarest phrases first, skipping those seen fewer than min_count times
for phrase in sorted((p for p in counts if counts[p] >= min_count), key=lambda p: counts[p]):
having = [c for c, phrases in captions.items() if phrase in phrases]
if len(having) > cap:
having = rng.sample(having, cap)
pool.update(having)
return sorted(pool)
Counter(p for phrases in captions.values() for p in phrases) # → Counter({'dog': 5, 'cat': 3, 'bus': 1})
filter_captions(captions, cap=3) # → [1, 2, 4, 5, 6, 7]
The paper says the filter keeps “a good coverage of concepts whose frequency is higher from 3” with far fewer pairs (its Figure 7), ending at about 595K of CC3M's roughly three million.
Why it matters today
Stage 1 only has to teach the projector a vocabulary, so breadth of concepts matters more than repetition. Selecting for coverage rather than size is a cheap way to cut data, and the same idea appears in deduplication and data-mixture choices for pre-training.
What changed since 2023
The recipe, encode with a pretrained vision encoder, project, splice into the prompt, tune on instructions, is still the backbone of open vision-language models. Around it:
| In the paper | Common since | Why | Lesson |
|---|---|---|---|
| One linear projection | A small two-layer MLP (from LLaVA-1.5 on) | A little more capacity in the adapter, at negligible cost | multimodal |
| 224-pixel images, 256 tokens | Higher resolutions, images cut into tiles, neighbouring tokens merged | Fine detail (brands, text) needs pixels; tokens grow with the square of the side | multimodal |
| 158K GPT-4-written samples from COCO | Much larger and more varied instruction mixes, including academic question answering | The ablation above: data variety drives quality | training stages |
| A text-only GPT-4 as the judge | Judges that can see the image, and benchmarks with fixed answers | A judge reading only a description can't check visual claims | evals |
Glossary
Every term with hover guidance on this page, in one place.