An annotated companion · AI Primer

Flamingo: a Visual Language Model for Few-Shot Learning, annotated

About this page. This is a companion, not a copy. It follows the paper (arXiv version 2, November 2022; published at NeurIPS 2022) section by section, quotes only a sentence or two per section (clearly marked), and explains everything in its own words. The equations are reproduced because mathematics is not copyrightable, and the pseudo-code of its Figures 4 and 5 is paraphrased as equations. The paper is distributed under arXiv's standard non-exclusive licence, so its tables are not reproduced: a few selected rows appear in this page's own tables and charts, each attributed, and its figures are redrawn from scratch. Worked numbers come from small examples built on this page, not from the paper, unless a sentence says otherwise; made-up numbers are labelled illustrative. Read the original alongside: every section links to it.

How to read this page

Nothing here assumes you already know the jargon. Three things help:

  • Any dotted word explains itself when you hover it, tab to it, or tap it.
  • Every symbol inside an equation does the same, and each equation is followed by a table of its symbols, a sentence reading it aloud, the numbers of a tiny example, and the same example in a few lines of Python you can run.
  • The pictures are live: hover or tap the parts of the three redrawn architecture figures, turn the gate of a new layer from closed to open, see which image each word is allowed to look at, and compare model sizes and shot counts on the paper's benchmarks.

The code that builds the pieces from scratch lives in the multimodal lesson. Its toy model takes the other common road, the one of the LLaVA companion: it splices image tokens into the prompt, where Flamingo adds new layers that look at the image from inside the language model. Comparing the two is one of the best ways to understand both.

Abstract

“We introduce Flamingo, a family of Visual Language Models (VLM) with this ability.”Alayrac et al. (2022), Abstract. Read the original

Everyday picture

“This ability” is learning a new task from a handful of examples. Show a friend three labelled photos of birds, then a fourth unlabelled one, and they will name it in the same style without being taught anything else. Large language models learned to do this with text (see the GPT-3 companion). Flamingo does it with pictures and video mixed into the text: a visual language model that reads images, videos and words in any order and writes words.

What the paper claims

  • Architecture: a way to join a pretrained, frozen vision model to a pretrained, frozen language model, to handle any interleaving of pictures and text, and to take images or video as input.
  • Data: training on web pages with images scattered through the text, plus image-caption and video-caption pairs, and no data annotated for machine learning.
  • Results: one model, prompted with task examples and never fine-tuned, sets a new few-shot state of the art across the 16 benchmarks it tests, and on 6 of them beats models fine-tuned on thousands of times more task-specific data.

Why it matters today

Flamingo showed that the “prompt it with examples” interface of large language models carries over to vision, and it popularised a recipe still in use: keep the big pretrained models frozen and train small new parts that connect them.

1 Introduction · original

“One key aspect of intelligence is the ability to quickly learn to perform a new task given a short instruction.”Alayrac et al. (2022), §1

Everyday picture

In 2022, there were two ways to point a vision model at a new task. The first is fine-tuning: collect thousands of labelled examples and keep training. It works, but it is slow and fiddly. The second came from contrastive models such as CLIP, which score how well a caption matches an image, so they can pick the best of a list of labels with no training at all (zero-shot). But a scorer can only choose from answers you give it. It cannot write a caption or answer an open question.

Tiny example

The paper's Figure 1 opens with a prompt of this shape, where each [photo] is an image:

[photo of a chinchilla] “This is a chinchilla. They are mainly found in Chile.” [photo of a shiba] “This is a shiba. They are very popular in Japan.” [photo of a flamingo] “This is”

and the model continues: “a flamingo. They are found in the Caribbean and South America.” Two examples were enough to set the pattern: name the animal, then add one fact about where it is found. A scorer like CLIP could only rank labels you supply; here the model writes the answer.

The contributions

  1. The Flamingo family of models, which take arbitrarily interleaved images, videos and text and generate text.
  2. A careful few-shot evaluation, with a large set of benchmarks held out from every design decision so the estimates are not flattered.
  3. A new few-shot state of the art on 16 tasks; on 6, beating the fine-tuned state of the art with only 32 examples, around 1,000 times less task-specific data.
  4. With fine-tuning, a new state of the art on five more: VQAv2, VATEX, VizWiz, MSRVTTQA and HatefulMemes.

Why it matters

Casting every vision task as “continue this text” is what lets one model handle captioning, question answering and classification without a separate head for each. The big-picture lesson explains in-context learning for text; this paper is where it reached pictures at scale.

2 Approach · original

“First, the Perceiver Resampler (Section 2.1) receives spatio-temporal features from the Vision Encoder (obtained from either an image or a video) and outputs a fixed number of visual tokens.”Alayrac et al. (2022), §2

Everyday picture

Picture a brilliant writer (the language model) and a sharp-eyed photographer (the vision encoder), both already trained and both told to stay exactly as they are. Flamingo adds two things between them. A note-taker (the Perceiver Resampler) boils each photo down to the same short set of notes, however big the photo or long the video. And a set of new glances (gated cross-attention layers) lets the writer look at those notes between its own thinking steps. The glances start with their eyes shut, so on day one the writer writes exactly as it always did.

Tiny example

The web page “[dog photo] This is a very cute dog. [cat photo] This is” becomes the text “<image> This is a very cute dog. <image> This is” plus two photos. Each photo becomes a fixed set of visual tokens. The language model reads only the text; whenever it reaches a word, its new cross-attention layers may look at the tokens of the photo that came just before that word. The model continues: “a very serious cat.”

Output: text“a very serious cat.” n-th LM block (frozen) n-th gated xattn-dense(trained) ⋯ more pairs ⋯ 1st LM block (frozen) 1st gated xattn-dense(trained) <image> This is a verycute dog. <image> This is a web page: two photosand the sentences around them each photo Vision Encoderfrozen NFNet-F6 PerceiverResampler 64 visual tokensper image frozen: the vision encoder and the LM blockstrained from scratch: the resampler, the gated layers

Hover or tap a part. Start at the bottom: the web page splits into photos (left) and text (right), and the two meet in the gated xattn-dense layers.

Figure 3 of the paper, redrawn: the Flamingo architecture. Based on Alayrac et al. (2022), Figure 3.

Reading it: read from the bottom up, in two columns. On the left, each photo passes through the frozen vision encoder and then the Perceiver Resampler, which hands back 64 visual tokens per image. On the right, the text (with an <image> tag where each photo was) climbs the language model's stack of frozen blocks. Before each chosen frozen block sits a new gated xattn-dense layer, and the wire from the left feeds the visual tokens into every one of them. The visual tokens never join the text sequence: the text stays as short as it was, and the pictures reach it only through these side doors. The colours follow the paper's split: the vision encoder and the language model blocks are frozen, the resampler and the gated layers are trained from scratch.

The math: equation 1

Flamingo is a text model that happens to see pictures: it gives a probability to each next word, and multiplies them for the whole text.

In words: “the probability of the whole text, given the pictures, is the product over positions of the probability of each word, given the words before it and the pictures that appeared before it.”

With the numbers (illustrative): after the cat photo, the model gives “a” probability 0.5, then “serious” 0.4, then “cat” 0.9. The three-word continuation has probability 0.5 × 0.4 × 0.9 = 0.18.

In Python:

import math
# p(y_ℓ | y_<ℓ, x_≤ℓ) for "a", "serious", "cat"
p = [0.5, 0.4, 0.9]
# Π over ℓ
prob = 1
for p_l in p:
    prob *= p_l
round(prob, 2)  # → 0.18
# the loss is the negative log of the same thing, a sum
round(-sum(math.log(p_l) for p_l in p), 3)  # → 1.715

The only difference from an ordinary language model is the x≤ℓ: each word may depend on the images and videos that came before it in the interleaved sequence (§2.3 says exactly which).

Why it matters

Because the output is just text probabilities, every vision task becomes “write the answer” or “score the candidate answers”, the same interface as a language model. The losses lesson covers the negative log-likelihood this product turns into.

2.1 Visual processing and the Perceiver Resampler · original

“It takes as input a variable number of image or video features from the vision encoder and produces a fixed number of visual outputs (64), reducing the computational complexity of the vision-text cross-attention.”Alayrac et al. (2022), §2.1

Everyday picture

A meeting has 64 note-takers, each with a standing question such as “what objects are there?” or “what is happening on the left?”. Hand them a single photo or a thirty-frame video: each note-taker skims all of it, writes one note, and the meeting always ends with exactly 64 notes. The questions themselves are learned. That is the Perceiver Resampler: a small transformer whose queries are a fixed set of learned vectors rather than the input.

The vision encoder first

  • A Normalizer-Free ResNet, the F6 model, pretrained by the authors with a contrastive loss on image-text pairs (the two-term loss of CLIP; see the CLIP companion), then frozen.
  • Its last stage gives a 2D grid of features, flattened into a list. For video, frames are sampled at 1 per second and encoded one by one, and a learned time embedding is added to each frame's features.

Tiny example

One learned latent query x = (1, 0) looks at three visual features, (2, 0), (0, 2) and (1, 1). A detail from the paper's Appendix A.1.1: the keys and values include the latents themselves, so the list it attends over is (2, 0), (0, 2), (1, 1) and (1, 0). With the projection matrices set to the identity to keep the numbers small, the scores are the dot products divided by √2: 1.414, 0, 0.707, 0.707. Softmax turns them into weights 0.449, 0.109, 0.221, 0.221. The weighted blend of the values is (1.340, 0.439), and the residual adds it to the query: the latent becomes (2.340, 0.439). It has soaked up mostly the first feature, the one that matched its question best.

R visual tokens outR = 64, whatever came in × 6 layers FFW, plus residualx = x + ffw(x) Attention, plus residualQ = X · K = V = [Xf, X] learned latentqueries X Xf: flattenT × S features → T·S + time embedding encoderencoderencoder t = 0t = 1t = 2

Hover or tap a part. Start at the bottom with the video frames, then follow the features and the learned queries into the attention step.

Figure 5 of the paper, redrawn: the Perceiver Resampler on a three-frame video. Based on Alayrac et al. (2022), Figure 5 and its pseudo-code.

Reading it: read from the bottom up. Each frame goes through the frozen vision encoder on its own; its grid of S spatial features gets that frame's time embedding added; then all T frames are flattened into one long list Xf of T × S features (an image is a video with T = 1). On the right sit R learned latent queries X. Inside the dashed box, which repeats 6 times, the latents ask the questions (Q) and the keys and values are the features and the latents together; a feed-forward layer follows, each step with a residual connection. Whatever the size of Xf, the output at the top has exactly R = 64 rows.

The math: one resampler layer

The paper gives this layer as pseudo-code; written as equations, with the attention of the Transformer companion:

In words: “each learned latent asks a question of the visual features and of the latents themselves, blends the answers by how well they match, and adds the blend to what it already holds.” (Then x ← x + ffw(x), and repeat 6 times.)

With the numbers: scores (1.414, 0, 0.707, 0.707); weights (0.449, 0.109, 0.221, 0.221); blend (1.340, 0.439); new latent (2.340, 0.439).

In Python:

import math
# one learned latent query, and three visual features
x = [1, 0]
X_f = [[2, 0], [0, 2], [1, 1]]
# K = V = [X_f; X]: the latents join the keys and values (W = identity here)
kv = X_f + [x]
d = 2
scores = [sum(a * b for a, b in zip(x, k)) / math.sqrt(d) for k in kv]
[round(s, 3) for s in scores]  # → [1.414, 0.0, 0.707, 0.707]
exps = [math.exp(s) for s in scores]
w = [e / sum(exps) for e in exps]
[round(v, 3) for v in w]  # → [0.449, 0.109, 0.221, 0.221]
blend = [sum(w_j * v[i] for w_j, v in zip(w, kv)) for i in range(d)]
[round(v, 3) for v in blend]  # → [1.34, 0.439]
# the residual: X ← X + attention
[round(a + b, 3) for a, b in zip(x, blend)]  # → [2.34, 0.439]

Some facts from the appendix: the resampler has 6 layers, width 1,536 and 16 heads (so each head works on 1,536 / 16 = 96 numbers), about 200 million parameters in every model size, and no spatial position encodings, only the time ones: they did not help, probably, the paper suggests, because a convolutional encoder already carries position in its channels. Concatenating the latents to the keys and values, which the original Perceiver does not do, was found “to perform slightly better”.

# width per head in the resampler: D / H
1536 // 16  # → 96

Why it matters today

A fixed number of tokens per image, however large, keeps the cost of looking at pictures predictable. Compressing an image to a small set of learned-query outputs reappears in later vision-language models; the multimodal lesson describes this family next to the splice-in-the-prompt design, and builds the attention with scaled_dot_product_attention.

2.2 Conditioning frozen language models on visual representations · original

“This multiplies the output of a newly added layer by tanh(α) before adding it to the input representation from the residual connection, where α is a layer-specific learnable scalar initialized to 0.”Alayrac et al. (2022), §2.2

Everyday picture

You are adding a new pipe into a building's water supply. If you open it fully on day one, whatever is in the new pipe floods the system. So you fit a valve, start it closed, and open it gradually as you learn what the new pipe carries. The gated cross-attention layer is the new pipe (visual information into the language model) and tanh(α), starting at tanh(0) = 0, is the valve.

Tiny example

A word's vector inside the language model is y = (1, 2). The new cross-attention layer, looking at the image, proposes the change a = (0.5, −1) (illustrative). At the start of training α = 0, so tanh(α) = 0 and y stays (1, 2): the language model behaves exactly as before. Once training has moved α to 0.5, tanh(0.5) = 0.462 and y becomes (1 + 0.462 × 0.5, 2 − 0.462 × 1) = (1.231, 1.538).

Y outvisually informed words LM layer(frozen) FFWy = y + frozen_ffw(y) self-attention on Y gatedxattn-dense(new, trained) × tanh(α_dense), then add FFWnew, squared ReLU × tanh(α_xattn), then add cross-attentionQ = Y · K = V = X language Y vision X

Hover or tap a part. Start at the bottom: the language input Y on the right, the visual tokens X on the left.

Figure 4 of the paper, redrawn: a gated xattn-dense layer inserted before a frozen language model layer. Based on Alayrac et al. (2022), Figure 4 and its pseudo-code.

Reading it: read from the bottom up. The word vectors Y enter the new block first. Its cross-attention takes its queries from the words and its keys and values from the visual tokens X, which arrive from the left. The result is multiplied by the gate tanh(αxattn) and added back to Y (a residual connection, as in every transformer layer); then a new feed-forward layer does the same with its own gate, tanh(αdense). Only then does the untouched, frozen language model layer run: its self-attention and feed-forward network, each with its usual residual. With both gates at zero, the new block adds nothing, and the stack computes exactly what the original language model computed.

The math: the two gated updates

In words: “let the words look at the image, scale what they find by a learned valve that starts closed, and add it on; then do the same with a new feed-forward layer and its own valve.”

With the numbers: y = (1, 2), attention output (0.5, −1). α = 0: y stays (1, 2). α = 0.5: tanh = 0.4621 and y = (1.231, 1.538). α = 3: tanh = 0.9951 and y = (1.498, 1.005), almost the full change.

In Python:

import math
y = [1, 2]
# attention(q = y, kv = x): what the words found in the image (illustrative)
a = [0.5, -1]
def gated(y, a, alpha):
    # y + tanh(α) · a
    return [round(y_i + math.tanh(alpha) * a_i, 3) for y_i, a_i in zip(y, a)]
gated(y, a, 0)  # → [1.0, 2.0]
round(math.tanh(0.5), 4), gated(y, a, 0.5)  # → (0.4621, [1.231, 1.538])
round(math.tanh(3), 4), gated(y, a, 3)  # → (0.9951, [1.498, 1.005])

Try it: open the valve

0.00

Reading it: the slider sets the learnable scalar α of one gated layer, and the top bar shows how open the valve is, |tanh(α)|, from 0 to 1. The two lower bars show the word's two numbers, y1 and y2, after the gated update; the readout prints them. At α = 0 they are the language model's own (1, 2), whatever the image says: this is why Flamingo, at the first step of training, is exactly the pretrained language model. Slide right and the image's contribution fades in; slide left and it enters with the opposite sign. tanh never goes past ±1, so no single gate can amplify the visual signal without limit. The paper's Appendix A.1.2 plots the real gates of Flamingo-3B: in every layer they move quickly away from 0 during training, and they seem to grow with depth, though the paper warns against reading much into that.

How many new layers?

A gated layer before every language-model block is best but expensive. The paper inserts one before every block for Flamingo-3B, every fourth for Flamingo-9B and every seventh for Flamingo-80B, starting from the first block. Their language models are the 1.4B, 7B and 70B Chinchilla models, with 24, 40 and 80 blocks (Table 4 of the paper).

Model sizes, from Tables 4 and 5 of Alayrac et al. (2022)
ModelGated layersFrozen LMTotal
Flamingo-3B24 of 24 blocks1.4B3.2B
Flamingo-9B10 of 40 (every 4th)7.1B9.3B
Flamingo (80B)12 of 80 (every 7th)70B80B
# one gated layer every k blocks, starting from the first
[len(range(0, n_blocks, k)) for n_blocks, k in [(24, 1), (40, 4), (80, 7)]]  # → [24, 10, 12]
# Table 5's parts, in billions: frozen LM + frozen vision + gated layers + resampler
[round(sum(parts), 1) for parts in [(1.4, 0.435, 1.2, 0.194), (7.1, 0.435, 1.6, 0.194), (70, 0.435, 10, 0.194)]]  # → [3.2, 9.3, 80.6]
# the share of Flamingo-80B that is frozen
round((70 + 0.435) / 80.629, 3)  # → 0.874

The gated layers are transformer layers of the same width as their language model, and use the squared ReLU activation; the frozen language model keeps its GELU.

Why it matters today

Gating a new branch so it starts as the identity is a general trick for adding parts to a trained network without breaking it on day one; the paper cites earlier uses of it. Here it is what lets a 70-billion-parameter language model take new inputs while staying frozen. The TransformerBlock of the transformer lesson is the kind of block the gated layers sit in front of.

2.3 Multi-visual input support: per-image masking · original

“At a given text token, the model attends to the visual tokens of the image that appeared just before it in the interleaved sequence, rather than to all previous images”Alayrac et al. (2022), §2.3

Everyday picture

A photo album with a caption under each photo. When you read a caption, you look at the photo just above it, not at every photo in the album. You still remember the earlier photos, because you read their captions on the way down.

Tiny example

The sequence “Cute pics of my pets! <image> My puppy. <image> My cat.” has two images. Number the images 1 and 2, and give every text token the number of the last image before it, or 0 if there is none: that numbering is the function φ. “Cute” gets 0, “puppy” gets 1, “cat” gets 2. In the cross-attention, each token may look only at the 64 visual tokens of image φ. “cat” cannot look directly at the puppy photo, but it has read “My puppy.” through the language model's ordinary self-attention, and those words did look at the puppy.

Cross-attend to:

Hover or tap a cell. Each row is a text token; the columns are the visual tokens of image 1 and image 2 (three drawn for each, standing for 64).

The masked cross-attention of Figure 7 of the paper, redrawn on a shorter sequence. Based on Alayrac et al. (2022), Figure 7.

Reading it: each row is one text token, with its φ value at the far left; each column is a visual token, the first three from image 1 and the next three from image 2. A filled cell with a tick means the row may attend to that column; an empty cell is masked out. With Flamingo's rule the allowed cells form a staircase: rows before any image see nothing (φ = 0), rows after the first image see only image 1, and rows after the second see only image 2. Switch to “all previous images” and the staircase fills in to the left: the last rows now see both images. The paper tried both, and the single-image rule scored 7.2 points higher overall (Appendix B.3.1). In this page's tokenisation the <image> tag itself counts as coming after its image, as the paper's Figure 7 draws it.

The math: φ and the images a token may use

In words: “φ gives every text position the number of the last image before it (0 if none); the images a token's prediction may depend on are those numbered up to φ of its position, but the cross-attention looks directly only at image φ itself.”

With the numbers: in “Cute pics of my pets! <image> My puppy. <image> My cat.”, φ(“Cute”) = 0, φ(“puppy”) = 1, φ(“cat”) = 2; the prediction of “cat” may use images {1, 2}, of which its cross-attention sees image 2.

In Python:

tokens = ["Cute", "pics", "<image>", "My", "puppy", "<image>", "My", "cat"]
# φ: the number of the last image at or before each position (the tag counts as its image)
phi, seen = [], 0
for tok in tokens:
    if tok == "<image>":
        seen += 1
    phi.append(seen)
phi  # → [0, 0, 1, 1, 1, 2, 2, 2]
# x_≤ℓ for "cat": every image numbered up to φ
list(range(1, phi[-1] + 1))  # → [1, 2]
# query-key pairs: 64 per token that has an image, against 64 per earlier image
sum(64 for p in phi if p > 0), sum(64 * p for p in phi)  # → (384, 576)

The last line counts the cross-attention's query-key pairs on this short sequence under the two rules; the gap grows with the number of images. On a training sequence of 256 tokens with 5 images, attending only to the last image costs at most 256 × 64 = 16,384 query-key scores per head, against up to 256 × 320 = 81,920 if every token could see all five.

Why it matters

This rule is what lets the model train on at most 5 images per sequence and then use up to 32 at test time: nothing in a token's cross-attention depends on how many images came before. The paper tried giving images explicit indices (<image 1>, <image 2>) or learned index embeddings so a token could see all of them; those did not cope as well when the number of images changed between training and test.

2.4 Training on a mixture of vision and language datasets · original

“The few-shot capabilities of Flamingo models rely on training on interleaved text and image data.”Alayrac et al. (2022), §2.4

Everyday picture

A child who learns to read from picture books, where the pictures sit in the middle of the story, learns something that flash cards alone do not teach: that a picture and the sentences around it belong together, and that a page can hold several of each. Flamingo's main diet is web pages kept in their original order, pictures and all, supplemented with captioned images and captioned videos.

The three kinds of data

Training datasets, from §2.4 and Appendix A.3 of Alayrac et al. (2022)
DatasetWhat it holdsSizeWeight λ
M3W (MultiModal MassiveWeb)text and images from about 43 million web pages, in page order43M pages1.0
ALIGNimages with their alt-text (noisy, short: 12.4 tokens on average)1.8B pairs0.2
LTIP (Long Text & Image Pairs)images with longer, better descriptions (20.5 tokens on average)312M pairs0.2
VTP (Video & Text Pairs)short videos, about 22 seconds on average, with a sentence each27M videos0.03

A web page becomes a training example by inserting an <image> tag where each image sat in the page's structure (the DOM) and a learned <EOC> (end of chunk) token before each image and at the end. From each page the paper samples 256 tokens and keeps at most the first 5 images in them. Caption pairs get the same format: <image>, the caption, <EOC>.

The math: equation 2

In words: “for each dataset, take the average negative log-likelihood of the text given its pictures, weight it by that dataset's λ, and add the four up.”

With the numbers (the losses are illustrative, the weights are the paper's): if the four average losses were 2.0 (M3W), 3.0 (ALIGN), 2.5 (LTIP) and 3.5 (VTP), the training objective would be 1.0 × 2.0 + 0.2 × 3.0 + 0.2 × 2.5 + 0.03 × 3.5 = 3.205.

In Python:

# λ_m for M3W, ALIGN, LTIP, VTP (Appendix B.1.2)
lam = [1.0, 0.2, 0.2, 0.03]
# E[−Σ log p] for each dataset (illustrative)
nll = [2.0, 3.0, 2.5, 3.5]
round(sum(l * n for l, n in zip(lam, nll)), 3)  # → 3.205

The paper says tuning the weights is “key to performance”, and that accumulating the gradients of all four datasets into each update beats alternating between them (“round-robin”): 70.7 against 62.9 on its overall score (Table 3, row ii).

A small data trick

On a web page, does the sentence after a picture describe it, or does it describe the next picture? Both happen. So for each page the paper flips a coin (pnext = ½) to decide whether text tokens are tied to the previous image or the next one. It is wrong about half the time on purpose, and it still scored a little better than always choosing one way (Appendix A.3.2).

Why it matters today

Interleaved web documents became a standard ingredient for training models that take several images in one conversation. The pretraining lesson covers the data mixture weights that this equation formalises.

2.5 Task adaptation with few-shot in-context learning · original

Everyday picture

To use Flamingo on a new task you do not train it; you write a prompt that shows the task. A few worked examples (the support examples), each an image followed by its answer, then the new image (the query) followed by the start of an answer for the model to finish.

Tiny example

Switch between a two-shot captioning prompt and the paper's zero-shot version, in which the two examples keep their text but lose their images:

Prompt:

Reading it: the chips are the prompt in order. Chips in square brackets are images, chips in angle brackets are special tokens, and the chip with the thick outline is the query image, after which the model writes. In the two-shot prompt each support image is followed by “Output:” and its caption, then <EOC>. In the zero-shot prompt the support images are gone and only their captions remain: the paper found that two text-only examples fix the answer's format without showing any images, and that a single example biased the model towards copying it (Appendix A.2).

Two ways to read the answer

  • Open-ended tasks (captioning, free-form questions): generate with beam search, beam size 3, stopping at the first <EOC>.
  • Close-ended tasks (multiple choice, classification): append each candidate answer to the prompt, score each with the model's log-likelihood, and rank them.

Tiny example (illustrative probabilities): the candidates are “a dog” and “a cat”. The model gives “a” probability 0.6 either way; then “dog” 0.9 and “cat” 0.3. Scores: log 0.6 + log 0.9 = −0.616 against log 0.6 + log 0.3 = −1.715. “a dog” wins.

import math
# p of each token of each candidate, after the prompt
candidates = {"a dog": [0.6, 0.9], "a cat": [0.6, 0.3]}
scores = {c: round(sum(math.log(p) for p in ps), 3) for c, ps in candidates.items()}
scores  # → {'a dog': -0.616, 'a cat': -1.715}
max(scores, key=scores.get)  # → 'a dog'

For tasks with many labelled examples, such as ImageNet classification, the prompt cannot hold them all. The paper then uses retrieval-based example selection (RICES): pick the support examples whose images look most like the query, and put the most similar one last, right before the query.

Why it matters today

This is the same prompt engineering a chat model uses with text (see the context lesson on few-shot prompts), with pictures as part of each example.

3 Experiments · original

“We emphasize that we do not validate any design decisions on these 11 benchmarks and use them solely to estimate unbiased few-shot learning performance of our models.”Alayrac et al. (2022), §3

Everyday picture

If you tune a recipe by tasting it every day, your own opinion of it is no longer a fair review. The paper splits its 16 benchmarks in two: 5 “dev” benchmarks (COCO, OKVQA, VQAv2, MSVDQA and VATEX) that it tasted while designing the model, and 11 it never looked at until the end. The 11 give the honest estimate.

Tiny example

Measuring few-shot skill needs two disjoint sets per benchmark: support examples to put in prompts and query examples to test on. A dev benchmark needs four (support and query for validation, and again for the final test), so that no final number was also used to make a design choice. All benchmarks share one set of evaluation settings, including the prompt formats “Output: {output}” and “Question: {question} Answer: {answer}”, with a few exceptions such as HatefulMemes, whose prompt includes the meme's text.

Why it matters

Few-shot results are easy to inflate by quietly choosing prompts or settings on the test set. Holding out 11 benchmarks from every decision is the paper's guard against that; see the benchmarks lesson for why the best of many tuned versions looks better than it is.

3.1 Few-shot learning on vision-language tasks · original

Everyday picture

Two dials: how big the model is, and how many examples its prompt shows. Turn either up and the scores tend to rise; the biggest model gets the most out of extra examples.

Hover or tap the chart, or focus it and use the arrow keys, to read the scores at 0, 4 and 32 shots.

Reading it: the x-axis is the number of examples in the prompt (0, 4 or 32); the y-axis is the benchmark's score (accuracy in percent for question answering, CIDEr for captioning, where values above 100 are normal). Each line is one model size, and the flat line is the best fine-tuned model at the time, trained on thousands of task-specific examples. All numbers are from Table 1 of the paper. On OKVQA, the three model sizes rise with shots and the 80B model passes the fine-tuned line; on VQAv2 and COCO, even 32 shots stay well below it. Pick VisDial: the 4-shot and 32-shot scores are identical; its dialogues make such long prompts that the paper capped it at 16 shots (§5). Pick STAR: more shots barely help at all. The dev benchmarks are the first five in the list.

Selected rows of Table 1 of Alayrac et al. (2022): Flamingo (80B) with 0, 4 and 32 shots, and the fine-tuned state of the art with its number of training examples
Benchmark0432Fine-tuned
OKVQA50.657.457.854.4 (10K)
VQAv256.363.167.680.2 (444K)
COCO (CIDEr)84.3103.2113.8143.3 (500K)
MSVDQA35.641.752.347.9 (27K)
Flickr30K (CIDEr)67.275.175.467.4 (30K)
iVQA40.744.145.335.4 (6K)
STAR39.742.442.236.7 (46K)
NextQA (WUPS)26.730.833.525.2 (38K)

The highlighted rows are the six benchmarks where 32 shots beat the fine-tuned state of the art. (The text says six; the caption of the paper's Table 1 says seven, but the table itself shows six.) The gap in training data is largest on COCO and VATEX, whose fine-tuned models saw 500,000 examples each.

# Flamingo (80B) at 32 shots against the fine-tuned SotA, all 15 scored benchmarks of Table 1
flamingo_32 = [57.8, 67.6, 113.8, 52.3, 65.1, 49.8, 75.4, 31.0, 45.3, 86.8, 42.2, 55.6, 37.9, 33.5, 70.0]
fine_tuned = [54.4, 80.2, 143.3, 47.9, 76.3, 57.2, 67.4, 46.8, 35.4, 138.7, 36.7, 75.2, 54.7, 25.2, 79.1]
sum(f > s for f, s in zip(flamingo_32, fine_tuned))  # → 6
# 32 examples against 10,000 (OKVQA's fine-tuned model)
10_000 // 32  # → 312

Why it matters

The fine-tuned models used from 6,000 to 500,000 task-specific examples against Flamingo's 32, which the paper summarises as around 1,000 times less data. The pattern that bigger models use more shots better mirrors what the GPT-3 companion shows for text.

3.2 Fine-tuning Flamingo as a pretrained vision-language model · original

Everyday picture

A generalist who is good at picking up tasks from a few examples can still get better with a proper training course. On the nine tasks where 32 shots fell short, the paper fine-tunes Flamingo on each task's full training set.

Tiny example

On VQAv2, Flamingo scores 67.6 with 32 shots. Fine-tuned, it scores 82.0, above the previous best of 81.3.

Selected entries of Table 2 of Alayrac et al. (2022): fine-tuned Flamingo against the state of the art
Benchmark32 shotsFine-tunedPrevious SotA
VQAv2 (test-dev)67.682.081.3
VATEX (CIDEr)65.184.281.4
VizWiz (test-dev)49.865.757.2
MSRVTTQA31.047.446.8
HatefulMemes (AUC)70.086.684.6
COCO (CIDEr)113.8138.1149.6

How: the language model stays frozen; the same layers as in pretraining train, plus, this time, the vision encoder, at a higher image resolution (480 × 480 instead of 320 × 320), with a short schedule and a small learning rate. Five new records (highlighted). On others it trails methods that optimise the test metric directly, such as CIDEr optimisation for COCO captioning, which it does not use.

Why it matters

The same weights serve both uses: prompt them for a quick start, fine-tune them when the data exists. That is how pretrained language models are used too; see the fine-tuning lesson.

3.3 Ablation studies · original

“Interestingly, fine-tuning our pretrained LM also leads to a drop in performance of -8.0%.”Alayrac et al. (2022), §3.3

Everyday picture

To learn which parts of a machine matter, take them out one at a time and see what breaks. Every ablation here is Flamingo-3B, trained for a shorter time, with one design choice changed, scored on the five dev benchmarks with 4 shots.

The overall score

To combine five benchmarks on different scales, the paper divides each score by the state of the art from Table 1 and averages the five ratios.

In words: “express each benchmark score as a fraction of the best known result, and average the fractions.”

With the numbers: the baseline Flamingo-3B scores 86.5 (COCO), 42.1 (OKVQA), 55.8 (VQAv2), 36.3 (MSVDQA) and 53.4 (VATEX); against the fine-tuned records 143.3, 54.4, 80.2, 47.9 and 76.3 the ratios average 0.706, which is 70.6. The paper reports 70.7; the small difference probably comes from rounding in the printed tables.

In Python:

# Table 3's baseline row and Table 1's fine-tuned records
s = [86.5, 42.1, 55.8, 36.3, 53.4]
sota = [143.3, 54.4, 80.2, 47.9, 76.3]
# (1/5) Σ_b s_b / SotA_b, as a percentage
round(100 * sum(a / b for a, b in zip(s, sota)) / 5, 1)  # → 70.6

Reading it: each bar is one ablation's overall score, from Table 3 of the paper; the top bar, striped, is the unchanged Flamingo-3B at 70.7, and the axis runs from 50 to 75 so the gaps are visible. Longer is better. The biggest single loss is removing the interleaved web data M3W (53.4, down 17.3 points), which is the data that teaches in-context learning. Training the language model instead of freezing it hurts twice: from random weights (57.8) and, more surprisingly, even from the pretrained weights (62.7), the catastrophic forgetting the quote refers to. Removing the tanh gate costs 4.2 points and made training unstable. The CLIP ViT-L/14 encoder scores 5.8 points below the authors' own NFNet-F6.

The findings, in short

  • Data: every dataset helps; M3W most. Swapping the image-text pairs for the public LAION-400M costs a little (66.4).
  • Architecture: gated xattn-dense beats plain cross-attention (66.9) and “grafting” new layers on top of the frozen model's output (63.1). The resampler beats a plain transformer (66.7) and an MLP (66.6) of similar size.
  • Cost: a gated layer every fourth block instead of every block drops the overall score only from 70.7 to 68.8, with fewer parameters (2.3B against 3.2B) and a step time of 1.02 s instead of 1.74 s. The paper calls this 66% faster training; the two step times give about 1.7 times as fast.
  • Frozen is better: freezing the language model beats training it, and is cheaper: no gradient updates for almost 90% of the largest model's weights.
# the cost trade-off of row (v): step times from Table 3
round(1.74 / 1.02, 2)  # → 1.71
# overall score lost by inserting a gated layer every 4th block
round(70.7 - 68.8, 1)  # → 1.9

Why it matters today

“Freeze the big pretrained models, train small new connectors” became a default for vision-language models because of results like row (viii): it was better as well as cheaper. The forgetting measure in the fine-tuning lesson quantifies the damage that unfreezing can do.

4 Related work · original

Everyday picture

Flamingo sits where three roads meet. Language models that learn from prompts (GPT-3, and Chinchilla, whose 70B model it builds on). Vision-language models, some trained BERT-style and fine-tuned per task, some contrastive like CLIP and ALIGN, some already generating text. And web-scale datasets scraped instead of annotated.

Tiny example

Ways to adapt a language model to a new task, as the paper lists them: small adapter modules, fine-tuning a small part, showing in-context examples, or optimising the prompt by gradient descent. Flamingo takes the third.

What is new

  • Earlier work already froze pretrained language models to avoid forgetting; Flamingo adds learnable layers within the frozen stack.
  • It claims the first language model to take arbitrarily interleaved images, video and text.
  • It shows training on whole web pages, pictures in place, matters, alongside paired image-text data. The concurrent CM3 also trained on web pages but generated HTML; Flamingo only generates plain text.

Why it matters

The contrastive road (see the CLIP companion) supplied Flamingo's eyes; the language-model road supplied its voice. The paper's contribution is the joint between them.

5 Discussion · original

“First, our models build on pretrained LMs, and as a side effect, directly inherit their weaknesses.”Alayrac et al. (2022), §5

Everyday picture

Borrow a brilliant writer and you also borrow their habits, good and bad: fluent guesses when unsure, and trouble with texts much longer than they are used to.

Tiny example: when prompts get too long

A VisDial example is an image followed by a dialogue of 21 sentences. Thirty-two of them make at least 32 × 21 = 672 sentences, a prompt of 4,096 to 8,192 tokens, while the language models were trained on sequences of at most 2,048. So VisDial results stop at 16 shots.

# 32 VisDial shots, 21 sentences each
32 * 21  # → 672
# how far past the training length the shortest such prompt goes
4096 / 2048  # → 2.0

The limitations the paper names

  • Inherited from the language model: occasional hallucinations and ungrounded guesses, poor generalisation to sequences longer than in training, poor sample efficiency.
  • Classification: it trails contrastive models, including its own vision encoder, because they optimise image-text matching directly.
  • In-context learning: sensitive to the order and format of examples; its cost grows linearly with the number of shots if the prompt's keys and values are cached and reused, and quadratically otherwise; and the gains flatten beyond about 32 shots.

The paper also discusses societal impacts: it inherits the language model's risks (offensive language, social biases, leaking private information), adds visual ones such as biases about the people in images, and reports early bias and toxicity checks on COCO captions.

Why it matters today

Every item on the list is still an active problem for vision-language models, and the “ungrounded guess”, confidently describing something that is not in the picture, is the one users meet most often.

Appendices, briefly · original

Everyday picture

The appendices are the workshop manual: the exact sizes, the training settings, the evaluation details, and a model card.

Training details worth knowing

  • Optimiser: AdamW, gradient clipping at norm 1, weight decay 0.1 (none on the resampler), learning rate warmed up linearly to 10−4 over 5,000 steps and then held constant; 500,000 steps.
  • Images: resized to 320 × 320, higher than the 288 × 288 used for the vision encoder's contrastive pretraining, which is affordable because nothing is back-propagated through the frozen encoder.
  • Video: 8 frames at 1 frame per second in training, 30 frames at 3 per second at test time, made possible by interpolating the learned time embeddings.
  • A tokenizer quirk: half the time, a space is prepended to paired captions, because the tokenizer splits a word differently with and without a leading space; the paper reports a substantial improvement.
  • Hardware: Flamingo-80B trained on 1,536 TPUv4 chips for 15 days, split 16 ways with Megatron-style model parallelism and ZeRO stage 1 for the optimiser state (see the ZeRO companion), trained weights in float32 and activations in bfloat16 (mixed precision).
# chip-days of the largest training run
1536 * 15  # → 23040

Two numbers that disagree

Appendix B.1.1 says the gated layers add 1.4B parameters to Flamingo-3B and 1.8B to Flamingo-9B; its Table 5 lists 1.2B and 1.6B. The table's figures add up to the stated totals (3.2B and 9.3B) and the text's do not, so this page uses the table's. Similarly, Appendix B.2.2 gives VQAv2 at 32 shots as 67.3%, where Tables 1 and 2 give 67.6.

# Flamingo-3B's total with the table's 1.2B, and with the text's 1.4B
round(1.4 + 0.435 + 1.2 + 0.194, 1), round(1.4 + 0.435 + 1.4 + 0.194, 1)  # → (3.2, 3.4)

Why it matters

The details that make a training run work (warmup, clipping, precision, sharding) are the same for a vision-language model as for a language model; the pretraining lesson builds them.

What changed since 2022

Everyday picture

Flamingo proved that a language model could learn to see without forgetting how to write. The field then split over how to plumb the connection.

DevelopmentWhat it changedBuild it or read it
Splice instead of cross-attendproject the image into a few hundred tokens and put them straight into the prompt, with no new layers inside the language model; simpler, and the default for many open modelsLLaVA companion, VisionLanguageModel
Open reproductionsOpenFlamingo rebuilt the design on the open LLaMA modelLLaVA companion, §2
Instruction datatuning on image conversations, not only captions and web pages, made models follow requests about picturesLLaVA companion
Transformer image encodersa CLIP-trained ViT became the usual eye, where Flamingo used a convolutional NFNetViT companion, CLIP companion

Why it matters

Whichever plumbing a model uses, Flamingo's lessons stand: keep the pretrained parts' knowledge intact, compress what the eye sends, train on documents where pictures and words interleave, and evaluate on benchmarks you did not tune on.

Glossary