An annotated companion · AI Primer

Mamba, annotated

About this page. This is a companion, not a copy. It follows the paper section by section, quotes only a sentence or two per section (clearly marked), and explains everything in its own words. The paper is released under the Creative Commons Attribution 4.0 licence; selected rows of its tables are reproduced here with attribution, its charts are replotted from those published numbers, and its diagrams are redrawn from scratch. Equations are reproduced with every symbol decoded; a few formulas that write out, in symbols, something the paper says in words are labelled as this page's own. Read the original alongside: every section links to it.

How to read this page

  • Any dotted word explains itself when you hover it, tab to it, or tap it.
  • Every symbol inside an equation does the same, and each equation comes with a table decoding it, a sentence reading it aloud, the numbers worked through, and the same numbers in Python.
  • The diagrams are live: tap any part to see what it holds and what shape it has. The sliders let you turn the one dial that makes Mamba selective.

Each idea climbs the same ladder: everyday picture, tiny example, diagram, the math, why it matters. The efficient architectures lesson builds every piece in NumPy: the recurrence, the convolution, the selective step size and the parallel scan. The tiny examples here reuse its numbers, so you can run them there.

Abstract · original

“We identify that a key weakness of such models is their inability to perform content-based reasoning, and make several improvements.”Gu & Dao (2023), Abstract

Everyday picture

A Transformer reading a long document keeps every page on the desk, so it can look anything up, but the desk (the KV cache) grows with every page and each new page is compared with all the old ones. A recurrent model reads with a single notebook page of fixed size: cheap, but it has to decide what to write down. Older fast models of this kind used the same note-taking rule for every word. Mamba lets each word decide how much to write and how much to erase.

What the paper claims

  • Making the parameters of a state space model depend on the current token (a selection mechanism) fixes their weakness on discrete data such as text.
  • That change rules out the fast convolution trick, so the authors write a hardware-aware parallel scan that computes the recurrence quickly on GPUs.
  • The resulting architecture, Mamba, has no attention and not even separate MLP blocks. It runs in time linear in sequence length, has about 5× higher generation throughput than a Transformer of similar size, and keeps improving on real data up to sequences a million long.
  • On language, Mamba-3B beats Transformers of the same size and matches Transformers twice its size.

Why it matters today

Mamba made recurrent models a serious alternative to attention for language, and it is the reason the phrase “state space model” now appears next to “Transformer” in discussions of long context. Its successor, Mamba-2, and hybrid models that mix a few attention layers with many Mamba layers, build directly on this paper; the efficient architectures lesson works out their memory.

1 Introduction · original

“The efficacy of self-attention is attributed to its ability to route information densely within a context window, allowing it to model complex data.”Gu & Dao (2023), §1

Everyday picture

Attention's strength and its cost are the same thing: every token can consult every other token directly. That means nothing outside the window can be seen, and the work grows with the square of the window. Many cheaper variants of attention exist, but, the authors note, none had been shown to work well at scale across domains. Structured state space models (SSMs) were a promising different road: they compute either as a recurrence or as a convolution, in time linear or nearly linear in length; they dominated long-range benchmarks such as the Long Range Arena and did well on continuous signals such as audio and vision. On text they lagged.

The three contributions

  • Selection. Let the SSM's parameters depend on the input, so the model can ignore irrelevant tokens and keep relevant ones indefinitely.
  • A hardware-aware algorithm. Compute the now input-dependent recurrence with a parallel scan, keeping the large expanded state in the GPU's fast on-chip memory. It is faster than the previous convolution-based SSMs both in theory (linear instead of nearly linear in length) and in practice, up to 3× on A100 GPUs.
  • A simpler architecture. Merge the usual SSM block and the Transformer's MLP block into one block, the Mamba block, and stack it.

Diagram: the overview (Figure 1)

GPU SRAM: small, fast GPU HBM: large, slower × Ā (keep) x_t × B (write) × C input x_tD = 5 channels Selectioncomputed from x_t Δ_t, B_t, C_t Discretize → Ā, B h_(t−1) h_t output y_t (D = 5)

Tap a part. Start with input x_t at the top left, then follow it into the state grid h_t.

Figure 1 of the paper, redrawn with the paper's example sizes, D = 5 channels and N = 4 state numbers per channel. Based on Gu & Dao (2023), Figure 1 (CC BY 4.0).

Reading it: one time step, top to bottom. The input xt has D = 5 channels. The selection box reads the input and produces this step's Δ, B and C; discretizing turns Δ (with the fixed A) into this step's keep factor Ā and write factor B̄. The two grids are the hidden state before and after the step: every channel has its own N = 4 numbers, so the state is 5 × 4, four times bigger than the input. The new state is the old one times Ā plus the input times B̄, and the output is read off with C, back to 5 numbers. The dashed frames are the GPU's two memories: the expanded state lives only in the small fast one (SRAM), while inputs and outputs travel to and from the large main memory (HBM). Older SSMs could skip building the big state because their parameters never changed with the input; selection brings the parameters back to life, and the memory trick in §3.3 keeps it affordable.

Why it matters

The authors list three properties that make selective SSMs a candidate backbone for foundation models: high quality on dense data like language and genomics; training cost linear in length and generation at constant time per step, since there is no cache of previous tokens; and quality that keeps improving with context up to a million tokens.

2 State space models · original

Everyday picture

A cup of tea cooling on a desk. Every minute it keeps some fraction of its heat and gains whatever hot water you pour in. Its temperature is a running summary of every pour, with older pours counting less. A state space model is that cup, many times over: a small state that fades a little each step and takes in the new input, read out as the output. Because the update is only multiply-and-add, it can be computed two ways, one step at a time or all at once.

The continuous system (equation 1) · original

Tiny example

One channel, a state of one number. Take A = −1 and B = 1. If the state is h = 2 and no input arrives (x = 0), the state is changing at rate −2: draining away. If an input x = 3 arrives, the rate is −2 + 3 = +1: filling up. The negative A is what makes the memory fade.

In words: “the state's rate of change is a multiple of the state itself (the fading) plus a multiple of the input (the pouring); the output is a fixed reading of the state.”

With the numbers: h′ = (−1)(2) + (1)(0) = −2 with no input, and (−1)(2) + (1)(3) = 1 with x = 3.

In Python:

A, B = -1.0, 1.0
h = 2.0
# h'(t) = A h(t) + B x(t), with no input and then with x = 3
[A * h + B * x for x in (0.0, 3.0)]  # → [-2.0, 1.0]

Discretization (equation 4) · original

Everyday picture

Equation 1 describes the tea continuously; a computer works in steps. Discretization answers: if the input holds steady for a step of length Δ, how much of the old state survives the step, and how much of the input gets in? The paper uses the zero-order hold rule, which assumes exactly that: the input is held constant through the step.

Tiny example

Same A = −1, B = 1, and a step Δ = 0.5. The state keeps e−0.5 = 0.607 of itself and takes in 0.393 of the input. A long step Δ = 5 keeps only 0.0067 and takes in 0.9933: the old state is essentially replaced. A tiny step keeps nearly everything and takes in nearly nothing. With A = −1 and B = 1 the two amounts always add up to 1.

In words: “the keep factor is e raised to the step length times the fade rate; the write factor is how much input accumulates over that step, which for a single number is (eΔA − 1) / A times B.”

With the numbers: Δ = 0.5: Ā = e−0.5 = 0.607, B̄ = (−0.5)−1 × (0.607 − 1) × 0.5 = (−2) × (−0.393) × 0.5 = 0.393. Δ = 5: Ā = 0.0067, B̄ = 0.9933.

In Python:

import math
A, B = -1.0, 1.0
def zoh(delta):
    # Ā = exp(ΔA)
    A_bar = math.exp(delta * A)
    # B̄ = (ΔA)^-1 (exp(ΔA) − 1) ΔB, one number instead of a matrix
    B_bar = (1 / (delta * A)) * (A_bar - 1) * delta * B
    return round(A_bar, 4), round(B_bar, 4)
zoh(0.5)  # → (0.6065, 0.3935)
zoh(5.0)  # → (0.0067, 0.9933)
(log scale, A = −1, B = 1)

Reading it: the slider sets the step size Δ on a log scale, from 0.001 to 10. The solid bar is Ā, how much of the old state survives the step; the striped bar is B̄, how much of this input is written. Slide left and the state is kept almost whole while the input barely registers: the step is so short that the input is a blip. Slide right and the state is wiped and replaced by the input. In an ordinary SSM, Δ is one learned number per channel, fixed for every token. Mamba's whole idea (§3) is to compute Δ from each token, so the model can slide this dial differently for every word. The lesson's discretize computes the same two numbers.

Why it matters

The paper notes that discretization links SSMs to continuous-time systems (for instance, the model can be applied at different resolutions) and to the gates of recurrent networks, a link §3.5 makes exact. Mechanically, it is just the first step of the forward pass.

Two ways to compute: recurrence and convolution (equations 2 and 3) · original

Tiny example

The state keeps half of itself each step and takes the input whole: Ā = 0.5, B̄ = 1, C = 1. Inputs x = (1, 0, 0, 2). Step by step the state is 1, then 0.5, then 0.25, then 0.5 × 0.25 + 2 = 2.125. Unrolled, the last output is every past input weighted by how much it has faded: 1 × 2 + 0.5 × 0 + 0.25 × 0 + 0.125 × 1 = 2.125. The same answer, with no loop.

In words: “each step, fade the state and add the new input, then read it out; equivalently, the output is the whole input sequence slid against a fixed list of weights, where the weight for an input k steps ago is C times Ā to the k times B̄.”

With the numbers: the loop gives y = (1, 0.5, 0.25, 2.125). The kernel is K̄ = (1, 0.5, 0.25, 0.125), and y4 = 1 × 2 + 0.5 × 0 + 0.25 × 0 + 0.125 × 1 = 2.125.

In Python:

A_bar, B_bar, C = 0.5, 1.0, 1.0
x = [1.0, 0.0, 0.0, 2.0]
# the recurrence, equation 2
h, y = 0.0, []
for x_t in x:
    h = A_bar * h + B_bar * x_t
    y.append(C * h)
y  # → [1.0, 0.5, 0.25, 2.125]
# the kernel K̄, equation 3a: C Ā^k B̄ for k = 0, 1, 2, 3
K = [C * A_bar ** k * B_bar for k in range(4)]
K  # → [1.0, 0.5, 0.25, 0.125]
# y_4 by convolution, equation 3b: newest input times K̄_0, and so on back
sum(K[k] * x[3 - k] for k in range(4))  # → 2.125

Reading it: the two equations are two roads to the same outputs. The recurrence is how the model generates: one token at a time, carrying a fixed-size state. The convolution is how older SSMs trained: the whole input is known in advance, the kernel is computed once, and every output is a weighted sum computed in parallel. The lesson builds both roads as ssm_recurrent and ssm_convolution, with the kernel from ssm_kernel, and checks they agree.

Time invariance, structure and size · original

Everyday picture

The convolution road has a catch: it only exists if the fading rule is the same at every step. If the cup of tea cooled faster on some minutes than others, there would be no single list of weights to slide along the input. The paper calls the “same rule at every step” property linear time invariance (LTI), and every structured SSM before Mamba had it, for efficiency. The paper's core insight is that LTI is also their weakness.

How big the state is

Structured SSMs use a diagonal A, so A, B and C are each just N numbers per channel. The model runs the same SSM independently on each of D channels, so the state for one token is D × N numbers. Over a batch of B sequences of length L, the full state is B × L × D × N numbers, N times more than the input and output (B × L × D). Computing all of it costs O(BLDN) time and memory, which the paper calls “the root of the fundamental efficiency bottleneck” that §3.3 attacks.

In words: “the full state has one number for every sequence, every position, every channel and every state slot.” (This page's notation for the paper's shape (B, L, D, N).)

With the numbers: the paper's speed benchmark uses D = 1,024 and N = 16. For one sequence of 2,048 tokens, the input is 1 × 2,048 × 1,024 = 2,097,152 numbers and the full state is 16 times that, 33,554,432.

In Python:

B, L, D, N = 1, 2048, 1024, 16
# the input and output: (B, L, D)
print(f"{B * L * D:,}")  # → 2,097,152
# the full state: (B, L, D, N)
print(f"{B * L * D * N:,}")  # → 33,554,432

SSM architectures the paper compares against

  • Linear attention: an approximation of self-attention with a recurrence that the paper views as a degenerate linear SSM.
  • H3: an S4 layer sandwiched between two gated connections, plus a short local convolution. It is the template for most SSM architectures.
  • Hyena: H3's architecture with S4 replaced by a long convolution whose weights come from a small network.
  • RetNet: an extra gate and a simpler SSM, computed in parallel with a variant of multi-head attention.
  • RWKV: a recurrent language model built on another linear-attention approximation; its main mechanism involves LTI recurrences.

Why it matters

“State space model” means many things in many fields (Kalman filters, hidden Markov models); in this paper it means only the structured, deep-learning kind. And the size equation is the whole engineering story: the state is where the model's memory lives, so a bigger N means a better memory, and N times the input is too big to write to GPU main memory for every token.

3 Selective state space models · original

3.1 Motivation: selection as a means of compression · original

“We argue that a fundamental problem of sequence modeling is compressing context into a smaller state.”Gu & Dao (2023), §3.1

Everyday picture

Attention is effective and expensive for the same reason: it doesn't compress at all. It keeps the whole context (the KV cache), which makes generation slow and training quadratic. A recurrent model is cheap because its state is fixed in size, and only as good as what it chose to put there. So the question is not “recurrence or attention?” but “how well does the state compress the context?”. The authors' answer is selectivity: the ability to focus on some inputs and filter out others, depending on what they are.

Diagram: the two tasks that expose LTI models (Figure 2)

Tap a row. Coloured squares are tokens to remember, blank squares are filler, and the squares after the bar are what the model must output.

Figure 2 of the paper, redrawn with this page's own token sequences. Based on Gu & Dao (2023), Figure 2 (CC BY 4.0).

Reading it: each row is a sequence read left to right; everything after the vertical bar is what the model must produce. In copying, the four tokens to remember always sit at the start and the gap before the answer is always the same length, so a model only needs to track time: “output what arrived exactly 12 steps ago” is a fixed convolution kernel, and LTI models solve it. In selective copying the same four tokens are scattered among filler at random positions, so the model has to look at content: remember the coloured ones, skip the blanks. No fixed kernel can do that, because the spacing changes every time. In induction heads, having seen “H then P” once, when H appears again the model must answer P: recall by association, which the paper calls a key ability of large language models.

Try it: selective against fixed step sizes

A tiny version of selective copying from the lesson: twelve inputs, one of which (a 7, in third place) is marked as worth remembering, the rest small noise. A one-channel SSM must still hold the 7 at the end. The selective model computes Δ from the marker (large Δ, so write, on the 7; tiny Δ, so keep, everywhere else). The fixed model uses one Δ for every token; choose it with the slider.

Hover or tap the chart to read the state at each position.

Reading it: the x-axis is the position in the sequence, the y-axis the value of the state after each token. The dotted line is the input itself, spiking to 7 at position 2. The solid line, the selective model, jumps to about 7 at the marked token and then barely moves, ending at 6.55, because every later token has a tiny Δ. The dashed line is the fixed Δ from the slider. It starts at 0.096, the best any fixed setting manages here, ending at only 0.26: a Δ small enough to protect the 7 from later noise is also too small to write the 7 in the first place. Slide it up and the 7 gets written, then washed out by the next nine tokens. Only input-dependent Δ can both write the important token and ignore the rest. The lesson computes these runs in selective_recall and time_invariant_recall.

Why it matters

The paper states the trade-off plainly: efficient models must have a small state, while effective models must have a state that contains everything needed from the context. Selection is how a small state can hold the right things.

3.2 Improving SSMs with selection · original

Everyday picture

A note-taker with a dial. For filler words (“um”, “so”) they barely touch their notes; for a name or a number they wipe the relevant line and write the new fact. An LTI SSM has the dial glued in one position. Mamba reads the dial setting off each token.

What changes, shape by shape (Algorithms 1 and 2)

Shapes before and after selection, from Gu & Dao (2023), Algorithms 1 and 2 (CC BY 4.0)
ParameterS4: fixedS6: selective
A(D, N)(D, N)
B = sB(x)(D, N)(B, L, N)
C = sC(x)(D, N)(B, L, N)
Δ(D)(B, L, D)
Ā, B̄(D, N)(B, L, D, N)
Computed asrecurrence or convolutionrecurrence (scan) only

In the shapes, B is the batch size, L the length, D the channels and N the state size (the paper reuses the letter B for batch and for the B parameter).

The highlighted rows are the change: B, C and Δ gain a length dimension L, one value per token. That is exactly what “time-varying” means, and it is why the convolution road closes. The paper chooses sB(x) = LinearN(x) and sC(x) = LinearN(x), where Lineard is a learned projection to d numbers, and sΔ(x) = BroadcastD(Linear1(x)): the input is squeezed to one number, then copied to all D channels, so that when a token should be ignored, every channel ignores it together.

Tiny example

From the lesson: suppose the projection gives 10 × marker − 5, so a marked token gets 5 and an ordinary one gets −5. Softplus turns these into step sizes: 5.007 for the marked token (write) and 0.0067 for an ordinary one (keep). Softplus is used because a step size must be positive.

In words: “add a learned offset to a one-number summary of this token, and pass the sum through softplus, a smooth ramp that is always positive; the result is this token's step size.”

With the numbers: z = 5: ln(1 + e5) = ln(149.4) = 5.007. z = −5: ln(1 + e−5) = ln(1.0067) = 0.0067.

In Python:

import math
def softplus(z):
    # ln(1 + e^z)
    return math.log(1 + math.exp(z))
# z = Parameter + s_Δ(x): 10 × marker − 5
[round(softplus(10 * marker - 5), 4) for marker in (1, 0)]  # → [5.0067, 0.0067]

Why it matters

This is the paper's entire modelling change: three small projections of the input. The lesson's step_sizes is this equation with a hand-set projection, and selective_ssm runs the per-token recurrence. The cost is that every token now has its own Ā and B̄, so the fixed kernel of equation 3 no longer exists; §3.3 is about paying that cost.

3.3 Efficient implementation: the hardware-aware scan · original

Everyday picture

A chef with a tiny counter next to the stove and a big pantry down the hall. Fetching from the pantry is the slow part. A sensible chef carries the few raw ingredients to the counter once, does all the chopping and mixing there, and walks only the finished dish back. The GPU is that kitchen: SRAM is the counter, HBM the pantry, and most operations other than matrix multiplication are limited by the walking (memory bandwidth), not the cooking.

Why earlier SSMs used convolutions (§3.3.1)

Recurrent models trade expressivity against speed: a bigger state remembers more but costs more. The recurrence is the more flexible road, but it builds the state of shape (B, L, D, N), N times bigger than the input. The convolution road skips the state entirely and only builds a kernel of shape (B, L, D). That trick is what let earlier LTI SSMs have states N ≈ 10 to 100 times larger than a traditional RNN's for free. Selection takes the trick away.

The fix, in three classic techniques (§3.3.2)

  • Kernel fusion. Instead of writing Ā and B̄, size (B, L, D, N), to HBM, read only Δ, A, B and C into SRAM, discretize and run the recurrence there, and write only the output, size (B, L, D), back.
  • Parallel scan. The recurrence looks sequential, but a work-efficient parallel scan computes all its states in about log₂ L rounds.
  • Recomputation. The intermediate states are needed for the backward pass, but instead of storing them the kernel recomputes them when it reloads the inputs (the same idea as activation checkpointing). The fused layer ends up needing the same memory as an optimized Transformer with FlashAttention.

Tiny example: how much walking is saved

The paper counts memory traffic in O(·) terms: O(BLDN) for the plain scan, O(BLD + DN) for the fused one, a saving of a factor of about N. With the benchmark's sizes, one sequence of 2,048 tokens, D = 1,024 and N = 16, those counts are 33.6 million numbers against 2.1 million: about 16 times less traffic, which the paper reports as a 20 to 40× speedup in practice.

In words: “the plain scan moves the whole expanded state through main memory; the fused scan moves only the inputs and outputs, plus the small fixed A.” (This page's counting of the paper's O(BLDN) and O(BLD + DN), with constants dropped.)

With the numbers: 1 × 2,048 × 1,024 × 16 = 33,554,432 against 2,097,152 + 16,384 = 2,113,536: a ratio of 15.9, close to N = 16.

In Python:

B, L, D, N = 1, 2048, 1024, 16
io_plain = B * L * D * N
io_fused = B * L * D + D * N
print(f"{io_plain:,} {io_fused:,}")  # → 33,554,432 2,113,536
round(io_plain / io_fused, 1)  # → 15.9

The scan, in one equation

Why can a recurrence run in parallel at all? Each step is “multiply the state by a, then add b”. Two such steps in a row are again one step of the same form, so steps can be merged in pairs, then pairs of pairs, like a knockout tournament. The paper cites the parallel scan without writing it out; this is the merge rule the lesson's parallel_scan uses.

In words: “doing step 1 and then step 2 is the same as one step that multiplies by both, and adds step 1's addition faded by step 2, plus step 2's own.”

With the numbers: in the tea example, steps 1–2 merge to (0.25, 0.5) and steps 3–4 to (0.25, 2). Merging those: (0.25 × 0.25, 0.25 × 0.5 + 2) = (0.0625, 2.125), and 2.125 is the final state from the loop in §2.

In Python:

def merge(first, then):
    a1, b1 = first
    a2, b2 = then
    return (a1 * a2, a2 * b1 + b2)
# the four tea steps: h becomes 0.5 h + x_t
s = [(0.5, 1.0), (0.5, 0.0), (0.5, 0.0), (0.5, 2.0)]
merge(merge(s[0], s[1]), merge(s[2], s[3]))  # → (0.0625, 2.125)

Why it matters

This is the same lesson FlashAttention taught for attention, by an overlapping set of authors: the bottleneck is memory traffic, not arithmetic, so the fastest algorithm is the one that keeps its big intermediate results on the chip. The hardware lesson explains the memory hierarchy behind it.

3.4 A simplified SSM architecture · original

Everyday picture

Earlier SSM networks alternated two kinds of floor in the building: an H3 floor (the sequence mixer) and an MLP floor (per-token processing), like a Transformer's attention and feed-forward layers. Mamba builds one kind of floor that does both, and stacks it.

input (D) output (D) Linear ↑ Linear ↑ Conv σ (SiLU) Selective SSM σ (SiLU) × Linear ↓

Tap a part. Start at the bottom with the two Linear ↑ projections, and follow each branch up.

The Mamba block from Figure 3 of the paper, redrawn. The optional normalization layer is left out. Based on Gu & Dao (2023), Figure 3 and §3.4 (CC BY 4.0).

Reading it: data enters at the bottom with D numbers per token and is widened twice, into two branches of E × D numbers each (E = 2). The left branch is the sequence mixer: a short convolution, the SiLU activation, then the selective SSM, the only place where tokens exchange information. The right branch is a gate: just SiLU, applied per token. The two meet at ×, where the gate scales the SSM's output number by number, and the final Linear narrows back to D. Remove the left branch's Conv and SSM and what remains is a gated MLP of the SwiGLU kind; the paper describes Mamba as exactly that block with a conv → SSM path added. Compared with H3, Mamba replaces H3's first multiplicative gate with an activation function.

The parameter count

In words: “almost all of a block's weights are in its three big projections: two that widen D numbers to E × D, and one that narrows E × D back to D; the SSM itself adds comparatively little.”

With the numbers: E = 2 gives 6D² per block. With D = 1,024 that is 6,291,456; two blocks make 12D² = 12,582,912, the same as one Transformer layer's attention (4D²) plus MLP (8D²). That is why the paper stacks two Mamba blocks for every Transformer layer it compares against.

In Python:

E, D = 2, 1024
# 2ED² to widen (two branches) + ED² to narrow
per_block = 2 * E * D ** 2 + E * D ** 2
print(f"{per_block:,}")  # → 6,291,456
# two Mamba blocks against one Transformer layer, 12D²
2 * per_block == 12 * D ** 2  # → True

Why it matters

One homogeneous block is simpler to build, tune and scale. The ablation in §4.6 (Table 2) finds the Mamba block performs about the same as H3 while being simpler, and slightly better with a selective layer inside.

3.5 Properties of selection: it is gating, made principled · original

“We highlight the most important connection: the classical gating mechanism of RNNs is an instance of our selection mechanism for SSMs.”Gu & Dao (2023), §3.5.1

Everyday picture

An LSTM or GRU has gates: dials between 0 and 1 that decide how much of the old memory to keep and how much new input to let in. They were invented by hand. Theorem 1 shows that a selective SSM in its simplest setting is such a gate, derived from discretization rather than invented.

Theorem 1

Take one state number (N = 1), A = −1, B = 1, Δ = softplus(Linear(x)). Then the selective recurrence becomes a gated update:

In words: “compute a gate between 0 and 1 from the current input; the new state is a blend of the old state and the input, with the gate deciding the mix.”

With the numbers: why this is the same thing: with A = −1, Ā = e−Δ, and e−softplus(z) = 1 / (1 + ez) = 1 − σ(z). Take z = Linear(xt) = 2. Then Δ = softplus(2) = 2.127 and Ā = e−2.127 = 0.119, while the gate is g = σ(2) = 0.881 and 1 − g = 0.119. The same keep factor, reached two ways. And B̄ = 1 − Ā = 0.881 = g.

In Python:

import math
z = 2.0
# the SSM road: Δ = softplus(z), Ā = e^(ΔA) with A = −1
delta = math.log(1 + math.exp(z))
A_bar = math.exp(-delta)
round(delta, 3), round(A_bar, 3)  # → (2.127, 0.119)
# the gate road: g = σ(z), keep 1 − g
g = 1 / (1 + math.exp(-z))
round(g, 3), round(1 - g, 3)  # → (0.881, 0.119)

Reading it: the slider is z = Linear(xt), what the model computes from the current token. The first two bars come from the SSM road (step size Δ, then keep factor Ā = e−Δ); the last two from the gate road (g = σ(z), and 1 − g). Drag anywhere: the keep bar and the 1 − g bar are always identical, which is Theorem 1. At very negative z the gate is shut, Δ → 0, and the token is ignored; at very positive z the gate is open, Δ is large, and the state is reset to this token. That is the paper's reading of Δ: large Δ focuses on the current input and forgets the state; small Δ keeps the state and ignores the input.

What selection buys, mechanically (§3.5.2)

  • Variable spacing. Filler between the important tokens (the “um”s of language) can be skipped, whatever its length.
  • Filtering context. Many sequence models don't get better with longer context, because they can't ignore what's irrelevant. A selective model can, so in principle its performance only improves with more context (§4.3 tests this).
  • Boundary resetting. When independent documents are packed into one sequence, an LTI model bleeds information across the boundary; a selective one can reset its state there (Δ → ∞, or g → 1).

And each parameter's role: Δ balances focusing on the current input against keeping the state; A matters only through ΔA, so making Δ selective already makes Ā selective, and A is left fixed for simplicity; selective B and C give finer control over whether an input enters the state and whether the state reaches the output.

3.6 Additional model details · original

  • Real, not complex. Earlier SSMs used complex numbers in the state, which helps on continuous signals. Mamba uses real numbers by default, and they work well on everything the paper tries except one audio task. The authors' guess: complex numbers help continuous modalities (audio, video) but not discrete ones (text, DNA).
  • Initialization. The n-th element of A starts at −(n + 1) in the real case (called S4D-Real) or −1/2 + ni in the complex case (S4D-Lin).
  • The Δ projection can project to R numbers instead of 1, a small fraction of D, for a negligible number of extra parameters.

Tiny example: where Δ starts

Δ's learned offset (the “Parameter”) is initialized so that softplus of it lands uniformly between 0.001 and 0.1. To get there you need the inverse of softplus. For a target step of 0.01, the offset is ln(e0.01 − 1) = −4.600; the range 0.001 to 0.1 maps to offsets from −6.907 to −2.252. Small initial steps mean the model starts out keeping its state and learns which tokens deserve a large Δ.

In words: “pick a starting step size at random between 0.001 and 0.1, and set the offset to whatever number softplus turns into that step size.”

With the numbers: u = 0.01 gives ln(0.01005) = −4.600; the ends of the range give −6.907 and −2.252.

In Python:

import math
def softplus_inverse(u):
    # τ_Δ^-1(u) = ln(e^u − 1)
    return math.log(math.exp(u) - 1)
[round(softplus_inverse(u), 3) for u in (0.001, 0.01, 0.1)]  # → [-6.907, -4.6, -2.252]
# check: softplus undoes it
round(math.log(1 + math.exp(softplus_inverse(0.01))), 6)  # → 0.01

The paper calls a selective SSM computed with a scan S6 for short: an S4 model plus selection plus a scan.

4 Empirical evaluation · original

Everyday picture

A new engine is tested on a track built to expose its old weakness (synthetic tasks), then on three real roads (language, DNA, audio), then timed on a dynamometer (speed and memory), and finally taken apart to see which part matters (ablations).

4.1 Synthetic tasks · original

Selective copying

Sequences of length 4,096 with 16 possible tokens, 16 of which must be remembered, and 2-layer models of width 64. The table pairs an architecture (the block around the sequence layer) with an inner layer. Architectural gating (H3's or Mamba's multiplications) helps a little; swapping the layer for a selective one (S6) solves the task.

Selective copying accuracy (%), from Gu & Dao (2023), Figure 4 table (CC BY 4.0)
ArchitectureLayerAccuracy
No gateS418.3
No gateS697.0
H3S457.0
H3Hyena30.1
H3S699.7
MambaS456.4
MambaHyena28.4
MambaS699.8

The paper argues that architectural gating is not selection: a multiplication applied to each token separately cannot change the spacing between tokens, which is what this task varies.

Induction heads

The induction head task: having seen “Harry Potter” earlier in a sequence, predict “Potter” the next time “Harry” appears. Models with 2 layers train on sequences of 256 tokens (a vocabulary of 16) and are tested on lengths from 64 up to 1,048,576.

Hover or tap the chart to read each model's accuracy at a test length.

Reading it: the x-axis is the test sequence length on a log scale, from 64 to 1,048,576; the models trained at 256, just right of the 128 label. The y-axis is accuracy. All three models score 100% at the length they trained on. Moving right, the attention model (multi-head attention with xPos positions, the best attention variant here) holds up at twice the length (99.6%), falls to 67.6% at four times and 25.4% at eight, then to about 7% to 9%; its line stops at 16,384 because longer sequences ran out of memory. H3, an LTI SSM, falls similarly. Mamba's line stays flat at 100% all the way to 1,048,576 tokens, 4,096 times longer than anything it trained on, with 74 thousand parameters. Values are from the paper's Table 3, where a check mark means perfect accuracy (drawn here as 100).

4.2 Language modeling · original

Scaling laws

Models from about 125 million to 1.3 billion parameters, trained on the Pile following the Chinchilla protocol, compared on perplexity. The baselines include the GPT-3 Transformer and a stronger recipe the paper calls Transformer++ (rotary positions, SwiGLU, RMSNorm, no bias terms, higher learning rates), plus H3, Hyena, RWKV and RetNet. Mamba is the first attention-free model to match Transformer++, and its advantage grows with sequence length (Figure 6).

Zero-shot evaluations (Table 1)

Models trained on 300 billion tokens of the Pile, scored on six common-sense tasks without any task-specific training (zero-shot). Pythia and RWKV used the same tokenizer, data and training length.

Selected rows of Table 1, from Gu & Dao (2023) (CC BY 4.0). Pile ppl is perplexity on the Pile validation set (lower is better); the average is over six accuracy columns (higher is better)
ModelPile pplLAMBADA accHellaSwagAverage acc
Pythia-1.4B7.5161.752.155.2
RWKV-1.5B7.7056.452.554.3
Mamba-1.4B6.8064.959.159.7
Pythia-2.8B6.7364.759.359.1
RWKV-3B7.0063.959.659.6
Mamba-2.8B6.2269.266.163.3
Pythia-6.9B6.5167.164.061.7
RWKV-7.4B6.3167.265.562.5

Reading it: each bar is one model's average accuracy over the six tasks, on an axis from 50% to 65% so the gaps are visible; Mamba's bars are solid, the others striped. Mamba-1.4B beats both models of its size and even edges past Pythia-2.8B. Mamba-2.8B, the paper's “Mamba-3B”, is about 4 points above Pythia-2.8B (63.3 against 59.1) and above both the 6.9B Pythia and the 7.4B RWKV. That is the abstract's “matches Transformers twice its size”.

4.3 DNA modeling · original

Everyday picture

DNA is a text with a four-letter alphabet and very long-range structure. The models are pretrained to predict the next base pair on one human genome (HG38, about 4.5 billion tokens).

  • Model size (Figure 7, left): at context 1,024, from about 200 thousand to 40 million parameters, Mamba scales better than HyenaDNA and Transformer++; at the largest size it matches them with roughly 3 to 4 times fewer parameters.
  • Context length (Figure 7, right): with model size fixed (6 layers, width 128, about 1.4 million parameters) and training tokens fixed, Mamba's perplexity keeps improving as context grows to a million base pairs, while HyenaDNA gets worse. The paper's explanation is §3.5's “filtering context”: a long fixed convolution averages in more and more noise.

Great apes classification

Pretrained models are fine-tuned to tell which of five great apes (human, chimpanzee, gorilla, orangutan, bonobo, which share 99% of their DNA) a random stretch of DNA came from. Random guessing scores 20%.

Hover or tap the chart to read accuracy at each context length.

Reading it: the x-axis is the context length the model was pretrained and fine-tuned on, from 1,024 to 1,048,576 base pairs (log scale); the y-axis is accuracy. At short contexts every model scores around 30%, only a little above the 20% of guessing, because a thousand base pairs barely distinguish species sharing 99% of their DNA. With more context, both Mamba models climb, reaching 71.7% (1.4 million parameters) and 81.3% (7 million) at a million base pairs, against 54.9% for HyenaDNA of the same small size. Not every step up in length helps (HyenaDNA drops at 262,144), but at the longest context the gap between the models is widest. Numbers from the paper's Table 5.

4.4 Audio modeling and generation · original

Long-context pretraining. On YouTubeMix, 4 hours of solo piano sampled 16,000 times a second, the paper swaps the S4 + MLP blocks of the SaShiMi audio model for Mamba blocks and trains with contexts from 8,192 up to about a million samples at fixed compute. Both improve with longer context and Mamba is better throughout, with the gap widening at longer lengths (Figure 9). The metric is bits per byte. This is the one experiment where the complex-valued SSM was used.

Speech generation. SC09 is a dataset of one-second clips of people saying the digits zero to nine. Models generate clips from scratch, scored by automatic measures including FID and IS (a score of how clear and varied the samples are; higher is better).

Selected rows of the SC09 table (Figure 10), from Gu & Dao (2023) (CC BY 4.0)
ModelParamsFID ↓IS ↑
WaveNet4.2M5.082.27
SaShiMi5.8M1.995.13
DiffWave (diffusion)24.1M1.925.26
DiffWave + SaShiMi23.0M1.425.94
Mamba6.1M0.946.26
Mamba24.3M0.677.33

A small Mamba beats much larger GAN and diffusion models, and halves SaShiMi's FID. An ablation (Figure 11) swaps block types in the U-Net's outer and centre stages: Mamba is better than S4 + MLP in the outer blocks, and in the centre Mamba beats S4 + MLP, which beats attention + MLP.

4.5 Speed and memory · original

  • The scan (state size N = 16, on an A100 GPU): faster than FlashAttention-2 beyond sequence length 2,048, and 20 to 40 times faster than a standard scan written in PyTorch. Appendix D adds: up to 7 times faster than attention at 32,768.
  • Generation: 4 to 5 times higher throughput than a Transformer of similar size, because with no KV cache it can serve much larger batches. An untrained 6.9B Mamba out-generates a Transformer five times smaller (1.3B).
  • Training memory (Table 7, 125M models, length 2,048): comparable to the most optimized Transformer, for example 4.8 GB against 4.6 GB at batch size 1 and 23.1 GB against 20.7 GB at batch size 16.

Why no cache matters: a Transformer must keep keys and values for every past token, so memory grows with the conversation; Mamba keeps a fixed state per layer. The efficient architectures lesson computes that fixed state (ssm_state_bytes) against a growing KV cache.

4.6 Model ablations · original

All on language modelling with models of about 350 million parameters; lower perplexity is better.

Selected ablations, from Gu & Dao (2023), Table 2 and Figure 13 table (CC BY 4.0)
VariantPerplexity
H3 block, S4 (complex) layer10.30
H3 block, S6 layer8.95
Mamba block, S4 (complex) layer10.54
Mamba block, S6 layer8.69
Nothing selective10.93
Only B selective10.15
Only C selective9.98
Only Δ selective9.81
Δ, B and C selective8.71

Three lessons: among time-invariant layers the choice barely matters (all about 10.2 to 10.6), while making the layer selective is a big gain; Δ is the single most important selective parameter, as Theorem 1 suggests, but the three together do far better than any one; and real-valued A initializations beat the standard complex one once the SSM is selective (8.71 against 9.16).

A bigger state only helps if it is selective

Hover or tap the chart to read perplexity at each state size.

Reading it: the x-axis is the state size N per channel (1 to 16, log scale); the y-axis is perplexity, lower being better. With B and C fixed (time-invariant), growing the state from 1 to 16 barely helps: 9.88 to 9.81. With B and C selective, the same growth takes perplexity from 9.73 to 8.71, more than a full point, for about 1% more parameters (367.1 million to 371.5 million). A bigger notebook is only useful to a note-taker who chooses what to write in it. Numbers from the paper's Figure 16 table.

5 Discussion and conclusion · original

“Our results suggest that Mamba is a strong candidate to be a general sequence model backbone.”Gu & Dao (2023), §6

The paper's own caveats

  • No free lunch. Selection fixes SSMs' weakness on discrete data but can hurt where LTI models shine. On raw audio, a smooth signal sampled at regular intervals, making the SSM selective hurt performance (Appendix E.4); the gap shrinks when only the inner layers, which see already-compressed audio, are made selective.
  • Downstream affordances. Transformer models come with a whole ecosystem: fine-tuning, prompting, in-context learning, instruction tuning, RLHF, quantization. Whether SSMs have the same properties was an open question.
  • Scale. The experiments stop at a few billion parameters, below the strongest open models and below the 7B scale at which RWKV and RetNet had been evaluated. Whether Mamba stays ahead at larger sizes was left open.

Conclusion

A selection mechanism lets structured state space models do context-dependent reasoning while still scaling linearly in sequence length, and a simple attention-free architecture built on it matches or exceeds strong Transformers across language, DNA and audio. The authors single out emerging long-context domains such as genomics, audio and video.

Appendices · original

Selection is not just “gating” (Appendix A)

“Gating” has come to mean any multiplication in a network. The paper gives an example to show why that is too loose: a layer y = σ(Wx) ∘ x (a GLU) technically counts as gated, as a hypernetwork and as data-dependent, yet it is little more than an activation function. The authors reserve “selection” for the older, narrower sense of RNN gating: a model choosing to take in or ignore inputs along the sequence. That is what separates it from the per-token multiplications in H3 or Hyena.

Where the memory goes (Appendix D)

During training in 16-bit precision, a FlashAttention layer stores about 12 bytes of activations per token and an MLP layer about 20, 32 in total; a selective SSM layer stores about 16. So two Mamba blocks, which replace one attention and one MLP layer, need about the same activation memory. When a sequence is too long to fit in SRAM, the fused scan is run chunk by chunk, carrying the state from one chunk to the next.

Proof of Theorem 1 (Appendix C)

With N = 1, A = −1 and B = 1, the continuous system is h′ = −h + x, a “leaky integrator”. Zero-order hold gives Ā = e−Δ and B̄ = 1 − Ā. With Δ = softplus(z), e−Δ = 1 / (1 + ez) = 1 − σ(z), so Ā = 1 − g and B̄ = g with g = σ(z). The §3.5 widget shows the two roads agreeing for any z.

Where it leads

Idea in the paperWhy it lastsWhere to build it
Compression is the problem: a state is a summaryExplains both why attention is costly (it never summarises) and why recurrent models forgetefficient architectures
Parameters that depend on the token (selection)The step size is a learned gate, so a fixed-size state can hold what mattersselective SSM
Parallel scan instead of convolutionAny recurrence of the form h = a·h + b trains in parallel, even when a changes every stepparallel scan
Keep the big intermediate on the chipMemory traffic, not arithmetic, is the usual bottleneck on GPUsFlashAttention companion
No KV cache, constant memory per tokenMakes long generation and big batches cheap; mixing in a few attention layers restores exact recallhybrid models
Selective SSMs as a form of linear attentionThe follow-up paper, Mamba-2, shows the two are views of one computationlinear attention

Glossary

Every term with hover guidance on this page, in one place.