Kurisutina

Mamba: linear-time sequence modeling with selective state spaces

arXiv:2312.00752v2 (31 May 2024), Carnegie Mellon University and Princeton University (cited as COLM 2024 in kang2025_state_offset's reference list). Read by researcher R4f for research batch R4 (recurrence for the base and the person slot), 5 October 2026. Provenance: papers/base/gu2023_mamba.provenance.json.

What was read

  • Read in full: every line of the pdftotext conversion (36 pages, 3,382 lines): abstract, Sections 1–6, Algorithms 1–2, Theorem 1, Tables 1–15, figure captions, references, Appendices A (selection against gating), B (related work), C (proof of Theorem 1), D (hardware-aware algorithm) and E (experimental details).
  • Not read: figure contents. The plot labels of Figures 4–10 and the graphic parts of Tables 1–2 come out of pdftotext as garbled symbol-font text, so I used the captions and the text. Tables 1 and 11 were extracted as text. Figure 9's caption is truncated. No number below is read off a plot.

What they did

  • Selective SSM ("S6"). The SSM parameters Δ, B and C become functions of the current input: B and C are linear projections of x, and Δ = softplus(parameter + Linear(x)), broadcast over channels.
    • A stays an input-independent parameter, acting through Ā = exp(ΔA).
    • Because the model now varies over time, it cannot be computed as a convolution. It runs as a recurrent scan, in a fused kernel that keeps the expanded state in fast on-chip memory and recomputes it in the backward pass.
  • Architecture. One homogeneous "Mamba block" merges the H3 block and a gated MLP (expansion E = 2).
    • Most parameters sit in the linear projections (3ED² per block); the SSM parameters are few.
    • Two Mamba blocks match one Transformer layer.
    • The recurrent state is D × N numbers per layer (N = 16 in the language models). It is built from the input and kept separately from the parameters.
  • Evaluations:
    • synthetic selective copying and induction heads;
    • language modelling on the Pile: scaling laws from 125M to 1.3B at Chinchilla token counts, and 300B-token models up to 2.8B against Pythia and RWKV;
    • DNA, audio, speed and memory;
    • ablations at about 350M.

Main results (verified)

  • Selection is gating. In Theorem 1, with N = 1, A = −1 and B = 1, the selective SSM becomes the gated RNN h_t = (1 − g_t)·h_{t−1} + g_t·x_t, with g_t = σ(Linear(x_t)).

    • A small Δ keeps the state and ignores the input; a large Δ resets the state and focuses on the input.
    • Selective models "can simply reset their state at any time to remove extraneous history", while time-invariant models "bleed information between the sequences".
    • The authors claim selection lets the model "filter out irrelevant information and remember relevant information indefinitely".
  • The trade-off is compression. Attention "explicitly does not compress context at all" (the KV cache). Recurrent models have a finite state, and "their effectiveness is limited by how well this state has compressed the context".

  • Selective copying (length 4,096, 16 tokens to memorise; accuracy %):

    Block S4 S6 Hyena
    No gate 18.3 97.0 —
    H3 57.0 99.7 30.1
    Mamba 56.4 99.8 28.4
  • Induction heads (2-layer models, vocabulary 16, trained at length 256): this recalls one association after long filler.

    • Mamba (74K parameters) was perfect at every length from 64 to 1,048,576 tokens, 4,000 times the training length.
    • Attention models (137K parameters) collapse beyond 2×. At 1,024 tokens: MHA-Abs 26.6%, RoPE 31.3%, xPos 67.6%; they run out of memory beyond 16,384. H3 and Hyena degrade in the same way.
  • Language modelling.

    • Scaling laws (125M on 2.5B tokens, 350M on 7B, 760M on 15B, 1.3B on 26B): Mamba is "the first attention-free model to match" a strong Transformer++ recipe. This is shown in a figure only.

    • Zero-shot averages over six tasks at 300B tokens:

      Mamba Avg Compared with Avg
      130M 44.7 Pythia-160M 40.6
      370M 50.0 Pythia-410M 48.2
      790M 57.1 Pythia-1B 51.9
      1.4B 59.7 Pythia-1.4B; RWKV-1.5B 55.2; 54.3
      2.8B 63.3 Pythia-2.8B; RWKV-3B; Pythia-6.9B 59.1; 59.6; 61.7
    • Pile perplexity: Mamba-2.8B 6.22 against Pythia-2.8B 6.73.

  • Ablations (about 350M, Chinchilla tokens; perplexity):

    Change Perplexity
    Mamba block: S6 / S4 complex / S4 real / Hyena 8.69 / 10.54 / 10.56 / 10.75
    H3 block with S6 8.95
    Selective Δ only / nothing selective / Δ, B and C selective 9.81 / 10.93 / 8.71
    State size N = 1 → 16, selective B and C 9.73 → 8.71 (about 1% more parameters)
    State size N = 1 → 16, constant B and C 9.88 → 9.81
  • Interleaving with attention. In the 2K-context ablation, a Mamba–attention interleave was "only slightly better" than pure Mamba, and a Mamba–MLP interleave slightly worse.

  • Efficiency.

    • The fused scan is faster than FlashAttention-2 beyond 2K tokens, up to 7× faster than attention at 32K, and 20–40× faster than a standard PyTorch scan.
    • Inference throughput is 4–5× that of a similar-sized Transformer, since there is no KV cache.
    • Training memory is comparable. For 125M models at length 2,048, Mamba used 4.8 GB at batch 1 and 38.2 GB at batch 32; a Transformer with FlashAttention-2 used 4.6 and 34.5 GB.
  • Not always better. On continuous audio waveforms, selection "significantly hampers performance".

  • The authors' open question. Whether SSMs have the Transformer ecosystem's "affordances" is untested: fine-tuning, adaptation, prompting, in-context learning, instruction tuning, RLHF and quantization. Models go only to 2.8B.

Limits

  • Small models, one language corpus (the Pile, 300B tokens); the scaling-law comparison is shown only as a figure.
  • The synthetic tasks recall one association, or 16 tokens. Exact recall of many items is not tested here; Jelassi 2024 and Waleffe 2024 test it.
  • Nothing on where knowledge is stored, on editing, or on conditioning by person or task.

What it means for the base (inference)

  • The state is per-sequence working memory; long-term knowledge is in the projections. Most parameters are linear projections, and the D × N state is rebuilt from the input for every sequence. Recurrence gives long-term knowledge no new home.
  • The base's weights decide what enters and survives the state, from the input.
    • Δ, B and C are computed by shared projections of the current token.
    • A person slot cannot reserve part of the state for itself. To protect content there, it would have to change those projections (an adapter), or bypass the state.
  • Gating can keep a little content indefinitely. One association survived a million tokens when the model had been trained to ignore everything else.
    • A small person code carried with closed gates is therefore possible in principle.
    • A large record is not (see jelassi2024_copying).
    • A gate that never opens is functionally a constant, which is a parameter.
  • R4b's ladder sizes match this paper's protocol. The scaling protocol used 125M on 2.5B tokens and 350M on 7B, exactly R4b's ladder, and at those sizes Mamba matched a strong Transformer in perplexity.
    • A 350M backbone comparison is therefore meaningful for language-modelling quality.
    • It needs separate recall probes, because perplexity hides recall (jelassi2024_copying).

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.