arXiv:2312.00752v2 (31 May 2024), Carnegie Mellon University and Princeton University (cited as COLM 2024 in kang2025_state_offset's reference list). Read by researcher R4f for research batch R4 (recurrence for the base and the person slot), 5 October 2026. Provenance: papers/base/gu2023_mamba.provenance.json.
What was read
- Read in full: every line of the pdftotext conversion (36 pages, 3,382 lines): abstract, Sections 1–6, Algorithms 1–2, Theorem 1, Tables 1–15, figure captions, references, Appendices A (selection against gating), B (related work), C (proof of Theorem 1), D (hardware-aware algorithm) and E (experimental details).
- Not read: figure contents. The plot labels of Figures 4–10 and the graphic parts of Tables 1–2 come out of pdftotext as garbled symbol-font text, so I used the captions and the text. Tables 1 and 11 were extracted as text. Figure 9's caption is truncated. No number below is read off a plot.
What they did
- Selective SSM ("S6"). The SSM parameters Δ, B and C become functions of the current input: B and C are linear projections of x, and Δ = softplus(parameter + Linear(x)), broadcast over channels.
- A stays an input-independent parameter, acting through Ā = exp(ΔA).
- Because the model now varies over time, it cannot be computed as a convolution. It runs as a recurrent scan, in a fused kernel that keeps the expanded state in fast on-chip memory and recomputes it in the backward pass.
- Architecture. One homogeneous "Mamba block" merges the H3 block and a gated MLP (expansion E = 2).
- Most parameters sit in the linear projections (3ED² per block); the SSM parameters are few.
- Two Mamba blocks match one Transformer layer.
- The recurrent state is D × N numbers per layer (N = 16 in the language models). It is built from the input and kept separately from the parameters.
- Evaluations:
- synthetic selective copying and induction heads;
- language modelling on the Pile: scaling laws from 125M to 1.3B at Chinchilla token counts, and 300B-token models up to 2.8B against Pythia and RWKV;
- DNA, audio, speed and memory;
- ablations at about 350M.
Main results (verified)
-
Selection is gating. In Theorem 1, with N = 1, A = −1 and B = 1, the selective SSM becomes the gated RNN h_t = (1 − g_t)·h_{t−1} + g_t·x_t, with g_t = σ(Linear(x_t)).
- A small Δ keeps the state and ignores the input; a large Δ resets the state and focuses on the input.
- Selective models "can simply reset their state at any time to remove extraneous history", while time-invariant models "bleed information between the sequences".
- The authors claim selection lets the model "filter out irrelevant information and remember relevant information indefinitely".
-
The trade-off is compression. Attention "explicitly does not compress context at all" (the KV cache). Recurrent models have a finite state, and "their effectiveness is limited by how well this state has compressed the context".
-
Selective copying (length 4,096, 16 tokens to memorise; accuracy %):
Block S4 S6 Hyena No gate 18.3 97.0 — H3 57.0 99.7 30.1 Mamba 56.4 99.8 28.4 -
Induction heads (2-layer models, vocabulary 16, trained at length 256): this recalls one association after long filler.
- Mamba (74K parameters) was perfect at every length from 64 to 1,048,576 tokens, 4,000 times the training length.
- Attention models (137K parameters) collapse beyond 2×. At 1,024 tokens: MHA-Abs 26.6%, RoPE 31.3%, xPos 67.6%; they run out of memory beyond 16,384. H3 and Hyena degrade in the same way.
-
Language modelling.
-
Scaling laws (125M on 2.5B tokens, 350M on 7B, 760M on 15B, 1.3B on 26B): Mamba is "the first attention-free model to match" a strong Transformer++ recipe. This is shown in a figure only.
-
Zero-shot averages over six tasks at 300B tokens:
Mamba Avg Compared with Avg 130M 44.7 Pythia-160M 40.6 370M 50.0 Pythia-410M 48.2 790M 57.1 Pythia-1B 51.9 1.4B 59.7 Pythia-1.4B; RWKV-1.5B 55.2; 54.3 2.8B 63.3 Pythia-2.8B; RWKV-3B; Pythia-6.9B 59.1; 59.6; 61.7 -
Pile perplexity: Mamba-2.8B 6.22 against Pythia-2.8B 6.73.
-
-
Ablations (about 350M, Chinchilla tokens; perplexity):
Change Perplexity Mamba block: S6 / S4 complex / S4 real / Hyena 8.69 / 10.54 / 10.56 / 10.75 H3 block with S6 8.95 Selective Δ only / nothing selective / Δ, B and C selective 9.81 / 10.93 / 8.71 State size N = 1 → 16, selective B and C 9.73 → 8.71 (about 1% more parameters) State size N = 1 → 16, constant B and C 9.88 → 9.81 -
Interleaving with attention. In the 2K-context ablation, a Mamba–attention interleave was "only slightly better" than pure Mamba, and a Mamba–MLP interleave slightly worse.
-
Efficiency.
- The fused scan is faster than FlashAttention-2 beyond 2K tokens, up to 7× faster than attention at 32K, and 20–40× faster than a standard PyTorch scan.
- Inference throughput is 4–5× that of a similar-sized Transformer, since there is no KV cache.
- Training memory is comparable. For 125M models at length 2,048, Mamba used 4.8 GB at batch 1 and 38.2 GB at batch 32; a Transformer with FlashAttention-2 used 4.6 and 34.5 GB.
-
Not always better. On continuous audio waveforms, selection "significantly hampers performance".
-
The authors' open question. Whether SSMs have the Transformer ecosystem's "affordances" is untested: fine-tuning, adaptation, prompting, in-context learning, instruction tuning, RLHF and quantization. Models go only to 2.8B.
Limits
- Small models, one language corpus (the Pile, 300B tokens); the scaling-law comparison is shown only as a figure.
- The synthetic tasks recall one association, or 16 tokens. Exact recall of many items is not tested here; Jelassi 2024 and Waleffe 2024 test it.
- Nothing on where knowledge is stored, on editing, or on conditioning by person or task.
What it means for the base (inference)
- The state is per-sequence working memory; long-term knowledge is in the projections. Most parameters are linear projections, and the D × N state is rebuilt from the input for every sequence. Recurrence gives long-term knowledge no new home.
- The base's weights decide what enters and survives the state, from the input.
- Δ, B and C are computed by shared projections of the current token.
- A person slot cannot reserve part of the state for itself. To protect content there, it would have to change those projections (an adapter), or bypass the state.
- Gating can keep a little content indefinitely. One association survived a million tokens when the model had been trained to ignore everything else.
- A small person code carried with closed gates is therefore possible in principle.
- A large record is not (see
jelassi2024_copying). - A gate that never opens is functionally a constant, which is a parameter.
- R4b's ladder sizes match this paper's protocol. The scaling protocol used 125M on 2.5B tokens and 350M on 7B, exactly R4b's ladder, and at those sizes Mamba matched a strong Transformer in perplexity.
- A 350M backbone comparison is therefore meaningful for language-modelling quality.
- It needs separate recall probes, because perplexity hides recall (
jelassi2024_copying).