arXiv:2406.07887v1 (12 June 2024), NVIDIA with Wisconsin–Madison, Princeton, Together AI, CMU and Cartesia AI (authors include Tri Dao and Albert Gu, the Mamba authors). Technical report. Read by researcher R4b for research batch R4 (scope addendum on recurrent architectures), 5 October 2026. Provenance: papers/base/waleffe2024_mamba_empirical.provenance.json.
What was read
- Read in full: every line of the pdftotext conversion (20 pages): abstract, Sections 1–6, all tables (1–11), figure captions, references, Appendix A (hybrid layer allocation algorithm).
- Not read: figures as images. Phonebook and MMLU curves (Figures 3, 6–8) are known only from their captions and the text; no number below is read off a plot.
What they did
- A controlled comparison at scale. 8B-parameter models trained on the same data with the same hyperparameters and evaluation:
- a Transformer (32 layers, width 4096, RoPE, SwiGLU);
- Mamba and Mamba-2 (56 layers, width 4096, SSM state dimension 128, no positional encoding);
- a hybrid, "Mamba-2-Hybrid": 56 layers, of which 24 Mamba-2 (43%), 4 self-attention (7%) and 28 MLP (50%), spread evenly, a Mamba-2 layer first, no positional encoding.
- Data: 1.1T and 3.5T tokens (70% English, 15% non-English, 15% code; a predecessor of the Nemotron-4 data), 256K SentencePiece vocabulary. Batch 256 or 1024, cosine schedule, BF16, Adam.
- Evaluation: 12 standard tasks (LM Evaluation Harness), MMLU in three formats, the synthetic Phonebook task (recall a phone number from a list in context), 9 natural long-context tasks, and 13 synthetic RULER tasks. Long-context variants: 16K, 32K and 128K, by continued pre-training on 50B tokens.
- Tooling: Mamba and Mamba-2 layers implemented in NVIDIA's Megatron-LM with tensor, sequence and (Mamba-2) pipeline parallelism; code and weights released.
Main results (verified)
- Pure SSMs model language as well as Transformers. Average over WinoGrande, PIQA, HellaSwag and ARC (without MMLU): 1.1T tokens, Transformer 67.87, Mamba 68.07, Mamba-2 68.56; 3.5T tokens, Transformer 68.14, Mamba-2 70.63.
- They lag on tasks that need copying or in-context learning.
- Five-shot MMLU at 1.1T tokens: Transformer 46.28, Mamba 28.00, Mamba-2 29.19 (about 17 points lower). At 3.5T tokens the gap closes to 1.37 points (50.07 against 48.7).
- In MMLU's cloze format, the SSMs match or beat the Transformer (Transformer 37.26/39.24 at 0/5-shot; Mamba 38.26/39.28; Mamba-2 37.68/38.17). The authors conclude that "pure SSM models contain the same knowledge as the Transformer" but need more training to handle the multiple-choice format, which they attribute to SSMs being unable to "route the knowledge of each answer into a single answer token".
- Phonebook: the Transformer answers near 100% up to its 4,096-token training context. Mamba and Mamba-2 start failing beyond about 500 tokens of phone book (1.1T tokens); Mamba-2 trained on 3.5T tokens still fails beyond about 1,000. More training does not fix this.
- The SSMs show "fuzzy memory": wrong answers share several correctly placed digits with the right number, so the entries are compressed into the running state, not lost.
- Telling the model in advance which name it will be asked about (the "reversed" phone book) helps the SSMs, but accuracy still degrades before 4,096 tokens: "it remains challenging for the SSM to decide which information to store exactly and which information to forget."
- A few attention layers fix recall.
- In 130M ablations (confirmed at 840M), validation loss was lowest with about 8% attention layers; 30–50% of layers can be MLPs without loss, and 50% MLPs train 20% faster than 5%. RoPE is unnecessary and slightly hurt after context extension; grouped-query attention costs about 0.04% perplexity.
- At 3.5T tokens, the 8B hybrid beat the Transformer on all 12 standard tasks: average 55.82 against 53.17 (+2.65); five-shot MMLU 53.60 against 50.07.
- The hybrid did Phonebook perfectly up to 5.5K tokens with a 4K training context; the 128K hybrid did it perfectly beyond 150K tokens.
- In-context learning remains weaker: moving from 0- to 5-shot MMLU gained the hybrid 1.45 points against 4.38 for the Transformer.
- Long context is mixed. On 9 natural long-context tasks the 16K/32K Transformers were about 1 point better on average, mainly on multi-document QA. The authors suspect the SSM states were "confused by documents irrelevant to the question", possibly because the extension data packed unrelated documents together. On the synthetic RULER tasks the hybrids were clearly better (for example 81.99 against 77.62 average at 4K; pure Mamba-2 52.14).
- Hybrids were more sensitive to prompt wording: on Musique, prompt changes moved the hybrid between 10.63 and 16.16 and the Transformer between 15.25 and 17.68.
- Training cost and tooling.
- Model FLOPs utilisation on 1,024 H100s: hybrid 29.9%, Transformer 30.7%, same parallel configuration.
- Pure Mamba (v1) at 8B trained almost 3× slower than Mamba-2 "due to the large state dimension"; the Mamba-2 scan is up to 8× faster than Mamba's. Tensor-parallel Mamba needs two all-reduces per layer against one for a Transformer; Mamba-2 needs one but uses GroupNorm (group size above 256).
- Hyperparameters transferred from Transformers unchanged ("Mamba network hyperparameters are similar to that of Transformers").
- Generation was predicted to be up to about 8× faster than the Transformer at long input contexts (batch 32).
Limits
- One scale (8B) and one data mixture; mostly base-model benchmarks. Pure Mamba (v1) was trained only on 1.1T tokens.
- The Phonebook and MMLU-format explanations are the authors' hypotheses; no fine-tuning was tried to fix pure-SSM copying.
- The long-context extension recipe (packed unrelated documents) may disadvantage SSM layers; the authors say so.
- No analysis of where knowledge is stored, of editing, or of carrying state across sessions.
What it means for the base (inference)
- Knowledge sits in the weights in SSMs too. Cloze-format parity suggests the same knowledge is stored; the recurrent state holds the current context, not long-term knowledge. Recurrence does not move knowledge out of the weights.
- A fixed-size state is lossy by construction. With a 128-dimensional state per channel, an 8B SSM could not reproduce exact entries from a list beyond about 500–1,000 tokens even when told in advance what to keep; it kept a compressed, "fuzzy" version. A person's verbatim episodic record cannot live in such a state; it needs an external store or attention over stored text.
- What the state keeps is decided by the shared network, from the input. Irrelevant documents disturbed the state; nothing in the architecture lets a slot protect state from the input stream.
- If recurrence is wanted, the evidence favours a hybrid (about 8% attention, about half MLP) over a pure SSM: same training cost per token as a Transformer, better short-context scores, exact recall restored, cheaper long-context generation.