Kurisutina

Locating and editing factual associations in Mamba

arXiv:2404.03646v2 (2 August 2024), Northeastern University (Khoury College of Computer Sciences). Published at COLM 2024, as printed. Code at romba.baulab.info (not read). Read by researcher R4f for research batch R4 (recurrence for the base and the person slot), 5 October 2026. Provenance: papers/base/sensharma2024_mamba_editing.provenance.json.

What was read

  • Read in full: every line of the pdftotext conversion (18 pages, 2,689 lines): abstract, Sections 1–8, ethics, reproducibility, references, and Appendices A (datasets), B (why attention knock-out is hard in Mamba), C (isolating the output projection), D (patching in Pythia), E (relation linearity in Pythia), F (its causality metric) and G (per-prompt patching).
  • Not read: figure contents. Every ROME score (Figure 6) and every indirect-effect curve exists only as a plot, and the text gives only axis ticks and labels. This summary uses what the text states; no score is read off a plot. Appendices F–G are plots with captions only.

What they did

  • Models: Mamba-2.8B, compared with the Transformer Pythia-2.8B.
  • The MambaBlock: h^(ℓ) = h^(ℓ−1) + o, with o = W_o(s ⊗ g), where s = selective-SSM(SiLU(Conv1D(W_a·h))) and g = SiLU(W_g·h).
  • Four analyses:
    1. Activation patching (causal tracing). Subjects were swapped, for example Michael Jordan for Pelé, on 400 facts from 6 relations of the RELATIONS dataset. Paths were blocked through the SSM output s, the gate g and the block output o.
    2. ROME rank-one edits to W_a, W_g or W_o at every layer, on the first 2,000 CounterFact records. Scores were efficacy, generalisation to paraphrases, and specificity (neighbouring facts unchanged).
    3. Linear relational embedding (LRE) faithfulness on 26 factual relations.
    4. An attention-knockout analogue: mean-ablating a token's input to the convolution and SSM path.

Main results (verified, from the text)

  • Facts localise as they do in a Transformer.
    • The indirect effect is high at an "early site" (early to middle layers, at the subject's last token) and a "late site" (later layers, at the prompt's last token), "consistent with" findings on GPT models.
    • The SSM outputs s matter only at the late site, as attention does in GPT models. Unlike Transformer MLPs, no Mamba state specialises in the early site alone.
    • The output projection W_o is the strongest mediator at both sites.
    • Residual states up to layer 46 matter only at the subject's last token; after layer 46, the effect at the prompt's last token jumps to 1.0. The separation between early-to-middle and later layers is cleaner than in Pythia.
  • ROME works on Mamba.
    • Rank-one edits to any of W_a, W_g or W_o reach "high scores for a range of early to middle layers", matching findings on Transformers. W_o gives the best overall results, with better generalisation and competitive specificity.
    • Weaker spots:
      • W_g and W_o edits lose score and generalisation after about layer 43;
      • W_a edits generalise poorly in early layers;
      • W_g shows "sudden drops" at some middle layers, making it "an unreliable mediator at some layers".
  • Relations decode about as linearly as in a Transformer. LRE exceeded 50% faithfulness for 10 of 26 relations in Mamba-2.8B and 11 in Pythia-2.8B. Both fail on relations with many possible answers.
  • How information flows.
    • Blocking the non-subject tokens' flow through the convolution and SSM in early-to-middle layers cut the probability of the answer by up to 50%.
    • Blocking the subject's flow showed two valleys: early layers, and layers 43–48. Blocking only the subject's last token mattered around layers 20–21.
    • The pattern matches Geva et al.'s GPT findings, with caveats.
  • The recurrent path cannot be cut surgically.
    • Subtracting one token's SSM contribution from a later state did not block that token: the 4-wide convolution leaks it into neighbouring states, and "Mamba-2.8b could often refer to the kth token from the qth token in copy and factual recall tasks".
    • Only blocking a token's flow to all later tokens, by mean ablation, worked.
  • The authors' conclusion: "when it comes to factual recall, the two architectures share many similarities". They speculate that autoregressive training itself induces localised factual recall, "independent of modeling architecture".

Limits

  • One model of each kind, at 2.8B, and Mamba-1 only; no hybrid.
  • Single edits only: no mass or sequential editing, no removal, no unlearning. Scores exist only in figures.
  • Patching locates causal effects, but the best place to edit need not be where patching points (the authors cite Hase et al. 2024).

What it means for the base (inference)

  • A recurrent backbone gives no new precision over knowledge.
    • Facts in Mamba live in projection weights, are recalled at the same moments as in a Transformer, and can be inserted by the same rank-one edits.
    • Editing is as feasible as in a Transformer, not more. Removal is untested; TOFU's entanglement result (maini2024_tofu) should be assumed to carry over.
  • Content in the recurrent path cannot be addressed item by item.
    • Even a known token's contribution could not be subtracted cleanly from the state. Removing something from a live state is all or nothing (a reset).
    • Adding, removing and gating single items therefore needs explicit entries outside the backbone: the store, the gate and the slot files.
  • Interpretability tools transfer. Causal tracing, ROME and LRE work on Mamba with modest changes, so the battery's knowledge probes would also apply to a hybrid base.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.