arXiv 2403.11901 v4 (21 August 2024), cs.LG. IBM Research (13 authors; Das and Chaudhury equal contribution). The
text read carries no venue statement, so peer-review status is not confirmed here. Code: github.com/IBM/larimar (not
read). Provenance: papers/base/das2024_larimar.provenance.json.
What was read
Every line of the pdftotext -layout text of the 18-page PDF (1,222 lines after folding): abstract, introduction,
model, memory operations, scope detector, results (wall-clock time, single, sequential and batch editing, selective
forgetting, leakage prevention, long context), related work, conclusions, impact statement, references, and appendices
A–G with Tables 7–14. Not read as images: Figures 1 (architecture; caption read), 2–6 (curves; captions and the
values stated in the text read). Values from figures are given only where the text states them.
Question
Can a language model be given a fast, one-shot, editable episodic memory, in the complementary-learning sense, so that facts can be written, updated and selectively forgotten at inference time without retraining or locating facts in the weights?
Model
- Parts. An encoder (BERT-large), a fixed-size memory matrix M (K × C; K = 512 rows, C = 768 unless stated), and a decoder (GPT-2 large for "Larimar-1.3B", GPT-J 6B for "Larimar-6B"). The memory readout is projected and broadcast as a KV cache to every decoder layer.
- Memory as least squares (after the Kanerva machine and Pham et al. 2021's generative pseudo-inverse memory). Write: M = W₀†Z (addressing weights from a learned prior memory; pseudo-inverse solution). Read: Z = W·M with W = Z_query·M†, optionally sampled. Sequential write and forget: a recursive least-squares update with α = +1 to add and α = −1 to remove a previously written encoding, which keeps M the exact least-squares solution with that item removed.
- Training. Encoder, memory and decoder trained jointly on 7.6 million 64-token WikiText chunks with a loss of memory-conditioned reconstruction, reconstruction without memory, a KL term, and a language-modelling term on pretraining data. Larimar-6B: 10 epochs, Adam, learning rate 5e-6, batch 32, eight A100-80GB. Edits are never trained on; they are only written at inference.
- Perplexity on 1,000 WikiText samples: 14.6 (1.3B) and 15.9 (6B); the authors read this as memory "barely" affecting the decoder (no no-memory comparison is printed).
- Scope detector. Decides whether a query concerns memory content; if not, the decoder runs unconditioned. External version: MiniLM sentence embeddings with 1-nearest-neighbour similarity (equal-error rate 2.9%, F1 0.974 on 3,800 samples). Internal version: a classifier on Larimar's own encodings.
Results
-
Single fact editing, CounterFact (first 2,000 records; baselines copied from earlier papers):
Editor Edit success Paraphrase Neighbourhood ROME on GPT-2 XL 100.0 96.4 75.4 Larimar-1.3B 100.0 83.5 74.7 ROME on GPT-J 99.9 99.1 78.9 Larimar-6B 99.6 88.4 80.4 Larimar-6B, two paraphrases also written 99.8 93.6 79.2 -
Without the scope detector the memory leaks into unrelated prompts. Larimar-6B: neighbourhood 80.4 with the detector, 13.7 without (paraphrase rises to 93.6). Larimar-1.3B ablations: neighbourhood 74.7–75.5 with the detector, 15.1–28.5 without.
-
ZsRE single edits (exact match): Larimar-6B 94.5 / 70.4 / 25.1 (82.2 paraphrase with two paraphrases written); ROME on GPT-J 99.8 / 95.9 / 27.2.
-
Speed (Table 1, single A100): Larimar 1.1 s (GPT-2) and 1.7 s (GPT-J) against ROME 4.8 s and 13.9 s and GRACE 13.9 s and 19.3 s. The text describes these as times "across 10 edits" in one place and "for a single edit (averaged over 10 edits)" in another. The abstract claims 8–10× speed-ups; the text says 4–10×.
-
Sequential editing, ZsRE, 1,000 edits (K = 1,000): edit retention rate (mean F1 after all edits) 0.97 (Larimar-1.3B) and 0.92 (Larimar-6B), against GRACE 0.93 and MEND 0.27. On a set with duplicated facts: Larimar 0.98, GRACE 0.96, SERAC 0.31. On unseen paraphrases, Larimar starts near F1 0 and passes GRACE after about 600 edits.
-
Capacity. Batch rewrite accuracy is near 100% up to 512 edits (= K) and 82% at 1,024. Appendix: 99% at N = K = 512; 94% when K = N > 512 (trained with K = 512); about 80% at N = 2K and 54% at N = 4K.
-
Selective forgetting (write N facts, remove one, write it back with answer "unknown"; K = 512):
Model Dataset Forgotten fact recalled Others retained Larimar-6B, N = K CounterFact 0.0 0.993 Larimar-6B, N = K ZsRE 0.03 0.86 Larimar-6B, N = 2K CounterFact / ZsRE 0.03 / 0.04 0.71 / 0.50 Llama-2-13B, 6-shot in context, N = 20 CounterFact / ZsRE 0.75 / 0.68 0.77 / 0.73 -
Leakage under paraphrase attack (about 300 CounterFact facts, 20 attempts each): attack success 17.6% (one fact written as "unknown") and 21.5% (batch), against ROME 29.0% and MEMIT 49.3%.
-
Long context by recursive memory search (CNN Fast Facts 2021–2023): recall stays 0.89→0.86 (1.3B) and 0.82→0.80 (6B) from 64 to 256 facts, while Mistral-7B 3-shot falls 0.98→0.42. Reading 128 facts took 0.36 s against 1.44 s for Mistral-7B.
-
Addressing. A Gaussian-kernel addressing against a reference memory beats the pseudo-inverse when facts have few paraphrases (with one paraphrase per fact, recall 0.69 against 0.33) and loses its edge with many.
Limits
- "Episodic" here means short factual sentences in a matrix: no time, place, source, sequence or retrieval history. The authors list short facts and sentence-completion training as limits; conversation is future work.
- Baseline numbers are copied from other papers, not rerun in one setting; no variance or significance is reported.
- Capacity is fixed by K; performance falls once writes exceed it.
- Forgetting is exact subtraction. Nothing here is fitted to, or compared with, human memory.
What it means for Kurisutina
- A working fast-write memory beside a frozen decoder. One-shot writes, sequential writes with retention 0.92–0.97 over 1,000 edits, and exact removal, all without gradient steps on the language model. That is a candidate mechanism for B2's fast store and for an access state (suppress or restore what a cue retrieves) without touching the person core.
- The read path must be gated. Without a scope detector, memory-conditioned decoding wrecked unrelated answers (neighbourhood 80.4 → 13.7). Any design where stored content conditions generation needs an explicit relevance gate, and a test that unrelated behaviour is unchanged.
- Deletion is not human forgetting. The project's episode store is append-only. Larimar-style removal fits an index over that store (what is currently accessible), never the record itself.
- Fixed capacity forces a policy. Past K the memory degrades, so something must decide what is consolidated, summarized or left only in the log. That is B3's job, and Go-CLS gives one rule for it.
Cross-references
summaries/carry_on/fountas2024_em_llm.md: a non-parametric KV-cache episodic memory, the other main LLM design.summaries/base/sun2023_go_cls.md: which memories should leave the fast store.summaries/memory/kumaran2016_complementary_learning.md: the CLS view Larimar cites.