Kurisutina

Improving language models by retrieving from trillions of tokens ("RETRO")

arXiv:2112.04426v3 (7 February 2022), DeepMind; ICML 2022. Read by researcher R4b for research batch R4 (how to train the vocabulary engine), 5 October 2026. Provenance: papers/base/borgeaud2022_retro.provenance.json.

What was read

  • Read in full: every line of the pdftotext conversion (43 pages): main text, Algorithm 1, Listing 1, all tables, references, Appendices A–F (datasets, architecture and retrofitting details, hyperparameters, ablations, qualitative samples in Tables 16–21, complementary results in Tables 14–15).
  • Not read: figures as images (captions and extracted labels only). The colour coding of the sample tables (which tokens were copied) is not visible in text.

What they did

  • A retrieval-enhanced transformer, trained from scratch with retrieval. Each 2,048-token training sequence is split into 64-token chunks. For each chunk, the 2 nearest neighbours (by a frozen BERT embedding, L2 distance, SCaNN index) are fetched from a database together with their 64-token continuations. A small bidirectional encoder (2 layers, width 896) encodes them; the decoder attends to them through chunked cross-attention in every third layer from layer 6.
  • Data: multilingual MassiveText (over 5T tokens). Training retrieves from 600B tokens; evaluation from 1.75T tokens (books sub-sampled to 4%); the abstract's "2 trillion token database". Training documents with 13-gram Jaccard similarity of 0.8 or more to any evaluation document were removed.
  • Sizes: baselines of 132M, 368M, 1.3B and 7.0B non-embedding parameters; RETRO adds 8–30% (172M, 425M, 1.5B, 7.5B). All trained on 419.4B tokens, AdamW, cosine schedule.
  • Evaluations: C4, Wikitext103, Curation Corpus, LAMBADA, the Pile, 23 Wikipedia articles written after the data was collected (September 2021), Natural Questions after fine-tuning. A leakage-aware metric scores only evaluation chunks whose longest overlap with training data is below a threshold.
  • "Retrofitting": freezing a trained baseline and training only the new cross-attention and encoder weights.

Main results (verified)

  • Retrieval is worth about 10× the parameters on language modelling. The gain is constant from 150M to 7B parameters (Figure 1). C4 bits-per-byte (Table 14), baseline → RETRO with retrieval on: 0.98 → 0.82 (172M), 0.92 → 0.77 (425M), 0.84 → 0.71 (1.5B), 0.78 → 0.66 (7.5B).
  • With retrieval switched off, a RETRO model performs like the baseline (Table 14). C4 bpb with retrieval off equals the baseline at every size (0.98, 0.92, 0.84, 0.78). At 7.5B: LAMBADA 0.70 off against 0.69 baseline (0.73 on); post-cutoff Wikipedia bpb 0.65 off against 0.65 baseline (0.61 on).
  • Larger databases help a lot; more neighbours help up to about 10 (172M) or 40 (7B), though training used only 2.
  • The Pile: the 7.5B RETRO beats Jurassic-1 (178B) and Gopher (280B) on a majority of subsets; it does not help on dm_mathematics or ubuntu_irc.
  • Part of the gain is copying leaked text, part is real.
    • With MassiveText as the retrieval set, Wikitext103 perplexity drops to 3.21 (valid) / 3.92 (test) against 21.53 / 22.96 for the baseline, largely through near-duplicates that survived deduplication (Table 19 shows near-verbatim copying of a Wikipedia article).
    • Even on evaluation chunks sharing fewer than 8 contiguous tokens with any training chunk, RETRO still beats the baseline (Figure 6).
  • Question answering: fine-tuned with 20 retrieved DPR passages, RETRO 7.5B reaches 45.5% exact match on Natural Questions against 30.4% for the closed-book 7B baseline (FiD: 51.4%; FiD + distillation 54.7%).
  • Retrofitting works cheaply and leaves the base untouched. Training only the new weights (under 10% of a 7B model) on 6M sequences (3% of pre-training) nearly reaches RETRO-from-scratch performance, and the retrieval-off performance stays exactly the baseline's. Unfreezing all weights degraded retrieval-off performance (Appendix B.2).
  • Ablations (247M, 157B tokens; C4 bpb): no retrieval 0.987; RETRO 0.822. Neighbours alone give 22% of the gain, their continuations 56%; training with 1 neighbour hurts (0.858); 4 neighbours adds little; cross-attention once at mid-depth is acceptable, at top or bottom only is poor.
  • Storage: chunk-level indexing takes 215 GB for Wikipedia and 93 TB for MassiveText; token-level kNN-LM would need 15 TB for Wikipedia's 4B tokens alone.
  • The authors' privacy note (suggestion, not tested): retrieval allows "obliteration of the retrievable data at inference time", and "individualisation on private data could be made by updating the retrieval database at inference time".

Limits

  • The baseline and RETRO were trained on the same full corpus, so nothing here tests whether retrieval reduces what the weights store. The retrieval-off results suggest it does not.
  • Gains are concentrated where the database overlaps the evaluation text; a large share of headline numbers depends on near-duplicates.
  • Closed-book knowledge of the RETRO model with retrieval off was not measured on QA.
  • No experiment removes data from the database and checks that the model then fails to produce it.

What it means for the vocabulary engine (inference)

  • Retrieval-augmented pretraining does not by itself keep knowledge out of the weights. A RETRO model with retrieval off is as good as a model trained without retrieval: its parameters learned what the corpus taught. To keep facts out of the weights, they must be kept out of (or made rare in) the training text, and put only in the database.
  • The architecture is the right shape for a gated world layer. A frozen base plus a separately trained cross-attention reader can read from any store; retrofitting leaves the base's own behaviour exactly unchanged. A person-specific store (the slot's holdings) or a shared factual record (the world layer) can be swapped at inference without retraining.
  • Database content is used even when only loosely related (the neighbours of new Wikipedia articles supplied correct names and dates). That is a gate risk: whatever the store holds will be used. Gating has to happen at what is retrievable, not at what the model chooses to use.
  • Small models benefit as much as large ones, so a small vocabulary engine with a reader is a reasonable design; at 172M parameters retrieval improved C4 bpb from 0.98 to 0.82.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.