arXiv:2203.15556v1 (29 March 2022), DeepMind. Read by researcher R4b for research batch R4 (how to train the vocabulary engine), 5 October 2026. Provenance: papers/base/hoffmann2022_compute_optimal.provenance.json.
What was read
- Read in full: every line of the pdftotext conversion (36 pages): main text, all tables, figure captions, Appendices A–J (dataset makeup, cosine cycle length, C4/GitHub replication, fitting details, curvature, FLOP computation, Adam vs AdamW, Pile/MMLU/BIG-bench tables, model card, the list of all trained models), references.
- Not read: figures as images (only their captions and the axis labels that pdftotext extracted).
What they did
- Question: for a fixed compute budget C, how should one split it between parameters N and training tokens D? Loss is modelled as L(N, D), minimised under FLOPs(N, D) = C.
- Over 400 runs, 70M to over 16B parameters, 5B to 500B tokens, all on MassiveText, all under one epoch of data. Three estimation methods:
- Fix model sizes, vary training length (4 cosine horizons per size, spanning 16×), take the minimum-loss envelope.
- IsoFLOP profiles: 9 fixed budgets from 6×10^18 to 3×10^21 FLOPs, vary model size, fit a parabola to find the loss valley.
- Fit a parametric loss L(N, D) = E + A/N^α + B/D^β with a Huber loss (δ = 10^-3) in log space.
- Test: train Chinchilla, 70B parameters on 1.4T tokens, with Gopher's compute (5.76×10^23 FLOPs; Gopher is 280B on 300B tokens), and compare.
Main results (verified)
-
Parameters and tokens should scale equally with compute. Exponents (N_opt ∝ C^a, D_opt ∝ C^b), 10th–90th bootstrap percentiles in brackets:
- Approach 1: a = 0.50 (0.488, 0.502), b = 0.50 (0.501, 0.512).
- Approach 2: a = 0.49 (0.462, 0.534), b = 0.51 (0.483, 0.529).
- Approach 3: a = 0.46 (0.454, 0.455), b = 0.54 (0.542, 0.543).
- Kaplan et al. 2020 had a = 0.73, b = 0.27.
- Replications on C4 (a = 0.50, b = 0.50) and GitHub code (a = 0.53, b = 0.47).
-
Compute-optimal sizes (Table 3, Approach 1):
Parameters FLOPs Tokens 400M 1.92×10^19 8.0B 1B 1.21×10^20 20.2B 10B 1.23×10^22 205.1B 67B 5.76×10^23 1.5T 175B 3.85×10^24 3.7T That is about 20 tokens per parameter. Approach 2 gives 7.7B tokens for 400M and 20.0B for 1B; Approach 3 gives 9.2B and 27.1B (Table A3).
-
The fitted loss (Approach 3): L(N, D) = 1.69 + 406.4/N^0.34 + 410.7/D^0.28. E = 1.69 is read as the entropy of natural text (in their tokenizer and data).
-
FLOPs ≈ 6ND. Their full count (embeddings, attention, dense blocks, logits; backward = 2× forward) is within 0.99–1.10 of 6ND for 73M–6.8B models (Table A4).
-
Schedule matters. The cosine cycle should match the training length; overshooting it by more than 25% clearly degrades the final loss (Figure A1). Decay by 10× is slightly better than to 0; 5× is clearly worse. This is why Kaplan's fixed-schedule estimates favoured larger models.
-
Chinchilla beats Gopher (4× larger) almost everywhere:
- MMLU 5-shot 67.6% against 60.0% (abstract: 67.5%); better on 51 of 57 subjects.
- BIG-bench average 65.1% against 54.4% (worse on 4 of 62 tasks).
- LAMBADA 77.4% against 74.5%; Pile bits-per-byte lower on every subset; Wikitext103 perplexity 7.16 against 7.75.
- Closed-book QA: Natural Questions 31.5% (5-shot) against 24.5%; TriviaQA unfiltered 73.2% against 63.6%.
-
Training details: AdamW beat Adam; bf16 forward/backward with a float32 weight copy in the optimiser state; batch 1.5M → 3M tokens; max LR 1×10^-4; SentencePiece vocabulary of 32,000 without NFKC normalisation.
-
Data epochs: in the 1.4T-token run, Wikipedia was seen 3.40 times and MassiveWeb 1.24 times (Table A1).
-
Toxicity is not a function of model quality: 25,000 unprompted samples gave mean PerspectiveAPI toxicity 0.087 (Chinchilla) against 0.081 (Gopher), 95th percentile 0.238 against 0.230.
-
Retrieval, in passing: the authors note that retrieval (Borgeaud et al.) "effectively increases the number of data tokens seen during training (by a factor of ∼10)".
Limits
- Only two large-scale comparison points (Chinchilla, Gopher), no intermediate-scale validation.
- The frontier is assumed to be a power law; the authors see concavity at high compute (Appendix E), so optimal sizes may be overestimated.
- All runs are under one epoch. The multi-epoch regime is explicitly left open (Muennighoff 2023 takes it up).
- The constants (E, A, B) belong to MassiveText and this tokenizer; they do not transfer to a different corpus. The exponents replicated on C4 and code.
- The model is not released; the evaluations may suffer train/test leakage (the authors flag this for the language-modelling benchmarks).
What it means for the vocabulary engine (inference)
- A first sizing rule: about 20 tokens per parameter for compute-optimal training, C ≈ 6ND. A 1B model wants about 20B tokens (1.2×10^20 FLOPs); a 400M model about 8B tokens (1.9×10^19 FLOPs).
- Compute-optimal is not the same as "best for us". Chinchilla optimises loss per training FLOP. A vocabulary engine that must not hold much knowledge may prefer a smaller or a differently trained model; this paper says nothing about what the extra parameters store.
- The model's knowledge rises with training tokens (closed-book QA gains of 7–10 points over Gopher at a quarter of the size). Training longer on a broad corpus is exactly what fills the weights with facts.
- Set the training length before starting and match the learning-rate schedule to it; a proof of concept should not plan to "extend later" without a new cooldown.