Kurisutina

Training compute-optimal large language models ("Chinchilla")

arXiv:2203.15556v1 (29 March 2022), DeepMind. Read by researcher R4b for research batch R4 (how to train the vocabulary engine), 5 October 2026. Provenance: papers/base/hoffmann2022_compute_optimal.provenance.json.

What was read

  • Read in full: every line of the pdftotext conversion (36 pages): main text, all tables, figure captions, Appendices A–J (dataset makeup, cosine cycle length, C4/GitHub replication, fitting details, curvature, FLOP computation, Adam vs AdamW, Pile/MMLU/BIG-bench tables, model card, the list of all trained models), references.
  • Not read: figures as images (only their captions and the axis labels that pdftotext extracted).

What they did

  • Question: for a fixed compute budget C, how should one split it between parameters N and training tokens D? Loss is modelled as L(N, D), minimised under FLOPs(N, D) = C.
  • Over 400 runs, 70M to over 16B parameters, 5B to 500B tokens, all on MassiveText, all under one epoch of data. Three estimation methods:
    1. Fix model sizes, vary training length (4 cosine horizons per size, spanning 16×), take the minimum-loss envelope.
    2. IsoFLOP profiles: 9 fixed budgets from 6×10^18 to 3×10^21 FLOPs, vary model size, fit a parabola to find the loss valley.
    3. Fit a parametric loss L(N, D) = E + A/N^α + B/D^β with a Huber loss (δ = 10^-3) in log space.
  • Test: train Chinchilla, 70B parameters on 1.4T tokens, with Gopher's compute (5.76×10^23 FLOPs; Gopher is 280B on 300B tokens), and compare.

Main results (verified)

  • Parameters and tokens should scale equally with compute. Exponents (N_opt ∝ C^a, D_opt ∝ C^b), 10th–90th bootstrap percentiles in brackets:

    • Approach 1: a = 0.50 (0.488, 0.502), b = 0.50 (0.501, 0.512).
    • Approach 2: a = 0.49 (0.462, 0.534), b = 0.51 (0.483, 0.529).
    • Approach 3: a = 0.46 (0.454, 0.455), b = 0.54 (0.542, 0.543).
    • Kaplan et al. 2020 had a = 0.73, b = 0.27.
    • Replications on C4 (a = 0.50, b = 0.50) and GitHub code (a = 0.53, b = 0.47).
  • Compute-optimal sizes (Table 3, Approach 1):

    Parameters FLOPs Tokens
    400M 1.92×10^19 8.0B
    1B 1.21×10^20 20.2B
    10B 1.23×10^22 205.1B
    67B 5.76×10^23 1.5T
    175B 3.85×10^24 3.7T

    That is about 20 tokens per parameter. Approach 2 gives 7.7B tokens for 400M and 20.0B for 1B; Approach 3 gives 9.2B and 27.1B (Table A3).

  • The fitted loss (Approach 3): L(N, D) = 1.69 + 406.4/N^0.34 + 410.7/D^0.28. E = 1.69 is read as the entropy of natural text (in their tokenizer and data).

  • FLOPs ≈ 6ND. Their full count (embeddings, attention, dense blocks, logits; backward = 2× forward) is within 0.99–1.10 of 6ND for 73M–6.8B models (Table A4).

  • Schedule matters. The cosine cycle should match the training length; overshooting it by more than 25% clearly degrades the final loss (Figure A1). Decay by 10× is slightly better than to 0; 5× is clearly worse. This is why Kaplan's fixed-schedule estimates favoured larger models.

  • Chinchilla beats Gopher (4× larger) almost everywhere:

    • MMLU 5-shot 67.6% against 60.0% (abstract: 67.5%); better on 51 of 57 subjects.
    • BIG-bench average 65.1% against 54.4% (worse on 4 of 62 tasks).
    • LAMBADA 77.4% against 74.5%; Pile bits-per-byte lower on every subset; Wikitext103 perplexity 7.16 against 7.75.
    • Closed-book QA: Natural Questions 31.5% (5-shot) against 24.5%; TriviaQA unfiltered 73.2% against 63.6%.
  • Training details: AdamW beat Adam; bf16 forward/backward with a float32 weight copy in the optimiser state; batch 1.5M → 3M tokens; max LR 1×10^-4; SentencePiece vocabulary of 32,000 without NFKC normalisation.

  • Data epochs: in the 1.4T-token run, Wikipedia was seen 3.40 times and MassiveWeb 1.24 times (Table A1).

  • Toxicity is not a function of model quality: 25,000 unprompted samples gave mean PerspectiveAPI toxicity 0.087 (Chinchilla) against 0.081 (Gopher), 95th percentile 0.238 against 0.230.

  • Retrieval, in passing: the authors note that retrieval (Borgeaud et al.) "effectively increases the number of data tokens seen during training (by a factor of ∼10)".

Limits

  • Only two large-scale comparison points (Chinchilla, Gopher), no intermediate-scale validation.
  • The frontier is assumed to be a power law; the authors see concavity at high compute (Appendix E), so optimal sizes may be overestimated.
  • All runs are under one epoch. The multi-epoch regime is explicitly left open (Muennighoff 2023 takes it up).
  • The constants (E, A, B) belong to MassiveText and this tokenizer; they do not transfer to a different corpus. The exponents replicated on C4 and code.
  • The model is not released; the evaluations may suffer train/test leakage (the authors flag this for the language-modelling benchmarks).

What it means for the vocabulary engine (inference)

  • A first sizing rule: about 20 tokens per parameter for compute-optimal training, C ≈ 6ND. A 1B model wants about 20B tokens (1.2×10^20 FLOPs); a 400M model about 8B tokens (1.9×10^19 FLOPs).
  • Compute-optimal is not the same as "best for us". Chinchilla optimises loss per training FLOP. A vocabulary engine that must not hold much knowledge may prefer a smaller or a differently trained model; this paper says nothing about what the extra parameters store.
  • The model's knowledge rises with training tokens (closed-book QA gains of 7–10 points over Gopher at a quarter of the size). Training longer on a broad corpus is exactly what fills the weights with facts.
  • Set the training length before starting and match the learning-rate schedule to it; a proof of concept should not plan to "extend later" without a new cooldown.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.