NeurIPS 2023; arXiv:2305.16264v5 (28 June 2025), Hugging Face, Harvard, Turku. Read by researcher R4b for research batch R4 (how to train the vocabulary engine), 5 October 2026. Provenance: papers/base/muennighoff2023_data_constrained.provenance.json.
What was read
- Read in full: every line of the pdftotext conversion (50 pages): main text, references, Appendices A–X (derivation of the fit, C4 coefficients, contour plots, double descent, deduplicated data, excess parameters, Galactica case study, training loss, OSCAR curves, validation loss by epoch, evaluation details, all downstream tables 4–14, filtering procedure and results, loss curves, limitations, contributions, hyperparameters and Table 15 architectures, prompt examples, other experiments, release, version control, broader impacts).
- Not read: figures as images (captions and extracted axis labels only).
What they did
- Question: when unique data is limited, how much is a repeated token worth, and how should compute be split between parameters and epochs?
- More than 400 GPT-2-architecture models, 10M to 9B parameters, up to 1,500 epochs, up to 900B training tokens, on subsets of C4 (also OSCAR; a few on the Pile). Repeated runs always repeat the whole subset; smaller subsets are nested in larger ones. Held-out test loss, not training loss (repeated data makes training loss overfit).
- Three protocols: fixed unique data (100M, 400M or 1.5B tokens) with varying parameters and epochs; fixed FLOPs (9.3×10^20, 2.1×10^21, 9.3×10^21) with 8 data budgets each; a parametric fit over all runs.
- The fit generalises Chinchilla: effective data D' = U_D + U_D·R*_D·(1 − e^(−R_D/R*_D)), and the same form for effective parameters, where R_D is the number of repeats (epochs − 1).
- Downstream: 19 tasks, 0–5-shot, scores rescaled so random = 0; five seeds for most configurations.
- Complementary strategies: filling missing data with Python code (The Stack), perplexity filtering, deduplication.
Main results (verified)
- Up to about 4 epochs, repeated data is almost as good as new data.
- The 8.7B model trained 4 epochs on 44B unique tokens ended with only 0.5% higher validation loss than 1 epoch on 178B unique tokens.
- Downstream differences up to about 4 epochs (25% of the data budget) were insignificant, then dropped. Example (4.2B, 84B tokens, C4): average 22.1 at 1 epoch, 22.8 at 4, 20.6 at 14, 15.2 at 44 epochs.
- The "half-life" of repetition is about 16 epochs. Fitted R*_D = 15.4 (15 repeats lose 1/e of the value of fresh tokens), R*_N = 5.3. Gains beyond about 16 epochs diminish extremely fast; Figure 1 marks repetition at 40 epochs as "worthless". Fit on 182 runs, R² = 0.77 (Table 1).
- Excess parameters lose value faster than repeated data, so with limited data spend extra compute on epochs before parameters.
- Tested at scale: for 9.3×10^21 FLOPs and 25B unique tokens, the predicted 6.3B model (9.7 epochs) beat the Chinchilla-sized 8.7B model (7.1 epochs): loss 2.359 against 2.376, downstream average 25.9 against 23.5, with 27% fewer parameters.
- Galactica (120B on 106B unique tokens, 4.25 epochs) should, by this fit, have been about 40B parameters trained for 1.35T tokens (12.75 epochs).
- Their C4 fit of the Chinchilla law: L(N, D) = 1.87 + 521/N^0.353 + 1488/D^0.353. With repetition: U_N = 0.051·U_D, R*_N = 5.3, R*_D = 15.4. It reproduces Chinchilla's C4 isoFLOP optimum (70.0B parameters, 1.37T tokens at Gopher's compute).
- Tiny data budgets: with 100M unique tokens the one-epoch compute-optimal model is about 7M parameters; the lowest loss came at about 20–60× more parameters and epochs (about 7,000× the FLOPs), a loss reduction of more than 50%. Models above 2B parameters on 100M tokens were worse than 200M–1B ones (Appendix F); double descent appeared around 200 epochs (Appendix D).
- Code as filler: replacing up to 50% of natural-language data with Python kept natural-language performance (4.2B: average 22.1 at 0% against 23.8 at 50%) and raised bAbI state-tracking from 0 to about 23. Code plus 4 epochs gives about 8× the training tokens of the unique natural-language data "just as good" as unique data.
- Filtering: perplexity filtering (top 25% by a Wikipedia-trained 5-gram model) helped, especially on noisy OSCAR; deduplication did not improve downstream scores on C4 (it may reduce memorisation, which they did not measure).
- Recipe details: cosine schedule decaying 10× (2×10^-4 → 2×10^-5), 1% warm-up, Adam (β2 0.999; 0.95 slightly better), dropout 0.1, weight decay 0.1, clipping 1.0, bf16, GPT-2 tokenizer (50,257), sequence 2,048. About 3 million GPU hours on AMD MI250X (LUMI).
Limits
- One architecture (GPT-2-style dense transformer), mostly one dataset (C4, checked on OSCAR); hyperparameters fixed rather than tuned per epoch count. The authors expect a higher learning rate to make diminishing returns start earlier.
- The fit does not allow excess epochs or parameters to hurt; failing runs (for example 44 epochs) are underestimated.
- Repeating the whole dataset was studied, not up-sampling part of it (Hernandez et al.: repeating 0.1% of the data 100 times degrades performance).
- Loss and broad few-shot benchmarks only; nothing on what the extra epochs store (facts, memorised strings).
What it means for the vocabulary engine (inference)
- A curated corpus can be small and repeated. If we filter the corpus hard (no person content, limited world facts), we can train on about 4 epochs at almost no cost, and up to about 16 with diminishing but real returns. Data scarcity is not a reason to fall back on a raw web corpus.
- With limited data, prefer a smaller model trained for more epochs rather than a bigger one.
- Repetition is also exposure. Every extra epoch is another exposure of every fact in the corpus; knowledge-capacity results (Allen-Zhu and Li 2024) say exposure counts decide what gets stored. Epochs therefore trade efficiency against the "holds too much" requirement; this paper does not measure that side.
- Code is a cheap filler that carries no person content and improves state tracking. That is attractive for a base that must not hold people's views, if code-like reasoning is acceptable in the vocabulary.