arXiv:2404.05405v1 (8 April 2024), Meta FAIR and MBZUAI (later published at ICLR 2025; the arXiv v1 was read). Read by researcher R4b for research batch R4 (how to train the vocabulary engine), 5 October 2026. Provenance: papers/base/allenzhu2024_knowledge_capacity.provenance.json.
What was read
- Read in full: every line of the pdftotext conversion (41 pages): main text (Sections 1–11), Appendices A (GPT-2 scaling details, memorisation against extraction, other biography datasets, parameterised laws), B (architectures), C (quantisation), D (mixture of experts), E (junk data), F (proof of the lower bound, including Lemma F.1 and the warm-up examples), G (textbook estimate), references.
- Not readable: the figures. pdftotext turned the scatter plots into thousands of lines of point labels; I passed over those lines and used the captions, which state each figure's conclusion. No number below comes from reading a plot.
What they did
- A controlled measure of stored knowledge. A "piece of knowledge" is a (name, attribute, value) tuple. Synthetic biographies of N people (bioS: six attributes, birth date, birth city, university, major, employer, work city, written with 50 random templates per attribute), up to N = 20M; a LLaMA2-rewritten realistic version (bioR, 40 rewrites per person, up to N = 1M, 22 GB); and a fully parametric family bioD(N, K, C, D, L, T).
- Bits, not loss. Theorem 3.2 gives an information-theoretic lower bound on the bits a model must hold to reach its measured losses on names and values. Capacity ratio = those bits divided by the parameter count. Each person in bioS carries about 47.6 bits (excluding the name).
- "Exposures", not epochs: how many times each fact is seen in training (1,000 exposures of bioS can happen within one pass, because every exposure is a new paraphrase).
- Models: GPT-2 with rotary embeddings and no dropout, 1M to 0.5B parameters, many depths and widths, trained from scratch (AdamW, cosine schedule to 0.1×, fp16 mixed precision; bf16 gave the same results); also LLaMA, Mistral, GPT-2 with a quarter-sized MLP or no MLP, mixture-of-experts, and post-training quantisation (GPTQ).
Main results (verified)
- 2 bits per parameter, after enough exposure.
- With 1,000 exposures, the peak capacity ratio was at least 2 across all GPT-2 sizes, depths, widths and data sizes (N = 10K–10M), and never above 2.3 (Result 1). Models with a maximum possible ratio up to 1.8 stored nearly everything: a dataset of B bits needs a model of about B/1.8 parameters.
- The same held across bioD's hyperparameter ranges (K and C 1–50, D 10–10,000, L 1–50, T 20–40,000) (Result 3) and for bioR, slightly lower because LLaMA2 adds irrelevant detail (Result 2).
- Their extrapolation: a 7B model could hold 14B bits, more than they estimate English Wikipedia (4.5B words) and all English textbooks (under 16B words) contain.
- Rare facts are stored at half the rate. With only 100 exposures, the capacity ratio falls to about 1 bit per parameter (Result 4). Facts seen less than once are effectively not learned: junk persons seen 0.05–0.2 times each contributed negligibly (footnote 20).
- The knowledge is extractable, not just memorised text. Fine-tuning with LoRA on question–answer pairs for half the people and testing on the other half recovered the stored facts; accuracy fell by only 1.2× for models exactly at the capacity boundary (Appendix A.2). Rewriting the same facts in diverse wording does not cost capacity and, in their earlier work, is what makes knowledge extractable at all (one fixed wording is memorised but nearly 0% extractable).
- Architecture matters only when training is short.
- At 1,000 exposures every architecture reached about 2 bits per parameter, even GPT-2 with no MLP layers: attention layers store knowledge too (Result 5).
- At 100 exposures, LLaMA and Mistral were 1.3× worse than GPT-2 even with tuned learning rates, because of the gated MLP; removing all MLPs cost more than 1.5×; a quarter-sized MLP cost nothing (Results 6–7).
- Quantisation: int8 kept the full capacity; int4 cut it by more than 2× (to about 0.7 bit per parameter) (Result 8).
- Mixture of experts (32 experts, top-1): only 1.3× (1,000 exposures) or 1.5× (100 exposures) below dense capacity per total parameter, while using 8.8% of parameters per token (Result 9).
- "Junk" data sharply reduces what is learned from useful data.
- With 7/8 of training tokens from random junk biographies, capacity for the useful data dropped 20× at 100 exposures, and still 3×, 1.5× and 1.3× at 300, 600 and 1,000 exposures (Result 10).
- Highly repetitive junk (a small set of facts repeated) did not hurt (Result 11).
- Prepending a special token to every useful document (like "wikipedia.org") cut the loss from 20× to 2× at 100 exposures, and at 300 exposures matched the no-junk law (Result 12). The model learns on its own which source is knowledge-rich; adding unique tokens to junk instead works too but is impractical.
- Where knowledge lives: not in one layer. Removing the last layer of an L-layer model near capacity removed much more than 1/L of the knowledge (an unpublished probing observation the authors mention).
- Cost scale: GPT2-20-16 on bioS(10M) for 1,000 exposures took 8.5 days on 64 A100s.
Limits
- Synthetic, uniform, independent facts. Real knowledge is correlated, unevenly frequent and mixed with competence; the 2-bit figure is an upper-bound style result for "sufficiently trained" models, not a measurement on natural text.
- Only knowledge tuples; nothing on how capacity is shared between knowledge and skills (grammar, reasoning, generic schemas) in the same weights.
- Capacity is measured with a lower bound; the junk and architecture comparisons deliberately tuned negative results harder than positive ones (stated by the authors).
- Single-author-group work on its own framework; arXiv v1.
What it means for the vocabulary engine (inference)
- Size bounds what a model can know. At 2 bits per parameter, a 125M model holds at most about 0.25B bits; a 1B model at most about 2B bits; 7B at most about 14B bits. A small base cannot hold "superhuman knowledge" on the scale of Wikipedia even if its corpus contains it. Capacity is a lever, but a weak one for our purpose: even 125M parameters can hold millions of facts.
- Exposure count is the stronger lever. Facts the model sees about 1,000 times are stored at full rate, 100 times at half, under one time almost not at all. A corpus engineered so that specific facts (names, dates, events, opinions tied to people) are rare and diverse, while general patterns repeat endlessly, should yield competence with little fact storage. This is untested in the paper for natural text (an inference to test).
- The flip side of Result 12: the model learns to prioritise knowledge from tagged sources. Tagging is therefore a two-edged tool: it can raise knowledge capture from tagged data, and a tag that marks "fact-heavy" text could also be used to steer what is learned. Whether a tag can make facts recallable only with the tag is not tested (open; see Gao et al. 2025 on metadata conditioning).
- Repetition-heavy filler is harmless to capacity (Result 11), so a base corpus dominated by repetitive generic material does not crowd out what the slot or a later stage must learn.