Francis Mollica (Rochester) and Steven T. Piantadosi (Berkeley). Royal Society Open Science 6:181393, doi 10.1098/rsos.181393 (open access; PMC6458406). Read by researcher R6 (capability), 6 October 2026, in place of Landauer (1986), Cognitive Science 10:477–493, whose only copies sit behind Wiley's Cloudflare check or a paywall (not circumvented). Landauer's numbers below are as this paper reports them. Provenance: papers/capability/mollica2019_language_megabytes.provenance.json.
What was read
- Read in full: every line of the full text from the PMC JATS XML (NCBI E-utilities), converted to text: abstract, introduction, results 2.1–2.5, Table 1, discussion, footnotes 1–9, ethics, data, and the reference list.
- Not read: the electronic supplementary material (Table S1 of assumptions, reviewer comments), Figures 1–3 as images (captions read), and the OSF code.
What they did
- Question: how many bits of information about their language does an adult English speaker have to learn? The method is theory-neutral Fermi estimation, with a lower bound, a best guess and an upper bound per level.
- Measure: bits learned = H[R] − H[R|D], the entropy over possible representations before learning minus the entropy after.
- Levels:
- phonemes, from the cue distributions of voice onset time, frication and formants;
- wordforms, from the surprisal of phone sequences;
- lexical semantics, from WordNet neighbour distances in an assumed semantic space;
- word frequency, from a new 2AFC experiment (MTurk, N = 251, 190 trials each);
- syntax, from the Catalan count of binary parses of 111 textbook sentences.
Main results (verified)
| Level | Lower | Best guess | Upper (bits) |
|---|---|---|---|
| Phonemes | 375 | 750 | 1,500 |
| Wordforms | 200,000 | 400,000 | 640,000 |
| Lexical semantics | 553,809 | 12,000,000 | 40,000,000 |
| Word frequency | 40,000 | 80,000 | 120,000 |
| Syntax | 134 | 697 | 1,394 |
| Total | 794,318 | 12,481,447 | 40,762,894 |
| Per day over 18 years | 121 | 1,900 | 6,204 |
- About 12.5 million bits, about 1.5 MB. Lexical semantics dominates. Syntax is only hundreds of bits, yet it still picks one of about 2^697 ≈ 10^210 possible systems.
- The assumptions behind the numbers:
- a lexicon of 40,000 words or idioms;
- 10 bits per wordform (from 43, 33, 24 and 16 bits under 1- to 4-phone models);
- 0.5–2.0 bits per semantic dimension, with 300 dimensions in the best guess and 500 at 2 bits in the upper bound;
- word frequency: people were 76.6% accurate, which fits about 4 stored levels, so 2 bits per word.
- Landauer (1986), as reported here. He estimated memory from behaviour, for example by converting recognition accuracy to bits.
- About 10–14 bits per remembered picture, and 30–50 bits per known word in a phonetic code.
- All his methods converged on a "functional capacity" of about 10^9 bits.
- Hunter (1988) criticised it, and Landauer replied (1988).
- Later large picture studies fit his per-item range: Brady 2008 estimated 2^13.8 unique items; Ferrara 2017 estimated 2^10.43 in children of 4–6.
- Neuroanatomical upper bounds cited in the paper: 10^13 bits from the number of cortical synapses, 10^20 bits from lifetime neural impulses.
- The authors' conclusion: "learners possess remarkable inferential mechanisms", extracting about 2,000 bits of language information a day for 18 years.
Limits
- Orders of magnitude only: each level rests on strong simplifying assumptions, and the semantic estimate depends on an assumed dimensionality of 100–500.
- Not estimated: word predictability, pragmatics, discourse relations, prosody, and "models of individual speakers and accents". The total is therefore a lower bound on what a speaker knows about language. World knowledge, skills and episodic memory are excluded altogether.
- Language knowledge is not language processing. The bits say what is stored, not what computation uses it.
- The word-frequency experiment is small and uses objective rankings. The authors say this underestimates resolution.
What it means for Kurisutina (inference)
- The culture layer's language content is small. A best guess of 1.25×10^7 bits, up to 4×10^7, would fit in roughly 6–20M parameters at Allen-Zhu's 2 bits per parameter (
summaries/base/allenzhu2024_knowledge_capacity.md). This assumes enough exposure; at about 100 exposures the figure is 1 bit per parameter, so double it. Language storage therefore does not set the engine's size. Processing, breadth of world knowledge, and the parts of language this paper leaves out do. - A whole person's learned content may be of the order of 10^9 bits (Landauer, secondhand). That is 5×10^8 to 10^9 parameters at 1–2 bits each. This is the size of everything a person learned, mostly shared with the culture. The divergences that belong in a slot are a fraction of it.
- The synapse bound, 10^13 bits, is four orders above the behavioural estimate. A brain does not store information at anything near its synaptic capacity; most of the substrate does something other than storing retrievable bits. "Humans are small, yet match large models" is therefore not a comparison of storage. Models are smaller than brains in parameters, and both hold far less retrievable information than their parameter counts allow.
- Learning rate reference: about 120–6,000 bits a day, best guess 1,900, of language knowledge during childhood. This is one of the few quantitative rates against which a replica's "keeps learning at the person's rate" could be sized. It covers acquisition, not adult learning.