Kurisutina

Humans store about 1.5 megabytes of information during language acquisition

Francis Mollica (Rochester) and Steven T. Piantadosi (Berkeley). Royal Society Open Science 6:181393, doi 10.1098/rsos.181393 (open access; PMC6458406). Read by researcher R6 (capability), 6 October 2026, in place of Landauer (1986), Cognitive Science 10:477–493, whose only copies sit behind Wiley's Cloudflare check or a paywall (not circumvented). Landauer's numbers below are as this paper reports them. Provenance: papers/capability/mollica2019_language_megabytes.provenance.json.

What was read

  • Read in full: every line of the full text from the PMC JATS XML (NCBI E-utilities), converted to text: abstract, introduction, results 2.1–2.5, Table 1, discussion, footnotes 1–9, ethics, data, and the reference list.
  • Not read: the electronic supplementary material (Table S1 of assumptions, reviewer comments), Figures 1–3 as images (captions read), and the OSF code.

What they did

  • Question: how many bits of information about their language does an adult English speaker have to learn? The method is theory-neutral Fermi estimation, with a lower bound, a best guess and an upper bound per level.
  • Measure: bits learned = H[R] − H[R|D], the entropy over possible representations before learning minus the entropy after.
  • Levels:
    • phonemes, from the cue distributions of voice onset time, frication and formants;
    • wordforms, from the surprisal of phone sequences;
    • lexical semantics, from WordNet neighbour distances in an assumed semantic space;
    • word frequency, from a new 2AFC experiment (MTurk, N = 251, 190 trials each);
    • syntax, from the Catalan count of binary parses of 111 textbook sentences.

Main results (verified)

Level Lower Best guess Upper (bits)
Phonemes 375 750 1,500
Wordforms 200,000 400,000 640,000
Lexical semantics 553,809 12,000,000 40,000,000
Word frequency 40,000 80,000 120,000
Syntax 134 697 1,394
Total 794,318 12,481,447 40,762,894
Per day over 18 years 121 1,900 6,204
  • About 12.5 million bits, about 1.5 MB. Lexical semantics dominates. Syntax is only hundreds of bits, yet it still picks one of about 2^697 ≈ 10^210 possible systems.
  • The assumptions behind the numbers:
    • a lexicon of 40,000 words or idioms;
    • 10 bits per wordform (from 43, 33, 24 and 16 bits under 1- to 4-phone models);
    • 0.5–2.0 bits per semantic dimension, with 300 dimensions in the best guess and 500 at 2 bits in the upper bound;
    • word frequency: people were 76.6% accurate, which fits about 4 stored levels, so 2 bits per word.
  • Landauer (1986), as reported here. He estimated memory from behaviour, for example by converting recognition accuracy to bits.
    • About 10–14 bits per remembered picture, and 30–50 bits per known word in a phonetic code.
    • All his methods converged on a "functional capacity" of about 10^9 bits.
    • Hunter (1988) criticised it, and Landauer replied (1988).
    • Later large picture studies fit his per-item range: Brady 2008 estimated 2^13.8 unique items; Ferrara 2017 estimated 2^10.43 in children of 4–6.
  • Neuroanatomical upper bounds cited in the paper: 10^13 bits from the number of cortical synapses, 10^20 bits from lifetime neural impulses.
  • The authors' conclusion: "learners possess remarkable inferential mechanisms", extracting about 2,000 bits of language information a day for 18 years.

Limits

  • Orders of magnitude only: each level rests on strong simplifying assumptions, and the semantic estimate depends on an assumed dimensionality of 100–500.
  • Not estimated: word predictability, pragmatics, discourse relations, prosody, and "models of individual speakers and accents". The total is therefore a lower bound on what a speaker knows about language. World knowledge, skills and episodic memory are excluded altogether.
  • Language knowledge is not language processing. The bits say what is stored, not what computation uses it.
  • The word-frequency experiment is small and uses objective rankings. The authors say this underestimates resolution.

What it means for Kurisutina (inference)

  • The culture layer's language content is small. A best guess of 1.25×10^7 bits, up to 4×10^7, would fit in roughly 6–20M parameters at Allen-Zhu's 2 bits per parameter (summaries/base/allenzhu2024_knowledge_capacity.md). This assumes enough exposure; at about 100 exposures the figure is 1 bit per parameter, so double it. Language storage therefore does not set the engine's size. Processing, breadth of world knowledge, and the parts of language this paper leaves out do.
  • A whole person's learned content may be of the order of 10^9 bits (Landauer, secondhand). That is 5×10^8 to 10^9 parameters at 1–2 bits each. This is the size of everything a person learned, mostly shared with the culture. The divergences that belong in a slot are a fraction of it.
  • The synapse bound, 10^13 bits, is four orders above the behavioural estimate. A brain does not store information at anything near its synaptic capacity; most of the substrate does something other than storing retrievable bits. "Humans are small, yet match large models" is therefore not a comparison of storage. Models are smaller than brains in parameters, and both hold far less retrievable information than their parameter counts allow.
  • Learning rate reference: about 120–6,000 bits a day, best guess 1,900, of language knowledge during childhood. This is one of the few quantitative rates against which a replica's "keeps learning at the person's rate" could be sized. It covers acquisition, not adult learning.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.