Kurisutina

Bridging the data gap between children and large language models

Trends in Cognitive Sciences 27(11):990–992 (2023), ScienceDirect PII S1364661323002036 (the journal version was not read, and its DOI was not checked). Read here: the author's PsyArXiv preprint (psyarxiv.com/qzbgx, file "Bridging the data gap (TiCS 2023) r1.pdf", i.e. revision 1 of the journal manuscript). Stanford University. Read by researcher R4a (Claude Opus 5.5) for research batch R4, 5 October 2026. Provenance: papers/base/frank2023_data_gap.provenance.json.

What was read

  • Read in full: all 199 lines of the pdftotext -layout text of the 5-page preprint: abstract, the three sections, Figure 1 (a log-scale plot; its axis labels and caption are in the text), acknowledgements and the 12 references.
  • Conversion artefact: pdftotext drops superscripts, so "1011" in the text is 10^11, "106 words per month" is 10^6, and so on. The exponents below were restored from context and checked against Figure 1's axis (10^6 to 10^12).
  • Not read: the published journal version (paywalled), which may differ in wording from revision 1.

What it is

An opinion and estimation piece (a Forum article), not an experiment. It estimates how much language a child receives, compares that with LLM training sets, and lists explanations for the gap.

Numbers (verified)

  • LLM training sets (tokens). GPT-3 5×10^11; Chinchilla 10^12; a leaked figure for one industry model 3.6×10^12.
  • Child input, upper bound (language produced by the people around the child): about 10^6 words per month (Roy et al. 2015; Dupoux 2018).
    • Age 5: 6×10^7 words. Age 20: 2×10^8 words.
    • Reading adds about 10^7 words a year (2–3 books of 10^5 words a week for 10–15 years), so a literate 20-year-old could reach about 4×10^8 words, "or even higher if they read constantly".
  • Child input, lower bound (limited-language environments): about 10^5 words per month, "up to an order of magnitude less" (Bergelson et al. 2019): about 6×10^6 words by age 5 and 3×10^7 by age 20, without literacy.
  • The gap. "Up to five orders of magnitude" between LLMs and children, and "at least three" between LLMs and the most literate adults. Figure 1's caption says 3–5 orders of magnitude.
  • The competence claim (author's judgement, 2023). Even a child receiving little language can learn a new board game, whereas "language models trained on human-like amounts of data can at best provide incoherent 'autocomplete'-like behaviors, with no in-context learning".

Explanations offered (none tested here)

  1. Prior structure: innate "core knowledge" of objects, agents and events, or architectural constraints that make fundamental structures emerge quickly.
  2. Grounding: multimodal sensory experience gives words concrete meanings; LLMs induce world knowledge from one stream of mostly language.
  3. The kind of input: interactive, social, partly simplified language in which the child takes part. The author suggests RLHF ("interaction training") may be why chat products converse well.
  4. Evaluation differences: LLMs are tested on complex reasoning, children on simple, scaffolded tasks; synchronized evaluation is needed (MEWL, the Baby Intuitions Benchmark).

Data remarks

  • BabyLM trains on curated 10^7- and 10^8-word sets. Holding data constant isolates architecture, but "without high-quality training data – for example, coherent, interactive dialogue about the here-and-now – even the best architecture might fail".
  • CHILDES is "too small to train an LLM"; multimodal datasets are smaller still and hold only vision and language.
  • TinyStories (LLM-generated child-appropriate stories) is named as a workaround; generating multimodal data the same way is suggested.

Limits

  • Order-of-magnitude estimates from three cited sources; no new data. The bounds count words heard (including overheard speech for the upper bound), not words attended to or understood.
  • English-speaking, mostly North American sources for the input estimates.
  • The "autocomplete" judgement about human-scale models predates the BabyLM results and is not backed by a measurement in this piece.

What it means for the base (inference)

  • Human-scale budget for B1: a person reaches adult language competence on roughly 3×10^7 to 4×10^8 words. A vocabulary engine trained on 10^8 words is inside the human range; one trained on 10^11 or more is 3 or more orders of magnitude beyond any person, which is where superhuman knowledge comes from.
  • The gap is not only data volume. If prior structure, grounding and interaction explain the human advantage, then a text-only model at human scale will be weaker than a person at the same input. The base may need more text than a person heard, which reopens the leak of person content that the human budget was meant to limit.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.