Proceedings of EMNLP 2024, pages 22055–22071 (ACL Anthology 2024.emnlp-main.1231). Stanford University. Code and data:
github.com/styfeng/TinyDialogues (not read). Read by researcher R4a (Claude Opus 5.5) for research batch R4,
5 October 2026, as a swap for a candidate that could not be obtained (see the provenance file). Provenance:
papers/base/feng2024_child_directed_speech.provenance.json.
What was read
- Read in full: all 921 lines of the
pdftotext -layouttext of the 17-page Anthology PDF: abstract, sections 1–5, limitations, ethics, acknowledgements, references, and appendices A–L (the TinyDialogues prompt and examples, data formats, dataset statistics, training details, Zorro details, significance tests, extra results, speaker labels, convergence, licences, code). - Figures 1–8 (word counts by age, loss curves) were read only through captions and text.
- Not read: the code and the released data.
Question
Is children's language input a uniquely good training signal (its developmental ordering, its conversational coherence), or is the child's learning algorithm what makes children efficient?
Method
- Learners. GPT-2 small (124M, causal) and RoBERTa-base (125M, masked), trained from scratch with a separate tokenizer per dataset. GPT-2: 20 epochs, 3 seeds, LR 1e-4. RoBERTa: 50 epochs, 2 seeds, LR 5e-5. The best epoch by validation loss was kept.
- Five datasets, each about 29M words (85/15 split: about 24.5M training, 4.5M validation words).
- CHILDES, English subset: about 11k conversations; children from birth to 13, but about 90% of the words are for ages 2–5. Mean utterance 3.85 words; about 74.5k unique words. Speaker-label frequencies in the training split (Table 10; apparently counts of turns, the unit is not stated): mother 45.7%, child 38.2%, investigator 4.5%, father 3.9%. The child's own utterances were kept, with speaker labels.
- TinyDialogues (TD), new and synthetic: 129,732 conversations (28.8M words) generated by GPT-4 (gpt-4-1106-preview) with a child of 2, 5, 10 or 15 as the central participant (about 7.2M words per age). Each prompt fixed a conversation type (explanatory, functional, narrative, argumentative: "conflict(s) or disagreement(s) … In most cases, the argument should be resolved, resulting in the [child] learning"), other participants (mom, dad, siblings, teacher, grandparent, neighbour, coach, tutor, boyfriend or girlfriend by age), and one noun, verb and adjective from age-appropriate word lists. Mean utterance 13.42 words; about 96k unique words.
- BabyLM (a 29M-word in-order sample of the 2023 BabyLM corpus, about 443k unique words), Wikipedia (about 12M Wikipedia plus 17M Simple Wikipedia words, about 644k unique), OpenSubtitles (about 301k unique words).
- Evaluation. Zorro (child-vocabulary grammatical minimal pairs, full suite, not vocabulary-filtered) and word similarity (Spearman correlation between model and human similarity on RG-65, WordSim-353, SimLex-999, SimVerb-3500 and MEN, best layer).
- Ordering experiments. Global: conversations in age order, reverse order or random order, trained as "repeated buckets" (each age section repeated n times before the next). Local: utterances within each conversation in original or shuffled order.
Results (verified)
-
Dataset comparison, GPT-2 (Table 1; Zorro accuracy, word similarity).
Data (29M words) Zorro WS CHILDES 78.29% 0.24 TinyDialogues 78.48% 0.42 Wikipedia 78.16% 0.32 OpenSubtitles 81.02% 0.38 BabyLM mix 82.90% 0.42 -
RoBERTa (Table 2): TD best on Zorro (78.52%, against 58.37–62.57% for the rest); Wikipedia best on word similarity (0.34); CHILDES worst on both (58.37%, 0.14).
-
Natural child-directed speech was the weakest training data. It was lowest on both measures for RoBERTa and on word similarity for GPT-2; on GPT-2's Zorro it sat in the low group (78.29%, against Wikipedia 78.16% and TD 78.48%). Synthetic child-directed conversation matched or beat it everywhere. The authors: "child language input is not uniquely valuable for training language models"; "more diverse datasets (e.g. general conversation data or a mixture of different data sources) may result in better learning than homogeneous child-directed data"; "conversational data seems essential for better grammar and syntax learning across architectures".
-
Global developmental ordering did not matter. GPT-2 on CHILDES: age order 75.62%, reverse 77.63%, random 76.87% on Zorro, word similarity 0.19–0.20 (no significant difference). On TD, random order was slightly best. RoBERTa failed to converge under repeated buckets (near chance) and is not interpreted.
-
Local coherence mattered for natural speech. Shuffling utterances within CHILDES conversations lowered GPT-2 word similarity from 0.24 to 0.19 and RoBERTa's from 0.14 to 0.04; TD was barely affected. Removing speaker labels lowered Zorro (CHILDES 78.29% → 76.61%) but not word similarity.
-
More repetitions helped (Table 17, single seed): CHILDES Zorro 68.89% at n = 3 to 77.01% at n = 10; TD 71.51% to 79.65% at n = 20.
-
Conclusion (authors): the results support the view that "rather than proceeding from better data, the child's learning algorithm is substantially more data-efficient than current language modeling techniques"; the source, composition and local properties of data still matter.
Limits
- Two small architectures, one data size (29M words), two benchmarks (Zorro, word similarity). No test of understanding, knowledge, opinions or persona.
- TinyDialogues was checked only for offensiveness. The authors "manually examined a large subset" for "profanities, racism, bias, offensive words", call it free of personal information, and say "as a safe and controlled language model, there is an incredibly low risk of offensive content". Its values, the resolution of its argumentative dialogues and its family scripts are GPT-4's and were not measured.
- CHILDES is skewed to ages 2–5 and short utterances, so the CHILDES-against-TD contrast confounds naturalness with age range, utterance length and transcription noise. The authors name this.
- Statistics. Every Zorro comparison is reported as p = 0, including differences of 0.2 points (78.29% against 78.48%), which suggests per-example outcomes pooled across seeds were treated as independent.
- Inconsistent word-similarity values. Appendix Tables 18–19 report CHILDES word similarity of 0.35 and 0.48–0.52 and TD 0.54 for settings that Tables 1, 3 and 17 report at 0.10–0.42. Unexplained.
- Licences (appendix K): CHILDES is CC BY-NC-SA 3.0 through TalkBank; TD to be released under MIT.
What it means for the base (inference)
- Developmental plausibility is not a reason to pick child-directed speech for B1. At 29M words it was the weakest data, and its developmental order added nothing. The choice of B1 data can be driven by what it leaks, not by its resemblance to a child's input.
- Synthetic conversation works, which is exactly the risk. TD beat natural speech, but every line of it is a GPT-4 completion, including how arguments end and who learns what. That imports the generating model's values and assistant register into the base. A synthetic corpus for B1 would need its generator's opinions measured, or its prompts stripped of anything opinion-bearing.
- Natural child-directed speech is mostly one adult per family talking (mothers 45.7% of CHILDES speaker turns): the speech of research families, with their routines and views. It is person content of a specific population, not neutral input.
- Mixtures beat single sources at this scale, so a curated mixture is the defensible default for B1.