Kurisutina

Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora

Proceedings of the BabyLM Challenge at CoNLL 2023, pages 1–34, doi 10.18653/v1/2023.conll-babylm.1 (ACL Anthology 2023.conll-babylm.1). ETH Zürich, Northeastern, Technion, MIT, IBM, MLCommons, Meta AI, UW, NYU. A later arXiv copy exists (2504.08165, not read). Read by researcher R4a (Claude Opus 5.5) for research batch R4, 5 October 2026. Provenance: papers/base/warstadt2023_babylm_findings.provenance.json.

What was read

  • Read in full: all 2,215 lines of the pdftotext -layout text of the 34-page Anthology PDF: sections 1–9, author contributions, references, and appendices A–G (data sources, evaluation data counts, BLiMP-supplement examples, subtask results for MSGS, BLiMP-supplement and (Super)GLUE, age-of-acquisition results, and a summary of every submission).
  • Figures are images: Figures 2–8 were read only through their captions and extracted labels. The BLiMP subtask table that section D.2 refers to ("Table ??") is not in the PDF.
  • Not read: the submissions themselves, the leaderboard, the call for papers.

What the challenge was

  • Tracks. Strict: train on the released 100M-word corpus only. Strict-Small: the released 10M-word subset only. Loose: 100M words of language plus unlimited non-linguistic data (audio, code, music, vision). Any language used to train auxiliary models (tokenizers, parsers, teachers) counted toward the budget; repeated epochs did not; text generated by a model trained only on the BabyLM corpus did not.

  • Why 100M. Children hear 2M–7M words a year (Gilkerson et al. 2017), counting overheard speech. Cut at age 12, that is 24M–84M words, rounded up to 100M. The 10M set corresponds to "the first two to five years". Figure 1, from Warstadt and Bowman (2022): a 12-year-old under 100M words; BERT 3B; RoBERTa 30B; GPT-3 200B; Llama 2 2T.

  • The corpus (Table 1; words in Strict-Small / Strict; share).

    Source Domain 10M 100M Share
    CHILDES (AO-CHILDES: American English, ages 0–6, child utterances removed) Child-directed speech 0.44M 4.21M 5%
    British National Corpus, dialogue portion (British English, second half of the 20th century) Dialogue 0.86M 8.16M 8%
    Children's Book Test (Gutenberg children's books) Children's books 0.57M 5.55M 6%
    Children's Stories Text Corpus (Gutenberg) Children's books 0.34M 3.22M 3%
    Standardized Project Gutenberg (English, authors born after 1850) Written English 0.99M 9.46M 10%
    OpenSubtitles (English) Film and TV subtitles 3.09M 31.28M 31%
    QED (volunteer subtitles of educational videos) Educational subtitles 1.04M 10.24M 11%
    Wikipedia (English, 2022-12-20 dump) Encyclopedia 0.99M 10.08M 10%
    Simple English Wikipedia (2022-12-01 dump) Encyclopedia, simplified 1.52M 14.66M 15%
    Switchboard Dialog Act Corpus (phone calls between strangers) Dialogue 0.12M 1.18M 1%
    Total 9.96M 98.04M 100%
    • About 56% is transcribed or scripted speech; about 40% is child-directed or child-appropriate.
    • "Fewer than 10M words of transcribed child-directed speech are available." Child-directed speech varies "by a factor of 10 or more across cultures and socio-economic groups" (Cristia et al. 2019).
    • Minimal preprocessing; splits about 83.3/8.3/8.3%; the 10M set is a random sample of the 100M set. Download about 240 MB zipped.
  • Evaluation. BLiMP (zero-shot grammatical minimal pairs); a new BLiMP supplement (hypernyms, subject–auxiliary inversion, turn-taking, question–answer congruence easy and tricky); a fine-tuned subset of (Super)GLUE; MSGS (whether a fine-tuned model generalises by linguistic or surface features; Matthews correlation, +1 linguistic, −1 surface); optional age-of-acquisition prediction. Aggregate: BLiMP plus supplement 50%, (Super)GLUE 30%, MSGS 20%. Task examples containing words seen fewer than twice in the 10M corpus were removed.

  • Baselines (20 epochs, context 128): OPT-125M, RoBERTa-base, T5-base. Skylines: Llama 2 70B (2T tokens; (Super)GLUE by in-context learning, MSGS fine-tuned) and the fully trained RoBERTa-base.

  • Participation. 31 papers, 162 models (Strict-Small 118 models from 29 participants; Strict 24 from 11; Loose 20 from 8), from universities or independent institutions in 16 countries.

Results (verified)

  • Top systems (Table 2; BLiMP, GLUE, MSGS, supplement, aggregate).

    Track System BLiMP GLUE MSGS Supp. Agg.
    Skyline Llama 2 (2T tokens) .84 .84 .26 .75 .71
    Skyline RoBERTa-base (30B words) .87 .79 .24 .76 .70
    Strict (100M) ELC-BERT .85 .78 .47 .77 .74
    Strict BootBERT .86 .79 .28 .72 .70
    Strict Best baseline (OPT-125M) .75 .70 .13 .68 .60
    Strict-Small (10M) ELC-BERT .80 .74 .29 .67 .66
    Strict-Small Best baseline (OPT-125M) .63 .62 .10 .53 .50
    Loose Contextualizer .86 .73 .58 .63 .73
  • 100M words reached the skylines on the aggregate. The LTG-BERT-based winners beat Llama 2 and RoBERTa-base on the aggregate (.74 against .71 and .70), on MSGS and on the supplement. On BLiMP, RoBERTa-base stays slightly ahead (.87 against .85–.86); on (Super)GLUE, Llama 2 does (macro .83 against .78 for ELC-BERT, Table 10). The best BLiMP score is "only about 3% shy of human performance".

  • 10× more data helped less than expected. Strict models "did not outperform those in Strict-Small by a large amount"; only two Strict models beat the best Strict-Small model on GLUE.

  • Extra modalities did not help. Loose submissions did worse in aggregate than Strict-Small ones. Vision–language co-training (FLAVA on Wikipedia with images) gave a slight grammar gain at 10M words and none at 100M. A text–audio model beat its text-only twin on 11 of 17 grammatical tasks but was "reported to be undertrained".

  • What worked. Architecture (LTG-BERT and its ELC-BERT variant), training objectives (distillation from teachers trained on the same corpus; Baby Llama's 58M student beat its 300M and 700M teachers), data formatting (sentences rather than documents, no sequence packing, shorter context: 32 tokens beat 128), data cleaning ("the clearest improvements come from cleaning the data of incoherent, ungrammatical, or non-linguistic strings"), and many epochs.

  • Epochs. ELC-BERT trained for over 450 epochs (100M) and over 2,000 epochs (10M); "the winning submission trained on about as many samples as BERT, despite having a training set only about 3% as large." Others found gains up to 10, 60 and 200 epochs. The organisers call hundreds of epochs "not cognitively plausible".

  • What did not work. Curriculum learning was the most popular approach (13 teams, 41.9%) and gave no consistent improvement; CLIMB's exhaustive comparison found none. One team linked performance to the amount of transcribed speech in the data.

  • Subtasks.

    • MSGS was "largely negative": at this scale models prefer surface features. ELC-BERT (macro 0.10 in Strict) and Contextualizer (0.24) were exceptions; Llama 2 scored −0.24 and RoBERTa −0.37. Earlier work found RoBERTa-like models need over a billion words to reach an overall linguistic bias.
    • Hypernyms: every model, skylines included, is near chance (.45–.50).
    • Turn-taking: ELC-BERT (Strict) .92 against Llama 2 .83 and RoBERTa .73, which the authors partly attribute to the corpus's large share of dialogue.
    • Age of acquisition: no Strict-Small submission beat the OPT-125M baseline (mean absolute deviation about 2.03–2.07 months).
    • Better scores on the challenge tasks went with worse prediction of human reading difficulty (Steuer et al.).
  • LLM-made data. Baby's CoThought used GPT-3.5-Turbo to rewrite unrelated corpus sentences into coherent paragraphs; BLiMP improved, but because the LLM saw far more than 100M words, the entry qualified for no track.

Limits

  • Grammar, (Super)GLUE classification after fine-tuning, and inductive bias only. Nothing measures opinions, values, persona, demographic defaults or factual holdings. Zero-shot (Super)GLUE was at or below chance for the baselines, so understanding was only tested after task fine-tuning.
  • Task filtering by the corpus vocabulary means scores are not comparable with prior full-dataset evaluations.
  • Single runs per submission; one entrant reports the RoBERTa baseline may not replicate across seeds.
  • Inconsistencies found.
    • In Table 2, the Strict-Small "McGill-BERT" row (.75, .70, .13, .68, .60) is identical to the Strict "Best Baseline (OPT-125M)" row: a possible copy error.
    • Table 2's task scores differ from the macro averages in Tables 9–10 by 0.01 in several cells (Llama 2 GLUE .84 against .83; RoBERTa .79 against .78), probably rounding against truncation.
    • Section D.2 cites "Table ??" and an empty citation ("?").
    • Section 7.2 says the LTG-BERT models beat both skylines "on all test suites except for (Super)GLUE"; Table 2 shows RoBERTa-base ahead on BLiMP (.87 against .85 and .86).
  • Not an inconsistency: Appendix A gives CHILDES as "about 5M words" and the BNC dialogue as "about 10M"; Table 1's 4.21M and 8.16M are the training splits (about 83%).

What it means for the base (inference)

  • Competence at the human budget is real for grammar. A 100M-word, mostly spoken corpus gives grammar at about skyline level. It is the closest published evidence that B1's language part does not need web-scale data.
  • The corpus is not person-free. By share: 31% film and TV subtitles (characters' opinions and the cultural defaults of film), 25% Wikipedia and Simple Wikipedia (encyclopedic factual holdings), 10% Gutenberg authors born after 1850, 8% BNC dialogue (British adults talking), 1% Switchboard. None of this was measured for opinion content; the challenge never asked.
  • Many epochs on a small corpus is the working recipe, which means every sentence in it is seen hundreds of times. Whatever person content the corpus carries is rehearsed, not diluted.
  • The evaluation gap is ours to fill: BabyLM's battery says nothing about P1. A B1 battery needs empty-slot opinion probes and bounded-knowledge probes beside BLiMP.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.