Cognitive Science 47(6):e13305, doi 10.1111/cogs.13305. New York University and Boston University. Read from the
copy on Brenden Lake's university homepage (cs.princeton.edu/~bl8144/papers/WangEtAl2023CognitiveScience.pdf). Code:
github.com/wkvong/multimodal-baby (not read). Read by researcher R4a (Claude Opus 5.5) for research batch R4,
5 October 2026, in place of Vong et al. (2024, Science), the same group's study of the same child, which could not be
obtained (see the provenance file). Provenance: papers/base/wang2023_one_child_structure.provenance.json.
What was read
- Read in full: all 939 lines of the
pdftotext -layouttext of the 23-page article: abstract, sections 1–5, notes 1–18, acknowledgements and references. Figures 1–7 were read through their captions and the text; Table 2 (cloze examples) in full. - Not read: the separate supporting-information appendix (A.1–A.8: preprocessing details, architectures, category labelling, cloze construction, the per-phenomenon Zorro results, Table 7, Figures 8–15), and the code.
Question
What can generic distributional learners acquire from a naturalistic subset of one child's language input, and does adding that child's visual input change it?
Data
- SAYCam (Sullivan et al. 2021): head-mounted camera recordings of three English-speaking children, several hours a week for about 2 years from 6–8 months. Only child S had much of his speech input transcribed (6–25 months).
- SAYCam-S: parents' utterances only (the child's own excluded), each paired with frames at 5 fps from the first 6.4 s of the utterance. 37,486 utterances, 249,774 tokens, 600,285 frames. Training split 33,737 utterances and 225,001 tokens; mean utterance 6.67 tokens; vocabulary 2,350 (tokens seen at least 3 times); out-of-vocabulary rate about 2%. Random 90/5/5 split; temporal order ignored.
- The authors estimate the 225K training tokens are 0.5%–4% of the child's input in the first 2 years, assuming 3M–20M words a year (Dupoux 2018).
- Access: SAYCam is on Databrary, available "to academic investigators through the Databrary authorization process".
Method
- Language only: a 1-layer LSTM (512 units) and CBOW (context of one word each side), trained from scratch on next- or missing-token prediction, 3 seeds.
- Vision and language: a captioning LSTM whose hidden state is initialised by a ResNeXt-50 encoder pretrained self-supervised on child S's own video (Orhan et al. 2020).
- Analyses: clustering of word embeddings (t-SNE, dendrograms), a noun-or-verb cloze test (2,412 clozes, 65% verbs), Zorro minimal pairs filtered to the model's vocabulary (15 subsets, 7 phenomena), per-category loss changes with vision, and representational similarity across networks.
Results (verified)
- Perplexity (validation): LSTM 24.80, CBOW 22.20, captioning LSTM 22.10.
- Categories emerge from one child's input. Nouns and verbs form separate clusters, verbs split into transitive and intransitive; adjectives and adverbs also cluster. Nouns cluster by semantic category (animals, body parts, clothing, food and others), "largely following a taxonomic organization mixed with some thematic influences".
- The family's own episodes shape the space. A cluster of "milk", "farm" and "cow" "can be directly traced back to a particular book in the training data", a farm picture book the parent read. In the cloze examples, "we might go to the ___ today" (truth: beach) gives the LSTM's top prediction "library" (61.2%), then "playground" (10.1%) and "beach" (8.8%); "food for us for breakfast" gives "bread" (56.2%).
- Noun–verb cloze: LSTM 97.96%, CBOW 91.20% (base rate 65%).
- Grammar is partial. Determiner–noun agreement: LSTM 67.7%, CBOW 61.1%. Subject–verb agreement: LSTM 55.7%, near chance. Some tests (quantifiers, case) were also passed by unigram and bigram baselines, so they do not discriminate. A Transformer trained on AO-CHILDES (5M words from many children; BabyBERTa) did better.
- Vision adds a little and changes little. Perplexity fell from 24.80 to 22.10, with the largest gains for nouns and verbs (significant for most categories). Representations stayed similar: LSTM against captioning LSTM r = .82; CBOW against either about .70.
Limits
- Scale: 225K tokens, a tiny fraction of one child's input, so complex phenomena cannot be tested. The authors cannot separate "data scale or data diversity due to more children across more ages and more environments" as the reason the multi-child model does better.
- Transcripts, not audio; frames, not video; loose alignment between utterances and frames; many passes over the data, not one.
- Passive learning: "The networks cannot choose their own actions …, do not have desires and goals, do not utilize social cues". The authors stress that clusters are far from human concepts.
- No test of knowledge, beliefs or anything person-level; one child, one family.
What it means for the base (inference)
- One child's stream gives categories, not a vocabulary engine. At realistic densities a single child's recorded input is a few hundred thousand words: enough for word classes and taxonomies, not for agreement, let alone the competence B1 needs.
- It is the most person-laden corpus possible. The model's defaults became this family's: the library trip, the farm book, bread for breakfast. A base trained on one household's speech would carry that household as its default person, which is the opposite of P1. This is also the clearest small-scale illustration of how person content enters a vocabulary: through co-occurrence of routine, not through stated opinions.
- Vision did not restructure language here, so a grounded stream is not a shortcut around the person-content problem at this scale.
- Not available to us without institutional access (Databrary authorisation); the project has none.