Kurisutina

TinyStories: How Small Can Language Models Be and Still Speak Coherent English?

arXiv 2305.07759 v2 (24 May 2023; dated "April 2023" on the title page). Microsoft Research. Preprint; no journal version was read. The dataset and models are on Hugging Face (not read). Read by researcher R4a (Claude Opus 5.5) for research batch R4, 5 October 2026. Provenance: papers/base/eldan2023_tinystories.provenance.json.

What was read

  • Read in full: all 1,612 lines of the pdftotext -layout text of the 27-page PDF: sections 1–8 and the references, including every completion table that extracted as text (Figures 1, 2, 5–13, 18, 21, 22, 24).
  • Figures that are only images were read through their captions: Figure 3 (loss and scores during training), Figure 4 (the table of scores by hidden size and depth), Figures 14–17 (overlap histograms), 19–20 (attention maps), 23 (scaling plot). The numbers in Figure 4 are not in the text and were not read.
  • The paper does not state the dataset's size (number of stories or tokens) anywhere in the text. Any size figure comes from the dataset page, which was not read.

Question

Is small models' failure to write coherent English due to their size, or to the breadth of their training corpora? Can a narrow corpus that keeps grammar, vocabulary, facts and reasoning let very small models speak coherently?

Method

  • The corpus. Short stories generated by GPT-3.5 and GPT-4, instructed to use "only very simple words that a 3 year old child would likely understand". Diversity came from a list of about 1,500 basic words (nouns, verbs, adjectives "of a typical 3-4 year-old child"): each prompt demanded one random noun, verb and adjective, plus a random subset of story features (dialogue, plot twist, bad ending, moral value). The authors: "prompting those models to produce stories, even if the temperature of generation is set to a high value, will still produce a very repetitive dataset". The corpus "aims to span the vocabulary and the factual knowledge base of a 3-4 year old child".
  • TinyStories-Instruct: each story preceded by instructions (words, a sentence, features including foreshadowing and conflict, and a 1–2 line summary written by GPT-3.5). An OOD variant never pairs summaries with word lists.
  • Models. GPT-Neo architecture (local attention window 256, context 512), GPT-Neo tokenizer cut to the 10K most common tokens. About 1M to 35M parameters for most analyses (largest about 80M), 1 to 8 layers (12 in the scaling runs). "All of the models can be trained on a single V100 GPU within at most 30 hours."
  • GPT-Eval. GPT-4 grades completions of about 50 hand-written story openings (10 completions each at temperature 1) for grammar, creativity, consistency and a guessed author age, "as if those were stories written by students". There is no human evaluation and no standard benchmark.

Results (verified)

  • Coherence at tiny scale. A 28M-parameter model writes a coherent, consistent continuation where GPT-2 XL (1.5B, trained on web text) drifts (Figure 1). Even 2.5M-parameter and one-layer models produce mostly sensible text.
  • Order of emergence: grammar plateaus first, then consistency, then creativity. Consistency appears when the hidden size goes from 64 to 128. The roughly 80M model reaches "almost perfect" grammar and consistency but "falls short of GPT-4's abilities in terms of creativity quite significantly".
  • Width holds facts, depth tracks context. A one-layer model gets some factual prompts right and no context-tracking prompts; a model of embedding size 64 gets no facts right but sometimes keeps context.
  • Child-level knowledge, unevenly. On "What language do they speak in France?", the 8.3M model answers French, the 28M model "English", the 21M one-layer model "Spanish", the 33M four-layer model French. "Can cows fly?" is answered correctly by the 28M and 33M models.
  • Instruction following needs at least 2 layers; the OOD Instruct model followed the unseen combination of summary and words.
  • Not copying. Rouge overlaps with the training stories are low and most generated 4- and 5-grams never appear in the training data. The authors cannot rule out "complex template matching".
  • Scaling: a polynomial relation between compute-optimal model size and training FLOPs, as in larger models (few points). More heads help when heads are few (Figure 24: 2-layer, width 768: loss 1.38, 1.34, 1.33 at 2, 4, 8 heads).
  • Interpretability: in small models, attention heads separate into distance-based and semantic heads, and MLP neurons fire on interpretable roles (subject pronouns, actions, adjectives, the protagonist's first mention). Neurons in GPT-2 XL had no apparent role.

Limits

  • Every training token, and every grade, comes from GPT-3.5 or GPT-4. The generators decide the stories' content, plots and morals; the evaluator is the same family. Nothing measures what values, social scripts or defaults the corpus carries.
  • "Moral value" is an explicit prompt feature. The completions in the figures are overwhelmingly prosocial (sharing, apologising, forgiving), and the stock phrase "I'm glad I could help" recurs across models in Figures 6 and 7.
  • No dataset size, no training-token counts, no epochs or learning rates in the text; reproducibility depends on the released artefacts.
  • Only one genre (simple stories) and one register; no test of general language competence, comprehension benchmarks or anything beyond story continuation.
  • Duplicate sentence in section 4.1 ("For the sake of replicability …" twice); cosmetic.

What it means for the base (inference)

  • Breadth, not just size, is what small models choke on. A narrow, simple corpus lets models of 10–30M parameters speak coherently. For B1 this is evidence that competence can be had from a deliberately narrow corpus, which is also how one would keep person content out.
  • But this corpus is a distillation of an assistant's worldview. A synthetic corpus written by a chat model carries that model's values and register ("I'm glad I could help", a moral in every story) into every weight of the base. Under P1 that is a trained character entering through the data, the exact thing the base must not hold.
  • The method is usable, the generator is the problem. Word-seeded generation solved diversity. A B1 corpus built this way would need a generator whose own opinions were measured, prompts with no moral or evaluative features, and an opinion audit of the output, none of which the paper does.
  • The children's-knowledge framing is an analogy, not a measurement. "The factual knowledge base of a 3-4 year old" was a design aim, not a tested property; the models' facts are patchy.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.