Kurisutina

Textbooks are all you need ("phi-1")

Microsoft Research. arXiv:2306.11644v2 [cs.CL], 2 October 2023 (preprint). Read by researcher R6 (capability), 6 October 2026, as the source on small narrow-domain models against large general ones in a language domain. Provenance: papers/capability/gunasekar2023_phi1_textbooks.provenance.json.

What was read

  • Read in full: every line of the pdftotext conversion (2,310 lines, 26 pages): sections 1–6, references, and Appendices A (more capability examples), B (limitations with failure examples) and C (examples of decontamination pairs).
  • Not read: Figures 2.1 and 3.1 as images; their numbers are restated in the captions and text.

What they did

  • Question: can data quality ("textbook quality") replace scale for a narrow skill, writing Python functions from docstrings (HumanEval, MBPP)?
  • Data:
    • CodeTextbook, about 7B tokens: about 6B tokens of The Stack and StackOverflow, filtered for "educational value" by a classifier trained on about 100K GPT-4 labels; plus under 1B tokens of synthetic Python textbooks written by GPT-3.5.
    • CodeExercises: about 180M tokens (879.5K exercises) of GPT-3.5 docstring exercises with solutions.
  • Models:
    • phi-1: 1.3B parameters (24 layers, width 2048). About 8 passes over CodeTextbook, about 51B tokens seen and 770 A100-hours (under 4 days on 8 A100s). Then fine-tuned on CodeExercises for 7 hours.
    • phi-1-small: 350M parameters, same pipeline.
  • Evaluation:
    • HumanEval and MBPP;
    • 50 new "unconventional" problems written by a separate team and graded by GPT-4;
    • a "strong decontamination" test: retraining after removing up to 40% of the exercises that resemble HumanEval.

Main results (verified)

  • A narrow small model can match or beat much larger general and code models on the narrow benchmark (Table 1, mostly self-reported scores; HumanEval pass@1):

    Model Parameters Training tokens HumanEval
    phi-1 1.3B 7B (≈ 51B seen) 50.6% (MBPP 55.5%)
    phi-1-small 350M 7B 45%
    phi-1-base (no fine-tuning) 1.3B 7B 29%
    StarCoder 15.5B 1T 33.6%
    StarCoder-Prompted 15.5B 1T 40.8%
    PaLM-Coder 540B 780B 35.9%
    GPT-3.5 175B n/a 47%
    WizardCoder 16B 1T 57.3%
    GPT-4 n/a n/a 67%
  • Data quality changes the curve.

    • 350M models trained on unfiltered Stack and StackOverflow saturated at 12.19% after about 200B tokens.
    • On the filtered subset they reached 17.68% after 36K steps, and 20.12% with the synthetic textbooks added.
  • A small fine-tune "unlocks" latent skills.

    • 180M tokens of simple exercises raised HumanEval from 29% to 51%.
    • It also improved tasks absent from the exercises: using PyGame, Tkinter and PyTorch APIs, and chat. The authors read this as "reorganizing and consolidating the knowledge acquired during pretraining".
  • Robustness checks.

    • On the 50 new problems (GPT-4 grading), the order matched HumanEval: phi-1 52%, StarCoder 51%, phi-1-small 45%, phi-1-base 37%, CodeGen-Mono-16.1B 38%.
    • After pruning up to 354K similar exercises, the retrained phi-1 scored 45.1%, still above StarCoder-Prompted's 41.5%.
  • Limits of the small specialist, in the authors' words:

    • "phi-1 lacks the domain-specific knowledge of larger models such as programming with specific APIs or using less common packages";
    • its "1.3B parameters trained with only 7B tokens … restricts our model's capacity to manage more complex tasks such as developing an intricate Flask application";
    • it is sensitive to prompt length and wording, and to grammatical errors;
    • it is weak at counting and spatial reasoning.
    • On phi-1-small: it "understands the logic but does not have enough capacity to learn the correct function calls".

Limits

  • The benchmarks are narrow (short Python functions), and comparison scores are self-reported.
  • The data are distilled from GPT-3.5 and filtered by GPT-4 labels. The small model inherits a large teacher's content and selection. It does not show skill acquired from scratch at small scale.
  • Generation details of the synthetic data are withheld "for proprietary reasons". Emergent-capability claims rest on selected qualitative examples.

What it means for Kurisutina (inference)

  • Narrow skill fits in about 10⁸–10⁹ parameters; breadth and robustness do not.
    • A 1.3B (even a 350M) specialist matches 15–540B general models on its narrow task, but lacks their library knowledge and robustness to how the task is posed.
    • The size of a large model buys breadth, robustness and content across many domains. A human expert at "GLM 5.3 level in one skill" is matched in that skill by a narrow model, not in everything.
  • A small fine-tune can surface skill already latent in the base (exercises unlocked API use present only in pretraining). It cannot supply content the base never saw.
    • In slot terms: a slot can shape and organise what the base holds cheaply.
    • Adding large domain content (APIs, facts) needs either base exposure or a larger store or adapter. This matches the "content must come from somewhere" side of the working answer.
  • Teacher distillation is how small models got there. For Amadeus, an expert slot could be trained on teacher-generated domain material, with existing models as instruments, plus the person's own behaviour. The person's level and errors must still come from the person (Maia).

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.