Microsoft Research. arXiv:2306.11644v2 [cs.CL], 2 October 2023 (preprint). Read by researcher R6 (capability), 6 October 2026, as the source on small narrow-domain models against large general ones in a language domain. Provenance: papers/capability/gunasekar2023_phi1_textbooks.provenance.json.
What was read
- Read in full: every line of the pdftotext conversion (2,310 lines, 26 pages): sections 1–6, references, and Appendices A (more capability examples), B (limitations with failure examples) and C (examples of decontamination pairs).
- Not read: Figures 2.1 and 3.1 as images; their numbers are restated in the captions and text.
What they did
- Question: can data quality ("textbook quality") replace scale for a narrow skill, writing Python functions from docstrings (HumanEval, MBPP)?
- Data:
- CodeTextbook, about 7B tokens: about 6B tokens of The Stack and StackOverflow, filtered for "educational value" by a classifier trained on about 100K GPT-4 labels; plus under 1B tokens of synthetic Python textbooks written by GPT-3.5.
- CodeExercises: about 180M tokens (879.5K exercises) of GPT-3.5 docstring exercises with solutions.
- Models:
- phi-1: 1.3B parameters (24 layers, width 2048). About 8 passes over CodeTextbook, about 51B tokens seen and 770 A100-hours (under 4 days on 8 A100s). Then fine-tuned on CodeExercises for 7 hours.
- phi-1-small: 350M parameters, same pipeline.
- Evaluation:
- HumanEval and MBPP;
- 50 new "unconventional" problems written by a separate team and graded by GPT-4;
- a "strong decontamination" test: retraining after removing up to 40% of the exercises that resemble HumanEval.
Main results (verified)
-
A narrow small model can match or beat much larger general and code models on the narrow benchmark (Table 1, mostly self-reported scores; HumanEval pass@1):
Model Parameters Training tokens HumanEval phi-1 1.3B 7B (≈ 51B seen) 50.6% (MBPP 55.5%) phi-1-small 350M 7B 45% phi-1-base (no fine-tuning) 1.3B 7B 29% StarCoder 15.5B 1T 33.6% StarCoder-Prompted 15.5B 1T 40.8% PaLM-Coder 540B 780B 35.9% GPT-3.5 175B n/a 47% WizardCoder 16B 1T 57.3% GPT-4 n/a n/a 67% -
Data quality changes the curve.
- 350M models trained on unfiltered Stack and StackOverflow saturated at 12.19% after about 200B tokens.
- On the filtered subset they reached 17.68% after 36K steps, and 20.12% with the synthetic textbooks added.
-
A small fine-tune "unlocks" latent skills.
- 180M tokens of simple exercises raised HumanEval from 29% to 51%.
- It also improved tasks absent from the exercises: using PyGame, Tkinter and PyTorch APIs, and chat. The authors read this as "reorganizing and consolidating the knowledge acquired during pretraining".
-
Robustness checks.
- On the 50 new problems (GPT-4 grading), the order matched HumanEval: phi-1 52%, StarCoder 51%, phi-1-small 45%, phi-1-base 37%, CodeGen-Mono-16.1B 38%.
- After pruning up to 354K similar exercises, the retrained phi-1 scored 45.1%, still above StarCoder-Prompted's 41.5%.
-
Limits of the small specialist, in the authors' words:
- "phi-1 lacks the domain-specific knowledge of larger models such as programming with specific APIs or using less common packages";
- its "1.3B parameters trained with only 7B tokens … restricts our model's capacity to manage more complex tasks such as developing an intricate Flask application";
- it is sensitive to prompt length and wording, and to grammatical errors;
- it is weak at counting and spatial reasoning.
- On phi-1-small: it "understands the logic but does not have enough capacity to learn the correct function calls".
Limits
- The benchmarks are narrow (short Python functions), and comparison scores are self-reported.
- The data are distilled from GPT-3.5 and filtered by GPT-4 labels. The small model inherits a large teacher's content and selection. It does not show skill acquired from scratch at small scale.
- Generation details of the synthetic data are withheld "for proprietary reasons". Emergent-capability claims rest on selected qualitative examples.
What it means for Kurisutina (inference)
- Narrow skill fits in about 10⁸–10⁹ parameters; breadth and robustness do not.
- A 1.3B (even a 350M) specialist matches 15–540B general models on its narrow task, but lacks their library knowledge and robustness to how the task is posed.
- The size of a large model buys breadth, robustness and content across many domains. A human expert at "GLM 5.3 level in one skill" is matched in that skill by a narrow model, not in everything.
- A small fine-tune can surface skill already latent in the base (exercises unlocked API use present only in pretraining). It cannot supply content the base never saw.
- In slot terms: a slot can shape and organise what the base holds cheaply.
- Adding large domain content (APIs, facts) needs either base exposure or a larger store or adapter. This matches the "content must come from somewhere" side of the working answer.
- Teacher distillation is how small models got there. For Amadeus, an expert slot could be trained on teacher-generated domain material, with existing models as instruments, plus the person's own behaviour. The person's level and errors must still come from the person (Maia).