arXiv:2401.06121v1 (11 January 2024), Carnegie Mellon University (later COLM 2024; the arXiv v1 was read). Read by researcher R4b for research batch R4 (how to train the vocabulary engine), 5 October 2026. Provenance: papers/base/maini2024_tofu.provenance.json.
What was read
- Read in full: every line of the pdftotext conversion (24 pages): main text, discussion and limitations, references, Appendices A (KS test), B (hyperparameters), C (the 100 "I don't know" strings), D (continued unlearning), E (sanity checks, Tables 3–4), F (knowledge-entanglement plots, captions only).
- Not read: figures as images; the per-method trajectories in Figures 5–33 exist only as plots, so numbers below come from the text and tables.
What they did
- A clean unlearning task. GPT-4 wrote 200 fictitious author profiles, 20 question–answer pairs each, seeded with attributes (birthplace, gender, year, genre, awards, parents' jobs, book-title words from Goodreads) to avoid stereotyped output. Because the authors do not exist, the fine-tuning set is the only source of the facts.
- Setup: fine-tune Llama-2-7B and Phi-1.5 on all 200 authors (5 epochs, AdamW, LR 10^-5, effective batch 32), then try to make the model forget 1%, 5% or 10% of authors (2, 10 or 20), with compute linear in the forget set.
- Gold standard: a "retain model" fine-tuned without the forget set.
- Forget quality: a Kolmogorov–Smirnov test comparing the distribution of a truth ratio (probability of perturbed wrong answers relative to a paraphrased correct answer) between the unlearned model and the retain model on the forget set. A high p-value means indistinguishable.
- Model utility: harmonic mean of probability, ROUGE-L and truth ratio on the retain set, real-author questions and world facts.
- Four baselines: gradient ascent on the forget set; gradient difference (ascent on forget, descent on retain); KL minimisation to the original model on the retain set; preference optimisation towards "I don't know"-style answers.
Main results (verified)
- Fine-tuning works: ROUGE on TOFU questions rose from 0.364 to 0.985 (Llama-2-7B) and from 0.440 to 0.869 (Phi-1.5) (Table 2).
- No method unlearns.
- Forgetting 1% of the data, all methods move towards better forget quality, but for Phi every p-value stays below 0.001, and for Llama-2-7B forget quality "is never higher than 0.01". Extending to 10 epochs does not cross p = 0.05 (Appendix D).
- Forgetting 5% or 10%, the models that reach high forget quality have lost most of their utility; models start producing gibberish after about two epochs of unlearning ("…Marais Marauders series behind running running running…").
- Unlearning damages neighbouring knowledge in order of closeness. With gradient difference on the 5% set, performance falls fastest on the retained fictitious authors, then on real authors, while world facts stay roughly unchanged ("knowledge entanglement").
- Access to retain data helps (gradient difference beats gradient ascent), but in real settings a suitable retain set is itself hard to find.
- Trajectories are not monotone: methods balancing two losses zig-zag between utility and forgetting.
- The metric is sound but can be fooled: retain models trained on different splits are indistinguishable from each other (p 0.77–0.99, Table 3), and a fine-tuned model is clearly distinguishable from a retain model on the forget set (p 1.1×10^-19 for the 90/10 split). But a broken model that assigns random probabilities can score well on forget quality, so utility must be read alongside it.
- The authors' interpretation: unlearning "requires overfitting"; current methods "modify LLMs just enough to produce slightly different output for specific prompts but they do not remove information or behavior from models on the whole", and they draw the same conclusion for alignment.
Limits
- The facts to forget were learned in a small fine-tuning stage, not in pre-training; forgetting pre-training knowledge is presumably harder (the authors call this a deliberate simplification).
- Only entity-level forgetting; no instance-level or behaviour-level unlearning.
- Four simple baselines from early 2024; later methods are not covered.
- Indistinguishability is approximated by one statistic, not by (ε, δ) guarantees.
What it means for the vocabulary engine (inference)
- Post-hoc removal cannot be the plan for "holds too much". If even 2 synthetic authors learned in a short fine-tune cannot be removed from a 7B model without collateral damage, then removing broad world knowledge, opinions or a persona from a pre-trained model is not a reliable route. Content must be kept out by construction (corpus design, conditioning, separate stores).
- Person content must sit in separable parameters. TOFU is exactly the case of "a person's facts written into shared weights": once written, they cannot be cleanly taken out. A slot that lives in its own adapter or embedding can be deleted exactly by dropping it; a slot merged into the base cannot.
- "I don't know" is not ignorance. Preference optimisation towards refusals changes outputs while the truth ratio still shows the knowledge; the bounded-knowledge test in the base battery (B9) must probe likelihoods or paraphrase ratios, not just generated answers.
- The truth-ratio-plus-KS design is reusable for the battery's "bounded knowledge" test: compare the base against a model that provably never saw a fact.