Transactions on Machine Learning Research (08/2024); arXiv:2405.09673v2 (20 September 2024). Columbia University and Databricks Mosaic Research. Read by researcher R4b for research batch R4 (how to train the vocabulary engine), 5 October 2026. Provenance: papers/base/biderman2024_lora_forgets_less.provenance.json.
What was read
- Read in full: every line of the pdftotext conversion (39 pages): main text, references, Appendices A (setup and hyperparameters), B (learning-rate and α sweeps), C (Tülu-v2-mix), D (supplementary Tables S1–S13), E (SVD figures, captions), F (pass@k), G (dataset examples), H (theoretical memory, Tables S14–S16), I (throughput and memory measurements).
- Not read: figures as images; numbers below come from text and tables.
What they did
- Question: does low-rank adaptation (LoRA: a frozen weight matrix plus a trained low-rank product) match full fine-tuning on hard domains, and does it forget less of what the base could do?
- Base model: Llama-2-7B. Two domains × two regimes:
- continued pre-training (CPT): StarCoder-Python (up to 20B tokens) and OpenWebMath (14.7B tokens, repeated to 20B);
- instruction fine-tuning (IFT): Magicoder-Evol-Instruct-110K (73M tokens) and MetaMathQA (103M tokens), 1–16 epochs.
- Arms: full fine-tuning against LoRA with rank 16, 64 or 256 on all attention and MLP matrices, α = 2r, each with a learning-rate sweep. 32 H100s, LionW optimiser.
- Learning: HumanEval pass@1 (code), GSM8K (math). Forgetting: average of HellaSwag, ARC-Challenge and WinoGrande (base model about 0.65).
- Also: comparison with weight decay and dropout as regularisers; output diversity; singular-value analysis of the full fine-tuning update; a general chat mixture (Tülu-v2-mix); memory and throughput.
Main results (verified)
- LoRA learns less, most clearly in continued pre-training.
- Code CPT: the best LoRA (r = 256) reached HumanEval 0.224 after 20B tokens, about what full fine-tuning reached after 4B (0.218); full fine-tuning peaked at 0.263 (Table S1).
- Math CPT: LoRA r = 256 peaked at GSM8K 0.203 (16B tokens), against 0.224 (4B) and 0.293 (20B) for full fine-tuning (Table S3).
- IFT: high rank closes the gap. Code: r = 256 reached 0.498 at epoch 4, full fine-tuning 0.497 at epoch 8; r = 16 and 64 reached 0.358 and 0.417. Math: r = 64 reached 0.624, r = 256 0.634, full 0.642.
- LoRA forgets less, and rank is the dial.
- Code CPT at 20B tokens: forgetting score 0.617 (LoRA r = 256) against 0.545 (full); LoRA r = 16 stayed at 0.635 (Table S2).
- Code IFT at epoch 16: 0.509 (r = 64) against 0.414 (full) (Table S6).
- Math (closer to the base's English pre-training): little difference (CPT 0.616 against 0.613; IFT 0.567 against 0.559).
- IFT forgets more than CPT; code forgets more than math; forgetting grows with training.
- LoRA beats ordinary regularisation: weight decay and attention dropout learned and forgot like full fine-tuning; LoRA r = 256 learned as much while forgetting less (code IFT).
- LoRA keeps output diversity closer to the base model; full fine-tuning collapsed the number of distinct HumanEval solutions.
- Full fine-tuning makes high-rank changes: the update needed to explain 90% of a weight matrix's variance has rank 10–100× typical LoRA ranks, already at 0.25B tokens, and grows with training; MLP updates are higher rank than attention updates.
- On a general chat mixture (Tülu-v2-mix), even r = 16 matched full fine-tuning on MT-bench within one standard error, and forgot less by epoch 6.
- Practical rules: use LoRA for instruction-style data, not continued pre-training; target all modules; r = 256 if memory allows; α = 2r (crucial at high rank); LoRA's best learning rate is about 10× full fine-tuning's (5×10^-5 to 5×10^-4) and LoRA is more sensitive to it.
- Memory (Appendix H): full fine-tuning with Adam needs about 16 bytes per parameter for weights, gradients and two moments in fp32 (112 GB for 7B, before activations); LoRA on 1% of parameters with the frozen base in bf16 needs about 15 GB for 7B. Measured on 8×H100: LoRA about 15% slower per token, peak memory about 40% lower at small batch, about 15% lower at large batch.
Limits
- One base model (7B), two domains; forgetting measured by three multiple-choice benchmarks only.
- The adapters are trained on 10^8–10^10 tokens; nothing about tiny per-person data or many adapters on one base.
- "Forgetting" means the adapted model's loss of base skills; the frozen base itself is unchanged by construction.
- The SVD result shows full fine-tuning finds high-rank changes, not that low-rank solutions do not exist.
What it means for the vocabulary engine (inference)
- A frozen base plus an adapter is the safe way to attach a slot. The base's weights never change, so other slots and the empty-slot behaviour are untouched; within the adapted model, forgetting of base skills is bounded and set by rank.
- The trade-off cuts against "the slot holds the knowledge". Low-rank slots learn less, especially when the new material is large and unlike the base's training data. A person's holdings are small in tokens, which is the regime (instruction-sized data) where adapters do match full fine-tuning; but if a slot must hold a lot of factual content, it needs high rank or an external store.
- Rank is a usable "slot strength" control, matching the architecture's declared slot-strength parameter: more rank, more person, more drift from the base.
- Memory rule for local training: about 16 bytes per trained parameter with Adam in fp32, plus activations. On a 12 GB card, full pre-training is limited to models of a few hundred million parameters; adapters on a frozen base fit much larger bases (my arithmetic).