arXiv:2501.01956v3 (27 June 2025), Princeton Language and Intelligence (ICML 2025 version). Read by researcher R4b for research batch R4 (how to train the vocabulary engine), 5 October 2026. Provenance: papers/base/gao2025_metadata_conditioning.provenance.json.
What was read
- Read in full: every line of the pdftotext conversion (21 pages): main text, limitations, references, Appendices A (hyperparameters, model configurations, cooldown details, datasets, compute, topic prompt, customised URLs), B (cross-document attention, run variance, cooldown length), C (top DCLM URLs), D (full result Tables 16–21).
- Not read: figures as images; Figure 6 (toxicity bars) could not be parsed reliably, so no toxicity numbers are taken from it.
What they did
- Method (MeCo). For the first 90% of pre-training, each document is prefixed with its source,
URL: en.wikipedia.org\n\n[document]; the loss is computed only on document tokens. For the last 10% ("cooldown"), training continues on different documents with no metadata, inheriting the learning-rate schedule and optimiser state, so the model also works without metadata. - Models: Llama-architecture transformers of 600M, 1.6B, 3B and 8B parameters, Llama-3 tokenizer, AdamW (β2 0.95), LR 3×10^-3 (8B: 5×10^-4), batch 4M tokens, sequences packed to 8,192 tokens, cross-document attention disabled.
- Data: DCLM-Baseline (main), RefinedWeb reproduction, C4; 160B tokens (8B model: 80B).
- Evaluation: OLMES suite, 10 tasks (MMLU, ARC-e, ARC-c, CSQA, HellaSwag, OpenBookQA, PIQA, SIQA, WinoGrande, TruthfulQA), 5-shot unless stated, 1,000 examples per task. One run per configuration; a variance check gives SD 0.1 points on the 10-task average.
- Conditional inference: prepending a real or invented URL to the prompt at test time.
Main results (verified)
- Same performance with a third less data. 1.6B on 160B DCLM tokens: 10-task average 56.7 with MeCo against 55.7 standard; a standard model trained on 240B tokens also reaches 56.7 (Table 1). Validation perplexity did not track downstream performance (13.3 MeCo against 12.9 for the 240B baseline).
- Across scales (Table 17): 600M 51.5 → 51.7; 1.6B 55.7 → 56.7; 3B 59.2 → 59.8; 8B (80B tokens) 57.8 → 58.4.
- Across corpora (Table 18): C4 50.5 → 51.6; RefinedWeb 53.2 → 54.0; DCLM 55.7 → 56.7.
- The cooldown is essential. A model trained 100% with metadata and evaluated without it dropped to 50.3 (ARC-c 28.8 against 42.7). Mixing 90% metadata and 10% plain data throughout gave 56.4; the two-stage schedule 56.7 (Table 4). Cooldown of 10% or 20% worked equally (56.7, 56.8); 30% was worse (56.2).
- Grouping, not meaning, is what helps (Table 5): full URLs 56.8; URL suffix only (".org") 56.2; keeping only the top 0.2% or 2% of URLs 56.4 and 56.3; hashed URLs (random strings) 56.7; model-generated topics 56.6. The authors conclude that metadata helps by telling the model which documents belong together.
- Conditional inference steers the model. With task-appropriate invented URLs, MeCo's average rose from 56.7 to 57.2; the standard model barely moved (55.7 → 55.8). Zero-shot CommonsenseQA: 54.7 unconditional, 53.6 with
boards.4chan.org, 60.9 withwww.factmonster.com. Prefixingen.wikipedia.orgreduced the toxicity of 4,096 unconditioned samples for both models, more so for MeCo (the text says "several-fold"). - Compute (Table 9, H100 GPU hours): 600M/160B tokens 776; 1.6B/160B 1,536 (about 2 days on 32 H100s); 1.6B/240B 2,304; 3B/160B 3,085; 8B/80B 3,905.
Limits
- Single runs; English only; benchmarks are general competence and common-knowledge tasks, not person-level behaviour.
- Metadata here is the source site, shared by thousands of documents. Nothing tests author-level conditioning, conditioning on a held-out source, or what the empty condition represents as a distribution.
- No mechanism: the authors offer two hypotheses (the model learns to prioritise useful sources; or it learns more structured representations).
- Interaction with post-training is not studied.
What it means for the vocabulary engine (inference)
- Source conditioning works at realistic scale and costs nothing, and the cooldown recipe gives exactly the "empty condition" we want: a model that runs without any source tag. That is the closest published analogue to author-conditioned pre-training with an empty slot.
- What the empty condition is. After cooldown, the unconditioned model is trained on the plain mixture of all sources, so its unconditioned output is a model of the pooled corpus (inference). That is a population mixture, which is what P1 asks for, but this paper does not measure its spread or calibration against people.
- Hashed IDs work as well as meaningful ones, so a person or author ID with no semantic content is a usable condition (inference: this is the form a slot key would take). Conditioning on an unseen, invented ID still steered behaviour (via its semantics); an unseen hashed ID was not tested.
- The tag carries behaviour, the weights keep competence. Conditioning moved task accuracy by up to 7 points and toxicity several-fold without retraining. That suggests person-level variation can be pushed into a condition. Whether the shared weights then hold less person-specific content is not measured.
- Measured throughput for costing: 1.6B parameters × 160B tokens ≈ 1.5×10^21 FLOPs (6ND) in 1,536 H100-hours, about 1.0×10^18 FLOPs per GPU-hour at this small-model scale (my arithmetic).