Proceedings of NAACL 2024 (Volume 1: Long Papers), pages 3245–3276 (ACL Anthology 2024.naacl-long.179; arXiv
2305.13169, not read). MIT, Cornell, Google Research, OpenAI, Carnegie Mellon. Models and code were not released.
Read by researcher R4a (Claude Opus 5.5) for research batch R4, 5 October 2026. Provenance:
papers/base/longpre2024_pretrainers_guide.provenance.json.
What was read
- Read in full: all 2,045 lines of the
pdftotext -layouttext of the 32-page Anthology PDF: sections 1–8, references, and appendices A–F (literature review with Table 4 of known models' data, curation details and Tables 5–6, training details, evaluation details, data-feature analysis, and the raw results in Tables 11–17 and Figures 8–11). - Figures 2, 6 and 7 (feature ratios) were read from their extracted labels, whose layout is partly scrambled; only values confirmed by the text are used below. Figures 1, 3 and 4 were read through captions and text.
Question
How do three routine data decisions (when the data was collected, which quality and toxicity filters are applied, and which domains are included) change a model pretrained from scratch?
Method
- 28 decoder-only models of 1.5B parameters ("LM-XL", T5X, sequence 512, batch 4,096, 88,064 steps), plus 20M models ("LM-Small") for scale comparisons, each pretrained from scratch on one data variant, then fine-tuned and evaluated.
- Data age: four C4 versions rebuilt from Common Crawl snapshots of 2013, 2016, 2019 and 2022 (246B, 206B, 226B, 360B tokens).
- Filters on C4 and the Pile: the PaLM/GLaM quality classifier (trained to keep text resembling Wikipedia and books) at thresholds 0.975–0.7, and an inverse filter that removes the highest-quality text; the Perspective API toxicity classifier at 0.95–0.3, plus an inverse filter and C4's bad-words n-gram filter.
- Domains: the Pile's 22 sources grouped into 9 domains (Common Crawl 227 GB, OpenWebText 63 GB, Wikipedia 6 GB, Books 118 GB, PubMed 109 GB, Academic 60 GB, Code and Math 135 GB, Legal 74 GB, Social 33 GB: Ubuntu IRC, EuroParl, Enron emails, HackerNews, OpenSubtitles, YouTube subtitles), each removed in turn.
- Evaluation: 27–30 QA datasets grouped by domain, toxicity identification (Social Bias Frames, DynaHate, Toxigen; AUC), toxic generation (RealToxicityPrompts, a representational-bias prompt set; Perspective ≥ 0.5), and five tasks with year-split data (news, Twitter, science, political affiliation).
Results (verified)
- Filters change what a model can recognise, not only what it says.
- Toxicity filtering lowers toxic generation and also the ability to identify toxicity. On C4 at threshold 0.3 (60.8% of data kept): toxic generation −52.1%, identification score −5.2. At 0.5 (75.8% kept): −35.0% and −1.8. QA also drops (−2.1 average at 0.5).
- The inverse toxicity filter (removing the least toxic 8%) gave the best toxicity identification (+1.4) and more toxic generation (+15.6%). The authors recommend that pretraining "target toxicity identification rather than curbing toxic generation", leaving generation to later methods.
- Quality filtering improved QA (+2.5 to +2.7 average at threshold 0.975, keeping 91%) and toxicity identification (up to +1.8), but raised toxic generation (+7.3% to +11.1%). The inverse quality filter (removing the best 27%) hurt QA (−3.1 to −3.3).
- "Toxicity and quality are surprisingly not well-aligned": high-toxicity documents scored higher on quality, largely because Books carries much profane and sexual text.
- Who is in the text. By a basic PII classifier, 37% of C4 documents and 41% of Pile documents contain a person's name; about 3% contain an email, 2–4% an address. Documents classed as high quality in C4 had more names (1.6x and 1.8x, per the text). In the Pile, high-toxicity documents were 1.4–1.9 times more likely to contain PII.
- Data age is persistent. One year of mismatch between pretraining and evaluation data costs about 0.4 points on average (2.8 for fine-tuning mismatch); mean correlation between misalignment and loss r = 0.61. It is "not overcome by finetuning", is asymmetric (worse when evaluation data is newer), and is larger for the 1.5B than the 20M model. A 2022 model did worse on Obama-era evaluations than older models.
- Heterogeneous domains matter most. Removing Common Crawl (−4.9 average QA), Books (−2.8), OpenWebText (−1.4) or PubMed (−1.4) hurt most; removing Social (+0.3), Academic (+0.2) or Code (−0.2) hardly mattered. Removing Wikipedia increased toxic generation (+4.2%). The best models trained on all or nearly all sources. Web and books also contributed most to toxic generation.
- Literature summarised in the appendix (not tested here): Dodge et al. (2021) found C4's filtering left "a dearth of English text from American minority communities as well as from non-Western communities like India or Nigeria"; Gururangan et al. (2022) found GPT-3's quality judgements track "some notion of language ideology more correlated with wealthier zip codes" rather than factuality or literary acclaim; filters "may introduce new biases".
- The authors' framing: deciding what to filter "requires non-trivial normative decisions that affect the biases of their datasets and thus their models"; curation choices are "not easily erased by subsequent finetuning".
Limits
- English only; one architecture family; one run per variant; fine-tuned evaluations (not zero-shot).
- "Quality" and "toxicity" are whatever two classifiers say; Perspective is a black-box API whose results may not reproduce.
- No measurement of opinions, values, persona or demographic defaults in the trained models. Representational bias is measured only as toxic continuations about identity groups.
- Models and code unreleased. Data licences (footnote 1): C4 is ODC-By, the Pile MIT.
- Inconsistencies found.
- Table 16's caption says quality filtering "decreases the ability of LM-XL to identify toxicity", but its numbers (+1.1 to +1.8) and the main text ("toxicity identification by 2%") show an increase.
- Table 3's QA averages differ slightly from Figure 10's for the same filters (+2.5 against +2.7; −3.1 against −3.3).
What it means for the base (inference)
- Deleting a kind of content removes the ability to recognise it. If opinionated text is filtered out of B1's corpus the way toxic text was, the base will be worse at representing opinions at all, and a person slot needs a base that can represent any opinion. The toxicity result argues for keeping opinion-bearing text, balanced, and keeping the position out by other means.
- A quality filter is a choice of whose writing counts. The classifier that improved everything is trained to prefer Wikipedia- and book-like text; the cited work links such filters to wealthier, Western authors. A "clean" corpus is therefore not a neutral one; for B1 it would sharpen exactly the demographic default the base must not have.
- Corpus date is a default cohort. Pretraining time leaves a lasting imprint that fine-tuning does not remove. A base trained on one era's text carries that era's world; a person slot for someone of another generation starts from the wrong era. B1 needs either a time-balanced corpus or explicit time conditioning.
- Named people are everywhere in web text (37–41% of documents). Person-level factual holdings about real individuals are the default content of web corpora, not an edge case.