Kurisutina

BLiMP: the Benchmark of Linguistic Minimal Pairs for English

Warstadt, Parrish, Liu, Mohananey, Peng, Wang, Bowman. Transactions of the ACL 8:377–392, doi 10.1162/tacl_a_00321; open access (CC BY 4.0). Data and generation scripts: github.com/alexwarstadt/blimp and /data_generation (not inspected). Read by researcher R4d (the base battery), 5 October 2026. Provenance: papers/base/warstadt2020_blimp.provenance.json.

What was read

  • Read in full: the ACL Anthology PDF (16 pages, pdftotext -layout): abstract, sections 1–7, Tables 1–4 (Table 4 lists all 67 paradigms with an example pair, model and human accuracy), figure captions, acknowledgements, appendix caveats, references.
  • Not read: Figures 1–6 as images (the correlation heatmap and learning curves; only values stated in the text are used), the dataset and code.

What it is

  • 67 paradigms × 1,000 minimal pairs of English sentences that differ in one place and contrast in acceptability, in 12 phenomena: anaphor agreement, argument structure, binding, control/raising, determiner–noun agreement, ellipsis, filler–gap, irregular forms, island effects, NPI licensing, quantifiers, subject–verb agreement. Example: The cats annoy Tim / *The cats annoys Tim.
  • Generated, not collected: linguist-written templates sample from a vocabulary of over 3,000 annotated items, so each paradigm isolates one contrast. Pairs have equal length (in lexicon entries) and differ by at most one item. Occasionally implausible sentences result (Sam ran around some glaciers), but both members of a pair are equally implausible, so world knowledge does not decide the choice.
  • Model score: the share of pairs where the model gives the acceptable sentence the higher total probability (forced choice, chance 50%). No fine-tuning is needed.

Human reference (verified)

  • Validation on Mechanical Turk: US-located self-reported native speakers; 20 validators per pair on 5 pairs per paradigm (6,700 judgments), with attention checks; $0.25 per five judgments.
  • Aggregate (majority-vote) agreement with the labels: 96.4%. A paradigm was kept only if the majority agreed on at least 4 of its 5 pairs; 2 candidate paradigms failed and were dropped.
  • Individual human agreement: 88.6% overall, reported as the conservative human score. By phenomenon: anaphor agreement 97.5, argument structure 90.0, binding 87.3, control/raising 83.9, determiner–noun 92.2, ellipsis 85.0, filler–gap 86.9, irregular forms 97.0, islands 84.9, NPI 88.1, quantifiers 86.6, subject–verb 90.9.
  • Some paradigms have low individual agreement (sentential subject island 61%, only-NPI scope 72%, wh-island 73%, tough-vs-raising 75%): the labels there are less secure.

Model results (verified)

  • Overall: 5-gram 60.5%, LSTM 68.9%, Transformer-XL 68.7%, GPT-2-large 80.1% (8 points below individual humans), all below humans.
  • Best and closest to humans on morphology: GPT-2 within 2.1 points on the three agreement categories. Hardest: islands (only GPT-2 well above chance, 20 points below humans), NPI licensing, quantifiers; argument structure weak for all models.
  • Across the 67 paradigms, model accuracy correlates "moderately" with human accuracy (GPT-2 most); neural models correlate more with each other than with humans (Transformer-XL with LSTM about 0.9), "suggesting neural networks share some biases that are not human-like".
  • Attractors and distance: an intervening adjective costs neural models 3–5 points; an attractor noun of the opposite number costs GPT-2 22 and the LSTM 20 points (Transformer-XL 5).
  • Data, not size: GPT-2 at 117M to 1,558M parameters all score about 84%, phenomenon SDs ≤ .03. Retraining the LSTM and Transformer-XL on 1/8M to 64M tokens shows phenomenon-specific learning curves (steepest for anaphor and determiner–noun agreement; flattest for NPIs and islands). Extrapolated log-linearly, human-like islands and NPIs would need well over 10²⁰ tokens for those architectures.
  • One- and two-prefix scoring methods give broadly similar results to full-sentence scoring, with some paradigm-level differences.

Caveats stated by the authors

  • Positive results do not prove human-like knowledge (e.g. the filler–gap paradigms: models detect a missing gap but not an illicit one); conclusions need several paradigms per concept, ideally nonce words or factorial designs.
  • Several anaphor and binding paradigms rely on stereotyped gender–name associations (Mary hugged himself is marked unacceptable; themself is never offered).
  • The labels encode mainstream US English; some dialects accept forms marked unacceptable (Suzy don't lie).

Limits of this reading

  • The human reference is small per paradigm (5 pairs × 20 raters) and from crowdworkers, with inattention included in the "individual" score.
  • The dataset has been public since 2020; any model trained on later web text may have seen it.

What it means for the base battery (inference)

  • A language-competence test that rewards neither encyclopaedic knowledge nor person content: the sentences use arbitrary names and often implausible events, and both members of a pair carry the same content, so only grammar decides.
  • A human reference exists at two levels: individual 88.6% (and per phenomenon) and majority 96.4%. A base meant to be a population listener should be non-inferior to the individual level overall and per phenomenon, and its paradigm profile should correlate with the human profile.
  • Contamination is controllable for our own base: exclude BLiMP from the vocabulary engine's training data and regenerate fresh pairs from the public templates with a new vocabulary and new names, scoring both.
  • Two items to fix before freezing: the gender–name paradigms test a population association (arguably base content, but it should be reported separately), and the dialect-specific paradigms should be flagged.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.