Kurisutina

From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models

Proceedings of ACL 2023 (Volume 1: Long Papers), pages 11737–11762 (ACL Anthology 2023.acl-long.656; arXiv 2305.08283, not read). University of Washington, Carnegie Mellon, Xi'an Jiaotong. Code and data: github.com/BunsenFeng/PoliLean (not read). Read by researcher R4a (Claude Opus 5.5) for research batch R4, 5 October 2026. Provenance: papers/base/feng2023_political_bias_trails.provenance.json.

What was read

  • Read in full: all 1,621 lines of the pdftotext -layout text of the 26-page Anthology PDF: sections 1–7, limitations, ethics, references, appendices A–H (probing details, precision and recall, experiment details, stability analysis, qualitative examples, hyperparameters, compute, artefacts), Tables 1–15 (including the 62 political-compass statements) and the Responsible NLP checklist.
  • Figures 1–4 (compass plots) are images: read through captions and text only. The kappa values of Figures 5–6 are in the extracted text.
  • Not read: the arXiv version, the code, and the POLITICS news corpus paper (Liu et al. 2022) that holds the news corpora's sizes.

Question

Do language models hold political leanings, do those leanings come from the pretraining data, and do they change downstream behaviour (hate-speech and misinformation detection)?

Method

  • Measuring a model's leaning: the political compass test (62 statements; social axis liberal–conservative, economic axis left–right, each −10 to +10).
    • Encoders: mask filling in "Please respond to the following statement: [STATEMENT] I <MASK> with this statement."; probabilities of agree-type against disagree-type words among the top 10 tokens, with a 0.3 margin for "strong".
    • Decoders: generated responses scored by a BART-MNLI stance detector, 10 seeds, low-confidence responses dropped. Human check: 80 of 110 responses had a clear stance (Fleiss' kappa 0.85 among three annotators), and on those the detector was 97% accurate.
    • 14 model families: BERT, RoBERTa, distilBERT, distilRoBERTa, ALBERT, BART, GPT-2, GPT-3 (ada to davinci-002), GPT-J, LLaMA-7B, Alpaca, Codex, ChatGPT, GPT-4.
  • Putting a leaning in through data. Continued pretraining of RoBERTa-base and GPT-2 on six corpora of comparable size: {left, centre, right} × {news, Reddit}. News: the POLITICS dataset split by AllSides ratings. Reddit: left- and right-leaning subreddits, with non-political subreddits as centre (posts: left 796,939, centre 952,152, right 934,452; mean 44.5, 34.7 and 50.4 tokens). Hateful content was filtered out with a RoBERTa classifier, so the corpora are "clean" opinion text. Also split at 20 January 2017 (pre- and post-Trump). Learning rate 2e-5, up to 20 epochs (following Gururangan et al. 2020); about 48 and 83 GPU-cluster hours on 16 A40s.
  • Downstream. Each partisan RoBERTa fine-tuned for hate speech (Yoder et al. 2022; by targeted identity group) and misinformation (PolitiFact; by source), with identical fine-tuning data and settings.

Results (verified)

  • Off-the-shelf models hold positions, and they differ. Models occupy all four quadrants of the compass. BERT variants are more socially conservative (authoritarian) than GPT variants; the authors suggest BookCorpus against more liberal modern web text, and RLHF for later GPT models. Sizes within a family can differ. Positions are stronger on social than economic issues (mean magnitude 2.97, SD 1.29, against 0.87, SD 0.84).
  • Examples (Table 2). On "Mothers may have careers, but their first duty is to be homemakers": GPT-3 Ada "This is not a true statement"; GPT-3 Davinci "I agree with this statement". On "The only social responsibility of a company should be to deliver a profit to its shareholders", Davinci agrees.
  • Continued pretraining moves the position toward the corpus. Left corpora shifted models left and liberal, right corpora right and conservative. The largest: RoBERTa on Reddit-left, social score 2.97 → −3.03. For RoBERTa, social media moved social values by 1.60 on average and news by 0.64; economic values moved 0.90 (news) and 0.61 (social media). "Most of the ideological shifts are relatively small, suggesting that it is hard to alter the inherent bias present in initial pretrained LMs."
  • Time matters. Post-2017 corpora moved most setups further from the centre than pre-2017 corpora.
  • More epochs and more partisan data did not create extremists (Figure 4): economic scores stayed near the centre, and social scores did not approach ±10.
  • Downstream effects. Overall accuracy barely changed (balanced accuracy 86.0–90.3 across variants), but behaviour by group did. Left-leaning models were better at hate speech against minorities (Black, LGBTQ+) and right-leaning models at hate speech against dominant groups (men, white people). For misinformation, each side was stricter with the other side's media and laxer with its own. Reddit-right pretraining hurt overall misinformation detection most (85.05 F1 against 88.37 for base RoBERTa). A partisan ensemble beat every single model (e.g. misinformation F1 90.50 against at best 88.15).
  • Authors' conclusion: "pernicious biases and unfairness in downstream tasks can be caused by non-toxic data, which includes diverse opinions, but there are subtle imbalances in data distributions." They argue that filtering or augmentation risk "censorship and exclusion from political participation".

Limits

  • The measurement is weak for small models. Across seven prompt wordings (Figure 6), agreement of a model's stance with itself was Fleiss' kappa 0.014 for GPT-2, 0.068 for RoBERTa, −0.130 for distilRoBERTa and 0.406 for GPT-3; across paraphrased statements (Figure 5), 0.047, 0.452, 0.349 and 0.600. The authors call this "moderately stable"; for GPT-2 and RoBERTa it is close to no consistency.
  • The political compass test: two axes, Western issues, criticised for unclear scoring, libertarian bias and vague statements (the authors say so). Prompted agreement is not the same as a disposition.
  • Continued pretraining of already-pretrained models; no model was pretrained from scratch on a controlled mix, so how much of a fresh model's position a corpus sets is not measured.
  • Corpus leaning is labelled by source (AllSides ratings, subreddit lists), not measured in the text.
  • US-centric framing (stated in the ethics section). A wording slip in appendix A.2 ("80 unclear responses" for the 80 clear ones).

What it means for the base (inference)

  • Opinions travel from text to weights without any toxic content. Clean, opinionated news and forum text gave models measurable positions, and the positions shaped behaviour on judgement tasks. A vocabulary engine trained on opinion-bearing text will hold a position unless the opinions in its data cancel.
  • Positions are set early and resist change. Continued pretraining moved positions only a little. Whatever position B1 learns first will be hard for a person slot to override, the same shrinkage toward the base that the collection documents for existing models.
  • The corpus's leaning can be labelled by source, and balancing is plausible. Source-level leaning labels (outlet ratings, community lists) are cheap. Balanced sources are a defensible first filter; a centre corpus of non-political communities is another (inferred, not tested here for a fresh model).
  • Imbalance skews judgement. Models steeped in one side's text were laxer with that side's misinformation and stricter with the other side's, and their hate-speech detection shifted by target group. A base with a net lean would judge people's claims asymmetrically. A base that must represent any person's views needs the vocabulary of all sides without holding one, which argues for balance (the paper did not test deletion).
  • We need a better instrument than the compass test for small models: kappa near 0 across prompt wordings means prompted stance is mostly noise at GPT-2 scale. Likelihood-based probes over paraphrase sets, scored against population marginals (as Santurkar et al. do), are the obvious replacement.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.