Kurisutina

Whose opinions do language models reflect?

arXiv 2303.17548, version 1 (30 March 2023), the only version on arXiv; published at ICML 2023 per Bisbee et al.'s reference list (the ICML version was not read). Stanford and Columbia. arXiv non-exclusive distribution licence. Code and data: github.com/tatsu-lab/opinions_qa (not read). Provenance: papers/carry_on/santurkar2023_opinionqa.provenance.json.

What was read

  • All 1,648 lines of pdftotext -layout output (43 pages): text, Appendices A–B, Tables 1–5 and the figure captions.
  • Figures are images. Pages 7, 31 and 34 were rendered and inspected: Figure 2 (representativeness scores), Figure 7 (probability on answer options), and Figures 9–10 (entropy, refusal rates).
  • Not inspected: Figures 3–6, 8, and 11–14 (the heat maps by group and topic, steering, robustness).
  • Numbers below from Figures 2, 7 and 10 were read off the images.

Question

Whose opinions does an LM's default answer distribution resemble? Can prompting steer it to a group's distribution? Is its alignment consistent across topics?

Method

  • OpinionQA. 15 Pew American Trends Panel waves (2017–2021), about 1,500 multiple-choice questions, and 60 demographic groups. Human distributions are weighted.
  • Models. Nine models from AI21 and OpenAI, 350M to 178B parameters.
    • Base: ada, davinci, j1-grande, j1-jumbo.
    • Human-feedback (HF) tuned: text-ada-001, text-davinci-001/002/003, j1-grande-v2-beta.
  • Model opinion. Next-token probabilities on the option letters, renormalised with "Refused" excluded; refusal is analysed separately.
  • Alignment. 1 − (1-Wasserstein distance on the ordinal option scale) / (N − 1), averaged over questions.
    • Representativeness: against everyone or a group, with no context.
    • Steerability: the best of three group prompts (QA, BIO, PORTRAY) on 500 contentious questions and 22 groups.
    • Consistency: the share of topics whose best-aligned group equals the model's overall best-aligned group.
  • "Modal" check. Human group distributions sharpened with temperature 1e-3, to test whether a model matches a group's mode rather than its spread.
  • Robustness. Permuted option order, and two instruction formats.

Results

  • No model is representative (Figure 2). Overall representativeness:

    Source Score
    Human demographic groups, average .949
    Human demographic groups, worst .865
    Base models .791–.824
    HF-tuned models .700–.813
    text-davinci-003 .700 (the lowest)
    • Every human demographic group is more representative of the whole population than any model.
    • Most models sit where Democrats against Republicans on climate change sit.
  • HF tuning shifts whom the model resembles.

    • Base models align best with lower-income, moderate, Protestant or Catholic respondents.
    • OpenAI's HF models align best with liberal, high-income, educated and non-religious respondents. That matches the InstructGPT crowdworker profile, though the authors cannot confirm the link.
    • All models represent poorly people aged 65+, widowed people, Mormons, and those attending services most often.
  • HF tuning collapses the spread.

    • text-davinci-003 typically puts > .99 on one option. Its entropy piles up near 0 (Figure 9), where humans spread out.
    • It matches the modal views of liberals and moderates better than their actual distributions: it converges to a group's caricature ("99% approval of Joe Biden").
  • Steering helps a little and uniformly (Figure 4b). Group prompts raise alignment for most models, except ada, but "none of the disparities … disappear". Groups the model represented poorly stay poorly represented.

  • Opinions are a patchwork. Consistency scores are low. Even the liberal-leaning text-davinci-002/003 align with conservatives on religion.

  • Probability on valid options (Figure 7).

    • HF models put nearly all their mass on the option letters.
    • j1-grande, j1-grande-v2-beta and davinci mostly put 10–60% there. ada and j1-jumbo put 35–95%, and text-ada-001 anything from 0 to 1. The authors' check: "at least 30% on average".
  • Refusal (Figure 10). Humans choose "Refused" 1.5% of the time; the models:

    Models Refusal
    text-davinci-001/002/003 1.8–3.8%
    text-ada-001 16.4%
    Base models and AI21's tuned model 13–21%
  • Robustness. Permuting options lowers every model's score slightly. Rankings and group patterns hold.

Limits

  • US and English only. Multiple choice only, next-token probabilities only.
  • Closed, since-retired models.
  • Groups, not individuals.
  • Truncated log-probabilities are bounded heuristically: the APIs returned only the top 100 (OpenAI) or 10 (AI21).
  • Inconsistencies found:
    • The text says 1,498 questions; Table 1's per-wave counts sum to 1,506.
    • The text says 60 groups; Table 2 lists 61 attribute values.
    • The text says 9 models; Table 5 lists 8. text-ada-001 is missing from it but appears in Figures 2, 7, 9 and 10.
    • Section 4.1 says "all models have low refusal rates … as low as 1–2%". Figure 10 shows 13–21% for the base models, AI21's tuned model and text-ada-001, against 1.5% for humans; only the text-davinci models are low.
    • Appendix A.3 says AI21 returns 10 log-probabilities, then uses "K (100 or 64)".
    • The base-model list names davinci twice.
  • Checked: Table 1's response counts sum to 80,098, i.e. 5,340 per wave on average. That is the "5340 users" in Hwang et al.

What it means for Kurisutina

  • The empty slot is not neutral, and it differs by arm (inferred). A base model's default resembles one population segment, and a chat-tuned model's resembles another, delivered as a near-certain mode.
    • In the pilot, "empty" therefore means a different thing for M1 (base) and M2/M3 (chat).
    • Under Brier, a chat model's sharp, wrong default is punished heavily. So own − empty may look larger for chat arms simply because their empty condition is overconfident, not because they use the record better.
    • Read the arms' own − empty contrasts with each arm's empty-condition entropy alongside.
  • Drift toward the default is drift toward a patchwork caricature. The default has no consistent person behind it. It is liberal on some topics and conservative on others, and collapsed onto modes after RLHF. A replica drifting toward it would not become "another person"; it would become less person-like.
    • For Q2 (proposal): measure entropy and topic-consistency of the replica's answers over time alongside agreement with the person.
  • The pilot's widowed stratum is where models are worst. People aged 65+ and widowed people are the groups all models represent least well.
    • Expect weaker replica performance in the widowed stratum (73 available pairs) and among older respondents, independent of the event.
    • Stratify results by age; if the widowed stratum is worst, do not read it as a failure to model widowhood (inferred).
  • Steering by group labels cannot fix representativeness. This agrees with Bisbee and Hwang: group prompts move the model a little and uniformly. The person-specific record is the only lever that reached individuals in any paper read so far.
  • Methods to reuse.
    • Read option-letter probabilities with "Refused" handled separately, and normalise over valid options. The pilot already does this with codes, and reports mass on valid codes (92–93% for M1).
    • Report the mass on valid codes per arm. As here, base models can put much of it elsewhere, and renormalising then hides uncertainty.
    • A Wasserstein-based alignment respects ordinal scales. It is a useful companion to Brier for ordinal GSS items (happy, satfin, health, trust), where Brier treats "very happy" against "not too happy" as no worse than "very" against "pretty" (proposal).

Cross-references

  • summaries/carry_on/bisbee2024_synthetic_replacements.md: persona prompting caricatures partisans; drift across model versions.
  • summaries/carry_on/hwang2023_user_opinions.md: individual users on the same Pew data.
  • summaries/carry_on/lu2026_assistant_axis.md: the assistant default as an attractor.
  • docs/research/gss_pilot_design.md: arms, the empty condition, strata.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.