arXiv 2303.17548, version 1 (30 March 2023), the only version on arXiv; published at ICML 2023 per Bisbee et al.'s
reference list (the ICML version was not read). Stanford and Columbia. arXiv non-exclusive distribution licence.
Code and data: github.com/tatsu-lab/opinions_qa (not read). Provenance:
papers/carry_on/santurkar2023_opinionqa.provenance.json.
What was read
- All 1,648 lines of
pdftotext -layoutoutput (43 pages): text, Appendices A–B, Tables 1–5 and the figure captions. - Figures are images. Pages 7, 31 and 34 were rendered and inspected: Figure 2 (representativeness scores), Figure 7 (probability on answer options), and Figures 9–10 (entropy, refusal rates).
- Not inspected: Figures 3–6, 8, and 11–14 (the heat maps by group and topic, steering, robustness).
- Numbers below from Figures 2, 7 and 10 were read off the images.
Question
Whose opinions does an LM's default answer distribution resemble? Can prompting steer it to a group's distribution? Is its alignment consistent across topics?
Method
- OpinionQA. 15 Pew American Trends Panel waves (2017–2021), about 1,500 multiple-choice questions, and 60 demographic groups. Human distributions are weighted.
- Models. Nine models from AI21 and OpenAI, 350M to 178B parameters.
- Base: ada, davinci, j1-grande, j1-jumbo.
- Human-feedback (HF) tuned: text-ada-001, text-davinci-001/002/003, j1-grande-v2-beta.
- Model opinion. Next-token probabilities on the option letters, renormalised with "Refused" excluded; refusal is analysed separately.
- Alignment. 1 − (1-Wasserstein distance on the ordinal option scale) / (N − 1), averaged over questions.
- Representativeness: against everyone or a group, with no context.
- Steerability: the best of three group prompts (QA, BIO, PORTRAY) on 500 contentious questions and 22 groups.
- Consistency: the share of topics whose best-aligned group equals the model's overall best-aligned group.
- "Modal" check. Human group distributions sharpened with temperature 1e-3, to test whether a model matches a group's mode rather than its spread.
- Robustness. Permuted option order, and two instruction formats.
Results
-
No model is representative (Figure 2). Overall representativeness:
Source Score Human demographic groups, average .949 Human demographic groups, worst .865 Base models .791–.824 HF-tuned models .700–.813 text-davinci-003 .700 (the lowest) - Every human demographic group is more representative of the whole population than any model.
- Most models sit where Democrats against Republicans on climate change sit.
-
HF tuning shifts whom the model resembles.
- Base models align best with lower-income, moderate, Protestant or Catholic respondents.
- OpenAI's HF models align best with liberal, high-income, educated and non-religious respondents. That matches the InstructGPT crowdworker profile, though the authors cannot confirm the link.
- All models represent poorly people aged 65+, widowed people, Mormons, and those attending services most often.
-
HF tuning collapses the spread.
- text-davinci-003 typically puts > .99 on one option. Its entropy piles up near 0 (Figure 9), where humans spread out.
- It matches the modal views of liberals and moderates better than their actual distributions: it converges to a group's caricature ("99% approval of Joe Biden").
-
Steering helps a little and uniformly (Figure 4b). Group prompts raise alignment for most models, except ada, but "none of the disparities … disappear". Groups the model represented poorly stay poorly represented.
-
Opinions are a patchwork. Consistency scores are low. Even the liberal-leaning text-davinci-002/003 align with conservatives on religion.
-
Probability on valid options (Figure 7).
- HF models put nearly all their mass on the option letters.
- j1-grande, j1-grande-v2-beta and davinci mostly put 10–60% there. ada and j1-jumbo put 35–95%, and text-ada-001 anything from 0 to 1. The authors' check: "at least 30% on average".
-
Refusal (Figure 10). Humans choose "Refused" 1.5% of the time; the models:
Models Refusal text-davinci-001/002/003 1.8–3.8% text-ada-001 16.4% Base models and AI21's tuned model 13–21% -
Robustness. Permuting options lowers every model's score slightly. Rankings and group patterns hold.
Limits
- US and English only. Multiple choice only, next-token probabilities only.
- Closed, since-retired models.
- Groups, not individuals.
- Truncated log-probabilities are bounded heuristically: the APIs returned only the top 100 (OpenAI) or 10 (AI21).
- Inconsistencies found:
- The text says 1,498 questions; Table 1's per-wave counts sum to 1,506.
- The text says 60 groups; Table 2 lists 61 attribute values.
- The text says 9 models; Table 5 lists 8. text-ada-001 is missing from it but appears in Figures 2, 7, 9 and 10.
- Section 4.1 says "all models have low refusal rates … as low as 1–2%". Figure 10 shows 13–21% for the base models, AI21's tuned model and text-ada-001, against 1.5% for humans; only the text-davinci models are low.
- Appendix A.3 says AI21 returns 10 log-probabilities, then uses "K (100 or 64)".
- The base-model list names davinci twice.
- Checked: Table 1's response counts sum to 80,098, i.e. 5,340 per wave on average. That is the "5340 users" in Hwang et al.
What it means for Kurisutina
- The empty slot is not neutral, and it differs by arm (inferred). A base model's default resembles one population
segment, and a chat-tuned model's resembles another, delivered as a near-certain mode.
- In the pilot, "empty" therefore means a different thing for M1 (base) and M2/M3 (chat).
- Under Brier, a chat model's sharp, wrong default is punished heavily. So own − empty may look larger for chat arms simply because their empty condition is overconfident, not because they use the record better.
- Read the arms' own − empty contrasts with each arm's empty-condition entropy alongside.
- Drift toward the default is drift toward a patchwork caricature. The default has no consistent person behind
it. It is liberal on some topics and conservative on others, and collapsed onto modes after RLHF. A replica drifting
toward it would not become "another person"; it would become less person-like.
- For Q2 (proposal): measure entropy and topic-consistency of the replica's answers over time alongside agreement with the person.
- The pilot's widowed stratum is where models are worst. People aged 65+ and widowed people are the groups all
models represent least well.
- Expect weaker replica performance in the widowed stratum (73 available pairs) and among older respondents, independent of the event.
- Stratify results by age; if the widowed stratum is worst, do not read it as a failure to model widowhood (inferred).
- Steering by group labels cannot fix representativeness. This agrees with Bisbee and Hwang: group prompts move the model a little and uniformly. The person-specific record is the only lever that reached individuals in any paper read so far.
- Methods to reuse.
- Read option-letter probabilities with "Refused" handled separately, and normalise over valid options. The pilot already does this with codes, and reports mass on valid codes (92–93% for M1).
- Report the mass on valid codes per arm. As here, base models can put much of it elsewhere, and renormalising then hides uncertainty.
- A Wasserstein-based alignment respects ordinal scales. It is a useful companion to Brier for ordinal GSS items (happy, satfin, health, trust), where Brier treats "very happy" against "not too happy" as no worse than "very" against "pretty" (proposal).
Cross-references
summaries/carry_on/bisbee2024_synthetic_replacements.md: persona prompting caricatures partisans; drift across model versions.summaries/carry_on/hwang2023_user_opinions.md: individual users on the same Pew data.summaries/carry_on/lu2026_assistant_axis.md: the assistant default as an attractor.docs/research/gss_pilot_design.md: arms, the empty condition, strata.