Kurisutina

Using language models to simulate human samples

Published in Political Analysis 31(3), 2023, DOI 10.1017/pan.2023.2; the published version was not read. Cambridge served only an abstract page. Read as the arXiv preprint 2209.06899 v1 (14 September 2022, the only arXiv version; CC BY-SA 4.0). Brigham Young University. Provenance: papers/carry_on/argyle2023_out_of_one_many.provenance.json.

What was read

All 2,530 lines of pdftotext -layout output (53 pages):

  • text and appendices A–E: GPT-3 usage, Study 1 survey instrument and regression tables, Study 2 metrics, ablation and model comparison, Study 3 template, missing data, descriptive statistics, alternative specifications, temperature, and cost;
  • Tables 1–17 and figure captions.

Figures are images; only captions and extracted labels were available. Differences from the published version are unknown.

Question

Can GPT-3, conditioned on real survey respondents' demographic "backstories", reproduce how human subgroups respond? The authors call this "algorithmic fidelity" and set four criteria:

  1. Turing test;
  2. backward continuity;
  3. forward continuity;
  4. pattern correspondence.

Method

  • Model. GPT-3 davinci via the API. Temperature 0.7 for text; next-token probabilities for votes.
  • Study 1. 2,107 respondents from "Pigeonholing Partisans".
    • First-person backstories: ideology, party, race, gender, income and age.
    • GPT-3 writes four words describing Democrats and four describing Republicans.
    • 2,873 Lucid raters judged 7,675 lists (3,592 human, 4,083 GPT-3), blind to source, on:
      • the writer's party;
      • tone and extremity;
      • traits, issues and groups;
      • human or computer.
  • Study 2. ANES 2012, 2016 and 2020 respondents.
    • Backstory of up to 10 attributes: race, gender, age, ideology, party, interest, church, discussion, flag patriotism and state.
    • P(Republican | backstory) from token sets for "In [year], I voted for…", dichotomised at .5.
    • Compared with each respondent's reported vote: tetrachoric r, κ, ICC and agreement.
  • Study 3. ANES 2016.
    • An interview-format context gives 11 of a person's answers; GPT-3 generates the 12th (5 tokens, temperature 0.7).
    • Cramér's V between every input variable and the generated output is compared with the same V in human data.
    • Only complete cases are kept: 1,782 of 4,270.

Results

  • Study 1.

    • Raters judged 61.7% of human lists and 61.2% of GPT-3 lists as human-written.
    • The writer's party was guessed correctly for 60.1% of human lists and 52.8% of GPT-3 lists; chance is 33%.
    • Traits were mentioned in 72.3% of human lists and 66.5% of GPT-3 lists; extremity 39.8% against 41.0%.
  • Study 2. Individual vote agreement:

    • tetrachoric r .90, .92 and .94 (2012, 2016, 2020); proportion agreement .85–.89;
    • strong partisans .99–1.00;
    • pure independents .31, .41 and .02, near chance.

    Overall P(Republican): GPT-3 .391 against ANES .404 (2012), .432 against .477 (2016), .472 against .412 (2020).

    • Ablation (Figure 8, 2016): party ID carries most of the accuracy, then ideology. Demographics alone do better than any single one of them.
    • GPT-Neo 6B came close to GPT-3.
  • Study 3.

    • Cramér's V in GPT-3 outputs tracks human associations; the mean difference is −0.026.
    • Some associations are much weaker in GPT-3 outputs, e.g. ideology→church .28 against .07, and political interest→discussing politics .40 against .16.
  • What the appendix shows but the main text does not discuss (Table 16). The marginals of GPT-3's generated answers are far from the humans':

    Variable GPT-3 ANES
    Mean age 35.5 50.1
    Male 75.9% 48.1%
    White 97.4% 80.3%
    Hispanic 0.1% 8.9%
    Graduate degree 0.2% 19.6%
    Some college 64.2% 34.8%
    Trump voter 24.5% 43.8%
    Clinton voter 23.3% 48.4%
    "Other" voter 52.3% 7.8%
    • GPT-3 also left more answers unusable (Table 15): 22.6% for ideology and 23.8% for vote.
    • At temperature 0.001, "GPT-3 identified all respondents as white".

Limits

  • One closed, since-retired model, US politics only.
  • Study 3 keeps 42% of cases (1,782 of 4,270 complete). Temperature comparisons use different case sets (N = 2,518, 1,782 and 1,022).
  • The criteria have no thresholds ("we do not propose specific metrics or numerical thresholds").
  • Inconsistencies found:
    • Study 1's regression (Table 2) shows GPT-3 lists differ from human lists:

      • traits −0.058;
      • issues +0.033;
      • groups +0.078;
      • all with SE 0.007.

      That is far beyond two standard errors (my computation). The text describes "a remarkable degree of consistency".

    • Study 3 is presented as pattern correspondence while GPT-3's answer distributions are heavily distorted (Table 16). Most strikingly, "Other" is chosen for 52% of 2016 votes. Associations can match while levels do not.

    • Appendix E: "Study 1 required 1,1471 model queries (one for each human subject)". The number is malformed, and Study 1 had 2,107 subjects. There is also "5,914 in 20112".

    • The appendix refers to "Figure 6" and "Figure 4 in the paper" for results shown in Figures 3–4.

    • The age variable is V161267 in the text and V161247 in Table 11.

  • Checked:
    • 3,592 + 4,083 = 7,675 lists.
    • 60.1 − 52.8 = 7.3 points, matching Table 3's −0.073.

What it means for Kurisutina

  • Associations are not calibration: check marginals.
    • GPT-3 reproduced which answers go together while predicting absurd levels: a 52% "other" vote, 0.2% graduate degrees, 97% white.
    • For the GSS pilot (proposal, exploratory): per item and condition, report the predicted marginal distribution against the observed later-wave distribution alongside Brier. A model can score reasonably on ranking while systematically over-predicting a middle or "other" category.
      • For finalter ("stayed the same"), happy ("pretty happy") and trust ("depends"), over-prediction of the middle category is the specific thing to look for (inferred from the "other" result).
  • Party carries the persona. The ablation shows party ID doing most of the work, and independents are at chance. This is the same pattern as Bisbee, Hwang and Kim & Lee: attribute backstories reproduce the political axis.
    • A person who is not well described by party is exactly where attribute conditioning fails, and where their own record has to carry the prediction.
  • Low temperature collapses to the mode. At temperature 0.001 everyone became white.
    • The pilot reads full next-token distributions rather than sampling, which avoids this.
    • A replica that answers at low temperature will present the modal answer as the person's (Santurkar's collapse).
  • The pilot's M1 design descends from this paper. Study 3's interview-format context, a person's other answers and then one target question, is close to M1's record continuation.
    • Argyle evaluated only group-level associations and explicitly not individual correspondence. The pilot's per-person Brier and own-against-other contrasts go further than anything here.
  • Individual vote agreement of .85–.89 is mostly party ID. Strong partisans reach .97, independents .53–.62. Headline "fidelity" numbers should always be broken down by how predictable the person is from one attribute (inferred).

Cross-references

  • summaries/carry_on/bisbee2024_synthetic_replacements.md: the critique and replication.
  • summaries/carry_on/santurkar2023_opinionqa.md: default opinions and mode collapse.
  • summaries/carry_on/hwang2023_user_opinions.md, summaries/carry_on/kim2024_ai_augmented_surveys.md: individual prediction from own answers.
  • docs/research/gss_pilot_design.md: M1 record continuation.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.