Political Analysis 32: 401–416 (2024), DOI 10.1017/pan.2024.5. Received May 2023, accepted January 2024, online 17 May
2024. Vanderbilt University. Open access under CC BY 4.0. Read as the published version of record from Cambridge
Core. Replication code: Harvard Dataverse doi:10.7910/DVN/VPN481 (not read). Provenance:
papers/carry_on/bisbee2024_synthetic_replacements.provenance.json.
What was read
-
All 964 lines of
pdftotext -layoutoutput (16 pages): abstract, text, footnotes 1–30, Table 1, the captions and extracted labels of Figures 1–5, and the references. -
Figures are images; only their captions, axis labels and the regression equations in Figure 5 were available.
-
Not read: the Supplementary Information, which is a separate file. It holds:
- the prompt development and the RLHF workaround;
- temperature;
- first- against second-person prompts;
- the default persona;
- other datasets, questions and countries;
- the ChatGPT 4 and Falcon-40B-Instruct replications.
Claims below that rest on the SI are the authors' words from the main text.
-
Note: the PDF carries a Cambridge download stamp with the downloading IP address. The local file is not for redistribution.
Question
Can a pretrained LLM, told to adopt the persona of a real survey respondent, reproduce that respondent's answers? Beyond overall means, does it reproduce their spread, the correlates of opinion, and the same results when rerun?
Method
- Model. ChatGPT 3.5 Turbo at temperature 1, the pre-25-June-2023 version. Parts were replicated with ChatGPT 4.0, the post-update 3.5, and Falcon-40B-Instruct (in the SI).
- Personas. 7,530 real ANES respondents from 2016 and 2020. The prompt is second person ("It is [YEAR]. You are a
[AGE] year-old …") and gives:
- age, marital status, race/ethnicity, gender, education and income;
- ideology, registration, party and political interest.
- Task. Feeling thermometers (0–100) for 16 groups, in one prompt returning a TSV with an explanation and a
confidence. 11 of the groups are in the ANES waves used.
- 30 draws per respondent gives 3,614,400 responses (7,530 × 30 × 16; verified).
- Analyses use the mean of the 30 draws, except for uncertainty, which uses the first draw only.
- Analyses.
- Means and SDs against the ANES, overall and by party and race.
- For each group and year (22 regressions), a regression with a data-source interaction. It tests whether each covariate's coefficient differs between synthetic and ANES data.
- Prompt sensitivity: demographics only, politics only, or both. MAE per respondent is modelled by prompt, group and party.
- Stability over time: a simpler prompt (year 2019, 20 draws per persona) run in April, June and July 2023. OpenAI updated 3.5 Turbo on 25 June.
Results
-
Means are close; spread is too small. Every synthetic mean lies within one ANES SD, and the rank order of groups is largely kept.
- Spread is much smaller, especially for racial and religious groups. That holds even for a single draw per person (footnote 20).
-
Partisans are caricatured. By party and race, the synthetic answers are more extreme by 0.5–1 ANES SD (10–20 thermometer points).
- Synthetic Democrats like liberals more and conservatives less than real Democrats. The same holds for Republicans, especially Black Republicans.
- "Synthetic responses would suggest that society is more politically hostile than it actually is."
-
Power analyses would be wrong by an order of magnitude (Table 1). Detecting a change in affective polarisation from 2012:
Source Effect SD n at 80% power n at 99% power ANES 7.8 31.4 129 299 ChatGPT 12.5 16.1 16 35 -
Correlates of opinion are often wrong.
- 48% of covariate coefficients differ significantly from the ANES ones, and 32% of those flip sign.
- Ideology is recovered best. Party effects are stronger than in people (an S-shape).
- The other covariates do worst: education, age, race, gender, marital status, income, registration and interest.
-
Politics carries the persona. MAE is "basically identical" for politics-only and full prompts. Dropping politics inflates error for the parties, ideological groups, gays and lesbians, and Muslims.
- Footnote 23 (SI): a first-person prompt ("I am a …") exaggerates polarisation less but has higher MAE overall.
-
The same prompt gives different data months later. Linear fits of later scores on April scores:
- June: y = 4.6 + 0.93x, not distinguishable from a placebo resample of April.
- July, after the model update: y = 19 + 0.75x. That is strong mean reversion: the coldest scores warmed.
The authors cannot say why: the change might come from RLHF, from new material, or from something else.
-
Elsewhere is worse. Other questions, another US online survey and other countries perform "significantly worse" (SI, not read). The authors call the main setting a best case.
Limits
- One closed model family in the main text, one question type (thermometers) and US data. The persona is demographics plus politics: no individual history.
- The regression comparison uses the 30-draw mean as each synthetic respondent's answer. Inferred: that shrinks the synthetic residual variance, which bears on the significance tests; the SD comparison uses single draws.
- No test of what the model would need to reproduce individuals, as opposed to groups.
- Inconsistencies found:
- The text says "just 33 partisan respondents" at 99% power; Table 1 says 35.
- My computation (one-sample t, α = .05, two-sided, d = 12.5/16.1) gives 16, 17, 20, 24 and 33 for 80–99% power. Table 1 has 16, 18, 21, 25 and 35.
- For the ANES column my computation gives 130/148/173/213/300 against the table's 129/147/172/212/299, i.e. off by one throughout, probably a rounding convention.
- So the ChatGPT column differs from the text and from the formula at 85–99% power.
- The Figure 1 caption says "by prompt type/timing (columns)", but the extracted figure has one panel.
- The text says "just 33 partisan respondents" at 99% power; Table 1 says 35.
- Checked:
- 7,530 × 30 × 16 = 3,614,400.
- 16 groups − 5 not asked in the ANES = 11.
What it means for Kurisutina
- Base-model drift is real and invisible from the outside. The same prompt on the "same" API model gave
systematically different answers after an unannounced update. The coldest attitudes moved toward the middle
(slope .75).
- A replica on a moving API "changes" without its slot changing. That is exactly the change the goal rules out.
- This supports the design choice already made: pinned local weights and a pinned llama.cpp commit, with hashes recorded.
- The same applies to any judge used in Q2 over time. A judge behind an API can drift between sessions, so a longitudinal drift measurement would mix the replica's drift with the judge's. Pin judge versions or use local judges, and re-score a fixed control set at every session (proposal).
- "The average person with her views" is a caricature, not an average. Persona prompting reproduces group means
but makes partisans more extreme and less varied than the real people with those attributes.
- This supports keeping the pilot's own-against-other contrast. It also shows that "other" (a donor's record) is a fairer baseline than a demographic persona, which would be a harder-to-beat but wrong reference (inferred).
- Politics carries most of what a persona prompt conveys. Demographics add almost nothing once ideology and party
are given. This agrees with Kim & Lee.
- For Kurisutina, a persona built from coarse attributes will reproduce the political axis and little else. What is person-specific has to come from the person's own record.
- Perspective matters (inferred, from footnote 23). Second-person prompts exaggerate polarisation, and
first-person prompts are less extreme but less accurate.
- The pilot's arms differ in perspective as well as model: M1 is a third-person record continuation; M2 and M3 are first person.
- Arm differences therefore mix perspective with base against chat. Worth noting when the results are read; a later run could cross them.
- Read distributions; don't sample. Synthetic variance tracks the sampling temperature (SI, not read). The pilot reads the full next-token distribution and scores it with Brier, which avoids confusing sampling noise with a person's variability. Keep it that way for Q2 measurements.
- Never size a human study from replica outputs. The replica's overconfidence would underpower it by an order of magnitude.
Cross-references
summaries/carry_on/kim2024_ai_augmented_surveys.md: a trained per-person slot, and homogenisation even after training.summaries/carry_on/kiley2020_personal_culture.md: the human baseline for change over time.summaries/carry_on/lu2026_assistant_axis.md,summaries/carry_on/chen2025_persona_vectors.md: drift toward the model's default.docs/research/gss_pilot_design.md: arms, prompts and elicitation.