arXiv 2305.14929, version 1 (24 May 2023), the only version. University of British Columbia and Vector
Institute, UC San Diego, and the Allen Institute for AI. CC BY 4.0. Code: github.com/eujhwang/personalized-llms (not
read). A short workshop-style paper; no venue is given on the abs page. Provenance:
papers/carry_on/hwang2023_user_opinions.provenance.json.
What was read
All 639 lines of pdftotext -layout output (13 pages): text, Tables 1–6, the captions of Figures 1–4, the prompts in
Appendix A (Figures 5–8), Appendix B and the references. Figures 1–2 are images; only their captions were read. The code
was not read.
Question
To predict one person's answer to a survey question with an LLM, what should the prompt hold: their demographics, their ideology, their own past answers, or all three?
Method
- Data. OpinionQA (Pew American Trends Panel): 15 topics, about 100 questions and 5,340 people per topic.
- Descriptive analysis.
- Pairs of people with identical demographics: agreement (Cohen's κ) on their answers.
- Pairs with at least 10 common questions: "similar opinions" means 70% of answers match. The analysis counts how often such pairs also share an ideology.
- Prediction. Zero-shot multiple choice with text-davinci-003.
- 100 people per topic. For each person, 20% of their answers serve as their "opinions" (at most 16), and the rest are test questions.
- Prompt variants: no persona; ideology; demographics + ideology; ideology + opinions; demographics + opinions; all three.
- Opinions are either all (up to 16) or the top-k most similar to the question, by ada-002 embeddings. k was 3, 5 or 8; results are reported for 3 and 8.
- Metrics: exact accuracy, and "collapsed" accuracy with the options merged into two sides.
- Group level. The same model is asked to predict the majority answer of each ideological group.
Results
-
Same demographics, different people. Pairs with identical demographics agree only moderately.
- The text says κ "around 0.5"; the Figure 2 caption says "around 0.4".
- Of pairs with similar opinions, 60–82% have different ideologies (Table 2).
-
A person's own relevant answers beat their attributes (Table 3, exact / collapsed accuracy):
Prompt Exact Collapsed no persona .43 .62 demographics + ideology .47 .65 top-3 own opinions only .51 .67 demographics + ideology + all 16 opinions .51 .69 demographics + ideology + top-8 opinions .54 .70 - Three well-chosen own answers alone match the attributes plus 16 random answers.
- Top-3 and top-8 perform alike: "a few of the most relevant opinions carry the most performance improvement".
- Demographics add a little on top of ideology and opinions (.53 to .54).
-
By topic (Table 4, exact). Gains from own opinions over demographics + ideology are largest for:
- gender and leadership (+.14);
- views on gender (+.13);
- guns (+.12).
Automation gets slightly worse (.49 to .48).
-
Typical error (Figure 4). A retrieved opinion containing the phrase "does not describe me well" made the model pick that option for an unrelated question. It had been right with demographics alone.
-
Groups are easier than individuals. Predicting an ideological group's majority answer reaches .52–.58 exact. Predicting individuals from demographics + ideology reaches .47. The targets differ (group majorities against individual answers), so the comparison is only indicative. The authors: LLMs are "good at modeling a representative individual of a sub-population".
Limits
- One closed, since-retired model (text-davinci-003), zero-shot. 100 people per topic, and the ±.01 intervals are pooled across them.
- Opinions and test questions come from the same survey wave. Nothing here concerns change over time.
- "Relevance" is embedding similarity to the question, which invites surface copying (the paper's own error example).
- There is no human test–retest ceiling, and no own-against-other-person control. The comparison is against attribute prompts only.
- Inconsistencies found:
- Agreement of demographically identical pairs: κ "around 0.5" in the text, "around 0.4" in the Figure 2 caption.
- Table 5: "Avg overall" (.566 / .667) is not the mean of the three ideology rows. That mean is .549 / .659, which is exactly the "Majority answer" row, the one the text describes as the model without ideology information. The rows or the averages are mislabelled.
- The abstract says "up to 7 points". That is the pooled gain (.47 to .54). Topic gains reach 14 points (Table 4).
- The method lists k ∈ {3, 5, 8}, but no top-5 result is reported. The Table 4 caption mentions collapsed accuracy, but only exact accuracy is shown.
- Checked:
- The topic means of Table 4 (.432, .468, .538) match Table 3's pooled .43, .47 and .54.
- Tables 2 and 6 rows sum to 100%.
What it means for Kurisutina
- Individual-level evidence for the own record. Given a person's own answers, especially the few most relevant
ones, an LLM predicts their other answers better than from their demographics and ideology. This is Kim & Lee's
result without training.
- For the slot: the person's own statements, retrieved by relevance, carry what attributes cannot.
- Demographics add little once the person's own answers are present. This agrees with Kim & Lee and Bisbee.
- The model gravitates to the group representative. It predicts a group's majority better than an individual
member from the same prompt, which is the pull toward "the average person with her views".
- The pilot's own-against-other contrast measures exactly that gap. A donor with the same event and earlier answer stands in for the group representative (inferred).
- Retrieval can contaminate. A retrieved memory that shares wording with an answer option can hijack the choice.
- For the replica's memory: rank by relevance of content, and test with probes where a stored statement shares surface wording with the wrong option (proposal).
- For the pilot (inferred): the background lists every earlier answer with its value label. Lexical overlap between those labels and the target's options could act the same way. It is worth checking in the results whether own-condition errors copy labels from the background.
- Exact against collapsed. Getting the side right (collapsed about .70) is much easier than the exact option
(about .54). This matches Kim & Lee's binarisation caveat.
- The pilot's Brier on the full code distribution is the stricter measure; a side-level score could be reported alongside it (proposal).
- Same time, not over time. Like Kim & Lee, this is cross-item prediction at one time point. It says nothing about carrying on.
Cross-references
summaries/carry_on/kim2024_ai_augmented_surveys.md: a trained per-person slot on GSS data.summaries/carry_on/bisbee2024_synthetic_replacements.md: persona prompts caricature groups; model versions drift.summaries/12_park2026.md: interview-grounded agents of 1,052 people.docs/research/gss_pilot_design.md: own, other and empty conditions.