What was read
- arXiv v3 PDF, 86 pages: main text (abstract, significance statement, introduction, methods, results, discussion, materials and methods summary, references) and the full Supplementary Materials (agent bank construction, AI interviewer, agent architecture and prompts, constructs, evaluation methods, supplementary results, pre-registration deviations, mechanism analysis, model comparisons, Tables 1–9 including the interview script and all regression and parity tables). Read in full.
Question
Can a large language model, given a person's own self-reports, predict that person's held-out attitudes, traits and behaviours without task-specific training, and which kind of self-report (interview, survey, both) works best?
Design
- Sample: 1,052 U.S. adults recruited through Bovitz to census quotas on age, gender, race, region, education and party; final sample slightly more educated, female, Democratic and Midwestern than target (Table 9). Paid $60 plus $30 for the retest plus game bonuses.
- Data per person: a two-hour voice-to-voice interview with an AI interviewer ("Isabella", GPT-4o with a reflection module, OpenAI TTS and Whisper) running an abbreviated American Voices Project script of 99 questions with adaptive follow-ups (mean 82 follow-ups; participants spoke 6,491 words on average); the GSS core (177 categorical and 6 numerical items); the BFI-44; five incentivised economic games; five text-based experiments from the Camerer et al. replication set; a self-written persona paragraph. All measures except the interview repeated two weeks later.
- Agents: GPT-4o prompted with the transcript, the survey responses, or both, plus stored "expert reflections" (up to 20 notes each from a psychologist, behavioural economist, political scientist and demographer persona); at query time the relevant expert's notes are appended and the model reasons step by step. Baselines: demographics-only (age, gender, race, ideology) and persona-paragraph agents.
- Leakage control: the outcome item is removed from survey-agent input (whole trait block for Big Five); 27 GSS items with near-duplicate interview questions were removed from the evaluation after a GPT-4.1 plus human screen of 54,694 question pairs (a deviation from the pre-registration; including them raises accuracy by 0.02).
- Metric: normalised accuracy = agent accuracy on wave-2 responses divided by the participant's own wave-1 to wave-2 consistency; 1.0 means the agent predicts the person as well as the person predicts themselves two weeks later.
Results (with numbers)
- GSS (150 items, chance 0.30). Participant self-consistency 79.5%. Raw / normalised accuracy: interview 65.7% / 0.83; survey 65.0% / 0.82; survey plus interview 68.3% / 0.86; demographics 58.1% / 0.74; persona 56.2% / 0.71. All three self-report agents beat both baselines (p < .001); interview and survey do not differ (p = .16); combined beats each alone (p < .001) by about 0.03–0.04.
- Big Five (normalised correlation): interview 0.80 (raw r = 0.78 vs self r = 0.95), survey 0.65, combined 0.77, demographics 0.61, persona 0.75. Survey-only agents do not beat the baselines here; combining adds nothing over the interview.
- Economic games: interview 0.66, survey 0.38, combined 0.49, demographics 0.48, persona 0.57; no significant differences (F = 1.63, p = 0.16). The authors call the games "a boundary case rather than a domain where self-reports clearly help".
- Experiments: agent-estimated effect sizes are larger than the human replication's but track them (r = 0.91–0.99 across five studies). Individual-level prediction of experimental responses is weak: normalised accuracy 0.50 for interview agents against a human test-retest accuracy of only 0.48 (Table 6 continued).
- Ablations (100-agent subsample, Table 6): removing 80% of the transcript at random (about 96 of 120 minutes) drops GSS normalised accuracy only from 0.82 to 0.79; a bullet-point factual summary of the interview scores 0.81, so content, not linguistic style, carries the prediction; maximal agents with everything score 0.87.
- Mechanisms (SI 8, 59 agents): performance falls as the GSS items most likely to have a verbatim answer in the transcript are removed, and again as items answerable by inference are removed; after removing the 30 most retrievable and 30 most inferable items, interview agents still lead demographics by 0.080. Both direct retrieval and inference contribute and neither fully explains the gap.
- Bias: demographic parity differences shrink with self-reports (GSS ideology gap 13.75 points with demographics → 8.60 interview → 6.22 survey). Regressions still show better prediction for strong liberals, strong Democrats and non-heterosexual participants.
- Models: on 50 agents, GPT-5, GPT-4.1, o1 and o3 all score 0.67 raw versus 0.66 for the 2024 GPT-4o; a 2025 GPT-4o run gives 0.64; mini models trail by 2–6 points. Fine-tuning GPT-4o on 500 agents' answers does not help (0.79 vs 0.84 normalised).
- Consent regime: participants told their data may leak despite pseudonymisation, that models "may become increasingly powerful over time", with withdrawal honoured for 25 years.
Authors' conclusions: self-report grounding recovers 82–86% of a person's own test-retest consistency on the GSS, "the practical ceiling"; gains from combining sources are modest, "pointing to diminishing returns once a person is described in sufficient depth"; quality of simulation "depends far less on model scale or synthetic persona engineering than on the depth and reliability of the data an agent is built from".
Limits (the paper's own and others)
- Outcomes are survey items, personality inventories, one-shot games and vignette experiments. Nothing about episodic memory, relationships, what comes to mind under a cue, or how the person changes.
- Why it works is not understood; retrieval and inference from the transcript are shown to matter, but a residual advantage remains unexplained.
- Individual-level accuracy only; no test that agents reproduce the covariance structure of attitudes across people (prior work says LLM responses over-correlate).
- Five experiments, underpowered for effect-size claims; agents over-react to treatments.
- One model family (OpenAI), English only, U.S. only.
- The normalising denominator is two-week test-retest on the same instrument, which is high for the GSS (79.5%) and low for experiments (48%); "83% of ceiling" means different absolute things in different domains.
- The AI interviewer was validated by the team's own judgement against 10 human-conducted pilots; no independent quality metric.
What the brief uses it for, and whether it holds
Brief section 0 and 4: 83%/82%/86% vs 74% demographics; gains from combining modest; benefit asymptotes once the model has seen enough evidence within a domain; attitude-type outcomes only; "the baseline this brief has to beat". Section 6.5: on attitude items, predictive benefit saturates within about two hours and the person-specific increment over demographics is small. Section 7: demographics-only agent as the non-personalised floor; person's own test-retest as the ceiling. Section 9: a two-hour interview baseline and a demographics-only floor must exist before neural data is collected.
- All numbers hold. "Asymptotes once the model has observed sufficient evidence within a domain" is the authors' wording.
- The brief's "saturates within about two hours" is, if anything, generous: the lesion analysis shows that a random 20% of the interview (about 24 minutes) recovers 0.79 of 0.82 on the GSS. The attitude curve is nearly flat well before two hours. The brief's per-domain efficiency curve (6.5) should start from that point, not from two hours.
- The person-specific increment over demographics is 0.09 normalised (about 7.5 raw points). The brief's framing, "small and saturates fast", is accurate for the GSS. For the Big Five the increment is larger (0.80 vs 0.61) and for games it is zero.
- The paper contains the closest existing analogue of the brief's functional-use test and it is the weakest result: individual responses in the five experiments are predicted at 0.50 of a 0.48 test-retest. That is evidence for the brief's claim that "the room is in ... change after new experience", and also a warning that the human ceiling in that domain is low, so scoring on predictive distributions (5.2) rather than exact matches is the right choice.
- Two details the brief does not use but should: (a) the persona paragraph a person writes about themselves (0.71) predicts worse than four demographic facts (0.74), a caution for any records-based "voice" component; (b) accuracy is higher for some ideological and sexual-orientation groups, so a base model's priors already personalise unevenly, which is relevant to the person-specificity test's controls.
- The AI interviewer is a working, published instance of the brief's step-1 agent, minus the selection rule: its follow-ups aim to "achieve the interview objective", not to discriminate between candidate person-models. Its reflection notes are a primitive of the brief's association-and-salience graph. The code is public (github.com/joonspk-research/generative_agent).
- The consent regime (25-year withdrawal, warning that future models may infer more) is a concrete template for section 10.2.
Cross-references
- [9] Verge: the impressionistic predecessor; Park is the measured version of "records or self-report → agent".
- [3] Marek: the small-n caution applies to any pilot that tries to beat 0.83 with two or three participants.
- [13] Anderson: the brief pairs Park's behavioural baseline with Anderson's neural zero-shot arm in step 3.