arXiv 2605.10659, version 1 (11 May 2026), the only version. York University (Toronto) and University Health
Network; funded by NSERC. A preprint with no venue stated, under the arXiv non-exclusive licence. The code is promised
on acceptance and is not released. Provenance: papers/carry_on/jia2026_liss_personas.provenance.json.
This is the paper chosen by the 2026-09-26 search: the closest open-access 2024–2026 study that predicts the same
people's later survey answers with LLMs. summaries/carry_on/wang2026_lifemem.md was read first and turned out to be
a same-wave design.
What was read
- All 1,613 lines of
pdftotext -layoutoutput (32 pages): main text, references, Appendices A–G and Tables 1–8. - Some pages were rendered and viewed; values below taken from figures are read off the images:
- page 3: Table 1, a colour matrix that pdftotext renders as names only;
- page 7: Figures 2–4. Figure 4's cell values are printed in the figure;
- page 23: Figure 9. The points were read approximately.
- Figures 1, 5–8 and 10–14 were not viewed; their captions were read.
- The code is not released.
Question
When can LLM "digital personas" stand in for real survey respondents? The test builds a persona for each LISS respondent from data before 2023 and predicts that respondent's answers from 2023 onward. It is scored per person and in aggregate.
Method
-
Data. The LISS panel (Centerdata, Netherlands), with a temporal cutoff of 2023.
- 6,276 respondents have valid records, each with 34 background variables.
- 2,923 standalone closed-ended question entries remain: ordinal, true/false, nominal and binary.
-
Two tasks, as defined in section 3.1:
- Single-wave prediction: the prior evidence is the respondent's most recent pre-cutoff core-study answers. The targets are single-wave (one-off) survey answers from 2023–2024.
- Core prediction: the prior evidence is the respondent's full pre-cutoff single-wave answer history. The targets are 2023 core-study answers.
- Neither task gives the persona the person's earlier answers to the target items (verified from the task definitions). Core modules repeat yearly, so those earlier answers exist, but they were not used. Single-wave targets are new questions.
-
Sample. 500 respondents per task, stratified proportionally by gender × age group × household stage. Within a stratum, respondents with more prior and target answers are preferred, which selects well-covered respondents. Eligible pools: 4,266 (core) and 5,785 (single-wave).
-
Four persona architectures:
- Background: the 34 variables only.
- Profile: background plus a GPT-5.4-written profile of the prior answers, at most 1,500 words under seven headings (personality traits, reasoning style, biases and heuristics and others). The raw answers are not included.
- Profile + lexical retrieval and Profile + semantic retrieval (text-embedding-3-small): the profile plus the K prior answer rows most related to the questions in the call, with K about 20.
-
Prediction.
- Three models: GPT-5.4, Gemini 3 Flash (preview) and Claude Haiku 4.5, with no chain-of-thought. That gives 12 settings.
- Target questions go in batches of about 20 per call, answered in JSON.
- One point answer per question. The instruction includes "If the answer is uncertain, choose the response most consistent with the available evidence."
- Calls that fail after the retries are scored as wrong.
-
No-context baseline. GPT-5.4 gets the question alone, 500 times per question. The temperature is not stated.
-
Metrics.
- respondent exact-match rate;
- question-level weighted F1;
- question-distribution Jensen–Shannon distance;
- respondent-distribution MMD;
- equity: the mean absolute deviation of subgroup/overall accuracy ratios;
- clustering ARI (k-means, k ≤ 7).
Uncertainty is a bootstrap SE over 100 respondent resamples.
-
Excluded baselines (B.7). Majority-class and demographic-cell majority baselines are excluded because they "would be oracle references rather than usable baselines": they need the held-out answers.
-
Section 6. Which predictions are right? For GPT-5.4 profile + lexical, 132,887 (core) and 30,725 (single-wave) respondent–question predictions are modelled on answer variability, domain, format and demographics. The methods are logistic regression, mixed-effects logistic regression, decision trees, random forests and XGBoost.
Results
-
Respondent exact-match rate (Tables 4–5):
Condition GPT core Claude core Gemini core GPT single Claude single Gemini single No context .478 — — .445 — — Background .507 .486 .515 .446 .430 .458 Profile .521 .497 .528 .455 .442 .467 Profile + lexical .529 .510 .536 .451 .442 .459 Profile + semantic .531 .510 .534 .452 .443 .462 My computations from these rows:
- the person's history over background alone: +0.021 to +0.024 (core) and +0.009 to +0.013 (single-wave);
- background over no context (GPT): +0.029 (core) and +0.001 (single-wave);
- best setting over no context (GPT): +0.053 and +0.010.
-
Question F1 (core, GPT): .306 with no context, .392 with background, .436 with profile + semantic.
-
Distributions gain more (core, GPT):
- question JSD .530 (no context) → .375 (background) → .306 (profile + semantic);
- respondent MMD .343 → .215 → .155.
-
Models. They differ little: Gemini ≥ GPT > Claude, and retrieval is best. The authors: "We therefore interpret model choice as secondary."
-
Domains.
- Closest: family and household, politics and values, religion and ethnicity.
- Farthest: social integration and leisure, and personality. The conclusion calls personas less reliable in domains that "depend on lived experience, self-assessment, or situational detail".
-
Where personas are right (Figure 4: Gemini 3 Flash, profile + semantic, single-wave; the F1 values are printed in the cells):
- low-variability questions: 0.75 for common respondents down to 0.65 for rare ones;
- very-high-variability questions: 0.25–0.28;
- medium variability with rare respondents: 0.29.
-
Variance is compressed (Figure 9, read off the plot).
- Persona answers carry about 0.45–0.75 of the human answer variance (single-wave) and about 0.5–0.7 (core).
- The no-context baseline carries about 0.16 and about 0.
-
Clustering. Human respondent clusters are not recovered. Table 4 lists core ARI as 0.000 for 9 of 13 settings (my count).
-
Equity. No large demographic gaps: DPI deviations are 0.007–0.032.
-
Section 6.
- Answer variability and the number of answer categories dominate whether a persona answer is right. Binary items are easier; demographics add little.
- XGBoost accuracy is .757 (core) and .761 (single-wave) in the behavioural layer.
-
Cost. About $3,000 of API calls and about 80 hours on a laptop.
Limits
- Not a forecast of change (verified from the design).
- The person's earlier answers to the same items are withheld, and the targets are other items after the cutoff.
- So there is no persistence (no-change) baseline and no transition table.
- B.7's "oracle" argument holds for new single-wave questions. It does not cover the core targets. For those, the pre-cutoff answer distribution and the person's own last answer existed before the cutoff (my inference).
- No other-person control (inferred).
- Nothing shows that the gain comes from this person's record rather than from any realistic record.
- Distributional gains exceed exact-match gains, which fits partly generic realism.
- Point answers under a "most consistent" instruction compress variance by construction. The distribution metrics then pool point predictions across people (inferred).
- Sample. Well-covered respondents are preferred, there are 500 per task, the items are closed-ended, and the panel is Dutch.
- Models. Closed API models. Model versions, dates and temperature are not stated, there are no repeat runs, and only GPT has a no-context baseline.
- Data handling.
- The answers of LISS respondents went to three commercial APIs.
- The paper's own data statement says: "Access is personal, and data users are not permitted to make copies of the data available to others."
- It does not say whether Centerdata approved sending data to model providers.
- Inconsistencies found:
- XGBoost. The text says "Since XGBoost achieves the strongest AUC". In Tables 2–3 the random forest has the higher AUC in 5 of the 6 task × layer cells, for example .819 against .817 (core, behavioural) and .773 against .762 (single-wave, behavioural). XGBoost leads on accuracy, not AUC.
- Tables 6 and 7 (my computation).
- Table 7's "Core" allocation column reproduces Table 6's single-wave sample exactly: 266 women and 234 men; ages 19.0/20.2/28.4/32.4%; household stages 39.4/35.0/25.6%.
- Table 7's "Single-Wave" column (262/238 by gender; ages 18.0/21.4/32.2/28.4%) matches neither sample.
- Table 8 gives identical MAD (0.2602) and MaxD (0.5275) for both samples.
- Recomputed from Table 7: 0.2812/0.4955 (single-wave column) and 0.2385/0.5490 (core column).
- Table 7's available counts sum to 3,867 and 5,727 (my computation), not to Table 8's eligible 5,785 and 4,266.
- Sparse strata. Section 3.2 says they are "merged with the nearest compatible stratum"; Appendix D says they are "excluded from allocation".
- Clustering. Figure 9 plots every single-wave setting at an ARI of about 0.035 or less, and core at about 0.07 or less. Tables 4–5 list 0.000 for most core settings and 0.044–0.316 for single-wave. The two cannot be the same statistic.
- "Baseline". B.2.1, on the Background persona, says "The baseline agent receives only the 34-variable background file". Everywhere else "Baseline" is the no-context condition.
- Checked:
- Table 6 totals: 266 + 234 = 271 + 229 = 500, and the percentages sum to 100;
- 4 architectures × 3 models = 12 settings;
- the cost arithmetic: $1,700 + $600 + $700 = $3,000; $3.40, $1.20 and $1.40 per respondent;
- predictions per respondent (my computation): about 266 (core) and about 61 (single-wave).
What it means for Kurisutina
- Question 1 and the pilot's own and empty conditions (inferred).
- The own history adds about 2 points of exact match over demographics (core) and about 1 point (single-wave).
- Against no information at all it adds 5.3 and 1.0 points (GPT).
- This is small, as the pilot's expectation E4 anticipates for the rest of the record.
- The pilot gives the target's earlier answer in every condition (part 2), which this paper withheld. So these numbers are neither a floor nor a ceiling for the pilot's D2 and D3. They measure transfer across items over time.
- The typical person is easy.
- Accuracy is highest for common respondents on low-variability items. It is lowest where the respondent is rare or the item splits people.
- That is where react like her, not like the average person has to show.
- Proposal: split the pilot's and the LISS test's scores by respondent answer rarity and by item entropy (Appendix C.1.2), next to the declared changed/unchanged split (X1). Run it on existing outputs, after the primary analysis.
- Homogenisation. Personas carry about half to three quarters of human answer variance, like Bisbee et al. and
Kim & Lee.
- The pilot's probability read-out keeps spread for scoring.
- A replica that generates answers open loop has no such protection (inferred).
- For the LISS test (
docs/research/liss_q1_design.md):- Outcomes (inferred). The Q1 outcomes (happiness, life satisfaction, mood, Big Five) are in the personality domain, which is where personas were worst. Expect low ceilings.
- Baselines. T2 must give the person's own earlier answers to the same items. It must be scored against no change and the population table. Jia et al. have neither, and adding them turns this design into a carry-on test.
- Conditions. The "other" condition, which T2 already has, is the control this paper lacks.
- Slot format. Raw answer rows, retrieved, beat a GPT-5.4 trait profile, though narrowly: .521 → .531 for GPT, core. This agrees with Kinzinger & Hartmann (raw answers beat summaries) and with the pilot's raw record. A trait-and-bias summary is a weaker slot than the person's own answers.
- Base-model defaults.
- The no-context baseline is the empty slot, but only for one model.
- With point answers it is nearly deterministic: its answer variance is about 0 in the core task (Figure 9). It shows the model's modal default, not a distribution.
- Nothing here controls drift across model versions. For that, pinned local weights remain the pilot's approach.
- Data handling. This paper is not evidence that Centerdata allows LISS data to go to model APIs. The project's rule of local models only until Centerdata answers stands.
Cross-references
summaries/carry_on/wang2026_lifemem.md: same search; a same-wave design with population metrics only.summaries/carry_on/lundberg2024_unpredictability.md: irreducible error; why rare patterns stay unpredictable.summaries/carry_on/kim2024_ai_augmented_surveys.md: GSS cross-item prediction with a trained per-person slot.summaries/carry_on/hwang2023_user_opinions.md: the person's own relevant answers beat attributes; retrieval hijacks.summaries/carry_on/kinzinger2026_soep_twins.md: raw answers beat summaries; the empty-persona baseline.summaries/carry_on/bisbee2024_synthetic_replacements.md: compressed variance.summaries/12_park2026.md: the two-week retest ceiling.docs/research/liss_q1_design.md: T1/T2 and data handling.docs/research/gss_pilot_design.md: the own, other and empty conditions and the tables.