arXiv 2509.19088 v5 (19 April 2026), preprint; earlier versions were titled "A Mega-Study of Digital Twins Reveals
Strengths, Weaknesses and Opportunities for Further Improvement". Columbia (22 authors), Yale, Yeshiva.
Provenance: papers/carry_on/peng2025_funhouse.provenance.json.
What was read
Everything in the extracted text of the 223-page v5 PDF, line by line: main text (pp. 1–27), Materials and Methods, all 94 references, acknowledgments, author contributions, data statement, and the whole Supporting Information: A (table of the 31 pre-registered comparisons), B (prompt template), C (persona constructions), D–F (figure pages; figures are images and were not inspected, only their captions), G (implementation variants), H (Centaur), I (expert forecast survey), J.1–J.19 (every sub-study's hypotheses, methods, pre-registered and non-pre-registered results, tables and discussion). Not read: the pre-registrations themselves (researchbox 4145, aspredicted 65ty-73sp), the released data and code, and the Twin-2K-500 dataset paper.
Question
Do LLM "digital twins" built from a person's own rich survey data reproduce that person's responses to new questions and experiments, and in what systematic ways do they distort them?
Design
- People. Members of the Twin-2K-500 panel (Toubia et al. 2025): ~2,000 US Prolific participants, representative on demographics, who had answered >500 questions over four waves (≈145 min): 14 demographics, 279 personality items (19 tests, 26 constructs), 85 cognitive items (11 measures), 34 economic-preference items (10 measures), 48 heuristics-and-biases items (16 experiments), 40 pricing items. Wave 4 repeated earlier items (test–retest).
- New studies. 19 pre-registered sub-studies designed by 23 co-authors, run on Prolific April–June 2025 (the persona data were collected from about February 2025): 164 outcomes, 13,506 human responses from 1,784 unique participants. Mix of classic paradigms, unpublished and newly designed stimuli. Each twin got exactly the survey its human got, including the same randomised condition.
- Twins. In-context only: GPT-4.1 (2025-04-14), temperature 0.7, prompt "answer … as if you are the person described … accounting for human cognitive limitations, uncertainty, and biases", JSON output validated and retried on failure. Full persona ≈ 30K tokens (128K characters) of question–answer text (wave-4 answers where repeated).
- Benchmarks. Uniform random; empty persona (same prompt, persona replaced by "[Empty Persona Profile]": the base model alone); demographics-only persona (the 14 demographic answers, a prefix of the full persona); persona summary (≈3K tokens, scores with percentiles); other base models (GPT-5, DeepSeek-R1, Gemini-2.5-flash, Gemini-3-pro, GPT-4.1 fine-tuned on Twin-2K-500); temperature 0; Centaur and Llama-3.1-70B; XGBoost trained per outcome on 50–650 of the humans' actual answers.
- Metrics per outcome. Individual accuracy = 1 − MAD/range, averaged over people; correlation across people between human and twin answers (Fisher-z averaged); |Glass's Δ| between means; SD ratio. Note: Park et al. correlate across questions within a person, a different and more flattering quantity.
Main results (numbers)
- Accuracy barely moves with personal data. Full persona 0.748; demographics-only 0.746 (p = .37); empty persona 0.734 (difference 0.014, significant); uniform random already 0.629.
- Correlation (sorting people) improves but stays weak. r = 0.197 full persona; 0.145 demographics; 0.080 empty; 0.001 random. Positive for 157/164 outcomes, significant for 97 (59%). Best configuration (GPT-4.1, temperature 0): r = 0.232, accuracy 0.752. No other base model, the fine-tuned GPT-4.1, or Centaur did better; the 3K-token summary performed like the full persona.
- Means. Twin means differ from human means by 0.352 SD on average; significant for 105/164 outcomes.
- Blueprint data carry little signal about these outcomes. XGBoost on the full persona, trained on 650 people's real answers for that outcome, stays below r = 0.29. Twins match XGBoost trained on ≈180 people (correlation) and ≈75 people (accuracy).
- Experts (68 academics, managers, students) predicted human treatment effects better than twin treatment effects (error −0.072, p < .001); the twin effect fell outside their 95% interval in 3 of 5 tasks.
The five distortions
- Insufficient individuation. Twin SD < human SD for 154/164 outcomes (140 significant). Full-persona twins are closer to the empty persona (MAD 0.175) than to their own human (MAD 0.252, p < .01): "overly shrunk towards a base model".
- Stereotyping. Full-persona twins are closer still to demographics-only twins (MAD 0.132). Worked example (Affective Primes, "lack of control"): r 0.105 → 0.555 with the full persona while accuracy only 0.892 → 0.907: personal data re-sort people without getting their levels right.
- Representation bias. Per-person twin accuracy (XGBoost on 61 demographic dummies, partial dependence) is higher for more educated, higher-income, politically moderate and moderately religious participants.
- Ideological bias. Twins are more "pro-human" (people are fair and trustworthy; pro-regulation; taxes for health care; kinder to people who donate to both parties) and more "pro-technology" (less averse to hiring algorithms, see targeting as less intrusive), and show human aversion when rating ideas with the source named.
- Hyper-rationality. Near-perfect factual knowledge (fee definitions 99.9% vs 51.8% for humans), normative answers (guesstimate 99.9% vs 59.4%; synthesise 99.8% vs 71.5%), no attraction effect on new stimuli, a default effect only in the classic (probably leaked) organ-donation paradigm, not in the new green-energy one where humans showed it.
Sub-studies that bear on a replica that changes over time
- Affective primes (J.2, N = 1,000). Writing about gratitude or lack of control shifted twins more than humans (gratitude check +1.33 vs +0.81; lack of control +1.00 vs +0.32), while some downstream effects were stronger in humans (fatigue, desire for predictability). The replica over-reacts to an induced state on the measure next to it and under-transmits it further out.
- Self-persuasion (J.13, N = 1,001). After arguing the other side of universal basic income, humans moved 6.17 points (SD 14.8), twins 2.80 (SD 2.4); the prompts differed slightly between humans and twins. In humans, more extreme attitudes moved less (b = −0.04); in twins, more extreme ones moved more (b = +0.06, regression to the mean). The authors read this as twins not reproducing their human's metacognition (attitude certainty). This is a direct case of a replica not updating like its person.
- Imagined scenarios (J.13). Twins were less responsive than humans to imagined conflict and time pressure, more to imagined pain; not more "obedient".
- Treatment effects (J.9, hiring). Twins recovered the human effect of algorithmic hiring on some outcomes (apply, hired, timely) but under-estimated the negative effect on sociality, fairness, communication and talent identification. Between-person slopes (human answer on twin answer) were 0.04–0.70.
- Knowledge and experience. Twins under-reported luxury ownership (25.8% vs 53.8%), reported almost no variation in father's education (~100% "high school"), and reported TikTok use as "never" for 51% while most humans used it monthly or weekly: the persona did not carry these facts, and the base model filled them in.
- Content from the person. Text written by a twin was only slightly more similar to its own human's text than to a random human's (BERTScore difference ≈ 0.01).
Limits and problems
- Snapshot, not carrying on. One-shot surveys and experiments within one session each; no memories, episodes or sequential experience; persona is questionnaire data, not narrative or autobiographical memory. The persona was collected ~2–5 months before the new studies, and no human retest exists for the new outcomes, so part of the error may be real change or noise in the humans.
- One construction method. Everything is in-context prompting of general models; fine-tuning was tried only conservatively (one example per participant) and lost instruction following.
- Metric. 1 − MAD/range gives 0.629 to random answers, so accuracy differences are compressed; the correlation metric is the informative one.
- Internal inconsistencies found in the SI (not affecting the headline numbers): J.2 reports own-twin vs random-twin BERTScore means (0.11 vs 0.13) opposite to the stated direction; J.8 gives r = 0.546/0.384 in results and 0.416/0.420 in its discussion; J.11's accuracy-versus-entertainment means (humans 2.83, twins 2.12 on 1 = correct … 6 = entertaining) contradict the sentence that humans put more weight on accuracy; J.18's twin means (targeted 7.30, broad 6.50) contradict the text saying twins rated targeting less fair; J.19's discussion says twins report more usage while the main text and Table 1 say they under-report (the data show less for TikTok, more for Netflix).
What it means for Kurisutina
- The user's drift criterion already has a static measurement template here. The distance of the replica from the empty-persona model versus its distance from the person (0.175 vs 0.252) and from a demographic stand-in (0.132) is exactly "is it the slot or the base model?". For a replica that carries on, the same three distances can be tracked after every new experience. Under-dispersion across people (SD ratio) is a second, cheap indicator.
- Personal updating (question 1) is not reproduced by in-context twins. Twins over-react to induced states, under-react to counter-argument, and reverse the human relation between attitude extremity and change. Knowing a person's answers does not give the model their way of changing; that has to be measured and represented, which fits the user's acceptance of several interviews.
- More generic psychometrics is not the answer. Five hundred well-validated items left even a supervised model below r = 0.29 on these outcomes. What predicts a person's behaviour in a domain has to be collected for that domain, which supports the brief's targeted, per-domain elicitation over broad batteries.
- The replica must not know more than the person. Hyper-rationality and perfect knowledge are the base model leaking through; the brief's explicit-ignorance rule (5.3) needs to cover general knowledge and skills, not only memories.
- Sorting is not being the person. Personal data mostly improved who-is-higher-than-whom, not the person's own level; a replica is judged on the latter.
Cross-references
- Park et al. 2026 [brief ref 12;
summaries/12_park2026.md]: the interview-based twins this paper contrasts with; different correlation metric. docs/archive_analysis.md: our own finding that the most common answer beats demographics on Park's data, and that interview agents do not beat it for 189 of 1,052 people, is the same "insufficient individuation" seen from another dataset.