Kurisutina

Synthetic Personalities: How Well Can LLMs Mimic Individual Respondents Using Socio-Economic Microdata?

arXiv 2606.04592 v1 (3 June 2026), cs.CY. TUM School of Management, Technical University of Munich. Preprint, no peer review indicated. Provenance: papers/carry_on/kinzinger2026_soep_twins.provenance.json.

What was read

All 1,554 lines of the pdftotext -layout text of the 41-page PDF: sections 1–5, Tables 1–3, figure captions and the references. Figures are images; their captions and extracted labels were read. The "Web Appendices" A–D that the text cites (filters, prompts, the empty-persona correlation panel, the random-selection ablation, per-cell tables) are not in this PDF and were not read.

Question

Can detailed individual digital twins be built from an existing, heterogeneous panel dataset rather than a purpose-built one? And which construction choices matter: model, amount of context, form of the context, reasoning mode?

Design

  • Data. German Socio-Economic Panel (SOEP-Core v40), 2023 wave only, used as a cross-sectional snapshot.
    • 16,055 participants after filters; 949 usable question-answer pairs each: 38 demographic, 728 persona-context and 183 held-out items (a random 80/20 split).
    • A fixed random subsample of 500 participants is used for the grid.
  • The 60-cell grid.
    • Three open-weights models on one H100 with vLLM: Qwen3-30B-A3B, Gemma-4-26B-A4B, Ministral-3-14B.
    • Five cumulative context depths: demographics, then quartiles of the persona items ranked by normalised Shannon entropy.
    • Two embeddings: a Chain-of-Density narrative summary, or the raw question-answer "dialog".
    • Two reasoning modes: thinking or not.
    • 2.1 million twin answers were scored. Sampling temperature is above 0; answers were normalised by an LLM judge.
  • Metrics, following Peng et al.:
    • accuracy, 1 − normalised distance (partial credit for ordinal and metric items);
    • Fisher-z-averaged Pearson correlation across participants per question;
    • the twin/human SD ratio.

Results

  • Best cells. Accuracy 0.788 (Gemma 4, dialog, thinking, 100% depth). Correlation r = 0.590 (Qwen 3, dialog, thinking, 100%). Qwen 3 averages 0.730 / 0.431 / SD ratio 0.99; Gemma 0.717 / 0.426 / 0.88; Ministral 0.684 / 0.388 / 1.07.
  • More context helps, concavely. From demographics to 100% depth, accuracy gains 5.2 points on average and correlation 21.9 points. Most of the gain comes in the middle quartiles; 75% depth is the stated cost-efficient point.
    • Ranking items by entropy made no difference against random items of the same number (p ≈ 0.32): volume matters, not selection.
  • The empty-persona baseline is flat at 0.65–0.66 accuracy. The personalisation gain grows from +4.2 points (demographics) to +10.8 points (100%), against +1.4 points in Peng et al.
    • Gains concentrate on "hard" items where the empty-persona model is wrong (+8.6 points). Easy items gain +2.2 points and are already at 88–91%: the person's answer is the modal one.
    • The authors: empty-persona ablations "should become standard hygiene".
  • Raw answers beat summaries. Dialog raised accuracy over the narrative summary in every model × reasoning cell at 100% depth (Gemma +6.2 points, Ministral +7.2, Qwen +0.8 on average), at 2.8 times the tokens (14,863 vs 5,354). The summary compresses graded answers into prose, for example the lowest point of a six-point scale into a plain negation.
  • Thinking mode raised correlation (Qwen +0.073, Gemma +0.036) but not accuracy, at 3–21 times the inference time.
  • Individuation varies by model. Under-individuation (SD ratio below 1) shows only for Gemma (0.88, raised to 0.93 by thinking). Qwen is near parity; Ministral is slightly over-dispersed.

Limits

  • Not longitudinal. Despite a 41-year panel, twins predict held-out items from the same 2023 wave. Nothing tests prediction of later answers or change after events.
  • The held-out items are a random sample of the same item pool, so near-duplicate items may be partly retrievable from the context.
  • The accuracy metric gives partial credit and no random baseline is reported; Peng et al.'s random baseline on the same metric family is 0.629. LLM-judge normalisation of answers; sampling above temperature 0.
  • Three models; German-language items; no closed models (licence); no test of ideological drift (acknowledged).
  • Inconsistencies found:
    • Figure 6 says "n = 4 models"; there are three.
    • The text gives per-model medians of the SD ratio as 0.99, 1.07 and 0.88, which are Table 3's means. Figure 7 prints medians of 0.88 (Qwen), 0.96 (Ministral) and 0.82 (Gemma).
    • Correlation gains are reported in "percentage points".

What it means for Kurisutina

  • It supports two choices already in the GSS pilot:
    • The pilot's slot is the person's raw earlier answers, one line per question with its code, not a summary. Here the raw form beat a narrative summary in every cell because summaries destroy graded answers.
    • The pilot's empty condition is the "standard hygiene" the authors call for; D3 is exactly their personalisation delta.
  • The gain lives where the person is not the mode. Context helps most on items where the empty model is wrong, and barely where the person gives the common answer. This is the item-level version of the changed/unchanged split (declared analysis X1): for carrying on, the informative cases are the ones where the person departs from what anyone would predict.
  • Panel microdata individuate more than psychometric batteries: +10.8 points against +1.4 points in Peng et al. That is cross-sectional, and our pilot asks the harder question of whether the record predicts the later answers after an event. This paper gives no evidence on that. The GSS pilot, and the carry-on tests after it, are where it gets answered.
  • Relevant to model choice for a rented run. A Qwen3-30B-A3B MoE on one H100 was the best all-round open model here. It is a candidate additional arm for the v2 run, alongside the dense 27B, to separate model size from architecture.

Cross-references

  • summaries/carry_on/peng2025_funhouse.md: the metrics and the empty-persona comparison this paper reuses.
  • summaries/carry_on/hullman2026_validating.md: what a held-out, randomly sampled evaluation can support.
  • docs/research/gss_pilot_design.md: the record format (part 1), the empty condition (D3), declared analysis X1.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.