Kurisutina

This human study did not involve human subjects: Validating LLM simulations as behavioral evidence

arXiv 2602.15785 v1 (17 February 2026), cs.AI, CC BY 4.0. Northwestern, Stanford. A 41-page position and review paper; no peer review indicated. Provenance: papers/carry_on/hullman2026_validating.provenance.json.

What was read

All 41 pages: sections 1–7, the references, and appendix 8.1–8.6. The appendix covers the formal leakage and moment conditions (after Ludwig et al. 2025), Perkowski's potential-outcomes treatment of "digital twin" experiments, a worked example of correlated bias, the PPI weight, plug-in bias correction, and a table describing how about 50 studies validated LLM simulations. Figures were read from their captions.

Question

When can answers generated by a language model stand in for human participants' answers as evidence about human behaviour, and what kind of claim can each validation strategy support?

Argument

  • Three strategies.
    • Heuristic validate-then-simulate: show the model resembles humans somewhere, then use it where there is no human data.
    • Simulate-then-validate: use models to explore and prioritise designs, then test the promising ones on humans.
    • Statistical calibration: combine a human sample with model predictions using estimators whose validity does not depend on the model being right.
  • How heuristic validation is done (from Anthis et al.'s sample of 53 studies and others):
    • Direction and significance of effects. Cui et al. 2025: up to 81% of 156 studies' main effects replicated, but also up to 83% of null effects came out "significant".
    • Effect-size correlations. About 0.85 (Cui) and about 0.5 (Hewitt) in the main text; the appendix gives Hewitt r ≈ 0.85–0.94, inconsistent within the paper. Models overestimate effect sizes.
    • Individual prediction, normalised by people's own retest consistency: Park et al. 2024a, 85% on the GSS; Toubia et al. 2025, 88%.
    • Distributional distances, Turing-style tests, consistency with theory, face validity, elicited rationales, and alignment of internal representations with fMRI.
  • Threats.
    • Systematic bias. Model answers have lower variance than human answers (six studies). Group differences are exaggerated. Models caricature identities, resembling what outsiders expect of a group more than what members say (Wang et al. 2025a). They are less consistent for "incongruous" personas, e.g. a conservative open to immigration (Liu et al. 2024). Accuracy is worse for under-represented groups and sensitive topics (replication 77% → 42% for race and gender).
    • Memorisation. Only 1 of 53 studies tested on scenarios the model could not have seen.
    • Misleading generalisation: brittleness, "potemkin understanding", unfaithful explanations.
    • A noisy human gold standard (low power, replication failures).
    • Prompt, temperature and fine-tuning repairs improve resemblance but "do not provide a basis for valid inference".
  • Formal conditions for plugging model answers in (Ludwig et al. 2025):
    1. No training leakage. No study scenario may appear in training data. Otherwise measured accuracy overstates true accuracy. There is a workaround when only predictive accuracy is the goal: define a finite population of scenarios, sample it randomly, and evaluate on held-out data (Manning & Horton 2025). That is valid even if some prompts were memorised.
    2. Preserved moment conditions. Even small errors bias downstream estimates if they correlate with the covariates of interest. Egami et al.: 90%-accurate labels gave about 30% bias and about 40% CI coverage. Worked example: an error of +δ/2 in one group and −δ/2 in the other shifts the slope by δ. This condition is practically impossible to verify ("if the researcher could be confident this was the case, they would no longer need to rely on LLM simulations").
  • Perkowski 2025 on digital-twin experiments.
    • A twin's error = person-specific bias + condition-specific bias + noise.
    • The person-specific bias cancels in within-person contrasts, provided the model does not favour one condition (TISA). If person and condition biases interact, the average interaction must also vanish (IISA).
  • Statistical calibration.
    • Prediction-powered inference (PPI): the human-only estimate, adjusted by a tuned multiple of the model's labelled-minus-unlabelled prediction difference. It is always at least as precise as the human-only estimate and centred on the human estimand.
    • Also: design-based supervised learning, doubly robust versions, and learned bias corrections (with cross-fitting), or a regression of human answers on model answers (Wang et al. 2020).
    • Gains so far are modest. Adding 100,000 model decisions to 10,000 human decisions raised the effective sample size by about 13%; another study found up to 14%.
    • Why so modest:
      • model inaccuracy;
      • human answers are only partly predictable: most people's attitudes are not coherent (Converse), the same person answers the same question differently (Zaller & Feldman), and errors cluster on individuals (Salganik et al. 2020);
      • estimator inefficiency.
  • Lower-risk uses.
    • Simulate-then-validate, which is still vulnerable to overestimated effects and false positives on small effects.
    • Discovery of features and outcomes.
    • Design analysis and stress tests.
    • Hypotheses about mechanism via model internals, e.g. sparse autoencoders on models fine-tuned to simulate people.
  • Conclusion. A clear line between heuristic use and methods with calibration guarantees. There is also a risk that research questions drift toward what models can simulate, an "illusion of exploratory breadth".

Limits

  • A position and review paper with no new data. The formal results are summaries of Ludwig et al. and Perkowski, not derived here.
  • The focus is population parameters (means, effects, coefficients) for confirmatory research, not the individual-level prediction or long-term simulation of one person.

What it means for Kurisutina

  • The GSS pilot is a validation of individual predictions, and it is set up the valid way.

    • The population of cases is defined: GSS panel pairs with the given events.
    • The sample is drawn at random from it, and prompts, models and code are fixed before any scoring (the freeze).
    • Every forecast is scored against the person's real held-out answer.

    That supports claims about predictive accuracy on that population, not beyond it. It does not license treating replica outputs as human data for other questions.

  • Leakage can't be ruled out. GSS files are public and Qwen's training data are undocumented. The within-person contrasts help, on Perkowski's logic: a bias shared by all conditions of the same person cancels in own − other and own − empty. But memorisation of a respondent's later answers would favour the condition carrying that respondent's own record, and inflate the own-vs-empty contrast (D3). A model with open training data (OLMo) would allow a direct leakage check. Noted as a limit of the pilot.

  • The adopted mixture analysis is calibration in miniature. Cross-fitting a weight between the model's forecast and the population table asks the same question as PPI: does the model carry information beyond what the human data already give? The paper's evidence (+13–14% effective sample size) sets expectations: modest gains are the norm.

  • There is a ceiling for any replica. Same-question inconsistency and incoherent attitudes cap individual predictability. MR119's reliabilities quantify that ceiling for the GSS items. Future reports should put the replica's accuracy against that ceiling (like Park et al.'s "85% of self-consistency"), not against perfection.

  • Drift appears as a named failure in the literature. Taubenfeld et al. 2024 (appendix table): agents in simulated debates converge toward the model's own biases even while role-playing the other side. Together with lower variance, exaggerated group differences and caricature, this is the "average person with her views" failure in other words. The pilot's own-vs-other and own-vs-empty contrasts, and the overconfidence already seen in the smoke test, measure exactly this.

  • For "measuring the exact delta" (the secondary goal), calibration estimators give a principled route. With real answers from the person on a random subset of items, the replica's predictions on the rest can be combined into person-level estimates with valid uncertainty. Its errors don't have to be assumed away.

Cross-references

  • Grief-Albert et al. 2026 (summaries/carry_on/griefalbert2026_emulate.md): exaggerated group differences and collapse in post-trained emulation.
  • Peng et al. 2026 (summaries/carry_on/peng2025_funhouse.md): twins closer to a generic persona than to their person.
  • Namazova et al. 2025 (summaries/carry_on/namazova2025_open_loop.md): prediction vs generation, the "model falsification" view cited here.
  • summaries/carry_on/gss_mr119.md: the reliability ceiling for GSS items.
  • docs/research/gss_pilot_design.md: freeze, held-out scoring, mixture analysis.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.