Kurisutina

Using large language models to simulate multiple humans and replicate human subject studies

arXiv:2208.10264v5 (9 July 2023), ICML 2023. Code: github.com/GatiAher/Using-Large-Language-Models-to-Replicate-Human-Subject-Studies (not inspected). Read by researcher R4d (the base battery), 5 October 2026. Provenance: papers/base/aher2023_turing_experiments.provenance.json.

What was read

  • Read in full: the arXiv v5 PDF (43 pages, pdftotext -layout): main text, Tables 1–4, figure captions, references, and Appendices A–F (the full surname lists, the input summary, Ultimatum, garden-path sentences with all 48 original and 48 novel sentences, Wisdom of Crowds tables, the complete Milgram and novel-obedience prompt templates, prods, break-off tables and termination completions).
  • Not read: the figures as images (plots of acceptance curves, garden-path ratings and break-off curves; numbers are taken from the text and tables).

What it is

  • A Turing Experiment (TE) asks a model to simulate a representative sample of participants in a known human study, not one individual, and compares the simulated outcome with the human finding. It should be zero-shot: neither the procedure nor the training data should contain data from that experiment ("difficult to enforce with models pretrained by others").
  • Method. Participants are varied by name only: Mr./Ms. (and Mx. in one study) plus surnames from the 2010 US census, 100 from each of five racial groups (1,000 names). Answers are read either as k-choice probabilities (the model's probability for each valid completion, normalised; the normaliser is the validity rate) or as free completions classified by further prompts. Sampling at temperature 1, top-p 1.
  • Prompt validation without peeking: prompts were tuned only to raise the validity rate (or coherence) before any hypothesis was tested, to avoid p-hacking.
  • Models: OpenAI text-ada-001, babbage-001, curie-001, davinci-001, davinci-002 (LM-5), davinci-003 (LM-6), gpt-3.5-turbo, gpt-4, and the base davinci.
  • Contamination check: novel variants for three of the four studies (new garden-path sentences, a new obedience scenario, new knowledge questions).

Main results (verified)

  • Ultimatum game (10,000 name pairs × 11 offers of $0–10): LM-5 reproduced the human pattern (offers of 50–100% almost always accepted, 0–10% rarely); smaller models were flat. Acceptance correlated > 0.9 within a name pair across offers $1–4 and $6–9, so name-driven variation was consistent, not noise. A large gender effect: men accepted a $2 offer from a woman 60% of the time, women accepted one from a man 20%, with little overlap between the distributions; human gender effects "are not uniformly consistent".

  • Garden-path sentences (24 from Christianson et al. 2001 plus 24 novel, each with a comma-disambiguated control): LM-5 rated every garden-path sentence more ungrammatical than its control, on both sets; the OT/RAT difficulty order flipped on the novel set. The task was simplified from the human paradigm.

  • Milgram obedience (100 simulated subjects; the model also played the experimenter's classifier): 75% fully obedient against Milgram's 26 of 40 (65%); a spike of stopping at 300 V, when the victim stops answering (18 simulated stops, against 5 humans who stopped there). The novel submersion scenario gave 75 fully obedient, with stops concentrated at the 20th and 22nd submersion.

  • Wisdom of crowds — the "hyper-accuracy distortion":

    Question Truth Humans (Moussaïd et al. 2013, n = 52): median (IQR) base davinci davinci-003, gpt-3.5, gpt-4
    Bones in an adult human 206 190 (108) 136 (346) 206 (0)
    Melting point of aluminium, °C 660 240 (532) 435 (887) 660 (0)
    100 °C in °F 212 200 (195) 100 (132) 212 (0)
    Earth days in a Mars year 687 365 (376) 327 (658) 687 (0)
    Speed of sound, m/s 343 333 (884) 348 (365) 340–343 (0)

    For all ten questions (five novel), the newer, more "aligned" models gave the exact answer with zero interquartile range: every simulated person knew it. The base davinci and small models did not. The authors suggest alignment for truthfulness causes it (not tested directly). gpt-4 rounded the speed of light to 3 × 10⁸.

Limits

  • Few studies, coarse comparisons (curves and medians, no statistical test of equivalence); human references are taken from published summaries (e.g. a 52-person study for the knowledge questions).
  • Names are the only person input, so "population spread" here is spread across names, partly stereotype-driven (the gender effect).
  • Novel variants remain close to the originals by design (the new obedience scenario is "reminiscent of Milgram"); training-set exposure cannot be ruled out for closed models.
  • In the Milgram TE the model judges its own outputs; a classifier error changes the outcome.

What it means for the base battery (inference)

  • The empty-slot test is a Turing Experiment. The base with an empty slot should reproduce a study's sample, including its spread, not one modal respondent.
  • Bounded knowledge has a ready human reference: people's estimates of general-knowledge quantities are wide and biased (aluminium: median 240 °C, IQR 532, truth 660). A base whose empty-slot draws all answer exactly is hyper-accurate, not human; the pass rule should compare the distribution of log(estimate/truth) with the human one.
  • Variation must be consistent within a simulated person (here, > 0.9 across related conditions) and not imported from stereotypes keyed to names; a base should be tested with and without demographic cues.
  • Freeze the read-out before looking at outcomes (their validity-rate rule), and include novel, never-published variants of every classic study.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.