PNAS 120(6): e2218523120, doi 10.1073/pnas.2218523120; open access (CC BY-NC-ND 4.0), PMC9963545. Code and data: github.com/marcelbinz/GPT3goesPsychology (not inspected). Read by researcher R4d (the base battery), 5 October 2026. Provenance: papers/base/binz2023_cognitive_psychology_gpt3.provenance.json.
What was read
- Read in full: the Europe PMC full-text XML, converted to text (significance statement, abstract, main text, figure captions, Materials and Methods, statements, references), and the SI Appendix PDF (12 pages,
pdftotext -layout): data-generating distributions, Figs S1–S2 captions and visible text, Tables S1–S7 (every vignette prompt with GPT-3's answer, the 17 gamble problems with GPT-3's choice probabilities, the 13 contrasts). - Not read: the figures as images (values read from the text only), the repository, the two linked commentaries.
What they did
GPT-3 ("Davinci" through the OpenAI API, temperature 0 unless stated) was run as a participant in canonical cognitive-psychology experiments of two kinds:
- Vignettes (12): fixed texts from famous studies (Linda, cab, hospital, Wason card selection, Cognitive Reflection Test, blickets, counterfactual "pills", and others), each answer scored as correct, a human-like mistake, or neither. Then adversarial vignettes: the same problems slightly altered.
- Tasks (4), procedurally generated so that the exact problems cannot be in the training data:
- decisions from descriptions: over 13,000 gamble pairs (Peterson et al. 2021), plus Kahneman and Tversky's 17 problems for six bias contrasts (human data from Ruggeri et al. 2020);
- the horizon task (Wilson et al. 2014): 3,200 games, "data from ten participants"; human data from Zaller et al. 2021;
- the two-step task (Daw et al. 2011): 200 runs of 20 rounds;
- seeing against doing in causal inference (Waldmann and Hagmayer 2005).
- Robustness: 12 prompt variants for gambles (instruction, currency, labels) and for the horizon task (currency, labels, a casino against an investment cover story); a second cover story for the two-step task.
Main results (verified)
- Vignettes: 6 of 12 correct, and all 12 either correct or a typical human error (it committed the conjunction fallacy and gave the intuitive wrong answer on all three CRT items; it solved the cab problem and Wason's selection task, which most people fail).
- Adversarial vignettes broke it in non-human ways: asked the probability that the cab was black, it said 0.2; reordering the Wason cards (4, 7, A, K) made it choose A and K; with "the bat costs $1.00 more than the bat" it still answered $0.10; a blicket version of the counterfactual task produced self-contradiction.
- Gambles: only Davinci beat chance; it stayed worse than humans in regret. Of six prospect-theory biases it showed three (framing, certainty, overweighting of small probabilities) and lacked three (reflection, isolation, magnitude perception), all of which the human data show.
- Horizon task: overall regret lower than humans' (GPT-3 M 2.72, SD 5.98; humans M 3.24, SD 10.26). It used random exploration (reward-difference β 0.18) but not strategically (no horizon interaction, β −0.02, p = .14) and no directed exploration (β −0.15, p = .58). Humans show both. Within a game humans improved more.
- Two-step task: model-based signatures (staying fell after a rewarded rare transition, χ²(1, N = 1984) = 36.53). Humans use a mixture of model-free and model-based learning.
- Causal reasoning: correct interventional inferences in the common-cause structure, but identical answers when the structure was a causal chain, so the earlier success was "purely accidental"; neither normative nor human-like.
- Prompt sensitivity: performance varied little across variants in gambles (10 of 12 better than chance, all worse than humans), but cover stories changed behaviour: an investment story made it risk-averse in the horizon task.
The authors' methodological points
- Vignettes are hard to interpret: the model has probably seen them, and small edits change its answers.
- Procedurally generated tasks avoid training-set contamination "by design" and allow analysing how a task is solved (bias contrasts, regression signatures), not only performance.
- An open question they raise: is a language model one participant or many, since it was trained on text from many people?
Limits
- One closed model, mostly at temperature 0; one run per vignette; no human data were collected (all human references come from published studies, with differing samples and formats).
- Comparisons are mostly significance tests of presence or absence of an effect, not equivalence tests of effect size.
What it means for the base battery (inference)
- Never use famous vignettes as pass/fail items. A base can pass them by having read the literature, and fail perturbed versions. Any battery item drawn from a known study must come with perturbed and procedurally generated variants, and the pass rule should require the same behaviour on the unseen variants.
- Score the signature, not the score. Directed exploration (horizon × information interaction), the model-based/model-free mixture in the two-step task, and the prospect-theory contrasts are parameterised signatures with human reference values; a base that outperforms humans (lower regret) without them is not human-like.
- Human spread is part of the reference: human regret SD 10.26 against GPT-3's 5.98 is the "flat population" pattern the base spec warns about.
- The one-or-many question is the battery's empty-slot question: the base must be tested as a population (many draws with an empty slot), not as one respondent.