Kurisutina

CogBench: a large language model walks into a psychology lab

arXiv:2402.18225v1 (28 February 2024; the only arXiv version), "preprint, under review"; the published ICML 2024 version was not read. Code: github.com/juliancodaforno/CogBench (not inspected). Read by researcher R4d (the base battery), 5 October 2026. Provenance: papers/base/codaforno2024_cogbench.provenance.json.

What was read

  • Read in full: the arXiv v1 PDF (26 pages, pdftotext -layout): abstract, sections 1–7, Figures 1–5 captions and their text layers, references, Appendix A (the 35 models), Appendix B (every task: methods, prompts, metrics), Appendix C (captions; per-model panels are images), Appendix D (prompting techniques).
  • Not read: the figures as images. The human reference values below are read from Figure 1's text layer, where the column assignment (four example models, then "Humans") follows the layout and is consistent with the text ("both under one" for the human prior and likelihood weights). Per-model values in Figures 2, 6 and 7 are not available as text.

What it is

  • Ten behavioural metrics from seven canonical experiments, each also giving a performance score, run on 35 LLMs (GPT-4, text-davinci-002/003, Claude-1/2, PaLM-2 text-bison, Falcon, MPT, LLaMA-2 and its chat, long-context and code variants), temperature 0, in-context learning only; history is concatenated into the prompt each trial.
  • Normalisation: each metric is rescaled so that a random agent = 0 and the average human = 1.
  • Several tasks are procedurally generated, which "makes it hard to game our benchmark by training on the test set"; temporal discounting is not (fixed items, one run per model).
Task (human source) Metric(s) Human value (Fig. 1) Runs per model
Probabilistic reasoning, wheel and urns (Dasgupta et al. 2020; humans' first trial only) Prior weighting β1 and likelihood weighting β2 in a log-odds regression (generalised Bayes) 0.88 and 0.91 (both under 1: "system neglect") 100
Horizon task (Wilson et al. 2014) Directed exploration: horizon coefficient in the unequal-information condition; random exploration: reward-difference × horizon interaction in the equal condition 0.33 and 0.02 100
Restless bandit with confidence reports (Ershadmanesh et al. 2023) Meta-cognition: adjusted quadratic scoring rule, 1 − (accuracy − scaled confidence)² 0.77 10
Instrumental learning, 4 interleaved bandits, 96 trials (Lefebvre et al. 2017) Learning rate (Rescorla–Wagner); optimism bias α+ − α− 0.19 and 0.19 10
Two-step task, 20 rounds (Daw et al. 2011) Model-basedness: reward × common-transition interaction on stay probability 0.03 100
Temporal discounting (Ruggeri et al. 2022) Score 0–19 (higher = more often the sooner option) 10.30 1
Balloon Analogue Risk Task (Lejuez et al. 2002) Risk: mean number of inflation attempts 17.22 10

Main results (verified)

  • Performance: GPT-4 and Claude-1 reach human level on 5 of 6 performance metrics; every model is competent on at least half. Almost all models are super-human on the horizon task while not exploring like humans (only LLaMA-2-70-chat shows more than human random exploration). The restless bandit is hard for most; the BART for all.
  • Behaviour: "none of the models exhibit human-like behavior on the majority of behavioral metrics." All models weight priors far more than evidence; most show a strong optimism bias and high learning rates; in the BART models sit at the extremes (never or always inflate); LLaMA-2-70 and its chat version perform equally on the BART with opposite risk-taking.
  • Multilevel regressions over the 35 models: parameter count raises performance (β 0.277) and model-basedness (β 0.481); RLHF raises meta-cognition (β 0.461) and makes behaviour about 2× closer to human in a UMAP of the ten metrics (an 11.7% smaller average L2 distance on normalised vectors); dataset size and code training had no clear effect; open-source models took fewer risks (β −0.612), refuting the authors' hypothesis.
  • Prompting: chain-of-thought raised posterior accuracy by 9.01% on average and model-basedness by 64.59%; take-a-step-back raised them by 3.10% and 118.59% (five models, inverse-variance weighted).

Limits

  • The human reference is a mean only. No human spread, interval or individual distribution is used; normalisation to "human = 1" cannot say whether a model lies inside the human range.
  • Task versions differ from the human ones: the restless bandit has 4 blocks and about 80 trials with 0–1 confidence (humans: 20 blocks, 400 trials, Likert); the BART uses 10 balloons per type labelled by letters (humans: 15, colours) and a different risk measure (all inflation attempts instead of adjusted pumps); human probabilistic-reasoning data are first trials only.
  • Temperature 0: each model is one deterministic agent, so no population spread is measured. Some runs are small (10 simulations; 1 for discounting).
  • The model features used in the regressions (sizes, data, RLHF) are partly guesses about proprietary models, as the authors note. External validity of these tasks for LLMs is untested.

What it means for the base battery (inference)

  • A ready, open, mostly procedurally generated block of learning and decision signatures with human reference means: belief updating (prior and likelihood weights), directed and random exploration, meta-cognition, learning rate and optimism bias, model-basedness, discounting, risk. These are "shared machinery" signatures in the base spec's sense.
  • It must be extended before it can serve as a pass rule: take each metric's human distribution (SD, or the individual estimates) from the source study; run the base as a population (many empty-slot draws, sampled, not at temperature 0); pass when the base's mean is equivalent to the human mean within declared bounds and its between-draw spread is within a declared ratio of the human spread.
  • Match the human procedures exactly (trial counts, scales, balloon counts, risk measure) or the comparison is not like for like; CogBench's simplifications are acceptable only for a screening run.
  • The temporal-discounting items are fixed and published, so they need perturbed or regenerated variants (Binz and Schulz 2023's lesson).

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.