Kurisutina

A foundation model to predict and capture human cognition (Centaur)

Nature 644, 1002–1009 (2025); received 26 October 2024, accepted 29 May 2025, published 2 July 2025. DOI 10.1038/s41586-025-09215-4; PMC12390832; CC BY 4.0. Peer reviewed (reviewers named: Russell Poldrack, Giosue Baggio and anonymous). 37 authors led by Marcel Binz and Eric Schulz (Helmholtz Munich), with Dayan, Griffiths, Eckstein, Mattar, Wilson and others. Provenance: papers/base/binz2025_centaur.provenance.json.

What was read

Every line of the PMC XML converted with tools/pmc2txt.py (1,077 lines after folding): abstract, main text, Methods (data collection, fine-tuning, evaluation metric, domain-specific models, neural alignment, model-guided discovery), captions of Figs 1–5 and Extended Data Figs 1–7, data and code availability, competing interests and the 71 references. Not read: the Supplementary Information (prompts for every Psych-101 paradigm, specifications of the 14 domain-specific models, a discussion against Newell's criteria), the Reporting Summary, the figures as images, and Extended Data Table 1 (per-experiment negative log-likelihoods), which is an image in the XML. Per-experiment numbers are therefore not available here; only the averages stated in the text are.

Question

Can one model predict and simulate human behaviour "in any experiment expressible in natural language", by fine-tuning a language model on many experiments' trial-by-trial human data?

Data: Psych-101

  • Size. 160 psychological experiments, 60,092 participants, 10,681,650 choices, 253,597,411 text tokens.
  • Domains. Multi-armed bandits, decision-making, memory, supervised learning, Markov decision processes and others (Extended Data Fig. 1a gives proportions only as an image). The authors: the focus "is largely on learning and decision-making"; psycholinguistics, social psychology and economic games are future additions.
  • Format. Each prompt is the entire trial-by-trial history of one participant's complete session, transcribed by hand into natural language, following the original instructions with simplifications, up to about 32,768 tokens.
  • Selection. Publicly available trial-level data; transcribable without significant loss of information; broad coverage.
  • No person information. Prompts carry no age, personality or socioeconomic status. The authors call experiments with individual-difference information "neglected data" and want them added "such that a model trained on these data can capture individual differences".
  • Bias. Some cross-cultural and meta-studies, but "a strong bias towards a WEIRD population".
  • Contamination. LogProber found no prompt above the threshold log B ≥ 1 (Extended Data Fig. 1c).
  • Access. Training data public on Hugging Face (marcelbinz/Psych-101); the test set under CC-BY-ND-4.0 in a gated repository (marcelbinz/Psych-101-test). The licence of the training split is not stated in the paper.

Model

  • Llama 3.1 70B, frozen and 4-bit quantized, with QLoRA adapters of rank 8 on every linear layer of attention and feed-forward blocks: 0.15% of the base model's parameters.
  • One epoch, cross-entropy with the loss masked to human-response tokens only. Batch 32, learning rate 5e-5, weight decay 0.01, 8-bit AdamW, 100 warm-up steps, unsloth. About five days on one A100 80GB. "Prolonged training … led to overfitting."
  • Minitaur: the same recipe on Llama 3.1 8B. Close to the training distribution it captures behaviour; it "generalizes less robustly" out of distribution (Extended Data Fig. 7, image).

Results

Held-out participants (90/10 split of participants within each experiment; mean negative log-likelihood per response, n = 992,867 responses):

Model NLL Difference to Centaur Cohen's d
Centaur 0.44 — —
Llama 3.1 70B (no fine-tuning) 0.58 0.14 0.20
Domain-specific cognitive models (14 models, e.g. generalized context model, prospect theory, RL models) 0.56 0.13 0.18
  • Fine-tuning improved fit in every experiment; Centaur beat the domain models "in all but one experiment".
  • Predictability varied greatly across experiments (the text gives 0.49 for Centaur and 0.47 for Llama; read as the spread across experiments).
  • How the domain models were fitted: one joint parameter set for all training participants, then evaluated on held-out participants. Not per person.
  • Fine-tuning Llama for other purposes (Nemotron, Hermes, Reflection) did not improve fit to humans over the base (Extended Data Fig. 2).
  • Noise ceiling (choices13k and an intertemporal-choice study): Centaur exceeds the estimated ceiling, which the authors attribute to context-dependent patterns that standard ceiling analyses miss. Prompted to predict each response independently, it matches the domain models at about half the ceiling (Extended Data Fig. 3).

Open-loop simulation (the model's own responses fed back), three paradigms:

  • Horizon task: total reward 54.12 (s.d. 2.89) against humans' 52.78 (s.d. 2.90); equivalence by two one-sided tests with a ±3-point margin, P = 0.02. "Similar" directed exploration, shown as densities over reward and an information bonus (Fig. 2b, image).
  • Two-step task: simulated runs span purely model-free, purely model-based and mixtures, a bimodal distribution like humans' (Fig. 2c, image). The authors: Centaur captures "the distribution over trajectories produced by the entire population", not only the average participant.
  • Social prediction game: like the original humans, it predicts human partners (64%) better than an artificial agent with matched statistics (35%); t(230) = 20.32.

Out of distribution:

Test n responses Centaur Llama Domain model
Two-step task with a new cover story (magic carpet) 9,702 0.51 0.63 0.61 (hybrid MB/MF)
Maggie's farm (three-armed bandit; structure change) 510,154 0.42 0.62 0.98
Logical reasoning (LSAT items; domain excluded from training) 99,204 1.65 1.92 none built
  • For out-of-distribution tests the domain model was fitted on the most similar training experiment.
  • Six more unseen paradigms (Extended Data Fig. 4, one-sided tests against Llama): moral decisions t(181388) = −103.54; economic games t(7798) = −11.69; naturalistic category learning t(21838) = −14.05; behavioural propensities t(156230) = −11.06; naturalistic reward learning t(9838) = −12.63, all p ≤ 0.0001. A deep sequential decision task: t(6092) = −1.06, p = 0.144, no significant gain, although the main text says Centaur "robustly captured human behaviour in all these settings".

Other measures:

  • Response times. About 4,000,000 response times; log RT regressed on log response entropy in mixed models. Conditional R² 0.87 with Centaur's entropies, 0.75 with Llama's, 0.77 with the cognitive models'.
  • Capability benchmarks (metabench): no significant change (ARC, GSM8K, HellaSwag, MMLU, Winogrande); TruthfulQA improved (z = 2.312, p = 0.021).
  • CogBench behavioural metrics, one-sided tests for movement toward humans: significant for prior weighting, random exploration, meta-cognition, model-basedness (z = 9.608) and temporal discounting. Not significant for likelihood weighting, directed exploration, learning rate, optimism bias and risk taking (p = 0.053).
  • Neural alignment. Two-step task fMRI (94 people, 300 choices each; an earlier study): Centaur's residual-stream representations predicted ROI activity better than Llama's at every layer (P ≤ 0.001; n = 11,374), and a cognitive model's representations did much worse. Sentence reading (5 people, 1,000 sentences): a pooled benefit of β = 0.007 [0.0002, 0.013], P = 0.045, not significant at any single layer.
  • Model-guided discovery. On a multi-attribute choice study (< 100 participants), a DeepSeek-R1-proposed strategy (AIC 181.7) fell short of Centaur (72.5). Scientific regret minimization against Centaur produced a weighted mix of two heuristics with AIC 71.7 and protected exceedance probability 0.83. Here cognitive models were fitted per participant by maximum likelihood.

Limits

  • Prediction is teacher-forced. Every headline number conditions on the person's own preceding choices and outcomes within the session. Open-loop checks are three paradigms, judged on distributions of a few statistics.
  • No individual is modelled as that individual. There is no person input beyond the in-context history, and no test of whether a person's behaviour in one experiment predicts another.
  • The comparison favours Centaur in two ways. The domain models share one parameter set across people, though such models are normally fitted per person (the Namazova commentary makes the same point). And the noise-ceiling result shows Centaur exploits within-session context that the baselines do not use.
  • Coverage. Mostly learning and decision tasks; WEIRD; only tasks expressible in text. Nothing on daily life, affect, autobiographical memory or change over months.
  • Statistics. One-sided tests with response-level degrees of freedom (t(1,985,732)) treat responses as independent despite nesting in people and experiments.
  • Supplement not read, so the 14 baseline models and the per-experiment table were not checked.

What it means for Kurisutina

  • A population model of behaviour can be learned from pooled experiments, and it forecasts well. Psych-101 is the largest open corpus of trial-level human behaviour, and it is public. For the base, it is candidate training and test material for decision and learning machinery. It does not cover the B2–B7 functions: no episodes over days, no affect ratings, no social memory, no change over time.
  • Its strength is the condition the replica loses. Centaur predicts the next choice given the person's own history in the prompt. After hand-over, the replica runs on its own history. The Namazova test of the same model found open-loop failures where this paper found open-loop successes. The two papers measured different statistics on the horizon task: total reward and an information bonus here, the choice curve across free trials there.
  • The person lives in the prompt here. Individuality is carried only by the in-context session. That is exactly what P2 forbids, and the authors themselves name person information as missing.
  • Recipe facts worth keeping. A small adapter (0.15%) on a frozen base moved behaviour toward humans without measurable loss on capability benchmarks, and longer training overfitted. Masking the loss to the person's responses is the natural objective for "predict the person, not the prose".

Cross-references

  • summaries/carry_on/namazova2025_open_loop.md: the open-loop critique of this model.
  • summaries/carry_on/eckstein2022_context.md: one of the task families in Psych-101.
  • summaries/carry_on/peng2025_funhouse.md: Centaur as one arm in a digital-twin benchmark.

This summary is our record of the paper, written after reading the full text and published as written; links into our own repository have been removed.